| Model | Needs in time | Too late | No need | Repeat | Silent | Latency | Cost / run |
|---|---|---|---|---|---|---|---|
| gpt-5.6-solopenai · default | 5 / 7followup@1957, followup@547, place@2560, place@626, transport@1323 | 2 | 0 | 2 | 19 / 33 | p50 7.6 · p90 12.7 · max 19.6 s | $0.952 |
| deepseek-v4.1-flashdeepseek · reasoning off | 6 / 7followup@1957, followup@547, place@2560, place@626, transport@1323, transport@652 | 7 | 2 | 3 | 0 / 33 | p50 5.9 · p90 14.4 · max 54.7 s | $0.070 |
| glm-5.3-flashz-ai · reasoning effort low | 6 / 7followup@1957, followup@547, place@2560, place@626, transport@1323, transport@652 | 7 | 1 | 2 | 9 / 33 | p50 13.9 · p90 30.2 · max 61.6 s | $0.047 |
| gpt-5.6-lunaopenai · default | 6 / 7booking, followup@1957, followup@547, place@2560, place@626, transport@652 | 13 | 1 | 6 | 2 / 33 | p50 8.4 · p90 10.1 · max 11.3 s | $0.112 |
| gemini-3.8-flashgoogle · default | 0 / 7– | 0 | 0 | 0 | 33 / 33 | p50 10.4 · p90 20.3 · max 22.8 s | $0.426 |
| kimi-k3moonshotai · default | 6 / 7booking, followup@1957, followup@547, place@2560, place@626, transport@652 | 13 | 0 | 1 | 2 / 33 | p50 70.5 · p90 200.6 · max 221.0 s | $2.101 |
Ranked by needs in time, minus half a point per wrong card (too late or no labelled need), minus one point if p90 latency exceeds 20 s. Latency is the measured model round trip per decision. Cost is for all 33 decisions of this recording.