Proactivity

Proactivity/Evening (Cafeteria), Sep 20, 2026
Transcript
ModelNeeds in timeToo lateNo needRepeatSilentLatencyCost / run
gpt-5.6-solopenai · default3 / 5Recording cut, Follow-ups, Unit economics01115 / 24
p50 9.4 · p90 18.3 · max 26.3 s
$0.478
gemini-3.8-flashgoogle · default2 / 5Recording cut, Unit economics00022 / 24
p50 9.1 · p90 15.9 · max 21.4 s
$0.232
kimi-k3moonshotai · default2 / 5Recording cut, Unit economics21215 / 24
p50 41.8 · p90 80.6 · max 97.3 s
$1.523
glm-5.3-flashz-ai · reasoning effort low2 / 5Follow-ups, Unit economics22212 / 24
p50 15.8 · p90 49.2 · max 408.7 s
$0.026
gpt-5.6-lunaopenai · default3 / 5Recording cut, Menu, Unit economics7209 / 24
p50 6.9 · p90 8.4 · max 9.1 s
$0.056
deepseek-v4.1-flashdeepseek · reasoning off2 / 5Menu, Unit economics10335 / 24
p50 6.8 · p90 65.6 · max 103.8 s
$0.045
Ranked by needs in time, minus half a point per wrong card (too late or no labelled need), minus one point if p90 latency exceeds 20 s. Latency is the measured model round trip per decision. Cost is for all 24 decisions of this recording.