Multi-turn conversation performance on Actions
93.7Average PerformanceFull
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| FullModel=GPT-4o-mini, Setting=Ideal instructions2026.02 | 93.7 | 92.4 | |
| FullModel=DeepSeek-v3.2-Thinking, Setting=Ideal instructions2026.02 | 92.2 | 88.6 | |
| FullModel=GPT-5.2, Setting=Ideal instructions2026.02 | 90.2 | 93.2 | |
| Experience-Driven MediatorModel=DeepSeek-v3.2-Thinking, Setting=Ambiguous user inputs with Mediator2026.02 | 88 | 71.6 | |
| Experience-Driven MediatorModel=GPT-4o-mini, Setting=Ambiguous user inputs with Mediator2026.02 | 85.7 | 81.2 | |
| Experience-Driven MediatorModel=GPT-5.2, Setting=Ambiguous user inputs with Mediator2026.02 | 76.2 | 65.2 | |
| ShardedModel=GPT-4o-mini, Setting=Ambiguous user inputs2026.02 | 45.5 | 60 | |
| ShardedModel=DeepSeek-v3.2-Thinking, Setting=Ambiguous user inputs2026.02 | 42.3 | 48.6 | |
| ShardedModel=GPT-5.2, Setting=Ambiguous user inputs2026.02 | 35.6 | 46.6 |