Multi-turn Inference Latency on IntentFlow T=13-24
1.4Average Latency (Demand Turns) (ms)Qwen3-30B-A3B
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Qwen3-30B-A3BParameters=30B2026.04 | 1.4 | 1.1 | 1.2 | |
| IntentFlow2026.04 | 1.6 | 1.2 | 1.4 | |
| Gemini-2.5-Flash-Lite2026.04 | 2.3 | 2.4 | 2.4 | |
| DeepSeek-V3.22026.04 | 3.8 | 3.1 | 3.4 | |
| Gemini-3-Flash2026.04 | 3.9 | 3.7 | 3.8 | |
| Claude-Haiku-4.52026.04 | 4 | 3.4 | 3.6 | |
| GPT-5-Nano2026.04 | 6.9 | 5.7 | 6.2 | |
| GPT-oss-120bParameters=120b2026.04 | 8.3 | 6.6 | 7.3 | |
| GPT-5-Mini2026.04 | 10.2 | 7.1 | 8.4 | |
| Qwen3.5-Flash2026.04 | 16.6 | 19.9 | 18.6 |