Multi-turn Inference Latency on IntentFlow T=49-60
1.2Average Latency (Demand Turns) (ms)Qwen3-30B-A3B
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Qwen3-30B-A3BParameters=30B2026.04 | 1.2 | 873 | 1 | |
| IntentFlow2026.04 | 1.8 | 1.3 | 1.5 | |
| Gemini-2.5-Flash-Lite2026.04 | 2 | 1.9 | 1.9 | |
| Claude-Haiku-4.52026.04 | 4 | 3.5 | 3.7 | |
| DeepSeek-V3.22026.04 | 4 | 3.1 | 3.5 | |
| Gemini-3-Flash2026.04 | 4.4 | 4.3 | 4.4 | |
| GPT-5-Nano2026.04 | 6.3 | 5.8 | 6 | |
| GPT-5-Mini2026.04 | 8.6 | 5.8 | 7.1 | |
| GPT-oss-120bParameters=120b2026.04 | 8.6 | 7.4 | 8 | |
| Qwen3.5-Flash2026.04 | 16.4 | 17.8 | 17.2 |