Multi-turn Inference Latency on IntentFlow T=1-12
1,400Average Latency (Demand Turns) (ms)Qwen3-30B-A3B
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Qwen3-30B-A3BParameters=30B2026.04 | 1,400 | 988,000 | 1,100 | |
| IntentFlow2026.04 | 1,600 | 1,200 | 1,300 | |
| Gemini-2.5-Flash-Lite2026.04 | 2,800 | 2,200 | 2,400 | |
| DeepSeek-V3.22026.04 | 3,400 | 3,100 | 3,200 | |
| Gemini-3-Flash2026.04 | 3,600 | 3,000 | 3,200 | |
| Claude-Haiku-4.52026.04 | 3,600 | 2,800 | 3,100 | |
| GPT-5-Nano2026.04 | 7,400 | 6,300 | 6,700 | |
| GPT-oss-120bParameters=120b2026.04 | 7,700 | 6,300 | 6,800 | |
| GPT-5-Mini2026.04 | 10,400 | 6,700 | 8,100 | |
| Qwen3.5-Flash2026.04 | 16,100 | 15,700 | 15,900 |