Multi-turn Inference Latency on IntentFlow T=25-36
1,500Average Latency (Demand Turns)Qwen3-30B-A3B
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Qwen3-30B-A3BParameters=30B2026.04 | 1,500 | 1,200 | 1,300 | |
| IntentFlow2026.04 | 1,700 | 1,200 | 1,400 | |
| Gemini-2.5-Flash-Lite2026.04 | 2,300 | 2,300 | 2,300 | |
| DeepSeek-V3.22026.04 | 3,700 | 3,100 | 3,400 | |
| Claude-Haiku-4.52026.04 | 4,100 | 3,200 | 3,600 | |
| Gemini-3-Flash2026.04 | 4,300 | 4,100 | 4,200 | |
| GPT-5-Nano2026.04 | 7,000 | 6,200 | 6,500 | |
| GPT-oss-120bParameters=120b2026.04 | 9,100 | 8,400 | 8,700 | |
| GPT-5-Mini2026.04 | 9,900 | 7,300 | 8,400 | |
| Qwen3.5-Flash2026.04 | 16,700 | 18,700 | 17,800 |