Multi-hop Reasoning on HotpotQA (Accuracy, Success Rate)
72.2AccuracyQwen3-14B
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Qwen3-14Bthinking mode=true2026.05 | 72.2 | — | — | — | |
| Qwen3-32Bthinking mode=true2026.05 | 71.4 | — | — | — | |
| Qwen3-32Bthinking mode=false2026.05 | 70.9 | — | — | — | |
| Qwen3-8Bthinking mode=true2026.05 | 70.3 | — | — | — | |
| gemma-3-27b-itthinking mode=false2026.05 | 69.6 | — | — | — | |
| Qwen3-8Bthinking mode=false2026.05 | 68.7 | — | — | — | |
| Qwen3-14Bthinking mode=false2026.05 | 68.3 | — | — | — | |
| Qwen3-4Bthinking mode=true2026.05 | 67.1 | — | — | — | |
| gemma-3-12b-itthinking mode=false2026.05 | 66.5 | — | — | — | |
| Qwen3-1.7Bthinking mode=true2026.05 | 60.9 | — | — | — | |
| OCC-RAG-1.7Bthinking mode=false2026.05 | 60.9 | — | — | — | |
| Qwen3-4Bthinking mode=false2026.05 | 60.6 | — | — | — | |
| OCC-RAG-0.6Bthinking mode=false2026.05 | 57.6 | — | — | — | |
| SmolLM3-3Bthinking mode=true2026.05 | 56.5 | — | — | — | |
| gemma-3-4b-itthinking mode=false2026.05 | 55.8 | — | — | — | |
| Baselinedecoding_mode=free-text2026.04 | 50 | 0.63 | — | — | |
| SmolLM3-3Bthinking mode=false2026.05 | 49.9 | — | — | — | |
| Pleias-RAG-1.2Bthinking mode=false2026.05 | 48.5 | — | — | — | |
| Qwen3-1.7Bthinking mode=false2026.05 | 47.7 | — | — | — | |
| Qwen3-0.6Bthinking mode=true2026.05 | 41.8 | — | — | — | |
| Constrained Reflectiondecoding_mode=constrained2026.04 | 38 | 0.41 | 0 | 80 | |
| Qwen3-0.6Bthinking mode=false2026.05 | 34.8 | — | — | — | |
| gemma-3-1b-itthinking mode=false2026.05 | 30.8 | — | — | — |