Multi-hop Question Answering on MuSiQue (LM, F1, Acc.)
47.55AccuracyGoldRouter (Multi-RAG-Agent)
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| GoldRouter (Multi-RAG-Agent)Models=Qwen-Plus2025.01 | 47.55 | 29.01 | 33.75 | |
| EfficientRAG (Single-RAG-Agent)Models=Qwen-Plus2025.01 | 46.42 | 26.62 | 31.41 | |
| TDA-RCBase LLM=DeepSeek-V32026.03 | 37.9 | — | — | |
| TDA-RCBase LLM=GPT-4o-mini2026.03 | 37.8 | — | — | |
| HoTBase LLM=DeepSeek-V32026.03 | 37.2 | — | — | |
| TDA-RCBase LLM=Qwen-Turbo2026.03 | 37 | — | — | |
| HoTBase LLM=GPT-4o-mini2026.03 | 36.9 | — | — | |
| Instruction InductionBase LLM=DeepSeek-V32026.03 | 36.7 | — | — | |
| HoTBase LLM=Qwen-Turbo2026.03 | 36.6 | — | — | |
| Instruction InductionBase LLM=GPT-4o-mini2026.03 | 36.5 | — | — | |
| Role / Persona PromptingBase LLM=DeepSeek-V32026.03 | 36 | — | — | |
| Role / Persona PromptingBase LLM=GPT-4o-mini2026.03 | 35.9 | — | — | |
| Instruction InductionBase LLM=Qwen-Turbo2026.03 | 35.8 | — | — | |
| Prompt CanvasBase LLM=GPT-4o-mini2026.03 | 35.5 | — | — | |
| Prompt CanvasBase LLM=DeepSeek-V32026.03 | 35.5 | — | — | |
| Role / Persona PromptingBase LLM=Qwen-Turbo2026.03 | 35.2 | — | — | |
| Prompt CanvasBase LLM=Qwen-Turbo2026.03 | 34.8 | — | — | |
| CoT (Without RAG)Models=Qwen-Plus2025.01 | 33.86 | 16.76 | 20.66 | |
| Analogical PromptingBase LLM=DeepSeek-V32026.03 | 33.1 | — | — | |
| Analogical PromptingBase LLM=GPT-4o-mini2026.03 | 32.8 | — | — | |
| Analogical PromptingBase LLM=Qwen-Turbo2026.03 | 32.6 | — | — | |
| GoldRouter (Multi-RAG-Agent)Models=LLaMA-3.1-8B2025.01 | 31.46 | 23.27 | 28.18 | |
| EfficientRAG (Single-RAG-Agent)Models=LLaMA-3.1-8B2025.01 | 29.13 | 21.38 | 26.22 | |
| CoT (Without RAG)Models=LLaMA-3.1-8B2025.01 | 21.86 | 13.3 | 17.76 |