Question Answering on Natural Questions (NQ) (LM and Accuracy Metrics)
83.02AccuracyGoldRouter (Multi-RAG-Agent)
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| GoldRouter (Multi-RAG-Agent)Models=Qwen-Plus2025.01 | 83.02 | 63.41 | |
| EfficientRAG (Single-RAG-Agent)Models=Qwen-Plus2025.01 | 82.61 | 62.73 | |
| CoT (Without RAG)Models=Qwen-Plus2025.01 | 74.55 | 52.99 | |
| EfficientRAG (Single-RAG-Agent)Models=LLaMA-3.1-8B2025.01 | 70.1 | 56.48 | |
| GoldRouter (Multi-RAG-Agent)Models=LLaMA-3.1-8B2025.01 | 69.7 | 55.13 | |
| CoT (Without RAG)Models=LLaMA-3.1-8B2025.01 | 58.6 | 44.41 | |
| Llama3Parameters=8B, Complexity Type=Quadratic, Evaluation Framework=vLLM, Evaluation Method=generation-based2025.09 | 30.97 | — | |
| MixtralParameters=13B/47B, Complexity Type=Quadratic, Evaluation Framework=vLLM, Evaluation Method=generation-based2025.09 | 28.48 | — | |
| SpikingBrain-76BParameters=12B/76B, Complexity Type=Hybrid, Evaluation Framework=vLLM, Evaluation Method=generation-based2025.09 | 21.55 | — | |
| SpikingBrain-7BParameters=7B, Complexity Type=Linear, Evaluation Framework=vLLM, Evaluation Method=generation-based2025.09 | 21.47 | — | |
| Qwen2.5Parameters=7B, Complexity Type=Quadratic, Evaluation Framework=vLLM, Evaluation Method=generation-based2025.09 | 17.67 | — |