Question Answering on AIW (test)
76AccuracyGemini-2.5-Pro
Evaluation Results
| Method | Links | |
|---|---|---|
| Gemini-2.5-ProModel Category=Proprietary LLM2026.01 | 76 | |
| SGR-Llama3.3-70BMethod=Self-Graph Reasoning2026.01 | 57.5 | |
| GPT-5.1Model Category=Proprietary LLM2026.01 | 57 | |
| GPT-4oModel Category=Proprietary LLM2026.01 | 32.5 | |
| LLaMA-3.3-70BModel Category=Open-source LLM2026.01 | 19.5 | |
| RwG-LLaMA3.1-70BMethod Category=Specialized Graph-based2026.01 | 12 | |
| LLaMA-3.1-8BModel Category=Open-source LLM2026.01 | 5 | |
| Qwen2.5-7BModel Category=Open-source LLM2026.01 | 5 | |
| Qwen2.5-72BModel Category=Open-source LLM2026.01 | 5 | |
| RwG-Claude-3-sonnetMethod Category=Specialized Graph-based2026.01 | 2.6 | |
| Claude-3.5-HaikuModel Category=Proprietary LLM2026.01 | 2.5 | |
| LLaMA-3.2-3BModel Category=Open-source LLM2026.01 | 1.7 |