Question Answering on LogiQA (test)
85.75AccuracyGemini-2.5-Pro
Evaluation Results
| Method | Links | |
|---|---|---|
| Gemini-2.5-ProModel Category=Proprietary LLM2026.01 | 85.75 | |
| Qwen2.5-72BModel Category=Open-source LLM2026.01 | 76.91 | |
| GPT-5.1Model Category=Proprietary LLM2026.01 | 76.34 | |
| GPT-4oModel Category=Proprietary LLM2026.01 | 74.01 | |
| SGR-Llama3.3-70BMethod=Self-Graph Reasoning2026.01 | 69.91 | |
| Claude-3.5-HaikuModel Category=Proprietary LLM2026.01 | 65.97 | |
| LLaMA-3.3-70BModel Category=Open-source LLM2026.01 | 64.01 | |
| RwG-LLaMA3.1-70BMethod Category=Specialized Graph-based2026.01 | 59.13 | |
| LLaMA-3.1-8BModel Category=Open-source LLM2026.01 | 49.17 | |
| RwG-Claude-3-sonnetMethod Category=Specialized Graph-based2026.01 | 45.16 | |
| LLaMA-3.2-3BModel Category=Open-source LLM2026.01 | 41.28 | |
| Qwen2.5-7BModel Category=Open-source LLM2026.01 | 34.1 |