Question Answering on AR-LSAT (test)
96.22AccuracyGemini-2.5-Pro
Evaluation Results
| Method | Links | |
|---|---|---|
| Gemini-2.5-ProModel Category=Proprietary LLM2026.01 | 96.22 | |
| Qwen2.5-72BModel Category=Open-source LLM2026.01 | 34.78 | |
| GPT-5.1Model Category=Proprietary LLM2026.01 | 33.33 | |
| GPT-4oModel Category=Proprietary LLM2026.01 | 31.75 | |
| SGR-Llama3.3-70BMethod=Self-Graph Reasoning2026.01 | 31.74 | |
| RwG-LLaMA3.1-70BMethod Category=Specialized Graph-based2026.01 | 31.73 | |
| LLaMA-3.3-70BModel Category=Open-source LLM2026.01 | 31.3 | |
| RwG-Claude-3-sonnetMethod Category=Specialized Graph-based2026.01 | 30.86 | |
| Claude-3.5-HaikuModel Category=Proprietary LLM2026.01 | 29.41 | |
| LLaMA-3.1-8BModel Category=Open-source LLM2026.01 | 26.96 | |
| LLaMA-3.2-3BModel Category=Open-source LLM2026.01 | 20 | |
| Qwen2.5-7BModel Category=Open-source LLM2026.01 | 17.39 |