Multi-hop reasoning alignment on ScienceQA and ARC fused
81.581-Hop Alignment ScoreGPT-4o-Mini
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| GPT-4o-MiniModel Category=Closed-Source Models (non-reasoning), Assessed Version Date=2024-112025.11 | 81.58 | 80.66 | 77.99 | 74.8 | 80.08 | |
| GPT-o1Model Category=Reasoning Models, Assessed Version Date=2024-092025.11 | 80.2 | 82.85 | 79.95 | 78.6 | 81 | |
| GPT-4oModel Category=Closed-Source Models (non-reasoning), Assessed Version Date=2025-032025.11 | 78.78 | 81.59 | 78.81 | 74.59 | 79.73 | |
| LLaMA2-13B-ChatModel Category=Open-Source Models (non-reasoning)2025.11 | 77.46 | 79.67 | 77.19 | 72.28 | 78.1 | |
| DeepSeek–R1Model Category=Reasoning Models, Assessed Version Date=2025-052025.11 | 77.3 | 87.82 | 86.91 | 82.87 | 84.01 | |
| GPT-3.5-TurboModel Category=Closed-Source Models (non-reasoning), Assessed Version Date=2023-112025.11 | 76.97 | 81 | 78.5 | 77.07 | 78.82 | |
| Qwen2.5-3B-InstructModel Category=Open-Source Models (non-reasoning)2025.11 | 75.93 | 81.64 | 79.51 | 77.6 | 79.03 | |
| Falcon-7B-InstructModel Category=Open-Source Models (non-reasoning)2025.11 | 75.23 | 79.44 | 73.38 | 70.44 | 76.02 |