Arithmetic Reasoning on Game of 24 (test)
90Success RateICRL Preset
Evaluation Results
| Method | Links | |
|---|---|---|
| ICRL PresetBase Model=GPT-4.1, Evaluation Protocol=running max success rate of last episode2025.05 | 90 | |
| ICRL AutonomousBase Model=GPT-4.1, Evaluation Protocol=running max success rate of last episode2025.05 | 84 | |
| HyperGuideBase model=Qwen2.52026.05 | 57 | |
| Best-of-NBase Model=GPT-4.1, Evaluation Protocol=running max success rate of last episode2025.05 | 49 | |
| Long-CoTBase Model=GPT-4.1, Evaluation Protocol=running max success rate of last episode2025.05 | 47 | |
| Self-RefineBase Model=GPT-4.1, Evaluation Protocol=running max success rate of last episode2025.05 | 47 | |
| ReflexionBase Model=GPT-4.1, Evaluation Protocol=running max success rate of last episode2025.05 | 44 | |
| HyperGuideBase model=Mistral2026.05 | 44 | |
| HyperGuideBase model=GPT-OSS2026.05 | 42 | |
| SoftCoTBase model=Qwen2.52026.05 | 27 | |
| Self-ConsistencyBase model=Qwen2.52026.05 | 21 | |
| SoftCoTBase model=GPT-OSS2026.05 | 21 | |
| SoftCoTBase model=Mistral2026.05 | 17 | |
| Self-ConsistencyBase model=GPT-OSS2026.05 | 16 | |
| OVMBase model=Qwen2.52026.05 | 15 | |
| Self-ConsistencyBase model=Mistral2026.05 | 15 | |
| OVMBase model=GPT-OSS2026.05 | 13 | |
| Few-shotBase model=Qwen2.52026.05 | 11 | |
| Tree of ThoughtsBase model=Qwen2.52026.05 | 10 | |
| Tree of ThoughtsBase model=GPT-OSS2026.05 | 9 | |
| OVMBase model=Mistral2026.05 | 9 | |
| Few-shotBase model=GPT-OSS2026.05 | 8 | |
| Few-shotBase model=Mistral2026.05 | 8 | |
| PT-SFTBase model=Qwen2.52026.05 | 7 | |
| PT-SFTBase model=GPT-OSS2026.05 | 7 | |
| PT-SFTBase model=Mistral2026.05 | 7 | |
| CoT-onlyBase Model=GPT-4.1, Evaluation Protocol=running max success rate of last episode2025.05 | 6 | |
| Tree of ThoughtsBase model=Mistral2026.05 | 6 |