Logical Reasoning on ProntoQA (test)
99.72AccuracyHBLR
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| HBLRModel=DeepSeek-R1, Prompting Strategy=HBLR2025.12 | 99.72 | — | |
| CoTModel=DeepSeek-R1, Prompting Strategy=CoT2025.12 | 99.57 | — | |
| HBLRModel=DeepSeek-V3, Prompting Strategy=HBLR2025.12 | 99.55 | — | |
| HBLRModel=GPT-4, Prompting Strategy=HBLR2025.12 | 99.36 | — | |
| DetermLRModel=GPT-42023.10 | 98.6 | 9.78 | |
| SymbCoTModel=DeepSeek-R1, Prompting Strategy=SymbCoT2025.12 | 98.47 | — | |
| SymbCoTModel=DeepSeek-V3, Prompting Strategy=SymbCoT2025.12 | 98.43 | — | |
| CRModel=GPT-42023.10 | 98.2 | 14.18 | |
| CoTModel=DeepSeek-V3, Prompting Strategy=CoT2025.12 | 97.67 | — | |
| ToTModel=GPT-42023.10 | 97.6 | 18.91 | |
| DirectModel=DeepSeek-R1, Prompting Strategy=Direct2025.12 | 97.28 | — | |
| SymbCoTModel=GPT-4, Prompting Strategy=SymbCoT2025.12 | 97.16 | — | |
| LAMBADAModel=GPT-42023.10 | 95.6 | 10.56 | |
| CoTModel=GPT-4, Prompting Strategy=CoT2025.12 | 94.79 | — | |
| SIModel=GPT-42023.10 | 93.8 | 11.38 | |
| COT-SCModel=GPT-4, n=162023.10 | 93.4 | 16 | |
| DetermLRModel=GPT-3.5-turbo2023.10 | 93.2 | 10.74 | |
| CRModel=GPT-3.5-turbo2023.10 | 92.4 | 16.93 | |
| ToTModel=GPT-3.5-turbo2023.10 | 91.2 | 19.3 | |
| COTModel=GPT-42023.10 | 91 | 1 | |
| LAMBADAModel=GPT-3.5-turbo2023.10 | 90.8 | 12.09 | |
| LogicModel=GPT-4, Prompting Strategy=Logic2025.12 | 90.5 | — | |
| HyperGuideBase model=Mistral2026.05 | 89 | — | |
| SIModel=GPT-3.5-turbo2023.10 | 88.6 | 13.58 | |
| LogicModel=DeepSeek-V3, Prompting Strategy=Logic2025.12 | 87.83 | — | |
| COT-SCModel=GPT-3.5-turbo, n=162023.10 | 86.8 | 16 | |
| LogicModel=DeepSeek-R1, Prompting Strategy=Logic2025.12 | 84.21 | — | |
| COTModel=GPT-3.5-turbo2023.10 | 84 | 1 | |
| PT-SFTBase model=Mistral2026.05 | 81.5 | — | |
| Few-shotBase model=Mistral2026.05 | 81 | — | |
| OVMBase model=Mistral2026.05 | 79 | — | |
| StandardModel=GPT-42023.10 | 77.4 | 1 | |
| DirectModel=GPT-4, Prompting Strategy=Direct2025.12 | 77.4 | — | |
| HBLRModel=GPT-3.5-Turbo, Prompting Strategy=HBLR2025.12 | 75.58 | — | |
| HyperGuideBase model=Qwen2.52026.05 | 75 | — | |
| DirectModel=DeepSeek-V3, Prompting Strategy=Direct2025.12 | 74.93 | — | |
| Self-ConsistencyBase model=Mistral2026.05 | 73 | — | |
| SoftCoTBase model=Qwen2.52026.05 | 72 | — | |
| SymbCoTModel=GPT-3.5-Turbo, Prompting Strategy=SymbCoT2025.12 | 71.95 | — | |
| Tree of ThoughtsBase model=Mistral2026.05 | 69 | — | |
| HyperGuideBase model=GPT-OSS2026.05 | 68 | — | |
| CoTModel=GPT-3.5-Turbo, Prompting Strategy=CoT2025.12 | 67.8 | — | |
| LogicModel=GPT-3.5-Turbo, Prompting Strategy=Logic2025.12 | 67.21 | — | |
| OVMBase model=Qwen2.52026.05 | 67 | — | |
| SoftCoTBase model=GPT-OSS2026.05 | 64 | — | |
| OVMBase model=GPT-OSS2026.05 | 62 | — | |
| SoftCoTBase model=Mistral2026.05 | 61 | — | |
| Few-shotBase model=Qwen2.52026.05 | 60 | — | |
| Self-ConsistencyBase model=Qwen2.52026.05 | 58 | — | |
| PT-SFTBase model=GPT-OSS2026.05 | 56.5 | — | |
| PT-SFTBase model=Qwen2.52026.05 | 52.5 | — | |
| StandardModel=GPT-3.5-turbo2023.10 | 51.8 | 1 | |
| Few-shotBase model=GPT-OSS2026.05 | 48 | — | |
| DirectModel=GPT-3.5-Turbo, Prompting Strategy=Direct2025.12 | 46.04 | — | |
| Tree of ThoughtsBase model=GPT-OSS2026.05 | 44 | — | |
| Tree of ThoughtsBase model=Qwen2.52026.05 | 41 | — | |
| Self-ConsistencyBase model=GPT-OSS2026.05 | 39 | — |