Logical Reasoning on ProofWriter (test)
92.32AccuracyHBLR
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| HBLRModel=DeepSeek-R1, Prompting Strategy=HBLR2025.12 | 92.32 | — | |
| HBLRModel=DeepSeek-V3, Prompting Strategy=HBLR2025.12 | 89.48 | — | |
| HBLRModel=GPT-4, Prompting Strategy=HBLR2025.12 | 89.41 | — | |
| SymbCoTModel=DeepSeek-R1, Prompting Strategy=SymbCoT2025.12 | 88.34 | — | |
| HyperGuideBase model=Mistral2026.05 | 87 | — | |
| CoTModel=DeepSeek-R1, Prompting Strategy=CoT2025.12 | 86.27 | — | |
| LogicModel=DeepSeek-R1, Prompting Strategy=Logic2025.12 | 84.35 | — | |
| SymbCoTModel=DeepSeek-V3, Prompting Strategy=SymbCoT2025.12 | 84.15 | — | |
| LogicModel=GPT-4, Prompting Strategy=Logic2025.12 | 83.38 | — | |
| DirectModel=DeepSeek-R1, Prompting Strategy=Direct2025.12 | 82.48 | — | |
| LogicModel=DeepSeek-V3, Prompting Strategy=Logic2025.12 | 82.11 | — | |
| SymbCoTModel=GPT-4, Prompting Strategy=SymbCoT2025.12 | 79.34 | — | |
| DetermLRModel=GPT-42023.10 | 79.17 | 14.63 | |
| HyperGuideBase model=Qwen2.52026.05 | 77.4 | — | |
| SoftCoTBase model=Qwen2.52026.05 | 75 | — | |
| Self-ConsistencyBase model=Qwen2.52026.05 | 74 | — | |
| LAMBADAModel=GPT-42023.10 | 72 | 15.04 | |
| OVMBase model=Mistral2026.05 | 72 | — | |
| CRModel=GPT-42023.10 | 71.67 | 16.76 | |
| SIModel=GPT-42023.10 | 70.67 | 17.46 | |
| Few-shotBase model=Qwen2.52026.05 | 70.4 | — | |
| ToTModel=GPT-42023.10 | 70.33 | 24.57 | |
| COT-SCModel=GPT-4, n=162023.10 | 69.33 | 16 | |
| Tree of ThoughtsBase model=Qwen2.52026.05 | 69 | — | |
| DetermLRModel=GPT-3.5-turbo2023.10 | 68.83 | 16.52 | |
| CoTModel=GPT-4, Prompting Strategy=CoT2025.12 | 68.11 | — | |
| Self-ConsistencyBase model=Mistral2026.05 | 67.6 | — | |
| COTModel=GPT-42023.10 | 67.41 | 1 | |
| HyperGuideBase model=GPT-OSS2026.05 | 67 | — | |
| Tree of ThoughtsBase model=Mistral2026.05 | 66 | — | |
| HBLRModel=GPT-3.5-Turbo, Prompting Strategy=HBLR2025.12 | 63.24 | — | |
| SoftCoTBase model=GPT-OSS2026.05 | 61.4 | — | |
| CRModel=GPT-3.5-turbo2023.10 | 59.16 | 18.81 | |
| SymbCoTModel=GPT-3.5-Turbo, Prompting Strategy=SymbCoT2025.12 | 59.03 | — | |
| LogicModel=GPT-3.5-Turbo, Prompting Strategy=Logic2025.12 | 58.62 | — | |
| Few-shotBase model=Mistral2026.05 | 57 | — | |
| SoftCoTBase model=Mistral2026.05 | 57 | — | |
| CoTModel=DeepSeek-V3, Prompting Strategy=CoT2025.12 | 56.84 | — | |
| LAMBADAModel=GPT-3.5-turbo2023.10 | 55.17 | 16.89 | |
| DirectModel=DeepSeek-V3, Prompting Strategy=Direct2025.12 | 54.46 | — | |
| ToTModel=GPT-3.5-turbo2023.10 | 54.16 | 24.88 | |
| DirectModel=GPT-4, Prompting Strategy=Direct2025.12 | 52.67 | — | |
| SIModel=GPT-3.5-turbo2023.10 | 50.17 | 18.49 | |
| CoTModel=GPT-3.5-Turbo, Prompting Strategy=CoT2025.12 | 49.17 | — | |
| PT-SFTBase model=Qwen2.52026.05 | 49 | — | |
| COT-SCModel=GPT-3.5-turbo, n=162023.10 | 48.67 | 16 | |
| Self-ConsistencyBase model=GPT-OSS2026.05 | 47 | — | |
| StandardModel=GPT-42023.10 | 46.83 | 1 | |
| Tree of ThoughtsBase model=GPT-OSS2026.05 | 46 | — | |
| COTModel=GPT-3.5-turbo2023.10 | 45 | 1 | |
| PT-SFTBase model=Mistral2026.05 | 42.6 | — | |
| PT-SFTBase model=GPT-OSS2026.05 | 42.4 | — | |
| OVMBase model=Qwen2.52026.05 | 38 | — | |
| DirectModel=GPT-3.5-Turbo, Prompting Strategy=Direct2025.12 | 36.53 | — | |
| StandardModel=GPT-3.5-turbo2023.10 | 36.17 | 1 | |
| OVMBase model=GPT-OSS2026.05 | 33 | — | |
| Few-shotBase model=GPT-OSS2026.05 | 29.6 | — |