Reasoning on StrategyQA
83.5AccuracyCoT-SC
Evaluation Results
| Method | Links | |
|---|---|---|
| CoT-SCModel=GPT-42023.11 | 83.5 | |
| StrategyLLM-SCModel=GPT-42023.11 | 83.5 | |
| JT-Safe-V2-35BParameters=35B2026.05 | 83.19 | |
| StrategyLLMModel=GPT-42023.11 | 81.5 | |
| CoTModel=GPT-42023.11 | 80.5 | |
| StrategyLLM-SCModel=Claude-3-Sonnet2023.11 | 77 | |
| SolutionLLMModel=GPT-42023.11 | 75.5 | |
| CoT-SCModel=Claude-3-Sonnet2023.11 | 75 | |
| StrategyLLMModel=Claude-3-Sonnet2023.11 | 75 | |
| SolutionLLMModel=Claude-3-Sonnet2023.11 | 73.5 | |
| RM-RegenBase Model=Llama 3.1-8B2026.03 | 73 | |
| RM-RegenBase Model=GPT-3.52026.03 | 69.75 | |
| CoTModel=Claude-3-Sonnet2023.11 | 69 | |
| ST CoTBase Model=Llama 3.1-8B, Iterations=22026.03 | 68.75 | |
| ST CoTBase Model=Llama 3.1-8B, Iterations=32026.03 | 68 | |
| Zero-shot-CPPrompting Strategy=Zero-shot-CP, Base Model=GPT-3.5-Turbo, Number of runs=52024.03 | 67.39 | |
| ST CoTBase Model=Llama 3.1-8B, Iterations=42026.03 | 66.5 | |
| SPEARModel=Qwen2.5-1.5B2026.05 | 66.08 | |
| ST CoTBase Model=Llama 3.1-8B, Iterations=52026.03 | 66 | |
| SPEARModel=Llama3.2-3B2026.05 | 65.84 | |
| Gen-SSD2026.04 | 65.68 | |
| TILRBackbone=GPT-2, Average generation tokens per inference=22026.06 | 65.5 | |
| Zero-shotPrompting Strategy=Zero-shot, Base Model=GPT-3.5-Turbo, Number of runs=52024.03 | 65.02 | |
| MoRSD2026.04 | 64.93 | |
| ProCoBase Model=Llama 3.1-8B, Iterations=42026.03 | 64.75 | |
| MCC-KD2026.04 | 64.37 | |
| Standard KD2026.04 | 63.23 | |
| RefinementBackbone=GPT-2, Average generation tokens per inference=22026.06 | 62.4 | |
| SOTA with Equivalent ParametersModel comparison=Equivalent Parameters2026.05 | 62.36 | |
| ST CoTBase Model=GPT-3.5, Iterations=22026.03 | 62 | |
| ST CoTBase Model=GPT-3.5, Iterations=32026.03 | 62 | |
| ProCoBase Model=Llama 3.1-8B, Iterations=32026.03 | 62 | |
| ProCoBase Model=Llama 3.1-8B, Iterations=52026.03 | 61.25 | |
| Zero-shot-CoTPrompting Strategy=Zero-shot-CoT, Base Model=GPT-3.5-Turbo, Number of runs=52024.03 | 60.74 | |
| ST CoTBase Model=GPT-3.5, Iterations=42026.03 | 60.25 | |
| ProCoBase Model=GPT-3.5, Iterations=32026.03 | 60 | |
| CoconutBackbone=GPT-2, Average generation tokens per inference=22026.06 | 60 | |
| ProCoBase Model=Llama 3.1-8B, Iterations=22026.03 | 59.75 | |
| ProCoBase Model=GPT-3.5, Iterations=22026.03 | 56.75 | |
| AdaAnchorBackbone=GPT-2, Average generation tokens per inference=22026.06 | 56.1 | |
| ProCoBase Model=GPT-3.5, Iterations=42026.03 | 55.75 | |
| CoTBackbone=GPT-2, Average generation tokens per inference=222026.06 | 55.2 | |
| Self-RefineBase Model=Llama 3.1-8B, Iterations=42026.03 | 55 | |
| RLTF-SDModel=Qwen2.5-1.5B2026.05 | 54.29 | |
| Self-RefineBase Model=Llama 3.1-8B, Iterations=22026.03 | 54.25 | |
| Self-RefineBase Model=Llama 3.1-8B, Iterations=32026.03 | 53.75 | |
| No-CoTBackbone=GPT-2, Average generation tokens per inference=12026.06 | 52.3 | |
| RLTF-SDModel=Llama3.2-3B2026.05 | 51.58 | |
| GRPOModel=Qwen2.5-1.5B2026.05 | 51.38 | |
| Self-RefineBase Model=Llama 3.1-8B, Iterations=52026.03 | 51.25 | |
| Feedback SFTModel=Qwen2.5-1.5B2026.05 | 50.51 | |
| Feedback SFTModel=Llama3.2-3B2026.05 | 49.49 | |
| GRPOModel=Llama3.2-3B2026.05 | 49.34 | |
| Self-RefineBase Model=GPT-3.5, Iterations=42026.03 | 44.75 | |
| Self-RefineBase Model=GPT-3.5, Iterations=22026.03 | 44 | |
| Self-RefineBase Model=GPT-3.5, Iterations=32026.03 | 41.75 | |
| OPSDModel=Llama3.2-3B2026.05 | 0 | |
| OPSDModel=Qwen2.5-1.5B2026.05 | 0 |