Deductive Reasoning on ProntoQA
99.2Pass@1Beam Search (BL)
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Beam Search (BL)Model=Gemma-3-12b-it2026.04 | 99.2 | — | |
| TTPOModel=Gemma-3-12b-it2026.04 | 99.2 | — | |
| Greedy (BL)Model=Gemma-3-12b-it2026.04 | 99 | — | |
| SteeringModel=Gemma-3-12b-it2026.04 | 99 | — | |
| BranchingModel=Gemma-3-12b-it2026.04 | 98.8 | — | |
| Greedy (BL)Model=Phi-4-reasoning-plus (13B)2026.04 | 97.8 | — | |
| SteeringModel=Phi-4-reasoning-plus (13B)2026.04 | 97.8 | — | |
| TTPOModel=Phi-4-reasoning-plus (13B)2026.04 | 97.6 | — | |
| TTPOModel=Phi-4-mini-instruct (4B)2026.04 | 96.6 | — | |
| BranchingModel=Phi-4-reasoning-plus (13B)2026.04 | 96 | — | |
| BranchingModel=Phi-4-mini-instruct (4B)2026.04 | 94.8 | — | |
| SteeringModel=Phi-4-mini-instruct (4B)2026.04 | 94.2 | — | |
| Beam Search (BL)Model=Phi-4-reasoning-plus (13B)2026.04 | 94.2 | — | |
| Greedy (BL)Model=Phi-4-mini-instruct (4B)2026.04 | 93.6 | — | |
| Beam Search (BL)Model=Phi-4-mini-instruct (4B)2026.04 | 93.2 | — | |
| Beam Search (BL)Model=Gemma-3-4b-it2026.04 | 90.6 | — | |
| TTPOModel=Gemma-3-4b-it2026.04 | 90.4 | — | |
| SteeringModel=Gemma-3-4b-it2026.04 | 90.2 | — | |
| BranchingModel=Gemma-3-4b-it2026.04 | 90.2 | — | |
| Greedy (BL)Model=Gemma-3-4b-it2026.04 | 90 | — | |
| RULEREASONER-8BMethod Category=RULEREASONER (Ours), Model Scale=8B2025.06 | 0.964 | 0.4 | |
| Easy-to-hard RLMethod Category=CURRICULUM LEARNING2025.06 | 0.962 | — | |
| SFT w/o CoTMethod Category=BEHAVIORAL CLONING, Reasoning Strategy=w/o CoT2025.06 | 0.96 | — | |
| Dr. GRPOMethod Category=ADVANCED RLVRS2025.06 | 0.96 | — | |
| DAPOMethod Category=ADVANCED RLVRS2025.06 | 0.96 | — | |
| ADARFTMethod Category=CURRICULUM LEARNING2025.06 | 0.96 | — | |
| Data-balance RLMethod Category=CURRICULUM LEARNING2025.06 | 0.958 | — | |
| SFT w/ Long CoTMethod Category=BEHAVIORAL CLONING, Reasoning Strategy=Long CoT2025.06 | 0.956 | — | |
| GRPOMethod Category=ADVANCED RLVRS2025.06 | 0.954 | — | |
| RULEREASONER-4BMethod Category=RULEREASONER (Ours), Model Scale=4B2025.06 | 0.95 | 0.6 | |
| RGFBMethod Category=PRIOR RBRS2025.06 | 0.94 | — | |
| OpenAI o3-miniMethod Category=FRONTIER REASONERS2025.06 | 0.94 | — | |
| Claude-3.7-SonnetMethod Category=FRONTIER REASONERS2025.06 | 0.928 | — | |
| SFT w/ Short CoTMethod Category=BEHAVIORAL CLONING, Reasoning Strategy=Short CoT2025.06 | 0.926 | — | |
| HtTMethod Category=PRIOR RBRS2025.06 | 0.92 | — | |
| Chain-of-LogicMethod Category=PRIOR RBRS2025.06 | 0.91 | — | |
| OpenAI o1Method Category=FRONTIER REASONERS2025.06 | 0.91 | — | |
| DeepSeek-R1Method Category=FRONTIER REASONERS2025.06 | 0.4 | — |