Logical Reasoning on LSAT-AR
74.35AccuracyQwen3-4B-Thinking-2507 + Overthink
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Qwen3-4B-Thinking-2507 + OverthinkBase Model=Qwen3-4B-Thinking-2507, Protocol=Overthink2026.05 | 74.35 | 8,480 | 0.876 | |
| Qwen3-4B-Thinking-2507 + BETBase Model=Qwen3-4B-Thinking-2507, Protocol=BET2026.05 | 74.35 | 3,572 | 2.079 | |
| Qwen3-4B-Thinking-2507 + DEERBase Model=Qwen3-4B-Thinking-2507, Protocol=DEER2026.05 | 72.17 | 6,190 | 1.164 | |
| Qwen3-4B-Thinking-2507 + DR.SAFBase Model=Qwen3-4B-Thinking-2507, Protocol=DR.SAF2026.05 | 72.17 | 3,929 | 1.834 | |
| Qwen3-4B-Thinking-2507Base Model=Qwen3-4B-Thinking-2507, Protocol=Vanilla2026.05 | 71.74 | 7,164 | 1 | |
| Qwen3-4B-Thinking-2507 + Length-PenaltyBase Model=Qwen3-4B-Thinking-2507, Protocol=Length-Penalty2026.05 | 69.13 | 4,462 | 1.547 | |
| Qwen3-4B-Thinking-2507 + ThinkSwitcherBase Model=Qwen3-4B-Thinking-2507, Protocol=ThinkSwitcher2026.05 | 63.91 | 6,256 | 1.02 | |
| Qwen3-4B-Thinking-2507 + VeriThinkerBase Model=Qwen3-4B-Thinking-2507, Protocol=VeriThinker2026.05 | 60 | 4,476 | 1.339 | |
| DeepSeek-R1-Distill-Qwen-14B + DR.SAFBase Model=DeepSeek-R1-Distill-Qwen-14B, Protocol=DR.SAF2026.05 | 49.13 | 3,508 | 1.946 | |
| DeepSeek-R1-Distill-Qwen-14B + BETBase Model=DeepSeek-R1-Distill-Qwen-14B, Protocol=BET2026.05 | 48.7 | 2,684 | 2.521 | |
| Qwen3-4B-Thinking-2507 + DiffAdaptBase Model=Qwen3-4B-Thinking-2507, Protocol=DiffAdapt2026.05 | 48.26 | 3,838 | 1.256 | |
| DeepSeek-R1-Distill-Qwen-14B + Length-PenaltyBase Model=DeepSeek-R1-Distill-Qwen-14B, Protocol=Length-Penalty2026.05 | 46.96 | 4,813 | 1.356 | |
| DeepSeek-R1-Distill-Qwen-14BBase Model=DeepSeek-R1-Distill-Qwen-14B, Protocol=Vanilla2026.05 | 42.61 | 5,921 | 1 | |
| DeepSeek-R1-Distill-Qwen-7B + DR.SAFBase Model=DeepSeek-R1-Distill-Qwen-7B, Protocol=DR.SAF2026.05 | 35.22 | 4,463 | 1.777 | |
| DeepSeek-R1-Distill-Qwen-7B + BETBase Model=DeepSeek-R1-Distill-Qwen-7B, Protocol=BET2026.05 | 34.78 | 2,146 | 3.65 | |
| DeepSeek-R1-Distill-Qwen-7B + Length-PenaltyBase Model=DeepSeek-R1-Distill-Qwen-7B, Protocol=Length-Penalty2026.05 | 32.6 | 5,871 | 1.251 | |
| DeepSeek-R1-Distill-Qwen-7BBase Model=DeepSeek-R1-Distill-Qwen-7B, Protocol=Vanilla2026.05 | 30.43 | 6,854 | 1 | |
| DeepSeek-R1-Distill-Qwen-7B + VeriThinkerBase Model=DeepSeek-R1-Distill-Qwen-7B, Protocol=VeriThinker2026.05 | 30.43 | 6,186 | 1.108 | |
| DeepSeek-R1-Distill-Qwen-7B + ThinkSwitcherBase Model=DeepSeek-R1-Distill-Qwen-7B, Protocol=ThinkSwitcher2026.05 | 29.57 | 6,641 | 1.003 | |
| DeepSeek-R1-Distill-Qwen-7B + OverthinkBase Model=DeepSeek-R1-Distill-Qwen-7B, Protocol=Overthink2026.05 | 29.57 | 6,479 | 1.028 | |
| DeepSeek-R1-Distill-Qwen-7B + DEERBase Model=DeepSeek-R1-Distill-Qwen-7B, Protocol=DEER2026.05 | 28.7 | 5,609 | 1.152 | |
| DeepSeek-R1-Distill-Qwen-7B + DiffAdaptBase Model=DeepSeek-R1-Distill-Qwen-7B, Protocol=DiffAdapt2026.05 | 23.48 | 5,042 | 1.049 |