Multi-step Narrative Reasoning on MUSR (ACC, TOK, η)
65.86AccuracyQwen3-4B-Thinking-2507 + DiffAdapt
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Qwen3-4B-Thinking-2507 + DiffAdaptBase Model=Qwen3-4B-Thinking-2507, Protocol=DiffAdapt2026.05 | 65.86 | 6,841 | 0.909 | |
| Qwen3-4B-Thinking-2507Base Model=Qwen3-4B-Thinking-2507, Protocol=Vanilla2026.05 | 64.42 | 6,081 | 1 | |
| Qwen3-4B-Thinking-2507 + BETBase Model=Qwen3-4B-Thinking-2507, Protocol=BET2026.05 | 63.84 | 1,188 | 5.073 | |
| Qwen3-4B-Thinking-2507 + DEERBase Model=Qwen3-4B-Thinking-2507, Protocol=DEER2026.05 | 62.02 | 5,351 | 1.094 | |
| Qwen3-4B-Thinking-2507 + DR.SAFBase Model=Qwen3-4B-Thinking-2507, Protocol=DR.SAF2026.05 | 61.31 | 2,118 | 2.732 | |
| Qwen3-4B-Thinking-2507 + Length-PenaltyBase Model=Qwen3-4B-Thinking-2507, Protocol=Length-Penalty2026.05 | 57.88 | 2,529 | 2.16 | |
| Qwen3-4B-Thinking-2507 + OverthinkBase Model=Qwen3-4B-Thinking-2507, Protocol=Overthink2026.05 | 57.78 | 4,103 | 1.329 | |
| Qwen3-4B-Thinking-2507 + VeriThinkerBase Model=Qwen3-4B-Thinking-2507, Protocol=VeriThinker2026.05 | 47.68 | 3,906 | 1.152 | |
| DeepSeek-R1-Distill-Qwen-14BBase Model=DeepSeek-R1-Distill-Qwen-14B, Protocol=Vanilla2026.05 | 46.3 | 4,633 | 1 | |
| DeepSeek-R1-Distill-Qwen-7BBase Model=DeepSeek-R1-Distill-Qwen-7B, Protocol=Vanilla2026.05 | 44.44 | 5,173 | 1 | |
| Qwen3-4B-Thinking-2507 + ThinkSwitcherBase Model=Qwen3-4B-Thinking-2507, Protocol=ThinkSwitcher2026.05 | 41.67 | 3,238 | 1.215 | |
| DeepSeek-R1-Distill-Qwen-14B + BETBase Model=DeepSeek-R1-Distill-Qwen-14B, Protocol=BET2026.05 | 41.01 | 2,216 | 1.852 | |
| DeepSeek-R1-Distill-Qwen-7B + BETBase Model=DeepSeek-R1-Distill-Qwen-7B, Protocol=BET2026.05 | 40.4 | 2,764 | 1.701 | |
| DeepSeek-R1-Distill-Qwen-7B + DR.SAFBase Model=DeepSeek-R1-Distill-Qwen-7B, Protocol=DR.SAF2026.05 | 40.2 | 4,229 | 1.107 | |
| DeepSeek-R1-Distill-Qwen-14B + DR.SAFBase Model=DeepSeek-R1-Distill-Qwen-14B, Protocol=DR.SAF2026.05 | 39.29 | 3,393 | 1.159 | |
| DeepSeek-R1-Distill-Qwen-7B + Length-PenaltyBase Model=DeepSeek-R1-Distill-Qwen-7B, Protocol=Length-Penalty2026.05 | 39.09 | 4,783 | 0.951 | |
| DeepSeek-R1-Distill-Qwen-7B + OverthinkBase Model=DeepSeek-R1-Distill-Qwen-7B, Protocol=Overthink2026.05 | 38.99 | 4,952 | 0.917 | |
| DeepSeek-R1-Distill-Qwen-14B + Length-PenaltyBase Model=DeepSeek-R1-Distill-Qwen-14B, Protocol=Length-Penalty2026.05 | 38.36 | 3,574 | 1.074 | |
| DeepSeek-R1-Distill-Qwen-7B + DiffAdaptBase Model=DeepSeek-R1-Distill-Qwen-7B, Protocol=DiffAdapt2026.05 | 36.16 | 5,716 | 0.736 | |
| DeepSeek-R1-Distill-Qwen-7B + DEERBase Model=DeepSeek-R1-Distill-Qwen-7B, Protocol=DEER2026.05 | 34.34 | 4,008 | 0.997 | |
| DeepSeek-R1-Distill-Qwen-7B + VeriThinkerBase Model=DeepSeek-R1-Distill-Qwen-7B, Protocol=VeriThinker2026.05 | 29.9 | 4,869 | 0.715 | |
| DeepSeek-R1-Distill-Qwen-7B + ThinkSwitcherBase Model=DeepSeek-R1-Distill-Qwen-7B, Protocol=ThinkSwitcher2026.05 | 23.13 | 4,619 | 0.583 |