Out-of-Domain Reasoning Aggregation on OOD Average
63.57AccuracyQwen3-4B-Thinking-2507 + BET
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Qwen3-4B-Thinking-2507 + BETBase Model=Qwen3-4B-Thinking-2507, Protocol=BET2026.05 | 63.57 | 2,865 | 2.338 | |
| Qwen3-4B-Thinking-2507 + DEERBase Model=Qwen3-4B-Thinking-2507, Protocol=DEER2026.05 | 62.41 | 5,434 | 1.224 | |
| Qwen3-4B-Thinking-2507Base Model=Qwen3-4B-Thinking-2507, Protocol=Vanilla2026.05 | 62.39 | 6,573 | 1 | |
| Qwen3-4B-Thinking-2507 + DR.SAFBase Model=Qwen3-4B-Thinking-2507, Protocol=DR.SAF2026.05 | 61.16 | 3,428 | 1.88 | |
| Qwen3-4B-Thinking-2507 + OverthinkBase Model=Qwen3-4B-Thinking-2507, Protocol=Overthink2026.05 | 60.37 | 6,717 | 0.947 | |
| Qwen3-4B-Thinking-2507 + Length-PenaltyBase Model=Qwen3-4B-Thinking-2507, Protocol=Length-Penalty2026.05 | 58.83 | 3,899 | 1.59 | |
| Qwen3-4B-Thinking-2507 + DiffAdaptBase Model=Qwen3-4B-Thinking-2507, Protocol=DiffAdapt2026.05 | 54.2 | 5,547 | 1.03 | |
| Qwen3-4B-Thinking-2507 + VeriThinkerBase Model=Qwen3-4B-Thinking-2507, Protocol=VeriThinker2026.05 | 52.59 | 4,789 | 1.157 | |
| Qwen3-4B-Thinking-2507 + ThinkSwitcherBase Model=Qwen3-4B-Thinking-2507, Protocol=ThinkSwitcher2026.05 | 50.34 | 5,305 | 1 | |
| DeepSeek-R1-Distill-Qwen-14BBase Model=DeepSeek-R1-Distill-Qwen-14B, Protocol=Vanilla2026.05 | 45.5 | 4,949 | 1 | |
| DeepSeek-R1-Distill-Qwen-14B + BETBase Model=DeepSeek-R1-Distill-Qwen-14B, Protocol=BET2026.05 | 45.43 | 2,308 | 2.141 | |
| DeepSeek-R1-Distill-Qwen-14B + DR.SAFBase Model=DeepSeek-R1-Distill-Qwen-14B, Protocol=DR.SAF2026.05 | 44.76 | 3,181 | 1.531 | |
| DeepSeek-R1-Distill-Qwen-14B + Length-PenaltyBase Model=DeepSeek-R1-Distill-Qwen-14B, Protocol=Length-Penalty2026.05 | 43.66 | 4,093 | 1.16 | |
| DeepSeek-R1-Distill-Qwen-7BBase Model=DeepSeek-R1-Distill-Qwen-7B, Protocol=Vanilla2026.05 | 38.73 | 5,540 | 1 | |
| DeepSeek-R1-Distill-Qwen-7B + BETBase Model=DeepSeek-R1-Distill-Qwen-7B, Protocol=BET2026.05 | 38.33 | 2,299 | 2.385 | |
| DeepSeek-R1-Distill-Qwen-7B + DR.SAFBase Model=DeepSeek-R1-Distill-Qwen-7B, Protocol=DR.SAF2026.05 | 38.24 | 3,836 | 1.426 | |
| DeepSeek-R1-Distill-Qwen-7B + OverthinkBase Model=DeepSeek-R1-Distill-Qwen-7B, Protocol=Overthink2026.05 | 36.79 | 5,261 | 1 | |
| DeepSeek-R1-Distill-Qwen-7B + Length-PenaltyBase Model=DeepSeek-R1-Distill-Qwen-7B, Protocol=Length-Penalty2026.05 | 36.66 | 5,065 | 1.035 | |
| DeepSeek-R1-Distill-Qwen-7B + DEERBase Model=DeepSeek-R1-Distill-Qwen-7B, Protocol=DEER2026.05 | 35.12 | 4,463 | 1.132 | |
| DeepSeek-R1-Distill-Qwen-7B + VeriThinkerBase Model=DeepSeek-R1-Distill-Qwen-7B, Protocol=VeriThinker2026.05 | 34.55 | 5,001 | 0.988 | |
| DeepSeek-R1-Distill-Qwen-7B + DiffAdaptBase Model=DeepSeek-R1-Distill-Qwen-7B, Protocol=DiffAdapt2026.05 | 33.31 | 4,923 | 0.968 | |
| DeepSeek-R1-Distill-Qwen-7B + ThinkSwitcherBase Model=DeepSeek-R1-Distill-Qwen-7B, Protocol=ThinkSwitcher2026.05 | 29.99 | 5,050 | 0.85 |