Scientific Question Answering on GPQA Diamond (ACC, TOK, η)
53.03Accuracy (ACC)Qwen3-4B-Thinking-2507 + DEER
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Qwen3-4B-Thinking-2507 + DEERBase Model=Qwen3-4B-Thinking-2507, Protocol=DEER2026.05 | 53.03 | 4,762 | 1.414 | |
| Qwen3-4B-Thinking-2507 + BETBase Model=Qwen3-4B-Thinking-2507, Protocol=BET2026.05 | 52.53 | 3,835 | 1.739 | |
| Qwen3-4B-Thinking-2507Base Model=Qwen3-4B-Thinking-2507, Protocol=Vanilla2026.05 | 51.01 | 6,475 | 1 | |
| Qwen3-4B-Thinking-2507 + VeriThinkerBase Model=Qwen3-4B-Thinking-2507, Protocol=VeriThinker2026.05 | 50.1 | 5,986 | 1.062 | |
| Qwen3-4B-Thinking-2507 + DR.SAFBase Model=Qwen3-4B-Thinking-2507, Protocol=DR.SAF2026.05 | 50 | 4,237 | 1.498 | |
| Qwen3-4B-Thinking-2507 + Length-PenaltyBase Model=Qwen3-4B-Thinking-2507, Protocol=Length-Penalty2026.05 | 49.49 | 4,705 | 1.335 | |
| Qwen3-4B-Thinking-2507 + OverthinkBase Model=Qwen3-4B-Thinking-2507, Protocol=Overthink2026.05 | 48.99 | 7,568 | 0.822 | |
| Qwen3-4B-Thinking-2507 + DiffAdaptBase Model=Qwen3-4B-Thinking-2507, Protocol=DiffAdapt2026.05 | 48.48 | 5,961 | 1.032 | |
| DeepSeek-R1-Distill-Qwen-14BBase Model=DeepSeek-R1-Distill-Qwen-14B, Protocol=Vanilla2026.05 | 47.58 | 4,294 | 1 | |
| DeepSeek-R1-Distill-Qwen-14B + BETBase Model=DeepSeek-R1-Distill-Qwen-14B, Protocol=BET2026.05 | 46.57 | 2,023 | 2.078 | |
| DeepSeek-R1-Distill-Qwen-14B + DR.SAFBase Model=DeepSeek-R1-Distill-Qwen-14B, Protocol=DR.SAF2026.05 | 45.86 | 2,643 | 1.566 | |
| DeepSeek-R1-Distill-Qwen-14B + Length-PenaltyBase Model=DeepSeek-R1-Distill-Qwen-14B, Protocol=Length-Penalty2026.05 | 45.66 | 3,891 | 1.059 | |
| Qwen3-4B-Thinking-2507 + ThinkSwitcherBase Model=Qwen3-4B-Thinking-2507, Protocol=ThinkSwitcher2026.05 | 45.45 | 6,420 | 0.899 | |
| DeepSeek-R1-Distill-Qwen-7B + VeriThinkerBase Model=DeepSeek-R1-Distill-Qwen-7B, Protocol=VeriThinker2026.05 | 43.33 | 3,948 | 1.22 | |
| DeepSeek-R1-Distill-Qwen-7B + DEERBase Model=DeepSeek-R1-Distill-Qwen-7B, Protocol=DEER2026.05 | 42.32 | 3,773 | 1.247 | |
| DeepSeek-R1-Distill-Qwen-7B + OverthinkBase Model=DeepSeek-R1-Distill-Qwen-7B, Protocol=Overthink2026.05 | 41.82 | 4,351 | 1.068 | |
| DeepSeek-R1-Distill-Qwen-7BBase Model=DeepSeek-R1-Distill-Qwen-7B, Protocol=Vanilla2026.05 | 41.31 | 4,592 | 1 | |
| DeepSeek-R1-Distill-Qwen-7B + DiffAdaptBase Model=DeepSeek-R1-Distill-Qwen-7B, Protocol=DiffAdapt2026.05 | 40.3 | 4,012 | 1.117 | |
| DeepSeek-R1-Distill-Qwen-7B + BETBase Model=DeepSeek-R1-Distill-Qwen-7B, Protocol=BET2026.05 | 39.8 | 1,987 | 2.227 | |
| DeepSeek-R1-Distill-Qwen-7B + DR.SAFBase Model=DeepSeek-R1-Distill-Qwen-7B, Protocol=DR.SAF2026.05 | 39.3 | 2,816 | 1.551 | |
| DeepSeek-R1-Distill-Qwen-7B + Length-PenaltyBase Model=DeepSeek-R1-Distill-Qwen-7B, Protocol=Length-Penalty2026.05 | 38.28 | 4,540 | 0.937 | |
| DeepSeek-R1-Distill-Qwen-7B + ThinkSwitcherBase Model=DeepSeek-R1-Distill-Qwen-7B, Protocol=ThinkSwitcher2026.05 | 37.27 | 3,889 | 1.065 |