Long-CoT Reasoning on GPQA
80.8AccuracySlow
Evaluation Results
| Method | Links | |
|---|---|---|
| SlowModel=Qwen3-235B-A22B-Thinking-2507, Decoding Strategy=Slow2026.03 | 80.8 | |
| SFIModel=Qwen3-235B-A22B-Thinking-2507, Decoding Strategy=SFI (Ours)2026.03 | 80.8 | |
| SFIModel=Qwen3-30B-A3B-Thinking-2507, Decoding Strategy=SFI (Ours)2026.03 | 71.21 | |
| SlowModel=Qwen3-30B-A3B-Thinking-2507, Decoding Strategy=Slow2026.03 | 69.7 | |
| SlowModel=Qwen3-4B-Thinking-2507, Decoding Strategy=Slow2026.03 | 64.14 | |
| SFIModel=Qwen3-4B-Thinking-2507, Decoding Strategy=SFI (Ours)2026.03 | 63.7 |