Mathematical Reasoning on AIME 25 (Accuracy, Tokens, Token Reduction %)
83.3AccuracyParallel Thinking
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Parallel ThinkingModel=DeepSeek-8B, Decoding Strategy=Parallel Thinking, Branching factor (n)=322026.06 | 83.3 | 25.2 | — | |
| Fork-thinkModel=DeepSeek-8B, Decoding Strategy=Fork-think, Branching factor (n)=32, Seed path length (l0)=2,0482026.06 | 82.4 | 21.4 | -14 | |
| GreedyModel=DeepSeek-8B, Decoding Strategy=Greedy2026.06 | 80 | 0.75 | — | |
| Fork-thinkModel=Qwen3-8B, Decoding Strategy=Fork-think, Branching factor (n)=32, Seed path length (l0)=2,0482026.06 | 75.6 | 15.9 | -27 | |
| Parallel ThinkingModel=Qwen3-8B, Decoding Strategy=Parallel Thinking, Branching factor (n)=322026.06 | 75.4 | 22 | — | |
| CoTModel=DeepSeek-8B, Decoding Strategy=CoT2026.06 | 75.3 | 0.8 | — | |
| Fork-thinkModel=Phi-4-RP-14B, Decoding Strategy=Fork-think, Branching factor (n)=32, Seed path length (l0)=2,0482026.06 | 66.6 | 12.5 | -7 | |
| CoTModel=Qwen3-8B, Decoding Strategy=CoT2026.06 | 64 | 0.7 | — | |
| Parallel ThinkingModel=Phi-4-RP-14B, Decoding Strategy=Parallel Thinking, Branching factor (n)=322026.06 | 62.6 | 13.5 | — | |
| GreedyModel=Qwen3-8B, Decoding Strategy=Greedy2026.06 | 60 | 0.62 | — | |
| CoTModel=Phi-4-RP-14B, Decoding Strategy=CoT2026.06 | 58.7 | 0.4 | — | |
| GreedyModel=Phi-4-RP-14B, Decoding Strategy=Greedy2026.06 | 46.6 | 0.66 | — |