Mathematical Reasoning on AIME24 (Accuracy, Tokens, Token Reduction %)
86AccuracyParallel Thinking
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Parallel ThinkingModel=DeepSeek-8B, Decoding Strategy=Parallel Thinking, Branching factor (n)=32, Aggregation Strategy=majority voting2026.06 | 86 | 22.6 | — | |
| Fork-thinkModel=DeepSeek-8B, Decoding Strategy=Fork-think, Branching factor (n)=32, Aggregation Strategy=majority voting, Seed path length (l0)=2,0482026.06 | 86 | 20.3 | -10 | |
| Fork-thinkModel=Qwen3-8B, Decoding Strategy=Fork-think, Branching factor (n)=32, Aggregation Strategy=majority voting, Seed path length (l0)=2,0482026.06 | 81.8 | 14.2 | -30 | |
| Parallel ThinkingModel=Qwen3-8B, Decoding Strategy=Parallel Thinking, Branching factor (n)=32, Aggregation Strategy=majority voting2026.06 | 80.1 | 20.2 | — | |
| GreedyModel=DeepSeek-8B, Decoding Strategy=Greedy, Aggregation Strategy=majority voting2026.06 | 80 | 0.68 | — | |
| CoTModel=DeepSeek-8B, Decoding Strategy=CoT, Aggregation Strategy=majority voting2026.06 | 80 | 0.7 | — | |
| Parallel ThinkingModel=Phi-4-RP-14B, Decoding Strategy=Parallel Thinking, Branching factor (n)=32, Aggregation Strategy=majority voting2026.06 | 77.8 | 12.2 | — | |
| Fork-thinkModel=Phi-4-RP-14B, Decoding Strategy=Fork-think, Branching factor (n)=32, Aggregation Strategy=majority voting, Seed path length (l0)=2,0482026.06 | 77 | 11.2 | -8 | |
| CoTModel=Qwen3-8B, Decoding Strategy=CoT, Aggregation Strategy=majority voting2026.06 | 68.7 | 0.6 | — | |
| CoTModel=Phi-4-RP-14B, Decoding Strategy=CoT, Aggregation Strategy=majority voting2026.06 | 67.3 | 0.4 | — | |
| GreedyModel=Qwen3-8B, Decoding Strategy=Greedy, Aggregation Strategy=majority voting2026.06 | 50 | 0.7 | — | |
| GreedyModel=Phi-4-RP-14B, Decoding Strategy=Greedy, Aggregation Strategy=majority voting2026.06 | 43.3 | 0.73 | — |