Question Answering on GPQA Diamond (Accuracy, Tokens, Token Reduction)
70.1AccuracyFork-think
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Fork-thinkModel=Phi-4-RP-14B, Decoding Strategy=Fork-think, Branching factor (n)=32, Seed path length (l0)=2,0482026.06 | 70.1 | 53.1 | -3 | |
| Parallel ThinkingModel=Phi-4-RP-14B, Decoding Strategy=Parallel Thinking, Branching factor (n)=322026.06 | 67.3 | 54.9 | — | |
| Parallel ThinkingModel=DeepSeek-8B, Decoding Strategy=Parallel Thinking, Branching factor (n)=322026.06 | 66.9 | 72.4 | — | |
| Fork-thinkModel=Qwen3-8B, Decoding Strategy=Fork-think, Branching factor (n)=32, Seed path length (l0)=2,0482026.06 | 63.2 | 60.4 | -27 | |
| Parallel ThinkingModel=Qwen3-8B, Decoding Strategy=Parallel Thinking, Branching factor (n)=322026.06 | 62.1 | 83.3 | — | |
| Fork-thinkModel=DeepSeek-8B, Decoding Strategy=Fork-think, Branching factor (n)=32, Seed path length (l0)=2,0482026.06 | 62.1 | 54.2 | -25 | |
| GreedyModel=DeepSeek-8B, Decoding Strategy=Greedy2026.06 | 61.1 | 2.3 | — | |
| CoTModel=DeepSeek-8B, Decoding Strategy=CoT2026.06 | 60.8 | 2.5 | — | |
| CoTModel=Qwen3-8B, Decoding Strategy=CoT2026.06 | 55.2 | 2.6 | — | |
| CoTModel=Phi-4-RP-14B, Decoding Strategy=CoT2026.06 | 53.9 | 1.7 | — | |
| GreedyModel=Qwen3-8B, Decoding Strategy=Greedy2026.06 | 39.8 | 3.3 | — | |
| GreedyModel=Phi-4-RP-14B, Decoding Strategy=Greedy2026.06 | 31.3 | 3.6 | — |