Reasoning on MathArena Apex (test)
0.3542AccuracyEmpirical-MCTS
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Empirical-MCTSBase Model=Gemini 3 Flash2026.02 | 0.3542 | 5.24 | |
| Gemini 3 Pro2026.02 | 0.2344 | 3.4 | |
| Gemini 3 Flash2026.02 | 0.1979 | 1.51 | |
| GPT-5.2 (High)2026.02 | 0.1354 | 12 | |
| Grok 4 Fast (Reasoning)2026.02 | 0.0521 | 0.16 | |
| Grok 42026.02 | 0.0208 | 6.21 | |
| Claude Sonnet 4.52026.02 | 0.0156 | 4.56 | |
| GPT-5.1 (High)2026.02 | 0.0104 | 6.58 | |
| DeepSeek-R1-05282026.02 | 0.0104 | 0.98 | |
| Gemini 2.5 Pro2026.02 | 0.0052 | 3.74 |