Proof writing on IMO-ProofBench
58.7Avg@3 Grade ScoreGemini 3 Pro
Evaluation Results
| Method | Links | |
|---|---|---|
| Gemini 3 ProTraining/Evaluation Protocol=proprietary2026.04 | 58.7 | |
| DeepSeek-Math-V2Model Size=685B2026.04 | 57.9 | |
| QED-Nano (+ RSA test-time scaffold)Model Size=4B, Training/Evaluation Protocol=RSA test-time scaffold2026.04 | 56.9 | |
| GPT-OSS-120BModel Size=120B2026.04 | 43.1 | |
| Nomos-12026.04 | 40.3 | |
| QED-NanoModel Size=4B2026.04 | 40 | |
| QED-Nano (SFT initialization only)Model Size=4B, Training/Evaluation Protocol=SFT initialization only2026.04 | 39.5 | |
| GPT-OSS-20BModel Size=20B2026.04 | 38.3 | |
| Qwen3-235B-A22B-Thinking-2507Model Size=235B-A22B2026.04 | 34.1 | |
| Qwen3-30B-A3B-Thinking-2507Model Size=30B-A3B2026.04 | 27.6 | |
| Qwen3-4B-Thinking-2507Model Size=4B2026.04 | 20.4 |