Mathematical Visual Reasoning on Math-Reasoning I
96.62AccuracyBenchmark Agent
Evaluation Results
| Method | Links | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Benchmark Agentevaluation_type=Human & LLM-as-Judge Quality Eval, backbone=GPT-5.12026.06 | 96.62 | 79.69 | 94.72 | 95.44 | 87.58 | 68.08 | 45.13 | 77.79 | — | — | — | — | |
| Qwen3.5evaluation_type=Consistency Evaluation2026.06 | — | — | — | — | — | — | — | — | 45.26 | 47.19 | 48.88 | 54.49 |