Mathematical Reasoning on MathVista (Pass@k Metrics)
86.1ALG Pass@1Canvas-of-thought
Evaluation Results
| Method | Links | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| Canvas-of-thoughtModel Name=Gemini-2.5-pro, # Budget=62026.02 | 86.1 | 87.6 | 89.9 | 92.3 | 70.9 | 72.1 | 84.6 | 87.3 | |
| Chain-of-thoughtModel Name=Gemini-2.5-pro, # Budget=12026.02 | 84.7 | 86.3 | 86.1 | 87.6 | 69.3 | 70.9 | 84.4 | 86.4 | |
| Tree-of-thoughtModel Name=Gemini-2.5-pro, # Budget=102026.02 | 84.5 | 86.2 | 86 | 87.7 | 69.9 | 71.4 | 84.5 | 86.6 | |
| Iterative ReflectionModel Name=Gemini-2.5-pro, # Budget=62026.02 | 84.2 | 86 | 85.7 | 87.6 | 70.8 | 71.2 | 84.6 | 87 | |
| Canvas-of-thoughtModel Name=GPT-5, # Budget=62026.02 | 83.9 | 87.1 | 82.2 | 87 | 62 | 66.5 | 72.5 | 77.7 | |
| Program-of-thoughtModel Name=Gemini-2.5-pro, # Budget=62026.02 | 83.9 | 85.8 | 85.8 | 87.6 | 70.2 | 71 | 84.7 | 86.8 | |
| Chain-of-thoughtModel Name=GPT-5, # Budget=12026.02 | 82.9 | 86.5 | 83.2 | 86.5 | 54.8 | 57.5 | 72.1 | 75.8 | |
| Tree-of-thoughtModel Name=GPT-5, # Budget=102026.02 | 82.8 | 86.4 | 83.1 | 86.8 | 58.4 | 62.9 | 72.2 | 76.2 | |
| Iterative ReflectionModel Name=GPT-5, # Budget=62026.02 | 82.4 | 86.3 | 82.8 | 86.9 | 59.1 | 63.3 | 72.3 | 76.9 | |
| Program-of-thoughtModel Name=GPT-5, # Budget=62026.02 | 82.2 | 86.3 | 83 | 86.8 | 57.8 | 61.1 | 72.4 | 76.5 |