Multi-step mathematical reasoning on We-Math (test)
72.8S1 ScoreGPT-4o
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| GPT-4o#Params=-2025.05 | 72.8 | 58.1 | 43.6 | |
| GPT-4V#Params=-2025.05 | 65.5 | 49.2 | 38.2 | |
| MathCoder-VL-8B#Params=8B2025.05 | 65.4 | 58.6 | 52.1 | |
| InternVL2-76B#Params=76B2025.05 | 65.2 | 49.4 | 49.1 | |
| InternVL2-8B#Params=8B2025.05 | 59.4 | 43.6 | 35.2 | |
| Qwen2-VL#Params=8B2025.05 | 59.1 | 43.6 | 26.7 | |
| MAVIS-7B#Params=7B2025.05 | 57.2 | 37.9 | 34.6 | |
| Gemini-1.5-Pro#Params=-2025.05 | 56.1 | 51.4 | 33.9 | |
| Math-PUMA-Qwen2#Params=8B2025.05 | 53.3 | 39.4 | 36.4 | |
| MathCoder-VL-2B#Params=2B2025.05 | 52 | 42.2 | 38.8 | |
| InternVL2-26B#Params=26B2025.05 | 51 | 39.2 | 46.1 | |
| IXC-2-VL#Params=7B2025.05 | 47 | 33.1 | 33.3 | |
| Math-PUMA-DS#Params=7B2025.05 | 45.6 | 38.1 | 33.9 | |
| IXC-2.5-Reward#Params=7B2025.05 | 44.4 | 35.3 | 27.9 | |
| Qwen-VL-Max#Params=-2025.05 | 40.8 | 30.3 | 20.6 | |
| Math-LLaVA-13B#Params=13B2025.05 | 38.7 | 34.2 | 34.6 | |
| LLaVA-1.5-13B#Params=13B2025.05 | 35.4 | 30 | 32.7 | |
| InternVL-Chat-2B-V1-5#Params=2B2025.05 | 34.3 | 26.1 | 20 | |
| Deepseek-VL#Params=8B2025.05 | 32.6 | 26.7 | 25.5 | |
| G-LLaVA-7B#Params=7B2025.05 | 32.4 | 30.6 | 32.7 |