Multimodal Mathematical Reasoning on DynaMath worst case
37.7AccuracyInternVL3.5-8B
Evaluation Results
| Method | Links | |
|---|---|---|
| InternVL3.5-8BModel Scale=8B, Optimization Method=N/A2026.06 | 37.7 | |
| InternVL3.5-4B (+OPD)Model Scale=4B, Optimization Method=OPD2026.06 | 32.7 | |
| InternVL3.5-4B (+GNDPO)Model Scale=4B, Optimization Method=GNDPO2026.06 | 32.5 | |
| InternVL3.5-4B (+GSPO)Model Scale=4B, Optimization Method=GSPO2026.06 | 31.1 | |
| InternVL3.5-4B (Base (Instruct))Model Scale=4B, Optimization Method=Base (Instruct)2026.06 | 30.7 | |
| InternVL3.5-2B (+OPD)Model Scale=2B, Optimization Method=OPD2026.06 | 24.4 | |
| InternVL3.5-2B (+GNDPO)Model Scale=2B, Optimization Method=GNDPO2026.06 | 23.8 | |
| InternVL3.5-2B (Base (Instruct))Model Scale=2B, Optimization Method=Base (Instruct)2026.06 | 19.8 | |
| InternVL3.5-2B (+GSPO)Model Scale=2B, Optimization Method=GSPO2026.06 | 19.8 | |
| InternVL3.5-1B (+GNDPO)Model Scale=1B, Optimization Method=GNDPO2026.06 | 12 | |
| InternVL3.5-1B (+OPD)Model Scale=1B, Optimization Method=OPD2026.06 | 11.8 | |
| InternVL3.5-1B (Base (Instruct))Model Scale=1B, Optimization Method=Base (Instruct)2026.06 | 8.4 | |
| InternVL3.5-1B (+GSPO)Model Scale=1B, Optimization Method=GSPO2026.06 | 8.4 |