Multimodal Reasoning on MathVision, MathVista, MMMU, MMMU-Pro, and WeMath
54.7RC Accuracy AvgGroupwise Ranking Reward
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Groupwise Ranking RewardReward Method=Groupwise Ranking Reward, Training Budget=Matched, Training Epochs=22026.04 | 54.7 | 55.9 | |
| Pointwise GRReward Method=Pointwise GR, Training Budget=Matched, Training Epochs=22026.04 | 52.3 | 53.8 | |
| RLVRReference Checkpoint=RLVR2026.04 | 47.4 | 53.6 | |
| Qwen2.5-VL-7B-ITReference Checkpoint=Qwen2.5-VL-7B-IT2026.04 | 47.3 | 49 | |
| PRMReward Method=PRM, Training Budget=Matched, Training Epochs=22026.04 | 46.8 | 49.1 |