Geometric Reasoning on Geo3K
55.2AccuracyNoisyRollout-7B
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| NoisyRollout-7BModel Scale=7B, Training Data overlap=trained on Geo3K2026.05 | 55.2 | — | — | |
| BPPOModel=Qwen3-VL-4B, Rollout Group Size (G)=32, Training Time (s)=18952 ± 376, Speedup (times)=2.08 ± 0.042026.05 | 49.85 | 1,438.18 | 52.1 | |
| PairModel=Qwen3-VL-4B, Rollout Group Size (G)=32, Training Time (s)=20403 ± 541, Speedup (times)=1.93 ± 0.052026.05 | 49.73 | 1,618.41 | 46.1 | |
| CPPOModel=Qwen3-VL-4B, Rollout Group Size (G)=32, Training Time (s)=22162 ± 253, Speedup (times)=1.78 ± 0.022026.05 | 49.56 | 4,005.67 | -33.5 | |
| GRPOModel=Qwen3-VL-4B, Rollout Group Size (G)=32, Training Time (s)=39447 ± 177, Speedup (times)=1.002026.05 | 49.34 | 3,001.24 | — | |
| EASE-7BModel Scale=7B2026.05 | 48.9 | — | — | |
| GRPO+FIRSTNModel=Qwen3-VL-4B, Rollout Group Size (G)=32, Training Time (s)=36419 ± 73, Speedup (times)=1.08 ± 0.012026.05 | 48.89 | 3,025.72 | -0.8 | |
| PAPO_D-7BModel Scale=7B2026.05 | 48.8 | — | — | |
| VPPO-RL-7BModel Scale=7B2026.05 | 46.9 | — | — | |
| VGPO-7BModel Scale=7B2026.05 | 45.8 | — | — | |
| ThinkLite-VL-7BModel Scale=7B2026.05 | 45.3 | — | — | |
| CFPOGBackbone=Qwen3-VL-2B-Thinking2026.06 | 44.8 | — | — | |
| VL-Rethinker-7BModel Scale=7B2026.05 | 44.4 | — | — | |
| MM-Eureka-7BModel Scale=7B2026.05 | 41.1 | — | — | |
| PAPOGBackbone=Qwen3-VL-2B-Thinking2026.06 | 41.08 | — | — | |
| GRPOBackbone=Qwen3-VL-2B-Thinking2026.06 | 39.29 | — | — |