Geometry Problem Solving on OlympiadBench
75.5AccuracyGPT-5
Evaluation Results
| Method | Links | |
|---|---|---|
| GPT-5Decoding=Greedy (τ=0), Normalization=Exact-match, Resolution=224x224, Prompting strategy=Identical prompts2026.02 | 75.5 | |
| Gemini-2.5-ProDecoding=Greedy (τ=0), Normalization=Exact-match, Resolution=224x224, Prompting strategy=Identical prompts2026.02 | 75.22 | |
| Qwen3-VL-7B-Instruct (GeoCode trained)Backbone=Qwen3-VL-7B-Instruct, Training Setting=Ours (GeoCode trained)2026.02 | 57.34 | |
| Qwen3-VLModel detail=Qwen3-VL-32B-Thinking, Decoding=Greedy (τ=0), Normalization=Exact-match, Resolution=224x224, Prompting strategy=Identical prompts2026.02 | 54.46 | |
| Qwen3-VL-7B-Instruct (Baseline)Backbone=Qwen3-VL-7B-Instruct, Training Setting=Baseline2026.02 | 45.82 | |
| GeometryZero-VL-7BTraining=GCPO2025.06 | 37.19 | |
| Qwen2.5-VL-7B-Instruct + ToRLTraining=ToRL2025.06 | 34.59 | |
| Qwen2.5-VL-7B-Instruct + SFTTraining=SFT2025.06 | 32.36 | |
| Qwen2.5-VL-7B-Instruct + GRPOTraining=GRPO2025.06 | 31.82 | |
| Qwen2.5-VL-7B-InstructBase Model=Qwen2.5-VL-7B-Instruct2025.06 | 30.83 | |
| Qwen2.5-VL-7B-Instruct (GeoCode trained)Backbone=Qwen2.5-VL-7B-Instruct, Training Setting=Ours (GeoCode trained)2026.02 | 6.63 | |
| Qwen2.5-VL-7B-Instruct (Baseline)Backbone=Qwen2.5-VL-7B-Instruct, Training Setting=Baseline2026.02 | 5.76 |