Vision-Language Reasoning on MMStar
74.7AccuracyDAPO + ReMind
Evaluation Results
| Method | Links | |
|---|---|---|
| DAPO + ReMindBase Model=Qwen3-VL-8B-Instruct2026.06 | 74.7 | |
| GRPO + ReMindBase Model=Qwen3-VL-8B-Instruct2026.06 | 73.4 | |
| DAPOBase Model=Qwen3-VL-8B-Instruct2026.06 | 72.77 | |
| RePOBase Model=Qwen3-VL-8B-Instruct2026.06 | 72.6 | |
| RLEPBase Model=Qwen3-VL-8B-Instruct2026.06 | 72.27 | |
| GRPOBase Model=Qwen3-VL-8B-Instruct2026.06 | 72.2 | |
| ExGRPOBase Model=Qwen3-VL-8B-Instruct2026.06 | 72 | |
| Base ModelBase Model=Qwen3-VL-8B-Instruct2026.06 | 71.83 | |
| ETTCAggregation Strategy=ETTC, Model Scale=7B-12B Ensemble2026.05 | 60.07 | |
| VotingAggregation Strategy=Majority Voting, Model Scale=7B-12B Ensemble2026.05 | 59.27 | |
| Qwen-7BModel Backbone=Qwen, Model Scale=7B, Aggregation Strategy=Single Model2026.05 | 56.77 | |
| Gemma-12BModel Backbone=Gemma, Model Scale=12B, Aggregation Strategy=Single Model2026.05 | 53.4 | |
| Pixtral-12BModel Backbone=Pixtral, Model Scale=12B, Aggregation Strategy=Single Model2026.05 | 50.35 | |
| LLaMA-11BModel Backbone=LLaMA, Model Scale=11B, Aggregation Strategy=Single Model2026.05 | 46.09 |