Multimodal Reasoning on MMMU (Acc, AUROC, ECE)
53AccuracyC3RL (Ours)
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| C3RL (Ours)Backbone=Qwen2.5VL-7B-Instruct, Averaged Over Seeds=true2026.07 | 53 | 65.5 | 9.6 | |
| C3RLBackbone=Qwen2.5VL-7B-Instruct, Evaluation Protocol=Averaged over 3 seeds (0, 42, 2025)2026.07 | 53 | 65.5 | 9.6 | |
| C3RL w/o RefBackbone=Qwen2.5VL-7B-Instruct, Averaged Over Seeds=true2026.07 | 52.4 | 61.3 | 17.3 | |
| C3RL w/o RefBackbone=Qwen2.5VL-7B-Instruct, Evaluation Protocol=Averaged over 3 seeds (0, 42, 2025)2026.07 | 52.4 | 61.3 | 17.3 | |
| RLVRBackbone=Qwen2.5VL-7B-Instruct, Averaged Over Seeds=true2026.07 | 52 | 52.5 | 41.7 | |
| RLVRBackbone=Qwen2.5VL-7B-Instruct, Evaluation Protocol=Averaged over 3 seeds (0, 42, 2025)2026.07 | 52 | 52.5 | 41.7 | |
| SCBackbone=Qwen2.5VL-7B-Instruct, Averaged Over Seeds=false2026.07 | 50.9 | 62.1 | 21.1 | |
| RLCRBackbone=Qwen2.5VL-7B-Instruct, Averaged Over Seeds=true2026.07 | 50.9 | 65.1 | 14.6 | |
| SCBackbone=Qwen2.5VL-7B-Instruct, Evaluation Protocol=Single test2026.07 | 50.9 | 62.1 | 21.1 | |
| RLCRBackbone=Qwen2.5VL-7B-Instruct, Evaluation Protocol=Averaged over 3 seeds (0, 42, 2025)2026.07 | 50.9 | 65.1 | 14.6 | |
| BaseBackbone=Qwen2.5VL-7B-Instruct, Averaged Over Seeds=true2026.07 | 47.3 | 57.1 | 38.1 | |
| BaseBackbone=Qwen2.5VL-7B-Instruct, Evaluation Protocol=Averaged over 3 seeds (0, 42, 2025)2026.07 | 47.3 | 57.1 | 38.1 | |
| SFT+RefBackbone=Qwen2.5VL-7B-Instruct, Averaged Over Seeds=true2026.07 | 36.8 | 65.1 | 9.5 | |
| SFT+RefBackbone=Qwen2.5VL-7B-Instruct, Evaluation Protocol=Averaged over 3 seeds (0, 42, 2025)2026.07 | 36.8 | 65.1 | 9.5 | |
| SaySelfBackbone=Qwen2.5VL-7B-Instruct, Averaged Over Seeds=true2026.07 | 33.1 | 78.4 | 16 | |
| SaySelfBackbone=Qwen2.5VL-7B-Instruct, Evaluation Protocol=Averaged over 3 seeds (0, 42, 2025)2026.07 | 33.1 | 78.4 | 16 |