Vision-Language Reward Model Evaluation on MMRewardBench
83.6AccuracySelective LWE
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Selective LWERelative Inference Cost (Input & Output Text)=3.9×2025.12 | 83.6 | 94.7 | 80.8 | |
| Majority VotingRelative Inference Cost (Input & Output Text)=5.0×2025.12 | 82.8 | 89.1 | 76.9 | |
| TextGrad*Relative Inference Cost (Input & Output Text)=4.4×2025.12 | 82.1 | 83.6 | 74.1 | |
| Sample-Specific PromptRelative Inference Cost (Input & Output Text)=2.5×2025.12 | 81.5 | 86.5 | 74.2 | |
| Dynamic CheatsheetRelative Inference Cost (Input & Output Text)=12.9×2025.12 | 81.1 | 90.1 | 76.4 | |
| VanillaRelative Inference Cost (Input & Output Text)=1.0×2025.12 | 80.8 | 86.3 | 74.7 | |
| CoTRelative Inference Cost (Input & Output Text)=1.2×2025.12 | 80.8 | 87.4 | 74.9 | |
| LWERelative Inference Cost (Input & Output Text)=10.9×2025.12 | 79.9 | 84.6 | 72.7 |