Vision-Language Reward Model Evaluation on VLRewardBench
74.5AccuracyLWE
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| LWERelative Inference Cost (Input & Output Text)=10.9×2025.12 | 74.5 | 80.5 | 64.6 | |
| TextGrad*Relative Inference Cost (Input & Output Text)=4.4×2025.12 | 73 | 74.9 | 61.5 | |
| Dynamic CheatsheetRelative Inference Cost (Input & Output Text)=12.9×2025.12 | 69.8 | 86.8 | 62.9 | |
| Selective LWERelative Inference Cost (Input & Output Text)=3.9×2025.12 | 67.6 | 94 | 64.8 | |
| Sample-Specific PromptRelative Inference Cost (Input & Output Text)=2.5×2025.12 | 66.1 | 72.7 | 52.9 | |
| CoTRelative Inference Cost (Input & Output Text)=1.2×2025.12 | 65.1 | 80.8 | 55.3 | |
| VanillaRelative Inference Cost (Input & Output Text)=1.0×2025.12 | 62.9 | 80.1 | 52.9 | |
| Majority VotingRelative Inference Cost (Input & Output Text)=5.0×2025.12 | 62.7 | 81 | 53.7 |