Reward Modeling (Accuracy) on ReasonEdit-Reward-113K 1.0 (test)
0.9323SRCCRE-Reward
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| RE-Reward2026.05 | 0.9323 | 0.7781 | 0.9316 | |
| Gemini-3-FlashCategory=closed-source MLLMs2026.05 | 0.8587 | 0.7209 | 0.8411 | |
| ChatGPT-5.4Category=closed-source MLLMs2026.05 | 0.8532 | 0.7284 | 0.8535 | |
| Gemma-4-31BCategory=middle-scale MLLMs2026.05 | 0.8264 | 0.6943 | 0.8251 | |
| Qwen-3.5-35B-A3BCategory=middle-scale MLLMs2026.05 | 0.7761 | 0.6598 | 0.7797 | |
| FLEURCategory=Vision-language metrics2026.05 | 0.4853 | 0.3612 | 0.5055 | |
| DeepSeek-VL2-SmallCategory=middle-scale MLLMs2026.05 | 0.4462 | 0.3525 | 0.3839 | |
| Qwen-3.5-9BCategory=small-scale MLLMs2026.05 | 0.4279 | 0.3544 | 0.432 | |
| Qwen3-VL-8BCategory=small-scale MLLMs2026.05 | 0.3947 | 0.3266 | 0.4175 | |
| mPLUG-Owl3-7BCategory=small-scale MLLMs2026.05 | 0.3835 | 0.3163 | 0.3894 | |
| MiniCPM-V-2.6-8BCategory=small-scale MLLMs2026.05 | 0.1548 | 0.1287 | 0.1538 | |
| Ovis2.5-9BCategory=small-scale MLLMs2026.05 | 0.0752 | 0.0725 | 0.0453 | |
| InternVL3-8BCategory=small-scale MLLMs2026.05 | 0.0687 | 0.0553 | 0.0929 | |
| CLIP-S-ViT-B/32Category=Vision-language metrics2026.05 | 0.0655 | 0.0452 | 0.0624 | |
| InternVL3.5-8BCategory=small-scale MLLMs2026.05 | 0.0497 | 0.041 | 0.043 | |
| PAC-S++Category=Vision-language metrics2026.05 | 0.005 | 0.0029 | 0.0062 |