Reward Modeling (Usefulness) on ReasonEdit-Reward-113K 1.0 (test)
0.9428SRCCRE-Reward
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| RE-Reward2026.05 | 0.9428 | 0.8015 | 0.9439 | |
| Gemini-3-FlashCategory=closed-source MLLMs2026.05 | 0.8676 | 0.7364 | 0.8609 | |
| ChatGPT-5.4Category=closed-source MLLMs2026.05 | 0.866 | 0.7467 | 0.8657 | |
| Gemma-4-31BCategory=middle-scale MLLMs2026.05 | 0.8371 | 0.7078 | 0.8417 | |
| Qwen-3.5-35B-A3BCategory=middle-scale MLLMs2026.05 | 0.7884 | 0.6729 | 0.7916 | |
| FLEURCategory=Vision-language metrics2026.05 | 0.5393 | 0.4064 | 0.5422 | |
| Qwen-3.5-9BCategory=small-scale MLLMs2026.05 | 0.3906 | 0.3258 | 0.4037 | |
| Qwen3-VL-8BCategory=small-scale MLLMs2026.05 | 0.3487 | 0.2907 | 0.3617 | |
| InternVL3-8BCategory=small-scale MLLMs2026.05 | 0.2841 | 0.241 | 0.2346 | |
| DeepSeek-VL2-SmallCategory=middle-scale MLLMs2026.05 | 0.1848 | 0.1531 | 0.1881 | |
| Ovis2.5-9BCategory=small-scale MLLMs2026.05 | 0.1118 | 0.0747 | 0.1698 | |
| InternVL3.5-8BCategory=small-scale MLLMs2026.05 | 0.0736 | 0.0614 | 0.0291 | |
| CLIP-S-ViT-B/32Category=Vision-language metrics2026.05 | 0.0729 | 0.0501 | 0.0677 | |
| PAC-S++Category=Vision-language metrics2026.05 | 0.0169 | 0.0112 | 0.0181 | |
| mPLUG-Owl3-7BCategory=small-scale MLLMs2026.05 | 0.0023 | 0.0019 | 0.003 | |
| MiniCPM-V-2.6-8BCategory=small-scale MLLMs2026.05 | 0 | 0 | 0 |