Reward Modeling (Logicality) on ReasonEdit-Reward 113K 1.0 (test)
0.8958SRCCRE-Reward
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| RE-Reward2026.05 | 0.8958 | 0.7315 | 0.9027 | |
| Gemini-3-FlashCategory=closed-source MLLMs2026.05 | 0.8438 | 0.7136 | 0.8173 | |
| ChatGPT-5.4Category=closed-source MLLMs2026.05 | 0.7539 | 0.6528 | 0.7572 | |
| Gemma-4-31BCategory=middle-scale MLLMs2026.05 | 0.7035 | 0.5902 | 0.7177 | |
| Qwen-3.5-35B-A3BCategory=middle-scale MLLMs2026.05 | 0.5728 | 0.486 | 0.5956 | |
| Qwen3-VL-8BCategory=small-scale MLLMs2026.05 | 0.5386 | 0.4545 | 0.5913 | |
| Qwen-3.5-9BCategory=small-scale MLLMs2026.05 | 0.4097 | 0.3463 | 0.4787 | |
| mPLUG-Owl3-7BCategory=small-scale MLLMs2026.05 | 0.3473 | 0.2819 | 0.3197 | |
| FLEURCategory=Vision-language metrics2026.05 | 0.3384 | 0.2496 | 0.3258 | |
| DeepSeek-VL2-SmallCategory=middle-scale MLLMs2026.05 | 0.3318 | 0.2593 | 0.2843 | |
| MiniCPM-V-2.6-8BCategory=small-scale MLLMs2026.05 | 0.2749 | 0.2285 | 0.2269 | |
| InternVL3-8BCategory=small-scale MLLMs2026.05 | 0.1697 | 0.144 | 0.148 | |
| InternVL3.5-8BCategory=small-scale MLLMs2026.05 | 0.0784 | 0.0612 | 0.1798 | |
| CLIP-S-ViT-B/32Category=Vision-language metrics2026.05 | 0.034 | 0.0238 | 0.0417 | |
| Ovis2.5-9BCategory=small-scale MLLMs2026.05 | 0.0314 | 0.0268 | 0.0419 | |
| PAC-S++Category=Vision-language metrics2026.05 | 0.0298 | 0.0209 | 0.0148 |