General Instruction Following on WildBench
92.6ScoreGenRM-R-Align-14B
Evaluation Results
| Method | Links | |
|---|---|---|
| GenRM-R-Align-14BReward Model=GenRM-R-Align-14B2026.02 | 92.6 | |
| GenRM-R-Align-8BReward Model=GenRM-R-Align-8B2026.02 | 89.2 | |
| GenRM-RLVR-14BReward Model=GenRM-RLVR-14B2026.02 | 88.1 | |
| GenRM-RLVR-8BReward Model=GenRM-RLVR-8B2026.02 | 84.3 | |
| Qwen3-14B-as-GenRMReward Model=Qwen3-14B-as-GenRM2026.02 | 83.7 | |
| Qwen3-8B-as-GenRMReward Model=Qwen3-8B-as-GenRM2026.02 | 79.4 | |
| Qwen3-8BReward Model=Baseline2026.02 | 72.8 | |
| Ministral 3Model Size=14B2026.01 | 68.5 | |
| Ministral 3Model Size=8B2026.01 | 66.8 | |
| Qwen3-VLModel Size=8B, Variant=Instruct2026.01 | 66.3 | |
| Qwen3Model Size=14B, Variant=Non-Thinking2026.01 | 65.1 | |
| Gemma3Model Size=12B, Variant=Instruct2026.01 | 63.2 | |
| Qwen3-VLModel Size=4B, Variant=Instruct2026.01 | 56.8 | |
| Ministral 3Model Size=3B2026.01 | 56.8 | |
| Gemma3Model Size=4B, Variant=Instruct2026.01 | 49.1 | |
| Qwen3-VLModel Size=2B, Variant=Instruct2026.01 | 42.2 | |
| GeniusTraining Data Source=Magpie, Training Sample Size=25K2025.04 | 2.68 | |
| GeniusTraining Data Source=OpenHermes, Training Sample Size=32K2025.04 | 1.44 | |
| LLaMA3.1-8B-Instruct2025.04 | -1.11 |