Reward Modeling on Arabic preference (test)
85.4AccuracyRM-Distiller-Qwen2.5-3B-Instruct
Evaluation Results
| Method | Links | |
|---|---|---|
| RM-Distiller-Qwen2.5-3B-InstructSample Num=10K2026.01 | 85.4 | |
| Llama-3-OffsetBias-8BSample Num=70K2026.01 | 83.2 | |
| Tulu-3-8B-RM-RB2Sample Num=350K2026.01 | 83.2 | |
| Skywork-Reward-V2-8BSample Num=40,000K2026.01 | 81.3 | |
| URM-LLaMA-3.1-8BSample Num=100K2026.01 | 76.7 | |
| Skywork-Reward-8B-v0.2Sample Num=80K2026.01 | 75.9 | |
| BT-Qwen2.5-3B-InstructSample Num=10K2026.01 | 72.2 |