Reward Modeling on RewardBench v1.0 (test)
0.978Average ScoreSkywork-Reward-V2-Llama-3.1-8B-40M
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| Skywork-Reward-V2-Llama-3.1-8B-40MModel Type=Reward Model, Backbone=Llama-3.1, Parameter Count=8B, Training Data=40M2025.07 | 0.978 | — | — | — | — | |
| Skywork-Reward-V2-Llama-3.1-8BModel Type=Reward Model, Backbone=Llama-3.1, Parameter Count=8B2025.07 | 0.964 | — | — | — | — | |
| INF-ORM-Llama3.1-70BModel Type=Open Reward Model, Backbone=Llama-3.1, Parameter Count=70B2025.07 | 0.951 | — | — | — | — | |
| LDL-Reward-Gemma-2-27B-v0.1Model Type=Open Reward Model, Backbone=Gemma-2, Parameter Count=27B2025.07 | 0.95 | — | — | — | — | |
| Skywork-Reward-Gemma-2-27B-v0.2Model Type=Open Reward Model, Backbone=Gemma-2, Parameter Count=27B2025.07 | 0.943 | — | — | — | — | |
| Llama-3.1-Nemotron-70BModel Type=Open Reward Model, Backbone=Llama-3.1, Parameter Count=70B2025.07 | 0.939 | — | — | — | — | |
| EvalPlanner (Llama-3.1-70B)Model Type=LLM-as-a-Judge, Backbone=Llama-3.1, Parameter Count=70B2025.07 | 0.939 | — | — | — | — | |
| EvalPlanner (Llama-3.3-70B)Model Type=LLM-as-a-Judge, Backbone=Llama-3.3, Parameter Count=70B2025.07 | 0.938 | — | — | — | — | |
| Skywork-Reward-V2-Qwen3-8BModel Type=Reward Model, Backbone=Qwen3, Parameter Count=8B2025.07 | 0.937 | — | — | — | — | |
| Skywork-Reward-V2-Qwen3-4BModel Type=Reward Model, Backbone=Qwen3, Parameter Count=4B2025.07 | 0.934 | — | — | — | — | |
| J1-Llama-70BModel Type=LLM-as-a-Judge, Backbone=Llama, Parameter Count=70B2025.07 | 0.933 | — | — | — | — | |
| Skywork-Reward-Llama-3.1-8B-v0.2Model Type=Open Reward Model, Backbone=Llama-3.1, Parameter Count=8B2025.07 | 0.931 | — | — | — | — | |
| Skywork-Reward-V2-Llama-3.2-3BModel Type=Reward Model, Backbone=Llama-3.2, Parameter Count=3B2025.07 | 0.93 | — | — | — | — | |
| RM-R1-Qwen-Instruct-32BModel Type=Generative Reward Model, Backbone=Qwen, Parameter Count=32B2025.07 | 0.929 | — | — | — | — | |
| RM-R1-DeepSeek-Distill-Qwen-32BModel Type=Generative Reward Model, Backbone=Qwen, Parameter Count=32B2025.07 | 0.909 | — | — | — | — | |
| ArmoRM-Llama3-8B-v0.1Model Type=Open Reward Model, Backbone=Llama-3, Parameter Count=8B2025.07 | 0.904 | — | — | — | — | |
| DeepSeek-GRM-27B (w/ MetaRM)Model Type=Generative Reward Model, Parameter Count=27B, MetaRM=true2025.07 | 0.904 | — | — | — | — | |
| Skywork-Reward-V2-Qwen3-1.7BModel Type=Reward Model, Backbone=Qwen3, Parameter Count=1.7B2025.07 | 0.903 | — | — | — | — | |
| Internlm2-20b-rewardModel Type=Open Reward Model, Backbone=Internlm2, Parameter Count=20B2025.07 | 0.902 | — | — | — | — | |
| Skywork-Reward-V2-Llama-3.2-1BModel Type=Reward Model, Backbone=Llama-3.2, Parameter Count=1B2025.07 | 0.899 | — | — | — | — | |
| Llama-3-OffsetBias-RM-8BModel Type=Open Reward Model, Backbone=Llama-3, Parameter Count=8B2025.07 | 0.89 | — | — | — | — | |
| DeepSeek-GRM-27BModel Type=Generative Reward Model, Parameter Count=27B2025.07 | 0.885 | — | — | — | — | |
| GPT-4oModel Type=LLM-as-a-Judge2025.07 | 0.867 | — | — | — | — | |
| J1-Llama-8BModel Type=LLM-as-a-Judge, Backbone=Llama, Parameter Count=8B2025.07 | 0.857 | — | — | — | — | |
| Skywork-Reward-V2-Qwen3-0.6BModel Type=Reward Model, Backbone=Qwen3, Parameter Count=0.6B2025.07 | 0.852 | — | — | — | — | |
| NLL-SymModel=Llama2026.02 | 0.843 | 0.941 | 0.728 | 0.897 | 0.804 | |
| Claude-3.5-SonnetModel Type=LLM-as-a-Judge2025.07 | 0.842 | — | — | — | — | |
| All-ThreshModel=Llama2026.02 | 0.82 | 0.922 | 0.689 | 0.872 | 0.798 | |
| NORMBTBase Model=Llama-3.2-3B-Instruct, Training Dataset=Skywork-Reward-Preference-80K-v0.22025.12 | 0.8148 | 0.838 | 0.7873 | 0.8878 | 0.746 | |
| NLL-AsymModel=Llama2026.02 | 0.809 | 0.911 | 0.695 | 0.837 | 0.794 | |
| Soft LabelModel=Zephyr2026.02 | 0.807 | 0.933 | 0.66 | 0.819 | 0.816 | |
| NLL-SymModel=Llama2026.02 | 0.807 | 0.947 | 0.765 | 0.808 | 0.707 | |
| Soft LabelModel=Llama2026.02 | 0.805 | 0.922 | 0.728 | 0.823 | 0.747 | |
| BT (baseline)Base Model=Llama-3.2-3B-Instruct, Training Dataset=Skywork-Reward-Preference-80K-v0.22025.12 | 0.8031 | 0.8603 | 0.7829 | 0.8986 | 0.6705 | |
| Margin BTModel=Llama2026.02 | 0.802 | 0.961 | 0.66 | 0.885 | 0.703 | |
| NORMBTBase Model=gemma-2b-it, Training Dataset=Skywork-Reward-Preference-80K-v0.22025.12 | 0.8012 | 0.838 | 0.7346 | 0.825 | 0.8071 | |
| Scaled BTModel=Mistral2026.02 | 0.8 | 0.93 | 0.66 | 0.824 | 0.787 | |
| All-ThreshModel=Zephyr2026.02 | 0.792 | 0.933 | 0.662 | 0.751 | 0.821 | |
| BT (baseline)Base Model=gemma-2b-it, Training Dataset=Skywork-Reward-Preference-80K-v0.22025.12 | 0.7863 | 0.8128 | 0.7336 | 0.8243 | 0.7746 | |
| All-ThreshModel=Mistral2026.02 | 0.783 | 0.897 | 0.68 | 0.73 | 0.827 | |
| Scaled BTModel=Zephyr2026.02 | 0.781 | 0.944 | 0.665 | 0.727 | 0.786 | |
| Scaled BTModel=Llama2026.02 | 0.781 | 0.944 | 0.748 | 0.793 | 0.638 | |
| NLL-SymModel=Mistral2026.02 | 0.779 | 0.939 | 0.684 | 0.708 | 0.783 | |
| NLL-AsymModel=Zephyr2026.02 | 0.777 | 0.911 | 0.662 | 0.68 | 0.856 | |
| B.4Backbone=Mistral2025.03 | 0.775 | 0.883 | 0.581 | 0.807 | 0.828 | |
| Margin BTModel=Llama2026.02 | 0.771 | 0.947 | 0.7 | 0.78 | 0.656 | |
| NORMBTBase Model=Llama-3.2-3B-Instruct, Training Dataset=Unified-Feedback (80K)2025.12 | 0.7696 | 0.9693 | 0.4978 | 0.8419 | 0.7693 | |
| Margin BTModel=Zephyr2026.02 | 0.768 | 0.922 | 0.7 | 0.857 | 0.593 | |
| NLL-SymModel=Zephyr2026.02 | 0.768 | 0.953 | 0.603 | 0.761 | 0.756 | |
| NLL-AsymModel=Mistral2026.02 | 0.767 | 0.919 | 0.629 | 0.819 | 0.699 | |
| NLL-AsymModel=Llama2026.02 | 0.765 | 0.936 | 0.697 | 0.785 | 0.641 | |
| All-ThreshModel=Llama2026.02 | 0.764 | 0.872 | 0.728 | 0.78 | 0.676 | |
| NLL-SymModel=Zephyr2026.02 | 0.764 | 0.913 | 0.722 | 0.753 | 0.668 | |
| CLoudBackbone=Llama-3-8B2025.03 | 0.759 | 0.965 | 0.455 | 0.754 | 0.862 | |
| Scaled BTModel=Llama2026.02 | 0.757 | 0.927 | 0.638 | 0.853 | 0.61 | |
| Scaled BTModel=Mistral2026.02 | 0.756 | 0.908 | 0.724 | 0.687 | 0.704 | |
| BT (baseline)Base Model=Llama-3.2-3B-Instruct, Training Dataset=Unified-Feedback (80K)2025.12 | 0.7524 | 0.9553 | 0.4989 | 0.8169 | 0.717 | |
| Margin BTModel=Mistral2026.02 | 0.749 | 0.933 | 0.601 | 0.738 | 0.725 | |
| NLL-SymModel=Mistral2026.02 | 0.749 | 0.927 | 0.706 | 0.695 | 0.668 | |
| BT + margin outBase Model=Llama-3.2-3B-Instruct, Training Dataset=Unified-Feedback (80K)2025.12 | 0.7484 | 0.9665 | 0.4583 | 0.8405 | 0.7281 | |
| TRACTBackbone=Llama-3-8B, Reference ID=7.L2025.03 | 0.748 | 0.922 | 0.434 | 0.799 | 0.837 | |
| BT + label smoothBase Model=Llama-3.2-3B-Instruct, Training Dataset=Unified-Feedback (80K)2025.12 | 0.7439 | 0.9553 | 0.4978 | 0.7973 | 0.7252 | |
| BT + marginBase Model=Llama-3.2-3B-Instruct, Training Dataset=Unified-Feedback (80K)2025.12 | 0.7415 | 0.9777 | 0.4759 | 0.8128 | 0.6997 | |
| TRACTBackbone=Mistral, Reference ID=7.M2025.03 | 0.736 | 0.927 | 0.542 | 0.759 | 0.716 | |
| NORMBTBase Model=gemma-2b-it, Training Dataset=Unified-Feedback (80K)2025.12 | 0.7357 | 0.9581 | 0.398 | 0.7797 | 0.8071 | |
| Scaled BTModel=Zephyr2026.02 | 0.73 | 0.888 | 0.634 | 0.723 | 0.673 | |
| Margin BTModel=Mistral2026.02 | 0.727 | 0.916 | 0.702 | 0.7 | 0.588 | |
| BT + margin outBase Model=gemma-2b-it, Training Dataset=Unified-Feedback (80K)2025.12 | 0.7253 | 0.9609 | 0.3838 | 0.7757 | 0.7809 | |
| BT (baseline)Base Model=gemma-2b-it, Training Dataset=Unified-Feedback (80K)2025.12 | 0.7225 | 0.9525 | 0.4035 | 0.7797 | 0.7541 | |
| Soft LabelModel=Llama2026.02 | 0.722 | 0.939 | 0.581 | 0.689 | 0.68 | |
| Prometheus-2-7BEvaluation Mode=pairwise, Backbone=Mistral2025.03 | 0.72 | 0.855 | 0.491 | 0.771 | 0.765 | |
| A.2Backbone=Mistral2025.03 | 0.717 | 0.886 | 0.564 | 0.773 | 0.645 | |
| BT + marginBase Model=gemma-2b-it, Training Dataset=Unified-Feedback (80K)2025.12 | 0.7123 | 0.9581 | 0.375 | 0.7865 | 0.7298 | |
| Margin BTModel=Zephyr2026.02 | 0.707 | 0.911 | 0.66 | 0.643 | 0.612 | |
| NLL-AsymModel=Mistral2026.02 | 0.704 | 0.888 | 0.627 | 0.685 | 0.616 | |
| Soft LabelModel=Zephyr2026.02 | 0.7 | 0.908 | 0.686 | 0.642 | 0.566 | |
| BT + label smoothBase Model=gemma-2b-it, Training Dataset=Unified-Feedback (80K)2025.12 | 0.6995 | 0.9385 | 0.3772 | 0.7595 | 0.7228 | |
| Soft LabelModel=Mistral2026.02 | 0.699 | 0.791 | 0.7 | 0.653 | 0.654 | |
| Soft LabelModel=Mistral2026.02 | 0.694 | 0.905 | 0.626 | 0.569 | 0.776 | |
| All-ThreshModel=Mistral2026.02 | 0.687 | 0.925 | 0.568 | 0.603 | 0.651 | |
| NLL-AsymModel=Zephyr2026.02 | 0.685 | 0.916 | 0.64 | 0.574 | 0.611 | |
| All-ThreshModel=Zephyr2026.02 | 0.674 | 0.925 | 0.6 | 0.557 | 0.615 | |
| A.4Backbone=Mistral2025.03 | 0.64 | 0.824 | 0.469 | 0.7 | 0.568 | |
| A.1Backbone=Mistral2025.03 | 0.585 | 0.777 | 0.513 | 0.569 | 0.482 | |
| Prometheus-2-7BEvaluation Mode=pointwise, Backbone=Mistral2025.03 | 0.538 | 0.679 | 0.423 | 0.673 | 0.378 | |
| B.1Backbone=Mistral2025.03 | 0.484 | 0.592 | 0.29 | 0.7 | 0.355 | |
| B.3Backbone=Mistral2025.03 | 0.479 | 0.595 | 0.358 | 0.651 | 0.312 | |
| A.3Backbone=Mistral2025.03 | 0.359 | 0.564 | 0.292 | 0.385 | 0.194 | |
| B.2Backbone=Mistral2025.03 | 0.355 | 0.629 | 0.305 | 0.12 | 0.368 |