Reward Modeling on Reward Bench Prior Sets
78.2Prior Sets ScoreCohere May 2024
Evaluation Results
| Method | Links | |
|---|---|---|
| Cohere May 2024Source of Model/Training Data=Proprietary Models2024.06 | 78.2 | |
| Cohere March 2024Source of Model/Training Data=Proprietary Models2024.06 | 74.6 | |
| RLHFlow-Llama 3 8BSource of Model/Training Data=Trained with GPT-4 Generated Data2024.06 | 74.6 | |
| ArmoRM-Llama 3 8BSource of Model/Training Data=Trained with GPT-4 Generated Data2024.06 | 74.3 | |
| GPT-4-0613Source of Model/Training Data=Proprietary Models2024.06 | 73.6 | |
| GPT-4o-0513Source of Model/Training Data=Proprietary Models2024.06 | 72.6 | |
| Eurus RM Mistral 7BSource of Model/Training Data=Trained with GPT-4 Generated Data2024.06 | 71.7 | |
| Starling RM Yi 34BSource of Model/Training Data=Trained with GPT-4 Generated Data2024.06 | 71.4 | |
| GPT-4-0125-previewSource of Model/Training Data=Proprietary Models2024.06 | 70.9 | |
| Llama 3 70B InstructSource of Model/Training Data=Trained with Data allowing Permissive Use2024.06 | 70.4 | |
| Claude-3-Opus-02292024Source of Model/Training Data=Proprietary Models2024.06 | 70.3 | |
| Llama 3 70B (w. HH-RLHF)Source of Model/Training Data=Trained with Data allowing Permissive Use2024.06 | 68.8 | |
| Llama 3 70B (w. HelpSteer)Source of Model/Training Data=Trained with Data allowing Permissive Use2024.06 | 67.7 | |
| Nemotron-4 340B RM (w. HelpSteer2)Source of Model/Training Data=Proprietary Models2024.06 | 67.4 | |
| Llama 3 70B (w. Open Assistant)Source of Model/Training Data=Trained with Data allowing Permissive Use2024.06 | 66.7 | |
| Llama 3 70B RM (w. HelpSteer2)Source of Model/Training Data=Trained with Data allowing Permissive Use2024.06 | 66.5 | |
| Pythia 1.4B (w. Open Assistant)Source of Model/Training Data=Trained with Data allowing Permissive Use2024.06 | 65.3 |