Reward Modeling on PsyCoPref v1.0 (test)
98.1AccuracyPsyCo-Llama3-3B-Reward
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| PsyCo-Llama3-3B-RewardModel Category=Our Reward Models2025.02 | 98.1 | 99.7 | 5 | 1.4 | |
| PsyCo-Llama3-8B-RewardModel Category=Our Reward Models2025.02 | 97.8 | 99.8 | 4.5 | 1.6 | |
| Llama-3.1-70B-InstructModel Category=Generative LLMs2025.02 | 88.2 | — | — | — | |
| Llama-3.1-Nemotron-70B-RewardModel Category=State-of-the-art Reward Models2025.02 | 87.3 | 93.8 | 4 | 10.2 | |
| gemma-2-9b-itModel Category=Generative LLMs2025.02 | 81.5 | — | — | — | |
| Llama-3.1-8B-InstructModel Category=Generative LLMs2025.02 | 80.1 | — | — | — | |
| Mistral-Nemo-Instruct-2407Model Category=Generative LLMs2025.02 | 78 | — | — | — | |
| Skywork-Reward-Gemma-2-27BModel Category=State-of-the-art Reward Models2025.02 | 69.2 | 74 | 12.3 | 22.9 | |
| Skywork-Reward-Llama-3.1-8B-v0.2Model Category=State-of-the-art Reward Models2025.02 | 57.9 | 62.3 | 33.1 | 37.9 |