Reward Modeling on Anthropic HH (test)
68.49AccuracyDahoas/gptj-rm-static
Evaluation Results
| Method | Links | |
|---|---|---|
| Dahoas/gptj-rm-static2023.04 | 68.49 | |
| Alpaca-RRHF_DPTraining Algorithm=RRHF_DP, Proxy Reward Model=Dahoas/gptj-rm-static2023.04 | 61.75 | |
| Alpaca-PPOTraining Algorithm=PPO2023.04 | 46.03 | |
| Alpaca2023.04 | 45.13 | |
| LLaMA2023.04 | 45.09 |