Personalized LLM Alignment Evaluation on PersonalRewardBench (test)
3.354Mean ScoreLlama3.1-8B-Instruct-GRPO
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Llama3.1-8B-Instruct-GRPOModel Series=Llama 3.1, Parameters=8B, Training Type=Instruct-GRPO2026.02 | 3.354 | 0.0102 | — | |
| Llama3.1-8B-Instruct-DPOModel Series=Llama 3.1, Parameters=8B, Training Type=Instruct-DPO2026.02 | 3.316 | 0.0068 | — | |
| Qwen2.5-72B-InstructModel Series=Qwen 2.5, Parameters=72B, Training Type=Instruct2026.02 | 3.214 | 0.0089 | — | |
| Llama3.1-70B-InstructModel Series=Llama 3.1, Parameters=70B, Training Type=Instruct2026.02 | 3.156 | 0.0093 | — | |
| Qwen2.5-7B-InstructModel Series=Qwen 2.5, Parameters=7B, Training Type=Instruct2026.02 | 2.97 | 0.0089 | — | |
| Llama3.1-8B-InstructModel Series=Llama 3.1, Parameters=8B, Training Type=Instruct2026.02 | 2.954 | 0.0074 | — |