LLM Alignment on Reddit
4.24HVSURF
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| SURFBackbone=Qwen/Qwen2.5-0.5B-Instruct, Fine-tuning method=PPO, PEFT method=LoRA, Evaluation objective=reward-only, Batch size=64, Number of batches=102026.05 | 4.24 | 0.29 | 2.89 | |
| SURFBackbone=Qwen/Qwen2.5-0.5B-Instruct, Fine-tuning method=PPO, PEFT method=LoRA, Evaluation objective=KL-regularized, Batch size=64, Number of batches=102026.05 | 4.03 | 0.3 | 2.93 | |
| Uniform-wBackbone=Qwen/Qwen2.5-0.5B-Instruct, Fine-tuning method=PPO, PEFT method=LoRA, Evaluation objective=reward-only, Batch size=64, Number of batches=102026.05 | 4 | 0.65 | 6.01 | |
| Uniform-wBackbone=Qwen/Qwen2.5-0.5B-Instruct, Fine-tuning method=PPO, PEFT method=LoRA, Evaluation objective=KL-regularized, Batch size=64, Number of batches=102026.05 | 3.76 | 0.65 | 5.89 | |
| SoupBackbone=Qwen/Qwen2.5-0.5B-Instruct, Fine-tuning method=PPO, PEFT method=LoRA, Evaluation objective=reward-only, Batch size=64, Number of batches=102026.05 | 0.95 | 1.69 | 159.83 | |
| SoupBackbone=Qwen/Qwen2.5-0.5B-Instruct, Fine-tuning method=PPO, PEFT method=LoRA, Evaluation objective=KL-regularized, Batch size=64, Number of batches=102026.05 | 0.75 | 1.69 | 161.2 |