Reward Modeling on UltraFeedback core250 (test)
1.103Reward Score Difference (TEA vs GRPO)TEA
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| TEAEndpoint (Best-of-N)=256, Backbone=Meta-Llama-3.1-8B-Instruct, Training Method=LoRA2026.05 | 1.103 | 0.821 | 1.384 | |
| TEAEndpoint (Best-of-N)=128, Backbone=Meta-Llama-3.1-8B-Instruct, Training Method=LoRA2026.05 | 1.067 | 0.839 | 1.3 | |
| TEAEndpoint (Best-of-N)=64, Backbone=Meta-Llama-3.1-8B-Instruct, Training Method=LoRA2026.05 | 0.835 | 0.659 | 1.001 | |
| TEAEndpoint (Best-of-N)=32, Backbone=Meta-Llama-3.1-8B-Instruct, Training Method=LoRA2026.05 | 0.701 | 0.565 | 0.832 |