Response Preference Evaluation on UltraFeedback (test)
83.42Win RateCLIPer
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| CLIPerComparison Baseline=Vanilla Baseline, Preference Dimension=1-dim2026.05 | 83.42 | — | — | 78 | 83.33 | — | |
| DPO-PoP-randomBackbone=Llama-3.1-8b, Data Type=synthetic, Judge=UltraRM2025.09 | 63 | — | — | — | — | 75 | |
| DPO-margin-gtBackbone=Llama-3.1-8b, Data Type=synthetic, Judge=UltraRM2025.09 | 59 | — | — | — | — | 50 | |
| CLIPerComparison Baseline=p-soup & Direct Fine-tuning, Preference Dimension=1-dim2026.05 | 58 | — | — | 71.33 | 42.33 | — | |
| MGDA-DECOUPLEDDecoding strategy=Nucleus sampling, Temperature=0.75, Top-p=0.9, Repetition penalty=1.05, Max new tokens=1024, Batch size=64, LLM-as-a-Judge model=GPT-4o, Max completion tokens=10, Concurrent requests=202026.04 | 56.3 | 3.4 | 40.3 | — | — | — | |
| DPO-PoP-iterBackbone=Llama-3.1-8b, Data Type=synthetic, Judge=UltraRM2025.09 | 56 | — | — | — | — | 34.96 | |
| GROUPDRODecoding strategy=Nucleus sampling, Temperature=0.75, Top-p=0.9, Repetition penalty=1.05, Max new tokens=1024, Batch size=64, LLM-as-a-Judge model=GPT-4o, Max completion tokens=10, Concurrent requests=202026.04 | 55.9 | 3.5 | 40.6 | — | — | — | |
| UNIFORMDecoding strategy=Nucleus sampling, Temperature=0.75, Top-p=0.9, Repetition penalty=1.05, Max new tokens=1024, Batch size=64, LLM-as-a-Judge model=GPT-4o, Max completion tokens=10, Concurrent requests=202026.04 | 55.7 | 3.4 | 40.9 | — | — | — | |
| MGDA-NORMALISEDDecoding strategy=Nucleus sampling, Temperature=0.75, Top-p=0.9, Repetition penalty=1.05, Max new tokens=1024, Batch size=64, LLM-as-a-Judge model=GPT-4o, Max completion tokens=10, Concurrent requests=202026.04 | 55.6 | 3.2 | 41.2 | — | — | — | |
| CDPODecoding strategy=Nucleus sampling, Temperature=0.75, Top-p=0.9, Repetition penalty=1.05, Max new tokens=1024, Batch size=64, LLM-as-a-Judge model=GPT-4o, Max completion tokens=10, Concurrent requests=202026.04 | 55.4 | 3.7 | 40.8 | — | — | — | |
| REFERENCEDecoding strategy=Nucleus sampling, Temperature=0.75, Top-p=0.9, Repetition penalty=1.05, Max new tokens=1024, Batch size=64, LLM-as-a-Judge model=GPT-4o, Max completion tokens=10, Concurrent requests=202026.04 | 55.3 | 3.2 | 41.5 | — | — | — | |
| DPO-margin-1Backbone=Llama-3.1-8b, Data Type=synthetic, Judge=UltraRM2025.09 | 55 | — | — | — | — | 28.13 | |
| CLIPerComparison Baseline=Direct Prompting, Preference Dimension=1-dim2026.05 | 53 | — | — | 53.67 | 49.67 | — | |
| DPO-margin-gt-scaledBackbone=Llama-3.1-8b, Data Type=synthetic, Judge=UltraRM2025.09 | 52 | — | — | — | — | 9.38 |