Transfer Learning Robustness on Combined Summarization & Email Writing
0.7805Max SubOpt ScoreDPO
Evaluation Results
| Method | Links | |
|---|---|---|
| DPOTest-time User Model=Llama-3.3-70B-Instruct, Online learning phase=true, Combined evaluation across tasks=true2026.01 | 0.7805 | |
| BaseTest-time User Model=Llama-3.3-70B-Instruct, Online learning phase=true, Combined evaluation across tasks=true2026.01 | 0.6654 | |
| SFTTest-time User Model=Llama-3.3-70B-Instruct, Online learning phase=true, Combined evaluation across tasks=true2026.01 | 0.5121 | |
| EarlyEnsembleTest-time User Model=Llama-3.3-70B-Instruct, Online learning phase=true, Combined evaluation across tasks=true2026.01 | 0.0955 | |
| LateEnsembleTest-time User Model=Llama-3.3-70B-Instruct, Online learning phase=true, Combined evaluation across tasks=true2026.01 | 0.0862 |