Response Quality Evaluation on IFLLM 1.0 (test)
53.54DPO Win RateRF + (IF - Gaze)
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| RF + (IF - Gaze)Alignment Method=rDPO + NLL, Reward Source=RF + (IF - Gaze), Judge=GPT4.1-mini, Base Model=Random Forest2026.06 | 53.54 | 41.96 | 0.345 | 202.5 | |
| mBERT base + Text + IFAlignment Method=rDPO + NLL, Reward Source=mBERT base + Text + IF, Judge=GPT4.1-mini, Base Model=ModernBERT2026.06 | 50.71 | 45 | 0.1958 | 221.1 | |
| RF + IFAlignment Method=rDPO + NLL, Reward Source=RF + IF, Judge=GPT4.1-mini, Base Model=Random Forest2026.06 | 50.08 | 45.92 | 0.1892 | 206.6 | |
| Explicit FeedbackAlignment Method=DPO + NLL, Reward Source=Explicit Feedback, Judge=GPT4.1-mini2026.06 | 49.5 | 46.25 | 0.1079 | 208.8 | |
| mBERT base + TextAlignment Method=rDPO + NLL, Reward Source=mBERT base + Text, Judge=GPT4.1-mini, Base Model=ModernBERT2026.06 | 49.42 | 45.54 | 0.1221 | 227.3 | |
| RF + (IF - Mouse)Alignment Method=rDPO + NLL, Reward Source=RF + (IF - Mouse), Judge=GPT4.1-mini, Base Model=Random Forest2026.06 | 48.71 | 46.67 | 0.1108 | 203.5 |