Real-world understanding on RealWorldQA (Score)
70.07ScoreDefender Iter. 3
Evaluation Results
| Method | Links | |
|---|---|---|
| Defender Iter. 3Training Strategy=Iterative Co-evolution2026.01 | 70.07 | |
| Ablation (Clean Data)Training Strategy=Ablation2026.01 | 69.28 | |
| Defender Iter. 1Training Strategy=Iterative Co-evolution2026.01 | 69.28 | |
| Defender Iter. 2Training Strategy=Iterative Co-evolution2026.01 | 69.28 | |
| Liu et al. (All)Training Strategy=Finite Augmentation Datasets2026.01 | 68.76 | |
| Yang et al.Training Strategy=Finite Augmentation Datasets2026.01 | 68.24 | |
| Liu et al. (Insert)Training Strategy=Finite Augmentation Datasets2026.01 | 68.24 | |
| Liu et al. (Add)Training Strategy=Finite Augmentation Datasets2026.01 | 67.97 | |
| Base (M_def^(0))Training Strategy=Base2026.01 | 67.71 |