Comparative Assessment on SummEval
68.9Coherence Accuracydebiased white-box
Evaluation Results
| Method | Links | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| debiased white-boxtype=teacher2024.03 | 68.9 | 0 | 79.7 | 0 | 61.4 | 0 | 67 | 0 | |
| RoBERTa-largeParameters=330M, type=error-correction2024.03 | 66.7 | 0.05 | 72 | 0.05 | 63.3 | 0.04 | 63.6 | 0.05 | |
| RoBERTa-baseParameters=110M, type=error-correction2024.03 | 66.4 | 0.04 | 71.7 | 0.06 | 61.6 | 0.03 | 61.9 | 0.05 | |
| DeBERTa-largeParameters=330M, type=error-correction2024.03 | 66.1 | 0.03 | 70.9 | 0.02 | 64.8 | 0.03 | 63.3 | 0.04 | |
| DeBERTa-baseParameters=110M, type=error-correction2024.03 | 66 | 0.03 | 71.1 | 0.04 | 64.1 | 0.03 | 62.1 | 0.06 | |
| DeBERTa-largeParameters=330M, type=distilled2024.03 | 65.1 | 0.04 | 71.5 | 0.04 | 64.9 | 0.03 | 63.2 | 0.03 | |
| RoBERTa-largeParameters=330M, type=distilled2024.03 | 64.8 | 0.05 | 67.6 | 0.06 | 62.7 | 0.05 | 62.1 | 0.05 | |
| DeBERTa-baseParameters=110M, type=distilled2024.03 | 62.6 | 0.03 | 67.1 | 0.04 | 62.1 | 0.03 | 63 | 0.05 | |
| biased white-boxtype=teacher2024.03 | 61.6 | 0.42 | 70.5 | 0.38 | 55.6 | 0.44 | 62.8 | 0.39 | |
| RoBERTa-baseParameters=110M, type=distilled2024.03 | 61.5 | 0.05 | 70.5 | 0.07 | 61.2 | 0.05 | 60.9 | 0.06 | |
| expected biased black-boxtype=teacher2024.03 | 58.5 | — | 65.5 | — | 54.3 | — | 58.8 | — | |
| BT-σaggregation=reliability-aware soft2026.02 | 57.38 | — | 47.47 | — | 42.99 | — | 54.15 | — | |
| BT-σ-asptype=aspect-dependent variant2026.02 | 57.36 | — | 47.56 | — | 43.08 | — | 54.56 | — | |
| Temp-BTtype=supervised baseline, calibration=temperature scaling2026.02 | 56.21 | — | 47.4 | — | 41.88 | — | 55.14 | — | |
| soft BT2026.02 | 53.94 | — | 47.86 | — | 42.69 | — | 53.11 | — | |
| hard BT-σaggregation=reliability-aware hard2026.02 | 53.02 | — | 47.08 | — | 40.44 | — | 52.69 | — | |
| Avg-Prob2026.02 | 52.55 | — | 41.75 | — | 36.21 | — | 50.09 | — | |
| hard BT2026.02 | 51.26 | — | 45.72 | — | 40.07 | — | 52.32 | — |