Dialogue Summarization on SAMSum Multiple Client (test)
49.99ROUGE-1 (Client 1)Conf
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| ConfDescription=Confidence-based weak-to-strong distillation2026.05 | 49.99 | 50.09 | 51.98 | 48.19 | 48.17 | 49.68 | 98.24 | |
| PTDescription=Ceiling performance, LLM fine-tuned directly on client data2026.05 | 49.51 | 49.98 | 48.91 | 50.21 | 50.11 | 49.74 | — | |
| VisSupDescription=Weak-to-strong distillation baseline2026.05 | 48.51 | 50.6 | 50.35 | 49.14 | 47.47 | 49.21 | 84.41 | |
| W2SDescription=Weak-to-strong distillation2026.05 | 48.02 | 49.26 | 50.11 | 48.06 | 46.39 | 48.37 | 59.71 | |
| GRAD-TRANSFORMERDescription=Learning to generate updates for LLMs2026.05 | 47.92 | 50.5 | 48.83 | 49.37 | 48.45 | 49.01 | 78.53 | |
| PSDescription=TinyLM fine-tuned from the client side2026.05 | 45.32 | 47.7 | 45.31 | 46.12 | 47.25 | 46.34 | — |