Mathematical Reasoning on AQuA-RAT Multiple Client (test)
64.92Client 1 AccuracyPT
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| PTDescription=Ceiling performance, LLM fine-tuned directly on client data2026.05 | 64.92 | 66.05 | 64.51 | 64.21 | 66.46 | 65.23 | — | |
| GRAD-TRANSFORMERDescription=Learning to generate updates for LLMs2026.05 | 61.74 | 64.31 | 63.9 | 62.87 | 65.03 | 63.57 | 85.15 | |
| ConfDescription=Confidence-based weak-to-strong distillation2026.05 | 55.38 | 58.15 | 56.82 | 59.79 | 59.49 | 57.93 | 34.7 | |
| VisSupDescription=Weak-to-strong distillation baseline2026.05 | 54.56 | 57.85 | 57.33 | 59.18 | 59.08 | 57.6 | 31.75 | |
| PSDescription=TinyLM fine-tuned from the client side2026.05 | 53.95 | 54.26 | 53.64 | 54.15 | 54.26 | 54.05 | — | |
| W2SDescription=Weak-to-strong distillation2026.05 | 52.31 | 55.18 | 56.1 | 59.28 | 56 | 55.77 | 15.38 |