Multitask Language Understanding on MMLU (Accuracy and Throughput)
63.34AccuracyUDS
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| UDSBackbone=Qwen-2.5-7B, Zero-shot Evaluation=true2025.10 | 63.34 | 3.41 | |
| GREATSBackbone=Qwen-2.5-7B, Zero-shot Evaluation=true2025.10 | 58.19 | 2.12 | |
| RHO-LossBackbone=Qwen-2.5-7B, Zero-shot Evaluation=true2025.10 | 57.08 | 1.94 | |
| RegularBackbone=Qwen-2.5-7B, Zero-shot Evaluation=true2025.10 | 55.32 | 2.27 | |
| MaxLossBackbone=Qwen-2.5-7B, Zero-shot Evaluation=true2025.10 | 54.51 | 5.93 | |
| MaxGradBackbone=Qwen-2.5-7B, Zero-shot Evaluation=true2025.10 | 54.33 | 0.31 | |
| RandomBackbone=Qwen-2.5-7B, Zero-shot Evaluation=true2025.10 | 54.26 | 9.29 | |
| UDSBackbone=Llama-3.1-8B, Zero-shot Evaluation=true2025.10 | 40.16 | 2.48 | |
| GREATSBackbone=Llama-3.1-8B, Zero-shot Evaluation=true2025.10 | 39.04 | 1.88 | |
| RegularBackbone=Llama-3.1-8B, Zero-shot Evaluation=true2025.10 | 38.24 | 2.09 | |
| RHO-LossBackbone=Llama-3.1-8B, Zero-shot Evaluation=true2025.10 | 37.63 | 1.4 | |
| MaxGradBackbone=Llama-3.1-8B, Zero-shot Evaluation=true2025.10 | 35.91 | 0.29 | |
| MaxLossBackbone=Llama-3.1-8B, Zero-shot Evaluation=true2025.10 | 35.62 | 2.75 | |
| RandomBackbone=Llama-3.1-8B, Zero-shot Evaluation=true2025.10 | 35.47 | 3.74 |