Multitask Language Understanding on MMLU (Average Accuracy)
79.64MMLU AccuracyFP16
Evaluation Results
| Method | Links | |
|---|---|---|
| FP16Backbone=Qwen3-30B-A3B, Evaluation Mode=Few-shot2026.06 | 79.64 | |
| SharQBackbone=Qwen3-30B-A3B, Evaluation Mode=Few-shot, Sparsity Pattern=4:8, Base Quantization=NVFP42026.06 | 78.79 | |
| NVFP4Backbone=Qwen3-30B-A3B, Evaluation Mode=Few-shot2026.06 | 77.72 | |
| FP16Backbone=Qwen2.5-7B, Evaluation Mode=Few-shot2026.06 | 74.16 | |
| SharQBackbone=Qwen2.5-7B, Evaluation Mode=Few-shot, Sparsity Pattern=4:8, Base Quantization=NVFP42026.06 | 72.83 | |
| NVFP4Backbone=Qwen2.5-7B, Evaluation Mode=Few-shot2026.06 | 72.06 | |
| FP16Backbone=Llama-3.1-8B, Evaluation Mode=Few-shot2026.06 | 65.24 | |
| SharQBackbone=Llama-3.1-8B, Evaluation Mode=Few-shot, Sparsity Pattern=4:8, Base Quantization=NVFP42026.06 | 63.76 | |
| MoD-Single RoundCategory=Text, Backbone=Qwen2.5VL-3b-Instruct, Protocol=single-round inference2026.06 | 63.73 | |
| Self-CorrectionCategory=Text, Backbone=Qwen2.5VL-3b-Instruct2026.06 | 63.7 | |
| MoD-Multi RoundCategory=Text, Backbone=Qwen2.5VL-3b-Instruct, Protocol=multi-turn dialectical reasoning2026.06 | 63.69 | |
| MoE-LoRACategory=Text, Backbone=Qwen2.5VL-3b-Instruct2026.06 | 63.59 | |
| Multi-Agent DebateCategory=Text, Backbone=Qwen2.5VL-3b-Instruct2026.06 | 63.21 | |
| Qwen2.5VL-3b-InstructCategory=Text2026.06 | 62.84 | |
| NVFP4Backbone=Llama-3.1-8B, Evaluation Mode=Few-shot2026.06 | 61.93 | |
| MoD-Multi RoundCategory=Text, Backbone=LLaVA-v1.6-13b, Protocol=multi-turn dialectical reasoning2026.06 | 56.35 | |
| Self-CorrectionCategory=Text, Backbone=LLaVA-v1.6-13b2026.06 | 56.06 | |
| Multi-Agent DebateCategory=Text, Backbone=LLaVA-v1.6-13b2026.06 | 55.97 | |
| MoD-Single RoundCategory=Text, Backbone=LLaVA-v1.6-13b, Protocol=single-round inference2026.06 | 55.95 | |
| MoE-LoRACategory=Text, Backbone=LLaVA-v1.6-13b2026.06 | 55.62 | |
| LLaVA-v1.6-13bCategory=Text2026.06 | 55.42 | |
| Qwen-VL-ChatCategory=Text2026.06 | 50.7 | |
| Original ModelBackbone=Qwen-3-4B, Avg. Sparsity (L/H)=0% / 0%, Mem (GB)=5.80, Time (s)=41002026.06 | 45.1 | |
| Learning to AllocateBackbone=Qwen-3-4B, Avg. Sparsity (L/H)=38% / 28%, Mem (GB)=5.30, Time (s)=28502026.06 | 44.5 | |
| TTT+CT-KVTraining time per task (s)=342025.07 | 44.1 | |
| CT-KVTraining time per task (s)=92025.07 | 43.7 | |
| TTTTraining time per task (s)=302025.07 | 43.6 | |
| CT-PromptTraining time per task (s)=332025.07 | 43.6 | |
| FlexiDepthBackbone=Qwen-3-4B, Avg. Sparsity (L/H)=34% / –, Mem (GB)=5.75, Time (s)=29102026.06 | 42.5 | |
| AdaSkipBackbone=Qwen-3-4B, Avg. Sparsity (L/H)=36% / –, Mem (GB)=5.65, Time (s)=28002026.06 | 41.2 | |
| ICLTraining time per task (s)=0, Evaluation protocol=In-Context Learning2025.07 | 41.2 | |
| DoRATraining time per task (s)=162025.07 | 40.3 | |
| LoRATraining time per task (s)=162025.07 | 40.1 | |
| Prefix TuningTraining time per task (s)=5, m=322025.07 | 39.9 | |
| Oracle Static PruningBackbone=Qwen-3-4B, Avg. Sparsity (L/H)=33% / 22%, Mem (GB)=4.95, Time (s)=26802026.06 | 39.8 | |
| Prompt TuningTraining time per task (s)=15, m=322025.07 | 39.2 | |
| Rank-Stabilized LoRATraining time per task (s)=172025.07 | 38.8 | |
| Prefix TuningTraining time per task (s)=8, m=# demo2025.07 | 38.8 | |
| Prompt TuningTraining time per task (s)=29, m=# demo2025.07 | 37.3 | |
| Zero-ShotTraining time per task (s)=0, Evaluation protocol=Zero-Shot2025.07 | 35.8 | |
| Static PruningBackbone=Qwen-3-4B, Avg. Sparsity (L/H)=33% / 0%, Mem (GB)=5.10, Time (s)=27502026.06 | 34.5 |