Question Answering on OpenbookQA (Accuracy and Normalized accuracy)
87.6AccuracyQwen3-8B + SFT + WeMask(TF)
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Qwen3-8B + SFT + WeMask(TF)shot=0-shot, mask rate=0.1, base model=Qwen3-8B2026.05 | 87.6 | — | |
| Qwen3-8B + WeMask(SFT)shot=0-shot, mask rate=0.1, base model=Qwen3-8B2026.05 | 87.2 | — | |
| Qwen3-8B + SFTshot=0-shot, mask rate=0.1, base model=Qwen3-8B2026.05 | 87 | — | |
| AdaFuseBackbone=Mistral-7B, Evaluation Protocol=domain-specific fine-tuning2026.03 | 86.6 | — | |
| PESC (block-wise)Backbone=Mistral-7B, Evaluation Protocol=domain-specific fine-tuning2026.03 | 86.4 | — | |
| MoRAL (layer-wise)Backbone=Mistral-7B, Evaluation Protocol=domain-specific fine-tuning2026.03 | 85.8 | — | |
| LoRABackbone=Mistral-7B, Evaluation Protocol=domain-specific fine-tuning2026.03 | 84.2 | — | |
| Qwen3-4B + SFT + WeMask(TF)shot=0-shot, mask rate=0.1, base model=Qwen3-4B2026.05 | 82.2 | — | |
| Qwen3-4B + SFTshot=0-shot, mask rate=0.1, base model=Qwen3-4B2026.05 | 81.4 | — | |
| Qwen3-4B + WeMask(SFT)shot=0-shot, mask rate=0.1, base model=Qwen3-4B2026.05 | 81 | — | |
| Qwen3-8B + Gated Attentionshot=0-shot, mask rate=0.1, base model=Qwen3-8B2026.05 | 79.4 | — | |
| Qwen3-4B + Gated Attentionshot=0-shot, mask rate=0.1, base model=Qwen3-4B2026.05 | 77.8 | — | |
| Mistral-7B (base)Backbone=Mistral-7B, Evaluation Protocol=Base2026.03 | 57.8 | — | |
| Nemotron-H-8BParameters=8B, FLOPs / token=1.51263 × 10^102026.03 | 35.4 | 47.4 | |
| Swimba-14BParameters=14B, Experts=4, FLOPs / token=1.51282 × 10^102026.03 | 34.9 | 46.7 | |
| TokAlign++Backbone=LLaMA38B, Initialization Method=TokAlign++, #GPU Hour=197.21, Shots=52026.05 | 33.8 | — | |
| TokAlign++Backbone=LLaMA38B, Initialization Method=TokAlign++, #GPU Hour=197.21, Shots=02026.05 | 31.2 | — | |
| Qwen3-8Bshot=0-shot, mask rate=0.1, base model=Qwen3-8B2026.05 | 27.6 | — | |
| TokAlign++Backbone=Pythia2.8B, Initialization Method=TokAlign++, #GPU Hour=38.96, Shots=52026.05 | 25.8 | — | |
| TokAlign++Backbone=Pythia2.8B, Initialization Method=TokAlign++, #GPU Hour=38.96, Shots=02026.05 | 24.8 | — | |
| TokAlign++Backbone=Pythia1B, Initialization Method=TokAlign++, #GPU Hour=19.94, Shots=52026.05 | 21.4 | — | |
| Qwen3-4Bshot=0-shot, mask rate=0.1, base model=Qwen3-4B2026.05 | 19.6 | — | |
| TokAlign++Backbone=Pythia1B, Initialization Method=TokAlign++, #GPU Hour=19.94, Shots=02026.05 | 19.6 | — |