Boolean Question Answering on BoolQ (Accuracy and Speed)
85.9AccuracyTALE
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| TALEBackbone=LLaMA 3.1 8B, Evaluation Protocol=0-shot, Layers Removed (#D)=32025.10 | 85.9 | 12 | -8.8 | |
| BaselineBackbone=LLaMA 3.1 8B, Evaluation Protocol=0-shot, Layers Removed (#D)=02025.10 | 85.4 | — | — | |
| BSBABackbone=LLaMA 3.1 8B, Evaluation Protocol=0-shot, Layers Removed (#D)=72025.10 | 85.4 | — | -17.6 | |
| TALEBackbone=Mistral 7B, Evaluation Protocol=0-shot, Layers Removed (#D)=42025.10 | 84.4 | 16 | -18.5 | |
| TALEBackbone=Qwen 2.5 7B, Evaluation Protocol=0-shot, Layers Removed (#D)=42025.10 | 83.9 | 14 | -13.3 | |
| OriginalBackbone=Mistral-7B, Pruning Ratio=0%2026.05 | 83.73 | — | — | |
| BSBABackbone=Qwen 2.5 7B, Evaluation Protocol=0-shot, Layers Removed (#D)=52025.10 | 82.7 | — | -23.2 | |
| TokAlign++Backbone=LLaMA38B, Initialization Method=TokAlign++, #GPU Hour=197.21, Shots=52026.05 | 82.6 | — | — | |
| AccBackbone=Mistral-7B, Pruning Ratio=25%, Relevance Estimation Source=Task-specific2026.05 | 82.29 | — | — | |
| BaselineBackbone=Qwen 2.5 7B, Evaluation Protocol=0-shot, Layers Removed (#D)=02025.10 | 81.9 | — | — | |
| BaselineBackbone=Mistral 7B, Evaluation Protocol=0-shot, Layers Removed (#D)=02025.10 | 81.3 | — | — | |
| BSBABackbone=Mistral 7B, Evaluation Protocol=0-shot, Layers Removed (#D)=52025.10 | 80.8 | — | -27.7 | |
| TokAlign++Backbone=LLaMA38B, Initialization Method=TokAlign++, #GPU Hour=197.21, Shots=02026.05 | 80.61 | — | — | |
| DenseBase Model=LLaMA-2-7B, Sparsity=Dense2025.06 | 77.68 | — | — | |
| FIXED_E=128Experts=1282026.05 | 74.89 | — | — | |
| TALEBackbone=Lucie 7B, Evaluation Protocol=0-shot, Layers Removed (#D)=52025.10 | 74 | 30 | -17.2 | |
| DenseBase Model=DeepSeek-7B, Sparsity=Dense2025.06 | 72.81 | — | — | |
| FIXED_E=32Experts=322026.05 | 71.71 | — | — | |
| MaskProBase Model=LLaMA-2-7B, Sparsity=2:42025.06 | 71.12 | — | — | |
| SparsegptBase Model=LLaMA-2-7B, Sparsity=2:42025.06 | 71.1 | — | — | |
| EMO (Stage 4)E=32→642026.05 | 70.52 | — | — | |
| EMO (Stage 5)E=64→1282026.05 | 70.4 | — | — | |
| EMO (Stage 3)E=16→322026.05 | 69.72 | — | — | |
| Pruner-ZBase Model=LLaMA-2-7B, Sparsity=2:42025.06 | 69.13 | — | — | |
| EMO (Stage 2)E=8→162026.05 | 67.8 | — | — | |
| MaskProBase Model=DeepSeek-7B, Sparsity=2:42025.06 | 67.77 | — | — | |
| SparsegptBase Model=DeepSeek-7B, Sparsity=2:42025.06 | 66.91 | — | — | |
| EMO (Stage 1)E=82026.05 | 66.85 | — | — | |
| Cosine SimilarityBackbone=Mistral-7B, Pruning Ratio=25%, Relevance Estimation Source=Task-specific2026.05 | 66.73 | — | — | |
| Out. Norm-SimBackbone=Mistral-7B, Pruning Ratio=25%, Relevance Estimation Source=Task-specific2026.05 | 66.54 | — | — | |
| Pruner-ZBase Model=DeepSeek-7B, Sparsity=2:42025.06 | 66.36 | — | — | |
| PerplexityBackbone=Mistral-7B, Pruning Ratio=25%, Relevance Estimation Source=Task-specific2026.05 | 64.86 | — | — | |
| Out. Cosine-SimBackbone=Mistral-7B, Pruning Ratio=25%, Relevance Estimation Source=Task-specific2026.05 | 64.62 | — | — | |
| TokAlign++Backbone=Pythia2.8B, Initialization Method=TokAlign++, #GPU Hour=38.96, Shots=02026.05 | 64.43 | — | — | |
| BSBABackbone=Lucie 7B, Evaluation Protocol=0-shot, Layers Removed (#D)=192025.10 | 63 | — | -54.2 | |
| Memory Grafting# Shots=0-shot, # Trainable Params=2.8B, # Activated (w/o token embed)=0.55B, # Trained Tokens=100B, # Experts (shared + routed, top-k)=1 + 47 (top-4)2026.05 | 62.54 | — | — | |
| TokAlign++Backbone=Pythia2.8B, Initialization Method=TokAlign++, #GPU Hour=38.96, Shots=52026.05 | 60.98 | — | — | |
| TokAlign++Backbone=Pythia1B, Initialization Method=TokAlign++, #GPU Hour=19.94, Shots=02026.05 | 60.55 | — | — | |
| Pythia1BBackbone=Pythia1B, Initialization Method=Vanilla, #GPU Hour=−, Shots=02026.05 | 60.43 | — | — | |
| Vanilla Engram# Shots=0-shot, # Trainable Params=2.8B, # Activated (w/o token embed)=0.55B, # Trained Tokens=100B, # Experts (shared + routed, top-k)=1 + 48 (top-4)2026.05 | 60.24 | — | — | |
| FIXED_E=16Experts=162026.05 | 59.02 | — | — | |
| Out. Divergence-SimBackbone=Mistral-7B, Pruning Ratio=25%, Relevance Estimation Source=Task-specific2026.05 | 58.2 | — | — | |
| MoE Baseline# Shots=0-shot, # Trainable Params=2.8B, # Activated (w/o token embed)=0.55B, # Trained Tokens=100B, # Experts (shared + routed, top-k)=1 + 64 (top-4)2026.05 | 57.92 | — | — | |
| PionArchitecture=LLaMA-1.3B, Training Tokens=54B2026.05 | 57.58 | — | — | |
| Pythia1BBackbone=Pythia1B, Initialization Method=Vanilla, #GPU Hour=−, Shots=52026.05 | 57.37 | — | — | |
| TaylorBackbone=Mistral-7B, Pruning Ratio=25%, Relevance Estimation Source=Task-specific2026.05 | 55.72 | — | — | |
| TokAlign++Backbone=Pythia1B, Initialization Method=TokAlign++, #GPU Hour=19.94, Shots=52026.05 | 54.25 | — | — | |
| BaselineBackbone=Lucie 7B, Evaluation Protocol=0-shot, Layers Removed (#D)=02025.10 | 53 | — | — | |
| MuonArchitecture=LLaMA-1.3B, Training Tokens=54B2026.05 | 51.56 | — | — | |
| CCDD-MMDiT# params.=144.1M, Evaluation Protocol=Zero-shot2025.10 | 51.1 | — | — | |
| GIDD# params.=92.1M, Evaluation Protocol=Zero-shot2025.10 | 50.43 | — | — | |
| CCDD-MoEDiT# params.=104.0M, Evaluation Protocol=Zero-shot2025.10 | 50.21 | — | — | |
| GIDD-base# params.=320M, Evaluation Protocol=Zero-shot2025.10 | 49.57 | — | — | |
| MDLM# params.=92.1M, Evaluation Protocol=Zero-shot2025.10 | 49.42 | — | — | |
| GPT2# params.=110M, Evaluation Protocol=Zero-shot2025.10 | 48.72 | — | — | |
| AdamWArchitecture=LLaMA-1.3B, Training Tokens=54B2026.05 | 46.3 | — | — | |
| Llama (retrain)# params.=117M, Evaluation Protocol=Zero-shot2025.10 | 46.21 | — | — |