Accuracy on Physical Commonsense Reasoning on PIQA
81.5Accuracy (PIQA)Thoughts-as-Planning
Evaluation Results
| Method | Links | |
|---|---|---|
| Thoughts-as-PlanningAvg Edits=472026.04 | 81.5 | |
| OriginalBackbone=LLaMA3-8B2026.05 | 81.28 | |
| DEEPSEEK-7BBase Model=DeepSeek-7B, Sparsity Pattern=Dense, Evaluation Protocol=Zero-shot2025.06 | 79.27 | |
| CoTGenAvg Edits=1802026.04 | 79.2 | |
| RLCoTAvg Edits=1502026.04 | 78.8 | |
| Accuracy (Arc-E)Backbone=LLaMA3-8B, Pruning Ratio=25%2026.05 | 78.51 | |
| LLAMA-2-7BBase Model=LLaMA-2-7B, Sparsity Pattern=Dense, Evaluation Protocol=Zero-shot2025.06 | 78.07 | |
| AutoCoTAvg Edits=200+2026.04 | 78 | |
| SoftCoT2026.04 | 77.9 | |
| TaylorBackbone=LLaMA3-8B, Pruning Ratio=25%2026.05 | 76.55 | |
| Accuracy (Ours)Backbone=LLaMA3-8B, Pruning Ratio=25%2026.05 | 76.55 | |
| Manual CoT2026.04 | 76.5 | |
| FIXED_E=128Experts=1282026.05 | 76.44 | |
| Memory Grafting# Shots=5-shot, # Trainable Params=2.8B, # Activated (w/o token embed)=0.55B, # Trained Tokens=100B, # Experts (shared + routed, top-k)=1 + 47 (top-4)2026.05 | 76.17 | |
| MaskProBase Model=DeepSeek-7B, Sparsity Pattern=(4:8)-sparsity, Evaluation Protocol=Zero-shot2025.06 | 75.87 | |
| GBLMBase Model=DeepSeek-7B, Sparsity Pattern=(4:8)-sparsity, Evaluation Protocol=Zero-shot2025.06 | 75.73 | |
| Vanilla Engram# Shots=5-shot, # Trainable Params=2.8B, # Activated (w/o token embed)=0.55B, # Trained Tokens=100B, # Experts (shared + routed, top-k)=1 + 48 (top-4)2026.05 | 75.63 | |
| WANDABase Model=DeepSeek-7B, Sparsity Pattern=(4:8)-sparsity, Evaluation Protocol=Zero-shot2025.06 | 75.46 | |
| MoE Baseline# Shots=5-shot, # Trainable Params=2.8B, # Activated (w/o token embed)=0.55B, # Trained Tokens=100B, # Experts (shared + routed, top-k)=1 + 64 (top-4)2026.05 | 75.41 | |
| Out. Cosine-SimBackbone=LLaMA3-8B, Pruning Ratio=25%2026.05 | 75.3 | |
| SPARSEGPTBase Model=DeepSeek-7B, Sparsity Pattern=(4:8)-sparsity, Evaluation Protocol=Zero-shot2025.06 | 75.24 | |
| Oryx-TM (Mamba-2)Parameter Scale=1.4B, Family=Oryx-TM2026.05 | 75.2 | |
| PRUNER-ZEROBase Model=DeepSeek-7B, Sparsity Pattern=(4:8)-sparsity, Evaluation Protocol=Zero-shot2025.06 | 75.12 | |
| Oryx-TG (Gated DeltaNet)Parameter Scale=1.4B, Family=Oryx-TG2026.05 | 75.1 | |
| Gated DeltaNetParameter Scale=1.4B, Family=Baseline2026.05 | 75 | |
| Oryx-TM (Transformer)Parameter Scale=1.4B, Family=Oryx-TM2026.05 | 75 | |
| Oryx-TG (Transformer)Parameter Scale=1.4B, Family=Oryx-TG2026.05 | 74.9 | |
| Out. Norm-SimBackbone=LLaMA3-8B, Pruning Ratio=25%2026.05 | 74.86 | |
| EMO (Stage 5)E=64→1282026.05 | 74.81 | |
| Mamba-2Parameter Scale=1.4B, Family=Baseline2026.05 | 74.7 | |
| MaskProBase Model=LLaMA-2-7B, Sparsity Pattern=(4:8)-sparsity, Evaluation Protocol=Zero-shot2025.06 | 74.65 | |
| GBLMBase Model=LLaMA-2-7B, Sparsity Pattern=(4:8)-sparsity, Evaluation Protocol=Zero-shot2025.06 | 74.16 | |
| WANDABase Model=LLaMA-2-7B, Sparsity Pattern=(4:8)-sparsity, Evaluation Protocol=Zero-shot2025.06 | 74.14 | |
| TransformerParameter Scale=1.4B, Family=Baseline2026.05 | 74.1 | |
| PRUNER-ZEROBase Model=LLaMA-2-7B, Sparsity Pattern=(4:8)-sparsity, Evaluation Protocol=Zero-shot2025.06 | 74.07 | |
| Oryx-TM (Mamba-2)Parameter Scale=810M, Family=Oryx-TM2026.05 | 73.9 | |
| Out. Divergence-SimBackbone=LLaMA3-8B, Pruning Ratio=25%2026.05 | 73.83 | |
| SPARSEGPTBase Model=LLaMA-2-7B, Sparsity Pattern=(4:8)-sparsity, Evaluation Protocol=Zero-shot2025.06 | 73.78 | |
| FIXED_E=32Experts=322026.05 | 73.78 | |
| Slice-GPTBackbone=LLaMA3-8B, Pruning Ratio=25%2026.05 | 73.66 | |
| Oryx-TM (Transformer)Parameter Scale=810M, Family=Oryx-TM2026.05 | 73.6 | |
| EMO (Stage 4)E=32→642026.05 | 73.5 | |
| Cosine SimilarityBackbone=LLaMA3-8B, Pruning Ratio=25%2026.05 | 73.23 | |
| TransformerParameter Scale=810M, Family=Baseline2026.05 | 73.1 | |
| Oryx-TG (Gated DeltaNet)Parameter Scale=810M, Family=Oryx-TG2026.05 | 73.1 | |
| FIXED_E=16Experts=162026.05 | 73.01 | |
| Oryx-TG (Transformer)Parameter Scale=810M, Family=Oryx-TG2026.05 | 72.8 | |
| Mamba-2Parameter Scale=810M, Family=Baseline2026.05 | 72.7 | |
| Accuracy (C4)Backbone=LLaMA3-8B, Pruning Ratio=25%2026.05 | 72.63 | |
| Gated DeltaNetParameter Scale=810M, Family=Baseline2026.05 | 72.5 | |
| MAGNITUDEBase Model=DeepSeek-7B, Sparsity Pattern=(4:8)-sparsity, Evaluation Protocol=Zero-shot2025.06 | 72.42 | |
| MuonArchitecture=LLaMA-1.3B, Training Tokens=54B2026.05 | 72.2 | |
| MAGNITUDEBase Model=LLaMA-2-7B, Sparsity Pattern=(4:8)-sparsity, Evaluation Protocol=Zero-shot2025.06 | 72.2 | |
| Mamba-2Parameter Scale=380M, Family=Baseline2026.05 | 72.1 | |
| Oryx-TG (Transformer)Parameter Scale=380M, Family=Oryx-TG2026.05 | 71.6 | |
| Oryx-TM (Mamba-2)Parameter Scale=380M, Family=Oryx-TM2026.05 | 71.3 | |
| Oryx-TG (Gated DeltaNet)Parameter Scale=380M, Family=Oryx-TG2026.05 | 71.3 | |
| AdamWArchitecture=LLaMA-1.3B, Training Tokens=54B2026.05 | 71.27 | |
| PionArchitecture=LLaMA-1.3B, Training Tokens=54B2026.05 | 71.27 | |
| EMO (Stage 3)E=16→322026.05 | 71.27 | |
| Gated DeltaNetParameter Scale=380M, Family=Baseline2026.05 | 71.2 | |
| Oryx-TM (Transformer)Parameter Scale=380M, Family=Oryx-TM2026.05 | 71.2 | |
| TransformerParameter Scale=380M, Family=Baseline2026.05 | 70.8 | |
| EMO (Stage 2)E=8→162026.05 | 70.13 | |
| EMO (Stage 1)E=82026.05 | 69.91 | |
| Muon32Model=LLaMA 1.1B, Evaluation Protocol=Zero-shot2026.05 | 69.6 | |
| Muon8Model=LLaMA 1.1B, Evaluation Protocol=Zero-shot2026.05 | 69.6 | |
| MOE+Lngram2026.05 | 69.26 | |
| MOE+Engram2026.05 | 68.44 | |
| Gated DeltaNetParameter Scale=130M, Family=Baseline2026.05 | 68.4 | |
| MuonQ4Model=LLaMA 1.1B, Evaluation Protocol=Zero-shot2026.05 | 68.3 | |
| MOE2026.05 | 68.01 | |
| Oryx-TM (Transformer)Parameter Scale=130M, Family=Oryx-TM2026.05 | 67.7 | |
| MOE+Lngram-23Lconfiguration=23L2026.05 | 67.4 | |
| Oryx-TM (Mamba-2)Parameter Scale=130M, Family=Oryx-TM2026.05 | 67.4 | |
| Oryx-TG (Transformer)Parameter Scale=130M, Family=Oryx-TG2026.05 | 67.3 | |
| Muon32Model=GPT-2 Large, Evaluation Protocol=Zero-shot2026.05 | 67.2 | |
| Mamba-2Parameter Scale=130M, Family=Baseline2026.05 | 67.1 | |
| Oryx-TG (Gated DeltaNet)Parameter Scale=130M, Family=Oryx-TG2026.05 | 67.1 | |
| TransformerParameter Scale=130M, Family=Baseline2026.05 | 66.8 | |
| PerplexityBackbone=LLaMA3-8B, Pruning Ratio=25%2026.05 | 66.76 | |
| Muon8Model=GPT-2 Large, Evaluation Protocol=Zero-shot2026.05 | 66.7 | |
| Muon32Model=LLaMA 350M, Evaluation Protocol=Zero-shot2026.05 | 65.7 | |
| MuonQ4Model=GPT-2 Large, Evaluation Protocol=Zero-shot2026.05 | 65.2 | |
| Muon32Model=GPT-2 Medium, Evaluation Protocol=Zero-shot2026.05 | 65.1 | |
| Muon8Model=GPT-2 Medium, Evaluation Protocol=Zero-shot2026.05 | 64.6 | |
| MuonQ4Model=LLaMA 350M, Evaluation Protocol=Zero-shot2026.05 | 63.7 | |
| Muon8Model=LLaMA 350M, Evaluation Protocol=Zero-shot2026.05 | 63.4 | |
| MuonQ4Model=GPT-2 Medium, Evaluation Protocol=Zero-shot2026.05 | 63.1 | |
| Gaussian-BaseCPT=Gaussian, SFT=Base2026.05 | 59.3 | |
| Gaussian-GaussianCPT=Gaussian, SFT=Gaussian2026.05 | 59.3 | |
| Muon4Model=LLaMA 1.1B, Evaluation Protocol=Zero-shot2026.05 | 58.7 | |
| Muon4Model=GPT-2 Medium, Evaluation Protocol=Zero-shot2026.05 | 58.2 | |
| Muon4Model=GPT-2 Large, Evaluation Protocol=Zero-shot2026.05 | 58.2 | |
| BaseCPT=Base, SFT=Base2026.05 | 56.8 | |
| Base-GaussianCPT=Base, SFT=Gaussian2026.05 | 56.47 | |
| Muon4Model=LLaMA 350M, Evaluation Protocol=Zero-shot2026.05 | 55.9 | |
| GaussianCPT=Gaussian, SFT=None2026.05 | 53.75 | |
| BaseCPT=Base, SFT=None2026.05 | 52.12 |