Accuracy on ARC-C (Question Answering)
86Accuracy (ARC-C)ExpGraph
Evaluation Results
| Method | Links | |
|---|---|---|
| ExpGraphBackbone Model=Llama-3.1-8B-Instruct (Large LLM), Method Category=LLM-Centric Experience Learning Baselines2026.05 | 86 | |
| MemRLBackbone Model=Llama-3.1-8B-Instruct (Large LLM), Method Category=Retrieval-Centric Experience Learning Baselines2026.05 | 84.34 | |
| Mem0Backbone Model=Llama-3.1-8B-Instruct (Large LLM), Method Category=Retrieval-Centric Experience Learning Baselines2026.05 | 83.33 | |
| IRCoTBackbone Model=Llama-3.1-8B-Instruct (Large LLM), Method Category=LLM-Centric Experience Learning Baselines2026.05 | 83.08 | |
| ExpeLBackbone Model=Llama-3.1-8B-Instruct (Large LLM), Method Category=Retrieval-Centric Experience Learning Baselines2026.05 | 82.58 | |
| ReasoningBankBackbone Model=Llama-3.1-8B-Instruct (Large LLM), Method Category=Retrieval-Centric Experience Learning Baselines2026.05 | 82.32 | |
| S3Backbone Model=Llama-3.1-8B-Instruct (Large LLM), Method Category=LLM-Centric Experience Learning Baselines2026.05 | 82 | |
| LightMemBackbone Model=Llama-3.1-8B-Instruct (Large LLM), Method Category=Retrieval-Centric Experience Learning Baselines2026.05 | 81.57 | |
| AWMBackbone Model=Llama-3.1-8B-Instruct (Large LLM), Method Category=Retrieval-Centric Experience Learning Baselines2026.05 | 81.57 | |
| Search-o1Backbone Model=Llama-3.1-8B-Instruct (Large LLM), Method Category=LLM-Centric Experience Learning Baselines2026.05 | 81.31 | |
| ExpGraphBackbone Model=Llama-3.2-3B-Instruct (Small LLM), Method Category=LLM-Centric Experience Learning Baselines2026.05 | 74 | |
| AWMBackbone Model=Llama-3.2-3B-Instruct (Small LLM), Method Category=Retrieval-Centric Experience Learning Baselines2026.05 | 72.47 | |
| ReasoningBankBackbone Model=Llama-3.2-3B-Instruct (Small LLM), Method Category=Retrieval-Centric Experience Learning Baselines2026.05 | 71.72 | |
| LightMemBackbone Model=Llama-3.2-3B-Instruct (Small LLM), Method Category=Retrieval-Centric Experience Learning Baselines2026.05 | 71.46 | |
| No MemoryBackbone Model=Llama-3.1-8B-Instruct (Large LLM), Method Category=No Memory2026.05 | 70.2 | |
| Mem0Backbone Model=Llama-3.2-3B-Instruct (Small LLM), Method Category=Retrieval-Centric Experience Learning Baselines2026.05 | 67.93 | |
| MemRLBackbone Model=Llama-3.2-3B-Instruct (Small LLM), Method Category=Retrieval-Centric Experience Learning Baselines2026.05 | 67.42 | |
| Search-o1Backbone Model=Llama-3.2-3B-Instruct (Small LLM), Method Category=LLM-Centric Experience Learning Baselines2026.05 | 66.67 | |
| S3Backbone Model=Llama-3.2-3B-Instruct (Small LLM), Method Category=LLM-Centric Experience Learning Baselines2026.05 | 65 | |
| Qwen3Model Version=4B, Architecture=Dense, # Params=4.0B, # Tokens=36T2025.10 | 63.65 | |
| DreamReasoner-8BType=Block Diffusion2026.06 | 63.5 | |
| IRCoTBackbone Model=Llama-3.2-3B-Instruct (Small LLM), Method Category=LLM-Centric Experience Learning Baselines2026.05 | 62.63 | |
| Gemma3Model Version=4B, Architecture=Dense, # Params=4.0B, # Tokens=4T2025.10 | 60.92 | |
| OuroModel Version=1.4B R4, Architecture=LoopLM, # Params=1.4B, # Tokens=7.7T, recurrent steps=42025.10 | 60.92 | |
| Dream-v0-7BType=Diffusion2026.06 | 59.8 | |
| Qwen3-8BType=AR2026.06 | 56.9 | |
| Qwen3Model Version=1.7B, Architecture=Dense, # Params=1.7B, # Tokens=36T2025.10 | 55.72 | |
| Qwen2.5Model Version=3B, Architecture=Dense, # Params=3.0B, # Tokens=18T2025.10 | 55.46 | |
| Qwen2.5Model Version=1.5B, Architecture=Dense, # Params=1.5B, # Tokens=18T2025.10 | 54.44 | |
| OLMo3-7BType=AR2026.06 | 53.5 | |
| Llama3.2Model Version=3B, Architecture=Dense, # Params=3.0B, # Tokens=9T2025.10 | 52.47 | |
| MiMo-7BType=AR2026.06 | 51.6 | |
| No MemoryBackbone Model=Llama-3.2-3B-Instruct (Small LLM), Method Category=No Memory2026.05 | 51.52 | |
| Qwen2.5-7BType=AR2026.06 | 51.5 | |
| LLaDA-8BType=Diffusion2026.06 | 47.5 | |
| ExpeLBackbone Model=Llama-3.2-3B-Instruct (Small LLM), Method Category=Retrieval-Centric Experience Learning Baselines2026.05 | 44.7 | |
| Llama3.2Model Version=1.2B, Architecture=Dense, # Params=1.0B, # Tokens=9T2025.10 | 41.98 | |
| Gemma3Model Version=1B, Architecture=Dense, # Params=1.0B, # Tokens=2T2025.10 | 39.25 | |
| TaDABackbone=Llama-2-7b, Number of samples=5002026.06 | 37.4 | |
| LinearBackbone=Llama-2-7b, Number of samples=5002026.06 | 36 | |
| Mag. PruningBackbone=Llama-2-7b, Number of samples=5002026.06 | 35.8 | |
| TIESBackbone=Llama-2-7b, Number of samples=5002026.06 | 34.4 | |
| Domain onlyBackbone=Llama-2-7b, Number of samples=5002026.06 | 33.6 | |
| DARE-TIESBackbone=Llama-2-7b, Number of samples=500, Evaluation protocol=mean over 3 random seeds2026.06 | 33.6 | |
| DeltaNetevaluation_mode=zero-shot, parameters=2.8B, context_length=< 2K tokens2025.11 | 32.85 | |
| Gated DeltaNetevaluation_mode=zero-shot, parameters=2.8B, context_length=< 2K tokens2025.11 | 32.59 | |
| Gated KalmaNetevaluation_mode=zero-shot, parameters=2.8B, context_length=< 2K tokens2025.11 | 32.51 | |
| Transformerevaluation_mode=zero-shot, parameters=2.8B, context_length=< 2K tokens2025.11 | 32.25 | |
| Mamba2evaluation_mode=zero-shot, parameters=2.8B, context_length=< 2K tokens2025.11 | 32.24 | |
| Task onlyBackbone=Llama-2-7b, Number of samples=5002026.06 | 31.6 | |
| Base modelBackbone=Llama-2-7b, Number of samples=5002026.06 | 31 | |
| SVD MergeBackbone=Llama-2-7b, Number of samples=5002026.06 | 28.4 | |
| Gated Linear Attentionevaluation_mode=zero-shot, parameters=2.8B, context_length=< 2K tokens2025.11 | 27.82 | |
| DCDM (MoE)Scale=1.5B, Active Params=1.2B (2.8B), In-context examples=0-shot2026.05 | 27.5 | |
| DARE-LinBackbone=Llama-2-7b, Number of samples=500, Evaluation protocol=mean over 3 random seeds2026.06 | 27.2 | |
| DCDMScale=1.5B, Active Params=1.5B, In-context examples=0-shot2026.05 | 27.13 | |
| MDLMScale=1.5B, Active Params=1.5B, In-context examples=0-shot2026.05 | 26.79 | |
| BDLMScale=1.5B, Active Params=1.5B, In-context examples=0-shot2026.05 | 26.71 | |
| Task Arith.Backbone=Llama-2-7b, Number of samples=5002026.06 | 26.4 | |
| DCDM (MoE)Scale=0.5B, Active Params=0.4B (0.8B), In-context examples=0-shot2026.05 | 24.74 | |
| DCDMScale=0.5B, Active Params=0.5B, In-context examples=0-shot2026.05 | 24.49 | |
| MDLMScale=0.5B, Active Params=0.5B, In-context examples=0-shot2026.05 | 23.29 | |
| BDLMScale=0.5B, Active Params=0.5B, In-context examples=0-shot2026.05 | 23.21 | |
| DLR +CoLAModel Scale=1B, Rank (r)=512, Evaluation Protocol=Zero-shot2026.06 | 22.29 | |
| Full-RankModel Scale=1B, Rank (r)=512, Evaluation Protocol=Zero-shot2026.06 | 22.06 | |
| DLR +Low-RankModel Scale=1B, Rank (r)=512, Evaluation Protocol=Zero-shot2026.06 | 21.59 | |
| Low-RankModel Scale=1B, Rank (r)=512, Evaluation Protocol=Zero-shot2026.06 | 19.78 |