Reading Comprehension on DROP
92.2F1 ScoreDeepSeek-R1
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| DeepSeek-R1Architecture=MoE, Activated Params=37B, Total Params=671B, shots=3-shot2025.01 | 92.2 | — | |
| DeepSeek-V3Architecture=MoE, Activated Params=37B, Total Params=671B, shots=3-shot2025.01 | 91.6 | — | |
| OpenAI-o1-1217shots=3-shot2025.01 | 90.2 | — | |
| Fine-tuned SOTASetting=Fine-tuned2020.05 | 89.1 | — | |
| QDGAT2023.05 | 88.4 | — | |
| Claude-3.5-Sonnet-1022shots=3-shot2025.01 | 88.3 | — | |
| Claude 3.5 Sonnet-1022Few-shot (3-shot)=true, Model Backbone=Claude 3.5 Sonnet2026.03 | 88.3 | — | |
| DeepSeek-V3.2 Exp Base# Shots=3-shot, # Activated Params=37B, # Total Params=671B2026.01 | 86.6 | — | |
| DeepSeek-V3.1 Base# Shots=3-shot, # Activated Params=37B, # Total Params=671B2026.01 | 86.3 | — | |
| R1-Distill-Qwen-14BFew-shot (3-shot)=true, Model Backbone=Qwen-14B2026.03 | 85.5 | — | |
| PaLM 2Number of exemplars (k-shot)=3, Model variant=Instruction-tuned2023.05 | 85 | — | |
| Llama 3 405B2024.07 | 84.8 | — | |
| MiMo-V2-Flash Base# Shots=3-shot, # Activated Params=15B, # Total Params=309B2026.01 | 84.7 | — | |
| MiMo-RL 7B-R-TAPFew-shot (3-shot)=true, Model Backbone=7B2026.03 | 84.5 | — | |
| OpenAI-o1-minishots=3-shot2025.01 | 83.9 | — | |
| OpenAI o1-miniFew-shot (3-shot)=true, Model Backbone=o1-mini2026.03 | 83.9 | — | |
| GPT-4o-0513shots=3-shot2025.01 | 83.7 | — | |
| GPT-4o 0513Few-shot (3-shot)=true, Model Backbone=GPT-4o2026.03 | 83.7 | — | |
| Kimi-K2 Base# Shots=3-shot, # Activated Params=32B, # Total Params=1043B2026.01 | 83.6 | — | |
| Gemini UltraEvaluation protocol=Variable shot2024.07 | 82.4 | — | |
| HRM-Text 1BArchitecture=Recurrent, FLOPs ($10^{21}$)=1, Tokens (T)=0.062026.05 | 82.2 | — | |
| GPT-4Number of exemplars (k-shot)=32023.05 | 80.9 | — | |
| GPT-42024.07 | 80.9 | — | |
| Llama 3 70B2024.07 | 79.6 | — | |
| QuanTAModel=LLaMA2-70B, Config=16-8-8-8, # Params (%)=0.014%2024.05 | 79.4 | — | |
| MiMo 7B-RLFew-shot (3-shot)=true, Model Backbone=7B2026.03 | 78.7 | — | |
| Mixtral 8x22B2024.07 | 77.5 | — | |
| R1-Distill-Qwen-7BFew-shot (3-shot)=true, Model Backbone=Qwen-7B2026.03 | 77 | — | |
| LoRAModel=LLaMA2-70B, rank=r=8, # Params (%)=0.024%2024.05 | 74.3 | — | |
| Olmo3 7BArchitecture=Dense, FLOPs ($10^{21}$)=252, Tokens (T)=62026.05 | 71.5 | — | |
| QwQ-32B PreviewFew-shot (3-shot)=true, Model Backbone=QwQ-32B2026.03 | 71.2 | — | |
| MEDALBackbone=LLaDA1.52025.12 | 71.1 | — | |
| MEDALBackbone=LLaDA2025.12 | 71 | — | |
| PaLMNumber of exemplars (k-shot)=12023.05 | 70.8 | — | |
| QuanTAModel=LLaMA2-13B, Config=16-8-8-5, # Params (%)=0.029%2024.05 | 69 | — | |
| MEDALBackbone=Dream2025.12 | 65.4 | — | |
| Bst5Backbone=LLaDA1.52025.12 | 65 | — | |
| Bst5Backbone=LLaDA2025.12 | 64.3 | — | |
| LoRAModel=LLaMA2-13B, rank=r=8, # Params (%)=0.050%2024.05 | 61 | — | |
| Engram-40BShots=1-shot2026.01 | 60.7 | — | |
| LLaDA1.5Backbone=LLaDA1.52025.12 | 60.5 | — | |
| Gemma3 4BArchitecture=Dense, FLOPs ($10^{21}$)=96, Tokens (T)=42026.05 | 60.1 | — | |
| LlamaBackbone=Llama2025.12 | 60 | — | |
| QuanTAModel=LLaMA2-7B, Config=16-16-16, # Params (%)=0.261%2024.05 | 59.6 | — | |
| QuanTAModel=LLaMA2-7B, Config=16-8-8-4, # Params (%)=0.041%2024.05 | 59.5 | — | |
| Llama 3 8B2024.07 | 59.5 | — | |
| Full Fine-TuningModel=LLaMA2-7B, # Params (%)=100%2024.05 | 59.4 | — | |
| Parallel AdapterModel=LLaMA2-7B, # Params (%)=0.747%2024.05 | 59 | — | |
| Engram-27BShots=1-shot2026.01 | 59 | — | |
| Series AdapterModel=LLaMA2-7B, # Params (%)=0.747%2024.05 | 58.8 | — | |
| LLaDABackbone=LLaDA2025.12 | 58.2 | — | |
| Bst5Backbone=Dream2025.12 | 57.7 | — | |
| Gemma 7B2024.07 | 56.3 | — | |
| LoRAModel=LLaMA2-7B, rank=r=128, # Params (%)=0.996%2024.05 | 56.2 | — | |
| MoE-27BShots=1-shot2026.01 | 55.7 | — | |
| DreamBackbone=Dream2025.12 | 55.4 | — | |
| LoRAModel=LLaMA2-7B, rank=r=32, # Params (%)=0.249%2024.05 | 54.8 | — | |
| LoRAModel=LLaMA2-7B, rank=r=8, # Params (%)=0.062%2024.05 | 54 | — | |
| 27B w/ mHC# Shots=3-shot2025.12 | 53.9 | — | |
| Mistral 7B2024.07 | 53 | — | |
| 27B w/ HC# Shots=3-shot2025.12 | 51.6 | — | |
| Ouro 1.4BArchitecture=Recurrent, FLOPs ($10^{21}$)=259, Tokens (T)=72026.05 | 49.7 | — | |
| 27B Baseline# Shots=3-shot2025.12 | 47 | — | |
| Llama3.2 3BArchitecture=Dense, FLOPs ($10^{21}$)=162, Tokens (T)=92026.05 | 45.2 | — | |
| Dense-4BShots=1-shot2026.01 | 41.6 | — | |
| SeCoBackbone=LLaMA-3.2-1B-Instruct, Compression constraint=16x2026.05 | 39.94 | 26.61 | |
| SeCoBackbone=LLaMA-3.2-1B-Instruct, Compression constraint=32x2026.05 | 39.07 | 25.95 | |
| GPT-3Setting=Few-Shot2020.05 | 36.5 | — | |
| GPT-3Setting=One-Shot2020.05 | 34.3 | — | |
| SAC (EPL)Backbone=LLaMA-3.2-1B-Instruct, Compression constraint=16x2026.05 | 33.89 | 24.55 | |
| 500x (EPL)Backbone=LLaMA-3.2-1B-Instruct, Compression constraint=16x2026.05 | 32.85 | 23.02 | |
| ICAE (EPL)Backbone=LLaMA-3.2-1B-Instruct, Compression constraint=16x2026.05 | 31.55 | 21.82 | |
| Qwen3.5 2BArchitecture=Dense, FLOPs ($10^{21}$)=432, Tokens (T)=362026.05 | 30.8 | — | |
| ICAE (EPL)Backbone=LLaMA-3.2-1B-Instruct, Compression constraint=32x2026.05 | 30.45 | 21.22 | |
| SAC (EPL)Backbone=LLaMA-3.2-1B-Instruct, Compression constraint=32x2026.05 | 30.42 | 20.83 | |
| 500x (DPL)Backbone=LLaMA-3.2-1B-Instruct, Compression constraint=16x2026.05 | 29.94 | 21.42 | |
| 500x (EPL)Backbone=LLaMA-3.2-1B-Instruct, Compression constraint=32x2026.05 | 29.43 | 20.96 | |
| 500x (DPL)Backbone=LLaMA-3.2-1B-Instruct, Compression constraint=32x2026.05 | 28.35 | 19.96 | |
| ConMeZOModel=OPT-1.3B2025.11 | 26.53 | — | |
| MeZOModel=OPT-1.3B2025.11 | 25.9 | — | |
| COMIBackbone=LLaMA-3.2-1B-Instruct, Compression constraint=16x2026.05 | 24.67 | 16.1 | |
| RAMBackbone=LLaMA-3.2-1B-Instruct, Compression constraint=32x2026.05 | 23.67 | 16.5 | |
| GPT-3Setting=Zero-Shot2020.05 | 23.6 | — | |
| COMIBackbone=LLaMA-3.2-1B-Instruct, Compression constraint=32x2026.05 | 21.38 | 13.91 | |
| RAMBackbone=LLaMA-3.2-1B-Instruct, Compression constraint=16x2026.05 | 18.39 | 11.11 | |
| Huginn 3.5BArchitecture=Recurrent, FLOPs ($10^{21}$)=127, Tokens (T)=0.82026.05 | 17.8 | — | |
| TinyLlama v1.0Few-shot (3-shot)=true, Training tokens=2T2026.03 | 15.34 | — | |
| TinyLlama v1.1Few-shot (3-shot)=true, Training tokens=2T2026.03 | 15.31 | — | |
| OPT-1.3B FFFFew-shot (3-shot)=true, Model Depth (d)=4, Training tokens=26B, Fine-tuning (FT)=true2026.03 | 15.29 | — | |
| OPT-1.3B FFFFew-shot (3-shot)=true, Model Depth (d)=6, Training tokens=26B, Fine-tuning (FT)=true2026.03 | 15.28 | — | |
| OPT-1.3B FFFew-shot (3-shot)=true, Training configuration=retrained2026.03 | 15.27 | — | |
| OPT-1.3B FFFFew-shot (3-shot)=true, Model Depth (d)=4, Training tokens=26B, Fine-tuning (FT)=false2026.03 | 15.25 | — | |
| OPT-1.3BFew-shot (3-shot)=true, Training tokens=300B2026.03 | 14.32 | — | |
| OPT-1.3B FFFFew-shot (3-shot)=true, Model Depth (d)=6, Training tokens=26B, Fine-tuning (FT)=false2026.03 | 14.25 | — | |
| Pythia-1.4BFew-shot (3-shot)=true, Training tokens=1T2026.03 | 12.27 | — | |
| Pythia-1.0BFew-shot (3-shot)=true, Training tokens=1T2026.03 | 4.25 | — |