Reading Comprehension on DROP (accuracy)
92.28DROP AccuracyFLOWBOT
Evaluation Results
| Method | Links | |
|---|---|---|
| FLOWBOTOptimizer LLM=GPT-4.1 mini, Executor LLM=GPT-4.1 mini2026.04 | 92.28 | |
| Direct Fine-tuningrank=1, additional_steps=02026.02 | 88.8 | |
| In-Squeezerank=128 to 1, schedule=Min steps2026.02 | 88.66 | |
| In-Squeezerank=128 to 1, schedule=Standard2026.02 | 88.25 | |
| Cont-Squeezerank=128 to 1, additional_steps=2002026.02 | 88.03 | |
| Direct Fine-tuningrank=1, additional_steps=2002026.02 | 87.97 | |
| Direct Fine-tuningrank=1, additional_steps=7002026.02 | 87.7 | |
| Cont-Squeezerank=128 to 1, additional_steps=7002026.02 | 87.67 | |
| Cont-Squeezerank=128 to 1, additional_steps=02026.02 | 86.41 | |
| OPERAdescription=Previous SOTA2022.10 | 84.9 | |
| LMSIPrompting Method=Self-Consistency2022.10 | 83 | |
| AFlowMethod Category=Workflow induction methods, Optimizer LLM=Claude-3.5 Sonnet, Executor LLM=GPT-4o mini2026.04 | 80.6 | |
| AFlow (our reimplementation)Optimizer LLM=GPT-4.1 mini, Executor LLM=GPT-4.1 mini2026.04 | 79.69 | |
| CoT + Self-ConsistencyMethod Category=Prompt-based methods, Executor LLM=GPT-4o mini2026.04 | 78.8 | |
| Chain-of-Thought (CoT)Method Category=Prompt-based methods, Executor LLM=GPT-4o mini2026.04 | 78.5 | |
| Direct Fine-tuningrank=1, training_steps=+0 steps2026.02 | 78.24 | |
| Self-ConsistencyLMSI=false2022.10 | 78.2 | |
| In-Squeezereduction=128 ... -> 1, strategy=Min steps2026.02 | 78.17 | |
| MedPromptMethod Category=Prompt-based methods, Executor LLM=GPT-4o mini2026.04 | 78 | |
| ADASMethod Category=Workflow induction methods, Optimizer LLM=Claude-3.5 Sonnet, Executor LLM=GPT-4o mini2026.04 | 76.6 | |
| In-Squeezereduction=128 ... -> 1, strategy=Standard2026.02 | 76.58 | |
| Cont-Squeezereduction=128 -> 1, training_steps=+700 steps2026.02 | 76.39 | |
| LMSIPrompting Method=CoT-Prompting2022.10 | 76.2 | |
| Direct Fine-tuningrank=1, training_steps=+200 steps2026.02 | 76.18 | |
| Cont-Squeezereduction=128 -> 1, training_steps=+200 steps2026.02 | 75.72 | |
| Direct Fine-tuningrank=1, training_steps=+700 steps2026.02 | 75.21 | |
| MultiPersonaMethod Category=Prompt-based methods, Executor LLM=GPT-4o mini2026.04 | 74.4 | |
| MEDALBackbone=LLaDA, Agentic Framework=ADAS2025.12 | 73 | |
| LMSIPrompting Method=Standard-Prompting2022.10 | 71.7 | |
| LLaDABackbone=LLaDA, Agentic Framework=ADAS2025.12 | 71.2 | |
| CoT-PromptingLMSI=false2022.10 | 70.6 | |
| Self RefineMethod Category=Prompt-based methods, Executor LLM=GPT-4o mini2026.04 | 70.2 | |
| Cont-Squeezereduction=128 -> 1, training_steps=+0 steps2026.02 | 69.83 | |
| Direct IO promptingMethod Category=Prompt-based methods, Executor LLM=GPT-4o mini2026.04 | 68.3 | |
| LlamaBackbone=Llama, Agentic Framework=ADAS2025.12 | 65.2 | |
| Standard-PromptingLMSI=false2022.10 | 60 | |
| LlamaBackbone=Llama, Agentic Framework=None2025.12 | 60 | |
| LLaDABackbone=LLaDA, Agentic Framework=None2025.12 | 58.2 | |
| Full-Attn# Shots=3-shot, Model Architecture=80B MoE (Hybrid 1:11), Attention Variant=Full-Attn2026.02 | 56.7 | |
| HySparse# Shots=3-shot, Model Architecture=80B MoE (Hybrid 1:11), Attention Variant=HySparse2026.02 | 56.5 | |
| Full-Attn# Shots=3-shot, Model Architecture=7B Dense (Hybrid 1:3), Attention Variant=Full-Attn2026.02 | 53.1 | |
| Zero-shotmode=0-S2026.02 | 52.6 | |
| HySparse# Shots=3-shot, Model Architecture=7B Dense (Hybrid 1:3), Attention Variant=HySparse2026.02 | 52.4 | |
| EnsembleAggregation Method=Byte-level Ensemble2025.06 | 47.9 | |
| Hybrid SWA# Shots=3-shot, Model Architecture=80B MoE (Hybrid 1:11), Attention Variant=Hybrid SWA2026.02 | 47.8 | |
| QWEN3Model=QWEN32025.06 | 47 | |
| FO (Adam)Base Model=LLaMA-3.2-1B, Optimization Regime=First-Order2025.10 | 45 | |
| Hybrid SWA# Shots=3-shot, Model Architecture=7B Dense (Hybrid 1:3), Attention Variant=Hybrid SWA2026.02 | 43.8 | |
| OLMO2Model=OLMO22025.06 | 40.9 | |
| AverageAggregation Method=Average2025.06 | 39.3 | |
| MERGEvolveSetting=single-task, Number of runs=52026.06 | 37 | |
| LoraHubSetting=single-task, Number of runs=52026.06 | 35.32 | |
| ZO Fine-tunerBase Model=LLaMA-3.2-1B, Optimization Regime=Zeroth-Order2025.10 | 32 | |
| Model SwarmsSetting=single-task, Number of runs=52026.06 | 31.7 | |
| Best Single ExpertSetting=single-task, Number of runs=52026.06 | 30.4 | |
| LLAMA3.2Model=LLAMA3.22025.06 | 29.9 | |
| MeZOBase Model=LLaMA-3.2-1B, Optimization Regime=Zeroth-Order2025.10 | 29 | |
| ShadowKV2025.12 | 28 | |
| Data MergeSetting=single-task, Number of runs=52026.06 | 27.4 | |
| EMMSetting=single-task, Number of runs=52026.06 | 25.9 | |
| Expert FusionSetting=single-task, Number of runs=52026.06 | 22.6 | |
| Pack of LLMsSetting=single-task, Number of runs=52026.06 | 21.56 | |
| CosineModel=1B, Pre-training Scheduler=Cosine, alpha_pre=0.1, alpha_mid=0.0, Pre-training tokens=2T, Mid-training tokens=500B, Training Stage=Supervised fine-tuned (SFT)2026.03 | 20.5 | |
| LinearModel=1B, Pre-training Scheduler=Linear, alpha_pre=0.1, alpha_mid=0.0, Pre-training tokens=2T, Mid-training tokens=500B, Training Stage=Supervised fine-tuned (SFT)2026.03 | 19.6 | |
| WSDModel=1B, Pre-training Scheduler=WSD, alpha_pre=0.1, alpha_mid=0.0, Pre-training tokens=2T, Mid-training tokens=500B, Training Stage=Supervised fine-tuned (SFT)2026.03 | 19.5 | |
| LinearModel=1B, Pre-training Scheduler=Linear, alpha_pre=0.1, alpha_mid=1.0, Pre-training tokens=2T, Mid-training tokens=500B, Training Stage=Supervised fine-tuned (SFT)2026.03 | 19.5 | |
| Warmup-Stable-Only (WSO)Model=1B, Pre-training Scheduler=Warmup-Stable-Only (WSO), alpha_pre=1.0, alpha_mid=1.0, Pre-training tokens=2T, Mid-training tokens=500B, Training Stage=Supervised fine-tuned (SFT)2026.03 | 19.4 | |
| TIESSetting=single-task, Number of runs=52026.06 | 18.8 | |
| CosineModel=1B, Pre-training Scheduler=Cosine, alpha_pre=0.1, alpha_mid=1.0, Pre-training tokens=2T, Mid-training tokens=500B, Training Stage=Supervised fine-tuned (SFT)2026.03 | 18.7 | |
| WSDModel=1B, Pre-training Scheduler=WSD, alpha_pre=1.0, alpha_mid=0.0, Pre-training tokens=2T, Mid-training tokens=500B, Training Stage=Supervised fine-tuned (SFT)2026.03 | 18.4 | |
| SnapKVBackbone=DeepSeek-R1-Distill-Llama-8B, Cache Budget=1282025.12 | 17 | |
| H2OKV cache budget=3842025.12 | 17 | |
| H2OKV cache budget=5122025.12 | 17 | |
| WSDModel=1B, Pre-training Scheduler=WSD, alpha_pre=0.1, alpha_mid=1.0, Pre-training tokens=2T, Mid-training tokens=500B, Training Stage=Supervised fine-tuned (SFT)2026.03 | 16.8 | |
| FullBackbone=Deepseek-R1-Distill-Qwen-7B2025.12 | 16 | |
| SnapKVBackbone=Deepseek-R1-Distill-Qwen-7B, Budget=5122025.12 | 16 | |
| StreamingLLMKV cache budget=5122025.12 | 16 | |
| Full2025.12 | 15 | |
| SnapKVKV cache budget=1282025.12 | 15 | |
| StreamingLLMKV cache budget=3842025.12 | 15 | |
| FullBackbone=DeepSeek-R1-Distill-Llama-8B2025.12 | 14 | |
| ShadowKVBackbone=Deepseek-R1-Distill-Qwen-7B2025.12 | 14 | |
| H2OKV cache budget=2562025.12 | 14 | |
| RKVKV cache budget=3842025.12 | 14 | |
| SnapKVBackbone=Deepseek-R1-Distill-Qwen-7B, Budget=1282025.12 | 13 | |
| StreamingLLMBackbone=Deepseek-R1-Distill-Qwen-7B, Budget=5122025.12 | 13 | |
| KnormKV cache budget=3842025.12 | 13 | |
| KnormKV cache budget=5122025.12 | 13 | |
| SnapKVBackbone=Deepseek-R1-Distill-Qwen-7B, Budget=3842025.12 | 12 | |
| H2OKV cache budget=1282025.12 | 12 | |
| SnapKVKV cache budget=2562025.12 | 12 | |
| SnapKVKV cache budget=5122025.12 | 12 | |
| SnapKVBackbone=Nemotron-Nano-8B, Budget=3842025.12 | 12 | |
| SnapKVBackbone=Deepseek-R1-Distill-Qwen-7B, Budget=2562025.12 | 11 | |
| RKVKV cache budget=5122025.12 | 11 | |
| SnapKVKV cache budget=3842025.12 | 11 | |
| StreamingLLMKV cache budget=2562025.12 | 11 | |
| FullBackbone=Nemotron-Nano-8B2025.12 | 11 | |
| ShadowKVBackbone=Nemotron-Nano-8B2025.12 | 11 | |
| SnapKVBackbone=Nemotron-Nano-8B, Budget=1282025.12 | 11 |