Science Question Answering on ARC Challenge (test)
94.3Accuracylarge teacher
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| large teacherModel=Qwen2.5-14B-Instruct2026.05 | 94.3 | — | — | — | — | |
| COPRO-MLMforward=Gemini 2.5 flash-lite, LMbackward=Gemini 2.5 flash, Optimization Granularity=Blockwise2025.05 | 94.02 | — | — | — | 0.56 | |
| TextGrad-MLMforward=Gemini 2.5 flash-lite, LMbackward=Gemini 2.5 flash, Optimization Granularity=Blockwise2025.05 | 93.91 | — | — | — | 0.55 | |
| COPRO-MLMforward=Gemini 2.5 flash-lite, LMbackward=Gemini 2.5 flash, Optimization Granularity=Promptwise2025.05 | 93.8 | — | — | — | 0.44 | |
| TextGrad-MLMforward=Gemini 2.5 flash-lite, LMbackward=Gemini 2.5 flash, Optimization Granularity=Promptwise2025.05 | 93.74 | — | — | — | 0.45 | |
| COPROLMforward=Gemini 2.5 flash-lite, LMbackward=Gemini 2.5 flash2025.05 | 93.53 | — | — | — | 0.19 | |
| small teacherModel=Qwen2.5-7B-Instruct2026.05 | 92.6 | — | — | — | — | |
| AdalFlow-MLMforward=Gemini 2.5 flash-lite, LMbackward=Gemini 2.5 flash, Optimization Granularity=Promptwise2025.05 | 92.53 | — | — | — | 0.005 | |
| GEPALMforward=Gemini 2.5 flash-lite, LMbackward=Gemini 2.5 flash2025.05 | 92.33 | — | — | — | 0.006 | |
| TextGradLMforward=Gemini 2.5 flash-lite, LMbackward=Gemini 2.5 flash, Validation Revert Strategy=w/2025.05 | 92.2 | — | — | — | 0 | |
| TextGradLMforward=Gemini 2.5 flash-lite, LMbackward=Gemini 2.5 flash, Validation Revert Strategy=w/o2025.05 | 91.56 | — | — | — | 0.68 | |
| AdalFlowLMforward=Gemini 2.5 flash-lite, LMbackward=Gemini 2.5 flash2025.05 | 91.1 | — | — | — | 0.01 | |
| Qwen 2.5 7B Instructk=32025.09 | 90.9 | 0 | 98.5 | 0.55 | — | |
| Qwen3-4B + SFT + WeMask(TF)Mask Rate=0.1, Training Protocol=TF, Base Model=Qwen3-4B2026.05 | 87.54 | — | — | — | — | |
| Qwen3-4B + SFT + WeMask(TF)Mask Rate=0.3, Training Protocol=TF, Base Model=Qwen3-4B2026.05 | 87.26 | — | — | — | — | |
| Qwen3-4B + WeMask(SFT)Mask Rate=0.5, Training Protocol=SFT, Base Model=Qwen3-4B2026.05 | 87.11 | — | — | — | — | |
| Qwen3-4B + WeMask(SFT)Mask Rate=0.7, Training Protocol=SFT, Base Model=Qwen3-4B2026.05 | 87 | — | — | — | — | |
| Qwen3-4B + WeMask(SFT)Mask Rate=0.3, Training Protocol=SFT, Base Model=Qwen3-4B2026.05 | 86.95 | — | — | — | — | |
| Qwen3-4B + SFT + WeMask(TF)Mask Rate=0.5, Training Protocol=TF, Base Model=Qwen3-4B2026.05 | 86.92 | — | — | — | — | |
| Qwen3-4B + SFTMask Rate=-, Training Protocol=SFT, Base Model=Qwen3-4B2026.05 | 86.69 | — | — | — | — | |
| Qwen3-4B + WeMask(SFT)Mask Rate=0.1, Training Protocol=SFT, Base Model=Qwen3-4B2026.05 | 86.69 | — | — | — | — | |
| Qwen3-4B + SFT + WeMask(TF)Mask Rate=0.7, Training Protocol=TF, Base Model=Qwen3-4B2026.05 | 86.25 | — | — | — | — | |
| Qwen3-4B + WeMask(SFT)Mask Rate=1.0, Training Protocol=SFT, Base Model=Qwen3-4B2026.05 | 85.07 | — | — | — | — | |
| Llama 3.1 8B Instructk=32025.09 | 85 | 0.2 | 98 | 0.43 | — | |
| Full RepetitionAvg KV Cache=471.1, Avg Prefill FLOPs=3.140 T, Backbone=Qwen 2.5-3B2026.07 | 82.8 | — | — | — | — | |
| Naïve SummaryAvg KV Cache=277.4, Avg Prefill FLOPs=3.676 T, Backbone=Qwen 2.5-3B, Repetition Type=Appending summary2026.07 | 81.7 | — | — | — | — | |
| LLMLinguaAvg KV Cache=277.4, Avg Prefill FLOPs=2.105 T, Backbone=Qwen 2.5-3B, Repetition Type=Appending summary2026.07 | 81.5 | — | — | — | — | |
| PARTREPAvg KV Cache=280.4, Avg Prefill FLOPs=2.481 T, Backbone=Qwen 2.5-3B, τ=0.152026.07 | 81.4 | — | — | — | — | |
| Echo EvictionAvg KV Cache=235.5, Avg Prefill FLOPs=3.152 T, Backbone=Qwen 2.5-3B, Repetition Type=Compressing full repetition2026.07 | 81.3 | — | — | — | — | |
| No RepetitionAvg KV Cache=235.5, Avg Prefill FLOPs=1.549 T, Backbone=Qwen 2.5-3B2026.07 | 79.8 | — | — | — | — | |
| H2O EvictionAvg KV Cache=270.8, Avg Prefill FLOPs=3.162 T, Backbone=Qwen 2.5-3B, Repetition Type=Compressing full repetition2026.07 | 75 | — | — | — | — | |
| Gemma-3-27B-Baseshot=3-shot, approach=likelihood-based, access_level=Open-weights, region=Non-European2026.02 | 70.5 | — | — | — | — | |
| Mistral-3.2-24B-Baseshot=3-shot, approach=likelihood-based, access_level=Open-weights, region=European2026.02 | 68.5 | — | — | — | — | |
| Llama-3.3-70B-Baseshot=3-shot, approach=likelihood-based, access_level=Open-weights, region=Non-European2026.02 | 68.4 | — | — | — | — | |
| Qwen-3-14B-Baseshot=3-shot, approach=likelihood-based, access_level=Open-weights, region=Non-European2026.02 | 68.3 | — | — | — | — | |
| Gemma-3-12B-Baseshot=3-shot, approach=likelihood-based, access_level=Open-weights, region=Non-European2026.02 | 68.2 | — | — | — | — | |
| OLMo-3-32B-Baseshot=3-shot, approach=likelihood-based, access_level=Fully-open, region=Non-European2026.02 | 67.9 | — | — | — | — | |
| Apertus-70B-Baseshot=3-shot, approach=likelihood-based, access_level=Fully-open, region=European2026.02 | 64 | — | — | — | — | |
| Apertus-8B-Baseshot=3-shot, approach=likelihood-based, access_level=Fully-open, region=European2026.02 | 63.1 | — | — | — | — | |
| EuroLLM-22B-Baseshot=3-shot, approach=likelihood-based, access_level=Fully-open, region=European2026.02 | 62.3 | — | — | — | — | |
| OLMo-3-7B-Baseshot=3-shot, approach=likelihood-based, access_level=Fully-open, region=Non-European2026.02 | 61.8 | — | — | — | — | |
| Qwen-3-30B-A3B-Baseshot=3-shot, approach=likelihood-based, access_level=Open-weights, region=Non-European2026.02 | 60.6 | — | — | — | — | |
| EuroLLM-9B-Baseshot=3-shot, approach=likelihood-based, access_level=Fully-open, region=European2026.02 | 58.8 | — | — | — | — | |
| Llama-3.1-8B-Baseshot=3-shot, approach=likelihood-based, access_level=Open-weights, region=Non-European2026.02 | 58.3 | — | — | — | — | |
| CLPDOrder=student-loss, Student Model=Qwen-0.5B2026.05 | 48.8 | — | — | — | — | |
| Curriculum learning onlyTeacher=small teacher, Order=student-loss, Student Model=Qwen-0.5B2026.05 | 46.9 | — | — | — | — | |
| Curriculum learning onlyTeacher=large teacher, Order=student-loss, Student Model=Qwen-0.5B2026.05 | 46.5 | — | — | — | — | |
| Progressive distillation onlyTeacher=small & large teacher, Student Model=Qwen-0.5B2026.05 | 46.3 | — | — | — | — | |
| Standard distillationTeacher=large teacher, Student Model=Qwen-0.5B2026.05 | 45.5 | — | — | — | — | |
| Standard distillationTeacher=small teacher, Student Model=Qwen-0.5B2026.05 | 45.4 | — | — | — | — | |
| CLPDOrder=student-loss, Student Model=Llama-1B2026.05 | 37.4 | — | — | — | — | |
| Progressive distillation onlyTeacher=small & large teacher, Student Model=Llama-1B2026.05 | 36.2 | — | — | — | — | |
| Curriculum learning onlyTeacher=large teacher, Order=student-loss, Student Model=Llama-1B2026.05 | 35.6 | — | — | — | — | |
| Curriculum learning onlyTeacher=small teacher, Order=student-loss, Student Model=Llama-1B2026.05 | 35.5 | — | — | — | — | |
| Standard distillationTeacher=small teacher, Student Model=Llama-1B2026.05 | 34.7 | — | — | — | — | |
| Standard distillationTeacher=large teacher, Student Model=Llama-1B2026.05 | 34.5 | — | — | — | — | |
| Qwen3-4B + SFT + WeMask(TF)Mask Rate=1.0, Training Protocol=TF, Base Model=Qwen3-4B2026.05 | 29.66 | — | — | — | — | |
| studentStudent Model=Qwen-0.5B2026.05 | 17.1 | — | — | — | — | |
| LLMLingua Comp.Avg KV Cache=78.6, Avg Prefill FLOPs=1.130 T, Backbone=Qwen 2.5-3B, Repetition Type=Compressing full repetition2026.07 | 15.2 | — | — | — | — | |
| studentStudent Model=Llama-1B2026.05 | 7.7 | — | — | — | — |