Reading Comprehension on C3
93.5AccuracyInternLM2-Chat-20B-SFT
Evaluation Results
| Method | Links | |
|---|---|---|
| InternLM2-Chat-20B-SFTParameter Scale=13~20B, Evaluation Setup=5-shot, Training Protocol=SFT2024.03 | 93.5 | |
| InternLM2-Chat-20BParameter Scale=13~20B, Evaluation Setup=5-shot, Training Protocol=RLHF2024.03 | 93.5 | |
| InternLM2-Chat-7B-SFTParameter Scale=< 7B, Evaluation Setup=5-shot, Training Protocol=SFT2024.03 | 91.7 | |
| InternLM2-Chat-7BParameter Scale=< 7B, Evaluation Setup=5-shot, Training Protocol=RLHF2024.03 | 91.5 | |
| Qwen-14B-ChatParameter Scale=13~20B, Evaluation Setup=5-shot2024.03 | 91.5 | |
| Qwen2.52025.12 | 84.7 | |
| Qwen-7B-ChatParameter Scale=< 7B, Evaluation Setup=5-shot2024.03 | 84.4 | |
| Baichuan2-13B-ChatParameter Scale=13~20B, Evaluation Setup=5-shot2024.03 | 84.4 | |
| Mixtral-8x7B-Instruct-v0.1Parameter Scale=13~20B, Evaluation Setup=5-shot2024.03 | 82.1 | |
| Qwen32025.12 | 81 | |
| Qwen-14BModel size=13 ~ 20B, shot=5-shot2024.03 | 80.6 | |
| ChatGLM3-6BParameter Scale=< 7B, Evaluation Setup=5-shot2024.03 | 79.3 | |
| InternLM2-20BModel size=13 ~ 20B, shot=5-shot2024.03 | 79.1 | |
| ChatGLM3-6B-BaseModel size=< 7B, shot=5-shot2024.03 | 78.9 | |
| Baichuan2-7B-ChatParameter Scale=< 7B, Evaluation Setup=5-shot2024.03 | 78.5 | |
| Qwen-1.5 14BModel Type=Teacher2024.07 | 77.38 | |
| Qwen-1.5 7BRole=Teacher2024.07 | 76.03 | |
| Gamayun2025.12 | 75.29 | |
| InternLM2-7BModel size=< 7B, shot=5-shot2024.03 | 74.1 | |
| Qwen-7BModel size=< 7B, shot=5-shot2024.03 | 71.4 | |
| KDTeacher Model=Qwen-1.5 14B, Student Model=Qwen-1.5 4B2024.07 | 70.31 | |
| DDKTeacher Model=Qwen-1.5 14B, Student Model=Qwen-1.5 4B2024.07 | 70.25 | |
| MiniLLMTeacher Model=Qwen-1.5 14B, Student Model=Qwen-1.5 4B2024.07 | 68.78 | |
| Llama3.22025.12 | 67.85 | |
| CPTTeacher Model=Qwen-1.5 14B, Student Model=Qwen-1.5 4B2024.07 | 67.72 | |
| InternLM2-20B-BaseModel size=13 ~ 20B, shot=5-shot2024.03 | 67.6 | |
| Mistral-7B-Instruct-v0.2Parameter Scale=< 7B, Evaluation Setup=5-shot2024.03 | 66.9 | |
| Baichuan2-13B-BaseModel size=13 ~ 20B, shot=5-shot2024.03 | 65.6 | |
| Qwen-1.5 4BModel Type=Student2024.07 | 65.26 | |
| Baichuan2-7B-BaseModel size=< 7B, shot=5-shot2024.03 | 64.6 | |
| NITPModel Scale=3B2026.05 | 63.67 | |
| Engram-27BShots=0-shot2026.01 | 63.6 | |
| DDKBase Model=Qwen-1.5 1.8B, Teacher=Qwen-1.5 7B2024.07 | 63.37 | |
| NITPModel scale=9bA1b, Evaluation=few-shot, Context length=81922026.05 | 63.01 | |
| InternLM2-7B-BaseModel size=< 7B, shot=5-shot2024.03 | 61.9 | |
| Engram-40BShots=0-shot2026.01 | 61.8 | |
| MiniLLMBase Model=Qwen-1.5 1.8B, Teacher=Qwen-1.5 7B2024.07 | 61.45 | |
| KDBase Model=Qwen-1.5 1.8B, Teacher=Qwen-1.5 7B2024.07 | 61.18 | |
| CPTBase Model=Qwen-1.5 1.8B, Teacher=Qwen-1.5 7B2024.07 | 60.3 | |
| MoE-27BShots=0-shot2026.01 | 60.1 | |
| Mixtral-8x7B-v0.1Model size=13 ~ 20B, shot=5-shot2024.03 | 59.1 | |
| NTPModel Scale=3B2026.05 | 59.01 | |
| Qwen-1.5 1.8BRole=Student2024.07 | 58.27 | |
| Dense-4BShots=0-shot2026.01 | 57.7 | |
| Gemma32025.12 | 57.31 | |
| Llama2-13B-ChatParameter Scale=13~20B, Evaluation Setup=5-shot2024.03 | 56.9 | |
| NTPModel scale=9bA1b, Evaluation=few-shot, Context length=81922026.05 | 56.65 | |
| Mistral-7B-v0.1Model size=< 7B, shot=5-shot2024.03 | 54.6 | |
| NITPModel Scale=2B2026.05 | 53.42 | |
| POEMParadigm=TTA2026.04 | 53.41 | |
| GPT-3.5Parameter Scale=API, Evaluation Setup=5-shot2024.03 | 52.5 | |
| Llama2-7B-ChatParameter Scale=< 7B, Evaluation Setup=5-shot2024.03 | 51.7 | |
| T^2ARD+MPTParadigm=MOA2026.04 | 51.2 | |
| ATTEMPTParadigm=MTL2026.04 | 51.12 | |
| SEEDBackbone=LLaMA2-7B, Selection Ratio=5%2026.05 | 51.1 | |
| SyCoParadigm=MOA2026.04 | 51 | |
| FOAParadigm=TTA2026.04 | 50.87 | |
| MPTParadigm=MTL2026.04 | 50.64 | |
| POEM+MPTParadigm=MOA2026.04 | 50.52 | |
| SPoTParadigm=MTL2026.04 | 50.3 | |
| ClusT3+ATTEMPTParadigm=MOA2026.04 | 50.18 | |
| ClusT3Paradigm=TTA2026.04 | 49.32 | |
| PTParadigm=Vanilla2026.04 | 49.27 | |
| NTPModel Scale=2B2026.05 | 49.26 | |
| TENT+SPoTParadigm=MOA2026.04 | 49.25 | |
| T^2ARDParadigm=TTA2026.04 | 49 | |
| TENTParadigm=TTA2026.04 | 48.18 | |
| FULLBackbone=LLaMA2-7B, Selection Ratio=100%2026.05 | 48.1 | |
| Llama2-13BModel size=13 ~ 20B, shot=5-shot2024.03 | 47.5 | |
| T^2ARD+TA-LoRAParadigm=MOA2026.04 | 47.38 | |
| POEM+TA-LoRAParadigm=MOA2026.04 | 46.58 | |
| RANDOMBackbone=LLaMA2-7B, Selection Ratio=5%2026.05 | 45.3 | |
| NITPModel scale=3bA0.5b, Evaluation=few-shot, Context length=81922026.05 | 44.38 | |
| Llama2-7BModel size=< 7B, shot=5-shot2024.03 | 43.8 | |
| Llama 2 7B (baseline)Base Model=Llama 2 7B, Compression Ratio=0%2025.05 | 43.8 | |
| LLM-StreamlineTrain-Free=false, Base Model=Llama 2 7B, Compression Ratio=25%2025.05 | 43.3 | |
| ReplaceMe (Cosine)Train-Free=true, Base Model=Llama 2 7B, Compression Ratio=25%2025.05 | 42.5 | |
| TA-LoRAParadigm=MTL2026.04 | 42.27 | |
| BASEBackbone=LLaMA2-7B, Selection Ratio=0%2026.05 | 42.1 | |
| UIDLTrain-Free=false, Base Model=Llama 2 7B, Compression Ratio=25%2025.05 | 40.2 | |
| LaCoTrain-Free=false, Base Model=Llama 2 7B, Compression Ratio=25%2025.05 | 39.7 | |
| ReplaceMe (LS)Train-Free=true, Base Model=Llama 2 7B, Compression Ratio=25%2025.05 | 39.4 | |
| NTPModel scale=3bA0.5b, Evaluation=few-shot, Context length=81922026.05 | 39.06 | |
| NTPModel scale=1.9bA0.3b, Evaluation=few-shot, Context length=81922026.05 | 32.21 | |
| NITPModel Scale=0.5B2026.05 | 32.1 | |
| SliceGPTTrain-Free=false, Base Model=Llama 2 7B, Compression Ratio=25%2025.05 | 31.5 | |
| LLMPrunerTrain-Free=false, Base Model=Llama 2 7B, Compression Ratio=25%2025.05 | 29.7 | |
| NITPModel scale=1.9bA0.3b, Evaluation=few-shot, Context length=81922026.05 | 29.69 | |
| NTPModel Scale=0.5B2026.05 | 28.21 |