Science Question Answering on ARC-c (test)
91.3AccuracyQwen3-8B + SFT
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Qwen3-8B + SFTBase Model=Qwen3-8B, Training=SFT2026.05 | 91.3 | — | |
| Qwen3-8B + SFT + WeMask(TF)Base Model=Qwen3-8B, Training=SFT, Variant=TF2026.05 | 91.3 | — | |
| Qwen3-8B + WeMask(SFT)Base Model=Qwen3-8B, Training=WeMask(SFT)2026.05 | 91.3 | — | |
| Trinity Large (MoE)Architecture=MoE, Size=Large2026.02 | 90 | — | |
| Qwen3 8BParameters=8B2026.02 | 88 | — | |
| Qwen3 4BParameters=4B2026.02 | 84 | — | |
| Qwen-1.5 14B (Teacher)Model Scale=14B2024.07 | 80.59 | — | |
| Llama3.1-8B-Instruct + WeMask(SFT)Base Model=Llama3.1-8B-Instruct, Training=WeMask(SFT)2026.05 | 79.69 | — | |
| Llama3.1-8B-Instruct + SFTBase Model=Llama3.1-8B-Instruct, Training=SFT2026.05 | 79.52 | — | |
| Llama3.1-8B-Instruct + SFT + WeMask(TF)Base Model=Llama3.1-8B-Instruct, Training=SFT, Variant=TF2026.05 | 79.35 | — | |
| Llama3.1-8B-InstructBase Model=Llama3.1-8B-Instruct2026.05 | 74.49 | — | |
| Datology 8BParameters=8B2026.02 | 72 | — | |
| Llama-3.1 8BParameters=8B2026.02 | 67 | — | |
| Granite-4.0 MicroSize=Micro2026.02 | 67 | — | |
| SmolLM3 3BParameters=3B2026.02 | 65 | — | |
| LFM2.5 1.2BParameters=1.2B2026.02 | 65 | — | |
| Datology 3BParameters=3B2026.02 | 63 | — | |
| FLAN-137Bsetting=zero-shot2026.03 | 61.7 | — | |
| Imagine-DeBERTa-v3-L (Retrieval)KB=Synthetic VQA+, inference_mode=Retrieval, setting=zero-shot2026.03 | 59.5 | — | |
| Imagine-DeBERTa-v3-LKB=Synthetic VQA+, setting=zero-shot2026.03 | 59.2 | — | |
| Llama-3.2 3BParameters=3B2026.02 | 58 | — | |
| Imagine-DeBERTa-v3-LKB=Synthetic VQA, setting=zero-shot2026.03 | 56.2 | — | |
| Qwen-1.5 1.8B + DDKDistillation=DDK2024.07 | 55.03 | — | |
| Qwen-1.5 1.8B + TEDDistillation=TED2024.07 | 55 | — | |
| Qwen-1.5 1.8B + KDDistillation=KD2024.07 | 54.56 | — | |
| Qwen-1.5 1.8B + MiniLLMDistillation=MiniLLM2024.07 | 53.92 | — | |
| GPT-3Model Size=175B, Evaluation Protocol=One-shot2021.12 | 53.2 | — | |
| CAR-DeBERTa-v3-LKB=AbsAT, setting=zero-shot2026.03 | 53.2 | — | |
| Qwen-1.5 1.8B + CPT & DoReMiDistillation=CPT & DoReMi2024.07 | 52.11 | — | |
| GLaMModel Size=64B/64E, Evaluation Protocol=Few-shot, Shots=32021.12 | 52 | — | |
| GPT-3Model Size=175B, Evaluation Protocol=Few-shot, Shots=502021.12 | 51.5 | — | |
| GPT-3Model Size=175B, Evaluation Protocol=Zero-shot2021.12 | 51.4 | — | |
| Qwen-1.5 1.8B + CPTDistillation=CPT2024.07 | 51.03 | — | |
| Qwen-1.5 1.8B (Student)Model Scale=1.8B2024.07 | 50.31 | — | |
| GLaMModel Size=64B/64E, Evaluation Protocol=One-shot2021.12 | 50.3 | — | |
| GLaMModel Size=64B/64E, Evaluation Protocol=Zero-shot2021.12 | 48 | — | |
| Imagine-RoBERTa-LKB=Synthetic VQA, setting=zero-shot2026.03 | 39.1 | — | |
| CAR-RoBERTa-LKB=AbsAT, setting=zero-shot2026.03 | 36.5 | — | |
| Z-LaVI (BART-L)setting=zero-shot2026.03 | 36.5 | — | |
| Imagine-GPT-2-LKB=Synthetic VQA, setting=zero-shot2026.03 | 35.1 | — | |
| GPT-J-6Bsetting=zero-shot2026.03 | 34.8 | — | |
| OPT-30Bsetting=zero-shot2026.03 | 34.8 | — | |
| Z-LaVI (OPT-30B)setting=zero-shot2026.03 | 34.1 | — | |
| Z-LaVI (RoBERTa-L)setting=zero-shot2026.03 | 33.4 | — | |
| Llama-3.2 1BParameters=1B2026.02 | 33 | — | |
| GPT-Neo-2.7Bsetting=zero-shot2026.03 | 31.8 | — | |
| SMLMKB=*, setting=zero-shot2026.03 | 28.4 | — | |
| Qwen3-8BBase Model=Qwen3-8B2026.05 | 22.78 | — |