Multiple Choice Question Answering on SciQ
100AccuracyQA
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| QACoverage=20%2022.09 | 100 | — | |
| QACoverage=50%2022.09 | 100 | — | |
| QA+E+Cmodel=RoBERTa-RACE, entailment=true, contradiction=true2022.09 | 98.21 | — | |
| SciQ-smallvariant=small2022.09 | 98.09 | — | |
| QA+Emodel=RoBERTa-RACE, entailment=true2022.09 | 98.09 | — | |
| QA+Cmodel=RoBERTa-RACE, contradiction=true2022.09 | 98.09 | — | |
| QAmodel=RoBERTa-RACE2022.09 | 97.96 | — | |
| MonoSoupSetting=M-1 (Linear)2026.02 | 95.3 | — | |
| MonoSoupSetting=M-2 (Cosine)2026.02 | 95.3 | — | |
| SciQ-basevariant=base2022.09 | 95.28 | — | |
| LINESSetting=M-2 (Cosine)2026.02 | 95.2 | — | |
| ModelStock (M-2, M-3)Setting=Pairwise2026.02 | 95.2 | — | |
| MonoSoup R = 0.8Setting=M-1 (Linear)2026.02 | 95.1 | — | |
| MonoSoup R = 0.8Setting=M-2 (Cosine)2026.02 | 95.1 | — | |
| K-stage RKSolver choice=Runge–Kutta, K=22026.05 | 95 | — | |
| LINESSetting=M-1 (Linear)2026.02 | 94.9 | — | |
| ModelStock (M-1, M-2)Setting=Pairwise2026.02 | 94.9 | — | |
| Heun K=1Solver choice=Heun, K=12026.05 | 94.9 | — | |
| ModelStock (M-1, M-3)Setting=Pairwise2026.02 | 94.8 | — | |
| Baseline2026.05 | 94.8 | — | |
| StandardSetting=M-1 (Linear)2026.02 | 94.6 | — | |
| StandardSetting=M-2 (Cosine)2026.02 | 94.5 | — | |
| MonoSoupSetting=M-3 (Ext.)2026.02 | 94.5 | — | |
| MonoSoup R = 0.8Setting=M-3 (Ext.)2026.02 | 94.2 | — | |
| ScaleOTSetting=Plug-in, Backbone=OPT-1.3B2025.07 | 94 | — | |
| GradOTSetting=Plug-in, Backbone=OPT-1.3B2025.07 | 93.9 | — | |
| LINESSetting=M-3 (Ext.)2026.02 | 93.8 | — | |
| StandardSetting=M-3 (Ext.)2026.02 | 93.5 | — | |
| CRaShSetting=Plug-in, Backbone=OPT-1.3B2025.07 | 93.1 | — | |
| OTSetting=Plug-in, Backbone=OPT-1.3B2025.07 | 92.9 | — | |
| QWEN3-0.6B-BASESetting=Reference2026.02 | 92.6 | — | |
| Full LLMSetting=Fine-tuning (FT), Backbone=OPT-1.3B2025.07 | 92.5 | — | |
| OTSetting=Emulator Fine-tuning (Emu. FT), Backbone=OPT-1.3B2025.07 | 92.2 | — | |
| OT†Setting=Plug-in, Backbone=OPT-1.3B2025.07 | 90.8 | — | |
| ScaleOTSetting=Emulator Fine-tuning (Emu. FT), Backbone=OPT-1.3B2025.07 | 90 | — | |
| FairSeqNumber of Parameters=13B, Evaluation Protocol=5-shot2022.04 | 89.9 | — | |
| OT†Setting=Emulator Fine-tuning (Emu. FT), Backbone=OPT-1.3B2025.07 | 89.4 | — | |
| CRaShSetting=Emulator Fine-tuning (Emu. FT), Backbone=OPT-1.3B2025.07 | 88.9 | — | |
| GradOTSetting=Emulator Fine-tuning (Emu. FT), Backbone=OPT-1.3B2025.07 | 88.2 | — | |
| FairSeqNumber of Parameters=2.7B, Evaluation Protocol=5-shot2022.04 | 87.5 | — | |
| NextLat (d=2)Horizon (d)=2, Parameters=1.3B, Training tokens=100B2025.11 | 87.5 | — | |
| JTP (d=1)Horizon (d)=1, Parameters=1.3B, Training tokens=100B2025.11 | 87.3 | — | |
| FairSeqNumber of Parameters=6.7B, Evaluation Protocol=5-shot2022.04 | 87.1 | — | |
| JTP (d=2)Horizon (d)=2, Parameters=1.3B, Training tokens=100B2025.11 | 86.7 | — | |
| MTP (d=2)Horizon (d)=2, Parameters=1.3B, Training tokens=100B2025.11 | 86.6 | — | |
| GPTParameters=1.3B, Training tokens=100B2025.11 | 86.1 | — | |
| NextLat (d=1)Horizon (d)=1, Parameters=1.3B, Training tokens=100B2025.11 | 86 | — | |
| FairSeqNumber of Parameters=1.3B, Evaluation Protocol=5-shot2022.04 | 85.9 | — | |
| MTP (d=1)Horizon (d)=1, Parameters=1.3B, Training tokens=100B2025.11 | 85.4 | — | |
| Full LLMSetting=Zero-shot (ZS), Backbone=OPT-1.3B2025.07 | 84.4 | — | |
| CRaShSetting=Emulator Zero-shot (Emu. ZS), Backbone=OPT-1.3B2025.07 | 84.3 | — | |
| HourglassSize=1074M2026.02 | 82.5 | — | |
| FairSeqNumber of Parameters=355M, Evaluation Protocol=5-shot2022.04 | 81.9 | — | |
| Conventional (OLMo-2)Size=1074M2026.02 | 81 | — | |
| OTSetting=Emulator Zero-shot (Emu. ZS), Backbone=OPT-1.3B2025.07 | 80.9 | — | |
| ConventionalSize=1074M2026.02 | 80.6 | — | |
| HourglassSize=906M2026.02 | 79.8 | — | |
| ConventionalSize=906M2026.02 | 78.8 | — | |
| HourglassSize=403M2026.02 | 77.7 | — | |
| ConventionalSize=403M2026.02 | 76.8 | — | |
| FairSeqNumber of Parameters=125M, Evaluation Protocol=5-shot2022.04 | 75.8 | — | |
| Z-LaVI (OPT-30B)evaluation_protocol=zero-shot, # Param.=30B, variant=Z-LaVI2022.10 | 74 | — | |
| Z-LaVI (GPT-J-6B)evaluation_protocol=zero-shot, # Param.=6B, variant=Z-LaVI2022.10 | 73.7 | — | |
| GPT-J-6Bevaluation_protocol=zero-shot, # Param.=6B, variant=Original2022.10 | 73.2 | — | |
| OPT-30Bevaluation_protocol=zero-shot, # Param.=30B, variant=Original2022.10 | 72.7 | — | |
| HourglassSize=113M2026.02 | 69.6 | — | |
| Base modelBackbone=Pythia-1B, Checkpoint=step10000, Freeze Layers=4, Seed=42, Samples=5002026.03 | 68.4 | — | |
| ConventionalSize=113M2026.02 | 68.3 | — | |
| Weight averageBackbone=Pythia-1B, Checkpoint=step10000, Freeze Layers=4, Seed=42, Samples=5002026.03 | 67.8 | — | |
| MonolithicBackbone=Pythia-1B, Checkpoint=step10000, Freeze Layers=4, Seed=42, Samples=5002026.03 | 67 | — | |
| Fict. specialistBackbone=Pythia-1B, Checkpoint=step10000, Freeze Layers=4, Seed=42, Samples=5002026.03 | 66.8 | — | |
| Code specialistBackbone=Pythia-1B, Checkpoint=step10000, Freeze Layers=4, Seed=42, Samples=5002026.03 | 65.8 | — | |
| Sci. specialistBackbone=Pythia-1B, Checkpoint=step10000, Freeze Layers=4, Seed=42, Samples=5002026.03 | 65.8 | — | |
| Z-LaVI (GPT-Neo-2.7B)evaluation_protocol=zero-shot, # Param.=2.7B, variant=Z-LaVI2022.10 | 64.9 | — | |
| KALAVAI MoEBackbone=Pythia-1B, Checkpoint=step10000, Freeze Layers=4, Seed=42, Samples=5002026.03 | 64.8 | — | |
| GPT-Neo-2.7Bevaluation_protocol=zero-shot, # Param.=2.7B, variant=Original2022.10 | 64 | — | |
| Z-LaVI (GPT-Neo-1.3B)evaluation_protocol=zero-shot, # Param.=1.3B, variant=Z-LaVI2022.10 | 60.8 | — | |
| ScaleOTSetting=Emulator Zero-shot (Emu. ZS), Backbone=OPT-1.3B2025.07 | 59.9 | — | |
| Z-LaVI (SBERT)evaluation_protocol=zero-shot, # Param.=110M, labeled_pretraining=true, variant=Z-LaVI2022.10 | 58.5 | — | |
| SBERTevaluation_protocol=zero-shot, # Param.=110M, labeled_pretraining=true, variant=Original2022.10 | 57.7 | — | |
| GPT-Neo-1.3Bevaluation_protocol=zero-shot, # Param.=1.3B, variant=Original2022.10 | 57.5 | — | |
| Z-LaVI (ROBERTa-L-mnli)evaluation_protocol=zero-shot, # Param.=355 M, labeled_pretraining=true, variant=Z-LaVI2022.10 | 51.3 | — | |
| Z-LaVI (BART-L-mnli)evaluation_protocol=zero-shot, # Param.=400 M, labeled_pretraining=true, variant=Z-LaVI2022.10 | 51 | — | |
| OT†Setting=Emulator Zero-shot (Emu. ZS), Backbone=OPT-1.3B2025.07 | 49.8 | — | |
| Z-LaVI w/o LMevaluation_protocol=zero-shot, # Param.=150M2022.10 | 49.5 | — | |
| GradOTSetting=Emulator Zero-shot (Emu. ZS), Backbone=OPT-1.3B2025.07 | 49.4 | — | |
| BART-L-mnlievaluation_protocol=zero-shot, # Param.=400 M, labeled_pretraining=true, variant=Original2022.10 | 48.8 | — | |
| Z-LaVI (SimCSE)evaluation_protocol=zero-shot, # Param.=355M, variant=Z-LaVI2022.10 | 48.6 | — | |
| ROBERTa-L-mnlievaluation_protocol=zero-shot, # Param.=355 M, labeled_pretraining=true, variant=Original2022.10 | 44.7 | — | |
| SimCSEevaluation_protocol=zero-shot, # Param.=355M, variant=Original2022.10 | 42.6 | — | |
| Randomevaluation_protocol=zero-shot2022.10 | 25 | — |