Question Answering on OpenBookQA
94.4AccuracyLMSI
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| LMSIPrompting Method=Self-Consistency2022.10 | 94.4 | — | — | |
| LMSIPrompting Method=CoT-Prompting2022.10 | 93 | — | — | |
| LMSIPrompting Method=Standard-Prompting2022.10 | 92 | — | — | |
| Wang et al. (2022a)description=Previous SOTA2022.10 | 91 | — | — | |
| Qwen3 14B BaseClassifier=Self-labeled2026.01 | 90.8 | — | — | |
| Qwen3 14B BaseClassifier=Majority-labeled2026.01 | 90.8 | — | — | |
| Qwen3 14BClassifier=Self-labeled2026.01 | 90.4 | — | — | |
| Self-ConsistencyLMSI=false2022.10 | 90 | — | — | |
| Qwen3 14BClassifier=Majority-labeled2026.01 | 90 | — | — | |
| Yi-34Bevaluation=5-shot ICL2024.09 | 89.8 | — | — | |
| Yi-34B + RTDevaluation=5-shot2024.09 | 88.8 | — | — | |
| Yi-34B + RTDevaluation=zero-shot2024.09 | 88.4 | — | — | |
| Self-rethinkingModel=PaLM2, Prompting Strategy=Self-rethinking, k=12024.03 | 87.71 | — | — | |
| Self-consistencyModel=PaLM2, Prompting Strategy=Self-consistency, inference times=32024.03 | 87.61 | — | — | |
| UNIFIEDQAModel Architecture=T5, Retrieval Context (IR)=Yes, Fine-tuning Protocol=Fine-tuned2020.05 | 87.2 | — | — | |
| MixDoRABase Model=LLaMA-3 8B, Trainable Parameters=3.0%2024.04 | 86.9 | — | — | |
| CoT-PromptingLMSI=false2022.10 | 86.4 | — | — | |
| UNIFIEDQAModel Architecture=T5, Retrieval Context (IR)=No, Fine-tuning Protocol=Fine-tuned2020.05 | 86 | — | — | |
| Qwen3 8BClassifier=Self-labeled2026.01 | 85.8 | — | — | |
| Qwen3 8BClassifier=Majority-labeled2026.01 | 85.6 | — | — | |
| LLaMA2-70B + RTDevaluation=5-shot2024.09 | 85.4 | — | — | |
| LoRABase Model=LLaMA-3 8B, Trainable Parameters=2.6%2024.04 | 85 | — | — | |
| MixLoRABase Model=LLaMA-3 8B, Trainable Parameters=3.0%2024.04 | 84.8 | — | — | |
| DoRABase Model=LLaMA-2 13B, Trainable Parameters=2.4%2024.04 | 84.5 | — | — | |
| Standard-PromptingLMSI=false2022.10 | 84.4 | — | — | |
| LLaMA2-70Bevaluation=5-shot ICL2024.09 | 84.4 | — | — | |
| Qwen3 8B BaseClassifier=Majority-labeled2026.01 | 84.4 | — | — | |
| T5Model Architecture=T5, Retrieval Context (IR)=No, Fine-tuning Protocol=Fine-tuned2020.05 | 84.2 | — | — | |
| T5Model Architecture=T5, Retrieval Context (IR)=Yes, Fine-tuning Protocol=Fine-tuned2020.05 | 84.2 | — | — | |
| Qwen3 8B BaseClassifier=Self-labeled2026.01 | 84.2 | — | — | |
| Yi-34Bevaluation=zero-shot2024.09 | 83.5 | — | — | |
| MixDoRABase Model=LLaMA-2 13B, Trainable Parameters=2.5%2024.04 | 83.4 | — | — | |
| LoRABase Model=LLaMA-2 13B, Trainable Parameters=2.4%2024.04 | 83.2 | — | — | |
| MixLoRABase Model=LLaMA-2 13B, Trainable Parameters=2.5%2024.04 | 83 | — | — | |
| CoTModel=PaLM2, Prompting Strategy=Chain-of-Thought2024.03 | 82.66 | — | — | |
| MixLoRABase Model=LLaMA-2 7B, Trainable Parameters=2.9%2024.04 | 81.6 | — | — | |
| Qwen3 4B BaseClassifier=Majority-labeled2026.01 | 81.6 | — | — | |
| Qwen3 4B BaseClassifier=Self-labeled2026.01 | 81.4 | — | — | |
| Standard PromptingModel=PaLM2, Prompting Strategy=Standard Prompting2024.03 | 80.92 | — | — | |
| MixDoRABase Model=LLaMA-2 7B, Trainable Parameters=2.9%2024.04 | 80.9 | — | — | |
| DoRABase Model=LLaMA-2 7B, Trainable Parameters=2.9%2024.04 | 80.6 | — | — | |
| DoRABase Model=LLaMA-3 8B, Trainable Parameters=2.6%2024.04 | 80.6 | — | — | |
| LoRABase Model=LLaMA-2 7B, Trainable Parameters=2.9%2024.04 | 80.4 | — | — | |
| KF+SIRModel Architecture=RoBERTa, Retrieval Context (IR)=Yes, Fine-tuning Protocol=standard2020.05 | 80 | — | — | |
| DrICLk=50, Backbone=Llama-2-7b-chat-hf2025.01 | 80 | — | — | |
| DrICLk=MAX, Backbone=Llama-2-7b-chat-hf2025.01 | 80 | — | — | |
| OPT-IML 175BEvaluation protocol=0-shot2022.12 | 79.9 | — | — | |
| TS (C)Algorithm Family=Contextual Linear2026.02 | 79.3 | — | — | |
| MetaICLk=60, Backbone=Llama-2-7b-chat-hf2025.01 | 79 | — | — | |
| MetaICLk=MAX, Backbone=Llama-2-7b-chat-hf2025.01 | 79 | — | — | |
| DrICLNumber of shots (k)=102025.01 | 79 | — | — | |
| DrICLNumber of shots (k)=302025.01 | 79 | — | — | |
| DrICLNumber of shots (k)=402025.01 | 79 | — | — | |
| DrICLNumber of shots (k)=502025.01 | 79 | — | — | |
| LLaMA3-8B + RTDevaluation=5-shot2024.09 | 78.6 | — | — | |
| Qwen3 4BClassifier=Self-labeled2026.01 | 78.4 | — | — | |
| MetaICLk=40, Backbone=Llama-2-7b-chat-hf2025.01 | 78 | — | — | |
| MetaICLk=50, Backbone=Llama-2-7b-chat-hf2025.01 | 78 | — | — | |
| DrICLk=10, Backbone=Llama-2-7b-chat-hf2025.01 | 78 | — | — | |
| DrICLNumber of shots (k)=202025.01 | 78 | — | — | |
| Qwen3 4BClassifier=Majority-labeled2026.01 | 78 | — | — | |
| LLaMA3-8Bevaluation=5-shot ICL2024.09 | 77.5 | — | — | |
| FLAN 137BEvaluation protocol=0-shot, Training strategy=leave-one-category-out2022.12 | 77.4 | — | — | |
| FLAN 137BEvaluation protocol=few-shot, k=variable, Training strategy=leave-one-category-out2022.12 | 77.2 | — | — | |
| MetaICLk=30, Backbone=Llama-2-7b-chat-hf2025.01 | 77 | — | — | |
| DrICLk=3, Backbone=Llama-2-7b-chat-hf2025.01 | 77 | — | — | |
| DrICLk=5, Backbone=Llama-2-7b-chat-hf2025.01 | 77 | — | — | |
| DrICLNumber of shots (k)=32025.01 | 77 | — | — | |
| DrICLNumber of shots (k)=52025.01 | 77 | — | — | |
| DrICLNumber of shots (k)=602025.01 | 77 | — | — | |
| DrICLNumber of shots (k)=702025.01 | 77 | — | — | |
| OPT-IML 30BEvaluation protocol=0-shot2022.12 | 76.7 | — | — | |
| OPT-IML 175BEvaluation protocol=few-shot, k=52022.12 | 76.5 | — | — | |
| FalconModel Size=180B2023.11 | 76.4 | — | — | |
| GFLOWPOPrompt LM=Gemma-7B, Target LM=Gemma-7B2026.02 | 76.2 | — | — | |
| DrICLk=20, Backbone=Llama-2-7b-chat-hf2025.01 | 76 | — | — | |
| DrICLk=30, Backbone=Llama-2-7b-chat-hf2025.01 | 76 | — | — | |
| DrICLk=40, Backbone=Llama-2-7b-chat-hf2025.01 | 76 | — | — | |
| DrICLk=60, Backbone=Llama-2-7b-chat-hf2025.01 | 76 | — | — | |
| DrICLk=70, Backbone=Llama-2-7b-chat-hf2025.01 | 76 | — | — | |
| DrICLk=AVG, Backbone=Llama-2-7b-chat-hf2025.01 | 76 | — | — | |
| MetaICLNumber of shots (k)=202025.01 | 76 | — | — | |
| MetaICLNumber of shots (k)=402025.01 | 76 | — | — | |
| MetaICLNumber of shots (k)=502025.01 | 76 | — | — | |
| MetaICLNumber of shots (k)=602025.01 | 76 | — | — | |
| BaselineFormat=Baseline, Bit width (b)=16.00, Quantization Aware Training (QAT)=true, Backbone=Llama 3.1 8B2025.05 | 76 | — | — | |
| MixLoRABase Model=Gemma 2B, Trainable Parameters=4.3%2024.04 | 75.8 | — | — | |
| FreeLB-RoBERTaModel Architecture=RoBERTa, Retrieval Context (IR)=No, Fine-tuning Protocol=standard2020.05 | 75.7 | — | — | |
| MixDoRABase Model=Gemma 2B, Trainable Parameters=4.3%2024.04 | 75.4 | — | — | |
| MetaICLk=20, Backbone=Llama-2-7b-chat-hf2025.01 | 75 | — | — | |
| MetaICLNumber of shots (k)=102025.01 | 75 | — | — | |
| MetaICLNumber of shots (k)=702025.01 | 74 | — | — | |
| DrICLNumber of shots (k)=02025.01 | 74 | — | — | |
| DrICLNumber of shots (k)=12025.01 | 74 | — | — | |
| No-RewriteAlgorithm Family=Base2026.02 | 73.5 | — | — | |
| MetaICLNumber of shots (k)=52025.01 | 73 | — | — | |
| Tensor RMS + CFormat=Tensor RMS + C, Bit width (b)=3.00, Quantization Aware Training (QAT)=true, Backbone=Llama 3.1 8B2025.05 | 72.8 | — | — | |
| FLIPPEDModel Size=11B, Evaluation Protocol=Zero-shot2022.10 | 72.54 | — | — | |
| STABLEPROMPTPrompt LM=Gemma-7B, Target LM=Gemma-7B2026.02 | 72.2 | — | — | |
| Qwen3 1.7B BaseClassifier=Self-labeled2026.01 | 72.2 | — | — |