Commonsense Reasoning on CommonSenseQA
91.2AccuracyFine-tuning
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Fine-tuning2023.06 | 91.2 | — | — | |
| High DiversityModel=Qwen-2.5-7B-Instruct2026.01 | 83.6 | — | — | |
| HD + ConfidenceModel=Qwen-2.5-7B-Instruct2026.01 | 82.8 | — | — | |
| SOTA (Fine-tuned)Mode=Fine-tuned2022.06 | 82.5 | — | — | |
| Self-ConsistencyModel=PaLM 540B2022.06 | 81.6 | — | — | |
| ConfidenceModel=Qwen-2.5-7B-Instruct2026.01 | 81 | — | — | |
| Majority VoteModel=Qwen-2.5-7B-Instruct2026.01 | 80.8 | — | — | |
| DIVERSEModel=code-davinci-0022022.06 | 79.9 | — | — | |
| Few-Shot SPBase Model=text-davinci-0022023.06 | 79.5 | — | — | |
| CoK + SC + F2-VBase Model=gpt-3.5-turbo2023.06 | 79.3 | — | — | |
| Qwen-1.5 7BRole=Teacher2024.07 | 79.28 | — | — | |
| DIVERSEModel=text-davinci-0022022.06 | 79.2 | — | — | |
| Greedy DecodeModel=PaLM 540B2022.06 | 79 | — | — | |
| CoK + SCBase Model=gpt-3.5-turbo2023.06 | 78.9 | — | — | |
| Single ModelModel=Qwen-2.5-7B-Instruct2026.01 | 78.4 | — | — | |
| Manual CoT + SCBase Model=gpt-3.5-turbo2023.06 | 78.2 | — | — | |
| CoK + F2-VBase Model=gpt-3.5-turbo2023.06 | 77.8 | — | — | |
| HD + Learn2AggModel=Qwen-2.5-7B-Instruct2026.01 | 77.6 | — | — | |
| Self-ConsistencyModel=code-davinci-0022022.06 | 77.3 | — | — | |
| CoK + F2-VBase Model=text-davinci-0022023.06 | 77.3 | — | — | |
| CoKBase Model=gpt-3.5-turbo2023.06 | 77.1 | — | — | |
| Manual CoTBase Model=gpt-3.5-turbo2023.06 | 76.5 | — | — | |
| ComplexCoT + SCBase Model=gpt-3.5-turbo2023.06 | 76 | — | — | |
| OK-TransformerLM=RoBERTa, Avg.=84.75, Speed-up=1.0x2023.05 | 75.92 | — | — | |
| FKL*Vocab size=32k, Number of shots=4-shot, Combined with SFT=true2025.12 | 75.7 | — | — | |
| Batch PartitioningLM=RoBERTa, Avg.=85.14, Speed-up=1.4x2023.05 | 75.59 | — | — | |
| FKL*Vocab size=64k, Number of shots=4-shot, Combined with SFT=true2025.12 | 75.4 | — | — | |
| CoKBase Model=text-davinci-0022023.06 | 75.4 | — | — | |
| ComplexCoTBase Model=gpt-3.5-turbo2023.06 | 75.4 | — | — | |
| ALMVocab size=64k, Number of shots=4-shot2025.12 | 75.3 | — | — | |
| ULDVocab size=64k, Number of shots=4-shot2025.12 | 75.3 | — | — | |
| HD + ConfidenceModel=Llama-3.1-8B-Instruct2026.01 | 75.2 | — | — | |
| FKLVocab size=64k, Number of shots=4-shot2025.12 | 75.2 | — | — | |
| DSKD*Vocab size=64k, Number of shots=4-shot, Combined with SFT=true2025.12 | 75.2 | — | — | |
| SFTVocab size=64k, Number of shots=4-shot2025.12 | 75.1 | — | — | |
| ULD*Vocab size=64k, Number of shots=4-shot, Combined with SFT=true2025.12 | 75.1 | — | — | |
| ALMVocab size=32k, Number of shots=4-shot2025.12 | 75.1 | — | — | |
| Frozen knowledgeLM=RoBERTa, Avg.=80.01, Speed-up=1.5x2023.05 | 75.02 | — | — | |
| DIVERSEModel=GPT-3 davinci (175B)2022.06 | 75 | — | — | |
| FKLVocab size=32k, Number of shots=4-shot2025.12 | 75 | — | — | |
| ALM*Vocab size=32k, Number of shots=4-shot, Combined with SFT=true2025.12 | 75 | — | — | |
| ALM*Vocab size=64k, Number of shots=4-shot, Combined with SFT=true2025.12 | 74.9 | — | — | |
| HD + Learn2AggModel=Llama-3.1-8B-Instruct2026.01 | 74.7 | — | — | |
| DSKDVocab size=64k, Number of shots=4-shot2025.12 | 74.7 | — | — | |
| SFTVocab size=16k, Number of shots=4-shot2025.12 | 74.7 | — | — | |
| SFTVocab size=32k, Number of shots=4-shot2025.12 | 74.5 | — | — | |
| ALMVocab size=16k, Number of shots=4-shot2025.12 | 74.4 | — | — | |
| Auto-CoTBase Model=text-davinci-0022023.06 | 74.4 | — | — | |
| FKLVocab size=16k, Number of shots=4-shot2025.12 | 74.3 | — | — | |
| ALM*Vocab size=16k, Number of shots=4-shot, Combined with SFT=true2025.12 | 74.3 | — | — | |
| DSKD*Vocab size=32k, Number of shots=4-shot, Combined with SFT=true2025.12 | 73.8 | — | — | |
| FKL*Vocab size=16k, Number of shots=4-shot, Combined with SFT=true2025.12 | 73.8 | — | — | |
| RoBERTaLM=RoBERTa, Avg.=83.952023.05 | 73.55 | — | — | |
| Manual CoTBase Model=text-davinci-0022023.06 | 73.5 | — | — | |
| Greedy DecodeModel=code-davinci-0022022.06 | 73.4 | — | — | |
| ConfidenceModel=Llama-3.1-8B-Instruct2026.01 | 73.4 | — | — | |
| ULD*Vocab size=32k, Number of shots=4-shot, Combined with SFT=true2025.12 | 73.4 | — | — | |
| Self-ConsistencyModel=text-davinci-0022022.06 | 73.3 | — | — | |
| Greedy DecodeModel=text-davinci-0022022.06 | 72.9 | — | — | |
| DSKDVocab size=32k, Number of shots=4-shot2025.12 | 72.9 | — | — | |
| DSKD*Vocab size=16k, Number of shots=4-shot, Combined with SFT=true2025.12 | 72.8 | — | — | |
| Debate 5 × 5Model=Qwen-2.5-7B-Instruct2026.01 | 72.7 | — | — | |
| ULDVocab size=32k, Number of shots=4-shot2025.12 | 72.4 | — | — | |
| High DiversityModel=Llama-3.1-8B-Instruct2026.01 | 72.1 | — | — | |
| Majority VoteModel=Llama-3.1-8B-Instruct2026.01 | 71.8 | — | — | |
| DSKDVocab size=16k, Number of shots=4-shot2025.12 | 71 | — | — | |
| Single ModelModel=Llama-3.1-8B-Instruct2026.01 | 70.3 | — | — | |
| ULD*Vocab size=16k, Number of shots=4-shot, Combined with SFT=true2025.12 | 70.3 | — | — | |
| ULDVocab size=16k, Number of shots=4-shot2025.12 | 70.1 | — | — | |
| KALEBackbone=Orca2 7B, Category=Augmented-based2026.01 | 69.62 | — | — | |
| Debate 5 × 5Model=Llama-3.1-8B-Instruct2026.01 | 68.8 | — | — | |
| Zero-Shot SPBase Model=text-davinci-0022023.06 | 68.8 | — | — | |
| KALEBackbone=Gemma2 9B, Category=Augmented-based2026.01 | 68.63 | — | — | |
| Self-ConsistencyModel=LaMDA 137B2022.06 | 67.8 | — | — | |
| DDKBase Model=Qwen-1.5 1.8B, Teacher=Qwen-1.5 7B2024.07 | 66.83 | — | — | |
| MeanLearnBackbone=Orca2 7B, Category=SFT-based2026.01 | 66.5 | — | — | |
| KDBase Model=Qwen-1.5 1.8B, Teacher=Qwen-1.5 7B2024.07 | 65.27 | — | — | |
| MiniLLMBase Model=Qwen-1.5 1.8B, Teacher=Qwen-1.5 7B2024.07 | 65.11 | — | — | |
| GraphRAGBackbone=Gemma2 9B, Category=Retrieval-based2026.01 | 65 | — | — | |
| CPTBase Model=Qwen-1.5 1.8B, Teacher=Qwen-1.5 7B2024.07 | 64.78 | — | — | |
| Qwen-1.5 1.8BRole=Student2024.07 | 64.7 | — | — | |
| Zero-Shot CoTBase Model=text-davinci-0022023.06 | 64.6 | — | — | |
| KG-SFTBackbone=Gemma2 9B, Category=SFT-based2026.01 | 64.54 | — | — | |
| STaRBackbone=Orca2 7B, Category=Augmented-based2026.01 | 64.53 | — | — | |
| GPT3MixBackbone=Gemma2 9B, Category=Augmented-based2026.01 | 64.29 | — | — | |
| LoTVerifier Training Dataset=AQuA2025.03 | 64 | — | — | |
| TOGBackbone=Gemma2 9B, Category=Retrieval-based2026.01 | 63.06 | — | — | |
| MeanLearnBackbone=Gemma2 9B, Category=SFT-based2026.01 | 63.06 | — | — | |
| TOGBackbone=Orca2 7B, Category=Retrieval-based2026.01 | 62.24 | — | — | |
| GLM-130Bshots=1, parameters=130B2022.10 | 62.2 | — | — | |
| DMTBackbone=Gemma2 9B, Category=SFT-based2026.01 | 62 | — | — | |
| GLM-130Bshots=0, parameters=130B2022.10 | 61.6 | — | — | |
| CoTBackbone=Gemma2 9B, Category=Prompt-based2026.01 | 61.43 | — | — | |
| GPT-3 (Davinci)shots=12022.10 | 61.2 | — | — | |
| SDFTBackbone=Gemma2 9B, Category=SFT-based2026.01 | 60.85 | — | — | |
| STaRBackbone=Gemma2 9B, Category=Augmented-based2026.01 | 60.2 | — | — | |
| StructGPTBackbone=Gemma2 9B, Category=Retrieval-based2026.01 | 59.3 | — | — | |
| SFTBackbone=Gemma2 9B, Category=SFT-based2026.01 | 58.97 | — | — | |
| KALEBackbone=OLMOE 7B, Category=Augmented-based2026.01 | 58.25 | — | — | |
| Greedy DecodeModel=LaMDA 137B2022.06 | 57.9 | — | — |