Commonsense Reasoning on CommonsenseQA (test)
90AccuracyHuman-Rater
Evaluation Results
| Method | Links | |
|---|---|---|
| Human-RaterModel=Human, Prompting Strategy=Average2023.05 | 90 | |
| ART-maxModel=Qwen2.5-32B-Instruct, Evaluation Protocol=D_test2026.04 | 84.9 | |
| ART-meanModel=Qwen2.5-32B-Instruct, Evaluation Protocol=D_test2026.04 | 84.8 | |
| VanillaModel=Qwen2.5-32B-Instruct, Evaluation Protocol=D_test2026.04 | 84.6 | |
| ART-maxModel=Qwen2.5-14B-Instruct, Evaluation Protocol=D_test2026.04 | 82.4 | |
| ART-meanModel=Qwen2.5-14B-Instruct, Evaluation Protocol=D_test2026.04 | 82 | |
| VanillaModel=Qwen2.5-14B-Instruct, Evaluation Protocol=D_test2026.04 | 81.5 | |
| SOTAModel=N/A, Prompting Strategy=N/A2023.05 | 80 | |
| ART-maxModel=Qwen2.5-7B-Instruct, Evaluation Protocol=D_test2026.04 | 79.6 | |
| ART-meanModel=Qwen2.5-7B-Instruct, Evaluation Protocol=D_test2026.04 | 79.2 | |
| VanillaModel=Qwen2.5-7B-Instruct, Evaluation Protocol=D_test2026.04 | 78.1 | |
| AMPLIFYModel=GPT-3.5, Prompting Strategy=AMPLIFY2023.05 | 77.9 | |
| ART-maxModel=Llama3.1-8B-Instruct, Evaluation Protocol=D_test2026.04 | 76.4 | |
| ART-meanModel=Llama3.1-8B-Instruct, Evaluation Protocol=D_test2026.04 | 76 | |
| GPT-3.5Model=GPT-3.5, Prompting Strategy=AO2023.05 | 75.7 | |
| VanillaModel=Llama3.1-8B-Instruct, Evaluation Protocol=D_test2026.04 | 75.5 | |
| GPT-3.5Model=GPT-3.5, Prompting Strategy=CoT2023.05 | 75.2 | |
| ART-meanModel=Ministral-8B-Instruct-2410, Evaluation Protocol=D_test2026.04 | 74.3 | |
| ART-maxModel=Ministral-8B-Instruct-2410, Evaluation Protocol=D_test2026.04 | 73.6 | |
| VanillaModel=Qwen2-7B-Instruct, Evaluation Protocol=D_test2026.04 | 73.6 | |
| AMPLIFYModel=GPT-3, Prompting Strategy=AMPLIFY2023.05 | 73.5 | |
| ART-maxModel=Qwen2-7B-Instruct, Evaluation Protocol=D_test2026.04 | 73.5 | |
| ART-meanModel=Qwen2-7B-Instruct, Evaluation Protocol=D_test2026.04 | 73.3 | |
| GPT-3Model=GPT-3, Prompting Strategy=CoT2023.05 | 72.6 | |
| VanillaModel=Ministral-8B-Instruct-2410, Evaluation Protocol=D_test2026.04 | 71.4 | |
| GPT-3Model=GPT-3, Prompting Strategy=AO2023.05 | 69.3 | |
| PASERPruning=SparseGPT (50%)2025.02 | 67.63 | |
| PASERPruning=Wanda (2:4)2025.02 | 66.79 | |
| PASERPruning=SliceGPT (25%)2025.02 | 66.51 | |
| PASERPruning=LLM-Pruner (25%)2025.02 | 66.44 | |
| Full DataPruning=Wanda (2:4)2025.02 | 65.04 | |
| Full DataPruning=SparseGPT (50%)2025.02 | 64.32 | |
| NuggetsPruning=SparseGPT (50%)2025.02 | 63.17 | |
| Instruction MiningPruning=SparseGPT (50%)2025.02 | 63.04 | |
| NuggetsPruning=Wanda (2:4)2025.02 | 62.91 | |
| NuggetsPruning=LLM-Pruner (25%)2025.02 | 62.59 | |
| IFDPruning=SparseGPT (50%)2025.02 | 62.48 | |
| IFDPruning=Wanda (2:4)2025.02 | 62.47 | |
| IFDPruning=SliceGPT (25%)2025.02 | 62.28 | |
| NuggetsPruning=SliceGPT (25%)2025.02 | 62.23 | |
| RandomPruning=SparseGPT (50%)2025.02 | 62.21 | |
| Instruction MiningPruning=Wanda (2:4)2025.02 | 62.13 | |
| RandomPruning=Wanda (2:4)2025.02 | 61.71 | |
| Full DataPruning=SliceGPT (25%)2025.02 | 61.32 | |
| IFDPruning=LLM-Pruner (25%)2025.02 | 61.22 | |
| RandomPruning=SliceGPT (25%)2025.02 | 61.09 | |
| Instruction MiningPruning=SliceGPT (25%)2025.02 | 60.81 | |
| Instruction MiningPruning=LLM-Pruner (25%)2025.02 | 60.65 | |
| RandomPruning=LLM-Pruner (25%)2025.02 | 60.21 | |
| Full DataPruning=LLM-Pruner (25%)2025.02 | 57.5 | |
| w/o TrainingPruning=SparseGPT (50%)2025.02 | 53.97 | |
| w/o TrainingPruning=Wanda (2:4)2025.02 | 52.65 | |
| VanillaModel=Llama2-7B-Chat, Evaluation Protocol=D_test2026.04 | 50.8 | |
| w/o TrainingPruning=SliceGPT (25%)2025.02 | 49.6 | |
| ART-meanModel=Llama2-7B-Chat, Evaluation Protocol=D_test2026.04 | 49.6 | |
| ART-maxModel=Llama2-7B-Chat, Evaluation Protocol=D_test2026.04 | 49.5 | |
| w/o TrainingPruning=LLM-Pruner (25%)2025.02 | 48.47 | |
| Muon+CWDOptimizer=Muon+CWD, Model Parameters=1B, Training Tokens=100B, Codebase=OLMo2025.10 | 33 | |
| AdamW+CWDOptimizer=AdamW+CWD, Model Parameters=1B, Training Tokens=100B, Codebase=OLMo2025.10 | 31 | |
| MuonOptimizer=Muon, Model Parameters=1B, Training Tokens=100B, Codebase=OLMo2025.10 | 30 | |
| AdamWOptimizer=AdamW, Model Parameters=1B, Training Tokens=100B, Codebase=OLMo2025.10 | 29 | |
| RandomModel=N/A, Prompting Strategy=Random Choice2023.05 | 20 |