Social Commonsense Reasoning on SIQA
86.9AccuracyHuman
Evaluation Results
| Method | Links | |
|---|---|---|
| HumanKB=-2026.03 | 86.9 | |
| In-Squeezereduction=128 ... -> 1, strategy=Min steps2026.02 | 85.24 | |
| In-Squeezereduction=128 ... -> 1, strategy=Standard2026.02 | 84.41 | |
| Direct Fine-tuningrank=1, training_steps=+700 steps2026.02 | 84.25 | |
| Cont-Squeezereduction=128 -> 1, training_steps=+700 steps2026.02 | 84.25 | |
| Direct Fine-tuningrank=1, training_steps=+0 steps2026.02 | 83.68 | |
| Cont-Squeezereduction=128 -> 1, training_steps=+200 steps2026.02 | 83.68 | |
| Direct Fine-tuningrank=1, training_steps=+200 steps2026.02 | 83.16 | |
| DeBERTa-v3-L (Supervised)KB=-, Evaluation Protocol=Supervised2026.03 | 80.1 | |
| RoBERTa-L (Supervised)KB=-, Evaluation Protocol=Supervised2026.03 | 76.6 | |
| Qwen3-4BParams=4B2025.12 | 75.59 | |
| S-MeZOBackbone=Mistral-7B, Evaluation Protocol=Fine-Tuning2024.02 | 70.2 | |
| ChatGPT (gpt-3.5-turbo)KB=-, Evaluation Protocol=Zero-shot2026.03 | 69.7 | |
| Qwen2.5-3BParams=3B2025.12 | 69.4 | |
| IMAGINE-DeBERTa-v3-LKB=Synthetic VQA+, Evaluation Protocol=Zero-shot2026.03 | 69 | |
| IMAGINE-DeBERTa-v3-L (Retrieval)KB=Synthetic VQA+, Inference Strategy=Retrieval, Evaluation Protocol=Zero-shot2026.03 | 69 | |
| Qwen3-1.7BParams=1.7B2025.12 | 68.58 | |
| MeZOBackbone=Mistral-7B, Evaluation Protocol=Fine-Tuning2024.02 | 68.5 | |
| GPT-3.5 (text-davinci-003)KB=-, Evaluation Protocol=Zero-shot2026.03 | 68 | |
| Zero-shot FusionKB=AT, CN, WD, WN, Evaluation Protocol=Zero-shot2026.03 | 66.6 | |
| IMAGINE-DeBERTa-v3-LKB=Synthetic VQA, Evaluation Protocol=Zero-shot2026.03 | 66.3 | |
| gemma2-2BParams=2B2025.12 | 65.92 | |
| CANDLE-DeBERTa-v3-LKB=CANDLE, Evaluation Protocol=Zero-shot2026.03 | 65.9 | |
| SmolLM3-3BParams=3B2025.12 | 65.25 | |
| Qwen2.5-1.5BParams=1.5B2025.12 | 64.94 | |
| llama-3.2-3BParams=3B2025.12 | 64.33 | |
| IMAGINE-RoBERTa-LKB=Synthetic VQA, Evaluation Protocol=Zero-shot2026.03 | 64.3 | |
| CAR-RoBERTa-LKB=AbsAT, Evaluation Protocol=Zero-shot2026.03 | 64 | |
| CAR-DeBERTa-v3-LKB=AbsAT, Evaluation Protocol=Zero-shot2026.03 | 64 | |
| Qwen2-1.5BParams=1.5B2025.12 | 63.46 | |
| YuLan-Mini-2.4BParams=2.4B2025.12 | 63.25 | |
| RoBERTa-L (MR)KB=AT, Evaluation Protocol=Zero-shot2026.03 | 63.1 | |
| PCMind-2.1-Kaiyuan-2BParams=2B2025.12 | 62.59 | |
| DeBERTa-v3-L (MR)KB=AT, Evaluation Protocol=Zero-shot2026.03 | 62.1 | |
| Qwen3-0.6BParams=0.6B2025.12 | 61.51 | |
| SmolLM2-1.7BParams=1.7B2025.12 | 60.18 | |
| CANDLE-VERA-T5-xxlKB=CANDLE, Evaluation Protocol=Zero-shot2026.03 | 59.4 | |
| VERA-T5-xxlKB=AT, Evaluation Protocol=Zero-shot2026.03 | 58.2 | |
| VERA-T5-xxlKB=AbsAT, Evaluation Protocol=Zero-shot2026.03 | 58.1 | |
| GPT-4 (gpt-4)KB=-, Evaluation Protocol=Zero-shot2026.03 | 57 | |
| MICOKB=AT, Evaluation Protocol=Zero-shot2026.03 | 56 | |
| OLMo-2-0425-1BParams=1B2025.12 | 55.53 | |
| GPT-2-L (MR)KB=AT, Evaluation Protocol=Zero-shot2026.03 | 53.6 | |
| IMAGINE-GPT-2-LKB=Synthetic VQA, Evaluation Protocol=Zero-shot2026.03 | 53 | |
| CAR-GPT-2-LKB=AbsAT, Evaluation Protocol=Zero-shot2026.03 | 52.3 | |
| llama-3.2-1BParams=1B2025.12 | 50.61 | |
| LLAMA2-13BKB=-, Evaluation Protocol=Zero-shot2026.03 | 50.3 | |
| COMET-DynGenKB=AT, Evaluation Protocol=Zero-shot2026.03 | 50.1 | |
| DeBERTa-v3-LKB=-, Evaluation Protocol=Zero-shot2026.03 | 47.8 | |
| RoBERTa-LKB=-, Evaluation Protocol=Zero-shot2026.03 | 47.3 | |
| AMALIA-9B-DPOModel Category=Fully open models, Training=DPO, Instruction-tuned=true2026.03 | 46.3 | |
| Self-talkKB=-, Evaluation Protocol=Zero-shot2026.03 | 46.2 | |
| MoHGE-14BTotal Parameters=14.122B, Activated Parameters of Experts=0.843B2026.04 | 45.62 | |
| Salamandra-7BModel Category=Fully open models, Instruction-tuned=true2026.03 | 44.8 | |
| Gervasio-8BModel Category=Open weight models, Instruction-tuned=true2026.03 | 44.8 | |
| AMALIA-9B-SFTModel Category=Fully open models, Training=SFT, Instruction-tuned=true2026.03 | 44.7 | |
| GPT-2-LKB=-, Evaluation Protocol=Zero-shot2026.03 | 44.6 | |
| MoE-14BTotal Parameters=16.760B, Activated Parameters of Experts=1.191B2026.04 | 44.28 | |
| Gemma 3-12BModel Category=Open weight models, Instruction-tuned=true2026.03 | 43.8 | |
| Llama 3.1-8BModel Category=Open weight models, Instruction-tuned=true2026.03 | 43.4 | |
| Apertus-8BModel Category=Fully open models, Instruction-tuned=true2026.03 | 43.2 | |
| PLDRv51-SOC-110M-5Zero-shot=true2026.03 | 43.09 | |
| Mistral-v0.1-7BKB=-, Evaluation Protocol=Zero-shot2026.03 | 42.9 | |
| FIXED_E=128Experts=1282026.05 | 42.63 | |
| EMO (Stage 5)E=64→1282026.05 | 42.48 | |
| Ministral-8BModel Category=Open weight models, Instruction-tuned=true2026.03 | 42.3 | |
| DenseTotal Parameters=1.672B2026.04 | 42.29 | |
| PLDRv51-SOC-110M-3Zero-shot=true2026.03 | 42.17 | |
| InstructBLIP-Vicuna-7BKB=-, Evaluation Protocol=Zero-shot2026.03 | 42.1 | |
| GPT-Neo-125MZero-shot=true2026.03 | 42.07 | |
| Gaussian-GaussianCPT=Gaussian, SFT=Gaussian2026.05 | 42.07 | |
| EMO (Stage 4)E=32→642026.05 | 42.07 | |
| PLDRv51-SOC-110M-2Zero-shot=true2026.03 | 41.91 | |
| FIXED_E=16Experts=162026.05 | 41.91 | |
| OLMo 2-7BModel Category=Fully open models, Instruction-tuned=true2026.03 | 41.9 | |
| ELSA3DModel Category=Unified 3D Models2026.07 | 41.8 | |
| PLDRv51-SOC-110M-1Zero-shot=true2026.03 | 41.66 | |
| PLDRv51-SOC-110M-4Zero-shot=true2026.03 | 41.66 | |
| FIXED_E=32Experts=322026.05 | 41.56 | |
| ShapeLLM-Omni-7BParameters=7B2025.12 | 41.5 | |
| CoRe3DParameters=7B2025.12 | 41.5 | |
| ShapeLLM-OmniModel Category=Unified 3D Models2026.07 | 41.5 | |
| CoRe3DModel Category=Unified 3D Models2026.07 | 41.5 | |
| Qwen 3-8BModel Category=Open weight models, Instruction-tuned=true2026.03 | 41.2 | |
| EMO (Stage 3)E=16→322026.05 | 41.15 | |
| Gaussian-BaseCPT=Gaussian, SFT=Base2026.05 | 41.04 | |
| Qwen2.5-vl-7BParameters=7B2025.12 | 41 | |
| Gemma 2-9BModel Category=Open weight models, Instruction-tuned=true2026.03 | 41 | |
| Qwen2.5-VLModel Category=General VLMs2026.07 | 41 | |
| Mistral-7BModel Category=Open weight models, Instruction-tuned=true2026.03 | 40.9 | |
| EuroLLM-9BModel Category=Fully open models, Instruction-tuned=true2026.03 | 40.8 | |
| ABL-SOC-110M-1Zero-shot=true2026.03 | 40.79 | |
| Qwen 2.5-7BModel Category=Open weight models, Instruction-tuned=true2026.03 | 40.7 | |
| LLaMA3.2-Vision-11BParameters=11B2025.12 | 40.6 | |
| LLaMA3.2-VisionModel Category=General VLMs2026.07 | 40.6 | |
| LLaMA-Mesh-8BParameters=8B2025.12 | 40.3 | |
| LLaMA-MeshModel Category=Mesh LLM2026.07 | 40.3 | |
| EMO (Stage 2)E=8→162026.05 | 40.23 | |
| EMO (Stage 1)E=82026.05 | 39.92 | |
| BaseCPT=Base, SFT=Base2026.05 | 38.95 |