Common Sense Reasoning on WG
94.1AccuracyHuman
Evaluation Results
| Method | Links | |
|---|---|---|
| HumanKB=-2026.03 | 94.1 | |
| DeBERTa-v3-L (Supervised)KB=-, Evaluation Protocol=Supervised2026.03 | 84.1 | |
| IMAGINE-DeBERTa-v3-LKB=Synthetic VQA+, Evaluation Protocol=Zero-shot2026.03 | 79.3 | |
| RoBERTa-L (Supervised)KB=-, Evaluation Protocol=Supervised2026.03 | 79.3 | |
| IMAGINE-DeBERTa-v3-L (Retrieval)KB=Synthetic VQA+, Inference Strategy=Retrieval, Evaluation Protocol=Zero-shot2026.03 | 79.2 | |
| Qwen3-235B-A22B#Bits=W16A16, Regime=PTQ, Zero-shot=true, Architecture=MoE2026.06 | 79.01 | |
| CANDLE-DeBERTa-v3-LKB=CANDLE, Evaluation Protocol=Zero-shot2026.03 | 78.3 | |
| CAR-DeBERTa-v3-LKB=AbsAT, Evaluation Protocol=Zero-shot2026.03 | 78.2 | |
| Llama2-70B#Bits=W16A16, Regime=PTQ, Zero-shot=true, Architecture=Dense2026.06 | 77.58 | |
| GPT-4 (gpt-4)KB=-, Evaluation Protocol=Zero-shot2026.03 | 77 | |
| IMAGINE-DeBERTa-v3-LKB=Synthetic VQA, Evaluation Protocol=Zero-shot2026.03 | 76.7 | |
| DeBERTa-v3-L (MR)KB=AT, Evaluation Protocol=Zero-shot2026.03 | 76 | |
| CAT-QBase Model=Llama2-70B, #Bits=W1.58A16, Regime=PTQ, Zero-shot=true, Architecture=Dense2026.06 | 75.34 | |
| Mistral-v0.1-7BKB=-, Evaluation Protocol=Zero-shot2026.03 | 75.3 | |
| Qwen3-32B#Bits=W16A16, Regime=PTQ, Zero-shot=true, Architecture=Dense2026.06 | 73.56 | |
| Qwen3-14B#Bits=W16A16, Regime=PTQ, Zero-shot=true, Architecture=Dense2026.06 | 73.16 | |
| LLAMA2-13BKB=-, Evaluation Protocol=Zero-shot2026.03 | 72.8 | |
| CAT-QBase Model=Qwen3-32B, #Bits=W1.58A16, Regime=PTQ, Zero-shot=true, Architecture=Dense2026.06 | 71.74 | |
| CANDLE-VERA-T5-xxlKB=CANDLE, Evaluation Protocol=Zero-shot2026.03 | 71.3 | |
| RF2-100B-A6.1B#Bits=W16A16, Regime=PTQ, Zero-shot=true, Architecture=MoE2026.06 | 70.32 | |
| Qwen3-30B-A3B#Bits=W16A16, Regime=PTQ, Zero-shot=true, Architecture=MoE2026.06 | 69.93 | |
| CAT-QBase Model=Qwen3-235B-A22B, #Bits=W1.58A16, Regime=PTQ, Zero-shot=true, Architecture=MoE2026.06 | 69.85 | |
| Llama2-7B#Bits=W16A16, Regime=PTQ, Zero-shot=true, Architecture=Dense2026.06 | 69.46 | |
| VERA-T5-xxlKB=AbsAT, Evaluation Protocol=Zero-shot2026.03 | 68.1 | |
| Qwen3-8B#Bits=W16A16, Regime=PTQ, Zero-shot=true, Architecture=Dense2026.06 | 68.03 | |
| CAT-QBase Model=Qwen3-14B, #Bits=W1.58A16, Regime=PTQ, Zero-shot=true, Architecture=Dense2026.06 | 68.03 | |
| VERA-T5-xxlKB=AT, Evaluation Protocol=Zero-shot2026.03 | 67.2 | |
| Qwen3-4B#Bits=W16A16, Regime=PTQ, Zero-shot=true, Architecture=Dense2026.06 | 65.51 | |
| CAT-QBase Model=Qwen3-8B, #Bits=W1.58A16, Regime=PTQ, Zero-shot=true, Architecture=Dense2026.06 | 65.43 | |
| ChatGPT (gpt-3.5-turbo)KB=-, Evaluation Protocol=Zero-shot2026.03 | 64.1 | |
| CAT-QBase Model=RF2-100B-A6.1B, #Bits=W1.58A16, Regime=PTQ, Zero-shot=true, Architecture=MoE2026.06 | 63.93 | |
| CAT-QBase Model=Qwen3-30B-A3B, #Bits=W1.58A16, Regime=PTQ, Zero-shot=true, Architecture=MoE2026.06 | 63.85 | |
| CAT-QBase Model=Qwen3-4B, #Bits=W1.58A8, Regime=PTQ, Zero-shot=true, Architecture=Dense2026.06 | 63.51 | |
| CAT-QBase Model=Qwen3-8B, #Bits=W1.58A8, Regime=PTQ, Zero-shot=true, Architecture=Dense2026.06 | 62.67 | |
| CAT-QBase Model=Qwen3-4B, #Bits=W1.58A16, Regime=PTQ, Zero-shot=true, Architecture=Dense2026.06 | 62.19 | |
| CAR-RoBERTa-LKB=AbsAT, Evaluation Protocol=Zero-shot2026.03 | 62 | |
| Qwen3-1.7B#Bits=W16A16, Regime=PTQ, Zero-shot=true, Architecture=Dense2026.06 | 61.88 | |
| GenKnowSubBase Model=Phi-2, Setting=En, Evaluation Protocol=zero-shot2025.05 | 61.24 | |
| IMAGINE-RoBERTa-LKB=Synthetic VQA, Evaluation Protocol=Zero-shot2026.03 | 61.2 | |
| Multi-hop Knowledge InjectionKB=AT, CN, WD, WN, Evaluation Protocol=Zero-shot2026.03 | 61 | |
| CAT-QBase Model=Llama2-7B, #Bits=W1.58A16, Regime=PTQ, Zero-shot=true, Architecture=Dense2026.06 | 60.93 | |
| ArrowBase Model=Phi-2, Evaluation Protocol=zero-shot2025.05 | 60.85 | |
| Zero-shot FusionKB=AT, CN, WD, WN, Evaluation Protocol=Zero-shot2026.03 | 60.8 | |
| GPT-3.5 (text-davinci-003)KB=-, Evaluation Protocol=Zero-shot2026.03 | 60.7 | |
| GenKnowSubBase Model=Phi-2, Setting=Avg, Evaluation Protocol=zero-shot2025.05 | 60.69 | |
| H128-MQA-CLA2Model Scale=3B, Head Dimension=128, Sharing Factor=22024.05 | 60.69 | |
| H128-MQAModel Scale=3B, Head Dimension=1282024.05 | 60.46 | |
| RoBERTa-L (MR)KB=AT, Evaluation Protocol=Zero-shot2026.03 | 59.6 | |
| InstructBLIP-Vicuna-7BKB=-, Evaluation Protocol=Zero-shot2026.03 | 59.6 | |
| H64-MQAModel Scale=3B, Head Dimension=642024.05 | 57.85 | |
| RoBERTa-LKB=-, Evaluation Protocol=Zero-shot2026.03 | 57.5 | |
| Phi-2Base Model=Phi-2, Evaluation Protocol=zero-shot2025.05 | 56.51 | |
| CAR-GPT-2-LKB=AbsAT, Evaluation Protocol=Zero-shot2026.03 | 55.2 | |
| IMAGINE-GPT-2-LKB=Synthetic VQA, Evaluation Protocol=Zero-shot2026.03 | 55.2 | |
| CAT-QBase Model=Qwen3-1.7B, #Bits=W1.58A16, Regime=PTQ, Zero-shot=true, Architecture=Dense2026.06 | 54.83 | |
| Self-talkKB=-, Evaluation Protocol=Zero-shot2026.03 | 54.7 | |
| GPT-2-L (MR)KB=AT, Evaluation Protocol=Zero-shot2026.03 | 54.7 | |
| CAT-QBase Model=Qwen3-1.7B, #Bits=W1.58A8, Regime=PTQ, Zero-shot=true, Architecture=Dense2026.06 | 54.62 | |
| LLaVA-1.5-7BKB=-, Evaluation Protocol=Zero-shot2026.03 | 54.5 | |
| GPT-2-LKB=-, Evaluation Protocol=Zero-shot2026.03 | 53.2 | |
| DeBERTa-v3-LKB=-, Evaluation Protocol=Zero-shot2026.03 | 50.3 |