Physical Commonsense Reasoning on PIQA
94.9AccuracyHuman
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| HumanKB=-2026.03 | 94.9 | — | — | — | |
| BF16Backbone=DeepSeek-V3.1 671B2026.02 | 92.93 | — | — | — | |
| HiF4Backbone=DeepSeek-V3.1 671B2026.02 | 92.44 | -0.49 | — | — | |
| BF16Backbone=LongCat 560B2026.02 | 91.46 | — | — | — | |
| NVFP4Backbone=DeepSeek-V3.1 671B2026.02 | 91.29 | -1.64 | — | — | |
| NVFP4+PTSBackbone=DeepSeek-V3.1 671B2026.02 | 91.29 | -1.64 | — | — | |
| HiF4Backbone=LongCat 560B2026.02 | 91.19 | -0.27 | — | — | |
| Qwen3Model=Qwen3, Number of Parameters=32B, Privacy Protocol=Plaintext2026.03 | 90.75 | — | — | — | |
| Hard-routing MoEBackbone=Qwen2.5-7B-Instruct, # Params (%)=14.2G (100%), # Trans. (%)=14.2G (100%), Fine-tuning setting=DFL, Rank (rl)=322026.02 | 90.46 | — | — | — | |
| AloePriModel=Qwen3, Number of Parameters=14B, Privacy Protocol=AloePri2026.03 | 90.26 | — | — | — | |
| NVFP4+PTSBackbone=LongCat 560B2026.02 | 89.83 | -1.63 | — | — | |
| Qwen3Model=Qwen3, Number of Parameters=14B, Privacy Protocol=Plaintext2026.03 | 89.72 | — | — | — | |
| Qwen3 14BClassifier=Self-labeled2026.01 | 89.7 | — | — | — | |
| Qwen3 14BClassifier=Majority-labeled2026.01 | 89.6 | — | — | — | |
| Qwen3-MoE-InstructModel=Qwen3-MoE-Instruct, Number of Parameters=30B-A3B, Privacy Protocol=Plaintext2026.03 | 89.55 | — | — | — | |
| AloePriModel=Qwen3, Number of Parameters=32B, Privacy Protocol=AloePri2026.03 | 89.5 | — | — | — | |
| NVFP4Backbone=LongCat 560B2026.02 | 89.39 | -2.07 | — | — | |
| Qwen3 14B BaseClassifier=Majority-labeled2026.01 | 89.3 | — | — | — | |
| Qwen3 8BClassifier=Self-labeled2026.01 | 88.8 | — | — | — | |
| AloePriModel=R1-Distill, Number of Parameters=14B, Privacy Protocol=AloePri2026.03 | 88.79 | — | — | — | |
| R1-DistillModel=R1-Distill, Number of Parameters=32B, Privacy Protocol=Plaintext2026.03 | 88.79 | — | — | — | |
| Qwen3 8BClassifier=Majority-labeled2026.01 | 88.7 | — | — | — | |
| Yi-34B + RTDevaluation=zero-shot2024.09 | 88.4 | — | — | — | |
| Yi-34Bevaluation=zero-shot2024.09 | 88.3 | — | — | — | |
| Qwen3 8B BaseClassifier=Self-labeled2026.01 | 88.2 | — | — | — | |
| In-Squeezereduction=128 ... -> 1, strategy=Standard2026.02 | 88.17 | — | — | — | |
| R1-DistillModel=R1-Distill, Number of Parameters=14B, Privacy Protocol=Plaintext2026.03 | 88.14 | — | — | — | |
| Qwen3 14B BaseClassifier=Self-labeled2026.01 | 88.1 | — | — | — | |
| AloePriModel=R1-Distill, Number of Parameters=32B, Privacy Protocol=AloePri2026.03 | 88.08 | — | — | — | |
| Qwen3 8B BaseClassifier=Majority-labeled2026.01 | 87.8 | — | — | — | |
| Yi-34B + RTDevaluation=5-shot2024.09 | 87.7 | — | — | — | |
| Cont-Squeezereduction=128 -> 1, training_steps=+700 steps2026.02 | 87.53 | — | — | — | |
| In-Squeezereduction=128 ... -> 1, strategy=Min steps2026.02 | 87.28 | — | — | — | |
| Cont-Squeezereduction=128 -> 1, training_steps=+200 steps2026.02 | 87.17 | — | — | — | |
| Direct Fine-tuningrank=1, training_steps=+0 steps2026.02 | 86.94 | — | — | — | |
| Sparse-and-Orthogonal LoRABackbone=Qwen2.5-7B-Instruct, # Params (%)=87M (0.60%), # Trans. (%)=42M (0.30%), Fine-tuning setting=DFL, Rank (rl)=322026.02 | 86.86 | — | — | — | |
| LLaMA2-70B + RTDevaluation=5-shot2024.09 | 86.6 | — | — | — | |
| Direct Fine-tuningrank=1, training_steps=+700 steps2026.02 | 86.47 | — | — | — | |
| Direct Fine-tuningrank=1, training_steps=+200 steps2026.02 | 86.44 | — | — | — | |
| LFPO (All Loss)Base Model=LLaDA 8B, Training Algorithm=LFPO (All Loss)2026.03 | 85.9 | — | — | — | |
| Sparse-and-Orthogonal LoRA (Single)Backbone=Qwen2.5-7B-Instruct, # Params (%)=87M (0.60%), # Trans. (%)=42M (0.30%), Fine-tuning setting=DFL, Rank (rl)=322026.02 | 85.74 | — | — | — | |
| FPFTBackbone=Qwen2.5-7B-Instruct, # Params (%)=14.2G (100%), # Trans. (%)=14.2G (100%), Fine-tuning setting=DFL, Rank (rl)=322026.02 | 85.65 | — | — | — | |
| AGRPOBase Model=LLaDA 8B, Training Algorithm=AGRPO2026.03 | 85.6 | — | — | — | |
| Qwen3 4BClassifier=Self-labeled2026.01 | 85.4 | — | — | — | |
| SQ-formatModel=DeepSeek-R1, Setting=W(SQ5)A8, Sparsity=0.875, Evaluation Protocol=Non-generative2025.12 | 85.31 | — | — | — | |
| LLaMA2-70Bevaluation=5-shot ICL2024.09 | 85.3 | — | — | — | |
| S-MeZOBackbone=Mistral-7B, Evaluation Protocol=Fine-Tuning2024.02 | 85.3 | — | — | — | |
| Qwen3 4BClassifier=Majority-labeled2026.01 | 85.3 | — | — | — | |
| GLM-4.5 BaseArchitecture=MoE, Activated Params=32B, Total Params=355B2026.02 | 85.3 | — | — | — | |
| Qwen3-4BParams=4B2025.12 | 84.98 | — | — | — | |
| DeepSeek-R1 (BF16)Model=DeepSeek-R1, Setting=W16A16, Sparsity=0, Evaluation Protocol=Non-generative2025.12 | 84.98 | — | — | — | |
| Hard-routing MoEBackbone=Qwen2.5-1.5B-Instruct, # Params (%)=3.2G (100%), # Trans. (%)=3.2G (100%), Fine-tuning setting=DFL, Rank (rl)=322026.02 | 84.87 | — | — | — | |
| DeepSeek-V3 BaseArchitecture=MoE, Activated Params=37B, Total Params=671B2026.02 | 84.7 | — | — | — | |
| GLM-5 BaseArchitecture=MoE, Activated Params=40B, Total Params=744B2026.02 | 84.6 | — | — | — | |
| DeBERTa-v3-L (Supervised)KB=-, Evaluation Protocol=Supervised2026.03 | 84.5 | — | — | — | |
| MeZOBackbone=Mistral-7B, Evaluation Protocol=Fine-Tuning2024.02 | 84.3 | — | — | — | |
| LoRABackbone=Qwen2.5-7B-Instruct, # Params (%)=362M (2.5%), # Trans. (%)=362M (2.5%), Fine-tuning setting=DFL, Rank (rl)=322026.02 | 83.88 | — | — | — | |
| DS-V3.1-Terminus (no_think)Model=DS-V3.1-Terminus (no_think), Number of Parameters=671B, Privacy Protocol=Plaintext2026.03 | 83.84 | — | — | — | |
| Qwen3 1.7BClassifier=Self-labeled2026.01 | 83.8 | — | — | — | |
| Qwen3 1.7BClassifier=Majority-labeled2026.01 | 83.6 | — | — | — | |
| Yi-34Bevaluation=5-shot ICL2024.09 | 83.5 | — | — | — | |
| Qwen3 4B BaseClassifier=Self-labeled2026.01 | 83.5 | — | — | — | |
| Sparse-and-Orthogonal LoRABackbone=Qwen2.5-1.5B-Instruct, # Params (%)=16M (0.54%), # Trans. (%)=8M (0.27%), Fine-tuning setting=DFL, Rank (rl)=322026.02 | 83.36 | — | — | — | |
| LoRIBackbone=Qwen2.5-7B-Instruct, # Params (%)=168M (1.16%), # Trans. (%)=168M (1.16%), Fine-tuning setting=DFL, Rank (rl)=322026.02 | 83.22 | — | — | — | |
| Our Trained ModelModel Size=52B, Data type=W8A8, Zero-shot=true2023.05 | 83.2 | — | — | — | |
| coupled-GRPOBase Model=LLaDA 8B, Training Algorithm=coupled-GRPO2026.03 | 83.2 | — | — | — | |
| Our Trained ModelModel Size=52B, Data type=FP16, Zero-shot=true2023.05 | 83.19 | — | — | — | |
| Mistral-v0.1-7BKB=-, Evaluation Protocol=Zero-shot2026.03 | 83 | — | — | — | |
| Qwen3 0.6BClassifier=Self-labeled2026.01 | 82.6 | — | — | — | |
| Qwen3 0.6BClassifier=Majority-labeled2026.01 | 82.5 | — | — | — | |
| Cont-Squeezereduction=128 -> 1, training_steps=+0 steps2026.02 | 82.39 | — | — | — | |
| AloePriModel=Qwen3-MoE-Instruct, Number of Parameters=30B-A3B, Privacy Protocol=AloePri2026.03 | 82.21 | — | — | — | |
| Qwen3 4B BaseClassifier=Majority-labeled2026.01 | 82.2 | — | — | — | |
| Qwen3 1.7B BaseClassifier=Self-labeled2026.01 | 82 | — | — | — | |
| Qwen3 1.7B BaseClassifier=Majority-labeled2026.01 | 82 | — | — | — | |
| LLaMA2-70B + RTDevaluation=zero-shot2024.09 | 81.9 | — | — | — | |
| ChatGPT (gpt-3.5-turbo)KB=-, Evaluation Protocol=Zero-shot2026.03 | 81.7 | — | — | — | |
| Sparse-and-Orthogonal LoRA (Single)Backbone=Qwen2.5-1.5B-Instruct, # Params (%)=16M (0.54%), # Trans. (%)=8M (0.27%), Fine-tuning setting=DFL, Rank (rl)=322026.02 | 81.69 | — | — | — | |
| LoRIBackbone=Qwen2.5-1.5B-Instruct, # Params (%)=32M (1.05%), # Trans. (%)=32M (1.05%), Fine-tuning setting=DFL, Rank (rl)=322026.02 | 81.48 | — | — | — | |
| Qwen2.5-3BParams=3B2025.12 | 81.45 | — | — | — | |
| IMAGINE-DeBERTa-v3-LKB=Synthetic VQA+, Evaluation Protocol=Zero-shot2026.03 | 81.4 | — | — | — | |
| Qwen2.5-vl-7BParameters=7B2025.12 | 81.3 | — | — | — | |
| IMAGINE-DeBERTa-v3-L (Retrieval)KB=Synthetic VQA+, Inference Strategy=Retrieval, Evaluation Protocol=Zero-shot2026.03 | 81.3 | — | — | — | |
| BitDelta (scalar)Model=Llama-3.1-8B-Instruct, Zero-shot protocol=true2025.12 | 81.22 | — | — | — | |
| UniGRPOBase Model=LLaDA 8B, Training Algorithm=UniGRPO2026.03 | 81.2 | — | — | — | |
| Evo 8BShots=0, E2E Latency (s)=8.6, Inference Speed (tokens/s)=522026.02 | 81.2 | — | — | — | |
| DPOBase model=Qwen-2 instruct 7B2025.05 | 81.12 | — | — | — | |
| SGDPOBase model=Qwen-2 instruct 7B2025.05 | 81.07 | — | — | — | |
| OLMO-2Model Type=Fully-open, Number of Parameters=7B2026.03 | 81.07 | — | — | — | |
| GPT-3Model Size=175B, Evaluation Protocol=Zero-shot2022.10 | 81 | — | — | — | |
| MambaSize=7B, Tokens=1.2T, Shot(s)=0, Training Strategy=trained from scratch2024.09 | 81 | — | — | — | |
| OriginalRatio=0%, Backbone=Qwen2-57B-A14B, Evaluation Protocol=zero-shot2025.09 | 81 | — | — | — | |
| HEAPrRatio=40%, Backbone=Qwen2-57B-A14B, Evaluation Protocol=zero-shot2025.09 | 81 | — | — | — | |
| SFTBase model=Qwen-2 instruct 7B2025.05 | 80.96 | — | — | — | |
| AloePriModel=DS-V3.1-Terminus (no_think), Number of Parameters=671B, Privacy Protocol=AloePri2026.03 | 80.96 | — | — | — | |
| SPGBase Model=LLaDA 8B, Training Algorithm=SPG2026.03 | 80.9 | — | — | — | |
| TDPOBase model=Qwen-2 instruct 7B2025.05 | 80.89 | — | — | — | |
| Vector (row/col)Model=Phi-4-reasoning, Zero-shot protocol=true2025.12 | 80.85 | — | — | — | |
| NCABase model=Qwen-2 instruct 7B2025.05 | 80.79 | — | — | — | |
| BCOBase model=Qwen-2 instruct 7B2025.05 | 80.74 | — | — | — |