Physical Commonsense Reasoning on PIQA (test)
90.7AccuracyUL20B
Evaluation Results
| Method | Links | |
|---|---|---|
| UL20Bmode=supervised finetuning, leaderboard submission=true2022.05 | 90.7 | |
| UNICORN 11BModel=UNICORN 11B2022.10 | 90.1 | |
| Lourie et al.status=previous state-of-the-art2022.05 | 90.1 | |
| UnifiedQA 11BModel=UnifiedQA 11B2022.10 | 89.5 | |
| ScoreBackbone=DeBERTa Large2022.10 | 87.41 | |
| Few-shot*Backbone=Qwen2.5-7B, Few-shot settings=5-shot ID2025.09 | 86 | |
| TEAMBackbone=DeBERTa Large2022.10 | 85.9 | |
| FP16Model=LLAMA3-70B, Evaluation Protocol=Zero-shot, Bit-width=Full-precision2024.03 | 84.66 | |
| Llama3.1-8B-Instruct + SFTBase Model=Llama3.1-8B-Instruct, Training=SFT2026.05 | 84.6 | |
| Qwen3-8B + SFT + WeMask(TF)Base Model=Qwen3-8B, Training=SFT, Variant=TF2026.05 | 84.44 | |
| Llama3.1-8B-Instruct + SFT + WeMask(TF)Base Model=Llama3.1-8B-Instruct, Training=SFT, Variant=TF2026.05 | 84.22 | |
| Llama3.1-8B-Instruct + WeMask(SFT)Base Model=Llama3.1-8B-Instruct, Training=WeMask(SFT)2026.05 | 84.22 | |
| Qwen3-8B + WeMask(SFT)Base Model=Qwen3-8B, Training=WeMask(SFT)2026.05 | 84.11 | |
| Qwen3-8B + SFTBase Model=Qwen3-8B, Training=SFT2026.05 | 84.06 | |
| LLAMA 1Size=65B2023.07 | 82.8 | |
| LLAMA 2Size=70B2023.07 | 82.8 | |
| ICRBackbone=Qwen2.5-7B2025.09 | 82.6 | |
| M²IVBackbone=Qwen2.5-7B2025.09 | 82.5 | |
| FalconSize=40B2023.07 | 82.4 | |
| LLAMA 1Size=33B2023.07 | 82.3 | |
| LIVEBackbone=Qwen2.5-7B2025.09 | 82 | |
| MPTSize=30B2023.07 | 81.9 | |
| LLAMA 2Size=34B2023.07 | 81.9 | |
| I2CLBackbone=Qwen2.5-7B2025.09 | 81.2 | |
| FP16Model=LLAMA3-8B, Evaluation Protocol=Zero-shot, Bit-width=Full-precision2024.03 | 80.74 | |
| MPTSize=7B2023.07 | 80.6 | |
| LLAMA 2Size=13B2023.07 | 80.5 | |
| LLAMA 1Size=13B2023.07 | 80.1 | |
| LLAMA 1Size=7B2023.07 | 79.8 | |
| ScoreBackbone=RoBERTa Large2022.10 | 79.4 | |
| Interaction-Aware Influence Function (Ours)Base Model=Llama-3.1-8B, Evaluation Protocol=Fine-tuned, Number of Seeds=52026.05 | 79.39 | |
| LLAMA 2Size=7B2023.07 | 78.8 | |
| QuaRotModel=LLAMA3-70B, Evaluation Protocol=Zero-shot, Bit-width=4-bit2024.03 | 78.07 | |
| OracleModel Size=Large, Expertise Distribution=Dist. 32026.04 | 77.8 | |
| REALMModel Size=Large, Noise Type=Systematic, Expertise Distribution=Dist. 32026.04 | 77.67 | |
| RandomBase Model=Llama-3.1-8B, Evaluation Protocol=Fine-tuned, Number of Seeds=52026.05 | 77.5 | |
| LESSBase Model=Llama-3.1-8B, Evaluation Protocol=Fine-tuned, Number of Seeds=52026.05 | 77.4 | |
| Additive IFBase Model=Llama-3.1-8B, Evaluation Protocol=Fine-tuned, Number of Seeds=52026.05 | 77.38 | |
| FalconSize=7B2023.07 | 76.7 | |
| REALMModel Size=Large, Noise Type=Uniform, Expertise Distribution=Dist. 32026.04 | 76.5 | |
| Zero-shotBackbone=Qwen2.5-7B2025.09 | 76.2 | |
| REALMModel Size=Large, Noise Type=Asymmetric, Expertise Distribution=Dist. 32026.04 | 76.17 | |
| NOISYModel Size=Large, Noise Type=Systematic, Expertise Distribution=Dist. 32026.04 | 75.6 | |
| QuaRotModel=LLAMA3-8B, Evaluation Protocol=Zero-shot, Bit-width=4-bit2024.03 | 75.14 | |
| Llama3.1-8B-InstructBase Model=Llama3.1-8B-Instruct2026.05 | 75.14 | |
| RDS+Base Model=Llama-3.1-8B, Evaluation Protocol=Fine-tuned, Number of Seeds=52026.05 | 75.14 | |
| NV-EmbedBase Model=Llama-3.1-8B, Evaluation Protocol=Fine-tuned, Number of Seeds=52026.05 | 75.13 | |
| TEAMBackbone=RoBERTa Large2022.10 | 74.55 | |
| OracleModel Size=Base, Expertise Distribution=Dist. 32026.04 | 67.68 | |
| REALMModel Size=Base, Noise Type=Systematic, Expertise Distribution=Dist. 32026.04 | 66.92 | |
| REALMModel Size=Base, Noise Type=Uniform, Expertise Distribution=Dist. 32026.04 | 65.79 | |
| REALMModel Size=Base, Noise Type=Asymmetric, Expertise Distribution=Dist. 32026.04 | 65.77 | |
| NOISYModel Size=Base, Noise Type=Systematic, Expertise Distribution=Dist. 32026.04 | 64.87 | |
| Qwen3-8BBase Model=Qwen3-8B2026.05 | 63.11 | |
| Few-shot*Backbone=Llama2-7B, Few-shot settings=5-shot ID2025.09 | 59.8 | |
| ICRBackbone=Llama2-7B2025.09 | 57 | |
| M²IVBackbone=Llama2-7B2025.09 | 56.8 | |
| LIVEBackbone=Llama2-7B2025.09 | 56.4 | |
| I2CLBackbone=Llama2-7B2025.09 | 55.6 | |
| NOISYModel Size=Large, Noise Type=Asymmetric, Expertise Distribution=Dist. 32026.04 | 53.47 | |
| OracleModel Size=Small, Expertise Distribution=Dist. 32026.04 | 52.88 | |
| REALMModel Size=Small, Noise Type=Uniform, Expertise Distribution=Dist. 32026.04 | 52.88 | |
| Zero-shotBackbone=Llama2-7B2025.09 | 52.2 | |
| NOISYModel Size=Base, Noise Type=Asymmetric, Expertise Distribution=Dist. 32026.04 | 52.1 | |
| NOISYModel Size=Large, Noise Type=Uniform, Expertise Distribution=Dist. 32026.04 | 51.99 | |
| NOISYModel Size=Base, Noise Type=Uniform, Expertise Distribution=Dist. 32026.04 | 51.47 | |
| REALMModel Size=Small, Noise Type=Systematic, Expertise Distribution=Dist. 32026.04 | 51.32 | |
| NOISYModel Size=Small, Noise Type=Systematic, Expertise Distribution=Dist. 32026.04 | 50.84 | |
| NOISYModel Size=Small, Noise Type=Asymmetric, Expertise Distribution=Dist. 32026.04 | 50.58 | |
| NOISYModel Size=Small, Noise Type=Uniform, Expertise Distribution=Dist. 32026.04 | 50.36 | |
| REALMModel Size=Small, Noise Type=Asymmetric, Expertise Distribution=Dist. 32026.04 | 50.25 |