Social Commonsense Reasoning on SocialIQA
87.11AccuracyHard-routing MoE
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Hard-routing MoEBackbone=Qwen2.5-7B-Instruct, # Params (%)=14.2G (100%), # Trans. (%)=14.2G (100%), Fine-tuning setting=DFL, Rank (rl)=322026.02 | 87.11 | — | — | |
| Sparse-and-Orthogonal LoRABackbone=Qwen2.5-7B-Instruct, # Params (%)=87M (0.60%), # Trans. (%)=42M (0.30%), Fine-tuning setting=DFL, Rank (rl)=322026.02 | 85.52 | — | — | |
| FPFTBackbone=Qwen2.5-7B-Instruct, # Params (%)=14.2G (100%), # Trans. (%)=14.2G (100%), Fine-tuning setting=DFL, Rank (rl)=322026.02 | 83.96 | — | — | |
| Sparse-and-Orthogonal LoRA (Single)Backbone=Qwen2.5-7B-Instruct, # Params (%)=87M (0.60%), # Trans. (%)=42M (0.30%), Fine-tuning setting=DFL, Rank (rl)=322026.02 | 83.74 | — | — | |
| LoRIBackbone=Qwen2.5-7B-Instruct, # Params (%)=168M (1.16%), # Trans. (%)=168M (1.16%), Fine-tuning setting=DFL, Rank (rl)=322026.02 | 83.4 | — | — | |
| Hard-routing MoEBackbone=Qwen2.5-1.5B-Instruct, # Params (%)=3.2G (100%), # Trans. (%)=3.2G (100%), Fine-tuning setting=DFL, Rank (rl)=322026.02 | 83.04 | — | — | |
| O32026.03 | 82.91 | — | — | |
| GPT-52026.03 | 82.69 | — | — | |
| GPT-5Mode=COT2026.03 | 81.63 | — | — | |
| LoRABackbone=Qwen2.5-7B-Instruct, # Params (%)=362M (2.5%), # Trans. (%)=362M (2.5%), Fine-tuning setting=DFL, Rank (rl)=322026.02 | 80.99 | — | — | |
| DeepSeek-R12026.03 | 80.6 | — | — | |
| Distill-Llama-70B2026.03 | 80.55 | — | — | |
| Qwen3 14B BaseClassifier=Self-labeled2026.01 | 79.9 | — | — | |
| Qwen3 14B BaseClassifier=Majority-labeled2026.01 | 79.8 | — | — | |
| O3Mode=COT2026.03 | 79.63 | — | — | |
| GPT-4oMode=COT2026.03 | 79.53 | — | — | |
| Sparse-and-Orthogonal LoRABackbone=Qwen2.5-1.5B-Instruct, # Params (%)=16M (0.54%), # Trans. (%)=8M (0.27%), Fine-tuning setting=DFL, Rank (rl)=322026.02 | 79.49 | — | — | |
| Mistral Small 24B Inst 2501Classifier=Self-labeled2026.01 | 79.2 | 20 | — | |
| Mistral Small 24B Inst 2501Classifier=Majority-labeled2026.01 | 79.2 | 20 | — | |
| LoRABackbone=Qwen2.5-1.5B-Instruct, # Params (%)=65.5M (2.01%), # Trans. (%)=65.5M (2.01%), Fine-tuning setting=DFL, Rank (rl)=322026.02 | 79.2 | — | — | |
| FPFTBackbone=Qwen2.5-1.5B-Instruct, # Params (%)=3.2G (100%), # Trans. (%)=3.2G (100%), Fine-tuning setting=DFL, Rank (rl)=322026.02 | 79.16 | — | — | |
| Qwen3 8B BaseClassifier=Self-labeled2026.01 | 78.8 | 20 | — | |
| Qwen3 8B BaseClassifier=Majority-labeled2026.01 | 78.8 | 20 | — | |
| Qwen3 8B BaseClassifier=Self-labeled2026.01 | 78.8 | — | — | |
| Qwen3 8B BaseClassifier=Majority-labeled2026.01 | 78.8 | — | — | |
| LoRIBackbone=Qwen2.5-1.5B-Instruct, # Params (%)=32M (1.05%), # Trans. (%)=32M (1.05%), Fine-tuning setting=DFL, Rank (rl)=322026.02 | 78.64 | — | — | |
| GPT-4o2026.03 | 78.4 | — | — | |
| Sparse-and-Orthogonal LoRA (Single)Backbone=Qwen2.5-1.5B-Instruct, # Params (%)=16M (0.54%), # Trans. (%)=8M (0.27%), Fine-tuning setting=DFL, Rank (rl)=322026.02 | 78.38 | — | — | |
| Qwen3-32B2026.03 | 77.74 | — | — | |
| SocialR1-8BAblation=only Rout2026.03 | 77.74 | — | — | |
| SocialR1-8BAblation=Full2026.03 | 77.53 | — | — | |
| Qwen3-8B2026.03 | 77.28 | — | — | |
| Qwen3-32BThinking=Disabled2026.03 | 76.15 | — | — | |
| Qwen3-8BThinking=Disabled2026.03 | 76 | — | — | |
| Qwen3 14BClassifier=Self-labeled2026.01 | 75.9 | — | — | |
| Qwen3 8BClassifier=Self-labeled2026.01 | 75.8 | — | — | |
| Qwen3 8BClassifier=Majority-labeled2026.01 | 75.8 | — | — | |
| Qwen3 14BClassifier=Majority-labeled2026.01 | 75.8 | — | — | |
| LLaMa3.1-70BMode=COT2026.03 | 75.74 | — | — | |
| SocialR1-8BAblation=w/o Rlen2026.03 | 75.74 | — | — | |
| SocialR1-4BAblation=Full2026.03 | 75.08 | — | — | |
| Qwen3-4B2026.03 | 74.51 | — | — | |
| SocialR1-4BAblation=w/o Rlen2026.03 | 74.51 | — | — | |
| Qwen3 4B BaseClassifier=Self-labeled2026.01 | 74 | — | — | |
| Qwen3 4B BaseClassifier=Majority-labeled2026.01 | 74 | — | — | |
| Mistral Nemo Inst 2407Classifier=Self-labeled2026.01 | 73.4 | -10 | — | |
| Mistral Nemo Inst 2407Classifier=Majority-labeled2026.01 | 73.4 | -10 | — | |
| SocialR1-8BAblation=w/o Rcont2026.03 | 73.23 | — | — | |
| SocialR1-4BAblation=only Rout2026.03 | 73.18 | — | — | |
| Qwen3-4BThinking=Disabled2026.03 | 73.13 | — | — | |
| SocialR1-8BAblation=w/o Rstruct2026.03 | 73.08 | — | — | |
| SocialR1-4BAblation=w/o Rcont2026.03 | 72.82 | — | — | |
| Qwen3 4BClassifier=Self-labeled2026.01 | 72.2 | — | — | |
| Qwen3 4BClassifier=Majority-labeled2026.01 | 72.2 | — | — | |
| SocialR1-4BAblation=w/o Rstruct2026.03 | 72.01 | — | — | |
| Llama 3.1 8B InstClassifier=Self-labeled2026.01 | 70.7 | 0 | — | |
| Llama 3.1 8B InstClassifier=Majority-labeled2026.01 | 70.6 | -20 | — | |
| Mistral Small 24B Base 2501Classifier=Self-labeled2026.01 | 70.2 | 20 | — | |
| Mistral Small 24B Base 2501Classifier=Majority-labeled2026.01 | 70.2 | 20 | — | |
| Qwen3 1.7B BaseClassifier=Self-labeled2026.01 | 69.4 | — | — | |
| Qwen3 1.7B BaseClassifier=Majority-labeled2026.01 | 69.4 | — | — | |
| Mistral 7B Inst v0.3Classifier=Self-labeled2026.01 | 68.9 | 30 | — | |
| Mistral 7B Inst v0.3Classifier=Majority-labeled2026.01 | 68.9 | 30 | — | |
| Llama 3.2 3B InstClassifier=Majority-labeled2026.01 | 68.1 | 40 | — | |
| Llama 3.2 3B InstClassifier=Self-labeled2026.01 | 68 | 30 | — | |
| Mistral Nemo Base 2407Classifier=Majority-labeled2026.01 | 64.9 | 10 | — | |
| Mistral Nemo Base 2407Classifier=Majority-labeled2026.01 | 64.9 | 10 | — | |
| Mistral Nemo Base 2407Classifier=Self-labeled2026.01 | 64.7 | -20 | — | |
| Mistral Nemo Base 2407Classifier=Self-labeled2026.01 | 64.7 | -20 | — | |
| Llama 3.1 8BClassifier=Self-labeled2026.01 | 62.7 | 0 | — | |
| Llama 3.1 8BClassifier=Self-labeled2026.01 | 62.7 | 0 | — | |
| Llama 3.1 8BClassifier=Majority-labeled2026.01 | 62.6 | -10 | — | |
| Llama 3.1 8BClassifier=Majority-labeled2026.01 | 62.6 | -10 | — | |
| Qwen3 1.7BClassifier=Self-labeled2026.01 | 62.5 | — | — | |
| Qwen3 1.7BClassifier=Majority-labeled2026.01 | 62.5 | — | — | |
| Mistral 7B v0.3Classifier=Majority-labeled2026.01 | 61.1 | -10 | — | |
| Mistral 7B v0.3Classifier=Majority-labeled2026.01 | 61.1 | -10 | — | |
| Mistral 7B v0.3Classifier=Self-labeled2026.01 | 60.8 | -30 | — | |
| Mistral 7B v0.3Classifier=Self-labeled2026.01 | 60.8 | -30 | — | |
| Llama 3.2 3BClassifier=Self-labeled2026.01 | 59.9 | 10 | — | |
| Llama 3.2 3BClassifier=Majority-labeled2026.01 | 59.9 | 20 | — | |
| Qwen3 0.6B BaseClassifier=Majority-labeled2026.01 | 58 | — | — | |
| Qwen3 0.6B BaseClassifier=Self-labeled2026.01 | 57.9 | — | — | |
| Llama 3.2 1B InstClassifier=Self-labeled2026.01 | 55.4 | 40 | — | |
| Llama 3.2 1B InstClassifier=Majority-labeled2026.01 | 55.1 | 10 | — | |
| Qwen3 0.6BClassifier=Self-labeled2026.01 | 52 | -30 | — | |
| Qwen3 0.6BClassifier=Majority-labeled2026.01 | 52 | -30 | — | |
| Qwen3 0.6BClassifier=Self-labeled2026.01 | 52 | — | — | |
| Qwen3 0.6BClassifier=Majority-labeled2026.01 | 52 | — | — | |
| Llama 3.2 1BClassifier=Self-labeled2026.01 | 48.5 | 50 | — | |
| Llama 3.2 1BClassifier=Self-labeled2026.01 | 48.5 | 50 | — | |
| Llama 3.2 1BClassifier=Majority-labeled2026.01 | 48.1 | 10 | — | |
| Llama 3.2 1BClassifier=Majority-labeled2026.01 | 48.1 | 10 | — | |
| DCLMPre-training Curation Strategy=Human Discovered, Model Parameters=3B, Training Token Budget=500B tokens2026.03 | 47.51 | — | — | |
| TRIMBackbone=LLAMA-3.2-1B, Coreset size (%)=5%, Fine-tuning protocol=Coreset fine-tuning, Evaluation protocol=EleutherAI lm-evaluation-harness2025.10 | 46.26 | — | — | |
| LLaMa3.1-70B2026.03 | 46.21 | — | — | |
| Full-data Fine-tuningBackbone=LLAMA-3.2-1B, Coreset size (%)=100%, Fine-tuning protocol=Full, Evaluation protocol=EleutherAI lm-evaluation-harness2025.10 | 45.04 | — | — | |
| Fineweb-EduPre-training Curation Strategy=Human Discovered, Model Parameters=3B, Training Token Budget=500B tokens2026.03 | 45.02 | — | — | |
| MoRTBackbone Scale=1.15B Scale, Cfg=4x, # Params=1.15B+14.94B, Evaluation Protocol=Zero/Few-shot2026.04 | 44.7 | — | — | |
| LESSBackbone=LLAMA-3.2-1B, Coreset size (%)=5%, Fine-tuning protocol=Coreset fine-tuning, Evaluation protocol=EleutherAI lm-evaluation-harness2025.10 | 44.52 | — | — |