Social Interaction Question Answering on SIQA
86.9AccuracyHuman Performance
Evaluation Results
| Method | Links | |
|---|---|---|
| Human Performance2024.01 | 86.9 | |
| T5Model Architecture=T52020.05 | 81.4 | |
| UNIFIEDQAModel Architecture=T52020.05 | 81.4 | |
| LoRA-RITEModel=3B2026.03 | 80.25 | |
| Stable-LoRAModel=3B2026.03 | 80.25 | |
| DeBERTa-v3-L (Supervised)Evaluation Protocol=Supervised2024.01 | 80.1 | |
| VERA-T5 (Multitask Supervised)Evaluation Protocol=Supervised2024.01 | 80.1 | |
| LoRA+Model=3B2026.03 | 79.94 | |
| AdamWModel=3B2026.03 | 79.89 | |
| RiemannModel=3B2026.03 | 79.89 | |
| RoBERTaModel Architecture=RoBERTa2020.05 | 78 | |
| Stable-LoRAModel=1.5B2026.03 | 77.64 | |
| LoRA+Model=1.5B2026.03 | 77.38 | |
| AdamWModel=1.5B2026.03 | 77.33 | |
| LoRA-RITEModel=1.5B2026.03 | 77.23 | |
| RiemannModel=1.5B2026.03 | 76.87 | |
| ROBERTa-L (Supervised)Evaluation Protocol=Supervised2024.01 | 76.6 | |
| BART-largeModel Architecture=BART-large2020.05 | 74 | |
| UnifiedQA-BARTModel Architecture=BART2020.05 | 73.2 | |
| Stable-LoRAModel=1B2026.03 | 72.26 | |
| AdamWModel=1B2026.03 | 71.7 | |
| LoRA-RITEModel=1B2026.03 | 71.6 | |
| LoRA+Model=1B2026.03 | 71.19 | |
| ChatGPT + Chain-of-thoughtVersion=gpt-3.5-turbo, Evaluation Protocol=Zero-shot2024.01 | 70.7 | |
| RiemannModel=1B2026.03 | 70.62 | |
| ChatGPT + Self-consistent chain-of-thoughtVersion=gpt-3.5-turbo, Evaluation Protocol=Zero-shot2024.01 | 69.7 | |
| ChatGPTVersion=gpt-3.5-turbo, Evaluation Protocol=Zero-shot2024.01 | 69.5 | |
| Stable-LoRAModel=0.5B2026.03 | 68.27 | |
| GPT-3.5Version=text-davinci-003, Evaluation Protocol=Zero-shot2024.01 | 68 | |
| LoRA+Model=0.5B2026.03 | 67.91 | |
| AdamWModel=0.5B2026.03 | 67.04 | |
| RiemannModel=0.5B2026.03 | 66.94 | |
| LoRA-RITEModel=0.5B2026.03 | 66.89 | |
| ZS-FusionCSKB=CSKG, Evaluation Protocol=Zero-shot2024.01 | 66.7 | |
| DeBERTa-v3-L (CANDLE Distilled)CSKB=CANDLE, Evaluation Protocol=Zero-shot2024.01 | 65.9 | |
| CAR-ROBERTa-LCSKB=ATOMIC, Evaluation Protocol=Zero-shot2024.01 | 64.8 | |
| BaselineBit width (b)=16.002025.05 | 64.8 | |
| STL-AdapterCSKB=CSKG, Evaluation Protocol=Zero-shot2024.01 | 64.7 | |
| STL-AdapterCSKB=ATOMIC, Evaluation Protocol=Zero-shot2024.01 | 64.4 | |
| CAR-ROBERTa-LCSKB=AbsATM, Evaluation Protocol=Zero-shot2024.01 | 64 | |
| CAR-DEBERTa-v3-LCSKB=AbsATM, Evaluation Protocol=Zero-shot2024.01 | 64 | |
| CAR-DEBERTa-v3-LCSKB=ATOMIC, Evaluation Protocol=Zero-shot2024.01 | 63.8 | |
| STL-PLMCSKB=ATOMIC, Evaluation Protocol=Zero-shot2024.01 | 63.2 | |
| ROBERTa-L (MR)CSKB=CSKG, Evaluation Protocol=Zero-shot2024.01 | 63.2 | |
| ROBERTa-L (MR)CSKB=ATOMIC, Evaluation Protocol=Zero-shot2024.01 | 63.1 | |
| DeBERTa-v3-L (MR)CSKB=ATOMIC, Evaluation Protocol=Zero-shot2024.01 | 62.1 | |
| MTLCSKB=CSKG, Evaluation Protocol=Zero-shot2024.01 | 61.9 | |
| SEEDBackbone=LLaMA2-7B, Selection Ratio=5%2026.05 | 61.8 | |
| ROBERTa-L (MR)CSKB=ATM10X, Evaluation Protocol=Zero-shot2024.01 | 61 | |
| DeBERTa-v3-L (MR)CSKB=ATM10X, Evaluation Protocol=Zero-shot2024.01 | 59.7 | |
| VERA-T5-xxl (CANDLE Distilled)CSKB=CANDLE, Evaluation Protocol=Zero-shot2024.01 | 59.4 | |
| FULLBackbone=LLaMA2-7B, Selection Ratio=100%2026.05 | 59.4 | |
| Tensor RMS + CBit width (b)=3.002025.05 | 58.6 | |
| Task-specific FTBase Model=Llama-3.1-8B, Fine-tuning Technique=LoRA, Merging Strategy=None2026.02 | 58.39 | |
| VERA-T5-xxlCSKB=ATM10X, Evaluation Protocol=Zero-shot2024.01 | 58.2 | |
| VERA-T5-xxlCSKB=AbsATM, Evaluation Protocol=Zero-shot2024.01 | 58.1 | |
| TSV-MBase Model=Llama-3.1-8B, Fine-tuning Technique=LoRA, Merging Strategy=TSV-M2026.02 | 58.03 | |
| VERA-T5-xxlCSKB=ATOMIC, Evaluation Protocol=Zero-shot2024.01 | 57.7 | |
| GPT-4Evaluation Protocol=Zero-shot2024.01 | 57 | |
| TIESBase Model=Llama-3.1-8B, Fine-tuning Technique=LoRA, Merging Strategy=TIES2026.02 | 56.96 | |
| Task-specific FTBase Model=Llama-3.1-8B, Fine-tuning Technique=OFT, Merging Strategy=None2026.02 | 56.81 | |
| MICOCSKB=ATOMIC, Evaluation Protocol=Zero-shot2024.01 | 56 | |
| RANDOMBackbone=LLaMA2-7B, Selection Ratio=5%2026.05 | 55.6 | |
| OrthoMergeBase Model=Llama-3.1-8B, Fine-tuning Technique=OFT, Merging Strategy=OrthoMerge2026.02 | 55.17 | |
| TSV-MBase Model=Llama-3.1-8B, Fine-tuning Technique=OFT, Merging Strategy=TSV-M2026.02 | 55.02 | |
| ROBERTa-L (MR)CSKB=CWWV, Evaluation Protocol=Zero-shot2024.01 | 54.8 | |
| ProQA2023.05 | 54.5 | |
| TABase Model=Llama-3.1-8B, Fine-tuning Technique=LoRA, Merging Strategy=TA2026.02 | 54.15 | |
| DAREBase Model=Llama-3.1-8B, Fine-tuning Technique=LoRA, Merging Strategy=DARE2026.02 | 53.84 | |
| TIESBase Model=Llama-3.1-8B, Fine-tuning Technique=OFT, Merging Strategy=TIES2026.02 | 53.74 | |
| ZS-FusionCSKB=CWWV, Evaluation Protocol=Zero-shot2024.01 | 53.7 | |
| Tensor RMS + SpBit width (b)=3.052025.05 | 53.6 | |
| KoCoModel Scale=1.6B Parameters2026.04 | 53.4 | |
| URL Prefix (MeCo)Model Scale=1.6B Parameters2026.04 | 52.9 | |
| Standard CPTModel Scale=1.6B Parameters2026.04 | 52.7 | |
| Muppet2023.05 | 52.63 | |
| Data SelectionModel Scale=1.6B Parameters2026.04 | 52.6 | |
| OLTQAConfiguration=Full2023.05 | 52.51 | |
| Hyperformer++2023.05 | 52.46 | |
| TABase Model=Llama-3.1-8B, Fine-tuning Technique=OFT, Merging Strategy=TA2026.02 | 52.2 | |
| DAREBase Model=Llama-3.1-8B, Fine-tuning Technique=OFT, Merging Strategy=DARE2026.02 | 52.1 | |
| MTLCSKB=CWWV, Evaluation Protocol=Zero-shot2024.01 | 52 | |
| OLTQAAblation=without Pk2023.05 | 51.76 | |
| OLTQAKnowledge Distillation=Static MKD2023.05 | 51.48 | |
| OLTQARetriever=EPR2023.05 | 51.23 | |
| OLTQAKnowledge Distillation=Back KD2023.05 | 50.72 | |
| OLTQARe-ranker=excluded2023.05 | 50.67 | |
| OLTQARetriever=BM252023.05 | 50.41 | |
| LLAMA2Parameters=13B, Evaluation Protocol=Zero-shot2024.01 | 50.3 | |
| UnifiedQA2023.05 | 50.15 | |
| COMET-DynGenCSKB=ATOMIC, Evaluation Protocol=Zero-shot2024.01 | 50.1 | |
| OLTQAKnowledge Distillation=without MKD2023.05 | 49.85 | |
| Distributed Lion-MaVoShots=3, Optimizer=D-Lion (MaVo), Model=LLaMA 7B, Instruction Finetuning=true2024.03 | 49.75 | |
| G-LionShots=3, Optimizer=Global Lion, Model=LLaMA 7B, Instruction Finetuning=true2024.03 | 49.64 | |
| OLTQAAblation=without Pm2023.05 | 49.64 | |
| G-AdamWShots=3, Optimizer=Global AdamW, Model=LLaMA 7B, Instruction Finetuning=true2024.03 | 48.97 | |
| LLAMA2Parameters=7B, Evaluation Protocol=Zero-shot2024.01 | 48.3 | |
| BASEBackbone=LLaMA2-7B, Selection Ratio=0%2026.05 | 48.3 | |
| Llama-3.1-8BBase Model=Llama-3.1-8B, Fine-tuning Technique=None, Merging Strategy=None2026.02 | 48.11 | |
| Distributed Lion-AvgShots=3, Optimizer=D-Lion (Avg), Model=LLaMA 7B, Instruction Finetuning=true2024.03 | 48.06 |