Accuracy on Natural Language Inference on CB
98.2AccuracyFADS-ICL
Evaluation Results
| Method | Links | |
|---|---|---|
| FADS-ICLBackbone=Llama-2 70B, Shots=128-shots2024.05 | 98.2 | |
| FADS-ICLBackbone=Llama-2 13B, Shots=128-shots2024.05 | 96.1 | |
| Individual2024.05 | 95.8 | |
| Traditional MTL2024.05 | 95.8 | |
| ICLBackbone=Llama-2 70B, Shots=128-shots2024.05 | 95.4 | |
| FADS-ICLBackbone=Llama-1 7B, Shots=128-shots2024.05 | 95 | |
| kNN-promptingBackbone=Llama-2 70B, Shots=128-shots2024.05 | 94.3 | |
| FADS-ICLBackbone=Llama-1 30B, Shots=128-shots2024.05 | 92.6 | |
| FADS-ICLBackbone=Llama-1 13B, Shots=128-shots2024.05 | 92.3 | |
| kNN-promptingBackbone=Llama-2 13B, Shots=128-shots2024.05 | 91.4 | |
| FADS-ICLBackbone=Llama-2 7B, Shots=128-shots2024.05 | 90.7 | |
| Falcon-180Bprotocol=1-shot2023.11 | 89.3 | |
| kNN-promptingBackbone=Llama-1 13B, Shots=128-shots2024.05 | 88.9 | |
| kNN-promptingBackbone=Llama-1 30B, Shots=128-shots2024.05 | 88.9 | |
| kNN-promptBackbone=Llama-2 70B, Shots=128-shots2024.05 | 87.9 | |
| PaLM 2-Lprompting=1-shot2023.05 | 87.5 | |
| PaLM-2 Lprotocol=1-shot2023.11 | 87.5 | |
| Ties-MergingValidation=false2024.05 | 87.5 | |
| EMR-MERGINGValidation=false2024.05 | 87.5 | |
| kNN-promptBackbone=Llama-2 13B, Shots=128-shots2024.05 | 86.1 | |
| FADS-ICLNumber of training shots (m)=32, LLM scale=1.5B2024.05 | 85.9 | |
| ColD-Fusion2022.12 | 85 | |
| kNN-promptingBackbone=Llama-1 7B, Shots=128-shots2024.05 | 84.3 | |
| UCPOFModel=Llama-3-8B-Instruct2026.02 | 83.93 | |
| gold shot2026.02 | 83.93 | |
| PaLMprompting=1-shot2023.05 | 83.9 | |
| PaLMprotocol=1-shot2023.11 | 83.9 | |
| Fisher MergingValidation=true2024.05 | 83.3 | |
| Task ArithmeticValidation=true2024.05 | 83.3 | |
| Ties-MergingValidation=true2024.05 | 83.3 | |
| FADS-ICLNumber of training shots (m)=128, LLM scale=1.5B2024.05 | 83.2 | |
| FADS-ICLBackbone=GPT-2 1.5B, Shots=128-shots2024.05 | 83.2 | |
| FADS-ICLcandidate pool size (m)=1282024.05 | 83.2 | |
| FADS-ICLcandidate pool size (m)=2562024.05 | 83.2 | |
| Multitask2022.12 | 82.86 | |
| Full RAGModel=Llama-3-8B-Instruct2026.02 | 82.14 | |
| UCPOF2026.02 | 82.14 | |
| PaLM 2-Sprompting=1-shot2023.05 | 82.1 | |
| PaLM-2 Sprotocol=1-shot2023.11 | 82.1 | |
| PaLM 2-Mprompting=1-shot2023.05 | 80.4 | |
| PaLM-2 Mprotocol=1-shot2023.11 | 80.4 | |
| kNN-promptBackbone=Llama-1 7B, Shots=128-shots2024.05 | 80.4 | |
| MUPPET2022.12 | 80.36 | |
| FADS-ICLBackbone=GPT-2 0.8B, Shots=128-shots2024.05 | 80 | |
| kNN-promptBackbone=Llama-2 7B, Shots=128-shots2024.05 | 80 | |
| kNN-promptBackbone=Llama-1 13B, Shots=128-shots2024.05 | 80 | |
| ICLBackbone=Llama-1 30B, Shots=128-shots2024.05 | 80 | |
| kNN-promptBackbone=Llama-1 30B, Shots=128-shots2024.05 | 80 | |
| Task ArithmeticValidation=false2024.05 | 79.2 | |
| FADS-ICLNumber of training shots (m)=64, LLM scale=1.5B2024.05 | 78.6 | |
| BERT Large FTBackbone=BERT Large, Shots=N/A (Fine-tuned)2024.05 | 78.6 | |
| FADS-ICLcandidate pool size (m)=642024.05 | 78.6 | |
| baseline2026.02 | 78.57 | |
| kNN-promptingBackbone=Llama-2 7B, Shots=128-shots2024.05 | 77.1 | |
| FO (Adam)Base Model=LLaMA-3.2-1B, Optimization Regime=First-Order2025.10 | 75 | |
| ICLBackbone=Llama-1 7B, Shots=128-shots2024.05 | 74.6 | |
| SimCSEcandidate pool size (m)=1282024.05 | 73.2 | |
| Trans-Encodercandidate pool size (m)=1282024.05 | 73.2 | |
| SimCSEcandidate pool size (m)=2562024.05 | 73.2 | |
| Trans-Encodercandidate pool size (m)=2562024.05 | 73.2 | |
| ZO Fine-tunerBase Model=LLaMA-3.2-1B, Optimization Regime=Zeroth-Order2025.10 | 73 | |
| FADS-ICLcandidate pool size (m)=322024.05 | 72.1 | |
| kNN-promptingNumber of training shots (m)=32, LLM scale=1.5B2024.05 | 71.8 | |
| SimCSEcandidate pool size (m)=642024.05 | 71.4 | |
| BM25candidate pool size (m)=1282024.05 | 71.4 | |
| BM25candidate pool size (m)=2562024.05 | 71.4 | |
| GPT Large FTBackbone=GPT Large, Shots=N/A (Fine-tuned)2024.05 | 70 | |
| ICLBackbone=Llama-2 7B, Shots=128-shots2024.05 | 70 | |
| MeZOBase Model=LLaMA-3.2-1B, Optimization Regime=Zeroth-Order2025.10 | 70 | |
| SBERTcandidate pool size (m)=1282024.05 | 69.6 | |
| SBERTcandidate pool size (m)=2562024.05 | 69.6 | |
| Trans-Encodercandidate pool size (m)=642024.05 | 69.3 | |
| BM25candidate pool size (m)=642024.05 | 68.6 | |
| kNN-promptingBackbone=GPT-2 0.8B, Shots=128-shots2024.05 | 67.5 | |
| SBERTcandidate pool size (m)=642024.05 | 66.8 | |
| SimCSEcandidate pool size (m)=322024.05 | 66.4 | |
| kNN-promptingNumber of training shots (m)=128, LLM scale=1.5B2024.05 | 65 | |
| kNN-promptingBackbone=GPT-2 1.5B, Shots=128-shots2024.05 | 65 | |
| BM25candidate pool size (m)=322024.05 | 65 | |
| GPT-3Evaluation Protocol=1-shot2023.03 | 64.3 | |
| Finetune2022.12 | 64.29 | |
| baselineModel=Llama-3-8B-Instruct2026.02 | 64.29 | |
| FADS-ICLNumber of training shots (m)=16, LLM scale=1.5B2024.05 | 63.9 | |
| SBERTcandidate pool size (m)=322024.05 | 62.9 | |
| ICLBackbone=Llama-1 13B, Shots=128-shots2024.05 | 62.5 | |
| ICLBackbone=Llama-2 13B, Shots=128-shots2024.05 | 62.5 | |
| gold shotModel=Llama-3-8B-Instruct2026.02 | 62.5 | |
| kNN-promptingNumber of training shots (m)=64, LLM scale=1.5B2024.05 | 61.4 | |
| Trans-Encodercandidate pool size (m)=322024.05 | 61.4 | |
| In-Context Learning2024.05 | 60.7 | |
| RegMeanValidation=true2024.05 | 58.3 | |
| Weight AveragingValidation=false2024.05 | 58.3 | |
| ICLNumber of training shots (m)=4, LLM scale=1.5B2024.05 | 57.9 | |
| ICLNumber of training shots (m)=8, LLM scale=1.5B2024.05 | 57.9 | |
| ICLNumber of training shots (m)=16, LLM scale=1.5B2024.05 | 57.9 | |
| ICLNumber of training shots (m)=32, LLM scale=1.5B2024.05 | 57.9 | |
| ICLNumber of training shots (m)=64, LLM scale=1.5B2024.05 | 57.9 | |
| ICLNumber of training shots (m)=128, LLM scale=1.5B2024.05 | 57.9 | |
| ICLBackbone=GPT-2 1.5B, Shots=128-shots2024.05 | 57.9 | |
| MaskProBase Model=LLaMA-2-7B, Sparsity=2:42025.06 | 57.14 |