Natural Language Inference on MNLI (MNLI Score)
87.4MNLI AccuracyDiSP
Evaluation Results
| Method | Links | |
|---|---|---|
| DiSPModel=Qwen2.5-7B, Setting=Few-shot2026.05 | 87.4 | |
| BM25Model=Qwen2.5-7B, Setting=Few-shot2026.05 | 86.1 | |
| RandomModel=Qwen2.5-7B, Setting=Few-shot2026.05 | 85 | |
| Se2Model=Qwen2.5-7B, Setting=Few-shot2026.05 | 82.1 | |
| SeDPOModel=Qwen2.5-7B, Setting=Few-shot2026.05 | 80.8 | |
| UpriseModel=Qwen2.5-7B, Setting=Few-shot2026.05 | 80.1 | |
| Zero-shotModel=Qwen2.5-7B, Setting=Zero-shot2026.05 | 76.8 | |
| ConMeZOBackbone=RoBERTa-Large, Evaluation Protocol=prompt-conditioned, Time Constraint=equal wall-clock time2025.11 | 73.3 | |
| LOZOBackbone=RoBERTa-Large, Evaluation Protocol=prompt-conditioned, Time Constraint=equal wall-clock time2025.11 | 70.5 | |
| LOZO-MBackbone=RoBERTa-Large, Evaluation Protocol=prompt-conditioned, Time Constraint=equal wall-clock time2025.11 | 69.5 | |
| DiSPModel=LLaMA3-8B, Setting=Few-shot2026.05 | 69.5 | |
| RandomModel=LLaMA3-8B, Setting=Few-shot2026.05 | 66.9 | |
| BM25Model=LLaMA3-8B, Setting=Few-shot2026.05 | 65.6 | |
| SeDPOModel=LLaMA3-8B, Setting=Few-shot2026.05 | 61.6 | |
| Se2Model=LLaMA3-8B, Setting=Few-shot2026.05 | 60.9 | |
| UpriseModel=LLaMA3-8B, Setting=Few-shot2026.05 | 58.9 | |
| Zero-shotModel=LLaMA3-8B, Setting=Zero-shot2026.05 | 52.3 | |
| Base modelRuntime (min)=N/A2026.04 | 42.8 | |
| SIFTRuntime (min)=215.62026.04 | 40.5 | |
| BLURRuntime (min)=114.72026.04 | 37.6 | |
| AdamWRuntime (min)=110.32026.04 | 36.2 | |
| MuonRuntime (min)=191.22026.04 | 32.8 | |
| POMERuntime (min)=115.42026.04 | 32.5 |