Comparative Reasoning on delta-SNLI
88.9AccuracyAlways Tell Me The Odds
Evaluation Results
| Method | Links | |
|---|---|---|
| Always Tell Me The OddsBackbone=Qwen2.5-14B-Instruct, Data Augmentation=+Syn, Training Strategy=+R2025.05 | 88.9 | |
| Always Tell Me The OddsBackbone=Qwen2.5-14B-Instruct, Data Augmentation=+Syn2025.05 | 86 | |
| Always Tell Me The OddsBackbone=Qwen2.5-14B-Instruct2025.05 | 85.1 | |
| Always Tell Me The OddsBackbone=Qwen2.5-7B-Instruct2025.05 | 84.4 | |
| Always Tell Me The OddsBackbone=Qwen2.5-8B-Instruct2025.05 | 83.3 | |
| Llama-3-InstructEvaluation Protocol=Probe, Size=14B2025.05 | 81.3 | |
| RoBERTa-LType=Encoder2025.05 | 77.9 | |
| DeepSeek-R1-Distill-Qwen-32BEvaluation Protocol=0-Shot2025.05 | 77.2 | |
| GPT-4oEvaluation Protocol=0-Shot2025.05 | 75.3 |