Intrinsic Reasoning on GNLI
0.843Spearman CorrelationLlama-3-Instruct
Evaluation Results
| Method | Links | |
|---|---|---|
| Llama-3-InstructEvaluation Protocol=Probe, Size=14B2025.05 | 0.843 | |
| Always Tell Me The OddsBackbone=Qwen2.5-14B-Instruct, Data Augmentation=+Syn2025.05 | 0.838 | |
| Always Tell Me The OddsBackbone=Qwen2.5-14B-Instruct, Data Augmentation=+Syn, Training Strategy=+R2025.05 | 0.82 | |
| Always Tell Me The OddsBackbone=Qwen2.5-14B-Instruct2025.05 | 0.814 | |
| Always Tell Me The OddsBackbone=Qwen2.5-7B-Instruct2025.05 | 0.811 | |
| GPT-4oEvaluation Protocol=0-Shot2025.05 | 0.796 | |
| Always Tell Me The OddsBackbone=Qwen2.5-8B-Instruct2025.05 | 0.789 | |
| DeepSeek-R1-Distill-Qwen-32BEvaluation Protocol=0-Shot2025.05 | 0.755 | |
| RoBERTa-LType=Encoder2025.05 | 0.586 |