Intrinsic Reasoning on circa
0.747Spearman CorrelationAlways Tell Me The Odds
Evaluation Results
| Method | Links | |
|---|---|---|
| Always Tell Me The OddsBackbone=Qwen2.5-14B-Instruct, Data Augmentation=+Syn, Training Strategy=+R2025.05 | 0.747 | |
| GPT-4oEvaluation Protocol=0-Shot2025.05 | 0.734 | |
| DeepSeek-R1-Distill-Qwen-32BEvaluation Protocol=0-Shot2025.05 | 0.663 | |
| Always Tell Me The OddsBackbone=Qwen2.5-14B-Instruct, Data Augmentation=+Syn2025.05 | 0.564 | |
| Llama-3-InstructEvaluation Protocol=Probe, Size=14B2025.05 | 0.553 | |
| Always Tell Me The OddsBackbone=Qwen2.5-7B-Instruct2025.05 | 0.544 | |
| Always Tell Me The OddsBackbone=Qwen2.5-14B-Instruct2025.05 | 0.536 | |
| Always Tell Me The OddsBackbone=Qwen2.5-8B-Instruct2025.05 | 0.474 | |
| RoBERTa-LType=Encoder2025.05 | 0.43 |