Logical Reasoning on ProntoQA (val)
98.01AccuracyCoT2
Evaluation Results
| Method | Links | |
|---|---|---|
| CoT2Backbone=4-layer, 4-head GPT2 (d=32), Budget=Full2025.05 | 98.01 | |
| COCONUTBackbone=4-layer, 4-head GPT2 (d=32)2025.05 | 96.94 | |
| Discrete CoTBackbone=4-layer, 4-head GPT2 (d=32), Budget=B=12025.05 | 82.47 | |
| Discrete no-CoTBackbone=4-layer, 4-head GPT2 (d=32)2025.05 | 73.65 |