Logical reasoning on ProsQA (test)
99.4AccuracySpiralThinker
Evaluation Results
| Method | Links | |
|---|---|---|
| SpiralThinkerBackbone=Llama-3.2-1B, Fine-tuning method=LoRA, Decoding strategy=Greedy2025.11 | 99.4 | |
| iCoT-SIBackbone=Llama-3.2-1B, Fine-tuning method=LoRA, Decoding strategy=Greedy2025.11 | 99 | |
| iCoT-KDBackbone=Llama-3.2-1B, Fine-tuning method=LoRA, Decoding strategy=Greedy2025.11 | 98 | |
| CoconutBackbone=Llama-3.2-1B, Fine-tuning method=LoRA, Decoding strategy=Greedy2025.11 | 97.8 | |
| Token AssortedBackbone=Llama-3.2-1B, Fine-tuning method=LoRA, Decoding strategy=Greedy2025.11 | 96.2 | |
| Pause TokenBackbone=Llama-3.2-1B, Fine-tuning method=LoRA, Decoding strategy=Greedy2025.11 | 95.8 | |
| CODIBackbone=Llama-3.2-1B, Fine-tuning method=LoRA, Decoding strategy=Greedy2025.11 | 80.8 |