Mathematical Reasoning on GSM8K-Aug (test)
57AccuracyTrained solver
Evaluation Results
| Method | Links | |
|---|---|---|
| Trained solver2026.02 | 57 | |
| DPVG2026.02 | 56.9 | |
| SpiralThinkerBackbone=Llama-3.2-1B, Fine-tuning method=LoRA, Decoding strategy=Greedy2025.11 | 56.56 | |
| Pause TokenBackbone=Llama-3.2-1B, Fine-tuning method=LoRA, Decoding strategy=Greedy2025.11 | 53.37 | |
| CODIBackbone=Llama-3.2-1B, Fine-tuning method=LoRA, Decoding strategy=Greedy2025.11 | 51.02 | |
| CoconutBackbone=Llama-3.2-1B, Fine-tuning method=LoRA, Decoding strategy=Greedy2025.11 | 49.85 | |
| Token AssortedBackbone=Llama-3.2-1B, Fine-tuning method=LoRA, Decoding strategy=Greedy2025.11 | 48.7 | |
| iCoT-SIBackbone=Llama-3.2-1B, Fine-tuning method=LoRA, Decoding strategy=Greedy2025.11 | 29.72 | |
| iCoT-KDBackbone=Llama-3.2-1B, Fine-tuning method=LoRA, Decoding strategy=Greedy2025.11 | 24.11 | |
| PVG2026.02 | 22.3 | |
| Base model2026.02 | 9.6 |