Mathematical Reasoning on GSM8K (test) (Accuracy, Avg.)
85.81AccuracySoftCoT
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| SoftCoTBackbone=Qwen2.5-7B-Instruct, Training Objective=Language Modeling Objective, Evaluation Protocol=Projection Module Trained2025.02 | 85.81 | 75.06 | |
| Zero-Shot Assist-CoTBackbone=Qwen2.5-7B-Instruct, Training Objective=None, Evaluation Protocol=Zero-Shot2025.02 | 84.85 | 71.64 | |
| Zero-Shot CoT-UnkBackbone=Qwen2.5-7B-Instruct, Training Objective=None, Evaluation Protocol=Zero-Shot2025.02 | 84.12 | 70.6 | |
| Zero-Shot CoTBackbone=Qwen2.5-7B-Instruct, Training Objective=None, Evaluation Protocol=Zero-Shot2025.02 | 83.7 | 70.29 | |
| CoconutBackbone=Qwen2.5-7B-Instruct, Training Objective=Language Modeling Objective, Evaluation Protocol=Finetuned2025.02 | 82.49 | — | |
| LoRA Fine-TuningBackbone=Qwen2.5-7B-Instruct, Training Objective=Language Modeling Objective, Evaluation Protocol=Fine-Tuning2025.02 | 81.8 | — |