Mathematical Reasoning on GSM8K (test) (Accuracy, Token Count, ACU)
96.5AccuracyO1-Pruner (Luo et al., 2025a)
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| O1-Pruner (Luo et al., 2025a)Train Set=PRM12K2025.02 | 96.5 | 534 | 0.56 | |
| CoT-Valve+P - MixChain-ZTrain Set=GSM8K, Backbone=QwQ-32B-Preview, Config Variant=Longer2025.02 | 96.1 | 317.1 | 0.95 | |
| Qwen2.5-Math-72B-InstructBackbone=Qwen2.5-Math-72B2025.02 | 95.8 | 312.1 | 0.43 | |
| O1-Pruner (Luo et al., 2025a) - SFTTrain Set=PRM12K, Protocol=SFT2025.02 | 95.7 | 717 | 0.42 | |
| Llama-3.1-405B-InstructBackbone=Llama-3.1-405B2025.02 | 95.6 | 186.7 | 0.13 | |
| Prompt (Ding et al., 2024)Control Method=Prompt-based2025.02 | 95.5 | 617.7 | 0.48 | |
| CoT-Valve+P - MixChain-ZTrain Set=PRM12K, Backbone=QwQ-32B-Preview2025.02 | 95.4 | 288.5 | 1.03 | |
| QwQ-32B-PreviewBackbone=QwQ-32B2025.02 | 95.1 | 741.1 | 0.4 | |
| CoT-Valve+P - MixChain-ZTrain Set=GSM8K, Backbone=QwQ-32B-Preview, Config Variant=Shorter2025.02 | 94.9 | 225.5 | 1.32 | |
| Overthink (Chen et al., 2024) - SFTTrain Set=PRM12K, Protocol=SFT2025.02 | 94.8 | 749.5 | 0.4 | |
| Overthink (Chen et al., 2024) - SimPOTrain Set=PRM12K, Protocol=SimPO2025.02 | 94.8 | 326.2 | 0.91 | |
| CoT-Valve++ - MixChain-CTrain Set=GSM8K, Backbone=QwQ-32B-Preview2025.02 | 94.4 | 276.3 | 1.07 | |
| CoT-Valve Ground-TruthTrain Set=GSM8K, Backbone=QwQ-32B-Preview2025.02 | 94 | 352.8 | 0.83 | |
| Prompt (Han et al., 2024)Control Method=Prompt-based2025.02 | 93.6 | 355.5 | 0.82 | |
| Qwen2.5-32B-InstructBackbone=Qwen2.5-32B2025.02 | 93.1 | 269.3 | 1.09 | |
| Llama-3.3-70B-InstructBackbone=Llama-3.3-70B2025.02 | 92.6 | 235.4 | 0.56 | |
| CoT-Valvetraining_dataset=QwQ Distill2025.02 | 77.5 | 569.8 | 1.7 | |
| CoT-Valve+Ptraining_dataset=MixChain-Z2025.02 | 77.1 | 371.2 | 2.596 | |
| SFT-LORAtraining_dataset=QwQ Distill2025.02 | 76.3 | 644.8 | 1.479 | |
| CoT-Valvetraining_dataset=MixChain-Z, strategy=Solution 12025.02 | 75.7 | 264.1 | 3.583 | |
| SFT-LORAtraining_dataset=GSM8k2025.02 | 59 | 191.9 | 3.843 | |
| CoT-Valvetraining=MixChain-Z, strategy=Solution 12025.02 | 58.9 | 275.4 | 21.387 | |
| SFTtraining=MixChain-Z, strategy=Solution 12025.02 | 57 | 288.4 | 19.764 | |
| LLaMA-3.1-8Bevaluation_protocol=8-shot2025.02 | 56.9 | 282.1 | 2.521 | |
| CoT-Valve+Ptraining=MixChain-Z2025.02 | 55.8 | 291 | 19.175 | |
| CoT-Valvetraining=QwQ Distill2025.02 | 55.5 | 267 | 20.786 | |
| SFTtraining=QwQ Distill2025.02 | 52.7 | 759.3 | 6.941 | |
| Prompt2025.02 | 46.7 | 209.9 | 22.249 | |
| SFT-Full Finetunetrain_set=GSM8K2025.02 | 46.1 | 139.4 | 33.07 | |
| LLaMA-3.2-1B-Instructshot=8-shot2025.02 | 45.9 | 104.3 | 44.008 | |
| LLaMA-3.2-1B-Instructshot=0-shot2025.02 | 45.9 | 199.8 | 22.973 | |
| SFTtrain_set=GSM8K2025.02 | 43.8 | 137.7 | 31.808 | |
| LLaMA-3.1-8Bevaluation_protocol=0-shot2025.02 | 15.7 | 915 | 0.214 |