Arithmetic Reasoning on Countdown (Accuracy)
63.1AccuracyGRPO + coupled rhythm credit
Evaluation Results
| Method | Links | |
|---|---|---|
| GRPO + coupled rhythm creditBackbone=Qwen3-4B-Base2025.10 | 63.1 | |
| GRPO + global-anchor creditBackbone=Qwen3-4B-Base2025.10 | 60.4 | |
| GRPO + local-chunk creditBackbone=Qwen3-4B-Base2025.10 | 59.9 | |
| GRPO + high-entropy creditBackbone=Qwen3-4B-Base2025.10 | 57.7 | |
| GRPO + random creditBackbone=Qwen3-4B-Base2025.10 | 55 | |
| GRPOBackbone=Qwen3-4B-Base2025.10 | 52.6 | |
| LIFT2Backbone=Dream-7B, H=22026.05 | 33.6 | |
| Instruct Model + GIFTShots=5-shot, Training Dataset=s1K, Tuning Method=LoRA Tuning2025.09 | 27.5 | |
| Instruct Model + GIFTShots=5-shot, Training Dataset=s1.1K, Tuning Method=LoRA Tuning2025.09 | 26 | |
| LIFT3Backbone=Dream-7B, H=32026.05 | 25.6 | |
| VanillaBackbone=Dream-7B2026.05 | 25 | |
| GIFTBackbone=Dream-7B2026.05 | 23.4 | |
| Instruct Model + SFTShots=5-shot, Training Dataset=s1K, Tuning Method=LoRA Tuning2025.09 | 23.2 | |
| CARTBackbone=Dream-7B2026.05 | 22.3 | |
| Instruct Model + SFTShots=5-shot, Training Dataset=s1.1K, Tuning Method=LoRA Tuning2025.09 | 21.7 | |
| InstructBackbone=Dream-7B2026.05 | 21.1 | |
| GIFTTraining Dataset=s1K, Tuning Method=LoRA Tuning, Base Model=LLaDA-8B-Instruct, Evaluation Protocol=0-shot2025.09 | 0.281 | |
| GIFTTraining Dataset=s1K-1.1, Tuning Method=LoRA Tuning, Base Model=LLaDA-8B-Instruct, Evaluation Protocol=0-shot2025.09 | 0.218 | |
| GIFTTraining Dataset=Tulu3-10k, Tuning Method=Full Parameter Tuning, Base Model=LLaDA-8B-Instruct, Evaluation Protocol=0-shot2025.09 | 0.213 | |
| SFTTraining Dataset=s1K, Tuning Method=LoRA Tuning, Base Model=LLaDA-8B-Instruct, Evaluation Protocol=0-shot2025.09 | 0.211 | |
| SFTTraining Dataset=s1K-1.1, Tuning Method=LoRA Tuning, Base Model=LLaDA-8B-Instruct, Evaluation Protocol=0-shot2025.09 | 0.207 | |
| GIFTTraining Dataset=openr1-3k, Tuning Method=LoRA Tuning, Base Model=LLaDA-8B-Instruct, Evaluation Protocol=0-shot2025.09 | 0.188 | |
| SFTTraining Dataset=Tulu3-10k, Tuning Method=Full Parameter Tuning, Base Model=LLaDA-8B-Instruct, Evaluation Protocol=0-shot2025.09 | 0.182 | |
| SFTTraining Dataset=openr1-3k, Tuning Method=LoRA Tuning, Base Model=LLaDA-8B-Instruct, Evaluation Protocol=0-shot2025.09 | 0.173 | |
| LLaDA-8B-InstructEvaluation Protocol=0-shot2025.09 | 0.17 |