Mathematical Reasoning on MATH (Strict Accuracy)
38.8Strict AccuracyDream-7B-Base + DyStruct
Evaluation Results
| Method | Links | |
|---|---|---|
| Dream-7B-Base + DyStructBackbone=Dream-7B, Decoding method=DyStruct, Few-shot examples=4, Unmasking iterations=256, Generation limit=2562026.05 | 38.8 | |
| Dream-7B-Base + DAEDALBackbone=Dream-7B, Decoding method=DAEDAL, Few-shot examples=4, Unmasking iterations=256, Generation limit=2562026.05 | 38.6 | |
| Dream-7B-BaseBackbone=Dream-7B, Decoding method=Base, Few-shot examples=4, Unmasking iterations=256, Generation limit=2562026.05 | 38.2 | |
| LESSBackbone=LLAMA-2-7B, Coreset Method=LESS, Coreset Fraction=50%2025.10 | 32.2 | |
| S2LBackbone=LLAMA-2-7B, Coreset Method=S2L, Coreset Fraction=50%2025.10 | 32.1 | |
| LLaDA-8B-Base + DyStructBackbone=LLaDA-8B, Decoding method=DyStruct, Few-shot examples=4, Unmasking iterations=256, Generation limit=2562026.05 | 31.4 | |
| TRIMBackbone=LLAMA-2-7B, Coreset Method=TRIM, Coreset Fraction=50%2025.10 | 31.4 | |
| LLaDA-8B-Base + DAEDALBackbone=LLaDA-8B, Decoding method=DAEDAL, Few-shot examples=4, Unmasking iterations=256, Generation limit=2562026.05 | 31.2 | |
| LLaDA-8B-BaseBackbone=LLaDA-8B, Decoding method=Base, Few-shot examples=4, Unmasking iterations=256, Generation limit=2562026.05 | 30.5 | |
| Full-data Fine-tuningBackbone=LLAMA-2-7B, Coreset Method=Full-data, Coreset Fraction=100%2025.10 | 29.45 | |
| TAGCOSBackbone=LLAMA-2-7B, Coreset Method=TAGCOS, Coreset Fraction=50%2025.10 | 28.28 | |
| RandomBackbone=LLAMA-2-7B, Coreset Method=Random, Coreset Fraction=50%2025.10 | 18.22 | |
| Pretrained (no Fine-tuning)Backbone=LLAMA-2-7B, Coreset Method=Pretrained2025.10 | 4.2 |