Mathematical Reasoning on Low-difficulty benchmark suite Average
58.32Pass@1Latent-GRPO
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Latent-GRPOBase Model=LLaMA-3.2-1B-Instruct, Reasoning Mode=Latent Reasoning (No Sampling)2026.04 | 58.32 | 21.2 | 7.86 | 4.44 | |
| GRPOBase Model=LLaMA-3.2-1B-Instruct, Reasoning Mode=Explicit Reasoning2026.04 | 57.98 | 94.2 | 5.28 | 1 | |
| SFTBase Model=LLaMA-3.2-1B-Instruct, Reasoning Mode=Explicit Reasoning2026.04 | 52.7 | 63.2 | — | — | |
| Soft-GRPOBase Model=LLaMA-3.2-1B-Instruct, Reasoning Mode=Latent Reasoning (No Sampling)2026.04 | 50.73 | 19.73 | 0.27 | 4.77 | |
| Latent-SFTBase Model=LLaMA-3.2-1B-Instruct, Reasoning Mode=Latent Reasoning (No Sampling)2026.04 | 50.46 | 19.78 | — | — |