General Reasoning on Open-Platypus (test)
78.06AccuracyLatent-GRPO
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Latent-GRPOModel Scale=Qwen3-4B2026.01 | 78.06 | 1,632.52 | |
| LLM-as-JudgeModel Scale=Qwen3-4B2026.01 | 65.21 | 3,522.18 | |
| Latent-GRPOModel Scale=Qwen3-1.7B2026.01 | 64.82 | 1,218.92 | |
| LLM-as-JudgeModel Scale=Qwen3-1.7B2026.01 | 56.69 | 2,573.41 | |
| Latent-GRPOModel Scale=Qwen3-0.6B2026.01 | 40.56 | 1,079.27 | |
| LLM-as-JudgeModel Scale=Qwen3-0.6B2026.01 | 34.45 | 1,937.82 |