Logical Reasoning on BIG-Bench Hard (BBH)
86.51AccuracyStep-GRPO
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Step-GRPOArchitecture=DeepSeek-R1-Distill-Llama-8B2026.04 | 86.51 | 1,097 | 82.1 | |
| GRPO-SOPArchitecture=DeepSeek-R1-Distill-Llama-8B2026.04 | 86.46 | 1,084 | 81.1 | |
| GRPO-λArchitecture=DeepSeek-R1-Distill-Llama-8B2026.04 | 86.3 | 989 | 74 | |
| VanillaArchitecture=DeepSeek-R1-Distill-Llama-8B2026.04 | 86.14 | 1,337 | 100 | |
| GRPOArchitecture=DeepSeek-R1-Distill-Llama-8B2026.04 | 85.96 | 1,247 | 93.2 | |
| DEER-SFTArchitecture=DeepSeek-R1-Distill-Llama-8B2026.04 | 84.89 | 1,520 | 113.6 | |
| GRPO-LPArchitecture=DeepSeek-R1-Distill-Llama-8B2026.04 | 84.81 | 811 | 60.7 |