Scientific Reasoning on GPQA Diamond (pass@1)
69.5pass@1SPLA
Evaluation Results
| Method | Links | |
|---|---|---|
| SPLAModel Size=14B, Temperature=0.6, Max output length=32k, Training context length=32k, Training tokens=1.2T, Sampling strategy=pass@12026.01 | 69.5 | |
| SPAModel Size=14B, Temperature=0.6, Max output length=32k, Training context length=32k, Training tokens=1.2T, Sampling strategy=pass@12026.01 | 69.2 | |
| InfLLM-v2Model Size=14B, Temperature=0.6, Max output length=32k, Training context length=32k, Training tokens=1.2T, Sampling strategy=pass@12026.01 | 68.7 | |
| Dense AttentionModel Size=14B, Temperature=0.6, Max output length=32k, Training context length=32k, Training tokens=1.2T, Sampling strategy=pass@12026.01 | 68.5 | |
| NSAModel Size=14B, Temperature=0.6, Max output length=32k, Training context length=32k, Training tokens=1.2T, Sampling strategy=pass@12026.01 | 59.6 | |
| Continual LUFFYBackbone=Qwen2.5-Math-7B, RLVR Setup=Continual RLVR2025.10 | 49 | |
| On-Policy (Continual)Backbone=Qwen2.5-Math-7B, RLVR Setup=Continual RLVR2025.10 | 47 | |
| ExGRPO (Continual)Backbone=Qwen2.5-Math-7B, RLVR Setup=Continual RLVR2025.10 | 42.4 | |
| GPG-ZeroBackbone=Qwen2.5-Math-7B, RLVR Setup=Previous Zero RLVR2025.10 | 40.4 | |
| LUFFYBackbone=Qwen2.5-Math-7B, RLVR Setup=Continual RLVR2025.10 | 39.9 | |
| On-PolicyBackbone=Qwen2.5-Math-7B, RLVR Setup=Zero RLVR2025.10 | 37.4 | |
| ExGRPOBackbone=Qwen2.5-Math-7B, RLVR Setup=Zero RLVR2025.10 | 37.4 | |
| Qwen-InstructBackbone=Qwen2.5-Math-7B2025.10 | 24.7 | |
| SFTBackbone=Qwen2.5-Math-7B, RLVR Setup=Off-policy Learning2025.10 | 24.7 | |
| RePO-ZeroBackbone=Qwen2.5-Math-7B, RLVR Setup=Previous Zero RLVR2025.10 | 24.2 | |
| SFT+RLBackbone=Qwen2.5-Math-7B, RLVR Setup=Off-policy Learning2025.10 | 24.2 | |
| Oat-ZeroBackbone=Qwen2.5-Math-7B, RLVR Setup=Previous Zero RLVR2025.10 | 23.7 | |
| PRIME-ZeroBackbone=Qwen2.5-Math-7B, RLVR Setup=Previous Zero RLVR2025.10 | 18.2 | |
| Qwen-BaseBackbone=Qwen2.5-Math-7B2025.10 | 11.1 |