Science Question Answering on GPQA (accuracy)
63.51AccuracySFT
Evaluation Results
| Method | Links | |
|---|---|---|
| SFTModel=DS-R1-Qwen-32B2026.04 | 63.51 | |
| D-QRELOModel=DS-R1-Qwen-32B2026.04 | 61.47 | |
| Delta-CoMeModel=DS-R1-Qwen-32B2026.04 | 58.46 | |
| SFTModel=DS-R1-Qwen-14B2026.04 | 56.94 | |
| Llama-InstructShots=0-shot2026.03 | 54.5 | |
| D-QRELOModel=DS-R1-Qwen-14B2026.04 | 53.27 | |
| Delta-CoMeModel=DS-R1-Qwen-14B2026.04 | 51.92 | |
| SVDModel=DS-R1-Qwen-32B2026.04 | 50.43 | |
| R1-Distill-Qwen-7BFT data=R1*, LoRA=-2026.03 | 49 | |
| Qwen-7BFT data=MoT, LoRA=-2026.03 | 48 | |
| Qwen-7BFT data=OT3 + MoT, LoRA=rk 1282026.03 | 48 | |
| Qwen-7BFT data=OT3, LoRA=-2026.03 | 47 | |
| Llama 3.1 InstructModel Scale=70B2025.04 | 46.7 | |
| SFTModel=DS-R1-LLaMA-8B2026.04 | 45.2 | |
| Qwen-7BFT data=OT3, LoRA=rk 642026.03 | 45 | |
| WandaModel=DS-R1-Qwen-32B2026.04 | 44.52 | |
| SVDModel=DS-R1-Qwen-14B2026.04 | 44.34 | |
| Qwen-7BFT data=OT3, LoRA=rk 1282026.03 | 43 | |
| MagnitudeModel=DS-R1-Qwen-32B2026.04 | 42.74 | |
| ParamΔModel Scale=70B2025.04 | 42.2 | |
| Llama 3 InstructModel Scale=70B2025.04 | 40.6 | |
| Qwen-3BFT data=OT3, LoRA=-2026.03 | 40 | |
| BitDeltaModel=DS-R1-Qwen-32B2026.04 | 39.43 | |
| WandaModel=DS-R1-Qwen-14B2026.04 | 39.04 | |
| Evo 8BShots=5, E2E Latency (s)=8.6, Inference Speed (tokens/s)=522026.02 | 38.4 | |
| MagnitudeModel=DS-R1-Qwen-14B2026.04 | 37.82 | |
| Qwen-3BFT data=MoT, LoRA=-2026.03 | 37 | |
| Qwen-7BFT data=-, LoRA=-2026.03 | 37 | |
| BackboneModel=DS-R1-Qwen-32B2026.04 | 36.36 | |
| Qwen2.5 7BShots=5, E2E Latency (s)=8.1, Inference Speed (tokens/s)=462026.02 | 36.1 | |
| HATified-SFTShots=0-shot2026.03 | 36 | |
| D-QRELOModel=DS-R1-LLaMA-8B2026.04 | 35.98 | |
| BitDeltaModel=DS-R1-Qwen-14B2026.04 | 35.35 | |
| M_RLP + PostPre-training Condition=Reinforcement Learning from Pretraining, Post-training=SFT + RLVR, Evaluation Protocol=@1[4]2025.09 | 34.97 | |
| RandomModel=DS-R1-Qwen-32B2026.04 | 34.93 | |
| M_RLP + PostPre-training Condition=Reinforcement Learning from Pretraining, Post-training=SFT + RLVR, Evaluation Protocol=Standard2025.09 | 33.33 | |
| Llama 3 InstructModel Scale=8B2025.04 | 32.8 | |
| BackboneModel=DS-R1-Qwen-14B2026.04 | 32.32 | |
| Qwen-3BFT data=OT3, LoRA=rk 1282026.03 | 32 | |
| M_base + PostPre-training Condition=Base, Post-training=SFT + RLVR, Evaluation Protocol=@1[4]2025.09 | 31.52 | |
| MagnitudeModel=DS-R1-LLaMA-8B2026.04 | 31.19 | |
| Qwen-3BFT data=-, LoRA=-2026.03 | 31 | |
| M_base + PostPre-training Condition=Base, Post-training=SFT + RLVR, Evaluation Protocol=Standard2025.09 | 30.93 | |
| WandaModel=DS-R1-LLaMA-8B2026.04 | 30.72 | |
| ParamΔModel Scale=8B2025.04 | 30.6 | |
| BitDeltaModel=DS-R1-LLaMA-8B2026.04 | 30.56 | |
| BD3-LM 7BShots=5, E2E Latency (s)=14.2, Inference Speed (tokens/s)=282026.02 | 30.5 | |
| RandomModel=DS-R1-Qwen-14B2026.04 | 30.39 | |
| SVDModel=DS-R1-LLaMA-8B2026.04 | 30.32 | |
| Delta-CoMeModel=DS-R1-LLaMA-8B2026.04 | 30.18 | |
| M_CPT + PostPre-training Condition=Continuous Pretraining, Post-training=SFT + RLVR, Evaluation Protocol=@1[4]2025.09 | 30.01 | |
| Llama 3.1 InstructModel Scale=8B2025.04 | 29.7 | |
| M_CPT + PostPre-training Condition=Continuous Pretraining, Post-training=SFT + RLVR, Evaluation Protocol=Standard2025.09 | 29.27 | |
| M_RLPPre-training Condition=Reinforcement Learning from Pretraining, Post-training=None, Evaluation Protocol=Standard2025.09 | 28.28 | |
| M_basePre-training Condition=Base, Post-training=None, Evaluation Protocol=@1[4]2025.09 | 27.52 | |
| M_RLPPre-training Condition=Reinforcement Learning from Pretraining, Post-training=None, Evaluation Protocol=@1[4]2025.09 | 27.02 | |
| LLaDA 8BShots=5, E2E Latency (s)=21.8, Inference Speed (tokens/s)=162026.02 | 26.4 | |
| M_CPTPre-training Condition=Continuous Pretraining, Post-training=None, Evaluation Protocol=Standard2025.09 | 26.26 | |
| LLaMA3 8BShots=5, E2E Latency (s)=7.4, Inference Speed (tokens/s)=582026.02 | 26.2 | |
| ARD 7BShots=5, E2E Latency (s)=32.5, Inference Speed (tokens/s)=122026.02 | 25.3 | |
| M_basePre-training Condition=Base, Post-training=None, Evaluation Protocol=Standard2025.09 | 25.25 | |
| MDLM 7BShots=5, E2E Latency (s)=18.9, Inference Speed (tokens/s)=222026.02 | 24.9 | |
| M_CPTPre-training Condition=Continuous Pretraining, Post-training=None, Evaluation Protocol=@1[4]2025.09 | 24.75 | |
| Llama 3.1 BaseModel Scale=70B2025.04 | 22.8 | |
| RandomModel=DS-R1-LLaMA-8B2026.04 | 22.22 | |
| BackboneModel=DS-R1-LLaMA-8B2026.04 | 20.13 | |
| Llama 3 BaseModel Scale=70B2025.04 | 14.3 | |
| Llama 3.1 BaseModel Scale=8B2025.04 | 6.3 | |
| Llama 3 BaseModel Scale=8B2025.04 | 5.1 |