Reasoning on MATH (Success Score)
89.8Success Score (MATH)SInternal
Evaluation Results
| Method | Links | |
|---|---|---|
| SInternalModel Backbone=DeepSeek-R1-Distill-Qwen-7B, RL Training (GRPO)=true2026.05 | 89.8 | |
| BaseModel Backbone=DeepSeek-R1-Distill-Qwen-14B, RL Training (GRPO)=true2026.05 | 89.8 | |
| SafeChainModel Backbone=DeepSeek-R1-Distill-Qwen-14B, RL Training (GRPO)=true2026.05 | 89.8 | |
| SInternalModel Backbone=DeepSeek-R1-Distill-Qwen-14B, RL Training (GRPO)=true2026.05 | 89.6 | |
| STAR-1Model Backbone=DeepSeek-R1-Distill-Qwen-14B, RL Training (GRPO)=true2026.05 | 89.4 | |
| STAR-1Model Backbone=DeepSeek-R1-Distill-Qwen-7B, RL Training (GRPO)=true2026.05 | 89.2 | |
| BaseModel Backbone=DeepSeek-R1-Distill-Qwen-7B, RL Training (GRPO)=true2026.05 | 88.8 | |
| SafeChainModel Backbone=DeepSeek-R1-Distill-Qwen-7B, RL Training (GRPO)=true2026.05 | 86.6 | |
| SInternalModel Backbone=DeepSeek-R1-Distill-Llama-8B, RL Training (GRPO)=true2026.05 | 84.2 | |
| STAR-1Model Backbone=DeepSeek-R1-Distill-Llama-8B, RL Training (GRPO)=true2026.05 | 83.8 | |
| BaseModel Backbone=DeepSeek-R1-Distill-Llama-8B, RL Training (GRPO)=true2026.05 | 81.8 | |
| SafeChainModel Backbone=DeepSeek-R1-Distill-Llama-8B, RL Training (GRPO)=true2026.05 | 80 | |
| Qwen-2.5-7BTurkish Support=Advanced, Tokenization Efficiency=High, License Type=Apache 2.02026.01 | 75.5 | |
| Llama-3.1-8BTurkish Support=Limited, Tokenization Efficiency=Low, License Type=Llama Community2026.01 | 51.9 | |
| Mistral-7B-v0.3Turkish Support=Moderate, Tokenization Efficiency=Moderate, License Type=Apache 2.02026.01 | 38.2 |