Coding on HumanEval (Score)
96.3ScoreJT-Safe-V2-35B
Evaluation Results
| Method | Links | |
|---|---|---|
| JT-Safe-V2-35BParameters=35B2026.05 | 96.3 | |
| SOTA with Equivalent ParametersModel comparison=Equivalent Parameters2026.05 | 94.5 | |
| TinyR1-32B-Preview2025.03 | 86.4 | |
| Math Expert2025.03 | 85.5 | |
| Science Expert2025.03 | 85.5 | |
| Code Expert2025.03 | 85.5 | |
| Base Model2025.03 | 83.5 | |
| SkillWeave#Params=1.42B, Backbone=Llama-3.2-1B-Instruct2026.05 | 42.5 | |
| FuseChat3.0#Params=1.48B, Backbone=Llama-3.2-1B-Instruct2026.05 | 41.4 | |
| Twin-merging#Params=1.48B, Backbone=Llama-3.2-1B-Instruct2026.05 | 39.6 | |
| PEFT#Params=1.42B, Backbone=Llama-3.2-1B-Instruct2026.05 | 39.1 | |
| Llama3.2-1B-Instruct#Params=1.15B, Backbone=Llama-3.2-1B-Instruct2026.05 | 38.9 | |
| self-rewarding#Params=1.15B, Backbone=Llama-3.2-1B-Instruct2026.05 | 38 | |
| LLaDA-8BParadigm=Masked Diffusion, Training Tokens=2.3T, Training Data=Not Released, Evaluation protocol=Reported by prior work, Shots=02026.06 | 35.4 | |
| Llama 3-8BParadigm=AR, Training Tokens=15T, Training Data=Not Released, Evaluation protocol=Reported by prior work, Shots=02026.06 | 34.8 | |
| Sumi-7BParadigm=Uniform Diffusion, Training Tokens=1.5T, Training Data=Fully Released, Evaluation protocol=Evaluated under our protocol, Shots=02026.06 | 22.6 | |
| OLMo-7BParadigm=AR, Training Tokens=2.5T, Training Data=Fully Released, Evaluation protocol=Evaluated under our protocol, Shots=02026.06 | 13.4 | |
| Llama 2-7BParadigm=AR, Training Tokens=2T, Training Data=Not Released, Evaluation protocol=Evaluated under our protocol, Shots=02026.06 | 12.8 | |
| Falcon-7BParadigm=AR, Training Tokens=1.5T, Training Data=Partially Released, Evaluation protocol=Evaluated under our protocol, Shots=02026.06 | 0 |