Code Generation and Functional Correctness on HumanEval
206.79Output ThroughputOur approach
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Our approachVocabulary size=13264, Hardware=Single A100-80G GPU, Target Model=Llama-3.1-8B-Instruct, Inference Engine=SGLang, Framework=SpecForge2026.03 | 206.79 | 2.2 | |
| 128k-vocabVocabulary size=128k, Hardware=Single A100-80G GPU, Target Model=Llama-3.1-8B-Instruct, Inference Engine=SGLang, Framework=SpecForge2026.03 | 202.31 | — |