Competition-level Mathematical Reasoning on AIME (Output throughput, Diff %)
224.83Output ThroughputOur approach
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Our approachVocabulary size=13264, Hardware=Single A100-80G GPU, Target Model=Llama-3.1-8B-Instruct, Inference Engine=SGLang, Framework=SpecForge2026.03 | 224.83 | 6.7 | |
| 128k-vocabVocabulary size=128k, Hardware=Single A100-80G GPU, Target Model=Llama-3.1-8B-Instruct, Inference Engine=SGLang, Framework=SpecForge2026.03 | 210.74 | — |