Large Language Model Inference on Chatbot Instruction Prompts and Finance Alpaca (test)
265.7Throughput (TPS)Aurora (Qwen3-Coder-Next (FP8))
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| Aurora (Qwen3-Coder-Next (FP8))Batch Size=1, Speculative Decoding=true, Precision=FP8, Lookahead=52026.02 | 265.7 | 264.8 | 208.7 | 320.5 | 1.51 | 3.06 | |
| Aurora (MiniMax M2.1 (FP8))Batch Size=1, Speculative Decoding=true, Precision=FP8, Lookahead=42026.02 | 211.8 | 210.6 | 163.1 | 270.3 | 1.57 | 2.62 | |
| Qwen3-Coder-Next (FP8)Batch Size=1, Speculative Decoding=false, Precision=FP8, Lookahead=52026.02 | 176.4 | 178 | 172.3 | 178.4 | — | — | |
| Aurora (Qwen3-Coder-Next (FP8))Batch Size=8, Speculative Decoding=true, Precision=FP8, Lookahead=52026.02 | 146.3 | 143.5 | 109.6 | 189.5 | 1.23 | 3.07 | |
| MiniMax M2.1 (FP8)Batch Size=1, Speculative Decoding=false, Precision=FP8, Lookahead=42026.02 | 134.9 | 136.4 | 130.6 | 136.9 | — | — | |
| Qwen3-Coder-Next (FP8)Batch Size=8, Speculative Decoding=false, Precision=FP8, Lookahead=52026.02 | 119.8 | 121.5 | 104.8 | 134.6 | — | — | |
| Aurora (Qwen3-Coder-Next (FP8))Batch Size=16, Speculative Decoding=true, Precision=FP8, Lookahead=52026.02 | 107.6 | 103.7 | 75.7 | 156.6 | 1.09 | 3.06 | |
| Aurora (MiniMax M2.1 (FP8))Batch Size=8, Speculative Decoding=true, Precision=FP8, Lookahead=42026.02 | 107.1 | 104.5 | 79.9 | 137.1 | 1.36 | 2.62 | |
| Qwen3-Coder-Next (FP8)Batch Size=16, Speculative Decoding=false, Precision=FP8, Lookahead=52026.02 | 99.6 | 102.1 | 74.5 | 119.2 | — | — | |
| Aurora (MiniMax M2.1 (FP8))Batch Size=16, Speculative Decoding=true, Precision=FP8, Lookahead=42026.02 | 83.1 | 82.9 | 60.9 | 112 | 1.29 | 2.62 | |
| MiniMax M2.1 (FP8)Batch Size=8, Speculative Decoding=false, Precision=FP8, Lookahead=42026.02 | 79 | 78.7 | 73.7 | 85.1 | — | — | |
| Aurora (MiniMax M2.1 (FP8))Batch Size=32, Speculative Decoding=true, Precision=FP8, Lookahead=42026.02 | 67.1 | 64.7 | 44 | 100.5 | 1.25 | 2.62 | |
| MiniMax M2.1 (FP8)Batch Size=16, Speculative Decoding=false, Precision=FP8, Lookahead=42026.02 | 64.5 | 63.7 | 58.9 | 72.3 | — | — | |
| MiniMax M2.1 (FP8)Batch Size=32, Speculative Decoding=false, Precision=FP8, Lookahead=42026.02 | 53.5 | 52.9 | 47.1 | 67.1 | — | — |