LLM Inference
Benchmarks
Dataset NameSOTA methodMetricTrendResultsLast Updated
5.56Speedup
57
Jun 9, 2026
11.09TTFT (ms)
33
Apr 28, 2026
46,782Throughput (tok/s)
30
Jun 2, 2026
61,106Throughput (tok/s)
28
Jun 2, 2026
106,324Throughput (tok/s)
26
Jun 2, 2026
2.44Throughput (tokens/s)
24
May 20, 2026
2.87Mean Speedup
21
Feb 26, 2026
3.09Speedup
21
Feb 26, 2026
2.84Speedup
21
Feb 26, 2026
2.81Speedup
21
Feb 26, 2026
80.3SLO Attainment (%)
18
Jun 2, 2026
1.16Model Load Time (s)
18
Feb 26, 2026
3.9Goodput (req/s)
18
Feb 26, 2026
1.16Goodput (req/s)
18
Feb 26, 2026
48,015Throughput (tok/s)
16
Jun 2, 2026
12.16Decode Latency (ms/token)
16
May 12, 2026
17.84Median Decode Latency (ms/step)
13
Jun 24, 2026
2,813.19Prefill Min Throughput (tokens/sec)
13
May 12, 2026
1,709.9Prefill Throughput (min)
12
May 12, 2026
5,463Memory Usage (MB)
12
May 7, 2026
75.17RPS (Requests/s)
11
Jun 2, 2026
530.02Prefill Throughput (min) (tokens/sec)
10
May 12, 2026
591.01Prefill Throughput (min, tokens/sec)
10
May 12, 2026
1,161.29Prefill Throughput (min, tokens/sec)
10
May 12, 2026
100SLO Attainment
9
Feb 26, 2026