Large Language Model Inference on GPQA
2,341.3ThroughputSpecBundle
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| SpecBundleTarget Model=Qwen-30B-A3B, Draft Model=SpecBundle, #GPUs=42026.03 | 2,341.3 | 1.66 | |
| SpecBundleTarget Model=Llama-4-Scout, Draft Model=SpecBundle, #GPUs=82026.03 | 1,502.2 | 2.78 | |
| Standard InferenceTarget Model=Qwen-30B-A3B, Draft Model=-, #GPUs=42026.03 | 1,410.4 | 1 | |
| SpecBundleTarget Model=Llama-3.3-70B, Draft Model=SpecBundle, #GPUs=42026.03 | 1,405.1 | 2.44 | |
| EAGLE3Target Model=Llama-4-Scout, Draft Model=Existing, #GPUs=82026.03 | 1,405.1 | 2.6 | |
| SpecBundleTarget Model=Ling-Flash-V2, Draft Model=SpecBundle, #GPUs=82026.03 | 1,185.7 | 1.49 | |
| EAGLE3Target Model=Llama-3.3-70B, Draft Model=Existing, #GPUs=42026.03 | 1,049 | 1.82 | |
| SpecBundleTarget Model=Qwen-235B-A22B, Draft Model=SpecBundle, #GPUs=82026.03 | 826.5 | 1.47 | |
| SpecBundleTarget Model=Kimi-K2, Draft Model=SpecBundle, #GPUs=82026.03 | 811.4 | 1.61 | |
| Standard InferenceTarget Model=Ling-Flash-V2, Draft Model=-, #GPUs=82026.03 | 794.1 | 1 | |
| EAGLE3Target Model=Qwen-235B-A22B, Draft Model=Existing, #GPUs=82026.03 | 716.7 | 1.27 | |
| Standard InferenceTarget Model=Llama-3.3-70B, Draft Model=-, #GPUs=42026.03 | 575.7 | 1 | |
| Standard InferenceTarget Model=Qwen-235B-A22B, Draft Model=-, #GPUs=82026.03 | 563.2 | 1 | |
| Standard InferenceTarget Model=Llama-4-Scout, Draft Model=-, #GPUs=82026.03 | 541 | 1 | |
| SpecBundleTarget Model=Llama-3.1-8B, Draft Model=SpecBundle, #GPUs=12026.03 | 514.2 | 2.7 | |
| Standard InferenceTarget Model=Kimi-K2, Draft Model=-, #GPUs=82026.03 | 505.4 | 1 | |
| EAGLE3Target Model=Llama-3.1-8B, Draft Model=Existing, #GPUs=12026.03 | 438.1 | 2.3 | |
| Standard InferenceTarget Model=Llama-3.1-8B, Draft Model=-, #GPUs=12026.03 | 190.5 | 1 |