Text Generation on benchmark suite 20-task
2,520Average End-to-End Latency (ms)Remote Speculate (RLM-Cascade)
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| Remote Speculate (RLM-Cascade)Serving Endpoint=Remote, Serving Strategy=Speculative2026.06 | 2,520 | 2,026 | 2,520 | 2.5 | 2,000 | |
| Remote Native OpusServing Endpoint=Remote, Serving Strategy=Native, Model=Opus2026.06 | 4,159 | 3,698 | 1,213 | 2.53 | 1,976 | |
| Local vLLM (4-bit 7B)Serving Endpoint=Local, Model Quantization=4-bit, Model Scale=7B, Serving Engine=vLLM2026.06 | 5,072 | 5,234 | 165 | 1.66 | 3,004 |