Language Model Inference on Llama-2 7B-Chat
76Latency (ms/token)ARC engine
Evaluation Results
| Method | Links | |
|---|---|---|
| ARC engineHardware Backend=GPU (M2 Ultra), Model Scale=7B, Deterministic=Yes (all platforms)2026.03 | 76 | |
| ARC engineHardware Backend=CPU (M2 Ultra), Model Scale=7B, Deterministic=Yes (all platforms)2026.03 | 139 | |
| Candle Q4 floatHardware Backend=M2 Ultra, Model Scale=7B, Deterministic=No2026.03 | 175 | |
| Candle Q4 floatHardware Backend=Vultr x86, Model Scale=7B, Deterministic=No2026.03 | 1,250 |