LLM Inference on Llama 70B 3.1 (Throughput and Context)
7.3Throughput (tok/s)mlx-lm
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| mlx-lmKV fmt=FP16, Hardware=M1 Max 64 GB2026.04 | 7.3 | 73,000 | — | |
| Open-TQ-MetalKV fmt=int4, Hardware=M1 Max 64 GB2026.04 | 6 | 236,000 | — | |
| llama.cppKV fmt=Mixed, Hardware=M1 Max 64 GB2026.04 | 5 | 50,000 | — | |
| OllamaKV fmt=GGUF, Hardware=M1 Max 64 GB2026.04 | 5 | 40,000 | — |