Large Language Model Inference Performance on Intel Ice Lake / AMD Zen4 AVX-512 VNNI
556Memory (MB)Litespark-Inference
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Litespark-InferenceSoftware Backend=AVX-512 VNNI2026.05 | 556 | 195 | 11.2 | 26.7 | |
| Binary AVX-512 + state cache2026.04 | 567 | — | 22 | — | |
| FP16 PyTorch2026.04 | 2,688 | — | 12 | — | |
| PyTorchSoftware Backend=PyTorch2026.05 | 7,800 | 2,450 | 0.42 | — |