HGEMM Performance Evaluation on 1,000 Configurations (Offline)
22Mean Speedup RatioCUDA-L2
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| CUDA-L2Baseline implementation=torch.matmul2025.12 | 22 | 19.2 | 0.211 | — | — | — | |
| CUDA-L2Baseline implementation=cuBLAS, Kernel configuration=TN2025.12 | 21.4 | 19.5 | 0.193 | — | — | — | |
| CUDA-L2Baseline implementation=cuBLAS, Kernel configuration=NN2025.12 | 20 | 17.5 | 0.197 | — | — | — | |
| CUDA-L2Baseline implementation=cuBLAS, Kernel configuration=max2025.12 | 19.2 | 16.4 | 0.191 | — | — | — | |
| CUDA-L2Baseline implementation=cuBLASLt-heuristic, Kernel configuration=TN2025.12 | 19.1 | 17.1 | 0.14 | — | — | — | |
| CUDA-L2Baseline implementation=cuBLASLt-heuristic, Kernel configuration=NN2025.12 | 17.3 | 15.6 | 0.143 | — | — | — | |
| CUDA-L2Baseline implementation=cuBLASLt-heuristic, Kernel configuration=max2025.12 | 16.8 | 15.3 | 0.14 | — | — | — | |
| CUDA-L2Baseline implementation=cuBLASLt-AutoTuning, Kernel configuration=TN2025.12 | 13.3 | 13.5 | 0.152 | — | — | — | |
| CUDA-L2Baseline implementation=cuBLASLt-AutoTuning, Kernel configuration=NN2025.12 | 12.1 | 11.4 | 0.157 | — | — | — | |
| CUDA-L2Baseline implementation=cuBLASLt-AutoTuning, Kernel configuration=max2025.12 | 11.4 | 11.2 | 0.152 | — | — | — | |
| CUDA-L2Baseline=torch.matmul2025.12 | — | — | — | — | 22 | 23.1 | |
| CUDA-L2Baseline=cuBLAS-max2025.12 | — | — | — | — | 19.2 | 20.2 | |
| CUDA-L2Baseline=cuBLASLt-heuristic-max2025.12 | — | — | — | — | 16.8 | 17 | |
| CUDA-L2Baseline=cuBLASLt-AutoTuning-max2025.12 | — | — | — | — | 11.4 | 13.2 |