HGEMM performance evaluation on HGEMM 1,000 configurations (Server)
30.2Mean Speedup RatioCUDA-L2
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| CUDA-L2Baseline implementation=cuBLAS, Kernel configuration=TN2025.12 | 30.2 | 26.3 | 0.275 | — | — | |
| CUDA-L2Baseline implementation=cuBLAS, Kernel configuration=NN2025.12 | 28.8 | 25.2 | 0.283 | — | — | |
| CUDA-L2Baseline implementation=torch.matmul2025.12 | 28.7 | 25.6 | 0.275 | — | — | |
| CUDA-L2Baseline=torch.matmul2025.12 | 28.7 | — | — | — | 29.8 | |
| CUDA-L2Baseline implementation=cuBLAS, Kernel configuration=max2025.12 | 26 | 22.9 | 0.26 | — | — | |
| CUDA-L2Baseline=cuBLAS-max2025.12 | 26 | — | — | — | 27.2 | |
| CUDA-L2Baseline implementation=cuBLASLt-heuristic, Kernel configuration=TN2025.12 | 25.9 | 24.1 | 0.198 | — | — | |
| CUDA-L2Baseline implementation=cuBLASLt-heuristic, Kernel configuration=NN2025.12 | 24.4 | 22.6 | 0.202 | — | — | |
| CUDA-L2Baseline implementation=cuBLASLt-heuristic, Kernel configuration=max2025.12 | 22.4 | 21.4 | 0.186 | — | — | |
| CUDA-L2Baseline=cuBLASLt-heuristic-max2025.12 | 22.4 | — | — | — | 22.7 | |
| CUDA-L2Baseline implementation=cuBLASLt-AutoTuning, Kernel configuration=TN2025.12 | 19.1 | 17.6 | 0.217 | — | — | |
| CUDA-L2Baseline implementation=cuBLASLt-AutoTuning, Kernel configuration=NN2025.12 | 17.9 | 15.9 | 0.22 | — | — | |
| CUDA-L2Baseline implementation=cuBLASLt-AutoTuning, Kernel configuration=max2025.12 | 15.9 | 14.4 | 0.207 | — | — | |
| CUDA-L2Baseline=cuBLASLt-AutoTuning-max2025.12 | 15.9 | — | — | — | 18.1 |