Matrix Multiplication on KernelBench FP4 Matmul
2,898Throughput (TF/s)AutoKernel (Triton)
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| AutoKernel (Triton)Shape (M×N×K)=2048×18432×3072, Precision=FP4, Hardware=H1002026.03 | 2,898 | 0.08 | 1.63 | |
| AutoKernel (Triton)Shape (M×N×K)=1024×18432×3072, Precision=FP4, Hardware=H1002026.03 | 2,777 | 0.042 | 1.73 | |
| AutoKernel (Triton)Shape (M×N×K)=4096×3072×3072, Precision=FP4, Hardware=H1002026.03 | 2,443 | 0.032 | 1.74 | |
| CUTLASSShape (M×N×K)=2048×18432×3072, Precision=FP4, Hardware=H1002026.03 | 1,777 | 0.13 | — | |
| AutoKernel (Triton)Shape (M×N×K)=2048×3072×3072, Precision=FP4, Hardware=H1002026.03 | 1,662 | 0.023 | 1.72 | |
| CUTLASSShape (M×N×K)=1024×18432×3072, Precision=FP4, Hardware=H1002026.03 | 1,609 | 0.072 | — | |
| AutoKernel (Triton)Shape (M×N×K)=1024×3072×3072, Precision=FP4, Hardware=H1002026.03 | 1,477 | 0.013 | 2.15 | |
| CUTLASSShape (M×N×K)=4096×3072×3072, Precision=FP4, Hardware=H1002026.03 | 1,405 | 0.055 | — | |
| AutoKernel (Triton)Shape (M×N×K)=128×18432×3072, Precision=FP4, Hardware=H1002026.03 | 1,105 | 0.013 | 2.01 | |
| CUTLASSShape (M×N×K)=2048×3072×3072, Precision=FP4, Hardware=H1002026.03 | 964 | 0.04 | — | |
| CUTLASSShape (M×N×K)=1024×3072×3072, Precision=FP4, Hardware=H1002026.03 | 686 | 0.028 | — | |
| CUTLASSShape (M×N×K)=128×18432×3072, Precision=FP4, Hardware=H1002026.03 | 550 | 0.026 | — | |
| AutoKernel (Triton)Shape (M×N×K)=128×3072×3072, Precision=FP4, Hardware=H1002026.03 | 186 | 0.013 | 1.86 | |
| CUTLASSShape (M×N×K)=128×3072×3072, Precision=FP4, Hardware=H1002026.03 | 100 | 0.024 | — |