Operator Fusion Performance Analysis on Qwen LLM Series
37.44Attention Latency Reductionoperator fusion strategy
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| operator fusion strategyModel=Qwen2.5-0.5B, Tensix Cores (Count)=64, Size=1 GB2026.06 | 37.44 | 12.04 | 7.91 | 99.94 | |
| operator fusion strategyModel=Qwen3-0.6B, Tensix Cores (Count)=64, Size=1.2 GB2026.06 | 18.53 | 5.66 | 4.58 | 99.57 | |
| operator fusion strategyModel=Qwen3-4B, Tensix Cores (Count)=128, Size=8 GB2026.06 | 10.63 | 15.89 | 3.58 | 98.75 |