Inference Efficiency Analysis on NVIDIA RTX 5090D GPU
17.56Latency (ms)AHA-WAM-Flash
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| AHA-WAM-FlashHardware=NVIDIA RTX 5090D GPU, Sampler=Distilled, Interface=Asynchronous planner–executor2026.06 | 17.56 | 56.95 | 10.82 | |
| AHA-WAMHardware=NVIDIA RTX 5090D GPU, Interface=Asynchronous planner–executor2026.06 | 41.37 | 24.17 | 4.59 | |
| Fast-WAMHardware=Official latency, Execution scheme=RTC-style non-blocking [38]2026.06 | 190 | 5.26 | 1 | |
| MotusHardware=NVIDIA RTX 5090D GPU, Execution scheme=RTC-style non-blocking [38]2026.06 | 1,866.1 | 0.54 | 0.1 |