Visual-Language Model Inference on VILA RTX 4090
168.1Throughput (tokens/sec)VILA-7B-AWQ
Evaluation Results
| Method | Links | |
|---|---|---|
| VILA-7B-AWQPrecision=W4A162023.06 | 168.1 | |
| VILA-13B-AWQPrecision=W4A162023.06 | 99 | |
| VILA-7BPrecision=FP162023.06 | 58.5 |
| Method | Links | |
|---|---|---|
| VILA-7B-AWQPrecision=W4A162023.06 | 168.1 | |
| VILA-13B-AWQPrecision=W4A162023.06 | 99 | |
| VILA-7BPrecision=FP162023.06 | 58.5 |