Frontier models with simple per-token pricing. Use supported open models through an OpenAI-compatible API. No GPU reservation, no minimums, and no infrastructure to manage.
Product typeModel Hosting & Inference
Pricing$0.05/M tokens (MiniMax M3 cache read) · Usage Based
Deployment
APIAvailable
What teams use Serverless Inference for
dedicated NVIDIA GPU infrastructure
OpenAI-compatible serverless inference
flexible options for committed capacity
Integrations and access
Platforms, integrations, and language support vary by plan and region. Confirm final requirements with the vendor.
Platforms
IntegrationsAPI and webhooks
Languages
Support
What reviewers say
Community reviewsReviews are owned and editable by their authors.