The fastest way to run open-source models on demand. Powered by cutting-edge inference research. No infrastructure to manage, no long-term commitments.
Product typeModel Hosting & Inference
Pricing · Usage Based
Deployment
APINot available
What teams use Serverless Inference for
Inference
Fine-tuning
GPU clusters
Model shaping
Pre-training
Integrations and access
Platforms, integrations, and language support vary by plan and region. Confirm final requirements with the vendor.
Platforms
IntegrationsAPI and webhooks
Languages
Support
What reviewers say
Community reviewsReviews are owned and editable by their authors.
Compliance claims are normalized from current vendor documentation and independently reviewed by SOTA2.
SOC 2 Type II
Zero data retention (ZDR): Together does not store inputs or outputs by default; temporary caching may be used for performance unless otherwise configured.
Training opt-in: Data sharing for training other models is opt-in and not enabled by default; configurable in Organization Settings under Privacy.
Organization-level privacy toggles: Admins can control storing prompts/responses, allowing data for training, and allowing passthrough models.
Passthrough model isolation: Non-passthrough models run on Together's own infrastructure with no call-out to model authors; passthrough is a separate organization-level toggle.
VPC-based deployments and private networking available for enterprise customers.
OIDC authentication supported for GPU Clusters.
Single sign-on (SSO) listed under Administration features.
Data retention:
Benchmarks and comparisons
Independent benchmarksStructured task results are being verified.
Side-by-side comparisonsComparison workspaces will be available in a later release.