Starter plan is free with $30/month compute credits. Team plan is $250/month plus compute. Enterprise plan is custom. Compute is billed per second for CPU, GPU, and memory with no charges for idle time.
Pay-per-token for language models, per-minute for audio, per-image for image generation, and per-hour for GPU rental. No long-term contracts or upfront costs.
Key features
Serve LLMs, image, video, and audio models
Isolated environments for coding agents and RL rollouts
Fine-tune and train models on GPUs
Autoscale compute from zero to thousands of GPUs
Distributed storage for models and weights
Global capacity across 20+ clouds
Serverless functions with HTTPS endpoints
Out-of-the-box observability and logs
100+ machine learning models
Developer-friendly APIs
Text generation models
Image generation with Flux
Automatic scaling
Custom model deployment
Zero data retention policy
SOC 2 and ISO 27001 certified
Dedicated GPU clusters
Multi-GPU setups with SXM connections
What makes it different
Custom infrastructure built for AI including container runtime, storage, and scheduler
Sub-second container startup times for GPU workloads
Pay only by the second with no idle charges
Programmatically create millions of concurrent sandboxes
Own cutting-edge inference-optimized infrastructure with secure US-based data centers
Zero retention policy with SOC 2 and ISO 27001 certification for data privacy
No long-term contracts or hidden fees with simple pay-as-you-go pricing
100+ production-ready models covering all major AI use cases
Customizable inference solutions tailored to cost, latency, throughput or scale priorities