Starter plan is free with $30/month compute credits. Team plan is $250/month plus compute. Enterprise plan is custom. Compute is billed per second for CPU, GPU, and memory with no charges for idle time.
Not given
Key features
Serve LLMs, image, video, and audio models
Isolated environments for coding agents and RL rollouts
Fine-tune and train models on GPUs
Autoscale compute from zero to thousands of GPUs
Distributed storage for models and weights
Global capacity across 20+ clouds
Serverless functions with HTTPS endpoints
Out-of-the-box observability and logs
Video captioning and detection
Image and video embeddings
Question answering on visual content
Video search and retrieval
Audio transcription and analysis
Real-time model inference
Model distillation for efficiency
Model quantization
What makes it different
Custom infrastructure built for AI including container runtime, storage, and scheduler
Sub-second container startup times for GPU workloads
Pay only by the second with no idle charges
Programmatically create millions of concurrent sandboxes
Natively multimodal architecture processing video, image, audio, and text in one model
Real-time inference optimized for existing hardware
Scalable video infrastructure for reasoning beyond detection
Enterprise deployment infrastructure with customization options