Starter plan is free with $30/month compute credits. Team plan is $250/month plus compute. Enterprise plan is custom. Compute is billed per second for CPU, GPU, and memory with no charges for idle time.
Free plan with pay-as-you-go model usage, 5M TPM / 100 RPM rate limits. Growth plan: $25/month with $25 in token credits included, 15M TPM / 3,000 RPM rate limits. Enterprise plan: custom, volume-based pricing with BYOC, SSO/SAML, custom SLAs, VPC and private deployments.
Key features
Serve LLMs, image, video, and audio models
Isolated environments for coding agents and RL rollouts
Fine-tune and train models on GPUs
Autoscale compute from zero to thousands of GPUs
Distributed storage for models and weights
Global capacity across 20+ clouds
Serverless functions with HTTPS endpoints
Out-of-the-box observability and logs
Managed Skills with instant deployment to all agents
Agent observability and debugging across all decisions
Globally distributed inference with low latency
Safety guardrails and agent evaluation
Multi-agent orchestration with stateful memory
Centralized security and access control
Support for multiple frontier LLM models
Real-time cost and performance insights
Agent Insights dashboard
What makes it different
Custom infrastructure built for AI including container runtime, storage, and scheduler
Sub-second container startup times for GPU workloads
Pay only by the second with no idle charges
Programmatically create millions of concurrent sandboxes
Ship agent improvements in minutes instead of release cycles
Centralized skill management without requiring redeploy
Complete visibility into agent decisions for fast debugging
Proven safe releases with guardrails before production
No infrastructure expertise required to deploy agents at scale