Starter plan is free with $30/month compute credits. Team plan is $250/month plus compute. Enterprise plan is custom. Compute is billed per second for CPU, GPU, and memory with no charges for idle time.
Usage-based pay-as-you-go pricing with hourly compute charges; committed contracts available for volume discounts. Hosted option with Anyscale-managed infrastructure or Bring Your Own Cloud deployment available.
Key features
Serve LLMs, image, video, and audio models
Isolated environments for coding agents and RL rollouts
Fine-tune and train models on GPUs
Autoscale compute from zero to thousands of GPUs
Distributed storage for models and weights
Global capacity across 20+ clouds
Serverless functions with HTTPS endpoints
Out-of-the-box observability and logs
Multimodal data curation at scale
Distributed model training with elastic scaling
Batch embedding generation at scale
Multi-cloud deployment and orchestration
Cluster-backed development environments
Workload-specific observability and debugging
Production-grade managed Ray clusters
Access control and governance
GPU budget controls and cost attribution
Post-training with SkyRL and veRL support
What makes it different
Custom infrastructure built for AI including container runtime, storage, and scheduler
Sub-second container startup times for GPU workloads
Pay only by the second with no idle charges
Programmatically create millions of concurrent sandboxes
Built by creators of Ray, the world's most widely adopted AI compute engine