Usage-based pay-as-you-go pricing with hourly compute charges; committed contracts available for volume discounts. Hosted option with Anyscale-managed infrastructure or Bring Your Own Cloud deployment available.
Free plan with pay-as-you-go model usage, 5M TPM / 100 RPM rate limits. Growth plan: $25/month with $25 in token credits included, 15M TPM / 3,000 RPM rate limits. Enterprise plan: custom, volume-based pricing with BYOC, SSO/SAML, custom SLAs, VPC and private deployments.
Key features
Multimodal data curation at scale
Distributed model training with elastic scaling
Batch embedding generation at scale
Multi-cloud deployment and orchestration
Cluster-backed development environments
Workload-specific observability and debugging
Production-grade managed Ray clusters
Access control and governance
GPU budget controls and cost attribution
Post-training with SkyRL and veRL support
Managed Skills with instant deployment to all agents
Agent observability and debugging across all decisions
Globally distributed inference with low latency
Safety guardrails and agent evaluation
Multi-agent orchestration with stateful memory
Centralized security and access control
Support for multiple frontier LLM models
Real-time cost and performance insights
Agent Insights dashboard
What makes it different
Built by creators of Ray, the world's most widely adopted AI compute engine
Multi-cloud deployment without code changes
Feels local but runs distributed
Unified GPU pooling across clouds and regions
Ship agent improvements in minutes instead of release cycles
Centralized skill management without requiring redeploy
Complete visibility into agent decisions for fast debugging
Proven safe releases with guardrails before production
No infrastructure expertise required to deploy agents at scale