Usage-based pay-as-you-go pricing with hourly compute charges; committed contracts available for volume discounts. Hosted option with Anyscale-managed infrastructure or Bring Your Own Cloud deployment available.
Pay-per-token for language models, per-minute for audio, per-image for image generation, and per-hour for GPU rental. No long-term contracts or upfront costs.
Key features
Multimodal data curation at scale
Distributed model training with elastic scaling
Batch embedding generation at scale
Multi-cloud deployment and orchestration
Cluster-backed development environments
Workload-specific observability and debugging
Production-grade managed Ray clusters
Access control and governance
GPU budget controls and cost attribution
Post-training with SkyRL and veRL support
100+ machine learning models
Developer-friendly APIs
Text generation models
Image generation with Flux
Automatic scaling
Custom model deployment
Zero data retention policy
SOC 2 and ISO 27001 certified
Dedicated GPU clusters
Multi-GPU setups with SXM connections
What makes it different
Built by creators of Ray, the world's most widely adopted AI compute engine
Multi-cloud deployment without code changes
Feels local but runs distributed
Unified GPU pooling across clouds and regions
Own cutting-edge inference-optimized infrastructure with secure US-based data centers
Zero retention policy with SOC 2 and ISO 27001 certification for data privacy
No long-term contracts or hidden fees with simple pay-as-you-go pricing
100+ production-ready models covering all major AI use cases
Customizable inference solutions tailored to cost, latency, throughput or scale priorities