Usage-based pay-as-you-go pricing with hourly compute charges; committed contracts available for volume discounts. Hosted option with Anyscale-managed infrastructure or Bring Your Own Cloud deployment available.
Not given
Key features
Multimodal data curation at scale
Distributed model training with elastic scaling
Batch embedding generation at scale
Multi-cloud deployment and orchestration
Cluster-backed development environments
Workload-specific observability and debugging
Production-grade managed Ray clusters
Access control and governance
GPU budget controls and cost attribution
Post-training with SkyRL and veRL support
Long-term memory for AI agents
Automatic fact extraction and reconciliation
Curated recall retrieval
Raw vector search capability
Model Context Protocol integration
Regional data residency
End-to-end encryption
Right to erasure
Bring-your-own-key model support
Recalld Chat assistant
What makes it different
Built by creators of Ray, the world's most widely adopted AI compute engine
Multi-cloud deployment without code changes
Feels local but runs distributed
Unified GPU pooling across clouds and regions
Curated recall uses 6.7× fewer tokens than raw search while maintaining accuracy
Automatic fact updating and reconciliation without manual deduplication pipelines
Verified on open benchmark maintained by independent company
Native MCP server requiring no SDK or glue code
GDPR-compliant by design with transparent data handling