Starter plan is free with $30/month compute credits. Team plan is $250/month plus compute. Enterprise plan is custom. Compute is billed per second for CPU, GPU, and memory with no charges for idle time.
Not given
Key features
Serve LLMs, image, video, and audio models
Isolated environments for coding agents and RL rollouts
Fine-tune and train models on GPUs
Autoscale compute from zero to thousands of GPUs
Distributed storage for models and weights
Global capacity across 20+ clouds
Serverless functions with HTTPS endpoints
Out-of-the-box observability and logs
Long-term memory for AI agents
Automatic fact extraction and reconciliation
Curated recall retrieval
Raw vector search capability
Model Context Protocol integration
Regional data residency
End-to-end encryption
Right to erasure
Bring-your-own-key model support
Recalld Chat assistant
What makes it different
Custom infrastructure built for AI including container runtime, storage, and scheduler
Sub-second container startup times for GPU workloads
Pay only by the second with no idle charges
Programmatically create millions of concurrent sandboxes
Curated recall uses 6.7× fewer tokens than raw search while maintaining accuracy
Automatic fact updating and reconciliation without manual deduplication pipelines
Verified on open benchmark maintained by independent company
Native MCP server requiring no SDK or glue code
GDPR-compliant by design with transparent data handling