Run and fine-tune models. Deploy custom models. All with one line of code.
The memory layer for AI Agents
Stage
Not given
Growing
Founded
Not given
Not given
Based in
Not given
Not given
Pricing model
usage-based
freemium
Pricing
Most models billed by time based on hardware used (CPU $0.000100/sec, T4 GPU $0.000225/sec, L40S GPU $0.000975/sec). Some models billed by input/output tokens or generated items. Private custom models billed for all uptime; fast-booting fine-tunes billed only for active processing.
Not given
Key features
Run thousands of community-contributed open-source models
Fine-tune models with custom data to create specialized versions
Deploy custom models using Cog, an open-source packaging tool
Automatic scaling based on demand
Pay only for compute time used
Image generation from text
Long-term memory for AI agents
Automatic fact extraction and reconciliation
Curated recall retrieval
Raw vector search capability
Model Context Protocol integration
Regional data residency
End-to-end encryption
Right to erasure
Bring-your-own-key model support
Recalld Chat assistant
What makes it different
One-line code interface for running models without ML expertise
Automatic infrastructure scaling without manual management
Open-source community with thousands of production-ready models
Fine-tuning capability to customize models for specific tasks
Curated recall uses 6.7× fewer tokens than raw search while maintaining accuracy
Automatic fact updating and reconciliation without manual deduplication pipelines
Verified on open benchmark maintained by independent company
Native MCP server requiring no SDK or glue code
GDPR-compliant by design with transparent data handling