Pay-per-token for language models, per-minute for audio, per-image for image generation, and per-hour for GPU rental. No long-term contracts or upfront costs.
Free plan with pay-as-you-go model usage, 5M TPM / 100 RPM rate limits. Growth plan: $25/month with $25 in token credits included, 15M TPM / 3,000 RPM rate limits. Enterprise plan: custom, volume-based pricing with BYOC, SSO/SAML, custom SLAs, VPC and private deployments.
Key features
100+ machine learning models
Developer-friendly APIs
Text generation models
Image generation with Flux
Automatic scaling
Custom model deployment
Zero data retention policy
SOC 2 and ISO 27001 certified
Dedicated GPU clusters
Multi-GPU setups with SXM connections
Managed Skills with instant deployment to all agents
Agent observability and debugging across all decisions
Globally distributed inference with low latency
Safety guardrails and agent evaluation
Multi-agent orchestration with stateful memory
Centralized security and access control
Support for multiple frontier LLM models
Real-time cost and performance insights
Agent Insights dashboard
What makes it different
Own cutting-edge inference-optimized infrastructure with secure US-based data centers
Zero retention policy with SOC 2 and ISO 27001 certification for data privacy
No long-term contracts or hidden fees with simple pay-as-you-go pricing
100+ production-ready models covering all major AI use cases
Customizable inference solutions tailored to cost, latency, throughput or scale priorities
Ship agent improvements in minutes instead of release cycles
Centralized skill management without requiring redeploy
Complete visibility into agent decisions for fast debugging
Proven safe releases with guardrails before production
No infrastructure expertise required to deploy agents at scale