Serverless inference with per-token pricing (pay for output, input, and cached tokens). Training pricing from $0.50 to $10.00 per 1M training tokens depending on model size and fine-tuning type. On-demand deployments at $0.134 to $0.334 per minute ($8 to $20 per hour) for various GPU types.
Key features
Find files across Finder, Gmail, Slack, and Google Calendar
Send emails and schedule meetings with voice commands
Review actions before execution or enable auto-execute
Handle multi-step workflows with a single voice request
Works across 20+ apps including Slack, Gmail, Cursor, and Notion
Automatic grammar correction and filler word removal
Guided training runs
Configuration-led training
Custom training logic
Serverless inference
On-demand dedicated deployments
Reserved capacity
OpenAI compatible API
Model library with latest open models
Reinforcement learning support
Multi-region deployments
What makes it different
Turn voice into action across multiple apps without switching
Execute multi-step workflows with a single voice command
User maintains full control with review-before-running option
Integrates seamlessly with existing productivity tools and workflows
Drop-in replacement for closed-model APIs with cost savings of 50-75%
Instant production deployment from any training checkpoint in seconds
Optimized inference engine for industry-leading throughput and latency
Elastic and global RL inference scaling
Own the complete learning loop for specialized intelligence