Serverless inference with per-token pricing (pay for output, input, and cached tokens). Training pricing from $0.50 to $10.00 per 1M training tokens depending on model size and fine-tuning type. On-demand deployments at $0.134 to $0.334 per minute ($8 to $20 per hour) for various GPU types.
Key features
Video captioning and detection
Image and video embeddings
Question answering on visual content
Video search and retrieval
Audio transcription and analysis
Real-time model inference
Model distillation for efficiency
Model quantization
Guided training runs
Configuration-led training
Custom training logic
Serverless inference
On-demand dedicated deployments
Reserved capacity
OpenAI compatible API
Model library with latest open models
Reinforcement learning support
Multi-region deployments
What makes it different
Natively multimodal architecture processing video, image, audio, and text in one model
Real-time inference optimized for existing hardware
Scalable video infrastructure for reasoning beyond detection
Enterprise deployment infrastructure with customization options
Drop-in replacement for closed-model APIs with cost savings of 50-75%
Instant production deployment from any training checkpoint in seconds
Optimized inference engine for industry-leading throughput and latency
Elastic and global RL inference scaling
Own the complete learning loop for specialized intelligence