A private System One Decision AI that runs locally
Stage
Not given
Beta
Founded
Not given
Not given
Based in
Not given
Not given
Pricing model
freemium
Not given
Pricing
Serverless inference with pay-per-use token pricing ($0.0015 to $4.50 per 1M tokens for chat/vision models). Provisioned throughput with reserved PTUs for committed capacity. Dedicated GPU endpoints at $5.49-$8.99 per hour. Fine-tuning from $0.34-$40 per 1M tokens. GPU clusters at $1.99-$9.99 per GP
Not given
Key features
GPU clusters at scale with autoscaling
Voice agent infrastructure with sub-second latency
Comprehensive model library with 25+ regions
Non-autoregressive with no free text generation
Multi-question call support in a single forward pass
Bounded output within supplied candidate set
Calibrated probabilities without temperature fitting
Local inference as a Rust binary executable
Deterministic inference with reproducible results
Support for choice, score, and yes-no decision types
CPU, Apple Metal, and CUDA hardware support
Trainable on your own decisions locally
What makes it different
2x faster inference powered by cutting-edge research optimization
60% lower cost with workload-specific optimization
90% faster pre-training with Together Kernel Collection
Research-optimized infrastructure with proprietary kernels
Single unified API across multiple model providers