Serverless inference with pay-per-use token pricing ($0.0015 to $4.50 per 1M tokens for chat/vision models). Provisioned throughput with reserved PTUs for committed capacity. Dedicated GPU endpoints at $5.49-$8.99 per hour. Fine-tuning from $0.34-$40 per 1M tokens. GPU clusters at $1.99-$9.99 per GP
Key features
GroqMetal: dedicated bare-metal infrastructure tuned for speed and reliability
GroqCore: production-ready inference stack with no infrastructure expertise required
GroqAssured: enterprise-grade governance, auditability, and control
256 LPUs per rack with 40 PB/s SRAM bandwidth
1,000 tokens per second per user throughput
128 GB on-chip SRAM per rack
315 PFLOPS of FP8 inference compute
13 globally distributed data centers across four continents
GPU clusters at scale with autoscaling
Voice agent infrastructure with sub-second latency
Comprehensive model library with 25+ regions
What makes it different
Fast inference at scale without sacrificing affordability
LPU pioneered technology integrated with NVIDIA GPUs
Fully integrated platform combining infrastructure, inference, and control
2x faster inference powered by cutting-edge research optimization
60% lower cost with workload-specific optimization
90% faster pre-training with Together Kernel Collection
Research-optimized infrastructure with proprietary kernels
Single unified API across multiple model providers