Pay-per-token for language models, per-minute for audio, per-image for image generation, and per-hour for GPU rental. No long-term contracts or upfront costs.
Key features
GroqMetal: dedicated bare-metal infrastructure tuned for speed and reliability
GroqCore: production-ready inference stack with no infrastructure expertise required
GroqAssured: enterprise-grade governance, auditability, and control
256 LPUs per rack with 40 PB/s SRAM bandwidth
1,000 tokens per second per user throughput
128 GB on-chip SRAM per rack
315 PFLOPS of FP8 inference compute
13 globally distributed data centers across four continents
100+ machine learning models
Developer-friendly APIs
Text generation models
Image generation with Flux
Automatic scaling
Custom model deployment
Zero data retention policy
SOC 2 and ISO 27001 certified
Dedicated GPU clusters
Multi-GPU setups with SXM connections
What makes it different
Fast inference at scale without sacrificing affordability
LPU pioneered technology integrated with NVIDIA GPUs
Fully integrated platform combining infrastructure, inference, and control
Own cutting-edge inference-optimized infrastructure with secure US-based data centers
Zero retention policy with SOC 2 and ISO 27001 certification for data privacy
No long-term contracts or hidden fees with simple pay-as-you-go pricing
100+ production-ready models covering all major AI use cases
Customizable inference solutions tailored to cost, latency, throughput or scale priorities