Per-token billing for text models and per-image/second for media models, with all models priced 20% below official provider rates. Free trial with test credits, no credit card required. Credits roll over indefinitely.
Key features
GroqMetal: dedicated bare-metal infrastructure tuned for speed and reliability
GroqCore: production-ready inference stack with no infrastructure expertise required
GroqAssured: enterprise-grade governance, auditability, and control
256 LPUs per rack with 40 PB/s SRAM bandwidth
1,000 tokens per second per user throughput
128 GB on-chip SRAM per rack
315 PFLOPS of FP8 inference compute
13 globally distributed data centers across four continents
OpenAI-compatible API
Unified billing and credentials
Real-time usage analytics
Budget alerts and cost control
Model comparison tools
Playground for testing
Enterprise-grade security
99.9% service uptime
Sub-400ms average latency
What makes it different
Fast inference at scale without sacrificing affordability
LPU pioneered technology integrated with NVIDIA GPUs
Fully integrated platform combining infrastructure, inference, and control