Usage-based pay-as-you-go pricing with hourly compute charges; committed contracts available for volume discounts. Hosted option with Anyscale-managed infrastructure or Bring Your Own Cloud deployment available.
Key features
GroqMetal: dedicated bare-metal infrastructure tuned for speed and reliability
GroqCore: production-ready inference stack with no infrastructure expertise required
GroqAssured: enterprise-grade governance, auditability, and control
256 LPUs per rack with 40 PB/s SRAM bandwidth
1,000 tokens per second per user throughput
128 GB on-chip SRAM per rack
315 PFLOPS of FP8 inference compute
13 globally distributed data centers across four continents
Multimodal data curation at scale
Distributed model training with elastic scaling
Batch embedding generation at scale
Multi-cloud deployment and orchestration
Cluster-backed development environments
Workload-specific observability and debugging
Production-grade managed Ray clusters
Access control and governance
GPU budget controls and cost attribution
Post-training with SkyRL and veRL support
What makes it different
Fast inference at scale without sacrificing affordability
LPU pioneered technology integrated with NVIDIA GPUs
Fully integrated platform combining infrastructure, inference, and control
Built by creators of Ray, the world's most widely adopted AI compute engine