Run and fine-tune models. Deploy custom models. All with one line of code.
Stage
Established
Not given
Founded
Not given
Not given
Based in
US
Not given
Pricing model
Not given
usage-based
Pricing
Not given
Most models billed by time based on hardware used (CPU $0.000100/sec, T4 GPU $0.000225/sec, L40S GPU $0.000975/sec). Some models billed by input/output tokens or generated items. Private custom models billed for all uptime; fast-booting fine-tunes billed only for active processing.
Key features
GroqMetal: dedicated bare-metal infrastructure tuned for speed and reliability
GroqCore: production-ready inference stack with no infrastructure expertise required
GroqAssured: enterprise-grade governance, auditability, and control
256 LPUs per rack with 40 PB/s SRAM bandwidth
1,000 tokens per second per user throughput
128 GB on-chip SRAM per rack
315 PFLOPS of FP8 inference compute
13 globally distributed data centers across four continents
Run thousands of community-contributed open-source models
Fine-tune models with custom data to create specialized versions
Deploy custom models using Cog, an open-source packaging tool
Automatic scaling based on demand
Pay only for compute time used
Image generation from text
What makes it different
Fast inference at scale without sacrificing affordability
LPU pioneered technology integrated with NVIDIA GPUs
Fully integrated platform combining infrastructure, inference, and control
One-line code interface for running models without ML expertise
Automatic infrastructure scaling without manual management
Open-source community with thousands of production-ready models
Fine-tuning capability to customize models for specific tasks