GR

Groq

Serverless Featured

Fastest LLM inference powered by custom LPU chips. OpenAI-compatible API with sub-second latency.

Features

OpenAI Compatible
Streaming
Batching
Fine-tuning
Embeddings
Vision
Audio
Function Calling
JSON Mode

Compliance

SOC 2
HIPAA
GDPR

Pricing Model

per token
Free tier: Available

Details

Models 8
API Base https://api.groq.com/openai/v1
Audio / MusicLLM Inference

Model Catalog (8)

Model Type Input $/1M Output $/1M Context Speed Status
GPT OSS 120B
OpenAI · 117B MoE (5.1B active)
llm $0.150 $0.600 131k —
Kimi K2
Moonshot AI · 1T MoE (32B active)
llm $1.00 $3.00 — —
Llama 3.3 70B
Meta · 70B
llm Enterprise only: Groq lists th see notes — —
Llama 4 Scout
Meta · 109B (17B active)
llm $0.110 $0.340 — —
Qwen 3 32B
Alibaba · 32B
llm $0.290 $0.590 — —
Qwen3.8 27B
Alibaba
llm $0.800 $4.00 131k —
Whisper Large V3
OpenAI · 1.5B
speech to_text $0.0000/sec — — —
gpt-oss-20b
OpenAI
llm $0.075 $0.300 131k —

Groq API FAQ

Is the Groq API free?

Yes, Groq has a free tier. Paid usage is billed per model.

What is the cheapest Groq model?

Gpt-oss-20b at $0.07 in / $0.30 out per 1M tokens, based on the 8 models we track for Groq.

How many models does Groq offer?

We track 8 Groq models: 7 LLM, 1 speech-to-text.

Is the Groq API OpenAI-compatible?

Yes. Groq exposes an OpenAI-compatible endpoint, so you can use the OpenAI SDKs by changing the base URL and API key.