CE
Cerebras
Serverless Featured
Ultra-fast AI inference on custom Wafer-Scale Engine chips. Up to 3000 tok/s output speed, 20x faster than GPU-based providers.
Features
OpenAI Compatible
Streaming
Batching
Fine-tuning
Embeddings
Vision
Audio
Function Calling
JSON Mode
Compliance
SOC 2
HIPAA
GDPR
Pricing Model
per token Free tier: Free tier with all models, no credit card required
Details
Models 4
API Base
https://api.cerebras.ai/v1 LLM Inference
Model Catalog (4)
| Model | Type | Input $/1M | Output $/1M | Context | Speed | Status |
|---|---|---|---|---|---|---|
| GLM 4.7 Zhipu AI | llm | $2.25 | $2.75 | — | — | |
| GPT OSS 120B OpenAI · 117B MoE (5.1B active) | llm | $0.350 | $0.750 | — | — | |
| Qwen 3 235B Alibaba · 235B MoE | llm | $0.600 | $1.20 | — | — | |
| Qwen3.8 27B Alibaba | llm | $0.990 | $1.49 | 66k | — |
Cerebras API FAQ
Is the Cerebras API free?
Yes, Cerebras has a free tier: Free tier with all models, no credit card required. Paid usage is billed per model.
What is the cheapest Cerebras model?
GPT OSS 120B at $0.35 in / $0.75 out per 1M tokens, based on the 4 models we track for Cerebras.
How many models does Cerebras offer?
We track 4 Cerebras models: 4 LLM.
Is the Cerebras API OpenAI-compatible?
Yes. Cerebras exposes an OpenAI-compatible endpoint, so you can use the OpenAI SDKs by changing the base URL and API key.