CO
Coral Bricks
Serverless
Fast inference for coding and research agents: open models (GLM 5.3, GLM 5.3 Flash, DeepSeek V4.1 Flash) served FP4-quantized on B200/B300 GPUs behind an OpenAI-compatible API with Chat Completions, Responses (server-side state) and background mode for long-running agents. Up to ~469 output tok/s, 1M-token context, cached input billed at $0. Prepaid per-token billing; private endpoints available.
Features
OpenAI Compatible
Streaming
Batching
Fine-tuning
Embeddings
Vision
Audio
Function Calling
JSON Mode
Compliance
SOC 2
HIPAA
GDPR
Pricing Model
per tokenDetails
Models 3
API Base
https://inference.coralbricks.ai/v1 LLM Inference
Model Catalog (3)
| Model | Type | Input $/1M | Output $/1M | Context | Speed | Status |
|---|---|---|---|---|---|---|
| DeepSeek V4.1 Flash DeepSeek MXFP4 | llm | $0.300 | $1.20 | 1M | — | |
| GLM 5.3 Z.ai NVFP4 | llm | $1.12 | $4.40 | 1M | — | |
| GLM 5.3 Flash Z.ai NVFP4 | llm | $0.150 | $0.500 | 1M | — |