CO

Coral Bricks

Serverless

Fast inference for coding and research agents: open models (GLM 5.3, GLM 5.3 Flash, DeepSeek V4.1 Flash) served FP4-quantized on B200/B300 GPUs behind an OpenAI-compatible API with Chat Completions, Responses (server-side state) and background mode for long-running agents. Up to ~469 output tok/s, 1M-token context, cached input billed at $0. Prepaid per-token billing; private endpoints available.

Features

OpenAI Compatible
Streaming
Batching
Fine-tuning
Embeddings
Vision
Audio
Function Calling
JSON Mode

Compliance

SOC 2
HIPAA
GDPR

Pricing Model

per token

Details

Models 3
API Base https://inference.coralbricks.ai/v1
LLM Inference

Model Catalog (3)

Model Type Input $/1M Output $/1M Context Speed Status
DeepSeek V4.1 Flash
DeepSeek
MXFP4
llm $0.300 $1.20 1M —
GLM 5.3
Z.ai
NVFP4
llm $1.12 $4.40 1M —
GLM 5.3 Flash
Z.ai
NVFP4
llm $0.150 $0.500 1M —