NE

NeuralWatt

Serverless

Hosted, OpenAI-compatible inference with energy-based pricing: pay a flat $5.00/kWh for actual GPU energy consumed (up to 95% cheaper on efficient models) instead of per-token, with real per-request energy metrics on every response. Standard per-token pricing and kWh-based monthly subscriptions ($20/$50/$100) are also available. Powered by Neuralwatt Optimize, also available for self-hosted vLLM via Neuralwatt Deploy.

Features

OpenAI Compatible
Streaming
Batching
Fine-tuning
Embeddings
Vision
Audio
Function Calling
JSON Mode

Compliance

SOC 2
HIPAA
GDPR

Pricing Model

per token

Details

Models 13
API Base https://api.neuralwatt.com/v1
LLM Inference

Model Catalog (13)

Model Type Input $/1M Output $/1M Context Speed Status
DeepSeek V4 Flash
DeepSeek
llm $0.140 $0.280 1.0M —
DeepSeek V4.1 Flash
DeepSeek
llm $0.150 $0.600 1.0M —
GLM 5.2
Z.ai
llm $1.45 $4.50 1.0M —
GLM 5.3
Z.ai
llm $1.45 $4.50 1.0M —
GLM 5.3 Flash
Z.ai
llm $0.150 $0.500 1.0M —
Gemma 4 31B
Google · 27B
llm $0.144 $0.420 262k —
Kimi K2.6
Moonshot AI · 1T MoE (32B active)
llm $0.690 $3.22 262k —
Kimi K2.7 Code
Moonshot AI · 1T MoE (32B active)
code $0.950 $4.00 262k —
Kimi K3
Moonshot AI
llm $3.00 $15.00 1.0M —
MiMo V2.6 Pro
Xiaomi
llm $0.870 $1.74 1.0M —
Qwen3.5 397B A17B
Alibaba
llm $0.690 $4.14 262k —
Qwen3.6 35B A3B
Alibaba
llm $0.290 $1.15 262k —
Qwen3.8 27B
Alibaba
llm $0.450 $3.20 262k —

NeuralWatt API FAQ

Is the NeuralWatt API free?

Not as far as we track: NeuralWatt has no free API tier on record. The cheapest option we track is DeepSeek V4 Flash at $0.14 in / $0.28 out per 1M tokens.

What is the cheapest NeuralWatt model?

DeepSeek V4 Flash at $0.14 in / $0.28 out per 1M tokens, based on the 13 models we track for NeuralWatt.

How many models does NeuralWatt offer?

We track 13 NeuralWatt models: 12 LLM, 1 coding.

Is the NeuralWatt API OpenAI-compatible?

Yes. NeuralWatt exposes an OpenAI-compatible endpoint, so you can use the OpenAI SDKs by changing the base URL and API key.