Pricing
The API bills per token. The table below lists the public rate for every callable model in US dollars, split into input, cached input and output, per 1M tokens; your balance is debited in Stars at 1 ★ = $0.01.
Billing unit
Rates are quoted in US dollars; balances settle in Stars, and one Star is fixed at $0.01. The Star figure returned with a response is exactly what your balance is debited, to four decimal places.
1 ★ = $0.01
Formula
Input tokens are split into a cache-miss part and a cache-hit part and priced separately at their dollar rates; output tokens are priced on their own. The three add up to the raw cost in dollars; dividing by the $0.01 Star value gives the Stars debited.
usd = (input_miss / 1e6) * in_rate_usd
+ (input_cached / 1e6) * cache_rate_usd
+ (output / 1e6) * out_rate_usd
input_miss = max(0, prompt_tokens - cached_tokens)
stars = usd / 0.01
charged = max(1, round(stars, 4)) # 单位 ★Every successful request carries a 1 Star minimum. If the raw cost lands below 1 Star you are charged 1 Star; above that you are charged the exact amount. A response with zero input and zero output tokens is not charged at all.
Public rates ($ / 1M tokens)
The model name is the exact string to put in the model field, identical to the id returned by GET /v1/models. Rates follow upstream changes; this page is the source of truth.
| Model | Direct / Third-party | Input | Cached input | Output | Long context |
|---|---|---|---|---|---|
| Loading live rates… | |||||
Image / Video / Audio (per call / per second)
Billed per call or per second. Final price = max(minimum charge, base × size multiplier × quality multiplier); each tier is pre-computed on the right.
| Model | Type | Direct / Third-party | Base | Size / resolution tiers | Quality tiers | Min charge |
|---|---|---|---|---|---|---|
| Loading live rates… | ||||||
Worked examples
Both token counts and Star amounts below are copied from real production calls, not estimated.
A ~64k-token long-context request where almost all of the input misses the cache.
model deepseek-v4-flash prompt_tokens 64012 (cached_tokens 2560) completion 1 input_miss 61452 / 1e6 * $0.28 = $0.01720656 input_cached 2560 / 1e6 * $0.056 = $0.00014336 output 1 / 1e6 * $0.56 = $0.00000056 ------------------------------------------------ usd $0.01735 X-Stars-Cost 0.01735 / 0.01 = 1.735 ★
The identical prompt sent again straight away; upstream reports all 64k input tokens as cache hits.
model deepseek-v4-flash (same prompt, resent) prompt_tokens 64012 (cached_tokens 64000) completion 1 input_miss 12 / 1e6 * $0.28 = $0.00000336 input_cached 64000 / 1e6 * $0.056 = $0.00358400 output 1 / 1e6 * $0.56 = $0.00000056 ------------------------------------------------ usd $0.003588 ( = 0.3588 ★ ) X-Stars-Cost max(1, 0.3588) = 1 ★ ( = $0.01 )
The cache cuts the raw cost from 1.735 Stars to 0.3588 Stars — about 80% off — but that lands under the 1 Star minimum, so this call is charged 1 Star. The saving shows up on the bill once contexts are long and reuse is heavy.
Reading what a call cost
Every successful response carries X-Stars-Cost, the exact number of Stars debited, plus X-Request-Id for reconciliation and support tickets.
X-Stars-Cost: 1.735 X-Request-Id: req_f08c8151fc95438f9c5c524511a16806
On /v1/messages and /anthropic/v1/messages the cost is also written straight into the usage object:
"usage": {
"input_tokens": 2203,
"output_tokens": 1,
"cost_stars": 1.5026,
"cost_usd": 0.015
}Billing rules
- Only successful responses are billed. Auth failures, unknown models, invalid parameters and rate-limit rejections cost nothing.
- Streaming is billed from the usage the upstream finally reports, at the same rates as non-streaming. A stream cut short is billed for the tokens actually produced.
- A worst-case amount is held before the call and reconciled against real usage the moment upstream reports it, so nothing stays reserved after the request finishes.
- An insufficient balance returns 402; there is no overdraft. Per-request detail is available on the usage page in the console, searchable by request id.