Editorial flat-vector illustration for a technology blog hero about model pricing: a large rounded model chip in deep blue with a soft cyan glow, a paper price tag clipped to its corner, three token streams flowing into a circular price gauge, and a slim calendar sheet sliding into frame, on a soft light-blue background with generous empty space in the lower third. No text.
Guides & Insights

DeepSeek V4 Pro API Pricing: $0.73/$2.18 Off-Peak, and Why the September 14 Shutdown Was Called Off

Author

Rowan Sterling

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

DeepSeek V4 Pro is not being retired. On September 11, 2026, Deep​Seek said it will keep serving the model through its API after September 14, with the billing method unchanged — reversing a plan announced two days earlier under which every deepseek-v4-pro request would have been rerouted to DeepSeek V4.1 Flash and billed at Flash rates. For anyone whose prompts, temperature settings and tool-call schemas are tuned against V4 Pro, that is the fact that matters this week: the model you built against stays where it is, and the migration you may have scheduled for Monday does not have to happen.

The price is a different story, and this page had it wrong for a month. DeepSeek V4 Pro no longer costs $0.44/$0.88 per million tokens. Off-peak it bills at $0.726 input and $2.178 output per 1M tokens on OrcaRouter, doubling to $1.452/$4.356 inside the peak windows, with cache reads at $0.024. Weekends are now entirely off-peak. And the vendor's own rate card — the one at Deep​Seek's API documentation — carries the continuation notice as a footnote to the same price table, which is about as close to a primary source as a pricing page gets.

What Deep​Seek actually said, and what it undoes

The timeline matters, because two of the three announcements were reversals of each other and the middle one is the one people remember.

• September 9 — Deep​Seek told users that until V4.1 Pro launches, requests to V4 Pro would be routed to DeepSeek V4.1 Flash and billed at V4.1 Flash rates.

• September 10 — V4.1 Flash shipped, and Deep​Seek moved the switch to a fixed moment: 12:00 Beijing time on September 14, 2026, after which V4 Pro would go offline and its traffic would be served by Flash.

• September 11 — Deep​Seek reversed. Its wording, now carried verbatim as a footnote on the Models & Pricing page of its own API docs: "In response to user demand, we have decided to continue providing API services for DeepSeek V4 Pro after September 14, 2026, with the billing method remaining unchanged. We will provide further notice should there be any changes."

That is the vendor speaking about its own product, and it is the strongest sourcing available here — but read what it does and does not settle. It settles continuity and billing: V4 Pro stays, at the rates already in force. It does not promise a timeframe. We will provide further notice is the whole of the commitment, and the same sentence leaves the door open for the plan to return. If you are making a multi-quarter commitment, plan for the model you can move off rather than the one you can keep.

The backlash that forced the reversal is worth understanding, because it maps onto a decision you may be about to make yourself. Developers objected not to V4.1 Flash — which Deep​Seek says beats V4 Pro on performance, price, speed and total task time — but to the automatic substitution. Swapping the model behind a stable model ID changes behaviour without changing the code: a tuned prompt, a working temperature, a tool-call schema that parses are all properties of a specific checkpoint. Coverage in CLS (科创板日报), 36Kr, IT之家 and ifeng, and in English at TheBlockBeats, describes Deep​Seek first delaying, then dropping the plan; Deep​Seek's own docs show the contrast neatly, because the legacy deepseek-v4-flash names were retired and now serve V4.1 Flash, and that switch stuck. The Pro switch did not.

The price today, checked against the vendor's own rate card

Price card titled 'DeepSeek V4 Pro — price per 1M tokens', sourced 'OrcaRouter directory, checked 2026-08-15'. Three highlight cells read Input $0.442, Output $0.884 and Cache read $0.060 USD at zero markup. A row reads DeepSeek official rate (CNY, today): ¥3 in, ¥6 out, cache hit ¥0.025. Two columns show the August 17 change: off-peak input ¥4.5, output ¥13.5, cache hit ¥0.15; peak input ¥9, output ¥27, cache hit ¥0.30. Footer: peak output is 4.5x today's ¥6, off-peak is half of peak, peak windows Beijing weekdays 09:00-12:00 and 14:00-18:00.

Two figures matter, and they are the same list price in two currencies.

• Off-peak (every hour outside the peak windows below) — input $0.726 per 1M on a cache miss, or ¥4.5; output $2.178, or ¥13.5; cache-hit input $0.024, or ¥0.15.

• Peak — exactly double. Input $1.452 / ¥9, output $4.356 / ¥27, cache-hit input $0.048 / ¥0.30.

• Deep​Seek's own dollar conversion — the vendor's docs print the same yuan list as $0.66 input / $1.98 output off-peak and $1.32 / $3.96 at peak, with cache hits at $0.022 and $0.044. The gap to OrcaRouter's dollar figures is the FX rate each side converts at, not an added token fee: OrcaRouter passes the provider rate through with no markup, so the number on the listing is the number you pay.

• No free tier for V4 Pro. A free Flash tier exists on the directory, but the flagship is paid-only.

• The spec sheet, for context — V4 Pro is an open-weights model under an MIT licence, 1.6T total parameters with 49B active per token, a 1M-token context window and up to 384K output tokens, at a concurrency limit of 500. Artificial Analysis lists 72.3 output tokens per second.

The model version behind all of this is DeepSeek-V4-Pro-0813, unchanged since the August 13 general availability release — which is the point. Nothing about the checkpoint moved on September 11; only the decision about its future did.

Peak, off-peak, and the weekend rule that changes the math

The peak/off-peak scheme the earlier version of this page described as arriving on August 17 is now a month old, and it has been amended once since.

• Peak windows — 01:00–04:00 and 06:00–10:00 UTC, Monday through Friday. In Beijing time that is 09:00–12:00 and 14:00–18:00, Deep​Seek's own working day. Off-peak rates are exactly half of peak, not a discount band on top of them.

• Weekends are entirely off-peak. Effective 00:00 Beijing time on August 23, 2026, Deep​Seek stopped applying peak rates on Saturday and Sunday. If your batch work can move, moving it to the weekend halves the bill — and for teams serving users in the Americas, most of the working week already lands off-peak without any scheduling at all.

• What the change cost, against the ¥3/¥6 the model launched with on August 13 — today's off-peak output is 2.25× the launch rate and peak output is 4.5×; peak cache-hit input is 12× the old ¥0.025.

There is an asymmetry here that explains why the migration question keeps coming up even after the shutdown was cancelled. V4 Pro's price went up in August while its successor's went down: Deep​Seek cut Flash-series API prices by up to 60% effective September 10. The pressure to move was never really about the September 14 notice; the notice just made it urgent.

What a request actually costs

Three worked examples at today's off-peak rates. Every one of them doubles inside a peak window.

• 100K input + 50K output — $0.0726 + $0.1089 ≈ $0.18.

• 1M input + 1M output — $0.726 + $2.178 = $2.90. That same call is $5.81 at peak, and it was $1.32 in mid-August.

• Agentic session, 500K input + 200K output — $0.363 + $0.4356 ≈ $0.80.

The monthly view is on the model page and is the honest test of whether the old headline still describes your bill: at 10M tokens a month with a 70% input share, the cost calculator lands on $11.62, or about $9.16 with prompt caching switched on. If your budget was built against $0.44, it now covers about 60% of the bill.

Cache is the lever nobody puts in the table

Prompt caching got more valuable, not less, when the rates went up — and the direction of the two numbers is the reason. On OrcaRouter a cache read now bills at $0.024 per 1M tokens against $0.726 uncached, a 96.7% discount on any input you resend. In the previous version of this page the cache rate was $0.060 against a $0.442 input price, an 86% discount. The cache-read price fell by more than half while the uncached input price rose by 64%, so the effective gap between a cached and an uncached token widened considerably.

For an agent that re-sends a large system prompt and a long context on every turn, cached input is the difference between a bill that scales with conversation length and one that does not. The practical rule is unchanged and worth repeating: put the stable system prompt and reference material first and the volatile content last, so the prefix you resend actually hits the cache. On this model that single structural choice is worth more than any routing or scheduling decision below it.

How V4 Pro pricing compares to the field

Comparison table card titled 'USD per 1M tokens — the field', sourced to the OrcaRouter directory checked 2026-08-15, with an input/output column pair per model: DeepSeek V4 Flash $0.15/$0.29; DeepSeek V4 Pro $0.44/$0.88 highlighted with a 'THIS PAGE' tag; MiniMax M3 $0.30/$1.20; Gemini 3.6 Flash $1.50/$7.50; Grok 4.5 $2.00/$6.00; Qwen 3.8 Max $2.00/$6.00; GPT-5.6 Terra $2.00/$12.00; Kimi K3 $3.00/$15.00; Claude Opus 5 $5.00/$25.00; GPT-5.6 Sol $5.00/$30.00; Claude Fable 5 $10.00/$50.00. Footer: V4 Pro is the cheapest 1M-context flagship, roughly 23x cheaper on input and 57x cheaper on output than Claude Fable 5.

Directory list prices for input/output per 1M tokens, checked 2026-09-12, set against V4 Pro's off-peak $0.726/$2.178. Two of these have moved since this page was first published, and one of them changes the conclusion.

• Claude Fable 5 — $10 / $50. V4 Pro is 13.8× cheaper on input and 23× on output.

• Claude Opus 5 — $5 / $25. 6.9× / 11.5×.

• GPT-5.6 Sol — $4 / $20 up to 272K of context, $8 / $30 beyond it. 5.5× / 9.2× on the shorter tier, 11× / 13.8× on the longer one.

• Kimi K3 — $3 / $15. 4.1× / 6.9×.

• Qwen3.8 Max — $2 / $6. 2.8× / 2.8×.

• Grok 4.5 — $2 / $6. 2.8× / 2.8×.

• Gemini 3.6 Flash — $0.75 / $3.75. Roughly parity on input at 1.03×, and 1.7× on output. Its price has halved since this page was written.

• MiniMax M3 — $0.30 / $1.20. Cheaper than V4 Pro on both lines now: V4 Pro costs 2.4× more on input and 1.8× more on output, with the same 1M-token window.

• DeepSeek V4 Flash — $0.242 / $0.726 on the directory. Roughly 3× cheaper on both lines, same 1M context.

So the honest framing has changed, and this page should say so rather than keep the old superlative. DeepSeek V4 Pro is no longer the cheapest 1M-context model on the board: Gemini 3.6 Flash and MiniMax M3 both undercut it, and MiniMax M3 does it while matching the context window. What V4 Pro still is, by more than an order of magnitude, is the cheapest thing in the frontier tier — every Western flagship-class model on that list costs multiples of it on both lines. If the question is "cheapest possible 1M-context call", the answer is no longer this model. If the question is "cheapest model that behaves like a frontier model", for now it still is.

When V4 Pro is the wrong choice

• If your traffic is heaviest during Beijing daytime. Peak hours double the rate, and the peak windows sit exactly on China's working day. If your users are in the Americas, most of your traffic already lands off-peak, which is a real structural advantage. If you are building for users in China, model the peak bill explicitly before you commit — and check whether the weekend rule helps you move anything.

• If your prompts are short. Below roughly 10K tokens of context the flagship tier is wasted capacity. DeepSeek V4 Flash at $0.242/$0.726 on the directory handles the same 1M context for about a third of the price, and for genuinely short tasks a cheaper model still wins.

• If you need vision. V4 Pro is text-only — no image or audio input — and Artificial Analysis lists its input modality as text. Gemini 3.6 Flash or Claude Fable 5 are the better fit if the modality is the point. Deep​Seek's own API is also worth noting here: it now serves its legacy flash model names with V4.1 Flash, which does accept images.

• If you are buying benchmark-topping raw score rather than price-per-context. On Artificial Analysis's current scale, V4 Pro scores 36 on the Intelligence Index — 8th of 113 models at the time of writing — and 69 on the Coding index as carried in the OrcaRouter directory. That is competitive, not leading: Claude Fable 5, Claude Opus 5 and GPT-5.6 Sol all outscore it on coding. V4 Pro's pitch is frontier-adjacent capability at a fraction of the per-token price, not the top of every chart.

How to call it on OrcaRouter

Screenshot of the OrcaRouter model page for DeepSeek V4 Pro (model ID deepseek/deepseek-v4-pro): FLAGSHIP FEATURED badge, '1M tokens' context and 384K max output, input and output price per 1M tokens (output shown at $0.88), /v1/chat/completions and /v1/responses endpoints, and an OpenAI SDK code sample pointing at api.orcarouter.ai/v1.

DeepSeek V4 Pro is hosted on OrcaRouter under the model ID deepseek/deepseek-v4-pro, served through the Open​AI-compatible endpoint at api.orcarouter.ai/v1 at the provider rate with zero markup. Migration from an existing Open​AI client is a base-URL and model-ID change and nothing else.

The reversal changes what you should do about it. If you had a September 14 migration on the calendar — repointing from V4 Pro to DeepSeek V4.1 Flash because Pro was going away — you can take it off the calendar, or at least downgrade it from a deadline to a test. Both models sit behind one OrcaRouter key, so comparing them is a model-ID change in a request rather than a second contract and a second SDK; and if you do want V4.1 Flash evaluated on real traffic without betting a production path on it, automatic failover is the low-risk way to give it live requests while V4 Pro stays as the fallback. That is the practical shape of the September 11 decision: the choice stays yours, and it stays reversible.

The honesty note that matters for a pricing page: we route to the model, we don't set its price. The $0.726/$2.178 off-peak figures are Deep​Seek's rates passed through, and they move when Deep​Seek moves them — which is exactly why this page is date-stamped, and why the version you are reading was rewritten rather than appended to.

Sources and date

Prices and specifications on this page were re-verified on 2026-09-12. Dollar figures for DeepSeek V4 Pro and for every model in the comparison are read from the OrcaRouter model directory, checked the same day; the parameter count, licence, context window, modality and speed figures come from Artificial Analysis's model page for DeepSeek V4 Pro 0813. The yuan rates, the peak windows and the September 11 continuation notice come from Deep​Seek's own Models & Pricing documentation, which carries the statement verbatim as a footnote to the price table. The three-step announcement sequence — the September 9 routing notice, the September 10 delay to 12:00 Beijing time on September 14, and the September 11 reversal — is reported by CLS (科创板日报), 36Kr, IT之家 and ifeng, and in English by TheBlockBeats. The August 23 weekend-billing rule and the September 10 Flash price cut come from Deep​Seek's announcements as covered by IT之家, The Paper and 21st Century Business Herald.

One closing note on trustworthiness, because this specific price has moved three times in a month — launch rates on August 13, peak/off-peak on August 17, the weekend amendment on August 23. Treat any "V4 Pro pricing" page without a date on it as unusable. The earlier version of this one quoted $0.44/$0.88 and an Artificial Analysis Intelligence index of 44.3; both were correct when they were written, and neither survived the month.

Compared in this article4

Detected from this article · Benchmarks: Artificial Analysis · updated daily