AI Tools Review
DeepSeek V4 Pro GA Review: 1.6T MoE, Tested

DeepSeek V4 Pro GA Review: 1.6T MoE, Tested

21 July 2026

Quick answer:

DeepSeek V4 went fully General Available on 19 July 2026, three months after its 24 April open-weight preview. It ships as two MIT-licensed Mixture-of-Experts models — V4 Pro (1.6 trillion total / 49 billion active parameters) and V4 Flash (284 billion / 13 billion active) — both with a 1-million-token context window, both compatible with OpenAI and Anthropic API formats, and both cheap: $0.435/$0.87 per million input/output tokens for Pro off-peak, doubling during Beijing business hours under a new peak-pricing scheme. DeepSeek's own chart shows V4 Pro roughly matching an older Claude model (Opus 4.6, not Fable 5) on SWE-bench Verified and clearly losing to GPT-5.4-xHigh on agentic tool-use tasks. Artificial Analysis currently scores it 44 on the Intelligence Index, and both tiers answer instead of declining in 94-96% of cases even when they shouldn't know the answer.

Two days after Alibaba's Qwen 3.8 Max preview and a week after Moonshot's Kimi K3 rattled semiconductor stocks, a much quieter but arguably more consequential release landed: DeepSeek V4 went General Available. Unlike its two rivals, this wasn't a splashy stage announcement with an unverified "second only to Fable 5" claim — it was a scheduled infrastructure change, flagged weeks in advance, that swapped DeepSeek's production API over to the V4-Pro and V4-Flash model IDs and introduced peak-time pricing for the first time.

It arrived into a YouTube news cycle already primed to overstate it. Videos with titles like "DeepSeek V4 Pro GA Leaks Just Destroyed Claude Fable?" pointed to an 80.6% vs 80.8% SWE-bench comparison as evidence DeepSeek had all but caught the Western frontier. The number is real — it comes straight from DeepSeek's own official release chart — but the comparison it belongs to is not the one the thumbnails implied. This article walks through what actually shipped, what DeepSeek's own data shows, what independent evaluators found separately, and exactly where that viral claim breaks down.

Note: this analysis draws on DeepSeek's official API documentation and Hugging Face model cards, Artificial Analysis's published evaluations and Intelligence Index scoring, the US National Institute of Standards and Technology's CAISI capability assessment, and this site's own previously verified pricing and benchmark data for Claude Fable 5, GPT-5.6 Sol and Kimi K3. Prices are billed in US dollars per DeepSeek's API pricing page.

Executive summary

  • General Availability on 19 July 2026 — three months after the 24 April open-weight preview, with legacy deepseek-chat and deepseek-reasoner endpoints retiring on 24 July 2026.
  • Two MIT-licensed MoE tiers: V4 Pro (1.6T total / 49B active parameters) and V4 Flash (284B / 13B active), both with a 1M-token context window — an 8x jump over V3.2's 128K.
  • New hybrid attention architecture (Compressed Sparse Attention + Heavily Compressed Attention) cuts long-context inference cost sharply versus V3.2, per DeepSeek's own efficiency charts.
  • Peak-time pricing, a first for DeepSeek: rates double during Beijing business hours (9am-12pm, 2pm-6pm), pushing V4 Pro from $0.435/$0.87 to $0.87/$1.74 per million input/output tokens.
  • The viral "beats Claude Fable 5" claim doesn't hold up. DeepSeek's own chart compares V4 Pro to Claude Opus 4.6 — not Fable 5 — and V4 Pro loses outright to GPT-5.4-xHigh on both agentic benchmarks in the same chart.
  • Artificial Analysis currently scores V4 Pro 44 on the Intelligence Index, ranked third among open-weight models, down from 52 (second place) at its April preview as newer rivals launched.
  • A genuine reliability gap: both tiers answer instead of abstaining in 94-96% of cases even when uncertain, and NIST's CAISI found real-world capability roughly eight months behind DeepSeek's own reported figures.

From preview to GA: the timeline

DeepSeek V4 didn't appear out of nowhere in July. The company shipped an open-weight preview on 24 April 2026 — DeepSeek-V4-Pro and DeepSeek-V4-Flash, both MIT-licensed, both published to Hugging Face on day one. That preview is what Artificial Analysis benchmarked first, scoring V4 Pro (Max reasoning effort) at 52 on the Intelligence Index, good for second place among open-weight reasoning models at the time, trailing only Moonshot's since-superseded Kimi K2.6 (54).

The path to General Availability was telegraphed well in advance. TechNode reported in late June that DeepSeek planned a mid-July official release alongside a new peak-time API pricing mechanism — the first time the company had introduced time-based surcharging. Model IDs for the GA build (deepseek-v4-pro-202606 and deepseek-v4-flash-202605) surfaced in leaks attributed to AI researcher @teortaxesTex on 4 July, roughly two weeks ahead of the actual switch. DeepSeek's official API documentation confirms the GA release landed on 19 July 2026, with the legacy deepseek-chat and deepseek-reasoner endpoints set to be fully retired on 24 July — a hard migration deadline for anyone still pointing at the old model names.

The timing put V4's GA squarely in the middle of the most crowded fortnight open-weight Chinese AI has had all year: Z.AI's GLM 5.2 in mid-June, Moonshot's Kimi K3 on 17 July, and Alibaba's Qwen 3.8 Max preview on 19 July — the same day DeepSeek's own GA quietly went live. Compared to the stage announcements from Moonshot and Alibaba, DeepSeek's release was almost understated: an API documentation update and a pricing change, not a keynote.

Architecture and training

Both V4 tiers are sparse Mixture-of-Experts models. V4 Pro carries 1.6 trillion total parameters with 49 billion active per token; V4 Flash carries 284 billion total with 13 billion active. DeepSeek's Hugging Face model card describes three specific architectural changes over V3.2: a hybrid attention mechanism combining Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA), Manifold-Constrained Hyper-Connections to strengthen residual signal propagation through the network, and a switch to the Muon optimizer for faster, more stable training convergence. Both models were pre-trained on more than 32 trillion tokens.

Diagram of the DeepSeek-V4 series transformer architecture, showing the embedding layer, stacked transformer blocks with CSA/HCA attention and DeepSeekMoE routing, residual and pre/post-block mixing, and the multi-token prediction (MTP) output heads
The DeepSeek-V4 series architecture: input tokens pass through stacked transformer blocks combining CSA/HCA attention with DeepSeekMoE routing, residual mixing, and multi-token prediction heads. Source: Wikimedia Commons, based on DeepSeek's published technical documentation.

The practical payoff of the hybrid attention change shows up at long context. DeepSeek's own efficiency charts, published alongside the GA release, show V4 Pro requiring roughly 3.7x lower single-token FLOPs than V3.2 at moderate context lengths, widening to nearly 10x lower as sequences approach the full 1M-token window — with V4 Flash cutting inference cost further still. The accumulated KV cache tells the same story: V4 Pro holds a 9.5x-to-13.7x smaller cache than V3.2 at comparable sequence lengths, which is what actually makes serving a 1M-token context window financially viable rather than a marketing number nobody uses in production.

Both tiers support two operating modes — Thinking and Non-Thinking — selectable per request, and both are compatible with OpenAI's ChatCompletions format and Anthropic's Messages API, meaning existing coding agents built against either ecosystem (Cursor, Cline, Claude Code, Codex-style harnesses) can point at DeepSeek V4 without rebuilding their integration layer. The 1M-token context window applies uniformly across both models and is now the default across DeepSeek's official consumer and API surfaces, not a separate paid tier.

Capabilities deep dive

Coding and agentic tool use

DeepSeek positions V4 Pro primarily as a coding and agentic model, and the strongest real result in its own release chart is Codeforces competitive-programming rating, where V4-Pro-Max posts 3,206 — ahead of GPT-5.4-xHigh's 3,168 and Gemini-3.1-Pro-High's 3,052, the highest of the three in that specific comparison. On the Apex Shortlist reasoning benchmark it scores 90.2%, again leading the chart. But the picture is not uniformly positive: on the two benchmarks DeepSeek's own chart labels "Agentic Capabilities" — Terminal-Bench 2.0 and Toolathlon, both of which test multi-step tool use rather than single-shot code generation — V4-Pro-Max is beaten outright by GPT-5.4-xHigh (75.1% vs 67.9% on Terminal-Bench 2.0; 54.6% vs 51.8% on Toolathlon). That is a genuine, DeepSeek-published gap, not a hostile third-party finding.

Knowledge and reasoning

On SimpleQA Verified, a factual-knowledge benchmark, V4-Pro-Max scores 57.9% against Claude-Opus-4.6-Max's 46.2% and GPT-5.4-xHigh's 45.3% — though Gemini-3.1-Pro-High leads the whole group at 75.6%. On Humanity's Last Exam (HLE), the hardest general-knowledge benchmark in the chart, V4-Pro-Max actually trails all three rivals at 37.7%, against 40.0% (Opus), 39.8% (GPT) and 44.4% (Gemini). Taken together, DeepSeek's own numbers describe a model that is genuinely strong at structured, verifiable tasks — competitive programming, factual recall with clear answers — and comparatively weaker at the kind of open-ended, expert-level reasoning HLE is designed to probe.

1M-token context in practice

The context window expansion from V3.2's 128K to 1M tokens is the single biggest usability change in this release, and — per the architecture section above — it is backed by a real efficiency improvement rather than just a bigger number on a spec sheet. For anyone running long-document analysis, large-codebase agentic work, or extended multi-turn agent sessions, that combination (long context plus a meaningfully lower KV-cache cost at that length) is arguably a bigger practical upgrade than any single benchmark score in this release.

Benchmarks: DeepSeek's chart vs independent scoring

Here is the chart at the centre of the viral claim, published directly by DeepSeek alongside the GA release and downloaded from DeepSeek's own Hugging Face model card:

BenchmarkDS-V4-Pro-MaxClaude-Opus-4.6-MaxGPT-5.4-xHighGemini-3.1-Pro-High
SimpleQA Verified57.9%46.2%45.3%75.6%
HLE37.7%40.0%39.8%44.4%
Apex Shortlist90.2%85.9%78.1%89.1%
Codeforces (rating)3,2063,1683,052
SWE-bench Verified80.6%80.8%80.6%
Terminal-Bench 2.067.9%65.4%75.1%68.5%
Toolathlon51.8%47.2%54.6%48.8%

Table transcribed directly from DeepSeek's official GA release chart (Hugging Face model card, deepseek-ai/DeepSeek-V4-Pro). Bold indicates the top score in each row. "—" means no bar was shown for that model on that benchmark in DeepSeek's chart.

This is the source of the "80.6 vs 80.8" number driving the "DeepSeek V4 destroyed Claude Fable" framing on YouTube. Read in context, three things are true simultaneously. First, the comparison is real and it is DeepSeek's own data. Second, the Claude model in that comparison is Claude Opus 4.6, not Claude Fable 5 — Anthropic's current Mythos-class flagship isn't in this chart at all. Third, on the specific benchmark used, V4-Pro-Max doesn't actually beat Opus-4.6-Max — it scores 0.2 points behind it (80.6% vs 80.8%), a rounding-error gap in either direction, not a win.

For an actual Fable 5 comparison, this site's own previously verified figures are the right reference point: Claude Fable 5 tops the Artificial Analysis Intelligence Index at 60 — the highest score of any model currently tracked — and leads SWE-bench Pro specifically at 80.3%. SWE-bench Pro is a separate, harder benchmark suite from SWE-bench Verified, and it does not appear anywhere in DeepSeek's release chart, so there is no direct, apples-to-apples DeepSeek V4 vs Fable 5 number available from either company's own published data at the time of writing. Anyone citing an "80.6 vs 80.8, DeepSeek nearly matches Fable 5" framing is, whether they realise it or not, comparing V4 to a different Claude generation on a benchmark Fable 5 hasn't published a score for.

Independent scoring tells a more modest story than DeepSeek's chart on its own. Artificial Analysis's current Intelligence Index — a broader composite across knowledge, reasoning, coding and agentic evaluations, recalculated as new models enter the field — places V4 Pro at 44, third among open-weight models, well below Fable 5's 60 and Kimi K3's 57. That 44 is also down from the 52 (second place among open-weight models) V4 Pro scored at its April preview, not because the model changed, but because Kimi K3, GLM 5.2 and other rivals have since launched and pushed the field forward — a reminder that a strong result on launch day tends to compress as the field catches up.

Independent evaluation: NIST CAISI

DeepSeek has not published a Western-style model card covering safety evaluation, red-teaming, or dual-use risk assessment for V4 in the way Anthropic, OpenAI or Google DeepMind do for comparable releases — a genuine transparency gap worth naming plainly rather than glossing over. The closest thing to independent, methodical scrutiny currently public comes from an unlikely source: the US National Institute of Standards and Technology's Center for AI Standards and Innovation (CAISI), which ran its own capability evaluation of DeepSeek V4 Pro using a non-public benchmark suite spanning cyber, software engineering, natural sciences, abstract reasoning and mathematics.

CAISI's headline finding: DeepSeek V4's real capabilities lag the frontier by roughly eight months, and using CAISI's own benchmark suite, V4 performed comparably to GPT-5 — a model released about eight months prior to CAISI's evaluation — rather than the more recent frontier tier DeepSeek's own published figures implied. That is a materially different conclusion from the "near-frontier" positioning in DeepSeek's own marketing, and it comes from a government evaluator with no commercial stake in either DeepSeek or its Western competitors.

CAISI's results weren't uniformly unflattering, though. On mathematics specifically, DeepSeek V4 scored strongly — 96-100% on some of CAISI's tests — while performing more weakly on abstract reasoning, notably 46% on the semi-private ARC-AGI-2 set. On cost, CAISI found DeepSeek V4 was cheaper than GPT-5.4 mini on five of seven benchmarks tested, with savings ranging from 53% cheaper up to 41% more expensive depending on the specific task — a more mixed and specific picture than the blanket "an order of magnitude cheaper" framing that often accompanies DeepSeek coverage. CAISI's report is explicitly scoped to capability and cost; it does not present dedicated safety or security findings, so it should not be read as a clean bill of health on those fronts either way.

Honesty and calibration: the hallucination problem

The most consistently flagged weakness across independent testing of DeepSeek V4 isn't a benchmark score — it's a behavioural pattern. Artificial Analysis's evaluation found that V4 Pro answers rather than declining or hedging in 94% of cases specifically designed to test whether a model should say "I don't know," and V4 Flash does so in 96% of cases — among the highest "always answers anyway" rates Artificial Analysis has recorded for a current-generation model. In practice, that means both tiers will produce a confident, well-formatted answer to a question they don't actually have reliable grounding for, rather than flagging uncertainty — the textbook definition of a hallucination risk, and a harder problem to notice than a wrong benchmark number because it looks identical to a correct answer until you check it.

This pairs uncomfortably with the token-efficiency finding from the same evaluation: V4 Pro generated 190 million output tokens over Artificial Analysis's test suite (V4 Flash used 240 million), well above typical usage for models of comparable capability. A model that is both verbose and reluctant to hedge is a specific combination worth planning around — expect longer, more confident-sounding responses that need more verification, not fewer, particularly on open-ended factual questions where SimpleQA-style benchmarks reward a plausible guess as much as a correct one.

Real-world use vs the benchmark chart

Put the pieces together and DeepSeek V4 Pro reads as a model that is genuinely excellent at structured, checkable work — competitive programming, factual lookup with a clear right answer, cost-constrained inference at long context — and comparatively weaker at exactly the tasks where a wrong answer is expensive and hard to catch: open-ended reasoning (HLE), multi-step agentic tool use (Terminal-Bench, Toolathlon, both of which DeepSeek's own chart shows GPT-5.4-xHigh winning), and anything where the model needs to recognise the limits of its own knowledge rather than answer anyway.

That is a coherent, useful profile for a specific set of workloads — high-volume coding assistance, competitive-programming-style problem sets, cost-sensitive long-context document processing — and a poor fit for others, particularly autonomous multi-step agents operating without a human checking each step, or any use case where a confidently wrong answer is worse than a slow correct one.

Pricing and availability

ModelInput (off-peak)Output (off-peak)Output (peak, 9am-12pm / 2pm-6pm)
V4 Pro$0.435 / 1M tokens$0.87 / 1M tokens$1.74 / 1M tokens
V4 Flash$0.14 / 1M tokens$0.28 / 1M tokens$0.56 / 1M tokens

Cache hits are billed separately at $0.004 per million tokens — a roughly 99% discount versus a fresh input token — which matters a great deal for agentic workloads that repeatedly re-send large amounts of shared context. Peak pricing is new to this GA release: rates double during Beijing business hours (9am-12pm and 2pm-6pm daily), the first time DeepSeek has introduced time-based surcharging on its API. Teams running batch or non-interactive workloads have a real, mechanical incentive to shift jobs outside those windows.

Both models are accessible via deepseek-v4-pro and deepseek-v4-flash model IDs through DeepSeek's official API, with OpenAI ChatCompletions and Anthropic Messages API compatibility, and via open weights on Hugging Face for self-hosting (DSpark variants are also published for those wanting a distilled, lighter-weight option). The legacy deepseek-chat and deepseek-reasoner endpoints are being fully retired on 24 July 2026 — three days after this article's publication — so anyone still targeting those model names needs to migrate before then or lose API access outright.

On raw off-peak output pricing, V4 Pro undercuts every closed frontier competitor by a wide margin: roughly 17x cheaper than Claude Fable 5's effective per-task cost of $2.75 (per Artificial Analysis's task-cost methodology) and meaningfully below Kimi K3's published $15 per million output tokens. That gap narrows somewhat in practice given V4's verbosity — more output tokens per task partially offsets the lower per-token rate — but the headline pricing advantage is real and, for cost-sensitive high-volume use, substantial even after accounting for it.

Limitations

  • High hallucination risk. 94-96% "always answers anyway" rate on questions designed to test appropriate uncertainty — one of the highest rates Artificial Analysis has recorded for a current model.
  • Loses on agentic tool-use benchmarks. DeepSeek's own chart shows GPT-5.4-xHigh beating V4-Pro-Max on both Terminal-Bench 2.0 and Toolathlon.
  • Trails on open-ended expert reasoning. V4-Pro-Max scores lowest of the four models compared on Humanity's Last Exam (37.7%).
  • No Western-style safety/model card. No published red-teaming, dual-use risk assessment, or safety evaluation comparable to what Anthropic, OpenAI or Google DeepMind publish for equivalent releases.
  • Independent capability scoring is more modest than DeepSeek's framing. NIST CAISI found real-world capability roughly eight months behind the frontier, using its own non-public benchmark suite.
  • Verbose by default. 180-240 million output tokens across Artificial Analysis's evaluation suite versus a 96 million median, which adds real cost even at V4's low per-token rate.
  • Peak-time pricing is new and untested at scale. How predictably DeepSeek enforces and communicates the 2x peak surcharge in practice is not yet established this early into GA.

How DeepSeek V4 Pro compares

Against Kimi K3, DeepSeek V4 Pro trails on the current Artificial Analysis Intelligence Index (44 vs 57) but wins clearly on price — V4 Pro's $0.87 per million output tokens off-peak is a fraction of Kimi K3's $15. See the Kimi K3 vs Claude Fable 5 scorecard for how Kimi's own claims held up under independent scoring — a useful parallel case study for reading any Chinese open-weight launch's day-one claims with appropriate caution.

Against Qwen 3.8 Max, the comparison is currently one-sided in DeepSeek's favour on transparency if nothing else: V4 shipped with a real benchmark chart, a named MIT licence, and open weights on day one, where Qwen 3.8 launched as a closed preview with no benchmark table and only a promise of open weights "coming soon." Neither model currently has an independently verified score that would settle which is actually stronger.

Against the closed Western frontier — Claude Fable 5 and GPT-5.6 Sol — DeepSeek V4 Pro is not a frontier-tier competitor on Artificial Analysis's composite Intelligence Index (44 vs Fable 5's 60), but it is a genuinely strong value option for coding-heavy, cost-sensitive workloads where the gap in capability matters less than a 15-20x difference in per-token cost. For the current scored picture across all of these models side by side, see this site's live Benchmarks page.

Who should use it — and who should wait

Worth using now: teams with high-volume, cost-sensitive coding workloads — competitive-programming-style problem sets, structured code generation, long-document processing at 1M-token context — where V4 Pro's low per-token price and genuinely improved long-context efficiency deliver real savings, and where outputs get checked (tests, compilers, human review) rather than trusted blind. Self-hosters who want a capable, genuinely open-weight MIT-licensed model with no vendor lock-in are also well served, given the DSpark variants offer a lighter-weight option for those without hyperscaler-class hardware.

Better to wait or look elsewhere: anyone building autonomous multi-step agents without a human checking each step, given DeepSeek's own chart shows it losing to GPT-5.4-xHigh on exactly that category of benchmark; anyone relying on the model for open-ended factual questions where a confident wrong answer is costly, given the 94-96% "always answers anyway" rate; and anyone who needs a published, Western-style safety and red-teaming evaluation before deploying — that documentation simply doesn't exist yet for V4 the way it does for Fable 5, GPT-5.6 Sol or Gemini 3.1.

The bottom line

DeepSeek V4 Pro's General Availability release is a real, substantive upgrade — a genuinely open-weight, MIT-licensed model with a legitimately improved 1M-token context architecture, priced far below anything in the closed frontier tier, backed by DeepSeek's own published chart showing real wins on competitive programming and structured reasoning. None of that requires exaggeration to be a good story on its own.

But it doesn't "destroy" Claude Fable 5, and DeepSeek's own chart never claimed it did — that framing came from YouTube commentary reading a Claude-Opus-4.6 comparison as a Fable 5 one. The real, sourced picture is more interesting than the viral headline: a cheap, capable, genuinely open coding model with a specific and well-documented weak spot around agentic tool use and answering questions it shouldn't be confident about — worth adopting for the right workload, and worth double-checking on anything else.

Sources for this article include DeepSeek's official API documentation, the DeepSeek-V4-Pro Hugging Face model card, Artificial Analysis's published evaluation of V4 Pro and V4 Flash, and NIST CAISI's independent capability evaluation.

Last updated: 21 July 2026, two days after DeepSeek V4's General Availability release. This article will be revised if DeepSeek publishes a direct Fable 5 comparison or if independent evaluators update their scoring.

Free Guide

Get the free guide: Claude vs ChatGPT, Gemini & Grok

A 20-page playbook covering everything you need to choose and use the big four AI models in 2026, full cost and feature comparisons, what each is best (and worst) at, and how-tos for images, vectors, building a website, Claude Code and more.

Pop your email in to get it free
Preview of the free guide: Claude vs ChatGPT, Gemini and Grok, 2026 features, pricing and what-you-can-do comparison.

Frequently Asked Questions

Is DeepSeek V4 actually open source?
Yes, both tiers are open-weight under the MIT licence. DeepSeek-V4-Pro (1.6 trillion total parameters, 49 billion active) and DeepSeek-V4-Flash (284 billion total, 13 billion active) are published on Hugging Face with weights anyone can download, fine-tune or self-host, subject to having the hardware for a model this size. This carried over unchanged from the open preview DeepSeek shipped on 24 April 2026 through to the General Availability release on 19 July 2026 — DeepSeek did not close the weights for GA, unlike Alibaba's Qwen 3.8 Max or Moonshot's early Kimi releases.
Does DeepSeek V4 Pro really beat Claude Fable 5?
No, and DeepSeek's own official benchmark chart doesn't even make that claim. The comparison chart DeepSeek published alongside the GA release benchmarks V4-Pro-Max against Claude-Opus-4.6-Max, GPT-5.4-xHigh and Gemini-3.1-Pro-High — not Fable 5. On SWE-bench Verified, V4-Pro-Max scores 80.6% against Opus-4.6-Max's 80.8%, a near-tie with an older Claude model, not Anthropic's current flagship. Claude Fable 5 itself tops the Artificial Analysis Intelligence Index at 60 (the highest score in the field) and leads SWE-bench Pro at 80.3% — a different, harder benchmark that doesn't appear in DeepSeek's chart at all. The 'DeepSeek V4 destroyed Claude Fable 5' framing circulating on YouTube conflates a different Claude generation with a different benchmark.
What's the difference between DeepSeek V4 Pro and V4 Flash?
V4 Pro is the larger, more capable tier: 1.6 trillion total parameters with 49 billion active per token via DeepSeek's Mixture-of-Experts routing. V4 Flash is the faster, cheaper tier: 284 billion total parameters with 13 billion active. Both share the same 1M-token context window, the same hybrid Compressed Sparse Attention / Heavily Compressed Attention mechanism, and the same Thinking / Non-Thinking dual-mode switch. Artificial Analysis's April evaluation of the preview scored V4 Pro (Max reasoning effort) at 52 on the Intelligence Index against V4 Flash (Max) at 47 — roughly Claude Sonnet 4.6 territory for the smaller model. Pricing follows the same gap: V4 Flash runs at $0.14/$0.28 per million input/output tokens off-peak versus V4 Pro's $0.435/$0.87.
How much does DeepSeek V4 cost to use via the API?
V4 Pro is $0.435 per million input tokens and $0.87 per million output tokens off-peak, with a 99% cache-hit discount down to $0.004 per million tokens for repeated context. V4 Flash is $0.14/$0.28 per million input/output tokens off-peak. The GA release introduced peak-time pricing for the first time: rates double during Beijing business hours (9am-12pm and 2pm-6pm daily), so V4 Pro effectively costs $0.87/$1.74 and V4 Flash $0.28/$0.56 during those windows. Off-peak, V4 Pro's output price is roughly 17 times cheaper than Claude Fable 5's per-task cost and around 34 times cheaper than GPT-5.6 Sol's published per-token output rate, though DeepSeek's own verbosity — it generated 180 million output tokens in Artificial Analysis's evaluation suite versus a 96 million median — eats into some of that headline saving in practice.
Is DeepSeek V4 reliable, or does it hallucinate?
This is the release's clearest weak point. Artificial Analysis's evaluation found both V4 Pro and V4 Flash answer instead of declining or hedging in 94% and 96% of cases respectively, even on questions specifically designed to test whether a model should say it doesn't know — among the highest 'always answers anyway' rates of any model in their comparison set. Separately, the US National Institute of Standards and Technology's CAISI unit found DeepSeek V4's real capabilities lag the frontier by roughly eight months once tested on a non-public benchmark suite, running noticeably behind the scores DeepSeek's own reporting implied. Treat confident-sounding answers from V4, particularly on open-ended factual questions, with more scepticism than you would apply to a frontier US lab model.
AI Tools Review Editorial Team

AI Tools Review Editorial Team Expert Verified

Our editorial team consists of veteran AI researchers, software engineers, and industry analysts. We spend hundreds of hours benchmarking frontier models natively to provide you with objective, actionable intelligence on agentic AI capabilities and cybersecurity landscapes.