DeepSeek V4
Compare and use this model
DeepSeek V4 has an article. Its structured comparison facts are awaiting review. Browse the reviewed catalog or suggest a sourced update.
DeepSeek V4 is a family of open-weight Mixture of Experts large language models developed by DeepSeek, a Hangzhou-based AI research lab. Released in preview on April 24, 2026, a launch covered by TechCrunch and Al Jazeera the same day, V4 comprises two variants:[10][16] DeepSeek-V4-Pro, with 1.6 trillion total parameters (49 billion active per token), and DeepSeek-V4-Flash, with 284 billion total parameters (13 billion active per token).[1] Both models support a one-million-token context window and are available as open weights under the MIT license.[1][2] The release introduced a hybrid attention architecture that dramatically reduces inference costs at long contexts, set new low price points for frontier-class models, and was the first DeepSeek release validated for inference on Huawei Ascend processor infrastructure.[8][12] DeepSeek's own technical report tempers the comparison, stating that V4 "falls marginally short of GPT-5.4 and Gemini-3.1-Pro, suggesting a developmental trajectory that trails state-of-the-art frontier models by approximately 3 to 6 months."[5][17] On July 31, 2026 DeepSeek put the first official (non-preview) build of the family into public beta, DeepSeek-V4-Flash-0731, and published its weights on Hugging Face the same day under the MIT license; that build is covered in detail at DeepSeek V4-Flash.[7][32][51] The official DeepSeek V4-Pro build, DeepSeek-V4-Pro-0813, followed on August 13, 2026 across DeepSeek's app, web service and API, and its MIT-licensed weights went up on Hugging Face the same day.[7][61] On September 10, 2026 DeepSeek released DeepSeek V4.1-Flash, which it describes as the smallest model in its new architecture family, retired V4-Flash from its API and routed the deepseek-v4-flash name to the new model; deepseek-v4-pro remains on the API and serves the 0813 build.[7][62]
Background
DeepSeek was founded in 2023 as a research arm of the Chinese quantitative hedge fund High-Flyer Capital Management. The lab became globally known after releasing DeepSeek-V3 in December 2024, a 671-billion-parameter MoE model that outperformed or matched several leading Western models at a fraction of the training cost. Its companion reasoning model, DeepSeek-R1, released in January 2025, demonstrated that strong chain-of-thought reasoning could be elicited through reinforcement learning without relying solely on supervised distillation from proprietary models. The two releases together caused a brief shock in global financial markets, sent Nvidia stock down sharply on January 27, 2025, and intensified discussions about AI export controls.
Subsequent updates extended the V3 line through 2025. DeepSeek-V3-0324 arrived in March 2025 with improvements to AIME (+19.8 points) and GPQA (+9.3 points). DeepSeek-R1-0528 in May 2025 pushed reasoning further through increased post-training compute. DeepSeek-V3.1 in August 2025 introduced a hybrid architecture supporting both thinking and non-thinking modes within a single model and lifted SWE-bench Verified scores to 66.0. DeepSeek-V3.1-Terminus followed in September 2025 with agent capability improvements. DeepSeek-V3.2-Exp, also released in September 2025, introduced DeepSeek Sparse Attention as an experimental long-context optimization. The official V3.2 launched on December 1, 2025, with 685 billion total parameters.[25] V3.2 was billed as the official successor to V3.2-Exp; a parallel release, V3.2-Speciale, attained gold-medal-level results on IMO, CMO, ICPC World Finals, and IOI 2025.[25][26]
V4 represents the first architectural ground-up redesign since V3. The total parameter count roughly doubles V3.2, active parameters grow from 37 billion to 49 billion (Pro) or 13 billion (Flash), and the context window goes to 1M tokens, against the 128K that DeepSeek documented as the context limit for the DeepSeek-V3.2 API. The trained maximum and the served limit are not the same number for V3.2: its released checkpoints set max_position_embeddings to 163,840.[56][57] The underlying attention mechanism is entirely new, replacing the Multi-head Latent Attention that had defined the V3 line.[5]
The release came after a publicly reported delay. 36Kr's launch-day commentary described V4 as fifteen months late, with the interval spent migrating the stack from CUDA to Huawei's CANN framework and extending the context window from 128K to 1M tokens.[30] ChinaTalk also reported on internal funding decisions that pushed multimodal capability to a future generation, and on departures of senior staff to Tencent, ByteDance, Xiaomi, and other Chinese tech firms during 2025.[15]
How is DeepSeek V4 built? Architecture
Model variants
DeepSeek V4 ships in two sizes:
| Variant | Total parameters | Active parameters | Context | Precision | Size on disk |
|---|---|---|---|---|---|
| DeepSeek-V4-Flash | 284B | 13B | 1M tokens | FP8 + FP4 mixed | 160 GB |
| DeepSeek-V4-Pro | 1.6T | 49B | 1M tokens | FP8 + FP4 mixed | 865 GB |
Both models also ship in base (pre-training checkpoint) variants. The instruct variants support three reasoning modes: Non-Think (fast, no extended reasoning), Think High (deliberate logical analysis), and Think Max (maximum reasoning depth). DeepSeek recommends a minimum 384K-token context window when using Think Max mode so that long chains of thought are not truncated.[2]
The "1M" figure is exact rather than approximate: the model catalog DeepSeek published for OpenAI Codex in July 2026 declares a context window of 1,048,576 tokens (2^20) for both deepseek-v4-flash and deepseek-v4-pro, and instructs Codex to treat 95 percent of that as effectively usable.[34][47] Maximum output is 384K tokens for both models.[6]
Hybrid attention: CSA and HCA
The central architectural innovation in V4 is a hybrid attention mechanism combining Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA). Both are departures from the Multi-head Latent Attention (MLA) used in V3 and from the DeepSeek Sparse Attention introduced in V3.2-Exp.[18]
In CSA, a learned token-level compressor consolidates every m tokens along the sequence dimension into a single key-value entry through softmax-gated pooling with a learned positional bias. Queries then attend over these compressed KV representations using DeepSeek Sparse Attention, which selects only the top-k most relevant compressed blocks via a Lightning Indexer component implemented in FP4 precision. A complementary sliding window layer covers the most recent n_win tokens for local context. The net effect is a 4x compression of the KV cache along the sequence axis.[5][18]
HCA is more aggressive. It consolidates every m' tokens (where m' is considerably larger than m, with a 128x compression ratio cited in the technical report) into a single KV entry and applies dense attention over those consolidated entries.[5] The high compression ratio alone delivers efficiency gains, without requiring the sparse selection step that CSA uses.
Layers in V4-Pro's 61-layer stack use these mechanisms in a specific pattern: layers 0 and 1 apply HCA only, layers 2 through 60 alternate between CSA and HCA, and a final multi-token-prediction (MTP) block uses sliding-window attention only.[5] This layered scheme avoids capacity waste by giving different layers different attention patterns suited to local and global context retrieval.
The result at 1M-token context is dramatic. According to the V4 technical report, DeepSeek-V4-Pro requires only 27% of the per-token inference FLOPs and 10% of the KV cache memory that V3.2 would need for the same context length. DeepSeek-V4-Flash is even leaner at 10% of FLOPs and 7% of KV cache relative to V3.2.[5] From V3.2 to V4-Pro, KV cache memory is reduced by approximately 9.5x to 13.7x across context lengths, according to figures cited by The Register.[8] Compared with a standard Grouped-Query Attention baseline (GQA-8), V4's KV cache is roughly 2% of the equivalent GQA-8 footprint.[5]
Manifold-Constrained Hyper-Connections (mHC)
V4 replaces standard residual connections with Manifold-Constrained Hyper-Connections (mHC). The residual stream width is expanded by a factor of four (n_hc = 4 in both model variants). The residual mapping matrices are constrained to the Birkhoff polytope, the set of doubly stochastic matrices, using the Sinkhorn-Knopp algorithm with a maximum of 20 iterations during training. This constraint bounds the spectral norm of the mapping matrices at 1, which prevents signal amplification through the network and improves training stability without sacrificing model expressivity.[5]
The abstract of the technical report, posted to arXiv on April 26, 2026 as arXiv:2606.19348, lists mHC as one of three headline upgrades alongside the CSA/HCA hybrid attention and the Muon optimizer, describing it as an enhancement of "conventional residual connections."[45] The Hugging Face model card uses the same framing, saying mHC is incorporated "to strengthen conventional residual connections, enhancing stability of signal propagation."[2] mHC has since drawn follow-on academic work of its own, including papers on accelerating the Birkhoff projection and on whether the 20 Sinkhorn-Knopp iterations are necessary.
Muon optimizer
DeepSeek-V4 adopts the Muon optimizer for the majority of its parameters. Muon orthogonalizes gradient updates using Newton-Schulz iterations. V4 uses a hybrid schedule: 8 convergence iterations followed by 2 stabilization iterations. Embeddings, prediction heads, biases, and RMSNorm weights use AdamW. According to DeepSeek's technical report, Muon speeds convergence and improves training stability at trillion-parameter scale, and the hybrid Muon plus AdamW configuration was chosen after ablation studies on smaller model sizes.[5]
Precision and quantization
Expert weights in the MoE layers use FP4 precision with quantization-aware training, roughly halving the weight storage footprint versus FP8. Non-expert parameters use FP8. The KV cache stores most entries in FP8, with BF16 reserved only for the rotary positional embedding (RoPE) dimensions where higher precision matters. Base model checkpoints use FP8 throughout, while instruct models combine FP4 (experts) with FP8 (everything else).[5] FP4 adoption at the expert level is a key reason V4-Pro's 865 GB download size is manageable despite 1.6 trillion nominal parameters; an equivalent FP8-only model would be roughly 60% larger.
The Register specifically noted V4's adoption of MXFP4 (microscaling FP4) as a step away from Nvidia-specific FP8 formats, framing it as a deliberate move toward hardware portability across accelerator vendors.[8]
Training stability mechanisms
Two techniques address instability specific to large-scale MoE training. Anticipatory Routing uses historical routing parameters from earlier in training (theta_{t minus delta_t}) to compute token assignments, decoupling backbone and router gradient updates so they do not interfere. SwiGLU Clamping constrains the linear components of SwiGLU activations to the range [-10, 10] and gates to a maximum of 10, preventing gradient explosions in expert layers when individual activations occasionally produce outlier values during training.[5]
Training methodology
Pre-training
The technical report states that DeepSeek-V4-Flash was pre-trained on 32 trillion tokens and DeepSeek-V4-Pro on 33 trillion, so the larger corpus belongs to Pro. Both are described as diverse and high-quality. Both corpora include long-document data to support the 1M-token context objective. The specific composition of the training data, beyond these summary figures, is not disclosed in the technical report.[5]
The V4 paper emphasizes that long-context training data is curated rather than synthesized, with emphasis on whole-document examples (codebases, books, legal corpora) rather than artificially concatenated short documents. The ratio of long-context to short-context examples is gradually increased during training.[5]
Post-training: specialist-then-distill
V4's post-training pipeline is a two-stage specialist-then-distill approach.
In the first stage, separate specialist models are trained for distinct domains: mathematics, coding, agentic tasks, and instruction following, among others. Each specialist starts from the shared pre-trained base and undergoes domain-specific Supervised Fine-Tuning (SFT) followed by Reinforcement Learning using Group Relative Policy Optimization (GRPO) with domain-tailored reward signals. This produces a set of more than ten domain-expert teacher models, each strong in its area.[5]
In the second stage, a single unified student model is distilled from all teachers simultaneously through On-Policy Distillation (OPD). The student generates its own outputs, then minimizes the reverse KL divergence against whichever teacher is most relevant to the current task's logit distribution. Full-vocabulary logit distillation, rather than top-k, is used for stable gradient estimates.[5] The result is a single inference model that combines the strengths of all domain specialists.
This approach contrasts with V3's more uniform reinforcement learning across domains and allows V4 to optimize more precisely for different task types without maintaining multiple inference models in production.
Post-training is also where DeepSeek made its first changes to shipped V4 models. The July 31, 2026 V4-Flash build carries the same architecture and the same parameter count as the April preview and was, in DeepSeek's words, "only re-post-trained."[7] The August 13 Pro build is "built on the DeepSeek-V4-Pro (Preview) model structure, with a DSpark speculative decoding module attached," so it keeps the preview's core network and adds a draft module.[61] The only V4 release that extended the core network was the experimental DeepSeek-V4-Flash-Vision-Exp, which its model card says builds on the V4-Flash architecture "by incorporating visual modules and undergoing continued training."[63]
DSec: training infrastructure for agents
A significant infrastructure component disclosed in the V4 technical report is DeepSeek Elastic Compute (DSec), a production sandbox platform that DeepSeek built for the execution demands of agentic AI during post-training and evaluation. The report describes DSec as three Rust components (an API gateway called Apiserver, a per-host agent called Edge, and a cluster monitor called Watcher) linked by a custom RPC protocol and scaling horizontally on DeepSeek's 3FS distributed filesystem, with a single DSec cluster managing hundreds of thousands of concurrent sandbox instances. A Python SDK, libdsec, exposes four execution substrates behind one interface: function calls dispatched to a pre-warmed container pool, Docker-compatible containers, microVMs built on Firecracker, and full VMs built on QEMU. Container images are stored as 3FS-backed read-only EROFS layers whose data blocks are fetched on demand, so startup does not wait for a full image download. A globally ordered trajectory log for each sandbox lets a preempted training task resume by replaying cached results for commands that had already completed, which also avoids re-executing non-idempotent operations.[5]
DeepSeek gave a fuller account in a separate report, "DeepSeek Elastic Compute (DSec): A Sandbox Infrastructure for Effective Agentic Training at Scale," posted to arXiv on September 19, 2026. It says DSec has served all sandbox workloads used in RL training and evaluation from DeepSeek V3.2 to V4.1, and that one production scale unit of about 160 nodes serves about 3 million sandboxes a day, with peak concurrency of about 380,000 and more than 5,000 sandbox creations per second.[58] The later documents also update the V4 description. Starting with V4.1, rollout execution moved onto DSec as an agent sandbox that hosts the scaffold and its tools plus a worker container that manages it, both outside the preemptible GPU pool, so a preempted training job can reconnect to the retained rollout state instead of reconstructing it through command-log replay.[58][59] Per-node density figures differ by report: the DeepSeek V4.1-Flash report says sub-NUMA partitioning raised supported density from roughly 1,000 to more than 2,500 concurrent live containers per physical node, while the DSec report says stable operation has been observed with at least 3,200 containers or 800 microVMs per node.[58][59] See DSec for the full system description.
Tool-call schema with dedicated tokens
V4 introduces a |DSML| special token that wraps an XML-based tool-call format. This was a deliberate move away from JSON tool-call formats (used by most Western models) because the XML form reduces escaping failures around quoted strings and structured parameters. Parameters are marked with string="true" or string="false" to distinguish quoted strings from numbers and booleans, which the technical report claims reduces parsing errors during multi-turn agent loops.[5]
Benchmark performance
The following scores are for DeepSeek-V4-Pro-Max (the highest-effort inference setting using Think Max mode) unless otherwise noted, taken from DeepSeek's technical report and Hugging Face model card.[2][5] The competitor columns reproduce DeepSeek's Table 6, with each rival at the reasoning effort DeepSeek labels (Claude Opus 4.6 at Max, GPT-5.4 at xHigh, Gemini 3.1 Pro at High, Kimi K2.6 in Thinking mode); the report's GLM-5.1 Thinking column is not reproduced here.[5] The report says DeepSeek re-ran Claude Opus 4.6 and Gemini 3.1 Pro itself on the two 1M-token tasks (MRCR and CorpusQA) to standardise the configuration, left GPT-5.4 out of those because its API failed to answer a large share of queries, and left some Kimi K2.6 and GLM-5.1 entries blank because those APIs were too busy.[5] Many of the other competitor cells are identical to the vendors' own launch figures, for example Claude Opus 4.6's 80.8% on SWE-bench Verified and 91.3% on GPQA Diamond, and Gemini 3.1 Pro's 80.6% and 94.3%.[84][87]
Coding
| Benchmark | V4-Pro-Max | Opus 4.6 Max | GPT-5.4 xHigh | Gemini 3.1 Pro High | K2.6 Thinking |
|---|---|---|---|---|---|
| LiveCodeBench Pass@1 | 93.5 | 88.8 | -- | 91.7 | 89.6 |
| Codeforces Rating | 3206 | -- | 3168 | 3052 | -- |
| SWE-bench Verified | 80.6% | 80.8% | -- | 80.6% | 80.2% |
| SWE-bench Pro | 55.4% | 57.3% | 57.7% | 54.2% | 58.6% |
| SWE-bench Multilingual | 76.2% | 77.5% | -- | -- | 76.7% |
| Terminal-Bench 2.0 | 67.9% | 65.4% | 75.1% | 68.5% | 66.7% |
| MCPAtlas Public | 73.6 | 73.8 | 67.2 | 69.2 | 66.6 |
| Toolathlon | 51.8 | 47.2 | 54.6 | 48.8 | 50.0 |
V4-Pro-Max's Codeforces rating of 3206 was the highest score achieved by any AI model at the time of release, surpassing GPT-5.4-xHigh's 3168 and Gemini's 3052.[2][5] Estimates put the score at roughly the equivalent of rank 23 on Codeforces globally. On SWE-bench Verified, V4-Pro is at near parity with Claude Opus 4.6 (80.8%), Gemini 3.1 Pro (80.6%) and Kimi K2.6 Thinking (80.2%); DeepSeek's table reports no figure for GPT-5.4 xHigh on that benchmark. V4-Pro-Max posts the highest LiveCodeBench score among the models DeepSeek scored on it, 93.5 (the table has no figure for GPT-5.4 or GLM-5.1), and trails GPT-5.4 xHigh by seven points on Terminal-Bench 2.0.[2]
Mathematics and science
| Benchmark | V4-Pro-Max | Opus 4.6 Max | GPT-5.4 xHigh | Gemini 3.1 Pro High |
|---|---|---|---|---|
| GPQA Diamond | 90.1% | 91.3% | 93.0% | 94.3% |
| HMMT 2026 Feb | 95.2% | 96.2% | 97.7% | 94.7% |
| IMOAnswerBench | 89.8% | 75.3% | 91.4% | 81.0% |
| Apex | 38.3% | 34.5% | 54.1% | 60.9% |
| Apex Shortlist | 90.2% | 85.9% | 78.1% | 89.1% |
| Putnam 2025 | 120/120 | -- | -- | -- |
| GSM8K (base) | 92.6% | -- | -- | -- |
| MATH (base) | 64.5% | -- | -- | -- |
V4 achieves a perfect score on Putnam 2025, the undergraduate mathematics competition dataset, though DeepSeek reports that result under a hybrid formal-informal regime with heavy compute scaling rather than as a standard evaluation.[5] On GPQA Diamond, a PhD-level science benchmark, V4 scores 90.1%, the lowest of the four, against Gemini 3.1 Pro's 94.3%.[2] The pattern across this block is that V4-Pro is competitive on competition mathematics but well behind on Apex, where GPT-5.4 xHigh and Gemini 3.1 Pro score 54.1 and 60.9 against its 38.3.[2]
General knowledge
| Benchmark | V4-Pro-Max | Opus 4.6 Max | GPT-5.4 xHigh | Gemini 3.1 Pro High |
|---|---|---|---|---|
| MMLU-Pro | 87.5% | 89.1% | 87.5% | 91.0% |
| SimpleQA Verified | 57.9% | 46.2% | 45.3% | 75.6% |
| Chinese SimpleQA | 84.4% | 76.4% | 76.8% | 85.9% |
| HLE | 37.7% | 40.0% | 39.8% | 44.4% |
On MMLU-Pro, V4-Pro ties GPT-5.4 xHigh at 87.5 and trails Claude Opus 4.6 and Gemini 3.1 Pro. On SimpleQA Verified (a factual accuracy benchmark), V4-Pro's 57.9% clears both Claude Opus 4.6 at 46.2% and GPT-5.4 xHigh at 45.3%, but Gemini 3.1 Pro leads the group by a wide margin at 75.6%.[2] DeepSeek's own technical paper acknowledges that V4 "trails state-of-the-art frontier models by approximately three to six months," placing it roughly at parity with mid-2025 frontier models such as GPT-5.2, Gemini 3.0 Pro, and Claude Opus 4.5.[5]
Long-context
| Benchmark | V4-Pro-Max | Opus 4.6 Max | Gemini 3.1 Pro High |
|---|---|---|---|
| MRCR 1M (MMR) | 83.5 | 92.9 | 76.3 |
| CorpusQA 1M | 62.0% | 71.7% | 53.8% |
| LongBench-V2 (base) | 51.5% | -- | -- |
| BrowseComp | 83.4% | 83.7% | 85.9% |
At 1M-token retrieval (MRCR), V4-Pro scores 83.5 MMR, trailing Claude Opus 4.6's 92.9 but ahead of Gemini 3.1 Pro's 76.3.[2] The Hugging Face blog reports that V4-Pro stays above 0.82 accuracy through 256K tokens on MRCR with 8 needles and holds at 0.59 at 1M tokens.[3]
The July 2026 agent-benchmark update
DeepSeek published nine agent-oriented scores alongside the V4-Flash-0731 release on July 31, 2026, and its framing was that the re-post-trained Flash build now beats the larger V4-Pro-Preview across every one of them.[7][32] A sample of the change:
| Benchmark | V4-Flash-0731 | V4-Flash-Preview | V4-Pro-Preview |
|---|---|---|---|
| Terminal-Bench 2.1 | 82.7 | 61.8 | 72.1 |
| DeepSWE | 54.4 | 7.3 | 12.8 |
| Toolathlon-Verified | 70.3 | 49.7 | 55.9 |
The full nine-benchmark table, the competitor columns, and the reception of the result are covered at DeepSeek V4-Flash. Note that the V4-Pro figures here are for the preview build that DeepSeek was serving on deepseek-v4-pro in July 2026, not for an official Pro release.[7] When the official Pro-0813 build shipped on August 13, DeepSeek's change log reported higher figures for it than for Flash-0731 on all nine benchmarks in the July list, including 87.9 on Terminal Bench 2.1 and 62.7 on DeepSWE; those are again DeepSeek's own runs.[7]
Reading DeepSeek's published scores
DeepSeek attaches two methodology notes to the July 2026 numbers, and both limit how far they can be compared with published leaderboards.[7]
First, the public code-agent tasks were run inside DeepSeek's own scaffold, which it calls the DeepSeek Harness in minimal mode and which it described as "to be released soon" on the day it published the scores, at the max effort level with top_p 0.95 and temperature 1.0.[7] DeepSeek published the harness on GitHub under the MIT license on August 13, 2026, as a developer preview whose README warns that "there will be compatibility-breaking changes."[64] Agent benchmark results depend heavily on the harness driving the model (retry policy, tool set, context management), so a score produced by an unreleased first-party framework is not interchangeable with one produced by a public one.
Second, two of the nine benchmarks are DeepSeek's own. DSBench-FullStack is described as an internal full-stack development test set and DSBench-Hard as an internal coding-agent hard-problem set, neither of which is published, so no outside party can reproduce those rows.[7] The Toolathlon row is likewise labelled "Toolathlon verified" rather than the full public suite.
How much does DeepSeek V4 cost? Pricing
DeepSeek charges per token on its official API. V4-Pro launched behind a 75% introductory discount, which DeepSeek extended twice and then made permanent: on May 22, 2026 the company announced it was "making our discount permanent," and the discounted rates became the standing list prices.[40][41] That flat schedule lasted until 16:00 UTC on August 16, 2026, when DeepSeek replaced it with peak and off-peak billing, announced alongside the official V4-Pro release on August 13, in which off-peak rates are half the peak rates.[7][65] Flash prices were then cut on September 10, 2026, when the Flash tier moved to DeepSeek-V4.1-Flash; the V4-Pro rates did not change.[7][62][75]
Official DeepSeek API pricing
Rates below are DeepSeek's first-party list prices as published on its pricing page on September 23, 2026. Figures are per one million tokens, and both models have a 384K-token maximum output.[62]
| Model id (build served) | Period | Input (cache hit) | Input (cache miss) | Output | Concurrency |
|---|---|---|---|---|---|
deepseek-v4-pro (DeepSeek-V4-Pro-0813) | Off-peak | $0.022 | $0.66 | $1.98 | 500 |
deepseek-v4-pro (DeepSeek-V4-Pro-0813) | Peak | $0.044 | $1.32 | $3.96 | 500 |
deepseek-flash (DeepSeek-V4.1-Flash) | Off-peak | $0.003 | $0.15 | $0.60 | 2,500 |
deepseek-flash (DeepSeek-V4.1-Flash) | Peak | $0.006 | $0.30 | $1.20 | 2,500 |
Peak hours are 01:00 to 04:00 and 06:00 to 10:00 UTC, Monday through Friday, excluding Chinese public holidays; all other hours, including weekends and Chinese public holidays in full, are off-peak.[62] The legacy names deepseek-v4-flash and deepseek-v4-flash-vision-exp are still accepted, but their requests are served by DeepSeek-V4.1-Flash and billed at the Flash price.[62] DeepSeek's Chinese-language page gives the same schedule in yuan: 0.02, 1 and 4 per million tokens off-peak for Flash (0.04, 2 and 8 at peak) and 0.15, 4.5 and 13.5 off-peak for Pro (0.30, 9 and 27 at peak), with peak hours stated as 09:00 to 12:00 and 14:00 to 18:00 Beijing time on weekdays.[66]
Historical snapshot: flat rates to August 16, 2026
The flat rates below, promotional from late April and standing list prices from the May 22 announcement, were in force until 15:59 UTC on August 16, 2026, and are shown as listed on the English and Chinese pricing pages on August 1, 2026.[6][31]
| Model | Input (cache hit) | Input (cache miss) | Output | Max output | Concurrency |
|---|---|---|---|---|---|
deepseek-v4-flash | $0.0028 / 1M | $0.14 / 1M | $0.28 / 1M | 384K tokens | 2,500 |
deepseek-v4-pro | $0.003625 / 1M | $0.435 / 1M | $0.87 / 1M | 384K tokens | 500 |
DeepSeek's Chinese-language page quoted the same rates in yuan: 0.02, 1 and 2 per million tokens for Flash, and 0.025, 3 and 6 for Pro.[31] Concurrency limits are counted per account rather than per API key, and requests beyond the limit return HTTP 429; DeepSeek offers capacity expansion on request at no extra charge, and a user_id parameter that isolates KV cache, scheduling, and content-safety handling between a customer's own end users.[35]
How V4-Pro's price has moved:
| Date | Event |
|---|---|
| April 24, 2026 | V4 preview launch. V4-Pro list price $1.74 input (cache miss) / $3.48 output per 1M tokens.[28] |
| April 25, 2026 | 75% discount announced, initially through May 5, 2026, 15:59 UTC.[49] |
| April 29, 2026 | Discount extended to May 31, 2026, 15:59 UTC.[49] |
| May 22, 2026 | DeepSeek announces the discount is permanent; $0.435 / $0.87 become the standing rates.[40][41][42] |
| June 30, 2026 | The South China Morning Post reports an email to API subscribers saying V4 prices will double during peak hours once the official V4 goes online, then expected in mid-July.[46] |
| August 13, 2026 | With the V4-Pro-0813 release, DeepSeek announces peak/off-peak pricing effective 16:00 UTC on August 16, with off-peak prices at half the peak prices.[7][65] |
| August 16, 2026, 16:00 UTC | Peak/off-peak billing takes effect. V4-Pro: $0.022 / $0.66 / $1.98 (cache hit / cache miss / output) off-peak and $0.044 / $1.32 / $3.96 peak. V4-Flash: $0.007 / $0.22 / $0.66 off-peak and $0.014 / $0.44 / $1.32 peak.[65][67] |
| August 23, 2026 | Weekends (Beijing time) become off-peak all day, effective 00:00 Beijing time.[68][75] |
| September 10, 2026 | Flash tier moves to DeepSeek-V4.1-Flash at lower prices; V4-Pro rates unchanged.[7][62][75] |
Cache-hit pricing represents automatic context caching. DeepSeek applies context caching transparently without requiring developers to declare cache keys or set TTLs, and the Responses API does not support the prompt_cache_key and prompt_cache_retention parameters for that reason.[6][33] Under the flat rates, Artificial Analysis calculated the resulting cache-hit discount at roughly 98 percent, against the 90 percent most of the industry offers.[37] Under the current schedule a V4-Pro cache hit costs one thirtieth of a miss (about 96.7 percent off) and a V4.1-Flash cache hit one fiftieth (98 percent off), in either period.[62]
Peak and off-peak pricing
DeepSeek first signalled the policy at the end of June 2026. According to the South China Morning Post, an email to API subscribers said V4 prices would double during peak hours, 09:00 to 12:00 and 14:00 to 18:00 Beijing time, attributed the change to the need for "better distribution of resources and [to] enhance service stability," and said it would take effect when the official V4 went online, which the email expected in mid-July.[46] Footnote (2) on DeepSeek's pricing page, present identically in the English and Chinese versions, said that the API "will soon adopt a peak/off-peak pricing policy" under which prices during peak hours would be 2x the regular prices, applicable to all billing items, with the effective date "subject to the official announcement."[6][31] That wording was unchanged on August 1, 2026, and the July 31 official Flash launch came and went without the policy being switched on.
DeepSeek set the date with the official V4-Pro release. Its August 13, 2026 change log entry said the company would "adopt peak/off-peak pricing, with off-peak prices set at half of the peak-hour prices," effective 16:00 UTC on August 16, 2026.[7] The pricing page restated peak hours in UTC as 01:00 to 04:00 and 06:00 to 10:00, the same windows as the Beijing-time definition.[65] The new rates were not the flat rates doubled at peak: V4-Pro's off-peak prices of $0.66 for cache-miss input and $1.98 for output were already above the $0.435 and $0.87 flat rates they replaced, and its off-peak cache-hit price rose from $0.003625 to $0.022.[6][65] From 00:00 Beijing time on August 23, 2026 weekends became off-peak all day, and the current footnote also treats Chinese public holidays as off-peak in full.[62][68]
The scheme inverts DeepSeek's earlier experiment with time-of-day pricing. From February 26, 2025 the company ran off-peak discounts instead, cutting DeepSeek-V3 by 50 percent and DeepSeek-R1 by 75 percent between 16:30 and 00:30 UTC daily.[43] That program ended on September 5, 2025, 16:00 UTC, when DeepSeek moved to a new flat price list.[48]
Comparison with competing models
The DeepSeek rows below are the flat rates DeepSeek charged until August 16, 2026, not its current peak and off-peak prices (see above). Competitor rows are each vendor's standard (non-batch) list price per 1M tokens as of mid-August 2026, from Wayback Machine snapshots of the vendors' pricing pages dated August 14 to 16, 2026.[78][79][82] The Gemini and OpenAI prices are the short-prompt tier: Google charged $4.00 input and $18.00 output for Gemini 3.1 Pro prompts over 200K tokens, and OpenAI billed GPT-5.4 and GPT-5.5 prompts over 272K input tokens at 2x input and 1.5x output for the whole session, while Anthropic billed Claude Opus 4.6 and 4.7 at the same rate across the full 1M-token window.[78][79][80][81][82] Context windows are from the vendors' model documentation.[78][80][81][83]
| Model | Input ($/1M) | Output ($/1M) | Context |
|---|---|---|---|
| DeepSeek V4-Flash | $0.14 | $0.28 | 1M tokens |
| DeepSeek V4-Pro | $0.435 | $0.87 | 1M tokens |
| Gemini 3.1 Pro | $2.00 | $12.00 | 1M tokens |
| GPT-5.4 | $2.50 | $15.00 | 1.05M tokens |
| GPT-5.5 | $5.00 | $30.00 | 1.05M tokens |
| Claude Opus 4.6 | $5.00 | $25.00 | 1M tokens |
| Claude Opus 4.7 | $5.00 | $25.00 | 1M tokens |
VentureBeat described V4 at launch as offering "near state-of-the-art intelligence at 1/6th the cost" of top Western models.[19] Artificial Analysis, running its full Intelligence Index suite on V4-Pro at first-party pricing after the cut, put the cost of one full run at about $268 in May 2026 (its published figure on the v4.1 index as of July 31 was $176.34), which it calculated as roughly 3x cheaper than Gemini 3.1 Pro Preview, 12x cheaper than GPT-5.5 at extra-high effort, and 19x cheaper than Claude Opus 4.7.[42]
Which API formats does DeepSeek V4 support?
The DeepSeek API accepts three wire formats. OpenAI ChatCompletions and Anthropic Messages requests both work against https://api.deepseek.com (Anthropic-format calls use the /anthropic path) and have done since the April 2026 launch.[1] The 0731 Flash build added a third: OpenAI's Responses API, the stateful, tool-oriented protocol that OpenAI's Codex clients speak.[7][33]
Support for the newer protocol started out narrower than the model family. On July 31, DeepSeek's documentation stated in three places, the pricing page footnote, the Responses API guide, and the Codex integration page, that the Responses API and the Codex configuration covered deepseek-v4-flash only, and that deepseek-v4-pro support was expected in early August 2026.[6][33][34] It arrived with the official Pro release: the August 13 change log says the API "now natively supports the OpenAI Responses API format and is specifically adapted for Codex," and the pricing page has marked the feature for both models since then.[7][65] DeepSeek's current Codex setup script configures either deepseek-flash or deepseek-v4-pro.[69]
DeepSeek publishes a ~/.codex/config.toml fragment that registers a [model_providers.deepseek] block with wire_api = "responses", plus a models.json catalog that declares the model's context window and reasoning levels so Codex treats it like a built-in model. A one-line setup script for macOS, Linux, and Windows writes both files, backing up the existing configuration first and validating the syntax before writing. The same configuration serves the Codex CLI, the ChatGPT desktop app, and the Codex extension for Visual Studio Code.[34]
As first documented on July 31, 2026, the implementation is a deliberate subset rather than a full clone. It is stateless, so previous_response_id and conversation are unsupported and store always returns false; background, metadata, include, prompt, service_tier, safety_identifier, and context_management are not implemented either. truncation is unsupported, so a request that exceeds the context window returns HTTP 400 instead of being silently trimmed. Among tools, function and web_search worked at that point, custom was accepted only for the apply_patch name that Codex needs, and file_search, code_interpreter, computer_use, and Model Context Protocol tools were ignored. Image and file inputs were not supported: an input_image part did not raise an error but was swapped for placeholder text. Unsupported parameters are silently ignored so that existing clients connect without modification.[33] The guide has changed since. As read on September 23, 2026 it names deepseek-flash as the model, processes input_image parts as real images for it, and lists web_search among the ignored built-in tools, while function and the apply_patch custom tool remain supported and file inputs remain unsupported.[70]
DeepSeek documents a separate integration for Claude Code at the /anthropic endpoint. As first published, its recommended configuration had V4-Pro fill the main, Opus and Sonnet roles and mapped V4-Flash to the Haiku and subagent roles.[77] The current guide's recommended configuration sets every role to deepseek-flash (with Claude Code's 1M-context setting for the main, Opus and Sonnet roles), while DeepSeek's server still maps any claude-opus model name to deepseek-v4-pro, billed at the Pro price, and claude-sonnet and claude-haiku names to deepseek-flash.[71]
Model variants
As of September 23, 2026 the DeepSeek API lists two models: deepseek-flash, served by DeepSeek-V4.1-Flash, which is not a V4 build, and deepseek-v4-pro, served by DeepSeek-V4-Pro-0813. It still accepts the legacy names deepseek-v4-flash and deepseek-v4-flash-vision-exp and routes both to V4.1-Flash.[62] Over the life of the family DeepSeek has released five instruct builds under the V4 name, plus two base checkpoints. The table below separates them, because DeepSeek's documentation and its benchmark charts use the longer names while the API uses the short ids.
| Name | What it is | API model id | Open weights | Status on 2026-09-23 |
|---|---|---|---|---|
| DeepSeek-V4-Flash-Preview | April 24, 2026 preview build, 284B total / 13B active | deepseek-v4-flash (until July 31) | Yes, MIT, on Hugging Face | Superseded on the API July 31, 2026; repository frozen at its June 22, 2026 state[7][44] |
| DeepSeek-V4-Flash-0731 | Official Flash build, re-post-trained, same architecture and size | deepseek-v4-flash (July 31 to September 10, 2026) | Yes, MIT, on Hugging Face from July 31, 2026 | Retired from the API on September 10, 2026; the name now routes to V4.1-Flash[7][51][62] |
| DeepSeek-V4-Pro-Preview | April 24, 2026 preview build, 1.6T total / 49B active | deepseek-v4-pro (until August 13) | Yes, MIT, on Hugging Face | Superseded on the API August 13, 2026[7][2] |
| DeepSeek-V4-Pro-0813 | Official (GA) Pro build: preview model structure with a DSpark speculative-decoding module attached | deepseek-v4-pro | Yes, MIT, on Hugging Face from August 13, 2026 | Served on the API, app and web; a retirement plan announced September 10 was withdrawn by September 11[7][61][62][74] |
| DeepSeek-V4-Flash-Vision-Exp | Experimental multimodal build: V4-Flash architecture plus visual modules, with continued training | deepseek-v4-flash-vision-exp (August 21 to September 10, 2026) | Yes, MIT, on Hugging Face from August 31, 2026 | Retired from the API on September 10, 2026; the name now routes to V4.1-Flash[7][62][63] |
| DeepSeek-V4-Flash-Base / V4-Pro-Base | Pre-training checkpoints, FP8 throughout | not served | Yes, MIT, on Hugging Face | Available[2][44] |
Two details are easy to misread. Until the August 13 release, DeepSeek's pricing page listed the model version for deepseek-v4-pro simply as "DeepSeek-V4-Pro," even though the July 31 change log entry stated that the V4-Pro API was not changed and that the official V4-Pro release would follow later, and DeepSeek's own July comparison chart labelled the Pro column "V4-Pro-Preview."[6][7][32] And because each update reused an existing model id, callers who had hard-coded deepseek-v4-flash picked up different behavior on July 31, and a different model on September 10, without changing a line of code; deepseek-v4-pro changed builds the same way on August 13. That matters when comparing evaluation runs across those dates.[7][62]
A third set of repository names is not a set of builds at all. DeepSeek-V4-Pro-DSpark and DeepSeek-V4-Flash-DSpark, published on Hugging Face on June 27, 2026, each open with the same disclaimer: the repository "is not a new model. It is the same checkpoint with an additional speculative decoding module attached."[52][53] DSpark is a draft-model method from DeepSeek's DeepSpec codebase, which trains and evaluates draft models for speculative decoding; vLLM and SGLang enable it with a flag rather than by loading different base weights.[52][54] The July 31 Flash build ships the module by default: DeepSeek-V4-Flash-0731 states that it "has the same model structure as DeepSeek-V4-Flash-DSpark, i.e. it comes with a speculative decoding module attached," and its config.json carries the dspark_* keys to prove it.[51]
DeepSeek-V4-Pro
V4-Pro is the larger of the two variants, with 1.6 trillion total parameters and 49 billion active per token. It uses FP4 for expert weights and FP8 for other layers, resulting in an 865 GB checkpoint.[2] It is aimed at demanding tasks: complex multi-step reasoning, advanced coding, scientific analysis, and long-document comprehension. On agentic benchmarks it rivals or approaches the performance of Claude Opus 4.6 and GPT-5.4. The V4-Pro Hugging Face page listed more than 1.06 million downloads in the first month after release; by July 31, 2026 the repository showed roughly 1.64 million downloads in the preceding 30 days and about 5,400 likes.[2]
Thinking-effort handling at first differed between the two models. On July 31, 2026 DeepSeek's documentation mapped a requested effort of low on deepseek-v4-pro up to an actual effort of high, so Pro had no genuine low-effort mode, and said the mapping would be revised in early August 2026.[36] The August 13 change log gave both V4 models three effort levels, low, high and max, and the current guide publishes a single mapping in which low stays low.[7][72]
Pro did not receive the July 31 update that Flash did: on August 1, 2026 the deepseek-v4-pro id was still answering with the April preview.[6][7] The official build, DeepSeek-V4-Pro-0813, replaced it on August 13 across the app, web service and API, together with Responses API support and published MIT weights.[7][61] On September 10 DeepSeek announced that from 12:00 Beijing time on September 14 it would route deepseek-v4-pro requests to V4.1-Flash and retire V4-Pro, saying V4.1-Flash outperformed it on "performance, cost, speed, and total time." By September 11 it had replaced that notice with a statement that, "in response to user demand," V4-Pro API service would continue after September 14 with billing unchanged.[7][73][74] The model is treated in full at DeepSeek V4-Pro.
DeepSeek-V4-Flash
V4-Flash has 284 billion total parameters with 13 billion active per token, weighing 160 GB.[1] According to DeepSeek, "reasoning capabilities closely approach V4-Pro" despite the smaller scale, and it runs at substantially lower cost and latency. V4-Flash is positioned as a drop-in replacement for the legacy deepseek-chat and deepseek-reasoner API endpoints, which route to V4-Flash's non-thinking and thinking modes respectively.[1] Those legacy endpoints were retired on July 24, 2026, 15:59 UTC, three months after the preview launch, leaving the two deepseek-v4-* ids as the only names on the API until DeepSeek introduced deepseek-flash for V4.1-Flash on September 10, 2026.[1][7]
One week later, on July 31, 2026, DeepSeek put the official Flash build into public beta as DeepSeek-V4-Flash-0731 and published its weights the same day. It is the main article subject at DeepSeek V4-Flash, which covers the benchmark jump, the independent measurements, the weights release, and the Codex integration in detail.[7][32][50][51] DeepSeek retired V4-Flash from its API on September 10, 2026, when it released V4.1-Flash; the deepseek-v4-flash name has since been routed to the new model, and the 0731 weights remain on Hugging Face.[7][51][62]
DeepSeek-V4-Flash-Vision-Exp
DeepSeek's first multimodal V4 build went live on the API on August 21, 2026 as an experimental model under the id deepseek-v4-flash-vision-exp.[7] DeepSeek's own table puts it level with Flash-0731 on text-agent tasks (83.9 against 82.7 on Terminal Bench 2.1) and ahead on multimodal agent tasks such as ApexBench (36.5 against 26.2) and Agents' Last Exam (27.3 against 25.2), where the text-only Flash-0731 ignores the image inputs.[63] Its weights were published on Hugging Face under the MIT license on August 31, 2026.[63] The model was retired from the API on September 10, 2026 together with V4-Flash.[7]
Base models
Both Pro and Flash ship with base (pre-training checkpoint) variants in addition to instruction-tuned versions. Base models are available as DeepSeek-V4-Flash-Base and DeepSeek-V4-Pro-Base and are intended for downstream fine-tuning. Both base checkpoints use FP8 precision throughout, without the FP4 expert quantization applied to instruct models.[5]
Is DeepSeek V4 open source? Release and licensing
DeepSeek released V4 weights under the MIT License, one of the most permissive open source licenses. Announcing the launch on April 24, 2026, DeepSeek wrote that "DeepSeek-V4 Preview is officially live and open-sourced," welcoming developers to "the era of cost-effective 1M context length."[1] The MIT terms allow commercial use, redistribution, and modification without royalties or restrictions.[2] The models are available for download on Hugging Face under the deepseek-ai organization.[2] V4-Pro accumulated more than one million downloads in its first month, and the V4 collection was the most-downloaded large language model collection on Hugging Face in late April and early May 2026.[4]
MIT is the license across all nine V4 repositories DeepSeek has published, but it is not declared uniformly. Seven of them, the two preview instruct checkpoints, the two DSpark repositories, the July 2026 Flash build, the August 2026 Pro build and the Vision-Exp build, carry the mit license tag in their Hub metadata and repeat it in their model cards. The two base checkpoints carry neither a Hub license tag nor a model card, and ship the licence only as a LICENSE file in the repository. Third-party listings that describe V4 as Apache 2.0 are wrong.[2][44][51][52][53][61][63]
The official Flash build briefly looked as though it would stay API-only, because DeepSeek did not publish it as a revision of the existing deepseek-ai/DeepSeek-V4-Flash repository, which is still frozen at its June 22, 2026 state and still opens with the sentence "We present a preview version of DeepSeek-V4 series."[44] It went into a new repository instead, deepseek-ai/DeepSeek-V4-Flash-0731, whose release commit landed on July 31, 2026, the same day as the public beta and the change log entry.[7][51] The file listing shows 48 safetensors shards totalling roughly 167 GB, an mit license tag, and the model card's own description of the build as "the official release of DeepSeek-V4-Flash, superseding the preview version."[51][55] Artificial Analysis had written on the day of the API launch that DeepSeek "is expected to release the model's full weights in the coming weeks"; the wait was hours, not weeks.[37]
Two figures on that repository do not match the specification table above, and readers comparing the two should know why. The Hub sidebar reports 304B parameters rather than 284B, and the repository is 167 GB across 48 shards where the preview repository is 160 GB across 46.[44][51][55] Only the storage step is the DSpark draft module. The parameter figure is an accounting artifact: DeepSeek-V4-Flash-DSpark carries the same module, the same 48 shards and byte-identical storage of 166,886,535,336 bytes, yet Hugging Face reports 165.3B for it against 304.2B for the 0731 repository. Two repositories holding the same bytes cannot differ by 139 billion parameters. The gap lies entirely in how each declares its tensors: the 0731 repository reports 296,352,743,424 INT8 elements and no E8M0 tensors, while the DSpark repository reports exactly half that many INT8 elements plus 9,261,408,000 E8M0 tensors, which the shared config.json identifies as quantisation scales through "scale_fmt": "ue8m0" rather than as weights.[44][51] Neither sidebar figure is a parameter count. DeepSeek's own statements agree that the model did not change: the change log says the 0731 build "keeps the same model architecture and size as DeepSeek-V4-Flash-Preview, and was only re-post-trained," the DeepSeek-V4-Flash-DSpark card gives 284B total and 13B activated for the same checkpoint, and the 0731 config.json still declares 256 routed experts, 6 experts per token, one shared expert, and a 1,048,576-token maximum position count.[7][51][53]
The practical consequence on August 1, 2026 was that the open-weight checkpoints and the served API models were the same artifacts again on both ids: deepseek-v4-flash served the 0731 build and deepseek-v4-pro the preview, and both had public weights. That held through the August 13 Pro update, whose 0813 weights were published the same day, and it still holds for deepseek-v4-pro. Since September 10 the Flash names have been served by V4.1-Flash, whose weights DeepSeek has also published under MIT.[7][62][76]
The technical report, titled "DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence," was first published as a PDF in the Hugging Face repository; a June 22, 2026 commit deleted the PDF and repointed the model card's Technical Report link to arXiv.[2][60] It was also posted to arXiv on April 26, 2026 as arXiv:2606.19348, credited to DeepSeek-AI and several hundred named authors.[45]
DeepSeek provides inference examples for multiple deployment frameworks including Hugging Face Transformers, vLLM, SGLang, and Docker Model Runner.[2] Quantized versions compatible with llama.cpp, Ollama, LM Studio, and Jan are available from the community. Unsloth published a fine-tuning fork on Hugging Face within days of release, and Nvidia published an optimized variant for its NIM inference platform.
Deployment at the 865 GB scale of V4-Pro requires significant hardware. For a rough estimate, running at FP4 plus FP8 mixed precision on a single node typically requires multiple high-memory accelerators, such as several H100 80GB GPUs. V4-Flash at 160 GB is more accessible for organizations with modest compute, fitting within a single 8-GPU H100 or H200 node at full precision and on smaller hardware after quantization.
Capabilities
Agentic coding
V4-Pro showed particular strength on agentic coding tasks. On SWE-bench Verified, it resolves 80.6% of GitHub issues, near parity with Claude Opus 4.6's 80.8% and Gemini 3.1 Pro's 80.6%. On BrowseComp (web browsing and research), it scores 83.4%.[2] V4-Pro-Max's Codeforces rating of 3206 represents the highest competitive programming score by an AI model at the time of release.[5] Developer community testing, as reported by MindStudio and independent developers, put V4-Pro in the top two or three choices for coding use cases in the months following release. A MindStudio survey reported by ghost.codersera found that 52% of surveyed DeepSeek developers were ready to replace their primary coding model with V4-Pro, and another 39% were leaning toward yes.
The technical report also disclosed an internal R&D coding evaluation in which V4-Pro scored 67% pass rate, between Claude Sonnet 4.5 (47%) and Claude Opus 4.5 (70%).[5]
By mid-2026 the ranking inside the family had inverted for agent work specifically. On the nine benchmarks DeepSeek published on July 31, 2026, the re-post-trained 13-billion-active Flash build scored above the 49-billion-active Pro preview on every one, a result DeepSeek presented as evidence that post-training, not parameter count, was the binding constraint on agent performance.[7][32]
Long-context comprehension
The 1M-token context window is fully operational at launch, rather than a theoretical maximum. On the MRCR 1M needle-in-a-haystack benchmark, V4-Pro scores 83.5 MMR, below Claude Opus 4.6 but ahead of other open models.[2] At 1M tokens, V4-Pro's FLOPs requirement is 27% of V3.2's, making long-context inference economically viable.[5] The Hugging Face blog article specifically frames V4 as "a million-token context that agents can actually use," noting that prior 1M-context releases tended to be theoretical maxima with degraded retrieval at high token counts.[3]
Reasoning modes
Both V4 variants support three inference modes. Non-Think mode produces fast responses without extended reasoning chains, suitable for routine queries. Think High mode engages deliberate analysis with a moderate token budget for thinking. Think Max mode pushes reasoning to its full extent, requiring a minimum 384K-token context window to accommodate extended chain-of-thought traces.[2]
In the two older wire formats the API splits this into two controls rather than one. Thinking mode is toggled with {"thinking": {"type": "enabled"}} or disabled, and it is enabled by default; effort within thinking mode is then set with reasoning_effort (OpenAI format) or output_config.effort (Anthropic format), each accepting low, high, or max, with high as the default. The Responses format collapses both into one field, {"reasoning": {"effort": ...}}, where an effort of none disables thinking altogether. As documented on July 31, 2026, temperature, top_p, presence_penalty, and frequency_penalty had no effect in thinking mode; they were accepted without error for client compatibility and then ignored.[36] The current guide makes one exception: top_p now takes effect in thinking mode within a range of 0.95 to 1.0, with lower values treated as 0.95.[72]
The interleaved-thinking design keeps all reasoning content when tools are in use, including across user message boundaries (a change from V3.2, which discarded it at each new user turn), but discards earlier reasoning at each new user message in conversation without tools. This avoids context pollution in chat-style use while supporting long-horizon agent loops.[5] The API enforces the rule. As documented on July 31, 2026, between two user messages reasoning_content from an earlier turn was ignored if the model made no tool call and had to be passed back if it did, and requests carrying the tools parameter that failed to pass it back returned an HTTP 400 error.[36] The current guide states the rule purely in terms of the request: when a request carries the tools parameter, the reasoning content of all previous turns must be passed back, even for turns without a tool call; when it does not, earlier reasoning content is ignored.[72]
Instruction following and multilingual
V4 shows competitive multilingual performance. Chinese SimpleQA scores 84.4%, and the model performs well on Chinese-language tasks. On SimpleQA Verified (English factual accuracy), V4-Pro outperforms Claude Opus 4.6 at 46.2% with a score of 57.9%.[2] Reviewer testing flagged that V4-Pro is less reliable than GPT-5.5 and Claude Opus 4.6 on prompts with many simultaneous structural constraints (precise output schemas, exact word counts, multi-format documents), where the larger Western models maintain higher consistency.
Hardware and infrastructure
Training hardware
DeepSeek's training infrastructure for V4 remains primarily Nvidia-based, according to reporting from ChinaTalk and The Register.[8][15] Earlier DeepSeek models were trained on Nvidia A100 clusters obtained before export controls, and subsequent models used A800 chips. V4's training used Nvidia GPUs, though DeepSeek has not publicly specified which generation. U.S. government officials separately alleged that DeepSeek obtained banned Nvidia Blackwell chips through third-party intermediaries, though this has not been independently confirmed.[15]
Huawei Ascend optimization
A significant V4 development is its validated support for Huawei Ascend 950PR processors for inference.[12] DeepSeek's technical report describes the MegaMoE fused kernel as successfully running on Huawei Ascend hardware, and the report mentions validating a fine-grained Expert Parallel scheme on both Nvidia GPUs and Ascend NPU platforms.[5] Huawei announced full Ascend platform support for V4 models at launch.[13] The company's TileLang domain-specific language reduces CUDA dependency on the inference path, enabling portability to non-Nvidia accelerators.
The Register notes that V4 adopts MXFP4 (microscaling FP4) for post-training and inference, reducing dependence on Nvidia-specific FP8 formats.[8] This was described by technical analysts as a deliberate step toward hardware portability.
For now, the Huawei integration covers inference, not training, with training still primarily on Nvidia hardware.
If V4's Huawei inference optimization proves reliable at scale, it could provide evidence that competitive frontier models can run on non-Nvidia infrastructure, with implications for the effectiveness of U.S. semiconductor export controls. Nvidia CEO Jensen Huang called the prospect of DeepSeek running on Huawei chips "a horrible outcome for America," in a statement reported by The Next Web.
How does DeepSeek V4 compare with frontier models?
The DeepSeek prices in this table are the flat rates in force until August 16, 2026; current DeepSeek rates are in the pricing section above.[6][62] Competitor prices are standard list prices per 1M tokens (cache-miss input) as of mid-August 2026, at the short-prompt tier where a vendor charges more for long prompts.[78][79][82][89][92] Benchmark scores are each lab's own published figure for its own model: DeepSeek's Table 6 (V4-Pro at Max effort), Anthropic's launch posts for Claude Opus 4.6 and 4.7, OpenAI's GPT-5.5 launch post (xhigh reasoning effort), Google DeepMind's Gemini 3.1 Pro model card, Moonshot's Kimi K2.6 model card (thinking mode) and Z.ai's GLM-5.1 model card. They come from different harnesses and settings, so the columns are not a controlled head-to-head.[5][84][85][86][87][88][90] "n/a" marks a benchmark the vendor's launch material does not score: OpenAI's GPT-5.5 post and Z.ai's GLM-5.1 card report SWE-Bench Pro but not SWE-bench Verified.[86][90]
| Model | Org | Open weights | Params (active) | Input $/1M | Output $/1M | SWE-bench Verified | GPQA Diamond |
|---|---|---|---|---|---|---|---|
| DeepSeek V4-Pro | DeepSeek | Yes (MIT) | 49B (1.6T total) | $0.435 | $0.87 | 80.6% | 90.1% |
| DeepSeek V4-Flash | DeepSeek | Yes (MIT) | 13B (284B total) | $0.14 | $0.28 | -- | -- |
| Claude Opus 4.6 | Anthropic | No | undisclosed | $5.00 | $25.00 | 80.8% | 91.3% |
| Claude Opus 4.7 | Anthropic | No | undisclosed | $5.00 | $25.00 | 87.6% | 94.2% |
| GPT-5.5 | OpenAI | No | undisclosed | $5.00 | $30.00 | n/a | 93.6% |
| Gemini 3.1 Pro | No | undisclosed | $2.00 | $12.00 | 80.6% | 94.3% | |
| Kimi K2.6 | Moonshot AI | Yes (modified MIT) | 32B (1T total) | $0.95 | $4.00 | 80.2% | 90.5% |
| GLM-5.1 | Zhipu AI | Yes (MIT) | 40B (744B total) | $1.40 | $4.40 | n/a | 86.2% |
V4-Pro is the largest open-weight model available as of its release date, surpassing Moonshot AI's Kimi K2.6 at 1T total parameters and Zhipu AI's GLM-5.1 at 744B.[88][91] Among Chinese open models, V4-Pro leads across math, coding, and STEM benchmarks.[20] That position eroded over the following quarter: by the end of July 2026 both labs had shipped successors (Kimi K3 and GLM-5.2) that Artificial Analysis rated above V4-Pro on its aggregate index.[37]
Industry impact
China-US AI competition
DeepSeek V4 landed on the same day Reuters reported the U.S. State Department had sent a diplomatic cable to embassies worldwide instructing staff to warn foreign governments about alleged IP theft by DeepSeek and other Chinese AI companies.[13] The concurrent timing, whether coincidental or deliberate, brought significant geopolitical attention to the release.
The Council on Foreign Relations published an analysis that same day arguing V4 "signals a new phase in the U.S.-China AI rivalry," with the competition shifting from raw frontier capability toward economic adoption and global influence, particularly in the Global South.[14]
CFR Senior Fellow Michael C. Horowitz emphasized the adoption race angle: "Success will not just be about having the best-performing models," he wrote, but about having good-enough solutions that deploy cheaply at scale. He noted that Chinese open models already had more downloads on Hugging Face than U.S. equivalents, a dynamic V4's open release was likely to amplify.[14]
CFR Senior Fellow Jessica Brandt raised concerns that V4's capabilities "reflect, at least in part, access to illicitly obtained U.S. intellectual property," citing alleged large-scale model distillation attacks through fake accounts.[14] Anthropic and OpenAI separately alleged that DeepSeek-affiliated actors had created tens of thousands of fake accounts conducting tens of millions of interactions to extract capabilities from frontier U.S. models. DeepSeek has not publicly responded to these allegations.
CFR Senior Fellow Chris McGuire offered a more measured technical assessment, noting that V4 "is not competitive with frontier U.S. models" and that DeepSeek remains significantly dependent on U.S. semiconductor technology.[14] He flagged that DeepSeek itself had admitted compute shortages limit V4 deployment at scale, with the V4-Pro model unavailable to most API customers in the launch days.
Market reaction
Following the V4 release, shares of SMIC (China's leading chip foundry) rose roughly 9% in Hong Kong trading, with Hua Hong Semiconductor up about 15%, reflecting investor expectations that Huawei Ascend chip demand would increase.[29] Competing Chinese AI startups MiniMax (HKG: 0100) and Knowledge Atlas (Zhipu, HKG: 2513) saw shares fall, with MiniMax sliding around 9-10% in the days after release and Zhipu down 3.4% in Monday trading following V4's launch.[27] The pattern reflected investor rotation out of model developers facing pricing pressure into chip suppliers benefiting from compute demand.
V4's preview release coincided with DeepSeek's first-ever external financing round. According to reporting from Bloomberg, The Information, and CnTechPost, DeepSeek had been in advanced talks since mid-April 2026 to raise at least $300 million at an initial $10 billion target valuation.[23] Within days of the V4 announcement, interest from investors including Tencent and Alibaba pushed the valuation discussion above $20 billion.[24] The financing round marks a sharp pivot for a company that had spent its first two and a half years rejecting venture capital offers in favor of a research-first culture funded by High-Flyer Capital Management.
Developer and commercial adoption
MIT Technology Review identified three reasons V4 matters beyond raw benchmark scores: the compressed attention architecture reduces inference costs to levels that make 1M-token context economically practical for the first time; the open MIT license enables commercial use without licensing negotiations; and the pricing sets a new benchmark for cost pressure on proprietary model providers.[9] Over 90% of developers surveyed by MindStudio included V4-Pro among their top coding model choices in the weeks following release. The legacy API endpoint transition (deepseek-chat to V4-Flash, deepseek-reasoner to V4-Flash thinking mode) signals that V4 is intended as a production replacement, not an experimental release.[7]
V4 was rapidly integrated into agent frameworks and coding tools. The DeepSeek API release notes specifically highlighted compatibility with Claude Code, OpenClaw, and OpenCode, alongside support for both OpenAI ChatCompletions and Anthropic-style API formats.[7] This dual-format support reduced switching costs for developers already using Western model APIs. The July 2026 addition of the Responses API and a published Codex configuration extended the same logic to OpenAI's own client tooling.[33][34]
Reception
Technical coverage was broadly positive but included specific criticisms. Simon Willison, a widely-read developer blogger, described V4 as "almost on the frontier, a fraction of the price," noting that the efficiency improvements at long contexts are genuine and reproducible.[17] He directly quoted DeepSeek's own admission that performance "falls marginally short of GPT-5.4 and Gemini-3.1-Pro, suggesting a developmental trajectory that trails state-of-the-art frontier models by approximately 3 to 6 months."[17] Willison also ran his pelican-on-bicycle SVG generation test, finding that V4-Flash produced a competent rendering while V4-Pro's pelican anatomy was somewhat off.[17]
The Register covered the architecture in depth, describing the KV cache reductions as the most significant technical contribution and emphasizing the MXFP4 adoption as a potential break from Nvidia hardware lock-in.[8] MIT Technology Review framed V4 as the moment when 1M-token context became economically practical for production use.[9] Bloomberg called V4 "DeepSeek's newest flagship a year after [its] AI breakthrough," referencing the V3 and R1 releases of late 2024 and early 2025.[21] CNBC and CNN both ran prominent international coverage on launch day.[11][22]
Developer community reaction on platforms including Reddit (r/DeepSeek, r/singularity, r/LocalLLaMA) and Hugging Face was enthusiastic about the context window, pricing, and Codeforces score. The 3206 Codeforces rating was immediately flagged as a landmark for competitive programming AI.
Criticisms concentrated on a few areas. At launch, some chat instances reportedly identified themselves as V3, suggesting incomplete deployment. Coverage of third-party benchmark results was incomplete in the first week. Developers testing frontend code generation found V4-Pro's output functionally correct but less visually polished than GPT-5.5 on UI tasks. The API initially had reliability problems for V4-Pro under load, with many users encountering rate limits and queuing during peak hours.
The 36Kr report and ChinaTalk analysis added context that V4's development was delayed by training migration failures, talent departures to Tencent, ByteDance, Xiaomi, and other Chinese tech firms, and internal funding decisions that postponed multimodal capability to a later release.[15][30] The 36Kr commentary framed V4 as the "singularity of the Cambrian explosion of AI applications in China," a foundational platform for downstream Chinese AI products rather than a frontier-capability moonshot.[30]
Independent evaluation by Artificial Analysis
Artificial Analysis, which runs its own harness rather than reprinting vendor numbers, gives the clearest outside read on the family. On its Intelligence Index v4.1, an aggregate of nine evaluations including GDPval-AA v2, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, and AA-Omniscience, the firm scored the V4-Pro preview at 44 and the July V4-Flash-0731 build at 50, up from 40 for the April Flash preview, as published on July 31, 2026.[37][38][39] The smaller model therefore rates six points above the larger one on the index, a gap Artificial Analysis attributed to the newer post-training run rather than to scale.[37]
Where DeepSeek and Artificial Analysis measured the same public benchmark, the numbers diverged. Artificial Analysis measured Terminal-Bench 2.1 for the 0731 Flash build at 79 percent, a 17-point rise over the preview but roughly four points below DeepSeek's reported 82.7, which had been produced with DeepSeek's own unreleased harness at max effort.[7][37] The firm also placed V4-Flash-0731 on its Pareto frontier for intelligence against cost per task and found that, even after an 80 percent price cut on a comparably scoring OpenAI model the day before, running the index through DeepSeek's first-party API cost about 60 percent less.[37]
Two of its findings cut against a simple reading of the gains. The improvement on AA-Omniscience came entirely from a lower hallucination rate, 12 points down to 84 percent, while the share of answers that were correct did not move. And the model is verbose: it consumed roughly 206 million output tokens to complete the index as first reported on July 31 (Artificial Analysis's model page later showed 210 million), down 12 percent from the April build but still far above the cross-model median.[37]
Limitations
Several limitations were noted at launch or in subsequent testing.
Performance gap with closed models. DeepSeek's own technical report states V4 trails state-of-the-art frontier models by approximately three to six months.[5] On MMLU-Pro (87.5% versus Gemini 3.1 Pro at 91.0%) and HLE (37.7% versus Claude Opus 4.6 at 40.0%), the gap to the leading models is measurable. On long-context retrieval (MRCR 1M: 83.5 versus Claude Opus 4.6's 92.9), V4 is competitive but not leading.[2]
Limited multimodal capability. V4 was text-only at release. Internal funding and compute constraints reportedly pushed multimodal features to a future release.[15] That was a notable gap for teams building agents that require visual validation of outputs, document understanding with figures and tables, or screen-reading workflows. In July 2026 the limitation was enforced at the protocol level: the Responses API declared text as the only input modality and replaced image parts with placeholder text rather than processing them.[33][47] The family's only multimodal build, the experimental V4-Flash-Vision-Exp, was served on the API from August 21 to September 10, 2026. Image input on DeepSeek's API now comes from V4.1-Flash, and the pricing page lists vision as not supported on deepseek-v4-pro.[7][62][63]
Complex multi-constraint instruction following. Reviewer testing found that V4-Pro, while strong on structured tasks, shows degraded reliability on complex prompts with many simultaneous constraints, where GPT-5.5 and Claude Opus 4.6 perform more consistently.
Long-horizon agentic reliability. On Terminal-Bench 2.0, V4-Pro scores 67.9% against GPT-5.4 xHigh's 75.1% on the same card, indicating a real gap in multi-step tool-use reliability over extended autonomous task runs.[2]
Scale deployment constraints. CFR analysis notes that compute shortages at DeepSeek itself limit V4 deployment at scale.[14] The full 865 GB V4-Pro model is demanding to self-host, and DeepSeek's own API was reported to be unable to serve V4-Pro to most customers in the days immediately after launch.
IP concerns. Multiple U.S. government officials and analysts have alleged that DeepSeek conducted large-scale distillation from proprietary U.S. frontier models during training. If accurate, this raises questions about the provenance of some capabilities. DeepSeek has not addressed these allegations publicly.
A staggered official release and short API lifetimes. As of August 1, 2026 only one of the two models had reached an official build, and that one was still in public beta; the announced official V4-Pro had no date, and Codex and Responses API support for deepseek-v4-pro was forecast for early August 2026.[6][7][33] Both gaps closed on August 13, when DeepSeek shipped V4-Pro-0813 with Responses API support and published its weights.[7][61] Less than a month later the official Flash build was itself retired from the API in favor of V4.1-Flash, and a plan to retire V4-Pro the same way was announced and then withdrawn. DeepSeek's documentation now promises only "further notice should there be any changes" to V4-Pro service.[7][73][74]
Benchmark reproducibility. DeepSeek's July 2026 agent scores were produced with an in-house harness that was unreleased at the time (DeepSeek published it as a developer preview on August 13, 2026),[64] and two of the nine reported benchmarks are internal test sets that no third party can run. The one directly comparable public figure came in about four points lower when Artificial Analysis measured it.[7][37]
See also
- DeepSeek
- DeepSeek V4-Flash
- DeepSeek V4-Pro
- DeepSeek V4.1-Flash
- DeepSeek-R1
- DeepSeek V3
- DeepSeek V3.2
- DSec
- Mixture of Experts
- Multi-head Latent Attention
- Speculative Decoding
- Artificial intelligence in China
- Reinforcement Learning from Human Feedback (RLHF)
- Group Relative Policy Optimization
- OpenAI Responses API
- OpenAI Codex
- Artificial Analysis
- LLM API pricing comparison
- Hugging Face
- vLLM
- SGLang
- llama.cpp
- Ollama
- LM Studio
- AI API cost and context planner
- AI model comparison finder
- AI model lifecycle calendar
References
- ^1 ^2 ^3 ^4 ^5 ^6 ^7DeepSeek V4 Preview Release, DeepSeek API Docs (April 24, 2026)
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11 ^12 ^13 ^14 ^15 ^16 ^17 ^18 ^19 ^20 ^21 ^22 ^23 ^24 ^25DeepSeek-V4-Pro model card, Hugging Face
- ^1 ^2DeepSeek-V4: a million-token context that agents can actually use, Hugging Face Blog (April 2026)
- ^DeepSeek-V4 Collection, Hugging Face
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11 ^12 ^13 ^14 ^15 ^16 ^17 ^18 ^19 ^20 ^21 ^22 ^23 ^24 ^25 ^26 ^27 ^28 ^29 ^30 ^31DeepSeek V4 technical report: "DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence" (2026), arXiv:2606.19348v1 PDF (previously hosted in the Hugging Face repository)
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10DeepSeek Models & Pricing, DeepSeek API Docs, Wayback Machine snapshot of August 1, 2026
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11 ^12 ^13 ^14 ^15 ^16 ^17 ^18 ^19 ^20 ^21 ^22 ^23 ^24 ^25 ^26 ^27 ^28 ^29 ^30 ^31 ^32 ^33 ^34 ^35 ^36 ^37 ^38 ^39 ^40 ^41 ^42 ^43 ^44 ^45DeepSeek API Changelog
- ^1 ^2 ^3 ^4 ^5 ^6DeepSeek's new models offer big inference cost savings, The Register (April 24, 2026)
- ^1 ^2Three reasons why DeepSeek's new model matters, MIT Technology Review (April 24, 2026)
- ^DeepSeek previews new AI model that 'closes the gap' with frontier models, TechCrunch (April 24, 2026)
- ^China's AI upstart DeepSeek drops new model, CNN Business (April 24, 2026)
- ^1 ^2DeepSeek unveils V4 model, with rock-bottom prices and close integration with Huawei's chips, Fortune (April 24, 2026)
- ^1 ^2DeepSeek launches 1.6 trillion parameter V4 on Huawei chips, Tom's Hardware (April 24, 2026)
- ^1 ^2 ^3 ^4 ^5DeepSeek V4 Signals a New Phase in the U.S.-China AI Rivalry, Council on Foreign Relations (April 24, 2026)
- ^1 ^2 ^3 ^4 ^5DeepSeek V4, ChinaTalk (April 2026)
- ^China's DeepSeek unveils latest models a year after upending global tech, Al Jazeera (April 24, 2026)
- ^1 ^2 ^3 ^4DeepSeek V4: almost on the frontier, a fraction of the price, Simon Willison (April 24, 2026)
- ^1 ^2DeepSeek AI Releases DeepSeek-V4: Compressed Sparse Attention and Heavily Compressed Attention Enable One-Million-Token Contexts, MarkTechPost (April 24, 2026)
- ^DeepSeek V4 arrives with near state-of-the-art intelligence, VentureBeat (April 2026)
- ^DeepSeek V4: Features, Benchmarks, and Comparisons, DataCamp (2026)
- ^DeepSeek Unveils Newest Flagship AI Model a Year after Upending Silicon Valley, Bloomberg (April 24, 2026)
- ^China's DeepSeek releases preview of long-awaited V4 model as AI race intensifies, CNBC (April 24, 2026)
- ^DeepSeek seeks outside funding for the first time at a $10B valuation, TechFundingNews (April 2026)
- ^Tencent, Alibaba Eye Investment in DeepSeek, Bloomberg via The Information (April 22, 2026)
- ^1 ^2DeepSeek-V3.2 Release, DeepSeek API Docs (December 1, 2025)
- ^DeepSeek-V3.2, Simon Willison (December 1, 2025)
- ^DeepSeek price slash fuels competition, hits Zhipu, Minimax, Invezz (April 27, 2026)
- ^DeepSeek Slashes V4-Pro API Pricing With Major Discount, Dataconomy (April 27, 2026)
- ^Who could gain from DeepSeek's V4 with China chips poised for stronger demand, South China Morning Post (April 2026)
- ^1 ^2 ^3DeepSeek-V4: The Singularity of the Cambrian Explosion of AI Applications in China Has Arrived, 36Kr (April 2026)
- ^1 ^2 ^3模型 & 价格 (Models & Pricing), DeepSeek API 文档, Wayback Machine snapshot of August 2, 2026
- ^1 ^2 ^3 ^4 ^5DeepSeek-V4-Flash Official API is now LIVE in public beta, @deepseek_ai on X (July 31, 2026)
- ^1 ^2 ^3 ^4 ^5 ^6 ^7Using the Responses API, DeepSeek API Docs, Wayback Machine snapshot of July 31, 2026
- ^1 ^2 ^3 ^4Integrate with Codex, DeepSeek API Docs (accessed July 31, 2026)
- ^Rate Limit & Isolation, DeepSeek API Docs (accessed July 31, 2026)
- ^1 ^2 ^3Thinking Mode, DeepSeek API Docs, Wayback Machine snapshot of July 31, 2026
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9DeepSeek V4 Flash 0731 scores 50 on the Artificial Analysis Intelligence Index, @ArtificialAnlys on X (July 31, 2026)
- ^DeepSeek V4 Flash: Intelligence, Performance & Price Analysis, Artificial Analysis (accessed July 31, 2026)
- ^DeepSeek V4 Pro: Intelligence, Performance & Price Analysis, Artificial Analysis (accessed July 31, 2026)
- ^1 ^2We are making our discount permanent, @deepseek_ai on X (May 22, 2026)
- ^1 ^2DeepSeek permanently reduces the price of its flagship V4 model by 75 percent, Engadget (May 23, 2026)
- ^1 ^2DeepSeek has made its temporary 75% price cut on the first-party V4 Pro API permanent, @ArtificialAnlys on X (May 23, 2026)
- ^Off-Peak Discounts Alert, @deepseek_ai on X (February 26, 2025)
- ^1 ^2 ^3 ^4 ^5 ^6deepseek-ai/DeepSeek-V4-Flash model card, Hugging Face (accessed July 31, 2026)
- ^1 ^2DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence, arXiv:2606.19348 (April 26, 2026)
- ^1 ^2Jiang, B., "After triggering price war, DeepSeek reverses course with surcharge on peak-hour API use," South China Morning Post (June 30, 2026)
- ^1 ^2Codex setup script and models.json model catalog for DeepSeek, DeepSeek CDN (accessed July 31, 2026)
- ^Pricing Changes: new pricing starts and off-peak discounts end Sep 5, 2025, @deepseek_ai on X (August 21, 2025)
- ^1 ^2The DeepSeek-V4-Pro discount has been extended until May 31, 2026, @deepseek_ai on X (April 29, 2026)
- ^DeepSeek puts V4-Flash API into public beta, TechNode (July 31, 2026)
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11deepseek-ai/DeepSeek-V4-Flash-0731 model card, Hugging Face (accessed August 1, 2026)
- ^1 ^2 ^3deepseek-ai/DeepSeek-V4-Pro-DSpark model card, Hugging Face (accessed August 1, 2026)
- ^1 ^2 ^3deepseek-ai/DeepSeek-V4-Flash-DSpark model card, Hugging Face (accessed August 1, 2026)
- ^DeepSpec: a full-stack codebase for training and evaluating draft models for speculative decoding, DeepSeek on GitHub
- ^1 ^2deepseek-ai/DeepSeek-V4-Flash-0731 file listing, Hugging Face (accessed August 1, 2026)
- ^Models & Pricing, DeepSeek API Docs, archived March 23, 2026: "deepseek-chat and deepseek-reasoner correspond to the model version DeepSeek-V3.2 (128K context limit)"
- ^deepseek-ai/DeepSeek-V3.2 config.json, Hugging Face
- ^1 ^2 ^3Huang, J. et al. (DeepSeek-AI), "DeepSeek Elastic Compute (DSec): A Sandbox Infrastructure for Effective Agentic Training at Scale," arXiv:2609.22978 (September 19, 2026)
- ^1 ^2DeepSeek-AI, "DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression," arXiv:2609.19969 (September 17, 2026), Sections 5.1.2-5.1.3
- ^"Update technical report" commit b5968e9 (June 22, 2026): DeepSeek_V4.pdf deleted, README changed, deepseek-ai/DeepSeek-V4-Pro, Hugging Face
- ^1 ^2 ^3 ^4 ^5 ^6deepseek-ai/DeepSeek-V4-Pro-0813 model card, Hugging Face (created August 13, 2026, accessed September 23, 2026)
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11 ^12 ^13 ^14 ^15 ^16 ^17Models & Pricing, DeepSeek API Docs (accessed September 23, 2026)
- ^1 ^2 ^3 ^4 ^5 ^6deepseek-ai/DeepSeek-V4-Flash-Vision-Exp model card, Hugging Face (created August 31, 2026, accessed September 23, 2026)
- ^1 ^2deepseek-ai/deepseek-harness repository, GitHub (created August 13, 2026; developer preview, MIT license)
- ^1 ^2 ^3 ^4 ^5 ^6Models & Pricing, DeepSeek API Docs, Wayback Machine snapshot of August 14, 2026 (scheduled peak/off-peak rates)
- ^模型 & 价格 (Models & Pricing), DeepSeek API 文档 (accessed September 23, 2026)
- ^Models & Pricing, DeepSeek API Docs, Wayback Machine snapshot of August 16, 2026, 17:14 UTC
- ^1 ^2Models & Pricing, DeepSeek API Docs, Wayback Machine snapshot of August 22, 2026 (weekend off-peak notice)
- ^Integrate with Codex, DeepSeek API Docs (accessed September 23, 2026)
- ^Using the Responses API, DeepSeek API Docs (accessed September 23, 2026)
- ^Integrate with Claude Code, DeepSeek API Docs (accessed September 23, 2026)
- ^1 ^2 ^3Thinking Mode, DeepSeek API Docs (accessed September 23, 2026)
- ^1 ^2Change Log, DeepSeek API Docs, Wayback Machine snapshot of September 10, 2026, 07:46 UTC (V4-Pro retirement notice)
- ^1 ^2 ^3Models & Pricing, DeepSeek API Docs, Wayback Machine snapshot of September 11, 2026, 16:55 UTC (retirement withdrawn)
- ^1 ^2 ^3Models & Pricing, DeepSeek API Docs, Wayback Machine snapshot of September 9, 2026 (last pre-V4.1 schedule)
- ^deepseek-ai/DeepSeek-V4.1-Flash model card, Hugging Face (created September 10, 2026, accessed September 23, 2026)
- ^Integrate with Claude Code, DeepSeek API Docs, Wayback Machine snapshot of July 26, 2026
- ^1 ^2 ^3 ^4Pricing, Claude API Docs, Wayback Machine snapshot of August 16, 2026
- ^1 ^2 ^3Pricing, OpenAI API documentation, Wayback Machine snapshot of August 15, 2026
- ^1 ^2GPT-5.5 model page, OpenAI API documentation (accessed September 23, 2026)
- ^1 ^2GPT-5.4 model page, OpenAI API documentation (accessed September 23, 2026)
- ^1 ^2 ^3Gemini Developer API pricing, Google AI for Developers, Wayback Machine snapshot of August 14, 2026
- ^Gemini 3.1 Pro Preview model page, Google AI for Developers (accessed September 23, 2026)
- ^1 ^2Introducing Claude Opus 4.6, Anthropic (February 5, 2026)
- ^Introducing Claude Opus 4.7, Anthropic (April 16, 2026)
- ^1 ^2Introducing GPT-5.5, OpenAI (April 23, 2026)
- ^1 ^2Gemini 3.1 Pro model card, Google DeepMind (February 2026)
- ^1 ^2moonshotai/Kimi-K2.6 model card, Hugging Face (accessed September 23, 2026)
- ^Kimi K2.6 Model Pricing, Kimi API Platform, Wayback Machine snapshot of August 19, 2026
- ^1 ^2zai-org/GLM-5.1 model card, Hugging Face (accessed September 23, 2026)
- ^zai-org/GLM-5 GitHub repository, model download index listing GLM-5.1 as 744B-A40B (accessed September 23, 2026)
- ^Pricing, Z.AI Developer Documentation, Wayback Machine snapshot of August 14, 2026
Improve this article
Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.
20 revisions · v21 · 12,275 words · full history
Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify
Research and drafting on this wiki are AI-assisted, under named human editorial standards. How AI is used here
Reviewer note: xg06 V1+V8+V11 independent verification: DSec section, staleness rewrite and every competitor cell of both comparison tables re-checked against vendor sources; all defects fixed 2026-09-23
Cite this page: AI Wiki. "DeepSeek V4." aiwiki.ai, updated 23 Sept 2026, fact-checked 23 Sept 2026. CC BY 4.0. https://aiwiki.ai/wiki/deepseek_v4