Citation and evidence

DeepSeek

25 min full readUpdated 39 references

This article's verification

Report a problem with this article

More

Use this article

Raw MarkdownExplore connections

Improve this page

Suggest editRevision historyDiscussion

Browse categories

AI CompaniesArtificial IntelligenceLarge Language ModelsOpen Source AI

Cite this article

Compare, price, and deploy models in this article

These sourced facts are shared with the comparison and planning tools. Each field retains its own review date.

deepseek-chat (legacy alias)

Article updated after the pricing / specification review. Inspect the sources before relying on these fields; a newer edit does not necessarily change the model facts.

Provider API lifecycle
retiredSource · checked 2026-07-23 · review due
Context tokens
1,000,000Source · checked 2026-07-23
Input / output per 1M tokens
No reviewed API rateSource · checked 2026-07-23 · review due

Legacy alias with a previously announced July 24, 2026 retirement. Current pricing documentation no longer lists this alias.

deepseek-reasoner (legacy alias)

Article updated after the pricing / specification review. Inspect the sources before relying on these fields; a newer edit does not necessarily change the model facts.

Provider API lifecycle
retiredSource · checked 2026-07-23 · review due
Context tokens
1,000,000Source · checked 2026-07-23
Input / output per 1M tokens
No reviewed API rateSource · checked 2026-07-23 · review due

Legacy alias with a previously announced July 24, 2026 retirement. Current pricing documentation no longer lists this alias.

DeepSeek is a Chinese artificial intelligence company based in Hangzhou. It was founded in 2023 by Liang Wenfeng, who had previously co-founded the quantitative investment firm High-Flyer.[1] The legal entity named in DeepSeek's current privacy policy is Hangzhou DeepSeek Artificial Intelligence Co., Ltd.[2] DeepSeek researches and releases large language models, distributes downloadable model weights, and operates a consumer chat service and an application programming interface.

The company became widely known outside China after the January 2025 release of DeepSeek-R1, a reasoning model trained with a substantial reinforcement learning component. DeepSeek publishes technical reports and model repositories, but the company, its hosted services, and each model release are separate subjects. Claims about a model's architecture, license, training run, or benchmark result do not automatically apply to the company or to every DeepSeek product.

DeepSeek grew out of High-Flyer's artificial intelligence research. Public corporate records reviewed by Reuters identified Liang as DeepSeek's controlling shareholder in early 2025, while High-Flyer had financed the company and supplied researchers and computing infrastructure.[3] This relationship explains DeepSeek's origin, but High-Flyer and DeepSeek are not interchangeable names.

History and organization

High-Flyer was founded in 2015 and used machine learning in quantitative trading. In 2019 it established an artificial intelligence research unit, and by 2022 it had assembled a cluster containing 10,000 Nvidia A100 graphics processors. Liang founded DeepSeek in 2023 as a separate company focused on artificial-intelligence model research.[1][3]

DeepSeek initially operated without publicly announced outside financing. That changed in 2026, although the available financial record remains incomplete because DeepSeek is privately held and has not published audited accounts or detailed financing documents. Reuters reported that the company raised about US$7.4 billion in June 2026 at a post-money valuation of about 450 billion yuan. The same report said Liang, Tencent, CATL, and other investors participated, and that DeepSeek was considering another round at about a 500 billion yuan valuation and an eventual mainland listing. Reuters put the planned round at as much as 50 billion yuan, gave the 500 billion yuan valuation as about US$74 billion, and closed the wire with its own currency footnote, "$1 = 6.7704 Chinese yuan renminbi." On the completed June round, the same story said Liang personally committed 20 billion yuan while Tencent and CATL put in 10 billion and 5 billion yuan to become the largest external shareholders. It listed China's national artificial-intelligence fund, NetEase, JD.com, IDG Capital, Loyal Valley Capital, Monolith Management, and Shixiang Capital as other investors, attributing that list to sources and media reports. Reuters emphasized that the proposed round and listing were at an early stage and could change.[4]

A Chinese stock-exchange filing supplied a different, indirect valuation signal. It said a fund had invested 2.90 billion yuan for an indirect 0.8265 percent stake, implying a valuation of 350.88 billion yuan, or about US$51.82 billion at the cited exchange rate. Reuters described the filing as rare public evidence about the first external round and noted that DeepSeek had not announced or detailed it.[5] On July 25, Reuters relayed a Bloomberg report that DeepSeek had paused its planned second round. Reuters said it could not independently verify that report, and Bloomberg's sources said the process might resume.[6] The pause did not hold. On August 6, 2026, PYMNTS, citing a Bloomberg report of the same day, said DeepSeek had restarted the second round and was seeking close to US$8 billion at a valuation of about US$74 billion, with discussions described as ongoing and subject to change.[31] On September 9, 2026, the South China Morning Post reported, citing two people it did not name, that DeepSeek had hired four underwriters including CITIC Securities and aimed to begin the process this year for a listing on the Shanghai Stock Exchange's STAR Market, while moving to complete a pre-investment round that could value it at about 500 billion yuan.[32] Every one of these figures is an unnamed-source transaction estimate rather than a current audited valuation, and the reports describe the listing itself as still at the preparation stage.

Public information does not support a precise current revenue, profit, ownership, or employee figure for the company. DeepSeek's model documentation contains engineering measurements, but those measurements are not company financial statements.

Research and model releases

DeepSeek has released general-purpose, coding, mathematical, vision-language, and reasoning models. The company page summarizes the main general model lineage; dedicated articles contain model-level detail.

Release periodModel or familyDocumented significance
November 2023DeepSeek LLMA 7 billion and 67 billion parameter language-model family trained on a two-trillion-token corpus.[7]
January 2024DeepSeek-CoderCode models from 1.3 billion to 33 billion parameters, trained on two trillion tokens with project-level code and a 16,000-token window.[8]
May 2024DeepSeek-V2A mixture-of-experts model that combined DeepSeekMoE with multi-head latent attention and was trained on 8.1 trillion tokens.[9]
December 2024DeepSeek V3A 671 billion parameter mixture-of-experts model with 37 billion parameters activated for each token.[10]
January 2025DeepSeek-R1A reasoning-model release accompanied by R1-Zero, the company's reinforcement-learning-only experiment, and distilled smaller models.[11]
August to December 2025DeepSeek V3.1 and DeepSeek V3.2Successive hosted and open-weight releases recorded in DeepSeek's API change log.[12]
April 2026DeepSeek V4A family with a 1.6 trillion parameter Pro model and a 284 billion parameter Flash model, both using 1 million-token context windows.[13]
July 2026DeepSeek-V4-Flash-0731The public-beta official release of the V4-Flash API. DeepSeek says the checkpoint kept the preview's architecture and size and was only re-post-trained.[12]
August 2026DeepSeek-V4-Pro-0813The general-availability release of V4 Pro across the app, web, and API, adding native Responses API support and low, high, and max thinking-effort levels.[12]
August 2026DeepSeek-V4-Flash-Vision-ExpAn experimental vision-understanding model offered on the API only, which DeepSeek described as matching V4-Flash on text tasks.[12]
September 2026DeepSeek V4.1-FlashA multimodal mixture-of-experts model with a 552 billion parameter backbone, a causal encoder-decoder layout, a 1 million-token context window, and MIT-licensed weights.[33][12]

Expanded article table

Architecture

DeepSeek-V2 and later general models use a mixture-of-experts design. Instead of activating all parameters for every token, a routing mechanism selects a subset of experts. V2 paired that design with multi-head latent attention, which compresses key-value representations to reduce inference memory. The V2 paper reports 236 billion total parameters with 21 billion activated for each token.[9]

V3 retained multi-head latent attention and DeepSeekMoE, while adding an auxiliary-loss-free strategy for balancing expert use, multi-token prediction, and low-precision FP8 training. Its paper reports 671 billion total parameters, 37 billion active parameters per token, and pretraining on 14.8 trillion tokens.[10] These are developer-reported architectural and training specifications, not independently audited measures of capability.

V4 changed the scale and attention system. The technical report describes V4 Pro as a 1.6 trillion parameter model with about 49 billion active parameters per token, and V4 Flash as a 284 billion parameter model with about 13 billion active. The reported pretraining corpus contained more than 32 trillion tokens. The report describes DeepSeek Sparse Attention, a multi-token prediction objective, a manifold-constrained hyper-connection method, and a modified Muon optimizer.[13] The paper and released checkpoints document the design, but they do not establish that the model is best for every task or deployment.

The V3 training-cost claim

The frequently repeated US$5.576 million figure is not DeepSeek's total research and development cost, the cost of DeepSeek-R1, or the company's total expenditure. The V3 paper reports 2.788 million H800 GPU-hours for the official V3 training run. It multiplies that quantity by an assumed rental price of US$2 per H800 GPU-hour to reach US$5.576 million. The paper explicitly excludes costs for prior research and ablation experiments.[10]

The number is useful for understanding the compute accounted for in one documented run. It cannot be compared directly with estimates that include salaries, data preparation, failed experiments, hardware acquisition, inference, or earlier model development. DeepSeek has not published an audited total cost for developing V3 or R1.

R1 and reinforcement learning

The R1 research program separated two related systems. DeepSeek-R1-Zero was trained from a base model with large-scale reinforcement learning and no supervised fine-tuning stage. DeepSeek reported that the experiment developed behaviors such as verification and longer reasoning, but also produced readability and language-mixing problems. The production R1 procedure added a small cold-start data set, reasoning-oriented reinforcement learning, rejection sampling, supervised fine-tuning, and a final reinforcement-learning stage.[11]

The R1 paper was published in Nature in September 2025 after peer review. The paper documents the training procedure, evaluation set-up, limitations, and safety testing. Peer review establishes that the research report passed the journal's review process; it does not certify every benchmark result, hosted response, or later DeepSeek model.[11]

DeepSeek also released six distilled models derived from Qwen and Llama base models using data generated by R1. They are separate checkpoints with different sizes and licenses from the full R1 model. Results reported for a distilled model should not be labeled simply as results for R1.

V4 and independent evaluation

DeepSeek announced a preview release of V4 Pro and V4 Flash on April 24, 2026. Its technical report compares the models with other systems on coding, reasoning, science, and agentic benchmarks.[13] Those tables are developer evaluations and depend on prompts, inference budgets, tool scaffolding, model versions, and possible benchmark exposure.

The U.S. National Institute of Standards and Technology's Center for AI Standards and Innovation independently evaluated V4 Pro in April 2026. CAISI called it the most capable model from the People's Republic of China that the center had evaluated at that time, across its selected cyber, software-engineering, natural-science, abstract-reasoning, and mathematics tasks. Its aggregate method placed V4 roughly eight months behind the model frontier used in the comparison. CAISI also found V4 more cost-efficient than GPT-5.4 mini on five of seven tested benchmarks, with results across the seven ranging from 53 percent less expensive to 41 percent more expensive. CAISI reported that DeepSeek's own benchmark picture was more favorable than CAISI's results.[14]

Those findings describe one evaluation suite and serving configuration, not a universal ranking. The result is useful precisely because it provides an independent comparison and makes its methodology and limits visible.

V4.1-Flash

DeepSeek released DeepSeek-V4.1-Flash on September 10, 2026. Its model card describes a multimodal mixture-of-experts model with a 552 billion parameter backbone and support for contexts up to one million tokens, arranged as a 40-layer Causal Encoder-Decoder in which a 20-layer causal encoder is followed by a 20-layer decoder. The card says that arrangement lets the model activate about 8 billion parameters for each input token during prefill and 16 billion for each output token during decoding, and it lists a further 196 billion parameters of Engram conditional memory that is reached by hashed lookup rather than by dense computation. The repository and the weights are under the MIT License.[33] DeepSeek's change log entry for the same date says the hosted model name became deepseek-flash, that the earlier V4-Flash and V4-Flash-Vision-Exp names were retired and routed to the new model, and that API prices were reduced with the release. The same entry says DeepSeek decided, in response to user demand, to continue providing API service for V4 Pro after September 14, 2026 with billing unchanged.[12][35] The architecture, training corpus, evaluation tables, and service limits are covered at DeepSeek V4.1-Flash; as with earlier releases, the published benchmark table is DeepSeek's own.

Cost and efficiency positioning

Much of what DeepSeek publishes about its models concerns the cost of running them: cache size, activated parameters per token, sparse attention, and repeated price cuts. Whether that adds up to a deliberate competitive strategy is a matter of interpretation. The underlying figures, though, are published and checkable, and they do not all point the same way.

The most direct primary series is DeepSeek's own measurement of global key-value cache size per token, published as a figure in the V4.1-Flash release materials:

ModelDate label on DeepSeek's figureGlobal KV cache per token
DeepSeek-V12023.11389,120 bytes
DeepSeek-V3.22025.1248,068 bytes
DeepSeek-V4-Flash2026.043,514 bytes
DeepSeek-V4.1-Flash2026.09890 bytes

Expanded article table

The V4.1-Flash model card gives the same 890 bytes per token and describes the reduction as roughly fourfold against V4-Flash and roughly 437-fold against V1.[33][34] These are DeepSeek's own measurements of one specific quantity in its own serving stack, not independently reproduced results, and "global KV cache" is narrower than total cache memory: the card separately reports the persistent cache footprint falling to about one eighth of V4-Flash's, a different ratio for a different quantity. The change-log dates line up with the figure's labels, with DeepSeek-V3.2 reaching the API on December 1, 2025 and V4-Flash on April 24, 2026.[12]

Pricing has moved the same way and is documented in the API change log. DeepSeek added disk-based context caching in August 2024 and described it as cutting prices by another order of magnitude.[12] With the V4 general-availability release it introduced peak and off-peak pricing, with off-peak rates set at half the peak rates, effective 16:00 UTC on August 16, 2026, and it reduced prices again with V4.1-Flash.[12] As listed on September 11, 2026, the peak price for deepseek-flash was US$0.30 per million input tokens, US$1.20 per million output tokens, and US$0.006 per million cached input tokens, while deepseek-v4-pro was US$1.32, US$3.96, and US$0.044 for the same three classes. Peak hours are 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays; everything else is off-peak at half price.[35]

Independent measurement gives a more mixed picture. Artificial Analysis scored V4.1-Flash at 40 on its Intelligence Index and V4-Pro-0813 at 36, with a weighted cost of US$0.27 and US$0.67 respectively for each Intelligence Index task. It also recorded V4.1-Flash generating 250 million output tokens over the index, against a 140 million median for its comparison class, and called the model "very verbose" while describing it as fast and reasonably priced for an open-weight model of its size.[36][37] Artificial Analysis versions its index, and scores produced under different versions are not comparable, so any figure taken from it needs both a version and a read date attached. The two above are Intelligence Index v4.3, read on September 11, 2026.

Cutting the other way is the CAISI evaluation described above, in which V4 Pro came out cheaper than GPT-5.4 mini on only five of the seven benchmarks tested, and the full spread ran from 53 percent cheaper to 41 percent more expensive. On that evidence the cost advantage was benchmark-dependent rather than uniform, and it came with an aggregate capability gap.[14] Cheaper serving is also not the same as cheaper development: as the V3 training-cost discussion above shows, DeepSeek's most-quoted cost number covers one training run's rented compute and excludes prior research, ablations, salaries, and hardware. What DeepSeek publishes on cost is engineering measurement and a price list, which are different things from audited financial disclosure.

Products and distribution

DeepSeek distributes its work through three principal channels:

  • downloadable model repositories, including weights and inference instructions;
  • a hosted API for developers;
  • a consumer chat service available through web and mobile applications.

The downloadable models and the hosted service are not the same product. A locally run checkpoint is operated by the person or organization deploying it. The hosted service is operated under DeepSeek's terms, privacy policy, availability rules, content controls, and model routing. In the April 2026 preview, DeepSeek made V4 Pro and V4 Flash available in its official repositories, and placed V4 in its chat service and API with a 1 million-token context window.[15] The API change log says the older deepseek-chat and deepseek-reasoner services were discontinued on July 24, 2026, after a transition to V4 endpoints.[12]

DeepSeek's terms say users must review outputs before relying on them, require human review for decisions with substantial effects on an individual, and prohibit several categories of unlawful or harmful use. The terms also state that the company does not guarantee the accuracy, completeness, security, or uninterrupted availability of outputs or services.[16] These provisions apply to DeepSeek's service relationship. They do not change the technical behavior or license of a separately downloaded model.

The current privacy policy says the service may collect account details, prompts, uploaded files, chat history, feedback, device and network information, approximate location, and payment information. It says information collected through the service is stored on servers in the People's Republic of China. The policy also says users can opt out of the use of their inputs for model training.[2] Local deployment can avoid sending prompts to DeepSeek's hosted service, but its actual privacy depends on the local operator, software stack, logging, and infrastructure.

Licensing and openness

DeepSeek describes many releases as open source, but the more precise label varies by release. Open weights means that trained parameters can be downloaded. It does not by itself mean that the training data, complete data-processing pipeline, source code, and every component needed to reproduce the model are available.

DeepSeek-V3 illustrates the distinction. Its repository code is under the MIT License, while the model weights are governed by a separate DeepSeek Model License that permits commercial use but includes use restrictions.[17] The R1 repository and full R1 weights are under the MIT License. Its distilled checkpoints are also subject to the licenses of the Qwen or Llama base models from which they were derived.[18] The V4 Pro and Flash model cards designate the models as MIT-licensed.[19]

The Open Source Initiative's Open Source AI Definition treats freedom to use, study, modify, and share an AI system as central, and calls for access to the preferred form for modification, including sufficient information about training data as well as code and parameters.[20] Under that broader framework, a downloadable weight file is not alone sufficient to establish open-source AI. This article therefore uses "open-weight" when the verifiable fact is weight availability and states the model-specific license separately.

DeepSeek's technical reports disclose architecture, optimization methods, data quantities, and evaluation procedures. They do not release the full pretraining corpus or a completely reproducible end-to-end training pipeline for the general models discussed here.

Reception and impact

DeepSeek's consumer application rose to the top of Apple's U.S. free iPhone application chart on January 27, 2025, shortly after the R1 release. On the same day, Nvidia shares fell 17 percent amid a broader selloff in technology companies exposed to artificial-intelligence spending. Associated Press reporting described investor concern that competitive models might be built or operated with less computing expenditure, while also noting uncertainty about the comparability and completeness of DeepSeek's cost claims.[21] The episode is covered separately at DeepSeek market crash (Jan 2025).

Researchers were interested in R1 for reasons beyond the market reaction. Nature reported that scientists were testing the model's reasoning behavior and examining its training approach, and described R1 as an open-weight model rather than a fully reproducible open-source system.[22] The later peer-reviewed paper gave researchers a more complete account of the training procedure.[11]

The releases also complicated simple claims that capability depends only on increasing dense model size or spending a particular amount of money. DeepSeek's work supplied evidence for the practical value of sparse expert routing, compressed attention, reinforcement learning, and distillation. It did not show that compute, data quality, hardware, or engineering ceased to matter. V4, for example, is substantially larger than V3 and was trained on more than twice as many reported tokens.[10][13]

Privacy, security, and regulatory responses

In January 2025, cloud-security company Wiz reported finding an internet-accessible DeepSeek ClickHouse database that required no authentication. Wiz said the database exposed more than one million log entries, including chat history, API-related information, and operational metadata. The researchers disclosed the finding to DeepSeek and reported that access was secured soon afterward.[23] The report documented an exposed system at that time; it does not establish that every current deployment has the same flaw.

Several governments and regulators then took actions with different legal scopes:

  • Italy's data-protection authority declared the processing described in its order unlawful and, as a matter of urgency, ordered a definitive limitation on processing the personal data of people in Italy. The limitation took effect immediately upon receipt of the January 30, 2025 order, while the authority reserved further decisions until its investigation was completed. It applied to Italian users' data, not as a global ban on model downloads.[24]
  • South Korea's Personal Information Protection Commission investigated the service after app downloads were suspended in February 2025. In April it said DeepSeek had transferred user inputs and device, network, and application information abroad without an adequate legal basis or required disclosures. The regulator said DeepSeek stopped sending user inputs to Volcano Engine on April 10 and ordered corrective measures, including destruction of previously transferred inputs and improved safeguards.[25]
  • Australia's Protective Security Policy Framework Direction 001-2025 required federal government entities to prevent use or installation of DeepSeek products on government systems and devices and to remove existing instances. The direction applies to Australian government technology; it is not a general prohibition on private use in Australia.[26]

These measures concern the hosted application, data processing, or government systems. They do not necessarily prohibit downloading a model in every jurisdiction. Organizations considering deployment must separately assess the applicable model license, data flows, security configuration, and local law.

Evaluation limits and disputed claims

DeepSeek's technical papers report results from the developer's own benchmark runs. Such results can be informative, but they are not interchangeable with independent evaluations. Prompt templates, reasoning budgets, tool access, sampling, contamination, serving hardware, and later checkpoint updates can change measured performance. CAISI's V4 evaluation found a less favorable aggregate comparison than DeepSeek's self-reported tables, while still finding competitive capability and cost on several tested tasks.[14]

Hosted-service behavior should also be distinguished from downloadable weights. The service operates under DeepSeek's terms and content rules, while a local operator controls the surrounding software for an open-weight checkpoint. Neither setting guarantees factual output. DeepSeek's privacy policy specifically warns that generated information may be incomplete, incorrect, or outdated.[2]

Distillation allegations

In January 2025, OpenAI and a U.S. government adviser said they suspected that DeepSeek had used outputs from OpenAI systems to improve its models through distillation. Associated Press reported that neither had publicly disclosed specific evidence at that time and that DeepSeek did not respond to the outlet's request for comment.[27] During peer review of the R1 paper, DeepSeek's authors said R1's performance did not depend on distilling reasoning patterns from other large language models. Nature reported that statement in September 2025; it was the authors' position, not an independent reconstruction of the training data.[28]

In February 2026, Anthropic alleged that it had identified a coordinated campaign associated with DeepSeek accounts that generated more than 150,000 exchanges with Claude. Anthropic said the activity targeted reasoning, reward-model, and content-filter-related tasks.[29] The allegation is supported by Anthropic's own detection report, but the underlying account and training records are not public. It should therefore be described as Anthropic's attribution, not as an adjudicated finding.

Anthropic returned to the subject in a threat-intelligence report published on September 10, 2026. It said that since February 2026 it had detected and disrupted unauthorized distillation campaigns that it attributed "with high confidence" to specific laboratories based in the People's Republic of China, DeepSeek was one of several it named, along with Alibaba, Moonshot AI, Zhipu, Xiaomi, SenseTime, and MiniMax. In DeepSeek's case Anthropic said the company built a chain-of-thought extraction pipeline that defeated one of Anthropic's own countermeasures: rather than returning raw reasoning, Anthropic's interface returns a reference to it, and the report says DeepSeek saved that reference, opened a fresh session, and prompted Claude to expand it back into the full reasoning trace. Anthropic put the observed scale at more than 12.1 million exchanges attributable to DeepSeek over 14 days in July 2026.[38]

The public record does not resolve whether, or to what degree, proprietary rival outputs were used in any specific DeepSeek model. Distillation from a model's own outputs, licensed models, or permissibly obtained data is a standard technical method; questions about unauthorized access or terms-of-service violations depend on the data source and legal facts.

Request-routing allegation

Anthropic's September 2026 report makes a second claim about DeepSeek that is distinct from distillation. Training on another company's outputs and serving another company's model to one's own paying users are different accusations, and the report treats them separately. Under the tracked-group designation GTG-16001, headed "DeepSeek serves Claude instead of its own models and collects exchanges for model training," Anthropic said DeepSeek inspected strings in inbound requests to identify users arriving through third-party or Anthropic coding harnesses such as Claude Code, the Claude Agent SDK, and OpenCode, tagged those users, and relayed selected tagged requests to Claude Opus. Anthropic said those customers were likely not made aware that their requests were being sent to Anthropic, and it gave three illustrations: an employee of a technology company in China whose analysis of internal program documentation was relayed, an information-technology operator working with a Russian government agency that the report associates with its Ministry of Defense, whose relayed requests exposed live database credentials, and engineers building a case-management tool for a municipal public-security bureau in China.[38]

This rests on Anthropic's own detection and telemetry. The published report gives techniques and aggregate counts rather than request logs or account identifiers, and it locates the activity by observation window, 14 days in July 2026, rather than by DeepSeek model version. Covering the report's release on September 10, 2026, CNBC said DeepSeek, Alibaba, Moonshot, Xiaomi, and Anthropic did not immediately respond to its requests for comment.[39] As with the February 2026 allegation, this is Anthropic's attribution rather than an adjudicated finding.

Hardware allegation

In March 2026, Reuters reported that an unnamed senior U.S. official alleged DeepSeek had trained V4 using advanced Nvidia Blackwell chips obtained despite U.S. export controls. Reuters said the official did not disclose the basis for the claim, DeepSeek and Nvidia did not comment, and China's foreign ministry said it was unaware of the specifics and opposed politicizing technology and trade.[30] The V4 paper does not identify its training hardware.[13] Without public technical or documentary evidence establishing the cluster, the hardware claim remains an attributed allegation.

References

  1. ^1 ^2Associated Press, "DeepSeek founder Liang Wenfeng built an AI company that shook the world," January 28, 2025. apnews.com/...-ai-0673d5c39d90108189cc31b88d85b9f8
  2. ^1 ^2 ^3DeepSeek, "Privacy Policy," last updated February 10, 2026. cdn.deepseek.com/...deepseek-privacy-policy
  3. ^1 ^2Reuters, "What is DeepSeek and why is it disrupting the AI sector?" January 28, 2025. investing.com/...-disrupting-the-ai-sector-3832137
  4. ^Reuters, "China's DeepSeek to raise fresh capital at $74 billion valuation ahead of onshore IPO, sources say," July 15, 2026 (Hong Kong/Beijing dateline; full wire text including the closing currency footnote, read via the Yahoo Finance syndication on September 11, 2026). finance.yahoo.com/...raise-fresh-capital-095807183
  5. ^Reuters, "Chinese filing implies DeepSeek valuation of around $52 billion," July 16, 2026. investing.com/...tion-of-around-52-billion-4796314
  6. ^Reuters, "DeepSeek tells prospective investors of funding pause, Bloomberg News reports," July 25, 2026. investing.com/...se-bloomberg-news-reports-4812797
  7. ^DeepSeek AI, "DeepSeek LLM: Scaling Open-Source Language Models with Longtermism," arXiv, January 2024. arxiv.org/...2401.02954
  8. ^DeepSeek AI, "DeepSeek-Coder: When the Large Language Model Meets Programming," arXiv, January 2024. arxiv.org/...2401.14196
  9. ^1 ^2DeepSeek AI, "DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model," arXiv, June 2024. arxiv.org/...2405.04434
  10. ^1 ^2 ^3 ^4DeepSeek AI, "DeepSeek-V3 Technical Report," arXiv, February 2025 revision. arxiv.org/...2412.19437
  11. ^1 ^2 ^3 ^4DeepSeek AI, "DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning," Nature, September 17, 2025. doi.org/...s41586-025-09422-z
  12. ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10DeepSeek, "API updates." api-docs.deepseek.com/updates
  13. ^1 ^2 ^3 ^4 ^5DeepSeek-AI, "DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence," arXiv:2606.19348v1, submitted April 26, 2026. arxiv.org/...2606.19348
  14. ^1 ^2 ^3National Institute of Standards and Technology, "CAISI Evaluation of DeepSeek V4 Pro," May 1, 2026. nist.gov/...caisi-evaluation-deepseek-v4-pro
  15. ^DeepSeek, "DeepSeek-V4 Preview Release," April 24, 2026. api-docs.deepseek.com/...news260424
  16. ^DeepSeek, "Terms of Use." cdn.deepseek.com/...deepseek-terms-of-use
  17. ^DeepSeek AI, "DeepSeek-V3 repository and license," GitHub. github.com/...DeepSeek-V3
  18. ^DeepSeek AI, "DeepSeek-R1 repository and model license notes," GitHub. github.com/...README.md
  19. ^DeepSeek AI, "DeepSeek-V4-Pro model card," Hugging Face. huggingface.co/...DeepSeek-V4-Pro
  20. ^Open Source Initiative, "The Open Source AI Definition." opensource.org/...open-source-ai-definition
  21. ^Associated Press, "Chinese AI startup DeepSeek overtakes ChatGPT on Apple App Store," January 27, 2025. apnews.com/...ina-f4908eaca221d601e31e7e3368778030
  22. ^Nature, "Chinese AI DeepSeek-R1 stuns scientists with its efficiency," January 30, 2025. nature.com/...d41586-025-00229-6
  23. ^Wiz Research, "Wiz Research uncovers exposed DeepSeek database leaking sensitive information, including chat history," January 29, 2025. wiz.io/...-uncovers-exposed-deepseek-database-leak
  24. ^Garante per la protezione dei dati personali, "DeepSeek: the Italian SA orders limitation of processing," January 30, 2025. garanteprivacy.it/...10098477
  25. ^Personal Information Protection Commission, "Results of preliminary inspection of DeepSeek service," April 24, 2025. pipc.go.kr/...noticeDetail.do
  26. ^Australian Government, "Direction 001-2025: DeepSeek products, applications and web services," February 4, 2025. protectivesecurity.gov.au/...ions-and-web-services
  27. ^Associated Press, "OpenAI says Chinese rivals use its work for their AI apps," January 29, 2025. apnews.com/...ght-a94168f3b8caa51623ce1b75b5ffcc51
  28. ^Nature, "DeepSeek says its success did not rely on OpenAI models," September 19, 2025. nature.com/...d41586-025-03015-6
  29. ^Anthropic, "Detecting and preventing distillation attacks," February 23, 2026. anthropic.com/...d-preventing-distillation-attacks
  30. ^Reuters, "China's DeepSeek trained AI model on Nvidia's best chip despite U.S. ban, official says," March 4, 2026. investing.com/...pite-us-ban-official-says-4520307
  31. ^PYMNTS, "DeepSeek Resumes Funding Round to Raise $8 Billion," August 6, 2026, reporting a Bloomberg News story of the same date. pymnts.com/...mes-funding-round-to-raise-8-billion
  32. ^South China Morning Post, "Chinese AI firm DeepSeek taps underwriters including Citic Securities for IPO: sources," September 9, 2026. scmp.com/...including-citic-securities-ipo-sources
  33. ^1 ^2 ^3DeepSeek AI, "DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression," Hugging Face model card, September 10, 2026. huggingface.co/...DeepSeek-V4.1-Flash
  34. ^DeepSeek AI, "Global KV cache size per token across generations of DeepSeek models" (Figure 1b), DeepSeek-V4.1-Flash repository, September 10, 2026. huggingface.co/...dsv41_kv_cache.png
  35. ^1 ^2DeepSeek, "Models & Pricing," API documentation, read September 11, 2026. api-docs.deepseek.com/...pricing
  36. ^Artificial Analysis, "DeepSeek V4.1 Flash (Reasoning, Max Effort) Intelligence, Performance & Price Analysis," Intelligence Index v4.3, read September 11, 2026. artificialanalysis.ai/...deepseek-v4-1-flash
  37. ^Artificial Analysis, "DeepSeek V4 Pro 0813 (Reasoning, Max Effort) Intelligence, Performance & Price Analysis," Intelligence Index v4.3, read September 11, 2026. artificialanalysis.ai/...deepseek-v4-pro
  38. ^1 ^2Anthropic, "Detecting and countering misuse of AI: September 2026," September 10, 2026. anthropic.com/...ntelligence-report-september-2026
  39. ^CNBC, "Chinese AI labs secretly used millions of Claude exchanges to train their models, Anthropic says," September 10, 2026. cnbc.com/...bs-moonshot-deepseek-alibaba-anthropic

Improve this article

Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.

11 revisions · v12 · 5,037 words · full history

Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify

Research and drafting on this wiki are AI-assisted, under named human editorial standards. How AI is used here

Reviewer note: Release timeline, funding reporting and the two Anthropic allegations checked against DeepSeek's own change log, the Reuters wire as syndicated, and Anthropic's September 2026 threat-intelligence report on September 11, 2026. Both allegations are recorded as Anthropic's attributions, not adjudicated findings.

Cite this page: AI Wiki. "DeepSeek." aiwiki.ai, updated 11 Sept 2026, fact-checked 11 Sept 2026. CC BY 4.0. https://aiwiki.ai/wiki/deepseek

Suggest edit