# DeepSeek V4.1-Flash

> Source: https://aiwiki.ai/wiki/deepseek_v4_1_flash
> Updated: 2026-09-30
> Fact-checked: 2026-09-23
> Categories: AI Models, Chinese AI, Large Language Models, Mixture of Experts, Multimodal AI, Open Source AI
> License: CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) - attribute to "AI Wiki (aiwiki.ai)"
> Cite as: AI Wiki. "DeepSeek V4.1-Flash." aiwiki.ai, 30 Sept 2026. https://aiwiki.ai/wiki/deepseek_v4_1_flash
> From AI Wiki (https://aiwiki.ai), the free encyclopedia of artificial intelligence. Reuse freely with attribution.

**DeepSeek V4.1-Flash** is an open-weight multimodal [mixture-of-experts model](https://aiwiki.ai/wiki/mixture_of_experts) released by [DeepSeek](https://aiwiki.ai/wiki/deepseek) on September 10, 2026. It accepts text and images and generates text. DeepSeek describes the model as a 552-billion-parameter language backbone supplemented by 196 billion parameters of Engram conditional memory. The backbone activates about 8 billion parameters per input token during prefill and 16 billion per generated token during decoding. Hugging Face's automatic parameter count for the complete published checkpoint, which also holds a separately trained speculative-decoding module, is about 763 billion. Its released configuration supports a [context window](https://aiwiki.ai/wiki/context_window) of 1,048,576 tokens [1][2][3][12].

V4.1-Flash introduced a 40-layer Causal Encoder-Decoder architecture, Compressed Sparse Attention 2, FP4 storage for the main [KV cache](https://aiwiki.ai/wiki/kv_cache), Engram lookup memory, and a DSpark [speculative decoding](https://aiwiki.ai/wiki/speculative_decoding) module. These mechanisms target the cost of repeatedly processing and caching long prompts, especially in tool-using agent workloads. DeepSeek reports that the global KV footprint is 890 bytes per token, about one-quarter that of the earlier [DeepSeek V4-Flash](https://aiwiki.ai/wiki/deepseek_v4_flash), and that its deployment method reduces persistent KV storage to about one-eighth of the earlier model's footprint at the same sequence length [1][2]. These comparisons are vendor measurements, not independently reproduced results.

The model weights and accompanying code are available from [Hugging Face](https://aiwiki.ai/wiki/hugging_face) under the [MIT License](https://aiwiki.ai/wiki/mit_license). The repository includes a prompt encoder, a readable inference implementation, an evaluation recipe, model configuration, and a technical report. DeepSeek labels the included inference code as a reference implementation rather than a production serving engine [1][10][12].

## Release and model identity

DeepSeek announced V4.1-Flash through its official social account and API update log on September 10, 2026. The hosted API identifier at launch was `deepseek-flash`. The same entry said the previous-generation V4-Flash and V4-Flash-Vision-Exp models had been retired, and that the `deepseek-v4-flash` and `deepseek-v4-flash-vision-exp` names would temporarily route to V4.1-Flash [4][5].

The launch-day entry also announced the retirement of V4-Pro. It said "extensive testing shows that V4.1 Flash now outperforms DeepSeek V4 Pro across performance, cost, speed, and total time, so we plan to retire V4 Pro in an orderly manner", and that after 12:00 Beijing time on September 14, 2026, which is 04:00 UTC, all `deepseek-v4-pro` requests would be routed to V4.1-Flash and billed at V4.1-Flash prices until a future V4.1-Pro release [21].

DeepSeek withdrew that plan the next day. A snapshot of the change log archived at 10:15 UTC on September 11 still carried the retirement notice; by later the same day DeepSeek had replaced it, in both the English and Chinese versions of the entry, with a statement that "[i]n response to user demand, we have decided to continue providing API services for DeepSeek V4 Pro after September 14, 2026, with the billing method remaining unchanged." The Models and Pricing page carries the same note and still lists `deepseek-v4-pro` as a separate model served by DeepSeek-V4-Pro-0813, with its own prices and a 500-request concurrency limit [5][6][21]. So the announced retirement was withdrawn before its own September 14 cutover date, and the routing change described on launch day is no longer scheduled.

V4.1-Flash is distinct from both the older V4-Flash family and [DeepSeek V4-Pro](https://aiwiki.ai/wiki/deepseek_v4_pro). The earlier V4-Flash model card described a 284-billion-parameter backbone with 13 billion active parameters per token. V4.1-Flash instead has a 552-billion-parameter backbone, separates prefill and decode activation at 8 billion and 16 billion parameters, incorporates a native vision pathway, and uses a new prompt format [1][2].

### Parameter accounting

Three parameter totals circulate for this model: 552 billion, 748 billion and 763 billion. Two of them describe the checkpoint; the third is just the first two added together.

DeepSeek's model card and technical report both describe 552 billion backbone parameters, and the report repeats the figure in its configuration summary after listing the layer, expert and vision-encoder dimensions. The same sources separately allocate 196 billion parameters to the two Engram conditional-memory modules, which are sparsely accessed lookup tables rather than dense computation [1][2]. Hugging Face derives its own count from the weight files, excluding quantization scale tensors, and reports 763,205,315,794 parameters for this repository [12]. That is roughly 15 billion more than 552 plus 196.

The gap is the DSpark speculative-decoding module. The repository stores DSpark as three blocks named `mtp.0` through `mtp.2`, each holding its own attention, mixture-of-experts and projection weights, and neither the model card nor the technical report folds them into a published parameter total [2]. Summing the tensor shapes recorded in the repository's safetensors headers, counting each FP4-packed routed-expert weight as one parameter rather than as half a byte and excluding the scale tensors that Hugging Face also excludes, gives this split:

| Component | Parameters, counted from the published weight files |
|---|---:|
| 40 backbone layers, embeddings, output head, vision encoder and projector | 552.05B |
| Engram conditional memory at layers 1 and 14 | 196.93B |
| DSpark modules `mtp.0` to `mtp.2` | 14.23B |
| Total | 763.21B |

Those three figures add to Hugging Face's total exactly. "552B" and "763B" are therefore both accurate descriptions of the same checkpoint: the first is DeepSeek's backbone count, the second is every weight tensor in the repository. The sum 552 + 196 = 748 matches no single artifact and erases the distinction the primary sources draw, so this article reports the components separately. The activation figures, about 8 billion parameters per prefill token and 16 billion per decoded token, describe the backbone and do not change with the choice of total [1][2].

## Architecture

V4.1-Flash is a multimodal [large language model](https://aiwiki.ai/wiki/large_language_model) built as a 40-layer causal Transformer. Its lower 20 layers form a causal encoder, and its upper 20 layers form a decoder. Unlike a conventional sequence-to-sequence encoder, the causal encoder cannot attend to future tokens. It prepares global key-value representations that the decoder can reuse while preserving autoregressive generation [1][2].

| Component | Released configuration |
|---|---|
| Language backbone | 40 layers, 5,120 hidden size, 64 attention heads, one KV head |
| Layer split | 20-layer causal encoder and 20-layer decoder |
| Context length | 1,048,576 positions |
| Local attention | 128-token sliding window |
| Experts | 384 routed experts plus one shared expert; six routed experts selected per token |
| Activated backbone parameters | About 8B per token during prefill, 16B during decode |
| Conditional memory | 196B Engram parameters across two modules |
| Vision encoder | 32 layers, 1,024 hidden size, 16 heads, 14-pixel patches |
| Sparse selection | Top 512 positions; later decoder indexers search a 16,384-position candidate pool in the released setup |

The table summarizes the public report and configuration file [2][3]. The context limit and parameter counts describe the released model, while achievable request length and throughput also depend on a serving implementation and available memory.

### Causal Encoder-Decoder

During ordinary Transformer prefill, every prompt token passes through every layer, and each layer produces its own global keys and values. V4.1-Flash's Causal Encoder-Decoder, or CED, instead projects the decoder's global keys and values from the final causal-encoder hidden states using decoder-layer-specific projection weights. Most prompt tokens therefore require full computation through only the lower half of the network for the global-attention path [2].

The decoder still computes layer-specific sliding-window attention. For a prompt of length `N`, a network depth `L`, and a local window `n`, the report models CED prefill complexity as approximately `O(NL/2 + nL/2)` rather than `O(NL)` when `N` is much larger than the local window. DeepSeek reports that this nearly halves prefill computation in its implementation while keeping performance close to its comparison baseline [2]. The design is related to the earlier YoCo architecture, which likewise shares cached representations across upper layers, but V4.1-Flash adds its own global-cache projections, local-attention path, and cache-reconstruction system [2][13].

### Compressed Sparse Attention 2

Compressed Sparse Attention 2, or CSA2, reduces global-attention storage and indexing work along the token and layer dimensions. Every CSA2 layer creates its own query and local sliding-window keys and values. Its three static modes differ in how the global cache and sparse selection are obtained [2].

- **Full mode** computes new main keys and values, index keys and queries, and a fresh top-k selection.
- **Reindex mode** reuses main keys and values plus index keys from an earlier layer, but computes a new query and new selection.
- **Reuse mode** reuses both the earlier global cache and the most recent compatible top-k selection.

This separation lets some layers update the attended positions without keeping another complete global cache, while other layers avoid both cache creation and rescoring. CSA2 also removes an overlapping compression pattern and absolute position embedding used by the preceding CSA design, and it derives index keys from main KV entries rather than through a separate hidden-state compression path [2].

The decoder uses a hierarchical sparse indexer. Its first Full-mode CSA2 layer scores all causally visible global positions, selects blocks by their maximum index score, and creates a shared candidate pool. In the released configuration, the pool comprises 2,048 blocks of eight positions, or 16,384 candidate positions. Later Reindex layers search that fixed pool before selecting 512 positions. The first global scan still grows with context length, but the later indexers' search size is bounded by the candidate-pool setting [2][3].

### KV-cache formats and bounded replay

The report distinguishes runtime global KV, which remains in accelerator memory, from persistent KV used to resume or reuse prefixes from SSD or host memory. V4.1-Flash stores the main global KV in FP4 and the local sliding-window KV in FP8. Along with cross-layer reuse, this produces a reported global-cache footprint of 890 bytes per token, about one-quarter of the V4-Flash footprint at equal sequence length in DeepSeek's stack [1][2][3].

Persisting every layer's local sliding-window state would weaken the storage saving. SWA Bounded Replay instead stores the long-lived global state and reconstructs the missing local state by replaying only the most recent window of tokens. The reconstruction is approximate: an exact reconstruction would need to replay the window through each relevant layer. DeepSeek reports negligible degradation in the conditions it tested and an overall persistent-cache footprint about one-eighth that of V4-Flash, but the report also identifies cache-resumption boundaries as an area requiring more stress testing [2].

### Experts, conditional memory, and decoding

Each feed-forward block uses DeepSeek's fine-grained MoE design. The released configuration has 384 routed experts, one shared expert, and six routed experts selected for each token. Image and text tokens use separate correction biases for load balancing, while the unadjusted routing scores determine how selected expert outputs are weighted [2][3].

Engram adds two sparsely accessed conditional-memory modules at backbone layers 1 and 14, using zero-based numbering. Each module hashes token n-grams of orders two through four into multiple embedding tables. Together the modules contain 196 billion parameters. Because the lookup address depends on the input sequence, embeddings can be prefetched from host memory while Transformer computation proceeds. The standalone Engram paper presents this as a way to separate some memorization capacity from dense computation through constant-time hashed lookup [2][14].

The architecture also includes Single-Pass mHC for residual-stream mixing and kernel fusion. DeepSeek's report says the implementation halves activation-memory traffic compared with its preceding four-kernel mHC implementation [2]. V4.1-Flash replaces the earlier multi-token-prediction module with DSpark, a separately trained speculative-decoding component that drafts groups of tokens and adjusts verification length according to confidence and serving load. DSpark is trained after the backbone pretraining stage and continues training alongside the backbone during post-training without sending its own objective gradients into the backbone [2][15].

## Multimodal pathway

V4.1-Flash accepts images through a vision encoder called DeepSeek-ViT. The encoder is based on a [Vision Transformer](https://aiwiki.ai/wiki/vision_transformer), but it replaces absolute position embeddings with two-dimensional rotary position encoding, uses a linear patch projection, RMS normalization, and SwiGLU activations. A 3 by 3 pixel-unshuffle step reduces the spatial token grid by a factor of nine. A two-layer MLP then projects the visual features into the language backbone's hidden dimension, where they are interleaved with text embeddings [1][2][3].

DeepSeek reports training the vision encoder from scratch in two preliminary stages. A contrastive stage used about 47 billion image-text pairs at resolutions no greater than 224 by 224 pixels. An autoregressive stage connected the encoder to a temporary 4-billion-parameter MoE language model and trained on 236 billion tokens drawn from image captions, alternative text, charts, and optical-character-recognition data at resolutions between 544 by 544 and 1,344 by 1,344 pixels. DeepSeek then discarded the temporary language model and retained the vision encoder for joint training with V4.1-Flash [2]. The report gives aggregate source types and counts but does not publish a source-by-source corpus inventory or license audit.

For the hosted API as documented on September 10, 2026, DeepSeek accepted JPEG, PNG, GIF, and WebP images supplied by URL, base64 data, or its Files API. The service limited an image to 1,024 visual tokens and documented request-level byte, dimension, and image-count limits. Those are service constraints and may differ from what a self-hosted runtime can support [7].

## Training and post-training

DeepSeek reports pretraining V4.1-Flash from scratch on 45 trillion tokens. Text-only and multimodal pipelines were combined with overlapping text-only samples replaced by their multimodal versions, producing a reported 7:1 ratio of text-only to multimodal tokens. Training began with [sparse attention](https://aiwiki.ai/wiki/sparse_attention) at a sequence length of 64,000 rather than using a dense-attention warmup. Context was extended to one million tokens after 34 trillion training tokens [1][2].

The disclosed optimizer mixture used AdamW for normalization weights and other non-matrix parameters, Muon for the main matrix parameters, a head-wise Muon variant for query and key matrices, and momentum updates followed by Sinkhorn balancing for Engram tables, token embeddings, and the prediction head. The training batch remained at 100.6 million tokens. These are author-reported details; neither the full training dataset nor a complete training system capable of reproducing the checkpoint was released [2].

Post-training followed supervised fine-tuning, reinforcement learning, and on-policy distillation. DeepSeek says the principal change from its earlier recipe was not a new optimization algorithm but a larger automated pipeline for constructing tool-use environments, coding tasks, evaluation points, and rollouts. Generated tasks were attempted by multiple agents and reviewed by a separate quality-inspection agent before entering training [1][2]. The report says V4.1 training raised demand on DeepSeek's [DSec](https://aiwiki.ai/wiki/dsec) sandbox platform to millions of concurrent sandbox instances. For RL across diverse agent scaffolds, each rollout was split into an agent sandbox that runs the scaffold and its tools and a worker container that orchestrates the rollout, normalizes interactions into a common trajectory schema and communicates with the trainer; both run on DSec outside the preemptible GPU training pool [2].

### Reasoning effort

The open-weight prompt format accepts an integer reasoning effort from 1 through 100. Its released encoder recognizes three string aliases: `low` maps to 50, `high` maps to 75, and `max` maps to 100; `high` is the default. DeepSeek's hosted API uses the same values for those three native tiers. Its compatibility layer also accepts additional requested labels: `minimal` and `low` route to native `low`; `medium`, `high`, and `xhigh` route to native `high`; and `max` and `ultra` route to native `max` [2][8][9]. The value 25 appears in DeepSeek's reasoning-effort ablation, but it is not a released string alias.

DeepSeek's ablation raised effort from 25 to 100. The report says average pass-at-one across eight reasoning-heavy benchmarks increased from 67.1 to 76.3, [DeepSWE](https://aiwiki.ai/wiki/deepswe) v1.1 increased from 66.0 to 74.2, and Terminal-Bench 2.1 increased from 82.4 to 90.6, while average output length grew by about 2.5 times. The authors found most gains between the low and middle settings, with diminishing returns near 100 [2]. These are internal results and do not establish the best setting for every application.

## DeepSeek-reported evaluation

DeepSeek published the following selected results for the instruction-tuned model at maximum reasoning effort. Its two primary evaluation descriptions disagree on one sampling setting. The current model card says every listed instruction-model result used temperature 1.0 and top-p 0.95. Section 5.3.1 of the technical report instead says GPQA Diamond, Humanity's Last Exam, Codeforces, and MathArena Apex used temperature 1.0 and top-p 1.0, while code-agent and visual-agent evaluations used temperature 1.0 and top-p 0.95 [1][2]. Neither source explains the difference, so the mixed table below should be read with that unresolved discrepancy. Code-agent tests generally used the Minimal mode of [DeepSeek Harness](https://aiwiki.ai/wiki/deepseek_harness) and a one-million-token context, while exceptions used benchmark-specific scaffolds. Visual-agent tests used a Claude Code scaffold and a 512,000-token context [1][2][11].

| Area | Benchmark and metric | DeepSeek-reported V4.1-Flash result |
|---|---|---:|
| Science reasoning | GPQA Diamond, pass@1 | 90.9 |
| Broad expert questions | Humanity's Last Exam, pass@1 | 36.8 overall; 39.1 on the text-only subset |
| Competitive programming | Codeforces rating | 3,471 |
| Terminal agent | Terminal-Bench 2.1, pass@1 | 90.6 |
| Terminal agent | Terminal-Bench 3.0, pass@1 | 30.0 |
| Terminal agent | Terminal-Bench 4.0, pass@1 | 31.2 |
| Software engineering | DeepSWE v1.1, resolved | 74.2 |
| Cybersecurity | CyberGym, pass@1 | 88.1 |
| General tool use | AutomationBench v1.0.6 public set, pass@1 | 54.8 |
| Visual chart agent | Chartography with tools, pass@1 | 78.9 |
| Visual reasoning | BabyVision with tools, pass@1 | 89.6 |
| Visual reasoning | ZeroBench-main with tools, pass@5 | 49.0 |

All values in the table come from DeepSeek's release evaluation [1][2]. The [AutomationBench](https://aiwiki.ai/wiki/automationbench) row uses the public evaluation set of AutomationBench version 1.0.6 under that benchmark's official scaffold, which matters when it is set beside a third-party run [2]. [Terminal-Bench](https://aiwiki.ai/wiki/terminal_bench) evaluates agents inside command-line environments, while DeepSWE uses repository-level software tasks. Their scores depend on the model, prompt, tools, sandbox, token budget, and agent loop rather than on the checkpoint alone [16][17]. CyberGym evaluates vulnerability reproduction against real software repositories [19]. Humanity's Last Exam is a multimodal collection of closed-ended expert questions spanning many academic subjects [20]. Third-party results for V4.1-Flash come from a different evaluation suite run under different harnesses, and are given under Independent evaluation below.

### Scaffold sensitivity

DeepSeek also ran the same checkpoint across several agent scaffolds. On DeepSWE v1.1, it reported results from 65.5 to 74.2 across the listed configurations. On Terminal-Bench 2.1, the range was 84.1 to 90.6 [1][2].

| Scaffold | DeepSWE v1.1 resolved | Terminal-Bench 2.1 pass@1 |
|---|---:|---:|
| Claude Code | 69.8 | 88.0 |
| Codex | 65.6 | 84.1 |
| OpenCode | 65.5 | 85.0 |
| Pi | 66.2 | 86.1 |
| mini-SWE | 74.2 | 90.3 |
| DeepSeek Harness Minimal | 72.6 | 90.6 |
| DeepSeek Harness Standard | 70.5 | 85.8 |
| DeepSeek Harness PTC | 67.6 | 85.8 |

DeepSeek used eight samples per DeepSWE task and three per Terminal-Bench task, Linux containers, temperature 1.0, top-p 0.95, a one-million-token context cap, and no more than 500 model-generation rounds. Terminal-Bench was run without network access [1][2]. The spread is evidence that an agent benchmark is a model-and-scaffold result; it should not be read as a single immutable property of the model.

### NL2Repo discrepancy

Two current DeepSeek-controlled sources disagree on one result. The live Hugging Face model card lists V4.1-Flash at **64.0** on NL2Repo-Bench, while Table 3 of the technical report and the API update log list **65.4** [1][2][5]. No revision note explains the difference. This article therefore does not select either number as authoritative. NL2Repo-Bench is a repository-level code-generation evaluation, and the inconsistency should be resolved before using this row in a comparison [18].

## Independent evaluation

[Artificial Analysis](https://aiwiki.ai/wiki/artificial_analysis) published an evaluation of V4.1-Flash at 20:36 UTC on September 10, 2026, about fourteen hours after the release [22]. It ran the model through DeepSeek's first-party API at the maximum reasoning-effort setting, and lists it as "DeepSeek V4.1 Flash (Reasoning, Max Effort)" [24].

Every figure below belongs to version 4.3 of the Artificial Analysis Intelligence Index, the version in force on the stated dates. Version 4.3 combines ten evaluations at fixed weights and differs from version 4.2 in three respects: AutomationBench-AA replaced the tau-3 Banking agent evaluation, Terminal-Bench moved from v2.1 to v4.0, and image handling changed in the GDP.pdf evaluation. Versions 4.1, 4.1.1, 4.2 and 4.3 all took effect between June and September 2026, so a score carrying one version number cannot be set against a score carrying another [23].

| Artificial Analysis, Intelligence Index v4.3, read September 11, 2026 | V4.1-Flash | V4-Flash 0731 | V4-Pro 0813 |
|---|---:|---:|---:|
| Intelligence Index | 39.5 | 34.5 | 36.3 |
| Terminal-Bench v4.0 | 26.8% | 12.1% | 14.1% |
| GDPval-AA v2, Elo | 1,632 | 1,468 | 1,493 |
| AA-LCR v1.1 | 84% | 80% | 80% |
| AutomationBench-AA | 68.9% | 54.0% | 56.7% |
| Output tokens per Index task | 88,574 | 62,054 | 55,154 |
| Cost per Index task | $0.27 | $0.22 | $0.67 |

All three models are run at their maximum reasoning-effort tier [24][25]. V4.1-Flash is DeepSeek's highest-scoring entry on that board, ahead of V4-Pro 0813 even though Artificial Analysis lists the Pro model at 1.6 trillion parameters against V4.1-Flash's 552 billion. Artificial Analysis summarized the result as V4.1-Flash overtaking V4-Pro 0813 as the company's flagship, and described the model as roughly 20% cheaper per token than V4-Flash 0731 and about four times cheaper than V4-Pro 0813 [22].

On AutomationBench-AA, read on September 11, 2026, Artificial Analysis placed V4.1-Flash first among the models it had run, at 68.9%, ahead of [GPT-6 Astra](https://aiwiki.ai/wiki/gpt_6_astra) at maximum effort (68.5%) and [Grok 4.6](https://aiwiki.ai/wiki/grok_4_6) at extra-high effort (67.0%) [26]. By September 30, 2026 it stood third, behind [Claude Sonnet 5.5](https://aiwiki.ai/wiki/claude_sonnet_5_5) (71.3%) and [Claude Opus 5.5](https://aiwiki.ai/wiki/claude_opus_5_5) (69.5%), both run at maximum effort with default fallback [31]. AutomationBench-AA is Artificial Analysis' own run of [Zapier](https://aiwiki.ai/wiki/zapier)'s AutomationBench. It uses a private 657-task held-out split of dataset version 1.0.6 across six business domains, in simulated environments for products including Gmail, Google Sheets, Slack, Salesforce, Zendesk, Jira and HubSpot, and scores a task zero if the agent breaks any guardrail [23][26].

One measurement runs the other way. Artificial Analysis recorded 88,574 output tokens per Intelligence Index task for V4.1-Flash, more than it recorded for [GLM-5.3](https://aiwiki.ai/wiki/glm_5_3) (71,128), [Claude Opus 5](https://aiwiki.ai/wiki/claude_opus_5) (72,511) or [Claude Fable 5.1](https://aiwiki.ai/wiki/claude_fable_5_1) (78,111), and wrote that this verbosity leaves the model "just short of the Intelligence vs. Cost Pareto frontier" [22][24][25]. Low prices absorb most of the cost of that verbosity: the same task set cost $0.27 per task against $2.01 for GLM-5.3, $2.00 for [Kimi K3](https://aiwiki.ai/wiki/kimi_k3) and $0.67 for V4-Pro 0813. Artificial Analysis also measured a median output speed of 199 tokens per second and a median time to first chunk of 1.13 seconds on DeepSeek's first-party API [24][25].

### Comparing the two evaluations

Artificial Analysis' numbers are not interchangeable with DeepSeek's table above, even where the same benchmark name appears in both. On Terminal-Bench 4.0, DeepSeek reports 31.2 using the Minimal mode of DeepSeek Harness with a one-million-token context, while Artificial Analysis reports 26.8% over 66 tasks with three repeats under the mini-SWE-agent harness. On AutomationBench, DeepSeek reports 54.8 on the public evaluation set of version 1.0.6 under the benchmark's official scaffold, while Artificial Analysis reports 68.9% on its own private held-out split of the same dataset version [1][2][23]. The split, the harness and the resulting number all differ.

### Errors and revisions in the launch post

Two statements in Artificial Analysis' launch post are contradicted by data the firm itself publishes.

The headline of the post reads "At just 552B parameters", while its "Additional model details" section reads "This is a 763B total parameters, 8B active parameters input and 16B active parameters output" [22]. Both numbers describe the same checkpoint, for the reason set out under Parameter accounting above, and Artificial Analysis' own model page records the size as 552 billion [24].

The post also states that V4.1-Flash "scores 27% in Terminal-Bench v4.0, more than double DeepSeek V4.0 Flash's 27%", which cannot hold as written [22]. Read on September 11, 2026, Artificial Analysis' model pages give 26.8% for V4.1-Flash and 12.1% for V4-Flash 0731 on Terminal-Bench v4.0, a ratio of about 2.2. The first figure in that sentence matches the live value; the second matches neither model's live value [24][25].

Two peer comparisons in the post have also drifted from the live pages. The post put Kimi K3 at 1,584 on GDPval-AA v2, while the Kimi K3 maximum-effort page showed 1,572 on September 11. GDPval-AA v2 reports Elo ratings, which move as the comparison pool changes. The post described V4.1-Flash's 84% on AA-LCR v1.1 as level with GPT-5.6 Sol and [Gemini 3.8 Flash](https://aiwiki.ai/wiki/gemini_3_8_flash); on September 11 that 84% matched GPT-5.6 Sol at maximum effort and Gemini 3.8 Flash at medium effort, while Gemini 3.8 Flash at high effort stood at 81% [22][25]. Artificial Analysis revises its boards without notice, which is why every figure here carries the date it was read.

## Availability, API limits, and pricing

At launch, DeepSeek offered V4.1-Flash through the `deepseek-flash` API model name and through the public Hugging Face repository. The hosted service documented a one-million-token context window, up to 384,000 output tokens, native vision input, tool calls, JSON output, context caching, fill-in-the-middle in non-thinking mode, an OpenAI-compatible Responses interface, and an Anthropic-compatible interface. It listed a maximum concurrency of 2,500 requests [5][6]. These are service capabilities and limits, not guarantees for third-party hosts.

The following API prices were listed by DeepSeek on September 10, 2026, were unchanged when the pricing page was read again on September 11, and are stated per one million tokens [5][6]:

| Token class | Peak price | Off-peak price |
|---|---:|---:|
| Cached input | $0.006 | $0.003 |
| Uncached input | $0.30 | $0.15 |
| Output | $1.20 | $0.60 |

DeepSeek defined peak periods as 01:00 to 04:00 UTC and 06:00 to 10:00 UTC on Monday through Friday, with all other times off-peak [6]. Pricing and limits are time-sensitive. The table records the launch state rather than promising that those terms remain current.

## Open weights and self-hosting

The Hugging Face repository is public and ungated, with weights and accompanying source files under MIT terms [1][12]. The configuration identifies the architecture as `DeepseekV41ForCausalLM`, stores routed-expert weights in FP4 within an FP8 quantization configuration, and defines a 1,048,576-position limit [3]. Because the checkpoint includes hundreds of billions of backbone and lookup-memory parameters, practical self-hosting requires substantial accelerator and host-memory resources even though only a small fraction of experts activate for each token.

The repository's minimal inference code covers the vision encoder and projector, CSA2, the hierarchical indexer, Engram lookup, MoE blocks, mHC, and the DSpark forward path. Its generation loop uses ordinary autoregressive sampling, and its own documentation says it is intended for readability rather than production serving [10]. The prompt-encoding reference is separate from the runtime. It documents the V4.1 conversation format, tool-call markup, image placeholders, mid-conversation system messages, and numeric reasoning-effort prefix [9].

The model card recommends temperature 1.0, top-p 0.95 or 1.0, a one-million-token context window, and a generation budget of at least 256,000 tokens for its documented evaluation style [1]. Those settings reflect DeepSeek's intended high-budget agent use and can be costly. They are not minimum requirements for ordinary short responses.

### Third-party serving and derivatives

Open-source inference stacks added support within a day. [vLLM](https://aiwiki.ai/wiki/vllm) merged three V4.1-Flash pull requests on September 10 and 11, 2026, covering the model definition and the Rust and Python frontends, and opened a tracking issue for kernel integration. Bug reports filed on September 11 describe failures on GB10-class hardware, a device-side assertion when DSpark speculative decoding runs on an MXFP4 mixture-of-experts backend, and an illegal memory access under high concurrency, so support was still settling [27]. [SGLang](https://aiwiki.ai/wiki/sglang) merged a V4.1-Flash cookbook on September 10 [28]. DeepSeek separately published deepseek-recipe, a set of MIT-licensed Rust libraries with Python bindings that convert Messages, Chat Completions and Responses API requests into DeepSeek V4 and V4.1 prompts and parse model output back; the repository was created on September 10, 2026 and leaves inference, tool execution and transport to the caller [1][29].

A Hugging Face API search for repositories matching DeepSeek-V4.1 returned 49 entries on September 11, 2026. DeepSeek's own repository accounted for 75,774 downloads and 1,762 likes in Hugging Face's counters at that time. The rest were community conversions: GGUF builds, [MLX](https://aiwiki.ai/wiki/mlx) quantizations from several publishers including mlx-community, NVFP4 and EXL3 packings, a separately packaged DSpark drafter, and repositories their authors describe as expert-pruned [12][30].

## Limitations

The technical report states that the new architecture's robustness boundaries are not fully characterized. It specifically identifies two possible failure sources: sparse-selection errors in CSA2 and capability degradation from the approximate local-state reconstruction used by SWA Bounded Replay. DeepSeek says these were not systematic in its internal tests, but finite testing cannot cover every extreme prompt, context pattern, or cache-resumption boundary [2].

DeepSeek also cautions that small benchmark gaps do not establish parity with leading proprietary systems on difficult reasoning problems and edge cases [2]. The release table uses maximum reasoning effort and large context budgets, and several values depend on particular agent scaffolds. It should therefore be interpreted as a disclosed vendor evaluation, not as independent evidence that the model is best for every workload.

The training report gives aggregate token counts, source types, and mixture ratios but not a source-by-source inventory of the 45-trillion-token corpus. It does not provide enough information to independently audit the corpus for licensing, contamination, consent, or geographic and linguistic balance [2]. The MIT license covers the released artifacts; it does not by itself answer those training-data questions.

## See also

- [Artificial Analysis](https://aiwiki.ai/wiki/artificial_analysis)
- [DeepSeek V4](https://aiwiki.ai/wiki/deepseek_v4)
- [Multimodal model](https://aiwiki.ai/wiki/multimodal_model)
- [Open weights](https://aiwiki.ai/wiki/open_weights)
- [Vision-language model](https://aiwiki.ai/wiki/vision_language_model)

## References

1. DeepSeek-AI. "DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression." Hugging Face model card, September 10, 2026. [https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash](https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash)
2. DeepSeek-AI. "DeepSeek-V4.1-Flash Technical Report." September 10, 2026. [https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/main/DeepSeek_V41_Tech_Report.pdf](https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/main/DeepSeek_V41_Tech_Report.pdf)
3. DeepSeek-AI. "DeepSeek-V4.1-Flash config.json." Hugging Face, September 10, 2026. [https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/main/config.json](https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/main/config.json)
4. DeepSeek. "Introducing DeepSeek-V4.1-Flash." Official announcement thread on X, September 10, 2026. [https://x.com/deepseek_ai/status/2097930608790167907](https://x.com/deepseek_ai/status/2097930608790167907)
5. DeepSeek. "DeepSeek API Updates." Accessed September 11, 2026. [https://api-docs.deepseek.com/updates/](https://api-docs.deepseek.com/updates/)
6. DeepSeek. "Models and Pricing." DeepSeek API documentation. Accessed September 11, 2026. [https://api-docs.deepseek.com/quick_start/pricing](https://api-docs.deepseek.com/quick_start/pricing)
7. DeepSeek. "Vision." DeepSeek API documentation. Accessed September 10, 2026. [https://api-docs.deepseek.com/guides/vision](https://api-docs.deepseek.com/guides/vision)
8. DeepSeek. "Thinking Mode." DeepSeek API documentation. Accessed September 10, 2026. [https://api-docs.deepseek.com/guides/thinking_mode](https://api-docs.deepseek.com/guides/thinking_mode)
9. DeepSeek-AI. "DeepSeek-V4.1 Text and Vision Encoding." Hugging Face repository, September 10, 2026. [https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/main/encoding/README.md](https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/main/encoding/README.md)
10. DeepSeek-AI. "Minimal Inference." Hugging Face repository, September 10, 2026. [https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/main/inference/README.md](https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/main/inference/README.md)
11. DeepSeek-AI. "Running DeepSWE with dsh-minimal and mini-swe-agent." Hugging Face repository, September 10, 2026. [https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/main/evaluation/README.md](https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/main/evaluation/README.md)
12. Hugging Face. "deepseek-ai/DeepSeek-V4.1-Flash model metadata." Accessed September 11, 2026. [https://huggingface.co/api/models/deepseek-ai/DeepSeek-V4.1-Flash](https://huggingface.co/api/models/deepseek-ai/DeepSeek-V4.1-Flash)
13. Sun, Yutao, et al. "You Only Cache Once: Decoder-Decoder Architectures for Language Models." Advances in Neural Information Processing Systems 37, 2024. [https://arxiv.org/abs/2405.05254](https://arxiv.org/abs/2405.05254)
14. Cheng, Xin, et al. "Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models." arXiv, 2026. [https://arxiv.org/abs/2601.07372](https://arxiv.org/abs/2601.07372)
15. Cheng, Xin, et al. "DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation." arXiv, 2026. [https://arxiv.org/abs/2607.05147](https://arxiv.org/abs/2607.05147)
16. Merrill, Mike A., et al. "Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces." arXiv, 2026. [https://arxiv.org/abs/2601.11868](https://arxiv.org/abs/2601.11868)
17. DataCurve. "DeepSWE v1.1." Accessed September 10, 2026. [https://deepswe.datacurve.ai/](https://deepswe.datacurve.ai/)
18. Ding, Jingzhe, et al. "NL2Repo-Bench: Towards Long-Horizon Repository Generation Evaluation of Coding Agents." arXiv, 2025. [https://arxiv.org/abs/2512.12730](https://arxiv.org/abs/2512.12730)
19. Wang, Zhun, et al. "CyberGym: Evaluating AI Agents' Real-World Cybersecurity Capabilities at Scale." International Conference on Learning Representations, 2026. [https://arxiv.org/abs/2506.02548](https://arxiv.org/abs/2506.02548)
20. Phan, Long, et al. "Humanity's Last Exam." arXiv, 2025. [https://arxiv.org/abs/2501.14249](https://arxiv.org/abs/2501.14249)
21. Internet Archive. "DeepSeek API Change Log," snapshot of September 10, 2026. Wayback Machine. [https://web.archive.org/web/20260910074600/https://api-docs.deepseek.com/updates/](https://web.archive.org/web/20260910074600/https://api-docs.deepseek.com/updates/)
22. Artificial Analysis. "DeepSeek V4.1 Flash overtakes DeepSeek V4 Pro 0813 as DeepSeek's new flagship model." Post on X, September 10, 2026. [https://x.com/ArtificialAnlys/status/2098148674203488422](https://x.com/ArtificialAnlys/status/2098148674203488422)
23. Artificial Analysis. "Intelligence Benchmarking Methodology: Artificial Analysis Intelligence Index v4.3." Accessed September 11, 2026. [https://artificialanalysis.ai/methodology/intelligence-benchmarking](https://artificialanalysis.ai/methodology/intelligence-benchmarking)
24. Artificial Analysis. "DeepSeek V4.1 Flash (Reasoning, Max Effort): Intelligence, Performance & Price Analysis." Accessed September 11, 2026. [https://artificialanalysis.ai/models/deepseek-v4-1-flash](https://artificialanalysis.ai/models/deepseek-v4-1-flash)
25. Artificial Analysis. "Models: comparison pages for DeepSeek V4 Flash 0731, DeepSeek V4 Pro 0813, Kimi K3, GLM-5.3, GPT-5.6 Sol, GPT-6 Astra, Grok 4.6, Gemini 3.8 Flash, Claude Opus 5 and Claude Fable 5.1." Accessed September 11, 2026. [https://artificialanalysis.ai/models](https://artificialanalysis.ai/models)
26. Artificial Analysis. "AutomationBench-AA." Accessed September 11, 2026. [https://artificialanalysis.ai/evaluations/automationbench-aa](https://artificialanalysis.ai/evaluations/automationbench-aa)
27. vLLM. "[Model] Support DeepSeek-V4.1-Flash," pull request 56214, merged September 11, 2026. GitHub. [https://github.com/vllm-project/vllm/pull/56214](https://github.com/vllm-project/vllm/pull/56214)
28. SGLang. "Add DeepSeek-V4.1 Flash cookbook," pull request 38802, merged September 10, 2026. GitHub. [https://github.com/sgl-project/sglang/pull/38802](https://github.com/sgl-project/sglang/pull/38802)
29. DeepSeek-AI. "deepseek-recipe." GitHub, created September 10, 2026. [https://github.com/deepseek-ai/deepseek-recipe](https://github.com/deepseek-ai/deepseek-recipe)
30. Hugging Face. "Model search: DeepSeek-V4.1." Hugging Face API, accessed September 11, 2026. [https://huggingface.co/api/models?search=DeepSeek-V4.1](https://huggingface.co/api/models?search=DeepSeek-V4.1)
31. Artificial Analysis. "AutomationBench-AA." Accessed September 30, 2026. [https://artificialanalysis.ai/evaluations/automationbench-aa](https://artificialanalysis.ai/evaluations/automationbench-aa)

