DeepSelect
DeepSelect is an open-source library of top-k selection kernels published by the Chinese AI company DeepSeek. It implements one narrow operation, "return the k largest values of each row of a matrix and their indices", for the two places where DeepSeek's own inference stack needs it: the token-selection step of DeepSeek Sparse Attention (DSA) and the sampler that picks the next token from the vocabulary distribution. The repository's README describes DeepSelect as a high-performance implementation of the top-k kernel used in DSA, which it says is used in the DeepSeek V3.2, DeepSeek V4 and DeepSeek V4.1 models, and in the sampler, and states that it supports NVIDIA CUDA and Huawei Ascend platforms with a 2x to 20x speedup over PyTorch's built-in torch.topk [1].
The GitHub repository deepseek-ai/DeepSelect was created on 9 September 2026 and the first release, version 1.0.0, was committed on 10 September 2026, the same day DeepSeek published its DeepSeek V4.1-Flash weights on Hugging Face [2][3][4]. Kernels for Huawei Ascend NPUs were merged on 30 September 2026, the day DeepSeek released a set of Ascend ports of its infrastructure libraries [3][5]. The code is under the MIT License [6]. As of 1 October 2026 the repository had 439 stars and 39 forks, and GitHub reported its languages as CUDA, Python, AGS Script and C++, the third of those being how GitHub labels the Ascend kernel file [2].
Why a separate top-k library
In DSA, a small scoring module that DeepSeek calls the lightning indexer produces a relevance score for every preceding token, and a selector keeps only the highest-scoring entries for the main attention computation. The scoring is cheap but the selection is not: it is a top-k over a row whose length is the whole key-value context, run at every layer and every decoding step. The sampler faces the same shape of problem at the output end, where top-k has to be taken over a vocabulary of roughly 128,000 logits. DeepSelect's design document states the motivation plainly: top-k is used in DSA to select the most relevant positions from a large number of context tokens and in sampling to select candidate tokens from a large vocabulary, and DeepSeek designed a new top-k algorithm to improve end-to-end inference performance [7].
The README is explicit that this is not a general-purpose top-k library. It says the fastest algorithm depends heavily on the input dtype, batch size, vocabulary size and k, and that the repository targets only two workload families [1].
| Scenario | Input dtype | Batch size | Row length | k |
|---|---|---|---|---|
| Lightning indexer | torch.bfloat16 | 1 upward, small and large both optimized | 1 upward, small and large both optimized | must be 4,096 or less; mainly optimized for 512, which the README gives as DeepSeek V4's k |
| Sampling | torch.float32 | 1 upward | around 128K | must be 4,096 or less |
Values of k above 4,096 are not supported at all [1]. The float32 path exists only in the CUDA implementation; the Ascend implementation accepts bfloat16 only [1].
Algorithm
DeepSeek published a short analysis of the algorithm and its implementation, in English and Chinese, on 10 September 2026 [3][7]. The method is a streaming threshold filter. It keeps a running threshold T, initialized to negative infinity, that always holds the k-th largest value seen so far, and repeats three steps over the input:
- Scan the input one block of size B at a time, taking the blocks in a random order.
- Filter each block against T, appending to a candidate buffer only the elements that could still enter the top k.
- Compact the candidate buffer, once it has grown past k plus a second parameter B2, by running a radix-select top-k inside fast on-chip memory and resetting T to the smallest of the k survivors.
Because T never decreases, later blocks pass fewer and fewer elements, so the candidate buffer grows more and more slowly. The document notes that the buffer never exceeds k + B + B2 entries, that every input element is read exactly once in contiguous blocks, and that the working set can therefore live in a small, fast memory space such as GPU shared memory. A reasonable configuration for k = 512 is given as B = B2 = 1024 [7].
The random block order is there to defend against adversarial inputs, and it is what makes an expected bound provable. Under a uniform random permutation of blocks, the document derives that the expected total number of elements passed through all of the inner top-k calls is bounded by (1 + k/B2) times (k + B + B2) times the harmonic number of the block count, giving an expected cost of order (1 + k/B2)(k + B + B2) log(N/B) for a row of length N. When B and B2 are chosen proportional to k, the extra work beyond the single streaming read is of order k log(N/k), far below the input size [7].
The exact result is preserved: this is an exact top-k, not an approximation. The write-up frames the gain as removing memory traffic rather than arithmetic, and the README makes the same point when it explains its benchmark metric, noting that top-k does no floating-point math so a FLOP rate would not be meaningful and that it reports effective memory bandwidth instead [1][7].
Implementation and hardware support
The CUDA side is written in CUDA C++ and leans on low-level PTX intrinsics and bit manipulation to cut the constant factor in the filter step, so that a thread can locate the elements above the threshold from a bitmask instead of rescanning them, and to accelerate prefix sums, byte extraction and histogram construction. DeepSeek also describes substituting floating-point addition for some integer addition, which is valid for non-negative integers no greater than 2^22, because comparison and bitwise instructions contend with integer addition for hardware resources in this kernel. For small batch sizes, or when the last wave leaves few streaming multiprocessors busy, a variant spreads one row across several SMs using a thread-block cluster: each block computes a local top-k and sends it to the cluster's first block, which reduces the partial results [7]. The source tree matches that description, with separate bfloat16, float32 and cluster kernel families, each shipped as a set of pre-generated instantiations [2].
The build script compiles for compute_100a and compute_103a by default, requires NVCC 12.9 or newer for the sm100 target, and carries the sm80 and sm90 targets commented out to keep compilation times down [8]. That default is the reason several of the open contributions on the repository and in downstream projects are requests for Hopper-class support [9][10].
The Ascend backend was added in pull request #23, merged on 30 September 2026 and credited to Yi Qian and Shengyu Liu of DeepSeek [5]. It is a separate kernel written in Ascend's own kernel language rather than a translation layer over the CUDA code [2]. The Python package picks a backend at import time by probing for an Ascend device node, loading either a CUDA or an NPU extension module and reporting which build target is missing if the extension is absent [11]. Building for Ascend requires the torch_npu package and a CANN installation providing Huawei's bisheng compiler; the script's error message points at a CANN 9.2.0 install path as the example, and the NPU architecture target is settable through an environment variable [8].
The Ascend implementation is more constrained than the CUDA one. Alongside the bfloat16-only input restriction, it supports only 32-bit output indices and does not offer the value-sorted output mode [1][11]. The two backends also differ in how they report a NaN: the check itself is always on, and by default the kernel traps and aborts, but the documented fallback behavior when aborting is disabled differs between CUDA and Ascend, and on CUDA rows no longer than k skip the check entirely [1][11].
The public interface is a single deep_select.topk function. Callers must align the row stride of the input to a backend-reported byte requirement and pad unaligned inputs themselves, and the outputs come back with their own stride alignment, so they may be non-contiguous. A per-row end argument gives variable-length rows an exclusive upper bound, with short rows padded by caller-specified fill values for both indices and values, which is what lets a serving engine hand the kernel a batch of different sequence lengths in one call. Two switches exist purely for speed: sorted_index, which the README advises leaving off unless the output genuinely has to be ordered by index, and return_value=False, which skips writing the values and which the README puts at roughly 10 percent faster [1][11].
Published performance
The README's headline figure is the 2x to 20x speedup over torch.topk on the same input, measured by the benchmark script in the repository, with effective memory bandwidth as the metric [1]. Three plots are published: the bfloat16 lightning-indexer case at k = 512 on CUDA and on Ascend, each for batch sizes 6, 512 and 4,096 across row lengths from 16K to 1M, and the float32 sampling case on CUDA at a vocabulary of 129,280 and k = 512. In the published plots DeepSelect's effective bandwidth climbs steadily with row length at the larger batch sizes, reaching several terabytes per second, while the torch.topk baseline stays nearly flat and far lower across the whole sweep [1]. DeepSeek does not name the specific GPU or NPU model used for those plots, so the absolute bandwidth figures cannot be tied to a device here.
Independent-of-DeepSeek numbers exist in the vLLM integration, whose author benchmarked the DSA indexer's decode top-k stage on an NVIDIA GB200 with CUPTI timing, CUDA-graph replay and a cold L2, at shapes taken from the DeepSeek V4.1-Flash lightning indexer with k = 512 and the indexer "vocabulary" equal to the key-value context length. Those are contributor measurements reported in a pull request, not vendor results [12].
| Shape (fp32 logits) | DeepSelect | Previous default for that shape | torch.topk |
|---|---|---|---|
| batch 256, context 1M | 191 us | 616 us (persistent) | 2,878 us |
| batch 64, context 1M | 88 us | 132 us (cooperative) | 905 us |
| batch 8, context 1M | 85 us | 40 us (cooperative) | 194 us |
The last row is the reason the integration kept a heuristic rather than switching unconditionally: at small batch sizes the existing cooperative kernel was still faster, so vLLM's automatic choice keeps it there [12].
Use in inference engines
DeepSelect was taken up by both major open-source serving engines within days of its release, in each case on the DSA indexer path for DeepSeek V4.1.
| Project | Change | Merged |
|---|---|---|
| SGLang | Exact bfloat16 consumer top-k adapted from DeepSelect | 13 September 2026 [13] |
| vLLM | DeepSelect vendored through CMake FetchContent, with every indexer top-k backend made selectable | 13 September 2026 [12] |
| vLLM | Bound DeepSelect sentinel columns in the sparse top-k remap | 24 September 2026 [14] |
| SGLang | DeepSelect moved onto SGLang's just-in-time compilation path so only the needed signature is built | 26 September 2026 [15] |
| SGLang | Page-table transform added to the top-k call and the layout contract tightened | 28 September 2026 [16] |
In vLLM the integration introduced a dedicated indexer top-k layer that hosts every backend entry point behind one module, plus a --sparse-indexer-topk-backend flag whose values include deep_select, several of vLLM's own kernels, a FlashInfer backend and a plain torch reference, defaulting to an automatic shape-based choice. Per-row sequence lengths map directly onto DeepSelect's end argument, and results are written into a preallocated buffer so the path stays CUDA-graph safe [12][17]. As of 1 October 2026 the DeepSelect wrapper in vLLM's main branch registers the kernels as a torch.ops.deep_select namespace and gates their use behind a support check on stride alignment and k [17]. The same backend selection code is reached from vLLM's GLM-5-next sparse indexer as well as the DeepSeek MLA path, so the kernels are not used only for DeepSeek models [18].
SGLang took a different route, first adapting the algorithm into its own kernel and later vendoring DeepSeek's headers under its just-in-time build tree while keeping the CUDA algorithm unchanged [13][15].
Part of the 30 September 2026 Ascend release
DeepSelect's Ascend kernels landed as one piece of a larger release on 30 September 2026, in which DeepSeek published Ascend ports or Ascend backends for the infrastructure libraries it had previously maintained for NVIDIA hardware. Companion repositories carry matching news entries dated the same day: TileKernels added an Ascend backend selected automatically at runtime behind the same Python APIs [19]; DeepGEMM Ascend shipped as an API-compatible port with support for Ascend 950 devices [20]; and FlashMLA released sparse attention prefill and decoding kernels for the Ascend 950 NPU, reporting up to 410 TFLOPS in prefill and 360 TFLOPS in decoding, which it puts at 95 percent and 83 percent of the hardware peak [21]. The same FlashMLA release dropped support for NVIDIA Hopper and for DeepSeek V3, V3.2 and V4.0, and changed the FP8 and FP4 key-value cache format [21]. DeepSelect's own Ascend notes do not name a chip generation, and its build script defaults the NPU architecture target to dav-3510, the architecture TileLang documents for the Ascend 950, and leaves it overridable through the ASCEND_NPU_ARCH environment variable [1][8].
Covering the release on 30 September 2026, the Indian trade publication Dataquest counted six components, TileLang, DeepGEMM, DeepEP, TileKernels, FlashMLA and DeepSelect, described DeepSelect's role as data filtering, and reported that DeepSeek presented the Ascend implementations as corresponding to components already available for NVIDIA hardware. It added that the two companies had jointly advanced a supernode configuration built from 128 Ascend 950 chips, and that Reuters independently reported the partnership the same day [22]. The Chinese technology outlet Zhidongxi, posting in English, framed the release as DeepSeek open-sourcing its AI infrastructure stack for Ascend NPUs with each component corresponding to a piece of its NVIDIA stack, and singled out TileLang as the key element [23]. That framing is the outlet's, not DeepSeek's.
Priority claim by the LiteTopK authors
On 11 September 2026, the day after DeepSelect's first release, an issue was opened on the repository on behalf of the authors of LiteTopK, asking DeepSeek to acknowledge and cite their earlier work. It is signed "Ziqi" and was filed from the GitHub account that hosts LiteTopK's reference implementation. LiteTopK is a fused indexer-and-top-k kernel for long-context sparse attention by Ziqi Yin, Jianyang Gao, Peiqi Yin, Jiangneng Li and Gao Cong, posted to arXiv on 13 July 2026; its abstract describes maintaining a tight approximate threshold online so that only promising candidates are written back, while preserving exact top-k correctness, and reports that LiteTopK combined with the paper's LiteDSA scheme gives a 1.35x prefill speedup for GLM 5.2 on eight B200 GPUs [24].
The issue argues that DeepSelect's streaming threshold-filtering component substantially overlaps LiteTopK's, notes that DeepSelect's design document introduces a "new TopK algorithm" without discussing LiteTopK, states that the authors had contacted DeepSeek about LiteTopK in August 2026, and lists prior integration activity around their implementation in vLLM, SGLang, ROCm's AITER and by a Zhipu AI engineer. It asks DeepSeek to clarify the relationship, credit LiteTopK in the README, rename the kernel to reflect a LiteTopK-derived origin if implementation reuse occurred, and discuss LiteTopK as prior work in the deep-dive documents [25]. As of 1 October 2026 the issue was still open, and the thread held a single comment, from another GitHub user, with no reply from the repository's committers [25]. The claim is the LiteTopK authors' and has not been adjudicated; DeepSeek's documentation as of 1 October 2026 does not mention LiteTopK [1][7]. Separately, vLLM's DeepSelect integration pull request refers to an earlier open pull request in that project as a custom fused LiteTopK kernel, treating the two as distinct pieces of work [12].
Release history
| Date | Event |
|---|---|
| 9 September 2026 | Repository deepseek-ai/DeepSelect created [2] |
| 10 September 2026 | Initial release, version 1.0.0; the do_check_nan argument dropped and NaN checking simplified and optimized the same day [3] |
| 10 September 2026 | Algorithm and implementation analysis published in English and Chinese [3][7] |
| 11 September 2026 | LiteTopK authors open a priority and citation issue [25] |
| 13 September 2026 | First downstream adoption merged, in SGLang and in vLLM [12][13] |
| 30 September 2026 | Top-k kernels for Huawei Ascend NPUs merged, alongside DeepSeek's wider Ascend release [3][5] |
The project publishes no GitHub releases or tags; the 1.0.0 version number appears only in the package's version file and the README's news list [1][2]. Authorship, by the citation block in the README, is Yi Qian, Shengyu Liu and Yichen Li [1]; the commit history to 1 October 2026 is the work of Yi Qian and Shengyu Liu [3].
References
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11 ^12 ^13 ^14DeepSelect README. deepseek-ai/DeepSelect, GitHub. github.com/...DeepSelect
- ^1 ^2 ^3 ^4 ^5 ^6deepseek-ai/DeepSelect repository metadata and file tree, GitHub REST API. api.github.com/...DeepSelect
- ^1 ^2 ^3 ^4 ^5 ^6 ^7Commit history, deepseek-ai/DeepSelect. github.com/...main
- ^deepseek-ai/DeepSeek-V4.1-Flash model metadata, Hugging Face API. huggingface.co/...DeepSeek-V4.1-Flash
- ^1 ^2 ^3"Add Kernels for Huawei Ascend NPU", pull request #23, deepseek-ai/DeepSelect. github.com/...23
- ^LICENSE (MIT License), deepseek-ai/DeepSelect. github.com/...LICENSE
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8"Algorithm Design Background", docs/DeepSelect-deep-dive.md, deepseek-ai/DeepSelect. github.com/...DeepSelect-deep-dive.md
- ^1 ^2 ^3setup.py, deepseek-ai/DeepSelect. github.com/...setup.py
- ^"[SM90] Add optimized SM90 support", pull request #14, deepseek-ai/DeepSelect. github.com/...14
- ^"[Kernel] Enable DeepSelect sparse indexer top-k on sm90", pull request #59204, vllm-project/vllm. github.com/...59204
- ^1 ^2 ^3 ^4deep_select/interface.py, deepseek-ai/DeepSelect. github.com/...interface.py
- ^1 ^2 ^3 ^4 ^5 ^6"[Perf][Kernel] Integrate DeepSelect TopK for the DSA sparse indexer", pull request #56464, vllm-project/vllm. github.com/...56464
- ^1 ^2 ^3"[dsv4.1] Exact bf16 consumer top-k adapted from DeepSelect", pull request #39305, sgl-project/sglang. github.com/...39305
- ^"[Bugfix][DSA] Bound DeepSelect sentinel columns in the sparse top-k remap", pull request #58215, vllm-project/vllm. github.com/...58215
- ^1 ^2"[DeepSeek V4.1] Add DeepSelect JIT kernel.", pull request #40556, sgl-project/sglang. github.com/...40556
- ^"[DeepSelect] Add page-table transform to top-k and tighten the layout contract", pull request #41364, sgl-project/sglang. github.com/...41364
- ^1 ^2vllm/model_executor/layers/indexer_topk.py, vllm-project/vllm. github.com/...indexer_topk.py
- ^vllm/models/glm5next/nvidia/sparse_indexer.py, vllm-project/vllm. github.com/...sparse_indexer.py
- ^Tile Kernels README. deepseek-ai/TileKernels, GitHub. github.com/...TileKernels
- ^DeepGEMM Ascend README. deepseek-ai/DeepGEMM-Ascend, GitHub. github.com/...DeepGEMM-Ascend
- ^1 ^2FlashMLA README. deepseek-ai/FlashMLA, GitHub. github.com/...FlashMLA
- ^"DeepSeek expands Huawei Ascend push with six open-source AI tools", Dataquest, 30 September 2026. dqindia.com/...h-six-open-source-ai-tools-12594158
- ^Zhidongxi (@Chinazhidx) post on X, 30 September 2026. x.com/...2105138013609369641
- ^Ziqi Yin, Jianyang Gao, Peiqi Yin, Jiangneng Li, Gao Cong. "LiteTopK: Exploiting the Curse of Dimensionality for a Fused Indexer-TopK Kernel in Long-Context Sparse Attention." arXiv:2607.11976, 13 July 2026. arxiv.org/...2607.11976
- ^1 ^2 ^3"Substantial design overlap with LiteTopK: request for public acknowledgment and citation", issue #13, deepseek-ai/DeepSelect. github.com/...13
Improve this article
Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.
1 revision · v2 · 3,134 words · full history
Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify
Research and drafting on this wiki are AI-assisted, under named human editorial standards. How AI is used here
Reviewer note: Independently fact-checked 1 Oct 2026 against the repo, design doc, vLLM and SGLang integrations; 4 defects corrected incl. the Ascend build default
Cite this page: AI Wiki. "DeepSelect." aiwiki.ai, updated 1 Oct 2026, fact-checked 1 Oct 2026. CC BY 4.0. https://aiwiki.ai/wiki/deepselect