Citation and evidence

TileKernels

13 min full readUpdated 22 references

This article's verification

Report a problem with this article

More

Use this article

Raw MarkdownExplore connections

Improve this page

Suggest editRevision historyDiscussion

Browse categories

AI InfrastructureChinese AIDeveloper ToolsOpen Source AI

Cite this article

TileKernels (also written Tile Kernels; Python package tile-kernels) is an open-source library of GPU and NPU kernels for large language model training and inference, published by DeepSeek on GitHub under the MIT License. Every kernel is written in TileLang, a Python-embedded domain-specific language for high-performance kernels. The library covers the small, mostly memory-bound operations that sit between the large matrix multiplications of a mixture-of-experts model: expert routing, FP8 and FP4 quantization casts, RMSNorm, SwiGLU, rotary position embedding, and kernels for two architecture components used in DeepSeek's 2026 models, Manifold-Constrained Hyper-Connections (mHC) and Engram conditional memory.[1][4]

The repository was created on 22 April 2026 and its first public commit and PyPI release (version 1.0.0) followed on 23 April 2026, the day before DeepSeek released the DeepSeek V4 preview on 24 April.[2][3][6][22] Version 2.0.0, released on 30 September 2026, added a second backend for Huawei Ascend NPUs that is selected automatically at runtime, so the same Python API runs on NVIDIA GPUs and on Ascend 950 hardware.[1][5] That release was part of a coordinated DeepSeek announcement of Ascend versions of its open infrastructure stack, alongside TileLang, DeepGEMM, DeepEP, FlashMLA and DeepSelect.[14][15] In Chinese-language reports of DeepSeek's announcement, TileKernels is described as the component that supplies the routine vector-compute and memory-access operators needed for data processing.[14]

Overview

The README describes TileKernels as "a library of dozens of highly optimized kernels implemented in TileLang" that provides kernels "for several common operations in LLM training and inference, including mixture-of-experts routing, Engram, quantization, and manifold hyper-connections." DeepSeek states that "most kernels achieve performance close to the hardware's compute or memory bandwidth limits" and that "all of these kernels have already been used in our internal training and inference workloads."[1] The README does not publish benchmark tables; performance can be measured locally with the repository's pytest benchmark mode.[1]

The library is a Python package. Its public operations are Python functions that validate tensor shapes and data types, then build and launch a TileLang kernel.[8][18] A tile_kernels.torch subpackage holds PyTorch reference implementations, and tile_kernels.modeling wraps the Engram gate kernels in a torch.autograd.Function so they can be used directly in a trainable layer.[1][4]

ItemDetail
DeveloperDeepSeek (DeepSeek-AI)
Repositorygithub.com/deepseek-ai/TileKernels
Repository created22 April 2026
First release1.0.0 on PyPI, 23 April 2026
Current release2.0.0, 30 September 2026
LanguagePython, kernels in TileLang
Backends (v2.0.0)NVIDIA CUDA (SM90, SM100); Huawei Ascend 950 (CANN)
LicenseMIT
Installpip install tile-kernels

Expanded article table

Sources: GitHub repository metadata, README and PyPI.[1][2][6]

Kernel families

The 2.0.0 README groups the kernels into seven feature areas, each a subpackage of tile_kernels.[1] The function names below are those exported by each subpackage's __init__.py in the 2.0.0 source.[4]

AreaSubpackageWhat it contains
MoE routingmoeTop-k expert selection and scoring: topk_gate, moe_topk_gate_forward and moe_topk_gate_backward, and normalize_weight
QuantizationquantPer-token, per-block and per-channel FP8/FP4 casting and dequantization ("cast back"), lossless per-block casting, fused SwiGLU forward/backward with per-token cast, RMSNorm forward/backward with optional fused FP8 cast, and a batched transpose for weights and their scale factors
EngramengramEngram hashing, gate forward and backward with fused RMSNorm, fused weight and weight-gradient reduction, and a set of Sinkhorn normalization kernels (momentum update, step, step reduce, finalize)
Manifold HyperConnectionmhcExpand, pre-norm, mix splitting and application, head mix computation, post, Sinkhorn normalization (forward and backward), a "big fuse" pre kernel and a multi-layer recomputation kernel
TransformtransformA rotary position embedding kernel (apply_rotary)
Randrandrandn, randn_like and seed handling
ModelingmodelingAn autograd wrapper (EngramGateFn) for the Engram gate

Expanded article table

Sources: README and source tree at version 2.0.0.[1][4]

MoE gating with Sqrt(Softplus)

The routing kernels match a specific DeepSeek architecture change. The DeepSeek V4 technical report says that, unlike DeepSeek V3, V4 changes "the activation function that computes the affinity scores from Sigmoid(·) into Sqrt(Softplus(·))."[11] The 2.0.0 moe_topk_gate_forward function accepts only this scoring function: its docstring says the scoring_func argument "must be 'sqrtsoftplus' (scores are sqrt(softplus(logits)))." The same function also takes required arguments for the number of shared experts, whether to treat shared experts as routed experts and the expert-parallel rank, and optional arguments for expert bias terms, a separate bias for image tokens, a routing mask and a logical-to-physical expert map.[8]

Quantization

Many of the library's kernels are casts between BF16 or FP32 activations and low-precision formats, often fused with the operation that precedes them. The per-token cast supports scale-factor groups of 16, 32, 64 or 128 channels and options such as power-of-two rounding of scale factors and packed UE8M0 scale factors.[18] The FlashMLA README uses tile_kernels.quant.per_token_cast in its example code to quantize attention projection weights to FP8 "with per-token scale factors, in DeepGEMM's layout" before handing them to FlashMLA's fused attention kernel, and it lists TileLang, TileKernels and DeepGEMM as requirements for that kernel's test script.[13]

mHC and Engram

The mhc kernels implement the pieces of Manifold-Constrained Hyper-Connections, which the DeepSeek V4 report introduces to "enhance conventional residual connections."[11] The DeepSeek V4 report says its training framework obtains "cost-effective mHC implementations via recomputation and fused kernels."[11] TileKernels includes a multi-layer recompute kernel and Sinkhorn forward and backward kernels among its mHC operations.[4]

The engram kernels serve Engram, a component of DeepSeek V4.1-Flash that the model's technical report says "adds sparsely accessed conditional memory."[12] The same report says that for Engram table updates, "Sinkhorn normalization maintains row and column scaling vectors across iterations to avoid repeated writes of the full normalized matrix." It adds: "Row normalization and the accumulation of partial column statistics are fused into a single kernel."[12] The Sinkhorn kernels added to TileKernels in version 2.0.0 have that shape: engram_sinkhorn_step takes per-row and per-column scale vectors, and its docstring reads "Normalize one Sinkhorn iterate and write per-CTA column partials."[20]

Use in DeepSeek models

DeepSeek's V4 technical report (April 2026) explains why the company writes such kernels in TileLang. It says the V4 architecture "would have resulted in hundreds of fine-grained Torch ATen operators," and that DeepSeek adopted TileLang "to develop a set of fused kernels to replace the vast majority of them."[11] The V4 report cites TileLang itself rather than naming the TileKernels repository.[11]

The DeepSeek V4.1-Flash technical report (September 2026) names the library directly. Describing the inference system, it says DeepSeek keeps hardware "fully pipelined inside a small number of fused kernels," listing a fused attention kernel in FlashMLA, the Mega-Gate, Mega-mHC and Mega-MoE kernels in DeepGEMM, "the kernels in TileKernels," and the TopK kernel in DeepSelect. According to the report, most V4.1-Flash Transformer layers "execute with only 15 kernels during prefill and 11 during decode."[12] The report's bibliography lists TileKernels as a 2026 GitHub release by Xiangwen Wang, Chenhao Xu, Huanqi Cao, Rui Tian, Weilin Zhao, Kuai Yu and Chenggang Zhao.[12]

According to IT Home's report of DeepSeek's 30 September announcement, DeepSeek said the TileLang approach now carries the implementation of most operators used in training the DeepSeek V4 series.[14]

History

Version 1.0.0 (April 2026)

The GitHub repository was created on 22 April 2026 with the description "A kernel library written in tilelang."[2] Its history begins with a single "Initial commit" dated 23 April 2026 that added about 14,200 lines across 127 files, followed the same day by a pull request from Rui Tian, one of the listed authors, revising comments in the Engram gate kernel.[3][4][21] Version 1.0.0 was uploaded to PyPI on 23 April 2026.[6]

The first README described the project more cautiously than the current one. It said of the kernels, "Some of them have already been used in internal training and inference scenarios," and added that the kernels "do not represent best practices and we are actively working on improving the code quality and documentation."[4] Version 1.0.0 targeted NVIDIA GPUs only. Its feature list included gating, a wider set of MoE routing kernels (token-to-expert mapping and fused expansion and reduction), quantization including an E5M6 format, batched transpose, Engram, mHC, and autograd wrappers for both the Engram gate and the mHC pipeline.[4]

Version 2.0.0 (September 2026)

Version 2.0.0 was merged as pull request #34 and published to PyPI on 30 September 2026.[5][6] The release commit changed 229 files, adding about 23,100 lines and deleting about 10,800.[5] It made these changes:

  • Two backends per kernel. Most kernel modules were split into a dispatcher (*_kernel.py) plus a CUDA implementation (*_cuda.py) and an Ascend implementation (*_asc.py).[5]
  • Runtime backend selection. tile_kernels/config.py gained an is_ascend() check and a get_device() helper that returns npu on Ascend systems and cuda otherwise; the dispatchers call the Ascend kernel builder when it returns true.[7][8]
  • New kernels. RMSNorm forward and backward (with an optional fused FP8 cast), a generic cast, per-token cast-and-cast-back, Engram Sinkhorn normalization, sqrt-softplus MoE top-k gating forward and backward, rotary position embedding, and a randn kernel.[5][4] A source comment describes the Ascend randn as a "temporary replacement for the extremely slow torch_npu.randn"; on CUDA it falls back to torch.randn.[9]
  • Removed kernels. The standalone transpose package, the E5M6 cast kernels, the mHC autograd modeling layer and several MoE routing kernels from version 1.0.0 (including fused mapping, fused expansion and fused reduction) were deleted.[5]
  • Test tooling. The test harness moved into tile_kernels/testing/pytest, with plugins for benchmarking, GPU memory tracking, kernel precompilation and fail-fast parallel runs.[5]
  • Wider author list. The package metadata grew from seven to sixteen named authors, all with DeepSeek email addresses, and the README citation was updated to match.[10][1]

The README gained a "News" section with a single entry: "[2026-09-30] Huawei Ascend support: Added Huawei Ascend support and updated the usage documentation. Following the NVIDIA path, the kernels now ship a second backend that is selected automatically at runtime, so the same Python APIs run on both NVIDIA GPUs and Huawei NPUs."[1] Its acknowledgement section now thanks "Huawei for its technical support and engineering expertise throughout the development of Tile Kernels' Ascend backend."[1]

Requirements by version

Requirement1.0.0 (April 2026)2.0.0 (September 2026)
Python3.10 or higher3.12 or higher
PyTorch2.10 or higher2.13 or higher
TileLang0.1.9 or higher0.1.15 or higher
NVIDIA backendSM90 or SM100 GPU, CUDA Toolkit 13.1 or higherSM90 or SM100 GPU, CUDA Toolkit 13.1 or higher
Ascend backendNot supportedAscend 950 NPU, CANN 9.2.0 or higher

Expanded article table

Sources: README at the initial commit and at version 2.0.0, and pyproject.toml.[4][1][10]

Ascend release of 30 September 2026

On the morning of 30 September 2026, DeepSeek announced through its official WeChat account that it was open-sourcing infrastructure components for Huawei's Ascend computing platform, covering TileLang, compute libraries and a distributed communication library, and that every component maps one to one onto its earlier open-source components for NVIDIA hardware.[14] IT Home's summary of the announcement lists the roles of the five libraries: DeepGEMM for general matrix multiplication, DeepEP for large-scale cross-device communication, TileKernels for conventional vector compute and memory-access operators, FlashMLA for sparse attention, and DeepSelect for efficient data selection.[14] DeepSeek said that "in multiple key test cases" the components' compute and communication performance was close to the hardware limit, and that Huawei's team had given it full support, including joint work on a 128-card supernode based on the Ascend 950.[14]

The Ascend support was delivered in different ways across the stack. DeepGEMM and DeepEP received separate -Ascend repositories, while TileKernels and FlashMLA added Ascend backends inside their existing repositories.[14][1][13] TileLang itself, which is developed by the tile-ai project rather than by DeepSeek, announced official support for Huawei Ascend 950 NPUs on the same date.[19]

Reuters, as relayed by ANI and The Tribune, reported that DeepSeek confirmed the release in a WeChat statement and that the two companies had jointly advanced the 128-chip Ascend 950 supernode.[17] DQIndia listed TileKernels as the component that "covers vector computation and memory access."[16] The Neuron noted that TileKernels "can select either an Nvidia or Ascend backend while exposing the same Python APIs to developers," but argued that "portability at the API layer is different from parity underneath it," pointing out that TileKernels "lists separate Nvidia CUDA and Huawei CANN requirements depending on which backend is running."[15]

Relationship to other DeepSeek libraries

TileKernels is one of several kernel libraries DeepSeek has open-sourced, each covering a different part of the model's compute.

LibraryRoleRelation to TileKernels
TileLangKernel DSL and compiler (tile-ai project)TileKernels is written entirely in TileLang and requires it as a dependency[1]
DeepGEMMMatrix multiplication (GEMM) kernels, including the fused Mega-MoE kernels[12]TileKernels quantization kernels can emit scale factors in DeepGEMM's layout[13]
FlashMLAAttention kernelsFlashMLA's fused-kernel tests require TileKernels for FP8 weight casting[13]
DeepEPExpert-parallel communicationSeparate library; covers communication rather than compute[14]
DeepSelectTop-k kernels for DeepSeek Sparse Attention and samplingSeparate library; listed with TileKernels among V4.1-Flash inference kernels[12]

Expanded article table

Community activity

As of 1 October 2026 the repository had 1,882 stars, 184 forks and 30 open issues and pull requests on GitHub.[2] Only two pull requests had been merged: the April comment fix (#1) and the 2.0.0 release (#34).[3][21] Open community pull requests include bug fixes for quantization, MoE and mHC kernels and a proposal to add AMD MI350 (gfx950, HIP) support, which had not been merged as of 1 October 2026.[21]

See also

References

  1. ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11 ^12 ^13 ^14DeepSeek-AI. "Tile Kernels" (README, version 2.0.0). GitHub, deepseek-ai/TileKernels. Accessed 1 October 2026. github.com/...TileKernels
  2. ^1 ^2 ^3 ^4GitHub REST API. "deepseek-ai/TileKernels" repository metadata (created_at, description, license, stars, forks). Accessed 1 October 2026. api.github.com/...TileKernels
  3. ^1 ^2 ^3GitHub. "Commits: deepseek-ai/TileKernels, main branch." Accessed 1 October 2026. github.com/...main
  4. ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10DeepSeek-AI. "Initial commit" (79461d72), 23 April 2026, including README.md at that commit; and source tree at version 2.0.0. GitHub. github.com/...1d72d6f97f91b89403705a841c9d4c1eddd6
  5. ^1 ^2 ^3 ^4 ^5 ^6 ^7DeepSeek-AI. "Release tile-kernels v2.0.0," pull request #34 (commit 828cc15d), merged 30 September 2026. GitHub. github.com/...34
  6. ^1 ^2 ^3 ^4Python Package Index. "tile-kernels" release history (1.0.0, 23 April 2026; 2.0.0, 30 September 2026). Accessed 1 October 2026. pypi.org/...tile-kernels
  7. ^DeepSeek-AI. `tile_kernels/config.py`, version 2.0.0. GitHub. github.com/...config.py
  8. ^1 ^2 ^3DeepSeek-AI. `tile_kernels/moe/moe_topk_gate_forward_kernel.py`, version 2.0.0. GitHub. github.com/...moe_topk_gate_forward_kernel.py
  9. ^DeepSeek-AI. `tile_kernels/rand/randn_kernel.py`, version 2.0.0. GitHub. github.com/...randn_kernel.py
  10. ^1 ^2DeepSeek-AI. `pyproject.toml` at the initial commit and at version 2.0.0. GitHub. github.com/...pyproject.toml
  11. ^1 ^2 ^3 ^4 ^5DeepSeek-AI. "DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence." arXiv:2606.19348v1, 26 April 2026. arxiv.org/...2606.19348
  12. ^1 ^2 ^3 ^4 ^5 ^6DeepSeek-AI. "DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression" (technical report). Hugging Face, 10 September 2026. huggingface.co/...DeepSeek_V41_Tech_Report.pdf
  13. ^1 ^2 ^3 ^4DeepSeek-AI. "FlashMLA" (README). GitHub, deepseek-ai/FlashMLA. Accessed 1 October 2026. github.com/...FlashMLA
  14. ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8清源. "DeepSeek 开源面向华为昇腾算力平台的基础设施组件,与英伟达平台一一对应" (DeepSeek open-sources infrastructure components for Huawei's Ascend computing platform, matching its NVIDIA components one to one). IT之家 (IT Home), 30 September 2026. ithome.com/...604
  15. ^1 ^2Eric Gerard Ruiz. "DeepSeek Opened the Code. Can Huawei Deliver the Compute?" The Neuron, 30 September 2026. theneuron.ai/...ode-can-huawei-deliver-the-compute
  16. ^"DeepSeek expands Huawei Ascend push with six open-source AI tools." DQIndia (Dataquest), 30 September 2026. dqindia.com/...h-six-open-source-ai-tools-12594158
  17. ^ANI. "DeepSeek joins hands with Huawei to build AI chip software, cutting Nvidia reliance." The Tribune, 30 September 2026. tribuneindia.com/...ftware-cutting-nvidia-reliance
  18. ^1 ^2DeepSeek-AI. `tile_kernels/quant/per_token_cast_kernel.py`, version 2.0.0. GitHub. github.com/...per_token_cast_kernel.py
  19. ^tile-ai. "Tile Language" (README, Latest News and Acknowledgements). GitHub, tile-ai/tilelang. Accessed 1 October 2026. github.com/...tilelang
  20. ^DeepSeek-AI. `tile_kernels/engram/engram_sinkhorn_kernel.py`, version 2.0.0. GitHub. github.com/...engram_sinkhorn_kernel.py
  21. ^1 ^2 ^3GitHub. "Pull requests: deepseek-ai/TileKernels" (all states). Accessed 1 October 2026. github.com/...pulls
  22. ^DeepSeek. "DeepSeek V4 Preview Release." DeepSeek API Docs, 24 April 2026. api-docs.deepseek.com/...news260424

Improve this article

Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.

1 revision · v2 · 2,630 words · full history

Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify

Research and drafting on this wiki are AI-assisted, under named human editorial standards. How AI is used here

Reviewer note: Independently fact-checked 1 Oct 2026 against the repo, GitHub API, PyPI and the DeepSeek V4 and V4.1-Flash tech reports; 1 minor defect corrected

Cite this page: AI Wiki. "TileKernels." aiwiki.ai, updated 1 Oct 2026, fact-checked 1 Oct 2026. CC BY 4.0. https://aiwiki.ai/wiki/tilekernels

Suggest edit