TileLang
TileLang (also written tile-lang, and self-described as "Tile Language") is an open-source domain-specific language for writing high-performance AI kernels such as GEMM, dequantized GEMM, FlashAttention and linear attention. Kernels are written in Python-like syntax and compiled through an infrastructure built on top of Apache TVM, which lets a kernel author describe the dataflow of tiles between memory levels while the compiler handles thread binding, memory layouts, use of hardware matrix instructions and software pipelining.[1][2] The project lives in the GitHub organization tile-ai; the main repository tile-ai/tilelang was created on 3 October 2024, was made publicly available on 20 January 2025, and had about 7,900 stars, more than 800 forks and more than 200 listed contributors as of 1 October 2026.[1][3] It is distributed under the MIT license and published on PyPI as tilelang, with documentation at tilelang.com.[4][5][6]
TileLang began as academic work: the paper describing it lists eleven authors, most of them at Peking University, with four at Microsoft Research in Beijing.[2] It became industrially prominent through DeepSeek, which published TileLang implementations of its sparse-attention kernels alongside hand-written CUDA ones from DeepSeek V3.2-Exp onward, maintains a TileLang kernel library (TileKernels), and, on 30 September 2026, released a set of components for Huawei Ascend NPUs in which TileLang is the programming layer.[7][8][9] On the same day, TileLang 0.1.15 added a Huawei Ascend 950 backend to the main repository, a backend whose initial version the project credits mainly to contributors from DeepSeek, with support from Huawei.[10][11]
Origins and the TileLang paper
The repository's own history places the first public release at 20 January 2025 (v0.0.1, a CUDA-only pre-release) and the first v0.1 release at 12 February 2025.[1][3] The project's acknowledgements state that the initial version was mainly developed by the GitHub users LeiWang1999, chengyupku and nox-410 "with supervision from Prof. Zhi Yang at Peking University", and that part of the work was carried out during an internship at Microsoft Research, where Lingxiao Ma, Yuqing Xia, Jilong Xue and Fan Yang gave advice and support.[1] The MIT license file carries an unusual note: "During the period from December 1, 2024, to Mar 14, 2025, this project is subject to additional collaboration terms with Microsoft Corporation."[4]
The reference paper is "TileLang: A Composable Tiled Programming Model for AI Systems" (arXiv:2504.17577), submitted on 24 April 2025 and revised on 27 April 2025, released under CC BY 4.0.[2][12] Its authors are Lei Wang, Yu Cheng, Yining Shi, Zhengju Tang, Zhiwen Mo, Wenhao Xie, Lingxiao Ma, Yuqing Xia, Jilong Xue, Fan Yang and Zhi Yang; the first three are marked as equal contributors. Affiliations as printed in the paper are Peking University for Wang, Cheng, Shi, Tang, Xie and Zhi Yang, Imperial College London for Zhiwen Mo, and Microsoft Research in Beijing for Ma, Xia, Xue and Fan Yang.[2] The stated thesis is that AI kernels follow clear dataflow patterns (move tiles between DRAM and SRAM, compute on them) but that writing fast versions of them requires hardware-centric tuning, so TileLang "decouples scheduling space (thread binding, layout, tensorize and pipeline) from dataflow" and exposes those choices as annotations and primitives.[12]
Programming model
A tile in TileLang is a first-class object: a shaped portion of data owned and manipulated by a warp, thread block or equivalent parallel unit. T.Kernel establishes the execution context (block indices and thread count), and buffers are placed explicitly in the memory hierarchy rather than left to opaque compiler passes: T.alloc_shared maps to on-chip shared memory and T.alloc_fragment to the register file, with a layout-inference pass deriving how a block-level fragment is partitioned across threads.[12] Data movement uses T.copy, initialization uses T.clear or T.fill, and elementwise work is expressed with T.Parallel.[12]
Above that sit tile operators, each of which must implement two interfaces: Lower, which turns the operator into lower-level IR (thread bindings, vectorized accesses), and InferLayout, which determines the buffer and loop layouts it needs. T.gemm, for example, requests swizzled shared-memory layouts for its inputs and a matrix-specific fragment layout for its output; T.reduce supports sum, min, max and product reductions using warp-level and block-level parallelism; T.atomic maps to native hardware atomics.[12] Scheduling is then expressed as annotations layered on that dataflow: T.Pipelined(..., num_stages=N) for software pipelining of a loop, T.Parallel for automatic loop-to-thread mapping and vectorization, T.annotate_layout to override the default (bank-conflict-minimizing) layout, and T.use_swizzle for L2-friendly rasterization.[12]
The quick-start kernel in the repository's README shows the style: an FP16 GEMM with FP32 accumulation and a fused ReLU epilogue, in which the author writes the tile loop and the copies and leaves vectorization, layouts and synchronization to the compiler.[1]
@tilelang.jit
def matmul_relu(A, B, block_M: int = 128, block_N: int = 128, block_K: int = 32):
M, N, K = T.const("M, N, K")
A: T.Tensor((M, K), T.float16)
B: T.Tensor((K, N), T.float16)
C = T.empty((M, N), T.float16)
with T.Kernel(T.ceildiv(N, block_N), T.ceildiv(M, block_M), threads=128) as (bx, by):
A_shared = T.alloc_shared((block_M, block_K), T.float16)
B_shared = T.alloc_shared((block_K, block_N), T.float16)
C_local = T.alloc_fragment((block_M, block_N), T.float32)
T.clear(C_local)
for k in T.Pipelined(T.ceildiv(K, block_K), num_stages=3):
T.copy(A[by * block_M, k * block_K], A_shared)
T.copy(B[k * block_K, bx * block_N], B_shared)
T.gemm(A_shared, B_shared, C_local)
for i, j in T.Parallel(block_M, block_N):
C_local[i, j] = T.max(C_local[i, j], 0)
T.copy(C_local, C[by * block_M, bx * block_N])
return C
Compiler and scheduling automation
The paper describes four scheduling spaces that TileLang automates on top of dataflow.[12]
| Area | Abstraction | What the compiler does |
|---|---|---|
| Memory layout | Layout (an algebraic index-to-address function over IterVars, inspired by TVM) and Fragment (a layout whose output is a thread index plus a register index) | Composes layouts, supports non-bijective transforms such as padding, applies built-in strategies such as swizzling to avoid shared-memory bank conflicts |
| Thread binding | A LayoutMap plus a priority order over tile operators | Infers layouts top-down from the strictest operator (GEMM on tensor cores) to the most flexible (elementwise), replicating or aligning buffers so dependent operators match |
| Tensorization | T.gemm, C++ source injection via T.import_source/T.call_extern, or inline PTX via T.ptx | Selects hardware instructions, by default through tile libraries such as NVIDIA's CuTe or AMD's composable kernel |
| Pipelining | T.Pipelined with a single num_stages argument | Analyzes dependencies between copy and compute blocks, interleaves them, and maps asynchronous work onto hardware units |
For pipelining, the paper reports that TileLang emits cp.async with the matching commit and wait instructions on Ampere, and on Hopper lowers copies to the TMA unit and matrix multiplication to wgmma.mma_async, performing warp specialization automatically by classifying statements as producers or consumers and inserting memory barriers guided by live-variable analysis.[12] The authors also note a cost of the tile-library route: using NVCC's trace tool, template expansion accounted for roughly 90 percent of compilation time for CUDA code generated by TileLang, which is part of why the project also supports implementing instructions in TileLang itself.[12]
Later engineering has moved parts of this infrastructure: the runtime interface was moved to apache-tvm-ffi in October 2025, SMT-based symbolic reasoning (Z3) was integrated into the TVM arithmetic analyzer in December 2025, and TileLang's IR usage migrated to TVM's TIRX representation in May 2026.[1] Release 0.1.15 added an opt-in role-based automatic warp-specialization scheduler for CUDA, which assigns TMA loads, MMA computation, TMA stores and worker operations to specialized warp groups.[10]
Backends and hardware support
The README describes TileLang as "evolving into a multi-backend compiler (TileLang-X)" organized around a modular backend abstraction, with Target objects selecting the compilation target and an auto target that detects CUDA, HIP, Metal and Ascend devices. Prebuilt wheels are published for Linux x86-64 and AArch64, Windows x86-64 and macOS arm64. Backends are graded: Primary and Supported and Experimental live in the main repository, while Ecosystem adapters live in separate repositories, are not in the release wheels, and may follow their own compatibility schedules.[1]
| Backend | Target | Hardware | Level |
|---|---|---|---|
| NVIDIA CUDA | cuda | SM70 through SM120 code paths | Primary |
| AMD ROCm/HIP | hip | CDNA and RDNA, including gfx942 and gfx950 paths | Supported |
| Huawei Ascend 950 | ascend | Ascend 950 NPUs | Supported |
| Apple Metal | metal | macOS on Apple silicon | Supported |
| LLVM CPU | llvm | Host CPUs | Experimental |
| NVIDIA CuTe DSL | cutedsl | NVIDIA GPUs | Experimental |
| WebGPU | webgpu | WebGPU runtimes | Experimental |
| Huawei Ascend A2/A3 | ascendc, pto, npuir | Ascend A2 and A3 | Ecosystem (tilelang-ascend, tilelang-mlir-ascend) |
| MetaX MACA | maca | MetaX C500 and C600 | Ecosystem (tilelang-metax) |
| Moore Threads MUSA | musa | Moore Threads S5000, S4000, M1000 | Ecosystem (tilelang-musa) |
| HYGON | hcu | BW1000, BW1100, BW150, K100_AI | Ecosystem (tilelang-hygon) |
| Sunrise-AI TANG | tang | Sunrise S2 and S3 | Ecosystem (tilelang-sunrise) |
The per-vendor adapter repositories in the tile-ai organization were created over 2025 and 2026: tilelang-ascend in September 2025, tilelang-metax in March 2026, tilelang-musa in March 2026, tilelang-mlir-ascend in April 2026, tilelang-hygon in June 2026 and tilelang-sunrise in August 2026.[3]
Releases
The project ships frequent point releases; 23 releases were published between January 2025 and 30 September 2026.[3] Selected ones, with highlights as described by the maintainers:
| Release | Date | Notes |
|---|---|---|
| v0.0.1 | 20 Jan 2025 | Pre-release, CUDA prebuilt wheels only |
| v0.1.0 | 12 Feb 2025 | First v0.1 public release |
| v0.1.6.post2 | 30 Oct 2025 | Final release compatible with Python 3.8 |
| v0.1.7 | 7 Dec 2025 | Followed in December 2025 by a CuTe DSL backend and Z3 integration |
| v0.1.8 | 16 Feb 2026 | Dynamic pipeline improvements, richer layout representations, AMD fixes |
| v0.1.9 | 22 Apr 2026 | CuTe DSL GEMM V2, Metal code generation, builds without a host toolchain |
| v0.1.10 | 25 May 2026 | Broader AMD and Blackwell support, initial Metal GEMM, Windows packaging |
| v0.1.11 | 8 Jun 2026 | Scan, pipeline, backend, CUDA, ROCm and Metal work |
| v0.1.12 | 8 Jul 2026 | LLVM backend, tile scheduler, backend registry, pass visualizer |
| v0.1.13 | 3 Aug 2026 | Multi-backend language dialect, source locations in diagnostics, new CUDA and Metal paths, removal of several legacy APIs |
| v0.1.14 | 2 Sep 2026 | Reducer v2, warp-specialization schedules, layout-inference cost models, faster cold compilation |
| v0.1.15 | 30 Sep 2026 | Native Huawei Ascend 950 backend, automatic CUDA warp specialization, unified block-scaled GEMM, more expressive Python frontend |
Tooling grew alongside the compiler: a pass visualizer and pass-diff tool (June and July 2026), an IR lower trace (July 2026), compiler-pass timing, an automatic delta-debugging tool (AutoDD), a Language Server Protocol implementation published in August 2026, and a set of ten teaching exercises published as tilelang-puzzles in February 2026.[1][3]
Reported performance
All of the following figures are reported by TileLang's own authors or maintainers and were measured by them.
The paper evaluates TileLang on NVIDIA H100 80GB, NVIDIA A100 80GB and AMD Instinct MI300X 192GB, against FlashAttention-3, Triton, cuBLAS, rocBLAS, PyTorch, BitsandBytes and Marlin.[12]
| Workload | Hardware | Reported result |
|---|---|---|
| Multi-head attention | H100 | 1.36x over FlashAttention-3, 1.41x over Triton, 1.70x over PyTorch; close to FlashAttention-3 at 8k sequence length |
| Linear attention (Mamba-2 chunk-scan and chunk-state) | H100 | Average 1.77x and 2.10x over Triton |
| Multi-head latent attention | H100 | 1075.9x over Torch, and up to 98 percent of hand-optimized FlashMLA performance, in about 70 lines of Python |
| Multi-head latent attention | MI300X | 129.2x over Torch, and 95 percent of the hand-written AITER library |
| GEMM | RTX 4090, A100, H100, MI300X | 1.10x, 0.97x, 1.00x and 1.04x versus vendor libraries; 1.08x, 1.03x, 1.13x and 1.25x versus Triton |
| Dequantized GEMM (via BitBLAS with a TileLang backend) | A100 | Up to 7.65x over cuBLAS FP16 at INT2 weights with INT8 activations; an average 1.04x over Marlin at INT4 weights with FP16 activations; an average 1.62x over BitsandBytes at NF4 weights |
The repository additionally reports that its manually pipelined DeepSeek V3.2 sparse MLA forward kernel, written to match FlashMLA's schedule, reaches close to 600 TFLOP/s on an H800 SXM, and that the sparse MLA backward kernel reaches roughly 100 TFLOP/s on H800 SXM and 115 TFLOP/s on H200 SXM.[13] For the Ascend 950 backend, the maintainers say they benchmark TileLang against Torch NPU on BF16 GEMM, FP8 casting and GQA backward across four shapes per operator, reporting compute-bound operators in TFLOP/s and the memory-bound cast in effective GB/s; the results are published as a figure rather than a table.[11]
Relation to Triton and CUDA
TileLang's authors position it against Triton rather than as a replacement for CUDA: both are Python-embedded kernel languages, but TileLang lets the author declare buffers explicitly at different levels of the memory hierarchy in the frontend and then uses layout inference to parallelize operations on them, while still permitting an expert to specify exact per-thread behavior.[12] The paper also contrasts TileLang with the C++ library ThunderKittens, arguing that TileLang keeps the author in Python and applies pipelining by default (async copies on Ampere, TMA on Hopper) while leaving manual pipelining available.[12]
On NVIDIA hardware TileLang is a CUDA generator rather than an alternative runtime: T.gemm dispatches by default through NVIDIA's CUTLASS tile library (and AMD's composable kernel on ROCm), kernels can embed raw CUDA via T.CUDASourceCodeKernel, compilation can go through NVRTC, and an experimental backend targets NVIDIA's CUTLASS CuTe DSL.[1][12] DeepSeek's own framing of the comparison is political as well as technical: in the WeChat post that accompanied its Ascend release, reported by Reuters, the company said that building "a new generation of independent, self-controlled GPU software ecosystems" requires first "establishing a high-level language that is universal, easy to program, and still capable of reaching the hardware's full performance potential", said "TileLang was created precisely to meet this need", and described it as offering "a simpler programming model" than NVIDIA's CUDA.[14]
The Ascend backends
TileLang's Ascend support is split in two, and the distinction matters because the two tracks target different NPU generations and live in different repositories.
Upstream Ascend 950 backend. Added in v0.1.15 on 30 September 2026, this is a backend inside tile-ai/tilelang, selected with target="ascend" (defaulting to the dav-3510 architecture) and used through the tilelang.ascend.language dialect. It reuses TileLang's shared frontend and adds Ascend-specific lowering, scheduling, synchronization and code generation for the Cube and Vector execution units, integrating the Bisheng compiler, tvm_ffi and Cython execution, PyTorch NPU tensors and streams, and NPU profiling. It must be built from source with USE_ASCEND=ON and requires a compatible Ascend driver and CANN toolkit plus torch_npu.[10][11] Ascend kernels allocate Unified Buffer storage with T.alloc_shared, L1 storage with T.alloc_l1 and Cube buffers with T.alloc_l0a, T.alloc_l0b and T.alloc_l0c; a single kernel can mix T.gemm on the Cube cores with SIMT regions (T.SimtVF(threads=...)) and SIMD regions (T.SimdVF()) on the Vector cores, with the compiler assigning cores, scheduling the overlap and inserting synchronization.[11]
What a higher-level programming model over Ascend's low-level stack means in practice is set out in the backend guide as a division of labour against hand-written Ascend C: the author writes allocations, copies, tile sizes, tiled operators such as T.gemm, scalar or SIMD compute regions and T.Pipelined annotations, and skips TPipe/TQue setup and DMA calls, low-level operator implementations, thread mapping and barrier placement, manual schedules and buffer-slot rotation, and set/wait flag management with flag IDs; the compiler plans storage, lowers tiled operations to hardware instructions, assigns tasks to AI Cube and AI Vector units, and infers and inserts paired synchronization flags. The same guide is careful to say that automatic scheduling "does not imply automatic tensorization of arbitrary T.Parallel code into SIMD instructions".[11] The project credits the initial version of this backend mainly to nine named GitHub contributors "from DeepSeek AI", thanks the main TileLang maintainer for help integrating it, and thanks "the Huawei team for their close collaboration and valuable support".[11]
TileLang-Ascend (A2 and A3). The separate repository tile-ai/tilelang-ascend was created on 25 September 2025 and announced as open source on 29 September 2025. It is an Ascend-specific variant of the language, MIT-licensed, tested and validated on A2 and A3 NPUs, and its compiler backend supports two technical routes held on different branches: Ascend C with PTO, and AscendNPU IR. It assumes CANN 8.3.RC1 or newer and torch-npu 2.6.0.RC1 or newer.[15] Its documentation maps the GPU memory hierarchy onto the NPU (global memory to global memory; shared memory to the L1 buffer on the Cube core and the Unified Buffer on the Vector core; registers to the L0A, L0B and L0C buffers), offers both automatic vectorization through T.Parallel and explicit tile primitives such as T.tile.add, and distinguishes a "Developer mode" in which the compiler separates Cube and Vector scopes and inserts synchronization from an "Expert mode" in which the author declares scopes with T.Scope("C")/T.Scope("V") and manages cross-core flags. Features arrived over 2025 and 2026 in that order: automatic synchronization insertion (November 2025), automatic buffer reuse (November 2025), T.Parallel (December 2025), T.Pipelined (January 2026), a PTO code-generation backend (January 2026), a programming guide (January 2026), shared-memory put/get primitives for inter-core communication (March 2026), wheel installation (March 2026), Flash Attention and sparse Flash Attention benchmarks (March 2026) and DeepSeek V4 kernels (April 2026). Its acknowledgements credit Huawei's HiSilicon, ICT and Compiler and Programming Language Lab, and the Peking University Kunpeng and Ascend Center.[15] A third repository, tile-ai/tilelang-mlir-ascend ("MLIR-based TileLang Ascend Adapter"), was created in April 2026.[3] The v0.1.15 release notes state that "Ascend A2/A3 support remains in the community-maintained TileLang-Ascend projects".[10]
Adoption by DeepSeek
DeepSeek is TileLang's most visible industrial user. The DeepSeek-V3.2-Exp repository points readers to the TileLang examples directory for "TileLang kernels with better readability and research-purpose design", listing it alongside the high-performance CUDA kernels in DeepGEMM and FlashMLA.[7] The TileLang repository carries example directories for DeepSeek MLA, DeepSeek V3.2, DeepSeek V4 and DeepSeek mHC, and its changelog records V3.2-specific work including a sparse MLA backward kernel and a top-k selector optimization in July 2026 that the maintainers say gave about 1.9x higher performance in their benchmark, plus "DeepSeek V4 operators" examples added in May 2026.[1]
DeepSeek also maintains TileKernels, created on 22 April 2026 and MIT-licensed, described as "a library of dozens of highly optimized kernels implemented in TileLang". It covers mixture-of-experts routing, Engram gating, FP8 and FP4 quantization casts and manifold hyper-connection (mHC) kernels including Sinkhorn normalization, plus RoPE and random-number kernels. The README says most kernels reach performance close to the hardware's compute or memory-bandwidth limits and that all of them have been used in DeepSeek's internal training and inference workloads. It requires TileLang 0.1.15 or newer, Python 3.12 or newer, PyTorch 2.13 or newer, and either an NVIDIA SM90 or SM100 GPU with CUDA 13.1 or newer, or an Ascend 950 NPU with CANN 9.2.0 or newer.[8] On 30 September 2026 TileKernels added Huawei Ascend support, with a second backend selected automatically at runtime so the same Python APIs run on NVIDIA GPUs and Huawei NPUs.[8] DeepGEMM-Ascend declares tilelang as a package dependency, used by its HC prenorm kernel, and its acknowledgements state that the mHC kernel "is backed by Tilelang"; FlashMLA's test script for the fused norm-RoPE-attention-RoPE-cast kernel requires TileLang, TileKernels and DeepGEMM.[9][16]
Outside DeepSeek, the SGLang project's day-zero support for DeepSeek V4, described on the LMSYS blog on 25 April 2026, uses TileLang kernels in several places: it extends a two-stage split-K TileLang mHC pre-GEMM kernel to partition the K dimension across CTAs for small-batch decoding, integrates a fused mhc_pre_big_fuse_tilelang path combining RMSNorm, Sinkhorn and residual mixing in one kernel, adapts TileLang DSA indexer kernels to V4's indexer, and extends a sparse-MLA TileLang kernel with per-head learnable attention-sink logits. The same post lists "tilelang attention" among the kernels used in its reinforcement-learning training path.[17]
The DeepSeek Ascend release of 30 September 2026
On 30 September 2026 DeepSeek published a set of Ascend-targeted infrastructure repositories that mirror its existing NVIDIA stack. Reuters, reporting on a post from DeepSeek's official WeChat account, wrote that DeepSeek said it had partnered with Huawei to develop programming tools optimized for Ascend chips, that Huawei "provided full support in developing the programming infrastructure", and that the two companies jointly advanced a "supernode" solution based on 128 Ascend 950 chips covering both computation and communication; the agency noted the announcement came two weeks after Huawei unveiled its next generation of AI processors and supernode systems.[14] The Chinese technology outlet Zhidongxi (智东西) summarized the release in English as bringing TileLang, DeepGEMM-Ascend, DeepEP-Ascend, TileKernels, FlashMLA and DeepSelect to Ascend, with components corresponding to DeepSeek's existing NVIDIA stack.[18]
What the repositories themselves state, as of 1 October 2026:
| Component | Status on 30 September 2026 | Primary-source detail |
|---|---|---|
| TileLang | Ascend 950 backend released upstream in v0.1.15 | Native code generation, automatic Cube/Vector scheduling and synchronization, mixed SIMD/SIMT programming, MXFP8 and MXFP4 block-scaled GEMM[10][11] |
| TileKernels | Huawei Ascend support added | Second backend chosen automatically at runtime; validated on Ascend 950 with CANN 9.2.0 or newer[8] |
| DeepGEMM-Ascend | Initial release, repository created 29 September 2026 | API-compatible port of DeepGEMM supporting BF16, FP8 and FP4 GEMM, MQA logits and MegaMoE on Ascend 950; CANN 9.20 with bisheng and ld.lld; depends on tilelang for the HC prenorm kernel[9] |
| DeepEP-Ascend | Repository created 30 September 2026 | Expert-parallel all-to-all dispatch and combine for MoE on Ascend NPUs; Ascend C kernels over HCCL/HCOMM, UBMEM and URMA, compiled at runtime by DeepJIT; measured 373-375 GB/s dispatch and 345-347 GB/s combine at EP8, falling to 313-320 and 272-278 GB/s at EP128 on Ascend 950DT with CANN 9.2.0[19] |
| FlashMLA | Ascend attention kernels released | Sparse attention prefill and decoding kernels for Ascend 950 reported at up to 410 TFLOP/s in prefill (95 percent of hardware peak) and 360 TFLOP/s in decoding (83 percent of peak), with a technical report; the same release dropped Hopper and earlier-model support and changed the FP8/FP4 KV cache format[16] |
| DeepSelect | Ascend TopK kernels released | TopK kernels for DeepSeek Sparse Attention and samplers, reported at 2x to 20x over torch.topk; the Ascend implementation supports bfloat16 only[20] |
The release is not a single-language port. DeepEP-Ascend describes its own kernels as Ascend C compiled at runtime through DeepJIT, and neither FlashMLA nor DeepEP-Ascend presents its Ascend kernels as TileLang implementations; TileLang is the high-level layer for part of the stack (TileKernels, and the HC prenorm and mHC kernels in DeepGEMM-Ascend) and the component DeepSeek singled out publicly.[9][14][16][19]
Ecosystem
Several projects in the tile-ai organization build on or around the language: TileOPs, an operator library for LLMs built on TileLang (created June 2025); TileRT, a tile-based runtime for low-latency LLM inference (November 2025); tilescale, a tile-based language aimed at computation "across all scales" (July 2025); AttentionEngine (May 2025); Liger-tilelang, TileLang kernels for LLM training (March 2025); tilelang-puzzles, a ten-exercise tutorial (repository created December 2025, announced February 2026); and tilelang-lsp, a language server with inlay hints for buffer shapes, dtypes, scopes and inferred layouts (open-sourced August 2026).[1][3] The organization also hosts a redesign of BitBLAS, the mixed-precision library whose backend the TileLang authors replaced with TileLang to run their dequantized-GEMM comparisons.[3][12]
References
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11tile-ai/tilelang: Tile Language (README)
- ^1 ^2 ^3 ^4TileLang: A Composable Tiled Programming Model for AI Systems (arXiv abstract page, arXiv:2504.17577)
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8tile-ai organization repositories (GitHub)
- ^1 ^2tile-ai/tilelang LICENSE
- ^tilelang on PyPI
- ^TileLang 0.1.15 documentation
- ^1 ^2deepseek-ai/DeepSeek-V3.2-Exp (README)
- ^1 ^2 ^3 ^4deepseek-ai/TileKernels: Tile Kernels (README)
- ^1 ^2 ^3 ^4deepseek-ai/DeepGEMM-Ascend: DeepGEMM Ascend (README)
- ^1 ^2 ^3 ^4 ^5TileLang release v0.1.15
- ^1 ^2 ^3 ^4 ^5 ^6 ^7TileLang Ascend 950 Backend guide
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11 ^12 ^13 ^14TileLang: A Composable Tiled Programming Model for AI Systems (full text, arXiv:2504.17577v2)
- ^TileLang DeepSeek V3.2 examples (README)
- ^1 ^2 ^3DeepSeek partners with Huawei to develop chip programming tools (Reuters, via Business Standard, 30 September 2026)
- ^1 ^2tile-ai/tilelang-ascend: TileLang-Ascend (README)
- ^1 ^2 ^3deepseek-ai/FlashMLA (README)
- ^DeepSeek-V4 on Day 0: From Fast Inference to Verified RL with SGLang and Miles (LMSYS Org blog, 25 April 2026)
- ^Zhidongxi (智东西 China AI News) post on DeepSeek's Ascend stack, 30 September 2026
- ^1 ^2deepseek-ai/DeepEP-Ascend (README)
- ^deepseek-ai/DeepSelect (README)
Improve this article
Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.
1 revision · v2 · 4,058 words · full history
Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify
Research and drafting on this wiki are AI-assisted, under named human editorial standards. How AI is used here
Reviewer note: Independently fact-checked 1 Oct 2026 against the tile-ai repos, arXiv:2504.17577 and the 30 Sep 2026 Ascend release; 4 defects corrected incl. the Triton speedup range
Cite this page: AI Wiki. "TileLang." aiwiki.ai, updated 1 Oct 2026, fact-checked 1 Oct 2026. CC BY 4.0. https://aiwiki.ai/wiki/tilelang