Zhijian Liu
Zhijian Liu is a computer scientist who works on efficient machine learning. He is an assistant professor at the University of California, San Diego, where he directs Z Lab, and a co-founder of Inco AI, an inference company. [1][2][3] He received his PhD from the Massachusetts Institute of Technology under Song Han and has worked as a research scientist at NVIDIA. [1][4][12] His research group describes its aim as making AI "smaller, faster, and more efficient" through work that spans algorithms, systems, and applications. [1][5] Liu is the senior author of DFlash, a block-diffusion method for speculative decoding published at ICML 2026, which Inco AI later extended as DFlash 2 and built into its Splash inference engine for Apple silicon. [6][7][8]
Education
Liu received a bachelor of engineering degree from Shanghai Jiao Tong University. [9] He then studied at MIT, where he was advised by Song Han. His 2020 master's thesis, "Hardware-efficient deep learning for 3D point cloud," combined point-based and voxel-based processing into a primitive called Point-Voxel Convolution and reported that the resulting model had been deployed on the autonomous racing vehicle of MIT Driverless. [10] His doctoral thesis, "Efficient Deep Learning with Sparsity: Algorithms, Systems, and Applications," was defended in April 2024 and issued in May 2024. [11][12] In the defense abstract he described sparsity as a way to close a gap between the growing computational demand of deep learning and the slowing supply of hardware performance, covering input sparsity, a system library for sparse computation, automated model compression, and applications in autonomous driving, language modeling, and high-energy physics. [11]
As a doctoral student he received the Qualcomm Innovation Fellowship and the NVIDIA Graduate Fellowship, and he was named a Rising Star in ML and Systems by MLCommons and a Rising Star in Data Science by the University of Chicago and UC San Diego. [9][11]
Career
NVIDIA Research lists Liu as a research scientist, and his personal homepage also describes that role. [1][4] By September 2026 his X profile listed NVIDIA as a previous position, alongside his current faculty post at UC San Diego and "building something new (@inco_ai)." [3] At UC San Diego he directs Z Lab, which says it is part of the university's ML Systems Group and Center for Visual Computing, and he recruits doctoral students through the Halıcıoğlu Data Science Institute and the Department of Computer Science and Engineering. [1][5]
NVIDIA's technical blog post on DFlash, published on 23 June 2026 with Liu among its seven authors, describes him as an assistant professor at UC San Diego who leads Z Lab and says, "He is also a co-founder of Inco AI." [2] Inco AI, whose site describes its work as "inference, reimagined for the agentic era," released DFlash 2 in August 2026, put its hosted inference platform into public beta in September 2026, and open-sourced the Splash engine for Macs in mid-September 2026. [7][8][13][24] Inco's DFlash 2 post says, "Our team released DFlash in January," referring to the Z Lab method. [7] Inco's blog posts are credited to the company rather than to named authors. [7][8]
Research
Model compression and efficient 3D deep learning
Much of Liu's doctoral work in Song Han's group concerned automated, hardware-aware model compression and efficient processing of 3D data. He was one of three equal-contribution first authors of HAQ, a framework for hardware-aware automated mixed-precision quantization published at CVPR 2019, and a co-author of AMC, which used automated machine learning for model compression on mobile devices. [14][15] He was co-first author of Point-Voxel CNN (NeurIPS 2019), which combined point-based and voxel-based representations for efficient 3D deep learning, and of TorchSparse (MLSys 2022), an inference engine for sparse point-cloud convolution. [16][17] He also co-authored "Deep Leakage from Gradients" (2019), which showed that shared training gradients can be used to reconstruct private training data. [18]
Autonomous driving and point-cloud transformers
Liu was co-first author of BEVFusion, a multi-task, multi-sensor fusion framework that unifies camera and LiDAR features in a bird's-eye-view representation, published at ICRA 2023, and of FlatFormer, an efficient point-cloud transformer based on flattened window attention, published at CVPR 2023. [19][20]
Vision-language models
Liu was co-first author, with Ligeng Zhu, of NVILA, a family of efficient vision-language models published at CVPR 2025 with 27 authors. [21] He is also a senior author of SparseVILA (ICCV 2025), which applies visual token sparsity to speed up vision-language model inference, and a co-author of VLASH, a method for real-time inference in vision-language-action models. [1][22]
Efficient LLM inference
Liu's recent work as a faculty member targets large language model inference. ParoQuant, by Yesheng Liang, Haisheng Chen, Zihan Zhang, Song Han, and Liu, proposes pairwise rotation quantization for reasoning models and was accepted at ICLR 2026. [1][23] SparseLoRA (ICML 2025) uses contextual sparsity to reduce the computation of fine-tuning. [5]
DFlash, by Jian Chen, Yesheng Liang, and Liu, replaces the autoregressive drafter in speculative decoding with a lightweight block-diffusion model that drafts a block of tokens in one forward pass, conditioned on features from the target model, which then verifies the block in parallel. [6] The paper was posted to arXiv on 5 February 2026 and accepted at ICML 2026; Z Lab reports up to 6 times lossless acceleration. [5][6] NVIDIA's June 2026 blog post reported that, on Blackwell Ultra GPUs running TensorRT-LLM, DFlash delivered up to 15 times the throughput of autoregressive decoding for gpt-oss-120b at the same interactivity level, and nearly doubled interactivity over EAGLE-3 speculative decoding for Llama 3.1 8B. [2] Inco AI's DFlash 2, announced on 18 August 2026, adds a path selector over the drafter's top candidates and a short dynamic convolution to the original design. [7]
Selected publications
| Year | Title | Venue | Liu's role |
|---|---|---|---|
| 2019 | HAQ: Hardware-Aware Automated Quantization with Mixed Precision | CVPR 2019 | Equal-contribution first author [14] |
| 2019 | Point-Voxel CNN for Efficient 3D Deep Learning | NeurIPS 2019 | Co-first author [16] |
| 2022 | TorchSparse: Efficient Point Cloud Inference Engine | MLSys 2022 | Equal-contribution first author [17] |
| 2023 | BEVFusion: Multi-Task Multi-Sensor Fusion with Unified Bird's-Eye View Representation | ICRA 2023 | Co-first author [19] |
| 2023 | FlatFormer: Flattened Window Attention for Efficient Point Cloud Transformer | CVPR 2023 | Co-first author [20] |
| 2025 | NVILA: Efficient Frontier Visual Language Models | CVPR 2025 | Co-first author [21] |
| 2025 | SparseVILA: Decoupling Visual Sparsity for Efficient VLM Inference | ICCV 2025 | Last author [1][22] |
| 2026 | ParoQuant: Pairwise Rotation Quantization for Efficient Reasoning LLM Inference | ICLR 2026 | Last author [1][23] |
| 2026 | DFlash: Block Diffusion for Flash Speculative Decoding | ICML 2026 | Last author [6] |
Awards and recognition
References
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11 ^12Zhijian Liu. Personal homepage. zhijianliu.com (accessed 23 September 2026)
- ^1 ^2 ^3Amr Elmeleegy, Benjamin Chislett, Fernando Xiong, Michael Iovine, Omri Almog, Hao Zhang, and Zhijian Liu. "Boost Inference Performance up to 15x on NVIDIA Blackwell Using DFlash Speculative Decoding." NVIDIA Technical Blog, 23 June 2026. developer.nvidia.com/...flash-speculative-decoding
- ^1 ^2Zhijian Liu (@zhijianliu_). X profile, via the fxtwitter API. x.com/zhijianliu_ (accessed 23 September 2026)
- ^1 ^2NVIDIA Research. "Zhijian Liu." research.nvidia.com/...zhijian-liu (accessed 23 September 2026)
- ^1 ^2 ^3 ^4Z Lab, UC San Diego. Homepage and news. z-lab.ai (accessed 23 September 2026)
- ^1 ^2 ^3 ^4Jian Chen, Yesheng Liang, and Zhijian Liu. "DFlash: Block Diffusion for Flash Speculative Decoding." arXiv:2602.06036, 5 February 2026 (v2 28 May 2026; ICML 2026). arxiv.org/...2602.06036
- ^1 ^2 ^3 ^4 ^5Inco AI. "DFlash 2: Keep Drafting Parallel." Inco AI blog, 18 August 2026. inco.ai/...dflash2
- ^1 ^2 ^3Inco AI. "Splash: A Local Engine Built Around the Model." Inco AI blog, 17 September 2026. inco.ai/...splash
- ^1 ^2 ^3 ^4University of Chicago Data Science Institute. "Zhijian Liu" (Rising Stars in Data Science profile). datascience.uchicago.edu/...zhijian-liu
- ^Zhijian Liu. "Hardware-efficient deep learning for 3D point cloud." S.M. thesis, Massachusetts Institute of Technology, Department of Electrical Engineering and Computer Science, 2020. hdl.handle.net/...127354
- ^1 ^2 ^3 ^4 ^5 ^6MIT EECS. Doctoral thesis defense announcement, "Efficient Deep Learning with Sparsity: Algorithms, Systems, and Applications," 30 April 2024. eecs.mit.edu/...lgorithms-systems-and-applications
- ^1 ^2Zhijian Liu. "Efficient Deep Learning with Sparsity: Algorithms, Systems, and Applications." Doctoral thesis, Massachusetts Institute of Technology, May 2024. DSpace@MIT. hdl.handle.net/...156615
- ^Inco AI. Homepage. inco.ai (accessed 23 September 2026)
- ^1 ^2Kuan Wang, Zhijian Liu, Yujun Lin, Ji Lin, and Song Han. "HAQ: Hardware-Aware Automated Quantization with Mixed Precision." arXiv:1811.08886, CVPR 2019. arxiv.org/...1811.08886
- ^Yihui He, Ji Lin, Zhijian Liu, Hanrui Wang, Li-Jia Li, and Song Han. "AMC: AutoML for Model Compression and Acceleration on Mobile Devices." arXiv:1802.03494, 2018. arxiv.org/...1802.03494
- ^1 ^2Zhijian Liu, Haotian Tang, Yujun Lin, and Song Han. "Point-Voxel CNN for Efficient 3D Deep Learning." arXiv:1907.03739, NeurIPS 2019. arxiv.org/...1907.03739
- ^1 ^2Haotian Tang, Zhijian Liu, Xiuyu Li, Yujun Lin, and Song Han. "TorchSparse: Efficient Point Cloud Inference Engine." arXiv:2204.10319, MLSys 2022. arxiv.org/...2204.10319
- ^Ligeng Zhu, Zhijian Liu, and Song Han. "Deep Leakage from Gradients." arXiv:1906.08935, 2019. arxiv.org/...1906.08935
- ^1 ^2Zhijian Liu, Haotian Tang, Alexander Amini, et al. "BEVFusion: Multi-Task Multi-Sensor Fusion with Unified Bird's-Eye View Representation." arXiv:2205.13542, ICRA 2023. arxiv.org/...2205.13542
- ^1 ^2Zhijian Liu, Xinyu Yang, Haotian Tang, et al. "FlatFormer: Flattened Window Attention for Efficient Point Cloud Transformer." arXiv:2301.08739, CVPR 2023. arxiv.org/...2301.08739
- ^1 ^2Zhijian Liu, Ligeng Zhu, Baifeng Shi, et al. "NVILA: Efficient Frontier Visual Language Models." arXiv:2412.04468, CVPR 2025. arxiv.org/...2412.04468
- ^1 ^2Samir Khaki, Junxian Guo, Jiaming Tang, et al. "SparseVILA: Decoupling Visual Sparsity for Efficient VLM Inference." arXiv:2510.17777, ICCV 2025. arxiv.org/...2510.17777
- ^1 ^2Yesheng Liang, Haisheng Chen, Zihan Zhang, Song Han, and Zhijian Liu. "ParoQuant: Pairwise Rotation Quantization for Efficient Reasoning LLM Inference." arXiv:2511.10645, ICLR 2026. arxiv.org/...2511.10645
- ^Inco AI. "Inco AI Launches Its Inference Platform, Leading Across Four Open Models on Artificial Analysis." Inco AI blog, 3 September 2026 (updated 8 September 2026). inco.ai/...inco-platform-aa
Improve this article
Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.
v1 · 1,673 words · full history
Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify
Research and drafting on this wiki are AI-assisted, under named human editorial standards. How AI is used here
Reviewer note: Independent verification 2026-09-23 (xg05 V12): authorships, roles, dates, NVIDIA byline and awards checked; no defects
Cite this page: AI Wiki. "Zhijian Liu." aiwiki.ai, updated 23 Sept 2026, fact-checked 23 Sept 2026. CC BY 4.0. https://aiwiki.ai/wiki/zhijian_liu