Kimi K3
Last edited
Fact-checked
Sources
25 citations
Revision
v1 · 3,837 words
Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify
Kimi K3 is a 2.8-trillion-parameter Mixture of Experts reasoning model developed by Moonshot AI and announced on July 16, 2026 (July 17 in Beijing) [3][4][7]. It is the flagship successor to the company's Kimi K2 line, and Moonshot describes it as the world's first open model in the 3-trillion-parameter class [1][3]. The model activates 16 of 896 experts per token, accepts images and video natively, keeps a reasoning mode permanently switched on, and runs a context window of 1,048,576 tokens [1]. Two architectural changes distinguish it from its predecessors: Kimi Delta Attention, a hybrid linear attention mechanism that cuts the KV cache by up to 75 percent at long context, and Attention Residuals, a replacement for standard residual connections that lets deeper layers retrieve representations from earlier ones selectively [4][15][16].
On Moonshot's own benchmark table, K3 beats Claude Opus 4.8 and GPT-5.5 on most published coding and agentic evaluations while trailing Claude Fable 5 and GPT-5.6 Sol overall, a gap the company acknowledged plainly in its launch material [5][7][13]. At launch the model was available only through Moonshot's API and the Kimi app; the company promised full open weights on Hugging Face "by July 27, 2026," and as of July 23 those weights had not yet been published [1][3][12]. Demand after release was heavy enough that Moonshot paused new consumer subscriptions on July 19, three days after launch, citing what Reuters reported as "unprecedented compute challenges" [8][9]. The release also rattled markets, with coverage in Fortune and elsewhere framing it as a second DeepSeek shock [21].
Background and release
Moonshot AI, founded in Beijing in 2023 by Yang Zhilin, spent 2025 and the first half of 2026 shipping open-weight models at a fast clip: Kimi K2 in July 2025, Kimi K2 Thinking in November 2025, Kimi K2.5 in January 2026, Kimi K2.6 in April, and the coding-focused Kimi K2.7 Code in June [23]. Moonshot's launch documentation for K3 claims that Kimi models held the open-source performance frontier "in 9 of the past 12 months (2025/07-2026/07)" [1]. By mid-2026 the company had, per Reuters, raised more than 2 billion dollars in a May funding round at a valuation that reached 30 billion dollars in June, and was reported to be preparing a Hong Kong IPO with Goldman Sachs and China International Capital Corporation advising [8].
K3 was announced on Thursday, July 16, 2026, United States time; because the unveiling came late in the day, outlets in Asia dated it July 17 [3][6][7]. The model went live immediately in the Kimi app at kimi.com, in the Kimi Code product, and through Moonshot's API [1][13]. The weights themselves were not part of the launch: Moonshot committed to releasing them within eleven days, by July 27 [1][3].
The scale claim drew most of the initial headlines. At 2.8 trillion total parameters, K3 was described across the press as the largest open-weight model ever announced, well past DeepSeek V4 Pro at 1.6 trillion parameters and Zhipu AI's GLM-5 series at 744 billion [7][20]. Moonshot president Yutong Zhang, speaking at the World Economic Forum earlier in the year, had framed the company's design philosophy around constraint rather than abundance: "We knew we didn't have the luxury to simply scale up compute. That forced us to focus on fundamental research and efficiency" [6].
Architecture
K3 is a sparse Mixture of Experts transformer. Moonshot calls the expert layer design Stable LatentMoE: each MoE layer holds 896 experts and the router activates 16 of them per token, a sparsity of roughly 98 percent [1][13][15]. The company has not published an active-parameter count. Outside analyses note that a naive 16/896 fraction of 2.8 trillion works out to about 50 billion active parameters per token, and independent estimates cluster in the 50 to 60 billion range, but the exact figure cannot be derived from the official specification because Moonshot has not disclosed the split between always-active components and routed experts [15][18]. Routing uses a scheme Moonshot calls Quantile Balancing, in which a token goes to an expert when its router score lands in the top quantile, which the company says keeps expert utilization even and avoids dead experts [15].
The two headline architectural changes are Kimi Delta Attention and Attention Residuals, and Moonshot credits the pair, together with revised training and data recipes, with roughly 2.5 times better scaling efficiency than K2: more capability per unit of training compute [4][10][15].
Kimi Delta Attention
Kimi Delta Attention (KDA) is a hybrid linear attention mechanism that Moonshot first published in its Kimi Linear research paper in October 2025, where it was validated on a 48-billion-parameter model [16]. Standard transformer attention compares every new token against a cache of keys and values for all previous tokens, so memory and decode cost grow with sequence length. Linear attention replaces that growing cache with a fixed-size recurrent state that each token updates and reads, which keeps memory flat but has historically lost accuracy on recall-heavy tasks. KDA builds on the gated DeltaNet family of linear attention methods and refines the gating: instead of one scalar decay applied uniformly across the whole state, KDA uses a per-channel diagonal decay, so each dimension of the memory can forget at its own rate. Information that matters over long ranges can persist while local detail gets overwritten [15][16].
Because pure linear attention still struggles with precise long-range retrieval, K3 interleaves KDA layers with conventional full-attention layers at a 3:1 ratio, three KDA layers for every full-attention layer, the ratio the Kimi Linear paper found gave the best quality-efficiency balance [16][22]. The payoff shows up at long context: Moonshot's reported figures are a KV cache reduction of up to 75 percent and decoding up to 6.3 times faster at million-token context lengths compared with full attention [4][15]; the Kimi Linear paper itself claims up to 6 times decoding throughput at a 1M context [16].
Attention Residuals
Attention Residuals (AttnRes) rework the residual stream that connects a transformer's layers. In a standard transformer, each layer's output is added onto an accumulating residual, so every layer receives the uniform sum of everything before it. With AttnRes, each layer instead attends selectively over the representations of earlier layers with learned weights, treating depth as a retrieval problem: a layer pulls the earlier representations it needs rather than inheriting an undifferentiated running total [4][15]. Moonshot describes it as a drop-in replacement for residual connections and reports that it delivers roughly 25 percent higher training efficiency for under 2 percent additional cost [4][5]. K3 uses a blocked variant, Block AttnRes, that reduces the extra computation; in ablations on the 48B Kimi Linear model, the technique produced a 7.5-point jump on GPQA-Diamond and matched a baseline trained with 1.25 times more compute [15].
Numerics and hardware footprint
K3 ships with MXFP4 weights and MXFP8 activations, a microscaling floating-point format in which each 4-bit weight carries per-block scaling factors [17][18]. Even so compressed, the checkpoint is estimated at roughly 1.4 terabytes, which puts self-hosting well beyond a single server: analysts estimate a practical minimum of around 64 accelerators in a rack-scale configuration with more than 1.5 terabytes of pooled high-bandwidth memory [12][17][22].
Multimodality and context window
K3 is natively multimodal on the input side. It accepts images, passed either base64-encoded or by ms:// file ID (public image URLs are not supported), and video, uploaded through Moonshot's Files API and referenced the same way [1]. Moonshot positions the visual capability less as a chat feature and more as part of the agentic loop: the model card language emphasizes iterating against images, logs, tests, and runtime feedback while working in large repositories [14]. Simon Willison, testing the vision path at launch, called the alt text it generated for his test image "very good" [3].
The context window is 1,048,576 tokens, and unusually, the output ceiling matches it: max_completion_tokens defaults to 131,072 and can be raised to the full 1,048,576 [1][2]. Moonshot charges the same rates at any context length, with no tiering by input size [1][2].
Reasoning modes
K3 has no non-reasoning mode. Thinking is always on, and developers control depth rather than presence through a reasoning_effort parameter documented at three levels: low, high, and max, with max the default [1]. At launch only max was actually live; early reviews noted a single thinking-effort level, with the lower settings appearing in the documentation in the days that followed [3][11]. This is a departure from the K2 generation, where instant and thinking variants were separate modes with separate sampling defaults. The always-on design fits the model's positioning for long-horizon agentic work, where Moonshot recommends it for complex coding, knowledge work, and multi-step tool use [14]. The cost of the approach is token volume: Willison's single SVG-drawing test consumed 13,241 reasoning tokens and cost about 25 cents [3].
Benchmarks and performance
Moonshot published a benchmark table at launch comparing K3 against Claude Fable 5, GPT-5.6 Sol, Claude Opus 4.8, and GPT-5.5, and its own framing was unusually direct for a launch: K3 wins narrowly on a set of coding and agentic evaluations while trailing Fable 5 and GPT-5.6 Sol on overall capability, and the model card concedes a noticeable user-experience gap against those two systems [5][7][10][13]. VentureBeat summarized the coding picture as K3 placing in the top three across all six coding benchmarks Moonshot tested, leading outright on SWE Marathon and Program Bench and trailing only GPT-5.6 Sol on Terminal-Bench 2.1, by half a point [5]. On the two hardest software-engineering suites in the table the picture splits: on DeepSWE, K3's 67.5 trails both GPT-5.6 Sol (73.0) and Fable 5 (70.0), while on FrontierSWE, K3's 81.2 trails only Fable 5 (86.6) and comfortably beats GPT-5.6 Sol (71.3) [13][24]. South China Morning Post reporting added that K3 consistently outperformed GPT-5.5, Claude Opus 4.8, and Zhipu AI's GLM-5.2 across the published evaluations [7].
The table below lists scores that appear consistently across multiple independent write-ups of Moonshot's launch materials. All figures are Moonshot's own reported results unless noted; independent replication was not possible before the weights release.
| Benchmark | Task type | Kimi K3 | Comparison (as reported by Moonshot) |
|---|---|---|---|
| Program Bench | Coding | 77.8 | GPT-5.6 Sol 77.6, Claude Fable 5 76.8 [4][24] |
| SWE Marathon | Long-horizon software engineering | 42.0 | Claude Opus 4.8 40.0, GPT-5.6 Sol 39.0, Fable 5 35.0 [4][24] |
| Terminal-Bench 2.1 | Agentic terminal use | 88.3 | GPT-5.6 Sol 88.8 (K3 second) [5][11][24] |
| BrowseComp | Agentic web browsing | 91.2 | GPT-5.6 Sol 90.4 [4][24] |
| OmniDocBench | Document understanding | 91.1 | Fable 5 89.8 [4][24] |
| GPQA Diamond | Graduate-level science | 93.5 | GPT-5.6 Sol 94.1 (K3 second), Fable 5 92.6 [4][24] |
| Humanity's Last Exam | Frontier reasoning | 56.0 | llm-stats leaderboard (standard variant): rank 8 overall, highest open-model entry; Fable 5 leads at 64.5 [19] |
Moonshot also reported narrow wins on Automation Bench (30.8 against GPT-5.6 Sol's 29.7) and SpreadsheetBench 2 (34.8 against Fable 5's 34.7), while its own full-set Humanity's Last Exam comparison has K3 at 43.5, behind Fable 5 (53.3), Opus 4.8 (49.8), and GPT-5.6 Sol (44.5) [24]. Notably absent from the launch framing were older suites such as SWE-bench Verified; the company's coding story leaned on newer long-horizon agentic evaluations instead [5][11].
Early independent signals were consistent with the self-reported positioning. Artificial Analysis measured K3 at 57 on its Intelligence Index, third among the models it tracked, behind only Claude Fable 5 and GPT-5.6 Sol, with throughput of 62 output tokens per second and 1.99 seconds to first token on the official API [10][12][13]. On the Vals AI index the model ranked second of 38 models at 74.7 percent [10][12]. Its most eye-catching independent result came from Arena.ai's Frontend Code Arena, where K3 debuted at number one with a preliminary Elo of 1,679, ahead of Claude Fable 5, a result Tom's Hardware put in its headline [12][20][24]. On LMArena's text leaderboard its early Elo was 1,486, still carrying a preliminary confidence interval of about 11 points [24].
API access and pricing
K3 is served through Moonshot's OpenAI-compatible API at https://api.moonshot.ai/v1 under the model ID kimi-k3, so code written against the OpenAI SDK can switch by changing the base URL and key [1][4]. Structured output with strict JSON schema is supported [11]. Context caching is automatic: once a prior request exceeds 256 prompt tokens, repeated prefixes bill at the cache-hit rate with no cache IDs or TTL management required [1]. Pricing is flat with respect to context length [1][2].
Moonshot argues the effective input cost is much closer to the cache-hit price in practice, claiming above 90 percent cache-hit rates in coding workloads [4]. Even so, the sticker price marked a break with the aggressive discounting Chinese labs had been known for. K3 costs about 3.2 times as much as K2.6 for fresh input and 3.75 times as much for output (K2.6 lists at $0.95 in, $4.00 out) [3][13], and Willison called it "the most expensive model released by a Chinese AI lab to date," noting that $3/$15 exactly matches Anthropic's Claude Sonnet tier [3]. It remains far below the top closed models; Fortune noted Anthropic charges $50 per million output tokens for Fable 5, against K3's $15 [6]. K3 is also listed on OpenRouter at the same $3/$15 passthrough rates, which in the first week carried a warning that upstream capacity was limited and requests might hit frequent 429 errors [14].
Open weights plan and license
At launch K3 was, strictly speaking, a promise of an open model rather than an open model. Moonshot's documentation states that full weights will be released "by July 27, 2026" [1][3], and as of July 23, 2026 no K3 repository had appeared under the moonshotai organization on Hugging Face; Artificial Analysis was still classifying the model as proprietary pending the release [12][13][17]. The open weights claim in the launch marketing, including the "first open 3T-class model" line, therefore rests on that commitment [3].
The license is likewise unpublished. Every recent model in the K2 line, including K2, K2.5, K2.6, and K2.7 Code, shipped under Moonshot's Modified MIT License, which is standard MIT plus a single attribution clause: commercial products built on the model that exceed 100 million monthly active users or 20 million dollars in monthly revenue must display the model name prominently in their user interface. Reporting ahead of the K3 weights release widely expects the same license, but Moonshot has not confirmed it for K3, and the exact text will only be knowable when the weights drop [17][25].
Demand and the subscription pause
Whatever the caveats, demand outran Moonshot's infrastructure almost immediately. On Sunday, July 19, three days after launch, the company said it was pausing new consumer subscriptions. "Kimi K3 has received far more love than we expected," the statement said. "Over the past 48 hours, demand has pushed close to the limits of our current capacity," and the company promised, "We're adding capacity as fast as we can and will reopen new subscription spots in batches" [9]. Reuters, reporting the pause on July 20, wrote that the release had drawn massive user interest leading to "unprecedented compute challenges," with user requests over 48 hours sharply exceeding forecasts and approaching the limits of existing clusters [8].
Existing subscribers kept access, and the API and OpenRouter listings stayed live through the pause [8][14]. Moonshot also said it would restructure its consumer offering into two membership plans, one of them dedicated to coding, to match compute allocation to how people actually use the model [8]. The episode became part of the story itself: a lab that had just shipped the largest model of the open-weight era could not find enough GPUs to sell access to it.
Reception and analysis
Written assessments in the first week converged on a few points. Nathan Lambert of Interconnects called K3 the strongest open model ever released and read it as evidence that leading Chinese labs are doing independent frontier research rather than riding distillation, while arguing that open-weight releases of this class are economically corrosive for closed frontier labs even as they accelerate AI diffusion through the broader economy [10]. Simon Willison was more wry, using his pelican SVG benchmark to note that a single drawing consumed 13,241 reasoning tokens, and cautioning that toy benchmarks say nothing about the agentic tool calling the model is actually built for [3]. Fortune's coverage argued that K3 pushed Chinese AI into "Fable-level territory" months before analysts expected; the magazine reported that observers, Anthropic CEO Dario Amodei among them, had not anticipated a Chinese model at this level until roughly half a year later [6][21].
Markets treated the release as a rerun of the January 2025 DeepSeek moment. On Friday, July 17, TSMC fell 7 percent despite reporting a 77 percent jump in quarterly operating profit, SoftBank dropped 9 percent, Nvidia slipped 1.2 percent, and Zhipu AI (Z.ai) fell nearly 30 percent in Hong Kong trading as investors repriced the domestic competition [21]. Paul Triolo of DGA-Albright Stonebridge Group told Fortune that "the AI ecosystem in China is probably much better than people thought" [21].
The counter-analysis came from the hardware side. SemiAnalysis argued that K3's linear attention should not be read as bearish for accelerator demand: the KV-cache savings cut per-token data movement by roughly an order of magnitude, but serving a 2.8-trillion-parameter model still requires 64 or more chips with over 1.5 terabytes of pooled HBM, and distributing 896 experts across GPUs adds weight-exchange traffic that claws back some of the bandwidth savings. On that reading, K3's efficiency gains are the classic Jevons paradox: cheaper intelligence per token drives more total consumption, not less [22]. The subscription pause days later made the same point empirically.
A separate thread of commentary focused on what K3 says about compute constraints. Under United States export controls, Moonshot could not simply match American training clusters, and both the architecture (attention that trades compute-heavy quadratic attention for memory-efficient recurrence) and president Yutong Zhang's own framing present K3 as engineering around that ceiling [6][22]. Whether the efficiency-first path keeps pace with the largest closed training runs is the open question the weights release, and the independent evaluations it enables, will start to answer.
Place in the Kimi model family
K3 caps a two-year progression in which Moonshot moved from closed consumer chat models to the largest open-weight systems available.
| Model | Release | Scale | Notes |
|---|---|---|---|
| Kimi K1.5 | January 2025 | undisclosed | Closed-weights multimodal reasoning model; detailed tech report, no weights |
| Kimi K2 | July 2025 | 1T total / 32B active | First open-weights release; text-only, agentic focus, Modified MIT |
| Kimi K2 Thinking | November 2025 | 1T / 32B | Reasoning variant; native INT4 via quantization-aware training |
| Kimi K2.5 | January 2026 | 1T / 32B | Native vision via MoonViT encoder; Agent Swarm; 256K context |
| Kimi K2.6 | April 20, 2026 | 1T / 32B | Agentic coding and larger swarms; leading open model at launch |
| Kimi K2.7 Code | June 12, 2026 | 1T-class | Coding-specialized K2.6 derivative; about 30 percent lower reasoning-token use [23] |
| Kimi K3 | July 16, 2026 | 2.8T / ~50B active (est.) | KDA + AttnRes architecture, 1M context, native image and video input; weights promised by July 27, 2026 |
The jump from K2.6 to K3 is architectural as much as it is scale. The K2 generation used Multi-head Latent Attention inherited from the DeepSeek lineage and a 384-expert MoE; K3 replaces the attention stack with the KDA hybrid validated in the Kimi Linear research line, more than doubles the expert count to 896, roughly triples total parameters, and quadruples the context window to a million tokens [1][15][16]. It is also the first Kimi flagship priced like a Western frontier model rather than undercutting one by an order of magnitude [3][13]. The Kimi assistant and Kimi Code products carried K3 from the day of launch [1][4].
References
- Moonshot AI. "Kimi K3 Quickstart." Kimi API Platform documentation. https://platform.kimi.ai/docs/guide/kimi-k3-quickstart ↩
- Moonshot AI. "Kimi K3 Pricing." Kimi API Platform documentation. https://platform.kimi.ai/docs/pricing/chat-k3 ↩
- Willison, Simon. "Kimi K3, and what we can still learn from the pelican benchmark." simonwillison.net, July 16, 2026. https://simonwillison.net/2026/Jul/16/kimi-k3/ ↩
- MarkTechPost. "Moonshot AI Releases Kimi K3: A 2.8 Trillion Parameter Open MoE Model With Kimi Delta Attention and 1M Context." July 16, 2026. https://www.marktechpost.com/2026/07/16/moonshot-ai-releases-kimi-k3-a-2-8-trillion-parameter-open-moe-model-with-kimi-delta-attention-and-1m-context/ ↩
- VentureBeat. "China's Moonshot AI releases Kimi K3, the largest open-source model ever, rivaling top U.S. systems." July 2026. https://venturebeat.com/technology/chinas-moonshot-ai-releases-kimi-k3-the-largest-open-source-model-ever-rivaling-top-u-s-systems ↩
- Fortune. "Moonshot's Kimi K3 pushes Chinese AI into Fable-level territory." July 16, 2026. https://fortune.com/2026/07/16/moonshots-kimi-k3-pushes-chinese-ai-into-fable-level-territory/ ↩
- South China Morning Post. "Moonshot AI unveils world's largest open-source AI model as China narrows gap with US rivals." July 17, 2026. https://www.scmp.com/tech/tech-trends/article/3360885/moonshot-ai-unveils-worlds-largest-open-source-ai-model-china-narrows-gap-us-rivals ↩
- Reuters (via Yahoo Finance). "China's Moonshot pauses Kimi subscriptions amid hot demand, IPO push." July 20, 2026. https://finance.yahoo.com/technology/ai/articles/chinas-moonshot-pauses-kimi-subscriptions-080317250.html ↩
- Associated Press (via ABC News). "China's new AI model halts new subscriptions as demand swamps capacity." July 20, 2026. https://abcnews.com/International/wireStory/chinas-new-ai-model-halts-new-subscriptions-demand-134909818 ↩
- Lambert, Nathan. "Kimi K3: The open-weights escalation." Interconnects, July 2026. https://www.interconnects.ai/p/kimi-k3-the-open-weights-escalation ↩
- DataCamp. "Kimi K3: Moonshot AI's Newest and Best Open-Source Model." July 2026. https://www.datacamp.com/blog/kimi-k3 ↩
- Northflank. "Kimi K3: benchmarks, pricing, hardware requirements, and self-hosting." July 2026. https://northflank.com/blog/what-is-kimi-k3-self-hosting ↩
- Trilogy AI. "Kimi K3 Is Live: Pricing, Benchmarks, and the Wait for Public Weights." July 2026. https://trilogyai.substack.com/p/kimi-k3-is-live-pricing-benchmarks ↩
- OpenRouter. "Kimi K3 - API Pricing & Benchmarks." https://openrouter.ai/moonshotai/kimi-k3 ↩
- Huang, Ken. "Demystifying Kimi K3: The Three Algorithms Behind the #1 Frontend Coding Model." July 2026. https://kenhuangus.substack.com/p/demystifying-kimi-k3-how-chinas-28t ↩
- Moonshot AI (Kimi Team). "Kimi Linear: An Expressive, Efficient Attention Architecture." arXiv:2510.26692, October 30, 2025. https://arxiv.org/abs/2510.26692 ↩
- Hugging Face community blog. "Kimi K3 Model Overview: 2.8T Parameters, MXFP4 Quantization, and What the Open Weights Mean for the Community." July 17, 2026. https://huggingface.co/blog/ResterChed/kimi-k3-model-overview-mxfp4-quantization-open-wei ↩
- Photon Capital. "Kimi K3's Active Set Is 50B-Class. Its Weights Are 2.8T." July 2026. https://photoncap.net/p/kimi-k3s-active-set-is-50b-class ↩
- llm-stats. "Humanity's Last Exam Leaderboard." Retrieved July 23, 2026. https://llm-stats.com/benchmarks/humanity's-last-exam ↩
- Tom's Hardware. "China's 2.8-trillion-parameter Kimi K3 beats Claude Fable 5 in Frontend Code Arena benchmark. Moonshot AI delivers largest open-weight AI model ever, as China works around U.S. compute limits." July 2026. https://www.tomshardware.com/tech-industry/artificial-intelligence/moonshot-releases-2-8-trillion-parameter-kimi-k3 ↩
- Fortune. "Markets experience new DeepSeek shock after Moonshot AI releases Kimi K3." July 17, 2026. https://fortune.com/2026/07/17/china-moonshot-kimi-k3-markets-china-ai/ ↩
- BigGo Finance. "SemiAnalysis: Kimi K3's Linear Attention Won't Weaken Nvidia Demand. It Actually Requires More GPUs and HBM." July 19, 2026. https://finance.biggo.com/news/0ef00b0f-75ac-42ae-994b-b4dc4df325ed ↩
- MarkTechPost. "Moonshot AI Releases Kimi K2.7-Code: a Coding Model Reporting +21.8% on Kimi Code Bench v2 Over K2.6." June 12, 2026. https://www.marktechpost.com/2026/06/12/moonshot-ai-releases-kimi-k2-7-code-a-coding-model-reporting-21-8-on-kimi-code-bench-v2-over-k2-6/ ↩
- Wan 2.7 blog. "Kimi K3 Benchmarks: Every Score, Every Comparison, Every Surprise (July 2026)." https://wan27.org/blog/kimi-k3-benchmarks ↩
- Wan 2.7 blog. "Is Kimi K3 Open Source? License, Weights, GitHub, and What You Can Actually Use Today (2026)." https://wan27.org/blog/kimi-k3-open-source ↩
Improve this article
Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.