Mixture of Experts

Explore Mixture of Experts through related topics and the articles other pages reference most.

Most referenced in this topic

Ranked by links from other AI Wiki pages.

Explore articles

Browse subtopics (25)

Articles that also belong to these categories. Counts cover all of Mixture of Experts.

Showing 1-53 of 53 articles

DeepSeek V3

DeepSeek-V3 is a 671-billion-parameter open-weights Mixture of Experts large language model from Chinese AI lab DeepSeek, released on December 26, 2024, that activates only 37 billion parameters per token and…

AI ModelsChinese AI

DeepSeek V3.1

DeepSeek V3.1 is a large language model developed by DeepSeek, released on August 19, 2025 and made broadly available via the official API on August 21, 2025.

AI ModelsChinese AI

DeepSeek V4

DeepSeek V4 is a family of open-weight Mixture of Experts large language models developed by DeepSeek, a Hangzhou-based AI research lab.

AI ModelsChinese AI

DeepSeek V4-Pro

DeepSeek V4-Pro is the flagship model of the DeepSeek V4 family: a Mixture of Experts large language model with 1.6 trillion total parameters, 49 billion of them activated per token, and a one-million-token…

AI ModelsChinese AI

DeepSeek V4.1-Flash

DeepSeek V4.1-Flash is an open-weight multimodal mixture-of-experts model released by DeepSeek on September 10, 2026. It accepts text and images and generates text.

AI ModelsChinese AI

DeepSeek-V2

DeepSeek-V2 is a 236-billion-parameter mixture-of-experts (MoE) large language model released in May 2024 by DeepSeek, the Chinese AI lab spun out of the quantitative hedge fund High-Flyer and led by Liang…

Chinese AILarge Language Models

DeepSeek-VL2

DeepSeek-VL2 is an open-weights family of Mixture-of-Experts (MoE) vision-language models released by the Chinese AI laboratory DeepSeek on December 13, 2024.

Chinese AIMultimodal AI

GLM-4.5

GLM-4.5 is an open-weights large language model released by Zhipu AI (operating internationally as Z.ai) on July 28, 2025, built on a 355-billion-parameter Mixture of Experts architecture that activates 32…

AI ModelsChinese AI

GLM-4.6

GLM-4.6 is a flagship open-weight large language model released by Zhipu AI under its international brand Z.ai on September 30, 2025, built on a sparse Mixture of Experts (MoE) architecture with roughly 357…

AI ModelsChinese AI

GLM-5

GLM-5 is an open-weight flagship large language model released by the Chinese AI company Zhipu AI, under its international brand Z.ai, on February 11, 2026.

AI ModelsChinese AI

GLM-5.3-Flash

GLM-5.3-Flash is an open-weight, natively multimodal mixture-of-experts large language model released by Z.ai on August 26, 2026.

AI ModelsChinese AI

Hunyuan

Hunyuan is Tencent's brand for its foundation models, an umbrella covering large language models, video generation, image synthesis, 3D asset creation, and multimodal reasoning.

AI ModelsChinese AI

Jamba

Jamba is a family of open-weight large language models from AI21 Labs, first released on March 28, 2024, and is the world's first production-grade language model built on a Mamba state space model (SSM)…

AI CompaniesLarge Language Models

Kimi K2

Kimi K2 is an open-weights Mixture of Experts language model from Moonshot AI, a Beijing startup, released on July 11, 2025 with 1.04 trillion total parameters and 32.6 billion activated per token

AI ModelsChinese AI

Kimi K2.5

Kimi K2.5 is an open-weights, natively multimodal large language model developed by Moonshot AI and released on January 27, 2026 .

AI ModelsChinese AI

Kimi K3

Kimi K3 is an open-weight multimodal reasoning model developed by Moonshot AI. Moonshot made the model available through its hosted products on July 16, 2026 and published the weights, code, configuration…

AI ModelsChinese AI

Ling-3.0-flash

Ling-3.0-flash is an open-weight mixture-of-experts language model from inclusionAI, the open-source AI initiative of Ant Group, and the first model of the Ling 3.0 generation.

AI ModelsChinese AI

Ling-3.0-flash-Sante

Ling-3.0-flash-Sante is a health and medicine variant of Ling-3.0-flash, the 124-billion-parameter mixture-of-experts language model from inclusionAI, the open-source AI initiative of Ant Group.

AI ModelsChinese AI

MiniMax M2

MiniMax M2 is an open-weight large language model released on October 27, 2025 by the Shanghai-based AI company MiniMax, built as a Mixture of Experts model with 230 billion total parameters and roughly 10…

AI AgentsAI Models

Mistral Large 3

Mistral Large 3 is a sparse mixture-of-experts large language model released on December 2, 2025 by the French AI company Mistral AI, distributed as open weights under the Apache 2.0 license with roughly 675…

AI CompaniesAI Models

Mixtral

Mixtral is a family of open-weight Sparse Mixture of Experts (SMoE) large language models developed by Mistral AI, a French artificial intelligence company founded in April 2023.

AI CompaniesLarge Language Models

MoonEP

MoonEP is an open-source expert-parallel communication library for mixture-of-experts models, released by Moonshot AI on July 27, 2026 under the MIT License.

AI InfrastructureChinese AI

OLMoE

OLMoE (Open Mixture-of-Experts) is a fully open sparse mixture of experts large language model released by the Allen Institute for AI (Ai2) on September 3, 2024 .

AI ModelsLarge Language Models

Qwen3

Qwen3 is the third-generation family of large language models developed by the Qwen Team at Alibaba Cloud (also known as Tongyi Qianwen lab).

AI ModelsChinese AI

Qwen3-Max

Qwen3-Max is the flagship large language model in Alibaba's Qwen series and the first Qwen model to cross one trillion parameters, released in preview on September 5, 2025 and formally launched at the Apsara…

AI ModelsChinese AI

Switch Transformer

The Switch Transformer is a sparsely activated Mixture of Experts (MoE) Transformer architecture introduced by William Fedus, Barret Zoph, and Noam Shazeer at Google in January 2021.

GoogleLarge Language Models

Yi-Lightning

Yi-Lightning is a closed-source large language model developed by Chinese artificial intelligence company 01.AI (零一万物, Língyī Wànwù), the company founded by Kai-Fu Lee.

AI ModelsChinese AI

ZAYA1-8B

ZAYA1-8B is an open-weight, reasoning-focused Mixture-of-Experts (MoE) large language model released by San Francisco-based AI research lab Zyphra on May 6, 2026.

AI ModelsLarge Language Models

gpt-oss

gpt-oss is a family of open-weight large language models released by OpenAI on August 5, 2025, and OpenAI's first open-weight language models since GPT-2 in 2019.

AI ModelsLarge Language Models