DeepSeek V3
DeepSeek-V3 is a 671-billion-parameter open-weights Mixture of Experts large language model from Chinese AI lab DeepSeek, released on December 26, 2024, that activates only 37 billion parameters per token and…
Explore AI Models through related topics and the articles other pages reference most.
Articles that also belong to these categories. Counts cover all of AI Models.
Showing 1-35 of 35 articles
DeepSeek-V3 is a 671-billion-parameter open-weights Mixture of Experts large language model from Chinese AI lab DeepSeek, released on December 26, 2024, that activates only 37 billion parameters per token and…
DeepSeek V3.1 is a large language model developed by DeepSeek, released on August 19, 2025 and made broadly available via the official API on August 21, 2025.
DeepSeek V4 is a family of open-weight Mixture of Experts large language models developed by DeepSeek, a Hangzhou-based AI research lab.
DeepSeek V4-Pro is the flagship model of the DeepSeek V4 family: a Mixture of Experts large language model with 1.6 trillion total parameters, 49 billion of them activated per token, and a one-million-token…
DeepSeek V4.1-Flash is an open-weight multimodal mixture-of-experts model released by DeepSeek on September 10, 2026. It accepts text and images and generates text.
GLM-4.5 is an open-weights large language model released by Zhipu AI (operating internationally as Z.ai) on July 28, 2025, built on a 355-billion-parameter Mixture of Experts architecture that activates 32…
GLM-4.6 is a flagship open-weight large language model released by Zhipu AI under its international brand Z.ai on September 30, 2025, built on a sparse Mixture of Experts (MoE) architecture with roughly 357…
GLM-5 is an open-weight flagship large language model released by the Chinese AI company Zhipu AI, under its international brand Z.ai, on February 11, 2026.
GLM-5.3-Flash is an open-weight, natively multimodal mixture-of-experts large language model released by Z.ai on August 26, 2026.
Hunyuan is Tencent's brand for its foundation models, an umbrella covering large language models, video generation, image synthesis, 3D asset creation, and multimodal reasoning.
Hy4 Preview is an open-weight large language model released by the Tencent Hy Team on August 28, 2026.
Jamba2 is the second generation of hybrid State Space Model and Transformer language models released by AI21 Labs on January 8, 2026.
Kimi K2 is an open-weights Mixture of Experts language model from Moonshot AI, a Beijing startup, released on July 11, 2025 with 1.04 trillion total parameters and 32.6 billion activated per token
Kimi K2.5 is an open-weights, natively multimodal large language model developed by Moonshot AI and released on January 27, 2026 .
Kimi K3 is an open-weight multimodal reasoning model developed by Moonshot AI. Moonshot made the model available through its hosted products on July 16, 2026 and published the weights, code, configuration…
Leanstral 1.5 is an open-weight model from Mistral AI for formal proof engineering in Lean 4.
Ling-3.0-flash is an open-weight mixture-of-experts language model from inclusionAI, the open-source AI initiative of Ant Group, and the first model of the Ling 3.0 generation.
Ling-3.0-flash-Sante is a health and medicine variant of Ling-3.0-flash, the 124-billion-parameter mixture-of-experts language model from inclusionAI, the open-source AI initiative of Ant Group.
Llama 4 Behemoth is the announced but never publicly released flagship model in the Llama 4 family from Meta AI.
Llama 4 Scout and Llama 4 Maverick are open-weight, natively multimodal AI large language models developed by Meta and released on April 5, 2025.
MAGI-2 Preview is a public research release of a unified audio-video generation model developed by Sand.ai.
MiniMax M2 is an open-weight large language model released on October 27, 2025 by the Shanghai-based AI company MiniMax, built as a Mixture of Experts model with 230 billion total parameters and roughly 10…
Mistral Large 3 is a sparse mixture-of-experts large language model released on December 2, 2025 by the French AI company Mistral AI, distributed as open weights under the Apache 2.0 license with roughly 675…
Mixtral 8x22B is a sparse mixture-of-experts (MoE) large language model released by the French AI company Mistral AI on April 17, 2024.
NVIDIA Nemotron 3.5 Lightning is an open-weights 30 billion parameter mixture-of-experts language model with 3 billion active parameters per token, released by NVIDIA on August 11
Nemotron-Labs-TwoTower is an open-weight diffusion language model released by NVIDIA in mid-2026.
North Mini Code is an open-weight large language model developed by Cohere for agentic software development and AI code generation.
OLMoE (Open Mixture-of-Experts) is a fully open sparse mixture of experts large language model released by the Allen Institute for AI (Ai2) on September 3, 2024 .
Qwen3 is the third-generation family of large language models developed by the Qwen Team at Alibaba Cloud (also known as Tongyi Qianwen lab).
Qwen3-Max is the flagship large language model in Alibaba's Qwen series and the first Qwen model to cross one trillion parameters, released in preview on September 5, 2025 and formally launched at the Apsara…
Qwen3.8-Flash-Next is an experimental open-weight multimodal large language model released by Alibaba Group's Qwen team on August 26, 2026.
Yi-Lightning is a closed-source large language model developed by Chinese artificial intelligence company 01.AI (零一万物, Língyī Wànwù), the company founded by Kai-Fu Lee.
ZAYA1-8B is an open-weight, reasoning-focused Mixture-of-Experts (MoE) large language model released by San Francisco-based AI research lab Zyphra on May 6, 2026.
dots3-note Preview is an open-weight multimodal model and large language model developed by dots studio, an AI team at Xiaohongshu.
gpt-oss is a family of open-weight large language models released by OpenAI on August 5, 2025, and OpenAI's first open-weight language models since GPT-2 in 2019.