LLaMA/Model Card
The Llama model card is the official documentation that Meta AI ships with each release of the Llama family of large language models.
Explore AI Models through related topics and the articles other pages reference most.
Articles that also belong to these categories. Counts cover all of AI Models.
Showing 121-180 of 184 articles
The Llama model card is the official documentation that Meta AI ships with each release of the Llama family of large language models.
Leanstral 1.5 is an open-weight model from Mistral AI for formal proof engineering in Lean 4.
Ling-3.0-flash is an open-weight mixture-of-experts language model from inclusionAI, the open-source AI initiative of Ant Group, and the first model of the Ling 3.0 generation.
Llama 3 is a family of open-weight large language models developed by Meta. Meta released the original Llama 3 checkpoints on April 18, 2024, in 8-billion-parameter and 70-billion-parameter sizes.
Llama 3.1 is a family of open-weight large language models released by Meta on July 23, 2024, in three sizes, 8 billion, 70 billion, and 405 billion parameters, each shipped in both a pre-trained base form and…
Llama 3.2 is a family of four open-weight large language models released by Meta on September 25, 2024, comprising lightweight 1 billion and 3 billion parameter text-only models for on-device AI and the 11…
Llama 3.3 is an instruction-tuned, text-only large language model with 70 billion parameters that Meta released on December 6, 2024
Llama 4 Behemoth is the announced but never publicly released flagship model in the Llama 4 family from Meta AI.
Llama 4 Scout and Llama 4 Maverick are open-weight, natively multimodal AI large language models developed by Meta and released on April 5, 2025.
Llama-3.1-Nemotron-70B-Instruct is a large language model released by NVIDIA in October 2024.
LongCat-Flash is an open-weight large language model developed by the LongCat team at Meituan, the Chinese on-demand local-services and food-delivery company.
MAI-1-preview is a large language model developed by Microsoft AI, the consumer artificial intelligence division of Microsoft led by Mustafa Suleyman.
MiniCPM is a family of compact, openly licensed language models published by OpenBMB, the shared open-source brand of Tsinghua University's natural language processing lab (THUNLP) and the Beijing company…
MiniCPM5-2B is an open-weights small language model published by OpenBMB on September 7, 2026 under the Apache License 2.0.
MiniMax M2 is an open-weight large language model released on October 27, 2025 by the Shanghai-based AI company MiniMax, built as a Mixture of Experts model with 230 billion total parameters and roughly 10…
Ministral is a family of two small language models released by the French artificial intelligence company Mistral AI on October 16, 2024.
Mistral Large is the family of flagship large language models developed by Mistral AI, the Paris-based AI laboratory founded in 2023, and is the company's most capable general-purpose model line.
Mistral Large 3 is a sparse mixture-of-experts large language model released on December 2, 2025 by the French AI company Mistral AI, distributed as open weights under the Apache 2.0 license with roughly 675…
Mistral Medium 3 is a proprietary multimodal large language model developed by Mistral AI and released on May 7, 2025.
Mixtral 8x22B is a sparse mixture-of-experts (MoE) large language model released by the French AI company Mistral AI on April 17, 2024.
Muse Glimmer is an open-weight text-and-image model developed by Meta AI for local agent and coding workloads. Meta released the model on August 10, 2026 under the identifier meta-models/Muse-Glimmer-30B.
Nemotron is NVIDIA's brand for its family of open large language models and the datasets, training recipes, and evaluation tools built around them.
Nemotron 3 is a family of open-weights large language model systems released by NVIDIA beginning on December 15, 2025, built for agentic AI and consisting of three sparse mixture-of-experts variants named…
NVIDIA Nemotron 3.5 Lightning is an open-weights 30 billion parameter mixture-of-experts language model with 3 billion active parameters per token, released by NVIDIA on August 11
Nemotron Nano 2 is a family of small, open-weight reasoning language models released by NVIDIA on August 18, 2025
Nemotron-4 is a family of decoder-only large language models developed by NVIDIA and documented in two technical reports released in 2024.
Nemotron-H is a family of open-weight large language models released by NVIDIA in April 2025 that replace most of the self-attention layers of a standard Transformer with Mamba-2 state-space layers, producing…
Nemotron-Labs-TwoTower is an open-weight diffusion language model released by NVIDIA in mid-2026.
North Mini Code is an open-weight large language model developed by Cohere for agentic software development and AI code generation.
OLMo (Open Language Model) is a family of fully open large language models built by the Allen Institute for AI (Ai2) and first released on February 1, 2024.
OLMo 2 is the second generation of fully open large language models released by the Allen Institute for AI (Ai2), spanning 7B, 13B, and 32B parameter sizes.
OLMo 3 is the third generation of fully open language models released by the Allen Institute for AI (Ai2).
OLMoE (Open Mixture-of-Experts) is a fully open sparse mixture of experts large language model released by the Allen Institute for AI (Ai2) on September 3, 2024 .
OpenAI o1 is a family of proprietary large language models developed by OpenAI and trained to use additional computation before returning an answer.
OpenAI o3 is a family of reasoning-focused large language models developed by OpenAI and the second generation of the company's o-series reasoning models, best known for scoring 87.5% on the ARC-AGI…
Ox Alpha was the anonymous preview alias for GLM-5.3-Flash, a natively multimodal large language model developed by Z.ai.
Phi-4 is a 14-billion-parameter small language model developed by Microsoft Research and released in December 2024, designed to match or beat models several times its size on reasoning tasks by training…
Phi-4-reasoning is a 14 billion parameter open weight reasoning model released by Microsoft Research on April 30, 2025.
Phi-4-mini is a 3.8 billion parameter open weight small language model released by Microsoft on February 26, 2025, under the permissive MIT license.
Phi-4-mini-flash-reasoning is a 3.8 billion parameter open weight reasoning model released by Microsoft in July 2025.
Pixtral is a family of multimodal vision-language models developed by Mistral AI, a French AI company founded in April 2023.
Pixtral Large is a 124-billion-parameter multimodal (vision-language) large language model released by Mistral AI on November 18, 2024.
Qwen3 is the third-generation family of large language models developed by the Qwen Team at Alibaba Cloud (also known as Tongyi Qianwen lab).
Qwen3-Max is the flagship large language model in Alibaba's Qwen series and the first Qwen model to cross one trillion parameters, released in preview on September 5, 2025 and formally launched at the Apsara…
Qwen3-Next is an efficiency-focused large language model and model architecture released in September 2025 by the Qwen team at Alibaba Cloud.
Qwen3.8 is the name Qwen uses for a model generation that includes hosted services and downloadable checkpoints.
Qwen3.8-Flash-Next is an experimental open-weight multimodal large language model released by Alibaba Group's Qwen team on August 26, 2026.
Reka Core is a frontier class multimodal foundation model developed by Reka AI, a research and product company founded in 2022 by former scientists from DeepMind, Google Brain, Meta FAIR, and Baidu.
Reka Edge is a 7-billion-parameter multimodal language model developed by Reka AI, introduced in April 2024 as the smallest member of the company's first publicly described model family.
Reka Flash is a family of multimodal large language models developed by Reka AI, a San Francisco Bay Area research company founded in 2022 by former researchers from Google DeepMind, Meta FAIR, and Google.
Ring-1T is an open-weight reasoning model released in October 2025 by inclusionAI, the open-source research initiative associated with Ant Group
SciBERT is a BERT-based language model pretrained from scratch on a large corpus of scientific papers, built by the Allen Institute for AI (AI2).
Skywork R1V is a family of open-weight multimodal vision-language models built for chain-of-thought reasoning, developed by Skywork AI
SmolLM is a family of small, fully open language models released by Hugging Face on July 16, 2024 in three sizes, 135 million, 360 million, and 1.7 billion parameters, all trained on a curated open dataset…
SmolLM 2 is a family of compact open-weight language models released by Hugging Face on November 1, 2024.
SmolLM 3 is a fully open 3 billion parameter language model released by Hugging Face on July 8, 2025, trained on 11.2 trillion tokens and designed as a small, multilingual, long-context reasoner.
Step-3 is an open-weight large multimodal mixture of experts (MoE) model released in July 2025 by StepFun, the Shanghai-based Chinese artificial intelligence startup also known as Jieyue Xingchen.
Yi-Large is a closed-source large language model developed by Chinese artificial intelligence company 01.AI (零一万物, Língyi Wànwù), founded by Kai-Fu Lee.
Yi-Lightning is a closed-source large language model developed by Chinese artificial intelligence company 01.AI (零一万物, Língyī Wànwù), the company founded by Kai-Fu Lee.