Large Language Models

Explore language models, how they work, and the techniques used to build applications with them.

Explore articles

Reset filters
Browse subtopics: Open Source AI

Articles that also belong to these categories. Counts cover all of Large Language Models.

Showing 121-165 of 165 articles

Phi-4

Phi-4 is a 14-billion-parameter small language model developed by Microsoft Research and released in December 2024, designed to match or beat models several times its size on reasoning tasks by training…

AI ModelsMicrosoft

Phi-4-mini

Phi-4-mini is a 3.8 billion parameter open weight small language model released by Microsoft on February 26, 2025, under the permissive MIT license.

AI ModelsOpen Source AI

Pixtral

Pixtral is a family of multimodal vision-language models developed by Mistral AI, a French AI company founded in April 2023.

AI CompaniesAI Models

QwQ

QwQ is a family of open-weight reasoning models from the Qwen team at Alibaba Cloud, built to compete with OpenAI's o1 and DeepSeek-R1 at a fraction of their size.

Chinese AIOpen Source AI

Qwen2

Qwen2 is the second major generation of open large language models developed by the Qwen team at Alibaba Cloud, released on 6 June 2024.

Chinese AIOpen Source AI

Qwen2.5

Qwen2.5 is a family of open-weight large language models that Alibaba Cloud's Qwen team released on 19 September 2024, spanning seven dense sizes from 0.5 billion to 72 billion parameters, pretrained on…

Chinese AIOpen Source AI

Qwen2.5-Math

Qwen2.5-Math is a family of mathematics-specialized large language models developed by the Qwen team at Alibaba Cloud and released in September 2024.

Chinese AIMathematics

Qwen3

Qwen3 is the third-generation family of large language models developed by the Qwen Team at Alibaba Cloud (also known as Tongyi Qianwen lab).

AI ModelsChinese AI

Qwen3-Coder

Qwen3-Coder is a family of open-weight large language models specialized for software engineering, developed by Alibaba's Qwen team (Tongyi Lab) and released under the Apache 2.0 license.

Chinese AIDeveloper Tools

Qwen3-Next

Qwen3-Next is an efficiency-focused large language model and model architecture released in September 2025 by the Qwen team at Alibaba Cloud.

AI ModelsOpen Source AI

Qwen3-Omni

Qwen3-Omni is a natively end-to-end omni-modal foundation model developed by the Qwen team at Alibaba Cloud, capable of understanding text, images, audio, and video and generating both text and natural speech…

Chinese AIMultimodal AI

Qwen3-VL

Qwen3-VL is a family of open-weight vision-language models built by the Qwen team at Alibaba Cloud, first released in September 2025.

Chinese AIMultimodal AI

Qwen3.5

Qwen3.5 is a family of open-weight large language models developed by the Qwen team at Alibaba, the successor to the Qwen3 series.

Chinese AIOpen Source AI

Qwen3.6

Qwen3.6 is a generation of large language models from the Qwen team at Alibaba, released in April 2026 as the successor to Qwen3.5.

Chinese AIOpen Source AI

RWKV-7 (Goose)

RWKV-7, codenamed Goose, is an attention-free, RNN-style large-language-model architecture introduced in March 2025 that runs inference in linear time with constant memory per token while still training in…

AI ResearchNeural Networks

RefinedWeb

RefinedWeb is a large-scale English pretraining dataset for large language models, built from filtered and deduplicated Common Crawl web data alone and released in June 2023 by the Technology Innovation…

Data & DatasetsOpen Source AI

Reka Flash

Reka Flash is a family of multimodal large language models developed by Reka AI, a San Francisco Bay Area research company founded in 2022 by former researchers from Google DeepMind, Meta FAIR, and Google.

AI ModelsMultimodal AI

Ring-1T

Ring-1T is an open-weight reasoning model released in October 2025 by inclusionAI, the open-source research initiative associated with Ant Group

AI ModelsOpen Source AI

SlimPajama

SlimPajama is a 627-billion-token English-language pre-training corpus for large language models, produced by Cerebras Systems in collaboration with the Opentensor Foundation by extensively cleaning and…

Data & DatasetsOpen Source AI

SmolLM

SmolLM is a family of small, fully open language models released by Hugging Face on July 16, 2024 in three sizes, 135 million, 360 million, and 1.7 billion parameters, all trained on a curated open dataset…

AI ModelsOpen Source AI

SmolLM 3

SmolLM 3 is a fully open 3 billion parameter language model released by Hugging Face on July 8, 2025, trained on 11.2 trillion tokens and designed as a small, multilingual, long-context reasoner.

AI ModelsOpen Source AI

StarCoder

StarCoder is a family of open-access large language models for code generation and code understanding, developed by the BigCode project, an open scientific collaboration led by Hugging Face and ServiceNow.

AI Code GenerationOpen Source AI

TxT360

TxT360 is an open large-scale pretraining corpus for large language models, released in October 2024 by the LLM360 project, a collaboration led by Petuum and the Mohamed bin Zayed University of Artificial…

Data & DatasetsOpen Source AI

Voxtral

Voxtral is a family of speech models from Mistral AI. The original open-weight speech-understanding release arrived on July 15, 2025 under the Apache 2.0 license.

Open Source AISpeech & Audio AI

WizardLM

WizardLM is a family of open-weights instruction-tuned LLaMA-derived large language models and an associated data-synthesis methodology, both produced by a research group at Microsoft led by Can Xu.

Open Source AI

Yi (language model)

Yi is a series of open, bilingual (English and Chinese) large language models developed by the Chinese startup 01.AI (Chinese: 零一万物, Lingyiwanwu), the company founded in March 2023 by Kai-Fu Lee.

Chinese AIOpen Source AI

ZAYA1-8B

ZAYA1-8B is an open-weight, reasoning-focused Mixture-of-Experts (MoE) large language model released by San Francisco-based AI research lab Zyphra on May 6, 2026.

AI ModelsMixture of Experts

Zyphra

Zyphra is an American artificial intelligence research and product company headquartered in San Francisco, California, with a secondary office in London.

AI CompaniesMixture of Experts

gpt-oss

gpt-oss is a family of open-weight large language models released by OpenAI on August 5, 2025, and OpenAI's first open-weight language models since GPT-2 in 2019.

AI ModelsMixture of Experts

llama.cpp

llama.cpp is an open-source large language model inference engine written in C and C++ by Bulgarian software engineer Georgi Gerganov that runs large language models on consumer-grade hardware without…

Developer ToolsMachine Learning

mT5

mT5 (multilingual T5) is a transformer-based encoder-decoder language model released by Google Research in October 2020 that covers 101 languages in a single model, pre-trained on a Common Crawl corpus called…

Natural Language ProcessingOpen Source AI