Large Language Models

Explore language models, how they work, and the techniques used to build applications with them.

Explore articles

Browse subtopics (65)

Articles that also belong to these categories. Counts cover all of Large Language Models.

Showing 121-180 of 545 articles

DeepSeek V4.1-Flash

DeepSeek V4.1-Flash is an open-weight multimodal mixture-of-experts model released by DeepSeek on September 10, 2026. It accepts text and images and generates text.

AI ModelsChinese AI

DeepSeek-Coder

DeepSeek-Coder is a family of open-weight code large language models built for code generation, completion, and infilling, developed by the Chinese AI lab DeepSeek (DeepSeek-AI).

AI Code GenerationChinese AI

DeepSeek-Prover

DeepSeek-Prover is a family of open-weight large language models developed by Chinese AI laboratory DeepSeek for formal theorem proving in the Lean 4 proof assistant.

Chinese AIReasoning Models

DeepSeek-R1

DeepSeek-R1 is an open-weight reasoning model and large language model family developed by the Chinese artificial intelligence laboratory DeepSeek. The original model was released on January 20, 2025.

Chinese AIReasoning Models

DeepSeek-R1-Distill

DeepSeek-R1-Distill is a family of six open-weight reasoning language models released by DeepSeek on January 20, 2025, alongside the flagship DeepSeek-R1 reasoning model.

AI ModelsChinese AI

DeepSeek-V2

DeepSeek-V2 is a 236-billion-parameter mixture-of-experts (MoE) large language model released in May 2024 by DeepSeek, the Chinese AI lab spun out of the quantitative hedge fund High-Flyer and led by Liang…

Chinese AIMixture of Experts

DeepSeekMath

DeepSeekMath is a family of open-weight large language models specialized for mathematical reasoning, released by Chinese AI laboratory DeepSeek in February 2024.

Chinese AIReasoning Models

Devstral

Devstral is a family of open-weight and API large language models specialized for agentic software engineering, developed by Mistral AI in collaboration with All Hands AI

AI ModelsOpen Source AI

Diffusion Language Models

Diffusion language models (DLMs, sometimes written dLLMs at frontier scale) are text generators that synthesize a sequence by iteratively denoising or unmasking many tokens in parallel

Diffusion Models

Dolma

Dolma is an open three-trillion-token English pretraining corpus released by the Allen Institute for AI (AI2) to power its fully open OLMo language models and to let researchers study how training data shapes…

Data & DatasetsOpen Source AI

Doubao

Doubao (豆包, literally "bean bun") is China's most-used artificial intelligence chatbot, developed by ByteDance, the parent company of TikTok.

Chinese AIConversational AI

Doubao Seed 1.6

Doubao Seed 1.6 is a family of general-purpose foundation models developed by the ByteDance Seed research team and released through Volcano Engine on 11 June 2025 at the company's Force Original Power…

AI ModelsChinese AI

EAGLE-2

EAGLE-2 ("Faster Inference of Language Models with Dynamic Draft Trees") is the second generation of the EAGLE family of speculative decoding methods for accelerating large language model inference, introduced…

AI Inference

ERNIE 4.5

ERNIE 4.5 is a family of large language models released by the Chinese technology company Baidu, open-sourced on June 30, 2025 under the Apache 2.0 license .

Chinese AIMultimodal AI

ERNIE 5.0

ERNIE 5.0 is a natively omni-modal foundation model from Baidu, unveiled at the company's annual Baidu World 2025 conference in Beijing on 13 November 2025 as the flagship in the ERNIE line at launch.

Chinese AIMultimodal AI

ERNIE X1

ERNIE X1 is a deep-reasoning large language model developed by Baidu, the Chinese search and artificial-intelligence company, as part of its ERNIE (Wenxin) family.

AI ModelsOpen Source AI

EXAONE

EXAONE (an acronym for EXpert AI for EveryONE) is the family of large language models and foundation models developed by LG AI Research

Open Source AI

EleutherAI

EleutherAI is a non-profit artificial intelligence research institute that builds and openly releases large language models, datasets, and evaluation tools, and studies their interpretability and alignment.

AI CompaniesAI Research

FACTS Grounding

FACTS Grounding is a factuality benchmark from Google DeepMind and Google Research that measures whether a large language model answers a request using only the information in a provided source document

AI BenchmarksModel Evaluation

Falcon 3

Falcon 3 is a family of open-weight large language models released on December 17, 2024 by the Technology Innovation Institute (TII), an applied research center based in Abu Dhabi, United Arab Emirates .

AI CompaniesAI Models

Fireworks AI

Fireworks AI is an artificial intelligence infrastructure company that runs a high-performance inference platform for deploying and serving open large language models (LLMs), image generation models, audio…

AI CompaniesAI Inference

Function calling

Function calling is a capability of large language models (LLMs) that lets a model decide, during generation, to invoke an external function or API and emit the function name plus its arguments as structured…

AI Agents

GGUF

GGUF (GPT-Generated Unified Format) is the standard binary file format for storing large language models for local inference, bundling a model's weights, tokenizer, and metadata into a single self-contained…

Developer ToolsMachine Learning

GLM

GLM, short for General Language Model, is both a pretraining framework for language understanding and generation and the name of a model family developed by researchers associated with Tsinghua University and…

AI ModelsChinese AI

GLM-130B

GLM-130B is a 130-billion-parameter bilingual (English and Chinese) large language model released in August 2022 by the Knowledge Engineering Group (KEG) and the Data Mining research group at Tsinghua…

Chinese AIOpen Source AI

GLM-4

GLM-4 is the fourth-generation foundation model family from Zhipu AI, a Beijing company spun out of the Knowledge Engineering Group at Tsinghua University and now trading internationally as Z.ai.

AI ModelsChinese AI

GLM-4.5

GLM-4.5 is an open-weights large language model released by Zhipu AI (operating internationally as Z.ai) on July 28, 2025, built on a 355-billion-parameter Mixture of Experts architecture that activates 32…

AI ModelsChinese AI

GLM-4.6

GLM-4.6 is a flagship open-weight large language model released by Zhipu AI under its international brand Z.ai on September 30, 2025, built on a sparse Mixture of Experts (MoE) architecture with roughly 357…

AI ModelsChinese AI

GLM-5

GLM-5 is an open-weight flagship large language model released by the Chinese AI company Zhipu AI, under its international brand Z.ai, on February 11, 2026.

AI ModelsChinese AI

GLM-5.1

GLM-5.1 is an open-weight large language model developed by the Chinese AI company Zhipu AI, which markets its products internationally under the brand Z.ai.

Chinese AIOpen Source AI

GLM-5.2

GLM-5.2 is an open-weight large language model developed by the Chinese company Zhipu AI, which sells its products internationally under the brand Z.ai.

Chinese AIOpen Source AI

GLM-5.3-Flash

GLM-5.3-Flash is an open-weight, natively multimodal mixture-of-experts large language model released by Z.ai on August 26, 2026.

AI ModelsChinese AI

GPT

GPT, short for Generative Pre-trained Transformer, is the name of a model family developed by OpenAI.

AI ModelsOpenAI

GPT-1

GPT-1 is the first model in the GPT (Generative Pre-trained Transformer) series, a 117-million-parameter, 12-layer decoder-only Transformer released by OpenAI on June 11, 2018 in the paper "Improving Language…

Natural Language ProcessingOpenAI

GPT-3.5

GPT-3.5 is a family of large language models developed by OpenAI. OpenAI used the name for the series from which it fine-tuned the model behind the original ChatGPT research preview, launched on November 30

AI ModelsOpenAI

GPT-4 Turbo

GPT-4 Turbo is a family of large language models released by OpenAI as a faster, cheaper, and longer-context variant of GPT-4.

AI ModelsOpenAI

GPT-4.1

GPT-4.1 is a family of multimodal large language models developed by OpenAI and announced on April 14, 2025, with a 1 million token context window and a SWE-bench Verified coding score of 54.6%

AI ModelsOpenAI

GPT-4.1 mini

GPT-4.1 mini is a large language model developed by OpenAI and released on April 14, 2025 as the mid-size member of the GPT-4.1 family.

AI ModelsOpenAI

GPT-4.5

GPT-4.5 is a large language model developed by OpenAI and released as a research preview on February 27, 2025, then deprecated and removed from the API on July 14, 2025, less than five months later.

AI ModelsOpenAI

GPT-4V (Vision)

GPT-4V, also written GPT-4V(ision) and read as "GPT-4 with vision," is the image-understanding capability that OpenAI added to its GPT-4 large language model, letting a user supply one or more images alongside…

Multimodal AIOpenAI

GPT-5

GPT-5 is a family of proprietary large language models released by OpenAI on August 7, 2025.

AI ModelsOpenAI

GPT-5 Codex

GPT-5-Codex is a coding-specialised variant of OpenAI's GPT-5 model, announced on September 15, 2025 and tuned for agentic software engineering inside the OpenAI Codex product family.

Developer ToolsOpenAI

GPT-5.1

GPT-5.1 is a family of large language models developed by OpenAI and released on November 12, 2025.

AI ModelsOpenAI