Large Language Models

Explore language models, how they work, and the techniques used to build applications with them.

Explore articles

Browse subtopics (65)

Articles that also belong to these categories. Counts cover all of Large Language Models.

Showing 61-120 of 545 articles

Chatbot Arena

Chatbot Arena (now branded simply as Arena, and previously known as LMArena) is a crowdsourced evaluation platform for large language models that ranks AI systems based on human preferences through anonymous…

AI Benchmarks

Chinchilla scaling laws

The Chinchilla scaling laws are a set of empirical findings published by DeepMind researchers in 2022 showing that, for a fixed compute budget, a large language model trains most efficiently when its number of…

AI ResearchDeep Learning

ChipNeMo

ChipNeMo is a research project and a family of domain-adapted large language models developed by Nvidia to assist with industrial semiconductor and chip-design tasks.

AI HardwareNVIDIA

Claude 1

Claude 1 was the first generation of Anthropic's Claude family of large language models, launched in March 2023 and offered in two variants: a full-sized model called simply Claude and a faster

AI ModelsAnthropic

Claude 2

Claude 2 is a large language model released by Anthropic on July 11, 2023, the second major version of the company's Claude family of conversational AI systems and the first Claude model offered directly to…

AI ModelsAnthropic

Claude 2.1

Claude 2.1 is a large language model released by Anthropic on November 21, 2023, as an incremental update to Claude 2.

AI ModelsAnthropic

Claude 3

Claude 3 is a family of three large language models released by Anthropic on March 4, 2024, comprising Claude 3 Haiku, Claude 3 Sonnet, and Claude 3 Opus in ascending order of capability and cost .

AI Models

Claude 3 Haiku

Claude 3 Haiku is the fastest and most compact model in the Claude 3 family, a generation of large language models released by Anthropic in March 2024.

AI ModelsAnthropic

Claude 3 Opus

Claude 3 Opus is a large language model developed by Anthropic and released on March 4, 2024 as the flagship of the Claude 3 family.

AI ModelsAnthropic

Claude 3 Sonnet

Claude 3 Sonnet is the balanced mid-tier large language model in Anthropic's Claude 3 family, released on March 4, 2024 alongside the more powerful Claude 3 Opus and the smaller, faster Claude 3 Haiku.

AI ModelsAnthropic

Claude 3.5 Haiku

Claude 3.5 Haiku is a fast, low-cost large language model developed by Anthropic, announced on October 22, 2024 and released through the Claude API on November 4, 2024 as the smallest and most affordable…

AI ModelsAnthropic

Claude 4

Claude 4 is the fourth-generation family of large language models developed by Anthropic, an American AI safety company.

AI ModelsAnthropic

Claude Fable 5

Claude Fable 5 is a large language model developed by Anthropic and released on June 9, 2026, as the first publicly available model in a new "Mythos class" that the company describes as "a tier of Claude…

AI ModelsAnthropic

Claude Fable 5.1

Claude Fable 5.1 is a large language model developed by Anthropic and released on September 1, 2026 as the successor to Claude Fable 5 in the company's Mythos-class tier

AI ModelsAnthropic

Claude Haiku 4.5

Claude Haiku 4.5 is a small, fast large language model developed by Anthropic and released on October 15, 2025, as the lightweight, high-speed member of the Claude 4.5 model family.

AI ModelsAnthropic

Claude Instant

Claude Instant was Anthropic's fast, low-cost line of large language models, offered through the Anthropic API from March 2023 to November 2024 as the lightweight counterpart to the flagship Claude models.

AI ModelsAnthropic

Claude Mythos 5.1

Claude Mythos 5.1 is the restricted-access deployment of the large language model that Anthropic released on September 1, 2026, alongside its generally available twin, Claude Fable 5.1.

AI ModelsAnthropic

Claude Opus 4

Claude Opus 4 is a large language model released by Anthropic on May 22, 2025 as the flagship of the Claude 4 generation, billed at launch as "the world's best coding model" and the first Anthropic model…

AI ModelsAnthropic

Claude Opus 4.6

Claude Opus 4.6 is a large language model developed by Anthropic and released on February 5, 2026, positioned at launch as the company's most capable generally available model for long-horizon agentic work…

AI ModelsAnthropic

Claude Opus 4.7

Claude Opus 4.7 is a hybrid reasoning large language model released by Anthropic on April 16, 2026 as the company's most capable generally available model at launch.

AI ModelsAnthropic

Claude Opus 4.8

Claude Opus 4.8 is a large language model developed by Anthropic, released on May 28, 2026 as the most capable generally available member of the company's Claude 4 line at launch.

AI ModelsAnthropic

Claude Opus 5

Claude Opus 5 is a large language model released by Anthropic on 24 July 2026, the newest entry in the Claude Opus line and the successor to Claude Opus 4.8.

AI ModelsAnthropic

Claude Sonnet 5

Claude Sonnet 5 is a large language model developed by Anthropic, released on June 30, 2026 as the mid-tier, default member of the current Claude lineup.

AI ModelsAnthropic

Code Llama

Code Llama is a family of open-weight large language models specialized for code generation and understanding, released by Meta AI on August 24, 2023.

AI Code GenerationMeta AI

Cohere Command A

Cohere Command A is a 111 billion parameter dense large language model released by Cohere on March 13, 2025, built for enterprise agents, Retrieval-Augmented Generation, and tool use across 23 languages.

AI ModelsOpen Source AI

Command R

Command R is a family of enterprise large language models from Cohere, launched in March 2024 and built specifically for retrieval-augmented generation (RAG), multi-step tool use, and grounded text generation…

AI CompaniesEnterprise AI

Command R+

Command R+ is a large language model developed by Cohere, the Toronto- and San Francisco-based enterprise artificial intelligence company, and released on April 4, 2024.

AI ModelsOpen Source AI

Context caching

Context caching is a large-language-model API feature that stores parts of a request's input (system prompts, instructions, attached documents, or earlier conversation turns) on the provider's infrastructure…

AI InferenceDeveloper Tools

Cosmopedia

Cosmopedia is an open synthetic pretraining dataset released by Hugging Face in February 2024, made up of textbooks, blog posts, stories, and WikiHow-style articles written entirely by a large language model.

Data & DatasetsOpen Source AI

CrewAI

CrewAI is an open-source multi-agent orchestration framework that enables developers to build teams of AI agents that collaborate to accomplish complex tasks.

AI AgentsDeveloper Tools

DBRX

DBRX is an open-weight mixture of experts large language model developed by Databricks and its Mosaic AI research team, released on March 27, 2024.

AI CompaniesMixture of Experts

DSPy

DSPy (short for Declarative Self-improving Python) is an open-source framework, developed at Stanford NLP, for programming rather than prompting large language models (LLMs).

Developer ToolsMachine Learning

DeepSeek

DeepSeek is a Chinese artificial intelligence company based in Hangzhou. It was founded in 2023 by Liang Wenfeng, who had previously co-founded the quantitative investment firm High-Flyer.

AI CompaniesArtificial Intelligence

DeepSeek V3

DeepSeek-V3 is a 671-billion-parameter open-weights Mixture of Experts large language model from Chinese AI lab DeepSeek, released on December 26, 2024, that activates only 37 billion parameters per token and…

AI ModelsChinese AI

DeepSeek V3.1

DeepSeek V3.1 is a large language model developed by DeepSeek, released on August 19, 2025 and made broadly available via the official API on August 21, 2025.

AI ModelsChinese AI

DeepSeek V3.2

DeepSeek-V3.2 is an open-weight Mixture of Experts large language model family developed by DeepSeek that introduces DeepSeek Sparse Attention (DSA)

Chinese AIOpen Source AI

DeepSeek V4

DeepSeek V4 is a family of open-weight Mixture of Experts large language models developed by DeepSeek, a Hangzhou-based AI research lab.

AI ModelsChinese AI

DeepSeek V4-Flash

DeepSeek V4-Flash is the smaller of the two large language models in the DeepSeek V4 family, a 284-billion-parameter Mixture of Experts model with 13 billion active parameters and a one-million-token context…

AI Code GenerationAI Models

DeepSeek V4-Pro

DeepSeek V4-Pro is the flagship model of the DeepSeek V4 family: a Mixture of Experts large language model with 1.6 trillion total parameters, 49 billion of them activated per token, and a one-million-token…

AI ModelsChinese AI