Chatbot Arena
Chatbot Arena (now branded simply as Arena, and previously known as LMArena) is a crowdsourced evaluation platform for large language models that ranks AI systems based on human preferences through anonymous…
Explore language models, how they work, and the techniques used to build applications with them.
Articles that also belong to these categories. Counts cover all of Large Language Models.
Showing 61-120 of 545 articles
Chatbot Arena (now branded simply as Arena, and previously known as LMArena) is a crowdsourced evaluation platform for large language models that ranks AI systems based on human preferences through anonymous…
The Chinchilla scaling laws are a set of empirical findings published by DeepMind researchers in 2022 showing that, for a fixed compute budget, a large language model trains most efficiently when its number of…
ChipNeMo is a research project and a family of domain-adapted large language models developed by Nvidia to assist with industrial semiconductor and chip-design tasks.
Claude is a family of large language models (LLMs) developed by Anthropic, an American artificial intelligence safety and research company.
Claude 1 was the first generation of Anthropic's Claude family of large language models, launched in March 2023 and offered in two variants: a full-sized model called simply Claude and a faster
Claude 2 is a large language model released by Anthropic on July 11, 2023, the second major version of the company's Claude family of conversational AI systems and the first Claude model offered directly to…
Claude 2.1 is a large language model released by Anthropic on November 21, 2023, as an incremental update to Claude 2.
Claude 3 is a family of three large language models released by Anthropic on March 4, 2024, comprising Claude 3 Haiku, Claude 3 Sonnet, and Claude 3 Opus in ascending order of capability and cost .
Claude 3 Haiku is the fastest and most compact model in the Claude 3 family, a generation of large language models released by Anthropic in March 2024.
Claude 3 Opus is a large language model developed by Anthropic and released on March 4, 2024 as the flagship of the Claude 3 family.
Claude 3 Sonnet is the balanced mid-tier large language model in Anthropic's Claude 3 family, released on March 4, 2024 alongside the more powerful Claude 3 Opus and the smaller, faster Claude 3 Haiku.
Claude 3.5 Haiku is a fast, low-cost large language model developed by Anthropic, announced on October 22, 2024 and released through the Claude API on November 4, 2024 as the smallest and most affordable…
Claude 3.5 Sonnet was a proprietary multimodal large language model developed by Anthropic.
Claude 3.7 Sonnet is a large language model developed by Anthropic and released on February 24, 2025
Claude 4 is the fourth-generation family of large language models developed by Anthropic, an American AI safety company.
Claude Fable 5 is a large language model developed by Anthropic and released on June 9, 2026, as the first publicly available model in a new "Mythos class" that the company describes as "a tier of Claude…
Claude Fable 5.1 is a large language model developed by Anthropic and released on September 1, 2026 as the successor to Claude Fable 5 in the company's Mythos-class tier
Claude Haiku 4.5 is a small, fast large language model developed by Anthropic and released on October 15, 2025, as the lightweight, high-speed member of the Claude 4.5 model family.
Claude Instant was Anthropic's fast, low-cost line of large language models, offered through the Anthropic API from March 2023 to November 2024 as the lightweight counterpart to the flagship Claude models.
Claude Mythos 5 is a large language model released by Anthropic on June 9, 2026 as the restricted top tier of its Claude family.
Claude Mythos 5.1 is the restricted-access deployment of the large language model that Anthropic released on September 1, 2026, alongside its generally available twin, Claude Fable 5.1.
Claude Mythos Preview, usually shortened to Mythos, is a frontier language model developed by Anthropic and announced on April 7, 2026.
Claude Opus 4 is a large language model released by Anthropic on May 22, 2025 as the flagship of the Claude 4 generation, billed at launch as "the world's best coding model" and the first Anthropic model…
Claude Opus 4.1 is a large language model developed by Anthropic and released on August 5, 2025.
Claude Opus 4.5 is a large language model developed by Anthropic and released on November 24, 2025.
Claude Opus 4.6 is a large language model developed by Anthropic and released on February 5, 2026, positioned at launch as the company's most capable generally available model for long-horizon agentic work…
Claude Opus 4.7 is a hybrid reasoning large language model released by Anthropic on April 16, 2026 as the company's most capable generally available model at launch.
Claude Opus 4.8 is a large language model developed by Anthropic, released on May 28, 2026 as the most capable generally available member of the company's Claude 4 line at launch.
Claude Opus 5 is a large language model released by Anthropic on 24 July 2026, the newest entry in the Claude Opus line and the successor to Claude Opus 4.8.
Claude Sonnet 4 is a multimodal hybrid-reasoning large language model developed by Anthropic and released on May 22, 2025.
Claude Sonnet 4.5 is a multimodal large language model (LLM) developed by Anthropic and released on September 29, 2025, which Anthropic described at launch as "the best coding model in the world." It is a…
Claude Sonnet 4.6 is a large language model developed by Anthropic and released on February 17, 2026.
Claude Sonnet 5 is a large language model developed by Anthropic, released on June 30, 2026 as the mid-tier, default member of the current Claude lineup.
As of July 2026, neither Claude nor ChatGPT is universally better: the honest answer depends on the task.
Code Llama is a family of open-weight large language models specialized for code generation and understanding, released by Meta AI on August 24, 2023.
Codestral is a family of code-specialized large language models developed by Mistral AI, beginning with Codestral 22B, released on May 29, 2024
Cohere is a Canadian artificial intelligence company that develops language, retrieval, speech, and multimodal models for businesses and public-sector organizations.
Cohere Command A is a 111 billion parameter dense large language model released by Cohere on March 13, 2025, built for enterprise agents, Retrieval-Augmented Generation, and tool use across 23 languages.
Command A Reasoning is an enterprise reasoning model released by Cohere on August 21, 2025, as part of the Command A family of large language models.
Command R is a family of enterprise large language models from Cohere, launched in March 2024 and built specifically for retrieval-augmented generation (RAG), multi-step tool use, and grounded text generation…
Command R+ is a large language model developed by Cohere, the Toronto- and San Francisco-based enterprise artificial intelligence company, and released on April 4, 2024.
Compass is a proprietary family of large language models developed by Shopee and its parent company, Sea, for Southeast Asian languages and e-commerce tasks.
A compound AI system is an AI system that achieves its objectives by combining multiple interacting components, such as large language models, retrieval mechanisms, external tools, guardrails, and…
Context caching is a large-language-model API feature that stores parts of a request's input (system prompts, instructions, attached documents, or earlier conversation turns) on the provider's infrastructure…
Context engineering is the practice of designing, building, and optimizing the full set of information that a large language model receives in its context window at inference time
A context window is the finite token sequence that a language model can process for one invocation.
Cosmopedia is an open synthetic pretraining dataset released by Hugging Face in February 2024, made up of textbooks, blog posts, stories, and WikiHow-style articles written entirely by a large language model.
CrewAI is an open-source multi-agent orchestration framework that enables developers to build teams of AI agents that collaborate to accomplish complex tasks.
Cross-model KV cache transfer is an experimental technique for converting the KV cache produced by one large language model into the cache representation expected by another model.
DBRX is an open-weight mixture of experts large language model developed by Databricks and its Mosaic AI research team, released on March 27, 2024.
DSPy (short for Declarative Self-improving Python) is an open-source framework, developed at Stanford NLP, for programming rather than prompting large language models (LLMs).
Decoding strategies are the algorithms that select output tokens from a language model's next-token probability distribution during text generation.
DeepSeek is a Chinese artificial intelligence company based in Hangzhou. It was founded in 2023 by Liang Wenfeng, who had previously co-founded the quantitative investment firm High-Flyer.
DeepSeek LLM is the first foundational large language model series released by the Chinese AI company DeepSeek.
DeepSeek-V3 is a 671-billion-parameter open-weights Mixture of Experts large language model from Chinese AI lab DeepSeek, released on December 26, 2024, that activates only 37 billion parameters per token and…
DeepSeek V3.1 is a large language model developed by DeepSeek, released on August 19, 2025 and made broadly available via the official API on August 21, 2025.
DeepSeek-V3.2 is an open-weight Mixture of Experts large language model family developed by DeepSeek that introduces DeepSeek Sparse Attention (DSA)
DeepSeek V4 is a family of open-weight Mixture of Experts large language models developed by DeepSeek, a Hangzhou-based AI research lab.
DeepSeek V4-Flash is the smaller of the two large language models in the DeepSeek V4 family, a 284-billion-parameter Mixture of Experts model with 13 billion active parameters and a one-million-token context…
DeepSeek V4-Pro is the flagship model of the DeepSeek V4 family: a Mixture of Experts large language model with 1.6 trillion total parameters, 49 billion of them activated per token, and a one-million-token…