LiteLLM
LiteLLM is an open-source AI gateway from BerriAI that lets developers call more than 100 large language model providers (including OpenAI, Anthropic, Google Gemini, Amazon Bedrock and Azure OpenAI) through a…
Explore language models, how they work, and the techniques used to build applications with them.
Articles that also belong to these categories. Counts cover all of Large Language Models.
Showing 301-360 of 545 articles
LiteLLM is an open-source AI gateway from BerriAI that lets developers call more than 100 large language model providers (including OpenAI, Anthropic, Google Gemini, Amazon Bedrock and Azure OpenAI) through a…
Llama 2 is a family of open-weight large language models developed by Meta AI. Meta released pretrained and dialogue-tuned checkpoints with 7 billion, 13 billion, and 70 billion parameters on July 18, 2023.
Llama 3 is a family of open-weight large language models developed by Meta. Meta released the original Llama 3 checkpoints on April 18, 2024, in 8-billion-parameter and 70-billion-parameter sizes.
Llama 3.1 is a family of open-weight large language models released by Meta on July 23, 2024, in three sizes, 8 billion, 70 billion, and 405 billion parameters, each shipped in both a pre-trained base form and…
Llama 3.2 is a family of four open-weight large language models released by Meta on September 25, 2024, comprising lightweight 1 billion and 3 billion parameter text-only models for on-device AI and the 11…
Llama 3.2 Vision is the set of multimodal (image-plus-text) models in Meta's Llama 3.2 family, released on September 25, 2024 at the Meta Connect 2024 developer conference.
Llama 3.3 is an instruction-tuned, text-only large language model with 70 billion parameters that Meta released on December 6, 2024
Llama 4 (Large Language Model Meta AI 4) is a family of natively multimodal large language models developed by Meta and released on April 5, 2025, the first generation of the Llama series to use a mixture of…
Llama 4 Behemoth is the announced but never publicly released flagship model in the Llama 4 family from Meta AI.
Llama 4 Scout and Llama 4 Maverick are open-weight, natively multimodal AI large language models developed by Meta and released on April 5, 2025.
Llama Nemotron is a family of open reasoning large language models built by Nvidia by post-training Meta's Llama models for math, coding, and agentic tasks.
Llama-3.1-Nemotron-70B-Instruct is a large language model released by NVIDIA in October 2024.
LoftQ (short for LoRA-Fine-Tuning-aware Quantization) is a quantization and initialization framework for large language models that jointly quantizes a pre-trained backbone and initializes the attached…
Long-context language models are large language models engineered to accept and reason over inputs far larger than the few-thousand-token windows used by early transformer systems, with frontier models in 2026…
LongBench is a benchmark suite for evaluating the long-context understanding capabilities of large language models (LLMs).
LongBench v2 is a benchmark for evaluating how well large language models understand and reason over long contexts.
LongCat-Flash is an open-weight large language model developed by the LongCat team at Meituan, the Chinese on-demand local-services and food-delivery company.
LongLoRA is a parameter-efficient fine-tuning technique that extends the context window of pre-trained large language models with substantially lower computation than full fine-tuning.
LongRoPE is a context-window extension technique for large language models (LLMs) that use rotary position embeddings (RoPE).
Longformer is a transformer architecture for processing long documents, introduced by Iz Beltagy, Matthew E. Peters
Lookahead Decoding is a parallel decoding algorithm for accelerating inference in large language models, introduced in November 2023 by Yichao Fu, Peter Bailis, Ion Stoica, and Hao Zhang from the Hao AI Lab at…
MAI-1-preview is a large language model developed by Microsoft AI, the consumer artificial intelligence division of Microsoft led by Mustafa Suleyman.
MBPP (Mostly Basic Python Problems) is a code generation benchmark of 974 crowd-sourced Python programming tasks designed to be solvable by entry-level programmers, introduced by Jacob Austin, Augustus Odena…
MMLU-Pro (Massive Multitask Language Understanding Professional) is an artificial intelligence benchmark of 12,032 ten-choice questions across 14 academic domains
MMLU-ProX is a multilingual benchmark for evaluating reasoning and knowledge in large language models
MMMU-Pro is a rigorous benchmark for evaluating multimodal AI systems on college-level, expert questions that genuinely require seeing an image, built as a harder and more robust version of the original MMMU…
MRCR (Multi-Round Co-reference Resolution) is a synthetic long-context evaluation that tests whether a large language model can locate and disambiguate among several near-identical "needles" buried inside a…
MT-Bench (Multi-Turn Benchmark) is a benchmark of 80 hand-written, two-turn questions that evaluates large language models (LLMs) on multi-turn conversation and instruction following by using a strong model…
Natural Language Processing (NLP) is the subfield of artificial intelligence and machine learning concerned with enabling computers to read, interpret, generate, and reason about human language in text and…
Magistral is the first family of reasoning models from Mistral AI, the French AI company, first released on June 10, 2025.
Mamba is a neural network architecture for sequence modeling that uses selective state space models (SSMs) to process sequential data in linear time with respect to sequence length
Many-shot jailbreaking is a technique for bypassing the safety training of a large language model by filling its context window with a long series of faux dialogue turns in which an AI assistant complies with…
Mark Chen is an American artificial intelligence researcher and research executive at OpenAI, where he serves as Chief Research Officer .
Med-PaLM is a large language model from Google Research, built in collaboration with DeepMind, that is specialized for answering medical questions.
Med-PaLM 2 is a medical large language model developed by Google Research and Google DeepMind, built on the PaLM 2 foundation model and tuned to answer questions about medicine and health.
Mercury is a family of commercial-scale diffusion-based large language models developed by Inception Labs, a Palo Alto startup co-founded in 2024 by Stanford professor Stefano Ermon together with his former…
Meta AI is the name Meta Platforms uses for two related but distinct things: the company's artificial intelligence research and engineering organization
Microsoft 365 Copilot is an artificial intelligence assistant developed by Microsoft that integrates large language models (LLMs) with the Microsoft 365 suite of productivity applications.
Minerva is a large language model developed by Google Research that specializes in quantitative reasoning, meaning it answers mathematics, science, and engineering questions by writing out step-by-step…
MiniCPM is a family of compact, openly licensed language models published by OpenBMB, the shared open-source brand of Tsinghua University's natural language processing lab (THUNLP) and the Beijing company…
MiniCPM5-2B is an open-weights small language model published by OpenBMB on September 7, 2026 under the Apache License 2.0.
MiniMax (Chinese: 上海稀宇科技有限公司, Shanghai Xiyu Technology Co., Ltd.) is a publicly traded artificial intelligence (AI) company headquartered in Shanghai, China that develops large language models, multimodal AI…
MiniMax M1 (stylised MiniMax-M1) is an open-weight large language reasoning model released on 16 June 2025 by the Shanghai-based artificial-intelligence company MiniMax
MiniMax M2 is an open-weight large language model released on October 27, 2025 by the Shanghai-based AI company MiniMax, built as a Mixture of Experts model with 230 billion total parameters and roughly 10…
MiniMax M2.7 is a large language model released by the Chinese AI company MiniMax on March 18, 2026.
MiniMax M3 is a large language model developed by the Shanghai-based AI company MiniMax, released on June 1, 2026 as the next flagship in the company's M-series.
MiniMax-Text-01 is an open-weights, large-scale mixture-of-experts (MoE) language model released by Shanghai-based AI company MiniMax on January 14, 2025.
Ministral is a family of two small language models released by the French artificial intelligence company Mistral AI on October 16, 2024.
Mistral 7B is a 7.3-billion-parameter, decoder-only large language model released by mistral ai on September 27, 2023, under the apache 2 license.
Mistral AI is a French artificial intelligence company that develops large language models, multimodal models, software for building and operating AI systems, and computing infrastructure.
Mistral Large is the family of flagship large language models developed by Mistral AI, the Paris-based AI laboratory founded in 2023, and is the company's most capable general-purpose model line.
Mistral Large 3 is a sparse mixture-of-experts large language model released on December 2, 2025 by the French AI company Mistral AI, distributed as open weights under the Apache 2.0 license with roughly 675…
Mistral Medium 3 is a proprietary multimodal large language model developed by Mistral AI and released on May 7, 2025.
Mistral Medium 3.5 is an open-weight large language model released by Mistral AI on 28 April 2026.
Mistral NeMo is a 12 billion parameter large language model released by Mistral AI in collaboration with NVIDIA on July 18, 2024 .
Mistral Small 4 is an open-weight large language model released by the French artificial intelligence company Mistral AI on March 16, 2026 under the Apache 2.0 license.
Mixtral is a family of open-weight Sparse Mixture of Experts (SMoE) large language models developed by Mistral AI, a French artificial intelligence company founded in April 2023.
Mixtral 8x22B is a sparse mixture-of-experts (MoE) large language model released by the French AI company Mistral AI on April 17, 2024.
Mixture of Agents (MoA) is a multi-model collaboration framework that combines multiple large language models (LLMs) in a layered architecture
Model Context Protocol (MCP) is an open protocol for exchanging context and invoking capabilities between applications that use large language models and external programs or data sources.