Large Language Models

Explore language models, how they work, and the techniques used to build applications with them.

Explore articles

Browse subtopics (65)

Articles that also belong to these categories. Counts cover all of Large Language Models.

Showing 481-540 of 545 articles

SmolLM 3

SmolLM 3 is a fully open 3 billion parameter language model released by Hugging Face on July 8, 2025, trained on 11.2 trillion tokens and designed as a small, multilingual, long-context reasoner.

AI ModelsOpen Source AI

SmoothQuant

SmoothQuant is a training-free, accuracy-preserving post-training quantization (PTQ) method that enables 8-bit weight and 8-bit activation (W8A8) integer inference for large language models without retraining…

AI Inference

Snowflake AI

Snowflake AI is the suite of artificial intelligence and machine learning capabilities built into the Snowflake AI Data Cloud, anchored by Cortex AI (managed generative AI services callable in SQL), the…

AI CompaniesAI Infrastructure

SparDA

SparDA (Sparse Decoupled Attention) is an add-on architecture for long-context large language model inference proposed by researchers at NVIDIA in a paper posted to arXiv on 3 June 2026.

AI InferenceModel Architecture

Speculative Decoding

Speculative decoding is a lossless inference acceleration technique for autoregressive transformer models in which a small, fast draft model proposes several future tokens at once and the larger target model…

AI InferenceDeep Learning

SpiRit-LM

SpiRit-LM (also written Spirit LM) is a large language model from Meta AI's Fundamental AI Research (FAIR) group that handles spoken and written language inside a single model.

Meta AISpeech & Audio AI

StarCoder

StarCoder is a family of open-access large language models for code generation and code understanding, developed by the BigCode project, an open scientific collaboration led by Hugging Face and ServiceNow.

AI Code GenerationOpen Source AI

Step-3

Step-3 is an open-weight large multimodal mixture of experts (MoE) model released in July 2025 by StepFun, the Shanghai-based Chinese artificial intelligence startup also known as Jieyue Xingchen.

AI Models

StepFun

StepFun (Chinese: 阶跃星辰, pinyin: Jiēyuè Xīngchén), formally Shanghai Jieyue Xingchen Intelligent Technology Co., Ltd., is a Shanghai-based Chinese artificial intelligence startup that builds the "Step" series…

AI CompaniesChinese AI

StreamingLLM

StreamingLLM is an inference-time technique that allows pretrained transformer language models, originally trained with a finite attention window

AI Inference

SubQ

SubQ is a large language model released on May 5, 2026 by Subquadratic, a Miami-based startup that emerged from stealth claiming to have built the first frontier model on a fully subquadratic attention…

AI CompaniesModel Architecture

Switch Transformer

The Switch Transformer is a sparsely activated Mixture of Experts (MoE) Transformer architecture introduced by William Fedus, Barret Zoph, and Noam Shazeer at Google in January 2021.

GoogleMixture of Experts

System prompt

A system prompt is a special set of instructions, guidelines, persona definitions, and contextual information given to a large language model (LLM) before any user input

AI SafetyPrompt Engineering

T5 (language model)

T5 (Text-to-Text Transfer Transformer) is a family of transformer-based language models released by Google in 2019-2020 that reframes every natural language processing (NLP) task, classification, translation…

Transformer Models

Tang Jie

Tang Jie (Chinese: 唐杰; born 1977), also published as Jie Tang, is a Chinese computer scientist, a chair professor at Tsinghua University, and the co-founder and chief scientist of Zhipu AI, the Beijing startup…

Chinese AIPeople

Tencent AI

Tencent AI is the artificial intelligence research, products, and services developed by Tencent Holdings Ltd., one of the world's largest technology companies, built around the Hunyuan family of foundation…

AI CompaniesChinese AI

Tokens per second

Tokens per second (TPS) is a key performance metric for measuring the speed of large language model (LLM) inference. It quantifies how many tokens a model can generate or process in one second.

AI BenchmarksAI Hardware

Tongyi Qianwen

Tongyi Qianwen (通义千问), the brand often glossed in English as "seeking truth by asking a thousand questions," is Alibaba Group's flagship large language model and conversational AI brand, launched by Alibaba…

Chinese AIConversational AI

Tool use

Tool use in artificial intelligence is the ability of a model-based system to request, coordinate, and use capabilities outside the model's ordinary token-generation process.

AI AgentsArtificial Intelligence

Toolformer

Toolformer is a research language model from Meta AI that learns, in a self-supervised way, to call external software tools through simple text-based API calls.

AI AgentsMeta AI

Top-k sampling

Top-k sampling is a decoding strategy for autoregressive language models that restricts each generation step to the k most probable next tokens.

AI InferenceAlgorithms

TxT360

TxT360 is an open large-scale pretraining corpus for large language models, released in October 2024 by the LLM360 project, a collaboration led by Petuum and the Mohamed bin Zayed University of Artificial…

Data & DatasetsOpen Source AI

UltraChat

UltraChat is a large-scale synthetic multi-turn instructional conversation dataset released in May 2023 by the OpenBMB group at Tsinghua University, comprising approximately 1.5 million dialogues generated by…

Chinese AIData & Datasets

Voxtral

Voxtral is a family of speech models from Mistral AI. The original open-weight speech-understanding release arrived on July 15, 2025 under the Apache 2.0 license.

Open Source AISpeech & Audio AI

WizardLM

WizardLM is a family of open-weights instruction-tuned LLaMA-derived large language models and an associated data-synthesis methodology, both produced by a research group at Microsoft led by Can Xu.

Open Source AI

WordPiece

WordPiece is a subword tokenization algorithm that builds a fixed-size vocabulary of word pieces by repeatedly merging the symbol pair whose combination most increases the likelihood of the training corpus…

Natural Language Processing

Wu Dao

Wu Dao (Chinese: 悟道, roughly "enlightenment" or "understanding the way") is a series of large pretrained models built by the Beijing Academy of Artificial Intelligence (BAAI), a non-profit research institute…

AI HistoryChinese AI

XLM-RoBERTa

XLM-RoBERTa (often abbreviated XLM-R) is a multilingual masked language model developed by Facebook AI Research (now Meta AI) and released in November 2019.

Natural Language Processing

Yang Zhilin

Yang Zhilin (Chinese: 杨植麟; pinyin: Yáng Zhílín) is the co-founder and chief executive officer of Moonshot AI (月之暗面), the Beijing startup that develops the Kimi chatbot and the Kimi family of large language…

Chinese AIPeople

Yi (language model)

Yi is a series of open, bilingual (English and Chinese) large language models developed by the Chinese startup 01.AI (Chinese: 零一万物, Lingyiwanwu), the company founded in March 2023 by Kai-Fu Lee.

Chinese AIOpen Source AI

Yi-Large

Yi-Large is a closed-source large language model developed by Chinese artificial intelligence company 01.AI (零一万物, Língyi Wànwù), founded by Kai-Fu Lee.

AI ModelsChinese AI

Yi-Lightning

Yi-Lightning is a closed-source large language model developed by Chinese artificial intelligence company 01.AI (零一万物, Língyī Wànwù), the company founded by Kai-Fu Lee.

AI ModelsChinese AI

ZAYA1-8B

ZAYA1-8B is an open-weight, reasoning-focused Mixture-of-Experts (MoE) large language model released by San Francisco-based AI research lab Zyphra on May 6, 2026.

AI ModelsMixture of Experts

Zhipu AI

Zhipu AI (智谱AI), now branded internationally as Z.ai, is a Chinese artificial intelligence company headquartered in Beijing.

AI CompaniesChinese AI

Zyphra

Zyphra is an American artificial intelligence research and product company headquartered in San Francisco, California, with a secondary office in London.

AI CompaniesMixture of Experts

gpt-oss

gpt-oss is a family of open-weight large language models released by OpenAI on August 5, 2025, and OpenAI's first open-weight language models since GPT-2 in 2019.

AI ModelsMixture of Experts

llama.cpp

llama.cpp is an open-source large language model inference engine written in C and C++ by Bulgarian software engineer Georgi Gerganov that runs large language models on consumer-grade hardware without…

Developer ToolsMachine Learning