Thursday, July 23, 2026
- Qwen2-Mathv3Qwen2-Math is a series of mathematics-specialized large language models released by the Qwen team at Alibaba on 8 August 2024.
- Qwen2-Audiov3Qwen2-Audio is an audio-language model developed by the Qwen team at Alibaba Cloud, released in August 2024 .
- Qwen-VLv3Qwen-VL is the first family of open vision-language (multimodal) models from the Qwen team at Alibaba Cloud, able to take images, text, and bounding boxes as input and produce text and bounding boxes as output.
- Qwen2.5-Coderv3Qwen2.5-Coder is the code-specialized series within the Qwen2.5 generation of large language models developed by the Qwen team at Alibaba.
- Yi (language model)v3Yi is a series of open, bilingual (English and Chinese) large language models developed by the Chinese startup 01.AI (Chinese: 零一万物, Lingyiwanwu), the company founded in March 2023 by Kai-Fu Lee.
- GLM-130Bv3GLM-130B is a 130-billion-parameter bilingual (English and Chinese) large language model released in August 2022 by the Knowledge Engineering Group (KEG) and the Data Mining research group at Tsinghua…
- GLM-4v3GLM-4 is the fourth-generation foundation model family from Zhipu AI, a Beijing company spun out of the Knowledge Engineering Group at Tsinghua University and now trading internationally as Z.ai.
- ChatGLMv3ChatGLM is a series of open, bilingual (Chinese and English) conversational large language models developed by Zhipu AI together with the Knowledge Engineering Group (KEG) lab at Tsinghua University.
- Qwen2v3Qwen2 is the second major generation of open large language models developed by the Qwen team at Alibaba Cloud, released on 6 June 2024.
- Qwen2.5v3Qwen2.5 is a family of open-weight large language models that Alibaba Cloud's Qwen team released on 19 September 2024, spanning seven dense sizes from 0.5 billion to 72 billion parameters, pretrained on…
- DeepSeek-VLv3DeepSeek-VL is the first open-source vision-language model series from DeepSeek, the Chinese AI company.
- DeepSeek LLMv3DeepSeek LLM is the first foundational large language model series released by the Chinese AI company DeepSeek.
- DeepSeek-V2v3DeepSeek-V2 is a 236-billion-parameter mixture-of-experts (MoE) large language model released in May 2024 by DeepSeek, the Chinese AI lab spun out of the quantitative hedge fund High-Flyer and led by Liang…
- Hierav3Hiera is a hierarchical vision transformer from Meta AI (FAIR), introduced in the paper "Hiera: A Hierarchical Vision Transformer without the Bells-and-Whistles" presented as an oral at the International…
- MEGABYTEv3MEGABYTE is a transformer architecture for autoregressive modeling of very long sequences directly at the byte level, introduced by researchers at Meta AI (FAIR) in May 2023.
- Self-Taught Evaluatorv3Self-Taught Evaluator is a method for training a strong LLM-as-a-judge without any human preference annotations, using synthetic training data and an iterative self-improvement loop.
- LLM Compiler (Meta)v3The Meta Large Language Model Compiler, usually shortened to LLM Compiler, is a family of pre-trained large language model models built by Meta AI for code and compiler optimization tasks.
- Perception Encoderv3Perception Encoder (PE) is a family of vision and vision-language encoders from Meta AI's Fundamental AI Research (FAIR) group, released in April 2025
- MetaCLIPv3MetaCLIP (Metadata-Curated Language-Image Pre-training) is a data curation recipe and a family of vision-language models from Meta AI, introduced in the 2023 paper "Demystifying CLIP Data" by Hu Xu, Saining…
- Droidletv3Droidlet is an open-source platform from Facebook AI Research (now Meta AI) for building embodied AI agents.
- Joelle Pineauv3Joelle Pineau (born 1974) is a Canadian computer scientist who is the first Chief AI Officer of Cohere, a professor and William Dawson Scholar at McGill University, and a core academic member of Mila.
- Catalina (Meta AI rack)v3Catalina is a high-power, liquid-cooled rack system designed by Meta for training and serving large AI models.
- Grand Teton (AI hardware)v4Grand Teton is an open GPU hardware platform designed by Meta for training and running large AI models.
- Research SuperCluster (RSC)v3The Research SuperCluster (RSC) is an AI supercomputer built by Meta AI, the artificial intelligence research division of Meta Platforms (the company formerly known as Facebook).
- Pluribus (poker AI)v3Pluribus is an artificial intelligence program that defeated elite human professionals at six-player no-limit Texas hold'em, the most popular form of poker played by people.
- CICERO (AI)v3CICERO is an AI agent built by Meta AI's Fundamental AI Research (FAIR) division that reached human-level performance in the board game Diplomacy.
- Open Catalyst Projectv3The Open Catalyst Project (OCP) is a research collaboration between Meta AI's Fundamental AI Research group (FAIR) and Carnegie Mellon University's Department of Chemical Engineering.
- Meta Motivov3Meta Motivo is a behavioral foundation model for controlling a simulated humanoid body, released by Meta AI's Fundamental AI Research (FAIR) group on December 12, 2024.
- AI Habitatv3AI Habitat (usually just Habitat) is an open-source simulation platform for embodied AI research, developed primarily by Meta AI (the group then known as Facebook AI Research, or FAIR) together with academic…
- Sapiens (computer vision)v3Sapiens is a family of human-centric computer vision foundation models developed by Meta (Reality Labs), introduced in 2024 and presented as an oral paper at the European Conference on Computer Vision (ECCV)…
- Nougat (model)v3Nougat (Neural Optical Understanding for Academic Documents) is a document-understanding model from Meta AI that converts the rendered image of a document page into structured markup text.
- Detectron2v3Detectron2 is an open-source software library for object detection and image segmentation, built on PyTorch and developed by Facebook AI Research (FAIR), the research group now part of Meta AI.
- I-JEPAv3I-JEPA (Image-based Joint-Embedding Predictive Architecture) is a self-supervised learning method for computer vision developed by Meta AI.
- SpiRit-LMv3SpiRit-LM (also written Spirit LM) is a large language model from Meta AI's Fundamental AI Research (FAIR) group that handles spoken and written language inside a single model.
- Audioboxv3Audiobox is a foundation research model for audio generation developed by Meta AI and its Fundamental AI Research (FAIR) group.
- data2vecv3data2vec is a self-supervised learning framework from Meta AI (then Facebook AI Research) that applies the same training method to three different input types: speech, computer vision, and text.
- Massively Multilingual Speech (MMS)v3Massively Multilingual Speech (MMS) is an open-source speech project released by Meta AI in May 2023 that performs speech recognition and text-to-speech synthesis in 1,107 languages and spoken language…
- No Language Left Behind (NLLB)v3No Language Left Behind (NLLB) is a machine translation research project and model family from Meta AI, announced in July 2022.
- SeamlessM4Tv3SeamlessM4T (short for Massively Multilingual and Multimodal Machine Translation) is a machine translation model released by Meta AI on August 22, 2023.
- ImageBindv3ImageBind is a multimodal model from Meta AI (its Fundamental AI Research lab) that learns a single joint embedding space across six different modalities: images and video, text, audio, depth, thermal…
- CM3leonv3CM3leon (pronounced "chameleon") is a multimodal generative model from Meta AI, introduced in July 2023, that handles both text-to-image and image-to-text generation in a single architecture.
- Make-A-Scenev3Make-A-Scene is a text-to-image generation model published by Meta AI (then Meta AI Research) in 2022.
- Make-A-Videov3Make-A-Video is a text-to-video generation system from Meta AI, announced on September 29, 2022
- Movie Genv3Movie Gen is a suite of media-generation foundation models from Meta AI, announced on October 4, 2024, that generates high-definition video with synchronized audio from text prompts.
- Emu Editv3Emu Edit is an instruction-based image editing model from Meta AI, announced on November 16, 2023 alongside the text-to-video model Emu Video.
- Emu Videov3Emu Video is a text-to-video generation model from Meta AI, announced on November 16, 2023, that creates short clips by first turning a text prompt into an image and then generating a video conditioned on both…
- Emu (Meta AI)v3Emu is a text-to-image generation foundation model developed by Meta AI and unveiled at the Meta Connect conference in September 2023.
- Meta AI Studiov3Meta AI Studio (often shortened to AI Studio) is a free platform from Meta that lets anyone build customizable AI characters and personas without writing code, then chat with them inside Instagram, Messenger…
- Large Concept Modelv3A Large Concept Model (LCM) is a research approach to language modeling, introduced by Meta AI's Fundamental AI Research (FAIR) group in December 2024
- Byte Latent Transformerv3The Byte Latent Transformer (BLT) is a tokenizer-free large language model architecture introduced by researchers at Meta AI's Fundamental AI Research (FAIR) group in December 2024.
- Llama APIv3The Llama API is Meta's first-party hosted cloud service for running Llama models.
- Llama Stackv3Llama Stack is an open standardized framework created by Meta for building generative AI applications.
- BlenderBotv3BlenderBot is a line of open-domain conversational agents built by Facebook AI Research (FAIR), the lab now known as Meta AI.
- Atlas (language model)v3Atlas is a retrieval-augmented language model developed by researchers at Meta AI (the group then known as Facebook AI Research, or FAIR).
- Toolformerv3Toolformer is a research language model from Meta AI that learns, in a self-supervised way, to call external software tools through simple text-based API calls.
- Chameleon (Meta AI)v3Chameleon is a family of early-fusion, token-based mixed-modal foundation models from Meta AI's Fundamental AI Research (FAIR) group that represents both images and text as discrete tokens in a single unified…
- Galactica (language model)v3Galactica is a large language model for science, built by the Papers with Code team at Meta AI and released on 15 November 2022.
- OPT (Open Pre-trained Transformer)v3OPT (Open Pre-trained Transformer) is a suite of decoder-only large language models released by Meta AI in May 2022, ranging from 125 million to 175 billion parameters and built to reproduce the scale and…
- TPU v4v3TPU v4 is Google's fourth-generation Tensor Processing Unit, a custom application-specific integrated circuit (ASIC) that accelerates machine learning workloads in Google's data centers and on Google Cloud.
- Veo 3.1v2Veo 3.1 is a video generation model released by Google DeepMind on October 15, 2025, as an incremental update to Veo 3.
- Whisk (Google Labs)v2Whisk is an experimental generative-AI tool from Google Labs that lets people create images by feeding it other images rather than long written prompts.
- Lyria 2v3Lyria 2 is a high-fidelity, text-to-music generation model built by Google DeepMind that turns text prompts into professional-grade instrumental audio.
- Magenta (project)v3Magenta is an open-source research project from Google that explores the role of machine learning in creating art and music.
- Parti (text-to-image model)v3Parti (Pathways Autoregressive Text-to-Image) is a text-to-image generation model from Google Research that produces images from natural-language descriptions by treating the task as a sequence-to-sequence…
- Imagen 2v2Imagen 2 is the second generation of Google's text-to-image diffusion model, developed by Google DeepMind and first announced for developers and enterprises on December 13, 2023.
- AudioLMv2AudioLM is a framework from Google Research for generating high-quality audio by treating the problem as a language-modeling task over discrete tokens.
- MusicLMv3MusicLM is a text-to-music generation model from Google Research that generates high-fidelity music at 24 kHz from natural language descriptions and keeps that audio consistent over several minutes.
- WaveNetv3WaveNet is a deep generative model for raw audio waveforms developed by DeepMind that synthesizes speech by predicting one waveform sample at a time, each conditioned on all the samples before it.
- Pathways (Google AI)v3Pathways is the name Google has used for two related but distinct things in its artificial-intelligence work: a research vision for a next-generation AI architecture, first articulated by Jeff Dean in October…
- Minerva (language model)v2Minerva is a large language model developed by Google Research that specializes in quantitative reasoning, meaning it answers mathematics, science, and engineering questions by writing out step-by-step…
- Med-PaLM 2v2Med-PaLM 2 is a medical large language model developed by Google Research and Google DeepMind, built on the PaLM 2 foundation model and tuned to answer questions about medicine and health.
- Med-PaLMv3Med-PaLM is a large language model from Google Research, built in collaboration with DeepMind, that is specialized for answering medical questions.
- PaLM 2v3PaLM 2 is the large language model Google announced on May 10, 2023, at its I/O developer conference as the successor to the original PaLM (Pathways Language Model).
- AI co-scientist (Google)v2The AI co-scientist is a multi-agent artificial intelligence system from Google, built on the Gemini 2.0 family of models, that is designed to help expert researchers generate novel scientific hypotheses…
- TacticAIv2TacticAI is an artificial intelligence system that gives football (soccer) coaches tactical advice on corner kicks.
- GNoMEv2GNoME (Graph Networks for Materials Exploration) is a deep-learning system from Google DeepMind that predicts the thermodynamic stability of inorganic crystals and uses those predictions to search for new…
- FunSearchv2FunSearch is a method from Google DeepMind that pairs a large language model with an automated evaluator to discover new solutions to hard problems in mathematics and computer science.
- GenCastv4GenCast is a probabilistic, ensemble-based machine-learning weather forecasting model developed by Google DeepMind that uses a diffusion model to generate large ensembles of possible 15-day weather…
- GraphCastv3GraphCast is a machine-learning weather-forecasting model from Google DeepMind that produces a 10-day global forecast at 0.25-degree resolution in under one minute on a single Google Cloud TPU v4 device
- Perceiverv2Perceiver is a family of general-purpose neural network architectures from DeepMind built around attention and a small latent bottleneck.
- Dreamer (reinforcement learning)v3Dreamer is a family of model-based reinforcement learning agents that learn a compact world model of their environment and then improve their behavior by "imagining" sequences of future outcomes inside that…
- RoboCatv2RoboCat is a self-improving foundation agent for robotic manipulation developed by Google DeepMind.
- Genie (DeepMind)v3Genie is an 11-billion-parameter generative interactive environment from Google DeepMind, described by its creators as the first foundation world model: it turns a single image, photo, or sketch into a…
- Sparrow (DeepMind)v2Sparrow is a research dialogue agent built by DeepMind and introduced on 22 September 2022.
- Flamingo (visual language model)v3Flamingo is a family of visual language models (VLMs) built by DeepMind and introduced in April 2022 that brought few-shot, in-context learning to multimodal inputs.
- Gato (DeepMind)v3Gato is a single generalist AI agent built by DeepMind and described in the May 2022 paper "A Generalist Agent" (arXiv:2205.06175).
- AlphaCode 2v2AlphaCode 2 is a competitive-programming system built by Google DeepMind that uses a fine-tuned version of the Gemini family of language models to generate, filter, and rank candidate solutions to algorithmic…
- AlphaChipv2AlphaChip is a reinforcement-learning method developed by Google DeepMind for designing the physical layout of computer chips, specifically the placement of large circuit components known as macros.
- AlphaQubitv2AlphaQubit is a neural-network decoder for quantum error correction developed jointly by Google DeepMind and Google Quantum AI.
- AlphaMissensev2AlphaMissense is a machine learning model from Google DeepMind that predicts whether a missense variant, a single amino-acid substitution in a protein, is likely to cause disease.
- AlphaDevv2AlphaDev is an artificial intelligence system built by Google DeepMind that used deep reinforcement learning to discover faster algorithms for common computing tasks, most notably small-scale sorting and…
- AlphaTensorv3AlphaTensor is an artificial-intelligence system from DeepMind that uses deep reinforcement learning to discover faster algorithms for matrix multiplication.
- AlphaGo Zerov3AlphaGo Zero is a Go-playing computer program developed by DeepMind that reached a superhuman level entirely through self-play reinforcement learning, starting from random play with no human game data.
- Gemma 3nv3Gemma 3n is an open, mobile-first multimodal model from Google built to run locally on phones, tablets, laptops, and other resource-constrained hardware, accepting text, image, audio, and video as input and…
- TxGemmav2TxGemma is a collection of open language models from Google built for therapeutic development and drug discovery.
- MedGemmav3MedGemma is a collection of open medical models from Google, built on the Gemma 3 architecture and tuned for understanding medical text and medical images.
- ShieldGemmav2ShieldGemma is a family of open safety-classifier models from Google, built on the Gemma family of lightweight open models.
- RecurrentGemmav2RecurrentGemma is a family of open-weight language models released by Google DeepMind that is built on the Griffin architecture rather than the standard Transformer.
- PaliGemmav4PaliGemma is an open vision-language model developed by Google that pairs the SigLIP image encoder with a Gemma language model, takes an image plus a text prompt as input, and produces text as output.
- CodeGemmav2CodeGemma is a family of open code-generation models that Google released in April 2024, built on the first generation of its lightweight Gemma models.