Synthesia 3.0
Synthesia 3.0 is a major release of the AI video generation platform from Synthesia, the London-based company co-founded in 2017 by Victor Riparbelli, Steffen Tjerrild, Lourdes Agapito and Matthias Niessner.
Explore AI Models through related topics and the articles other pages reference most.
Articles that also belong to these categories. Counts cover all of AI Models.
Showing 361-408 of 408 articles
Synthesia 3.0 is a major release of the AI video generation platform from Synthesia, the London-based company co-founded in 2017 by Victor Riparbelli, Steffen Tjerrild, Lourdes Agapito and Matthias Niessner.
Table question answering models (TableQA models) are machine learning systems that answer natural language questions over structured tabular data such as spreadsheets, database tables, and HTML tables…
Tabular classification models predict a discrete label, or a probability distribution over labels, from rows of a table.
Tabular regression models are machine learning systems that predict a continuous numeric target from a vector of tabular features, where rows are samples and columns are heterogeneous attributes (numeric…
Text classification models are machine learning systems that assign one or more predefined categorical labels to a span of natural language text, such as positive vs. negative, spam vs. ham
Text generation models are language models trained to produce coherent natural-language text by predicting tokens one at a time, each conditioned on the preceding context.
Text-to-image models are generative artificial intelligence systems that synthesize a new image from a natural-language description, called a prompt.
Text-to-speech (TTS) models are machine learning systems that convert written text into spoken audio.
Text-to-text (text2text) generation models are a family of neural network systems that frame many natural language processing tasks as a single problem: given an input text string
Toast 1 is a hosted text model from Mixedbread that is specialized for multistep search and evidence synthesis.
Token classification models are natural language processing systems that assign a discrete label to every token in an input sequence, where a token is typically a word, subword piece, or character.
Translation models are computational systems that convert text or speech from a source language into a target language.
Tripo P1, marketed in full as Tripo Smart Mesh P1.0, is a production grade native 3D diffusion model that generates clean, engine ready 3D meshes from text or image prompts in as little as two seconds.
UMA (Universal Model for Atoms) is a family of machine-learning interatomic potentials released in 2025 by the FAIR Chemistry team at Meta AI.
Unconditional image generation models are generative neural networks that learn the marginal distribution p(x) of a set of training images and produce new samples from that learned distribution, with no extra…
The Universal Speech Model (USM) is a family of large multilingual speech models developed by Google Research that performs automatic speech recognition (ASR) and speech-to-text translation across more than…
V-JEPA (Video Joint Embedding Predictive Architecture) is a self-supervised video model from Meta AI that learns by predicting masked regions of a video in an abstract latent representation space rather than…
V-JEPA 2 (Video Joint Embedding Predictive Architecture 2) is an open-source video world model released by Meta AI on June 11, 2025 that learns to understand, predict, and plan in the physical world by…
VGGNet is a deep convolutional neural network architecture, introduced in 2014 by Karen Simonyan and Andrew Zisserman of the Visual Geometry Group at the University of Oxford, that classifies images using a…
Veo 3 is a video generation model developed by Google DeepMind and announced at Google I/O on May 20, 2025, and it is the first commercially available video generation model to natively produce synchronized…
Veo 3.1 is a video generation model released by Google DeepMind on October 15, 2025, as an incremental update to Veo 3.
Video classification models are machine learning systems that assign one or more category labels to a video clip, typically describing the human action depicted.
Vidu is an AI video generation model and platform developed by Shengshu Technology (Chinese: 生数科技), a Beijing startup that grew out of research at Tsinghua University.
Visual question answering models are AI systems that take an image and a natural language question about that image and return a natural language answer.
Voice activity detection (VAD), also called speech activity detection (SAD), is the task of deciding which segments of an audio signal contain human speech and which contain only silence, background noise…
Voyage-3 is a family of general-purpose text embedding models developed by Voyage AI, launched in September 2024 with voyage-3 and voyage-3-lite , expanded in January 2025 with voyage-3-large , and refreshed…
Wan 2.1 (also written Wan2.1, from the Chinese Tongyi Wanxiang or 通义万象) is a family of open-weights text-to-video and image-to-video generation models that Alibaba's Tongyi Wanxiang team released and…
Wan 2.1-VACE (also written Wan2.1-VACE) is an open-weights video creation and editing model released by Alibaba's Tongyi Lab on May 14, 2025 .
Wan 2.5 is a natively multimodal AI video generation model developed by Alibaba Cloud's Tongyi Lab and previewed at the company's Apsara 2025 conference in Hangzhou on September 24, 2025.
WeMM-Embedding (WeChat Multi-Modal Embedding) is a family of open-weight universal multimodal embedding models built by the WeChat Vision team at Tencent.
AI in weather forecasting refers to the use of machine learning, and especially deep learning, to predict the state of the atmosphere.
WeatherNext 2 is an artificial intelligence weather forecasting model from Google DeepMind and Google Research, announced on November 17, 2025
A world action model (WAM) is a robot policy design that builds action generation on a video world model backbone rather than on a vision-language model, so that a single network jointly predicts how a scene…
Xiaomi MiMo-V2.5 is an open-weights model family released by Xiaomi in April 2026, made up of two siblings that share a name but solve different problems.
YOLOv8 is a family of one-stage computer vision models released by Ultralytics on January 10, 2023.
Yi-Large is a closed-source large language model developed by Chinese artificial intelligence company 01.AI (零一万物, Língyi Wànwù), founded by Kai-Fu Lee.
Yi-Lightning is a closed-source large language model developed by Chinese artificial intelligence company 01.AI (零一万物, Língyī Wànwù), the company founded by Kai-Fu Lee.
ZAYA1-8B is an open-weight, reasoning-focused Mixture-of-Experts (MoE) large language model released by San Francisco-based AI research lab Zyphra on May 6, 2026.
Zero-shot classification models are machine learning systems that assign input text to a set of candidate categories without having seen labeled training examples for those specific categories.
Zero-shot image classification models are vision systems that assign images to categories the model has never encountered as labeled training examples.
dots3-note Preview is an open-weight multimodal model and large language model developed by dots studio, an AI team at Xiaohongshu.
gpt-oss is a family of open-weight large language models released by OpenAI on August 5, 2025, and OpenAI's first open-weight language models since GPT-2 in 2019.
OpenAI o4-mini is a compact reasoning model developed by OpenAI, released on April 16, 2025.
rBio (styled rbio1 in the accompanying preprint) is a biological reasoning large language model released by the Chan Zuckerberg Initiative (CZI) on August 21, 2025.
π*0.6 (written "Pi-star-0.6") is a vision-language-action robot foundation model developed by Physical Intelligence, a San Francisco robotics startup.
π0 (pronounced "pi-zero") is a vision-language-action model for general-purpose robot control developed by Physical Intelligence, a San Francisco-based robotics startup, and introduced on October 31, 2024.
π0.5 (also written pi0.5, pi 0.5, or π₀.₅, and pronounced "pi zero point five") is a vision-language-action model developed by the robotics company Physical Intelligence and released on April 22, 2025.
π₀ (pronounced pi-zero and sometimes written pi0 or pizero) is a vision-language-action model (VLA) developed by the robotics foundation-model startup Physical Intelligence