Stable Audio 2.5
Stable Audio 2.5 is an enterprise focused text-to-audio generation model released by Stability AI on September 10, 2025.
Explore Speech & Audio AI through related topics and the articles other pages reference most.
Articles that also belong to these categories. Counts cover all of Speech & Audio AI.
Showing 61-79 of 79 articles
Stable Audio 2.5 is an enterprise focused text-to-audio generation model released by Stability AI on September 10, 2025.
Suno is a generative artificial intelligence company that develops a text-to-music platform capable of producing complete songs, including vocals, instrumentals, and lyrics, from simple text prompts.
Suno v5 is the fifth-generation AI music generation model from Suno Inc., the Cambridge, Massachusetts startup, released on September 23, 2025 to Pro and Premier subscribers as what Suno called "the world's…
Superwhisper is a system-wide voice-to-text dictation application for macOS, Windows, and iOS, built by SuperUltra, Inc., a bootstrapped Toronto company founded by Neil Chudleigh.
Text-to-speech (TTS) models are machine learning systems that convert written text into spoken audio.
The Universal Speech Model (USM) is a family of large multilingual speech models developed by Google Research that performs automatic speech recognition (ASR) and speech-to-text translation across more than…
Voice activity detection (VAD), also called speech activity detection (SAD), is the task of deciding which segments of an audio signal contain human speech and which contain only silence, background noise…
Voice Engine is a speech-generation and voice-cloning model developed by OpenAI that can produce natural-sounding speech resembling a specific person from a single audio sample as short as 15 seconds.
A voice assistant is a software agent whose primary interface is spoken language: it listens for a trigger, converts speech to a machine-readable request, decides what the speaker wants, and answers with…
Voice cloning is the use of machine learning to generate synthetic speech in the voice of a specific real person (the target speaker) from a sample of their recorded audio.
Voicebox is a non-autoregressive, text-conditioned generative model for speech developed by Meta AI Research and announced on June 16, 2023.
Voxtral is a family of speech models from Mistral AI. The original open-weight speech-understanding release arrived on July 15, 2025 under the Apache 2.0 license.
Wav2Vec is a family of self-supervised learning models from Meta AI (formerly Facebook AI Research) that learn speech representations directly from raw audio waveforms
Wav2Vec 2.0 is a self-supervised learning framework for speech representation, developed by the Facebook AI Research (FAIR) group at Meta and introduced in 2020.
WaveNet is a deep generative model for raw audio waveforms developed by DeepMind that synthesizes speech by predicting one waveform sample at a time, each conditioned on all the samples before it.
Whisper is an open-source family of automatic speech recognition (ASR) models developed by OpenAI and first released on September 21, 2022.
Word error rate (WER) is the standard metric for measuring the accuracy of an automatic speech recognition (ASR) system
XTTS (sometimes stylized ⓍTTS, short for "cross-lingual text-to-speech") is an open-weights multilingual text-to-speech model developed by Coqui AI that performs zero-shot voice cloning from short reference…
iFlytek (Chinese: 科大讯飞, formally Anhui USTC iFlytek Co., Ltd.) is a partially state-owned Chinese artificial intelligence and information technology company, founded on December 30, 1999 in Hefei, Anhui…