Audio Models
Audio models are machine learning systems that take audio as input, produce audio as output, or both, spanning speech recognition, speech synthesis, music generation, sound effect generation, voice conversion…
Explore AI Models through related topics and the articles other pages reference most.
Articles that also belong to these categories. Counts cover all of AI Models.
Showing 1-23 of 23 articles
Audio models are machine learning systems that take audio as input, produce audio as output, or both, spanning speech recognition, speech synthesis, music generation, sound effect generation, voice conversion…
Audio-to-audio models are machine learning systems that take an audio waveform as input and produce a different audio waveform as output.
Automatic speech recognition (ASR) models, also called speech-to-text systems, are machine learning systems that convert spoken audio into written text.
Cartesia is a San Francisco-based AI company focused on real-time voice synthesis, speech recognition, and state space model (SSM) research.
Cohere Transcribe is a family of automatic speech recognition models developed by Cohere and Cohere Labs.
Deepgram Nova-3 is the third-generation automatic speech recognition (ASR) model developed by Deepgram, a San Francisco-based voice AI company.
ElevenLabs Music (also marketed as Eleven Music) is an AI music generation product developed by ElevenLabs, the voice AI company founded in 2022 by Piotr Dabkowski and Mati Staniszewski.
Eleven v3, marketed by ElevenLabs as Eleven v3 (alpha), is a third-generation text-to-speech model that ElevenLabs released in public alpha on June 5, 2025 and described as "the most expressive Text to Speech…
F5-TTS (short for "A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching") is an open-source text-to-speech and zero-shot voice cloning model released in October 2024 by researchers from…
GPT-Live is a family of full-duplex voice models developed by OpenAI for continuous spoken interaction.
GPT-Transcribe (gpt-transcribe) and GPT-Live-Transcribe (gpt-live-transcribe) are two closed-weight speech-to-text models that OpenAI added to its API on July 28, 2026.
Gemini 3.5 Transcribe is a family of speech recognition models developed by Google DeepMind and introduced by Google on August 26, 2026.
Hume Octave 2 is a multilingual emotional text-to-speech model released by Hume AI on October 1, 2025.
Lyria is a family of AI music generation models developed by Google DeepMind, spanning text-to-music synthesis, real-time interactive music performance, and full-length song composition.
Moshi is a full-duplex speech-to-speech foundation model developed by Kyutai, a French nonprofit artificial intelligence research laboratory.
Muse Voice Transcribe is a hosted speech recognition model developed by Meta Superintelligence Labs. Meta released it on September 1, 2026 as the lab's first real-time audio perception model.
Canary is a family of open speech models developed by Nvidia as part of its NeMo conversational AI toolkit.
Sesame CSM (Conversational Speech Model) is an open weights speech generation model from Sesame AI, a San Francisco startup co-founded by former Oculus chief executive Brendan Iribe.
Stable Audio 2.5 is an enterprise focused text-to-audio generation model released by Stability AI on September 10, 2025.
Suno v5 is the fifth-generation AI music generation model from Suno Inc., the Cambridge, Massachusetts startup, released on September 23, 2025 to Pro and Premier subscribers as what Suno called "the world's…
Text-to-speech (TTS) models are machine learning systems that convert written text into spoken audio.
The Universal Speech Model (USM) is a family of large multilingual speech models developed by Google Research that performs automatic speech recognition (ASR) and speech-to-text translation across more than…
Voice activity detection (VAD), also called speech activity detection (SAD), is the task of deciding which segments of an audio signal contain human speech and which contain only silence, background noise…