Speech & Audio AI

Explore Speech & Audio AI through related topics and the articles other pages reference most.

Explore articles

Reset filters
Browse subtopics: AI Models

Articles that also belong to these categories. Counts cover all of Speech & Audio AI.

Showing 1-23 of 23 articles

Audio Models

Audio models are machine learning systems that take audio as input, produce audio as output, or both, spanning speech recognition, speech synthesis, music generation, sound effect generation, voice conversion…

AI Models

Cartesia

Cartesia is a San Francisco-based AI company focused on real-time voice synthesis, speech recognition, and state space model (SSM) research.

AI CompaniesAI Models

Deepgram Nova-3

Deepgram Nova-3 is the third-generation automatic speech recognition (ASR) model developed by Deepgram, a San Francisco-based voice AI company.

AI ModelsVoice AI

ElevenLabs Music

ElevenLabs Music (also marketed as Eleven Music) is an AI music generation product developed by ElevenLabs, the voice AI company founded in 2022 by Piotr Dabkowski and Mati Staniszewski.

AI ModelsGenerative AI

ElevenLabs v3

Eleven v3, marketed by ElevenLabs as Eleven v3 (alpha), is a third-generation text-to-speech model that ElevenLabs released in public alpha on June 5, 2025 and described as "the most expressive Text to Speech…

AI ModelsGenerative AI

F5-TTS

F5-TTS (short for "A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching") is an open-source text-to-speech and zero-shot voice cloning model released in October 2024 by researchers from…

AI ModelsOpen Source AI

GPT-Live

GPT-Live is a family of full-duplex voice models developed by OpenAI for continuous spoken interaction.

AI ModelsOpenAI

GPT-Transcribe

GPT-Transcribe (gpt-transcribe) and GPT-Live-Transcribe (gpt-live-transcribe) are two closed-weight speech-to-text models that OpenAI added to its API on July 28, 2026.

AI ModelsOpenAI

Lyria

Lyria is a family of AI music generation models developed by Google DeepMind, spanning text-to-music synthesis, real-time interactive music performance, and full-length song composition.

AI ModelsGenerative AI

Moshi

Moshi is a full-duplex speech-to-speech foundation model developed by Kyutai, a French nonprofit artificial intelligence research laboratory.

AI ModelsConversational AI

Muse Voice Transcribe

Muse Voice Transcribe is a hosted speech recognition model developed by Meta Superintelligence Labs. Meta released it on September 1, 2026 as the lab's first real-time audio perception model.

AI ModelsMeta AI

Sesame CSM

Sesame CSM (Conversational Speech Model) is an open weights speech generation model from Sesame AI, a San Francisco startup co-founded by former Oculus chief executive Brendan Iribe.

AI ModelsGenerative AI

Suno v5

Suno v5 is the fifth-generation AI music generation model from Suno Inc., the Cambridge, Massachusetts startup, released on September 23, 2025 to Pro and Premier subscribers as what Suno called "the world's…

AI ModelsGenerative AI

Voice Activity Detection Models

Voice activity detection (VAD), also called speech activity detection (SAD), is the task of deciding which segments of an audio signal contain human speech and which contain only silence, background noise…

AI Models