AI Models

Explore AI Models through related topics and the articles other pages reference most.

Explore articles

Reset filters
Browse subtopics: Speech & Audio AI

Articles that also belong to these categories. Counts cover all of AI Models.

Showing 1-23 of 23 articles

Audio Models

Audio models are machine learning systems that take audio as input, produce audio as output, or both, spanning speech recognition, speech synthesis, music generation, sound effect generation, voice conversion…

Speech & Audio AI

ElevenLabs v3

Eleven v3, marketed by ElevenLabs as Eleven v3 (alpha), is a third-generation text-to-speech model that ElevenLabs released in public alpha on June 5, 2025 and described as "the most expressive Text to Speech…

Generative AISpeech & Audio AI

F5-TTS

F5-TTS (short for "A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching") is an open-source text-to-speech and zero-shot voice cloning model released in October 2024 by researchers from…

Open Source AISpeech & Audio AI

GPT-Transcribe

GPT-Transcribe (gpt-transcribe) and GPT-Live-Transcribe (gpt-live-transcribe) are two closed-weight speech-to-text models that OpenAI added to its API on July 28, 2026.

OpenAISpeech & Audio AI

Lyria

Lyria is a family of AI music generation models developed by Google DeepMind, spanning text-to-music synthesis, real-time interactive music performance, and full-length song composition.

Generative AIGoogle DeepMind

Muse Voice Transcribe

Muse Voice Transcribe is a hosted speech recognition model developed by Meta Superintelligence Labs. Meta released it on September 1, 2026 as the lab's first real-time audio perception model.

Meta AISpeech & Audio AI

Sesame CSM

Sesame CSM (Conversational Speech Model) is an open weights speech generation model from Sesame AI, a San Francisco startup co-founded by former Oculus chief executive Brendan Iribe.

Generative AIOpen Source AI

Suno v5

Suno v5 is the fifth-generation AI music generation model from Suno Inc., the Cambridge, Massachusetts startup, released on September 23, 2025 to Pro and Premier subscribers as what Suno called "the world's…

Generative AIMusic & Audio Generation

Voice Activity Detection Models

Voice activity detection (VAD), also called speech activity detection (SAD), is the task of deciding which segments of an audio signal contain human speech and which contain only silence, background noise…

Speech & Audio AI