AudioCraft
AudioCraft is an open-source generative-audio library released by Meta AI (Fundamental AI Research, FAIR) on August 2, 2023 that generates high-quality music and sound from text prompts using a single…
Explore Speech & Audio AI through related topics and the articles other pages reference most.
Articles that also belong to these categories. Counts cover all of Speech & Audio AI.
Showing 1-18 of 18 articles
AudioCraft is an open-source generative-audio library released by Meta AI (Fundamental AI Research, FAIR) on August 2, 2023 that generates high-quality music and sound from text prompts using a single…
As of July 2026, there is no single best AI voice generator: the right text-to-speech (TTS) tool depends on the job.
ElevenLabs is a voice and audio artificial intelligence company that builds text-to-speech AI, voice cloning, AI dubbing, generative sound effects, music synthesis, and conversational voice agent technology.
ElevenLabs Music (also marketed as Eleven Music) is an AI music generation product developed by ElevenLabs, the voice AI company founded in 2022 by Piotr Dabkowski and Mati Staniszewski.
Eleven v3, marketed by ElevenLabs as Eleven v3 (alpha), is a third-generation text-to-speech model that ElevenLabs released in public alpha on June 5, 2025 and described as "the most expressive Text to Speech…
Hume Octave 2 is a multilingual emotional text-to-speech model released by Hume AI on October 1, 2025.
Lyria is a family of AI music generation models developed by Google DeepMind, spanning text-to-music synthesis, real-time interactive music performance, and full-length song composition.
AI in music is the use of artificial intelligence, especially machine learning and generative AI, to compose, perform, mix, master, transcribe, voice-clone, and reproduce music.
VALL-E is a zero-shot learning text-to-speech (TTS) system from Microsoft Research that clones a target voice from a 3-second recording and synthesizes new speech in that voice without any per-speaker training.
PlayHT, later rebranded PlayAI (and reachable at play.ht and play.ai), was an American generative AI voice company that built text-to-speech models, voice cloning tools, and a platform for conversational AI…
Resemble AI is a generative voice and AI security company based in San Francisco, California, and originally founded in Toronto, Canada.
Rime (also styled Rime Labs, and reachable at rime.ai) is an American artificial intelligence company that builds text-to-speech and spoken-language models tuned specifically for business voice agents…
Sesame CSM (Conversational Speech Model) is an open weights speech generation model from Sesame AI, a San Francisco startup co-founded by former Oculus chief executive Brendan Iribe.
Stable Audio 2.5 is an enterprise focused text-to-audio generation model released by Stability AI on September 10, 2025.
Suno is a generative artificial intelligence company that develops a text-to-music platform capable of producing complete songs, including vocals, instrumentals, and lyrics, from simple text prompts.
Suno v5 is the fifth-generation AI music generation model from Suno Inc., the Cambridge, Massachusetts startup, released on September 23, 2025 to Pro and Premier subscribers as what Suno called "the world's…
Voice cloning is the use of machine learning to generate synthetic speech in the voice of a specific real person (the target speaker) from a sample of their recorded audio.
Voicebox is a non-autoregressive, text-conditioned generative model for speech developed by Meta AI Research and announced on June 16, 2023.