Audio-to-Audio Models
Audio-to-audio models are machine learning systems that take an audio waveform as input and produce a different audio waveform as output.
Explore Speech & Audio AI through related topics and the articles other pages reference most.
Articles that also belong to these categories. Counts cover all of Speech & Audio AI.
Showing 1-7 of 7 articles
Audio-to-audio models are machine learning systems that take an audio waveform as input and produce a different audio waveform as output.
AudioLM is a framework from Google Research for generating high-quality audio by treating the problem as a language-modeling task over discrete tokens.
ElevenLabs Music (also marketed as Eleven Music) is an AI music generation product developed by ElevenLabs, the voice AI company founded in 2022 by Piotr Dabkowski and Mati Staniszewski.
Lyria is a family of AI music generation models developed by Google DeepMind, spanning text-to-music synthesis, real-time interactive music performance, and full-length song composition.
Stable Audio 2.5 is an enterprise focused text-to-audio generation model released by Stability AI on September 10, 2025.
Suno is a generative artificial intelligence company that develops a text-to-music platform capable of producing complete songs, including vocals, instrumentals, and lyrics, from simple text prompts.
Suno v5 is the fifth-generation AI music generation model from Suno Inc., the Cambridge, Massachusetts startup, released on September 23, 2025 to Pro and Premier subscribers as what Suno called "the world's…