AI Voice Agent
An AI voice agent is a conversational artificial intelligence system that communicates with users through spoken language in real time, holding fluid telephone or in-app conversations, interpreting intent…
Explore Speech & Audio AI through related topics and the articles other pages reference most.
Ranked by links from other AI Wiki pages.
Articles that also belong to these categories. Counts cover all of Speech & Audio AI.
Showing 1-60 of 79 articles
An AI voice agent is a conversational artificial intelligence system that communicates with users through spoken language in real time, holding fluid telephone or in-app conversations, interpreting intent…
AssemblyAI is an American Speech AI company, founded in 2017 by Dylan Fox in San Francisco, that trains its own automatic speech recognition models and sells them to developers and enterprises through a single…
Audio classification models are machine learning systems for audio classification, the task of assigning one or more labels to an audio recording or to short fragments of it.
Audio models are machine learning systems that take audio as input, produce audio as output, or both, spanning speech recognition, speech synthesis, music generation, sound effect generation, voice conversion…
Audio-to-audio models are machine learning systems that take an audio waveform as input and produce a different audio waveform as output.
AudioCraft is an open-source generative-audio library released by Meta AI (Fundamental AI Research, FAIR) on August 2, 2023 that generates high-quality music and sound from text prompts using a single…
AudioLM is a framework from Google Research for generating high-quality audio by treating the problem as a language-modeling task over discrete tokens.
Automatic speech recognition (ASR) models, also called speech-to-text systems, are machine learning systems that convert spoken audio into written text.
As of July 2026, there is no single best AI voice generator: the right text-to-speech (TTS) tool depends on the job.
Cartesia is a San Francisco-based AI company focused on real-time voice synthesis, speech recognition, and state space model (SSM) research.
Cohere Transcribe is a family of automatic speech recognition models developed by Cohere and Cohere Labs.
Connectionist temporal classification (CTC) is a loss function and output layer design for training neural networks to label unsegmented sequences, such as transcribing an audio recording into characters when…
CosyVoice is a family of open-source multilingual neural text-to-speech (TTS) and voice cloning models developed by the Tongyi Speech Lab (Tongyi SpeechTeam) at Alibaba Group and released under the Apache 2.0…
Deepgram is an American voice artificial intelligence company, founded in 2015 and headquartered in San Francisco, that builds proprietary deep learning models for speech recognition, text-to-speech synthesis…
Deepgram Nova-3 is the third-generation automatic speech recognition (ASR) model developed by Deepgram, a San Francisco-based voice AI company.
Descript is an artificial intelligence-powered audio and video editing platform that lets users edit media by editing a text transcript: delete a word from the transcript and the matching audio and video…
DolphinGemma is an audio language model developed by Google to help scientists analyze the vocalizations of wild dolphins.
ElevenLabs is a voice and audio artificial intelligence company that builds text-to-speech AI, voice cloning, AI dubbing, generative sound effects, music synthesis, and conversational voice agent technology.
ElevenLabs Music (also marketed as Eleven Music) is an AI music generation product developed by ElevenLabs, the voice AI company founded in 2022 by Piotr Dabkowski and Mati Staniszewski.
Eleven v3, marketed by ElevenLabs as Eleven v3 (alpha), is a third-generation text-to-speech model that ElevenLabs released in public alpha on June 5, 2025 and described as "the most expressive Text to Speech…
EnCodec is a real-time neural audio codec developed by Meta AI's FAIR (Fundamental AI Research) team that compresses speech, ambient sound, and music into a compact stream of discrete tokens
F5-TTS (short for "A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching") is an open-source text-to-speech and zero-shot voice cloning model released in October 2024 by researchers from…
Fireflies.ai is an artificial intelligence meeting assistant platform that automatically records, transcribes, summarizes, and analyzes conversations across video conferencing tools such as Zoom, Microsoft…
GLM-4-Voice is an open-weights end-to-end speech-to-speech large language model released in October 2024 by Zhipu AI together with the Knowledge Engineering Group (KEG) at Tsinghua University.
GPT-Live is a family of full-duplex voice models developed by OpenAI for continuous spoken interaction.
GPT-Realtime is a family of speech-to-speech models exposed through the OpenAI Realtime API, a low-latency interface that lets developers build voice agents which take spoken audio in and return spoken audio…
GPT-Transcribe (gpt-transcribe) and GPT-Live-Transcribe (gpt-live-transcribe) are two closed-weight speech-to-text models that OpenAI added to its API on July 28, 2026.
Gemini 3.5 Transcribe is a family of speech recognition models developed by Google DeepMind and introduced by Google on August 26, 2026.
Gladia is a French artificial-intelligence company that builds audio infrastructure for developers and voice-product teams, centered on a speech-to-text (transcription) API.
HuBERT (Hidden-Unit BERT) is a self-supervised learning model for speech representation, introduced by researchers at Meta AI (then Facebook AI Research) in 2021 .
Hume AI is a New York-based artificial intelligence research company and API platform, founded in March 2021 by Alan Cowen, that builds "emotionally intelligent" voice AI: models trained to measure human…
Hume Octave 2 is a multilingual emotional text-to-speech model released by Hume AI on October 1, 2025.
Inworld AI is an American artificial intelligence company headquartered in Mountain View, California, that builds real-time voice and character AI infrastructure for games, interactive applications, and voice…
Kai-Fu Lee (Chinese: 李開復; born December 3, 1961) is a Taiwanese-American computer scientist, venture capitalist, and author who is the founder and CEO of 01.AI, the Beijing large language model startup behind…
Krisp AI (formerly 2Hz) is an artificial intelligence company that develops real-time voice AI products, including noise cancellation, accent conversion, voice translation, meeting transcription, and…
LibriSpeech is a freely available corpus of approximately 1,000 hours of 16 kHz read English speech that serves as the standard benchmark for training and evaluating automatic speech recognition (ASR) systems.
Lyria is a family of AI music generation models developed by Google DeepMind, spanning text-to-music synthesis, real-time interactive music performance, and full-length song composition.
Massively Multilingual Speech (MMS) is an open-source speech project released by Meta AI in May 2023 that performs speech recognition and text-to-speech synthesis in 1,107 languages and spoken language…
Moshi is a full-duplex speech-to-speech foundation model developed by Kyutai, a French nonprofit artificial intelligence research laboratory.
Murf AI is an artificial intelligence company that develops a text-to-speech platform for generating realistic synthetic voiceovers.
Muse Voice Transcribe is a hosted speech recognition model developed by Meta Superintelligence Labs. Meta released it on September 1, 2026 as the lab's first real-time audio perception model.
AI in music is the use of artificial intelligence, especially machine learning and generative AI, to compose, perform, mix, master, transcribe, voice-clone, and reproduce music.
Canary is a family of open speech models developed by Nvidia as part of its NeMo conversational AI toolkit.
Parakeet is a family of open automatic speech recognition (ASR) models developed by NVIDIA as part of the NeMo conversational AI toolkit.
NVIDIA Riva is a GPU-accelerated software development kit and family of containerized inference services for speech and translation AI, built by NVIDIA.
VALL-E is a zero-shot learning text-to-speech (TTS) system from Microsoft Research that clones a target voice from a 3-second recording and synthesizes new speech in that voice without any per-speaker training.
The OpenAI Realtime API is a speech-to-speech interface from OpenAI that lets developers build low-latency, bidirectional voice AI applications powered by GPT-4o and later by the purpose-built gpt-realtime…
Otter.ai is an American artificial intelligence company that builds AI meeting assistants that automatically record, transcribe, summarize, and answer questions about spoken conversations in real time.
PlayHT, later rebranded PlayAI (and reachable at play.ht and play.ai), was an American generative AI voice company that built text-to-speech models, voice cloning tools, and a platform for conversational AI…
Qwen2-Audio is an audio-language model developed by the Qwen team at Alibaba Cloud, released in August 2024 .
Resemble AI is a generative voice and AI security company based in San Francisco, California, and originally founded in Toronto, Canada.
Rime (also styled Rime Labs, and reachable at rime.ai) is an American artificial intelligence company that builds text-to-speech and spoken-language models tuned specifically for business voice agents…
SUPERB, which stands for Speech processing Universal PERformance Benchmark, is a comprehensive evaluation framework designed to measure how well self-supervised learning (SSL) models generalize across a…
SeamlessM4T (short for Massively Multilingual and Multimodal Machine Translation) is a machine translation model released by Meta AI on August 22, 2023.
Sesame (formally Sesame AI Labs) is a San Francisco-based artificial intelligence company founded in June 2023, best known for developing the Conversational Speech Model (CSM) and the Maya and Miles voice…
Sesame CSM (Conversational Speech Model) is an open weights speech generation model from Sesame AI, a San Francisco startup co-founded by former Oculus chief executive Brendan Iribe.
SoundStream is an end-to-end neural audio codec introduced by Google Research in July 2021 that compresses speech, music, and general audio at low-to-medium bitrates ranging from 3 kbps to 18 kbps on 24 kHz…
Speech recognition, usually called automatic speech recognition (ASR), is the computational task of converting a spoken-language signal into a sequence of written symbols.
Speechmatics is a British artificial intelligence company that develops automatic speech recognition (ASR) and voice AI technology for enterprise customers.
SpiRit-LM (also written Spirit LM) is a large language model from Meta AI's Fundamental AI Research (FAIR) group that handles spoken and written language inside a single model.