AssemblyAI
AssemblyAI is an American Speech AI company, founded in 2017 by Dylan Fox in San Francisco, that trains its own automatic speech recognition models and sells them to developers and enterprises through a single…
Explore Voice AI through related topics and the articles other pages reference most.
Articles that also belong to these categories. Counts cover all of Voice AI.
Showing 1-26 of 26 articles
AssemblyAI is an American Speech AI company, founded in 2017 by Dylan Fox in San Francisco, that trains its own automatic speech recognition models and sells them to developers and enterprises through a single…
Cartesia is a San Francisco-based AI company focused on real-time voice synthesis, speech recognition, and state space model (SSM) research.
CosyVoice is a family of open-source multilingual neural text-to-speech (TTS) and voice cloning models developed by the Tongyi Speech Lab (Tongyi SpeechTeam) at Alibaba Group and released under the Apache 2.0…
Deepgram is an American voice artificial intelligence company, founded in 2015 and headquartered in San Francisco, that builds proprietary deep learning models for speech recognition, text-to-speech synthesis…
Deepgram Nova-3 is the third-generation automatic speech recognition (ASR) model developed by Deepgram, a San Francisco-based voice AI company.
ElevenLabs is a voice and audio artificial intelligence company that builds text-to-speech AI, voice cloning, AI dubbing, generative sound effects, music synthesis, and conversational voice agent technology.
GLM-4-Voice is an open-weights end-to-end speech-to-speech large language model released in October 2024 by Zhipu AI together with the Knowledge Engineering Group (KEG) at Tsinghua University.
GPT-Live is a family of full-duplex voice models developed by OpenAI for continuous spoken interaction.
GPT-Realtime is a family of speech-to-speech models exposed through the OpenAI Realtime API, a low-latency interface that lets developers build voice agents which take spoken audio in and return spoken audio…
Gladia is a French artificial-intelligence company that builds audio infrastructure for developers and voice-product teams, centered on a speech-to-text (transcription) API.
Hume AI is a New York-based artificial intelligence research company and API platform, founded in March 2021 by Alan Cowen, that builds "emotionally intelligent" voice AI: models trained to measure human…
Inworld AI is an American artificial intelligence company headquartered in Mountain View, California, that builds real-time voice and character AI infrastructure for games, interactive applications, and voice…
Moshi is a full-duplex speech-to-speech foundation model developed by Kyutai, a French nonprofit artificial intelligence research laboratory.
Murf AI is an artificial intelligence company that develops a text-to-speech platform for generating realistic synthetic voiceovers.
Muse Voice Transcribe is a hosted speech recognition model developed by Meta Superintelligence Labs. Meta released it on September 1, 2026 as the lab's first real-time audio perception model.
The OpenAI Realtime API is a speech-to-speech interface from OpenAI that lets developers build low-latency, bidirectional voice AI applications powered by GPT-4o and later by the purpose-built gpt-realtime…
PlayHT, later rebranded PlayAI (and reachable at play.ht and play.ai), was an American generative AI voice company that built text-to-speech models, voice cloning tools, and a platform for conversational AI…
Resemble AI is a generative voice and AI security company based in San Francisco, California, and originally founded in Toronto, Canada.
Rime (also styled Rime Labs, and reachable at rime.ai) is an American artificial intelligence company that builds text-to-speech and spoken-language models tuned specifically for business voice agents…
Sesame (formally Sesame AI Labs) is a San Francisco-based artificial intelligence company founded in June 2023, best known for developing the Conversational Speech Model (CSM) and the Maya and Miles voice…
Speechmatics is a British artificial intelligence company that develops automatic speech recognition (ASR) and voice AI technology for enterprise customers.
Superwhisper is a system-wide voice-to-text dictation application for macOS, Windows, and iOS, built by SuperUltra, Inc., a bootstrapped Toronto company founded by Neil Chudleigh.
Voice Engine is a speech-generation and voice-cloning model developed by OpenAI that can produce natural-sounding speech resembling a specific person from a single audio sample as short as 15 seconds.
A voice assistant is a software agent whose primary interface is spoken language: it listens for a trigger, converts speech to a machine-readable request, decides what the speaker wants, and answers with…
Voice cloning is the use of machine learning to generate synthetic speech in the voice of a specific real person (the target speaker) from a sample of their recorded audio.
XTTS (sometimes stylized ⓍTTS, short for "cross-lingual text-to-speech") is an open-weights multilingual text-to-speech model developed by Coqui AI that performs zero-shot voice cloning from short reference…