Voice AI

Explore Voice AI through related topics and the articles other pages reference most.

Most referenced in this topic

Ranked by links from other AI Wiki pages.

Explore articles

Browse subtopics (21)

Articles that also belong to these categories. Counts cover all of Voice AI.

Showing 1-41 of 41 articles

Advanced Voice Mode

Advanced Voice Mode is a real-time, spoken conversation feature in ChatGPT, developed by OpenAI, that lets users talk with the assistant using natural speech and receive spoken replies.

ChatGPTOpenAI

Amazon Alexa

Amazon Alexa is a virtual assistant developed by Amazon, first introduced on November 6, 2014, alongside the Amazon Echo smart speaker, and rebuilt in 2025 as a generative AI assistant called Alexa+ that runs…

AI CompaniesAI Tools & Products

AssemblyAI

AssemblyAI is an American Speech AI company, founded in 2017 by Dylan Fox in San Francisco, that trains its own automatic speech recognition models and sells them to developers and enterprises through a single…

AI CompaniesSpeech & Audio AI

Bland AI

Bland AI is an enterprise voice AI platform that enables businesses to deploy, manage, and scale AI-powered phone agents for inbound and outbound calling.

AI AgentsAI Companies

Cartesia

Cartesia is a San Francisco-based AI company focused on real-time voice synthesis, speech recognition, and state space model (SSM) research.

AI CompaniesAI Models

CosyVoice

CosyVoice is a family of open-source multilingual neural text-to-speech (TTS) and voice cloning models developed by the Tongyi Speech Lab (Tongyi SpeechTeam) at Alibaba Group and released under the Apache 2.0…

Chinese AISpeech & Audio AI

Deepgram

Deepgram is an American voice artificial intelligence company, founded in 2015 and headquartered in San Francisco, that builds proprietary deep learning models for speech recognition, text-to-speech synthesis…

AI CompaniesNatural Language Processing

ElevenLabs

ElevenLabs is a voice and audio artificial intelligence company that builds text-to-speech AI, voice cloning, AI dubbing, generative sound effects, music synthesis, and conversational voice agent technology.

AI CompaniesGenerative AI

GLM-4-Voice

GLM-4-Voice is an open-weights end-to-end speech-to-speech large language model released in October 2024 by Zhipu AI together with the Knowledge Engineering Group (KEG) at Tsinghua University.

Chinese AISpeech & Audio AI

GPT-Live

GPT-Live is a family of full-duplex voice models developed by OpenAI for continuous spoken interaction.

AI ModelsOpenAI

Gemini Live

Gemini Live is Google's natural-voice conversational mode in the Gemini app: a hands-free, interruptible spoken interface that lets a person talk to Google's AI assistant in real time, the way they would on a…

Conversational AIGoogle

Gladia

Gladia is a French artificial-intelligence company that builds audio infrastructure for developers and voice-product teams, centered on a speech-to-text (transcription) API.

AI CompaniesSpeech & Audio AI

HeyGen

HeyGen is an American artificial intelligence company headquartered in Los Angeles, California, that develops AI-powered video generation software specializing in digital avatars, voice cloning, and…

AI CompaniesGenerative AI

Hume AI

Hume AI is a New York-based artificial intelligence research company and API platform, founded in March 2021 by Alan Cowen, that builds "emotionally intelligent" voice AI: models trained to measure human…

AI CompaniesConversational AI

Inworld AI

Inworld AI is an American artificial intelligence company headquartered in Mountain View, California, that builds real-time voice and character AI infrastructure for games, interactive applications, and voice…

AI CompaniesAI in Gaming

Moshi

Moshi is a full-duplex speech-to-speech foundation model developed by Kyutai, a French nonprofit artificial intelligence research laboratory.

AI ModelsConversational AI

Muse Voice Transcribe

Muse Voice Transcribe is a hosted speech recognition model developed by Meta Superintelligence Labs. Meta released it on September 1, 2026 as the lab's first real-time audio perception model.

AI ModelsMeta AI

OpenAI Realtime API

The OpenAI Realtime API is a speech-to-speech interface from OpenAI that lets developers build low-latency, bidirectional voice AI applications powered by GPT-4o and later by the purpose-built gpt-realtime…

Conversational AIDeveloper Tools

PlayHT

PlayHT, later rebranded PlayAI (and reachable at play.ht and play.ai), was an American generative AI voice company that built text-to-speech models, voice cloning tools, and a platform for conversational AI…

AI CompaniesGenerative AI

Retell AI

Retell AI is a voice AI agent platform that enables businesses to build, deploy, and manage AI-powered phone agents for inbound and outbound call automation.

AI AgentsAI Companies

Rime (company)

Rime (also styled Rime Labs, and reachable at rime.ai) is an American artificial intelligence company that builds text-to-speech and spoken-language models tuned specifically for business voice agents…

AI CompaniesGenerative AI

SAG-AFTRA

SAG-AFTRA (the Screen Actors Guild-American Federation of Television and Radio Artists) is an American labor union representing performers in film, scripted television, streaming, video games, commercials, and…

AI EthicsAI Policy & Regulation

Sesame (AI company)

Sesame (formally Sesame AI Labs) is a San Francisco-based artificial intelligence company founded in June 2023, best known for developing the Conversational Speech Model (CSM) and the Maya and Miles voice…

AI CompaniesOpen Source AI

Sindarin

Sindarin (stylized sindarin. and operating as Sindarin Persona) is a small American startup building a real-time conversational voice AI platform for developers.

AI CompaniesConversational AI

Siri

Siri is the voice assistant developed by Apple Inc. that launched on the iPhone 4S on October 14, 2011, making it the first mainstream voice assistant shipped on a smartphone .

AI CompaniesAI Tools & Products

Tavus

Tavus is an American generative AI research company headquartered in San Francisco, California, that develops video generation models and real-time conversational video technology.

AI CompaniesConversational AI

Vapi

Vapi is a voice AI orchestration platform that lets software developers build, deploy, and scale AI phone agents through a programmable API.

AI AgentsAI Companies

Voice Engine (OpenAI)

Voice Engine is a speech-generation and voice-cloning model developed by OpenAI that can produce natural-sounding speech resembling a specific person from a single audio sample as short as 15 seconds.

OpenAISpeech & Audio AI

Voice assistant

A voice assistant is a software agent whose primary interface is spoken language: it listens for a trigger, converts speech to a machine-readable request, decides what the speaker wants, and answers with…

AI Tools & ProductsConversational AI

Voice cloning

Voice cloning is the use of machine learning to generate synthetic speech in the voice of a specific real person (the target speaker) from a sample of their recorded audio.

Generative AISpeech & Audio AI

XTTS (Coqui XTTS)

XTTS (sometimes stylized ⓍTTS, short for "cross-lingual text-to-speech") is an open-weights multilingual text-to-speech model developed by Coqui AI that performs zero-shot voice cloning from short reference…

Open Source AISpeech & Audio AI