Category

Speech & Audio AI

77 AI Wiki articles on Speech & Audio AI. The most referenced are Speech Recognition, Whisper and ElevenLabs.

77 articlesRSS

Showing 1-60 of 77 articles

AI Voice Agent

An AI voice agent is a conversational artificial intelligence system that communicates with users through spoken language in real time, holding fluid telephone...

AI AgentsArtificial Intelligence

AssemblyAI

AssemblyAI is an American Speech AI company, founded in 2017 by Dylan Fox in San Francisco, that trains its own automatic speech recognition models and sells...

AI CompaniesVoice AI

Audio Classification Models

See also: Audio Models and Audio Audio classification models are machine learning systems for audio classification, the task of assigning one or more labels to...

Deep LearningMachine Learning

Audio Models

See also: Audio and Models Audio models are machine learning systems that take audio as input, produce audio as output, or both, spanning speech recognition,...

AI Models

Audio-to-Audio Models

See also: Audio Models and Tasks Audio-to-audio models are machine learning systems that take an audio waveform as input and produce a different audio waveform...

AI ModelsMusic & Audio Generation

AudioCraft

See also: Generative AI, Meta AI, and Deep Learning AudioCraft is an open-source generative-audio library released by Meta AI (Fundamental AI Research, FAIR)...

Deep LearningGenerative AI

AudioLM

AudioLM is a framework from Google Research for generating high-quality audio by treating the problem as a language-modeling task over discrete tokens....

GoogleMusic & Audio Generation

Automatic Speech Recognition Models

See also: Audio Models and Speech recognition Automatic speech recognition (ASR) models, also called speech-to-text systems, are machine learning systems that...

AI Models

Best AI Voice Generators (Text-to-Speech)

As of July 2026, there is no single best AI voice generator: the right text-to-speech (TTS) tool depends on the job. For the most natural narration plus the...

AI Tools & ProductsGenerative AI

Cartesia

Cartesia is a San Francisco-based AI company focused on real-time voice synthesis, speech recognition, and state space model (SSM) research.[1] Founded in 2023...

AI CompaniesAI Models

Cohere Transcribe

Cohere Transcribe is a family of automatic speech recognition models developed by Cohere and Cohere Labs. The first release, cohere-transcribe-03-2026, is a...

AI ModelsOpen Source AI

Connectionist Temporal Classification

Connectionist temporal classification (CTC) is a loss function and output layer design for training neural networks to label unsegmented sequences, such as...

Deep LearningMachine Learning

CosyVoice

CosyVoice is a family of open-source multilingual neural text-to-speech (TTS) and voice cloning models developed by the Tongyi Speech Lab (Tongyi SpeechTeam)...

Chinese AIVoice AI

Deepgram

Deepgram is an American voice artificial intelligence company, founded in 2015 and headquartered in San Francisco, that builds proprietary deep learning models...

AI CompaniesNatural Language Processing

Deepgram Nova-3

Deepgram Nova-3 is the third-generation automatic speech recognition (ASR) model developed by Deepgram, a San Francisco-based voice AI company. The model...

AI ModelsVoice AI

Descript

Descript is an artificial intelligence-powered audio and video editing platform that lets users edit media by editing a text transcript: delete a word from the...

AI Tools & ProductsArtificial Intelligence

DolphinGemma

DolphinGemma is an audio language model developed by Google to help scientists analyze the vocalizations of wild dolphins. Built on the same research that...

AI for ScienceGoogle

ElevenLabs

--- Private Industry 2022 Founders London, United Kingdom; New York, United States Key people Text-to-speech models (Eleven v3, Multilingual v2,...

AI CompaniesGenerative AI

ElevenLabs Music

ElevenLabs Music (also marketed as Eleven Music) is an AI music generation product developed by ElevenLabs, the voice AI company founded in 2022 by Piotr...

AI ModelsGenerative AI

ElevenLabs v3

--- ElevenLabs Type June 5, 2025[1] General availability Successor to Eleven Multilingual v2; new model with deeper text understanding[3] Languages ...

AI ModelsGenerative AI

EnCodec

EnCodec is a real-time neural audio codec developed by Meta AI's FAIR (Fundamental AI Research) team that compresses speech, ambient sound, and music into a...

Meta AI

F5-TTS

F5-TTS (short for "A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching") is an open-source text-to-speech and zero-shot voice cloning model...

AI ModelsOpen Source AI

Fireflies.ai

Fireflies.ai is an artificial intelligence meeting assistant platform that automatically records, transcribes, summarizes, and analyzes conversations across...

AI Tools & ProductsArtificial Intelligence

GLM-4-Voice

GLM-4-Voice is an open-weights end-to-end speech-to-speech large language model released in October 2024 by Zhipu AI together with the Knowledge Engineering...

Chinese AIVoice AI

GPT-Live

GPT-Live is a family of full-duplex voice models developed by OpenAI for continuous spoken interaction. OpenAI released GPT-Live-1 and GPT-Live-1 mini on July...

AI ModelsOpenAI

GPT-Realtime / OpenAI Realtime API

GPT-Realtime is a family of speech-to-speech models exposed through the OpenAI Realtime API, a low-latency interface that lets developers build voice agents...

OpenAIVoice AI

GPT-Transcribe

GPT-Transcribe (gpt-transcribe) and GPT-Live-Transcribe (gpt-live-transcribe) are two closed-weight speech-to-text models that OpenAI added to its API on July...

AI ModelsOpenAI

Gladia

Gladia is a French artificial-intelligence company that builds audio infrastructure for developers and voice-product teams, centered on a speech-to-text...

AI CompaniesVoice AI

HuBERT

HuBERT (Hidden-Unit BERT) is a self-supervised learning model for speech representation, introduced by researchers at Meta AI (then Facebook AI Research) in...

Machine Learning

Hume AI

Hume AI is a New York-based artificial intelligence research company and API platform, founded in March 2021 by Alan Cowen, that builds "emotionally...

AI CompaniesConversational AI

Hume Octave 2

Hume Octave 2 is a multilingual emotional text-to-speech model released by Hume AI on October 1, 2025.[1] It is the second generation of the company's Octave...

AI ModelsGenerative AI

Inworld AI

Inworld AI is an American artificial intelligence company headquartered in Mountain View, California, that builds real-time voice and character AI...

AI CompaniesAI in Gaming

Kai-Fu Lee

Kai-Fu Lee (Chinese: 李開復; born December 3, 1961) is a Taiwanese-American computer scientist, venture capitalist, and author who is the founder and CEO of...

Chinese AIPeople

Krisp AI

Krisp AI (formerly 2Hz) is an artificial intelligence company that develops real-time voice AI products, including noise cancellation, accent conversion, voice...

AI Tools & Products

LibriSpeech

LibriSpeech is a freely available corpus of approximately 1,000 hours of 16 kHz read English speech that serves as the standard benchmark for training and...

AI BenchmarksNatural Language Processing

Lyria

Lyria is a family of AI music generation models developed by Google DeepMind, spanning text-to-music synthesis, real-time interactive music performance, and...

AI ModelsGenerative AI

Massively Multilingual Speech (MMS)

Massively Multilingual Speech (MMS) is an open-source speech project released by Meta AI in May 2023 that performs speech recognition and text-to-speech...

Meta AIOpen Source AI

Moshi

Moshi is a full-duplex speech-to-speech foundation model developed by Kyutai, a French nonprofit artificial intelligence research laboratory. Announced on July...

AI ModelsConversational AI

Murf AI

Murf AI is an artificial intelligence company that develops a text-to-speech platform for generating realistic synthetic voiceovers. Headquartered in San...

AI Tools & ProductsVoice AI

Music

See also: Music ChatGPT Plugins AI in music is the use of artificial intelligence, especially machine learning and generative AI, to compose, perform, mix,...

AI Tools & ProductsGenerative AI

NVIDIA Canary

Canary is a family of open speech models developed by Nvidia as part of its NeMo conversational AI toolkit. Canary models perform both automatic speech...

AI ModelsNVIDIA

NVIDIA Parakeet

Parakeet is a family of open automatic speech recognition (ASR) models developed by NVIDIA as part of the NeMo conversational AI toolkit. The models transcribe...

NVIDIAOpen Source AI

NVIDIA Riva

NVIDIA Riva is a GPU-accelerated software development kit and family of containerized inference services for speech and translation AI, built by NVIDIA. Riva...

AI InferenceDeveloper Tools

Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers (VALL-E)

See also: Papers, Text-to-Speech, Microsoft Research VALL-E is a zero-shot learning text-to-speech (TTS) system from Microsoft Research that clones a target...

Generative AIMicrosoft

OpenAI Realtime API

The OpenAI Realtime API is a speech-to-speech interface from OpenAI that lets developers build low-latency, bidirectional voice AI applications powered by...

Conversational AIDeveloper Tools

Otter.ai

Otter.ai is an American artificial intelligence company that builds AI meeting assistants that automatically record, transcribe, summarize, and answer...

AI Tools & ProductsArtificial Intelligence

PlayHT

PlayHT, later rebranded PlayAI (and reachable at play.ht and play.ai), was an American generative AI voice company that built text-to-speech models, voice...

AI CompaniesGenerative AI

Qwen2-Audio

Qwen2-Audio is an audio-language model developed by the Qwen team at Alibaba Cloud, released in August 2024 [1][2]. It accepts audio inputs (human speech,...

Chinese AIMultimodal AI

Resemble AI

Resemble AI is a generative voice and AI security company based in San Francisco, California, and originally founded in Toronto, Canada. It builds tools for...

AI CompaniesGenerative AI

Rime (company)

Rime (also styled Rime Labs, and reachable at rime.ai) is an American artificial intelligence company that builds text-to-speech and spoken-language models...

AI CompaniesGenerative AI

SUPERB

SUPERB, which stands for Speech processing Universal PERformance Benchmark, is a comprehensive evaluation framework designed to measure how well...

AI BenchmarksMachine Learning

SeamlessM4T

SeamlessM4T (short for Massively Multilingual and Multimodal Machine Translation) is a machine translation model released by Meta AI on August 22, 2023. It was...

Meta AINatural Language Processing

Sesame (AI company)

Sesame (formally Sesame AI Labs) is a San Francisco-based artificial intelligence company founded in June 2023, best known for developing the Conversational...

AI CompaniesOpen Source AI

Sesame CSM

Sesame CSM (Conversational Speech Model) is an open weights speech generation model from Sesame AI, a San Francisco startup co-founded by former Oculus chief...

AI ModelsGenerative AI

SoundStream

SoundStream is an end-to-end neural audio codec introduced by Google Research in July 2021 that compresses speech, music, and general audio at low-to-medium...

Google

Speech Recognition

Speech recognition, usually called automatic speech recognition (ASR), is the computational task of converting a spoken-language signal into a sequence of...

Deep LearningMachine Learning

Speechmatics

Speechmatics is a British artificial intelligence company that develops automatic speech recognition (ASR) and voice AI technology for enterprise customers. It...

AI CompaniesVoice AI

SpiRit-LM

SpiRit-LM (also written Spirit LM) is a large language model from Meta AI's Fundamental AI Research (FAIR) group that handles spoken and written language...

Large Language ModelsMeta AI

Stable Audio 2.5

Stable Audio 2.5 is an enterprise focused text-to-audio generation model released by Stability AI on September 10, 2025. It generates production ready music...

AI ModelsGenerative AI

Suno

Suno is a generative artificial intelligence company that develops a text-to-music platform capable of producing complete songs, including vocals,...

AI CompaniesGenerative AI