Cohere Transcribe
Cohere Transcribe is a family of automatic speech recognition models developed by Cohere and Cohere Labs.
Explore Speech & Audio AI through related topics and the articles other pages reference most.
Articles that also belong to these categories. Counts cover all of Speech & Audio AI.
Showing 1-9 of 9 articles
Cohere Transcribe is a family of automatic speech recognition models developed by Cohere and Cohere Labs.
F5-TTS (short for "A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching") is an open-source text-to-speech and zero-shot voice cloning model released in October 2024 by researchers from…
Massively Multilingual Speech (MMS) is an open-source speech project released by Meta AI in May 2023 that performs speech recognition and text-to-speech synthesis in 1,107 languages and spoken language…
Moshi is a full-duplex speech-to-speech foundation model developed by Kyutai, a French nonprofit artificial intelligence research laboratory.
Parakeet is a family of open automatic speech recognition (ASR) models developed by NVIDIA as part of the NeMo conversational AI toolkit.
Sesame (formally Sesame AI Labs) is a San Francisco-based artificial intelligence company founded in June 2023, best known for developing the Conversational Speech Model (CSM) and the Maya and Miles voice…
Sesame CSM (Conversational Speech Model) is an open weights speech generation model from Sesame AI, a San Francisco startup co-founded by former Oculus chief executive Brendan Iribe.
Voxtral is a family of speech models from Mistral AI. The original open-weight speech-understanding release arrived on July 15, 2025 under the Apache 2.0 license.
XTTS (sometimes stylized ⓍTTS, short for "cross-lingual text-to-speech") is an open-weights multilingual text-to-speech model developed by Coqui AI that performs zero-shot voice cloning from short reference…