Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers (VALL-E)
VALL-E is a zero-shot learning text-to-speech (TTS) system from Microsoft Research that clones a target voice from a 3-second recording and synthesizes new speech in that voice without any per-speaker training.