Citation and evidence

Tavus Griffin

4 min full readUpdated 6 references

This article's verification

Report a problem with this article

More

Use this article

Raw MarkdownExplore connections

Improve this page

Suggest editRevision historyDiscussion

Browse categories

AI ModelsGenerative AIMultimodal AIVideo Generation

Cite this article

Griffin is Tavus's full-duplex audiovisual model for face-to-face conversation, announced on October 1, 2026. Tavus calls it a Human Interaction Model; Griffin-Lite is the research preview.[1]

It is distinct from Google DeepMind's Griffin, a language-model architecture introduced in 2024 that combines gated linear recurrences with local attention.[2]

Design

Its conversational controller interprets incoming audio and video continuously, directing speech and visible behavior at sub-second intervals. Streaming generators produce the response while perception continues.[1]

ComponentPreview specification
SpeechAutoregressive diffusion transformer; Tavec codec: 48 kHz, 40 values/frame, 100 frames/second
Video720p, 25 fps; 320 ms chunks
Audio-to-video latency0.43 seconds average on H100 GPUs

Expanded article table

These specifications are reported by Tavus. Chunk duration and audio-to-video latency measure different things: video represented per chunk and delay until received audio affects the picture, respectively.[1]

Video training combines Distribution Matching Distillation, teacher forcing, and Self Forcing.[1] Distribution Matching Distillation is a knowledge distillation method for training a faster generator to approximate a diffusion model's output distribution. The original method uses estimates of the target and generated distributions to direct learning, together with a regression loss. Its published image-generation results describe the underlying technique, rather than an evaluation of Griffin.[3]

Self Forcing addresses a different problem in autoregressive models: during generation, errors in earlier outputs become part of the context used for subsequent outputs. Its training procedure makes the model generate sequences from its own previous outputs, with a loss that evaluates the resulting video. This reduces the mismatch between training on reference sequences and inference on generated sequences. The Self Forcing paper provides the research context for the named training method; its performance figures are not Griffin-Lite results.[4]

VideoFDB evaluation

VideoFDB evaluates conversational behavior using 237 clips of two-person video calls, covering 11 kinds of conversational dynamics. Its perception track tests interpretation of a partner's behavior; its generation track tests the agent's own audiovisual responses. A language-model judge assigns scores using defined rubrics. Perception includes fluency, conversational flow, and visual grounding; generation includes fluency, matching the partner's affect, and appropriate nonverbal cues.[5]

The official NVIDIA leaderboard lists the following results. Overall scores use a 0-5 scale; TOR-Alignment is a timing measure.[6]

TrackSystemOverallTOR-AlignmentMedian latency
PerceptionTavus Griffin Lite3.7373.8%2232 ms
PerceptionHuman reference4.2090%1400 ms
PerceptionMiniCPM-o 4.5, audio and video3.4073%720 ms
PerceptionMiniCPM-o 4.5, audio only3.4472%920 ms
GenerationTavus Griffin Lite3.8362.8%1892 ms
GenerationHuman reference3.9278%900 ms
GenerationGemini 2.5 + Anam2.8044%2840 ms
GenerationGemini 2.5 + Keyframe2.3931%3520 ms

Expanded article table

Griffin-Lite has the highest non-human overall score on both tracks in this comparison. Its higher overall scores do not imply the shortest latency: the perception results include faster systems with lower overall scores. The table also distinguishes the two MiniCPM-o configurations, which receive different scores.[6]

The generation subscores differ: Griffin-Lite scores 4.40 for affect matching against the human reference's 4.14, but 2.83 for nonverbal-cue appropriateness against 3.18. The overall score therefore combines strengths and weaknesses across separate criteria.[6]

TOR-Alignment measures the fraction of examples satisfying the expected timing behavior. Depending on the event, success can mean remaining silent, continuing to speak, yielding, taking a turn, or producing a brief acknowledgment. Consequently, low latency alone does not establish correct conversational timing.[5]

VideoFDB's scope is English-language webcam conversations recorded in the United States and Canada. The paper restricts its evaluation claims to individual conversational events; it does not establish performance in extended, multilingual, or multiparty conversations.[5]

Participant study and preview access

In Tavus's study, 26 of 54 participants (48%) believed Griffin-Lite was human after a one-minute call, compared with 1 of 41 (2.4%) for Phoenix-4.5. Participants expected another participant and learned afterward that their partner was AI.[1]

Access is restricted to selected testers while Tavus develops disclosure and other safety measures.[1]

References

  1. ^1 ^2 ^3 ^4 ^5 ^6Tavus Research. "Griffin: The First Human Interaction Model". October 1, 2026. Accessed October 2, 2026.
  2. ^Soham De et al. "Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models". arXiv:2402.19427, February 29, 2024.
  3. ^Tianwei Yin et al. "One-step Diffusion with Distribution Matching Distillation". CVPR 2024; arXiv:2311.18828, first submitted November 30, 2023.
  4. ^Xun Huang et al. "Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion". arXiv:2506.08009, June 9, 2025.
  5. ^1 ^2 ^3Amrita Mazumdar et al. "VideoFDB: Evaluating Full-Duplex Vision-Speech Capabilities in Conversational Agents". arXiv:2605.30256v1, May 28, 2026, sections 3-4 and Appendix B.
  6. ^1 ^2 ^3NVIDIA Research. "VideoFDB: Evaluating Full-Duplex Vision-Speech Capabilities in Conversational Agents," project page and leaderboard. Accessed October 2, 2026.

Improve this article

Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.

v1 · 827 words · full history

Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify

Research and drafting on this wiki are AI-assisted, under named human editorial standards. How AI is used here

Reviewer note: Independently checked Tavus announcement, NVIDIA VideoFDB leaderboard and paper, Griffin namesake, DMD and Self Forcing papers. Vendor study and preview limits retained; benchmarks separated from architectural latency.

Cite this page: AI Wiki. "Tavus Griffin." aiwiki.ai, updated 2 Oct 2026, fact-checked 2 Oct 2026. CC BY 4.0. https://aiwiki.ai/wiki/tavus_griffin

Suggest edit