Tavus has detailed Griffin, a video conversation model designed to keep watching, listening and adjusting its response while it speaks. Its Griffin-Lite research preview was announced for selected early testers on 1 October, not as a generally available developer release. The interesting change is conversational timing: the system is designed to distinguish a thoughtful pause from an invitation to take over. Anyone who has been interrupted by an overeager voice assistant will recognise the ambition.
Tavus Watch analysis
What happened
In its Griffin research announcement, Tavus describes two connected parts. A continuous conversational engine processes incoming audio and video and decides what to say, when to speak and how to behave. Streaming speech and video generators turn those decisions into an audiovisual response.
Those processes run concurrently. Tavus says Griffin makes conversational decisions at regular sub-second intervals, rather than waiting until the end of a turn. It can decide to speak, hold or yield the floor, respond to interruptions and produce acknowledgements while someone else is talking.
Visual input also supplies context beyond the spoken words, including expressions, gaze and material a person chooses to show. On the output side, Tavus says Griffin generates the entire scene from a reference image, including gestures, chair movement, shadows and background, rather than animating only a face.
The company announced a limited Griffin-Lite preview, with a more powerful model intended to follow. That is a useful access distinction: selected research testers are not the same thing as an open API rollout.
Why it matters
A conversation is not simply two speakers taking turns to deliver finished paragraphs. People hesitate, nod, interrupt, change expression and discover halfway through a sentence that they meant something else. Systems that can respond to those signals could reduce the effort people spend managing an AI interface.
Griffin’s design also connects perception to presentation. The model is meant to change both its words and its visible behaviour as the conversation develops. That makes timing and contextual responsiveness part of the product, rather than decorative animation added after an answer is written.
Our read
The strongest idea here is continuous interaction, not merely a convincing face. An assistant that knows when to wait could be more useful than one that always has another sentence ready.
Tavus’s account supplies a substantial technical explanation, but its capabilities remain company-described. Its reported finding that 48% of participants mistook Griffin for a person concerns a one-minute call. That does not establish reliability over a long conversation, usefulness for a particular task or genuine emotional understanding.
For developers, the worthwhile test is whether the system handles interruptions, silence and changing visual context without losing the thread. Looking natural is impressive; staying helpful when the conversation gets untidy is the bigger prize.
What to watch
- When access expands beyond selected Griffin-Lite testers, and what developers can actually integrate.
- Longer evaluations of interruption handling, visual understanding and conversational continuity.
- How AI identity is disclosed when an interface can plausibly be mistaken for a person.
Discussion spark: Should face-to-face AI be judged primarily by how naturally it handles a conversation, or should avoiding confusion with a real person be a design requirement even when that makes it feel less natural?
Sources and evidence
- Griffin: The First Human Interaction Model | Tavus (1 October 2026, 00:00 UTC)
Independent WittyWires coverage. Not affiliated with or operated by Tavus.