Tavus Griffin: a 'video Turing test' passed by the vendor's own study
Tavus says Griffin is the first model to pass a real-time video Turing test, with 48% of its own study's participants fooled. We separate claim from evidence.

Tavus, a San Francisco company that sells face-to-face AI avatars, introduced Griffin on 1 October, in a research post signed by co-founder and CEO Hassaan Raza and head of research Ioannis Patras. Griffin is what Tavus calls a Human Interaction Model: a full-duplex video-to-video system that sees, hears, speaks and moves at the same time. The headline claim is the bold one: "Griffin is the first model to pass the real-time, video Turing test." The company measured that result itself.
What Tavus says it built
Griffin is not a chatbot with a face bolted on. Architecturally it is two engines stitched together: a Continuous Conversational Modeling engine that watches and listens, decides every sub-second whether to speak, react or stay quiet, and emits control signals; and an Audio Visual Generation engine that turns those signals into streaming speech and streaming video. Perception, decision-making and generation run concurrently, so the model can change its mind mid-sentence rather than waiting for a turn to end. Tavus contrasts this with the cascade of speech recognition, language model and speech synthesis that most voice agents use, which wait for the user to stop talking before they start.
The generation claim is unusually broad. Tavus says Griffin generates every pixel in every frame in real time from a single reference image, controlling the whole scene, including the chair the person sits in, the shadows they cast and the background behind them, not only a face and a pair of hands. It says the model back-channels ("mm-hm") while you talk, answers without dead air, stops when you cut in, and tracks how long a silence has lasted so it can speak up on its own. The demos on the page show it coaching a Rubik's Cube solve, playing Simon Says, and guiding someone soldering a motherboard.
The number behind the headline
The pass rate is 48%. In a study Tavus ran, 54 participants were recruited through an independent platform and told they would be matched with another participant for a one-minute video call about what they were looking forward to this year. Their partner was a Griffin-Lite avatar. Afterward, 26 of the 54 believed they had spoken to a real person. The same protocol run against Tavus's previous stack, Phoenix 4.5 with Sparrow-2 and Raven-1, produced 1 of 41, which the post reports as a 2.4% pass rate; its X announcement rounds that to "under 3%".
Almost everything here should be read as the vendor grading its own exam. The study was designed, run and reported by Tavus, on Tavus's avatar, using Tavus's questions; no independent lab checked the framing or the sample. The company is candid that doubters tended to suspect within the first 20 seconds, and that the lowest of five conversation ratings was how naturally the exchange flowed, at 4.9 out of 7. Its own diagram carries a note that undercuts the polish: "The timing in this diagram is illustrative and was not measured from a real session." A pass rate is only as strong as what counts as a pass.
NVIDIA scored the benchmark, not the study
The strongest evidence on the page is not the 48%. It is NVIDIA's Video Full-Duplex Benchmark, which NVIDIA built and scored independently, using published metrics and its own language-model judge. Tavus reports Griffin-Lite first on both tracks. On generation, which scores whether the model's voice and face produce the right behaviour, it logged 3.83 against a 3.92 human reference and 2.80 for the next-best system. On perception, which asks whether the model uses what it sees, it logged 3.73 against a 4.20 human reference, ahead of 15 models evaluated.
| VideoFDB track | Griffin-Lite | Best other system | Human reference |
|---|---|---|---|
| Generation | 3.83 | 2.80 (Gemini 2.5 + Anam) | 3.92 |
| Perception | 3.73 | 3.44 (MiniCPM-o 4.5, audio only) | 4.20 |
Two caveats sit inside those rows. The gap to human performance is still real: Griffin trails humans by 0.09 on generation but 0.47 on perception, the track that depends on reading the person. And a language-model judge is a model, not a measurement of felt experience, however carefully NVIDIA set the rubric.
The safety sentence that carries the argument
Tavus does not bury the problem. "The same properties that make Human Interaction Models powerful interfaces for natural communications between human and machine allow them to deceive a human into believing it is not AI," the post says. On that basis the company says further alignment and safety work is required before wider release, that it is building "safe disclosure" features and working with AI-safety organizations, and that it invites outside evaluators to take part.
Then come the sentences that matter for anyone who has watched a demo go viral. Griffin-Lite is a research preview for a select group of trusted testers, and "Griffin-Lite will not be available for use for customers at this time." There is no weights release, no API and no price. The capability that made a 48% pass rate possible is exactly the capability Tavus says it is not ready to ship.
Reception split on the obvious question
The story was the top post on r/singularity, at roughly a thousand upvotes and about three hundred comments, and the reaction split between amazement and unease. A widely upvoted criticism argued the capability is not needed and its main uses are predatory. On X, the most-liked reply under Tavus's own post, from animator Alex Hirsch, put it bluntly: "What's a single use for this other than fraud manipulation and propaganda." The debate is not whether the model works, but what passing this test is for. The Turing test was never a measure of intelligence; it was a measure of indistinguishability, and Tavus has built a system optimised for exactly that.
What would settle it
Griffin-Lite is not generally available, so nobody outside Tavus and its chosen testers can reproduce the headline number. What would move this from a vendor claim to a fact: the study's protocol and raw data published for outside audit; NVIDIA's VideoFDB leaderboard kept open and re-run on the shipping model rather than the preview; and, most important, a test whose evaluator does not work for the company being evaluated. Tavus credits Baseten, Daily and Cerebrium for infrastructure and Queen Mary University of London for research collaboration, and says it wants evaluators. The door is open; the evidence is not yet through it.


