Tavus announced Griffin on October 1, 2026, and says it is the first model to pass a real-time video Turing test: 26 of 54 participants (48%) believed they had spoken with a real person during a one-minute video call with a Griffin-Lite-powered persona. This piece examines what the video Turing test actually measured, puts the result in a statistical context, explains the full-duplex design behind it, and looks at the fraud and disclosure questions it raises. Everything below comes from Tavus's own materials and press coverage, since the study is a vendor study and no independent replication has been published.
NVIDIA
MetaWhat Tavus announced
According to Tavus's Griffin page, Griffin is a "Human Interaction Model": a full-duplex video-to-video system that combines continuous conversational modeling with audiovisual generation, so it can perceive and respond at the same time instead of waiting for a turn to end. Business Today describes it as a pipeline that is not "text transcription + LLM + separate synthesis."
Availability is narrow. The Tavus page says Griffin-Lite is a research preview for select testers, is not yet on the Tavus platform, and that a full release is expected after safety concerns are addressed. Traictory's summary notes that no API, weights release or pricing has been announced.
How the video Turing test was run
Per the Tavus page and Traictory, the study had 54 participants recruited through an independent platform, who held one-minute video calls on the topic of what they looked forward to that year. Of the 54, 26 said they believed the persona was a real person. Those who believed it was human averaged 79% confidence, and those who identified it as AI averaged 81%. Tavus reports that over half of participants never suspected AI, and those who did typically suspected within about 20 seconds.
The baseline was Tavus's own previous stack, Phoenix-4.5 for rendering, Sparrow-2 for turn-taking and Raven-1 for perception, which Tavus reports convinced 1 of 41 participants (2.4%). Traictory flags the main caveats directly: the study was designed and run by Tavus, and it is a self-evaluation.

Reading 48% with statistics
Sample size matters here. The following is derived arithmetic by Metir from the reported counts, not a figure from Tavus. A 95% Wilson score interval for 26 of 54 runs from about 35.4% to 61.1%. For the previous system's 1 of 41, the same method gives about 0.4% to 12.6%.
Share of participants who believed the persona was human
Point estimate with a 95% Wilson score interval, derived from the reported counts. The dashed line marks 50%.
Counts reported by Tavus. Intervals are Metir arithmetic, not figures published by Tavus.
Two readings follow from that arithmetic:
- The gap to the old system is large. The two intervals do not overlap, so the improvement over the previous Tavus stack is unlikely to be noise, even at these sample sizes.
- The 50% line sits inside the Griffin interval. The data is consistent with true rates both well below and above one half. A single 48% figure should be read as "roughly a coin flip, with wide error bars."
The design also differs from the classic format. In a standard three-party test, an interrogator speaks with a person and a machine at once and must pick which is human, so chance is 50%. The Tavus study, as described, asked individuals about a single persona. In a single-persona test, a "believed human" rate is not directly comparable to a pick-the-human win rate, and the sources fetched for this piece do not report a matched human-control arm. That is a question to ask, not a finding.
How it compares with text Turing tests
The best-known recent text result is from Jones and Bergen at UC San Diego. In a randomized three-party test with five-minute conversations, GPT-4.5 given a humanlike persona prompt was judged to be the human 73% of the time, and LLaMa-3.1-405B 56%, while GPT-4o and ELIZA scored 21% and 23%.
The numbers look inconsistent with Griffin's 48%, but the setups differ on every axis: text versus video, three-party versus single persona, five minutes versus one, and a pre-registered academic design versus a company study. Video also adds channels such as gaze, lip timing and micro-expressions, which are plausibly harder to fake, though that is an inference and neither study isolates it.
Griffin is the first model to pass the real-time, video Turing test.
Tavus, on its Griffin page
Full-duplex versus turn-based pipelines
Most video agents today are cascaded. Speech is transcribed, a language model writes a reply, a voice is synthesized, and a renderer animates a face. Each stage waits for the one before it, and turn-taking is handled by a separate component. Tavus's prior stack followed this modular pattern, with separate rendering, perception and turn-taking models.
Full-duplex systems are meant to listen and generate at once, so they can handle interruptions, backchannels such as a nod, and silences. Business Today lists facial movement, gestures, shadow control and awareness of silences among Griffin's features.
Tavus also reports figures on the NVIDIA VideoFDB benchmark, which scores full-duplex video interaction: Griffin-Lite at 3.83 out of 5 on generation (human reference 3.92) and 3.73 on perception (human reference 4.20). Per Traictory, the next-best systems scored 2.80 and 3.44 on those tracks. By Metir's arithmetic, Griffin-Lite sits 0.09 below the human reference on generation and 0.47 below on perception, so the perception side, understanding the human, is where the reported gap remains largest. Tavus also reports 0.43 seconds average video latency, which it says is 50% faster than the next-best streaming diffusion model.
Fraud and deepfake implications
Real-time video is the channel where impersonation has already cost real money. In the Arup case reported by CNN, a finance worker joined a video call with what appeared to be the chief financial officer and colleagues, all deepfake recreations, and sent HK$200 million (about $25.6 million) across 15 transactions.
A model that responds interactively, rather than replaying pre-rendered clips, changes the economics of that attack in principle, because a target can ask unscripted questions and still get plausible answers. Tavus itself ties the staged rollout to safety concerns, and Griffin-Lite is not publicly available, so the practical risk today is about the trajectory, not this model. The controls that matter for organizations are procedural: out-of-band callbacks, payment approval limits and verification phrases do not depend on spotting visual artifacts.
Disclosure norms
The EU AI Act's Article 50 is one of the clearest written norms. It requires providers to design AI systems that interact directly with people so that those people are informed they are talking to AI, unless that is obvious to a reasonably well-informed person. It also requires deployers of systems that generate deepfake video to disclose that the content is artificially generated, with a lighter regime for evidently artistic or fictional works. How an interactive, lifelike persona fits those exceptions in practice is an open question regulators have not answered for Griffin specifically.
For teams building with real-time avatars, the design choice is explicit: disclose at the start of the call, not after a user has already formed a belief.
What to watch
- Replication. An independent lab running a larger sample, with a human control and a three-party format, would settle how comparable 48% is to text results.
- Longer calls. Tavus reports that suspicion usually arose within about 20 seconds in the one-minute window. Whether the effect holds over ten minutes is untested in the sources reviewed.
- Release terms. Watch for the safeguards that accompany any wider rollout, including consent for likeness and disclosure defaults.
- Detection. Benchmarks like VideoFDB score generation quality. A matching public benchmark for detecting such output would show how the defensive side keeps up.
For teams comparing assistants across models and modalities, Metir's model-agnostic workspace lets you test claims like these against several systems side by side.
Sources:
- Tavus: Griffin
- Business Today: Tavus's Griffin model passes video Turing test with 48% human success rate
- Traictory: Tavus Griffin video Turing test
- Jones and Bergen, Large Language Models Pass the Turing Test (arXiv:2503.23674)
- CNN: Arup revealed as victim of $25 million deepfake scam
- EU AI Act, Article 50
Image credits
- Alan Turing, 1951 portrait: Elliott and Fry, public domain, Wikimedia Commons.
- Alan Turing at Princeton University, 1936 (hero image): unknown photographer, public domain, Wikimedia Commons.
