metir
metir
Docs
Download on App StoreGet it on Google PlayLog inSign up
Back to Blog
Tavus
Video Turing Test
Real-Time Video AI
Deepfakes
AI Disclosure
Full-Duplex

Tavus Griffin Video Turing Test: What 48% Really Means

Tavus says Griffin-Lite fooled 26 of 54 people in one-minute video calls. We check the video Turing test method, the confidence interval and fraud implications.

Metir AI TeamOctober 5, 20268 min read
Tavus Griffin Video Turing Test: What 48% Really Means

Tavus announced Griffin on October 1, 2026, and says it is the first model to pass a real-time video Turing test: 26 of 54 participants (48%) believed they had spoken with a real person during a one-minute video call with a Griffin-Lite-powered persona. This piece examines what the video Turing test actually measured, puts the result in a statistical context, explains the full-duplex design behind it, and looks at the fraud and disclosure questions it raises. Everything below comes from Tavus's own materials and press coverage, since the study is a vendor study and no independent replication has been published.

48%Believed it was human26 of 54 participants
2.4%Previous Tavus stack1 of 41 participants
1 minuteCall lengthPer participant
3.83 vs 3.92NVIDIA VideoFDB generationGriffin-Lite vs human reference
NVIDIA logoNVIDIA
OpenAI logoOpenAI
Meta logoMeta
Companies and systems referenced in this analysis: NVIDIA (VideoFDB benchmark), OpenAI and Meta (text Turing test comparison models).

What Tavus announced

According to Tavus's Griffin page, Griffin is a "Human Interaction Model": a full-duplex video-to-video system that combines continuous conversational modeling with audiovisual generation, so it can perceive and respond at the same time instead of waiting for a turn to end. Business Today describes it as a pipeline that is not "text transcription + LLM + separate synthesis."

Availability is narrow. The Tavus page says Griffin-Lite is a research preview for select testers, is not yet on the Tavus platform, and that a full release is expected after safety concerns are addressed. Traictory's summary notes that no API, weights release or pricing has been announced.

How the video Turing test was run

Per the Tavus page and Traictory, the study had 54 participants recruited through an independent platform, who held one-minute video calls on the topic of what they looked forward to that year. Of the 54, 26 said they believed the persona was a real person. Those who believed it was human averaged 79% confidence, and those who identified it as AI averaged 81%. Tavus reports that over half of participants never suspected AI, and those who did typically suspected within about 20 seconds.

The baseline was Tavus's own previous stack, Phoenix-4.5 for rendering, Sparrow-2 for turn-taking and Raven-1 for perception, which Tavus reports convinced 1 of 41 participants (2.4%). Traictory flags the main caveats directly: the study was designed and run by Tavus, and it is a self-evaluation.

Black and white studio portrait of the mathematician Alan Turing, photographed in 1951
Alan Turing, photographed on 29 March 1951 by Elliott and Fry. The test that carries his name asks whether a machine can be told apart from a person; this photo is of Turing himself and does not depict Griffin or the Tavus study. Public domain, via Wikimedia Commons.

Reading 48% with statistics

Sample size matters here. The following is derived arithmetic by Metir from the reported counts, not a figure from Tavus. A 95% Wilson score interval for 26 of 54 runs from about 35.4% to 61.1%. For the previous system's 1 of 41, the same method gives about 0.4% to 12.6%.

Share of participants who believed the persona was human

Point estimate with a 95% Wilson score interval, derived from the reported counts. The dashed line marks 50%.

Griffin-Lite48.1% (26 of 54), interval 35.4% to 61.1%
Previous Tavus stack2.4% (1 of 41), interval 0.4% to 12.6%
0%35%70%

Counts reported by Tavus. Intervals are Metir arithmetic, not figures published by Tavus.

Two readings follow from that arithmetic:

  • The gap to the old system is large. The two intervals do not overlap, so the improvement over the previous Tavus stack is unlikely to be noise, even at these sample sizes.
  • The 50% line sits inside the Griffin interval. The data is consistent with true rates both well below and above one half. A single 48% figure should be read as "roughly a coin flip, with wide error bars."

The design also differs from the classic format. In a standard three-party test, an interrogator speaks with a person and a machine at once and must pick which is human, so chance is 50%. The Tavus study, as described, asked individuals about a single persona. In a single-persona test, a "believed human" rate is not directly comparable to a pick-the-human win rate, and the sources fetched for this piece do not report a matched human-control arm. That is a question to ask, not a finding.

How it compares with text Turing tests

The best-known recent text result is from Jones and Bergen at UC San Diego. In a randomized three-party test with five-minute conversations, GPT-4.5 given a humanlike persona prompt was judged to be the human 73% of the time, and LLaMa-3.1-405B 56%, while GPT-4o and ELIZA scored 21% and 23%.

The numbers look inconsistent with Griffin's 48%, but the setups differ on every axis: text versus video, three-party versus single persona, five minutes versus one, and a pre-registered academic design versus a company study. Video also adds channels such as gaze, lip timing and micro-expressions, which are plausibly harder to fake, though that is an inference and neither study isolates it.

“

Griffin is the first model to pass the real-time, video Turing test.

Tavus, on its Griffin page

Full-duplex versus turn-based pipelines

Most video agents today are cascaded. Speech is transcribed, a language model writes a reply, a voice is synthesized, and a renderer animates a face. Each stage waits for the one before it, and turn-taking is handled by a separate component. Tavus's prior stack followed this modular pattern, with separate rendering, perception and turn-taking models.

Full-duplex systems are meant to listen and generate at once, so they can handle interruptions, backchannels such as a nod, and silences. Business Today lists facial movement, gestures, shadow control and awareness of silences among Griffin's features.

Tavus also reports figures on the NVIDIA VideoFDB benchmark, which scores full-duplex video interaction: Griffin-Lite at 3.83 out of 5 on generation (human reference 3.92) and 3.73 on perception (human reference 4.20). Per Traictory, the next-best systems scored 2.80 and 3.44 on those tracks. By Metir's arithmetic, Griffin-Lite sits 0.09 below the human reference on generation and 0.47 below on perception, so the perception side, understanding the human, is where the reported gap remains largest. Tavus also reports 0.43 seconds average video latency, which it says is 50% faster than the next-best streaming diffusion model.

Fraud and deepfake implications

Real-time video is the channel where impersonation has already cost real money. In the Arup case reported by CNN, a finance worker joined a video call with what appeared to be the chief financial officer and colleagues, all deepfake recreations, and sent HK$200 million (about $25.6 million) across 15 transactions.

A model that responds interactively, rather than replaying pre-rendered clips, changes the economics of that attack in principle, because a target can ask unscripted questions and still get plausible answers. Tavus itself ties the staged rollout to safety concerns, and Griffin-Lite is not publicly available, so the practical risk today is about the trajectory, not this model. The controls that matter for organizations are procedural: out-of-band callbacks, payment approval limits and verification phrases do not depend on spotting visual artifacts.

Disclosure norms

The EU AI Act's Article 50 is one of the clearest written norms. It requires providers to design AI systems that interact directly with people so that those people are informed they are talking to AI, unless that is obvious to a reasonably well-informed person. It also requires deployers of systems that generate deepfake video to disclose that the content is artificially generated, with a lighter regime for evidently artistic or fictional works. How an interactive, lifelike persona fits those exceptions in practice is an open question regulators have not answered for Griffin specifically.

For teams building with real-time avatars, the design choice is explicit: disclose at the start of the call, not after a user has already formed a belief.

What to watch

  • Replication. An independent lab running a larger sample, with a human control and a three-party format, would settle how comparable 48% is to text results.
  • Longer calls. Tavus reports that suspicion usually arose within about 20 seconds in the one-minute window. Whether the effect holds over ten minutes is untested in the sources reviewed.
  • Release terms. Watch for the safeguards that accompany any wider rollout, including consent for likeness and disclosure defaults.
  • Detection. Benchmarks like VideoFDB score generation quality. A matching public benchmark for detecting such output would show how the defensive side keeps up.

For teams comparing assistants across models and modalities, Metir's model-agnostic workspace lets you test claims like these against several systems side by side.

Sources:

  • Tavus: Griffin
  • Business Today: Tavus's Griffin model passes video Turing test with 48% human success rate
  • Traictory: Tavus Griffin video Turing test
  • Jones and Bergen, Large Language Models Pass the Turing Test (arXiv:2503.23674)
  • CNN: Arup revealed as victim of $25 million deepfake scam
  • EU AI Act, Article 50

Image credits

  • Alan Turing, 1951 portrait: Elliott and Fry, public domain, Wikimedia Commons.
  • Alan Turing at Princeton University, 1936 (hero image): unknown photographer, public domain, Wikimedia Commons.

Ready to experience AI that adapts to you?

metir brings together the world's best AI models in one seamless experience. Start for free today.

Get Started Free
metir

Agentic Operating System for Professionals buried in meetings, emails and docs.

© 2026 metir. All rights reserved.

Product

  • Features
  • Pricing
  • Research
  • Docs
  • Blog
  • Enterprise

Company

  • Docs
  • Support
  • Careers

Legal

  • Terms of service
  • Privacy policy

Personalisation is powerful. Privacy is non-negotiable.

Status: All systems operational