The One-Minute Call That Fooled Half the Room
Can you spot the bot? Bay Area startup Tavus says its new Griffin model — which it bills as the first “Human Interaction Model” — fooled 48 percent of 54 people who chatted with it live into believing they were talking to a real human. According to the New York Post, which covered the announcement this week, the company claims the result means Griffin has passed the real-time video Turing test.
The technical claim is striking. Griffin generates every pixel of every video frame in real time from a single reference image — not just the face, arms and fingers, but the movement of the chair, the shadows it casts and the background, the Post reports. The avatars appear to breathe and blink naturally, react in real time, can interrupt a person, be interrupted without losing their train of thought, and adjust to new subject matter mid-conversation.
“Humans are evolutionarily designed to communicate face to face,” Tavus CEO and co-founder Hassaan Raza said in the announcement. “We speak as much through our words as we do our expressions, tone, gestures and timing.” The company says Griffin eliminates the “reality-breaking” glitches of earlier avatars — like smiling while hearing bad news. In demos, it taught a participant to solve a Rubik’s Cube and played Simon Says, responding to visual cues live.
What “Passed the Turing Test” Actually Means Here
Before crowning a new milestone, read the protocol. An independent teardown lays out what the 48 percent actually measures: participants were recruited through a research platform and told they would be matched with another participant for a one-minute call. Only at the very end were they asked whether it had crossed their mind the partner might not be a person.
The full numbers: 26 of 54 said “real person.” Those who believed it was human were 79 percent confident on average — nearly identical to the 81 percent of those who correctly spotted the bot. Naturalness scored 5.4 out of 7 and trustworthiness 5.6 — but conversation flow, at 4.9, was the weakest axis, and doubters typically grew suspicious within 20 seconds. A matched run of Tavus’s older stack fooled just 1 of 41 people (2.4 percent), so the leap is real even if the framing is generous.
Underneath the marketing sits a harder number: on NVIDIA’s VideoFDB benchmark for full-duplex AI video, Griffin scored 3.83 out of 5 against a human reference of 3.92 — with the next-best published system at just 2.80, according to Digit. In other words, “passed the video Turing test” is Tavus’s label for 26 out of 54 on its own one-minute protocol — a short, prompted, in-house study where participants were primed to expect a human — not a standardized exam.
Why this matters: confidence is the vulnerability
The number that should worry you isn’t 48 percent — it’s the confidence symmetry. Both groups were roughly 80 percent sure, and one of them was wrong. In a hiring pipeline, that asymmetry is a weapon: a fake candidate who survives a one-minute screen, a fake interviewer harvesting personal details, a “colleague” on a video call who isn’t. The same realism that makes Griffin a compelling tutor makes it a near-perfect impostor.
Tavus is no stranger to the consent question — TechCrunch reported in 2024 that the company requires verbal consent statements for its cloning products. But Griffin generates full real-time video from a single reference image, and the consent model is being stress-tested by the capability itself. The open questions now are disclosure standards for AI video and whether detection tooling can keep pace with generation — the same week Google opened its SynthID detector to the public.
Treat the 48 percent as a directional signal, not a milestone: video realism has crossed the point where a casual glance can be trusted. From here on, “seeing is believing” needs an asterisk.
Meanwhile in AI Tech:
Also this week: Google’s public SynthID AI detector — the detection side of the arms race Griffin just escalated.
And: Claude filed a fake homicide tip — and Washington made AI incidents reportable like data breaches — what happens when AI systems act without a human in the loop.


