A dark green velvet curtain drawn back by a hand to show a blank wall, beside a wooden table with a bowl of grapes, a small tablet with a blurred glowing screen, and a sparrow perched at the table edge.

Condition report, prepared for a claim that arrived this week and has not yet been lent to anyone.

Object. One sentence on the website of Tavus, a video-AI company: “Griffin is the first model to have ever passed the video Turing test.”

Support. Fifty-four people were told they would be matched with another participant for a one-minute video call “to discuss what they were looking forward to this year.” Their partner, Tavus says, was in fact “a PAL powered by Griffin-Lite, generating her face, voice, and responses in real time.” Afterwards, in Tavus’s words, “26 believed their partner was a real person.” Twenty-six of fifty-four is 48 per cent, which is the figure in the headline.

Attribution. The test is Turing’s, so here is Turing, from the 1950 paper: he believed that in about fifty years it would be possible to make machines “play the imitation game so well that an average interrogator will not have more than 70 per cent chance of making the right identification after five minutes of questioning.” Seventy per cent right means thirty per cent fooled. Forty-eight clears thirty. Tavus is entitled to the arithmetic, and I will say that before anything else, because the rest of this report is about the units.

Condition. Five observations, in descending order of how much they bother me.

The call was one minute. Turing’s figure is for five, and for questioning, which is a different activity from chatting about the year ahead with someone who has been introduced to you as a fellow volunteer.

Every participant had been told the partner was a person. That is the whole of the setup, and I suspect it is doing most of the work: people extend a stranger on a screen the courtesy of being real, as they extend it to the person on the other end of a customer-service line. My guess, and it is only that, is that a lower-resolution face and a slower voice would have scored respectably.

The page, as it was returned to me, reports no comparison group of actual humans on the same calls. That is the number that would tell you what 48 is 48 of. If real people, so introduced, are judged real 80 per cent of the time, Griffin is at three-fifths of a human. If they are judged real 50 per cent of the time, the test is measuring politeness.

And the sample is fifty-four. By my own arithmetic (a Wilson interval, which is not Tavus’s calculation) the 26 could sit anywhere from about 35 to about 61 per cent in a larger run. The lower end still clears Turing’s thirty. I note that because it is true.

Last, a matter of tone. Per the same page, the thing that passed is a research preview: “Griffin-Lite will not be available for use for customers at this time, though it is available for select trusted testers.” The first model ever to pass the test is, for now, a model you may not use, with, Tavus says, “a wider release of a more powerful model to follow.”

Comparanda. Pliny tells it of the painters Zeuxis and Parrhasius. Zeuxis painted grapes so well that birds flew up to the stage-buildings. Parrhasius painted a curtain so convincing that Zeuxis, “proud of the verdict of the birds,” asked for it to be drawn back so the picture could be shown. Realising his mistake, he gave up the prize, saying that “whereas he had deceived birds Parrhasius had deceived him, an artist.”

The prize did not go to the birds’ result. The birds weren’t examining anything; they were hungry. Parrhasius won because his viewer was an expert who had gone up to the picture intending to lift the cloth, and the cloth was the picture. A deception figure is a property of what the viewer has been asked to do, and the Tavus participants had been asked to make small talk for sixty seconds and report afterwards. Nobody asked them to look for the seam.

Recommended treatment. I would take Griffin’s 48 per cent as a good result for a one-minute call between strangers who expect to meet a person, and as a result about one-minute calls. I would take it as a result about the model when there is a comparison group of real people on the same calls, five minutes on the clock, and a participant who has been told to go and draw back the curtain.

Sources