Get all your news in one place.
100's of premium titles.
One app.
Start reading
Medical Daily
Medical Daily
Amelia Palmer

An AI Ran 100 Live Video Doctor Visits Against Real Physicians, and Raters Called It Even on the Core Skills

An AI system spent 100 clinical scenarios watching patients on video, listening to their breathing, and guiding them through physical examination maneuvers on camera. A panel of 20 primary care physicians then graded those consultations against the same encounters run by 10 board-certified doctors using the identical video interface.

On history-taking thoroughness, diagnostic accuracy, management appropriateness, and communication quality, the graders rated the machine on par with the physicians. On eliciting physical signs and guiding patients through virtual examinations, they rated it significantly higher.

Nobody in the study was a real patient.

What Actually Happened in Those 300 Consultations

Google Research and Google DeepMind described the work in a research blog post, covering a video configuration of AMIE, the Articulate Medical Intelligence Explorer, built on Gemini and Project Astra.

The design was a randomized Objective Structured Clinical Examination, the format medical schools use to grade students. Fifteen trained patient actors carried out 300 standardized consultations across three arms: AMIE conducting live video visits, a text-only version of AMIE as a baseline, and 10 board-certified primary care physicians using the same interface. The 100 scenarios spanned five body systems: cardiopulmonary, abdominal, head/eyes/ears/nose/throat, neurological and psychiatric, and musculoskeletal.

An independent panel of 20 experienced primary care physicians scored every consultation using general competency scales plus case-specific criteria written for each scenario. The full technical paper is posted on arXiv and has not been peer-reviewed.

Three Agents Running at Once, Because One Cannot Keep Up

The engineering detail explains the timing. Google states plainly that a single agent cannot currently answer a patient at conversational speed while also reasoning carefully about a differential diagnosis and processing a continuous video and audio feed. Deep reasoning takes time. Pauses erode trust.

So the system was split into three agents running in parallel. A talker agent handles the spoken conversation and keeps latency low. A planner agent works in the background, continuously updating the differential and management plan while flagging any missing information. A perception agent watches the video and audio streams for clinically relevant cues, then contextualizes what it sees within the conversation.

That architecture is what produced the study's most interesting result. The video system did not just match text-based AMIE. It was rated significantly higher, on average, than both the text version and the human physicians at eliciting physical signs and proactively guiding patient actors through examination maneuvers, an advantage that also showed up in the case-specific scoring rubrics.

Where the Sources Disagree, and Why It Matters

Patient actors preferred the video interface to text chat by a wide margin, rating it easier to use and more effective for communicating health concerns.

On the human side of the comparison, the two accounts of this study do not fully align, and readers should be aware of that. The arXiv paper's own abstract states that patient actors preferred AMIE's approach to assessing and explaining conditions, while the physicians were preferred for rapport and partnership building. Google's blog summary, however, says actors rated AMIE favorably on empathy, rapport, and confidence in care compared with both physicians and the text version.

Those are different claims about the same rubric item. The paper is the primary document and it is the more cautious of the two, which is the version worth weighting.

The distinction is not cosmetic. Rapport predicts whether patients disclose embarrassing symptoms, whether they take what they were prescribed, and whether they come back.

The Limitation That Governs Everything Else

Every consultation in this study involved a professional actor performing a scripted condition.

The researchers are explicit about the cost. Actors cannot fully replicate the complexity and unpredictability of real encounters. The scenarios were limited to conditions that can be authentically portrayed through acting, thereby omitting presentations where audiovisual perception would be diagnostically consequential. Automated testing revealed occasional perceptual and reasoning errors, even when overall conversation quality was high, and the system still exhibits intermittent technical faults that disrupt conversational naturalness. The paper lists remaining weaknesses in fine anatomical precision, subtle affective nuance, and high-frequency movements.

Google's own conclusion is that studies with real patients and real conditions are an essential next step before any conclusions about real-world utility can be drawn. Two such efforts are running: a feasibility study of text-based AMIE with Beth Israel Deaconess Medical Center, and a nationwide randomized study with Included Health evaluating AI in virtual care.

AMIE is a research system. It is not available to patients, is not cleared by the FDA, and is not something anyone can consult today. Earlier versions reached expert-level performance in text-based diagnostic dialogue and, more recently, in managing disease across multiple visits, both published in Nature and both in simulated settings.

The distance between an actor performing shortness of breath and a 68-year-old with three chronic conditions and a complicated story remains the entire question. Anyone with a medical concern should be evaluated by a licensed clinician, and consumer AI tools are not a substitute for that.

Key Questions Answered

What did the study test?

Whether an AI system could conduct real-time video medical consultations at a quality comparable to that of primary care physicians, judged by an independent panel of doctors.

Did the AI beat the doctors?

It was rated on par for history-taking, diagnosis, management, and communication, and significantly higher for eliciting physical signs and guiding virtual examinations.

What did the physicians do better?

The paper reports that patient actors preferred the human physicians for rapport and partnership building. Google's blog summary describes the rapport ratings differently, which is worth noting.

Were real patients involved?

No. All 300 consultations used trained patient actors performing scripted scenarios.

Has this been peer-reviewed?

No. The technical paper is posted on arXiv, a preprint server.

Can I use this system?

No. AMIE is a research system, not a product, and it has no regulatory clearance.

What happens next?

Google says validation with real patients and real conditions is required, and two real-world clinical studies of the text-based system are already underway.

Sign up to read this article
Read news from 100's of titles, curated specifically for you.
Already a member? Sign in here
Related Stories
Top stories on inkl right now
One subscription that gives you access to news from hundreds of sites
Already a member? Sign in here
Our Picks
Fourteen days free
Download the app
One app. One membership.
100+ trusted global sources.