Skip to content
Newsroom
Models 1d ago by Rajat Jain

Google AMIE Matches Doctors in Video Consultation Trial

In a randomized 300-session study, AMIE's real-time video consultations matched board-certified primary care physicians on clinical ratings.

Google AMIE Matches Doctors in Video Consultation Trial

Google published August 11 the first-of-its-kind demonstration that its medical AI system, AMIE, conducts real-time video consultations at expert level: in a randomized study of 300 simulated consultations, clinical evaluators rated AMIE (Video) on par with board-certified primary care physicians across core competencies, and patient actors preferred the video experience over text chat.

Key facts

  • Three-arm randomized study: AMIE in video configuration, AMIE in text-only mode (a baseline isolating the audio-visual component), and ten board-certified primary care physicians consulting through the same video interface.
  • 100 clinical scenarios across five body systems: cardiopulmonary, abdominal, head/eyes/ears/nose/throat (HEENT), neurological/psychiatric, and musculoskeletal.
  • Fifteen trained patient actors performed 300 standardized consultations; an independent panel of 20 experienced physicians scored everything against clinical rubrics.
  • AMIE (Video) matched PCPs on history-taking thoroughness, diagnostic accuracy, management appropriateness, and communication quality.
  • AMIE (Video) was rated higher than both PCPs and AMIE (Text) at eliciting physical signs and guiding patient actors through virtual examination maneuvers.
  • Patient actors rated video higher than text on ease of use, effectiveness, empathy, rapport, and confidence in care.

A three-agent architecture

The engineering problem AMIE (Video) solves is latency. A single agent cannot respond at conversational speed while performing deep clinical reasoning and continuously processing audio and visual streams — pauses destroy rapport. Google split the work across three agents running in parallel:

  • a Talker agent handling low-latency spoken interaction with the patient,
  • a Planner agent refining differential diagnoses and management plans in the background,
  • a Perception agent continuously reviewing audio and visual streams for clinically relevant non-verbal cues — signs of distress, physical findings, auditory signals.

Ablations confirmed each agent contributes to the clinical and dialogue-quality metrics. The perception agent is the novel part: it lets the system respond to things the patient shows, not just things the patient says.

Designed for evaluation

Google built an automated evaluation suite around a taxonomy of clinical audio-visual competencies drawn from the medical literature — non-verbal visual cues, auditory signals, and physical examination maneuvers. It combines targeted single-turn audio-visual assessments with multi-turn simulated consultations, including scenarios where the simulation injects visual cues as text (a Parkinson’s patient actor “showing handwriting” is described as holding paper with cramped script). That infrastructure let the team iterate on system design before the expensive human study.

Why the results matter

The shift from text-based to audio-visual clinical AI is where telehealth meets physical examination. Physicians observe gait, register visible discomfort, note breathing, and guide patients through maneuvers — cues that text-based AI medic software (and most telehealth tooling) cannot see. If AMIE’s results hold with real patients, the addressable space expands from triage chats to consultation support that includes observation.

AMIE’s lineage makes the result credible. Text-based AMIE previously demonstrated expert-level diagnostic dialogue and effectiveness as a differential diagnosis aid, both published in Nature, and has been extended toward longitudinal disease management and specialist-level evaluations in oncology. The August 11 milestone is the first time the audio-visual modality — not just the dialogue — is shown to reach expert level in a randomized design.

The right caution: this is a simulated-patient study with professional actors, not a clinical trial. Google is explicit that findings must be validated with real patients, expanded to presentations actors cannot enact, and backed by safety frameworks. Its real-world feasibility work with Beth Israel Deaconess Medical Center (text-based AMIE) and a nationwide randomized study with Included Health are early steps in that direction.

What to watch

  • Replication with real patients and the Included Health randomized virtual-care study.
  • Whether audio-visual AMIE surfaces in Google’s consumer health surface (Health in ChatGPT shows rivals are shipping too).
  • Regulation: an expert-level autonomous agent in clinical dialogue raises licensure and liability questions that benchmarks won’t settle.

Related: Daily AI Brief: August 10, Open-weight model landscape 2026.

#Models #Health #Google DeepMind #Research