One email each morning — yesterday's AI, sortedGet it in your inbox

METAL LAB

Google adds video and voice consultation capabilities to medical AI 'AMIE'

AMIE, previously limited to text conversation, expands to real-time video consultations, demonstrating expert-level performance in simulated consultations

여성이 손동작을 하며 노트북 속 AI 수어 번역과 소통한다

이미지: Google

Summary

  • Google Research has advanced AMIE, its medical AI research system, to enable real-time audio-visual consultation.
  • It is presented as the first case of expert-level performance in randomized-controlled-trial-style simulated consultations.
  • The system focuses on reading nonverbal cues such as the patient's facial expressions, breathing, and movements.
Advancing AMIE Towards Expert-Level Video Consultations
시스템
AMIE(Articulate Medical Intelligence Explorer) — 구글 리서치의 의료 AI 연구 시스템
발표일
2026년 8월 11일, Google Research 블로그
확장 범위
텍스트 기반 대화 → 실시간 음성·영상(오디오-비주얼) 진료 상담
시험 규모
임상 시나리오 100개(5개 신체 계통) · 환자 배우 15명 · 표준화 상담 300건
시스템 구조
Talker·Planner·Perception 3개 에이전트 병렬 — Gemini·Project Astra 기반
핵심 결과
진단 정확도 등은 1차 진료의와 동급, 신체 징후 끌어내기·공감 평가는 상회

An AI that examines patients through the screen

"Could you show me your wrist on the camera?" What if it's not the doctor but the AI asking this during a video consultation? As the patient rotates their wrist, the system observes the movement and facial expressions, and decides what to check next. This is exactly what the new version of AMIE (Articulate Medical Intelligence Explorer), the medical AI research system Google Research unveiled on August 11, does. AMIE, which had been confined to text chat, has now been extended to real-time voice and video consultations, and was validated in a trial involving 100 clinical scenarios, 15 trained patient actors, and 300 standardized consultations. The research team described this as "the first demonstration of expert-level performance in video consultations."

How the trial was designed

The research team created 100 clinical scenarios spanning five body systems: cardiopulmonary, abdominal, ENT/head and neck, neurological/psychiatric, and musculoskeletal. Fifteen trained patient actors performed according to the scenarios, conducting a total of 300 standardized consultations, which were compared across three arms: video-based AMIE, text-based AMIE, and 10 primary care physicians conducting consultations through the same video interface. Who performed better was judged by an independent panel of 20 experienced primary care physicians who were not involved in the consultations themselves. The design mirrors a randomized controlled trial (RCT) approach, notably having participants compete over the entire consultation process rather than the "problem-solving benchmarks" common in medical AI research.

Internally, the system operates through three specialized agents running asynchronously in parallel. Talker, which converses with the patient in real time; Planner, which handles clinical reasoning; and Perception, which reads clinical signals from the screen and audio, each operate independently while exchanging information. This is essentially a division-of-labor approach that mimics how a human doctor simultaneously speaks, thinks, and observes during a consultation. It is built on Gemini and Project Astra, the real-time multimodal technology project.

Results — where it matched and where it exceeded

Evaluation axisResult
History-taking, diagnostic accuracy, management planning, communicationOn par with primary care physicians
Eliciting physical signs, guiding virtual physical examinationsExceeded both physicians and text-based AMIE
Patient actors' preference for ease of use and effectivenessExceeded both physicians and text-based AMIE
Ratings of empathy, rapport, and trust in the consultationExceeded both physicians and text-based AMIE

What's notable is the nature of the areas where it excelled. Eliciting physical signs through the screen and guiding patients through examination movements were areas that the earlier text-based AMIE was fundamentally incapable of performing—and it was precisely in these areas that it received higher ratings than human physicians. The fact that patient actors also gave the AI higher scores for empathy and trust echoes a pattern that has recurred since the earlier text-based AMIE research.

What this means

Anil Palepu, senior research scientist, and Mike Schaekermann, research lead, who led the study, explained that "a consultation is far more than the words exchanged." In an actual examination room, a doctor observes how a patient walks, traces of pain visible on their face, and the rhythm of their breathing, and directly guides physical examination movements. Text-based conversation alone had no way to capture these nonverbal cues—this had been the limitation until now.

Technically, this represents a step up in difficulty. In text-based diagnostic dialogue, exchanging one sentence at a time suffices, but video consultation requires listening to the patient while simultaneously watching the screen and continuing to reason throughout. This signals a shift in medical AI—from competing purely on language model reasoning ability to integrating multimodal signals such as facial expressions, movement, and voice tone in real time. This also connects to Google's recent broader push to expand research infrastructure, including academic tools (PaperVizAgent, ScholarPeer) and climate model evaluation (AIMIP).

Google unveils two AI agents for drawing and reviewing academic paper figures

So what changes now

Video consultation has already become the standard form of telemedicine. The fact that AMIE achieved expert-level performance in exactly this format, even if only by simulation standards, means a new benchmark has opened up to test whether it can move beyond a text chatbot into actual clinical workflows. Still, the limitations are clear. Because this result comes from simulations performed by professional patient actors, the scope is limited to conditions that can be portrayed through acting, and the system also showed intermittent recognition and reasoning errors as well as technical issues. Clinical validation involving actual patients has not yet been conducted. An AI that guides examination movements through a screen has now arrived at the threshold of the lab—and the key to opening the next door will be validation with real patients.