
이미지: Google
Summary
- Google Research has advanced AMIE, its medical AI research system, to enable real-time audio-visual consultation.
- It is presented as the first case of expert-level performance in randomized-controlled-trial-style simulated consultations.
- The system focuses on reading nonverbal cues such as the patient's facial expressions, breathing, and movements.
- 시스템
- AMIE(Articulate Medical Intelligence Explorer) — 구글 리서치의 의료 AI 연구 시스템
- 발표일
- 2026년 8월 11일, Google Research 블로그
- 확장 범위
- 텍스트 기반 대화 → 실시간 음성·영상(오디오-비주얼) 진료 상담
- 시험 규모
- 임상 시나리오 100개(5개 신체 계통) · 환자 배우 15명 · 표준화 상담 300건
- 시스템 구조
- Talker·Planner·Perception 3개 에이전트 병렬 — Gemini·Project Astra 기반
- 핵심 결과
- 진단 정확도 등은 1차 진료의와 동급, 신체 징후 끌어내기·공감 평가는 상회
An AI that examines patients through the screen
"Could you show me your wrist on the camera?" What if it's not the doctor but the AI asking this during a video consultation? As the patient rotates their wrist, the system observes the movement and facial expressions, and decides what to check next. This is exactly what the new version of AMIE (Articulate Medical Intelligence Explorer), the medical AI research system Google Research unveiled on August 11, does. AMIE, which had been confined to text chat, has now been extended to real-time voice and video consultations, and was validated in a trial involving 100 clinical scenarios, 15 trained patient actors, and 300 standardized consultations. The research team described this as "the first demonstration of expert-level performance in video consultations."
How the trial was designed
The research team created 100 clinical scenarios spanning five body systems: cardiopulmonary, abdominal, ENT/head and neck, neurological/psychiatric, and musculoskeletal. Fifteen trained patient actors performed according to the scenarios, conducting a total of 300 standardized consultations, which were compared across three arms: video-based AMIE, text-based AMIE, and 10 primary care physicians conducting consultations through the same video interface. Who performed better was judged by an independent panel of 20 experienced primary care physicians who were not involved in the consultations themselves. The design mirrors a randomized controlled trial (RCT) approach, notably having participants compete over the entire consultation process rather than the "problem-solving benchmarks" common in medical AI research.
Internally, the system operates through three specialized agents running asynchronously in parallel. Talker, which converses with the patient in real time; Planner, which handles clinical reasoning; and Perception, which reads clinical signals from the screen and audio, each operate independently while exchanging information. This is essentially a division-of-labor approach that mimics how a human doctor simultaneously speaks, thinks, and observes during a consultation. It is built on Gemini and Project Astra, the real-time multimodal technology project.
Results — where it matched and where it exceeded
| Evaluation axis | Result |
|---|---|
| History-taking, diagnostic accuracy, management planning, communication | On par with primary care physicians |
| Eliciting physical signs, guiding virtual physical examinations | Exceeded both physicians and text-based AMIE |
| Patient actors' preference for ease of use and effectiveness | Exceeded both physicians and text-based AMIE |
| Ratings of empathy, rapport, and trust in the consultation | Exceeded both physicians and text-based AMIE |
What's notable is the nature of the areas where it excelled. Eliciting physical signs through the screen and guiding patients through examination movements were areas that the earlier text-based AMIE was fundamentally incapable of performing—and it was precisely in these areas that it received higher ratings than human physicians. The fact that patient actors also gave the AI higher scores for empathy and trust echoes a pattern that has recurred since the earlier text-based AMIE research.
What this means
Anil Palepu, senior research scientist, and Mike Schaekermann, research lead, who led the study, explained that "a consultation is far more than the words exchanged." In an actual examination room, a doctor observes how a patient walks, traces of pain visible on their face, and the rhythm of their breathing, and directly guides physical examination movements. Text-based conversation alone had no way to capture these nonverbal cues—this had been the limitation until now.
Technically, this represents a step up in difficulty. In text-based diagnostic dialogue, exchanging one sentence at a time suffices, but video consultation requires listening to the patient while simultaneously watching the screen and continuing to reason throughout. This signals a shift in medical AI—from competing purely on language model reasoning ability to integrating multimodal signals such as facial expressions, movement, and voice tone in real time. This also connects to Google's recent broader push to expand research infrastructure, including academic tools (PaperVizAgent, ScholarPeer) and climate model evaluation (AIMIP).
Google unveils two AI agents for drawing and reviewing academic paper figures
So what changes now
Video consultation has already become the standard form of telemedicine. The fact that AMIE achieved expert-level performance in exactly this format, even if only by simulation standards, means a new benchmark has opened up to test whether it can move beyond a text chatbot into actual clinical workflows. Still, the limitations are clear. Because this result comes from simulations performed by professional patient actors, the scope is limited to conditions that can be portrayed through acting, and the system also showed intermittent recognition and reasoning errors as well as technical issues. Clinical validation involving actual patients has not yet been conducted. An AI that guides examination movements through a screen has now arrived at the threshold of the lab—and the key to opening the next door will be validation with real patients.


