
이미지: The Decoder
Summary
- The authors of a JAMA opinion piece argue that autonomous AI will soon surpass doctor-AI collaboration, and that regulations should not lock in a "human final decision-maker" requirement
- ChatGPT o3 got the first diagnosis right in 60% of 377 complex cases, versus an average of 15.9% for 20 internists, while GPT-4 alone reasoned diagnoses correctly 92% of the time, higher than the 76% scored by physicians using the same model
- The authors themselves acknowledged that most of the evidence comes from simulations rather than actual patient care, and that physical procedures like surgery and childbirth will remain in human hands for the foreseeable future
- 게재 매체
- 의학저널 JAMA 오피니언
- 주 저자
- 이지키얼 이매뉴얼 — 펜실베이니아대 생명윤리학자, 오바마 건강보험개혁 설계자
- 공동 저자
- 닐 코슬라 — AI 원격의료 업체 Curai Health CEO, 부친 비노드 코슬라는 오픈AI·Curai Health 투자자
- ChatGPT o3 성적
- 복잡 사례 377건 중 60% 첫 진단 정답 (내과의 20명 평균 15.9%)
- GPT-4 실제 환자 사례
- AI 단독 진단추론 92% vs 같은 모델 접근한 의사 76%
- 예측 시점
- 2030년까지 일부 인지 업무에서 자율 AI 실전 배치 가능 전망
- 한계 인정
- 근거 대부분이 단일 과제 시뮬레이션, 실제 환자 진료 데이터 아님
Start with the numbers from the diagnostic showdown
In a contest over 377 complex cases, OpenAI's ChatGPT o3 went head-to-head with 20 internists on getting the first diagnosis right. The result: o3 hit 60%, while the internists averaged just 15.9%. In a separate study using real patient cases, GPT-4's standalone diagnostic reasoning accuracy was 92% — but physicians who had access to the same model actually scored lower, at 76%. These figures form the core evidence behind an opinion piece published in the medical journal JAMA.
Who wrote this opinion piece
Knowing who wrote it helps explain the tone. Lead author Ezekiel Emanuel is a bioethicist at the University of Pennsylvania and one of the architects of the Affordable Care Act under the Obama administration. Co-author Neal Khosla is CEO of Curai Health, an AI-driven telehealth company, and his father, Vinod Khosla, has invested in both OpenAI and Curai Health. Two of the authors, in other words, stand to benefit directly from the future this piece envisions.
Their argument runs squarely against the position of physician groups like the American Medical Association (AMA) and the American College of Physicians (ACP), which have insisted that AI should assist doctors, not replace them. Medical professor Robert Wachter went even further in his book, calling AI-only care the "economy class" of medicine. The JAMA authors treat this hierarchy as an unproven assumption and push back with two counterarguments.
Evidence that AI leads on five tasks
The first counterargument is empirical. The authors say that in most studies published since 2024, AI performance alone has matched or exceeded that of physicians across five core reasoning tasks in medicine: taking patient histories, diagnosis, test selection, guideline-based treatment, and chronic disease management.
Google's conversational system AMIE scored higher than primary care physicians on nearly every measure in simulated patient consultations. Microsoft's diagnostic orchestrator identified the correct diagnosis roughly four times more often than physicians under budget constraints, while also costing less. The authors dismiss most studies with contrary findings as outdated, arguing they either excluded the latest models or used weak methodology.
| Comparison | AI Alone | Human (or Human+AI) |
|---|---|---|
| ChatGPT o3, 377 complex cases | 60% | Average of 20 internists: 15.9% |
| GPT-4, real patient cases | 92% | Physicians using same model: 76% |
A chess analogy
The second counterargument is forward-looking, and it's really the heart of the piece's message. The authors note that while models are improving rapidly, a study published in The Lancet on colonoscopy found that physicians actually lose skill when relying on AI. The implication: the gap will only widen from here.
Once machines clearly surpass humans, the authors argue, a physician acting as a check becomes a source of error rather than a safeguard. They cite a meta-analysis of 106 experiments to support this. When humans are better, collaboration helps — but when AI is better, humans override correct system judgments at the wrong moments, making outcomes worse.
The authors point to chess as a historical parallel. After Deep Blue defeated Kasparov in 1997, human-machine collaborative teams retained an edge for years — but starting around 2017, AI began to surpass even those collaborative teams. If medicine follows the same trajectory, the authors conclude, then enshrining a physician's final say into regulation would lock in a mode of care destined to fall behind. That's why they argue that by 2030, autonomous AI will be ready for real-world use in some — perhaps many — cognitive tasks, and that liability, reimbursement, regulation, and medical education need to be redesigned now.
Limitations the authors themselves acknowledge
Still, the authors concede weaknesses in their own argument. Nearly all the evidence comes from simulations of single tasks rather than data from actual patient care, and the process of handing information back and forth between humans and models is itself a point of vulnerability.
Physical procedures — surgery, childbirth, colonoscopy — remain in human hands for now, since robotics hasn't caught up. Autonomous systems can also fail in ways physicians never do, such as hallucination, internet outages, or cyberattacks. The authors acknowledge these risks need to be weighed against gains in accuracy.
Separately, Google Research announced on August 11 that it had added real-time audio-visual consultation capability to AMIE, its medical AI research system, which showed expert-level performance in simulated consultations across 100 clinical scenarios. A day later, on August 12, results emerged from training Gemini 3.5 Flash using a reinforcement learning technique called ResidencyRL, putting it through a simulated medical residency. It outperformed on measures like information-gathering completeness, but scored lower on hallucination-free responses (12%) and prescription safety (42%). In other words, data in the same vein as the simulation-based scorecards cited in the JAMA opinion piece has been emerging in a steady stream over the past two weeks.
Editor's take
The first thing that strikes me reading this piece is: who wrote it? Emanuel is someone who can shift the weight of U.S. health policy, and the Khosla father-son pair stand to profit directly if this future comes to pass. That doesn't mean their data is wrong. The numbers — o3's 60% versus 15.9%, GPT-4's 92% versus 76% — come from reproducible benchmarks, and the authors honestly disclose their own limitations: reliance on simulations, lack of real patient data. The problem is that readers need to separate the data from the interests behind it, and this piece isn't written in a way that makes that separation easy.
The chess analogy is likely to keep recurring in medical AI debates going forward. The fact that human-machine collaboration held an edge for two decades after Deep Blue suggests that today's "AI-assisted" phase in hospitals may not be a permanent equilibrium but a transitional stage. But there's a crucial difference between chess and medicine: no one dies from a misplayed chess move. The authors don't fully erase that difference either — their own admission of missing real-world patient data is proof of that.
Practically speaking, what hospitals and healthcare startups here should be focused on right now isn't preemptive regulatory positioning, but data. The real weak point of this opinion piece is the gap between simulation and actual clinical practice, and whoever closes that gap with real clinical data first will hold the stronger hand in the next round of regulatory debate. Google's back-to-back releases over the past two weeks — the AMIE audio-visual expansion and the ResidencyRL training results — read the same way: an effort to stockpile simulation-based scorecards ahead of time.
In the coming months, expect formal pushback from physician groups like the AMA and ACP, and expect the fight to center on a single regulatory phrase: whether to enshrine "physician final verification" into law.



