
Episode #17
Dr. Adam Rodman on What Happens When AI Outdiagnoses Doctors
What happens when an AI outperforms doctors at diagnosis — and sounds just as confident when it’s wrong? In this episode of Good Medicine, Dr. Rohan Ramakrishna sits down with Dr. Adam Rodman, internal medicine physician at Beth Israel Deaconess, medical educator at Harvard, and creator of the podcast Bedside Rounds. Rodman is a rare hybrid: a historian of medicine who also runs clinical trials of large language models. He traces diagnosis from the ancient Greeks through Tim de Dombal’s 1970s decision-support system to his recent paper in Science, in which a reasoning model outperformed physicians on complex diagnostic puzzles — and then to a silent trial in a real emergency room. Rodman explains why an AI’s growing confidence doesn’t track its accuracy, why doctors working alongside AI sometimes do worse than AI alone, and why building a trustworthy AI doctor looks less like a chatbot and more like Waymo. A clear-eyed conversation about benchmarks, deskilling, medical education, and what physicians are actually for. New episodes are released every other week, wherever you get your podcasts. For more from Roon, visit: https://www.roon.com/ Sign up for our substack: https://rohanramakrishna.substack.com/ Find us on Instagram and X: @roondoctors If you have a question, comment, or suggestion for a future guest, please email us: jane@roon.care. -- (00:00) Intro (02:42) Meet Dr. Adam Rodman (and the pineapple pizza question) (04:17) Podcasting before it was cool: Bedside Rounds (06:16) The bizarre autopsy of Charles II (07:41) Epistemology: how doctors know what they know (09:53) The woman of Thasos and 2,500 years of diagnosis (12:16) Arthur Elstein and what actually makes a good diagnostician (15:31) What is clinical reasoning, exactly? (16:48) Tim de Dombal, AAPHelp, and the first AI in medicine (19:24) Radiology, ground truth, and Geoff Hinton's prediction (21:15) Why AI feels different from the stethoscope (23:34) The Paris Clinical School and the numerical method (25:21) Richard Cabot and the autopsy as the original benchmark (25:47) What makes a good benchmark? (29:14) Regulating generative AI: ARPA-H and ADVOCATE (31:16) The cost of inference and per-token benchmarks (33:22) Inside the Science paper: CPCs and the ER silent trial (38:13) Dr. CaBot, AMIE, and confidence that doesn't track accuracy (43:15) How AI fails differently than humans do (46:46) Where AI outperformance actually comes from (48:41) Why doctors plus AI did worse than AI alone (50:58) New workflows, not superhuman performance (51:59) Is AI the first point of access? (52:56) OpenEvidence vs. the frontier models (59:33) Physician deskilling: real risk or moral panic? (1:04:23) Waymo as the blueprint for an AI doctor (1:05:48) Redesigning medical school for 2035 (1:08:01) Is writing thinking? (1:12:05) Is AI improving care today? (1:15:48) Quick Hits






