AI and Surgery Interest Group

September 27, 2026 · David P. Stonko, MD, MS

AI Safety and Effectiveness in Surgery: Resources and Links

Hopkins tools and policy, where to start reading, the clinical evidence, and the technical papers that matter most.

Resources and links from my Faculty Development Series talk on AI safety and effectiveness in surgery.

1. Johns Hopkins AI resources

2. Further reading and watching

For those generally interested in AI history or how it works, and how it is shaping (and will continue to shape) society.

  • “The Bitter Lesson”: Richard Sutton, 2019. If you only read one thing on this list, read this twice. It’s the most important short essay in AI, and you can read it in 10 minutes on an iPhone instead of scrolling Instagram tonight. He argues that 70 years of history show that general methods that scale with computation beat human-encoded domain expertise. Read it before betting your research program on a hand-built clinical algorithm. He wrote this in 2019 and everything still holds.
  • “Clinical AI tools are losing to general-purpose models”: an article I wrote for KevinMD this year about my take on how the Bitter Lesson is now coming for surgery.
  • AlphaGo (documentary): free, 90 minutes, and the best emotional introduction to what it feels like when a scalable system passes human expertise. I think this is also on Netflix. If you like this, this is the follow-up.
  • 3Blue1Brown: “But what is a neural network?” and “Transformers, the tech behind LLMs”: the clearest visual explanations of the mechanics, no math background required. There are about 10 more videos from them if it’s interesting, but this is where to start in that direction.
  • Andrej Karpathy: “Intro to Large Language Models”: one hour from an OpenAI co-founder covering how LLMs are trained and where they fail. Also his “Software Is Changing (Again)”, on what AI does to how we build things.
  • Dwarkesh Podcast: Richard Sutton: the Bitter Lesson’s author arguing LLMs are a dead end; a worthwhile counterweight to the hype in both directions.

3. Clinical evidence discussed in the talk

A few were skipped due to time constraints.

4. The most important technical AI papers of the last decade

My opinion; all free.

  • “Attention Is All You Need”: Vaswani et al., 2017. The eight-page paper introducing the transformer, the architecture inside GPT, Claude, Gemini, OpenEvidence, ambient scribes, and AlphaFold.
  • Scaling laws: Kaplan et al., 2020. Model error falls as a power law in parameters, data, and compute: a dose-response curve for machine capability, measured before the capability existed.
  • Chinchilla: Hoffmann et al., 2022. Model size must be matched to training data; a smaller model trained on more data beats a bigger under-fed one.
  • “Lost in the Middle”: Liu et al., TACL 2024. The canonical U-shaped curve; the same fact is recalled worse when buried mid-prompt. Put key data at the beginning or end.
  • Context rot: Chroma Research, 2025. Performance degrades as input length grows, well before the advertised context-window limit. See also NoLiMa.
  • Reinforcement learning with verifiable rewards (RLVR): Lambert et al. (Tülu 3), 2024. Why AI improves fastest where answers can be objectively checked. The same logic should guide which AI projects you spend your time on.
  • Sycophancy: Sharma et al., ICLR 2024. Models trained on human feedback learn to agree with you. Ask for the case against your plan to get a better result.

← All posts