<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="3.10.0">Jekyll</generator><link href="https://localminimum.us/feed.xml" rel="self" type="application/atom+xml" /><link href="https://localminimum.us/" rel="alternate" type="text/html" /><updated>2026-09-28T18:33:27+00:00</updated><id>https://localminimum.us/feed.xml</id><title type="html">LocalMinimum.us</title><subtitle>Resources, reading and notes on AI in surgical care.</subtitle><author><name>David P. Stonko, MD, MS</name></author><entry><title type="html">Machines of Loving Grace, two years later</title><link href="https://localminimum.us/library/2026/09/machines-of-loving-grace-two-years-later/" rel="alternate" type="text/html" title="Machines of Loving Grace, two years later" /><published>2026-09-28T00:00:00+00:00</published><updated>2026-09-28T00:00:00+00:00</updated><id>https://localminimum.us/library/2026/09/machines-of-loving-grace-two-years-later</id><content type="html" xml:base="https://localminimum.us/library/2026/09/machines-of-loving-grace-two-years-later/"><![CDATA[<p>This first feature covers an essay Dario Amodei published in October 2024, <a href="https://darioamodei.com/essay/machines-of-loving-grace">Machines of Loving Grace</a>. If you only have time for one thing this week, read the essay instead of this post.</p>

<p>Amodei is the CEO of Anthropic, the company that makes Claude, and he usually writes about AI risk. This essay was his attempt to describe what happens if things go right. The title comes from a poem Richard Brautigan wrote in 1967, while he was poet-in-residence at Caltech.</p>

<p>Early in the essay, Amodei defines what he calls powerful AI. He avoids the term artificial general intelligence (AGI), which he considers imprecise and weighed down by science fiction and hype. (I don’t like the term either.) His definition is specific. The model would:</p>

<ul>
  <li>be smarter than a Nobel laureate across most fields</li>
  <li>use a computer the way a remote worker does</li>
  <li>work on its own on tasks lasting days or weeks</li>
  <li>run as millions of copies in parallel</li>
</ul>

<p>He calls this a “country of geniuses in a datacenter.” He wrote that it could arrive as early as 2026. In a <a href="https://darioamodei.com/essay/the-adolescence-of-technology">follow-up essay</a> this January, he repeated that it could be one to two years away, or considerably further out.</p>

<p>The first reason I chose to start with this is because I think his prediction around the emergence of powerful AI is now coming true. The second reason is the tone. Much of the current conversation about AI is pessimistic, in medicine and elsewhere, and I have added to it myself. This essay is a detailed optimistic case, and its sections on biology and neuroscience are the most relevant to physicians and surgeons.</p>

<h2 id="the-argument">The argument</h2>

<p>Consider how far surgery, anesthesia, and medicine have come in the last 100 years. In 1925, life expectancy in the US was 59 years; in 2024 it reached 79, the highest on record. Penicillin was not discovered until 1928, so a surgeon of that era operated without antibiotics. A <a href="https://www.thelancet.com/journals/lancet/article/PIIS0140-6736%2812%2960990-8/abstract">meta-analysis</a> of more than 21 million anesthetics found that deaths caused solely by anesthesia fell from 357 per million before the 1970s to 34 per million in the 1990s and 2000s (Bainbridge et al., Lancet 2012).</p>

<blockquote>
  <p><strong>What if the next 100 years of medical progress arrived in the next 10?</strong></p>
</blockquote>

<p>Now think about where surgery might be 100 years from now. Robots might operate semi-autonomously or on their own. New drugs might replace some operations entirely, and cures might exist for many diseases we now manage for life. Amodei’s central claim is that AI could increase the rate of discovery at least tenfold, compressing 50 to 100 years of biomedical progress into 5 to 10. He calls this the compressed 21st century. He has a quite lengthy discussion on what the holdups will be, namely regulatory and clinical trial timelines. Elements of biological experimentation and science are not reducible beyond some minimum time that AI can help us reach but probably not surpass. But he has reasons for why this is more of a speed bump than might first be assumed.</p>

<p>His reasoning is that a small number of tools account for much of the progress in biology: CRISPR, mRNA vaccines, CAR-T, advanced microscopy, and cheap genome sequencing. By his estimate, about one such tool appears each year. He argues that finding them depends mostly on how many talented people are working on the problem. His example is CRISPR, which was known as part of bacterial immunity for about 25 years before anyone saw that it could edit genes.</p>

<p>He then lists outcomes he considers plausible (on variable timelines and probabilities of success):</p>

<ul>
  <li>prevention of nearly all natural infectious disease</li>
  <li>a 95% or greater reduction in cancer mortality and incidence</li>
  <li>cures for most genetic disease</li>
  <li>prevention of Alzheimer’s</li>
  <li>a doubling of the human lifespan</li>
</ul>

<p>He calls these educated guesses, meant to show the scale of change.</p>

<h2 id="what-has-happened-since">What has happened since</h2>

<p><strong>Autonomy.</strong> METR, a nonprofit evaluation group, measures how long a task, in expert human hours, an AI agent can complete on its own. That task length has roughly doubled every seven months for six years. In February 2026, Claude Opus 4.6 reached about 14.5 hours at a 50% success rate. As of May, Claude Mythos measured at least 16 hours, which METR says is beyond what its current task suite can reliably measure. METR’s tasks are mostly software, so these figures do not transfer directly to other professional work.</p>

<p><strong>Research.</strong> In May, Nature published two systems that carry out much of the reasoning in a research project: Robin from FutureHouse and Co-Scientist from Google DeepMind. Robin generated the hypotheses, chose the experiments, and analyzed the data, while human researchers did the bench work. It identified ripasudil, a glaucoma drug, as a candidate for dry macular degeneration. The finding still needs preclinical work before any trial in patients.</p>

<p><strong>Drugs.</strong> In July, Insilico started a Phase III trial of rentosertib for idiopathic pulmonary fibrosis. AI identified both its target and its molecule. In the Phase IIa study, the highest-dose arm gained a mean 98.4 mL of FVC at 12 weeks. An analysis presented at ASCO this year counted 117 AI-enabled drugs that had entered human trials. Of those, 8 had completed Phase 2.</p>

<p><strong>Surgery.</strong> Axel Krieger’s group at Johns Hopkins built SRT-H, a surgical robot based on a transformer architecture similar to what is used in ChatGPT. It performed the clipping and cutting phase of cholecystectomy on eight ex vivo gallbladders with a 100% success rate and no human intervention (<a href="https://www.science.org/doi/10.1126/scirobotics.adt5254">Kim et al., Science Robotics 2025</a>). The authors described this as step-level autonomy and noted that a living patient involves problems this experiment did not address.</p>

<h2 id="why-this-matters-for-surgeons">Why this matters for surgeons</h2>

<p>First, consider the role Amodei gives AI. In surgical research, many of us still use AI as a better regression: we point it at NSQIP or VQI and hope for a higher AUC. Amodei describes AI acting as the principal investigator, and Robin is an early version of that. The idea is close to Richard Sutton’s 2019 essay <a href="http://www.incompleteideas.net/IncIdeas/BitterLesson.html">The Bitter Lesson</a>, which argues that general methods that scale with computing power eventually beat approaches built on hand-coded expert knowledge.</p>

<p>Second, anyone who has run a power calculation will recognize his point about clinical trials. Trials are slow mainly because most new therapies have small effects, and small effects require large samples. Therapies with large effects move faster; his example is the COVID mRNA vaccines, approved in nine months. Rentosertib is an example. AI shortened discovery, but the drug still needs a randomized, double-blind, placebo-controlled Phase III.</p>

<p>Third, he states the limits. Intelligence does not make cells divide faster, and experiments take as long as they take. Surgery happens in that physical world, which is one reason SRT-H has not reached patients. Among his speculative examples are better implanted devices and the ability to regrow or reshape tissue. Anyone who works on grafts and limb salvage should pay attention to both.</p>

<h2 id="the-counterargument">The counterargument</h2>

<p>There are two main counterarguments. First is Niko McCarty’s <a href="https://www.lesswrong.com/posts/2zmxYKKsSaWWjALg2/levers-for-biological-progress-a-response-to-machines-of">response</a> in Asimov Press. He argues that most of what slows biology is biophysical rather than computational, and that there may not be enough high-quality biological data for even a superintelligent model to reach its potential. Early results fit that view: AI has sped up discovery, but of the 117 AI-enabled drugs that had entered human trials by the end of 2025, only 8 had completed Phase 2 (<a href="https://doi.org/10.1200/JCO.2026.44.16_suppl.11072">ASCO 2026</a>). The second is Amodei’s more recent essay, <a href="https://darioamodei.com/essay/the-adolescence-of-technology">The Adolescence of Technology</a> (January 2026), which he frames as the disquieting counterpart to Machines of Loving Grace and which exposes the risks of all of this. He lays out five: AI systems that develop goals of their own, misuse by individuals to cause mass harm (bioweapons especially), misuse by governments to seize power, economic disruption and job loss, and broader destabilization from how fast things change. For each he proposes defenses, from interpretability research and gene synthesis screening to chip export controls and economic policy.</p>

<p>The regulatory question is now open. On September 19, President Trump said he would create an <a href="https://www.forbes.com/sites/maryroeloffs/2026/09/19/trump-vows-to-launch-ai-force-to-safeguard-us-dominance/">“AI Force”</a> and name an AI czar, and rejected calls, including Amodei’s, to slow frontier development, arguing that it would cede ground to China. Supporters of regulation argue that systems this capable need binding safety standards and independent oversight before wide deployment, much as drugs and devices do. Critics counter that when the largest AI companies ask for regulation, they are seeking regulatory capture: licensing and compliance rules that incumbents can afford and startups and open-source developers cannot, which builds a moat around the leaders. How this is resolved will shape how quickly any of Amodei’s predictions reach patients.</p>

<h2 id="my-thoughts">My thoughts</h2>

<p>You don’t need to accept Amodei’s timeline to learn from his essay. He writes with long prose, which can be a little tedious, but it has more clarity on his position than the sound bites you might get from CNBC about AI risks. My summary leaves out most of what makes it worth reading: the reasoning behind each prediction, the limits he acknowledges, and sections on neuroscience, economics, and governance that I didn’t cover at all. The introduction and the biology section take about 20 minutes.</p>

<p><strong><a href="https://darioamodei.com/essay/machines-of-loving-grace">Read Machines of Loving Grace</a></strong></p>

<p><em>From <a href="https://localminimum.us/">The Local Minimum</a>, Issue 1.</em></p>]]></content><author><name>David P. Stonko, MD, MS</name></author><summary type="html"><![CDATA[This first feature covers an essay Dario Amodei published in October 2024, Machines of Loving Grace. If you only have time for one thing this week, read the essay instead of this post.]]></summary></entry><entry><title type="html">AI Safety and Effectiveness in Surgery: Resources and Links</title><link href="https://localminimum.us/library/2026/09/ai-safety-effectiveness-surgery-resources/" rel="alternate" type="text/html" title="AI Safety and Effectiveness in Surgery: Resources and Links" /><published>2026-09-27T00:00:00+00:00</published><updated>2026-09-27T00:00:00+00:00</updated><id>https://localminimum.us/library/2026/09/ai-safety-effectiveness-surgery-resources</id><content type="html" xml:base="https://localminimum.us/library/2026/09/ai-safety-effectiveness-surgery-resources/"><![CDATA[<p>Resources and links from my Faculty Development Series talk on AI safety and effectiveness in surgery.</p>

<h2 id="1-johns-hopkins-ai-resources">1. Johns Hopkins AI resources</h2>

<ul>
  <li><strong>HopGPT</strong>: <a href="https://it.johnshopkins.edu/ai/hopgpt/">it.johnshopkins.edu/ai/hopgpt</a> (or just go through <a href="https://my.jh.edu/">my.jh.edu</a>). Hopkins’ secure, JHED-authenticated AI platform; free to JHU/JHM faculty, staff, and students; frontier models behind Hopkins controls; approved for sensitive data including PHI/PII. The sanctioned (and encouraged by administration) avenue.</li>
  <li><strong>JHU Guidelines for Responsible Use of AI</strong>: <a href="https://it.johnshopkins.edu/ai/guidelines-for-responsible-use-of-ai/">it.johnshopkins.edu/ai/guidelines-for-responsible-use-of-ai</a>. The institutional ground rules for everyday AI use.</li>
  <li><strong>Research IT: Artificial Intelligence</strong>: <a href="https://researchit.jhu.edu/artificial-intelligence/">researchit.jhu.edu/artificial-intelligence</a>. Where to start for research uses; remember IRB approval for PHI/PII work and SIP approval for clinical research.</li>
  <li><strong>“Generative AI at Johns Hopkins: What Faculty Need to Know”</strong>: <a href="https://medicine-matters.blogs.hopkinsmedicine.org/2026/05/generative-ai-at-johns-hopkins-what-faculty-need-to-know/">Medicine Matters, May 2026</a>. The faculty-facing summary of current policy.</li>
</ul>

<h2 id="2-further-reading-and-watching">2. Further reading and watching</h2>

<p>For those generally interested in AI history or how it works, and how it is shaping (and will continue to shape) society.</p>

<ul>
  <li><strong><a href="http://www.incompleteideas.net/IncIdeas/BitterLesson.html">“The Bitter Lesson”</a></strong>: Richard Sutton, 2019. If you only read one thing on this list, read this twice. It’s the most important short essay in AI, and you can read it in 10 minutes on an iPhone instead of scrolling Instagram tonight. He argues that 70 years of history show that general methods that scale with computation beat human-encoded domain expertise. Read it before betting your research program on a hand-built clinical algorithm. He wrote this in 2019 and everything still holds.</li>
  <li><strong><a href="https://kevinmd.com/2026/07/clinical-ai-tools-are-losing-to-general-purpose-models.html">“Clinical AI tools are losing to general-purpose models”</a></strong>: an article I wrote for KevinMD this year about my take on how the Bitter Lesson is now coming for surgery.</li>
  <li><strong><a href="https://youtu.be/WXuK6gekU1Y">AlphaGo (documentary)</a></strong>: free, 90 minutes, and the best emotional introduction to what it feels like when a scalable system passes human expertise. I think this is also on Netflix. If you like this, <a href="https://www.youtube.com/watch?v=d95J8yzvjbQ">this is the follow-up</a>.</li>
  <li><strong>3Blue1Brown: <a href="https://youtu.be/aircAruvnKk">“But what is a neural network?”</a> and <a href="https://youtu.be/wjZofJX0v4M">“Transformers, the tech behind LLMs”</a></strong>: the clearest visual explanations of the mechanics, no math background required. There are about 10 more videos from them if it’s interesting, but this is where to start in that direction.</li>
  <li><strong>Andrej Karpathy: <a href="https://youtu.be/zjkBMFhNj_g">“Intro to Large Language Models”</a></strong>: one hour from an OpenAI co-founder covering how LLMs are trained and where they fail. Also his <a href="https://youtu.be/LCEmiRjPEtQ">“Software Is Changing (Again)”</a>, on what AI does to how we build things.</li>
  <li><strong><a href="https://www.dwarkesh.com/p/richard-sutton">Dwarkesh Podcast: Richard Sutton</a></strong>: the Bitter Lesson’s author arguing LLMs are a dead end; a worthwhile counterweight to the hype in both directions.</li>
</ul>

<h2 id="3-clinical-evidence-discussed-in-the-talk">3. Clinical evidence discussed in the talk</h2>

<p>A few were skipped due to time constraints.</p>

<ul>
  <li><strong><a href="https://doi.org/10.1001/jamanetworkopen.2024.40969">GPT-4 alone outperformed physicians using GPT-4</a></strong>: randomized trial (Goh et al., <em>JAMA Network Open</em> 2024). Physicians with GPT-4 scored 76% on diagnostic reasoning vs. 74% without; GPT-4 alone scored 92%. Access to AI is not the same as skill with it.</li>
  <li><strong><a href="https://www.nature.com/articles/s41591-024-03456-y">The follow-up: the gap is closable</a></strong>: Goh et al., <em>Nature Medicine</em> 2025. In a second RCT on management reasoning, physicians using GPT-4 outperformed those without it. Using AI well is a learnable skill.</li>
  <li><strong><a href="https://www.nature.com/articles/s41591-022-01894-0">TREWS: AI sepsis detection that worked (built at Hopkins)</a></strong>: Adams et al., <em>Nature Medicine</em> 2022. Prospective five-hospital study of ~590,000 patients; when clinicians engaged alerts within 3 hours, in-hospital sepsis mortality fell 18.7% (relative).</li>
  <li><strong><a href="https://jamanetwork.com/journals/jamainternalmedicine/fullarticle/2781307">The Epic sepsis model: same task, opposite outcome</a></strong>: Wong et al., <em>JAMA Internal Medicine</em> 2021. On external validation the widely deployed proprietary model showed AUC 0.63 (vs. reported 0.76 to 0.83) and missed two-thirds of sepsis cases. External validation and implementation decide outcomes.</li>
  <li><strong><a href="https://doi.org/10.1016/S2468-1253(25)00133-5">AI de-skilling is measurable</a></strong>: Budzyń et al., <em>Lancet Gastroenterology &amp; Hepatology</em> 2025. After routine AI-assisted colonoscopy, experienced endoscopists’ unassisted adenoma detection fell from 28.4% to 22.4%. Keep your unassisted reps.</li>
  <li><strong><a href="https://doi.org/10.1038/s41591-026-04431-5">General-purpose LLMs outperform specialized clinical AI tools</a></strong>: Vishwanath et al., <em>Nature Medicine</em> 2026. Benchmark comparison showing frontier general models beating purpose-built clinical AI. <a href="https://www.nature.com/articles/s41591-026-04457-9">Companion paper on physicians’ real-world questions</a>.</li>
</ul>

<h2 id="4-the-most-important-technical-ai-papers-of-the-last-decade">4. The most important technical AI papers of the last decade</h2>

<p>My opinion; all free.</p>

<ul>
  <li><strong><a href="https://arxiv.org/abs/1706.03762">“Attention Is All You Need”</a></strong>: Vaswani et al., 2017. The eight-page paper introducing the transformer, the architecture inside GPT, Claude, Gemini, OpenEvidence, ambient scribes, and AlphaFold.</li>
  <li><strong><a href="https://arxiv.org/abs/2001.08361">Scaling laws</a></strong>: Kaplan et al., 2020. Model error falls as a power law in parameters, data, and compute: a dose-response curve for machine capability, measured before the capability existed.</li>
  <li><strong><a href="https://arxiv.org/abs/2203.15556">Chinchilla</a></strong>: Hoffmann et al., 2022. Model size must be matched to training data; a smaller model trained on more data beats a bigger under-fed one.</li>
  <li><strong><a href="https://arxiv.org/abs/2307.03172">“Lost in the Middle”</a></strong>: Liu et al., <em>TACL</em> 2024. The canonical U-shaped curve; the same fact is recalled worse when buried mid-prompt. Put key data at the beginning or end.</li>
  <li><strong><a href="https://www.trychroma.com/research/context-rot">Context rot</a></strong>: Chroma Research, 2025. Performance degrades as input length grows, well before the advertised context-window limit. See also <a href="https://arxiv.org/abs/2502.05167">NoLiMa</a>.</li>
  <li><strong><a href="https://arxiv.org/abs/2411.15124">Reinforcement learning with verifiable rewards (RLVR)</a></strong>: Lambert et al. (Tülu 3), 2024. Why AI improves fastest where answers can be objectively checked. The same logic should guide which AI projects you spend your time on.</li>
  <li><strong><a href="https://arxiv.org/abs/2310.13548">Sycophancy</a></strong>: Sharma et al., <em>ICLR</em> 2024. Models trained on human feedback learn to agree with you. Ask for the case against your plan to get a better result.</li>
</ul>]]></content><author><name>David P. Stonko, MD, MS</name></author><summary type="html"><![CDATA[Resources and links from my Faculty Development Series talk on AI safety and effectiveness in surgery.]]></summary></entry></feed>