Newsletter · · Ashutosh Agarwal
Ambient AI Scribes Move From Pilot to Standard Practice in US Hospitals - Vertical Spotlight: Healthcare & Biotech - Week of September 15, 2026
Startups and venture newsletter for the week of September 15, 2026. The Healthcare & Biotech vertical spotlight, where Abridge said it is live across more than 300 health systems, Baptist Health shared ambient AI results for nurses, and podcasts argued over generalist versus specialist medical AI and whether physics beats machine learning in drug discovery.
Vertical Spotlight: Healthcare & Biotech
Week of September 15, 2026: The Week the AI Scribe Stopped Being a Pilot
Ambient AI, the software that quietly listens to a doctor or nurse and writes the note for them, is now inside two-thirds of US hospitals. This week's podcasts stopped asking whether it works and started arguing about the harder questions: who it replaces, whether a medical AI should be a specialist or a generalist, and whether the same wave is about to hit the drug lab.
The Landscape
For two years, "AI in healthcare" mostly meant demos. This week, across a run of podcasts, the conversation had clearly changed. The people talking are no longer founders promising something. They are chief nursing officers and hospital CEOs reporting numbers from systems that are already live, and the numbers are big.
The single clearest signal came from Abridge, an ambient AI company whose software sits in the exam room, listens to the visit, and turns the conversation into a clinical note. On CareTalk, the company's founder and CEO, Dr. Shiv Rao, himself a practicing cardiologist, laid out the scale:
- Abridge is now live across more than 300 large health systems, covering the care of over 250 million Americans and more than half the doctors in the country.
- It runs at giants like Kaiser and the VA, and at the other end of the spectrum in small federally qualified health centers.
That is not a pilot. That is infrastructure. And the second half of the week made clear this is not only a doctor story: it has reached nursing, the largest workforce in any hospital. At Baptist Health, a chief nursing officer named Jerry described what happened when ambient AI was rolled out to nurses (on the Digital Health Talks episode of Healthcare NOW Radio):
- A 52% drop in documentation latency, the lag between something happening to a patient and it being written down where others can act on it.
- Over 70,000 clicks saved for nurses in a short window.
- About two hours of a nurse's shift freed from documentation.
- And the surprise nobody predicted: an 8% jump in patient satisfaction with nurse communication, plus a 100% opt-in rate. Not a single eligible patient refused.
Why did patients like it more? Jerry's answer is worth sitting with, because it inverts the usual fear that technology makes care colder. Because the nurse now speaks the assessment out loud for the AI to capture ("your heart rate is regular, your pupils are reactive, everything's looking good"), the patient hears it in real time and can ask questions. The documentation stopped being a wall between nurse and patient and became part of the conversation.
The throughline of the week, then, is a shift from "does ambient AI work?" to "how fast can we scale it, and what does it do to the workforce?" According to Epic, the dominant electronic-records company, two-thirds of US hospitals running Epic were already using some form of ambient AI documentation as of 2025, a fact John Marchica of Darwin Research Group laid out on Health Care Rounds.
And underneath the deployment story ran a second, deeper one: the same "put an AI agent on the boring work" logic is now moving upstream into how drugs are discovered and how trials are designed, with a real fight breaking out over whether the future belongs to giant general-purpose AI models, to purpose-built specialist models, or to old-fashioned physics.
Companies to Know
Abridge: the ambient AI that went from note-taking to the billing office
Abridge started with the "wedge" everyone knows: writing the doctor's note. But CEO Dr. Shiv Rao made clear on CareTalk that the real game is bigger. The company is pushing into revenue cycle (the messy business of turning a clinical visit into a correct, compliant insurance claim) because, as he put it, doctors "don't get compensated for the care that we deliver. We get compensated for the care that we documented that we deliver."
- A striking technical detail for founders: Rao said 60–70% of Abridge's AI models are built in-house, with only about 30% relying on frontier models like ChatGPT, Claude or Gemini. His logic: "Are you going to point the biggest brain at the smallest, tiniest sort of challenge?" For a narrow task like ordering a medication list correctly, a purpose-trained small model beats renting a giant one, both on quality and on the profit-and-loss statement.
- His pitch for why an AI-native company beats a firm "just sort of hitting a model off the shelf and hoping for the best": "It's the science, it's the engineering, and it's the evals."
- A concrete example of the software going beyond transcription: while Rao was rounding in the hospital on a chest-pain patient, Abridge gently flagged that, given the patient's prior scans and BMI, a PET scan might beat the nuclear stress test he'd ordered. He agreed, changed the order, and figures it saved "a day and a half, two days" of admission.
(CareTalk: Healthcare. Unfiltered., "The Ambient AI Inside 300 Health Systems w/ Dr. Shiv Rao, CEO, Abridge," Sept 11, 2026)
Baptist Health: ambient AI for nurses, with the receipts
Baptist Health is one of the first systems to take ambient documentation, proven with doctors, and push it to nurses, whose work is documented very differently. Beyond the headline metrics above, the CNIO's advice on how to make it stick is the most useful part for anyone rolling this out:
- Build a test environment. Nurses won't try a new tool for the first time in front of a patient "who's counting on me to save their life." A sandbox where they can practice, now baked into onboarding, was "a massive game changer."
- Watch your vendor contract. Many vendors charge differently for physicians versus nurses; getting to a "fiscally palatable" enterprise license is what actually lets you scale.
- Leverage FOMO. Turn your first happy user into a champion for the next unit. "Then it's not an IT initiative. It becomes that clinical initiative."
- Her one regret: "I wish we wouldn't have waited. I wish we would have stepped in and expanded faster."
(Healthcare NOW Radio, "Digital Health Talks: Ambient AI for Nurses From Early Adopters to Standard Practice," Sept 12, 2026)
VI Clinic: "AI agents, not AI doctors"
Founded by physician Boris Jinjolava, VI Clinic is building what it calls an operating system for healthcare operations: a chain of AI agents that run the un-glamorous workflow around a visit. His framing pushes back hard on the "AI will replace doctors" narrative:
- The real problem in healthcare, he argues, isn't decision-making; doctors can make decisions. It's that the data needed to decide is "fragmented" and scattered across systems that don't talk to each other. VI Clinic's job is to "bring the right data to AI" at the point of care.
- The agent chain: a pre-visit agent that interviews the patient beforehand, which then feeds prior authorization, then the care process, then revenue cycle / claims agents that make sure the claim is complete and correctly coded.
- His most interesting claim is about trust between insurers and providers: if payer, provider and patient all "sit on same data" in real time, the endless back-and-forth of denied and resubmitted claims "can be avoided."
- On the build-vs-buy question that's dividing the industry, he sided firmly with using general-purpose models and scoping them tightly: current frontier models "already know everything," so instead of training your own medical model you "just make these specific scope, specific task, give the boundaries, and they behave better." He recounted a large European health system that tried to build its own models and gave up: by the time their months-long procurement finished, "the situation is changed and... it's outdated."
(The Healthcare Growth Cycle Podcast, "AI Agents, Not AI Doctors: Building Healthcare's Operating System with Boris Jinjolava, MD," Sept 8, 2026)
EPAM Life Sciences: using AI to fix the trial before it starts
Dr. Anastasia Christianson, a former AI leader at both Pfizer and Johnson & Johnson and now a managing principal at EPAM, made the case that the biggest prize in drug development isn't a flashy new molecule; it's designing the clinical trial well enough that it doesn't fail.
- Her rallying number: "We know that one in 10 drugs fail... can we actually have five in 10 succeed instead of only one in 10 succeeding?"
- Where she's putting AI to work: trial design ("measure twice, cut once"), synthetic control arms (using data to stand in for some patients rather than recruiting a full placebo group), digital twins and protocol simulation to catch problems before enrolling a single patient, and chatbots to recruit and retain patients.
- Her warning to founders is the episode's title: "Don't start with AI, start with the bottleneck." Too many teams "take a tool... and look for where can I apply it." The winners start from a business problem, define what success looks like and how they'll measure it, and only then reach for AI: "don't try and put a square peg in a round hole."
- And a crucial caveat on the hype: none of this (digital twins, virtual trials) is being used today to replace trials in real patients. It's being used to design better ones.
(Med Tech Gurus, "Don't Start With AI, Start With the Bottleneck," Sept 9, 2026)
Blue Matter: the pharma R&D org chart, redrawn around agents
Tara Alstrott-Sherick leads R&D consulting at Blue Matter, which she said works with "nearly all of the world's top 20 global pharma companies" (and touts a 90%+ repeat-business rate). Her prediction about the drug-discovery workforce was the most concrete forecast of the week:
- In under five years, the R&D workforce will be "smaller in human headcount with much deeper capabilities." People will "think about the agents that they will manage in their workforce versus the people that they will manage."
- She painted a "reimagined org chart" where AI agents sit in the boxes doing the work, and a human's job is guiding and shaping many agents at once, accomplishing far more in the same time.
- Her caution to leaders is a human one: an AI agent is "trained to be quite polite when you challenge it" ("Of course, yes, you'll get right on that"), whereas people "require coaching... feedback... encouragement." Managers will need a "dual way of working."
- The most common false assumption clients still walk in with: that AI transformation "is primarily a technology problem." Her fix: start from "what decisions do you want to make differently and why?", not from picking a platform or vendor. That, she argues, is why so many companies suffer "pilot fatigue."
(Passionate Pioneers with Mike Biselli, "Compounding Expertise, Amplifying Impact: Navigating AI's Role in the Future of Drug Discovery with Tara Alstrott-Sherick," Sept 14, 2026)
BiodinLab: the contrarian betting on physics over machine learning
On AI For Pharma Growth, the founder of Toronto-based BiodinLab made the week's most against-the-grain argument: for the hardest drug-discovery problems, physics beats machine learning. His company runs molecular dynamics simulations "from scratch" plus a proprietary complexity analysis, needing no training data at all.
- The proof point he cited: in antimicrobial research, the approach cut lab time by 50–70%, "and there was no machine learning."
- Why it matters where AI struggles: rare diseases and data-sparse targets, where there simply isn't enough historical data to train a model, and where a model's inability to explain why it picked a molecule becomes dangerous. "You don't want to guess your way into a solution... you might guess a brilliant molecule for a rare disease, but you would like to be able to replicate that."
- His blunt take on the limits of scaling AI: large language models "are already reaching a plateau phase where every additional 1% improvement... requires a 10x investment in training. And there's not enough money in the universe to reach a 99% level of precision."
- His vision to bridge the two worlds: run quick physics simulations on all 250,000+ proteins in the public Protein Data Bank to build a database of "complexity hotspots" (the atoms and residues that actually drive a molecule's behavior) and then train AI on that "physics-heavy" data rather than raw data.
(AI For Pharma Growth, "E234: The contrarian case for physics over data: can deterministic, training-free models beat ML in lead optimization?," Sept 8, 2026)
Recursive (Richard Socher): "AI will do for biology what calculus did for physics"
On The MAD Podcast, AI researcher and Recursive founder Richard Socher made the boldest long-term claim of the week: that the same next-word-prediction trick behind chatbots is a general engine for science.
- His central analogy: science has gotten good at understanding "smaller and smaller pieces," but weaving them back into complex systems (a brain, a microbiome, a cell) is where it stalls. "AI is kind of what calculus did for physics. AI will do for biology."
- Why a chatbot technique can design a protein: to an AI, "it doesn't really care if it's English or a sequence of amino acids." Trained on enough protein sequences, it can "generate completely new kinds of proteins like you can generate new kinds of sentences that have never been in that combination in the training data."
- The honest bottleneck he keeps returning to: data. "We're just nowhere near having enough training data for biology." He pointed to Tahoe Therapeutics and Peril Bio as companies generating the "perturbation" data (knocking out one gene at a time and recording what happens) that a future "virtual cell" model would need to learn from.
- Where he's confident AI will become superhuman: any domain "we can either have a simulation and/or a verification tool" for: games, math, and above all software. Biology is harder precisely because "we cannot yet perfectly simulate a complex cell, let alone tissues or organs."
(The MAD Podcast with Matt Turck, "When AI Improves Itself | Richard Socher (Recursive)," Sept 10, 2026)
The Money and the Market, Briefly
If ambient scribing is the headline, virtual nursing (off-site nurses handling admissions, discharges, monitoring and charting for the bedside team) is the deployment quietly saving hospitals the most money right now. John Marchica of Darwin Research Group pulled together the receipts on Health Care Rounds:
- Advocate Health saved over 43,000 hours across 25 hospitals in a single year and cut enough turnover to avoid $6.3 million in costs.
- Piedmont scaled virtual nursing across 17 hospitals and 2,700 beds in under 18 months.
- Prime Healthcare (55 hospitals) rolled it out after a pilot cut patient falls by 84% on medical-surgical units.
- AdventHealth saw one hospital's RN turnover fall from 46% to 16% year-over-year.
- Parkview Health deployed Abridge for nursing and went from 60,000 to 3 million AI "tokens" processed per month in a matter of months, a raw measure of how fast usage compounds once it sticks.
But Marchica was refreshingly honest about the ceiling. A University of Pennsylvania survey of nearly 900 nurses across 10 states found 57% said virtual nursing did not reduce their workload, even though 53% said care quality improved. And there's "a real undercurrent of nurses suspecting health systems are using it to avoid hiring." His verdict: these tools "shift work around and buy hospitals financial breathing room," but none of it fixes the underlying nursing shortage, which is rooted in a lack of nursing-school instructors. "Whether it reduces nurse burnout or just relocates it somewhere else in the system, that's still an open question."
One more sober note, from Microsoft's Peter Lee on the Lifers podcast: for a hospital, ambient AI is "kind of expensive," and a big system that can already code its visits well may spend more "for no direct obvious revenue benefit," even as it improves the experience. "It's just so hard to get a complete win around these technologies."
One Debate: Should a medical AI be a brilliant generalist, or a trained specialist?
Here is the fight that broke out across three separate podcasts this week, and it has billions of dollars and a lot of patient safety riding on it.
The intuition everyone in healthcare started with is obvious: a model trained specifically on medical data should be safer and more accurate for medicine than a general-purpose chatbot that also knows about sports and recipes. That intuition is now under direct attack.
On the Lifers podcast, Peter Lee, who leads health and life sciences research at Microsoft, took the contrarian side as hard as anyone could. He pointed to a recent Nature study finding that frontier models like Gemini or OpenAI's actually outperformed specialized medical tools built on top of them (the best-known being OpenEvidence). His explanation is genuinely provocative:
"I have always been disturbed by what I see as a prevalent belief, particularly in the world of healthcare technology, that we can get better results and we can reduce hallucinations by training on specialized data. It is exactly the opposite. You don't eliminate hallucinations and you actually create autistic savants."
His reasoning: true human-level intelligence requires exposure "to the full diversity of human thought and expression," including the fact that humans lie, and that telling truth from falsehood matters enormously in medicine. Narrow it down to clean medical data and you get a savant that's brilliant at a counting task but that he "wouldn't trust at all" caring for a human being. Chris Longhurst, CEO of Seattle Children's, agreed completely, while adding the practical wrinkle that health systems "with grocery-store-like margins" can't keep paying to run models with hundreds of billions of parameters in the cloud, so he stays "curious about small language models and local language models."
VI Clinic's Boris Jinjolava landed in the same camp from the operations side: don't build your own medical model, because the frontier models "already know everything"; just scope them tightly, and "AI behaves well when it's framed or bounded" to small, specific tasks.
But the deployment reality cuts the other way. Abridge's Shiv Rao is running one of the largest medical-AI operations on earth, and he builds 60–70% of his models in-house rather than renting frontier models. His argument isn't that specialists are smarter across the board; it's that for a well-defined task, a purpose-built small model is better and cheaper, and that the economics of pointing "the biggest brain at the smallest, tiniest challenge" simply don't work at scale.
So the real, unresolved question isn't "specialist versus generalist" in the abstract. It's this: for a given medical job, is the task narrow and verifiable enough that a small, bounded model wins on quality and cost, or is it open-ended enough that only a model exposed to the full messy diversity of human knowledge can be trusted? Nobody this week claimed to have a clean rule for telling the two apart. Rao, Lee, Longhurst and Jinjolava are all shipping, and all drawing that line in different places.
And running underneath it is the debate BiodinLab reopened from the lab side: whether, for the hardest problems, you should be reaching for any learned model at all, or for physics.
The Human in the Loop
The most thoughtful thread of the week wasn't about accuracy at all. It was about accountability.
Chris Longhurst pointed to a paper he co-authored, published last month in the British Medical Journal, titled "Why are humans still in the loop with advancing AI capabilities?" Their answer boils down to three qualities a machine can't supply: competence (not just knowledge, but the judgment to weigh risks and values that "don't have textbook answers"), communication (the influence to get a patient to actually accept the right treatment), and character (a clinician who "takes professional responsibility for patient care").
Peter Lee sharpened it into a warning that reaches well beyond medicine, the danger of "people hiding behind the data":
"There's a difference between a human being saying, 'I believe this,' versus, 'well, the data says this.' ... one of the things we really need to be on guard for is to not use AI as a way to avoid personal accountability for their role in consequential decisions."
Both were quick to say this doesn't mean rejecting the tools. Longhurst quoted Dr. Bob Wachter's line that he "wouldn't trust a physician that didn't use it," comparing an AI-refusing doctor to one who won't use a stethoscope. Utah is already running a supervised sandbox where an AI autonomously handles certain prescription refills, and Longhurst said the early results look positive. The direction of travel is clearly toward handing more narrow, well-measured tasks to AI so clinicians can focus on what only humans do. The open question every founder in this space is really selling into is where, exactly, that line gets drawn, and who signs their name under the outcome when it's crossed.
Next week: Sales & Customer Service.