Newsletter · · Ashutosh Agarwal
AI Starts Running the Lab as GPT-5 Beats the Bench by 40% - AI Drug Discovery Weekly - Week of July 16, 2026
AI drug discovery newsletter for the week of July 9 to 16, 2026. Ginkgo Bioworks let GPT-5 run a full experiment end to end in its robot cloud lab and beat the best published result by 40 percent, Formation Bio laid out how an AI-native firm could become the next drug giant, and the wet lab, not the model, emerged as the field's real bottleneck.
AI Drug Discovery Weekly
Week of July 16, 2026: AI Starts Running the Lab as GPT-5 Beats the Bench by 40%
TL;DR
-
The most striking result of the week came from Ginkgo Bioworks (DNA). In a project with OpenAI, the company let GPT-5 run a full scientific experiment end to end, design it, read the results, design the next round, inside its robot-run "cloud lab." After six cycles, the AI beat the previous best published result by 40%. That is a shift from AI suggesting molecules to AI doing the science. (The Times Tech Podcast, Jul 9)
-
Formation Bio's CEO laid out the clearest version yet of this year's big question: can an AI-native company become the next Eli Lilly, rather than selling its drugs to the current one? He explained, with real numbers, exactly why no new drug giant has been born since the 1980s, and how he plans to break that pattern. (The Heart of Healthcare, Jul 13)
-
New industry-wide numbers surfaced: there are now 173 AI-originated drugs in clinical trials, up from just 24 in late 2023, and AI-designed drugs are reportedly clearing early trials at 80–90% versus 40–65% the old way. (Equity Mates, Jul 8)
-
The bottleneck keeps moving. It is no longer the AI's intelligence, it is how fast you can actually run experiments in the real world to feed it. LabGenius can dream up 2 million antibody designs but physically test only about 3,000 every five to six weeks. (The Bio Report, Jul 8)
-
The stocks did not care. RXRX, SDGR and LLY all fell week-over-week, Recursion hardest, back down near its 52-week low, even as Lilly collected a fresh wave of analyst price-target hikes on obesity-drug optimism.
What's new
The single most important thing that happened this week: an AI ran a real experiment, not a simulation, and won.
Ginkgo Bioworks co-founder and CEO Jason Kelly walked through it on the Times Tech Podcast. First, the setup. Ginkgo runs what it calls "cloud labs", real laboratories where robots do the physical work (moving tiny amounts of liquid, running samples through machines) so that a scientist can control the whole thing from a keyboard, the way you'd control a website. Kelly's own description of old-fashioned lab work: "five years of standing at a lab bench, moving clear liquids around very, very carefully with things called pipettes... basically like very fancy straws." The cloud lab takes the human hands out of that. (The Times Tech Podcast, Jul 9)
The experiment with OpenAI asked a bigger question: could an AI model, GPT-5 in this case, do the thinking part of science, not just the manual part? Real scientists work in a loop: have an idea, run experiments, look at the surprising bits in the data, design a smarter next round, repeat until you have an answer. Ginkgo let the model run that loop on its own, in an area called cell-free protein synthesis (a clean way of making proteins in a tube). Each round, the model designed roughly 30,000 tiny experiments. Humans set only the "guardrails", always run four copies of each recipe, always include control samples, and otherwise let the model "have at it."
The result, in Kelly's words: "After four rounds, we had beat state of the art. And after six rounds, we beat it by 40 percent." The benchmark was the best "protein per dollar" figure a human academic lab had previously published. Kelly's takeaway is the part worth dwelling on: a scientist "could spin up an agent, like a GPT-5 agent, and tell it, go work for a few weeks... And you could maybe have 10 or 100 or 1,000 of those running at the same time, pursuing lots of different hypotheses." That is a genuinely different way to do science, many tireless digital researchers grinding through experiments in parallel.
He also made an economic argument that matters for the whole field. Automated labs, he claims, are not just faster but cheaper: robots let you pack equipment in tightly (about 3x less floor space) and run 24 hours a day (about 3x more hours), which he rounds to a "nine to 10x improvement" in efficiency doing the exact same work. And a pointed geopolitical note: Western biotech has been shifting lab work to China purely because human scientists are cheaper there. Kelly's bet is that automating the lab is how the West keeps its science industries at all.
The week's second big theme: who ends up owning the drugs. On The Heart of Healthcare, Formation Bio CEO Ben Liu gave the most concrete answer yet to the question flagged last week, whether AI companies will build their own drug empires or just sell tools to old pharma. His starting point is a genuine puzzle: the last drug giants to be founded (Vertex, Amgen, Regeneron) all date to the 1980s. Why none since? (The Heart of Healthcare, Jul 13)
His explanation is all about the math of risk. A venture-backed biotech takes three to five years just to get a drug into human trials, then two to four more years for a mid-stage ("Phase 2") result, and only about 30% of those come back positive. When one does work, a big pharma will typically offer around $3 billion to buy it (roughly one year's peak sales). The alternative, running the expensive final trial yourself, costs a few hundred million and has about a 40% chance of going to zero, against a 60% chance of a $6–9 billion payday. Do the risk-adjusted math and, Liu argues, the rational move is almost always to sell. So promising young drug companies get swallowed before they can grow up.
Formation's answer is structural. It is built as a "hub and spoke" so it can sell a single drug without selling the whole company, and it insists on taking "at least 10 shots", because with a 30% success rate, ten tries gives you an expected three winners, and that portfolio math finally makes it rational to keep building rather than cash out. Liu wants a company that looks like "this generation's Eli Lilly or Novartis" but run by "a few hundred people augmented by these AI systems" instead of 100,000. The company began life as a trials-running service called TrialSpark (the "crawl, walk, run" approach, build the plumbing first), is backed by Sam Altman and Lockheed among others, and has teamed with former Novartis CEO Joe Jimenez on a drug fund called Adidam Bio. Its AI tools have names, "Forge" designs trials and "Apollo" runs them, and it even records its own internal meetings and investment memos to train its models on how the firm makes decisions.
Two more useful data points on the field's momentum, from an investing-focused conversation on Equity Mates: there are now 173 AI-originated drugs in clinical development worldwide, up from just 24 in late 2023, and AI-designed drugs are reportedly passing their first human trials at 80–90%, versus 40–65% for traditionally discovered drugs. The hosts pointed to Insilico Medicine as the proof point, its INS018-055, the first fully AI-designed drug to finish a Phase 2 trial, reached that milestone for about $6 million in 18 months, against an estimated $100 million-plus and six-plus years the old way. Treat the exact percentages as directional (early trials are a small, self-selected sample), but the direction is real. (Equity Mates, Jul 8)
The debate
Where is the real bottleneck now, the AI, or the wet lab that feeds it?
For two years the running argument in this field has been whether the constraint is data, compute, or the raw intelligence of the models. This week's episodes point hard in one direction: the models are no longer the weak link, the physical experiments are.
The clearest illustration came from LabGenius, whose Chief Scientific Officer Angus Sinclair described building antibodies that hit two targets at once (useful, but dangerous if they attack healthy cells or trigger an immune overreaction). His key point: for this kind of design there is barely any public data to learn from, "a paucity of information", so the company has to generate its own in the lab. And here is the pinch: LabGenius's AI can propose up to 2 million designs, but the wet lab can physically test only about 3,000 every five to six weeks. The imagination is effectively infinite; the hands are not. (The Bio Report, Jul 8)
Sinclair was refreshingly blunt about hype, too. He said AI is genuinely delivering in de novo protein and antibody design, the areas with lots of public data to train on, but "still needs to catch up and grow" for the harder, data-poor problems like multi-specific antibodies. His verdict on what it takes to fix that: mostly more data.
This is exactly why Ginkgo's cloud-lab pitch and LabGenius's throughput problem are two sides of the same coin. If the binding constraint is "how many real experiments can I run per week," then whoever industrializes the wet lab, makes it faster, cheaper, and runnable by an AI overnight, controls the pace of the whole field. It also connects to last week's point from DARPA's Mike Koeris, that biology suffers from a shortage of the right kinds of data. The emerging consensus across recent episodes: the frontier is no longer a smarter model, it is a faster, cheaper way to manufacture the experimental data that model needs.
There's a second debate riding underneath, sharpened by Formation Bio: tools or drugs? Ben Liu explicitly chose to own the drugs rather than sell software to pharma, because "if you are running trials for other companies... it's not always how do you do things as cheap and fast. It's how do you extract as much revenue from doing it." That is a direct challenge to the "sell shovels" business model, and a reminder that the most ambitious AI-native players increasingly want to compete with their would-be customers, not serve them.
Stocks in play
Recursion Pharmaceuticals (RXRX), $3.02, down about 20% from last week's $3.76 and back near its 52-week low of $2.77 (range $2.77–$7.18; market cap ~$1.4B). It fell 8.8% on the snapshot day alone, on heavy volume (34.6 million shares versus a 19.3 million average), real selling pressure, not a quiet drift. There was no AI-drug-discovery podcast coverage of Recursion this week and no company news in the window, so this looks like sentiment and positioning rather than a specific event. Worth watching whether the slide continues into its next pipeline update, still expected later this year.
Schrodinger (SDGR), $15.60, down about 7% from $16.79 last week (52-week range $10.95–$23.75; market cap ~$1.2B), off 4.1% on the day. Again, no fresh podcast mentions and no company news this week, the physics-based drug-design pure-play remains quiet on the airwaves.
Eli Lilly (LLY), $1,170.23, down about 4% from $1,216.95 last week even as the analyst community piled on price-target increases (52-week range $623.78–$1,249.45; market cap ~$1.10 trillion). This week's target hikes: BofA to $1,334, Guggenheim to $1,273, UBS to $1,425, Bernstein to $1,385, and Citi to $1,600, all Buy or Outperform, all framed around a strong Q2 and the ongoing rotation out of AI names and into durable-growth pharma. As noted last week, this bullishness is about obesity and GLP-1 drugs, not AI drug discovery. Bernstein did flag a "weaker-than-expected" U.S. launch of one product as a caution.
Lilly also made real news this week, though none of it is an AI-discovery story:
- It agreed to acquire AtaiBeckley for $6.75/share in cash plus up to $2.50/share in milestone-based payments, about $2.8 billion upfront and roughly $1 billion more in potential payments, a ~40% premium. The prize is BPL-003, an intranasal 5-MeO-DMT (a psychedelic) for treatment-resistant depression. This is a neuroscience bet, not an AI bet.
- Its cancer drug Retevmo won full FDA approval for RET fusion-positive solid tumors (precision oncology, tied to a genetic test).
- Small but on-theme: VivoSim (VIVS) received a $5 million milestone from Lilly for dosing the first patient in a Phase 2 study of a drug Lilly bought from it. VivoSim builds lab-grown human-tissue models used to test drugs before human trials, a "generate better preclinical data" tool, and a quiet reminder that the data-generation layer this newsletter keeps circling is starting to show up in real cash milestones.
Read-throughs
- The wet-lab layer may be where value quietly accrues. If the field's true bottleneck is running experiments fast and cheap (Ginkgo, LabGenius), then the companies that industrialize experimentation, automated and cloud labs, high-throughput screening, tissue models like VivoSim's, sit on a choke point everyone else needs. This is the same "picks and shovels" logic, but pointed at the lab, not the model.
- Watch for the "own the drug" pivot spreading. Formation Bio's decision to become a drug company rather than a tool vendor echoes Anthropic's move (covered in prior weeks) into its own drug programs. If more AI-native players choose to compete with pharma instead of supplying it, the eventual buyers of AI drug-discovery software (big pharma) may find their vendors turning into rivals, a strategic risk that isn't yet in anyone's model.
- Tempus (TEM) reframes the opportunity as a data-plumbing problem. Eric Lefkofsky's argument, that up to 20% of U.S. cancer deaths come from patients on the "wrong therapeutic path," and that fixing the data flow could save 100,000+ lives a year, is a reminder that a lot of near-term AI value in medicine is about routing existing patients correctly, not inventing new molecules. Tempus says it is only "20% of the way" through connecting the roughly 8,000–9,000 U.S. hospitals (it reaches ~5,500 today), which frames a long runway. (AI and Healthcare, Jul 10)
- The macro numbers cut both ways. 173 drugs in trials (from 24 two years ago) and 80–90% early-trial success rates are genuinely encouraging, but early-stage success is the easy part, and the hard, expensive failures still happen in late-stage trials. Keep the optimism anchored: cheaper, faster discovery does not yet prove cheaper, faster approval.
What changed vs last week
- Last week's swing question got a concrete answer. The open question was whether the moat is proprietary wet-lab data (the CZI/pharma view) or cheaply manufactured synthetic data (the compute-rich AI-lab view). This week reframed it: LabGenius shows the moat is proprietary data precisely because you can't buy it publicly and can't yet make enough of it, the wet-lab throughput ceiling (2M designs, 3,000 tests) is the real wall. The debate moved from "data vs. compute vs. model IQ" to "the model is fine; the lab is the bottleneck."
- The "AI-lab-as-drug-company" thread advanced from Anthropic to a fuller playbook. Last week it was Anthropic turning Claude Science on its own pipeline. This week Formation Bio supplied the actual business logic, the risk math of why drug startups sell, and the hub-and-spoke and "10 shots" structure designed to beat it. The vertical-integration story now has numbers behind it, not just a headline.
- A brand-new proof point on economics. The Insilico figure ($6M / 18 months for a completed Phase 2 vs. $100M+ / 6+ years) is the sharpest single cost comparison logged so far, and the 24 to 173 clinical-pipeline jump quantifies how fast the approach is scaling.
- Prices diverged from the narrative again. For a second straight week the excitement lives in the podcasts while the tradable names (RXRX, SDGR) drift or fall; Recursion's ~20% weekly drop toward its 52-week low is the notable move. Lilly's strength remains an obesity and rotation story, not an AI one, unchanged from last week.
- New name on the radar: VivoSim (VIVS), via its Lilly milestone, a preclinical human-tissue simulation company, another entrant in the "generate better data" layer.