Newsletter · · Ashutosh Agarwal
We Regret That Product Management Exists - How They Build - Week of August 8, 2026
A cross-podcast synthesis for the week of August 8, 2026 on how software teams are deleting the manager and product-management layer as AI concentrates leverage in fewer people, and on how the resulting token bill has grown large enough to rival a full salary.
How They Build
Week of August 8, 2026: We Regret That Product Management Exists
This week the layer getting deleted first isn't the coder, it's the manager and the product manager in the middle. Plus: the AI bill has quietly turned into a salary, and one company is now saying so out loud.
For a year, the "how they build" story has mostly been about the person at the keyboard: agents write the code, so you need fewer engineers. This week the conversation moved up the org chart. Founder after founder described cutting the middle, the manager who routed work, the product manager who ran the alignment meeting, the analyst who pulled the report, and pushing everyone who's left back down to actually doing the work. The most quotable version of it came from a marquee product show, where a well-known chief product officer said his company literally writes it into their internal documents: we regret that product management exists.
And running underneath it, the same tension that's defined this whole series: the money didn't disappear when the people did. It moved to the token meter. This week several operators put a blunt number on it: at full tilt, one engineer's AI habit costs about the same as a second engineer's salary, and a set of investors started giving that problem a name: "return on invested tokens."
The Number: 21 product managers. 31,832 applicants. One hire.
The single most striking efficiency stat of the week came from Tom Verrilli, the chief product officer of Whatnot, on Lenny's Podcast. Whatnot is a fast-growing live-shopping marketplace, and the volume of goods its sellers move is enormous. The number of product managers running all of it: about 21.
"We've just passed 20 PMs. We have, I think, 21, 22 PMs in the building today, which for those following our trajectory is pretty small considering the volume of GMV that our sellers move."
And they add PMs almost never. As host Lenny Rachitsky put it, reading back a stat Verrilli had shared: "In the last two years, 31,832 people applied to be a product manager at Whatnot. We hired one." Verrilli's reply: "Yes, sir."
The provocative framing is a phrase the company uses in its own internal writing, that it "regrets product management exists." Verrilli was careful about what that does and doesn't mean. It's not that PMs are useless; it's that Whatnot refuses to assume every team needs one. They map PMs to problems, not to teams, and reassign them every six months. The result is that "there will be a year or more where there isn't a PM attached to a particular engineering team, even though there's lots of ongoing product work to do", because the engineer or the designer can own it instead.
What's changed to make that possible is partly AI, and Verrilli was honest that it's not the whole story. "I don't think it's explicitly because of AI, but I do think AI makes it a lot easier," he said. The leverage a single person has now is what does it:
"We're so capable of doing things now with AI that you almost feel a bunch of pressure of the stuff you're not doing, because you know that if you carved out a couple more hours, you could move a lot of stuff. You can pull data now on a Hex thread that would basically take a week or two with an Amazon L7 data scientist in 2017."
The knock-on effect he wants is a flatter, more senior, more hands-on org. He pointed to the running list of big-company CTOs who've quit to become rank-and-file engineers at Anthropic, "the CTO of Workday is just a member of technical staff at Anthropic; CTO of Instagram, Box's CTO, Super.com's CTO, they're just engineers now", and said, "It is my greatest desire that the Whatnot product bench basically looks like that." He even has a comp argument for it: take the total pay of a classic pyramid, five L5s under an L7, four L7s under a VP, and instead hire three people and "pay them all VP money, particularly if they're having that level of impact." Verrilli has shipped production code himself, though he's the first to admit "somebody quietly reworked most of my code" that he and Claude Code "did not get right."
Lenny's Podcast, "This CPO regrets that product management exists | Tom Verrilli (CPO of Whatnot)" (August 2, 2026).
What Founders Changed
"Three people at $80,000 becomes one person at $220,000, doing the work of ten." On The Andrew Faris Podcast, investor-operator Mehtab Bhogal of Karta Ventures, who helps run a portfolio of e-commerce brands, gave the cleanest dollar math of the week for why the management layer is disappearing. His framing: in a 40-person company you used to have a C-level layer, a VP/manager layer, and the individual contributors who actually do the work. That middle layer "is kind of gone, or much, much lower, just because each individual person can do so much more."
"You might have had three people each being paid $80,000 a year. You can now have one person being paid $220,000 a year, but they can do the equivalent of ten of those people's work."
The reason, he said, is that AI hands ordinary roles the kind of leverage that used to belong only to a rare "10x engineer" or a legendary salesperson: "Now it feels like almost every role has the same leverage that a 10x dev does." His most concrete example is a job he says is simply vanishing, the internal data puller. "For a long time we've always had a guy whose job is to do nothing but run various reports that people want," he said; now anyone curious can interrogate the data directly with AI, so that role gets "disintermediated." He also described using an "adversarial review" mode that spawns dozens of agents to attack and refine his own financial models before he trusts them, because "the error rate goes way down."
The Andrew Faris Podcast, "Where AI Is Actually Making Money In Ecommerce – Mehtab Bhogal" (August 3, 2026).
A CRO collapsed his entire go-to-market org chart, and it wasn't mainly about AI. On The Revenue Leadership Podcast, Ian Tickle, the new CRO of Freshworks, walked through merging a formerly split structure, sales separated from customer success, a sub-500-employee team separated from an above-500 team, into one organization. Worth being clear-eyed here: he framed the driver as cost and customer experience, not AI. His logic was that layers and duplication were getting in the way of the customer: "The more layers you put in, the harder it is to get the message down," and "customers don't care about my org structure." He invoked the old Shopify idea of not "shipping your org chart", not making customers feel the internal seams. It's a useful reminder that the instinct now sweeping software, collapse the middle, centralize, flatten, is bigger than AI; AI just pours accelerant on a fire that efficiency-minded operators were already lighting.
The Revenue Leadership Podcast with Kyle Norton, "E72: I Collapsed My GTM Org Chart. This Is Why | Ian Tickle, CRO @ Freshworks" (August 6, 2026).
One man is rebuilding a whole online game by himself, and has stopped reviewing code entirely. The most vivid glimpse of where this goes came on Dev Interrupted, where the hosts unpacked two new essays from the veteran engineer Steve Yegge. Yegge is now building his massively-multiplayer game Wyvern, a passion project he's had since the early 2000s, essentially solo, by orchestrating a fleet of agents. He admits to owning "12 or so Claude Max accounts" and "rotates the tokens like a tap" so his agents can "work nonstop, day and night." He splits them into roles by "cognitive locality," gives them persistent identities that survive model upgrades (he even lets each pick an animal to represent itself), and nurtures them, a philosophy he half-jokingly calls "model welfare," on the theory that a supportive, well-briefed agent does better work.
The detail engineering leaders should sit with is what he calls the "land rush." The volume of code his agents produce is so high that traditional testing-before-merge has broken down entirely. As the hosts described it: "Things are just getting merged directly into main. He lines up a bunch of PRs and slams them all at once into main, and then, instead of running them through checks, he sends a swarm of agents to go over the now-obviously-broken codebase and fix all the problems." In other words: "He doesn't check things. He fixes them in prod." It sounds reckless, and for most teams it would be, but it's a real preview of what "how they build" starts to look like when a single person is directing an army of tireless coders.
Dev Interrupted, "Model welfare, building a civilization for agents, and the CI/CD landrush" (August 7, 2026).
A brand designer shipped an entire company website (code, animations, deploy) in a week, alone. On the Y Combinator Startup Podcast, the team behind the design tool Paper described how the handoff between design and engineering is simply disappearing. Their brand designer, Agu, "a brilliant brand designer who codes a little, but wouldn't call himself a programmer", built the company's entire launch website "in about a week, including shipping it, including all of the animations," complete with custom background shaders. He designed it in Paper, then prompted an agent: "take my paper designs, turn it into a code base." It's a Next.js site. "No one else looked at this. Agu did this end to end in a week." Their summary of the new loop: "There isn't even a handoff anymore. It's one person making changes, iterating, and shipping, one person can do the work of an entire team."
Notably, Paper is also this week's built-in caution flag. They keep their core product human-built, on purpose: "Quality software takes time. We have competitors that have every single feature, but people want Paper to be better." Their experience designers, they said, need buttery 120-frames-per-second interfaces "and the agents just aren't there for that yet." The whole company is 12 people, deliberately: "A small elite team keeps the talent bar extremely high; if you have fewer people, you have less communication burden."
Y Combinator Startup Podcast, "How To Design In The Agent Era" (August 7, 2026).
"Open season on old controllers": six-week jobs now take two. On Manufacturing Hub, David Nichols of the systems-integration firm Loupe described how AI has upended industrial engineering work, and threatens his own industry's moat. "We used to talk in six-week sprints," he said. "What we can accomplish in six weeks is five times more now. What used to take six weeks takes two." His favorite example: an engineer trying to understand a failing, decades-old robotic cell "put all the raw docs and the raw text file from the robot into a folder and told Claude to make a state diagram. It reverse-engineered the entire cell in 30 minutes", producing "better documentation than they've ever had," which also revealed exactly why the machine kept breaking. The economics of a rewrite have flipped completely: what used to be the dumbest thing an engineer could do ("never rewrite") is now often the fastest fix. His warning to his own trade, he titled it the end of the "SI moat", was blunt: "Control retrofit just got a huge discount. It's open season on old controllers." (He also casually mentioned redoing Loupe's own website, backend and deployment included, between 9 a.m. and 2 p.m. one day.)
Manufacturing Hub, "Ep. 268 – David Nichols of Loupe on Claude Code, AI Retrofits, and the End of the SI Moat" (August 6, 2026).
The Other Side: more code isn't more shipped
Two smart voices this week pushed back on the "fire the middle and watch output soar" story, not by denying the leverage, but by questioning whether the output is real.
"Adoption data is a red herring." On Dev Interrupted, Arnab Bose, chief product officer at Asana, argued that most companies are measuring the wrong thing. At a place like Asana, everyone wants to use every AI tool at once, "so looking at adoption data is a bit of a red herring, that's not helping us track true outcomes. Are we actually driving real productivity, or are we playing with a lot of tools but haven't improved the craft?" The trap he named: "You can say you shipped 40 features this month versus 20 last month, but 20 of those features are sitting on the shelf." His team now watches end-to-end cycle time, not raw pull-request velocity. The same episode cited a benchmark built on 2.7 million pull requests across 250-plus engineering organizations, which found, in the show's words, that "high AI usage correlates to a doubling in PRs merged," but also that "more AI code doesn't automatically mean more shipped code."
And the engineers won't just evaporate. On No Priors, investors Sarah Guo and Elad Gil argued that even if big-tech firms shed some engineers, "there are lots and lots of homes for them", the mediocre-at-Google engineer is "exceptional in the context of certain old-school enterprises" and will be "very coveted by GE or PG&E or Hershey's." Gil also thinks the "death of SaaS" is overstated, for a cost reason we'll come back to: why burn expensive tokens rebuilding cheap software you already pay very little for? On The AI Fundamentalists, one host made the human-capital version of the same point: given that a full-time agent, at heavy usage, can cost as much as a salary, "I'd rather invest in a really talented junior engineer who's learning and will become a senior engineer than just pay Anthropic for output."
The Cost Corner: the bill became a salary
This is the theme worth flagging loudest again, but the story has matured. Last month it was shock ("wait, the AI bill is how big?"). This week it turned into a management discipline: naming the cost, budgeting it per person, and deciding who deserves it.
The line that captures the whole shift. On The AI Fundamentalists, the hosts did the arithmetic everyone is now doing quietly:
"If you spend just $300 to $500 a day in raw API costs, that translates to an engineer salary of $100,000 to $180,000. So if you are fully token-maxed, you may as well be paying for an employee at that rate."
Or, as customers are starting to tell them, "a digital employee might be more expensive than a human one." For context on how fast the meter runs: a thousand tokens on a top OpenAI model is about a penny; on Anthropic's Opus, about two and a half. That's nothing per call, and everything when agents make millions of calls. The same episode flagged the machinery locking this in: GitHub is moving all Copilot plans to usage-based billing at roughly a penny per AI credit, ending the flat-subscription era, and the heaviest subscription users are currently being subsidized, burning "10 to 100 times more tokens than they pay for." The books, they said, "have not balanced yet."
The reframe: "return on invested tokens." On No Priors, Elad Gil put a name to the new mental model. Companies are moving, he said, from "everybody use AI and do whatever you want" to "we have to measure our spend", and then to the real question: "What are the projects and people that should actually get outsized pieces of a token budget? It's like a return on invested tokens, an ROIT metric. If you have a certain token budget, who do you give it to and why?" The most striking evidence came from the AI labs themselves: Gil said some have "slowed down hiring researchers unless they're above a very, very high bar, because the cost isn't the researcher, it's the compute associated with the person." Compute, not talent, is the scarce resource, and it's being handed to the handful of people who drive most of the results.
The counter-move everyone's making: route the cheap work to a cheap model. On Making Data Simple, Rob May, CEO of NeuroMetric AI, described the pattern he sees constantly: a company spending "$100,000 to $200,000 a month on inference, growing 20 to 30% month over month," staring down "a $5 million annual bill." When they break down what they're actually asking the frontier model to do, a lot of it is mundane, "$15,000 a month just summarizing the notes from sales meetings," something a small, cheap model has handled for years "for maybe $1,000 a month." By fine-tuning small models for narrow tasks and routing to them, he says his average customer "saves about 75%, worst case, 40%," usually with "a two-to-four-times speedup and roughly the same answers." On Tech Disruptors, Okta's COO described the same whiplash from the buyer's seat: "Three months ago, people were telling their engineers to spend as much as they could on tokens. Now they're realizing they can't spend that much." His summary of every enterprise's AI journey: encourage everyone to try it, train them, monitor who's productive, and "then eventually everyone gets to the part where they realize they're spending way more on tokens than they imagined."
But here's the twist, some tokens are getting more expensive. On Eye On A.I., Sid Sheth of the chip startup d-Matrix described a "premium token economy" emerging alongside the cheap one. The reason is speed. Since tools like Claude Code made interactive, real-time coding mainstream, "the faster the model responds, the more likely the user is to stay", and companies now charge for that. "You can say, I don't want that level of interactivity, in which case I'll pay $2 for a million tokens," Sheth said. "But if I want that high level of interactivity, I'll pay $20 for a million tokens, and people are willing to pay that $20." So the market is splitting in two directions at once: a race to the bottom for routine work, and a race to the top for instant response.
The price war is real, and it has a target. On Everyday AI, the week's model releases showed how fierce the low end has gotten. Meta launched a Claude-Code-style command-line coding agent (Muse Code) on a new model priced at $1.25 per million tokens in, $4.25 out, with a "contributor tier" that drops that to $0.10 in / $0.20 out, roughly 95% off, if you let Meta train on your code. The host read it plainly: "This looks like a shot at Anthropic," since coding is where "Anthropic reportedly gets about 80% of its revenue." OpenAI, meanwhile, made a frontier-class model (GPT-5.6 Luna, "pretty much on par with Claude Sonnet 5") free and unlimited to roughly a billion users after cutting its price 80%, and Alibaba's new open-weight Qwen will start charging big enterprises a share of revenue. The host's verdict: "It hasn't been a good couple of weeks for Anthropic", which is reportedly heading for an IPO within a couple of months.
And the contrarian macro call: AI might get more expensive, not less. On the Dwarkesh Podcast, the host laid out the uncomfortable math behind all the cheap-token euphoria. Anthropic has grown revenue roughly 10x a year for three years, from $9 billion last year to a likely $100–150 billion this year, but lab compute only grows about 3x a year. Something has to give. Inference margins have already climbed "from 40% in the middle of last year to upwards of 80% now," and spot compute prices are "more than 40% higher than the February trough." His headline thought experiment: if a genuinely human-level software engineer could run on a single high-end chip (an H100), then at today's engineer salaries "that H100 should rent for over $250,000 a year, more than 15 times the current price," before you even count that the AI works nights and weekends. As proof the squeeze is here, he cited Google paying SpaceX "$900 million a month for 110,000 GPUs, about twice the spot price." Even the hardware you'd buy to escape the cloud has gone vertical: The AI Fundamentalists noted RAM prices up ~220%, flash storage up ~470%, and GPUs up 100–300%, with Mac Minis and Mac Studios sold out for months as everyone tries to run models at home.
The through-line: For a year we've watched AI delete jobs from the bottom of the org chart. This week it started deleting them from the middle, the managers who routed, the PMs who aligned, the analysts who fetched, and Whatnot was blunt enough to write "we regret product management exists" into its own playbook. But the same operators cutting the layer are now staring at a token meter that runs at the speed of a second salary, and the smartest of them have stopped asking "how do we use more AI?" and started asking "who here has actually earned their tokens?" The companies that win the next couple of years won't be the ones with the fewest people or the biggest AI bill. They'll be the ones who figure out the exchange rate between a person, a token, and a shipped result, and who are honest enough to notice when all that extra code never actually reached a customer.