Newsletter · · Ashutosh Agarwal

Anthropic Launches a Cheap Small Model as Developers Show No Loyalty to Any Lab - Platform Watch - Week of October 9, 2026

Platform Watch for the week of October 2 to 9, 2026. Podcast synthesis on Anthropic's Claude Haiku 5.5 taking aim at the cheap end of the market, OpenAI's comeback as developers switch tools, Meta and Microsoft cutting back on Claude, finance chiefs pushing work to cheaper models, ChatGPT plugin revenue sharing, the Factory and Cognition fight, and Salesforce's $2 billion Listen Labs deal.

Platform Watch

Week of October 9, 2026: Anthropic Launches a Cheap Small Model as Developers Show No Loyalty to Any Lab


Anthropic shipped a small model that beats OpenAI's cheap option and costs pennies per task. That pushes the labs right into the low-cost space that a wave of startups, and China's open models, had been living in. Meanwhile OpenAI is back near a $70 billion run rate because developers switched to it, Meta and Microsoft are cutting back on Claude in favor of their own tools, and Salesforce paid $2 billion for a three-year-old app company. This week's lesson: switching costs are close to zero at every layer of AI, and that cuts both ways for founders.

Last week in this newsletter, OpenAI copied a startup's product in about seven days. The startup was Jev, which makes a fast, cheap model that picks an answer from a short list instead of writing paragraphs.

This week brought a twist. New spending data from the corporate card company Ramp shows Jev is still winning customers anyway. It also brought a bigger move from the other side of the market. Anthropic decided it wants to compete at the cheap end too.

Put those two facts next to each other and you get the main tension of the week. The labs are now competing on price at every level, from the most expensive models down to the cheapest. And customers, from solo developers to Meta, are proving they will switch whenever something better or cheaper shows up.

This Week's Platform Move: Anthropic's Haiku 5.5 takes over the cheap end of the market

Anthropic released Claude Haiku 5.5, the smallest model in its new lineup. Anthropic called it the cheapest, fastest and most capable small model it has ever released. The podcasts that covered it focused on one thing: price.

The numbers:

  • Price per token. For prompts under 100,000 tokens, Haiku 5.5 costs 10 cents per million input tokens and 50 cents per million output tokens. (A token is a small chunk of text, about three-quarters of a word. AI usage is billed in tokens.) Above 100,000 tokens, both rates jump fivefold, to 50 cents and $2.50 (Elon Musk Podcast, "Anthropic Haiku 5.5 and Subagent Economics," October 8, 2026).

  • Compared with Anthropic's other models. The AI Daily Brief said Anthropic is offering it at "a quarter of the cost of Sonnet 5.5, and half the cost of Haiku 4.5 per token." Anthropic is also "discounting tokens by 80% for smaller prompts below 100,000 tokens." The Elon Musk Podcast, quoting Dario Amodei and Anthropic CFO Krishna Rao, put it as costing 75% less to run than Haiku 4.5.

  • Against OpenAI's cheap model. According to The AI Daily Brief, Haiku beats GPT-6 Luna across the board. It "scored more than double on Terminal Bench 4.0 for agentic coding." On computer use, meaning the AI operates a computer the way a person would, it scored 72.4% on OSWorld versus 48.9% for Luna. Artificial Analysis gave it a 43 on its Intelligence Index.

  • Cost per task. Price per token can mislead, because some models use far more tokens than others to finish the same job. Artificial Analysis found Haiku cost 21 cents per task on its maximum setting, or 12 cents on "extra high." That's "right in line with GPT-6 Luna, a few cents cheaper than GLM-5.3 Flash and DeepSeek v4.1 Flash, and one-sixth the price of GPT-6.1 Sol." Haiku "was one of the most token-hungry models Artificial Analysis has ever tested, but it more than made up for it by being extremely cheap" (The AI Daily Brief, "The Most Important Trends Showing Up in New AI Products," October 9, 2026).

The cost picture now looks different. Investor Haseeb Qureshi summed it up, as quoted on The AI Daily Brief:

"Back in June, both GLM and DeepSeek were on the cost and intelligence Pareto frontier. At their price points, nothing could beat them. Then the cost crisis hit, and in the last few months, the US labs have responded with aggressive price cuts and efficiency improvements. Now, with the new Haiku, the frontier is fully reshaped. The entire Pareto frontier is now owned by the US labs. Anthropic, OpenAI, and TypeSafe's Jev. Your move, China!"

In plain English, the "Pareto frontier" is the set of models where you can't get more intelligence for the same price. Three months ago, Chinese open models owned the cheap end of that curve. Now American labs, plus one startup, own all of it.

What builders are doing with it. The Elon Musk Podcast described how customers use Haiku as a "subagent," a cheap worker model that a bigger, more expensive model hands simple jobs to:

  • Rogo, the AI tool for finance, has a large model build a presentation deck. It hands off the tedious part, digging specific segment numbers out of a company's annual report (10-K), to Haiku 5.5. Rogo found it "accurate enough to trust for the data extraction, and fast and cheap enough to run constantly."

  • Joel Everett uses it to pull "the supplier name, the date, the amount, the VAT, and the PO number" from invoices, and to sort bank transactions into accounting categories. The hosts' rough math: classifying 10,000 transaction lines might cost "50 bucks" on a top model and "a few cents" on Haiku 5.5. Everett also noted it still shows "old AI stupidness" and can miss things in the material it's given.

  • Mohamed Adil Hussain told Haiku to find, install and play any game it liked. It got stuck, worked out its own fix, and installed a Linux game called Pingus. The total cost was "roughly six cents."

The catch: a price cliff at 100,000 tokens. The natural fix when a cheap model misses context is to give it more context. With Haiku 5.5, that can wipe out the savings. The podcast walked through a customer service agent. By turn 20 of a chat, the stored conversation might be 50,000 tokens. By turn 40, it passes 100,000, and "the discount that made the model attractive in the first place vanishes." The hosts compared it to old cell phone plans where 1,000 texts cost $10 but text 1,001 cost a dollar. Their advice, credited to Robert Daniel Menor, is to track token usage for each subagent "from day one," and to trim or summarize old context before it reaches the model.

Why this matters for startups

1. Building got cheaper. If your product runs dozens of small AI steps in the background, like reading documents, sorting things or watching inboxes, your cost per customer just dropped. For most application companies, that's good news.

2. Selling "cheaper than the labs" got harder. Last week, the threat to Jev came from OpenAI copying its features. This week it comes from Anthropic pricing a general model low enough to be "in line with" the cheap specialists. When a lab decides to fight at the bottom of the market, a startup whose main pitch is low price is fighting the company that owns the model.

3. Jev's numbers show the other side. Ramp's October data, discussed on The AI Daily Brief, ranked TypeSafe AI, Jev's maker, first in growth relative to size among software vendors and second in market-share growth. Ramp puts Jev in the foundation model category, and it "gained a full point of market share within a month," growing faster than every frontier lab except Anthropic. Ramp economist Eric Karazian wrote that "the introduction of highly-performance standard and light models and highly token-efficient models designed for specific enterprise use cases like JEV will threaten OpenAI and Anthropic's revenue growth among businesses." The host's point: "Startups and some advanced enterprises are actually signing up and using the model in a significant way." Getting copied didn't kill Jev, at least not yet.

4. China's pitch shrinks to one argument. The AI Daily Brief noted that US labs "have increasingly solved the cost side" but haven't addressed sovereignty, meaning the ability to run models on your own hardware and keep your data from ever being used for training. If you're building on Chinese open-weight models to save money, that advantage just got smaller.

The race behind it: OpenAI is back, and developers moved quickly

The most useful discussion of the week came on 20VC, with Harry Stebbings, Jason Lemkin, Rory O'Driscoll and guest Dev Ittycheria, back as CEO of MongoDB (20VC, "Cognition vs Factory: Vinod Khosla Creates a Storm | OpenAI Nears $70B Run Rate: Anthropic Under Threat | ElevenLabs Doubles Its Valuation to $22B & Salesforce Buys Listen Labs for $2B," October 8, 2026).

The comeback in numbers:

  • OpenAI is said to be nearing a $70 billion revenue run rate and raising $30 billion at a $1.4 trillion valuation. That would be the first private round ever above a trillion dollars. The panel said OpenAI's IPO would wait until 2027, while Anthropic's is expected in the middle of November.

  • Rory O'Driscoll laid out how much the race has swung. OpenAI's revenue grew just 18% from Q1 to Q2. He said Q2 was "the first time Anthropic" pulled ahead on reported revenue, citing "11 billion versus 6.8 billion, growing 10x versus growing 2x." Now OpenAI is reported to have grown 70% quarter-on-quarter in Q3. "If gap revenue grows 70% Q1, Q or even 60%, up from 18%... that's a massive reacceleration."

  • He also explained why the number mattered so much to OpenAI: "given the OpenAI forward purchases of just about everything from compute to data centers to land in Virginia, if they didn't grow at a hyper speed, a lot of things come unstuck."

Where the growth came from: developers switching. Harry Stebbings: "Every team that I speak to has at least deviated some meaningful portion from Anthropic to Codex and everyone says how amazing it is." Jason Lemkin: "even coding at the start of this year only worked on Anthropic," and now OpenAI made something "truly epically competitive to the point where people are switching out their models for real because it's better and cheaper."

Dev Ittycheria put it most clearly, drawing on more than 12 years of selling to developers:

"There's not a lot of loyalty. Developers are very quick to use to switch from one tool to another or frankly use multiple tools... developers are quick to find the shiny new toy."

He added one exception. Where the lab is deeply embedded in an enterprise, with signed contracts, thousands of trained employees and workflows built around it, "it may be harder to switch."

The question for founders: Rory framed it as a "rotating share game." Did OpenAI's growth come out of Anthropic, or did both grow? The answer arrives with Anthropic's Q3 numbers, which he noted will be known by the time it goes public. "What happens if Anthropic's numbers for Q3 aren't blowout?" Harry asked.

The labs' biggest customers are building their own tools

Startups worry about the labs copying them. This week, the labs got a taste of the same thing from their biggest customers.

Tech Brew Ride Home, citing The Information, reported that Meta and Microsoft are trying to "wean themselves off of using Claude internally" (Tech Brew Ride Home, "Le Chonk," October 6, 2026):

  • At Meta, employees using Claude Code fell to about 30,000 from about 60,000 earlier this year. Part of that came from spring layoffs that cut about 10% of staff, but "a bigger factor" was Meta's own tools. Muse Code, a Claude Code competitor running on Meta's models that Meta began testing with outside customers in August, "had more than 60,000 employee users recently." Even so, in a recent 28-day period Meta spent "more than $105 million on Claude Code."

  • At Microsoft, executives had projected at least $1 billion in internal spending on Anthropic this year. Microsoft has since cut that figure "by more than a third" after asking staff to "use less Claude to save on costs and spend more time using Microsoft's homegrown AI tools."

The same show noted that Anthropic's leaked IPO filing showed two customers making up about 25% of its revenue. That's the same risk founders face: when your best customer is also capable of building what you sell, its size is a risk as well as a source of revenue.

Finance chiefs are coming for the AI bill

Several podcasts pointed to the same shift. Companies have stopped spending on AI without limits, and cheaper options are winning work.

  • Rory O'Driscoll's math on 20VC. "If OpenAI and Anthropic are each making $70 billion in ARR, that's $140 billion. That's more than Microsoft taxed for the longest time across both Windows and Office." At that size, he said, a CFO tells the CIO: "I don't care that it's easier to use Anthropic, dude. We need to shave $3 million off this token bill. So start using JEV for your 20% of your decisions and start using Reflection or something for anything that's not frontier." He compared the labs to the old man in The Old Man and the Sea, bringing home a huge fish while sharks take bites out of it.

  • The open-model argument. Dev Ittycheria (who disclosed that Sequoia, which he is affiliated with, is an investor in Reflection) said Reflection's new US-made open model is near frontier quality and "three to four times more cost-effective." Jason Lemkin said that at Dreamforce, "nobody I talked to really wanted to use a Chinese-source model," so a solid US alternative "could be a torrent of demand." Rory added that frontier models on OpenRouter, a service that routes requests across many models, "get 20, 30% of tokens and get 90% of the dollars."

  • "The enterprise is moving away from frontier models." That was a co-host's claim on Big Technology. Dots-style tasks such as email, research and pulling in information are "not frontier anymore," and "enterprises are no longer just pouring money only into the frontier. It's getting distributed across open weight models. People are building in-house" (Big Technology Podcast, "Anthropic's IPO Leak, OpenAI's Dots vs. Meta's Muse, Visual Turing Test," October 2, 2026).

  • "Token maxing is over." Steve Eisman, the investor made famous by The Big Short, said that earlier this year employees were told to "use AI till you're ill," companies "blew through their budgets within months," and "sometime this summer, I think token maxing probably ended." His harder claim is that the labs "realize that there are no moats around their business whatsoever" as open-weight models take share (Prof G Markets, "Steve Eisman: One Company Could Break The AI Boom," October 2, 2026).

  • "Token shock." On Business of Tech, Anurag Agrawal said bills spiking because agents use tokens faster than per-token prices fall is the top frustration for 42% of buyers. He said 78% of small and mid-sized businesses are prioritizing private, hybrid or edge setups to turn unpredictable AI costs into fixed monthly fees (Business of Tech, "Vendor Rebates Move to AI Growth: Anurag Agrawal Details the Margin Fallout for MSPs," October 8, 2026).

Not everyone agreed that cheaper is all that matters. Dev Ittycheria expects the "Jevons paradox": as costs fall, companies use AI for more and more work, so "I don't see this as being a zero sum game." Lemkin said every executive he met at Dreamforce "was just overloaded with the amount of internal demand there is for tokens."

For startups: If your margins depend on reselling one lab's top model, your customers' finance teams are about to look hard at that line item. If your product sends each task to the cheapest model that can do it well, you're on the right side of this shift.

The labs want to be the platform, and they're starting to pay for it

OpenAI's DevDay was last week. This week its leaders filled in details that matter to anyone building inside ChatGPT.

Revenue sharing for plugins. On Lenny's Podcast, Tibo Sottiaux, OpenAI's head of ChatGPT, called the ecosystem the "sleeper hit" of DevDay and shared something "we didn't actually talk [about] in the keynote":

"We will pay, you know, our plugins that are popular and, you know, are seeing a lot of usage. They're going to get part of like the revenue share as well."

How it works: ChatGPT subscribers get a certain amount of usage. When they spend that usage inside a partner product through "Sign in with ChatGPT," OpenAI will share economics with the partner. "Sign in with ChatGPT" now has 16 partners. Sottiaux said it started informally with the makers of Pi and OpenCode: "We kind of like shook hands, you know, like virtually."

How plugins get recommended: retention. Asked how to get a plugin discovered, Sottiaux said: "Build a good plugin." OpenAI looks at "retention numbers... how successful the plugin, the quality, and then, you know, that's what we then start to recommend to users in conversations." If a plugin isn't good, "it will stop being recommended" (Lenny's Podcast, "OpenAI's Head of ChatGPT: We're entering a new era of AI (again) | Tibo Sottiaux," October 4, 2026).

His advice to builders is worth noting: "if you were just like pushing yourself and just really imagining like all of this being roughly 10 times better than it is today, you know, like in a year, it's like, you know, you would build in a different way." He also said "the majority of actions on the internet will be taken by agents." Notion saw "a ton of traffic" once it built an interface agents could use, and "you have to figure out the economics of that."

The labs are hedging on their own models. The AI Daily Brief pointed to two moves:

  • OpenAI is now selling open-model inference through its B2B marketplace. When an enterprise makes a big spending commitment with OpenAI, "you can use a meaningful portion of that on not OpenAI models, but on other open models."

  • Elon Musk said GrokBot "will use the best backend model for any given task, including Claude Opus 5.5, Midjourney, Suno, and other leading APIs," with simple requests going to a fast version of Grok. The host framed this as a bet that the advantage sits in the "harness," the product and context wrapped around a model, rather than in the model alone. Investor Thomas Tunguz: "No company can provide the best model for every use case."

That's the strategy many startups have used: own the customer relationship and route each task to whichever model works best. Now the labs and the big platforms are doing it too.

The labs now care about product design. OpenAI launched Intelligent UI, which lets ChatGPT answer with charts, interactive visuals and other layouts instead of walls of text. It also brought a new GPT-6 model to ChatGPT's free tier for its 1.2 billion weekly users. The AI Daily Brief's take is a warning for app builders: "So far, they've gotten away with having fairly bad, unintuitive products, because the intelligence you get access to through them is so powerful... Now that everyone has that sort of intelligence on tap, product differentiation is going to really matter." For years, the case for many AI apps was "a nicer interface on top of the model." The labs are now working on that themselves.

Dots come for vertical agents. On Decoder, Nilay Patel noted that OpenAI is "even offering specialist Dots for marketing, legal work, and accounting." The Verge's Hayden Field said Dots sit in the $100 to $200 a month tier and are being marketed to knowledge workers, with Sam Altman pitching a Dot as "your chief of staff" (Decoder with Nilay Patel, "OpenAI has a Muse problem," October 8, 2026). Not everyone was impressed. On AI For Humans, the hosts said they were already doing much of what Dots does with Codex voice mode, and that Dots "doesn't make phone calls," unlike two free agents they're testing (AI For Humans, "AI put Minecraft in Skyrim (and OpenAI had a Dev Day)," October 2, 2026). The same hosts, on what AI can now clone: "Nobody is safe."

Coding tools: a crowded market turns personal

The coding-agent market is crowded enough that the competition is now spilling into public fights.

On 20VC, the panel covered the Factory vs. Cognition dispute. Chris Degnan, the longtime Snowflake sales leader, was a board observer at Factory and, according to Factory founder Matan, spoke with him daily. He then took the chief revenue officer job at Cognition, a direct competitor. Matan objected publicly. Then Vinod Khosla, whose firm led Factory's Series C, went on Twitter and, as Harry Stebbings described it, called Factory a "struggling second tier competitor."

  • Dev Ittycheria (who disclosed he is an angel investor in Factory, that Sequoia is also a Factory investor, and that MongoDB partners with Cognition) was "just totally flummoxed," saying Khosla's tweet "basically handed every other firm a weapon when they're competing on a deal saying, is this the partner you want when things go bad?"

  • Rory O'Driscoll: "It was very unhelpful for Cognition, very unhelpful for both companies... it was not venture's finest hour."

  • Jason Lemkin argued loyalty has "changed permanently" in AI: "Half of the Frontier Labs folks are rotating back and forth." He said the old "rule of two," where a departing executive could take at most two people, has died. "Now I'm seeing folks take like eight or 10 with them the first week."

Why it belongs in Platform Watch: Two well-funded coding-agent startups are now fighting over sales leaders and investor loyalty while the labs win developers with their own tools. That's the 20VC story about developers moving to Codex. Meta's in-house Muse Code shows the same pressure. Startups in this space are competing with each other and with the companies that supply their models.

Exits: sell the feature, keep the franchise

Salesforce buys Listen Labs for $2 billion. Listen Labs uses AI to run market research interviews. It's about three years old. Per Harry Stebbings, it had earlier offers at $1.5 billion, and Sequoia returned "about $850 million back to investors on 96 invested." Rory O'Driscoll called market research a good category for AI because "voice as a modality is a solved LLM problem," and AI can adjust questions based on answers, unlike "canned, stupid last generation surveys." Jason Lemkin was skeptical about the fit: he didn't see how "20 million of revenue of next generation AI survey moves the needle" at Salesforce's size (20VC, October 8, 2026).

Dev Ittycheria gave the line founders should remember:

"I think a lot of AI application companies will sell because a clever product on top of someone else's platform is frankly a feature... the companies that are more durable are the ones that... are creating a data loop where... the usage creates data that no one else has, the data makes the product better, the product attracts more usage... if someone tries to copy you, they can copy the features, but they don't have all the interactions that you have."

Rory's prediction: "In five years from now, maybe 10% of these apps companies will have sold early. They'll have got their 30x revenue and the rest of them will be grinding it out for six times revenue at 300 million with 3,000 employees." Lemkin's advice: if you can sell for $2 billion in the first five years "and you're not building Mongo or better... I would take it."

Nvidia and Perplexity. On The Information's TITV, reporters said Perplexity proposed that Nvidia buy it outright. The two sides also discussed a license-and-hire deal (where a big company licenses the technology and hires the team without a full acquisition) worth "more than $20 billion or more." They landed on a partnership in which Nvidia invests $3 billion, valuing Perplexity at "more than $35 billion before the investments." Part of the appeal for Nvidia is that Perplexity "wants to run, like, lots of different models and especially open source models," which fits Jensen Huang's vision of "10,000 models" rather than "one or two, like, God models" (The Information's TITV, "Inside Nvidia's $100B Dealmaking Spree, Elon Musk's Massive Chip Ambitions," October 7, 2026). A startup that doesn't depend on any single lab turned out to be very attractive to the company that sells chips to all of them.

ElevenLabs doubles to $22 billion. On 20VC, Jason Lemkin said he'd put "up to 10% of the fund" in at that price. He said the company "could be approaching a billion in revenue," its margins are "very impressive," and "this isn't GEV-level risk." His reason ties directly to platform risk: in voice, "quality here is in many ways more important" than in typical AI work. "When that flower dealer is picking up the phone, it has to work." His summary: "I don't think all models are fungible... This is where the LLM matters."

Exposed vs. Defensible

Exposed this week Defensible this week
Startups whose main pitch is "cheaper than the labs." Haiku 5.5 now matches GPT-6 Luna on cost per task at a sixth of Sol's price (The AI Daily Brief, Oct 9). Voice infrastructure with a clear quality lead. ElevenLabs: "not a fungible use case today" (20VC, Oct 8).
Products built on Chinese open models mainly to save money. US labs "now own" the low-cost frontier, leaving sovereignty as the main remaining argument (The AI Daily Brief, Oct 9). Companies with a data loop. Usage creates data no one else has, as Dev Ittycheria described (20VC, Oct 8).
Single-model coding tools. Developers show "not a lot of loyalty" and have moved meaningful work to Codex (20VC, Oct 8). Meta's in-house Muse Code passed 60,000 internal users (Tech Brew Ride Home, Oct 6). Vertical AI with private data. EvenUp in personal injury, where "99% of that data is in the private domain" (This Week in Startups, Oct 8). Leo in procurement, where the price benchmarks are "proprietary data" models can't train on (The a16z Show, Oct 2).
Generic personal-agent startups. Up against free Meta Muse, $100–$200 OpenAI Dots and GrokBot. One investor called Instinct's roughly $10 billion price "a big uphill battle" against big-company distribution, ads and wallets (This Week in Startups, Oct 8). Multi-model products that own the customer. Perplexity's model-neutral approach drew a $3B Nvidia investment at a $35B+ valuation (The Information's TITV, Oct 7). GrokBot and OpenAI's marketplace are copying the routing approach (The AI Daily Brief, Oct 9).
Vertical "assistant" wrappers in marketing, legal and accounting. OpenAI now offers specialist Dots in those fields (Decoder, Oct 8). Agents that do the whole job, including the messy exceptions. Leo found the routine invoice process is "only like 20% of the work"; the other 80% is exceptions (The a16z Show, Oct 2).
Apps whose edge is a nicer interface on a lab model. OpenAI's Intelligent UI shows the labs now "actually care about the product experience" (The AI Daily Brief, Oct 9). Distribution partners inside ChatGPT. OpenAI will share revenue with popular plugins, ranked by retention (Lenny's Podcast, Oct 4).

Why the right-hand column holds up

The a16z Show's episode "Why AI Agents Can Beat the Incumbents" (October 2, 2026) gave the clearest explanation of what lasts. Vlad, founder of the procurement AI company Leo, said his company treats lab models "as a commodity" and uses several providers. The value is in what models can't do alone:

  • The messy 80%. Traditional invoice software covers the clean process, but "it's actually only like 20% of the work... 80% of the problem is like, what if the invoice is fraudulent? What if there's like a mismatch?"

  • Data the labs can't see. For price benchmarking, "those are like all proprietary data based on like one enterprise, like across multiple enterprises. So like general purpose models can't train their models on that." Leo is starting to train its own outcome-focused models, citing Jev as an example of the approach.

  • Engineers, not consultants. "85% of the people are engineers," and the forward-deployed engineers who customize the product for each client have one key goal: "literally your job is like to automate yourself."

a16z partner Seema described the incumbents' first move as the "slap on a chatbot strategy" and said the durable vertical company owns "the end to end work," builds a data asset, and grows customer dependence: "the overall customer is dependent on that product."

On This Week in Startups, an investor named Chris said his firm went from about 80% horizontal software to 80% vertical AI. "I'm not at all a believer that the models are just going to own everything... how you apply that intelligence and how you give it context and permissioning is going to be critical" (This Week in Startups, "What VCs Really Think About Personal AI Agents | E2347," October 8, 2026).

A note on consumer AI. a16z's seventh Top 100 Consumer AI Apps report, discussed on The a16z Show, found only about 4.5% of US consumers pay for an AI subscription. The top 1% of payers spend $903 a month and the median spends $25. Some personal-assistant founders report costs of "hundreds or thousands of dollars a month to serve a user," mostly because power users run coding jobs through them. Meanwhile OpenAI has reached about a $1 billion annual run rate in advertising, a level that "would previously take a company years and years and years." Olivia Moore's point: consumer startups paying high model costs can't use the old "grow free, monetize later" playbook, "given how high COGS are" (COGS means cost of goods sold, here mostly AI usage) (The a16z Show, "The Top 100 Consumer AI Apps: Who's Actually Paying?," October 5, 2026). OpenAI can subsidize free users with ads. Most startups can't.

Founder Takeaway

Build so you can switch models easily, and make sure your customers can't easily switch away from you.

This week showed that switching costs for AI models are close to zero at every level. Developers moved to Codex. Meta moved to its own tools. CFOs are pushing work to Jev and open models. Even GrokBot and OpenAI's marketplace now route to other companies' models. You can't count on loyalty to any lab, and the labs can't count on loyalty from you. Use that:

  1. Send each task to the cheapest model that does it well. Haiku 5.5 at 12–21 cents a task changes the math on background work. Put an expensive model in charge and cheap subagents underneath, as Rogo does. If you charge a flat fee and sell a top model at cost, your margin is at the mercy of whichever lab sets your price.
  2. Watch for price cliffs. Haiku's fivefold jump above 100,000 tokens is a design issue, not a footnote. Track token usage for each subagent from the start, and trim context before it gets to the model.
  3. Don't make "cheaper than the labs" your whole pitch. The labs are now competing on price at the cheap end too. Jev is surviving on speed, measured accuracy and real customer adoption, not just price.
  4. Build the part a lab can't copy in a week. That means the messy exceptions, the private data, the trust it takes for a customer to let your agent negotiate on their behalf, and a data loop that improves with every user. As Dev Ittycheria put it: "they can copy the features, but they don't have all the interactions that you have."
  5. Treat ChatGPT as a sales channel, not where your product lives. Revenue sharing for plugins ranked by retention is a real opportunity. But OpenAI controls the recommendations and the customer's wallet, so keep a direct relationship with your users too.
  6. If you're a feature, decide early. Listen Labs got $2 billion in three years. Rory O'Driscoll expects maybe 10% of app companies to sell early at 30x revenue and the rest to grind it out. Be honest about which one you're building.