Skip to main content
Live
Main content

GitHub Copilot's per-token pricing shift signals end of subsidized AI era

Microsoft's move to charge per token instead of a flat rate exposes the gap between AI's true cost and what customers will pay.

Jaeden Schafer
Editor in Chief · · 5 min read
Microsoft logo

Microsoft is restructuring GitHub Copilot pricing to charge per token rather than a flat monthly rate, a shift drastic enough that customers have started calling it the Tokenpocalypse. The change lands as Anthropic prepares to file for an IPO and as enterprises confront the gap between subsidized AI pricing and the real cost of inference. Uber, one of the loudest corporate adopters, burned through its annual AI budget in roughly a month and a half this year before imposing internal usage caps.

The Copilot repricing is the clearest signal yet that the era of flat-rate, all-you-can-eat AI is closing. Microsoft is passing token costs through to customers because the underlying economics don't work at a fixed subscription. The $20-per-month price OpenAI set for ChatGPT Plus when the product launched was, by the company's own admission later, less a pricing strategy than a guess — and even at higher tiers for advanced models, the revenue still doesn't cover the cost of serving a heavy user.

That gap has been filled by investor capital. Foundation model labs and the platforms built on top of them have been operating on a subsidy that consumers never saw on the invoice. As that subsidy thins out, the cost shows up — first in enterprise contracts, then in consumer pricing, then in usage caps that look a lot like the rate limits Uber just imposed internally.

This whole ecosystem is heavily, heavily subsidized by investor money. And so stuff that seems like it has no cost is, in fact, incredibly expensive.
Anthony Ha, TechCrunch weekend editor

Key facts

  • 01Microsoft is moving GitHub Copilot from a flat rate to per-token pricing, a change Reddit users have dubbed the 'Tokenpocalypse.'
  • 02Uber blew through its annual AI budget in roughly a month and a half, then imposed usage caps inside the company.
  • 03The $20-per-month [ChatGPT](/openai) Plus price still doesn't close the gap to the true cost of inference, even on advanced models.
  • 04[Anthropic](/claude) is preparing IPO filings that will need to disclose token-cost exposure as a risk factor.
  • 05Google agreed to pay SpaceX $920M per month to rent xAI compute capacity, a benchmark for how expensive the underlying infrastructure has become.

The Uber example is instructive in two directions. On one hand, it shows how fast even disciplined enterprise buyers can blow past AI budgets when token consumption isn't metered tightly. On the other, Uber is the standard counterexample bulls cite when skeptics call AI unprofitable — the company was wildly unprofitable for years before scale closed the gap. The catch is that closing that gap required Uber to transform its business model repeatedly and squeeze drivers and riders along the way.

Whether AI labs have the same levers is the open question. The cost structure is less squishy than ride-hailing — GPUs, power, and cooling are hard inputs. Google's agreement to pay SpaceX $920M per month to rent xAI compute capacity, which AI Chat Daily covered last week, is a useful benchmark for how expensive the underlying infrastructure has become. There's only so much margin to extract before the labs have to either raise prices materially or shrink what a token buys.

Anthropic's forthcoming S-1 will be the first real stress test of how a frontier lab discloses this exposure to public-market investors. Token costs, inference margin, customer concentration, and the volatility of compute supply agreements all have to be written into risk factors at a moment when the underlying economics are shifting month to month. The same applies to OpenAI whenever it files.

Can these AI labs collapse that cost [and] progress the tech enough in a way that it eventually meets in the middle with customers' appetite for spending?
Sean O'Kane, TechCrunch reporter

The behavioral cycle is moving faster than the disclosure cycle. The industry obsession with 'tokenmaxxxing' — pushing models to consume as many tokens as possible per task to improve output quality — peaked and reversed within roughly six months as enterprises looked at the bills. Companies that built workflows around aggressive token consumption are now rebuilding them around efficiency. Vendors who sold customers on unlimited usage are now metering it.

The skeptical read is that this is the moment the AI bubble narrative gets its first hard data point. If GitHub Copilot, one of the most successful AI products by revenue and adoption, can't sustain flat-rate pricing, the assumption that consumer and prosumer AI tools will stay cheap is wrong. Enterprises that built ROI models on 2024-era pricing will need to redo them. Startups that priced their AI features assuming token costs would fall faster than usage would rise are exposed.

Related · from this week
GitHub Copilot moves to usage-based billing as inference costs double
Jaeden Schafer · 5 min read →

The pro-innovation read is that this is exactly the cost discipline the market needed. Subsidized pricing distorted demand signals, encouraged wasteful prompting, and made it impossible to tell which AI features were genuinely valuable. Per-token pricing forces a real conversation about what's worth paying for, and it pushes labs to focus engineering effort on inference efficiency rather than benchmark-chasing.

For the AI market, the Copilot shift is the leading indicator that pricing power is moving from the hyperscaler distribution layer back to the model labs that own the inference cost. Microsoft can't absorb the token bill on Copilot any more than its customers can, so it's passing the cost through. Expect Google, Amazon, and every SaaS vendor with an AI feature to do the same over the next two quarters — and expect the IPO filings from Anthropic and eventually OpenAI to read very differently than the marketing did.

ShareXLinkedInEmail
AI Box

Every AI model. One chat.

The latest models from ChatGPT, Claude, Gemini, Sora, ElevenLabs — 80+ models in a single chat. Compare answers side by side. Pick the best one every time.

  • ChatGPT, Claude, Gemini, Grok, DeepSeek — in one chat
  • Generate images & video with Sora, Veo, Ideogram
  • Compare any two models side by side
  • From $8.99/mo · 80+ models, all included
Try AI Boxaibox.ai
Trusted by 3,000+ teams
Got a tip?

Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.

Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.

AI Box Daily briefingFree · Daily · No fluff

Stay ahead of everyone in AI.

The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.

Loved by 10,000+ AI professionals
Free forever. Unsubscribe with one click.

The briefing read inside teams at

Keep reading

More from Business

Microsoft logo
Business

GitHub Copilot moves to usage-based billing as inference costs double

Starting June 1, Copilot subscribers get AI Credits matching their monthly fee — overages billed by the token at API rates up to $30 per million.

Jaeden Schafer5 min read
Palantir posts $1.9B quarter, Karp calls AI frontier labs 'Marxist'
Business

Palantir posts $1.9B quarter, Karp calls AI frontier labs 'Marxist'

Revenue jumped 93% year-over-year as CEO Alex Karp accused OpenAI and Anthropic of trying to capture their customers' means of production.

Jaeden Schafer5 min read
Microsoft logo
Business

Microsoft shifts Excel and Word AI prompts to its own MAI models

The company is routing a share of Office 365 prompts away from OpenAI and Anthropic to homegrown MAI models as AI costs climb.

Jaeden Schafer4 min read