Google has rewritten how Gemini meters usage across every tier, moving from a per-request quota to a compute-weighted model where the same prompt can cost different amounts depending on which model runs it and how hard that model has to think. The change hits Free, AI Plus at $8 a month, AI Pro at $20 a month, and AI Ultra at either $100 or $200 a month. Users who previously counted on rules of thumb like three video generations a day are now watching their credit deplete at variable rates, with no published conversion table.
The two axes that determine cost are the model tier — Flash-Lite, Flash, or Pro — and the reasoning depth: Standard, Extended, or Deep Think. A quick weather query on Flash costs a fraction of what a Deep Think coding session on Pro consumes. Google frames this as a fairer alignment between what users pay and what its data centers actually spend, but the trade-off is opacity: prompt costs are no longer legible before the fact.
Google will not publish a numeric baseline for the Free tier, describing its limits only as "standard." AI Plus subscribers get 2x those undisclosed standard limits, AI Pro gets 4x, and AI Ultra gets either 5x or 20x above AI Pro depending on whether the user is on the $100 or $200 plan. The multipliers are precise; the base they multiply is not.
Key facts
- 01Gemini paid tiers now cost $8/month (AI Plus), $20/month (AI Pro), and $100 or $200/month (AI Ultra).
- 02AI Plus gets 2x the free tier's undisclosed baseline, AI Pro gets 4x, and AI Ultra runs 5x or 20x above AI Pro.
- 03Context windows scale from 32K tokens (free) to 128K (AI Plus) to 1 million tokens on AI Pro and Ultra.
- 04Usage now resets on two clocks: a rolling 5-hour bucket and a weekly cap, both visible in the app.
- 05Paid users who hit their cap are demoted to the most basic model until the next reset, not cut off entirely.
Context window size is the one clean number. Free users get 32K tokens per conversation thread, roughly 24,000 words. AI Plus quadruples that to 128K tokens, or about 96,000 words. AI Pro and AI Ultra jump to 1 million tokens, roughly 750,000 words — enough to load an entire novel or a large codebase into a single prompt.
Usage resets on two clocks. The primary bucket refills every 5 hours, and a secondary weekly cap resets once a week. Both are visible inside the Gemini app on web, Android, and iOS by opening the settings cog and selecting Usage limits, where two progress bars show current consumption and the timestamp of the next reset.
Paid subscribers who exhaust their allocation are not locked out. Google demotes them to the most basic model available and lets them keep prompting at that reduced capability until the next reset window opens. Free users get no such fallback, and Google's support documentation warns that free-tier access may be throttled first when the company faces capacity constraints.
The formal disclaimer in Google's documentation is that "access is subject to change or may be limited based on testing, experimentation or availability," which as Wired's David Nield notes translates in practice to "some days may be different from others" for a given quota. That flexibility gives Google room to manage data-center load in real time; it also means a subscriber cannot budget their week with any precision.
The shift matches how the underlying economics actually work. A Deep Think reasoning trace on Pro consumes vastly more inference compute than a one-shot Flash-Lite completion, and flat per-request quotas forced Google to either overprovision the light users or under-serve the heavy ones. Compute-based metering is what OpenAI and Anthropic have effectively done on their APIs for years — Google is bringing it to consumer subscriptions.
The friction is that consumer products historically hide compute cost from the user. When ChatGPT Plus or Claude Pro throttles a heavy user, the response is usually a hard weekly ceiling on the most expensive model, with everything else uncapped. Google's approach — variable cost per prompt, undisclosed base rate, two overlapping reset windows — is more honest to the underlying math but harder to reason about at the moment of use.
There is a competitive angle. AI Ultra at $200 offers up to 20x AI Pro's cap and the full 1 million token context, positioning it against ChatGPT Pro at $200 and Claude's Max tiers. AI Plus at $8 is the aggressive entry point, undercutting the $20 tier that has become the default price of paid consumer AI. Google is using the low end to convert free users and the high end to keep power users from defecting to competitors' unlimited-feeling plans.
For subscribers, the practical shift is that model choice is now a budgeting decision, not just a quality one. Running Pro with Deep Think on every prompt will drain a weekly cap fast; leaving routine queries on Flash preserves the ceiling for work that actually needs it. The Usage limits screen is where that budgeting has to happen, and it is where Google will keep surfacing upgrade offers whenever a user starts pushing against a cap.
The bigger read is that consumer AI pricing is converging on the same variable-cost logic that governs API billing, minus the transparency. Google can now dial supply against demand without changing sticker prices, using model-tier and reasoning-depth pricing as the release valve. That gives Google's margins room to breathe as Gemini usage grows, but it also erodes the one thing subscription pricing traditionally offered — predictability — and hands more of that pricing power to whichever competitor is willing to publish a hard, legible cap first.
Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.
Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.



