Google Switches Gemini to Compute-Based Usage Limits

Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- Google switched Gemini usage metering from per-request counting to compute-power-based quotas, so a complex video generation costs more than a simple text prompt even if both count as one request.
- Gemini subscriptions in the US now come in four tiers — Free, AI Plus ($8/mo), AI Pro ($20/mo), and AI Ultra ($100 or $200/mo) — with paid plans offering 2x, 4x, or 5x–20x the free tier's unspecified "standard" limits.
- All Gemini users can access the Flash-Lite, Flash, and Pro models, each available with Standard, Extended, and Deep Think thinking levels that affect response quality and how much of a user's quota is consumed.
- Context windows scale dramatically by tier: 32K tokens (~24,000 words) for free users, 128K for AI Plus, and 1 million tokens (~750,000 words) for AI Pro and Ultra subscribers.
- Google's Gemini app displays two usage bars — one resetting every five hours and one weekly — and paid users who hit their limits are automatically demoted to the most basic model until the next reset.
- Google's support documents reserve the right to change limits without notice and note that free users may be affected first when the company needs to manage AI capacity.
Why it matters: Google just made Gemini pricing opaque for everyday users. Free and lower-tier subscribers can no longer rely on simple rules like "5 image generations a day," and they're explicitly first in line for throttling when Google's data centers are strained — meaning heavy or complex AI work increasingly requires a $20–$200/month subscription to remain reliable.