The macro trend: from scarcity to abundance
Industry trackers estimate per-token frontier-model pricing dropped roughly 60-80% between early 2025 and mid-2026, with budget-tier models pushing the floor down to a fraction of a cent per thousand tokens. Four forces get the credit: more compute-efficient model architectures that activate a fraction of their total parameters per request, a new generation of GPUs delivering multiples of the prior generation's inference throughput, API volumes growing fast enough to spread fixed infrastructure costs across far more tokens, and aggressively priced entrants forcing incumbents to respond rather than compete on price differentiation alone.
None of that made anyone's total AI bill smaller. Cheaper inference made it economical to run far more of it — the same pattern economists call the Jevons paradox — and reported enterprise AI API spend grew substantially even as the per-token price fell. A lower rate card is not the same thing as a lower invoice, which is exactly the gap a pricing-change monitor like VendorSpies is built to help a team see coming.
Vendor by vendor: what actually happened
OpenAI split its flagship line into three explicit price/performance tiers this generation, spanning roughly $0.20 to $5 per million input tokens depending on tier, with cached input priced at a steep discount off the standard rate and its Batch API still offering 50% off both input and output for workloads that can tolerate asynchronous processing.
Anthropic priced its current-generation Sonnet tier below its own immediate predecessor — $2 / $10 per million input/output tokens against the prior generation's $3 / $15 — a rare instance of a vendor cutting the price of its mainstream model outright rather than only shipping cheaper tiers alongside an unchanged flagship. Prompt caching remains multiplier-based (a 1.25x premium to write a 5-minute cache, 2x to write a 1-hour cache, and a 0.1x rate on a cache hit), stacked on top of the same 50% Batch API discount.
Google kept its Gemini Flash lineup priced from roughly $0.30 to $1.50 per million input tokens across tiers — and already has a price increase publicly scheduled for January 1, 2027, on its newer Flash models and on context caching. Pre-announcing a future price increase months ahead of when it takes effect is unusual; most vendors change a rate card quietly and let a monitor or a surprised customer notice.
What still drives your bill up even when the sticker price falls
A cheaper rate card doesn't protect a team from the usual ways an AI bill gets away from them: reasoning tokens billed as output even when never shown, an account crossing into a higher usage tier and unlocking throughput nobody budgeted for, retry storms after a rate-limit error, or caching that's technically enabled but never actually hitting because the leading context changes on every call. The guides below cover each of these in detail, vendor by vendor and vendor-agnostic:
Already locked in for 2027
Google's January 2027 Gemini Flash and caching price increase is public today, which means it's exactly the kind of change a free watch is built for — not to catch it after the fact, but to make sure it's on your radar well before it hits a January invoice. See the Google pricing calculator for the current numbers and add a free watch so any further change reaches you first.