Google's 'Frozen v2' Chip Bakes Gemini into Silicon — What It Means When Your AI Vendor Owns the Full Stack
Written 22 July 2026. Primary reporting: Reuters, Google plans new chip to run Gemini models more efficiently, The Information reports (20 July 2026). Additional coverage: Bloomberg, TechCrunch, Crypto Briefing on the Frozen v2 codename. This post is my analysis of the strategic implications for enterprise AI buyers.
The Information broke a story on 20 July that Google is developing a new server chip that would bake components of its Gemini AI models directly into hardware. According to the reporting, the chip is codenamed “Frozen v2”, is being designed to run 6 to 10 times more efficiently than Google’s existing AI chips (measured in tokens processed per unit of power), and is being aimed at deployment by 2028. Alphabet shares moved on the news.
The reporting also flagged the why: Google Cloud has reportedly been declining deals with external customers because of AI computing capacity constraints. Frozen v2 is one of Google’s plays to unblock that.
That’s the reporting. Now the analysis I actually want you to take away.
The interesting story isn’t the chip
If you’re a CTO who read the Frozen v2 story and thought “great, my Gemini bill will get cheaper” — you’re reading it as a customer of an efficient vendor. That’s not wrong, but it’s incomplete.
The pattern that matters is what Google is becoming. The company already owns:
- The model (Gemini)
- The training chip (TPU 8t, codename Sunfish, Broadcom-designed)
- The inference chip today (TPU 7 / Ironwood — 192 GB HBM3E, 7.37 TB/s memory bandwidth, 4,614 FP8 TFLOPS, deployed internally for Gemini API traffic)
- The inference chip tomorrow (TPU 8i, codename Zebrafish, MediaTek-designed, 288 GB HBM, 10.1 petaflops FP4)
- The model-specific inference chip after that (Frozen v2, 6-10× more efficient, 2028)
- The data centre (Google’s own)
- The network fabric (Google’s own)
- The cloud contract (Google Cloud)
- The developer tooling (Vertex AI, Firebase, Gemini API)
That is not “we’re using Nvidia and also Google TPUs.” That is full-stack vertical integration on a single vendor’s rails. When your provider owns every layer from silicon to service, three things happen at once — and they don’t happen in your favour.
1. Their unit economics improve without their unit price having to
Custom silicon aimed at your specific model can reasonably promise 6-10× more tokens per watt versus general-purpose accelerators. That efficiency gain shows up somewhere. It could show up in your per-token price. It could show up in Google’s margin. There’s no market mechanism guaranteeing the former.
For an enterprise buyer: when your vendor makes a 10× efficiency gain, the price cut you actually see is a matter of vendor discretion, not vendor economics. Unless you’re one of the accounts big enough to renegotiate on that basis, the improvement passes above your head.
2. Your ability to compare shrinks
The Ironwood story is public — GA in late 2025, real spec numbers, benchmarks against Nvidia GB200. That transparency exists partly because Google Cloud wants to sell TPU-hosted workloads to third parties.
Frozen v2 is different. If Gemini components are hardwired into the chip, the workload it runs is Gemini. There’s no meaningful “run our model on your Frozen v2” story to tell. Which means the price of a Gemini token in 2028 becomes harder to compare to a Claude token or a GPT token, because the underlying compute is invisible and non-portable. Comparison shopping gets structurally weaker.
3. The AU capacity story bites first
The Reuters piece flags Google Cloud declining external customer deals today because of capacity. Read that carefully: a hyperscaler with global data centres is turning away paying customers because they don’t have enough compute in the right places. That constraint hits smaller regions first. If you’re running an AI-heavy Australian workload on Google Cloud right now, your growth ceiling may be being set by Mountain View’s capacity plan, not by your business plan.
That’s not hypothetical. That’s this week.
What actually to do about it
Three concrete moves worth making this quarter, regardless of your current cloud:
Draw the real switching diagram
Not the marketing “multi-cloud ready” slide. The actual diagram. Which components of your AI stack are:
- Vendor-neutral (Python, containers, open-source libraries)
- Portable with rework (prompt library, retrieval logic, evaluation harness)
- Genuinely locked (fine-tuned models, vector store schema, vendor-specific auth, vendor tooling like Vertex AI or Bedrock guardrails)
Every arrow that only points at one vendor is a lock-in the Frozen v2 story just made more valuable to them and more expensive to you.
Renegotiate on unit price, not total contract
Vertical-integration efficiency shows up in vendor pricing 6-12 months after the hardware ships. If Frozen v2 lands in 2028 and you’re on a three-year commit signed today, the negotiation you want to open in 2028 is per-token pricing tied to hardware generation — not total contract size. Get that clause negotiated NOW. It’s much cheaper before you need it.
Instrument for portability, not just observability
The same fail-closed gate discipline I’ve written about for adversarial verification of AI content pipelines applies to vendor risk. The gate you want: “if the primary vendor’s pricing, latency, or region availability moves against us, can we redirect the workload inside our SLA window?” If the answer is “we’d need a three-month project,” that’s not multi-cloud — that’s optimism.
The five-year read for an Australian executive
Between now and 2028, three things will happen at once:
- Ironwood-class inference gets commoditised. Every hyperscaler will have a TPU 7-equivalent capability at broadly comparable price/performance. This year and next is when the market is still legible.
- The eighth-generation TPUs (Sunfish training, Zebrafish inference) ship into Google Cloud and Google’s own internal Gemini workloads. The gap between “you use Nvidia GB200” and “we use our own silicon plus Nvidia” widens for Google.
- Frozen v2 lands in 2028 — model-specific silicon that only makes sense if you’re running Gemini. From that point forward, “cheaper to run Gemini” and “impossible to compare Gemini” are the same statement.
For an Australian executive planning AI infrastructure across that window: the discipline that matters is architectural — not vendor selection. Pick the vendor whose current capability fits, but architect as if you’ll have to move once inside the five-year window. Because you might.
The bottom line
Frozen v2 is not a strategic event on its own. It is the most visible surface right now of a market in which the largest AI providers are becoming more vertically integrated every quarter. The question the reporting should trigger in a CTO’s mind is not “what does this chip do?” It’s “am I still architected for a market where I can move?”
That is the conversation to have with your AI leadership team this week.
Related reading
- Primary source: Reuters — Google plans new chip to run Gemini models more efficiently, The Information reports (20 July 2026)
- Coverage: Bloomberg · TechCrunch · Bloomberg-cited AI Insider write-up · Crypto Briefing — Frozen v2 codename detail
- Google’s own TPU context: Ironwood: The first Google TPU for the age of inference · Google previews TPU 8t/8i (Sunfish/Zebrafish)
- My adjacent post: Adversarial Verification and Fail-Closed Gates in a Production AI Pipeline — same “assume-adversarial” discipline applied to vendor-lock-in risk
For technical deep-dives on the cloud and IT topics I cover strategically, visit Cloud Geeks — our specialist IT infrastructure blog.
Ganda Tech Services is my technology consultancy, bringing together cloud infrastructure, web development, and mobile expertise for Australian businesses.
AI Strategy Primer for Australian Business Leaders
A practical framework for AI adoption in 2026 — cut through the hype and start with what matters.