Frosted green AI chip glowing inside a dark Google data-center server hall

Google Frozen v2 Chip: 6-10x More Efficient Than TPUs

Google’s Frozen v2 Chip Could Quietly Reset What AI Costs

Every so often, a single leak moves billions of dollars before breakfast. The Google Frozen v2 chip did exactly that.

On July 20, 2026, The Information reported that Google is quietly building a new server chip – nicknamed Frozen v2 – that its own engineers think could be six to ten times more power-efficient than the company’s newest TPUs. Naturally, Alphabet’s stock popped the same morning. Consequently, the Google Frozen v2 chip went from zero public awareness to a market-moving story within hours.

Then the caveats started piling up. Google hasn’t confirmed anything. The chip isn’t expected until around 2028. And “six to ten times more efficient” doesn’t mean what most people assume.

So let’s slow down and do the thing the newsletters skipped. What is this chip, really? How does that efficiency claim actually work? And what could it do to your company’s AI bill and to Nvidia? Most importantly, how much of it should you believe today?

Short version: it’s a big deal if it’s real. That “if” is doing a lot of work.

In this article: What Frozen v2 is · what “6–10x efficient” means · how it works · what it means for AI costs · the Nvidia question · why you shouldn’t bet on it yet · the verdict · FAQ.

What Is the Google Frozen v2 Chip?

The Google Frozen v2 chip is a reported custom server processor that permanently embeds parts of the Gemini AI model directly into the silicon. According to The Information (July 2026), Google engineers project it could serve six to ten times more tokens per unit of power than the company’s current TPUs. Google hasn’t officially confirmed the chip.

The name is the clue. “Frozen” refers to freezing a chunk of Gemini’s architecture into the hardware itself, instead of loading it as flexible software every time. It’s reportedly a separate project from Google’s Tensor Processing Units – a new chip family meant to sit alongside TPUs, not replace them. In short, think specialist, not all-rounder.

What “6-10x More Efficient” Actually Means

Here’s the part every one-line summary gets wrong. Efficiency isn’t raw speed, and it isn’t your electricity bill at home. The report measures it in tokens per unit of power – how much AI output you get for every watt burned.

In fact, a token is a small chunk of text, roughly three-quarters of a word. Every time Gemini answers a question, writes code, or summarizes a document, it spits out tokens, and each one costs a sliver of electricity. Multiply that by billions of queries a day and power becomes the single biggest lever on cost.

So “6–10x more efficient” is a claim about cost per token at scale, not a faster laptop or a greener gadget. If a data center can serve ten times the tokens from the same power draw, the whole economics of running AI shift.

Here’s how the reported chip stacks up against what’s shipping today:

Frozen v2 (reported)Google TPU 8t / 8iNvidia GPU
Built forGemini onlyGoogle’s AI workloadsAlmost any AI model
FlexibilityVery low (hard-coded)MediumHigh
Efficiency edge6-10x tokens/watt vs TPU*2.8x Ironwood (training)General-purpose
Who can use itGoogle, internalGoogle + some cloud usersAnyone who can buy or rent
StatusReported, ~2028ShippingShipping

*Projected by Google engineers, per The Information. Unconfirmed.

Diagram comparing Frozen v2, Google TPU, and Nvidia GPU on efficiency and flexibility

Read that “flexibility” row carefully – it’s the whole trade-off. Frozen v2 buys efficiency by giving up versatility.

How the Google Frozen v2 Chip Works: Gemini Baked Into the Silicon

To begin with, most AI chips are generalists. They’re flexible slabs of math that can run today’s model, next year’s model, and a rival’s model too. That flexibility costs power, because the chip constantly shuffles model weights and instructions in and out of memory.

Frozen v2 reportedly takes the opposite bet. It hard-codes parts of Gemini’s structure straight into the circuits. Less shuffling. Fewer bits moving around. And far less wasted energy. As a result, the chip does one thing – run Gemini – and does it with far less overhead.

If that sounds familiar, it’s the same logic Apple used by designing its own chips for its own software. Tighten the loop between hardware and model, and you squeeze out efficiency a general-purpose chip can’t match.

Still, there’s an obvious catch, and it’s baked right into the name. Freeze a model into silicon and you’d better be sure that model isn’t about to change. We’ll come back to that.

Why This Could Change What US and UK Companies Pay for AI

Here’s where it gets real for businesses in New York, London, and everywhere between.

AI’s dirty secret in 2026 isn’t that intelligence is expensive to build it’s that it’s expensive to run. Inference, the everyday act of answering prompts, now eats roughly two-thirds of AI compute spending. And those bills are climbing fast: one 2026 industry report pegged the average enterprise AI budget jumping from about $1.2 million in 2024 to $7 million this year.

Now layer on the energy math. Data centers already burned around 415 terawatt-hours of electricity in 2024 – about 1.5% of the world’s power and the International Energy Agency projects that could roughly double to 950 TWh by 2030, with AI as the main driver. Power is the ceiling on how much AI the world can afford.

A chip that serves 6-10x more tokens per watt attacks that ceiling directly. If Google runs Gemini far cheaper, it can pass savings through Google Cloud, undercut rivals, or simply pocket fatter margins. Either way, the price you pay for AI features – in your CRM, your help desk, your coding tools – sits downstream of what it costs the people running the models. Cheaper silicon eventually lands on your invoice.

For example, picture a mid-size UK software firm paying five figures a month for AI features it now can’t live without. A step-change in efficiency upstream is the difference between that bill ballooning and that bill finally easing.

What the Google Frozen v2 Chip Means for Nvidia

Let’s address the name on everyone’s lips. Nvidia.

Indeed, Nvidia still owns this market. Bloomberg Intelligence projects it’ll hold 70–75% of the AI chip market through 2030, and its general-purpose GPUs remain the default for training and running almost any model. One rumored, Gemini-only chip arriving in 2028 doesn’t dent that.

But the direction of travel matters. The Google Frozen v2 chip is another sign the biggest AI players want off the Nvidia treadmill. After all, Google already builds TPUs. Likewise, Amazon has Trainium, Microsoft has Maia, and Meta has MTIA. Every custom chip a hyperscaler builds is compute it no longer rents from Nvidia. Analysts have floated scenarios where Nvidia’s inference share slides toward 20–30% by 2028 as this captive silicon scales.

Here’s the nuance the headlines miss: chips like Frozen v2 aren’t for sale. They’re captive built by Google, for Google. If you’re not inside that ecosystem, the chip doesn’t exist for you, and you’re still buying Nvidia. So Frozen v2 threatens Nvidia’s volume at the very top of the market, not its product for everyone else.

That’s why the stock story cut both ways: good for Alphabet, mildly annoying for Nvidia, and nowhere near the death blow some posts implied.

Enjoying the no-hype version? Bookmark Nexvolu’s AI coverage – we break down each twist in the chip-war saga as it lands.

The Catch: Why You Shouldn’t Bet on This Yet

Time for the reality check the hype cycle skipped.

It’s unconfirmed. The whole story rests on The Information‘s reporting and anonymous sources. Google gave a carefully worded non-answer about co-designing hardware and software, and didn’t confirm the specifics. Reuters, CNBC, and TechCrunch all framed it as a report, not a fact. So should you.

It’s years away. The report pegs deployment at around 2028. In AI terms, that’s practically a geological age. Nvidia will have new architectures by then, and Google’s own TPUs will have moved on too.

The numbers are projections. “6-10x” is an internal engineering estimate for a chip that may not exist in final form. Real-world efficiency has a habit of landing below the lab pitch.

The “frozen” bet could backfire. Hard-code Gemini into silicon and you’re betting the model’s shape won’t change much. If Gemini’s architecture shifts significantly, those chips risk turning into expensive paperweights. Reporting notes it’s essentially a technical test platform for now.

None of this means the story’s fake. It simply means it’s early. So treat the Google Frozen v2 chip as a strong signal of where Google’s headed, not a product you can plan around.

Nexvolu’s Verdict on the Google Frozen v2 Chip

The verdict: A genuinely exciting efficiency bet that could reshape AI economics – if the 6-10x claim survives three years and ships as promised.

Best for: readers tracking the AI-chip arms race and Google Cloud’s long game. Skip it if: you need something actionable this quarter – this is a 2028 story.

Pros: attacks AI’s real bottleneck (power, not speed) · smart vertical-integration play, Apple-style · pressures Nvidia’s long-term grip.

Cons: entirely unconfirmed · years from deployment · the “frozen” design is brittle if Gemini changes.

Standout: the report measures efficiency in tokens per watt – a cost metric, not a speed metric. That single distinction is what most coverage botched.

Nexvolu Editorial Score: 7/10 – a high-upside, high-uncertainty story; thrilling as a signal, unprovable as a spec. (Editorial opinion based on public reporting, not hands-on testing.)

Frequently Asked Questions

What is the Google Frozen v2 chip?

The Google Frozen v2 chip is a reported custom chip that embeds parts of the Gemini AI model directly into silicon, so it runs Gemini using far less power. The Information reported it in July 2026, with engineers projecting 6-10x better efficiency than current TPUs. Google hasn’t confirmed it.

In practice, that makes it a specialist chip, not a general one. It’s reportedly a separate family from Google’s TPUs and is meant to complement them rather than replace them. Because it’s tuned for a single model, it trades flexibility for efficiency – great for running Gemini at scale, useless for anything else. Expected deployment is around 2028.

Is the Frozen v2 chip confirmed by Google?

No. As of July 2026, Frozen v2 remains unconfirmed. The claim comes from The Information, citing anonymous sources; Reuters, CNBC, and TechCrunch then picked it up. Google offered a general statement about co-designing hardware and software but didn’t confirm the chip or its specs.

For instance, a Google Cloud spokesperson said the company constantly researches new innovations and optimizes hardware and software together – a non-denial that also isn’t a confirmation. TechCrunch noted Google “didn’t deny it either.” Until Google publishes official specs, treat every figure as a projection. Reported, not confirmed.

When will Frozen v2 be released?

The report suggests Google aims to deploy Frozen v2 as soon as 2028. That’s a target, not a promise, and chip timelines slip often. Right now it reportedly functions as a technical test platform rather than a finished product.

Between now and 2028, Google’s TPUs will keep advancing and Nvidia will ship new architectures, so Frozen v2’s real-world advantage will be measured against 2028’s competition – not today’s hardware. Plan around it only loosely, if at all.

What does “6-10x more efficient” actually mean?

It means tokens per unit of power, not raw speed. Google engineers reportedly project Frozen v2 could serve six to ten times more AI tokens for the same electricity as current TPUs. It’s a cost-and-energy claim, measured at data-center scale.

A token is roughly three-quarters of a word. Because power is the biggest ongoing cost of running AI, squeezing more tokens from each watt lowers the cost per query. If accurate, that reshapes the economics of serving Gemini – but it says nothing about your personal device, your internet speed, or how “smart” the model feels.

Will Frozen v2 hurt Nvidia?

Not much in the short term. Nvidia is projected to hold 70-75% of the AI chip market through 2030, and Frozen v2 is a single Gemini-only chip that’s years away. It’s a pressure signal, not a threat to Nvidia’s core business.

The bigger picture is that every hyperscaler – Google, Amazon, Microsoft, Meta – is building custom silicon to rent less from Nvidia. Some analysts see Nvidia’s inference share sliding toward 20–30% by 2028 as this captive hardware scales. But these chips aren’t sold externally, so most companies will keep buying Nvidia regardless.

Can I buy or rent a Frozen v2 chip?

No. Frozen v2 is captive silicon – built by Google, for Google’s own Gemini workloads. There’s no indication it will ever be sold or rented directly, and it doesn’t exist as a shipping product yet.

You’d only feel its effect indirectly. If Frozen v2 makes running Gemini cheaper, those savings could show up as lower prices or better margins across Google Cloud and Gemini-powered products. For chips you can actually rent today, Google’s TPUs and Nvidia’s GPUs remain the realistic options.

Professional at night studying rising AI cost charts on a glowing office monitor

The Bottom Line

So where does that leave us?

The Google Frozen v2 chip is the most interesting AI-chip story of the summer, and also the least certain. Of course, the idea – freeze a model into silicon, chase 6–10x efficiency, attack AI’s real bottleneck is genuinely clever. The evidence is a single report, a stock bump, and a 2028 timeline.

Two takeaways worth keeping. First, the industry now fights for AI’s future on efficiency, not raw power – so tokens per watt is the number to watch. Second, the chip that quietly changes your AI bill probably won’t be one you ever buy; it’ll be one a hyperscaler builds for itself.

Either way, watch this space – but don’t rearrange your roadmap around a rumor.

Found this useful? Share it with whoever on your team keeps forwarding AI-chip headlines and subscribe to Nexvolu for the next explainer.

What’s your take is baking a model into silicon a masterstroke, or a bet that ages badly by 2028? Drop your prediction in the comments.

References

  • Google Plans New ‘Frozen’ Chip to Run Its AI Models Much More Efficiently (Erin Woo & Qianer Liu, Jul 20, 2026): The Information
  • Google plans new chip to run Gemini models more efficiently, The Information reports (Jul 20, 2026): Reuters
  • Alphabet stock pops on report it’s developing a more efficient AI chip (Jul 20, 2026): CNBC
  • Google is working on a new AI chip designed to make Gemini more efficient (Jul 20, 2026): TechCrunch
  • Energy and AI: Energy demand from AI (2025/2026): IEA
  • Our eighth generation TPUs (8t and 8i): Google Cloud / The Keyword
  • Nvidia owns the AI chips market (Bloomberg Intelligence 70-75% through 2030): Quartz

⚠️ The enterprise AI-budget figure ($1.2M→$7M) comes from a secondary 2026 industry report and is presented as “one industry report,” not a primary source. Verify or drop before publishing if you want only primary citations.

Explore More AI & Technolgy Insights

Leave a Reply

Your email address will not be published. Required fields are marked *