Raleigh News Today

collapse
Home / Daily News Analysis / Google is working on a new AI chip designed to make Gemini more efficient

Google is working on a new AI chip designed to make Gemini more efficient

Jul 24, 2026  Twila Rosenbaum  5 views
Google is working on a new AI chip designed to make Gemini more efficient

Alphabet, Google's parent company, is designing a new server chip to help its in-house Gemini models operate more efficiently. The chip, internally dubbed "Frozen v2," is slated to be released sometime in 2028, according to a report from The Information, which cited anonymous sources. The new chip could be between six and ten times more efficient than Google's existing AI chips, measured by the number of tokens generated per unit of power.

Google did not directly confirm the report but also did not deny it. In a statement to TechCrunch, the company said: "Our teams are constantly researching and experimenting with new innovations to deliver maximum performance and efficiency for our users and customers. While not every project moves into production, this rigorous exploration is central to our full stack approach. By co-designing our hardware and software from the ground up, we ensure our systems are integrated and highly optimized for real-world workloads."

Background on Google's AI Chip Efforts

Google has been developing custom chips for artificial intelligence for years. Its Tensor Processing Units (TPUs) were first introduced in 2016 and have been a cornerstone of the company's AI infrastructure. The TPU line has evolved through several generations, with the latest TPU v5p offering significant performance gains for training and inference. However, the Frozen v2 appears to be a separate effort specifically optimized for running Gemini models, focusing on inference efficiency rather than general-purpose AI workloads.

The reported efficiency gains of six to ten times would represent a major leap over current silicon. Industry experts note that improving token-per-watt ratios is critical as AI models grow larger and consume more energy. Google's Gemini models, which power everything from search to cloud services, require massive computational resources. A more efficient chip could dramatically reduce operational costs and environmental impact.

The Industry Shift Toward Custom Silicon

AI companies have increasingly sought to produce their own chips as a way to make their in-house models run more efficiently and to address global shortages in AI computing capacity. Such efficiency has become a key selling point for tech companies as concerns about AI spend have dampened the market euphoria that previously characterized the industry. At the same time, firms are engaged in an ongoing attempt to wean themselves off chipmaker Nvidia, which has historically dominated the AI chip market and whose dominance has left major AI makers dependent on its hardware.

Nvidia's GPUs, particularly the H100 and upcoming B100, are the gold standard for AI training and inference. But their high cost and limited availability have pushed companies like Google, Amazon, Microsoft, and OpenAI to design their own accelerators. In June, OpenAI announced its first custom chip, an inference processor dubbed Jalapeño. Earlier this month, it was reported that Anthropic was discussing a new chipmaking partnership with Samsung. These moves signal a broader trend toward vertical integration in the AI industry.

Financial Implications for Alphabet

Investors have previously worried about Alphabet's massive planned expenditures designed to help it build out its AI strategy. Earlier this year, Google said that it plans to spend between $180 billion and $190 billion on capital expenditures through 2025. With so much money at stake, the company needs to prove that those investments will pay off. News of the more efficient Frozen v2 chip appears to have assuaged investors, giving Google a boost ahead of its earnings report later this week. Following publication of The Information's report, the company's stock climbed some 3% on Monday morning.

Analysts suggest that a dedicated inference chip could improve Google's profit margins on AI services. Currently, much of the cost of running Gemini comes from compute time on third-party hardware or older TPUs. If Frozen v2 delivers on its promise, it could reduce the cost per query significantly, enabling Google to offer competitive pricing while maintaining profitability.

Technical Details and Performance Metrics

While full specifications have not been disclosed, the report indicates that Frozen v2 is designed to excel at token generation, the process by which AI models produce text, code, or other outputs. Efficiency is measured in tokens per watt, meaning the chip can generate more responses using less electricity. Given that inference workloads are becoming a larger share of AI compute demand, such optimizations are increasingly valuable.

Google's existing AI chips, like the Edge TPU and Cloud TPU v5p, are already considered among the most efficient in the industry. However, they were designed for a broader range of tasks. Frozen v2 appears to be a purpose-built chip for large language models, similar to how specialized ASICs (application-specific integrated circuits) have been used for cryptocurrency mining or video transcoding. By tailoring the architecture to the specific mathematical operations required by Transformer models, Google can achieve higher performance per watt.

Competitive Landscape

The race to develop custom AI chips is heating up. Amazon has its Trainium and Inferentia chips, used in AWS. Microsoft is reportedly working on a chip code-named "Athena" in collaboration with AMD. Meta has also been designing custom silicon for AI. Even Apple has its Neural Engine for on-device AI. However, Google's advantage lies in its deep expertise with TPUs and its ability to co-design hardware and software. The company's TensorFlow and JAX frameworks are tightly integrated with its hardware, offering unique optimizations.

Meanwhile, Nvidia is not standing still. The company recently announced plans to accelerate its chip release schedule to a yearly cadence, with new architectures like Blackwell and Rubin. Nvidia's CUDA software ecosystem remains a powerful moat, making it difficult for customers to switch. However, cloud giants are motivated to reduce their reliance on a single supplier, which gives custom chips like Frozen v2 a strategic value beyond pure performance.

The success of Google's chip will also depend on manufacturing. Most custom AI chips are fabricated by Taiwan Semiconductor Manufacturing Company (TSMC). Geopolitical tensions and supply chain constraints continue to pose risks. Google has reportedly reserved capacity at TSMC's advanced nodes, but the timeline to 2028 leaves room for unforeseen delays.

Impact on Gemini and Google's AI Strategy

Gemini is Google's flagship family of AI models, spanning sizes from Nano to Ultra. The models are used in products like Google Search, Bard (now Gemini chatbot), Google Workspace, and cloud APIs. As competition with OpenAI, Anthropic, and Meta intensifies, Google needs to continually improve model quality while controlling costs. A custom chip that offers a step-change in efficiency could give Google a pricing advantage and enable more complex models to be deployed at scale.

Additionally, having proprietary hardware allows Google to optimize the entire stack. The company can design the chip in parallel with model development, ensuring that Gemini's architecture aligns with hardware capabilities. This holistic approach is similar to what Apple does with its A-series and M-series chips. Early leaks suggest that Gemini 3 may incorporate features that specifically leverage Frozen v2's architecture, such as sparse attention mechanisms or mixture-of-experts routing.

The chip may also have implications for Google's data center expansion. If each chip consumes less power for the same output, Google can reduce cooling requirements, lower energy costs, and fit more compute into the same physical footprint. This could help the company meet its sustainability goals while scaling up AI capacity. Google has committed to operating on 24/7 carbon-free energy by 2030, and more efficient chips are a key part of that plan.

News of Frozen v2 has also sparked speculation about potential third-party availability. While Google primarily uses its own chips internally, it has sold access to TPUs through Google Cloud. If Frozen v2 offers dramatic efficiency improvements, it could attract external customers looking to run large language models more cheaply. However, given the strategic importance of Gemini, Google may keep the chip exclusive initially.

The timeline of 2028 means that Google will continue to rely on external suppliers and its existing TPU line for the next few years. During that period, Nvidia will likely release new generations, and other competitors will also bring custom chips to market. Nevertheless, the early disclosure of Frozen v2 appears to be a signal to both investors and rivals that Google is serious about hardware innovation. The stock market's positive reaction indicates that Wall Street sees this as a prudent long-term bet.

In the broader context, the AI industry is undergoing a fundamental shift from relying on general-purpose hardware to developing specialized silicon for specific models. This trend mirrors the evolution of other computing domains, such as graphics processing and mobile computing. As models become more diverse and application-specific, hardware differentiation will become a key competitive advantage. Google's Frozen v2 project is a bet that by tightly integrating chip design with model architecture, it can achieve performance gains that far outpace general-purpose alternatives.


Source: TechCrunch News


Share:

Your experience on this site will be improved by allowing cookies Cookie Policy