Google reportedly developing ‘Frozen v2’ chip with Gemini’s architecture etched into the silicon —…

Google is actively developing a highly specialized server chip, informally known as "Frozen v2," designed to embed critical components of its Gemini large language model’s architecture directly into the silicon. This ambitious project, first reported by The Information based on insights from sources familiar with the matter, represents a significant strategic move by the tech giant to address the burgeoning demand for artificial intelligence compute power and enhance the efficiency of its AI services. Engineers involved in the "Frozen v2" initiative project that this custom silicon could deliver a remarkable six to ten times more tokens per unit of power compared to the latest generation of Google’s Tensor Processing Units (TPUs). The company is reportedly targeting deployment for this innovative chip as early as 2028, underscoring the urgency and strategic importance of the endeavor. The development comes as Google Cloud faces an acute shortage of AI compute resources, a situation so severe that it has reportedly led to the rejection of potential deals with external clients.
The Genesis of Specialized AI Hardware
The drive toward highly specialized silicon like "Frozen v2" is a direct consequence of the explosive growth in artificial intelligence, particularly generative AI models such as Google’s Gemini. These models, characterized by their massive parameter counts and complex architectures, demand unprecedented levels of computational power for both training and inference. While general-purpose graphics processing units (GPUs) from companies like Nvidia have dominated the AI landscape, their architecture, designed for parallel graphics processing, is not always optimally efficient for the specific matrix multiplication and tensor operations central to neural networks. This mismatch has led to an industry-wide scramble for more efficient, purpose-built hardware.
Google has been at the forefront of this movement for nearly a decade, pioneering its Tensor Processing Units (TPUs) in 2016. TPUs were designed from the ground up to accelerate machine learning workloads, initially focusing on inference for services like Google Search and then evolving to handle complex training tasks for models like AlphaGo. Successive generations of TPUs have consistently pushed the boundaries of AI performance per watt and cost efficiency. However, even with its custom TPU advantage, Google, much like its cloud competitors Amazon Web Services (AWS) with Inferentia and Trainium, and Microsoft with Maia and Athena, finds itself grappling with an insatiable demand for AI compute. The current shortage is not merely an inconvenience; it represents a significant bottleneck for innovation and scalability in the AI sector, impacting cloud providers’ ability to serve both internal projects and external enterprise customers. The "Frozen v2" project is Google’s latest, and perhaps most aggressive, response to this escalating challenge.
Unpacking Frozen v2: A Leap in Inference Efficiency
At its core, "Frozen v2" represents a paradigm shift in how AI models interact with hardware during inference. Traditional AI accelerators, including Google’s own TPUs and most GPUs, are designed to be programmable. This means they can load and execute any compatible AI model, making runtime decisions about data flow, memory access, and computational operations as they process each query. While this flexibility is invaluable for versatility, it introduces overhead in terms of power consumption and latency. Each decision made by the hardware during execution consumes time and energy, as data is constantly shuttled between various computational units and memory banks.
"Frozen v2" aims to circumvent some of this overhead by "freezing" or embedding certain architectural aspects of the Gemini model directly into the chip’s transistors. This means that a portion of the model’s structure – the way its layers connect, the specific types of mathematical operations it performs, and even optimized data pathways – would be hardwired into the silicon itself. By fixing these decisions in the hardware, the chip can eliminate many of the dynamic runtime steps that general-purpose accelerators must undertake. This results in a more streamlined execution path, significantly reducing the number of computational steps and the volume of data that needs to be moved around for each query or "token."
The projected efficiency gains are substantial: six to ten times more tokens per unit of power. This metric is crucial in the context of large-scale AI deployment. "Tokens" represent the fundamental units of text or data processed by an AI model, and increasing the number of tokens processed per unit of power directly translates to lower operational costs, reduced energy consumption in data centers, and a smaller carbon footprint. Furthermore, by cutting down response latency, "Frozen v2" could unlock new categories of real-time AI applications that are currently constrained by processing delays. Imagine conversational AI systems with instantaneous responses, autonomous vehicles making split-second decisions with enhanced perception models, or highly responsive virtual assistants that feel indistinguishable from human interaction. These applications demand minimal latency, and specialized inference chips like "Frozen v2" could be the key to their widespread realization.
From "Frozen" to "Frozen v2": An Iterative Design Philosophy
The concept of baking AI model elements into silicon is not entirely new for Google. The "Frozen v2" project is, as its name suggests, a successor to an earlier internal initiative, the original "Frozen" design. This initial concept, reportedly spearheaded by Google DeepMind chief scientist Jeff Dean, was even more radical: it proposed to bake the actual weights of the Gemini model directly into the chip. The weights are the numerical parameters that define what an AI model has learned, essentially its "knowledge." Embedding these directly would have created an incredibly fast and efficient chip for that specific version of the Gemini model.
However, Google ultimately set aside the original "Frozen" proposal. The primary concern was the rapid pace of AI model development. Large language models like Gemini are constantly being updated, refined, and improved. A chip hardwired with the weights of a single model version would quickly become obsolete, rendering the expensive custom silicon largely useless as new, superior iterations of Gemini emerged. The lifecycle of such a chip would be too short to justify the enormous investment in research, development, and manufacturing.

The pivot to "Frozen v2" reflects a more pragmatic and sustainable approach. Instead of locking in the model’s weights, which are dynamic, "Frozen v2" focuses on freezing the model’s architecture. The architecture defines the fundamental structure of the neural network – the types of layers, how they connect, and the flow of data. While model architectures do evolve, they tend to do so at a slower pace than the model weights themselves. This allows "Frozen v2" to remain useful across multiple releases of the Gemini model, as long as those releases are built upon the same underlying architectural principles. The weights, crucial for the model’s specific knowledge, would remain updatable and programmable, providing the necessary flexibility. Sources indicate that the precise extent to which the model’s architecture will be locked into the silicon is still under active deliberation, balancing the desire for maximum efficiency with the need for future adaptability. This iterative design process highlights the complexities of hardware-software co-design in a field as dynamic as AI.
Google’s Holistic Hardware Ecosystem
The introduction of "Frozen v2" should not be seen as a replacement for Google’s existing Tensor Processing Unit (TPU) line but rather as an augmentation and a step towards even greater specialization. Google’s TPU strategy has been continually evolving, with the eighth generation, announced at Cloud Next in April, splitting into distinct variants optimized specifically for training and inference. This segmentation itself represents a move towards greater efficiency, acknowledging that the computational demands of teaching an AI model (training) are different from those of using it to make predictions (inference).
"Frozen v2" is envisioned to sit alongside this evolving TPU ecosystem. Google reportedly does not plan to produce "Frozen v2" at the same massive volumes as its general-purpose TPUs. Instead, this generation is seen partly as a "trial run" for more specialized silicon, a proving ground for highly tailored architectures as AI model designs gradually mature and stabilize. This strategy suggests Google is exploring the boundaries of hardware specialization, reserving the most extreme optimizations for high-impact, internal applications initially, before potentially scaling them more broadly.
Google’s commitment to custom silicon is also evident in its supply chain strategies. Reports indicate that Google has already booked Intel to package more than 3 million TPUs in 2028, the same year "Frozen v2" is targeted for deployment. This massive procurement deal underscores Google’s broader "full-stack approach" to AI, wherein it seeks to control and optimize every layer of the technology stack, from the underlying silicon to the software frameworks and the AI models themselves. By designing its own chips, Google can ensure that its hardware is perfectly tailored to its software and models, theoretically achieving performance and efficiency levels that would be difficult to match with off-the-shelf components.
The Competitive Landscape: A Race for AI Silicon Dominance
Google’s "Frozen v2" project emerges within a fiercely competitive landscape where every major tech player is vying for an edge in AI compute. The market for AI chips, encompassing GPUs, TPUs, and other specialized accelerators, is projected to grow from tens of billions of dollars today to hundreds of billions within the next decade. Nvidia currently dominates this market, especially for AI training, with its powerful GPUs and comprehensive CUDA software ecosystem. However, the immense costs and power consumption associated with large-scale AI inference are driving innovation across the board, leading to a proliferation of custom and specialized solutions.
Several companies are already demonstrating or deploying model-hardwired inference silicon, validating the general direction Google is pursuing. Taalas, a Toronto-based startup that has secured over $200 million in funding, launched its HC1 chip in February. This chip features the Llama 3.1 8B model permanently wired into an 815mm-squared die, manufactured on TSMC’s N6 process. Taalas claims impressive performance figures, boasting 17,000 tokens per second per user without the need for high-bandwidth memory (HBM) on the package, a significant factor in reducing power and cost. This approach represents the extreme end of specialization, where an entire model is etched into silicon.
Another notable player is Groq, which has developed a Language Processing Unit (LPU) architecture known for its low latency and high throughput, particularly for large language model inference. Nvidia, recognizing the potential of such specialized architectures, struck a significant $20 billion deal in December to license technology from Groq. This move highlights Nvidia’s strategy to not only dominate with its general-purpose GPUs but also to incorporate or acquire promising specialized inference technologies to maintain its market leadership.
Beyond these examples, other tech giants are also heavily invested in custom silicon. AWS continues to expand its Inferentia (for inference) and Trainium (for training) chip families, custom-designed for its cloud workloads. Microsoft has introduced its own custom AI chips, Maia (for AI acceleration) and Athena (for cloud infrastructure), to power its Azure AI services and reduce reliance on external vendors. The trend is clear: as AI models become more pervasive and complex, the industry is moving towards a diverse ecosystem of highly optimized, domain-specific architectures tailored to particular AI tasks, rather than a one-size-fits-all solution. "Frozen v2" positions Google firmly within this leading edge of hardware-software co-design.
Strategic Implications and Future Outlook
The successful deployment of "Frozen v2" could have profound implications across several domains.

For Google Cloud: It would significantly enhance Google Cloud’s competitive offering for AI workloads, particularly for inference-heavy applications requiring low latency and high throughput. By providing a dramatically more efficient solution for Gemini-based services, Google Cloud could attract enterprises seeking to deploy large-scale, real-time AI. It would also alleviate the internal compute shortage, allowing Google to scale its own AI products and services more rapidly and reliably, potentially reducing the operational costs associated with running its vast AI infrastructure. This differentiation could be a crucial factor in the increasingly crowded cloud market.
For AI Development: The availability of highly optimized inference hardware could spur the creation of entirely new categories of AI applications. Developers might no longer be constrained by the latency and cost of general-purpose compute, enabling more ambitious real-time AI interactions, more complex agentic systems, and more immersive AI experiences. It also reinforces the trend of hardware-software co-design becoming paramount in AI, pushing developers and researchers to consider the underlying silicon from the outset of model development.
For the Semiconductor Industry: "Frozen v2" signals a continued diversification of the AI chip market. While Nvidia’s dominance in training may persist, the inference market is ripe for disruption by specialized architectures. This could lead to increased competition, foster innovation in chip design and manufacturing processes, and create opportunities for new players focusing on specific AI model types or applications. It also highlights the growing importance of foundry services (like TSMC) capable of producing such complex, custom silicon.
Challenges and Risks: Despite the promising projections, the "Frozen v2" project faces inherent challenges. The primary risk remains the rapid evolution of AI models. While freezing architecture offers more longevity than freezing weights, architectural paradigms in AI are not entirely static. A radical shift in model design in the coming years could still render a highly specialized chip less optimal or even obsolete. Manufacturing complex custom silicon at scale is also an enormous undertaking, fraught with technical and financial risks. Furthermore, balancing specialization with versatility will be a continuous tightrope walk for Google and other developers of custom AI hardware.
Official Stance and Industry Perspective
Google has yet to officially confirm the specifics of the "Frozen v2" project, aligning with its typical stance on unannounced product development. A spokesperson for Google, in response to inquiries from The Information, stated that "not every project moves into production" and that "this rigorous exploration is central to our full stack approach." This statement, while non-committal, underscores Google’s ongoing commitment to exploring innovative hardware solutions across its entire technology stack, from fundamental research to deployment.
The industry, meanwhile, broadly recognizes the imperative for such innovation. The current trajectory of AI development, with ever-larger and more complex models, is unsustainable without significant advancements in computational efficiency. Companies that can design and deploy chips tailored precisely to their AI models will gain a substantial strategic advantage in terms of performance, cost, and power consumption. "Frozen v2," if successful, could solidify Google’s position as a leader not just in AI software, but also in the foundational hardware that powers the AI revolution.
In conclusion, Google’s "Frozen v2" project represents a bold step towards an era of extreme specialization in AI hardware. By embedding the very architecture of its Gemini model into silicon, Google aims to unlock unprecedented levels of efficiency and performance for AI inference. While challenges remain, this initiative highlights Google’s unwavering commitment to its full-stack AI strategy and its proactive response to the escalating global demand for AI compute, potentially reshaping the future of AI applications and the semiconductor industry.







