Google is developing a new server chip, internally known by code name Frozen v2, designed to engrave the core structure of the Gemini model directly onto the hardware.
According to rumors reported by The Information, the solution could increase the efficiency of the AI infrastructure up to 10 times compared to the current generation of proprietary components, processing a significantly higher number of tokens per watt consumed.
The project, which aims for an initial rollout by 2028, is designed to overcome the current bottlenecks in the data center network and has already drawn positive reactions on Wall Street, where Alphabet’s stock has risen notably.
The peculiarity of Frozen v2 lies in the choice to partially surpass the flexibility of generic chips such as Nvidia GPUs or Google’s own classic Tensor Processing Units.
In traditional processors, each operation requires continuous real-time logic checks and constant data transfers between memory and compute cores. The new project, born from the idea of chief scientist Jeff Dean, instead chooses to fix the Gemini family’s computational logic inside the physical circuit.
Wiring the model’s structure directly into the semiconductors drastically reduces intermediate steps and data movements. Technicians’ estimates indicate that, at the same power consumption, the volume of tokens processed could be between 6 and 10 times more efficient than existing TPUs.
This approach does not aim to replace the current line of proprietary accelerators, such as the TPU 8i series used for general inference, but aims to accompany it by creating a highly specialized product line to handle the high traffic of user queries.
The push toward such advanced solutions arises from the need to address a strong saturation of compute power inside. The constant growth in model sizes and the increase in requests has created strong tensions on the global network’s operational capacity, leading Google Cloud in some cases to have to decline orders from external clients.
In this context, optimizing yield per watt is the key to extending service delivery without necessarily waiting for the construction of new energy infrastructure.
Integrating software directly into the semiconductors does however come with precise constraints. The silicon design and production timelines take several years, while AI models evolve at monthly cadences.
If in the future the Gemini structure were to undergo substantial modifications, tailored chips could risk losing effectiveness or require complex adaptations.
To mitigate this factor, management is keeping Frozen v2 for now in an experimental platform state, still ensuring the possibility of updating the weights and parameters of the model integrated into the matrix.
Company spokespeople have confirmed that research teams continually explore new synergies between hardware and software to optimize real workloads, noting that not all development projects necessarily reach the mass production phase.
If field performance is confirmed by 2028, the technology will enable Gemini’s responses to be delivered with unprecedented efficiency, leaving traditional TPUs to handle training activities and more flexible workloads.
The first render images of Samsung Galaxy S27 Pro and Ultra, circulated in the past…
Xiaomi's brand partner has officially launched in India POCO X8 Power, a mid-range smartphone that…
OpenAI unveiled GPT-6 Astra, the new model that the company describes as the most advanced…
Amazon has announced a new feature designed to fight scams: starting today (but only in…
Midea porta a IFA 2026 la propria visione "Simply ideal" per la casa connessa: no…
Ecovacs has just earned recognition as the world's No. 1 brand for home robotics in…