Google is developing a new server chip, internally known by code name Frozen v2, designed to engrave the core structure of the Gemini model directly onto the hardware.
According to rumors reported by The Information, the solution could increase the efficiency of the AI infrastructure up to 10 times compared to the current generation of proprietary components, processing a significantly higher number of tokens per watt consumed.
The project, which aims for an initial rollout by 2028, is designed to overcome the current bottlenecks in the data center network and has already drawn positive reactions on Wall Street, where Alphabet’s stock has risen notably.
The peculiarity of Frozen v2 lies in the choice to partially surpass the flexibility of generic chips such as Nvidia GPUs or Google’s own classic Tensor Processing Units.
In traditional processors, each operation requires continuous real-time logic checks and constant data transfers between memory and compute cores. The new project, born from the idea of chief scientist Jeff Dean, instead chooses to fix the Gemini family’s computational logic inside the physical circuit.
Wiring the model’s structure directly into the semiconductors drastically reduces intermediate steps and data movements. Technicians’ estimates indicate that, at the same power consumption, the volume of tokens processed could be between 6 and 10 times more efficient than existing TPUs.
This approach does not aim to replace the current line of proprietary accelerators, such as the TPU 8i series used for general inference, but aims to accompany it by creating a highly specialized product line to handle the high traffic of user queries.
The push toward such advanced solutions arises from the need to address a strong saturation of compute power inside. The constant growth in model sizes and the increase in requests has created strong tensions on the global network’s operational capacity, leading Google Cloud in some cases to have to decline orders from external clients.
In this context, optimizing yield per watt is the key to extending service delivery without necessarily waiting for the construction of new energy infrastructure.
Integrating software directly into the semiconductors does however come with precise constraints. The silicon design and production timelines take several years, while AI models evolve at monthly cadences.
If in the future the Gemini structure were to undergo substantial modifications, tailored chips could risk losing effectiveness or require complex adaptations.
To mitigate this factor, management is keeping Frozen v2 for now in an experimental platform state, still ensuring the possibility of updating the weights and parameters of the model integrated into the matrix.
Company spokespeople have confirmed that research teams continually explore new synergies between hardware and software to optimize real workloads, noting that not all development projects necessarily reach the mass production phase.
If field performance is confirmed by 2028, the technology will enable Gemini’s responses to be delivered with unprecedented efficiency, leaving traditional TPUs to handle training activities and more flexible workloads.
iQOO pushes gaming smartphone concept even further and internally unveils the new concept phone iQOO…
Clearly, underneath the new graphical changes there will still be Android 17. It should be…
Xiaomi officially unveiled HyperOS 4 on August 13, and, just a few hours after the…
The rumors about a "split" launch for the new iPhone lineup are becoming increasingly tangible,…
One of the most anticipated events for those who keep an eye on MediaWorld's offers…
POCO, the Xiaomi sub-brand that built its identity around the formula "top-of-the-line specs, mid-range price",…