Google is developing a new server chip, internally known by code name Frozen v2, designed to engrave the core structure of the Gemini model directly onto the hardware.
According to rumors reported by The Information, the solution could increase the efficiency of the AI infrastructure up to 10 times compared to the current generation of proprietary components, processing a significantly higher number of tokens per watt consumed.
The project, which aims for an initial rollout by 2028, is designed to overcome the current bottlenecks in the data center network and has already drawn positive reactions on Wall Street, where Alphabet’s stock has risen notably.
The peculiarity of Frozen v2 lies in the choice to partially surpass the flexibility of generic chips such as Nvidia GPUs or Google’s own classic Tensor Processing Units.
In traditional processors, each operation requires continuous real-time logic checks and constant data transfers between memory and compute cores. The new project, born from the idea of chief scientist Jeff Dean, instead chooses to fix the Gemini family’s computational logic inside the physical circuit.
Wiring the model’s structure directly into the semiconductors drastically reduces intermediate steps and data movements. Technicians’ estimates indicate that, at the same power consumption, the volume of tokens processed could be between 6 and 10 times more efficient than existing TPUs.
This approach does not aim to replace the current line of proprietary accelerators, such as the TPU 8i series used for general inference, but aims to accompany it by creating a highly specialized product line to handle the high traffic of user queries.
The push toward such advanced solutions arises from the need to address a strong saturation of compute power inside. The constant growth in model sizes and the increase in requests has created strong tensions on the global network’s operational capacity, leading Google Cloud in some cases to have to decline orders from external clients.
In this context, optimizing yield per watt is the key to extending service delivery without necessarily waiting for the construction of new energy infrastructure.
Integrating software directly into the semiconductors does however come with precise constraints. The silicon design and production timelines take several years, while AI models evolve at monthly cadences.
If in the future the Gemini structure were to undergo substantial modifications, tailored chips could risk losing effectiveness or require complex adaptations.
To mitigate this factor, management is keeping Frozen v2 for now in an experimental platform state, still ensuring the possibility of updating the weights and parameters of the model integrated into the matrix.
Company spokespeople have confirmed that research teams continually explore new synergies between hardware and software to optimize real workloads, noting that not all development projects necessarily reach the mass production phase.
If field performance is confirmed by 2028, the technology will enable Gemini’s responses to be delivered with unprecedented efficiency, leaving traditional TPUs to handle training activities and more flexible workloads.
The future of OPPO will be in line with that of OnePlus and Realme: according…
The next flagship series from the Chinese company should arrive fully in international markets as…
The need to hide private details before sharing an image is an increasingly common requirement,…
Steam becomes the real protagonist of household cleaning with Roborock F25 Steam, the new vacuum-mop…
Currently the tech sector is facing a period of constant price increases, largely justified by…
Instagram has decided to intervene firmly against a worrying trend: the spread of videos filmed…