Google is expanding its line of AI models by offering users increasingly scalable and cost-efficient tools from a computational cost perspective.
In addition to having started the pre-training phase of Gemini 4, the Mountain View company has just made available two optimized versions, named Gemini 3.6 Flash and 3.5 Flash-Lite, designed specifically for those developing complex agents or managing enormous volumes of data.
In addition to these public variants, there is a cybersecurity-specialized version, Gemini 3.5 Flash Cyber, highlighting a strategy that decisively targets reducing processing costs and improving response speed.
Gemini 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyber official and already available
The main novelty for developers and businesses is represented by Gemini 3.6 Flash, described as a genuine daily work tool. Its primary advantage lies in the saving of the resources required to complete structured operations.
According to data provided by Google, this model consumes 17% fewer output tokens than its direct predecessor, with reduction peaks reaching 65% in specific tests evaluating code-writing performance.
From a business perspective, the cost of input remains unchanged at $1.50 per million tokens, but the output rate drops to $7.50. Despite the price cut, measured performance shows objective improvements.
The model supports up to 1 million input tokens and 64,000 output tokens, directly integrating native tools for navigation operations and interaction with the user interface.
Velocità e volumi con 3.5 Flash-Lite
For applications that require ultra-low latency and high concurrent processing capacity, Google has introduced Gemini 3.5 Flash-Lite. This version is designed to generate up to 350 tokens per second, making it the optimal solution for repetitive tasks such as processing long textual documents or rapid information classification.
The prices for this variant are set at $0.30 per million tokens in input and $2.50 in output. Although these figures represent an increase compared to the previous generation, the qualitative leap in logical reasoning justifies the tariff repositioning.
The results of tests for performing prolonged tasks and using the terminal show a nearly doubled score compared with the past. Companies can also modulate the level of system analysis, reducing it to prioritize pure speed or increasing it to handle processes divided into numerous logical steps.
Sicurezza informatica e il ritardo della versione Pro
A particular focus has been devoted to Gemini 3.5 Flash Cyber, a variant trained specifically for the discovery, verification and correction of vulnerabilities of code.
Instead of relying on a single, complex run, the system uses a structure called CodeMender to repeatedly interrogate this lightweight model, exploring in parallel several avenues of investigation.
Due to potential risks related to misuse for cyberattacks, this technology will not be open to the general public, but will be licensed exclusively to governments and trusted partners.

In the meantime, the update to the flagship model, namely Gemini 3.5 Pro, is months behind schedule. Initially planned for last June, it is currently confined to a private testing phase.
Rumor has it that technicians are spending extra time to refine programming-related abilities, without providing a new firm release date at this time.
Lo sguardo verso la prossima generazione
While the more agile variants populate the current market through application programming interfaces and mobile applications, the company is already investing in the architectures of tomorrow.
Google executives have confirmed the official start of the pre-training phase for Gemini 4, defining it as the most extensive and complex computing initiative ever undertaken in their labs.
No specific timelines or estimates of the technical parameters that will comprise the new system have been provided.
The wait is now focused on how these imminent iterations will manage to balance a processing power that is markedly higher with the same optimization of costs and energy consumption that today characterizes the family of models just introduced.



