Google has released Gemini 3.6 Flash and 3.5 Flash-Lite as new workhorses designed to cut latency and token costs for enterprise AI agents.
The economics of running autonomous software agents inside a production environment come down to a fixed equation few vendors advertise directly. A model needs to reason through a multi-step task competently, but every extra token it generates while doing so adds cost and delay to a workflow that might run thousands of times an hour.
Teams building background agents rather than chat interfaces need throughput first and parameter count second. Google’s answer, announced this week, splits that trade-off across three models: Gemini 3.6 Flash for coding and multimodal reasoning, Gemini 3.5 Flash-Lite for high-volume, low-latency work, and a restricted Gemini 3.5 Flash Cyber variant built for vulnerability remediation.
The math behind Gemini 3.6 Flash
Google’s developer documentation for 3.6 Flash centres on one figure: 17 percent fewer output tokens than the prior 3.5 Flash version, based on measurements from the Artificial Analysis Index.
In specific synthetic tests, including the Datacurve DeepSWE benchmark, Google reports drops in token usage of up to 65 percent. Pricing sits at $1.50/1M input tokens and $7.50/1M output tokens, positioning the model for reasoning loops that run continuously rather than on-demand.
On DeepSWE, the company records a 49 percent success rate for 3.6 Flash against 37 percent for its predecessor. On MLE Bench, the score moves from 49.7 percent to 63.9 percent, and on Google’s GDPval-AA v2 test – which attempts to measure real-world knowledge work rather than coding puzzles – 3.6 Flash scores 1421 against 1349 for the older model.
Figma, Hebbia, and Harvey put the model to work
Figma has integrated 3.6 Flash into its prototyping infrastructure, and according to Matt Colyer, the company’s Director of Product Engineering, the model gives developers a faster route through design iterations without a drop in output quality.
Legal technology platform Harvey and research tool Hebbia route data through the model for multimodal document work: ingesting raw financial filings, parsing document structure, reading embedded charts, and producing draft reports for review.
Google also folded a client-side computer-use tool directly into the Gemini API and Gemini Enterprise platforms, removing the custom intermediary software engineers previously built to let models operate on top of an operating system.
The company reports an OSWorld-Verified score of 83.0 percent, up from 78.4 percent, and says updated safeguards against chemical, biological, radiological, and nuclear misuse improve resistance to jailbreaking without raising refusal rates for benign requests.
A cheaper tier for high-volume background agents
Gemini 3.5 Flash-Lite targets a different job: document processing and agentic search running at volume rather than reasoning depth. The Artificial Analysis Index measured the model at 350…
Source link
Disclaimer
We strive to uphold the highest ethical standards in all of our reporting and coverage. We blogs.grocliq.com want to be transparent with our readers about any potential conflicts of interest that may arise in our work. It’s possible that some of the investors we feature may have connections to other businesses, including competitors or companies we write about. However, we want to assure our readers that this will not have any impact on the integrity or impartiality of our reporting. We are committed to delivering accurate, unbiased news and information to our audience, and we will continue to uphold our ethics and principles in all of our work. Thank you for your trust and support.
Website Upgradation is going on for any glitch kindly connect at [email protected]