|

Google’s Gemini 3.6 Flash targets enterprise agent token costs

Banner for the AI & Big Data Expo event series.

Google has launched Gemini 3.6 Flash and three.5 Flash-Lite as new workhorses designed to chop latency and token costs for enterprise AI brokers.

The economics of working autonomous software program brokers inside a manufacturing atmosphere come right down to a set equation few distributors promote immediately. A mannequin must cause by means of a multi-step activity competently, however each additional token it generates whereas doing so provides price and delay to a workflow which may run 1000’s of instances an hour.

Teams constructing background brokers fairly than chat interfaces want throughput first and parameter depend second. Google’s reply, introduced this week, splits that trade-off throughout three fashions: Gemini 3.6 Flash for coding and multimodal reasoning, Gemini 3.5 Flash-Lite for high-volume, low-latency work, and a restricted Gemini 3.5 Flash Cyber variant constructed for vulnerability remediation.

The math behind Gemini 3.6 Flash

Google’s developer documentation for 3.6 Flash centres on one determine: 17 p.c fewer output tokens than the prior 3.5 Flash model, based mostly on measurements from the Artificial Analysis Index.

In particular artificial assessments, together with the Datacurve DeepSWE benchmark, Google reviews drops in token utilization of as much as 65 p.c. Pricing sits at $1.50/1M enter tokens and $7.50/1M output tokens, positioning the mannequin for reasoning loops that run constantly fairly than on-demand.

On DeepSWE, the corporate information a 49 p.c success fee for 3.6 Flash in opposition to 37 p.c for its predecessor. On MLE Bench, the rating strikes from 49.7 p.c to 63.9 p.c, and on Google’s GDPval-AA v2 check – which makes an attempt to measure real-world information work fairly than coding puzzles – 3.6 Flash scores 1421 in opposition to 1349 for the older mannequin.

Figma, Hebbia, and Harvey put the mannequin to work

Figma has built-in 3.6 Flash into its prototyping infrastructure, and in keeping with Matt Colyer, the corporate’s Director of Product Engineering, the mannequin provides builders a quicker route by means of design iterations with no drop in output high quality.

Legal expertise platform Harvey and analysis software Hebbia route knowledge by means of the mannequin for multimodal doc work: ingesting uncooked monetary filings, parsing doc construction, studying embedded charts, and producing draft reviews for assessment.

Google additionally folded a client-side computer-use software immediately into the Gemini API and Gemini Enterprise platforms, eradicating the customized middleman software program engineers beforehand constructed to let fashions function on prime of an working system.

The firm reviews an OSWorld-Verified rating of 83.0 p.c, up from 78.4 p.c, and says up to date safeguards in opposition to chemical, organic, radiological, and nuclear misuse enhance resistance to jailbreaking with out elevating refusal charges for benign requests.

A less expensive tier for high-volume background brokers

Gemini 3.5 Flash-Lite targets a distinct job: doc processing and agentic search working at quantity fairly than reasoning depth. The Artificial Analysis Index measured the mannequin at 350 output tokens per second, the quickest within the 3.5 collection in keeping with Google.

Pricing runs at $0.3/1M enter tokens and $2.5/1M output tokens, low cost sufficient that engineering groups can route easy, high-volume subagent requests to a minimal pondering stage and reserve greater pondering ranges for multi-step work.

On Google’s GDM-MRCR v2 long-context check, Gemini 3.5 Flash-Lite recorded a 72.2 p.c success fee in opposition to 60.1 p.c for its predecessor, and its GDPval-AA v2 rating almost doubled, from 642 to 1140. The mannequin carries the identical native computer-use software as 3.6 Flash.

Separately, Google says Gemini 3.5 Pro stays in accomplice testing forward of a full launch, and pre-training for the subsequent Gemini 4 structure is already underway.

Gemini 3.5 Flash Cyber: A restricted mannequin for patching code

Automated vulnerability scanners now floor flaws quicker than most safety groups can patch them, and that hole is the place Google positions Gemini 3.5 Flash Cyber.

The mannequin is constructed to validate and remediate code vulnerabilities, and Google reviews efficiency on the CyberGymnasium benchmark aggressive with frontier fashions (although it hasn’t made these figures public in the identical element as its consumer-facing releases.)

Distribution stays restricted to governments and vetted companions by means of a pilot programme, a limitation Google frames as a safeguard in opposition to the mannequin producing exploit code for offensive use.

Inside Google’s CodeMender safety agent, a number of situations of three.5 Flash Cyber run in parallel, cross-checking each other’s findings earlier than producing a single remediation report a human reviewer indicators off on.

Engineering groups in search of to combine these new fashions can entry them by means of the Gemini API through Google AI Studio, Android Studio, or the Gemini Enterprise Agent Platform. Consumers can even entry the brand new fashions within the Gemini app and three.5 Flash-Lite can be rolling out in Google Search.

See additionally: Bristol Myers Squibb buys Nvidia AI system for drug discovery

Banner for the AI & Big Data Expo event series.

Want to be taught extra about AI and massive knowledge from trade leaders? Check out AI & Big Data Expo going down in Amsterdam, California, and London. The complete occasion is a part of TechEx and is co-located with different main expertise occasions together with the Cyber Security & Cloud Expo. Click here for extra info.

AI News is powered by TechForge Media. Explore different upcoming enterprise expertise occasions and webinars here.

The put up Google’s Gemini 3.6 Flash targets enterprise agent token costs appeared first on AI News.

Similar Posts