Jul 22, 2026
ManyPress

Advertisement

Artificial Intelligence

Google has launched new AI models designed to reduce latency and token costs for enterprise software agents, alongside a specialized security variant.

ManyPress

ManyPress

ManyPress Editorial

2 min readSource:Artificial Intelligence News
Google Releases Gemini 3.6 Flash and 3.5 Flash-Lite for Enterprise AI Agents

Key facts

  • Gemini 3.6 Flash showed a 14 percent improvement on the GDPval-AA v2 test compared to the previous model.
  • Figma has integrated Gemini 3.6 Flash into its prototyping infrastructure to accelerate design iterations.
  • Gemini 3.5 Flash-Lite recorded a 72.2 percent success rate on the GDM-MRCR v2 long-context test.
  • Google reports that Gemini 3.5 Pro remains in partner testing, with pre-training for the Gemini 4 architecture underway.
  • The new models are accessible via the Gemini API through Google AI Studio, Android Studio, and the Gemini Enterprise Agent Platform.

Google has introduced Gemini 3.6 Flash and Gemini 3.5 Flash-Lite, two new models engineered to lower latency and token costs for enterprise AI agents. The release aims to support high-volume, autonomous workflows by balancing reasoning capabilities with operational efficiency. Additionally, Google unveiled a restricted Gemini 3.5 Flash Cyber variant focused on vulnerability remediation, while integrating new computer-use tools directly into its API and enterprise platforms.

By the numbers

$1.50/1M
Gemini 3.6 Flash input token price
$7.50/1M
Gemini 3.6 Flash output token price
$0.3/1M
Gemini 3.5 Flash-Lite input token price
$2.5/1M
Gemini 3.5 Flash-Lite output token price
350
Gemini 3.5 Flash-Lite output tokens per second

Performance and Pricing Metrics

Gemini 3.6 Flash features 17 percent fewer output tokens than the 3.5 Flash version, with pricing set at $1.50 per million input tokens and $7.50 per million output tokens. In benchmarks like DeepSWE, the model achieved a 49 percent success rate compared to 37 percent for its predecessor. Gemini 3.5 Flash-Lite is optimized for high-volume tasks, offering speeds of 350 output tokens per second at a cost of $0.30 per million input tokens and $2.50 per million output tokens.

Specialized Security and Integration

The Gemini 3.5 Flash Cyber variant is currently restricted to governments and vetted partners to prevent the generation of offensive exploit code. It is designed to validate and remediate code vulnerabilities, utilizing parallel instances to cross-check findings. Meanwhile, Google has integrated a client-side computer-use tool directly into its API, allowing models to operate on operating systems without custom intermediary software.

Advertisement

This article was independently rewritten by ManyPress editorial AI from reporting originally published by Artificial Intelligence News.

Artificial Intelligence