Google has released Gemini 3.6 Flash and 3.5 Flash-Lite, two new models built specifically to lower token costs and latency for enterprise AI agents. The announcement targets a core problem for businesses running autonomous software agents in production: every extra token a model generates adds cost and delay to workflows that may run thousands of times per hour.
New Models Target Enterprise Agent Economics
The economics of running autonomous software agents in production come down to a simple equation. A model needs to reason through multi-step tasks competently, but every extra token it generates adds cost and delay. Google's answer splits the trade-off across three models: Gemini 3.6 Flash for coding and multimodal reasoning, Gemini 3.5 Flash-Lite for high-volume low-latency work, and a reserved model for specific use cases.
According to Artificial Intelligence News, pricing for Gemini 3.6 Flash sits at $1.50 per 1 million input tokens and $7.50 per 1 million output tokens. This positions the model for reasoning loops that run continuously rather than on-demand.
Token Efficiency and Latency Improvements
Gemini 3.6 Flash uses up to 17% fewer tokens and costs less per token, according to CNBC. The model is designed for teams building background agents rather than chat interfaces, where throughput matters more than parameter count.
Gemini 3.5 Flash-Lite targets faster, high-volume workloads, making it suitable for enterprise environments where speed and cost efficiency are critical.
Our Take: A Practical Move for Production AI
In our view, Google is addressing a real pain point for enterprises. Many companies have found that powerful AI models are too expensive to run at scale in production. By releasing models specifically optimized for token cost and latency, Google is making it more feasible for businesses to deploy autonomous agents that run continuously.
The pricing structure — $1.50 per million input tokens and $7.50 per million output tokens — is competitive for reasoning-heavy workflows. The 17% token reduction means enterprises can run more tasks for the same cost. This is a practical step toward making AI agents a standard part of enterprise operations rather than an expensive experiment.