Gemini 3.8 Flash now available on AI Gateway
Gemini 3.8 Flash brings multimodal input, tool calling, web search, and default reasoning to coding agents through Vercel. Its temporary 50% discount runs through December 31.
Google’s **Gemini 3.8 Flash** is now on Vercel AI Gateway with a **1M-token context window** and a 65,536-token output limit. It accepts text, images, PDFs, and video, and supports tool calls and web search.
Builders can select google/gemini-3.8-flash in connected coding agents. Thinking is enabled by default, so test its latency, token use, and behavior on representative repository tasks before changing a production routing default.
Google’s **Gemini 3.8 Flash** is now on Vercel AI Gateway with a **1M-token context window** and a 65,536-token output limit. It accepts text, images, PDFs, and video, and supports tool calls and web search. Builders can select google/gemini-3.8-flash in connected coding agents. Thinking is enabled by default, so test its latency, token use, and behavior on representative repository tasks before changing a production routing default. Vercel says it improves software engineering, agent work, and multi-step reasoning at the prior Flash model’s speed and cost, but supplies no evaluation results here. The advertised **50% discount ends December 31**.
This adds another million-token, multimodal coding-agent route but does not narrow the model decision through evidence. Its default thinking behavior introduces a concrete evaluation variable—latency, token consumption, and repository-task behavior—while the temporary discount can distort cost comparisons. Capacity, modality support, and provider claims remain inputs to a matched workload trial rather than grounds for changing production defaults.