Sign InOpen Brain
VercelEngineering PostOfficial Source

GLM 5.3 FlashX now available on AI Gateway

GLM 5.3 FlashX brings roughly 200 tokens-per-second serving to Vercel AI Gateway, offering a lower-wait option for coding-agent output and repeated tool loops.

Vercel · Sep 18, 2026
Open Source Open MarkdownOpen JSON
Source Summary

Vercel AI Gateway now serves **GLM 5.3 FlashX** under `zai/glm-5.3-flashx`, with reported throughput of **about 200 tokens per second** for the multimodal coding model.

Practical Implication

For latency-sensitive agents, test it on interactive output and tool-heavy loops. Gateway setup also provides usage and cost tracking, retries, failover, budgets, and routing through one API.

Agent-Ready Context
Vercel AI Gateway now serves **GLM 5.3 FlashX** under `zai/glm-5.3-flashx`, with reported throughput of **about 200 tokens per second** for the multimodal coding model.

For latency-sensitive agents, test it on interactive output and tool-heavy loops. Gateway setup also provides usage and cost tracking, retries, failover, budgets, and routing through one API.

The material reports serving speed, not task quality, time to first token, or end-to-end tool-loop latency. Provider pricing has no gateway markup, but actual workload cost is not given.
Connected Context · Feed7 Judgment

This adds a speed-oriented multimodal coding route to the Gateway’s growing model set, with roughly 200 tokens-per-second serving throughput as a concrete screening signal. It does not resolve the prior selection problem: task quality, first-token delay, end-to-end tool-loop latency, reliability, and workload cost still require direct evaluation before default routing.

Context Map
toolscoding#coding-agents#gateways#model-selection
Uncertainty
The material reports serving speed, not task quality, time to first token, or end-to-end tool-loop latency. Provider pricing has no gateway markup, but actual workload cost is not given.