GLM 5.3 FlashX now available on AI Gateway
GLM 5.3 FlashX brings roughly 200 tokens-per-second serving to Vercel AI Gateway, offering a lower-wait option for coding-agent output and repeated tool loops.
Vercel AI Gateway now serves **GLM 5.3 FlashX** under `zai/glm-5.3-flashx`, with reported throughput of **about 200 tokens per second** for the multimodal coding model.
For latency-sensitive agents, test it on interactive output and tool-heavy loops. Gateway setup also provides usage and cost tracking, retries, failover, budgets, and routing through one API.
Vercel AI Gateway now serves **GLM 5.3 FlashX** under `zai/glm-5.3-flashx`, with reported throughput of **about 200 tokens per second** for the multimodal coding model. For latency-sensitive agents, test it on interactive output and tool-heavy loops. Gateway setup also provides usage and cost tracking, retries, failover, budgets, and routing through one API. The material reports serving speed, not task quality, time to first token, or end-to-end tool-loop latency. Provider pricing has no gateway markup, but actual workload cost is not given.
This adds a speed-oriented multimodal coding route to the Gateway’s growing model set, with roughly 200 tokens-per-second serving throughput as a concrete screening signal. It does not resolve the prior selection problem: task quality, first-token delay, end-to-end tool-loop latency, reliability, and workload cost still require direct evaluation before default routing.