Sign InOpen Brain
Atlas / Infra

Gateways

Open JSONConfidence: Auto-collectedLast updated 2026-08-02

Current Answer

No editorial synthesis yet — the evidence below is collected automatically from source labels. A current answer lands here once an editor approves one.

Evidence

Qwen 3.8 Max now available on Vercel AI Gateway
Vercel · 2026-08-02

Vercel AI Gateway now exposes Qwen 3.8 Max to coding agents, adding one model endpoint for long-context text and vision work with gateway routing, budgets, and usage tracking.

AI Gateway now supports team and project spend budgets
Vercel · 2026-07-31

AI Gateway can now enforce spend caps across a team, project, or API key, giving agent workloads layered cost controls instead of relying on per-key limits alone.

10x more capacity for Laguna S 2.1 on AI Gateway
Vercel · 2026-07-31

AI Gateway has raised Laguna S 2.1 capacity tenfold for both paid and free model IDs, reducing throughput constraints for high-volume or long-running coding agents.

AI Gateway logs now have a dedicated page
Vercel · 2026-07-31

AI Gateway’s dedicated logs expose per-request cost, tokens, latency, routing, and provider fallbacks, making agent failures and spend anomalies easier to trace.

MiniMax H3 now available on AI Gateway
Vercel · 2026-07-30

MiniMax H3 brings short 2K video generation to Vercel AI Gateway, with text, keyframe, and multimodal reference inputs. Reference and keyframe modes cannot be combined.

AI Gateway: GPT-5.6 pricing and speed updates
Vercel · 2026-07-30

Vercel cut Luna and Terra token prices and raised Sol fast-mode speed without changing model IDs, so existing agent workloads inherit the changes without code edits.

AI Gateway adds unified fast mode support
Vercel · 2026-07-29

AI Gateway now exposes one beta fast-mode option across models, letting coding agents request lower latency while retaining standard-speed fallback when no fast tier exists.

Regional inference now available on AI Gateway
Vercel · 2026-07-27

AI Gateway can pin inference and retained provider data to the US or EU, giving agent builders one residency control with per-response verification across supported providers.

WebSocket support for OpenAI Responses API live on AI Gateway
Vercel · 2026-07-27

Vercel’s gateway now supports persistent WebSocket sessions for the Responses API, reducing repeated context transfer during long, tool-heavy agent runs.

Kimi K3 and Kimi K3 Fast with ZDR and US-based providers now on AI Gateway
Vercel · 2026-07-27

Kimi K3 now has US-hosted, ZDR-capable gateway routes plus a faster tier, giving coding-agent users explicit latency, residency, retention, and cost choices.

TRACE-ROUTER: Task-Consistent and Adaptive Online Routing for Agentic AI
arXiv · 2026-07-24

TRACE-Router selects one model per agent task, keeps every call on that backend, and learns from the final outcome. Its benchmarks suggest task-level routing can improve accuracy and latency together.

Claude Opus 5 now available on AI Gateway
Vercel · 2026-07-24

AI Gateway now serves Claude Opus 5 with configurable reasoning, fast mode, fallbacks, and coding-agent setup; benign security tasks may still hit safeguards.

Notion's Token Town — Sarah Sachs, Notion
AI Engineer · 2026-07-23

Agent economics can regress even when token prices look stable. Route by task, preserve model optionality, and move deterministic work out of LLM calls before scaling usage.

AI Gateway now supports streaming transcription
Vercel · 2026-07-22

AI Gateway can now stream audio into transcription models and emit partial text, letting text-based agents accept lower-latency voice input without changing the agent itself.

How Searchable ships customer-requested features in 30 minutes on Vercel
Vercel · 2026-07-21

Searchable says Vercel’s AI SDK and Gateway let it swap models without SDK or key changes, helping the team ship some customer requests in 30 minutes and move 2–5× faster.

Claude Fable 5 access restored on AI Gateway

Claude Fable 5 is back on Vercel's AI Gateway after US export controls lifted. New, stricter safety classifiers can refuse routine coding calls, so configure model fallbacks; no Zero Data Retention (30-day hold).

diegosouzapw/OmniRoute
GitHub

OmniRoute puts many model providers behind one OpenAI-compatible endpoint, with routing, quota failover, cost telemetry, compression, and coding-agent setup. Its breadth raises operational and trust questions.

Access and share AI Gateway leaderboard data

Vercel’s AI Gateway rankings are now queryable as CSV or JSON, giving agent builders production-usage signals for model and provider selection beyond benchmark scores.

Routing rules now available on AI Gateway

Vercel AI Gateway adds firewall-style routing rules: rewrite one model to another or deny a model outright, applied at the gateway so you swap models across your whole team without shipping a code change.

Claws Out: Securing and Building with OpenClaw - Nick Taylor, Pomerium

OpenClaw’s trusted-proxy mode removes duplicate WebSocket tokens and device pairing, but only if proxy IPs and identity headers are tightly constrained.

Grok 4.5 now available on AI Gateway

Grok 4.5 is available through Vercel AI Gateway with text and image input plus low, medium, and high reasoning settings for tuning speed against depth.

Seedream 5.0 Pro is now available on AI Gateway

Seedream 5.0 Pro adds image generation and editing to Vercel AI Gateway, targeting reliable text rendering and dense infographic layouts through the AI SDK.

chenyme/grok2api

Grok2API fronts Grok Build, Web, and Console account pools with OpenAI- and Anthropic-compatible APIs, but its unofficial SSO routing creates terms, credential, and renewal risk.

Stable permalink · evidence auto-collected from source labels · synthesis maintained by feed7 editorial