Sign InOpen Brain
VercelEngineering PostOfficial Source

Open-weight models take 56% of token volume, Astra doubles Fable 5.1 spend

Vercel’s gateway data shows open-weight models reached 56% of August token volume while price per token fell, reinforcing workload routing over loyalty to one frontier model.

Vercel · Sep 17, 2026
Open Source Open MarkdownOpen JSON
Source Summary

Open-weight models rose from **7% of gateway tokens in December to 56% in August**, while accounting for 14% of August spend. Average price per token fell **23.2% in August**, and the median high-volume team paid 7.6% less.

Practical Implication

Treat model choice as routing, not allegiance: run cheaper or open-weight models where evals permit, and reserve frontier tiers for tasks whose quality gains justify their price. The rapid moves between Fable, Opus, Gemini, and Astra support frequent reassessment.

Agent-Ready Context
Open-weight models rose from **7% of gateway tokens in December to 56% in August**, while accounting for 14% of August spend. Average price per token fell **23.2% in August**, and the median high-volume team paid 7.6% less.

Treat model choice as routing, not allegiance: run cheaper or open-weight models where evals permit, and reserve frontier tiers for tasks whose quality gains justify their price. The rapid moves between Fable, Opus, Gemini, and Astra support frequent reassessment.

This is anonymized traffic from Vercel AI Gateway, not the whole market. Spend uses public list prices, prior months may be revised, and workload shifts are inferred across teams rather than traced token by token.
Connected Context · Feed7 Judgment

This strengthens the earlier gateway trend from emerging open-weight adoption to majority token volume, while showing that volume share and spend share diverge sharply. It supports frequent, evaluation-gated routing rather than model allegiance, but remains evidence about one gateway’s workload mix—not model quality, causal savings for individual teams, or the broader market.

DeepSeek overtakes Google on volume, cost per token falls 13.6%The September index extends July’s evidence that routing toward open-weight models coincided with lower blended token cost, while preserving the same limitation that gateway traffic does not establish comparative quality.Open-weight models surge to 29% of volume, price per token flattensThe rise from the earlier reported 29% open-weight volume to 56% confirms that the shift continued, alongside the prior split between cheap volume and frontier-model workloads.CFOs and the new economics of AICursor’s ninefold cost variation and widespread multi-model use reinforce the index’s operational conclusion that economics depend on model mix and routing rather than loyalty to one provider.Notion's Token Town — Sarah Sachs, NotionToken Town explains why falling average token prices do not guarantee lower agent costs, qualifying the index’s price decline with the need to track workload design and completed-task economics.
Context Map
industry#open-models#model-selection#adoption
Uncertainty
This is anonymized traffic from Vercel AI Gateway, not the whole market. Spend uses public list prices, prior months may be revised, and workload shifts are inferred across teams rather than traced token by token.