AI Gateway adds unified fast mode support
AI Gateway now offers a unified fast mode abstraction that allows you to request lower latency or higher throughput across multiple models by setting speed to 'fast'. This feature automatically falls back to standard speed when fast tier isn't available, and is now in beta.
AI Gateway has a unified fast mode abstraction, now in beta.
You can now request the same way for every model on AI Gateway. Set speed to , and the gateway serves the fast tier when it's available and falls back to standard speed when it isn't.fast modefast
Fast mode trades a higher per-token cost for lower latency or higher throughput. You get it wherever it is available, without pinning a provider or wiring anything up yourself.
Set to under to upgrade the model to its fast serving path when one is routable:speedfastproviderOptions.gateway
You can also address a fast variant directly with its fast slug, such as , which is the same as setting on the base model.anthropic/claude-opus-5-fastspeed: 'fast'
Use the option when you want to stay on the base model ID and fall back to standard speed. Use a fast slug when you want to name the fast variant directly, for example in a fallback list. The option works across every AI Gateway API format.speedgateway.modelsspeed
Fast mode is available for a growing set of models. Find them in the with a lightning bolt next to the slug. Requesting fast mode on a model with no fast tier has no effect, and the request runs at standard speed. Fast variants of models are usually more expensive than the base model: check for more details.models listpricing
In Claude Code, toggle fast mode with for Anthropic Opus models after . For other agents, select a fast slug in your model configuration, such as ./fastanthropic/claude-opus-5-fastsetup
For more information about fast mode, read the .documentation
Requesting fast mode
Using fast mode in coding agents
Source: original entry ↗
More from Vercel
Follow Vercel to get its new changes in your feed and email digest.
OpenAI Decisions API now available on AI Gateway
OpenAI's Decisions API is now accessible through Vercel's AI Gateway with an OpenAI-compatible endpoint, enabling decision models to answer typed questions and return probabilities, choices, and scores for routing, triage, and guardrails use cases. Support is available across the OpenAI SDK, AI SDK, HTTP API, and CLI with the latest versions.
Timestamp attributes now supported in Vercel Flags
Vercel Flags now supports timestamp attributes for entities, allowing you to create time-based targeting rules. Use this feature to run limited-time campaigns, show content between specific dates, or target users based on registration date.
Glyph Cluster now available in stealth on AI Gateway
Glyph Cluster, a reasoning model for coding and long-context analysis, is now available as a stealth model on Vercel's AI Gateway for Pro and Enterprise plan teams with purchased AI Gateway credits at no cost during the stealth period. The model supports function calling, streams responses, and can be accessed via AI SDK, OpenAI-compatible APIs, and coding agents.