AI Gateway is now generally available
AI Gateway is now generally available, providing a unified API to access hundreds of AI models with transparent pricing and built-in observability. It delivers sub-20ms latency routing, automatic failover, and detailed analytics.
is now generally available, providing a single unified API to access hundreds of AI models with transparent pricing and built-in observability.AI Gateway
With sub-20ms latency routing across multiple inference providers, AI Gateway delivers:
You can use AI Gateway with the or through the OpenAI-compatible endpoint. With the AI SDK, it’s just a simple model string switch.AI SDK
Get started with a single API call:
Read more about the , learn more about , or .announcementAI Gatewayget started now
Transparent pricing with no markup on tokens (including Bring Your Own Keys)
Automatic failover for higher availability
High rate limits
Detailed cost and usage analytics
Source: original entry ↗
More from Vercel
Follow Vercel to get its new changes in your feed and email digest.
OpenAI Decisions API now available on AI Gateway
OpenAI's Decisions API is now accessible through Vercel's AI Gateway with an OpenAI-compatible endpoint, enabling decision models to answer typed questions and return probabilities, choices, and scores for routing, triage, and guardrails use cases. Support is available across the OpenAI SDK, AI SDK, HTTP API, and CLI with the latest versions.
Timestamp attributes now supported in Vercel Flags
Vercel Flags now supports timestamp attributes for entities, allowing you to create time-based targeting rules. Use this feature to run limited-time campaigns, show content between specific dates, or target users based on registration date.
Glyph Cluster now available in stealth on AI Gateway
Glyph Cluster, a reasoning model for coding and long-context analysis, is now available as a stealth model on Vercel's AI Gateway for Pro and Enterprise plan teams with purchased AI Gateway credits at no cost during the stealth period. The model supports function calling, streams responses, and can be accessed via AI SDK, OpenAI-compatible APIs, and coding agents.