AI Gateway: GPT-5.6 pricing and speed updates
GPT-5.6 Luna and Terra models are now cheaper with reduced token pricing, and GPT-5.6 Sol's fast mode is now 2.5x faster. AI Gateway passes upstream rate changes directly to users with no code changes required for existing requests.
On , and are now cheaper and is faster.AI GatewayGPT-5.6 LunaGPT-5.6 TerraGPT-5.6 Sol
AI Gateway adds no markup on token pricing, so these changes reach you at the upstream rate.
The changes apply to both short and long context pricing.
GPT-5.6 Sol keeps the same price; its fast mode now runs 2.5x faster, up from 1.5x. Model IDs are unchanged, so existing requests get the new rates and speed with no code change.
See full pricing detail on .AI Gateway
Model | Change | Input: Short context (per 1M tokens) | Output: Short context (per 1M tokens) |
|---|---|---|---|
| 80% price reduction | $0.2 | $1.2 |
| 20% price reduction | $2 | $12 |
| Same price, fast mode now 2.5x faster (up from 1.5x) | Unchanged | Unchanged |
Source: original entry ↗
More from Vercel
Follow Vercel to get its new changes in your feed and email digest.
OpenAI Decisions API now available on AI Gateway
OpenAI's Decisions API is now accessible through Vercel's AI Gateway with an OpenAI-compatible endpoint, enabling decision models to answer typed questions and return probabilities, choices, and scores for routing, triage, and guardrails use cases. Support is available across the OpenAI SDK, AI SDK, HTTP API, and CLI with the latest versions.
Timestamp attributes now supported in Vercel Flags
Vercel Flags now supports timestamp attributes for entities, allowing you to create time-based targeting rules. Use this feature to run limited-time campaigns, show content between specific dates, or target users based on registration date.
Glyph Cluster now available in stealth on AI Gateway
Glyph Cluster, a reasoning model for coding and long-context analysis, is now available as a stealth model on Vercel's AI Gateway for Pro and Enterprise plan teams with purchased AI Gateway credits at no cost during the stealth period. The model supports function calling, streams responses, and can be accessed via AI SDK, OpenAI-compatible APIs, and coding agents.