Live model performance metrics accessible via AI Gateway
AI Gateway now displays live throughput and latency metrics across hundreds of models, updated hourly, to help you choose the best model and provider based on real performance data. Metrics are available in the model list, individual model pages, and programmatically via REST API.
AI Gateway now displays throughput and latency metrics across hundreds of models, helping you choose the right model based on live performance data.
Metrics appear in three places and are updated every hour:
The AI Gateway now includes sortable columns for latency and throughput. Each row displays the best P50 metrics (lowest latency, highest throughput) for that model across all its available providers. Metrics are updated every hour and based on live AI Gateway customer requests.model list
Sort by throughput to find the fastest token generation, or by latency to find models with the quickest time-to-first-token.
On the individual model pages, you can see P50 latency and throughput for each provider that has recorded usage. This helps you compare provider performance for the same model and choose the best option for your use case.
To access these pages, click on any model in the to get a more detailed view of the breakdown across all the providers that carry the model in AI Gateway. Metrics are refreshed hourly and only appear for providers with sufficient traffic.model list
Here is an example for :openai/gpt-oss-120b
Similar to the overall model list, you can sort by latency and throughput across providers on the model detail pages.
These metrics are also available programmatically via the endpoints REST API. To use this, replace with the for the model of interest.[ai-gateway-string]creator/model-name
This returns live hourly P50 and P95 latency (ms TTFT) and throughput (T/s) for the specified model, by provider. Here is an example output from the endpoint for the Cerebras provider for .zai/glm-4.7
If you want to query the full list of models, you can also use the model metrics endpoint in conjunction with .https://ai-gateway.vercel.sh/v1/models
: Best performance per model (P50 latency and throughput)Model list
: Provider-level performance breakdownModel detail pages
: Rolling endpoint performance aggregates (latency and throughput, P50/P95)REST API
Model list
Model detail pages
REST API
Source: original entry ↗
More from Vercel
Follow Vercel to get its new changes in your feed and email digest.
OpenAI Decisions API now available on AI Gateway
OpenAI's Decisions API is now accessible through Vercel's AI Gateway with an OpenAI-compatible endpoint, enabling decision models to answer typed questions and return probabilities, choices, and scores for routing, triage, and guardrails use cases. Support is available across the OpenAI SDK, AI SDK, HTTP API, and CLI with the latest versions.
Timestamp attributes now supported in Vercel Flags
Vercel Flags now supports timestamp attributes for entities, allowing you to create time-based targeting rules. Use this feature to run limited-time campaigns, show content between specific dates, or target users based on registration date.
Glyph Cluster now available in stealth on AI Gateway
Glyph Cluster, a reasoning model for coding and long-context analysis, is now available as a stealth model on Vercel's AI Gateway for Pro and Enterprise plan teams with purchased AI Gateway credits at no cost during the stealth period. The model supports function calling, streams responses, and can be accessed via AI SDK, OpenAI-compatible APIs, and coding agents.