megachangelog
Feature

Service tiers now available on AI Gateway

AI Gateway now supports service tiering to optimize for latency, throughput, and cost per request. Users can select from default, priority (faster but higher cost), or flex (lower cost but higher latency) tiers across OpenAI, Gemini, and other supported models, with billing automatically adjusted based on the tier each request actually uses.

AI Gateway now supports service tiering. Service tiers let you optimize for latency, throughput, and cost per request to match your use case. Pick a faster tier for interactive workloads (less queueing, higher token throughput), or a lower cost tier for background jobs that can tolerate more latency.

At launch, service tiering is available for OpenAI and Gemini models.

work across every AI Gateway API format: , , , , and . AI Gateway adjusts billing based on the tier each request used.Service tiersAI SDKChat Completions APIAnthropic Messages APIOpenAI Responses APIOpenResponses API

If a service tier is not specified, requests run on the default tier.

Set under . The same option works across all models and providers that support service tiers, so you can swap the model without restructuring provider-specific options:serviceTierproviderOptions.gateway

The applied tier is returned in the provider metadata so you can confirm which tier served the request. This is useful when a priority request falls back to default capacity, or when comparing observed latency across tiers.

Service tier is serviced on a best-effort basis: if a tier can't be applied, the request runs on the default tier at the default rate, and only an invalid service tier value fails the request.

If a model can be served by multiple providers and you only want a service tier on some of them or to configure different service tiers across each, set the tier on the provider's own namespace instead.

AI Gateway applies per-tier rates automatically based on the tier each request actually used. If a request gets downgraded to default capacity, billing reflects the default rate, not the priority rate.priority

More details about providers and tiering are in the .service tiers reference

Read more

Tiers

Basic usage

Pricing

  • : Standard processingdefault

  • : Faster processing at increased costpriority

  • : Lower cost with potentially higher latencyflex

Per-provider control

Tier

default

priority

flex

Cost relative to default

Baseline

~1.8-2x default pricing

~0.5x default pricing

ai-gatewayfeaturepricingperformanceopenaigemini

Source: original entry ↗