Kimi K3 and Kimi K3 Fast now available on AI Gateway with US providers and ZDR
Kimi K3 and Kimi K3 Fast models from Moonshot AI are now available on AI Gateway through US-based providers including Baseten and Fireworks, with support for Zero Data Retention (ZDR). Users can route requests to US data centers for compliance and leverage automatic failover across multiple providers for improved uptime and throughput.
and its faster serving path, , are now available from US-based providers on AI Gateway, including Baseten and Fireworks. (ZDR) is also supported for both models. Kimi K3 from Moonshot AIKimi K3 FastZero Data Retention
Running Kimi K3 on US-based providers lets teams with data residency and compliance requirements use the model on US infrastructure. Because AI Gateway serves the models from multiple providers, it automatically routes across them for failover, higher uptime, and more available throughput than any single provider offers. You call the same model ID, and the gateway handles provider selection and fallback.moonshotai/kimi-k3
Kimi K3 Fast trades a higher per-token cost for lower latency. Request it with the option on the base model, which stays on and falls back to standard speed when the fast tier is unavailable. Alternatively, use . The fast variant costs ~50% more than the base model. speedmoonshotai/kimi-k3moonshotai/kimi-k3-fast
To use Kimi K3, set to in the :modelmoonshotai/kimi-k3AI SDK
To route Kimi K3 requests to use only US data centers for inference, set . Regional pricing is ~10% more than the regular variant.inferenceRegion
Zero Data Retention for Kimi K3 is also available. Turn on Zero Data Retention for every request from the , or set it per request with :AI Gateway dashboard settingszeroDataRetention
To see every provider serving Kimi K3, along with per-provider pricing, supported parameters, uptime, throughput, and latency, call the model endpoints API:
Model prices vary by provider and variant type.
Run and select Kimi K3. This will detect the agents on your machine, provision an AI Gateway key, and write their config. See how to set it up via the .vercel ai-gateway coding-agents setupVercel CLI
Try Kimi K3 in the .model playground
US inference
Zero Data Retention
Providers and endpoints
Use Kimi K3 in your coding agent
Source: original entry ↗
More from Vercel
Follow Vercel to get its new changes in your feed and email digest.
OpenAI Decisions API now available on AI Gateway
OpenAI's Decisions API is now accessible through Vercel's AI Gateway with an OpenAI-compatible endpoint, enabling decision models to answer typed questions and return probabilities, choices, and scores for routing, triage, and guardrails use cases. Support is available across the OpenAI SDK, AI SDK, HTTP API, and CLI with the latest versions.
Timestamp attributes now supported in Vercel Flags
Vercel Flags now supports timestamp attributes for entities, allowing you to create time-based targeting rules. Use this feature to run limited-time campaigns, show content between specific dates, or target users based on registration date.
Glyph Cluster now available in stealth on AI Gateway
Glyph Cluster, a reasoning model for coding and long-context analysis, is now available as a stealth model on Vercel's AI Gateway for Pro and Enterprise plan teams with purchased AI Gateway credits at no cost during the stealth period. The model supports function calling, streams responses, and can be accessed via AI SDK, OpenAI-compatible APIs, and coding agents.