GLM 5.3 FlashX now available on AI Gateway
GLM 5.3 FlashX, a high-speed serving option for Z.ai's multimodal coding model, is now available on AI Gateway, delivering inference at ~200 tokens per second. This enables faster streamed responses for coding agents, tool loops, and interactive applications.
is now available on AI Gateway.GLM 5.3 FlashX
GLM 5.3 FlashX is a high-speed serving option for Z.ai's multimodal coding model, delivering inference at ~200 tokens per second for faster streamed responses.
The higher serving speed is useful for coding agents, tool loops, and interactive applications where users wait on generated output.
Use across API formats and in coding agents:zai/glm-5.3-flashx
To use it in a coding agent, see the , then run to create a key and configure your supported agents. Select inside the agent.coding agents guidevercel ai-gateway setupzai/glm-5.3-flashx
AI Gateway provides a unified API for calling models, tracking usage and cost, and configuring retries, failover, and performance optimizations for higher-than-provider uptime. It includes built-in , , , and more.custom reportingbudgets for API keysrouting rules
AI Gateway reflects provider pricing with no markup and does not charge a platform fee on inference, including on (BYOK) requests.Bring Your Own Key
Try , or available on AI Gateway.GLM-5.3-FlashX in the model playgroundview all language models
Source: original entry ↗