Ling 3.0 Flash now available on AI Gateway
Ant Group's Ling 3.0 Flash model is now available on Vercel's AI Gateway, offering token-efficient agentic inference with a 256K context window and free access through August 3rd. The model is optimized for high-frequency agentic workflows, coding agents, and long-context interactions without any platform fee or markup on inference.
from Ant Group is now available on AI Gateway.Ling 3.0 Flash
The model is free to use for the next three weeks, through August 3rd.
Ling 3.0 Flash is a Mixture-of-Experts model with 124B total parameters and about 5.1B active per token. It has a 256K token context window and runs in thinking and non-thinking modes.
Ling 3.0 Flash is built for token-efficient agentic inference at production scale, doing more work within tighter token, latency, and cost budgets across multi-step agent runs. The model targets high-frequency agentic workflows, coding agents, document work, and long-context multi-turn interactions.
To use Ling 3.0 Flash, set model to in the :inclusionai/ling-3.0-flash-freeAI SDK
AI Gateway provides a unified API for calling models, tracking usage and cost, and configuring retries, failover, and performance optimizations for higher-than-provider uptime. It includes built-in , , , , and more.custom reportingZero Data Retention supportbudgets for API keysrouting rules
AI Gateway reflects provider pricing with no markup and does not charge a platform fee on inference, including on (BYOK) requests.Bring Your Own Key
Try Ling 3.0 Flash in the .model playground
Source: original entry ↗