Ling-2.6-flash is an instant (instruct) model from inclusionAI with 104B total parameters and 7.4B active parameters, designed for real-world agents that require fast responses, strong execution, and high token efficiency. It delivers performance comparable to state-of-the-art models at a similar scale while significantly reducing token usage across coding, document processing, and lightweight agent workflows.
Modalities
Price
Free
Context
262K
Released
Apr 21, 2026
Token volume and request traffic to this model over time.
Ling-2.6-flash is an instant (instruct) model from inclusionAI with 104B total parameters and 7.4B active parameters, designed for real-world agents that require fast responses, strong execution, and high token efficiency.
Yes. The pricing shown on this page for Ling-2.6-flash is zero, so you are not charged for prompt or completion tokens. Free endpoints are rate limited — see the rate limit docs.
Ling-2.6-flash has a 262,144 token context window.
Ling-3.0-flash, Ring-2.6-1T and Ling-2.6-1T are other text models from inclusionAI.
Ling-2.6-flash was released on April 21, 2026.