
Kimi-VL is a lightweight Mixture-of-Experts vision-language model that activates only 2.8B parameters per step while delivering strong performance on multimodal reasoning and long-context tasks. The Kimi-VL-A3B-Thinking variant, fine-tuned with chain-of-thought and reinforcement learning, excels in math and visual reasoning benchmarks like MathVision, MMMU, and MathVista, rivaling much larger models such as Qwen2.5-VL-7B and Gemma-3-12B. It supports 128K context and high-resolution input via its MoonViT encoder.
Modalities
Context
131K
Released
Apr 10, 2025
Knowledge Cutoff
Dec 2024
Token volume and request traffic to this model over time.
Kimi-VL is a lightweight Mixture-of-Experts vision-language model that activates only 2.8B parameters per step while delivering strong performance on multimodal reasoning and long-context tasks. The Kimi-VL-A3B-Thinking variant, fine-tuned with chain-of-thought and reinforcement learning, excels in math and visual reasoning benchmarks like MathVision, MMMU, and MathVista, rivaling much larger...
Kimi VL A3B Thinking has a 131,072 token context window.
Kimi VL A3B Thinking accepts images and text as input and returns text.
Kimi K3, Kimi K2.7 Code, Kimi K2.6 and 4 more are other text models from MoonshotAI.
Kimi VL A3B Thinking was released on April 10, 2025. Its knowledge cutoff is December 31, 2024.