OpenRouter for AI crawlers
OpenRouter is a unified API for every major LLM: one endpoint, hundreds of models, with routing, fallbacks, and cost tracking across providers. This page is a plain, text-only index of the site for AI crawlers, agents, and low-bandwidth access. It lists the top 100 models by weekly usage for every output modality, every provider on the network, and every model collection, each with a short description and a canonical link.
For a shorter machine-readable overview see /llms.txt. For the complete list of URLs see /sitemap.xml. Models with a guide link to it at /{author}/{model}/llms.txt.
Site
- Models Compare every model on OpenRouter by pricing, context length, and benchmarks, all through one API.
- Rankings LLM leaderboards by real-world usage, ranked by tokens processed through the OpenRouter API, with per-modality views.
- Pricing How OpenRouter pricing works, with pay-as-you-go and enterprise plans, plus answers to common billing questions.
- Providers The network of model providers available on OpenRouter, with a page per provider.
- Compare Compare models by benchmarks, price, context length, latency, uptime, and throughput.
- Collections Curated collections of models, including free models and image generation models.
- Apps App and agent rankings, listing the apps built on OpenRouter and the top apps by category.
- State of AI An empirical study of real-world LLM usage across more than 100 trillion tokens on OpenRouter.
- Data OpenRouter's empirical AI usage data for researchers, journalists, and academics.
- Docs Developer documentation for the OpenRouter API, SDKs, and features.
- Support Frequently asked questions about using OpenRouter, and how to reach support.
Top Image models
Ranked by tokens processed on OpenRouter over the past week. Free variants and OpenRouter routers are listed as their own entries.
- Auto Router by OpenRouter. The Auto Router automatically selects the best model for your prompt, powered by the wisdom of the market. It routes you based on what the OpenRouter community collectively spends on for exactly the kind of task your prompt represents. It looks at spend over a trailing 7-day window, so it stays up to date with new model releases automatically. Your response will be priced at the same rate as the routed model. The cost_tier setting (low, medium, high, xhigh, or max) selects how much you're willing to pay. The default is low, giving you a highly cost-efficient set of models. When selecting models, the auto router respects any guardrails, privacy/ZDR policies, and model restrictions you have in the request or account settings. Multi-turn conversations stick to one model until it is no longer a leading choice for your task. Learn more in our [docs](/docs/guides/routing/routers/auto-router). For another way to route, see [Pareto Code](/openrouter/pareto-code). llms.txt
- Auto Router (Beta) by OpenRouter. The experimental version of our Auto Router where we test new improvements. Use it to get the latest and greatest version of our general purpose auto router, but expect beta quality. When new improvements are proven in auto-beta, we graduate them to the auto router. We've love to hear your feedback on whether auto-beta is working for you, or requests for routing improvements. To see which model was used, visit [Activity](/activity), or read the `model` attribute of the response. Your response will be priced at the same rate as the routed model. Learn more, including how to customize the models for routing, in our [docs](/docs/guides/routing/routers/auto-router). llms.txt
- Google: Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image) by Google. Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image) is Google's fastest, most cost-efficient Gemini image model, built for high-velocity developer pipelines and rapid-fire visual exploration. It delivers text-to-image generation in roughly 4 seconds — about 2.7× faster than Gemini 3.1 Flash Image — while keeping the character consistency, precise editing, and real-world knowledge of the Nano Banana family. A single drop-in API handles text-to-image, image editing, and multi-image composition. As a multimodal model it also returns text alongside images. Outputs are generated at 1K resolution across 14 aspect ratios and carry an invisible SynthID watermark so they can be identified as AI-generated. Positioned as the best balance of quality and speed in the Nano Banana 2 line, it lets you generate thousands of images at a fraction of the cost of heavier production models — ideal for prototyping, real-time apps, and visual workflows at scale. llms.txt
- Google: Nano Banana 2 (Gemini 3.1 Flash Image) by Google. Gemini 3.1 Flash Image, a.k.a. "Nano Banana 2," is Google’s latest state of the art image generation and editing model, delivering Pro-level visual quality at Flash speed. It combines advanced contextual understanding with fast, cost-efficient inference, making complex image generation and iterative edits significantly more accessible. Aspect ratios can be controlled with the [image_config API Parameter](https://openrouter.ai/docs/features/multimodal/image-generation#image-aspect-ratio-configuration) llms.txt
- ByteDance Seed: Seedream 4.5 by bytedance-seed. Seedream 4.5 is the latest in-house image generation model developed by ByteDance. Compared with Seedream 4.0, it delivers comprehensive improvements, especially in editing consistency, including better preservation of subject details, lighting, and color tone. It also enhances portrait refinement and small-text rendering. The model’s multi-image composition capabilities have been significantly strengthened, and both reasoning performance and visual aesthetics continue to advance, enabling more accurate and artistically expressive image generation. Pricing is $0.04 per output image, regardless of size. llms.txt
- OpenAI: GPT Image 2.5 Sunburst by OpenAI. GPT Image 2.5 Sunburst is an image generation and editing model from OpenAI, positioned as the precision-oriented tier of the GPT Image 2.5 series. It is suited to detailed creative work where editing accuracy matters more than generation speed, via the dedicated Images API. llms.txt
- Google: Nano Banana (Gemini 2.5 Flash Image) by Google. Gemini 2.5 Flash Image, a.k.a. "Nano Banana," is now generally available. It is a state of the art image generation model with contextual understanding. It is capable of image generation, edits, and multi-turn conversations. Aspect ratios can be controlled with the [image_config API Parameter](https://openrouter.ai/docs/features/multimodal/image-generation#image-aspect-ratio-configuration) llms.txt
- OpenAI: GPT Image 2 by OpenAI. OpenAI's latest image generation model. Supports high-fidelity image generation and editing via the dedicated Images API. llms.txt
- OpenAI: GPT-5.4 Image 2 by OpenAI. [GPT-5.4](https://openrouter.ai/openai/gpt-5.4) Image 2 combines OpenAI's GPT-5.4 model with state-of-the-art image generation capabilities from GPT Image 2. It enables rich multimodal workflows, allowing users to seamlessly move between reasoning, coding, and visual generation within the same interaction. llms.txt
- Google: Nano Banana 2 (Gemini 3.1 Flash Image Preview) by Google. Gemini 3.1 Flash Image Preview, a.k.a. "Nano Banana 2," is Google’s latest state of the art image generation and editing model, delivering Pro-level visual quality at Flash speed. It combines advanced contextual understanding with fast, cost-efficient inference, making complex image generation and iterative edits significantly more accessible. Aspect ratios can be controlled with the [image_config API Parameter](https://openrouter.ai/docs/features/multimodal/image-generation#image-aspect-ratio-configuration) llms.txt
- ByteDance Seed: Seedream 5.0 Lite by bytedance-seed. Seedream 5.0 Lite is an image generation model from ByteDance Seed. It is suited for professional visual creation that benefits from web-connected retrieval, complex-prompt comprehension, visual references, and broad knowledge coverage. llms.txt
- Google: Nano Banana Pro (Gemini 3 Pro Image) by Google. Nano Banana Pro is Google’s most advanced image-generation and editing model, built on Gemini 3 Pro. It extends the original Nano Banana with significantly improved multimodal reasoning, real-world grounding, and high-fidelity visual synthesis. The model generates context-rich graphics, from infographics and diagrams to cinematic composites, and can incorporate real-time information via Search grounding. It offers industry-leading text rendering in images (including long passages and multilingual layouts), consistent multi-image blending, and accurate identity preservation across up to five subjects. Nano Banana Pro adds fine-grained creative controls such as localized edits, lighting and focus adjustments, camera transformations, and support for 2K/4K outputs and flexible aspect ratios. It is designed for professional-grade design, product visualization, storyboarding, and complex multi-element compositions while remaining efficient for general image creation workflows. llms.txt
- Meta: Muse Image by Meta. Muse Image is an agentic image generation model from Meta that generates and edits images from text and reference images. Unlike single-pass image models, it reasons before it renders, breaking down multi-part prompts and refining its output within the chain of thought, and invokes web search for factual accuracy on knowledge-intensive prompts. The model supports text-to-image generation, targeted image editing, multi-image composition, reference-image conditioning for style and subject consistency across a series, and precise text rendering within generated images. Iterative editing works by passing the previous output image back with a new instruction. llms.txt
- OpenAI: GPT Image 2.5 Flare by OpenAI. GPT Image 2.5 Flare is an image generation and editing model from OpenAI, positioned as the speed-oriented tier of the GPT Image 2.5 series. It is suited to high-volume everyday generation, creator content, and rapid prototyping via the dedicated Images API. llms.txt
- Google: Nano Banana Pro (Gemini 3 Pro Image Preview) by Google. Nano Banana Pro is Google’s most advanced image-generation and editing model, built on Gemini 3 Pro. It extends the original Nano Banana with significantly improved multimodal reasoning, real-world grounding, and high-fidelity visual synthesis. The model generates context-rich graphics, from infographics and diagrams to cinematic composites, and can incorporate real-time information via Search grounding. It offers industry-leading text rendering in images (including long passages and multilingual layouts), consistent multi-image blending, and accurate identity preservation across up to five subjects. Nano Banana Pro adds fine-grained creative controls such as localized edits, lighting and focus adjustments, camera transformations, and support for 2K/4K outputs and flexible aspect ratios. It is designed for professional-grade design, product visualization, storyboarding, and complex multi-element compositions while remaining efficient for general image creation workflows. llms.txt
- ByteDance Seed: Seedream 5.0 Pro by bytedance-seed. Seedream 5.0 Pro is an image generation and editing model from ByteDance Seed. It is suited for commercial visual-production workflows that require precise editing control, lifelike scenes, and natural rendering. llms.txt
- Black Forest Labs: FLUX.2 Pro by black-forest-labs. A high-end image generation and editing model focused on frontier-level visual quality and reliability. It delivers strong prompt adherence, stable lighting, sharp textures, and consistent character/style reproduction across multi-reference inputs. Designed for production workloads, it balances speed and quality while supporting text-to-image and image editing up to 4 MP resolution. Pricing is as follows, [per the docs](https://bfl.ai/pricing?category=flux.2): Input: We charge $0.015 for each megapixel on the input (i.e. reference images for editing) Output: The first megapixel is charged $0.03 and then each subsequent MP will be charged $0.015. llms.txt
- xAI: Grok Imagine Image 2.0 by SpaceXAI. Grok Imagine Image 2.0 is an image generation and editing model from xAI. It is suited for creating images from text prompts and editing images from references, with low and medium quality modes. llms.txt
- Black Forest Labs: FLUX.2 Klein 4B by black-forest-labs. FLUX.2 [klein] 4B is the fastest and most cost-effective model in the FLUX.2 family, optimized for high-throughput use cases while maintaining excellent image quality. Pricing is based on the output image. The first generated megapixel is charged $0.014. Each subsequent megapixel is charged $0.001. llms.txt
- SpaceXAI: Grok Imagine Image Quality by SpaceXAI. Grok Imagine Image Quality is SpaceXAI's fast, high-fidelity image generation and editing model. It accepts text prompts and optional reference images, producing photorealistic outputs at 1K or 2K across a range of aspect ratios, including flexible adjustment of reference images. The model emphasizes realistic detail — natural lighting and physics, accurate textures, and consistent rendering of named entities such as brands, public figures, and specific locations. It supports clean multilingual text rendering inside images, making it the top choice for posters, packaging, ads, menus, and social graphics. When given reference images, it preserves identity and structure for product placement, brand-aligned variations, and character continuity across scenes. llms.txt
- inclusionAI: Ming Image 0.1 Design by inclusionai. Ming Image 0.1 Design is a text-to-image model from inclusionAI aimed at graphic-design output, with an emphasis on legible text rendering inside the generated image. It generates from a prompt only and does not accept reference images. Output format can be requested as PNG, JPEG, or WebP. Image dimensions are chosen by the model rather than by the request, so explicit sizes and aspect ratios are rejected instead of silently reshaped. llms.txt
- OpenAI: GPT-5 Image Mini by OpenAI. GPT-5 Image Mini combines OpenAI's advanced language capabilities, powered by [GPT-5 Mini](https://openrouter.ai/openai/gpt-5-mini), with GPT Image 1 Mini for efficient image generation. This natively multimodal model features superior instruction following, text rendering, and detailed image editing with reduced latency and cost. It excels at high-quality visual creation while maintaining strong text understanding, making it ideal for applications that require both efficient image generation and text processing at scale. llms.txt
- OpenAI: GPT-5 Image by OpenAI. [GPT-5](https://openrouter.ai/openai/gpt-5) Image combines OpenAI's GPT-5 model with state-of-the-art image generation capabilities. It offers major improvements in reasoning, code quality, and user experience while incorporating GPT Image 1's superior instruction following, text rendering, and detailed image editing. llms.txt
- Microsoft AI: MAI-Image-2.6 Flash by Microsoft. MAI-Image-2.6 Flash is the lower-latency, lower-cost member of the [MAI-Image-2.6](/microsoft/mai-image-2.6) family from Microsoft AI, built for latency-sensitive, high-throughput production image generation and editing at comparable quality to the precision tier. It supports the same multi-reference editing with up to five input images and the same fixed or model-chosen aspect ratios. llms.txt
- Black Forest Labs: FLUX.2 Max by black-forest-labs. FLUX.2 [max] is the new top-tier image model from Black Forest Labs, pushing image quality, prompt understanding, and editing consistency to the highest level yet. Pricing is as follows, [per the docs](https://bfl.ai/pricing?category=flux.2): Input: We charge $0.03 for each megapixel on the input (i.e. reference images for editing) Output: The first generated megapixel is charged $0.07. Each subsequent megapixel is charged $0.03. llms.txt
- Qwen: Qwen Image 3 by Qwen. Qwen Image 3 is a unified image generation and editing model from Qwen. It supports precise rendering of text and details as small as 10px, along with a richer world knowledge base than previous generations. llms.txt
- Qwen: Qwen Image 3 Pro by Qwen. Qwen Image 3 Pro is an image generation and editing model from Qwen. It supports precise rendering of text and details as small as 10px, along with richer world knowledge compared to previous generations. llms.txt
- Recraft: Recraft V4.1 by Recraft. Recraft V4.1 is an image generation model from Recraft tuned for high aesthetics. It supports text and image inputs with image output at ~1K resolution across multiple aspect ratios, with typical generation around 6 seconds. Compared to V4, photorealism feels more natural with quieter backgrounds and more purposeful lighting, 3D rendering and soft gradients are smoother, and the model follows shorter prompts more reliably. Suited for exploration, concept work, and everyday creative work where speed and cost matter. Supports the following `image_config` parameters: `strength` (controls how much the output deviates from the source image), `rgb_colors` (sets a color palette), and `background_rgb_color` (sets the background color). See the image generation docs for details: https://openrouter.ai/docs/features/multimodal/image-generation Note: only one input image is supported. llms.txt
- Krea: Krea 2 Medium Turbo by krea. Krea 2 Medium Turbo is a distilled, speed-focused variant of Krea 2 Medium from Krea. It is designed for rapid iteration and graphic design exploration where fast generation is the priority. llms.txt
- Recraft: Recraft V4.1 Flash by Recraft. Recraft V4.1 Flash is a text-to-image model from Recraft, the speed and cost tier of the V4.1 family. It generates ~1K raster images in about 1.5 seconds end to end, at roughly a fifth of the cost of [Recraft V4.1](/recraft/recraft-v4.1), and is suited for rapid iteration, prototyping, and high-volume generation where turnaround matters more than peak fidelity. It is generation only: image inputs, styles, and style references are not supported. Supports the same aspect ratios as V4.1 and the `image_config` parameters `rgb_colors` (sets a color palette) and `background_rgb_color` (sets the background color). See the image generation docs for details: https://openrouter.ai/docs/features/multimodal/image-generation llms.txt
- OpenAI: GPT Image 1 Mini by OpenAI. A cost-efficient variant of GPT Image 1 for high-quality image generation at reduced latency and cost via OpenAI's dedicated Images API. llms.txt
- OpenAI: GPT Image 1 by OpenAI. OpenAI's GPT Image 1 generates and edits images via the dedicated Images API. Features accurate text rendering, transparent backgrounds, and up to 16 reference images for edits. llms.txt
- Sourceful: Riverflow V2.5 Pro by Sourceful. Riverflow V2.5 Pro is the most powerful variant of Sourceful's Riverflow 2.5 lineup, best for top-tier control and quality-sensitive outputs. The Riverflow 2.5 series is a unified text-to-image and image-to-image family that treats generation as a production workflow, using an integrated reasoning model to plan multi-step edits and judge candidates before accepting a result. Riverflow 2.5 combines their reasoning with a mix of closed and open image diffusion models to provide greater accuracy and steerability.Reasoning effort is controllable via the reasoning parameter (low/medium/high/xhigh) - higher levels do more editing passes and apply a stricter internal judge, with xhigh suited to batch runs that need high repeatability. It generates at 1K, 2K, and 4K resolution and accepts up to 10 input images for editing. Pricing is dynamic: cost is finalized per job at completion based on billable processing, so it scales with reasoning effort, resolution, and editing complexity rather than a fixed per-image rate. Additional features (via image_config): - Custom font rendering via font_inputs (max 2) to match brand lettering, spacing, and weight - Custom scoring via scoring_prompt and scoring_rubric, so the reasoning model evaluates and steers each candidate against the criteria you care about - Background control via background_mode (original, transparent, solid) and background_hex_color See the image generation docs for details: https://openrouter.ai/docs/features/multimodal/image-generation Note: Sourceful imposes a 4.5MB request size limit, therefore it is highly recommended to pass image URLs instead of Base64 data. llms.txt
- Microsoft AI: MAI-Image-2.6 by Microsoft. MAI-Image-2.6 is an image generation and editing model from Microsoft AI, the precision tier of the MAI-Image-2.6 family alongside the faster [MAI-Image-2.6 Flash](/microsoft/mai-image-2.6-flash). It is suited for design-ready visuals and controlled, iterative editing, and is particularly strong at multi-reference editing that combines people, products, styles, and scenes from up to five input images while preserving composition, style, and detail. Aspect ratio can be fixed or left to the model to choose for the composition. llms.txt
- Microsoft AI: MAI-Image-2.5 by Microsoft. Microsoft AI's MAI-Image-2.5 is a high-quality image generation model available via Azure AI Foundry. It produces photorealistic and artistic images from text prompts with support for various aspect ratios. llms.txt
- inclusionAI: Ming Image 0.1 Design Layer by inclusionai. Ming Image 0.1 Design Layer is an image-to-image model from inclusionAI that decomposes a flattened design image into separate RGBA layers, such as a background layer and foreground elements, and returns one image per layer. It requires exactly one reference image and a prompt describing the layer plan. Output format can be requested as PNG or WebP. Output dimensions follow the input image, so explicit sizes and aspect ratios are rejected instead of silently reshaped. llms.txt
- Black Forest Labs: FLUX.2 Flex by black-forest-labs. FLUX.2 [flex] excels at rendering complex text, typography, and fine details, and supports multi-reference editing in the same unified architecture. Pricing is as follows, [per the docs](https://bfl.ai/pricing?category=flux.2): We charge $0.06 for each megapixel on both input and output side. llms.txt
- Sourceful: Riverflow V2.5 Fast by Sourceful. Riverflow V2.5 Fast is the speed-optimized variant of Sourceful's Riverflow 2.5 lineup, best for production deployments and latency-critical workflows. The Riverflow 2.5 series is a unified text-to-image and image-to-image family that treats generation as a production workflow, using an integrated reasoning model to plan multi-step edits and judge candidates before accepting a result. Riverflow 2.5 combines their reasoning with a mix of open image diffusion models to provide greater accuracy and steerability. Reasoning effort is controllable via the reasoning parameter (low/medium/high) - higher levels do more editing passes and apply a stricter internal judge, while lower levels return faster for early exploration. It generates at 1K and 2K resolution (no 4K) and accepts up to 4 input images for editing. Pricing is dynamic: cost is finalized per job at completion based on billable processing, so it scales with reasoning effort, resolution, and editing complexity rather than a fixed per-image rate. Additional features (via image_config): - Custom font rendering via font_inputs (max 2) to match brand lettering, spacing, and weight - Custom scoring via scoring_prompt and scoring_rubric, so the reasoning model evaluates and steers each candidate against the criteria you care about - Background control via background_mode (original, transparent, solid) and background_hex_color See the image generation docs for details: https://openrouter.ai/docs/features/multimodal/image-generation Note: Sourceful imposes a 4.5MB request size limit, therefore it is highly recommended to pass image URLs instead of Base64 data. llms.txt
- Krea: Krea 2 Large by krea. Krea 2 Large is Krea's high-capability image generation model, more than twice the size of Krea 2 Medium. Its lighter post-training gives images a rawer, more textured, and flexible character, making it particularly suited for photorealism, expressive artistic styles, motion blur, grain, and low-dynamic-range looks. llms.txt
- Sourceful: Riverflow V2 Fast by Sourceful. Riverflow V2 Fast is the fastest variant of Sourceful's Riverflow 2.0 lineup, best for production deployments and latency-critical workflows. The Riverflow 2.0 series represents SOTA performance on image generation and editing tasks, using an integrated reasoning model to boost reliability and tackle complex challenges. Pricing is $0.02 per 1K output image and $0.04 per 2K output image. Does not support 4K image output. Additional features: - Custom font rendering via font_inputs ($0.03/font, max 2) - Image enhancement via super_resolution_references ($0.20/reference, max 4) See the image generation docs for details: https://openrouter.ai/docs/features/multimodal/image-generation Note: Sourceful imposes a 4.5MB request size limit, therefore it is highly recommended to pass image URLs instead of Base64 data. llms.txt
- Recraft: Recraft V4.1 Vector by Recraft. Recraft V4.1 Vector is the vector (SVG) variant of Recraft V4.1, tuned for high aesthetics. It supports text and image inputs and produces SVG image output across multiple aspect ratios, with typical generation around 10 seconds. Output scales cleanly, making it suitable for icons, logos, and other graphics. V4.1 brings more personality to text and illustrations, smoother gradients, and stronger short-prompt adherence compared to V4. Suited for everyday illustration work where output should be designed rather than photographed. Supports the following `image_config` parameters: `strength` (controls how much the output deviates from the source image), `rgb_colors` (sets a color palette), and `background_rgb_color` (sets the background color). See the image generation docs for details: https://openrouter.ai/docs/features/multimodal/image-generation Note: only one input image is supported. llms.txt
- Recraft: Recraft V4 Styles by Recraft. Recraft V4 Styles is a style-consistent image generation model from Recraft. Every request requires at least one style reference image and generates a new image that reproduces the reference's rendering technique, colour, texture, and composition rather than editing it. It produces raster images at approximately 1K resolution. Note: 1 to 10 reference images in PNG, JPG or WEBP, each at least 256 px on its shortest edge. Pricing has two parts: $0.035 per image plus a one-time $0.005 style-creation charge per request, regardless of how many images that request returns. A single-image request costs $0.040, and a six-image request costs $0.215. llms.txt
- Krea: Krea 2 Medium by krea. Krea 2 Medium is Krea's balanced, cost-efficient image generation model and a practical starting point for a broad range of use cases. Its extensive post-training supports stable, consistent generations, with particular strengths in illustration, anime, painting, and other expressive artistic styles. llms.txt
- Recraft: Recraft V4 Styles Vector by Recraft. Recraft V4 Styles Vector is a style-consistent image generation model from Recraft. Every request requires at least one style reference image and generates a new image that reproduces the reference's rendering technique, colour, texture, and composition rather than editing it. It produces SVG output for graphics that need to scale cleanly. Note: 1 to 10 reference images in PNG, JPG or WEBP, each at least 256 px on its shortest edge. Pricing has two parts: $0.05 per image plus a one-time $0.005 style-creation charge per request, regardless of how many images that request returns. A single-image request costs $0.055, and a six-image request costs $0.305. llms.txt
- Microsoft AI: MAI-Image-2.5 Pro by Microsoft. Microsoft AI's MAI-Image-2.5 is a high-quality image generation model available via Azure AI Foundry. It produces photorealistic and artistic images from text prompts with support for various aspect ratios. llms.txt
- Recraft: Recraft V4.1 Pro Vector by Recraft. Recraft V4.1 Pro Vector is the vector (SVG) variant of Recraft V4.1 Pro, tuned for high aesthetics. It supports text and image inputs and produces higher-resolution SVG image output across multiple aspect ratios, with typical generation around 15 seconds. Output scales cleanly, making it suitable for icons, logos, and other graphics. V4.1 brings more personality to text and illustrations, smoother gradients, and stronger short-prompt adherence compared to V4 Pro. Suited for higher-resolution illustration work and production graphics where output should be designed rather than photographed. Supports the following `image_config` parameters: `strength` (controls how much the output deviates from the source image), `rgb_colors` (sets a color palette), and `background_rgb_color` (sets the background color). See the image generation docs for details: https://openrouter.ai/docs/features/multimodal/image-generation Note: only one input image is supported. llms.txt
- Recraft: Recraft V4 by Recraft. Recraft V4 is an image generation model from Recraft. It supports text and image inputs with image output at ~1K resolution across multiple aspect ratios. It delivers stronger compositional judgment, color coherence, and legible embedded text compared to V3, making it suited for infographics, signage, and packaging. Supports the following `image_config` parameters: `strength` (controls how much the output deviates from the source image), `rgb_colors` (sets a color palette), and `background_rgb_color` (sets the background color). See the image generation docs for details: https://openrouter.ai/docs/features/multimodal/image-generation Note: only one input image is supported. llms.txt
- Recraft: Recraft V4.1 Utility by Recraft. Recraft V4.1 Utility is a general-purpose image generation model from Recraft. It supports text and image inputs with image output at ~1K resolution across multiple aspect ratios, with typical generation around 6 seconds. The Utility line is designed for restraint as the aesthetic choice - flat lighting, front-facing composition, and simple, controlled scenes - making it a practical fit for product imagery, mockups, and structured visuals where the high-aesthetic V4.1 line would be too expressive. V4.1 improvements over V4 include more natural object understanding, sharper mockups, cleaner default icons, and stronger prompt adherence with shorter prompts. Supports the following `image_config` parameters: `strength` (controls how much the output deviates from the source image), `rgb_colors` (sets a color palette), and `background_rgb_color` (sets the background color). See the image generation docs for details: https://openrouter.ai/docs/features/multimodal/image-generation Note: only one input image is supported. llms.txt
- Recraft: Recraft V4.1 Pro by Recraft. Recraft V4.1 Pro is an image generation model from Recraft tuned for high aesthetics. It supports text and image inputs with image output at ~2K resolution across multiple aspect ratios - double the resolution of V4.1 - with typical generation around 10 seconds. It shares the V4.1 visual sensibility at higher fidelity and detail density, with more natural photorealism, smoother 3D rendering and gradients, and stronger short-prompt adherence than V4 Pro. Suited for production work where image quality is the priority and the idea benefits from more resolution to breathe. Supports the following `image_config` parameters: `strength` (controls how much the output deviates from the source image), `rgb_colors` (sets a color palette), and `background_rgb_color` (sets the background color). See the image generation docs for details: https://openrouter.ai/docs/features/multimodal/image-generation Note: only one input image is supported. llms.txt
- Recraft: Recraft V4 Vector by Recraft. Recraft V4 Vector is the vector (SVG) variant of Recraft V4. It supports text and image inputs and produces vector image output across multiple aspect ratios. Compared to the raster V4, output is delivered as SVG, suitable for icons, logos, and other graphics that need to scale cleanly. V4 delivers stronger compositional judgment, color coherence, and legible embedded text compared to V3. Supports the following `image_config` parameters: `strength` (controls how much the output deviates from the source image), `rgb_colors` (sets a color palette), and `background_rgb_color` (sets the background color). See the image generation docs for details: https://openrouter.ai/docs/features/multimodal/image-generation Note: only one input image is supported. llms.txt
- Sourceful: Riverflow V2 Pro by Sourceful. Riverflow V2 Pro is the most powerful variant of Sourceful's Riverflow 2.0 lineup, best for top-tier control and perfect text rendering. The Riverflow 2.0 series represents SOTA performance on image generation and editing tasks, using an integrated reasoning model to boost reliability and tackle complex challenges. Pricing is $0.15 per 1K/2K output image and $0.33 per 4K output image. Additional features: - Custom font rendering via font_inputs ($0.03/font, max 2) - Image enhancement via super_resolution_references ($0.20/reference, max 4) See the image generation docs for details: https://openrouter.ai/docs/features/multimodal/image-generation Note: Sourceful imposes a 4.5MB request size limit, therefore it is highly recommended to pass image URLs instead of Base64 data. llms.txt
- Recraft: Recraft V4 Styles Pro by Recraft. Recraft V4 Styles Pro is a style-consistent image generation model from Recraft. Every request requires at least one style reference image and generates a new image that reproduces the reference's rendering technique, colour, texture, and composition rather than editing it. It produces raster images at approximately 2K resolution. Note: 1 to 10 reference images in PNG, JPG or WEBP, each at least 256 px on its shortest edge. Pricing has two parts: $0.10 per image plus a one-time $0.005 style-creation charge per request, regardless of how many images that request returns. A single-image request costs $0.105, and a six-image request costs $0.605. llms.txt
- Recraft: Recraft V4 Pro by Recraft. Recraft V4 Pro is an image generation model from Recraft. It supports text and image inputs with image output at ~2K resolution across multiple aspect ratios, double the resolution of V4. It offers higher fidelity and detail density than V4, suited for production use cases where image quality is a priority. Supports the following `image_config` parameters: `strength` (controls how much the output deviates from the source image), `rgb_colors` (sets a color palette), and `background_rgb_color` (sets the background color). See the image generation docs for details: https://openrouter.ai/docs/features/multimodal/image-generation Note: only one input image is supported. llms.txt
- Recraft: Recraft V3 by Recraft. Recraft V3 is an image generation model from Recraft. It supports text and image inputs with image output at ~1K resolution across multiple aspect ratios. Supports the following `image_config` parameters: `strength` (controls how much the output deviates from the source image), `style` (applies an artistic style), `text_layout` (places text at specific positions), `rgb_colors` (sets a color palette), and `background_rgb_color` (sets the background color). See the image generation docs for details: https://openrouter.ai/docs/features/multimodal/image-generation Note: only one input image is supported. llms.txt
- Recraft: Recraft V4.1 Utility Pro by Recraft. Recraft V4.1 Utility Pro is a general-purpose image generation model from Recraft. It supports text and image inputs with image output at ~2K resolution across multiple aspect ratios — double the resolution of V4.1 Utility - with typical generation around 10 seconds. Like V4.1 Utility, it is designed for restraint as the aesthetic choice - flat lighting, front-facing composition, and simple, controlled scenes - at higher fidelity for production use cases. V4.1 improvements over V4 Pro include more natural object understanding, sharper mockups, cleaner default icons, and stronger prompt adherence with shorter prompts. Suited for product imagery, e-commerce mockups, and structured visuals where image quality matters and the high-aesthetic V4.1 Pro line would be too expressive. Supports the following `image_config` parameters: `strength` (controls how much the output deviates from the source image), `rgb_colors` (sets a color palette), and `background_rgb_color` (sets the background color). See the image generation docs for details: https://openrouter.ai/docs/features/multimodal/image-generation Note: only one input image is supported. llms.txt
- Recraft: Recraft V4 Pro Vector by Recraft. Recraft V4 Pro Vector is the vector (SVG) variant of Recraft V4 Pro. It supports text and image inputs and produces vector image output across multiple aspect ratios at the higher fidelity Pro tier. Output is delivered as SVG, suitable for icons, logos, and other graphics that need to scale cleanly. V4 Pro offers higher fidelity and detail density than V4, with stronger compositional judgment, color coherence, and legible embedded text compared to V3. Supports the following `image_config` parameters: `strength` (controls how much the output deviates from the source image), `rgb_colors` (sets a color palette), and `background_rgb_color` (sets the background color). See the image generation docs for details: https://openrouter.ai/docs/features/multimodal/image-generation Note: only one input image is supported. llms.txt
- Recraft: Recraft V4 Styles Pro Vector by Recraft. Recraft V4 Styles Pro Vector is a style-consistent image generation model from Recraft. Every request requires at least one style reference image and generates a new image that reproduces the reference's rendering technique, colour, texture, and composition rather than editing it. It produces SVG output for graphics that need to scale cleanly. Note: 1 to 10 reference images in PNG, JPG or WEBP, each at least 256 px on its shortest edge. Pricing has two parts: $0.12 per image plus a one-time $0.005 style-creation charge per request, regardless of how many images that request returns. A single-image request costs $0.125, and a six-image request costs $0.725. llms.txt
Top Video models
Ranked by tokens processed on OpenRouter over the past week, with generated video seconds breaking ties and ranking models without token usage. Free variants and OpenRouter routers are listed as their own entries.
- ByteDance: Seedance 2.5 by bytedance. Seedance 2.5 is a video generation model from ByteDance. It is suited for long-form storytelling, multimodal reference-based generation, video editing, and video extension. It supports first-frame and first-and-last-frame control, up to 50 image, video, and audio reference assets, optional generated audio, and multilingual audiovisual generation. llms.txt
- Google: Veo 3.1 Lite by Google. Google's most cost-effective video generation model, designed for high-volume applications and rapid iteration. Veo 3.1 Lite generates 720p and 1080p video from text or image prompts with native synchronized audio at less than 50% of the cost of Veo 3.1 Fast. Supports 4–8 second clips in landscape (16:9) and portrait (9:16) formats, with SynthID watermarking. Ideal for content platforms, short-form video creation, and automated media generation. llms.txt
- ByteDance: Seedance 2.0 Fast by bytedance. Seedance 2.0 Fast is a video generation model from ByteDance. It supports text-to-video, image-to-video with first and last frame control, and multimodal reference-to-video. It prioritizes generation speed and lower cost over maximum output quality. The number of tokens is given by (height of output video * width of output video * duration * 24) / 1024 llms.txt
- ByteDance: Seedance 2.0 Mini by bytedance. Seedance 2.0 Mini is a video generation model from ByteDance. It supports text-to-video, image-to-video with first and last frame control, and multimodal reference-to-video with image, video, and audio inputs. It supports 480p and 720p output for 4-15 second videos. The number of tokens is given by (height of output video * width of output video * duration * 24) / 1024 llms.txt
- ByteDance: Seedance 2.0 by bytedance. Seedance 2.0 is a video generation model from ByteDance. It supports text-to-video, image-to-video with first and last frame control, and multimodal reference-to-video. It is particularly strong at preserving character consistency, visual style, and camera movement from reference material. The number of tokens is given by (height of output video * width of output video * duration * 24) / 1024 llms.txt
- Alibaba: Wan 3.0 by alibaba. Wan 3.0 is a video generation model from Alibaba for text-to-video, image-to-video, and reference-guided video generation. It produces 480p, 720p, or 1080p video with durations from 2 to 30 seconds. llms.txt
- SpaceXAI: Grok Imagine Video by SpaceXAI. Grok Imagine Video is SpaceXAI's fast, text-, image-, and reference-conditioned video generation model. It produces short videos (1–15 seconds, 24 fps) at 480p or 720p across seven aspect ratios - 1:1, 16:9, 9:16, 4:3, 3:4, 3:2, and 2:3. The model supports three generation modes: text-to-video from a prompt alone, image-to-video that animates a still input, and reference-to-video that grounds the output in up to seven reference images for consistent characters, styles, or settings. llms.txt
- Google: Veo 3.1 Fast by Google. Google's mid-tier video generation model balancing speed and quality. Veo 3.1 Fast generates high-quality video from text or image prompts with native synchronized audio, offering faster turnaround than Veo 3.1 at lower cost. Supports first-frame and last-frame conditioning, multiple resolutions and aspect ratios, and SynthID watermarking. llms.txt
- SpaceXAI: Grok Imagine Video 1.5 by SpaceXAI. Grok Imagine Video 1.5 is a video generation model from SpaceXAI. It creates videos from text prompts, with an optional starting image to guide the scene. It can direct subject and camera motion, pacing, atmosphere, and physical behavior while maintaining visual continuity, and can generate synchronized sound effects, ambience, and dialogue. llms.txt
- MiniMax: H3 Max by MiniMax. MiniMax H3 Max is a video-generation model from MiniMax, jointly released with fal.ai. Derived through additional training from MiniMax H3, it is designed for faster text-to-video and image-to-video generation with controlled first-frame or last-frame keyframes. llms.txt
- MiniMax: H3 by MiniMax. MiniMax H3 is a lightweight, open-weights video generation model from MiniMax. It is designed for precise multimodal editing and controlled content generation, including instruction-guided edits, text and brand rendering, and video-to-video motion transfer. The model is suited for commercial creative workflows across advertising, e-commerce, gaming, and interface design, with native audiovisual output for reference-driven generation. llms.txt
- Kling: Video v3.0 Pro by kwaivgi. Kling v3.0 Pro is Kuaishou's premium video generation model, offering higher visual quality than the Standard tier. It supports text-to-video and image-to-video workflows, with first-frame and last-frame control for precise scene composition. Clips range from 3 to 15 seconds in 16:9, 9:16, or 1:1 aspect ratios. Native audio generation is available as an option. llms.txt
- Kling: Video v3.0 Standard by kwaivgi. Kling v3.0 Standard is a video generation model from Kuaishou. It supports text-to-video and image-to-video workflows, with first-frame and last-frame control for guided scene composition. Clips range from 3 to 15 seconds in 16:9, 9:16, or 1:1 aspect ratios. Native audio generation is available as an option. llms.txt
- ByteDance: Seedance 1.5 Pro by bytedance. ByteDance's next-generation audio-visual generation model with a 4.5B parameter Dual-Branch Diffusion Transformer architecture. Seedance 1.5 Pro generates video and audio simultaneously in a single unified pass — eliminating the timing issues of sequential audio dubbing. Supports multi-language lip-sync (English, Mandarin, Japanese, Korean, Spanish, and more), cinematic camera control (pan, tilt, zoom, orbit), multi-character dialogue, and character consistency across shots. Produces clips from 4–12 seconds at up to 1080p. The number of tokens is given by (height of output video * width of output video * duration * 24) / 1024 llms.txt
- Alibaba: Wan 3.0 Prime by alibaba. Wan 3.0 Prime is a fast-mode variant of [Wan 3.0](https://openrouter.ai/alibaba/wan-3.0) from Alibaba. It supports text-to-video and first-frame image-to-video generation. llms.txt
- HeyGen: Avatar IV by heygen. HeyGen: Avatar IV is an image-to-video model that animates a single photo into an expressive, lip-synced talking-head video. Rather than only matching mouth shapes to words, it interprets the vocal tone, rhythm, and emotion of the audio to drive head motion, facial expression, and gestures, producing output at up to 1080p. The spoken audio comes from one of two inputs: a text script, which the model voices with HeyGen text-to-speech, or a supplied audio track, which the image is lip-synced to directly. Passthrough parameters let you choose a voice, tune voice settings, set expressiveness, prompt specific motion, replace or remove the background, add captions, and title the video. llms.txt
- Alibaba: HappyHorse 1.1 by alibaba. HappyHorse 1.1 is a video generation model from Alibaba. It generates short videos from a text prompt, a single starting image, or a set of reference images, with output up to 1080p and durations of 3 to 15 seconds. It is suited for creative content, social media clips, and image-driven animation, and improves on the prior version with stronger prompt adherence, smoother motion, and more consistent characters across frames. llms.txt
- Google: Veo 3.1 by Google. Google's state-of-the-art video generation model, built for maximum visual fidelity in final production cuts. Veo 3.1 generates high-quality 1080p video from text or image prompts with native synchronized audio — including dialogue, ambient effects, and background sound. Supports scene extension (up to 20 chained clips for 140+ second narratives), frames-to-video transitions between two images, vertical video for Shorts, and 4K upscaling. llms.txt
- Alibaba: Wan 2.7 by alibaba. Wan 2.7 is a video generation model from Alibaba. It supports text-to-video, image-to-video with first and last frame control, and reference-to-video, where multiple reference images guide the style and content of the generated scene. llms.txt
- Black Forest Labs: FLUX.3 Video by black-forest-labs. FLUX.3 Video is a video generation model from Black Forest Labs. It supports text-to-video, image-guided generation with opening and closing keyframes, and video continuation workflows, making it suited for controlled creative production, scene animation, and extending existing clips. It can generate synchronized audio alongside the video. llms.txt
- Kling: Video O1 by kwaivgi. Kling Video O1 is a video generation model from Kuaishou. It supports text and image inputs with video output, enabling text-to-video and image-to-video workflows. It is suited for cinematic content production, with first-frame and last-frame control for precise scene composition. It generates 5 or 10 second clips in 16:9, 9:16, or 1:1 aspect ratios. llms.txt
- Black Forest Labs: FLUX Video Edit by black-forest-labs. FLUX Video Edit [fast] takes a source video and an edit prompt and returns a precisely edited video. Add, remove, or replace objects and characters, rebuild the setting, edit on-screen text, change colors, materials and visual effects, change or translate dialogue with lip sync, or change the events of a clip. Output preserves the source duration, aspect ratio, and audio; inputs above 720p are downscaled to 720p. Source clips up to 15 seconds and 50 MiB. llms.txt
- Alibaba: Wan 2.6 by alibaba. Alibaba's most advanced video generation model, supporting over 10 visual creation capabilities in a unified system. Wan 2.6 generates 1080p video at 24fps from text, images, reference videos, or audio, with native audio-visual synchronization and precise lip-sync. Key features include reference-to-video (insert a character's appearance and voice into new scenes), multi-shot storytelling from simple prompts, synchronized sound effects and music, and support for 16:9, 9:16, and 1:1 aspect ratios with clips up to 15 seconds. llms.txt
- MiniMax: Hailuo 2.3 by MiniMax. Hailuo 2.3 is a video generation model from MiniMax. It accepts text prompts and reference images as input and generates video output, supporting both text-to-video and image-to-video workflows. It is suited for creative content production, cinematic scene generation, and character animation, with a focus on realistic motion and expressive character rendering. llms.txt
- Runway: Gen-4.5 by Runway. Runway Gen-4.5 is a video generation model from Runway for text-to-video and image-to-video workflows. It is designed for cinematic scene creation with strong motion quality, visual fidelity, and prompt adherence. The model can coordinate complex action sequences, detailed camera movement, precise event timing, and atmospheric changes, making it suited for narrative shots, product reveals, and controlled image animation. llms.txt
- Alibaba: HappyHorse 1.0 by alibaba. HappyHorse 1.0 is a video generation model from Alibaba. It generates short videos from a text prompt, a single starting image, or a set of reference images, with output up to 1080p and durations of 3 to 15 seconds. It is suited for creative content, social media clips, and image-driven animation across a range of aspect ratios. llms.txt
- OpenAI: Sora 2 Pro by OpenAI. OpenAI's flagship video generation model, delivering production-quality video with physics-accurate motion, synchronized audio, and world-state persistence across shots. Sora 2 Pro follows intricate multi-shot instructions while maintaining consistent spatial relationships — objects don't disappear or change shape between cuts. Supports text-to-video and image-to-video, with synchronized background soundscapes, speech, and sound effects. Includes advanced content safety with C2PA metadata provenance and SynthID-style watermarking. llms.txt
- Runway: Aleph 2.0 by Runway. Runway Aleph 2.0 is an in-context video editing model from Runway. It applies text instructions and keyframe-guided edits across existing footage while preserving details that are not meant to change. It is suited for adding, removing, or replacing subjects and objects, changing backgrounds or outfits, relighting scenes, shifting weather or time of day, restyling footage, and expanding a shot to a new aspect ratio. llms.txt
- Black Forest Labs: FLUX Video Upscale by black-forest-labs. FLUX Video Upscale is a video upscaling model from Black Forest Labs. It enlarges a single source video by 1.5× to 3× while preserving its duration, with an optional prompt and precise or creative processing modes. llms.txt
Top Speech models
Ranked by tokens processed on OpenRouter over the past week. Free variants and OpenRouter routers are listed as their own entries.
- Google: Gemini 3.1 Flash TTS Preview by Google. Gemini 3.1 Flash TTS Preview is a text-to-speech model from Google, and a substantial generational step up from Gemini 2.5 Flash TTS. It takes text input and produces audio output across 70+ languages — nearly 3× the language coverage of its predecessor. The headline addition is a system of 200+ inline audio tags (e.g. `[whispers]`, `[laughs]`, `[excited]`) that let developers steer delivery, emotion, and pacing mid-sentence, alongside a "director's chair" workflow in Google AI Studio for defining per-character Audio Profiles and scene-level context. It supports up to two speakers with independent voice and style configuration per speaker, outputs PCM audio at 24 kHz / 16-bit mono, and automatically watermarks all output with SynthID. Context window is 32k tokens. llms.txt
- hexgrad: Kokoro 82M by hexgrad. Kokoro 82M is a lightweight, open-weight text-to-speech model from hexgrad. It converts text to speech across 8 languages (American and British English, Spanish, French, Hindi, Italian, Japanese, Portuguese, and Chinese) using 54 preset voices organized by language and gender. At 82M parameters, it is well-suited for multilingual TTS deployments where footprint and cost efficiency matter. llms.txt
- Google: Gemini 3.8 Flash TTS by Google. Gemini 3.8 Flash TTS is a text-to-speech model from Google and the successor to [Gemini 3.1 Flash TTS Preview](https://openrouter.ai/google/gemini-3.1-flash-tts-preview). It is the creative tier of the 3.8 TTS family, suited for narration, character work, and multi-speaker dialogue where voice fidelity, acting nuance, and regional dialect coverage matter more than latency. It accepts the same speech configuration as its predecessor, including the 30 prebuilt studio voices, and supports Google's Voice Design (custom voices from a text description) and Voice Replication (voices cloned from a short sample) through `voice_...` identifiers. llms.txt
- SpaceXAI: Grok Voice TTS 1.0 by SpaceXAI. Grok Voice TTS 1.0 is a text-to-speech model from SpaceXAI. It converts text into spoken audio across 20+ languages with automatic language detection, and offers five built-in voices (Eve, Ara, Rex, Sal, Leo) covering a range of tones. Inline speech tags allow control over pauses, emphasis, pitch, speed, and vocal style. Output is available in MP3, WAV, PCM, μ-law, and A-law formats at sample rates from 8 kHz to 48 kHz, with up to 15,000 characters per request. llms.txt
- Fish Audio: S2.1 Pro Free (free) by fish-audio. S2.1 Pro Free is the no-cost variant of Fish Audio S2.1 Pro, intended for testing, prototyping, and low-volume applications. It provides the same synthesis capabilities without production latency or availability guarantees. llms.txt
- Google: Gemini 3.8 Flash Lite TTS by Google. Gemini 3.8 Flash Lite TTS is a text-to-speech model from Google and the fast, high-throughput member of the 3.8 TTS family alongside [Gemini 3.8 Flash TTS](https://openrouter.ai/google/gemini-3.8-flash-tts). It is suited for conversational voice-agent pipelines, high-volume read-aloud, and voice replication workloads where latency and cost matter more than maximum expressiveness. It accepts the same speech configuration as [Gemini 3.1 Flash TTS Preview](https://openrouter.ai/google/gemini-3.1-flash-tts-preview), including the 30 prebuilt studio voices, and supports Google's Voice Design and Voice Replication through `voice_...` identifiers. llms.txt
- Qwen: Qwen-Audio-3.0-TTS Plus by Qwen. Qwen-Audio-3.0-TTS Plus is Alibaba's higher-quality text-to-speech model, generating spoken audio from text via the DashScope Speech Synthesizer API. llms.txt
- Fish Audio: S2.1 Pro by fish-audio. S2.1 Pro is a production-oriented text-to-speech model from Fish Audio. It is suited for multilingual voice applications, expressive narration, and dialogue synthesis, with open-ended natural-language controls for speaking style and emotion. llms.txt
- Deepgram: Flux TTS (free) by deepgram. Flux TTS is a text-to-speech model from Deepgram. It is suited for natural, expressive English speech synthesis across Deepgram's Flux voice catalog. llms.txt
- Microsoft AI: MAI-Voice-2 by Microsoft. MAI-Voice-2 is an expressive text-to-speech model from Microsoft AI. It is suited for conversational assistants, media narration, accessibility, education, and other long-form voice applications. It supports 15 languages across 18 locales, fine-grained control of tone and delivery, multi-speaker generation, and voice prompting from short audio clips without fine-tuning. The model prioritizes naturalness and expressivity over latency-critical generation. llms.txt
- Qwen: Qwen-Audio-3.0-TTS Flash by Qwen. Qwen-Audio-3.0-TTS Flash is Alibaba's fast, cost-efficient text-to-speech model, generating spoken audio from text via the DashScope Speech Synthesizer API. llms.txt
- Microsoft AI: MAI-Voice-2-Flash by Microsoft. MAI-Voice-2-Flash is a low-latency text-to-speech model from Microsoft AI for voice agents, assistants, call centers, accessibility, narration, and other interactive applications. It generates expressive 24 kHz mono speech across 15 languages and 18 locales, with fine-grained control over tone and delivery. Voice prompting and cloning require Microsoft AI-approved access and appropriate speaker consent. llms.txt
- Deepgram: Aura-2 by deepgram. Aura-2 is a multilingual text-to-speech model from Deepgram. It supports Deepgram’s canonical Aura-2 voice catalog for speech synthesis across multiple languages. llms.txt
- MiniMax: Speech 2.8 HD by MiniMax. MiniMax Speech 2.8 HD is a text-to-speech model from MiniMax. It is suited for applications that generate spoken audio from text and accepts arbitrary MiniMax voice IDs. llms.txt
- MiniMax: Speech 2.8 Turbo by MiniMax. MiniMax Speech 2.8 Turbo is a text-to-speech model from MiniMax. It is suited for applications that generate spoken audio from text and accepts arbitrary MiniMax voice IDs. llms.txt
- Mistral: Voxtral Mini TTS by Mistral AI. Voxtral Mini TTS is Mistral's text-to-speech model featuring zero-shot voice cloning and multilingual support. It converts text input into natural-sounding audio output. llms.txt
- Canopy Labs: Orpheus 3B by canopylabs. Orpheus 3B is an English text-to-speech model from Canopy Labs, fine-tuned for natural prosody and expressive delivery. It offers 7 preset voices and is suited for narration, voice assistants, and interactive applications where naturalistic speech is a priority. llms.txt
- Fish Audio: S2 Pro by fish-audio. S2 Pro is a multilingual text-to-speech model from Fish Audio. It is suited for expressive narration and multi-speaker dialogue, with natural-language controls for speaking style and emotion. llms.txt
- Fish Audio: S1 by fish-audio. S1 is a multilingual text-to-speech model from Fish Audio. It is suited for voice applications that need broad emotional expression, using parenthetical controls to guide speaking style across its supported languages. llms.txt
- Sesame: CSM 1B by sesame. CSM 1B is a conversational speech model from Sesame. It accepts text input and produces English speech output, with voice options spanning conversational and read-speech styles. At 1B parameters, it is suited for dialogue-oriented applications such as voice assistants and interactive agents. llms.txt
Top Transcription models
Ranked by characters transcribed on OpenRouter over the past week. Free variants and OpenRouter routers are listed as their own entries.
- OpenAI: Whisper Large V3 Turbo by OpenAI. Whisper Large V3 Turbo is an optimized version of OpenAI's Whisper Large V3 speech recognition model, designed for speed and cost efficiency. It supports transcription across 99+ languages with a 12% word error rate, and accepts common audio formats including mp3, mp4, wav, webm, flac, and ogg. Achieves real-time speed factors up to 216x, making it well-suited for latency-sensitive and high-throughput transcription workloads. llms.txt
- OpenAI: Whisper Large V3 by OpenAI. Whisper Large V3 is OpenAI's open-source automatic speech recognition model offering both audio transcription and translation. It supports 99+ languages and accepts common audio formats including mp3, mp4, wav, webm, flac, and ogg. With 1,550M parameters, it achieves a 10.3% word error rate and is well-suited for noise-robust, multilingual transcription in demanding conditions. Supports timestamp granularities at word and segment levels. llms.txt
- Microsoft AI: MAI-Transcribe 2 by Microsoft. MAI-Transcribe 2 is a multilingual speech-to-text model from Microsoft AI, ranked #1 on the FLEURS multilingual benchmark. It supports 60 languages with automatic language identification, code switching for mixed-language speech, speaker diarization, word-level timestamps, keyword biasing for domain-specific terminology, and configurable verbatim or clean transcription styles. It is suited for captions, call transcription, subtitling, accessibility, and other voice-enabled applications, and is faster than MAI-Transcribe-1.5 on long-form audio. On OpenRouter, set `response_format` to `"verbose_json"` for segment timestamps, and add `timestamp_granularities: ["word"]` for word-level timestamps. Set `provider.options.azure.diarization.enabled` to `true` for speaker labels, provide keywords through `provider.options.azure.phraseList.phrases`, and select a transcription style through `provider.options.azure.enhancedMode.modelOptions.transcribeStyle`. See the [speech-to-text guide](https://openrouter.ai/docs/guides/overview/multimodal/stt). llms.txt
- OpenAI: GPT-4o Mini Transcribe by OpenAI. GPT-4o Mini Transcribe is OpenAI's smaller, cost-efficient speech-to-text model built on GPT-4o Mini audio capabilities. It's priced per token (input and output), making it suitable for high-volume transcription workflows that benefit from token-level billing transparency at a lower cost point. llms.txt
- Qwen: Qwen3 ASR 1.7B by Qwen. Qwen3 ASR 1.7B is an automatic speech recognition model from Qwen. It supports multilingual language identification and transcription across 30 languages and 22 Chinese dialects, with streaming and offline inference plus segment-level and word-level timestamps. llms.txt
- Qwen: Qwen3 ASR 0.6B by Qwen. Qwen3 ASR 0.6B is a compact automatic speech recognition model from Qwen. It supports multilingual language identification and transcription across 30 languages and 22 Chinese dialects, with streaming and offline inference plus segment-level and word-level timestamps. llms.txt
- OpenAI: Whisper 1 by OpenAI. Whisper is OpenAI's open-source automatic speech recognition model, available via API as `whisper-1`. It supports transcription and translation across 50+ languages from audio files up to 25 MB. Accepts formats including mp3, mp4, wav, and webm. Priced per minute of audio duration, billed to the nearest second. llms.txt
- SpaceXAI: Grok STT 1.0 by SpaceXAI. Grok STT is SpaceXAI's speech-to-text model, available via the REST /v1/stt endpoint. It supports transcription with word-level timestamps, optional speaker diarization, and multichannel audio. llms.txt
- Mistral: Voxtral Mini Transcribe by Mistral AI. Voxtral Mini Transcribe is Mistral's speech-to-text model, derived from the Voxtral Mini family. It accepts audio input and returns transcribed text via the standard transcription API. Suited for transcribing meetings, voice notes, podcasts, and other spoken content. llms.txt
- NVIDIA: Parakeet TDT 0.6B v3 by Nvidia. Parakeet TDT 0.6B v3 is NVIDIA's 600M-parameter multilingual speech-to-text model built on the FastConformer-TDT architecture. Trained on the Granary dataset (670,000+ hours of audio), it supports automatic language detection across all official EU languages and achieves a 6.34% average word error rate on the HuggingFace Open ASR Leaderboard. Returns transcribed text with punctuation and segment timestamps. llms.txt
- OpenAI: GPT Transcribe by OpenAI. GPT Transcribe is a high-accuracy speech-to-text model from OpenAI. It is suited for recorded audio, streamed file transcription, and committed Realtime turns, with free-form context, keyword hints, and multiple language hints for specialized terms and multilingual speech. llms.txt
- Qwen: Qwen3 ASR Flash by Qwen. Qwen3-ASR-Flash is Alibaba's automatic speech recognition service, built on the Qwen3-Omni foundation and trained on tens of millions of hours of multimodal speech data. The model handles 11 languages — including Chinese (with Cantonese, Sichuanese, Minnan, and Wu dialects), English, Arabic, French, German, Spanish, Italian, Portuguese, Russian, Japanese, and Korean — with automatic language detection so no manual configuration is needed for mixed-language audio. The model is designed for difficult acoustic conditions: it transcribes lyrics over background music, handles noisy and far-field recordings, filters silence and non-speech audio, and accepts arbitrary context text (names, jargon, domain terminology) to bias recognition toward specific vocabulary. llms.txt
- OpenAI: GPT-4o Transcribe by OpenAI. GPT-4o Transcribe is OpenAI's high-quality speech-to-text model built on GPT-4o audio capabilities. It's priced per token (input and output), making it suitable for workflows that benefit from token-level billing transparency. llms.txt
- Deepgram: Nova-3 by deepgram. Deepgram Nova-3 general-purpose speech-to-text model with monolingual and multilingual transcription support. llms.txt
- Meta: Muse Voice Transcribe 1.0 by Meta. Muse Voice Transcribe 1.0 is a synchronous speech-to-text model from Meta. It is suited for push-to-talk, endpointing, and speaker-aware transcription, with keyword biasing for domain terms and language biasing through language-name hints. It accepts mono 16-bit PCM WAV audio at 16 kHz or 24 kHz for recordings up to 10 minutes. It does not provide word-level timestamps or confidence scores, and other audio formats must be converted to WAV before upload. llms.txt
- Microsoft AI: MAI-Transcribe 1.5 by Microsoft. MAI-Transcribe 1.5 is a multilingual speech-to-text model from Microsoft AI. It is suited for captions, call transcription, subtitling, accessibility, and other voice-enabled applications, with reliable transcription across 43 languages, diverse accents, and noisy real-world audio. It supports automatic language identification and keyword biasing for domain-specific terminology, and improves long-form transcription speed over MAI-Transcribe-1. Speaker diarization is not supported. llms.txt
- Google: Chirp 3 by Google. Chirp 3 is Google's latest multilingual speech-to-text model. It offers enhanced transcription accuracy across 24 GA languages and 77+ preview languages, with support for automatic language detection, automatic punctuation, and a built-in denoiser for cleaner audio processing. llms.txt
- NVIDIA: Nemotron 3.5 ASR Streaming Multilingual 0.6B by Nvidia. Nemotron 3.5 ASR Streaming Multilingual 0.6B is a speech recognition model from NVIDIA. Its prompt-conditioned, cache-aware FastConformer-RNNT design targets low-latency transcription across more than 40 languages for real-time captioning, voice agents, and multilingual transcription pipelines. llms.txt
- AssemblyAI: Universal-3.5 Pro by Assemblyai. Universal-3.5 Pro is AssemblyAI's speech-to-text model served through its Sync API, returning a complete transcript with word-level timestamps in a single synchronous response for audio clips up to 120 seconds. It accepts 16-bit WAV input and supports free-text prompting, keyterms, and conversation context to steer transcription toward domain vocabulary. llms.txt
- Mistral: Voxtral Small 24B 2507 STT by Mistral AI. Voxtral Small 24B 2507 STT is a speech transcription model from Mistral AI. It is suited for transcription, translation, and audio understanding workloads that benefit from its larger model capacity. llms.txt
- Mistral: Voxtral Mini 3B 2507 by Mistral AI. Voxtral Mini 3B 2507 is a speech and audio understanding model from Mistral AI. It is suited for transcription, translation, and compact audio processing workloads. llms.txt
- Fish Audio: Transcribe 1 by fish-audio. Transcribe 1 is a speech-to-text model from Fish Audio. It is suited for audio transcription with automatic language detection and can return timestamped word-level segments when alignment details are requested. llms.txt
- Fish Audio: Transcribe 1 Pro by fish-audio. Transcribe 1 Pro is a speech-to-text model from Fish Audio tuned for interviews, meetings, and podcasts. It labels speakers with inline `speaker` markers, preserves emotion and vocal-event cues such as `[laughter]`, detects language automatically, and can return timestamped word-level segments. llms.txt
Top Audio models
Ranked by tokens processed on OpenRouter over the past week. Free variants and OpenRouter routers are listed as their own entries.
- OpenAI: GPT Audio Mini by OpenAI. A cost-efficient version of GPT Audio. The new snapshot features an upgraded decoder for more natural sounding voices and maintains better voice consistency. Input is priced at $0.60 per million tokens and output is priced at $2.40 per million tokens. llms.txt
- Google: Lyria 3 Pro Preview by Google. Full-length songs are priced at $0.08 per song. Lyria 3 is Google's family of music generation models, available through the Gemini API. With Lyria 3, you can generate high-quality, 48kHz stereo audio from text prompts or from images. These models deliver structural coherence, including vocals, timed lyrics, and full instrumental arrangements. Lyria 3 Pro can generate full-length songs with verses, choruses, bridges. llms.txt
- OpenAI: GPT Audio by OpenAI. The gpt-audio model is OpenAI's first generally available audio model. The new snapshot features an upgraded decoder for more natural sounding voices and maintains better voice consistency. Audio is priced at $32 per million input tokens and $64 per million output tokens. llms.txt
- Google: Lyria 3 Clip Preview by Google. 30 second duration clips are priced at $0.04 per clip. Lyria 3 is Google's family of music generation models, available through the Gemini API. With Lyria 3, you can generate high-quality, 48kHz stereo audio from text prompts or from images. These models deliver structural coherence, including vocals, timed lyrics, and full instrumental arrangements. Lyria 3 Clip can generate short clips, loops, previews. llms.txt
Top Embeddings models
Ranked by tokens processed on OpenRouter over the past week. Free variants and OpenRouter routers are listed as their own entries.
- OpenAI: Text Embedding 3 Small by OpenAI. text-embedding-3-small is OpenAI's improved, more performant version of the ada embedding model. Embeddings are a numerical representation of text that can be used to measure the relatedness between two pieces of text. Embeddings are useful for search, clustering, recommendations, anomaly detection, and classification tasks.
- Qwen: Qwen3 Embedding 8B by Qwen. The Qwen3 Embedding model series is the latest proprietary model of the Qwen family, specifically designed for text embedding and ranking tasks. This series inherits the exceptional multilingual capabilities, long-text understanding, and reasoning skills of its foundational model. The Qwen3 Embedding series represents significant advancements in multiple text embedding and ranking tasks, including text retrieval, code retrieval, text classification, text clustering, and bitext mining.
- OpenAI: Text Embedding 3 Large by OpenAI. text-embedding-3-large is OpenAI's most capable embedding model for both english and non-english tasks. Embeddings are a numerical representation of text that can be used to measure the relatedness between two pieces of text. Embeddings are useful for search, clustering, recommendations, anomaly detection, and classification tasks.
- BAAI: bge-m3 by baai. The bge-m3 embedding model encodes sentences, paragraphs, and long documents into a 1024-dimensional dense vector space, delivering high-quality semantic embeddings optimized for multilingual retrieval, semantic search, and large-context applications.
- Google: Gemini Embedding 2 by Google. Gemini Embedding 2 is Google's first multimodal embedding model. We currently support mapping text and images into a unified vector space for semantic search and retrieval-augmented generation (RAG). It supports input context up to 8,192 tokens and flexible output dimensions from 128 to 3,072 (recommended: 768, 1536, or 3,072). Designed for cross-modal similarity — you can embed a text query and retrieve the most relevant images, or vice versa — making it well-suited for multimodal search, recommendation, and document understanding pipelines.
- Perplexity: Embed V1 0.6B by Perplexity. pplx-embed-v1-0.6B is one of Perplexity's state-of-the-art text embedding models built for real-world, web-scale retrieval. pplx-embed-v1 is optimized for standard dense text retrieval with the 0.6B parameter model targeting lightweight, low-latency embedding generation.
- Qwen: Qwen3 Embedding 4B by Qwen. The Qwen3 Embedding model series is the latest proprietary model of the Qwen family, specifically designed for text embedding and ranking tasks. This series inherits the exceptional multilingual capabilities, long-text understanding, and reasoning skills of its foundational model. The Qwen3 Embedding series represents significant advancements in multiple text embedding and ranking tasks, including text retrieval, code retrieval, text classification, text clustering, and bitext mining.
- Google: Gemini Embedding 001 by Google. gemini-embedding-001 provides a unified cutting edge experience across domains, including science, legal, finance, and coding. This embedding model has consistently held a top spot on the Massive Text Embedding Benchmark (MTEB) Multilingual leaderboard since the experimental launch in March.
- Mistral: Mistral Embed 2312 by Mistral AI. Mistral Embed is a specialized embedding model for text data, optimized for semantic search and RAG applications. Developed by Mistral AI in late 2023, it produces 1024-dimensional vectors that effectively capture semantic relationships in text.
- VoyageAI by MongoDB: voyage-4-lite by voyageai. voyage-4-lite is a lightweight, general-purpose embedding model optimized for low latency and cost. Enabled by Matryoshka learning and quantization-aware training, voyage-4-lite supports embeddings in 2048, 1024, 512, and 256 dimensions, with multiple quantization options. Learn more about voyage-4-lite here: [blog.voyageai.com/2026/01/15/voyage-4](https://blog.voyageai.com/2026/01/15/voyage-4)
- Google: Gemini Embedding 2 Preview by Google. Gemini Embedding 2 Preview is Google's first multimodal embedding model. We currently support mapping text and images into a unified vector space for semantic search and retrieval-augmented generation (RAG). It supports input context up to 8,192 tokens and flexible output dimensions from 128 to 3,072 (recommended: 768, 1536, or 3,072). Designed for cross-modal similarity — you can embed a text query and retrieve the most relevant images, or vice versa — making it well-suited for multimodal search, recommendation, and document understanding pipelines.
- VoyageAI by MongoDB: voyage-4 by voyageai. voyage-4 is a general-purpose (including multilingual) embedding model optimized for retrieval/search and AI applications. voyage-4 supports embeddings in 2048, 1024, 512, and 256 dimensions, with multiple quantization options. Learn more about voyage-4 here: [blog.voyageai.com/2026/01/15/voyage-4](https://blog.voyageai.com/2026/01/15/voyage-4)
- Perplexity: Embed V1 4B by Perplexity. pplx-embed-v1 -4B is one of Perplexity's state-of-the-art text embedding models built for real-world, web-scale retrieval. pplx-embed-v1 is optimized for standard dense text retrieval with the 4B parameter model maximizing retrieval quality.
- OpenAI: Text Embedding Ada 002 by OpenAI. text-embedding-ada-002 is OpenAI's legacy text embedding model.
- NVIDIA: Llama Nemotron Embed VL 1B V2 (free) by Nvidia. The Llama Nemotron Embed VL 1B V2 embedding model is optimized for multimodal question-answering retrieval. The model can embed 'documents' in the form of image, text, or image and text combined. Documents can be retrieved given a user query in text form. The model supports images containing text, tables, charts, and infographics.
- VoyageAI by MongoDB: voyage-4-large by voyageai. voyage-4-large is a state-of-the-art general-purpose and multilingual embedding optimized for retrieval quality. Enabled by Matryoshka learning and quantization-aware training, voyage-4-large supports embeddings in 2048, 1024, 512, and 256 dimensions, with multiple quantization options. Learn more about voyage-4-large here: [blog.voyageai.com/2026/01/15/voyage-4](https://blog.voyageai.com/2026/01/15/voyage-4)
- Intfloat: Multilingual-E5-Large by intfloat. The multilingual-e5-large embedding model encodes sentences, paragraphs, and documents across over 90 languages into a 1024-dimensional dense vector space, delivering robust semantic embeddings optimized for multilingual retrieval, cross-language similarity, and large-scale data search.
- Sentence Transformers: all-MiniLM-L6-v2 by sentence-transformers. The all-MiniLM-L6-v2 embedding model maps sentences and short paragraphs into a 384-dimensional dense vector space, enabling high-quality semantic representations that are ideal for downstream tasks such as information retrieval, clustering, similarity scoring, and text ranking.
- NVIDIA: Nemotron 3 Embed 1B (free) by Nvidia. NVIDIA Nemotron 3 Embed 1B is an open text embedding model from NVIDIA, optimized for high-throughput, low-latency retrieval. It is suited for enterprise search, RAG, code retrieval, and agentic retrieval workflows, retaining more than 95% of the 8B model’s accuracy in a smaller deployment footprint.
- BAAI: bge-base-en-v1.5 by baai. The bge-base-en-v1.5 embedding model converts English sentences and paragraphs into 768-dimensional dense vectors, delivering efficient, high-quality semantic embeddings optimized for retrieval, semantic search, and document-matching workflows. This version (v1.5) features improved similarity-score distribution and stronger retrieval performance out of the box.
- VoyageAI by MongoDB: voyage-multimodal-3.5 by voyageai. voyage-multimodal-3.5 is a state-of-the-art multimodal embedding model capable of vectorizing not only text, images, and video individually, but also content that interleaves all three modalities. It delivers excellent performance for mixed-modality searches involving text and visual content such as PDF screenshots, figures, tables, videos, and more. Enabled by Matryoshka learning and quantization-aware training, voyage-multimodal-3.5 supports embeddings in 2048, 1024, 512, and 256 dimensions, with multiple quantization options. Learn more about voyage-multimodal-3.5 here: [blog.voyageai.com/2026/01/15/voyage-multimodal-3-5](https://blog.voyageai.com/2026/01/15/voyage-multimodal-3-5)
- Mistral: Codestral Embed 2505 by Mistral AI. Mistral Codestral Embed is specially designed for code, perfect for embedding code databases, repositories, and powering coding assistants with state-of-the-art retrieval.
- Sentence Transformers: all-mpnet-base-v2 by sentence-transformers. The all-mpnet-base-v2 embedding model encodes sentences and short paragraphs into a 768-dimensional dense vector space, providing high-fidelity semantic embeddings well suited for tasks like information retrieval, clustering, similarity scoring, and text ranking.
- BAAI: bge-large-en-v1.5 by baai. The bge-large-en-v1.5 embedding model maps English sentences, paragraphs, and documents into a 1024-dimensional dense vector space, delivering high-fidelity semantic embeddings optimized for semantic search, document retrieval, and downstream NLP tasks in English.
- VoyageAI by MongoDB: voyage-code-4 by voyageai. voyage-code-4 is a code embedding model from Voyage AI, a MongoDB company. It is designed for coding agents and code retrieval, with Matryoshka embeddings at 2048, 1024, 512, and 256 dimensions and multiple quantization options. Learn more about voyage-code-4 here: [blog.voyageai.com/2026/08/13/voyage-code-4](https://blog.voyageai.com/2026/08/13/voyage-code-4)
- Intfloat: E5-Large-v2 by intfloat. The e5-large-v2 embedding model maps English sentences, paragraphs, and documents into a 1024-dimensional dense vector space, delivering high-accuracy semantic embeddings optimized for retrieval, semantic search, reranking, and similarity-scoring tasks.
- Thenlper: GTE-Base by thenlper. The gte-base embedding model encodes English sentences and paragraphs into a 768-dimensional dense vector space, delivering efficient and effective semantic embeddings optimized for textual similarity, semantic search, and clustering applications.
- Sentence Transformers: all-MiniLM-L12-v2 by sentence-transformers. The all-MiniLM-L12-v2 embedding model maps sentences and short paragraphs into a 384-dimensional dense vector space, producing efficient and high-quality semantic embeddings optimized for tasks such as semantic search, clustering, and similarity-scoring.
- LiquidAI: LFM2.5-Embedding-350M (free) by Liquid. LFM2.5-Embedding-350M is a text embedding model from Liquid AI. It produces 1,024-dimensional embeddings for retrieval and semantic search. Successful OpenRouter requests and embeddings may be retained and used to train Liquid models.
- Google: Gemini Embedding 2 (batch) by Google. Gemini Embedding 2 is Google's first multimodal embedding model. We currently support mapping text and images into a unified vector space for semantic search and retrieval-augmented generation (RAG). It supports input context up to 8,192 tokens and flexible output dimensions from 128 to 3,072 (recommended: 768, 1536, or 3,072). Designed for cross-modal similarity — you can embed a text query and retrieve the most relevant images, or vice versa — making it well-suited for multimodal search, recommendation, and document understanding pipelines.
- Thenlper: GTE-Large by thenlper. The gte-large embedding model converts English sentences, paragraphs and moderate-length documents into a 1024-dimensional dense vector space, delivering high-quality semantic embeddings optimized for information retrieval, semantic textual similarity, reranking and clustering tasks. Trained via multi-stage contrastive learning on a large domain-diverse relevance corpus, it offers excellent performance across general-purpose embedding use-cases.
- Intfloat: E5-Base-v2 by intfloat. The e5-base-v2 embedding model encodes English sentences and paragraphs into a 768-dimensional dense vector space, producing efficient and high-quality semantic embeddings optimized for tasks such as semantic search, similarity scoring, retrieval and clustering.
- Sentence Transformers: multi-qa-mpnet-base-dot-v1 by sentence-transformers. The multi-qa-mpnet-base-dot-v1 embedding model transforms sentences and short paragraphs into a 768-dimensional dense vector space, generating high-quality semantic embeddings optimized for question-and-answer retrieval, semantic search, and similarity-scoring across diverse content.
- Sentence Transformers: paraphrase-MiniLM-L6-v2 by sentence-transformers. The paraphrase-MiniLM-L6-v2 embedding model converts sentences and short paragraphs into a 384-dimensional dense vector space, producing high-quality semantic embeddings optimized for paraphrase detection, semantic similarity scoring, clustering, and lightweight retrieval tasks.
- OpenAI: Text Embedding Ada 002 (batch) by OpenAI. text-embedding-ada-002 is OpenAI's legacy text embedding model.
- OpenAI: Text Embedding 3 Large (batch) by OpenAI. text-embedding-3-large is OpenAI's most capable embedding model for both english and non-english tasks. Embeddings are a numerical representation of text that can be used to measure the relatedness between two pieces of text. Embeddings are useful for search, clustering, recommendations, anomaly detection, and classification tasks.
- OpenAI: Text Embedding 3 Small (batch) by OpenAI. text-embedding-3-small is OpenAI's improved, more performant version of the ada embedding model. Embeddings are a numerical representation of text that can be used to measure the relatedness between two pieces of text. Embeddings are useful for search, clustering, recommendations, anomaly detection, and classification tasks.
Top Rerank models
Ranked by tokens processed on OpenRouter over the past week. Free variants and OpenRouter routers are listed as their own entries.
- VoyageAI by MongoDB: rerank-2.5-lite by voyageai. rerank-2.5-lite is a reranker optimized for both latency and quality, delivering a 7.16% improvement in retrieval accuracy over Cohere Rerank v3.5 across 93 datasets. It also outperformed Cohere Rerank v3.5 by 10.36% on the Massive Instructed Retrieval Benchmark (MAIR). The model supports a combined context length of 32K tokens per query–document pair, including up to 8K tokens for the query, enabling more accurate retrieval over longer documents. Additionally, rerank-2.5-lite supports instruction following, allowing users to guide relevance scoring through natural language prompts. Learn more about rerank-2.5-lite here: [blog.voyageai.com/2025/08/11/rerank-2-5](https://blog.voyageai.com/2025/08/11/rerank-2-5)
- VoyageAI by MongoDB: rerank-2.5 by voyageai. rerank-2.5 is a cutting-edge reranker optimized for quality, delivering a 7.94% improvement in retrieval accuracy over Cohere Rerank v3.5 across 93 datasets. It also outperformed Cohere Rerank v3.5 by 12.70% on the Massive Instructed Retrieval Benchmark (MAIR). The model supports a combined context length of 32K tokens per query–document pair, including up to 8K tokens for the query, enabling more accurate retrieval over longer documents. Additionally, rerank-2.5 supports instruction following, allowing users to guide relevance scoring through natural language prompts. Learn more about rerank-2.5 here:[ https://blog.voyageai.com/2025/08/11/rerank-2-5](https://blog.voyageai.com/2025/08/11/rerank-2-5)
- Qwen3 Reranker 8B by Qwen. Qwen3 Reranker 8B is a text reranking model from Alibaba Cloud built on the Qwen3 architecture. It evaluates query-document pairs to produce relevance scores for use in retrieval and RAG pipelines. Supports 100+ languages and programming languages, with instruction-aware reranking that allows customizing scoring criteria per task. Offers strong performance on multilingual benchmarks including MTEB, CMTEB, and MMTEB.
- NVIDIA: Llama Nemotron Rerank VL 1B V2 (free) by Nvidia. Llama Nemotron Rerank VL 1B V2 is a 1.7B multimodal reranking model from NVIDIA. It evaluates the relevance of document images and text against user queries, designed for vision RAG pipelines handling charts, tables, infographics, and mixed-media documents. Functions as a cross-encoder that accepts text queries paired with image, text, or combined document inputs, delivering approximately 6-7% recall improvements over embedding-only baselines on visual document retrieval benchmarks.
- Cohere: Rerank 4 Pro by Cohere. Cohere's AI search foundation model for enhancing the relevance of information surfaced within search and RAG systems. Features a 32K context window, multilingual support across 100+ languages, no data pre-processing required, and state of the art performance with low latency.
- Cohere: Rerank 4 Fast by Cohere. Cohere's AI search foundation model for enhancing the relevance of information surfaced within search and RAG systems. Features a 32K context window, multilingual support across 100+ languages, no data pre-processing required, and high performance with lowest latency.
- Cohere: Rerank v3.5 by Cohere. Rerank v3.5 is designed to reorder search results for improved relevance. It supports multi-aspect and semi-structured data reranking over 100+ languages. Ideal for refining results from semantic or keyword search pipelines.
Top Decisions models
Ranked by tokens processed on OpenRouter over the past week. Free variants and OpenRouter routers are listed as their own entries.
- TypeSafe: Jev 1.13 by Typesafe. Jev is a structured decision model from TypeSafe, and the first of its System One models. System One models make fast, structured decisions for software, returning a typed choice rather than free-form text. It is suited for routing, classification, and other decision points inside an application where a fast, predictable answer matters more than generated prose. Learn more in TypeSafe's docs: https://docs.typesafe.ai/concepts/system-one
- TypeSafe: Jev Latest by Typesafe. This model always redirects to the latest model in the Jev family.
Top Text models
Ranked by tokens processed on OpenRouter over the past week. Free variants and OpenRouter routers are listed as their own entries.
- DeepSeek: DeepSeek V4.1 Flash by DeepSeek. DeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, and the first built on the company's Causal Encoder-Decoder (CED) architecture. It activates 8B parameters on input and 16B on output from a 552B-parameter backbone, an asymmetric split that keeps per-token compute low relative to the model's total size. Image understanding is native to the architecture, with visual and text embeddings trained jointly from the start of pre-training rather than added afterward as in the earlier experimental [V4 Flash Vision Exp](https://openrouter.ai/deepseek/deepseek-v4-flash-vision-exp). It is suited for coding, terminal, and computer-use agents, along with long-horizon tasks that must run to completion across many steps and long-context analysis. Compressed KV caching cuts cache memory to roughly a quarter of the previous Flash generation, significantly reducing costs on agentic workloads. DeepSeek positions it as the cost-efficient tier of the V4.1 family and reports that it exceeds [V4 Pro](https://openrouter.ai/deepseek/deepseek-v4-pro-0813) on performance, speed, and task completion time. llms.txt
- Z.ai: GLM 5.3 Flash by Z.ai. GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while reducing compute overhead. llms.txt
- Tencent: Hy4 preview by tencent. Tencent: Hy4 preview is a mixture-of-experts model from Tencent, with 49B active parameters out of 770B total. It is designed for coding agents, complex tool-use workflows, and productivity tasks that require planning, context continuity, and sustained multi-step execution. llms.txt
- OpenAI: GPT-5.6 Luna by OpenAI. GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as chat, classification, and lightweight agentic workflows, providing capable reasoning for its price tier. llms.txt
- DeepSeek: DeepSeek V4 Flash 0731 by DeepSeek. DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows. This is the GA release of DeepSeek V4 Flash. llms.txt
- Xiaomi: MiMo-V2.5 by Xiaomi. MiMo-V2.5 is a native omnimodal model by Xiaomi. It delivers Pro-level agentic performance at roughly half the inference cost, while surpassing MiMo-V2-Omni in multimodal perception across image and video understanding tasks. Its 1M context window supports complete documents, extended conversations, and complex task contexts in a single pass, making it ideal for integration with agent frameworks where strong reasoning, rich perception, and cost efficiency all matter. llms.txt
- NVIDIA: Nemotron 3 Ultra (free) by Nvidia. NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it supports text input and output with a context window of up to 1M tokens. It is suited for long-running agentic workflows, including agent orchestration, coding agents, deep research, and complex enterprise tasks. It is particularly strong at multi-step reasoning and planning, with high-throughput inference designed for high-volume agent pipelines. It is part of the NVIDIA Nemotron family of open models for agentic AI. llms.txt
- Tencent: Hy3 by tencent. Hy3 is a 295B-parameter Mixture-of-Experts model from Tencent (21B active, 192 experts with top-8 routing) built for reasoning, agentic workflows, and real-world production use. It supports a configurable reasoning effort: a direct no-think mode by default, plus low and high chain-of-thought modes for complex math, coding, and multi-step problems. With a 256K context window, Hy3 targets long-horizon tasks, including improved coreference resolution, multi-turn constraint tracking, and stable tool-calling that generalizes across agent scaffoldings. Tencent positions it as a reliable, cost-effective option across coding, document processing, financial analysis, game development, and frontend design, with a strong emphasis on grounded, anti-hallucination behavior that answers when grounded and flags when evidence is missing rather than fabricating. llms.txt
- DeepSeek: DeepSeek V4 Flash 0423 by DeepSeek. DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supporting a 1M-token context window. It is designed for fast inference and high-throughput workloads, while maintaining strong reasoning and coding performance. The model includes hybrid attention for efficient long-context processing. Reasoning efforts `high` and `xhigh` are supported; `xhigh` maps to max reasoning. It is well suited for applications such as coding assistants, chat systems, and agent workflows where responsiveness and cost efficiency are important. llms.txt
- Z.ai: GLM 5.3 by Z.ai. GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text input and output with a 1M-token context window, and improves on GLM-5.2 in coding and in the balance between performance and token efficiency. Reasoning is always on and cannot be disabled. Reasoning efforts `low`, `high`, and `max` are supported; `max` is the default. llms.txt
- Space Bunny Alpha by Stealth. Space Bunny Alpha is an anonymous large model with blazing-fast inference, strong coding capabilities and native multimodal input support. It delivers adjustable reasoning effort, and a 1M-token context window. Space Bunny Alpha is a stealth model. It is developed and operated by a third-party provider who has chosen to remain anonymous during this preview. OpenRouter routes requests to it and is not its developer, owner, or provider. Prompts and completions may be retained by the provider but are not used for training; all other use is governed by the [Stealth Model Terms](https://openrouter.ai/terms/stealth). llms.txt
- Google: Gemini 3.8 Flash by Google. Gemini 3.8 Flash is Google's most intelligent Flash model with significant gains from 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning. llms.txt
- Meta: Muse Spark 1.3 Contributor by Meta. Muse Spark 1.3 Contributor is the cost-efficient contributor tier of Meta’s multimodal reasoning model for experimentation, learning, and early-stage agentic, multi-agent, and coding workflows. It is designed to track information across extended tasks, work through conflicting inputs, and request clarification or confirmation when needed. Prompts and outputs may be used to improve Meta’s products. llms.txt
- OpenAI: GPT-5.6 Sol by OpenAI. GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series. It is suited for complex reasoning, coding, and agentic workflows, and is particularly strong at command-line and multi-step coding tasks and long-horizon problem solving. llms.txt
- Upstage: Solar Pro 4 by upstage. Solar Pro 4 is Upstage's cost-efficient large language model, featuring a 524K context window. It is built for long-horizon tasks and agentic workflows, with strong capabilities in office productivity, document-intensive work, and coding. When using BYOK with ZDR enforcement, the Console API key must belong to a ZDR-enabled Upstage organization. Contact Upstage to enable ZDR on your org. llms.txt
- OpenAI: GPT-6 Astra by OpenAI. GPT-6 Astra is OpenAI's flagship model for demanding end-to-end work. It is suited for advanced analysis, software engineering, deep research, scientific work, and document creation, with particular strengths in long-horizon agentic tasks that involve computer and browser use. llms.txt
- Z.ai: GLM 5.2 by Z.ai. GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workflows, project-level software engineering, and complex multi-step automation. Reasoning efforts `high` and `xhigh` are supported; `xhigh` maps to max reasoning. It is particularly strong at coding and tool use across long-running tasks, able to maintain engineering context and follow standards consistently through a full development workflow, from requirements to multi-platform deployment, in a single task. llms.txt
- Xiaomi: MiMo-V2.6-Flash by Xiaomi. MiMo-V2.6-Flash is an open-source foundation model developed by Xiaomi. Built on a Mixture-of-Experts architecture with 309B total parameters and 15B activated per token, it employs a hybrid attention mechanism for greater computational efficiency. The model features a 1M-token context window and native multimodal capabilities. Optimized for agentic workflows, it delivers strong performance across coding, visual, general, and research scenarios, excelling at complex, long-horizon tasks with robust generalization across a diverse range of agent harnesses. llms.txt
- Anthropic: Claude Sonnet 5 by Anthropic. Sonnet 5 is Anthropic's most capable Sonnet-class model, with frontier performance across coding, agents, and professional work. It supports adaptive thinking with selectable reasoning effort levels (low, medium, high, max, and x-high), a 1M-token context window, and text, image, and file inputs. Sonnet 5 uses an updated tokenizer and includes real-time cyber safeguards that block certain high-risk dual-use activities. llms.txt
- MiniMax: MiniMax M3 by MiniMax. MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding, and tool use. It is built on MiniMax Sparse Attention (MSA), which replaces full attention with KV-block selection to cut per-token compute at long context — roughly 1/20 the cost of the previous generation at 1M tokens, with substantially faster prefill and decode while retaining quality across most tasks. Trained as a native multimodal model on interleaved data and tuned for multi-turn, production-like collaboration via an interactive user-simulator framework, the model is oriented toward sustained, multi-step tasks rather than single-turn execution. llms.txt
- MoonshotAI: Kimi K3 by moonshotai. Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic workflows, and is particularly strong at navigating large repositories, using tools, debugging, and iterating against images, logs, tests, and runtime feedback. Its architecture uses KDA and Attention Residuals for computational efficiency. llms.txt
- inclusionAI: Ling 3.0 Flash Fin (free) by inclusionai. Ling 3.0 Flash Fin is a finance-focused mixture-of-experts model from InclusionAI, built on Ling 3.0 Flash with 5.1B active parameters out of 124B total. It is designed for real-world investment workflows that require complex multi-step tasks and long-horizon planning and execution, while retaining general capabilities in reasoning, coding, and mathematics. llms.txt
- Poolside: Laguna S 2.1 (free) by poolside. Laguna S 2.1 is the latest coding agent model from [Poolside](<https://poolside.ai/>). Laguna S 2.1 is a 118B total parameter model with 8B active parameters, scoring 70.2% on Terminal-Bench 2.1 and 40.4% on DeepSWE, making it one of the strongest coding models in its category. Open-weight under the OpenMDW-1.1 license. Laguna S 2.1 is designed for software engineering and agentic coding use cases, and you are responsible for confirming that it is appropriate for your intended application. Laguna S 2.1 is subject to the [OpenMDW-1.1 License](<https://openmdw.ai/license/1-1/>), and should be used consistently with Poolside's [Acceptable Use Policy](<https://poolside.ai/legal/acceptable-use-policy>). We advise against circumventing Laguna S 2.1 safety guardrails without implementing substantially equivalent mitigations appropriate for your use case. Please report security vulnerabilities or safety concerns to [[email protected]](<mailto:[email protected]>). If you are using Laguna S 2.1 for free, we may use your inputs and outputs to train and improve our models. llms.txt
- Anthropic: Claude Opus 5 by Anthropic. Claude Opus 5 is Anthropic’s flagship model for demanding reasoning, coding, and long-horizon agentic work. It is particularly strong at end-to-end software tasks, code review and bug finding, visual analysis of charts and documents, complex office deliverables, and coordinating parallel subagents. The model maintains strong instruction following and tool use across extended tasks, while remaining effective at lower effort settings for workloads that prioritize latency and token efficiency. llms.txt
- DeepSeek: DeepSeek V4 Pro 0813 by DeepSeek. DeepSeek V4 Pro 0813 is a large-scale mixture-of-experts model from DeepSeek. This is the GA release of DeepSeek V4 Pro. llms.txt
- OpenAI: GPT-6 Luna by OpenAI. GPT-6 Luna is the fast, cost-efficient model in OpenAI's GPT-6 series, positioned below GPT-6 Sol. It is suited for high-volume and latency-sensitive workloads such as chat, classification, and lightweight agentic tasks, and at higher reasoning effort it can take on complex software engineering and computer-use tasks that previously called for a Sol-tier model. It shares the GPT-6 family's gains in factual reliability and its clearer, more concise communication style. llms.txt
- DeepSeek: DeepSeek V4 Pro 0423 by DeepSeek. DeepSeek V4 Pro is a large-scale Mixture-of-Experts model from DeepSeek with 1.6T total parameters and 49B activated parameters, supporting a 1M-token context window. It is designed for advanced reasoning, coding, and long-horizon agent workflows, with strong performance across knowledge, math, and software engineering benchmarks. Built on the same architecture as DeepSeek V4 Flash, it introduces a hybrid attention system for efficient long-context processing. Reasoning efforts high and xhigh are supported; xhigh maps to max reasoning. It is well suited for complex workloads such as full-codebase analysis, multi-step automation, and large-scale information synthesis, where both capability and efficiency are critical llms.txt
- Google: Gemini 3 Flash Preview by Google. Gemini 3 Flash Preview is a high speed, high value thinking model designed for agentic workflows, multi turn chat, and coding assistance. It delivers near Pro level reasoning and tool use performance with substantially lower latency than larger Gemini variants, making it well suited for interactive development, long running agent loops, and collaborative coding tasks. Compared to Gemini 2.5 Flash, it provides broad quality improvements across reasoning, multimodal understanding, and reliability. The model supports a 1M token context window and multimodal inputs including text, images, audio, video, and PDFs, with text output. It includes configurable reasoning via thinking levels (minimal, low, medium, high), structured output, tool use, and automatic context caching. Gemini 3 Flash Preview is optimized for users who want strong reasoning and agentic behavior without the cost or latency of full scale frontier models. llms.txt
- DeepSeek: DeepSeek V4 Flash Latest by DeepSeek. This model always redirects to the latest model in the DeepSeek V4 Flash family. llms.txt
- Google: Gemini 2.5 Flash Lite by Google. Gemini 2.5 Flash-Lite is a lightweight reasoning model in the Gemini 2.5 family, optimized for ultra-low latency and cost efficiency. It offers improved throughput, faster token generation, and better performance across common benchmarks compared to earlier Flash models. By default, "thinking" (i.e. multi-pass reasoning) is disabled to prioritize speed, but developers can enable it via the [Reasoning API parameter](https://openrouter.ai/docs/use-cases/reasoning-tokens) to selectively trade off cost for intelligence. llms.txt
- Dots Studio: Dots3-Note Preview (free) by Dots Studio. Dots3-Note Preview is an open-weight mixture-of-experts model from Dots Studio, with 16B active parameters out of 280B total. It is the lightest model in the Dots 3 family and is suited for reasoning, coding, multimodal understanding, long-context processing, and multi-step agent workflows. llms.txt
- OpenAI: GPT-5.6 Terra by OpenAI. GPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier. It is suited for everyday coding, reasoning, and agentic tasks where capability and cost need to be balanced, offering strong performance at roughly half the cost of Sol. llms.txt
- OpenAI: gpt-oss-120b by OpenAI. gpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic, and general-purpose production use cases. It activates 5.1B parameters per forward pass and is optimized to run on a single H100 GPU with native MXFP4 quantization. The model supports configurable reasoning depth, full chain-of-thought access, and native tool use, including function calling, browsing, and structured output generation. llms.txt
- DeepSeek: DeepSeek Flash Latest by DeepSeek. This model always redirects to the latest model in the DeepSeek Flash family. llms.txt
- NVIDIA: Nemotron 3.5 Lightning (free) by Nvidia. NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total. It is suited for high-throughput agentic workloads and specialized tasks that benefit from domain-specific customization. llms.txt
- Google: Gemini 3.7 Flash by Google. Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for tasks that require responsive performance and reliable multi-step problem solving. llms.txt
- Google: Gemini 3.1 Flash Lite by Google. Gemini 3.1 Flash Lite is Google’s GA high-efficiency multimodal model optimized for low-latency, high-volume workloads. It supports text, image, video, audio, and PDF inputs, and is designed for lightweight agentic workflows, simple data extraction, and applications where responsiveness and API cost are the primary constraints. Supports full thinking levels (minimal, low, medium, high) for fine-grained cost/performance trade-offs. Priced at half the cost of Gemini 3 Flash. llms.txt
- Anthropic: Claude Fable 5.1 by Anthropic. Claude Fable 5.1 improves on Claude Fable 5 across the board, with the biggest gains in agentic coding, long-running agentic workflows, and knowledge work: long code refactors, front-end and visual code generation, and finance and analysis tasks in particular. It also tends to be more concise than Fable 5 in its plans and summaries. We recommend testing it as a direct upgrade wherever you use Fable 5 today, and alongside Opus 5 on reasoning-heavy tasks. llms.txt
- Nex AGI: Nex-N2.5-Pro (free) by nex-agi. Nex-N2.5 is an agentic model built to turn goals into working, verified outcomes. Its core strength is agentic coding within a visual feedback loop: it can explore codebases, implement multi-file changes, run commands, launch applications, interact with browser and desktop interfaces, and test software from the user's perspective. When observed behavior does not match the intended result, Nex-N2.5 can diagnose the issue, revise its implementation, and test again. This makes it especially effective for autonomous software engineering, GUI-based QA, computer-use automation, deep research, and scientific workflows where success must be demonstrated in the environment—not merely inferred from generated code. llms.txt
- Google: Gemini 2.5 Flash by Google. Gemini 2.5 Flash is Google's state-of-the-art workhorse model, specifically designed for advanced reasoning, coding, mathematics, and scientific tasks. It includes built-in "thinking" capabilities, enabling it to provide responses with greater accuracy and nuanced context handling. Additionally, Gemini 2.5 Flash is configurable through the "max tokens for reasoning" parameter, as described in the documentation (https://openrouter.ai/docs/use-cases/reasoning-tokens#max-tokens-for-reasoning). llms.txt
- Qwen: Qwen3.8 Flash by Qwen. Qwen3.8 Flash is a multimodal reasoning model from Alibaba. It is suited for coding assistance, agentic workflows, visual understanding, document and codebase analysis, desktop interaction, chart analysis, and long-video analysis. llms.txt
- Xiaomi: MiMo-V2.6-Pro by Xiaomi. MiMo-V2.6-Pro is the flagship foundation model developed by Xiaomi. Built at a scale of over 1T parameters, it is designed to push the ceiling of capability for the most demanding workloads. The model features a 1M-token context window and native multimodal capabilities. Optimized for agentic workflows, it delivers top-tier performance across coding, visual, general, and research scenarios, excelling at complex, long-horizon tasks with robust generalization across a diverse range of agent harnesses. llms.txt
- Qwen: Qwen3.7 Flash by Qwen. Qwen3.7 Flash is a vision-language reasoning model from Alibaba. It is suited for multimodal agents, visual coding, search, and computer interaction, with strengths in object recognition, spatial understanding, and real-world visual perception. llms.txt
- Google: Gemma 4 31B by Google. Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function calling, and multilingual support across 140+ languages. Strong on coding, reasoning, and document understanding tasks. Apache 2.0 license. llms.txt
- Anthropic: Claude Opus 5.5 by Anthropic. Claude Opus 5.5 is Anthropic's flagship model for demanding reasoning, coding, and long-horizon agentic work, succeeding Claude Opus 5. It is particularly strong at multi-step changes in large codebases, code review and bug finding, financial and scientific analysis, and reading dense charts, diagrams, and screenshots, and it is more careful than its predecessor about only stating figures and citing sources it can back up. The model completes comparable tasks in fewer steps and with fewer tokens than Opus 5, and reports on its work in plainer language, with clear updates on what it did, what it found, and what it needs from the user. Thinking is always adaptive, so effort is the main lever for trading off depth, latency, and cost, and lower effort settings remain effective for latency-sensitive workloads. llms.txt
- DeepSeek: DeepSeek V3.2 by DeepSeek. DeepSeek-V3.2 is a large language model designed to harmonize high computational efficiency with strong reasoning and agentic tool-use performance. It introduces DeepSeek Sparse Attention (DSA), a fine-grained sparse attention mechanism that reduces training and inference cost while preserving quality in long-context scenarios. A scalable reinforcement learning post-training framework further improves reasoning, with reported performance in the GPT-5 class, and the model has demonstrated gold-medal results on the 2025 IMO and IOI. V3.2 also uses a large-scale agentic task synthesis pipeline to better integrate reasoning into tool-use settings, boosting compliance and generalization in interactive environments. Users can control the reasoning behaviour with the `reasoning` `enabled` boolean. [Learn more in our docs](https://openrouter.ai/docs/use-cases/reasoning-tokens#enable-reasoning-with-default-config) llms.txt
- Anthropic: Claude Opus 4.8 by Anthropic. Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family. It supports text, image, and file inputs with text output, with reasoning support and a 1M-token context window. It is suited for highly autonomous agents, long-horizon agentic work, knowledge work, and memory-driven tasks where coherence over extended sessions matters. It is particularly strong on multi-step reasoning, complex coding, and end-to-end project orchestration - large codebases, multi-stage debugging, and long-running asynchronous agent pipelines. Beyond coding, it handles knowledge work such as drafting documents, building presentations, and analyzing data, maintaining quality across very long outputs. llms.txt
- Qwen: Qwen3.8 Max (0902) by Qwen. Qwen3.8 Max 0902 is an updated snapshot of Qwen3.8 Max from Alibaba's Qwen team. It is a 2.4-trillion-parameter mixture-of-experts model that accepts text, image, and video input and returns text, with a 1M-token context window and reasoning enabled by default. This snapshot is post-trained for coding and agentic work, including multi-step software projects, multi-tool orchestration, and long-horizon task execution. It also targets chart reasoning, document parsing, and multimodal understanding over long documents and extended video. Tool calling, structured outputs, and configurable reasoning effort are supported. llms.txt
- Meta: Muse Spark 1.3 by Meta. Muse Spark 1.3 is a multimodal reasoning model from Meta for long-running agentic, multi-agent, and coding workflows. It is designed to keep track of information across extended tasks, work through conflicting inputs, and request clarification or confirmation when needed, with an emphasis on concise execution. llms.txt
- Qwen: Qwen3.8 27B by Qwen. Qwen3.8 27B is an open-weight dense vision-language model from Qwen. It is suited for coding, professional workflows, research, multimodal interaction, and long-running agent tasks, with flexible thinking that can be enabled or disabled. llms.txt
- Anthropic: Claude Sonnet 4.6 by Anthropic. Sonnet 4.6 is Anthropic's most capable Sonnet-class model yet, with frontier performance across coding, agents, and professional work. It excels at iterative development, complex codebase navigation, end-to-end project management with memory, polished document creation, and confident computer use for web QA and workflow automation. llms.txt
- Thinking Machines: Inkling (free) by thinkingmachines. Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems, retrieval-augmented generation, instruction following, and multilingual conversational applications. Its native image and audio understanding supports multimodal analysis alongside text. llms.txt
- NVIDIA: Nemotron 3 Super (free) by Nvidia. NVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model, activating just 12B parameters for maximum compute efficiency and accuracy in complex multi-agent applications. Built on a hybrid Mamba-Transformer Mixture-of-Experts architecture with multi-token prediction (MTP), it delivers over 50% higher token generation compared to leading open models. The model features a 1M token context window for long-term agent coherence, cross-document reasoning, and multi-step task planning. Latent MoE enables calling 4 experts for the inference cost of only one, improving intelligence and generalization. Multi-environment RL training across 10+ environments delivers leading accuracy on benchmarks including AIME 2025, TerminalBench, and SWE-Bench Verified. Fully open with weights, datasets, and recipes under the NVIDIA Open License, Nemotron 3 Super allows easy customization and secure deployment anywhere — from workstation to cloud. llms.txt
- Google: Gemma 4 26B A4B by Google. Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind. Despite 25.2B total parameters, only 3.8B activate per token during inference — delivering near-31B quality at a fraction of the compute cost. Supports multimodal input including text, images, and video (up to 60s at 1fps). Features a 256K token context window, native function calling, configurable thinking/reasoning mode, and structured output support. Released under Apache 2.0. llms.txt
- SpaceXAI: Grok 4.6 by SpaceXAI. Grok 4.6 is a model from SpaceXAI with frontier performance on coding, knowledge work, and STEM. It is succeeded by [Grok 4.7](/x-ai/grok-4.7). llms.txt
- inclusionAI: Ling 3.0 Flash Sante (free) by inclusionai. Ling 3.0 Flash Sante is a health and medicine-focused mixture-of-experts model from InclusionAI, built on Ling 3.0 Flash with 5.1B active parameters out of 124B total. It is designed for medical knowledge reasoning, clinical safety, evidence-based retrieval, and long-horizon medical tasks, while retaining general capabilities in reasoning, coding, and agentic tasks. llms.txt
- OpenAI: GPT-6 Sol by OpenAI. GPT-6 Sol is the cost-efficient high-end model in OpenAI's GPT-6 series, positioned below the flagship GPT-6 Astra and above the fast GPT-6 Luna tier. It is suited for demanding professional work, agentic coding, business workflow automation, and computer use, and is particularly strong at long-horizon software engineering tasks in real codebases. It approaches Astra-level factual reliability at a much lower cost and shares Astra's clearer, more concise communication style in technical and coding conversations. llms.txt
- OpenAI: GPT-5.6 Luna Pro by OpenAI. GPT-5.6 Luna Pro is the same underlying model as [GPT-5.6 Luna](https://openrouter.ai/openai/gpt-5.6-luna), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode llms.txt
- Z.ai: GLM Flash Latest by Z.ai. This model always redirects to the latest model in the GLM Flash family. llms.txt
- Xiaomi: MiMo-V2.5-Pro by Xiaomi. MiMo-V2.5-Pro is Xiaomi’s flagship model, delivering strong performance in general agentic capabilities, complex software engineering, and long-horizon tasks, with top rankings on benchmarks such as ClawEval, GDPVal, and SWE-bench Pro. It can independently and autonomously complete professional tasks that would take human experts days or weeks, involving more than a thousand tool calls. Its context length of up to 1M makes it well suited for integration with a wide range of agent frameworks. llms.txt
- Mistral: Mistral Nemo by Mistral AI. A 12B parameter model with a 128k token context length built by Mistral in collaboration with NVIDIA. The model is multilingual, supporting English, French, German, Spanish, Italian, Portuguese, Chinese, Japanese, Korean, Arabic, and Hindi. It supports function calling and is released under the Apache 2.0 license. llms.txt
- Google: Gemini 3.1 Pro Preview by Google. Gemini 3.1 Pro Preview is Google’s frontier reasoning model, delivering enhanced software engineering performance, improved agentic reliability, and more efficient token usage across complex workflows. Building on the multimodal foundation of the Gemini 3 series, it combines high-precision reasoning across text, image, video, audio, and code with a 1M-token context window. Reasoning Details must be preserved when using multi-turn tool calling, see our docs here: https://openrouter.ai/docs/use-cases/reasoning-tokens#preserving-reasoning. The 3.1 update introduces measurable gains in SWE benchmarks and real-world coding environments, along with stronger autonomous task execution in structured domains such as finance and spreadsheet-based workflows. Designed for advanced development and agentic systems, Gemini 3.1 Pro Preview improves long-horizon stability and tool orchestration while increasing token efficiency. It introduces a new medium thinking level to better balance cost, speed, and performance. The model excels in agentic coding, structured planning, multimodal analysis, and workflow automation, making it well-suited for autonomous agents, financial modeling, spreadsheet automation, and high-context enterprise tasks. llms.txt
- Anthropic: Claude Haiku 4.5 by Anthropic. Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latency of larger Claude models. Matching Claude Sonnet 4’s performance across reasoning, coding, and computer-use tasks, Haiku 4.5 brings frontier-level capability to real-time and high-volume applications. It introduces extended thinking to the Haiku line; enabling controllable reasoning depth, summarized or interleaved thought output, and tool-assisted workflows with full support for coding, bash, web search, and computer-use tools. Scoring >73% on SWE-bench Verified, Haiku 4.5 ranks among the world’s best coding models while maintaining exceptional responsiveness for sub-agents, parallelized execution, and scaled deployment. llms.txt
- Google: Gemini 3.5 Flash Lite by Google. Gemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities. It is suited for subagents that execute focused tasks within complex, multi-agent workflows. llms.txt
- OpenAI: gpt-oss-20b by OpenAI. gpt-oss-20b is an open-weight 21B parameter model released by OpenAI under the Apache 2.0 license. It uses a Mixture-of-Experts (MoE) architecture with 3.6B active parameters per forward pass, optimized for lower-latency inference and deployability on consumer or single-GPU hardware. The model is trained in OpenAI’s Harmony response format and supports reasoning level configuration, fine-tuning, and agentic capabilities including function calling, tool use, and structured outputs. llms.txt
- OpenAI: GPT-4o-mini by OpenAI. GPT-4o mini is OpenAI's newest model after [GPT-4 Omni](/models/openai/gpt-4o), supporting both text and image inputs with text outputs. As their most advanced small model, it is many multiples more affordable than other recent frontier models, and more than 60% cheaper than [GPT-3.5 Turbo](/models/openai/gpt-3.5-turbo). It maintains SOTA intelligence, while being significantly more cost-effective. GPT-4o mini achieves an 82% score on MMLU and presently ranks higher than GPT-4 on chat preferences [common leaderboards](https://arena.lmsys.org/). Check out the [launch announcement](https://openai.com/index/gpt-4o-mini-advancing-cost-efficient-intelligence/) to learn more. #multimodal llms.txt
- DeepSeek: DeepSeek V4 Flash Vision Exp by DeepSeek. DeepSeek V4 Flash Vision Exp is an experimental vision-enabled version of [DeepSeek V4 Flash 0731](https://openrouter.ai/deepseek/deepseek-v4-flash-0731) from DeepSeek, adding image understanding while matching the base model on text capabilities including agents, reasoning, and world knowledge. It is a sparse mixture-of-experts model with 13B active parameters out of 284B total. It is suited for document and chart understanding, visual question answering, and multimodal agent workflows that interleave text and images. llms.txt
- Z.ai: GLM 5.3 FlashX by Z.ai. GLM-5.3-FlashX is the high-speed variant of Z.ai's GLM-5.3-Flash, a native multimodal model delivering inference speeds of up to 200 tokens/s. Built on the same hybrid sparse and linear attention architecture (320B total parameters, 18B active), it is suited for efficient coding, visual understanding, and long-horizon agent tasks with a 1M-token context window. llms.txt
- Google: Gemini 3.6 Flash by Google. Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produce polished outputs with fewer unnecessary edits and less hedging, while reducing token use and the number of model calls needed to complete a task. llms.txt
- MoonshotAI: Kimi K2.6 by moonshotai. Kimi K2.6 is Moonshot AI's next-generation multimodal model, designed for long-horizon coding, coding-driven UI/UX generation, and multi-agent orchestration. It handles complex end-to-end coding tasks across Python, Rust, and Go, and can convert prompts and visual inputs into production-ready interfaces. Its agent swarm architecture scales to hundreds of parallel sub-agents for autonomous task decomposition - delivering documents, websites, and spreadsheets in a single run without human oversight. llms.txt
- OpenAI: GPT-5 Mini by OpenAI. GPT-5 Mini is a compact version of GPT-5, designed to handle lighter-weight reasoning tasks. It provides the same instruction-following and safety-tuning benefits as GPT-5, but with reduced latency and cost. GPT-5 Mini is the successor to OpenAI's o4-mini model. llms.txt
- OpenAI: GPT-5 Nano by OpenAI. GPT-5-Nano is the smallest and fastest variant in the GPT-5 system, optimized for developer tools, rapid interactions, and ultra-low latency environments. While limited in reasoning depth compared to its larger counterparts, it retains key instruction-following and safety features. It is the successor to GPT-4.1-nano and offers a lightweight option for cost-sensitive or real-time applications. llms.txt
- Nex AGI: Nex-N2.5-Mini (free) by nex-agi. Nex-N2.5 is an agentic model built to turn goals into working, verified outcomes. Its core strength is agentic coding within a visual feedback loop: it can explore codebases, implement multi-file changes, run commands, launch applications, interact with browser and desktop interfaces, and test software from the user's perspective. When observed behavior does not match the intended result, Nex-N2.5 can diagnose the issue, revise its implementation, and test again. This makes it especially effective for autonomous software engineering, GUI-based QA, computer-use automation, deep research, and scientific workflows where success must be demonstrated in the environment—not merely inferred from generated code. llms.txt
- SpaceXAI: Grok 4.7 by SpaceXAI. Grok 4.7 is SpaceXAI's flagship model for coding, agentic tasks, and knowledge work, succeeding Grok 4.6. It is particularly strong at long-running software engineering tasks, verifying its own work, and managing long context, and it improves on its predecessor at professional knowledge work such as drafting documents and presentations. The model was trained with a longer reinforcement learning run weighted toward problems that take many hours to complete, and natively understands the Grok Bot harness for conversational tasks. It ships with a new safeguard stack that pairs strong jailbreak resistance with low refusal rates for legitimate cybersecurity and biology work. SpaceXAI's reported benchmark results use the xhigh reasoning effort. llms.txt
- Thinking Machines: Inkling Small (free) by thinkingmachines. Inkling Small is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 12B active parameters out of 276B total. It is positioned as the smaller, more efficient member of the Inkling family and is suited for reasoning, coding, agentic workflows, retrieval-augmented generation, instruction following, and multilingual conversation. llms.txt
- Cohere: North Mini Code (free) by Cohere. North Mini Code is Cohere's first agentic coding model and the debut of its North family. A sparse mixture-of-experts model with 30B total parameters and 3B active, it is optimized for code generation, agentic software engineering, and terminal tasks, and is trained to generalize across agent harnesses such as OpenCode and SWE-Agent. It offers a 256K-token context window with up to 64K tokens of output, supports interleaved reasoning and tool use via JSON schema, and is released open-weight under the Apache 2.0 license. Its small active-parameter footprint enables low-latency inference, including on local hardware. llms.txt
- Anthropic: Claude Opus 4.6 by Anthropic. Opus 4.6 is Anthropic’s strongest model for coding and long-running professional tasks. It is built for agents that operate across entire workflows rather than single prompts, making it especially effective for large codebases, complex refactors, and multi-step debugging that unfolds over time. The model shows deeper contextual understanding, stronger problem decomposition, and greater reliability on hard engineering tasks than prior generations. Beyond coding, Opus 4.6 excels at sustained knowledge work. It produces near-production-ready documents, plans, and analyses in a single pass, and maintains coherence across very long outputs and extended sessions. This makes it a strong default for tasks that require persistence, judgment, and follow-through, such as technical design, migration planning, and end-to-end project execution. For users upgrading from earlier Opus versions, see our [official migration guide here](https://openrouter.ai/docs/guides/guides/model-migrations/claude-4-6-opus) llms.txt
- OpenAI: GPT-5.4 Nano by OpenAI. GPT-5.4 nano is the most lightweight and cost-efficient variant of the GPT-5.4 family, optimized for speed-critical and high-volume tasks. It supports text and image inputs and is designed for low-latency use cases such as classification, data extraction, ranking, and sub-agent execution. The model prioritizes responsiveness and efficiency over deep reasoning, making it ideal for pipelines that require fast, reliable outputs at scale. GPT-5.4 nano is well suited for background tasks, real-time systems, and distributed agent architectures where minimizing cost and latency is essential. llms.txt
- OpenAI: GPT Luna Latest by OpenAI. This model always redirects to the latest model in the GPT Luna family. llms.txt
- DeepSeek: DeepSeek Pro Latest by DeepSeek. This model always redirects to the latest model in the DeepSeek Pro family. llms.txt
- OpenAI: GPT-6 Luna Pro by OpenAI. GPT-6 Luna Pro is the same underlying model as [GPT-6 Luna](https://openrouter.ai/openai/gpt-6-luna), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode llms.txt
- inclusionAI: Ling 3.0 Flash by inclusionai. *Ling-3.0-flash* is a *124B-parameter Mixture-of-Experts (MoE) model*, with approximately *5.1B parameters activated per token*. The model is designed with *token efficiency and production-scale agentic inference* as key priorities, enabling developers to complete more useful work within constrained token, latency, and serving-cost budgets. llms.txt
- Google: Gemini 3.5 Flash by Google. Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed. It is highly optimized for coding proficiency and parallel agentic execution loops, supporting text, image, video, audio, and PDF inputs. Defaults to medium thinking effort for faster and more cost-efficient responses, with full support for thinking levels (minimal, low, medium, high) for fine-grained cost/performance trade-offs. llms.txt
- Google: Gemini 3.1 Flash Lite Preview by Google. Gemini 3.1 Flash Lite Preview is Google's high-efficiency model optimized for high-volume use cases. It outperforms Gemini 2.5 Flash Lite on overall quality and approaches Gemini 2.5 Flash performance across key capabilities. Improvements span audio input/ASR, RAG snippet ranking, translation, data extraction, and code completion. Supports full thinking levels (minimal, low, medium, high) for fine-grained cost/performance trade-offs. Priced at half the cost of Gemini 3 Flash. llms.txt
- OpenAI: GPT-5.4 Mini by OpenAI. GPT-5.4 mini brings the core capabilities of GPT-5.4 to a faster, more efficient model optimized for high-throughput workloads. It supports text and image inputs with strong performance across reasoning, coding, and tool use, while reducing latency and cost for large-scale deployments. The model is designed for production environments that require a balance of capability and efficiency, making it well suited for chat applications, coding assistants, and agent workflows that operate at scale. GPT-5.4 mini delivers reliable instruction following, solid multi-step reasoning, and consistent performance across diverse tasks with improved cost efficiency. llms.txt
- OpenAI: GPT-4.1 Mini by OpenAI. GPT-4.1 Mini is a mid-sized model delivering performance competitive with GPT-4o at substantially lower latency and cost. It retains a 1 million token context window and scores 45.1% on hard instruction evals, 35.8% on MultiChallenge, and 84.1% on IFEval. Mini also shows strong coding ability (e.g., 31.6% on Aider’s polyglot diff benchmark) and vision understanding, making it suitable for interactive applications with tight performance constraints. llms.txt
- Anthropic: Claude Sonnet Latest by Anthropic. This model always redirects to the latest model in the Claude Sonnet family. llms.txt
- Poolside: Laguna XS 2.1 (free) by poolside. Laguna XS 2.1 is the latest coding agent model in the 33B-A3B category from [Poolside](https://poolside.ai/) and a step forward from their Laguna XS.2 model (released in April 2026). It combines tool calling and reasoning capabilities with a compact footprint, offering a 256K context window and up to 32K output tokens. Quantized to FP8 for fast, cost-efficient agentic coding workflows. Laguna XS 2.1 is designed for software engineering and agentic coding use cases, and you are responsible for confirming that it is appropriate for your intended application. Laguna XS 2.1 is subject to the [OpenMDW-1.1 License](https://openmdw.ai/license/1-1/), and should be used consistently with Poolside's [Acceptable Use Policy](https://poolside.ai/legal/acceptable-use-policy). We advise against circumventing Laguna XS 2.1 safety guardrails without implementing substantially equivalent mitigations appropriate for your use case. Please report security vulnerabilities or safety concerns to [[email protected]](mailto:[email protected]). If you are using Laguna XS 2.1 for free, we may use your inputs and outputs to train and improve our models. llms.txt
- Anthropic: Claude Sonnet 4.5 by Anthropic. Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers state-of-the-art performance on coding benchmarks such as SWE-bench Verified, with improvements across system design, code security, and specification adherence. The model is designed for extended autonomous operation, maintaining task continuity across sessions and providing fact-based progress tracking. Sonnet 4.5 also introduces stronger agentic capabilities, including improved tool orchestration, speculative parallel execution, and more efficient context and memory management. With enhanced context tracking and awareness of token usage across tool calls, it is particularly well-suited for multi-context and long-running workflows. Use cases span software engineering, cybersecurity, financial analysis, research agents, and other domains requiring sustained reasoning and tool use. llms.txt
- Qwen: Qwen3.6 35B A3B by Qwen. Qwen3.6-35B-A3B is an open-weight multimodal model from Alibaba Cloud with 35 billion total parameters and 3 billion active parameters per token. It uses a hybrid sparse mixture-of-experts architecture combining Gated DeltaNet linear attention with standard gated attention layers, enabling efficient inference at a fraction of the compute cost. The model supports a 262K token native context window (extensible to 1M via YaRN) and accepts text, image, and video inputs. It includes integrated thinking mode with reasoning traces preserved across multi-turn conversations, function calling, and structured output. Released under the Apache 2.0 license. llms.txt
- Qwen: Qwen3.7 Plus by Qwen. Qwen3.7-Plus is a cost-effective model in Alibaba's Qwen3.7 series. It supports text and image input with text output, building on the series' text capabilities with a comprehensive upgrade to its vision-language abilities while retaining full-stack, agent-level intelligence for coding, tool use, and productivity workflows. Its distinguishing trait is multi-modal interactive hybrid agent capability: it can perceive real-world scenes, read screens and interact with GUIs, generate code from visual references, and perform end-to-end navigation within mobile apps. llms.txt
- OpenAI: GPT-4.1 Nano by OpenAI. For tasks that demand low latency, GPT‑4.1 nano is the fastest and cheapest model in the GPT-4.1 series. It delivers exceptional performance at a small size with its 1 million token context window, and scores 80.1% on MMLU, 50.3% on GPQA, and 9.8% on Aider polyglot coding – even higher than GPT‑4o mini. It’s ideal for tasks like classification or autocompletion. llms.txt
- SpaceXAI: Grok 4.5 by SpaceXAI. Grok 4.5 is a model from SpaceXAI with frontier performance on coding, knowledge work, and STEM. llms.txt
- MiniMax: MiniMax M2.7 by MiniMax. MiniMax-M2.7 is a next-generation large language model designed for autonomous, real-world productivity and continuous improvement. Built to actively participate in its own evolution, M2.7 integrates advanced agentic capabilities through multi-agent collaboration, enabling it to plan, execute, and refine complex tasks across dynamic environments. Trained for production-grade performance, M2.7 handles workflows such as live debugging, root cause analysis, financial modeling, and full document generation across Word, Excel, and PowerPoint. It delivers strong results on benchmarks including 56.2% on SWE-Pro and 57.0% on Terminal Bench 2, while achieving a 1495 ELO on GDPval-AA, setting a new standard for multi-agent systems operating in real-world digital workflows. llms.txt
- Z.ai: GLM 5.1 by Z.ai. GLM-5.1 delivers a major leap in coding capability, with particularly significant gains in handling long-horizon tasks. Unlike previous models built around minute-level interactions, GLM-5.1 can work independently and continuously on a single task for more than 8 hours, autonomously planning, executing, and improving itself throughout the process, ultimately delivering complete, engineering-grade results. llms.txt
- Qwen: Qwen3 235B A22B Instruct 2507 by Qwen. Qwen3-235B-A22B-Instruct-2507 is a multilingual, instruction-tuned mixture-of-experts language model based on the Qwen3-235B architecture, with 22B active parameters per forward pass. It is optimized for general-purpose text generation, including instruction following, logical reasoning, math, code, and tool usage. The model supports a native 262K context length and does not implement "thinking mode" (<think> blocks). Compared to its base variant, this version delivers significant gains in knowledge coverage, long-context reasoning, coding benchmarks, and alignment with open-ended tasks. It is particularly strong on multilingual understanding, math reasoning (e.g., AIME, HMMT), and alignment evaluations like Arena-Hard and WritingBench. llms.txt
- Google: Gemini Flash Latest by Google. This model always redirects to the latest model in the Gemini Flash family. llms.txt
- Qwen: Qwen3.5-Flash by Qwen. The Qwen3.5 native vision-language Flash models are built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. Compared to the 3 series, these models deliver a leap forward in performance for both pure text and multimodal tasks, offering fast response times while balancing inference speed and overall performance. llms.txt
- OpenAI: GPT Sol Latest by OpenAI. This model always redirects to the latest model in the GPT Sol family. llms.txt
- OpenAI: GPT-5.5 by OpenAI. GPT-5.5 is OpenAI’s frontier model designed for complex professional workloads, building on GPT-5.4 with stronger reasoning, higher reliability, and improved token efficiency on hard tasks. It features a 1M+ token context window (922K input, 128K output) with support for text and image inputs, enabling large-scale reasoning, coding, and multimodal workflows within a single system. llms.txt
Providers
- OpenAI 105 models on OpenRouter.
- Tencent Cloud 5 models on OpenRouter.
- Together 18 models on OpenRouter.
- Xiaomi 5 models on OpenRouter.
- NovitaAI 77 models on OpenRouter.
- Relace 7 models on OpenRouter.
- Google Vertex 62 models on OpenRouter.
- DeepInfra 111 models on OpenRouter.
- Wafer 8 models on OpenRouter.
- NVIDIA 8 models on OpenRouter.
- Open Inference 4 models on OpenRouter.
- DeepSeek 2 models on OpenRouter.
- Alibaba Cloud Int. 67 models on OpenRouter.
- StreamLake 22 models on OpenRouter.
- TypeSafe 1 models on OpenRouter.
- Meta 7 models on OpenRouter.
- Azure 62 models on OpenRouter.
- Claude Platform on AWS 10 models on OpenRouter.
- Stealth 1 models on OpenRouter.
- Z.ai 14 models on OpenRouter.
- Amazon Bedrock 36 models on OpenRouter.
- GMICloud 23 models on OpenRouter.
- Anthropic 25 models on OpenRouter.
- CoreWeave 19 models on OpenRouter.
- Upstage 3 models on OpenRouter.
- Parasail 39 models on OpenRouter.
- Google AI Studio 29 models on OpenRouter.
- AtlasCloud 34 models on OpenRouter.
- Baidu Qianfan 10 models on OpenRouter.
- Poolside 4 models on OpenRouter.
- MiniMax 13 models on OpenRouter.
- Fireworks 13 models on OpenRouter.
- Sail Research 7 models on OpenRouter.
- SiliconFlow 41 models on OpenRouter.
- SpaceXAI 14 models on OpenRouter.
- inference.net 5 models on OpenRouter.
- Baseten 14 models on OpenRouter.
- Venice 36 models on OpenRouter.
- Phala 20 models on OpenRouter.
- DigitalOcean 16 models on OpenRouter.
- Modal 5 models on OpenRouter.
- Nex AGI 2 models on OpenRouter.
- Thinking Machines 2 models on OpenRouter.
- NextBit 11 models on OpenRouter.
- Morph 7 models on OpenRouter.
- Groq 8 models on OpenRouter.
- Moonshot AI 3 models on OpenRouter.
- Reka AI 6 models on OpenRouter.
- Inceptron 6 models on OpenRouter.
- Cohere 10 models on OpenRouter.
- Friendli 7 models on OpenRouter.
- Makora 5 models on OpenRouter.
- Nebius Token Factory 10 models on OpenRouter.
- DekaLLM 8 models on OpenRouter.
- Decart 5 models on OpenRouter.
- AkashML 6 models on OpenRouter.
- Cloudflare 18 models on OpenRouter.
- Mistral 28 models on OpenRouter.
- Crusoe 6 models on OpenRouter.
- Darkbloom 8 models on OpenRouter.
- ModelRun [by Modular] 4 models on OpenRouter.
- Chutes 6 models on OpenRouter.
- Voyage AI by MongoDB 7 models on OpenRouter.
- Ionstream 2 models on OpenRouter.
- NEAR AI 1 models on OpenRouter.
- Inception 2 models on OpenRouter.
- Cerebras 1 models on OpenRouter.
- io.net 3 models on OpenRouter.
- Mancer 9 models on OpenRouter.
- Perplexity 7 models on OpenRouter.
- Liquid 2 models on OpenRouter.
- Krea 4 models on OpenRouter.
- Prime Intellect 1 models on OpenRouter.
- Seed 14 models on OpenRouter.
- StepFun 1 models on OpenRouter.
- MARA 5 models on OpenRouter.
- AionLabs 6 models on OpenRouter.
- SambaNova 7 models on OpenRouter.
- Unbiased 1 models on OpenRouter.
- Arcee AI 1 models on OpenRouter.
- Black Forest Labs 7 models on OpenRouter.
- Sakana 4 models on OpenRouter.
- Perceptron 1 models on OpenRouter.
- Recraft 16 models on OpenRouter.
- Sourceful 4 models on OpenRouter.
- Fish Audio 6 models on OpenRouter.
- Deepgram 3 models on OpenRouter.
- AssemblyAI 1 models on OpenRouter.
- HeyGen 1 models on OpenRouter.
- Runway 2 models on OpenRouter.
Collections
- Image Generation Models Browse and compare top image generation models on OpenRouter. Create images from text prompts through one unified API.
- Free AI Models on OpenRouter Access powerful AI models at zero cost. Experiment, learn, and build with free AI models and LLMs. OpenRouter is committed to keeping AI accessible for everyone.
- Discounted AI Model Prices: Current Provider Deals AI models with a promotional discount on their cheapest provider right now.
- Best AI Models for Tool Calling and Function Calling The best AI models for tool calling and function calling, ranked by real usage on OpenRouter this week.
- AI Models That Allow Distillation and Training AI models that explicitly allow distillation and training on their outputs, ranked by weekly usage.
- Best AI Models for Coding, Ranked by Real Usage The best AI models for coding, ranked by real developer usage on OpenRouter over the past week.
- Best Roleplay (RP) & Creative Writing AI Models by Usage The best AI models for roleplay, character chat, and creative writing, ranked by real usage this week.
- AI Models with Vision: Multimodal LLMs for Image Understanding Compare vision models and multimodal LLMs on OpenRouter. Find AI models that analyze images, read documents, interpret charts and answer questions about visual content through a single API.
- Top AI Models Used by OpenClaw Discover the most popular AI models to use with OpenClaw, the open-source autonomous AI agent. See top LLMs used in OpenClaw, ranked by real usage data on OpenRouter.
- Text Embedding Models Explore text embedding models on OpenRouter. Find top embedding APIs for semantic search, RAG pipelines, clustering, and similarity matching through one unified API.
- Video Generation Models Browse video generation models available on OpenRouter. Create videos from text and image prompts using leading video generation models through one API gateway.
- Best Audio Generation Models Compare the best AI audio generation models on OpenRouter. Find top models for music generation, sound generation, and audio-output applications through one unified API.
- Best Text-to-Speech Models Explore the best text-to-speech models on OpenRouter. Compare top TTS APIs for voice generation, narration, assistants, accessibility, and audio apps.
- Best Speech-to-Text and Transcription Models Find the best speech-to-text and transcription models on OpenRouter. Compare top audio transcription APIs for meetings, calls, captions, and speech recognition.
- Best Rerank Models for Search and RAG Compare the best rerank models on OpenRouter for semantic search, RAG pipelines, recommendation systems, and retrieval quality through one API.