Skip to content
Not available in this workspace
OpenRouterOpenRouter
© 2026 OpenRouter, Inc

Product

  • Chat
  • Rankings
  • Benchmarks
  • Apps
  • Discover
  • Models
  • Collections
  • Providers
  • Pricing
  • Enterprise
  • Labs

Company

  • About
  • Blog
  • Careers
    Hiring
  • Privacy
  • Terms of Service
  • Support
  • Works With OR
  • Data

Developer

  • Documentation
  • API Reference
  • Developer Platform
  • Status

Connect

  • Discord
  • GitHub
  • LinkedIn
  • X
  • YouTube

This is a base model trained for raw text prediction, not instruction-following. Prompts should be written as examples, not simple requests.

Favicon for deepseek

DeepSeek: DeepSeek V3.1 Base

deepseek/deepseek-v3.1-base

Model weights

This is a base model, trained only for raw next-token prediction. Unlike instruct/chat models, it has not been fine-tuned to follow user instructions. Prompts need to be written more like training text or examples rather than simple requests (e.g., “Translate the following sentence…” instead of just “Translate this”).

DeepSeek-V3.1 Base is a 671B parameter open Mixture-of-Experts (MoE) language model with 37B active parameters per forward pass and a context length of 128K tokens. Trained on 14.8T tokens using FP8 mixed precision, it achieves high training efficiency and stability, with strong performance across language, reasoning, math, and coding tasks.

Modalities

Context

164K

Released

Aug 20, 2025

Knowledge Cutoff

Mar 2025

About DeepSeek: DeepSeek V3.1 Base

OpenRouter makes DeepSeek: DeepSeek V3.1 Base available through a unified, OpenAI-compatible API using the model ID deepseek/deepseek-v3.1-base.

DeepSeek: DeepSeek V3.1 Base accepts text and returns text. It has a 163,840-token context window.

It was released on August 20, 2025; its knowledge cutoff is March 31, 2025.

More models from DeepSeek

  • DeepSeek V4 Pro 0813
  • DeepSeek V4 Flash 0731
  • DeepSeek V4 Pro 0423

Frequently asked questions

This is a base model, trained only for raw next-token prediction. Unlike instruct/chat models, it has not been fine-tuned to follow user instructions. Prompts need to be written more like training text or examples rather than simple requests (e.g., “Translate the following sentence…” instead of just “Translate this”).

DeepSeek V3.1 Base has a 163,840 token context window.

DeepSeek V4 Pro 0813, DeepSeek V4 Flash 0731, DeepSeek V4 Pro 0423 and 9 more are other text models from DeepSeek.

DeepSeek V3.1 Base was released on August 20, 2025. Its knowledge cutoff is March 31, 2025.

ActivityFAQ

Activity

Token volume and request traffic to this model over time.