Skip to content
Not available in this workspace
OpenRouterOpenRouter
© 2026 OpenRouter, Inc

Product

  • Chat
  • Rankings
  • Benchmarks
  • Apps
  • Discover
  • Models
  • Collections
  • Providers
  • Pricing
  • Enterprise
  • Labs

Company

  • About
  • Blog
  • Careers
    Hiring
  • Privacy
  • Terms of Service
  • Support
  • Works With OR
  • Data

Developer

  • Documentation
  • API Reference
  • Developer Platform
  • Status

Connect

  • Discord
  • GitHub
  • LinkedIn
  • X
  • YouTube
Favicon for deepseek

DeepSeek: R1 Distill Llama 8B

deepseek/deepseek-r1-distill-llama-8b

Model weights

DeepSeek R1 Distill Llama 8B is a distilled large language model based on Llama-3.1-8B-Instruct, using outputs from DeepSeek R1. The model combines advanced distillation techniques to achieve high performance across multiple benchmarks, including:

  • AIME 2024 pass@1: 50.4
  • MATH-500 pass@1: 89.1
  • CodeForces Rating: 1205

The model leverages fine-tuning from DeepSeek R1's outputs, enabling competitive performance comparable to larger frontier models.

Hugging Face:

  • Llama-3.1-8B(opens in new tab)
  • DeepSeek-R1-Distill-Llama-8B(opens in new tab) |

Modalities

Released

Feb 7, 2025

Knowledge Cutoff

Jul 2024

About DeepSeek: R1 Distill Llama 8B

OpenRouter makes DeepSeek: R1 Distill Llama 8B available through a unified, OpenAI-compatible API using the model ID deepseek/deepseek-r1-distill-llama-8b.

DeepSeek: R1 Distill Llama 8B accepts text and returns text.

It was released on February 7, 2025; its knowledge cutoff is July 31, 2024.

More models from DeepSeek

  • DeepSeek V4 Pro 0813
  • DeepSeek V4 Flash 0731
  • DeepSeek V4 Pro 0423

Frequently asked questions

DeepSeek R1 Distill Llama 8B is a distilled large language model based on Llama-3.1-8B-Instruct, using outputs from DeepSeek R1. The model combines advanced distillation techniques to achieve high performance across multiple benchmarks, including: - AIME 2024 pass@1: 50.4 - MATH-500 pass@1: 89.1 - CodeForces Rating: 1205 The model leverages fine-tuning from DeepSeek R1's outputs, enabling...

DeepSeek V4 Pro 0813, DeepSeek V4 Flash 0731, DeepSeek V4 Pro 0423 and 9 more are other text models from DeepSeek.

R1 Distill Llama 8B was released on February 7, 2025. Its knowledge cutoff is July 31, 2024.

ActivityFAQ

Activity

Token volume and request traffic to this model over time.