DeepSeek Open Weights (MIT License)Status: active

DeepSeek-V3

Frontier open-weights 671B MoE base and chat model with multi-head latent attention (MLA) and dual-pipe training.

Released: Dec 26, 2024
Last Verified: Aug 10, 2026
Context Window

64k

Tokens

Max Output

8.192k

Output limit

Architecture

MoE (Multi-Head Latent Attention)

Model family

Parameters

671B (37B active)

Total / Active

Input Token Cost

$0.14

Per 1M tokens

Output Token Cost

$0.28

Per 1M tokens

Model Overview

DeepSeek-V3 is an open-weights 671B parameter Mixture of Experts model activating 37B parameters per token. Featuring Multi-Head Latent Attention (MLA) and Multi-Token Prediction (MTP), it achieves state-of-the-art general language, coding, and mathematical benchmark performance with industry-low API pricing ($0.14/$0.28 per MTok).

Developer Implementation Notes

Unbeatable API price-performance ratio. Uses Multi-Token Prediction (MTP) for speculative decoding speedup.

Key Strengths

  • MIT licensed open-weights model
  • Ultra-low API cost ($0.14 input / $0.28 output per MTok)
  • Excellent tool calling and coding benchmark scores
  • MLA architecture drastically reduces KV cache memory consumption

Limitations & Boundaries

  • 64k context window limit
  • Local deployment of full 671B weights requires high-memory cluster

Capabilities & Modalities

Modalities:text
Reasoning Chains:No
Tool / Function Calling:Supported
Structured Outputs:Supported
Fine-Tuning:Supported
Local Deployment:Yes
Min VRAM:40 GB

Best Production Use Cases

  • General enterprise LLM workloads at minimal cost
  • Code generation and API translation
  • Multilingual translation and summarization

Verified Benchmark Results

Standardized evaluations with methodology notes and authoritative citation links.

View all industry benchmarks →
Coding Official
LiveCodeBench

49.2%

Pass@1, 0-shot code generation

Reasoning & Math Official
GPQA Diamond

59.1%

Zero-shot Chain-of-Thought

General Knowledge & Frontier Official
MMLU-Pro

75.9%

5-shot standard evaluation

Tool & Agentic Official
BFCL (Berkeley Function Calling)

88.9%

AST accuracy on tool calling tasks

Pricing & Inference Cost Calculator

Token Cost Estimator – DeepSeek-V3

Calculate projected inference spend with prompt caching

Quick Workload Presets

Input Tokens / Req2,000
100100k200k
Output Tokens / Req500
5016k32k
Number of Requests1,000
125k50k
Cache Hit Rate0%
0% (No cache)50%90% (Max)
Total Input Spend

$0.2800

2.00M input tokens

Total Output Spend

$0.1400

0.50M output tokens

Estimated Total Cost

$0.42

Avg: $0.00042 / req

Rates: $0.14 in / $0.28 out per million tokens (DeepSeek Official API)Last verified: Aug 10, 2026

Compare with Similar Models

DeepSeek

DeepSeek-R1

Open-weights frontier reasoning model trained via large-scale reinforcement learning without supervised cold start.

Compare DeepSeek-V3 vs DeepSeek-R1
Mistral AI

Codestral 2501

Mistral AI's dedicated code generation model with 256k context window and fill-in-the-middle (FIM) capabilities.

Compare DeepSeek-V3 vs Codestral 2501
Meta AI

Llama 3.3 70B Instruct

Meta's open-weights 70B flagship matching Llama 3.1 405B capabilities on industry benchmarks at 1/5th the compute.

Compare DeepSeek-V3 vs Llama 3.3 70B Instruct

Data Accuracy & Verification Notice

AI model specifications, pricing records, and benchmark metrics published on this platform are compiled directly from authoritative sources (official provider documentation, research papers, and verified evaluation harnesses). Benchmark results reflect specific test harnesses and prompting methodologies; scores are not directly comparable across differing evaluation setups.