← All Models

DeepSeek V4 Flash

18.5K Downloads

DeepSeek V4 Flash 0731 is a 284B Mixture-of-Experts model built for coding, tool use, and agentic workflows, with 13B active parameters and a 1M-token context window.

Run with LM Studio Cloud
Input$0.13Cached$0.028Output$0.26
Download models
Updated 4 hours ago

Memory Requirements

To run the smallest DeepSeek V4 Flash, you need at least 156 GB of RAM.

Capabilities

DeepSeek V4 Flash models support tool use and reasoning. They are available in gguf.

About DeepSeek V4 Flash

DeepSeek

DeepSeek-V4-Flash-0731 is the official DeepSeek V4 Flash release, replacing the preview with substantially enhanced agentic capabilities. It is available in LM Studio both as a download and as a Cloud model.

Highlights

  • Strong long-horizon coding, terminal, tool-use, and automation performance
  • Three reasoning-effort levels: low, high, and max
  • MIT-licensed model weights
  • Local downloads and hosted LM Studio Cloud inference on one page

Benchmarks

DeepSeek reports that the official Flash release outperforms both the Flash preview and V4 Pro preview across the benchmarks below despite using a much smaller activated parameter count.

BenchmarkDeepSeek V4 Flash 0731V4 Flash PreviewV4 Pro PreviewGLM-5.2Opus-4.8
Terminal Bench 2.182.761.872.181.085.0
NL2Repo54.239.438.548.969.7
Cybergym76.738.752.7β€”83.1
DeepSWE54.47.312.846.258.0
Toolathlon-Verified70.349.755.959.976.2
Agents' Last Exam25.215.816.523.825.7
AutomationBench Public25.110.812.812.927.2
DSBench-FullStack †68.737.041.861.871.6
DSBench-Hard †59.625.831.154.571.7

Source: DeepSeek V4 Flash 0731 official model card. For public code-agent tasks, DeepSeek evaluated the model with the minimal mode of its DeepSeek Harness at max reasoning effort, temperature = 1.0, and top_p = 0.95. † DSBench-FullStack and DSBench-Hard are internal DeepSeek test sets. Scores and methodology are reproduced as reported by DeepSeek.

Reasoning and sampling

Use reasoning_effort to choose between low, high, and max deliberation. For local agentic workloads, DeepSeek recommends temperature = 1.0 and top_p = 0.95; for other workloads, it recommends top_p = 1.0. DeepSeek recommends allowing up to 384K output tokens for high and max reasoning.

Download or use in the Cloud

Choose a downloadable quantization that fits your hardware, or select DeepSeek V4 Flash from the Cloud model picker in LM Studio Bionic. LM Studio Cloud models run with Zero Data Retention enabled by default.

License and sources

DeepSeek V4 Flash is released under the MIT License.

Official model card Β· DeepSeek V4 technical report

DeepSeek V4 Flash