Run DeepSeek V4 Flash locally or in the cloud (US-hosted)

Aug 2, 2026·
DeepSeek

DeepSeek V4 Flash is now available in LM Studio Bionic. Run it in the cloud for agentic work, or download the model to run locally in LM Studio.

DeepSeek V4 Flash cloud inference in LM Studio is hosted on US-based servers, with Zero Data Retention (ZDR) enabled by default.

Leap in capabilities for the size class, and for the cost

DeepSeek reports that V4 Flash outperforms GLM 5.2 on every benchmark below where both models have a score. It does so with 284B total parameters versus GLM 5.2's 753B—about 62% fewer parameters—bringing capable coding and tool-use performance at a remarkably low inference cost.

Built for agentic work

DeepSeek V4 Flash 0731 is the official release of DeepSeek V4 Flash, replacing the preview with substantially stronger agentic capabilities.

This makes it a strong fit for Bionic workflows that require sustained tool use: navigating a repository, implementing changes across files, debugging, or working through complex research and document tasks.

Benchmark results

The following results are reported by DeepSeek in the official model card:

BenchmarkDeepSeek V4 Flash 0731V4 Flash PreviewV4 Pro PreviewGLM-5.2Opus-4.8
Terminal Bench 2.182.761.872.181.085.0
NL2Repo54.239.438.548.969.7
Cybergym76.738.752.783.1
DeepSWE54.47.312.846.258.0
Toolathlon-Verified70.349.755.959.976.2
Agents' Last Exam25.215.816.523.825.7
AutomationBench Public25.110.812.812.927.2
DSBench-FullStack †68.737.041.861.871.6
DSBench-Hard †59.625.831.154.571.7

For public code-agent tasks, DeepSeek evaluated V4 Flash 0731 using the minimal mode of DeepSeek Harness, max reasoning effort, temperature = 1.0, and top_p = 0.95. † DSBench-FullStack and DSBench-Hard are internal DeepSeek test sets. See the official model card for methodology and details.

Run it locally or in the cloud

DeepSeek V4 Flash is a 284B-parameter Mixture-of-Experts model. You can download it and run it on your own hardware in LM Studio, but plan for at least 156 GB of system memory. The model and its weights are released under the MIT License.

If you would rather not manage that hardware, select the cloud model in Bionic. It delivers highly capable coding and agentic performance while remaining incredibly cost-efficient. LM Studio hosts it on US-based infrastructure, with Zero Data Retention enabled by default.

Get started

Download LM Studio Bionic, create an LM Studio account, and select DeepSeek V4 Flash from the cloud model picker. Or open the DeepSeek V4 Flash model page to download it for local use.