← All models

DeepSeek V4 Flash

Fast general chat

Chat readyMaciPhone

Mac: macOS 15 or later on Apple silicon, and it updates itself. iPhone: through TestFlight, iPhone 15 Pro or later on iOS 18.

  1. Download Minirun
  2. Models > Find Models — pick DeepSeek V4 Flash
  3. Verify all files
  4. New chat
The Minirun instruments panel on macOS during a DeepSeek V4 Flash run: the decode rate, the data read per token, and the footprint against the budget.The same model on an iPhone 16 Pro: 0:15 / token, the data read for that token, the footprint against a 3.76 GB budget, and where the time went.
DeepSeek V4 Flash part-way through a turn, with the instruments Minirun shows while it answers.

About

DeepSeek's mixture-of-experts text model, packaged for Minirun. It is the one to start with: the smallest container that still answers like a large model, and the quickest to a first reply on both Mac and iPhone.

On a Mac it answers at about 1.7 s / token at the 10.7 GB Balanced budget. On an iPhone 16 Pro: about 15 s / token at 3.8 GB, and it runs at 2 GB. The pace depends on the drive, the cable and the budget.

What you need

About 2.0 GB of memory on Mac, 2.0 GB on iPhone 16 Pro, and 167 GB free on your SSD.

Model requirements
Technical details
Stored precision
FP4 experts · FP8 matrices · BF16/F32 remainder
Container size
167 GB
License
MIT
Repository
nanguoyu/DeepSeek-V4-Flash-0731-minirun
Source model
deepseek-ai/DeepSeek-V4-Flash-0731
DeepSeek V4 Flash — Minirun