DeepSeek-V4.1-Flash
DeepSeek-V4.1-Flash is an open-weight language model from DeepSeek, released on September 10, 2026. It is a mixture-of-experts model with a 552B-parameter backbone plus 196B parameters of Engram conditional memory. During prefill 8B parameters are active per token, and during decoding 16B. DeepSeek lists a context window of 1M tokens. The model accepts text and image input.
- Developer
- DeepSeek
- Released
- September 10, 2026
- Total parameters
- 552B + 196B memory
- Active parameters per token
- 8B–16B
- Architecture
- MoE
- Context window
- 1M
- Input
- Text, image
- License
- MIT
- Commercial use
- Yes
- Official weights, GB
- 510.3 GB
- Hugging Face repository
- huggingface.co
- Checked on
The weights are released under the MIT license, which allows commercial use.
The official weights on Hugging Face take about 510 GB. Running the model needs at least that much memory across GPUs and system RAM, plus room for the context cache; quantized versions need less.
- Total parameters, billions
- 748B
- Active parameters, billions
- 8B
- Context window, tokens
- 1M
- License type
- MIT
- License conditions
- MIT: keep the copyright and permission notice.
- Weights precision
- Mixed FP8 + BF16
- License text
- huggingface.co
- Sources
Page URL Hugging Face model page https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash License text on Hugging Face https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/main/LICENSE Hugging Face weights index https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/raw/main/model.safetensors.index.json DeepSeek API docs https://api-docs.deepseek.com/updates