DeepSeek-V4.1-Flash
MultimodalDeepSeek-V4.1-Flash is DeepSeek's open-weights multimodal Mixture-of-Experts release of September 2026. Its 552B-parameter backbone activates 8B parameters per token during prefill and 16B during decode, and it supports up to 1M tokens of context. The architecture pairs a 20-layer causal encoder with a 20-layer decoder, Compressed Sparse Attention 2 with a Hierarchical Sparse Indexer, an FP4 main KV cache (about 890 bytes per token), SWA Bounded Replay, Single-Pass mHC and Engram conditional memory (196B parameters), plus DSpark speculative decoding. It accepts native image input through DeepSeek-ViT and was pretrained from scratch on a 45T-token multimodal corpus with reasoning effort continuously controllable from 1 to 100.
Key Specifications
Timeline
Technical Specifications
Pricing & Availability
Benchmark Results
Model performance metrics across various tests and benchmarks
Programming
Mathematics
Reasoning
Other Tests
License & Metadata
Compare DeepSeek-V4.1-Flash
All comparisonsArticles about DeepSeek-V4.1-Flash

DeepSeek V4.1 Flash vs GPT-6 Astra: 98% of the score, 1.4% of the cost
Arena-measured data put DeepSeek V4.1 Flash within two points of GPT-6 Astra on design output at 1.4% of the cost, and the ranking flips once speed is weighted.

DeepSeek V4.1-Flash ships open weights built around a smaller KV cache
DeepSeek published V4.1-Flash open weights on Sept 10: a 552B multimodal MoE with an 890-byte KV cache per token, and vendor numbers that show where it loses.
Similar Models
All ModelsDeepSeek-V4-Flash-0423
DeepSeek
DeepSeek-V4-Flash-Max
DeepSeek
DeepSeek-V4-Pro-Max
DeepSeek
DeepSeek-V2.5
DeepSeek
DeepSeek-V3.2 (Thinking)
DeepSeek
DeepSeek-V3.2-Exp
DeepSeek
DeepSeek-R1
DeepSeek
Command A+
Cohere
Recommendations are based on similarity of characteristics: developer organization, multimodality, parameter size, and benchmark performance. Choose a model to compare or go to the full catalog to browse all available AI models.