DeepSeek AI released DeepSeek-V4.1-Flash, a multimodal Mixture-of-Experts model featuring 552 billion parameters in its core architecture, 196 billion additional Engram parameters, and a context window of up to 1 million tokens.

According to MarkTechPost, the release is also associated with FP4 compression of the KV cache and cross-layer attention reuse. The publication describes the development as a response to the load caused by reprocessing long input data and storing KV caches.

The practical significance depends on whether the claimed characteristics are confirmed in independent tests. The available package contains no results from such tests, information on model availability, or comparisons with alternatives.