DeepSeek AI Released DeepSeek-V4.1-Flash with 1M Context, FP4 KV Cache, and Cross-Layer Attention Reuse
Long-horizon brokers have turned LLM serving into an input-heavy workload. Repeated prefills and million-token contexts depart KV caches that pressure HBM, SSD capability, and bandwidth. DeepSeek AI constructed its latest launch round that actual bottleneck. DeepSeek-V4.1-Flash is a multimodal Mixture-of-Experts mannequin with 552B spine parameters, 196B further Engram parameters, and a 1M-token context window. It…
