[AINews] DeepSeek v4.1-Flash: 763B-P8B-D16B novel causal Encoder–Decoder architecture with vision marks the Return of the Whale
DeepSeek’s V4.1-Flash is framed here as a major architecture break, not a minor point release.
Latent Space says the model introduces a causal encoder-decoder setup with separate active-parameter paths for prefill and decode: 8B for input tokens and 16B for output tokens. The piece emphasizes efficiency more than benchmark rank, pointing to lower KV-cache footprint, long-context economics, native vision, and cheaper serving as the real advance. It also notes DeepSeek is effectively moving past V4 Pro while keeping the new release named only “v4.1 Flash,” a mismatch the authors treat as part of the story. Latent Space's note
Latent Space says the model introduces a causal encoder-decoder setup with separate active-parameter paths for prefill and decode: 8B for input tokens and 16B for output tokens. The piece emphasizes efficiency more than benchmark rank, pointing to lower KV-cache footprint, long-context economics, native vision, and cheaper serving as the real advance. It also notes DeepSeek is effectively moving past V4 Pro while keeping the new release named only “v4.1 Flash,” a mismatch the authors treat as part of the story. Latent Space's note
score 8