DeepSeek V4.1 Flash explained: 552B MoE with small active parameters, MIT license, KV-cache savings, benchmarks, and how to self-host or call it by API.
DeepSeek V4 explained: the 1.6T-parameter MoE architecture, Compressed Sparse Attention, 1M-token context, SWE-bench and reasoning benchmarks, pricing, and how to deploy it.
DeepSeek V4 explained: the 1.6T-parameter MoE architecture, Compressed Sparse Attention, 1M-token context, SWE-bench and reasoning benchmarks, pricing, and how to deploy it.
In two weeks of June 2026, ~12 frontier open-weight models shipped — GLM-5.2, MiniMax M3, DeepSeek V4.1, Qwen 3.7. What it means for cost, moats, and strategy.