1 link tagged with all of: inference-efficiency + language-models + deepseek
Click any tag below to further narrow down your results
Links
DeepSeek released V4.1-Flash, a 552B-parameter model that uses only 8B active parameters for input processing and 16B for output, cutting KV cache requirements to 1/4 the memory and 1/8 the storage of the previous generation. The company is retiring V4-Pro and routing all its traffic to V4.1-Flash at lower prices starting September 14, 2026.
- V4.1-Flash outperforms V4-Pro on benchmarks while using asymmetric encoder-decoder architecture that dramatically reduces active parameters and cache overhead
- KV cache compression cuts memory by 75% and storage by 87.5%, directly lowering inference costs for agents and long-running tasks
- Pricing drops on September 10, 2026, with off-peak rates at 50% of peak rates; V4-Pro requests automatically migrate to V4.1-Flash at the new lower rates