2 links tagged with all of: language-models + api-pricing
Click any tag below to further narrow down your results
Links
Sakana AI released Fugu Ultra v2, a model that orchestrates tasks across multiple specialized models rather than relying on a single monolithic architecture. It's designed for complex reasoning, autonomous research, and software development with a 1M token context window and costs $5/$30 per million input/output tokens.
- Fugu Ultra v2 uses learned multi-agent orchestration to route work across open and specialized models, avoiding dependence on proprietary frontier models
- The model supports configurable reasoning effort levels, function calling, structured outputs, and integrated web search
- Real-world performance shows 31 tokens/second throughput, 10.6 second latency, and 87.47% availability across providers over the past 3 days
DeepSeek released V4.1-Flash, a 552B-parameter model that uses only 8B active parameters for input processing and 16B for output, cutting KV cache requirements to 1/4 the memory and 1/8 the storage of the previous generation. The company is retiring V4-Pro and routing all its traffic to V4.1-Flash at lower prices starting September 14, 2026.
- V4.1-Flash outperforms V4-Pro on benchmarks while using asymmetric encoder-decoder architecture that dramatically reduces active parameters and cache overhead
- KV cache compression cuts memory by 75% and storage by 87.5%, directly lowering inference costs for agents and long-running tasks
- Pricing drops on September 10, 2026, with off-peak rates at 50% of peak rates; V4-Pro requests automatically migrate to V4.1-Flash at the new lower rates