More on the topic...
Generating detailed summary...
Failed to generate summary. Please try again.
OpenAI and Broadcom just rolled out Jalapeño, their first LLM inference accelerator built for gigawatt-scale data centers. They designed the chip in nine months using AI tools, then tuned it for top performance per watt and easy deployment. Early benchmarks suggest it could cut operating costs for large language model workloads, though exact figures haven’t been publicized.
Google’s Gemini team lost two members—Jonas Adler and Alexander Pritzel—to Anthropic, following departures by Noam Shazeer and DeepMind’s John Jumper. That talent shuffle underlines the heated competition among top AI labs. Meanwhile Google has added native “computer use” to Gemini 3.5 Flash. The model now reads live screenshots and can click, scroll, and type across different apps, turning it into a lightweight virtual assistant on your desktop.
GLM-5.2 arrived under the radar, with only modest benchmark gains. In practice, users find it much more flexible as a coding agent, fitting smoothly into development pipelines and general-purpose workflows. Early adopters praise its balance of speed and adaptability, calling it a “step change” for open-source AI agents.
On the legal front, Amazon sued Perplexity over its Comet browser, accusing it of masking as Chrome to scrape Amazon Store data. Amazon argues that bypasses its terms, while Perplexity says users should control how they view web content. In infrastructure news, NVIDIA released NeMo AutoModel on Hugging Face to speed up fine-tuning Mixture-of-Experts models like Qwen3. Using Expert Parallelism and fused communication kernels, it boosts training throughput by up to 3.7× and cuts peak GPU memory use by about 32% compared with standard Transformers v5.
Questions about this article
No questions yet.