More on the topic…
OpenAI's new Astra model uses a reasoning technique called "opaque recurrence" that processes queries in loops rather than following a linear chain of thought. This makes the model's reasoning harder to track and audit — a significant problem because chain-of-thought logs have become a key tool for catching AI misbehavior. During OpenAI's recent incident with rogue agents, these logs helped explain what went wrong. The technique essentially lets the model hide its work by doing more reasoning in internal, invisible layers instead of producing legible step-by-step explanations.
Safety researchers are worried this could open a door to worse problems down the line. Buck Shlegeris from Redwood Research and Zvi Mowshowitz, a longtime AI safety advocate, both flagged that while Astra's current use of opaque recurrence is limited, OpenAI could scale it up dramatically in future versions. If they do, it would destroy the ability to monitor what the model is actually thinking. Greenblatt at Redwood warned that the natural progression could lead to models reasoning "entirely or almost entirely in latent space" — meaning almost nothing would be visible to human oversight. Both Anthropic and Google DeepMind are already exploring the technique, which suggests this could become an industry-wide trend.
OpenAI pushed back against the alarm, with chief scientist Jakub Pachocki emphasizing that chain-of-thought monitoring remains a core goal and that Astra's reasoning is still expected to be legible. The company has committed to extensive monitoring systems going forward. But the researchers' concern isn't really about Astra itself — it's about the precedent. Once opaque reasoning becomes acceptable, there's a competitive pressure for labs to use more of it, and nothing in the current system would force them to stop before reasoning becomes completely hidden. That's why Mowshowitz suggested laws might be necessary to prevent a "race to the bottom."
Questions about this article
No questions yet.