OpenAI's Astra architecture raises concerns about AI oversight, experts say

AI safety researchers have expressed concern that OpenAI's Astra model uses a novel architecture that could obscure its reasoning process. The method, called recurrent depth, reduces computing costs but makes chain-of-thought monitoring more difficult. OpenAI's chief scientist says the company has limited its use and remains committed to interpretability.
The technique, known as recurrent depth or looped Transformers, compresses part of the model’s processing into repeated internal cycles rather than explicit, step-by-step language. This cuts computing costs per prompt, a key selling point as businesses face steep fees for advanced AI. However, it also bypasses the natural-language “chain of thought” that safety teams currently rely on to trace a model’s decisions, making oversight harder.
OpenAI’s chief scientist, Jakub Pachoki, pushed back against the initial report, insisting the company has limited the technique’s use to keep reasoning legible. He acknowledged that monitoring may become more difficult over time, but argued this challenge is not tied to architecture changes. Meanwhile, outside researchers warn that even limited adoption could normalize the approach, prompting other firms to expand it until models become fully opaque.
If recurrent depth becomes standard, regulators and third-party auditors may lose a key tool for verifying AI behavior, potentially eroding public trust in automated systems. Businesses deploying such models could face greater liability if they cannot explain an AI’s harmful action. However, the efficiency gains may also lower costs, making advanced AI more accessible. The balance between transparency and performance will likely shape future governance debates.