MobbleOpen in Mobble ⇢
Technology · Artificial intelligence · published 2026-09-01 · via VentureBeat

Longer reasoning helps top AI models retrieve up to 65% of seemingly forgotten facts

A study from Google Research and Technion reveals that large language models often have facts stored in their parameters but fail to output them during generation. Experiments with frontier models like GPT-5 and Gemini-3 show they encode 95-98% of tested facts, yet can recover up to 65% of initially unrecalled facts simply by reasoning for longer. This suggests that hallucination may stem from retrieval failures rather than missing knowledge, challenging common assumptions about model limitations.

Expanded Detail

The study, conducted jointly by Google Research and Technion, examined why large language models sometimes produce incorrect or fabricated answers despite having the correct information embedded in their neural networks. By testing frontier systems such as GPT-5 and Gemini-3, researchers found that these models successfully encode 95–98% of the facts they were queried about, yet frequently fail to surface them during standard generation.

When the same models were allowed to reason for extended periods, they recovered up to 65% of the facts they had initially missed. This suggests that many apparent hallucinations are not due to missing knowledge but rather to inefficient retrieval during output. The finding challenges the common assumption that a model's failure to answer correctly indicates a gap in its stored information, pointing instead toward deeper issues in how reasoning processes access latent knowledge.

Context

This research could reshape how developers and users perceive AI reliability. If hallucinations often stem from retrieval failures rather than knowledge gaps, then improving reasoning strategies or inference-time techniques may offer a practical path to reducing errors without retraining massive models. Businesses relying on AI for factual tasks—such as customer support, research, or coding—might see fewer misleading outputs, while regulators and the public may gain a more nuanced view of model limitations. However, the remaining 35% of unrecovered facts still highlight persistent risks, meaning caution remains warranted in high-stakes applications.

Expanded detail and Context are AI-generated analysis; the linked article remains the authoritative source.
Read the full article at VentureBeat →
This summary is AI-generated and original to Mobble; the linked article is the authoritative source. Original headline: “Frontier models can recover up to 65% of facts they can't directly recall — just by thinking longer.” Browse more stories.