AI Models
New frontier model reasons for hours — and knows when it might be wrong
A major AI lab unveiled a system that sustains chains of reasoning across multi-day projects and flags its own uncertainty, a step researchers called decisive for trustworthy AI agents.
By Maya Okafor, Senior AI Correspondent — SAN FRANCISCO
SAN FRANCISCO — A leading AI lab on Friday unveiled a frontier model it says can sustain coherent reasoning across tasks lasting hours or even days, while explicitly signalling when its own conclusions are uncertain — a combination researchers have described as the missing ingredient for dependable AI agents.
In demonstrations, the system planned and executed a week-long software migration, paused to ask clarifying questions at genuinely ambiguous points, and attached calibrated confidence estimates to each of its intermediate conclusions. On public benchmarks measuring long-horizon task completion, it posted the largest single-generation jump since agent evaluations began.
The release intensifies a year-long shift in the industry away from chatbots and toward autonomous agents that carry out complete units of work. Rival labs are expected to answer within months, and enterprise customers began receiving early access on Friday.
Enable JavaScript to read the full story on Neural Daily News.