AI Models
Open-weights model matches closed frontier systems on reasoning benchmarks
A freely downloadable model has drawn level with the best proprietary systems on independent reasoning evaluations, narrowing the open–closed gap to months.
By Maya Okafor, Senior AI Correspondent — PARIS
PARIS — An open-weights language model released this month has matched the leading proprietary systems on independent reasoning benchmarks, according to results verified by three external evaluation groups — the closest the open ecosystem has come to the closed frontier since the current generation began.
The model, downloadable and fine-tunable by anyone, trails flagship commercial systems only on the longest-horizon agentic tasks. On mathematics, coding and scientific reasoning suites, the gap is within the margin of error.
Proprietary labs argue the answer is reliability engineering, safety tooling and the surrounding agent infrastructure rather than raw model quality. Policymakers, meanwhile, are re-examining assumptions: several pending regulations were drafted on the premise that frontier capability would remain concentrated in a few companies.
Enable JavaScript to read the full story on Neural Daily News.