Lead
OpenAI has confirmed that its next major model family is called Astra, and said an internal version of the system solved ten open problems in mathematics and theoretical computer science that had remained unsolved for at least a decade, according to The Decoder, which published its report between Aug. 1 and Aug. 2, 2026. The confirmed problems span high-dimensional geometry, coding theory, group theory, quantum complexity, lattice cryptography, and extremal combinatorics, per The Decoder. The disclosure marks the first time OpenAI has officially named the model line that succeeds its prior generation of reasoning systems, and it arrives paired with a separate announcement that Astra will be the first OpenAI model to clear a planned U.S. government review process before public release.
The two announcements together signal that OpenAI is positioning Astra as both a research instrument and a regulated product. The Decoder, citing three people familiar with the plans who were originally reported by The Information, said the model is designed to coordinate multiple agents over hours or days of continuous work, a shift from the minutes-long reasoning windows that characterize earlier reasoning systems.
The Details
The ten unsolved problems are described in The Decoder’s Aug. 1-2 reporting as drawn from a working set maintained by Thomas Bloom, a mathematician at the University of Manchester who runs erdosproblems.com, a public catalogue of open questions traceable to Paul Erdős’s tradition. Bloom called the results “big news” in a post on X and said the constructions uncovered by Astra were bigger than the May 2026 counterexample to the unit distance conjecture, a separate high-profile result he referenced in the same post. Bloom also pushed back on the idea that the system replaces mathematicians, writing that the model “drew on more than a century of mathematical theory,” according to The Decoder.
One of the ten proofs establishes the existence of non-sofic groups, resolving a major open question in group theory, per The Decoder. The remaining results cover the five other fields cited in the announcement. OpenAI published a walkthrough of the model’s reasoning process for each solution, and the company said human researchers worked with the same model to convert the underlying arguments into research papers suitable for publication. Each proof was also formalized in Lean, a proof assistant that produces machine-checkable certificates of mathematical correctness.
The compute footprint was modest by frontier-model standards. The Decoder reported that the tokens used to generate all ten solutions would cost roughly $2,000 at published API rates, a figure the publication framed as evidence that test-time compute, rather than raw training scale, was the relevant input. Noam Brown, an OpenAI researcher who has published on test-time reasoning techniques, said on X that the company has tried and failed on the Millennium Prize Problems and called Astra a “major step for scientific reasoning,” per The Decoder. The Clay Mathematics Institute offers $1 million per Millennium Prize; only one of the seven problems on the list has been solved since 2000.
On credit and authorship norms, OpenAI cited the Leiden Declaration on AI and Mathematics, a document that lays out conventions for crediting AI systems in mathematical research, as a reference point for how the work should be attributed. The Decoder also reported that Sam Altman demonstrated Astra to politicians and regulators in Washington, D.C., a step that has been corroborated by Bleeping Computer (Aug. 2, 2026) and AnalyticsIndia Mag (Aug. 3, 2026), with SiliconANGLE (Aug. 2, 2026) confirming that the proofs have been published.
According to The Decoder, Astra is the first OpenAI model to go through the planned U.S. government review process, a procedure that requires official approval before the system can be made publicly available. The review sits alongside, rather than in place of, the company’s existing internal safety evaluations.
Why It Matters
The technical claim worth weighing is the move from minutes-long reasoning to multi-day coordination. Earlier reasoning systems were tuned to answer a question in a single, bounded session. Astra, per The Decoder, is built to coordinate multiple agents over hours or days, a different architectural bet that assumes long-horizon tasks can be decomposed into roles and supervised at the protocol layer rather than the token layer. Brown suggested in his X comments that the ceiling is far higher, noting that it is “possible to push test-time compute much further,” according to The Decoder.
The math results matter less for any individual theorem than for the genre of output they represent. The Decoder’s account describes publishable, machine-verifiable artifacts: Lean formalizations that any external reviewer can re-check, paired with reasoning walkthroughs that disclose the system’s intermediate steps. That combination is closer to a software release than to a chatbot answer, and it creates a template other labs can copy or contest.
The regulatory layer is the structural change. The government review requirement, as reported by The Decoder, introduces an approval gate between internal capability demonstrations and public deployment. The historical parallel is GPT-3 in 2020, when OpenAI moved a research demo into a product surface, but the 2026 path adds regulator pre-clearance as an explicit step. Bleeping Computer’s coverage of the Senate demo indicates that legislative staff have already seen the system in person, which puts the timeline for a written framework into sharper focus than an unscheduled briefing would.
For peers, the open question is whether DeepMind and Anthropic publish comparable artifacts under comparable credit norms. The Leiden Declaration reference suggests OpenAI wants a shared vocabulary for AI-assisted mathematical work before competing labs set their own defaults.
What to Watch Next
Three milestones will test the announcement. First, the U.S. government review window: Altman demonstrated Astra to senators before any formal review framework had landed, per The Decoder, and the public version of that framework is the next concrete artifact to watch. Second, the trajectory of test-time compute: Brown’s remark that it is “possible to push test-time compute much further” sets up an obvious next experiment, and any additional unsolved problems tackled with a larger compute budget would be a measurable signal of how far the technique scales. Third, peer replication: whether DeepMind, Anthropic, or academic groups publish their own Lean-formalized solutions to long-standing open problems using a comparable multi-agent setup will indicate whether Astra’s pattern is a one-off demonstration or the start of a publishing convention.
One longer-term marker remains out of reach. The Millennium Prize Problems are still unsolved, and per The Decoder’s account of Brown’s comments, OpenAI has tried and failed on them. The gap between ten

