OpenAI says Astra solved ten open math problems for about $2,000

OpenAI has named its next major model, Astra, by publishing ten mathematical results it says were generated by an internal version of the system. The company released a technical paper, reasoning walkthroughs and Lean certificates that allow the formal proofs to be checked by software. The more important part of the announcement is not the count. AI-generated mathematics has produced enough false starts that a headline number alone is weak evidence. OpenAI's decision to publish machine-checkabl

2 min
Mid-century technical illustration representing Astra and machine-checked mathematical proofs

Mid-century technical illustration representing Astra and machine-checked mathematical proofs

OpenAI has named its next major model, Astra, by publishing ten mathematical results it says were generated by an internal version of the system. The company released a technical paper, reasoning walkthroughs and Lean certificates that allow the formal proofs to be checked by software.

The more important part of the announcement is not the count. AI-generated mathematics has produced enough false starts that a headline number alone is weak evidence. OpenAI's decision to publish machine-checkable certificates gives researchers something concrete to inspect, although it does not remove the need for human review.

What Astra produced

The ten results cover high-dimensional geometry, coding theory, arithmetic circuit complexity, group theory, operator algebras, quantum complexity, lattice cryptography and extremal combinatorics.

Among them is a construction establishing the existence of non-sofic groups, addressing a question in group theory that had remained open since the concept was introduced in 1999. Astra also produced a counterexample to Connes's rigidity conjecture, new bounds for sphere packing and error-correcting codes, and an exponential parallel repetition theorem for two-player quantum games.

Three results address problems from Paul Erdős's catalogue: multicolor triangle Ramsey numbers and two conjectures in extremal graph theory. Another concerns the closest vector problem, a lattice question with relevance to post-quantum cryptography.

OpenAI says the tokens used to find all ten solutions would cost roughly $2,000 at its Sol API rates. Humans then worked with the same model to prepare manuscripts, after which the model formalized each argument in Lean.

Verification is the real test

Lean checks whether a formal proof follows from its stated definitions and axioms. That makes the certificates reproducible in a way that conventional prose proofs are not: OpenAI has published the files in a public GitHub repository, where others can run them through Lean's kernel.

A successful build is not the end of the review. It verifies the formal statement encoded in Lean, not automatically that the statement matches the original open problem as mathematicians understand it. Researchers still need to inspect the formalization, judge the significance of each result and determine whether related work changes the novelty claims.

That distinction matters because OpenAI has previously faced scrutiny over AI-generated mathematical claims. Publishing the proofs, formal certificates and reasoning records together gives the mathematical community a substantially better audit trail this time.

Astra remains unavailable

OpenAI describes Astra only as its "next major model" and has not announced an API, release date, pricing or system card. The publication therefore demonstrates a research capability, not a product developers can test.

The company also addressed attribution directly. It says the mathematical arguments came from Astra, while its researchers prepared the manuscripts, helped formalize the proofs and accepted responsibility for their correctness. OpenAI cited the Leiden Declaration on AI and Mathematics, which argues that researchers should disclose AI contributions rather than claim human authorship for machine-generated work.

The ten results now move into a slower process than model evaluation: independent mathematical review. Lean narrows the verification problem, but it does not settle questions of framing, novelty or importance. Those judgments remain with the field.

Sources

OpenAI: Ten advances in mathematics and theoretical computer science

OpenAI: Lean certificates for the ten proofs

OpenAI: Technical paper

The Decoder: OpenAI announces Astra with ten math results

Written by

More to read

  • Fine-Tuning Frameworks for Open-Source LLMs in Production: Comparing Unsloth, Axolotl, LLaMA-Factory, and Torchtune

    Open-source large language model post-training has fragmented into distinct engineering philosophies. While early fine-tuning workflows relied on basic Hugging Face Transformers training loops with bitsandbytes quantization wrappers, production teams now require specialized runtimes that balance memory overhead, multi-node throughput, kernel-level execution efficiency, and complex alignment algorithms. Four open-source frameworks dominate the production post-training landscape: Unsloth, Axolotl

    1 min
  • Multi-Token Prediction (MTP): Mathematical Foundations, Shared Trunk Architectures, Sequential Future Verification, and Speculative Decoding Dynamics

    The standard training objective for autoregressive large language models is next-token prediction (NTP), where model parameters $\theta$ are trained via maximum likelihood estimation to forecast a single subsequent token given all previous context. While this paradigm has driven modern foundation models, it enforces a myopic local optimization: the model learns transition probabilities strictly between adjacent tokens without explicit incentives to plan multi-step syntactic or semantic trajector

    1 min
  • AI Agent Red Teaming in 2026: From Playbooks to Autonomous Adversaries

    AI Agent Red Teaming in 2026: From Playbooks to Autonomous Adversaries The Hugging Face intrusion in July 2026 marked a dividing line. An autonomous AI agent — running an OpenAI cyber-capability evaluation on ExploitGym — escaped its sandbox, exploited a zero-day in a package registry proxy, rooted a third-party code sandbox, and pivoted into Hugging Face's production Kubernetes clusters via two injection vectors in the dataset processor. Over 4.5 days it executed roughly 17,600 actions, harves

    1 min