- Meta released six papers on October 2 in which mathematicians used Muse Spark 1.1 and 1.2 through the ordinary meta.ai chat, with AI-drafted passages labeled and a separate group of mathematicians reviewing each result.
- Meta says five papers answer previously open questions, and it discloses that three results were also found independently, one of them by an AI agent called Nilradical on September 16.
- The release contrasts with OpenAI's September claim of more than 100 solved problems without published proofs, and it lands as arXiv caps submitters at two papers a month.
Meta's six mathematics papers written with its Muse Spark model are the most transparent public record yet of how a frontier AI system contributes to research, and Meta's own footnotes show that three of the six results were reached by other teams within weeks. Open problems within AI's reach are now being cleared by several groups at once, which moves the contest in mathematics toward credit and checking.
Mathematicians used Muse Spark through the ordinary meta.ai chat to write six papers
The work ran on the consumer product. According to Meta AI Research, teams of mathematicians picked the problems and steered the work while using Muse Spark 1.1 and 1.2 in Thinking Mode, the same interface any meta.ai user sees. A second, separate group of mathematicians then reviewed each finished paper, and every paper marks which passages the researchers drafted and which the model drafted.
The model's role varied by paper. In group theory, Muse Spark wrote a search program in the GAP algebra system that found a counterexample inside a group of 384 elements, disproving a 2024 conjecture by M. Kida. In arithmetic physics, it generated candidate proofs and drafted three core technical sections of a paper extending Yuri Manin's link between p-adic string theory and number theory. In differential equations, the collaboration proved that a class of waves in the biharmonic nonlinear Schrödinger equation collapses in finite time, settling a question open since 2015 and confirming simulations published in 2002.
“Our goal here wasn't to mass-produce papers, but to empower researchers and help them develop mathematical insights that others can understand and build on.”
Meta AI Research, Solving Open Research Problems Together, October 2, 2026
Alexandr Wang, who leads Meta's AI effort, announced the six titles on X the same evening.
Three of the six results were reached by other teams within weeks
Meta's blog credits concurrent work in half the cases. Three independent groups published the same sharp threshold for fitting random Gaussian points to an ellipsoid in August 2026. An AI agent called Nilradical reported a different counterexample to Kida's conjecture on September 16. Researchers named Hu and Wen independently found counterexamples to the evolution-algebra conjecture that Meta's sixth paper disproves. Machine learning theorist Jason Lee summarized the point on X on October 3, writing that “actually 3 of 6 were already resolved.”
The overlap is mostly a statement about the problems. The results that collided share a profile: questions posed recently, often within the last two years, where the answer turned on finding a concrete counterexample or a sharp numerical boundary that a model can search for. Imperial College mathematician Kevin Buzzard described the same pattern in July, writing that human mathematicians are being “outcounterexampled” by AI systems. When many labs and many human-plus-model teams point similar tools at the same recent conjectures, they arrive at the same answers in the same month.
| Papers released | 6, on October 2, 2026 |
| Answer previously open questions | 5, by Meta's count |
| Also reached independently | 3 of 6, disclosed by Meta |
| Kida counterexample | A group of 384 elements, found by a Muse Spark search program |
| Models used | Muse Spark 1.1 and 1.2, Thinking Mode, via meta.ai chat |
Meta's labeled-passage method answers the verification problem OpenAI left open
Two weeks earlier, OpenAI said an internal model had resolved more than 100 open problems, including a claimed Navier-Stokes solution, and released no proofs for checking. Meta took the opposite route on every count that matters to working mathematicians. It published full papers, named the human authors and reviewers, labeled machine-drafted text and cited competing results that weaken its own priority claims.
A paper that marks which paragraphs a model wrote can be checked, argued with and built on. A count of solved problems can only be believed or doubted.
The method still has limits. Meta selected the reviewers, so their sign-off is closer to an internal audit than to journal peer review, and the company has not released the chat logs, which leaves the number of failed attempts and wrong turns unknown. The papers show what succeeded and leave out how often the model misled the researchers along the way.
Credit and checking become the scarce resources in AI-assisted mathematics
For mathematicians, the practical consequence arrives first in priority. A result on a recent conjecture now has a short shelf life, because another team with a capable model may post the same answer within weeks, and researchers will need to search preprints and agent reports before they claim anything as new. That pressure meets a publishing system already straining under generative AI output: arXiv has limited each submitter to two papers a month since October 1 after submissions doubled in two years.
For AI labs, mathematics has become a public benchmark of research ability, and the six papers show which format holds up. Disclosed collaboration, labeled authorship and credited concurrent work produce results that survive scrutiny, including scrutiny that cuts Meta's headline from six solved problems to a smaller number of uncontested ones.
The lasting value of Meta's six papers lies in the record they leave of how each answer was reached. As capable models make recent open problems solvable by several teams at once, the lab whose results mathematicians can trace line by line will hold more standing in the field than the lab that announces the largest count.
Santage is committed to independent, transparent journalism. This article is produced in accordance with Santage's Editorial Standards and aims to provide accurate and timely information. Readers are encouraged to verify information independently.