Columbia Researcher Claims GPT-5.6 Sol Solved Six Open Erdos Problems

0
100

A PhD candidate at Columbia University says he solved six previously open problems from the Erdos problems collection in about five days, using OpenAI’s GPT-5.6 Sol together with the Codex coding workflow. The results have not been peer reviewed or independently verified, and Shouqiao Wang, the researcher behind the claim, has asked others to check his work rather than take it on faith.

Wang, an alumnus of Peking University who is now studying at Columbia University, posted the claims on the official Erdos Problems forum sometime around July 14 to 20, 2026. He says he attempted about 13 problems and got what he considers full solutions on six of them, a success rate of roughly 46 percent. The six he lists are problems 390, 486, 536, 788, 1002, and 1038 on that forum. He also described the run in a post on X, where it drew quick attention from mathematicians and AI researchers who wanted to see the receipts.

The tooling is a big part of why the story spread. Wang says he leaned on OpenAI‘s GPT-5.6 Sol paired with the Codex workflow, the same kind of setup a software developer might use to write, run, and test code. In his own words on X, Wang said he has a math background but that the Codex workflow he used does not require deep mathematical knowledge. That claim is what raised eyebrows, because it points at a general model and a coding harness doing the heavy lifting rather than a specialized theorem prover built by mathematicians for exactly this task.

What sets Wang’s post apart from a bare announcement is how much he published. He posted detailed materials to GitHub, including the workflows and prompts he used, proof PDFs, LaTeX sources, Python code, and Lean formalizations. Lean is a proof assistant that makes a computer check each logical step, so a clean Lean formalization gives reviewers a mechanical way to test at least parts of an argument. By putting the prompts and the code in the open, Wang has made it possible for other people to try to reproduce the runs on their own machines. The Chinese technology outlet 36Kr walked through the episode in a report published on 36Kr, which helped push the story beyond math circles.

None of that makes the six results correct. That is the part worth sitting with. Not one of the proofs has been confirmed, and the reaction across the math community runs a wide range. Some researchers are openly excited about what a general model might do for everyday research work. Others are firm that the truth of each result has to be established before anyone celebrates. Posting proofs to a forum and to GitHub is the opening move in that process, not the finish. Peer review, independent checking, and a careful read of the Lean files all still have to happen, and any one of them could turn up a gap.

There is also a claim floating around that needs to be kept well away from this one. Separately, a viral story circulated that GPT-5.6 Sol had solved Erdos Problem 119. That is a different and unrelated claim, and it stayed listed as open and unverified. Wang’s post says nothing about problem 119, and his six problems have nothing to do with it. Blending the two together would misstate what Wang actually claimed, so it is best to treat them as two distinct threads that happen to involve the same model.

Some background helps explain the stakes. The Erdos problems are a curated set of open questions tied to Paul Erdos, one of the most prolific mathematicians of the twentieth century. Many of them are easy to state and stubbornly hard to settle, which is exactly what makes them a useful test for any new method. If even a portion of Wang’s six proofs holds up under scrutiny, it would be an unusually concrete case of a general purpose model helping to close open questions in pure mathematics, and it would do so with a paper trail that anyone can inspect. Wang’s own account stresses the workflow more than any single clever step, which is part of what makes the release interesting to people who study how these models reason.

For now the honest summary is narrow. A Columbia researcher has made a striking claim, shown his full working, and asked the field to verify it. Wang has done the part that a lot of viral AI claims skip, which is to hand over the prompts, the code, and the formal proofs so that skeptics can dig in. Whether the six proofs survive that digging is the open question, and it is the one that will decide how much this episode really tells us about AI and mathematics. Until the checking is done, the safe reading is that these are claims, carefully documented, and not yet confirmed.

EntrelligenceFree guide
Your First 10 AI Skills

Your First 10 AI Skills

10 practical AI skills, copy-paste prompts and a 7-day plan to start using AI with confidence.

Download the guide →
0 0 votes
Article Rating
Subscribe
Notify of
guest
0 Comments
Oldest
Newest Most Voted