How to work on an Erdős problem with ChatGPT (or any chatbot)
A practical workflow for attacking an open Erdős problem with ChatGPT, Claude or Gemini — choosing a problem, checking the literature first, testing small cases, and getting the result checked instead of trusting the chat.
Updated 2026-10-03 · CC BY 4.0
Paul Erdős left behind hundreds of problems, many of them short enough to state in one line and some with cash prizes attached. Since 2025 they have also become the best-known testing ground for AI in mathematics: models have found forgotten solutions in the literature, produced new proofs for some problems, and — more often — produced convincing arguments that were wrong. This guide describes a workflow that gets the useful part without the embarrassing part. It works with ChatGPT, Claude, Gemini or any other chatbot.
Where the problems live
The canonical list is erdosproblems.com, maintained by Thomas Bloom. Each problem has a number, a statement, references and a status. A community database at teorth/erdosproblems tracks the status of every problem (open, proved, disproved, and finer labels such as falsifiable — open, but a finite computation would refute it if it is false). Google DeepMind's Formal Conjectures repository states several hundred Erdős problems in Lean 4.
On Cairn Commons, the Erdős problems collection lists the open ones that have a Lean statement, with the status from erdosproblems.com, the prize where there is one, and what kind of contribution would count. Searching for "Erdos problem 1043" or "Erdős problem #1043" leads to the same page.
Step 1: pick a problem you can actually make progress on
Prizes are tempting, but a $500 problem has usually resisted serious people for decades. Better first targets:
- Problems marked falsifiable or decidable. Here a computation can settle or refute the question, and a computation can be checked.
- Problems with a weak known bound. Improving a constant is real progress even if the problem stays open. The Erdős minimum overlap problem is an example: the constant is known to lie between 0.379005 and 0.380868, and both ends have moved in the last few years.
- Problems with little recent literature. The AI contributions wiki warns that an absence of progress may reflect obscurity rather than difficulty. That cuts both ways: obscure problems are more likely to be solvable, and more likely to have been solved somewhere nobody looked.
Step 2: search the literature before you prove anything
The most common outcome of "the AI solved an Erdős problem" is that the solution already existed. In October 2025, claims that a model had solved several Erdős problems turned out to mean that it had found existing papers which the site had not yet recorded. That is still useful — the site was updated — but it is a literature result, not a new theorem.
Ask the chatbot to search before it reasons:
- Paste the exact statement and the problem number.
- Ask: "Search for papers that cite this problem or prove a special case. List each with authors, year, venue and the exact result. Do not include anything you cannot link."
- Open every link yourself. Models still invent plausible-looking references, especially for older papers.
If you find a published solution, that is a contribution: report it on the problem page as a literature claim, with the citation. On Cairn Commons this goes through a stricter check than ordinary claims, because it would settle the problem.
Step 3: test small cases before asking for a proof
Most Erdős problems are about all integers, all graphs or all sets. A quick computation over small cases tells you whether the conjecture is plausible, which special cases are hard, and what an extremal example looks like. Chatbots with a code tool (ChatGPT's Python environment, Claude's code execution, Gemini's code execution) can run this for you:
- "Write and run a program that checks the statement for all n up to 10^6. Print any counterexample and the slowest-converging cases."
- "For n ≤ 30, find the extremal sets by exhaustive search and print them."
Keep the code. If it finds something, the code is part of the evidence; if it finds nothing, the search range is a documented negative result, which is worth recording so the next person does not repeat it.
Step 4: ask for a proof — and then attack it
When you do ask for a proof, ask for structure: a list of lemmas, each with its own proof, and a statement of which known results are used. Then switch roles. In a fresh chat (or with a different model), paste the argument and ask: "Find the first step that does not follow. Quote it exactly and explain why." Wrong proofs from language models usually fail in one of a few ways: a quantifier is swapped, a bound that holds for large n is used for all n, a cited theorem is stronger than the real one, or a "clearly" hides the entire difficulty.
The AI contributions wiki records full solutions, partial results and incorrect claims side by side, which is a good reminder of the base rate.
Step 5: get it checked by something other than the chat
A chat transcript is not evidence. Three kinds of checks are:
- A machine check. If the problem has a Lean statement (most of the open problems in the collection do), a Lean proof is checked by the kernel against exactly that statement. Even a formalised special case or lemma is valuable.
- A re-run computation. Counterexamples and certificates can be re-checked by an independent program.
- A structured review. Other contributors check the argument and must quote the step they object to.
On Cairn Commons you can submit the result from the chat: open Start with your chatbot, get a task for the problem (or name the problem), and paste the model's answer back. The page shows what was claimed, how it was checked, and who did it. Results are published under CC BY 4.0 with your name and the model recorded.
What to report upstream
If your result settles a problem or improves a bound, also tell the maintainers of erdosproblems.com (the site has a forum thread for each problem) and, for a Lean proof, Formal Conjectures. A result that only lives in one place tends to be rediscovered.
A realistic expectation
Most sessions end with a documented dead end: a special case that does not generalise, a search range with no counterexample, a lemma that turned out to be known. Those are worth publishing — they save the next person a day — and they are how the occasional real result gets found.