Article

Who Checks the Whole Map?

An AI-assisted counterexample overturned the three-dimensional form of an 87-year-old mathematical conjecture. The more revealing part was the expert questioning that followed.

In July 2026, a result landed that would have sounded like science fiction not long ago.

The Jacobian Conjecture dates to 1939. Over the complex numbers, it asks whether a polynomial map with a non-zero constant Jacobian determinant must have a polynomial inverse. Levent Alpöge announced a counterexample in three dimensions: a map that sends three different points to the same place. He credited Akhil Mathew with prompting the question and Anthropic’s Claude Fable 5 with work leading to the map. The result has since been independently formalised and verified. It refutes the conjecture in three dimensions and higher; the two-dimensional case remains open.

The obvious headline is that an AI helped overturn an 87-year-old mathematical conjecture in three dimensions.

That is astonishing. But it is not the most useful part of the story.

For me, the more revealing artefact is what happened next. Terence Tao published a detailed mathematical digestion of the counterexample, together with a public conversation with ChatGPT in which he used GPT-5.6 to interrogate its structure, reconstruct its geometry, compare alternative formulations, and check the calculations.

The transcript is a near-perfect example of why human expertise is not going away in the age of AI.

As raw analytical capacity becomes abundant, human judgement becomes the scarce input.

There was no magic prompt

Read the conversation and one thing becomes clear: this was not a clever sentence typed into a box, followed by a miracle.

Tao asks about triple covers, ramification, affine varieties, Schur complements, dilation weights, birational changes of variables, and the difference between local and global injectivity. He proposes interpretations, tests them, notices where an explanation still contains a “miracle”, and asks the model to make the non-injectivity transparent.

At one point, he asks whether the non-injectivity can be seen “in a transparent way”.

That looks like a simple question. It is not.

The question compresses an enormous amount of expertise. It assumes that the algebra has already been understood well enough to suspect a simpler geometric structure. It identifies what remains unsatisfactory. It sets a standard for the answer: not merely a correct calculation, but an explanation that makes the result intelligible.

Someone without the relevant expertise could copy those exact words and learn almost nothing. They would not know why that was the right question, whether the answer was adequate, which hidden assumptions mattered, or what to ask next.

This is what Every recently called the “expertise trap”: the prompt is not the source of the expertise, but its compressed output.

Processing is no longer the scarce input

If our defence of human relevance is that we can still process more information, calculate faster, or produce more words than a machine, we have already lost the argument.

Human processing is not the point.

AI can expand the polynomial. It can calculate the Jacobian determinant. It can search the literature, compare constructions, generate alternative explanations, and work through symbolic checks far faster than an unaided human.

What it does not remove is the need to decide:

  1. Which problem is worth solving? There are infinitely many questions an intelligent system could pursue. Expertise identifies the consequential one.
  2. How should the problem be framed? A bad representation can make a simple truth look impossible. Expertise changes the frame.
  3. What counts as an answer? A calculation, a proof, an explanation, and a useful decision are not the same thing.
  4. Where is the answer likely to be wrong? Plausibility is not truth. The ability to find the fragile step is earned through experience.
  5. What does the result mean? Facts do not arrive with their own significance attached. Judgement supplies it.

The machine supplied extraordinary horsepower. Tao supplied direction, standards, and steering.

That is expertise.

Judgement checks the whole map

There is an almost too-perfect metaphor hiding inside the mathematics.

The counterexample is locally invertible everywhere, yet globally it is not one-to-one. Every small neighbourhood looks right. The failure appears only when you understand the whole map.

AI output can fail in much the same way. Every sentence can be fluent. Every step can look plausible. Every local transition can appear coherent. And the answer can still be globally wrong.

Expertise supplies the domain model: what matters, what constraints apply, and where failure tends to hide. Judgement uses that model to evaluate the result against reality and evidence. Alignment connects both to purpose: whether the system is pursuing the outcome we actually intend.

Together, they keep asking:

  • Are we solving the right problem?
  • Is this actually true?
  • What would disprove it?
  • What has the model assumed without saying?
  • What are we now entitled to conclude—and what are we not?

Expertise is not simply knowing more facts. It is knowing where to point intelligence, how to constrain it, when to distrust it, and when the evidence is finally strong enough to act.

The data is beginning to show the same thing

This pattern is not limited to frontier mathematics.

Anthropic recently analysed roughly 400,000 Claude Code sessions. In a typical session, people made about 70% of the planning decisions—what to do, which approach to take, and what counted as done—while the model made about 80% of the execution decisions.

The more relevant domain expertise the person brought, the more the model did with each instruction. In sessions classified as expert, each prompt triggered about 12 agent actions and 3,200 words of output, compared with five actions and 600 words in sessions classified as novice. Sessions rated intermediate or above reached Anthropic’s strictest measure of verified success 28–33% of the time, compared with 15% for novice sessions.

This was an observational, classifier-based study, not a controlled experiment, so it does not prove that expertise alone caused the difference. But the pattern is exactly what the Jacobian conversation suggests: AI can absorb more of the execution while increasing the return on knowing what the work is for.

The person who understands the domain does not merely get a slightly better answer. They unlock more of the machine.

The apprenticeship paradox

There is, however, a harder problem.

If expertise becomes more valuable, how do we continue to produce experts when AI can remove so much of the friction through which expertise is formed?

Anthropic ran a separate randomised controlled trial with 52 mostly junior software engineers learning an unfamiliar Python library through two coding exercises. The AI-assisted group scored 50% on a later comprehension test, compared with 67% for the group that coded by hand. The largest gap was in debugging: precisely the skill required to recognise when generated code is wrong and understand why.

The AI group finished only about two minutes faster, and that difference was not statistically significant. The study was small, measured immediate comprehension, and examined two coding exercises using one unfamiliar library; it would be wrong to generalise it into a universal law. But it is a useful warning.

The danger is not that using AI automatically makes us less capable. Among the AI-assisted participants, the stronger performers asked conceptual questions, requested explanations, and checked their own understanding. The danger is using AI to escape the difficulty through which judgement develops.

Getting stuck is not always wasted time. Debugging is not merely a delay on the way to the answer. The struggle changes the person doing the work. It builds the internal models that later let them recognise when a confident machine is confidently wrong.

If we delegate every formative difficulty, we may gain short-term output while accumulating long-term competence debt: more AI-generated work, and fewer people able to judge it.

Use AI to deepen judgement, not replace it

The answer is not to use less AI. The gains are too real, and the frontier is moving too quickly.

The answer is to use it differently.

Before asking for an answer, form a view. State what you think is happening and why. Ask the model to identify the weakest assumption, produce the strongest counterargument, and tell you what evidence would change the conclusion.

When learning, ask for explanations rather than only outputs. When the model finds an error, make sure you can explain the error yourself. When a result is surprising, slow down. Ask for boundary cases, alternative derivations, and independent checks.

Most importantly, never outsource the definition of done.

The final responsibility for truth, relevance, and consequence remains with the person directing the system. That responsibility cannot be delegated by typing a better prompt.

Keith Rabois’s “barrels and ammunition” framework, as explained by Conor Dewey, argues that organisations are constrained not by the amount of raw talent they can assemble, but by the people who can take responsibility for an idea and carry it all the way to completion. AI makes that distinction sharper. It places vast amounts of intelligence on tap, but it does not decide what is worth building, what trade-offs are acceptable, or what reality requires.

As implementation becomes abundant, judgement becomes the constraint.

The scarce human input

The people who thrive in the age of AI will not be those who try to out-process the machine. Nor will they be those who merely learn to ask it for more.

They will be the people who understand a domain well enough to recognise the important question; who can distinguish a locally plausible answer from a globally true one; who know when to trust the tool, when to challenge it, and when to go back to first principles.

The machine can produce an answer. Expertise is knowing whether it answered the question.

artificial intelligenceexpertisejudgementlearning