Blog/Article

The Frontier Becomes an Eval

OpenAI dumped solutions to hundreds of open maths problems and mathematicians aren't thrilled. Ted Chiang saw elements of this in 2000, the same feelings are coming for many of us, the future is going to get weird

·4 min read
The Frontier Becomes an Eval

OpenAI recently released solutions for hundreds of major open mathematical problems across multiple fields. These solutions (or at least proposed solutions) were generated by a new internal model - averaging just 3 hours GPTPro time each. This is in addition to, but in contrast with, the proposed proof of the Millennium Prize Navier-Stokes problem - that used 10,000 concurrent agents over 88 hours.

OpenAI's Navier-Stokes announcement

This "mathocalypse" has triggered surprising, funny, and worried reactions. Some other professions, notably software development, have been going through something similar. In these highly verifiable domains the power of modern AI systems is expanding beyond the human frontier - arguably not in brilliance or beauty, but clearly in sheer speed and usefulness.

Some mathematicians have called this one of the biggest single days in mathematical history, others a nadir.

"Releasing over 700 files at once is not a demonstration of scholarship, but a demonstration of power."

"Mathematicians did not ask for this work to be done." – Association for Human Mathematics, Communications Working Group

The Association for Human Mathematics statement on OpenAI's release

The problems of their (incredibly difficult) profession are being automated away by models and their possibly non-mathematical prompters. The frontier is becoming an eval. You can understand their sadness and frustration - it is not only the solutions, but also how this is being done. Terence Tao, often called the world's preeminent mathematician, has already spoken of Maths 2.0, a need to recentre maths away from solutions and towards holistic understanding and explanation.

I read Ted Chiang's short story collection "Stories of Your Life and Others" recently, it has a one page story framed as a journal article from 25 years after the last original human submission. In the story it is metahumans - genetically advanced humans - who progress science and technology far beyond "normal" human comprehension. In ours it is AI.

"No one denies the many benefits of metahuman science, but one of its costs to human researchers was the realization that they would probably never make an original contribution to science again. Some left the field altogether, but those who stayed shifted their attentions away from original research and toward hermeneutics: interpreting the scientific work of metahumans." – The Evolution of Human Science, Ted Chiang

In the story humans struggle to even comprehend never mind explain the advances. Those who try find some solace.

"…the scientific tradition is a vital part of that culture. Hermeneutics is a legitimate method of scientific inquiry and increases the body of human knowledge just as original research did. Moreover, human researchers may discern applications overlooked by metahumans, whose advantages tend to make them unaware of our concerns." – The Evolution of Human Science, Ted Chiang

I must say I side against our mathematical scholars here. They sound like a medieval guild protecting their space. They have been caught out by the incredibly rapid AI advance in their domain. Some of the points are fair (papers no one can follow, retractions already), but this seems to be a verification issue and then the "hermeneutics" job Chiang describes, for sure not a reason to stop solving problems. Can you imagine biology or medical science responding in this way to solutions to their problems? AlphaFold already earned a Nobel.

As I have written about before, the AI labs know that solutions in human biology are the Holy Grail. They solve disease and worries over revenue and data centres evaporate. Anthropic has a biology division, Demis Hassabis has Isomorphic Labs. This is a tougher domain to crack, but the capital and thinking power focused on it makes me very bullish that we see incredible advances over the coming years.

AI is coming for all domains, and past the frontier we may lose full understanding of the how, but if the why remains hugely beneficial for humanity this is a price worth paying. (Of course, if we can't understand their science, other questions of the existential nature arise that I won't go into right now :) )

"We need not be intimidated by the accomplishments of metahuman science. We should always remember that the technologies that made metahumans possible were originally invented by humans, and they were no smarter than we." – The Evolution of Human Science, Ted Chiang