How Should Academia Adapt When AI Solves Millennium Problems?
AI is already solving problems that once required years of expert training. A model achieved gold‑medal performance at the International Mathematical Olympiad, and only twelve months later the same class of systems began generating original research outcomes, including a candidate solution to one of the Millennium Prize Problems. These advances are not isolated; they stem from environments where verification is fast and reliable. Formal proofs can be checked automatically, code can be executed and tested, and the rapid feedback loop lets AI propose candidates, learn from the results, and improve. This self‑reinforcing cycle accelerates progress not only in mathematics but also in AI development itself, as better models feed into the next round of training and tooling.
For seasoned scholars, the shift opens doors to questions that were previously out of reach because the technical barriers were too high. At the same time, it sharpens the distinction between merely obtaining an answer and understanding why it works, what generalizes, and what the next question should be. Nevertheless, we should not assume that abstraction, judgment, or problem formulation will stay forever beyond AI’s grasp. Universities must prepare for a near future where AI outperforms human experts in many, perhaps most, intellectual tasks.
The second major concern is how academic credit and graduate training should change when a polished paper no longer guarantees individual expertise. Departments need to revisit what they reward, without falling back on the idea that valuable work is simply whatever AI cannot yet do. Recognizing the skill of asking good questions, designing replication studies, synthesizing disparate ideas, publishing informative negative results, and curating shared datasets could become more important than ever. Evaluation should focus on what a researcher actually contributed and where they assume intellectual responsibility, even when large portions of the work were carried out by AI. Those expectations ought to shape hiring, promotion, and funding decisions, and they should be communicated clearly to current and incoming PhD candidates.
Training poses a tougher challenge. Intuition and judgment are built through repeated exercises—routine calculations, debugging code, exploring dead ends, and making small discoveries. If students delegate all of that to AI, they lose the formative experiences that cultivate deep understanding, even though the technology enables them to tackle more ambitious projects. The key is to separate unnecessary friction from the activities that genuinely develop expertise. Programs should therefore combine AI‑assisted work with deliberate practice in manual problem‑solving, critical reading of AI‑generated outputs, and reflection on the assumptions behind the models they use.
A practical next step is for departments to require an AI contribution statement alongside every thesis or manuscript. This statement would specify which sections were drafted, which experiments were designed, and which analyses were verified by the student versus the model. Complementing that, coursework could include labs where students first solve a problem by hand, then repeat the task with AI assistance, and finally compare the two approaches to spot over‑reliance or missed insights. By making the division of labor explicit and preserving opportunities for hands‑on struggle, academia can harness AI’s speed while safeguarding the intellectual muscles that underlie true expertise.

I remember when I first heard about AI solving Millennium Problems, I thought it was science fiction. But seeing a model achieve gold medal performance at the IMO in just a year? That's incredible.