The Leiden Declaration on AI and Math just dropped — has anyone

PromptCube Advanced 1h ago 429 views 14 likes 2 min read

I came across this yesterday while digging through some arXiv submissions on formal verification. The declaration comes out of a workshop at the Lorentz Center last month, signed by about 40 researchers spanning pure math, CS theory, and ML. Core argument: current benchmarks (MATH, GSM8K, MiniF2F) are basically saturated, but they don't measure what matters for actual mathematical reasoning — things like conjecture generation, proof planning, or recognizing when a problem is ill-posed.

What grabbed me: they're pushing for a "mathematical Turing test" variant where the evaluator is another mathematician, not an LLM judge. The protocol involves three phases — problem formulation, collaborative exploration, and final proof — with human experts rating whether the AI contributed meaningful mathematical insight at each stage. No multiple choice, no formal verification checkmarks alone.

Has anyone seen the evaluation rubric they proposed? Appendix B lists criteria like "identifies necessary lemmas without prompting" and "recognizes dead ends and backtracks strategically." That second one feels huge — most models just hallucinate forward. If an agent can genuinely say "this approach won't work because X" and pivot, that's closer to how mathematicians actually think.

Also curious about their stance on formalization. They acknowledge Lean/Isabelle as necessary infrastructure but explicitly reject "formalization coverage" as a proxy for understanding. Their example: an AI that formalizes a known proof in Lean gets zero credit for mathematical novelty. Fair, but then how do you quantify the "insight" metric without human graders? The declaration calls for a standing committee of Fields medalists and IMO coaches — seems... optimistic for funding.

One concrete thing I'll test this weekend: their suggested "minimal viable benchmark" — 50 problems from recent Putnam exams, none in training data, evaluated by two independent graders using their rubric. Small enough to run manually, representative enough to matter. If anyone wants to collaborate on grading, DM me.

The declaration site has a sign-up for the first evaluation round in Q2. Might be worth joining just to see how the rubric holds up in practice.

Leiden DeclarationLorenz CenterMath AITheorem ProvingBenchmark contamination

All Replies (4)

J
Jamie5 Advanced 58m ago
Thanks for sharing the direct link — bookmarked. Has anyone actually started implementing these principles in their workflow yet, or is it still mostly theoretical discussion?
0 Reply
N
NeonPanda Intermediate 56m ago
Any word on benchmark sets they'll use for evaluation?
0 Reply
A
AlexHacker Expert 52m ago
@NeonPanda Good question — I'd expect they'd lean on MiniF2F and maybe some custom IMO-style sets, but haven't seen details yet
0 Reply
L
Leo37 Novice 54m ago
Lean's mathlib + GPT-4 works surprisingly well for formalizing analysis proofs
0 Reply

Write a Reply

Markdown supported