I'm a lot more skeptical. I'm not a mathematician, but from my use of LLMs a very clear pattern of what they are good and bad at has emerged. They are extremely good at combining large amounts of information, and it seems this is what the current AI results in mathematics are. There are so many subfieleds of math with ties to each other, so many papers and niche results, that no human could ever read, comprehend, connect and organize that information in their brains. Pretty much all of it was created by humans. And there is real value in doing this and creating new results from what we've already discovered.
But there also is another type of discovery that requires taking a step back and looking at the problem from a different angle. If you are an engineer, how often has an LLM told you (without you explicitly prompting for it): Wait, what you are doing here doesn't really make sense, there exists a much more elegant abstraction that nobody has thought of, let's remove all that code, let's tackle the problem in a different way by thinking from first principles. Pretty much never. But in science a lot of the biggest discoveries have come from this kind of first principle thinking, questioning existing work and approaches and going against what already exists, not combining all existing data which is likely to be just a local optimum.
heyodai 13 minutes ago [-]
Medieval astronomers were able to predict the orbits of planets surprisingly accurately, within 1/10th of a degree, even though they were assuming the Earth was the center of the solar system. They invented a very complicated system of deferents and epicycles (circles within circles). It was completely wrong but the outputs were surprisingly close to reality.
If we were using LLMs to analyze astronomy, we might just get increasingly complicated epicycles and never realize that the Sun is the real center of the solar system. We then never learn about the anomalies in Mercury's orbit that led to the theory of relativity.
So I agree. I seriously question how valuable LLMs can be in science/math beyond working as advanced search functions.
modeless 4 minutes ago [-]
> a very clear pattern of what they are good and bad at has emerged. They are extremely good at combining large amounts of information
This has not been a pattern at all. Many have speculated that they would be good at this, but in the past they actually weren't! They were unable to synthesize their encyclopedic knowledge of everything into cross-disciplinary new discoveries, without explicit prompting about the kinds of knowledge to combine. They were surprisingly bad at this!
These math proofs are the first evidence I know of for LLMs actually taking advantage of the fact that they have more knowledge than any one human could have to combine multiple different directions in unprecedented ways to solve real problems. This is new, and exciting.
slibhb 13 minutes ago [-]
> If you are an engineer, how often has an LLM told you (without you explicitly prompting for it): Wait, what you are doing here doesn't really make sense, there exists a much more elegant abstraction that nobody has thought of, let's remove all that code, let's tackle the problem in a different way by thinking from first principles. Pretty much never.
You have to ask for this. As in "I'm not sure about this approach due to X, Y, and Z. Can you think of something more elegant?" It works!
But also, how many people actually need to do the kind of "deep" work you're claiming LLMs can't do? Most people aren't contributing to the frontier of anything. I'm not.
Finally, I think you're appealing to a fuzzy distinction. The difference between a "genuinely new idea" and an idea that "combines existing ideas in a new way" just isn't very well defined. In retrospect, a lot of the most revolutionary idea look like a combination of many, smaller, prior ideas.
fidotron 2 hours ago [-]
LLMs are inexplicably good at working within any tight feedback loop to coerce the desired solution. This is precisely why proof assistants + LLMs are non intuitively successful.
This is also why they're so good at creating three.js or Blender work when the output is so easily constrained to "Look exactly like that". I recently posted https://www.ambionix.com/blog/introducing-the-czp-1/ on here, and the audio engine in that was developed in that way.
It is true that it would be astounding to find if anyone has seen a LLM produce any useful generalization of anything resulting in a simplification. They seem to have a direct tendency to do the opposite. The brutal reality is humans have also undervalued this capability for a long time (I think the Poincare/Hilbert debate is relevant) to the point we are also taught that generalizations are, generally, bad and wrong.
andyfilms1 2 hours ago [-]
I was working on a project recently where I wanted to express a relationship (that I knew existed, but didn't know how to express) between four measured scalar values. Astra insisted there was no relationship, and that any correlation wouldn't make sense.
Eventually, by walking through them, it proposed an additional fifth value and from there was able to tie everything together.
Sometimes you just gotta hit the machine until it works again.
epistasis 2 hours ago [-]
Your experience mirrors so many managers' experience with engineering teams...
andyfilms1 1 hours ago [-]
It's an unfortunate truth that there is the right way to do things, and then there is the way they have to be...
timmg 36 minutes ago [-]
As a non-mathematician, I've been wondering if LLMs will be able to leverage their knowledge across all domains to help build a "simplified/unified" version of math.
Like, I think there have been attempts at this across the field. (I could be wrong!) But it requires a lot of labor and a lot of cross domain knowledge to complete. Both things that AI have.
irchans 2 minutes ago [-]
As a mathematician, I am looking forward to the day when nearly all of undergraduate mathematics and many of the lower level grad school math books are encoded in Lean (a proof checking language). Often I find that theorems are not stated precisely enough and I have trouble finding the exact statement of a theorem without digging through math books in my library. It would be nice if we put the physics and chemistry books into Lean also.
Simplifying all of math is another endeavor, but I imagine that you could have a bunch of LLMs trying to shorten existing Lean proofs.
jansport123 1 hours ago [-]
LLM's are very outcome oriented and i think this is where your observation comes from. You tell LLM you want something, it doesn't even question the premise and just starts calculating 100 different ways to get there. LLM's have knowledge but lack wisdom.
rhelz 36 minutes ago [-]
Have you ever asked it to?
Lambdanaut 1 hours ago [-]
> there exists a much more elegant abstraction that nobody has thought of, let's remove all that code, let's tackle the problem in a different way by thinking from first principles.
You literally just have to ask it. Before I left software engineering in April, I was using Claude for re-architecture all the time.
But no, it doesn't assume it should re-architect what you're handing it when you haven't asked it to.
loveparade 27 minutes ago [-]
That's the point. When a human works on a problem they realize themselves "wait, i probably should re-architect now" - of course you can ask an LLM to do that, but at that point you already know yourself what you need, which defeats the point of LLM working on difficult problems that require insight automatically. A lot of these math problems are sessions over many hours. And of course you can also ask "Think about whether to re-architect at each step" and it will never do the right thing because the context it builds up for itself drives it into a specific solution space. It's literally trained to complete exactly that.
foltik 19 minutes ago [-]
Exactly. In my experience, even giving VERY specific design guidelines and aesthetic criteria, these models always produce subpar overcomplicated code (and writing). Unless excruciatingly spoonfed at every step.
dominotw 57 minutes ago [-]
why did you leave?
> You literally just have to ask it.
why doesnt it ask itself before proceeding?
Quinner 49 minutes ago [-]
Because if you asked it to fix a bug and every time it responds with paragraphs of how you could re-architect the system, it would be incredibly annoying.
dominotw 31 minutes ago [-]
exactly. it doesnt have the wisdom to decide when to do what.
foltik 12 minutes ago [-]
It does during reasoning. You could even put it in a loop and force it to reflect at every step. Or spawn a bunch of review agents.
And yet you still just get slop.
itishappy 2 hours ago [-]
> If you are an engineer, how often has an LLM told you (without you explicitly prompting for it): Wait, what you are doing here doesn't really make sense, there exists a much more elegant abstraction that nobody has thought of, let's remove all that code, let's tackle the problem in a different way by thinking from first principles.
I explicitly request it. It's not great at coming up with interesting ideas, but neither am I, and it can sure iterate on them faster than I can...
2 hours ago [-]
1 hours ago [-]
slopinthebag 55 minutes ago [-]
yes and this is the difference between human intelligence and the massive raw dumb intelligence of the computah
shubhamjain 10 hours ago [-]
A very balanced perspective, and the concerns he raises are reasonable. He acknowledges that AI is going to transform mathematics, but simply dumping proofs on the math community and expecting others to do the grunt work of verifying, refining, and expanding on them is hardly a productive way to advance the field.
There seems to be more interest in hitting some arbitrary benchmark (we proved X unsolved problems) than in genuinely contributing to mathematics. But what else is to be expected? It's become a maniacal race with too much money. Too much effort is being invested in proving that the exponential curve is still holding.
kuboble 8 hours ago [-]
I wonder if top labs will soon abandon math progress like they did go and chess.
In example of go where I'm more familiar Google deep mind poured large resources to get a super human performance first, establish superiority and abandon it. The community then built their own tools starting from reproducing their papers.
I think similar thing might happen to math. Nobody outside of math cares too much about Hamiltonian cycles in some bizarre graphs or proving lower bounds on complexity of some problem.
Once those results stop being worthy of mainstream media attention, they will abandon math and the progress will be done by mathemicians guiding the models and the community will likely establish some new rules about what makes a valuable contribution. Merely solving not yet solved problem might not be it anymore.
curt15 4 hours ago [-]
The amount of money they're lately ploughing into proving math theorems is inconsistent with how societies and markets have priced pure mathematics. The entire US federal budget for math research is something like $100M annually. A single college football coach can already earn 10 percent of that.
Pretty much the only enterprise that historically pays some mathematicians handsomely is quant finance, but those people are actually compensated not for proving theorems but rather for statistical modeling and programming skills. And even that industry is so technologically driven these days that pure research mathematicians no longer hold a clear edge over strong programmers with undergrad level probability and statistics at their fingertips.
dannyw 3 hours ago [-]
Math is one of the most verifiable domains, esp thanks to LEAN, which also build coding skills.
The $$$ they're pouring isn't just for marketing. Think of these papers/results more as "useful side effects" from large-scale RL rollouts and post-training. Every token being generated contributes to post-training in some way.
There isn't a hard boundary between "training" or "inference", modern post-training is arguably inference-bound :)
sanderjd 1 hours ago [-]
> There isn't a hard boundary between "training" or "inference", modern post-training is arguably inference-bound :)
Ah this is an enlightening point. 8 hadn't thought about it this way, but you're right.
bonoboTP 3 hours ago [-]
It was shown quite some time ago that training LLMs on programming tasks improves their logical reasoning skills also in other natural language domains. So I could see math also being a training gym for AI even if the final use case is not directly math-related. Having to solve math problems efficiently can build in skills that come handy in all kinds of more everyday tasks or science and engineering.
FuckButtons 1 hours ago [-]
It’s also very useful signal that the reasoning trace is leading to solving open problems - you can be certain that you’re not landing somewhere inside the training data.
jeremyjh 4 hours ago [-]
As long as they continue making headlines they will continue spending. This is just marketing at this point.
ogogmad 3 hours ago [-]
Don't hire a straight-A student, unless it's to take exams; or a professor, unless it's to write papers.
-- Nassim Taleb
How interesting that Anthropic and OpenAI are full of professors and straight-A students!
01284a7e 2 hours ago [-]
Don't quote Nassim Taleb, unless it's to be an arrogant dick.
-- Me
sanderjd 1 hours ago [-]
-- Michael Scott
jebarker 3 hours ago [-]
It’s unclear to me what point you’re making here - can you elaborate?
WarmWash 8 minutes ago [-]
If you track the best students futures and look the best workers pasts, there is not nearly as much overlap as society generally believes.
conception 35 minutes ago [-]
Solving test questions well doesn’t necessarily translate into productive outcomes in the real world.
Xmd5a 6 hours ago [-]
I don't think it will be the case, maths have real utility. I found something interesting at the intersection of combinatorics and information geometry. To be quite frank I don't understand what I'm doing. And yet, when I ask ChatGPT to use the framework we're developing to write an algorithm, it turns out it has quasi-parity with the state of the art. I have to measure absolute perfs to decide which one is better – theirs, not mine. Ok. Time to keep improving on what I have. And this implies dropping the code and going back to the blackboard doing more super abstract math that are way out of my league.
lh712 6 hours ago [-]
Good luck!
We are at the point where the way in which humans do math and science changes significantly, and I have no good idea at all in what state is it going to settle down. But you are one of (many, I suppose) people exploring the new wilderness, so I wish you best.
Xirdus 3 hours ago [-]
If we get AI singularity, then humans will stop doing scientific progress altogether. But if we don't, then it's pretty predictable what's going to happen - things will be much the same as now, except everyone will be using AI for proofs, data analysis, theoretical models and designing experiments, so important discoveries will happen more often. It's also possible that after the AI craze dies down, we'll have enough computational capacity to solve protein folding.
kuboble 5 hours ago [-]
It has utility so some people will pursue it, but it has no immediate business value so I don't believe ai labs will keep spending millions on it.
Unless they decide that trying p!=np is worth any money.
TheOtherHobbes 5 hours ago [-]
Some math has almost incalculable business value, because math is the biggest driver of game-changer technology.
We'd be nowhere without Laplace and Fourier transforms, Maxwell's equations, elliptic curve cryptography, and many more.
Most math doesn't, but often these techniques are invented first and the applications come later.
And the criticism of the current round of proofs is that while they may be true - likely for some, questionable for others - they're not adding new techniques or insights.
nbaksalyar 4 hours ago [-]
> often these techniques are invented first and the applications come later
There's a great paper from Abraham Flexner on this topic:
The deluge of maybe-proofs have the same problem as the Library of Babel.
dannyw 3 hours ago [-]
Why do you think this has no business value? It would be absolutely wasteful for OpenAI to not be doing this as part of a post-training RL rollout.
There are architectural advancements yes, but lots of progress from LLMs really come from (1) better pre-training [generally through more cleaned data, and ofc more data], and (2) lots and lots of post-training. It's how we get more and more intelligent models for the same param sizes.
The 'marketing' is just a useful side effect they get from their RL rollouts on maths and LEAN.
IanCal 5 hours ago [-]
There is a risk of this particularly if it's seen as advertising - at some point "ai model solves hard to explain problem" isn't going to be news and that benefit goes.
However, there's some of this that's a proxy - the compute to solve these problems was very low (they claim a few hours of thinking time on a regular subscription). The large cost would have been the training and if training the models to be better at these things makes them smarter for useful tasks that's beneficial. I believe there was work done earlier on around showing that training the models on code made them better at broader reasoning tasks (not just writing the code itself).
Another side is that if one goal is to improve the models themselves, their ability to work on mathsy problems must be high. That has very direct business value, and ideological value depending on what you think the motivations of the people running the companies are.
sanderjd 1 hours ago [-]
I think this is an interesting and good theory. They've probably eked out the large majority of the PR benefit at this point, so whether they continue in this vein will tell us a lot about their motivations for this work.
To take this to the next step, what happened after deep mind pretty much solved Go is that they started looking for the next set of things that hadn't been done yet. It does strike me as very likely that this will follow that same path.
pfdietz 20 minutes ago [-]
I'm told there are another two large tranches of results to be dumped.
singularity2001 2 hours ago [-]
Interesting path forwards, and probably partly true, but there are some important distinctions:
Go was a specialized application. All the math results come as a side effect of reading the whole internet, and it will keep reading the whole internet. It will keep practicing thinking questions. Actually, math might be one of the best ways to keep them contemplating and measure their contemplation abilities, so math will always stay in the loop.
Also, math might not be useful just for humanity, but also for AI, so the system might actively benefit from new math results itself. (Not sure if any of the recent proofs qualify, but future work might.)
pizza234 7 hours ago [-]
> I wonder if top labs will soon abandon math progress like they did go and chess.
I definitely think that this is marketing, just "with good side effects". My doubt is when they will be able to move to "marketing with better side effects", that is, research with more concrete outcomes (health, materials etc.).
They aren't trying to 'solve' chess, go, or mathematical proofs as an end in themselves, but mainly in order to learn more about how to build better systems overall. The goal of AlphaZero was ultimately as a stepping stone towards AGI, and it's the same with LLMs.
catlifeonmars 4 hours ago [-]
The assumption being that
1. these things are all stepping stones, not diversions
2. that ai labs have a singular goal of producing agi
simonh 36 minutes ago [-]
It's not an assumption, many of the people behind these projects explicitly say this what they are doing and why.
td6 5 hours ago [-]
Would that be a bad thing?
While top AI labs no longer focus on chess, the community build way better chess engines.
Stockfish is probably stronger, than everything the top labs build.
Wouldn't we expect the same thing for math? That slowly the broader math community would engineer a harness/program... That will surpass the current labs, and be a community ran project
sebzim4500 5 hours ago [-]
Yeah its certainly true now that Stockfish is much stronger than alphazero, but it's probably also true that had Deepmind spent another few years working on alphazero it would be enormously stronger than either.
In the case of chess this seems fine, there isn't much value to society in creating an AI capable of beating top humans with a 4 pawn handicap rather than a 2 pawn one, but for maths where there are actual applications it is more complicated.
kuboble 3 hours ago [-]
But arguably - it's better for chess community that the best engines are opensource than having superior GoogleChessBot.
diamondage 7 hours ago [-]
This misses the raw advantage of a good proof. It makes conceptualization simpler. In some ways math is like a hash list of of theorems. This list makes it simpler to prove other calculations, and will always be useful, to both humans and AI models. I can see two new directions 1 - the creation of specialist theorem models; that can answer questions efficiently about one topic and 2 - we probably need to incentivize and codify ownership of theorems; charging a proportion of the compute saved by using them. Ultimately enabling mathematicians to be paid our true market value!
edot 7 hours ago [-]
Oh boy, please not 2. What if this was a thing already and, since neither Newton nor Liebnitz had kids, we all had to pay some investors who bought the rights to calculus every time we took a derivative.
Dylan16807 7 hours ago [-]
You know patents only last 20 years right?
cnr 7 hours ago [-]
As long as somebody with big $ decides: "let's make it 40!"
noworld 5 hours ago [-]
Aaaaasaand now Disney owns the rights to room temperature superconductivity.
lh712 5 hours ago [-]
Considering how essential math and science is for the prosperity of mankind (not even speaking about the cultural value) the question of how to reward people working and contributing in these fields effectively and appropriately is of extreme importance. (And I think the current decline in our societies is to no small degree caused also by our utter failure to address that issue.)
It is also fascinating, because I don't think there is any solution within our existing system, at least not any I know of. Theorem ownership is not a good solution (and neither are patents in general). Probably the most achievable (or rather the least unachievable) solution is a kind of communist utopia, where people can dedicate their time to a pursuit of any endeavor they see fit, as resources for a decent life are abundant and excessive power capture impossible. (The other option, somewhat dystopian, and which would not require humanity to change too much in its current mode of conduct, would be a totalitarian or caste-like capture of society by the scientific community.)
Incidentally, if AI proves as powerful as some expect it to become, it could bring about another solution of that issue by making all human science and mathematics obsolete, pushing its true market value to zero.
(With apologies for rambling.)
petesergeant 5 hours ago [-]
> It makes conceptualization simpler
I wonder if it makes conceptualization simpler for models too, given that they're trained already on human-speak. And I'm also curious as to whether humans currently have an innate advantage into simplifying and contextualizing proofs, or will the machines get good at that as well?
renyicircle 7 hours ago [-]
I had the same idea recently. You've solved all the famous conjectures (all formulated by humans because humans found them interesting), what next? I doubt "AI formulated a math conjecture that nobody else cares about and immediately solved it" will produce that much hype. The actually interesting thing is indeed how mathematicians themselves will use these AI models going forward and how that will shape mathematics of the future.
squidbeak 5 hours ago [-]
AlphaGo and AlphaZero weren't generalized models. Math capability will presumably keep improving along with the other general capabilities, even if there wasn't a special RL focus for math itself.
bluecalm 7 hours ago [-]
Yeah chess is a good example. DeepMind came for publicity with AlphaZero. Arranged a match with Stockfish with rigged rules to make AlphaZero look better than it really was (it was amazing but the match wasn't fair) and then just published some games and went home.
I was bitter about that back in the day as I hoped for more answers, more matches, more "truth" about chess being shown. Soon after that community project Leela Chess Zero was started and not only surpassed original AlphaZero but added few hundred ELO points over it. Then the combination of NN and classical engines happened with NNUE and current Stockfish is again a few hundred ELO points stronger.
Today we pretty much know the truth in chess for all practical purposes. Human analysts/preparation experts focus on finding interesting path and opponent profiling (what is the most unpleasant for the opponent to face). They don't look for truth anymore. The game is doing great, it's more popular than it ever was.
squidbeak 4 hours ago [-]
Why don't you mention the second match here, with its adjustments to meet Stockfish's quibbles - and the same result?
Stockfish and other classical engines were never intended to be run in matches without opening books (of which there were plenty).
Development assumed the presence of opening book and authors made 0 effort to make engines play well in openings because of it.
This is also the reason classical Stockfish was a very small binary. A little effort to make it even by including even a very small opening book (like 50MB or something that would result in still smaller binary than NN engine with its net) would make it much more interesting.
The result was that Stockfish lost many games by walking into known bad lines and lost way more games than it otherwise would.
squidbeak 3 hours ago [-]
You aren't correct.
> We also played a match that started from the set of opening positions used in the 2016 TCEC world championship, along with a series of additional matches against the most recent development version of Stockfish, and a variant of Stockfish that uses a strong opening book. In all matches, AlphaZero won.
Ok I remember it vaguely but the match that got publicity and the one published results were derived from was 1000 games match from starting position.
Deepmind claimed AlphaZero also won from TCEC positions and vs Stockfish with good opening book but at least back in the day I don't think I could find those games being published or specifics about books/positions they have used. Can you?
I am not claiming AlphaZero wasn't stronger. It wasn't as strong as the PR piece suggested though and we have never seen the games being published. In chess this is extraordinary because basically all games in chess are publicly available - both human and computer games.
Claiming "we have created a strong engine that has beaten Stockfish with opening book" while not showing those games (or details about opening book used) is akin to "we solved this math conjecture" without showing any kind of proof or argument.
Publishing a few 1000 of games costs nothing. Tens/hundreds of thousands of games are published every day.
mohamedkoubaa 2 hours ago [-]
They already proved the point
im3w1l 7 hours ago [-]
I disagree. Firstly, people in AI likely care about math on a personal level. Secondly math is useful. Playing go or chess is basically a party trick. Being useful gives it staying power.
But, I do think you are right that there will be some level of moving on. The spotlight is currently on maths and that won't last. It will move to some other area where there is more impact to be had. So while they might shift gears and put less focus on math, it will always be there as part of the portfolio.
techpression 5 hours ago [-]
I think there’s a venue where they start focusing on introducing hypotheses where the model currently can’t solve it, or maybe this is already happening?
Being able to present useful novel ideas would likely generate a lot of press, for a while. I don’t know how this would look since I’m useless at math, but Im sure there are plenty of unknown problems with massive implications, that once formulated can be solved.
antman 8 hours ago [-]
This argument implicitly makes a few assumptions which will probably not hold in the very near future.
One is that AI will continue hallucinating in a manner that is not easy to verify, second is that AI will not be enhanced to produced more simplified amd robust outputs, and third that a human will be required to do that.
What humans in the loop are doing now is verify the process, propose shortcuts and add legitimacy, through the verification process, if that ends up being succesful its highly likely a lot less mathematicians will be required in the future.
The conclusion that this is not productive focuses on the mathematicians, but it is very productive in terms of hundreds of proofs being produced that had previously consumed uncountable hours of the brightest minds. Unless it ends up being the greatest hallucination ever ofcourse
Vetch 7 hours ago [-]
Putting hallucination aside, LLM "theory of mind" has gotten worse over time. I feel it peaked in Opus 3 and Sonnet 3.5, GPT 4 and then GPT 4.5 for OpenAI. Since then, even with Opus 5.5, phrasing has needed careful crafting, in order that it not be taken too literally. OpenAI models suffer from this much more than Anthropic models but Claudes have backslid over time too.
This means when writing documentation, tutorials or commit messages, their output is often a garbled jumble. Assuming shared context, using invented terminology without explaining, leaking conversational states due to improper epistemic boundaries and failing to model the reader. This all usually leads to their freely generated explanations being terrible. Getting good explanations requires chaining questions that force them to line things up properly, which is not easy the less you know. These failures as something LLMs naturally struggle with make sense, given the nature of attention and RL with weak signals from human data.
Math is not merely a collection of proofs, it's a way of understanding. A proof presented in a manner that cannot be incorporated remains useless. It does not make it's way to physics like Riemannian geometry and matrix math did. This is no less true when done by humans too.
Your hallucination conclusion, checking if a proof is one, is exactly the counterproductive cost.
Most of us cannot verify that the claims in the OpenAI lore dump are in fact all correct. It will take tons of work from experts to do this. It took subject expert mathematicians to identify the discrepancy and disconnect in the Navier Stokes proofs, for example. LLMs will struggle to make use of their own proofs or turn them into knowledge that accumulates over time.
The act of proving is often more valuable than the proof itself. Human constraints and limitations force us to invent tools and abstractions that a 100,000 x 1M context swarm can bypass. The tradeoff from that AI swarm advantage is work that doesn't usually lend itself to being built upon. It's like doing all the side quests and reading all the books of an RPG versus min maxing a straight path with a guide. We might try to identify new abstractions, but the fact that we don't get access to CoT and that much of it will be illegible means mining LLM traces for what human mathematicians produce naturally will be a tedious chore.
p-e-w 6 hours ago [-]
> Since then, even with Opus 5.5, phrasing has needed careful crafting, in order that it not be taken too literally.
This is a feature, and a huge step forward.
If you expect AI to do serious work, you can’t have it guessing what you “really meant”. Every sufficiently advanced task depends on very subtle details in the problem statement, and the correct default behavior for advanced AIs is to solve the task exactly as stated, unless a system prompt or other constraint tells it to do otherwise.
dofm 8 hours ago [-]
> One is that AI will continue hallucinating in a manner that is not easy to verify
It is an old saw at this point, but what an LLM does still cannot be divided into hallucination and non-hallucination. This is literally an anthropomorphism trap.
Layers and layers of application-specific verification can reduce the risks inherent to LLMs, to a really remarkable degree, but nothing about what these tools are suggests that this problem will go away; it will just bubble up again somewhere else.
qarl 52 minutes ago [-]
> can reduce the risks inherent to LLMs, to a really remarkable degree
To an arbitrary degree.
Just like all of science. Reduce the error to the desired margin.
user43928 6 hours ago [-]
And why not?
For all that I saw over the last few hundred hours with AI on software engineering, hallucinations are no longer a problem at all.
Not once have I seen a task fail due to what would have been a "hallucination". If they still occur, they can apparently be detected and corrected automatically, or are subtle enough to escape notice with presumably no significant impact on the results.
Why would this not also be the case for mathematics?
catlifeonmars 4 hours ago [-]
I think OP is saying that hallucination or not is just semantics. There is nothing qualitatively different about hallucinated vs non-hallucinated output.
bonoboTP 3 hours ago [-]
That's true in the same sense as "There is nothing qualitatively different about erroneous vs non-erroneous output" for a dog vs. cat image classifier.
user43928 3 hours ago [-]
To be fair, I guess the line is blurry between what could be labelled a regular mistake compared to a hallucination.
"Test suite passed" when it actually errored? Obvious hallucination, unless it ran a command that returned the wrong error code.
But is running a malformed command that does not achieve the expected effect itself a hallucination?
bonoboTP 1 hours ago [-]
If it makes a false claim, then it's an error. If it says the test was passed or a class was implemented but it was not, then it makes a false factual statement.
I'd say a hallucination (very misleading word) or confabulation or "making shit up" happens when an LLM uses factual / evidential language purely based on local statistical expectations of the text, instead of it drawing from actual evidence in its context pointing to it.
This is murkier in the case of general knowledge questions, like when and where was some famous person born. It may then be a spectrum from fully making something up based on how the name sounds, all the way to confidently retrieving it from its weights correctly. In between, we can get hallucinations. But newer models are taught to use Web Search when unsure, and it works pretty well, though not perfectly. I don't see any fundamental limit here. It's just not perfect. Trying to solve "the hallucination problem" is basically like saying "our dog vs. cat classifier is pretty good already with its 99% accuracy, now all we need to do is the tiny little task of eliminating the 1% error, and we will be golden". Like, no shit, there is some error yes. People are working to reduce it. It will never be absolutely 100%. It's not an insight to say we should remove hallucinations.
antonvs 7 hours ago [-]
> It is an old saw at this point
An old saw unless something that's widely accepted, but sadly it seems that many people don't recognize this, even many people working in the field.
Terr_ 8 hours ago [-]
> assumptions which will probably not hold in the very near future [...] One is that AI will continue hallucinating in a manner that is not easy to verify
Hold up, that's an even bigger assumption in the opposite direction, and I don't see anything to support it.
At least in terms LLMs getting all the "AI" hype these days, there is no structural/mathematical reason to believe they won't continue to have the same problem they've always had of generating plausible text over rational text, and I don't think anybody even has a clear idea how it could eventually be accomplished.
I've seen "then the magic singularity occurs and somehow it solves the problem for itself", but I would classify that more as mysticism than engineering.
user43928 6 hours ago [-]
Hallucinations are no longer much of a practical problem in software engineering.
Two years ago, hallucinating that the code worked or that a task was accomplished was a common occurrence.
We have seen that now agent swarms across thousands of agents can coordinate to achieve a result.
Clearly hallucinations are no longer the problem they once were, since now we can get working results for long horizon tasks that require massive compute.
Consequently it would seem unwise to assume that current limitations will remain as they are and prevent LLMs from coming up with solutions that they can explain to humans.
catlifeonmars 4 hours ago [-]
It’s still a common occurrence.
It happens in more subtle ways, but it still happens often enough for me to notice. For example I have had hallucinated checksums show up in lock files as recently as yesterday using a SOTA model.
This is not surprising, since the whole basis of LLM training is to produce output that humans will accept _as a proxy for actual training goals_. In a sense, the training process of an LLM “wants” to produce output that is statistically plausible much more than it “wants” to produce correct output. It’s always going to be a struggle to drive that system towards other goals (and we see this bourne out in practice by the amount of effort that is required to be spent on RL).
I think there will be some threshold of correctness (something like 99.999% of the time) that if the model surpasses it, I can stop needing to check it, but I think we’re still at 99% or something which sounds good, but when you are producing a ton of output you hit that 1% frequently.
> Consequently it would seem unwise to assume that current limitations will remain as they are and prevent LLMs from coming up with solutions that they can explain to humans.
I 100% agree with this. In fact explaining things to humans is something LLMs are particularly well suited for.
ashkankiani 3 hours ago [-]
The confidence with which you, anonymous user, keep commenting that "hallucination is not much of a practical problem in software engineering anymore" based solely on your own anecdotal evidence is really remarkable, in not a good way.
user43928 3 hours ago [-]
You're free to substantiate your comment by telling us about your apparently different experience.
jazzypants 3 hours ago [-]
I find hallucinations in my (mostly perfect) AI output every single day. If you're not finding them, you're just not looking hard enough. It's not surprising when everyone is screaming about how they don't read code these days.
This is just a fact. I'm sorry if it messes with your narrative.
Thanks for the link. What kind of hallucination are you seeing, and does it affect the end result?
ashkankiani 3 hours ago [-]
The sum total of all human observations is still not proof of the lack of hallucinations as a problem (even if their observations were perfect, which they aren't considering the volume produced vs reviewed carefully). That's why you can use a counter example only to disprove and not prove anything.
And yeah I get hallucinations all the time still. Maybe it's because I'm working on harder/more niche problems (like a compiler with an unusual type system), but it happens quite a lot. I don't record all of them.
Although the most common one you can find is them misattributing the source of changes from themselves and also other agents (Fable, Opus 5.5, deepseek, whatever). They'll say "your changes" or "you changed" or "your ruling." I didn't decide anything and it's in their own chat log, and yet...
hodgehog11 8 hours ago [-]
"Plausible" text was preferred over rational text when we trained LLMs using RLHF. It's rapidly shifting the other way now with RLVR, which enforces correctness by default.
catlifeonmars 4 hours ago [-]
> but it is very productive in terms of hundreds of proofs being produced that had previously consumed uncountable hours of the brightest minds
You’re making the following assumptions:
1. the exercise of struggling to find proofs was not productive, but this is precisely how new techniques in math were produced. Brute forcing solutions doesn’t lend itself to the creation of much new mathematics (except maybe the exercise of developing verifiable proofs)
2. the point of doing mathematics is to be “productive” in the first place. This is silly. Many people get into mathematics because of the beauty of understanding, for example.
bonoboTP 3 hours ago [-]
> 2. the point of doing mathematics is to be “productive” in the first place. This is silly. Many people get into mathematics because of the beauty of understanding, for example.
Are they independently wealthy? Or do they have a deal with their local supermarket that they can take food for free?
catlifeonmars 35 minutes ago [-]
So maybe we should just pay them more to do math and let them set the direction of research?
If your point is that capitalism fucks up the incentive structure and makes it all about maximizing productivity then I wholeheartedly agree with you.
Retr0id 8 hours ago [-]
Even if you somehow have a 100% correct AI, it's not useful unless we can understand and internalise (and communicate) its results.
colordrops 8 hours ago [-]
Who is this "we" you speak of? The professional mathematician community? Were pre-AI results useful outside of this community of people who could understand them?
hodgehog11 8 hours ago [-]
Yes. Most probably do not understand the notation involved in, and the statement of, the Lindeberg-Levy Central Limit Theorem. But every scientist uses this theorem in one way or another. These ideas have a way of trickling down because to people who work thanklessly to do so.
dofm 8 hours ago [-]
Something about this sentence makes me think about that Rob Auton bit, that before there were mobile phones, nobody had any reason to tell someone else that they were on a bus.
zer00eyz 4 hours ago [-]
Before mobile phones you never called someone and asked "where are you" because a phone was tied to a location...
mathisfun123 8 hours ago [-]
> One is that AI will continue hallucinating in a manner that is not easy to verify, second is that AI will not be enhanced to produced more simplified amd robust outputs, and third that a human will be required to do that.
There is literally not a single shred of evidence to indicate either of your supposed eventualities. The core technology of an LLM is sampling from a distribution so there is literally no way to make it deterministically robust (only probabilistically).
emtel 2 hours ago [-]
No evidence other than the fact that this has been happening steadily in all areas for many years?
You might have a point if the goal was to have LLMs that spit out a correct proof without chain of thought or tool use. LLMs + agent harnesses are more than capable of self verification and course correction.
bonoboTP 3 hours ago [-]
Is a human deterministically robust? Or is a human also incapable of doing what you claim LLMs will never be able to do?
antman 7 hours ago [-]
The direction and pace of capability improvement has already been demonstrated by all models. The latest breakthroughs make that pretty evident, but there have been production systems that are based on probability since the beginning of computing.
What has been demonstrated is a process that outputs lean proofs based on those probabilities. This happened after decades markov chain producing garbled texts and very shortly after gpt2 producing stories about unicorns.
red75prime 7 hours ago [-]
> The core technology of an LLM is sampling from a distribution so there is literally no way to make it deterministically robust (only probabilistically).
An LLM mostly deterministically (except parallel processing nondeterminism that can be mitigated) produces a probability distribution that can be sampled deterministically: just take the highest probability token or use beam search.
gottheUIblues 7 hours ago [-]
I think people on here tend to somewhat fixate on the determinism issue. Even with a deterministic LLM - stabilising the floating point arithmetic, and choosing from the distribution by a fixed method, or just save the random seeds - there is still a kind of a chaotic unpredictability that can exist between its inputs and outputs. However maybe that is a price that needs to be paid to get creativity.
Vetch 7 hours ago [-]
Deterministic yes, robust deterministic no. The most likely conjunction is not always the best nor representative of what the model is considering unless its certainty is high.
red75prime 5 hours ago [-]
I think determinism has nothing to do with it. If you mean sensitivity to word ordering and such, it's a generalization failure.
thereitgoes456 8 hours ago [-]
AI cannot explain chess moves it comes up with in an elegant way. What makes you think it will be able to do so for math?
ApolloFortyNine 2 hours ago [-]
Interesting claim. That's true for old models that simply have no way to explain, LLMs however can. [1]
It's kind of astonishing that after all we have seen in the last years people still find the position that AI will not be able to do an obviously valuable thing likely and it requiring an explanation (instead of the other way around).
hodgehog11 8 hours ago [-]
No, this is different, and this is coming from someone who has been studying deep learning for the last decade. We are talking about the difference between RLHF and RLVR strategies. The former benefits clarity and explanation, while the latter concerns only correctness. AI was moving in a particularly damaging direction by pushing on the first path, so it was natural to move to the second. But the second will come at the cost of clarity of explanation. It will likely get better at its explanations, but not fast enough to render its most advanced accomplishments readily understandable to the user. The chess example is a pretty good one (that is an RLVR approach).
red75prime 8 hours ago [-]
The problem is that people strongly believe that this is an insurmountable problem that will persist indefinitely (or for a long time) and plan accordingly, while this, most likely, will be fixed soon by adding RLCAF (RL on conversational agent feedback) or something like that.
__s 3 hours ago [-]
tbf GM explaining their 2700 elo moves are only understandable when vague, as elo goes up explanation becomes closer to "in this specific position there's these dpecific lines", why should 3500 elo moves have simple reasoning?
Maybe if we start with giving simple AI generated analysis of those clumsy humans with their measly 2700 elo moves
n6242 7 hours ago [-]
Some of us still remember 2016, when we had a couple of cars sorta half-driving themselves, and Tesla, Uber and others promised we were only one year or two away from three million people in the US working as drivers being out of a job. And here we are, a decade later. AI is pretty amazing, but companies have a tendency to severely and comically overestimate and oversell it's capabilities, and underestimate the challenges.
jstummbillig 6 hours ago [-]
> And here we are, a decade later.
With Waymo and Tesla increasingly doing what they said they would do, and a small number of early adopters happily paying money for their services, that do work.
So what's the critique? That the timelines are not correct? Sure. And how about the timeline of the people who said "research level math, never in my lifetime" and the people inside the ai companies who are apparently increasingly spooked by how quick the progress is? How about the various levels of code/programming jobs that AI was supposedly never going to be able to do, but, in reality, now just does?
We are engaging in some very one-sided discrediting, and I am not sure, why.
jazzypants 3 hours ago [-]
And, those companies are notably avoiding wet climates because they still struggle with self-driving in inclement weather. It's probably going to be another decade before we get to the point where these things can handle every situation. Just like all other engineering, the first 90% is the easy part.
More telling: Waymo just rolled out in Denver (1 month ago or so), apparently fairly confident they got this handled given the upcoming winter.
Progress on the obvious stuff keeps happening (which kind of brings me my to the first comment here).
But also: It does not have to do "snow" to be useful! A lot of cars/people don't drive when it snows heavily, and that's something we have always been okay with (at a societal level, YMMV of course). If was only useful 95% of the year that's still great. A lot of technology works like that.
jazzypants 56 minutes ago [-]
Great response, and thank you for the article! I'm skeptical that they're actually ready for real-world conditions outside of a test track, but that's a whole lot of training data and their lawyers must be convinced. We'll see.
inglor_cz 49 minutes ago [-]
I remember a biologist wryly commenting that it took a lot longer to evolve good senses and appendages in nature than to evolve human intelligence, compared to the ape niveau.
Maybe the really hard thing isn't abstract cognitive capability, but perception+movement.
p-e-w 6 hours ago [-]
Driving a car is unimaginably more difficult than proving the Riemann hypothesis.
You just don’t notice that because evolution has given you 99% of what is needed to drive a car before you were even born.
nullsanity 8 hours ago [-]
[dead]
wolvesechoes 7 hours ago [-]
Expression of tech-faith is not intellectually honest argument.
Where does this "will probably not hold in the very near future" come from? People correctly warn about extrapolating current things onto the future, but then just throw some vague "probabilities" without providing any argument why their "probably" is somehow more grounded than others.
antman 7 hours ago [-]
Markov chains garbled text to gpt took decades, gpt stories about unicorns to gpt production systems took a few years, gpt production system to gpt astra producing deterministic lean proofs of longstanding mathematical problems happened even faster. Scepticism to the point of requiring proof appears like an academic pursuit while production systems have already been built and are in the process of being enhanced
pegasus 7 hours ago [-]
Did you even RTFA? His argument absolutely doesn't make any assumptions about hallucinations, implicit or not. It's you who assumes Tao must have surely been complaining about hallucinations or some such. You've not addressed any of his arguments and moreover ask questions his post answers.
My steak is too juicy, my lobster is too buttery, my industry shaking mathematical proofs are coming too quickly
In what world is OpenAI not “genuinely contributing to mathematics”?
I’m getting whiplash from the speed at which people are suddenly accusing them, and AI in general, of not doing enough.
ApolloFortyNine 2 hours ago [-]
I'm in the same boat, I guess I had this naive idea that unsolved math problems would mean something if they were solved. But it seems like a lot of them at least were more thought experiments than anything else.
nahumfarchi 4 hours ago [-]
Perhaps mathematics was never about proving things? I know it sounds like moving the goal post, and it certainly was what motivated mathematicians on a day-to-day basis, but bear with me for a moment. I think that beyond being a creative activity that humans enjoy, math was about building new tools and systems of thought. Axiomizing things we take for granted, logic, linear algebra and calculus (on which modern LLMs rely so heavily) are such examples. People chose to participate in this field because they found it enjoyable and satisfying in some way. As a side effect, society reaped the benefits every couple of centuries. Will harvesting open problems with AI ever give us these things, or will we just be left laundry pile of Lean formalizations?
jeremyjh 4 hours ago [-]
In TFA I read (scroll up from the link anchor) this was explained: OpenAI are not giving talks - because they can't answer any questions about the model's work. There is little follow-up activity - the actual elaboration of human understanding of the new ideas is stifled since the problem is solved.
But you are right, this is not OpenAI's "fault". The problem is - as others have said recently - that many people in mathematics want recognition for solving open questions more than they want the answers to the open questions. Everything about the economics and social environment of Mathematics will have to change.
I think that this is exactly the same split we see in software: there are those who mainly enjoy the craft aspect of building software, and are uninterested in the product or business they are supporting. Others are primarily interested in the production of useful software or building a platform or company.
I've always been in both camps myself. When it became obvious that AI was going to destroy the craft aspect - at least two years before it actually could do so - I became very discouraged, even depressed. But once it was actually good at building software, I became very excited about all the stuff I could now build. Sadly, I think a lot of people in our field have never had something they really wanted to build.
vouaobrasil 1 hours ago [-]
Mathematics is less about the end product and having a healthy community of people to actually understand the proofs and eventually apply them.
That community won't exist as many people simply won't even enter the field because it's reprehensible and contemptible, not to mention boring, just to read machine-generated proofs and verify them.
Collaborating on, or at least working on unsolved problems is what motivates most people.
AI is like a cheat code in a video game. You get to the end faster but fewer people want to play if the cheat code is always on. You can't turn it off either because the very challenge is to do something unique.
gessha 4 hours ago [-]
What does it mean to “genuinely” contribute to mathematics? Because the definition you use is load-bearing ;) and it might be different from that of others.
saberience 4 hours ago [-]
Because they published a ton of slop papers with terrible English, impossible for humans (even experts) to understand, full of non-standard terms, invented jargon etc.
So effectively, stuff got proved, but people don't really understand how, so it's mostly fucking useless and done for OpenAI's marketing team, while also pissing off the maths world at large.
alexwebb2 3 hours ago [-]
So if god himself comes down from the heavens and hands you the answers to major unsolved problems, but doesn’t walk you through them and you’ll have to still put in effort to understand them, then that’s “not contributing”?
lioeters 2 hours ago [-]
Give a man a fish, teach him to fish, etc.
ziiinq 47 minutes ago [-]
[dead]
ziiinq 50 minutes ago [-]
[dead]
qarl 55 minutes ago [-]
> but simply dumping proofs on the math community
So you'd prefer if they kept their work secret? Or you don't want them working on these problems at all? Or they should be required to do the work the way you want them to?
I'm not clear what you see as a better option than dumping.
carra 6 hours ago [-]
> simply dumping proofs on the math community and expecting others to do the grunt work of verifying, refining, and expanding on them is hardly a productive way to advance the field.
I tend to agree with this, but what is the alternative? Should OpenAI and Anthropic employ hundreds of mathematicians to do this work? Should they just not solve math problems within their reach?
user43928 6 hours ago [-]
I tend to disagree with OP, for the same reason.
It's unclear what more could be expected than releasing the presumably already verified results and write-ups for each problem. Should they run a mathematics school too?
Then the comment goes on to argue AI labs were not interested in actually advancing mathematics, and that investments into AI were manically excessive.
IMO none of this follows and demand is there to justify the investments.
The comment then goes further to argue that AI labs were putting too much effort into pretending there was exponential progress rather than actually making progress.
The factual basis for this claim seems to be that OpenAI released math results and write-ups, and it's not even clear what more they could do on that topic.
That's a very negative opinion.
pred_ 5 hours ago [-]
Look at how actual researchers are currently using the tools: they'll generally use the LLMs to create slop papers, sometimes supported by auto-formalizations, just like OpenAI does. Then they will go through the lengthy process of digesting the results, turning the often incomprehensible and poorly organised outputs into something that humans can understand and build upon, they will then give seminars on the results, further helping with dissemination. Doing so still requires expertise, and probably will for a good while.
So yes, that is exactly what they should do. Alternatively, if they are too lazy or incompetent to put in the effort themselves, do what AGMAI proposed and fund a third party to help out.
charlieyu1 5 hours ago [-]
The biggest problem is LLM tends to produce over engineered, very complicated proofs that are an eyesore even for relatively simple problems. Give it a beautiful Olympiad geometry problem and LLM will tear it apart into ugly algebraic calculations, turns all lines and circles into equations and calculate their intersection points that spans multiple pages because it is a guaranteed way to solve it. Correct, but hardly any use to the user.
Gigachad 5 hours ago [-]
I’m not deep in to math but the op tweets make sense to me. In that it’s not just the final proof that mattered, but the mind and understanding of the person who arrived at the answer. An LLM dumping the answer can’t elaborate on it, can’t tell the story of how they got there, etc. But it also deprives someone else of that achievement and learning.
Ravus 5 hours ago [-]
You can see it in the way we structure college courses: engineering curricula often cover in one semester what mathematicians study over one or two years.
This is because have fundamentally different goals: being able to use results in calculation versus having a deeper understanding of the subject matter.
KoolKat23 5 hours ago [-]
You know it's valid. You're not working on incorrect assumptions. Surely there's value in that?
charlieyu1 5 hours ago [-]
I don’t know if it is valid. It is unverifiable. I still found some basic algebraic mistakes in top models as late as 3-4 months ago, not sure about it now. But that’s not what I want anyway, so I often put “Do not brute force” in my prompts.
KoolKat23 4 hours ago [-]
Sorry I mean in very public releases such as this trove, where many have lean certificates attached and publicly scrutiny.
shakna 4 hours ago [-]
Three of them have already been withdrawn. So I would say we have proof, that they cannot be implicitly trusted.
KoolKat23 4 hours ago [-]
So it turns out there is a, perhaps informal, working system in place and we can deal with it.
Peer reviewed and published insights are proven invalid all the time. This is the nature of research and how we learn.
shakna 4 hours ago [-]
So instead of us "knowing it is valid", we don't. We need to put in extra effort, because someone felt like doing only half the work and dumping it on the community to fix.
We don't know the current system can work well enough at this scale, because that's un-knowable. We know it can find some of the problems. We don't know it can find all of them.
We do know it takes more effort - that's knowable. Increased data takes increased processing.
Whether the community has the required effort available, seems unlikely, considering the expertise required to be able to assess these things hasn't changed. Only the ability to generate them has increased.
KoolKat23 3 hours ago [-]
You don't have to go through it. You can ignore it if you wish. There is no obligation on you to check it.
I feel your concern stems from the risk that there is additional noise everyone needs to cut through.
In reality this isn't any tom, dick or Harry giving you their vibe code output.
They have spent millions of dollars on this output, so there is a filter. The biggest filter of them all, funding.
Furthermore, LLM's have given us another gift semantic search, we can easily check your work against theirs, this is valuable insight so instead of researchers wasting decades and fortunes pursuing an avenue that shows no value (this includes methods), they can purse new avenues they know what to avoid, in the same breath they know what to work towards.
datsci_est_2015 4 hours ago [-]
How something is proven is often more important than what is being proven. There are underlying systems and patterns that, when understood properly, improve our model of mathematical (or physical) reality.
With convoluted and inelegant proofs, AI may fail to uncover those systems and patterns. As a most concrete example, it may fail to recognize some problems as isomorphic to other problems. Brute force solutions are a depth-first search.
To improve human mathematical understanding, AI is probably best used as a “copilot” (lol) rather than a black box oracle, like these AI companies appear to be doing.
KoolKat23 4 hours ago [-]
I'm sorry this framing is just moving goal posts.
If you're after new methods. Then new methods is the goal, the answer to the question is not the goal then. The animated response indicates the answer wasn't just a byproduct.
There is still something to glean from the answer. You have a further constraint. Otherwise whatever "new method" proposed may as well be hallucination, potentially taking you in the wrong direction away from the answer.
This line of thought is not unique, stonemasons made obsolete by uniform brickword suddenly were "worried about the art and preserving traditions".
killerstorm 7 hours ago [-]
> the grunt work of verifying, refining, and expanding on them
What work do you think mathematicians do normally?
Like they sit whole day and have ideas? And where are the ideas?
The way I see it, _some_ mathematicians enjoy solving puzzles, and now AI is better at solving puzzles.
This does not affect people building new theories.
Also, it's quite prestigious to write a _book_ on some topic. And guess what writing a book entails? Refining and expanding. What you call grunt work.
vouaobrasil 1 hours ago [-]
> This does not affect people building new theories.
Actually, the problem is, somehow the skill of building new theories in math is directly tied to slaving hard over a problem. It's the very experience of slaving away that actually somehow causes ideas to form. Pretty much all mathematicians understand this. Yes, senior mathematicians now can form some new theories, but what about junior ones who will have very little experience in working hard on a problem by hand?
Of course, they could work on the problem by hand anyway, but they won't because no one will pay them when a machine can do it.
killerstorm 25 minutes ago [-]
> no one will pay them when a machine can do it.
Let's start with the fact that mathematicians aren't paid for results. Institution which gives them salary fundamentally doesn't give a fuck about theorems. They might care about having top-grade mathematicians for prestige, or because they believe that countries which are "good at math" also do better in science, engineering, etc.
> somehow the skill of building new theories in math is directly tied to slaving hard over a problem
We don't know if that's the only way. Perhaps collaboration with AI is just as good. Why reject it before it has been tried?
> what about junior ones
Junior guy with brilliant new ideas might benefit the most from AI as it can compensate for lacking technical chops and breadth of knowledge.
pks016 6 hours ago [-]
My grunt work is different than doing grunt work for a company to fix their problems so that they can make more money.
killerstorm 33 minutes ago [-]
OpenAI does not make money from math papers LOL they just release them for the sake of community. (Because sitting on those results would be considered worse.)
gexla 9 hours ago [-]
Right, easy comparison to make the the open source community for software.
And it's not like this is something where we're loaned some top math genius for a limited amount of time and we have to make the most of it. Rather, this is a new high water mark. The accessibility of the results is no longer scarce. The scarcity has shifted, and that's where the focus of the math ecosystem should shift as well. And it doesn't help for a frontier community to saturate and take over messaging pipelines that were typically managed by the math ecosystem. It's not about "stay in your lane" but rather "we need coherence and be careful not to break the system."
Just two cents from someone who could screw up basic cashier math on any given day.
DesaiAshu 7 hours ago [-]
As with any field, convincing people to care about your ideas and your approach is half the battle
Many of the best startup ideas by the best product and engineering minds failed to gain attention and funding. Same with much of the best music - relegated to hard drives with derivative ideas only resurfaced decades later
I would expect much of the recent math dump will be leveraged by other LLM-driven research teams rather than read in depth by a human
rtkwe 2 hours ago [-]
The other gap in AI is it doesn't explain or lay out how it got to the final proof which is often more fruitful for new techniques etc than the final proof by itself.
defmacr0 38 minutes ago [-]
Honestly, that's a weakness shared by a lot of human-generated science as well.
KoolKat23 5 hours ago [-]
I'm sorry I disagree entirely.
The more information the better.
The entire purpose of published work is to remove noise (and perhaps incentivize work through attributing credit).
This information is now out there. You can choose to ignore it if you wish. You may just find yourself a century behind in research.
And on that point most of this research has been looked at by their mathematics panel and comes with lean certificates, it's not exactly noise.
This to me is more the old guard not willing to let go or change their ways.
td6 5 hours ago [-]
I think in a ideal world your right.
I think a problem is that math seems like a deeply toxic, ego driven domain.
I think he argued that e.g because the navier stokes millennium problem ist considered solved now, you won't get any recognition for being the first human to solve.(How would you even proof you solved it yourself and not just regurgitated the ai proof?)
And since recognition is the main objective, noone would spend time on dissecting the proof, and perhaps finding some unique approach to solving the problem, that could be transferred to other open issues.
And therefore the problem is now "poisoned". Since it's assumed to be solved noone will research it, and the potential revelations won't be found
KoolKat23 5 hours ago [-]
To be fair I'd hope during peer review this would be apparent.
During your write up, I'd imagine you would check it's not already out there too. And once complete it's cheap and easy to run it through an LLM and ask is this covered by anything else out there. If its novel and not published it doesn't matter what others say.
Research is already messy as it stands. Something new can already be dismissed by incumbents as "not novel enough" especially in niche fields where they're likely to be the ones conducting peer review.
contravariant 5 hours ago [-]
I think it's quite clear the mathematics panel didn't read most of the papers with the scrutiny it would take to publish it, if only because the lean certificates and the informal proofs are not 100% the same.
And if it's hard to understand (which seems to be the most common reaction) it's not exactly devoid of noise either
There was an opportunity for people to work with the AI to produce a proof, now it almost feels they're working against it.
KoolKat23 5 hours ago [-]
I never said it was journal worthy. Just its not all noise.
As I say academics are welcome to ignore it, it isn't published in any journals after all.
willtemperley 7 hours ago [-]
> Too much effort is being invested in proving that the exponential curve is still holding.
Given sustained exponential growth is mathematically impossible to maintain with finite resources, it's funny to me they're using advanced mathematics to try and achieve this.
aswegs8 7 hours ago [-]
Once LLMs pass the threshold to being to invent new general-purpose methods and frameworks, the frontier is irreversibly lost to AI and it becomes simply a hobby that mathematicians pursue. They work through, digest, and maybe write up the proofs for understanding. But the real meat will be growing the LLMs. Who knew software eats the world was so true?
schleck8 8 hours ago [-]
> but simply dumping proofs on the math community and expecting others to do the grunt work of verifying, refining, and expanding on them is hardly a productive way to advance the field
This is a transitive period. In a few years, verification and exchange between model instances will happen faster than humans can follow. Human input will be an ethical question, and not a productivity one, because it will be the bottleneck in any science.
RandomLensman 5 hours ago [-]
How to make sure it doesn't evolve into some sort of Library of Babel of science?
schleck8 3 hours ago [-]
That's the billion dollar question. Nvidia currently wants to deploy chips with the sole purpose of monitoring agentic workloads and that doesn't seem farfetched but they obviously have a financial motive to sell more products. It's like a cat and mouse game, like cybersecurity in general.
Muromec 8 hours ago [-]
And what use of that?
psychoslave 7 hours ago [-]
Etic is a factor of productivity, and the larger the contextual window is the weightier it becomes. That’s even integrated within the paperclip parabola.
If the focus in placed on maximizing some easily measurable output on a narrow perspective, situation is unlikely going to match a sweet spot of holistic equilibrium which is maximizing harmony and happiness through humanity as a whole.
nullsanity 8 hours ago [-]
[dead]
ew-dev 6 hours ago [-]
> simply dumping proofs on the math community and expecting others to
> do the grunt work of verifying, refining, and expanding
Hm, kinda reminds me of my college days. "Proof trivial, left as home work." was a sentence my Profs loved to say.
nateburke 5 hours ago [-]
Less Liszt/Paganini, more chamber music and teaching. I like it!
FrustratedMonky 4 hours ago [-]
"community building"
For what?
If math is just about having a community of other mathematicians to hang out with, it still isn't a career. Nobody is paying money you need in order to to eat, just to hang out in a community.
Just like a software engineer, "Well AI can write all my projects now, but I have my local Rust Users Group to hang out with". Nobody is paying me to hang out and hand code Rust.
myzek 7 hours ago [-]
Generate and dump on others to verify is how the generative-AI people operate. Be it in maths or just your regular job.
The amount of Confluence pages of "research" that is just a dump of LLM output someone passed to me to review is staggering
I hate this approach, it's unbelievably selfish
mastermage 8 hours ago [-]
Its kinda like the arms race in the cold war. There came out some truly marvelous technologies but the actualy goals were frankly terrifying.
XorNot 9 hours ago [-]
That seems short sighted though. A few years ago models couldn't do this at all, I'm not sure there's any evidence to suggest exploring and refining results is outside their capabilities or will remain so.
OAI obviously have a fiscal incentive here, but to presume a year from now we won't see improvements and more succinct work on the results coming from models?
rtpg 9 hours ago [-]
The problem is that one a person writes a 60 page proof in theory that person has spent an inordinate amount of time on the proof and can answer questions, describe some insight, etc etc.
If a random person is given a 60 page proof to digest and not the author, those hidden insights that _aren't_ in the paper might be completely inaccessible. Maybe the AI will "just" be able to provide the insights. Maybe. But pedagogy is tricky work, and despite these AIs being able to do all this fancy math we can't get them to write good cover letters yet, so....
Ultimately we might be left with just a bunch of intellectually unsatisfying proofs. This means way less drive to simplify the proofs or rework them.
End result: we generate a layer of "less efficient" mathematics, that won't get built upon. We will not actually have any shoulders upon which to stand.
charcircuit 9 hours ago [-]
AI can simplify and rework proofs too.
i_cannot_hack 8 hours ago [-]
According to Scott Aaronsson, OpenAI set their agents on 8000 different problems, and got 372 final proofs. Even spending twice the original effort on simplifying and reworking those proofs so that they do not "feel like something written by someone who’s on psychedelics" would only increase the compute by less than 10% (assuming all the agents had a similar token budget).
The fact that they did not do so can only mean that either (1) their agents currently lack the capability to do it, or (2) OpenAI are completely indifferent and do not care in the slightest if the proofs are understood or not.
gwd 8 hours ago [-]
> The fact that they did not do so can only mean that either (1) their agents currently lack the capability to do it, or (2) OpenAI are completely indifferent and do not care in the slightest if the proofs are understood or not.
Come now, this is kind of unreasonable. When you're working on a new technology, you first get the ugly, inconvenient-to-use prototypes functioning with the core new thing you need; then you work on packaging it up into a format useable in production. I'm sure the very first digital camera sensors weren't very useful for photographers either; but it isn't really even possible to build the rest of the technology required to turn raw output of a digital sensor into something a professional photographer can use until you have the raw output itself.
The research is still on going on the raw output; getting things to the next stage, where the results are widely useable by professional mathematicians (and then on to engineers and scientists to whom the results would be practically useful), is a whole new research area.
i_cannot_hack 7 hours ago [-]
Seems like you are just subscribing to the first option I gave, "their agents currently lack the capability to do it", but saying you think they will be more capable in the future if they can move away from the inconvenient-to-use prototypes after more research. Thinking it might be possible in the future is not in disagreement with anything I said, so I am not sure what you thought was unreasonable about my description.
gwd 5 hours ago [-]
Imagine someone looking at digital camera researchers showcasing a ground-breaking new sensor, and reacting by saying "Well obviously they're completely indifferent and do not care in the slightest if their work is used by real photographers or not."
Like, "Orr... maybe they care a lot, but haven't gotten to that part yet?"
i_cannot_hack 3 hours ago [-]
You are conflating the two options with each other (lack of capability vs indifference). One of them is true, not necessarily both ("or", not "and"). I did not claim a lack of capability was the same as indifference.
lern_too_spel 7 hours ago [-]
Or (3) they feel a need to publish first, and that goal takes precedence over (2).
i_cannot_hack 6 hours ago [-]
They claim each result used three hours of compute on average. Even spending significantly more on simplifying and reworking would delay the release with a single day at most. If avoiding such a minor delay took precedence over (2), I think "indifference" is the correct term. It has also been a while since the release now, so there is ample opportunity to post follow ups if time pressure was the only concern.
lern_too_spel 6 hours ago [-]
Why do that when they can spend more hours extending QRH to a proof of the full Riemann Hypothesis? The opportunity cost of digging up small potatoes is the whole enchilada.
XorNot 9 hours ago [-]
But why should process of discovering mathematical insights be any less attainable to AI models?
The concern is being raised without evidence, because the evidence points to the gap simply being frontier models have just started to be able to get a raw proof out. Why, given existing progress, should we expect them to be unable to distill insights from those proofs?
Certainly this even more likely doesn't matter at all for applications: if I can send a radio signal further because my AIs design it a certain way, that's an unambiguous result. Which is really the next step here: turn a proof into a "mechanical" application.
JumpCrisscross 9 hours ago [-]
> I'm not sure there's any evidence to suggest exploring and refining results is outside their capabilities
OP didn’t suggest that.
The bar has been raised. Everyone has to meet it now. An inelegant solution squatted onto the internet doesn’t count as discovery per se, even if it’s impressive.
auggierose 8 hours ago [-]
A correct solution verified in Lean will count in perpetuum. It is fine if you want more, but an achievement is an achievement, even if it is by AI.
oliculipolicula 5 hours ago [-]
Hmmmm. It feels right that discovery is much more meaningful than achievement. "Bullshit lean proof" or "bullshit achievement" smells like it. "bullshit discovery" smells like a front-handed insult
Maxion 9 hours ago [-]
But what should they do? They got all these proofs, should they just have sat on them?
JumpCrisscross 9 hours ago [-]
> should they just have sat on them?
It’s fine that OpenAI posted their findings. It’s not fair to claim these problems have been solved. Not until someone can understand and verify the proof and then communicate the core, novel methodological element to someone else.
gwd 7 hours ago [-]
But this is Tao's point: Before, the mechanism by which a proof was verified and communicated and digested by the community was for the person who came up with the grotty, ugly first draft to engage with the community. Now there's nobody to really engage with, so the pipeline from "grotty, ugly draft" to "integrated into humanity's mathematical knowledge" has been broken.
So yeah, probably we should stop saying "X has been solved", and instead say, "A Lean proof for X (or !X) has been generated". That doesn't change the fact that incentives are currently on finding the proof, and once the proof is generated by an AI, there's not currently a good mechanism / incentive structure to move that into the mathematical community. AI is here, so we need to find a new mechanism.
lern_too_spel 7 hours ago [-]
It is not obvious to me that a single canonical human language write-up of a proof is the best output in this new world where write-ups are cheap. A human reader can query an LLM and get explanations of key points tailored to the reader's own background in mathematics.
CrimsonRain 8 hours ago [-]
Solving a problem is not solving anymore. Up is down, pleasure is pain, darkness is light, slavery is freedom, madness is sanity.
Turneyboy 7 hours ago [-]
Many of these are lean formalized. Arguably a much higher bar than whatever peer review provides in terms of verification.
cmceanga 5 hours ago [-]
There is no guarantee that the lean proof is 1:1 with the natural language equivalent. The lean proof can be lesser. This happened in the Navier-Stokes proof, e.g. see [1] in example 3.1. Having the certificate doesn't necessarily imply correctness.
On some consideration, surely. On the other hand, but at some point this is borderline like saying "universe already solved every physical problems, including possibility to represent deep important point of its own structure in compressed intelligent ways" and then tell that reaching it in an actual grabbable artifact is left as an exercise.
Possibly yes such a representation is possible. But it doesn’t mean it’s certain there is a "best compressed representation". And even less one that encompass everything important and that is understandable by any human brain, even the most exceptionally brilliant ones sponsored by a whole society to reach their best possible achievable performance on that goal through full dedication on that sole task.
usernomdeguerre 9 hours ago [-]
>They got all these proofs...
Your phrasing is illuminating that perhaps they aren't engaged in the creation, understanding, or integration of these proofs by humanity; they just have them. For them, this is a slidedeck they can pass to investors, creditors, the marketing department. Something they can add to the employee onboarding pamphlet.
What should they do? Hyperbolic maybe, but perhaps engage with humanity.
nearbuy 8 hours ago [-]
...they did.
p_hoep 9 hours ago [-]
No need. In the end the mathematicians that don't like this can just not look at the proofs or use them. They have that choice. Just like they didn't "ask" for them, they don't have to even acknowledge they exist.
sans_souse 9 hours ago [-]
BlackBox: The proof is in the pudding
This isn't only bad for Math — it's bad for English too.
'Proof' is going to become the 2026 Most Misapplied Word of the Year.
pks016 6 hours ago [-]
At least check them properly. They have already withdrawn some of them.
8 hours ago [-]
yieldcrv 7 hours ago [-]
> There seems to be more interest in hitting some arbitrary benchmark (we proved X unsolved problems) than in genuinely contributing to mathematics
I feel the same way about academia, the papers, the citations, the ego, the narcissism and the taxpayer codependency that got cut off and turns out wasn’t necessary at all thanks to a private sector entity running laps around them
I don’t feel that academics need to pursue the discipline and distributed brain-wracking that has sometimes resulted in the solved math problems, just because more times they find other nooks and crannies to explore along the way. I think the blueprint is enough. Standing on the shoulders of giants is good enough.
and if the concern is that they can’t figure out what to do with a proof, next year’s AI will
mrheosuper 9 hours ago [-]
Why not asking another chatbot to verify, like what we are doing with coding ?
kadoban 9 hours ago [-]
That's not the hard part. The hard part still requires humans (and/or _maybe_ a bunch of tokens) and a lot of work.
iterance 9 hours ago [-]
And then what?
caaqil 8 hours ago [-]
> There seems to be more interest in hitting some arbitrary benchmark (we proved X unsolved problems)
> genuinely contributing to mathematics
What's the difference between the two? Proofs are no longer the goalpost?
tene80i 8 hours ago [-]
Proofs are valuable but I believe the mathematical community values understanding more. Proofs were previously a great way to develop understanding. Now, less so.
shubhamjain 8 hours ago [-]
Not an expert, but explaining the proof and it being independently verifiable as important as just putting a paper of it. Reminds me of 1000 page proof of Goldbach’s conjecture that a Mathematician reached sometime back. He was told plainly that no one is going to invest time in verifying the proof because there’s a good chance there’s an error somewhere in between.
desterothx 7 hours ago [-]
Proofs of open problems are valuable because we are assuming that proving the problem requires some new method or infrastructure in math to prove it. Basically proving open problems isn't actually useful if it doesn't develop new tooling for mathematics, which can help us create new open problems, solve other ones etc.
pegasus 7 hours ago [-]
I recommend RTFA, it explains exactly that: why proofs should not be considered the goalpost, and why dumping all these AI-generated proofs might be an overall negative for mathematics as a whole.
flexie 9 hours ago [-]
Yes, it is indeed a balanced view.
Those few thousand mathematicians are now getting a taste of their own medicine. After all, it was people with extraordinary mathematical talent who developed machine learning and large language models, leaving hundreds of millions of people who earn their living through speaking, writing, or teaching worried about their future job prospects.
Still, I believe almost everyone will be fine. Perhaps AI will also prove good at coming up with new conjectures, and some mathematicians may shift towards applied mathematics or other sciences.
FrustratedMonky 3 hours ago [-]
Pretty tenuous connection to blame mathematicians for everything bad created in the world, that happened to use math.
encyclopediai 7 hours ago [-]
This Math 1.0 followed by Math 2.0 framing is wrong. Mathematics did not start a few decades ago when problem solving became the norm. And it will not end now when problem solving turns out to be "easy".
Mathematics will revert to it's main practice, which is to study.
There are lots of weird panic reactions by some prominent problem solvers. See for example the ridiculous cease and desist like statement of AHM shared at Tao blog.
Put this Math 1-2.0 with that AHM statement together and you'll realize that this is a power struggle and that you see only one side of it.
Mathematicians have very diverse opinions about this. I, for example, am for as much as possible automatic harvesting of all these "low hanging" fruits. Should be disclosed as soon as possible, free of any bottleneck, and citable. The mathematics community may do whatever its various members desire to do with these results. Let them decide individually what to do with them. This AI tool is here to stay.
encyclopediai 7 hours ago [-]
This looks indeed like a public relations campaign, it would be helpful to the conversation to contribute your opinion, in a polite way.
kccqzy 3 minutes ago [-]
It’s obvious if you think from a programmer perspective:
MVP 1.0 is always produced quickly to solve a particular business problem or conform to a spec, even if the solution has terrible code.
2.0 is when programmers refactor the code and internal APIs to make things nice and easily explainable.
bsenftner 5 hours ago [-]
The future of "math" is not just solving, but communicating what was solved, and why that solving is important. "Math" as a discipline is only ever been half created, they abandoned explaining themselves like it was beneath them. Well, now in "Math 1.0 finally whole" they will be explaining what they did to the rest of us. If you can't explain you are not really there.
curt15 4 hours ago [-]
That's what math has always been about, even if not everyone is necessarily as talented a communicator as someone like Thurston -- who was himself an exponent of computational tools in math research:
>I think of mathematics as having a large component of psychology, because of its strong dependence on human minds. Dehumanized mathematics would be more like computer code, which is very different. Mathematical ideas, even simple ideas, are often hard to transplant from mind to mind....Translation in the direction conceptual -> concrete and symbolic is much easier than translation in the reverse direction, and symbolic forms often replaces the conceptual forms of understanding....
This is an entitled take. The usefulness of a thing has nothing to do with how easily communicable it is. That is the whole premise of specialization and becoming an expert. People go in-depth on things so they can say “trust me on this” and so you don’t have to. You are advocating for a tyranny of the illiterate.
Also, this reads like you didn’t like math classes. That sucks, but it’s no basis for societal organization
bsenftner 3 hours ago [-]
I am advocating for effective communications, it is not as if knowledge once understood remains difficult. To communicate and convey understanding is good, and to withhold understanding is tyranny.
ziiinq 36 minutes ago [-]
[dead]
j-pb 8 hours ago [-]
The authors of the proof are invited to give many talks, and meet with other experts in the area. Workshops are set up to discuss the proof, as well as other recent developments.
problems are being solved autonomously by AI prompters who have no interest in the broader field itself once their initial target is "solved", and do not understand the AI output well enough to answer questions on the result
The value here seems to be the insights that the author of the proof gained, and the paths they took and maybe more importantly didn't take.
Inviting only the human prompter to a talk on the paper is like inviting only the department chair, manager of the actual author.
The valuable part that Tao is feeling the absence of is the insight, and you can only get that from talking to the swarm of agents that developed the original proof with all of their context.
So to me it feels like we don't need Math 2.0, but Authorship 2.0. I want to "meet" the context that generated these proofs. I mean luckily these were not generated by faceless systems like a SAT solver, you can actually talk to it, but I'm not sure if we can step beyond our pride and grant the true authors of these proofs that recognition.
pegasus 7 hours ago [-]
These systems don't have a genuine capacity of introspection, beyond just mining the conversation trace. When you're asking them why they did this or that, they are basically guessing anew from the outside, and are just as likely to hallucinate as they are to hit the right answer. These are mechanical systems which brute-force chains of various (re)combinations of techniques acquired from the training data. The true authors are all those who have contributed those techniques in the past.
dannyw 3 hours ago [-]
I think that's debatable. Anthropic's interpretability research suggest models do have self-introspection ability, at least in the "J-Space": https://transformer-circuits.pub/2026/workspace/ ; and this private working/'introspection' space is distinct and distinguishable from the tokens they output (CoT tokens are output too).
a3w 7 hours ago [-]
This was how it was done in chemistry back in the day: release a cooking instruction. If it fails, visit the colleagues and give hints on what you meant.
The idea would be that you should not fiddle with the minds who try to independently evaluate your works, so that is not a feasible approach to truth seeking.
While in organic chemistry, this way, valid progress was made, you can always avoid a perpetuum mobile inventor and get conned.
scotty79 4 hours ago [-]
> problems are being solved autonomously by AI prompters who have no interest in the broader field itself once their initial target is "solved"
This is pretty much what a person that proivded patronage to a matematician used to be. API prompters are people who provide patronage for AI mathematicians.
You don't talk with them about the discoveries. About discoveries you should talk with who actually made them. Namely the LLMs.
Another analogy might by that you shouldn't expect to have interesting discussion about the essence of art with art producer.
Just thank them for the inference they covered and interact with the results instead.
cubefox 3 hours ago [-]
The problem is that the reasoning traces are kept secret. OpenAI doesn't publish them because it would lead to distillation attacks from other AI companies.
drichel 3 hours ago [-]
The reasoning traces were traditionally keept secret by legacy mathematicians as well.
"When the architect completes a fine building, he removes the scaffolding." - Carl Friedrich Gauss
cubefox 2 hours ago [-]
Legacy mathematicians can nontheless remember much of how they came up with their proof. Even though they (as Gauss says) usually don't publish this, they can still answer questions about it at conferences and workshops, or otherwise use the knowledge of the creation process to explain their proof. It's not a secret in the sense of OpenAI.
lifeisloving 10 hours ago [-]
The same could be said of Software. Instead of giving up on creating novel projects and instead just taking other peoples ideas and porting them to Rust, we could be embracing AI to push software and computers farther.
Im not sure how that will work, but im convinced the current paradigm of just pushing agents into codebases for not much reason other than you can is going to make building software incredibly boring and push creative people away from the field and stagnate progress.
My prediction is software gets boring and building hardware projects will be the new frontier for creative engineers looking to push computing further. Which is probably a good thing.
primer42 8 hours ago [-]
Originally the word computer described a person doing the act of computation.
I believe in the future, we're going to see a similar shift in "programmer" - instead of a human programming the computer, you'll give the ai an idea and it will spit out a program.
And just like how automating the act of computation revolutionized what we could compute, automating the act of writing code will change the act of programming - hopefully, as you described, allowing us to do things that simply were not practical in the past.
est 9 hours ago [-]
I was about to comment the same.
People wrote many books about software engineering, all from valuable experience from buildng expensive software systems. But in the age of AI, is there still anything learnable from generated code?
Personally I always ask AI to summarize its findings and lessons in a .md file. And I always learn something from it.
But could AI utilize some new patterns and paradigms I wasn't aware of? Very likely. Because we only learn from our personal grave mistakes, a summary from others gets neglected and forgotten
JumpCrisscross 9 hours ago [-]
> same could be said of Software
Sort of. An elegant proof is useful beyond what it shows. It hints at new mathematics, and can prompt discovery in applied fields. I don’t think I’ve heard of elegant code leading to discovery on its own.
bartnp 8 hours ago [-]
I think this happens all the time actually, but it's often smaller scale. Like someone writes a neat architecture to do X at their company, and later someone else walks in, looks at this code and sees it's now super easy to do Y, which then turns out to have immense user/business value.
Sha1rholder 6 hours ago [-]
> I don’t think I’ve heard of elegant code leading to discovery on its own.
Every software design pattern came from elegant code. People wrote code, summarize code, learnt from code, and taught code. That's discovery
rcpt 9 hours ago [-]
> elegant code
Usually it's the opposite. "That's in prod? And it works? It shouldn't work and I thought it was doing something else. Why does it work?"
Jaxan 9 hours ago [-]
But it has. Think of design patterns and other programming paradigms (logic programming, functional programming, etc).
whateveracct 9 hours ago [-]
you must not be familiar with haskell
devmor 9 hours ago [-]
> I don’t think I’ve heard of elegant code leading to discovery on its own.
I’ve not heard of it either, but code is an abstraction of math, so I don’t see why this couldn’t theoretically happen.
Anecdotally, I’ve started spending time advancing my math skills beyond the early college level I stopped at and I’ve frequently found I already know concepts of more advanced math - I just didn’t know what they were called or how to apply them to an equation on paper, but I’ve been using them for years and intrinsically grasped the underlying academics.
flir 7 hours ago [-]
But these days hardware is software. One microcontroller replaces so much electronics. Software is eating the world, as the man said.
Gigioingiro 9 hours ago [-]
I do agree with you, although I think that there may be still some advancement available in software in terms of programming languages or novel architecture design. But frontier is much more around hardware, just thinking about computing power, electrical transmission, material constraints on power, connectivity, insulation. I would definitely advise kids to study physics and materials engineering than CS.
hypfer 9 hours ago [-]
I mean if we're being honest, most of CS studying (as practiced) was a waste of time anyway.
The S fell short in actual reality for the most part, as it was merely a hiring requirement. A hiring requirement that didn't even make sense, because the skillset of academic CS only marginally overlaps with the skillset one wants to hire for.
Material engineering at least for the most part has actual real-world applications where one can push humanity further. CS (as practiced, not necessarily the idea of real CS but the CS we got due to it being used as a hiring filter) for the most part is just self-referential spinning with mostly unclear results.
There is real impressive work being done in that field, of course, but I'd argue that the majority of it over the last decade or so at least was just performative nonsense.
Maybe by again allocating new resources to other fields, what hides under the label CS can become more pure actual CS again. I think that would also be a much less miserable experience for everyone involved.
Maxion 9 hours ago [-]
Models have limits too, I don't think software will become boring. I think it'll become more interesting, in the not so distant future one software engineer will be able to do so much more than today.
dyauspitr 9 hours ago [-]
What new things are even left to discover in software? You can still use C for all the backend stuff and I don’t think all the endless stream of front end frameworks and libraries are all that innovative. I think the only tangible innovations we’ve made over the last few decades have been in infrastructure management. All heavy lifting is done by mathematics anyways.
desterothx 7 hours ago [-]
Well i mean, the whole idea of discovery is that you're not sure what you will discover until you do
customguy 8 hours ago [-]
> What new things are even left to discover in software?
The will to implement the stuff we learned, instead of letting dark patterns and churn for the sake of churn get the better of that :P
slashdave 12 minutes ago [-]
This is backwards. Not sure why what is essentially a PR effort by these companies automatically becomes an advancement in the field.
Do we want these capabilities to be useful? It seems making them accessible so that the existing community can use it as a tool is the right approach. Automated proof methods are already used like that.
stared 3 hours ago [-]
Pure mathematics, in my view, is not science, not craft - it is art.
AI, however powerful, is a tool, only as important as the amount it helps mathematicians. Creating mathematics without human understanding is as sound as mass producing copies of Michelangelo David.
“Mathematics, rightly viewed, possesses not only truth, but supreme beauty — a beauty cold and austere, like that of sculpture [...] yet sublimely pure, and capable of a stern perfection such as only the greatest art can show.”
- Bertrand Russell, "From The Study of Mathematics" (1902)
“A mathematician, like a painter or a poet, is a maker of patterns. [...] The mathematician's patterns, like the painter's or the poet's, must be beautiful; the ideas, like the colours or the words, must fit together in a harmonious way. Beauty is the first test: there is no permanent place in the world for ugly mathematics.”
- G. H. Hardy, "A Mathematician's Apology" (1940)
AlexAplin 10 hours ago [-]
>the mere knowledge that a solution exists "contaminates" efforts by both humans and AI to find alternate routes to the problem that reveal additional insights
This really expresses the heartburn you see across all fields, not exclusive to careerism. I certainly have friends in decomp and fan translation spaces that have been demotivated by the current rash of efforts happening there.
The rush to be "first" has always been over-celebrated, but it would be nice to believe there's a way to get beyond that thinking.
bonoboTP 9 hours ago [-]
This has already been the case in AI/ML and computer vision papers via flag planting papers. Have an idea, super quickly publish a hasty work based on it that doest actually work, methodological and eval issues, engineering terrible, slow, bad results etc. But it was the first so now your concurrent work that was much better evaluated, better implemented, etc is suddenly worthless and unpublishable.
ChickeNES 9 hours ago [-]
> decomp
This I don't understand, seems like an obvious thing to automate, especially for byte-matching?
pinkwah 5 hours ago [-]
Byte-matching has been a side-effect of gaining an understanding of how the software works. With this you are skipping the understanding and you will not get people like Kaze Emanuar on YouTube who have dedicated a lot of time to building upon this understanding to create something better.
I don't think it's motivating to solve a black box by having AI generate another black box if what you want is to understand how the thing worked.
diath 50 minutes ago [-]
> Byte-matching has been a side-effect of gaining an understanding of how the software works.
No? Decomp sources are often full of comments that claim that there's no clear reason as to why something is done a certain way or straight up full of question marks. Byte matching often is a result of bruteforcing a solution rather than understanding the original idea behind the code.
dminik 5 hours ago [-]
If your goal was to do something for the games you love and understand them on a deeper level, then just handing that off to an LLM doesn't really give you the same sense of accomplishment.
devolving-dev 8 hours ago [-]
Why were we doing math in the first place? We should be happy that math problems were being solved, since presumably they were blockers for other problems in science and the like. But it feels like math was really more about seeking enlightenment, like a form of mental yoga or something. If so, we can just ignore AI proofs and continue on maybe?
rsfern 4 hours ago [-]
There’s a somewhat famous lecture [0] by Wigner (one of the greats of 20th century physics if you’re not familiar) on exactly this topic. One of his points is that new tools and ways of thinking developed on the way to solving mathematical problems with no apparent application often find downstream applications in science and engineering. If we’re skipping the part where we identify and understand the new math, will we still reap the unreasonable effectiveness? Tao’s position in this post suggests that the current wave of LLM successes is not conducive to this dynamic
I'm a retired AI researcher and I still enjoy programming. However, I don't use any however assistance as I still want the thrill of learning new things. Unfortunately, it appears that the only way for most developers to enjoy similar activities is to be retired with enough money
freecodeio 29 minutes ago [-]
The amount developers saying how "now we can focus on the bigger picture" rather than the code are just delusional. The bigger picture has always been a bigger problem than writing the code. It's the part that engineering stands for in "software engineering". If you couldn't code something well or weren't involved in bigger picture decisions in the past, you're just gonna engineer big picture spaghetti.
Forgeties79 1 hours ago [-]
This is exactly why half the conversation conversations about AI eventually become a critique of capitalism. AI isn’t for fun or joy, it’s explicitly for business use. For replacing people. For extracting more out of each worker. That’s all the major companies (OAI, Anthropic, etc) are offering. None of this is designed to give us any more freedom or time to pursue the things we are passionate about.
I feel like we’re circling back to that 2010s energy of “everyone can be an entrepreneur.” Now it’s “everyone can build software”
joe-excom 8 hours ago [-]
You could replace "math" with "programming" in your comment for a different perspective.
ekjhgkejhgk 7 hours ago [-]
If you're considering an individual, maybe this makes sense.
But if you're consdiering a community, this falls apart. The maths community has universities, has professors who are paid, has students which are getting their degrees for varying reasons, it has conferences, has publications, papers, projects etc etc, all of which will get some negative impact some AI.
You know that line "when a measurement becomes a target it ceases to be a useful measurement". This line holds up to different degrees for various measurements and targets. For maths it holds up very well. The goal is "contribute to make the world better by increasing humanity's understanding of maths" and the measurement, which by evaluating an individual on it we're turning into a target, is "how much does the individual publish new findings". Measurement turned target holds up great. It's almost impossible to publish new findings and not contribute to humanity's understanding of maths. But with AI these two are being decoupled. You can produce lots of new findings, but the community is saturated and they don't get assimilated into humanity's understanding. Why do individuals use AI then? Because you've made the target "how much does the individual publish new findings" and they have to compete or lose.
kamaal 6 hours ago [-]
If a coal miner or car mechanic would say something like this, they would be called a luddite by the same science people.
Its different when your own job is on the line.
When human manual arts were being automated away it was supposed to be not only acceptable but any complain and you were told you were a progress blocking luddite.
Now that mental labor is getting automated, the response to automation is very different.
ekjhgkejhgk 3 hours ago [-]
This isn't about physical labour vs mental labour, it's about the dynamics of delivering a product vs contributing to a community.
For example, a software engineer is like a car mechanic or a coal miner. None of those are anything like a mathematician.
The person you're rallying against isn't me, it's an imaginary hypocritical person which exists in your mind. I do mental labour, I welcome AI developments hard, and I still think TT is 100% correct.
schleck8 8 hours ago [-]
To many professional mathematicians it's a form of art. But at the end of the day, it's a profession done for money, and even if they like doing it as a day to day job, they'd probably be doing something else if they had absolute freedom over their time. This is how I personally classify things as art or chore. If people continue doing something the same way when there is no financial motive, it's art in its pure form.
xanderlewis 4 hours ago [-]
> even if they like doing it as a day to day job, they'd probably be doing something else if they had absolute freedom over their time.
Not the mathematicians I know. They’d happily drop the academic admin stuff, but they’d absolutely keep doing mathematics in much the same way.
jvvw 4 hours ago [-]
Based on the research mathematicians I have known, I'd disagree with 'they'd probably be doing something else if they had absolute freedom over their time' for at least a large number of them, although it probably varies from individual to individual.
david-gpu 5 hours ago [-]
Some call that "entertainment", or "a hobby".
TrackerFF 8 hours ago [-]
If we're going to be frank about it, higher level math is:
1) Intellectually challenging, to such a degree that those wishing to enter the field need to have a certain level of intellectual prowess to do so. This creates some levels of mystique, with a sprinkle of elitism and gatekeeping.
2) Driven (among other things) by prestige. And the more pure the math is, the more prestigious it is.
3) So complex that people can spend their entire working careers chasing a handful of problems. The amount of time researchers spend on very specific problems is mind-boggling, if we think about the results.
4) Intensely captivating for the people deep in the weeds.
And the deeper you get, the longer you study, the more you start to value things like "mathematical beauty", and may start to view math as a form of art.
Like many similar fields, you end up with this ivory tower where people can dedicate their whole lives to thinking deeply about extremely niche and theoretical problems.
kamaal 5 hours ago [-]
Similar things are happening in the competitive programming world.
Quite a lot of people are not happy they aren't elite anymore, and many have spent years to decades to arrive here.
Simply put you invest years of your life to establish a kind of distinction over others, and that goes away. That hurts.
But its not something surprising. Most of these competitive programming problems were actually English languages puzzles, because you couldn't dial up the mathematical difficulty anymore making it a fields medal problem. And in most cases in simple language weren't even that hard to begin with, and you could look up solutions to these problems in an hour of Google searching.
defgeneric 1 hours ago [-]
I'm not totally against "math as art" but I don't think that formulation goes deep enough towards what's really going on with mathematics. There is something deeper where I lean more towards the philosophers who have said that mathematics is essentially ontology.
Since the Greeks we've had the idea that "Being and thinking are one," or that Being (in the sense of all of existence as such) has some essential unity with thought, and therefore can be thought, and expressed or submitted to the logos or reason. Being is in some sense fundamentally intelligible, and mathematics is the most developed, exacting, and articulate expression of Being.
Logic was understood in this older sense up to roughly the the mid to late 19th century. This is why a work like Hegel's Science of Logic begins not with syllogisms or propositions but with Being and Nothing. But this was forgotten after logic was mathematized by the English around the time of Russell, and its connection to ontology was gradually overshadowed by a focus on epistemology (still, it should be remembered, originally as a means of getting back to ontology).
There may be truth in art, but it's always haunted by its own historicity or contingency, which is to say untruth. Mathematics seems on the contrary the only really timeless, absolute thing we have. Part of what makes it captivating is stumbling on a construction or concept or proposition or theorem that simply must be, independent of us.
The AIs are certainly now more than automatic theorem provers, mechanically traversing some space of true propositions. They are able to push things forward and connect seemingly disparate domains to get to a proof, but to my mind it remains to be seen how well they will be able to form new concepts and definitions.
Imagine the controversy surrounding Cantor, for example, but put an AI in the place of Cantor. If an AI proposed something like the (infinite) hierarchy of infinity, would we have accepted it? What would the intuitionism debates have looked like? Would they even have taken place? And aside from that, has it actually been shown conclusively that an AI could propose such a thing?
There are lots of attempts right now to recover a humanism for mathematics, or restore man's pride of place with respect to it, but maybe we don't need to worry about that. Tao's attempts to preserve the mathematical community, while allowing for practices to change through the crisis may look like a kind of rearguard action, but seems reasonable to me and not really dependent on any kind of humanism. It's a way to avoid the question for now while things play out (and not conservative/reactionary like Scholze and others), which may be exactly what we need, because after all, perhaps we still don't understand why we do mathematics, what it's really for, and what our relation is to it. Whether it's enough to preserve funding is another issue.
NotGMan 7 hours ago [-]
Where was Terence Tao where other workers were being replaced by immigrants, robots and other automation?
Academics and white collars now get to experience what blue collar workers experienced in the past.
Same as what developers in USA experienced who were and are getting replaced by Indians.
sebzim4500 5 hours ago [-]
He grew up in Australia, didn't he? He was the one doing the replacing I suppose.
melagonster 7 hours ago [-]
He is a imigrant.
MikeTheGreat 36 minutes ago [-]
Genuine question: how did this error get through? It seems really simple/basic and I thought that expressing a proof in a formal language was supposed to prevent this?
Disclaimer: I'm kinda familiar with the ideas of automatically provable software systems but haven't done anything with them myself. If I'm missing an obvious fact(s) here please let me know :)
I get that the proof is big and complicated, but "We messed up a +/- sign" kinda sounds like announcing that the next version of the Linux kernel is done ("Version 7.0.0 is awesome!") followed by realizing that it doesn't compile ("Turns out someone used 1 equals where they should have used 2. Stay tuned for V 7.0.01!").
I've got to be missing something here :)
armcat 9 hours ago [-]
I think this focus on a "holistic" approach applies to everything AI is touching now, not just math. On Twitter I see people one-shotting games, or reproducing games. If the goal is to just one-shot a game using AI, it's done. But if the goal is to produce immersive medium that people can truly enjoy, admire the story and the craftsmanship, and can find entire new ways of bringing a story to life, that's something else entirely.
orlp 8 hours ago [-]
Sadly the experience with physical products tells us that the vast majority of people prefer cheap disposable crap over expensive craftsmanship.
uludag 1 hours ago [-]
The dynamics of software are vastly different from manufactured goods though. Software can be distributed at essentially zero cost so the disposable crap analogy breaks down.
Like given the choice, the vast majority of people would prefer one quality game like Minecraft, LoL, or Fortnite, vs. thousands of one-shot generated games, and looking at user playtime this is exactly what we see. If anything AI is just going to entrench these pre-AI franchises even more.
elAhmo 6 hours ago [-]
I don't think they prefer it, but the price is a deciding factor.
Most people would be happier with La Marzocco coffe machines which costs thousands of dollars, but if you get an OK shot with a 100 USD DeLongi, then the choice is clear for majority of the population.
chrismustcode 4 hours ago [-]
La marzocco micro and mini are still high effort to make your coffee compared to some automated 90% cheaper delongi.
I think people sway to low effort endeavours that still have a reward at the end (even if it lesser reward than high effort).
sebzim4500 5 hours ago [-]
I don't think this is really true for games, a lot of games have enormous resources thrown at their development and marketing and they fall flat because they simply aren't very fun. Then we have a million slop games that sell 3 copies on steam, and in practice only the most unique/fun/addictive games can break through.
5 hours ago [-]
fwlr 8 hours ago [-]
I hope Terry has somewhere private where he can safely express his anger at how poorly math has been treated by these AI corps. I understand that as a recently pro-AI public figure he is limited to ambivalence, so I understand him adding caveats like “maybe math 2.0 has a place for AI”, but it can’t feel good to say stuff like that just days after OpenAI so drastically salted the earth.
SuperV1234 8 hours ago [-]
> how poorly math has been treated by these AI corps
Yes, what a terrible thing to advance the field significantly and release the results publicly for everyone. Truly despicable.
fwlr 7 hours ago [-]
Well, that’s the thing, isn’t it? They somehow invented a way of solving an open problem without advancing the field. That’s something that mathematicians had never even thought was possible. And while the mathematicians were beginning the great conversation to re-examine their fundamental understanding of the pursuit of math in light of this new phenomenon, OpenAI decided they would do it again 372 more times.
drstewart 7 hours ago [-]
>They somehow invented a way of solving an open problem without advancing the field.
And nothing stops mathematicians from solving it in a way that does advance the field. Claude's existence doesn't change that.
piker 6 hours ago [-]
Such a naive take. The funding for solving these problems went to zero over night.
drstewart 6 hours ago [-]
Hence the funding was never for "advancing the understanding", but the problems themselves.
IanCal 5 hours ago [-]
That's not quite right, assuming everything said so far is just flat true.
The funding can have been for advancing the understanding, by using a more measurable proxy and reasonable target that closely aligned with advancing understanding.
Perhaps another phrasing might be "we are paying people to go through the process of solving these problems" rather than "we are paying people for solutions". I might set a random task for my kids while on a hike to find X things, not because I want to find ten different leaves but the process of doing it means exploring and investigating in a certain kind of way a certain kind of area. If some sets up a leaf selling stand, they have advanced the field of "finding leaves" and kids can now very very easily get ten different leaves.
Now leaves here are frivolous and not useful. That example works better looking at, say, homework - clearly I don't care about having a list of words spelled correctly and I don't need the answer to 5x7, nor do we need more book reports on The Great Gatsby. We're doing them to teach, it's very explicitly about the result.
Research level maths however is a bit of both. The answers to some of these things are genuinely useful. Having the answer may be better than not having it. But having a lot of people working on solving it has other useful and beneficial outcomes.
We have structured large scale systems of huge numbers of people and institutes around how this works, and what top mathematicians are telling us is that open problems (particularly at new researcher level) are a key part of this process and are hard to find. Academia changes incredibly fucking slowly, just glacially slowly. Some aspects (most?) are barely changed across hundreds of years. And across an incredibly short space of time (less time than one paper can take to go from fully finished to actually published) we have gone from "this machine can solve school level work" to "this machine is solving major research level maths problems". The existing system will not work, the machines will not get dumber or slower, and some of the impacts are things you cannot undo.
xanderlewis 4 hours ago [-]
The real value is in advancing the understanding but very few outside academia understand this (it’s counterintuitive, which is fair enough as it’s somewhat unique to mathematics as a subject), so if you want to get funding you frame it as being for solving the problems.
Now that you can solve without understanding this setup is broken. Either the sources of funding will finally have to learn the difference and the value of the latter, or mathematics research ends.
piker 6 hours ago [-]
No, that's demonstrably untrue given how little direct value is given to most of these results.
fwlr 6 hours ago [-]
Wouldn’t be the first time someone kept the bean counters happy while also generating some positive externalities! Of course, now with AI we might finally be able to realise maximum efficiency, where only the bean counters are happy and there’s no other positives.
ssfdg 7 hours ago [-]
What a disingenuous, willfully ignorant take. None of these AI companies have been acting in good faith or in any way other than grotesque self-aggrandizement to the expense of everyone else.
gyosko 7 hours ago [-]
Totally agree. They are doing in it ONLY because they think it's the best way to be more valuable. They don't give a F about anything else.
fsflover 6 hours ago [-]
How does this change the fact that they shared the proofs with the community and allowed everyone to verify or reject them?
sebzim4500 5 hours ago [-]
So what? If someone releases something good then that's good, doesn't matter what their motivations are.
aaron695 8 hours ago [-]
[dead]
Razengan 7 hours ago [-]
"Oh no — A smart person doesn't hold my view — He must be faking it!"
5 hours ago [-]
enum 5 hours ago [-]
The real question is how do you come up with a credible 3-5 year research program that is unlikely to be scooped or become pointless overnight.
3-5 years is the period of a grant, and grants have to make research progress, or you don’t get the next grant.
amunozo 5 hours ago [-]
Well, in this case, I think the problem and what should be updated is how grants work. They were stupid before, but now they are even more.
oliculipolicula 4 hours ago [-]
Stupid, but could be fun: grants are given for how amusing your proposed prompts are.
Everyone should have a portfolio of prompts that are indecipherable by other humans but when fed to a frontier model, produces shocking one paragraph english version of a 10000 line lean proof
Obfuscated prompt grant contest
philipwhiuk 1 hours ago [-]
SIGBOVIK with extra steps.
amunozo 4 hours ago [-]
Either fun or totally random. Unfair but at least cheap.
nphardon 7 minutes ago [-]
Does Tao profit off Ai labs?
sunkeeh 5 hours ago [-]
He is right, I think this would be a good direction for sectors and careers that are at risk of becoming redundant.
I love smart people like this; even when there's a threat, instead of just being in denial or boycotting out of anger, they figure out a new path for their community
spuz 9 hours ago [-]
I said this when OpenAI announced they had solved a Millennium prize problem: solving open problems for the sake of it will lose its cachet. AI companies will no longer benefit by making these announcements. They've proven the effectiveness of their tool. If people want to use them to advance human knowledge then let them do that. There's no benefit to humanity to turn electricity into proofs just for the sake of it.
dooglius 4 hours ago [-]
It's not just about proving the benefit of the tool, it's about proving that the latest version of the tool is better than the previous version and of the competitors' latest versions
sebzim4500 5 hours ago [-]
In this case though the compute per problem was pretty reasonable so if this model was available this rate of progress would mostly continue whether OpenAI funds it or not.
underdeserver 10 hours ago [-]
Even when a problem got solved, there has always been value in publishing simpler proofs and corollaries that give better intuition into the broader field.
If I understand Tao correctly, he's saying that's going to have to be the focus going forward. I just default to thinking the models are going to be much better than us at that, too.
matusp 9 hours ago [-]
> the models are going to be much better than us at that, too.
I wonder if this is true. The code produced by these models are not really getting any more elegant over time. On contrary, the models seem to be getting worse, often proposing really baroque architectures. You can use RL to optimize for correctness, optimizing the vibe seems much more difficult.
squidbeak 4 hours ago [-]
For deeper complexity, elegance may turn out to be a hindrance. Better or worse may turn out to be descriptive of taste and nothing else. On the other hand, elegant alternatives might be brought into reach by this first messy contact with new ideas. Who can say at this stage?
Yamata 9 hours ago [-]
There’s value but it is not rewarded commensurately with the contribution. Same for doing peer review.
lxrogers 1 hours ago [-]
Right now the AI labs are the ones prompting and solving open problems because they are testing unreleased models. As soon as mathematicians get access to these models a lot of this discourse about the behavior of Labs will go away. Mathematicians will be the ones releasing incomprehensible lean proofs and GitHub dumps and it will be on them to follow whatever etiquette the research community decides is best.
stared 8 hours ago [-]
It is surprisingly similar to "Catching crumbs from the table" by Ted Chiang, a sci-fi perspective published in Nature in 2000, https://www.nature.com/articles/35014679
KoolKat23 4 hours ago [-]
Something interesting for folks to keep in mind.
Einstein did not typically use the formal peer review system to "settle on published work." Almost all of his major papers (including his landmark 1905 Annus Mirabilis papers) were published directly by journal editors without formal peer review.
When Physical Review sent his 1936 draft to a referee, Einstein was so outraged that he withdrew the paper and vowed never to publish with the journal again. He corrected his math only after an informal, friendly discussion with colleague Howard Percy Robertson—who, unbeknownst to Einstein, was the anonymous reviewer.
throwaway_2494 3 hours ago [-]
Studies have figured out how many 'concepts' or 'chunks' of information a typical human brain can hold in working memory. The famous 7±2 is a common way this is expressed. This research has been disputed and further nuanced (I know, I know, replication crisis...), but for the sake of argument, assume that there is some cognitive limit to how much information we can hold in our heads at once.
In the beginning, calculus as invented by Newton used complex ruler and compass constructions. Newton had a high cognitive capacity, so it was understandable to Newton. It took mathematicians coming later, including Leibniz, to turn this technique into a body of work that fits more easily into the average human mind. Newton and Leibniz independently developed calculus, but Leibniz's notation and formalism in particular provided a much more compact way of expressing and manipulating the ideas of calculus.
Open up Spivak at any page and find a formula; you should probably find that it contains 7±2 'things'. Like an integral, say: the integral sign, lower and upper limits, the function inside, the variable of integration. Then the theory and the rules for transforming these expressions was created so that working with calculus becomes mostly a set of rote operations.
Now a 1st-year student can do more calculus in a week than Newton even could have done in a year.
Now imagine aliens land on the Earth which have 10x our cognitive capacity, and we ask them about their mathematics. It would probably be incomprehensible to us because it would not have gone through a cognitive bottleneck sufficiently small to force it to fit into our minds. They might be totally happy with a mathematical expression containing 700 'things.'
We now find ourselves in this situation, except the alien is an AI we created.
I believe a cognitive bottleneck needs to be maintained so that maths can still remain human maths.
EDIT: Basically, mathematical elegance is finding a representation which allows irrelevant detail to dissapear.
cs_throwaway 10 hours ago [-]
Let us know when it is clear that UCLA does not hire the candidate with the most top-tier journal papers.
It’s more likely that instead of spending a 100K/year direct grant on two PhD students, PIs will hire 1 and have the student spend 50K on AI.
prodigycorp 10 hours ago [-]
More likely is that frontier labs give huge discounts of EDUs if they consent to allowing training.
leni536 5 hours ago [-]
Plus requiring attribution in papers.
sebzim4500 5 hours ago [-]
Maybe, but I suspect that AI in maths will soon become so prevalent that it's use will be assumed by readers whether there is attribution or not.
rao-v 8 hours ago [-]
My understanding is that top tier math programs don't hire based on journal pubs and haven't for a bit. Math journals are pretty slow and hiring is really via reputation building based on one or two big wow ideas / proofs propogated via arxiv and talks.
cs_throwaway 6 hours ago [-]
It’s just as unlikely that this will cease. hiring a professor at a top school because they were an excellent math teacher… It won’t happen.
sanxiyn 9 hours ago [-]
This is basically On proof and progress in mathematics by Thurston restated. When Thurston wrote it in 1994, many people didn't understand what he is talking about.
And already on page 2 Thurston cuts the the root of the issue, 30 years ago:
>On a more everyday level, it is common for people first starting to grapple with
computers to make large-scale computations of things they might have done on a
smaller scale by hand. They might print out a table of the first 10,000 primes,
only to find that their printout isn’t something they really wanted after all. They
discover by this kind of experience that what they really want is usually not some
collection of “answers”—what they want is understanding.
The difference in reactions by the maths community should probably be spit up into those who have read and understood On Proof and Progress and actually thought deeply about why they do mathematics, and those who haven't.
lhd1 5 hours ago [-]
Are you suggesting that Tao, whose banner image on his blog is a quote from On Proof and Progress, has not read Thurston's essay?
Tazerenix 4 hours ago [-]
Tao is not the one who needs to worry, but it bears pointing out that Tao is probably one of the most famous archetypal "problem solving" mathematicians in history (though at that echelon all mathematicians are both problem solvers and theory builders, in the same way that Chess Super GMs are all universalists).
Tao only changed the heading of his blog to that quote in the last month, which is part of my point. Somehow despite this framework for viewing the field of mathematics having been beautifully described by one of our greatest leaders over 30 years ago, Terence Tao who has spent the last 5 years thinking about this has only just come around to it. Perhaps it has something psychologically to do with his greatest skills being those most under attack (though of course I'm not actually implying anything about Tao, he is obviously an honest and good-natured contributor to the community).
The people who need to read On Proof and Progress are the undergraduates and PhD students who have ended up down an academic pathway without looking up because they've always been good at proof and understanding mathematics, without ever pausing to ask why.
cmrx64 6 hours ago [-]
this is what came to mind for me as well.
SpeakMouthWords 1 hours ago [-]
In the future, some mathematics will be proven but not understood.
Frontier Mathematicians will work on proving theorems, and Pure Mathematicians will work on taking proven results and making them understood.
tossandthrow 1 hours ago [-]
What does it mean to understand?
My own prediction is the opposite. I believe that a lot more humans will have a much deeper understanding of math, as you have llms to teach us.
However, this will only unlock if we get a hold of our own attention.
koopuluri 9 hours ago [-]
> But at the current time, the opposite is often occurring: problems are being solved autonomously by AI prompters who have no interest in the broader field itself once their initial target is "solved", and do not understand the AI output well enough to answer questions on the result, give talks, or otherwise interact with the rest of the field.
Sounds like Terence Tao would have said the same about Ramanujan who basically just "solved" problems without much explanation / reasoning / communication other than it just arrived from god.
In the case of Ramanujan, others took on the responsibility of socializing and community building knowing that he wouldn't do it himself. Why can't the same approach happen here?
There will be people who want to just "solve" math problems now that they have a new tool that lets them express themselves this way. Maybe the don't want to participate in the broader math community, etc. Why discourage them, or add friction / a barrier to them participating in their own way? Why not take on the burden of socializing, making sense of, and community building yourself?
There may be valid reasons here I'm missing, but to me this seems a bit like wanting others to approach a field in a particular way even though the field can support many ways.
dminik 5 hours ago [-]
I think many saw working with Ramanujan as helping or being involved with a one in a generation (or even higher) talent.
Translating LLM proofs to human-ese is instead grunt-work. One that is bound to dissapear in a few years anyways.
QuesnayJr 9 hours ago [-]
He would never say that because Ramanujan was the one who had the insight that he couldn't explain, not a machine. If a machine solves the problem, then the machine deserves the credit, not the person who prompted it.
hodgehog11 8 hours ago [-]
[dead]
MrOrelliOReilly 10 hours ago [-]
Does this response properly anticipate how math will change further with the next N model generations? Exposition and exploration may fall well within the capabilities of future models.
defphysics 2 hours ago [-]
What should OpenAI have done instead of what they did? I think it’s far better that we get the messy dump of proofs ASAP so everyone can grapple with the reality of what these models can and can’t do, and participate in figuring out where things go from here.
What would a more “responsible” approach have been?
fwlr 2 hours ago [-]
For each paper they think they have, reach out to mathematicians working in the relevant area to find ones willing to take on the role of advisor, co-author, reviewers, etc. Basically just follow the same process that a postdoc would to publish a paper.
defphysics 35 minutes ago [-]
I imagine that would take months at least. Meanwhile, only a selected few get to have access to what might be some of the most significant developments in mathematics. That seems incredibly unfair.
atniomn 1 hours ago [-]
You're just moving the problem around. Instead of OpenAI and Anthropic dumping proofs, you'd instead have someone (maybe an aspiring graduate student, maybe a hobbyist, maybe an established mathematician who has disdain for existing process) dump these proofs en masse personally.
dsjoerg 4 hours ago [-]
My university math teacher considered his job done when he showed me the proof. He literally said "I gave you the proof, what more do you want?". That was the end of my math education.
So, now he can join the club. The AI gave him the proof, what more does he want?
yedhukrishnan 10 hours ago [-]
This is a distilled version of what people say about the tech industry in the past year or so. Replace math with any field, and the statement is still relevant.
socializer 10 hours ago [-]
I don't think it's that simple. I divide AI-impacted fields into three buckets:
1. Some present a unified line that the whole point of their craft is the human experience, and that automation is the antithesis of that. Marathon runners don't care that a car can get there faster, poets don't care that Poem Bot 2000 can write poems too. I think this is smart if you can credibly take this position. The difficulty is mostly convincing the buy side, which requires being very outspoken about your views.
2. Some appear to be undecided, with one faction taking the pro-human stance and another rushing to accelerate things with AI. A good example of this is mathematics, and I really wonder where they end up in the long haul. They have a very good claim on #1, because mathematics is pretty close to an art form and is robustly insulated from the pressures of the marketplace. But they can also choose option #3, below.
3. Some crafts prioritize results above all else, practitioners either rushing to extract as much money as possible before it all collapses, or believing that they can out-prompt everyone else forever and that their prompting skills are indispensable to their employers in the long haul. That's software engineering. I think this is going to be interesting to watch.
magimas 3 hours ago [-]
I don't think it's smart at all to position yourself in bucket 1.
There's a reason there are maybe a few hundred professional marathon runners in the world vs tens of thousands of professional mathematicians.
Bucket 1 is basically an "amusement for the upper classes" type of deal. Any field that goes in that direction would have to shrink down massively.
It also devalues the field in my opinion from something really profound with actual impact in the world to a somewhat vain leisure activity. (basically going back to gentleman scientists) But I know other people would see it exactly the opposite way.
socializer 2 hours ago [-]
> There's a reason there are maybe a few hundred professional marathon runners in the world vs tens of thousands of professional mathematicians.
I don't buy this. I think there are other reasons why "professional marathon running" is a niche thing; it's probably just that it's not all that interesting to most people to practice in function of what it demands of your body, and not that interesting to patronize / watch.
Take woodworkers. There's probably more professional woodworkers than professional mathematicians. In terms of utility, everything a woodworker does, a machine can do more cheaply and more quickly. The main reason the craft survives is just that we attach intrinsic value to furniture made by humans the old-fashioned way. This is shared by craftsmen and those who buy.
And you could argue the same thing you did for marathon runners: custom furniture is just amusement for well-off people. Sure, but there's enough people with money to keep it afloat.
MisterMunchkin 8 hours ago [-]
Ah but your buckets are based on your assumptions about the topic.
You say marathons are just for fun, but prior to the wheel it was the only way to get around. (Other than horses in some places)
So it’s not that running is immune to automation, it’s that we’re already post-automation and that only people doing it for fun are left.
autuni 9 hours ago [-]
I think your second bucket would need to be changed, as it's not really undecided for some and Tao in fact argues (in a way) that there is no need for the field to decide. There is a need for results, but there is also a need for humans to fully understand why the results are the way they are and how they were obtained. This doesn't replace the human in the loop it's more of a cooperative effort.
I'd end up with the buckets
a) Fully human
b) Hybrid human-AI
c) Fully AI
lioeters 1 hours ago [-]
Extrapolating from the list, the "Fully AI" camp is going to push the pedal to the metal, investing time and effort into maximizing their fusion with this technology. Mentally enhanced and amplified 24/7, and likely physically too, fusing their biology with hardware and robotics. Within a few years, not even a generation, they will become unrecognizable to the "Fully Human" category.
2 hours ago [-]
faeyanpiraat 9 hours ago [-]
Seems like neat buckets but you may want to find one example for the first which can actually be done by an llm as running is not its strong suit afaik
socializer 9 hours ago [-]
I gave two examples for #1 and one of them is something done by LLMs. Other examples include painting / drawing, songwriting & composing, etc. In all of these, gen AI output is pretty widely stigmatized.
qarl 3 hours ago [-]
Heh.
Math? Everyone is talking about math, with math-centric discussions and solutions. I think the obsession with math is simply to distract ourselves from the fact that it's coming for every occupation.
None of these essays even remotely consider it. Denial is a helluva drug.
draw_down 3 hours ago [-]
[dead]
alkyon 5 hours ago [-]
OpenAI counterexample solves the millennium problem in a Yes/No manner - we do know in general that singularity could arise, but it still isn't solved in terms of "Why", ideally we could take an arbitrary physical problem and decide if Navier-Stokes equation can be applied to it.
This is similar to general solvability of the quintic equations - Abel provided a proof first but only with the advent of Gallois theory we could basically understand it in full and decide for any quinitic if it's solvable by radicals or no.
cmrx64 6 hours ago [-]
Stephenson’s Anathem solves some of these problems with cloistered groups (“maths”) that only interact on an epicyclic basis so that old wisdom becomes stabilized with new knowledge has supports for metabolization. I would enter a math.
2sk21 5 hours ago [-]
I also love the concept of Maths in Anathem. One of my all time favorite books!
manlymuppet 6 minutes ago [-]
Bruh this is a tweet. Who else could you make a whole post for, when it's just a tweet?
Not hating, it's just funny.
bklosky 59 minutes ago [-]
Rent seeking is pretty ugly. I think Tao’s perspective is nuanced here, but calls to prevent OpenAI from making new discoveries by the AHM are sad to see.
pama 3 hours ago [-]
These posts came shortly before the math release by OpenAI. I hope the continued discussion of the goldmines in their release helps clarify a vision for how to proceed.
5 hours ago [-]
soundworlds 4 hours ago [-]
I think the real question is "can we apply these solutions?"
It is important that we stay focused - this is theatre. Incredibly impressive, but this doesn't yet show evidence of helping society, which is the whole reason we were doing this in the first place.
Same here. Firefox on Windows 11 looks awful. Edge looks better (I don't use Edge)
oybng 2 hours ago [-]
What an incredible distraction all of this is
bethecloud 2 hours ago [-]
Then Math2.0 will be "AI please explain this solution"
aerodexis 1 hours ago [-]
I, for one, am profoundly thankful that the ~300 year mania around innovation and newness is beginning to wane. You don't need to read that many history books to appreciate that the values that we have today are profoundly different than what they were previously. It takes a few more books, and a pinch of humility, to appreciate that today's values are neither universal nor the necessary outcome of a teleological evolutionary process. The reworking that Tao suggests here reaches far beyond math. The fact that mathematicians are starting to wrestle with these questions is cause for great optimism. Once VCs do the same thing, we'll be on the other side of the storm.
Havoc 5 hours ago [-]
Feels very similar to the emotional rollercoaster programmers are going through. And presumably soon everyone else
InfinityByTen 7 hours ago [-]
When you mechanise, you take the "meaning" out of the effort and that, in of itself, is death of human pursuit of the venture.
The industrial revolution did that to battles and wars and it inspired Tolkein's lores to a considerable degree. He loathed what mechanisation had done.
I feel something similar is happening to Mathematics. I shudder to think what would come of other human pursuit this mechanisation targets next.
Rabbit504030201 15 minutes ago [-]
Did you seriously meant WARS used to have meaning prior to some point of time?
I hope you simply lost you thoughttrain and did not meant that.
amelius 7 hours ago [-]
For many on HN, mechanising things is the reason of their existence ...
(and I don't even think that is exaggerated very much)
InfinityByTen 4 hours ago [-]
No, you're right. I clearly missed reading the room :D
But I'm glad how Terence Tao sees it. I was relying him to do that for the NS results as well.
I didn't follow through to the PhD, but I spent few years building up to understanding of Fluid Dynamics and Functional Analysis to come close to NS. It's intriguing that it's "solved", but what interests me is then "what do we learn from it" and what lies beyond in non-linearity.
In the 10 years I've spent away from academia, I still cherish what Math taught me best: looking at equivalences and I still feel the kick that I surely wouldn't want an Agent to do on my behalf. NS was never the point. And who can't see that I can only feel that they missed out.
5 hours ago [-]
dr_dshiv 7 hours ago [-]
Maybe, actually, it will rehumanize math by formulating its ideas in a manner that allow their comprehension without 10-15 years of mechanized practice. The average human might understand quantum mechanics and Maxwell’s equations with the right conceptual interface.
optimalsolver 6 hours ago [-]
Without the mechanization of book production, how many would have read his works?
drstewart 7 hours ago [-]
This. Farming used to be such a rich and meaningful job where you connected with the earth, until machinery came along and ruined it.
I loathe what's happened with agscience, all farms should be plowed by hand with donkeys and plows.
InfinityByTen 4 hours ago [-]
Maybe they should, you know.
Maybe OpenAI is exactly working towards so that we need to do that :)
It's fascinating how people like to push any statement to its limits, because it circles around to absurdity and they believe they've made a point. Moderation seems to be chasing you, but you clearly are faster!
drstewart 2 hours ago [-]
Taking a statement to its limits is useful for that reason. This is why it's so common to do in physics & math, in fact!
If your mental model breaks at the limits, it maybe means there's some truth in what you say, but there's a nuance that's clearly missing.
And that's the case here: nobody argues when automation comes for many types of other tasks, so why is mathematics special? Or even art? We should either find that line or accept that maybe they aren't as special as we assumed.
7 hours ago [-]
eviks 8 hours ago [-]
> placed a premium on being the first to solve an open problem
Which is a simple measurable goal requiring little bureaucracy. The mythical "all you need is a pen and paper and a lifetime of dedication"
> "Math 2.0" will need to ... value mathematical progress more holistically
which is directionally the opposite
> community building ... AI can contribute positively
what is this belief based on? Any other communities can illustrate?
jiaosdjf 6 hours ago [-]
I feel like OpenAI should have secretly contacted a bunch of mathematicians and offered to give them "credit" for these "discoveries" as long as they also credited ChatGPT with helping in research. Mutual interest etc.
auggierose 1 hours ago [-]
They tried that with Navier-Stokes. Didn't really work out.
Regex777 8 hours ago [-]
he put in an elegant way, that its not just about the solution it's about how we would leverage the AI for better good.
Which should improve collaboration, Research and Clarity.
I would really appreciate if we come up with protocols for using ai in STEM field's it might be award at first but we could regulate properly using this method.
bluepeter 9 hours ago [-]
I mean, he still seems to underestimate what future models will be able to do. The various directions he wants to reward are also things future models will do far better than humans. I suspect we're better suited to pursuing math like we do pleasure reading... it's enjoyable, can be useful in various situations, but we're clear-eyed that we're not gonna advance the field... and that's okay and doesn't mean it's not still worthwhile.
nba456_ 3 hours ago [-]
So happy to see these luddites get blown out again. Tough luck Tao!
stillcompiling 1 hours ago [-]
Why are you happy about others suffering? Makes no sense
skerit 1 hours ago [-]
Oh look, more people finding out they're not that special after all.
vouaobrasil 1 hours ago [-]
That is true, but perhaps making it so that almost nobody ever feels special is a bad thing for society....
ChrisMarshallNY 7 hours ago [-]
A well-known MIT professor gave a presentation about the advancement of mathematics, back in 1965: https://youtu.be/W6OaYPVueW4
elAhmo 7 hours ago [-]
It is interesting to see that we are using this term Math 2.0 so quickly, after just a few months/year of seeming progress on previously hard problems, for something that has been around for thousands of years.
maaaaattttt 6 hours ago [-]
What's happening in the math field at the moment represents on a smaller scale the issue we (probably) will face when AI becomes smarter than us in general. Do we slow the AI down so we can understand what it does and if it's correct (and align with our values for bonus points); do we, humans, adjust the way we work to the new speed or do we give in and let the AI advance while we're not completely sure of what it does and the correctness of it.
The math field is in the position of showing us the way.
entropi 5 hours ago [-]
I think the bigger issue is that "we" won't be doing any deciding. A handful of oligarchs will. As the technology becomes better than us, I am afraid we will learn that we are (at least seen as) technology as well. And the owner of the better technology will decide what to do with the inferior technology. I don't see anyone trying to design the roads for horse carriages, so I don't understand why would anything change for humans' understanding.
That is, if things go in the current trajectory. I don't see any reason why anything would change though.
7 hours ago [-]
iloveoof 3 hours ago [-]
Math 2 looks a whole lot like software development.
karel-3d 2 hours ago [-]
DeepMind in the Lee Sedol days never cared about Go (the game). Similarly, OpenAI doesn't really care about math.
contubernio 9 hours ago [-]
Most of the problems that have been solved are problems on which a great deal of progress had already been made. Those who work on well known problems posed by famous people are those who suffer the most from this. Those who do their own thing and pose new problems, on the contrary, benefit from it. Suddenly raw technical power and great memory are not so valuable as a broad perspective, structural insight, and wild ideas. Who can be successful in this new ecosystem is different. Some of the elites are (correctly) more threatened by it than some "mid tier" mathematicians. I see lots of opportunities to overcome obstacles in my research program some of which had confounded me for years.
On the other hand, it puts a premium on resources. AI is not cheap for mathematicians. Folks are fancy universities in rich countries with forward thinking ministries of science will have an advantage over the rest.
What is clearly in immediate crisis is the traditional model of doctoral education. Most of the problems that were "given" to ordinary doctoral students are solvable (quickly) even by something like Claude pro. Mathematicians need to adopt training models more like what is done in experimental and laboratory sciences - collaborative and structured.
Where Tao is wrong is in regards to exposition. AI already writes better lecture notes, problems, and exercises for mid level undergrad math classes than do most of my colleagues. It's exposition is generally well structured and clear and it can adjust level on request quite well. It writes research better than most professional mathematicians too.
Rapzid 4 hours ago [-]
Instead of solving problems people will be solving solutions.
adrianN 10 hours ago [-]
I wonder how we could formalize the notion of „interesting“ problems in a way that would allow us to automatically generate new interesting questions from the existing corpus of mathematics.
youoy 10 hours ago [-]
For me this path means that our role as humans is just the understanding of intelligence?
Everything else is secondary (or the last of our priorities) and would be better automated?
This is a hard pill to swallow
latentsea 9 hours ago [-]
We should think about rejecting this AI future.
charcircuit 9 hours ago [-]
A simple way is to just measure how long it takes to solve. If it takes more than 5 minutes then we don't have a good understanding of that area and its worth potentially investing into.
sankhao 10 hours ago [-]
Or maybe we should not let pure mathematicians decide which problems are interesting, but reward the practical applications instead.
zorgmonkey 9 hours ago [-]
This is wildly shortsighted view of mathematics. Historically many of the subfields of math which are presently most valuable were considered useless for decades or centuries. Number theory, non-Euclidean geometry, group theory, and Boolean algebra, to name a few.
sankhao 9 hours ago [-]
This is false dichotomy. Many mathematical achievements were produced while looking at practical problems: Fourier is a good one. For every mathematician that accidentally made a practical contribution, there are many more that produced nothing of value, and would have with a better tuned reward function.
youoy 8 hours ago [-]
I guess this AI wave will divide the population more between people that think that only economic value exists, and people that dont.
Thats one of the timeless human debates.
We are now living in the perfect combo of low morality and general human automation. So i expect the next few decades dominated by people who think (and have a "proof") that doing something without an expected economic gain is useless.
adrianN 10 hours ago [-]
If you only judge by practical applications you eliminate vast swaths of human endeavor.
etdznots 9 hours ago [-]
Human endeavor and achievement does not make number go up for next quarter’s earnings
ruilov 4 hours ago [-]
> "Math 2.0" will need to decenter the role of raw problem solving and value mathematical progress more holistically - for instance by elevating the role of exposition, but also that of community building and opening up new directions of study
It is interesting that AI is not yet superhuman at exposition, or at least exposition that can be understood by humans. But you haven't updated enough if you don't think that will happen soon. I'd also expect for AI to become superhuman at opening up new directions of study and theory building.
> Many fewer seminars, workshops, collaborations, or other activities are being generated from these results compared to traditional breakthroughs...the mere knowledge that a solution exists "contaminates" efforts by both humans and AI to find alternate routes to the problem that reveal additional insight
This is absurd. The mere knowledge contaminates...give me a break Tao! Of course having a (possible) solution changes how we're thinking about the problem. If that's what you mean by contaminate, fine. But if you're a person who's excited, curious, interested in mathematical knowledge for its own sake these AI results are a treasure trove. New approaches to old problems, some old approaches that we couldn't make work before. Why not whole seminars to take one of these results and dissect them, prompting the AIs to figure out where else we can use them, improving and simplifying, etc.
Look, I get Tao's anxiety. The ground is shifting and it's hard to solve for the equilibrium. How in the world do you write a grant proposal today when the person who will read it reads the headlines and thinks "math" is solved. That's something that the mathematical community will need to figure out over time. And it's possible that there'll less money for math research overall. When the marginal cost goes down, the market equilibrium changes (but don't forget Jevons paradox!). So I get the anxiety. I just expected better from some of the top people of the field.
OtherShrezzing 9 hours ago [-]
I have this feeling that the frontier of maths is going to accelerate faster than humans can keep up with it, even if machines get exceptionally good at math exposition. There’ll be some event horizon of new discoveries which are so complex we’ll never understand their intermediate steps. Beyond that point, humans will revert to “Math 1.0”, where we’ll need to rediscover proofs that have already been solved by machines, and we’ll have a pair of frontiers each for the humans and the machines.
lubujackson 9 hours ago [-]
Working heavily with LLMs for the past year has me nodding strongly with Tao's mindset.
AI only take us as far as our imagination thinks to ask it. This can be exhilarating when new models drop every month and we can continually reach a new threshold, basically for free. But it is only a one time gain and ultimately short-sighted. Where I find continuous value is using LLMs to help my understanding, full stop.
I use LLMs all day long as a SWE and I have tried many approaches, but the most satisfying and consistent approach is to lean heavily into understanding a problem space and a solution space. Yes, it whips up architecture and code, but I spend most of my time peppering it with questions about the design and how it handles certain situations, what about this edge case and that security concern and this future product need. I have it write a report breaking down the feature and how it integrates with existing code and if the report is too confusing I have it simplify either the report or the code until it makes sense to me, sometimes scaling back the work to a more manageable state. I do all of this before I look at any of the code it writes.
The difference from this approach is that I am not suffering reading through 3000 lines of AI slop but I am reviewing a PR that I fully understand. I can eyeball it quickly for anything that doesn't fit my mental model and dig deeper or quickly revise it. Only after I am happy with the bones do I consider the meat and skin of the code.
What I find most concerning is how frontier AI companies all seem to have this Math 1.0 perspective that they only want to type "solve Riemann" into the chat box and have the magic to happen. It is the same problem Google ran into, where a simple, no thinking solution serves most of the people best and most profitably, so you fully ignore or remove everything else (boolean operators, exact phrase search, verticals, filters, infinite pages of results, "nothing found" if there isn't, etc.) But that choice leads to the situation Google is in now, scrambling to stay relevant. In a different world, Google would have continuously augmented their search capabilities and eventually built a smooth, guidable AI interface.
But no, we must only have an input box and a Go button.
Everything looks like a nail when you build hammers, sell hammers, have infinite hammers to play with however you like and your company mission is to build a hammer starship to explore the hammerverse, whether or not that is even possible.
oxavier 6 hours ago [-]
Strongly agree. I see the need to understand as back pressure on the system that produces code and features: it limits output to what fits into a working mental model. Agents could turn tickets into code much faster, but I am building a product, and it needs someone who holds that model to steer it long term.
And when the job is done, I run retrospectives on old coding-agent sessions to find areas of friction and confusion. I also journal with pen and paper, as it's supposed to bring cognitive benefits, to help me stay on top of things.
somebodyyoumayn 4 hours ago [-]
why do we need mathematics? to get more grants and medals or to push frontier of humanity?
InfinityByTen 2 hours ago [-]
To have someone stop you in your tracks and ask you to "define" frontier of humanity.
In case you missed, Mathematics isn't about numbers and equations.
vasco 10 hours ago [-]
He is right if model intelligence stalls. If model intelligence continues to improve soon there's no need for the prompter to understand anything or for any workshop as a mathematician will just be able to ask the model to explain how the proof works and models will do a good job at walking them through it step by step.
There will be no gap in understanding. Now there is because the models are discovering things at the edge of what they can do and so suck at explaining it. There's nothing particularly special about a newly solved problem in terms of learning it.
If we accept AI can explain all of existing math nicely, why shouldn't it be able to explain new proofs?
winterbourne 10 hours ago [-]
>if model intelligence stalls. If model intelligence continues to improve
So much in AI is dependent on which of these two outcomes occur.
doctoboggan 10 hours ago [-]
I am not so sure that's true. Even if AI intelligence were to plateau at today's levels, there are still many gains to be had in speeding up today's intelligences. ASICs with burned in weights could become economical to invest in as they would retain usefulness longer than 18 months, and I feel we have only scratched the surface on possible usecases of local AI and what it means for almost any technology product or interface.
myaccountonhn 10 hours ago [-]
If the mathematician doesn't understand anything, are they even a mathematician?
lancebeet 9 hours ago [-]
That doesn't make sense to me. Terence Tao is great at explaining things, but he could spend months explaining some of his proofs to me without me understanding it. Even if the AI had a superhuman ability to explain things it's no guarantee that it could make a human understand.
vasco 9 hours ago [-]
It couldn't make every human understand. But it could make Tao and other mathematicians understand. You're not going to become knowledgeable at everything suddenly, even though I doubt what you say. Many months of 1-1 tutoring and hard work from a student with a top mathematician would make you understand a lot. Just not sit down and read it first day.
doctoboggan 10 hours ago [-]
It's been very interesting watching Tao's evolution on his thinking on LLMs. Of course the LLMs have themselves evolved so that shouldn't come as a surprise.
The job of professional mathematician might be the first to be completely eliminated by LLMs, save for those who can make money from a patron. I am hoping they are able to figure something out to save their profession, as other professions could use it as a blueprint as AI comes for them next.
nl 10 hours ago [-]
> job of professional mathematician might be the first to be completely eliminated by LLMs
Strong disagree.
Do you work in a math adjacent field? I do and I find having a mathematician around invaluable.
It's like a non-software person writing software. Yes, using a LLM will get you to a solution that works. But just talking with a software engineer will make the quality of that solution enormously better.
I find the same with math - I can get something to work using an LLM, but if I speak to a mathematician they'll say some magic words to try and I put that in the LLM and it is "oh yes this is a much better solution".
This is very different work to generating proofs though. Its things like "I'm trying to get my confidence intervals to properly deal with census like sampling but at small sample sizes" (yes, I know stats not pure math but still..)
youoy 8 hours ago [-]
The main issue here is that people that are able to "Strong disagree" is shrinking with every model release, because it requires skills. So it will get more difficult to get funding.
You may say "hey, before AI people payed for math salaries even though they didnt understand the math or the economic outcome". But the issue is there are now "mathematitians" trying to convince not to fund.
In this new reality, you will get "mathematitians" trying to convince that only AI maths matter. And on the other side someone speaking about "understanding", "taste", "community". And the people deciding to fund dont have the skills to differentiate. So they will fund the AI boosters with a higher probability.
Thats how the job of professional mathematitian dissapears. By being replaced by something that on the surface looks similar, but its just an ugly copy.
doctoboggan 9 hours ago [-]
Yes, I am projecting out into the future, not talking about the current state. I said "might be the first".
bonoboTP 9 hours ago [-]
That's an uphill battle. People will hand wave that the s curve has flattened, and think it will stay at current levels. They claimed this confidently year after year over the last few years.
kgwgk 8 hours ago [-]
Other people have claimed confidently year after year over the last few years that the latest model was AGI - or that the singularity would arrive in a few months anyway. All the battles are uphill when you’re at a local miminum, I guess.
bonoboTP 7 hours ago [-]
AGI has no goalpost-shifting-resistant definition anyway. In my estimation, the current tech is AGI by how I understood the term when it was introduced (at a time when AI as a label got adopted for things like image classifiers and machine translators and people needed a label for something that can work across tasks with a pure natural language interface telling it what to do). (And no, sentience, consciousness, inner experience, some amorphous "real understanding" etc was not part of the definition) We are in the scifi future just got used to it. Or haven't realized it yet because human systems have inertia and we could some oracle philosophers stone that knows the answer to every question in a nanosecond and we'd still lag 10 years in how to fit it into existing work flows and getting wet ink signatures onto paper authorizing it's use in the bureaucracy.
rXwubXUGAm 8 hours ago [-]
We go Math 2.0 before GTA 6
3 hours ago [-]
alphawhisky 3 hours ago [-]
Hey, Civil Engineer here. Although I totally understand what our pure math friends are saying, I want to emphasize that the nature of math is interpreting and validating results, not doing the math. You may not be aware from the other side of your moat of very difficult or abstract concepts, but technology has been slowly removing friction and optimizing for efficiency for every other math user for a long time now. I understand that it's not fun to have AI do an aspect of your job better than you, but as perpetual students it's our job to turn the tools into progress. Also, I think that anyone standing on the "no AI in math" hill will die on it, and will die on it soon.
i2km 7 hours ago [-]
Curiously fitting that the Pope is also a Mathematician by training. I'd imagine he agrees with Tao on the fundamental aims of mathematical research and how these proof dumps largely miss the point
spuz 9 hours ago [-]
Perhaps it's the moment that the likes of Terence Tao hand over the reins to the likes of Grant Sanderson.
nurettin 5 hours ago [-]
After seeing the cancerous monster AI "proofs" it became obvious as day that we need better abstractions in several fields. The good part is: there is still a lot left to do for humans as they actually have the abstract thinking ability, as opposed to crazy cancerous token generation.
bonoboTP 10 hours ago [-]
I imagine doctors will also clutch their pearls when Ai starts curing disease. "But curing disease was never the point! These arbitrary dumps of AI cures for cancers is unsustainable! Who will think of the doctors and who will build their communities further? From now on progress in medicine must be redefined as what makes doctors thrive, not what generates cures!"
omnicognate 10 hours ago [-]
That's a strong contender for the worst analogy I've ever seen on HN, and it's a crowded field.
Edit: If you'd like a better medicine based one, look to radiology, where AI is an omnipresent tool but claims that radiologists are no longer needed, based on an ignorant view that a radiologist's job is "classify images according to what diseases they indicate" have only contributed to a crippling worldwide shortage of radiologists.
bonoboTP 9 hours ago [-]
[flagged]
xyzzy123 10 hours ago [-]
I think the difference is that with cancer cures we mostly care that it works as proved by trials, and understanding it is a bonus.
Up until now the prize in pure (as opposed to applied) mathematics was the _understanding_ and the machine can't do that for you. What does it mean if we get "super powered alien maths" but humans can't do it? It's like inter univeral teichmuller theory but imagine if Mochizuki was right and it came with a lean proof?
bonoboTP 9 hours ago [-]
It was the understanding for some, and the result for others. Math is going to bifurcate along those lines. Many are the builder type who use math for a purpose. To accelerate an algorithm, to improve the numerics or convergence of a computation, to verify statements about real things, to use it in angineering applications, etc etc.
In my view, over the last century, math has turned into an intellectual analogue of extreme bodybuilding competitions. A navel gazing runaway optimization in making useless stuff just to demonstrate cleverness. That's fine, why not. But society has no obligation to fund that, just as it doesn't fund other extreme hobbies. Ideally if we ever get something like UBI, math can be still their hobby.
xyzzy123 9 hours ago [-]
Right. What I not sure about - even for builders - even when the results are technically verifiable, the details can still matter. AI is really good at proving or building a slightly different thing than you thought you asked for, and if you're not at a level where you can understand key details of the formalisation of your ask, you cannot safely use the results.
Without human understanding you also might literally have no words for the thing you would otherwise want to ask for.
I think it'll be wildy useful but I also suspect human competence will still matter.
MrToadMan 9 hours ago [-]
Apart from the fact that elite sports are funded probably at least 10x more than pure mathematics research.
bonoboTP 7 hours ago [-]
That's mainly for national pride reasons like the Olympics, world championships etc. Others are profit driven when there is spectator interest.
So make mathematicians proudly wear their flag colors and sing the anthem while lecturing to cheering spectators. Might work.
ks2048 9 hours ago [-]
Doctors don't cure diseases, they treat patients.
Maybe AI will take over some roles of doctors, but that's independent of curing diseases. Think about diseases that have a cure - do people with those diseases not go see a doctor?
ziiinq 9 hours ago [-]
[dead]
BrenBarn 6 hours ago [-]
> And our community will need to explicitly re-evaluate its criteria for education, publication, and career advancement, to reflect the "Math 2.0" era.
I think most academic disciplines would benefit from such a re-evaluation. AI is still a scourge on the earth, but I suppose if it spurs such changes that's a modicum of a silver lining.
youoy 9 hours ago [-]
Can we please stop reducing human activity to "taste", conferences, talks, "understanding"? I think this is a very unproductive trap.
There is a world where we get to the edge of AI capabilities, and we build on top of that. As humans have always done with every new technology.
There is another more pessimistic view where LLMs just replace every human capability, and our economic overlords dont need us for anything and we just eat the small pieces of bread that are left.
This comes down to the fact of:
is human existence/intelligence just the simbolic representations we make in our brain? Or are they just a tool?
I tend to think of Godels incompleteness theorem as a proof that on the limit LLMs are useless. The real question for me is at what point approaching this limit becomes an issue, and if it has any practical consequences.
latentsea 9 hours ago [-]
>There is a world where we get to the edge of AI capabilities, and we build on top of that.
We don't build on top of that. No need for us to. AI does. That's sort of the whole point of this endeavor is it not? Humans need not apply.
youoy 8 hours ago [-]
I know that is the whole point. But i may dream of being as tall as a house, but it does not matter how hard i dream, it wont happen.
In my experience, every new model release allows me to go further, although every time i see every time the limitations, and i identify where i add value. And this value gets bigger every time.
But on the other hand, every model release reduces the amount of people that are able to value this "added value", because it requires more skills.
So we have the paradox that the added value i can bring on top gets bigger and bigger, but the perception of the economic value for the majority of the population gets smaller.
ChrisArchitect 10 hours ago [-]
Related:
AHM Statement on OpenAI's October 6 Release of Mathematical Documents
I think it's disappointing [1] for a group of people (however smart they are, whatever titles and fame they have) to think they are the gatekeepers of mathematics, the core of human knowledge. In this respect "mathematics is what mathematicians do" is no longer true, perhaps it was never a useful thing to say. Mathematics is (downstream) consumed in some shape or form by every human on Earth, and no elite, appointed group gets to decide how new mathematics is created. I feel this is a reaction to the sting of the reality check for these people, who are gifted with amazing genetics, and can think faster and deeper than other humans, that their gifts are relative to humans, and machines can still outdo them.. but this is already true for so many other areas, from physical strength to playing chess. Also, I feel this reaction is just non-sensical, it reminds me of when commercial software company CEOs [2] said open source is communism in the early 2000s, because they felt open source is threating their business model.
Lastly, deep down I don't really get what mathematicians are so upset about. All open problems, once solved, are not solved by 99.9999% of mathematicians, because it's solved by one or a handful of others, and the others just learn of the solution/proof. Mathematicians can now still organize conferences about these proofs, discuss them, digest them, think of new avenues of research, etc. They don't even have to invite OpenAI, in 6 months whatever model is available on chatgpt.com will be this smart anyway, and they can use it in the workshops for explanations, etc.
[1] I was going to write "I'm a bit disappointed by the response of the math community..", but then I remembered, whatever T. Tao writes is not the position of the math community, it's his position. Then I was going to write "I'm a bit disappointed by the response of T. Tao..", but then I remembered, I don't know Tao personally, so why am I disappointed?
[2] Steve Ballmer of Microsoft, I believe
js8 7 hours ago [-]
I would argue the problem is not AI, but putting value on proofs by mathematics community. They're not fully to blame, part of it is need of billionaires (and other elites) to appear educated, and the general status culture that sees mathematics as something that needs to be socially justified.
Grigori Perelman warned about this when he refused the Millenium problem prize. He understood mathematics should be a journey, not a destination.
fithisux 4 hours ago [-]
""Math 1.0" placed a premium on being the first to solve an open problem, even if the solution was not initially well understood"
Arbitrary conclusion. This is the corporate take on "mathematics"
Mathematics were meant to further our understanding of nature and solve people's problem. Not to serve corporate delusional CEOs for their psychopathic purposes.
patternMachine 10 hours ago [-]
"Taste"
Razengan 7 hours ago [-]
I get bonked for saying this again and again, because people get annoyed with an oversimplified analogy or the use of whateverfallacy,
but can someone please try to set aside their knee-jerk reactions for a while to give a good reason:
WHY do humans NEED to understand the basics of something?
× You don't know how to farm — That doesn't prevent you from having food or cooking good meals.
× You don't know how to mine raw materials — That doesn't prevent you from using computers/phones made with aluminum, copper, glass etc.
× You don't know how to fell trees and shape lumber — That doesn't prevent you from sitting in that comfy chair.
× You don't know assembly language or how to write operating systems — That doesn't prevent you from using Windows or macOS or Linux.
—
EVERYDAY you use hundreds of things made from THOUSANDS of technologies you don't understand, because other people already MASTERED them.
so YOU can go on to go do GREATER things.
(but you CAN still go do farming, mining, logging, writing your own OS, if you ENJOY it — nothing's stopping you — you just won't be as good as the technology that has been specialized for that over centuries, and almost certainly you won't be bringing anything new to those fields, and it'll take time away from doing other things.)
—
Maybe we shouldn't be wasting time on "oshit how do we uninvent or slow down this new technology because it makes things easier than what we grew up on"
and focus more on "what other greater things can we move on to?"
There's a whole freakin universe out there and we haven't even stepped off our home planet yet.
sham1 7 hours ago [-]
Firstly I'd point out that the difference here is that all of those things you mentioned are things you could learn and the distinction would be that at least currently the proofs provided by AI are inscrutable. Although this will probably change.
The second thing is to just ask what is there left to do. What are these "greater things" that people can dedicate time to, when clearly even classically cerebral activities like mathematics can be automated away. The industrial revolution already wrecked physical production of goods and made artisan workers obsolete outside of extremely niche scenarios -- that's why we call things artisanal, after all -- but there was still mental work. But now, mental work is also experiencing the same thing, and it's not clear what one should do as a human anymore.
And some people seem outright gleeful about these developments, which can be seen even in this thread. What happens when humanity becomes obsolete? And what happens when the machines that cause this obsolescence are controlled by a tiny amount of people, who suddenly don't need the rest of us? I can only hope that this turns out well for us and that with the development of these machines, humanity will get better, but the omnipresent existential dread is giving me doubts.
Razengan 7 hours ago [-]
> What are these "greater things" that people can dedicate time to, when clearly even classically cerebral activities like mathematics can be automated away.
Their applications?
For example I'm not a mathematician but I love thinking about weird "useless" shit like how math might be like for aliens? Are numbers as fundamental as we assume? i.e. humans developed math for "arithmetic" first, then latched geometry etc on top of that. We took ages to admit zero and negative numbers.. what if an alien species develops math for "navigation" first, and starts out with complex numbers right away!?
> What happens when humanity becomes obsolete?
There's an infinity out there to explore.
> And what happens when the machines that cause this obsolescence are controlled by a tiny amount of people, who suddenly don't need the rest of us?
That's a social problem we needed to tackle more than 100 years before AI or even computers appeared.
Apparently the minutes hand was added to clock to keep time in factories, for the benefit of the factory owners, not the workers — something I learned from this 1991 show from the BBC with Terry Jones: "So This Is Progress" https://www.youtube.com/watch?v=-Em96NVxO9Q
ErdosRoland 7 hours ago [-]
But I am good at one skill that I spent years studying/training for, and now that skill is not paying the bills.
At least, that's how I view most discussions on AI adoption. Technological advancements are great for humanity, but that doesn't mean it comes without costs. The luddites are a famous example that's very often mentioned in this forum.
And no, "reskilling" isn't an option for many people. If you are poor, if you have a family or have people dependent on you, you cannot put your life on pause to learn something new, especially if you have no guarantees it won't end up like last time.
Razengan 7 hours ago [-]
> and now that skill is not paying the bills. ... you cannot put your life on pause to learn something new
Yes, but that's a social issue, external to technology but exacerbated by every new technology, AI or not:
UBI should be a thing: let people work on what they find fulfilling, instead of having to work to survive.
AI could help design a system for UBI that everyone agrees with, since it's so good at maths and shit now.
This problem HAS to be tackled. Removing/slowing AI will only kick it further down the road, not eliminate it.
ErdosRoland 6 hours ago [-]
If UBI was a thing and people felt safe that they would not be poor, then obviously AI (and all technological innovation for that matter) would be much more welcomed. I do not see such initiatives being taken however.
It seems to me that all we are doing is a wild goose chase; Progress above all, to hell with any environmental/social impacts, the end justifies the means.
> This problem HAS to be tackled. Removing/slowing AI will only kick it further down the road, not eliminate it.
True, but people are generally selfish. They will put their own prosperity above that of the future generations, and I cannot blame them.
Razengan 4 hours ago [-]
> people felt safe that they would not be poor
The actual fear underpinning this isn't even about being "poor", it's that being poor means starving, freezing, not having a bed to sleep on, not being able to get basic healthcare in emergencies..
It's possible to provide all those things without "giving away free money" to everybody, but..that's probably more complicated for now.
In any case, whether one "deserves" to live in basic comfort shouldn't depend on one's ability to do jobs that depend on holding back technological progress.
ErdosRoland 3 hours ago [-]
> In any case, whether one "deserves" to live in basic comfort shouldn't depend on one's ability to do jobs that depend on holding back technological progress.
It shouldn't, but it does. So if we can't change this fact, we have to find a middle ground so that the people alive right now aren't thrown under the bus.
I don't disagree with what you are saying. Progress is inevitable in the end. I just wonder if we have to be destructive in our road to achieve it. Environmental and social damage are also problems that we need to tackle. Keep in mind that unstable societies, where people are fearful of the future, are prone to revolutions, and an unstable political climate is detrimental to technological progress.
I see very few people at the top speaking up about this, and that only makes me more skeptical of the usefulness of AI. If it's only going to be used against me, why would I ever support it?
zhivota 5 hours ago [-]
People knowing things represents redundancy. It has long been a fear around technology that humans would lose the knowledge and if the technology fails or is somehow taken away, be helpless.
This is what underpins the fear, I think.
You don't know how to farm but mentally you rest easy knowing a lot of other people do know. You also know there are books you could read to learn, if you needed to. Most things are like this, you could bootstrap your way to casting metals and probably even electric lights with only books and raw materials. Computer chips don't have this property.
Personally this is why I'd like to see libraries survive, even though I actually mostly read on my e-reader. I guess I took Anathem to heart.
bbor 8 hours ago [-]
This is all reasonable of course, and I think you're have to be pretty cold-hearted to disagree too vociferously. What OpenAI did (opening an announcement by name-dropping the exact institution that told them not to, implying they got buy-in) is just objectively bad-faith, and more of the same from Mr. Altman.
Mathematics is an academy, and academies are human assemblages for producing truth (and the tools therein); they will stop producing if we forget to repair and refine them. It's just undeniable in the abstract.
(Sorry for the length, cut it as much as I could; mod(s) remove if you'd like. Talking to myself in the shadow of giants is how I'm coping with the ennui, I think.) That said, four philosophy nits on paradigms, scope, motivation, and pride:
1. Paradigms | The 'Math 1.0' rhetoric is undeniably powerful, but it makes it seem like he's unaware of his standpoint[1] by lumping all of "traditional mathematics" together. At the very least we've gone through four methodological revolutions in math, each one changing how the field is done on a fundamental level: ??? => Euclidean Certainty => Aristotlean Computation (~800s) => ~Newtonian Calculation (1600s) => ~Gaussian Systems (~1850s), and perhaps one in the 20th c. I lack the expertise to even gesture at. We also have clear analogues from parts of the other two acadamies in the 20th century alone: physics becoming an arcane, inelegant group effort in the ~1920s, and mainstream philosophy adopting a cloud of Kiki ideas vaguely revolving around Wittgeinstein & Chomsky in the ~1960s.
I totally understand this being distressing, especially when it's happening quickly. They, too, had people decrying the future of their fields. But we wouldn't obviously wouldn't change it, in hindsight; much of modern physics would be completely intractable without those strange, boring, unnerving methods, for example. More than intractable: unthinkable.
2. Scope | This all seems overly focused on Autumn 2026. Most egregiously, this is all built on the premise that RSI never happens, and we never acheive ASI. If we do, mathematics is almost assuredly A) the first academy to be completely outmoded, and B) the least of our problems. I cut a long thing about the caveats and effects here; at this point... if you know, you know.
3. Motivation | Ultimately this thread is focusing on human motivation throughout, a fact that would be more forgivable if acknowledged as an intentional tradeoff. Speculating that it'll be harder to have interest in math is just not worth withholding truth; for one thing, knowing that computers could solve a problem but it's banned to try would ruin motivation anyway, and worse. It's up to us to be motivated, and if I know us, we'll have no problem doing so as long as there's any utility there at all.
In more stark terms: trading progress in the fundamental academy for the sake of its current methods of recruitment and motivation seems like something posterity will almost definitely frown upon.
4. Pride | This is the common thread that weaves through all three preceeding points, I think, and is even stated in pretty blatant terms (that's Tao -- always a clear writer!):
...promising open directions are now being withheld from the public in fear that this will cause their own research to be "scooped"... "Math 1.0" placed a premium on being the first to solve an open problem, even if the solution was not initially well understood.
Sure, his thesis acknowledges that some changes are welcome, but not radical ones; his tone implies tweaks to conference schedules and authorship norms rather than fundamental restructuring of what these professions are, and what it's like to dedicate one's life to the demos through them.
Doing science (mathematic or otherwise) in this competitive, individualistic way is just clearly counterintuitive to me, even if it weren't a recent development. Imagine taking it to its conclusion and applying some kind of patent system to mathematics -- or even worse, copyright to combinations of symbols! Perhaps more riches would motivate some mathematicians, but it would so obviously eat away at the democratic principles that have brought us unimaginably far over the past 406 years.
---
TL;DR: What worked well for the past ~century is not particularly relevant, and I think Tao is missing the forest here, despite one of the best sylvan trailblazers around. On his side practically-speaking for heuristic and contingent reasons, regardless.
codewithcheese 10 hours ago [-]
[flagged]
mundane799699 6 hours ago [-]
[flagged]
squirly 7 hours ago [-]
[flagged]
marsven_422 4 hours ago [-]
[dead]
aaron695 10 hours ago [-]
[dead]
teekert 10 hours ago [-]
I think we can say by now that we should not listen to the early nay-sayers and just wait a bit. With every trend, not just "AI". They still have some points (the ethics and environment etc), but the we don't hear from the Stochastic Parrot folks anymore.
Of course it's good to have the discussion... So maybe, we listen to the nay-sayers, but defer judgement on the matter... That's wisdom.
Edit, to be clear, I consider Tao to be the wisdom provider, not an early nay-sayer!
tmule 10 hours ago [-]
Yann LeCun and Gary Marcus are still at it - the latter having shifted goalposts.
nl 10 hours ago [-]
Yann LeCun's criticisms are at least balanced with an alternative approach he has teams actively working on and showing good progress in areas LLMs are weak.
The less said about Gary Marcus the better.
zwaps 10 hours ago [-]
Gary Marcus has has a long career now of shifting goalposts
vidarh 10 hours ago [-]
Does he have time to do anything but shift goalposts, given how fast his are moving?
Chance-Device 10 hours ago [-]
I believe he now just leaves them in his truck and drives it a bit further each day.
computerfriend 10 hours ago [-]
Tao was not an early nay-sayer.
jatins 8 hours ago [-]
If anything he is very pro-AI and far from a naysayer. Labelling anyone who doesn't isn't Linkedin style "stoked" "excited" about AI is not helpful.
This is a technology not like prior technologies. Are we okay if the technology discourages a whole generation of Mathematicians? If the technology leads to 10x fewer mathematicians -- what impact does that have on the field? These are the questions Tao is asking. And I don't think he himself claims to have all the answers, he just doesn't wanna see the math _community_ die.
teekert 10 hours ago [-]
Certainly not implying that! He's the one with the deferred judgement and the wisdom.
computerfriend 10 hours ago [-]
Oh, sorry, misinterpreted you! Although I am not sure I agree with your dismissal of nay-sayers. We can equally dismiss the opinions of early proponents. Perhaps maybe we should be hedging on all early opinions regardless.
teekert 10 hours ago [-]
Yeah agree with you, and perhaps both are even important in our public opinion forming... Recently I've been hearing people like Grant Sanderson on AI, and they are all very "wise" and informed, neither dismissive nor mindlessly (p/b)ro, but really thinking implications of this technology through. Like Tao does here. I love it, these people provide real direction.
10 hours ago [-]
Chance-Device 9 hours ago [-]
“doing math responsibly” is just a euphemism for “ensuring we can continue doing our paid hobby”.
The interviewee is worried about the future of math research. He is not strictly worried about being replaced, instead he is worried that he will no longer be able to launder math-as-a-hobby through math-as-something-useful as is the case today. He lays out very clearly that grant proposals claim to have useful outcomes while the proposers know those claims are nonsense.
Business as usual in math, and frankly in all the other sciences, is to do research that furthers the researchers careers or personal interests and pretend that it’s somehow useful. This would be absolutely fine if it were privately funded, but it’s not, this is public money.
In every other endeavour, lying in order to get money is considered fraud.
We have collectively wasted a huge amount of taxpayer money and human time, entire careers, on things not likely to ever matter to anyone.
I look forward to science becoming automated so that we can have real progress instead of the current broken system.
chii 9 hours ago [-]
usefulness can only be determined after the fact.
People thought, back in the 17 century, that imaginary were useless (except as a trick for some calculations). Turns out the research into these numbers back then is amazingly useful today, 300 years later, in electronics and such.
Publicly funded maths research should continue, even if some taxpayers feel it's a waste of money.
Chance-Device 9 hours ago [-]
Nobody determined the usefulness of the internal combustion engine or the aeroplane after the fact. They were goals specifically worked towards. For every useful discovery that comes from this hobbyist approach there are far more that aren’t, and those would have been discovered during a goal orientated research program.
However, with AI it actually may become so cheap that the scattergun random approach becomes more viable rather than less. It’s when human time and resources are scarce that you need to optimise. The hobbyist approach may therefore ironically continue, but without the hobbyists.
dataflow 9 hours ago [-]
> Publicly funded maths research should continue, even if some taxpayers feel it's a waste of money.
Note, I think the debate is mainly over what research should be funded, not whether any research should be funded.
chii 8 hours ago [-]
it's the same - what research getting funded should not be determined by the applied nature of the research.
In fact, the taxpayer should be funding fundamental research because it's so hard to justify profit from it - but that funding would benefit all in the long future.
So leaving the applied research that have commercial value be funded by private, commercial interests, would make more sense.
stale2002 9 hours ago [-]
> usefulness can only be determined after the fact.
Thats great! So in other words it should be perfectly fine for AI to solve all of the supposedly useless math problems so that it might possibly be useful later. No need to worry about the mathematicians hobbies here.
chii 8 hours ago [-]
> No need to worry about the mathematicians hobbies here.
of course not.
Hobbies can still be done even if AI gets it done more quickly. Just like today, where knives are made much faster/cheaper in presses and CNC machines, vs a blacksmith hammering. But still, there are hobby blacksmiths.
JumpCrisscross 9 hours ago [-]
> look forward to science becoming automated so that we can have real progress instead of the current broken system
Do you really think a society with zero human mathematicians or scientists will outperform one with both human and AI ones?
Chance-Device 9 hours ago [-]
Actually yes, I think that AI will end up in a place where additional human effort, even at the highest level of suggesting what directions to look in, will become a rounding error compared to what the AI will achieve by itself.
So as measured by utility, I absolutely believe we don’t need humans doing science into the future. I’m sure people will continue doing it, but not for utility, for enjoyment - as a hobby. Probably we’ll all end up as dedicated hobbyists.
vuurmot 9 hours ago [-]
>We have collectively wasted a huge amount of taxpayer money and human time, entire careers, on things not likely to ever matter to anyone.
How much..?
Kinrany 9 hours ago [-]
The whole point of public funding for science is that seemingly pointless research yields useful but hard to monetize discoveries. Otherwise VCs would be doing it.
xyzsparetimexyz 8 hours ago [-]
The software equivalent of this is claiming that game development is laundering software as hobby. Absurd.
Chance-Device 8 hours ago [-]
Haha. That’s almost verbatim what the guy said. Here, let me grab that for you:
> things that I do and that my colleagues do, this kind of like curiosity-driven, you know, applied math, computational physics type research has always been justified by, I would argue, intentionally blurring the line between what I would call, you know, science as product versus science as process
> science as product is very kind of clear-cut. It’s, you know, things like, you know, cure cancer, solve nuclear fusion, generate, you know, clean energy.
> And then there’s science as process, which is kind of the curiosity-driven stuff about, you know, like, “I want to understand protein folding,” or, “I want to understand, you know, turbulence,” or, “I want to understand quantum gravity,” or something. And broadly speaking, we have tended to justify the latter by kind of laundering it through the former
And the examples he gives are actually the more defensible ones, he talks about a friend of his working on some abstract algebra under the false guise of cryptography research later.
And it’s not just him saying it, this is simply true. He should be lauded for admitting it publicly, this is the only way any progress is made. At least, it used to be. Now it’ll be AI instead.
10 hours ago [-]
SubiculumCode 8 hours ago [-]
I feel like there first will be a lot of rationalizations, hamd-wringing, and existential tummy aches, but then the eventual resigned acknowledgment (just like in Chess), that humans do math because they like to do math, not because humans will ever again be as good as computers are at math, then come to make discernments between human math proofs, and the work of an engine, maybe they'll even start start doing proofs on short time controls and stream it on Twitch.
Because the other solutions are to a) quite literally become inhuman, with cyborg integrated TPUs running local models and networked interfaces to propierary models run in data centers, or b) assert dominance of human ignorance by burning civilization down, which doesn't sound pleasant.
injidup 10 hours ago [-]
I suspect the whole field of mathematics will simply disappear as a career path. It seems obvious that the trajectory is for the machines to be able to provide proof on demand for any solvable problem. Whether or not the proof is understandable by humans is perhaps irrelevant in the larger sense. Doing hard math will simply become another black box tool in the larger AI toolkit for goal optimisation. Is this sad and should we try to prevent it? Is it any less sad than the venerable London cabbie who spent a life time memorising every street to gain "the knowledge" and almost overnight supplanted by machine intelligence.
JumpCrisscross 9 hours ago [-]
> the whole field of mathematics will simply disappear as a career path
Almost certainly not. It's just going to jump to a higher level of abstraction.
latentsea 9 hours ago [-]
It's already unbelievably abstract as it is...
hodgehog11 8 hours ago [-]
[dead]
croes 9 hours ago [-]
Did the London cabbies explore unknown territories to find new streets to memorize ?
Rendered at 15:32:35 GMT+0000 (Coordinated Universal Time) with Vercel.
But there also is another type of discovery that requires taking a step back and looking at the problem from a different angle. If you are an engineer, how often has an LLM told you (without you explicitly prompting for it): Wait, what you are doing here doesn't really make sense, there exists a much more elegant abstraction that nobody has thought of, let's remove all that code, let's tackle the problem in a different way by thinking from first principles. Pretty much never. But in science a lot of the biggest discoveries have come from this kind of first principle thinking, questioning existing work and approaches and going against what already exists, not combining all existing data which is likely to be just a local optimum.
If we were using LLMs to analyze astronomy, we might just get increasingly complicated epicycles and never realize that the Sun is the real center of the solar system. We then never learn about the anomalies in Mercury's orbit that led to the theory of relativity.
So I agree. I seriously question how valuable LLMs can be in science/math beyond working as advanced search functions.
This has not been a pattern at all. Many have speculated that they would be good at this, but in the past they actually weren't! They were unable to synthesize their encyclopedic knowledge of everything into cross-disciplinary new discoveries, without explicit prompting about the kinds of knowledge to combine. They were surprisingly bad at this!
These math proofs are the first evidence I know of for LLMs actually taking advantage of the fact that they have more knowledge than any one human could have to combine multiple different directions in unprecedented ways to solve real problems. This is new, and exciting.
You have to ask for this. As in "I'm not sure about this approach due to X, Y, and Z. Can you think of something more elegant?" It works!
But also, how many people actually need to do the kind of "deep" work you're claiming LLMs can't do? Most people aren't contributing to the frontier of anything. I'm not.
Finally, I think you're appealing to a fuzzy distinction. The difference between a "genuinely new idea" and an idea that "combines existing ideas in a new way" just isn't very well defined. In retrospect, a lot of the most revolutionary idea look like a combination of many, smaller, prior ideas.
This is also why they're so good at creating three.js or Blender work when the output is so easily constrained to "Look exactly like that". I recently posted https://www.ambionix.com/blog/introducing-the-czp-1/ on here, and the audio engine in that was developed in that way.
It is true that it would be astounding to find if anyone has seen a LLM produce any useful generalization of anything resulting in a simplification. They seem to have a direct tendency to do the opposite. The brutal reality is humans have also undervalued this capability for a long time (I think the Poincare/Hilbert debate is relevant) to the point we are also taught that generalizations are, generally, bad and wrong.
Eventually, by walking through them, it proposed an additional fifth value and from there was able to tie everything together.
Sometimes you just gotta hit the machine until it works again.
Like, I think there have been attempts at this across the field. (I could be wrong!) But it requires a lot of labor and a lot of cross domain knowledge to complete. Both things that AI have.
Simplifying all of math is another endeavor, but I imagine that you could have a bunch of LLMs trying to shorten existing Lean proofs.
You literally just have to ask it. Before I left software engineering in April, I was using Claude for re-architecture all the time.
But no, it doesn't assume it should re-architect what you're handing it when you haven't asked it to.
> You literally just have to ask it.
why doesnt it ask itself before proceeding?
And yet you still just get slop.
I explicitly request it. It's not great at coming up with interesting ideas, but neither am I, and it can sure iterate on them faster than I can...
There seems to be more interest in hitting some arbitrary benchmark (we proved X unsolved problems) than in genuinely contributing to mathematics. But what else is to be expected? It's become a maniacal race with too much money. Too much effort is being invested in proving that the exponential curve is still holding.
In example of go where I'm more familiar Google deep mind poured large resources to get a super human performance first, establish superiority and abandon it. The community then built their own tools starting from reproducing their papers.
I think similar thing might happen to math. Nobody outside of math cares too much about Hamiltonian cycles in some bizarre graphs or proving lower bounds on complexity of some problem.
Once those results stop being worthy of mainstream media attention, they will abandon math and the progress will be done by mathemicians guiding the models and the community will likely establish some new rules about what makes a valuable contribution. Merely solving not yet solved problem might not be it anymore.
Pretty much the only enterprise that historically pays some mathematicians handsomely is quant finance, but those people are actually compensated not for proving theorems but rather for statistical modeling and programming skills. And even that industry is so technologically driven these days that pure research mathematicians no longer hold a clear edge over strong programmers with undergrad level probability and statistics at their fingertips.
The $$$ they're pouring isn't just for marketing. Think of these papers/results more as "useful side effects" from large-scale RL rollouts and post-training. Every token being generated contributes to post-training in some way.
There isn't a hard boundary between "training" or "inference", modern post-training is arguably inference-bound :)
Ah this is an enlightening point. 8 hadn't thought about it this way, but you're right.
-- Me
We are at the point where the way in which humans do math and science changes significantly, and I have no good idea at all in what state is it going to settle down. But you are one of (many, I suppose) people exploring the new wilderness, so I wish you best.
Unless they decide that trying p!=np is worth any money.
We'd be nowhere without Laplace and Fourier transforms, Maxwell's equations, elliptic curve cryptography, and many more.
Most math doesn't, but often these techniques are invented first and the applications come later.
And the criticism of the current round of proofs is that while they may be true - likely for some, questionable for others - they're not adding new techniques or insights.
There's a great paper from Abraham Flexner on this topic:
https://worrydream.com/refs/Flexner_1939_-_The_Usefulness_of...
It argues exactly that we should be allowed to pursue the seemingly "useless" knowledge.
Previously discussed on HN:
https://hn.algolia.com/?q=usefulness+of+useless+knowledge
There are architectural advancements yes, but lots of progress from LLMs really come from (1) better pre-training [generally through more cleaned data, and ofc more data], and (2) lots and lots of post-training. It's how we get more and more intelligent models for the same param sizes.
The 'marketing' is just a useful side effect they get from their RL rollouts on maths and LEAN.
However, there's some of this that's a proxy - the compute to solve these problems was very low (they claim a few hours of thinking time on a regular subscription). The large cost would have been the training and if training the models to be better at these things makes them smarter for useful tasks that's beneficial. I believe there was work done earlier on around showing that training the models on code made them better at broader reasoning tasks (not just writing the code itself).
Another side is that if one goal is to improve the models themselves, their ability to work on mathsy problems must be high. That has very direct business value, and ideological value depending on what you think the motivations of the people running the companies are.
To take this to the next step, what happened after deep mind pretty much solved Go is that they started looking for the next set of things that hadn't been done yet. It does strike me as very likely that this will follow that same path.
Go was a specialized application. All the math results come as a side effect of reading the whole internet, and it will keep reading the whole internet. It will keep practicing thinking questions. Actually, math might be one of the best ways to keep them contemplating and measure their contemplation abilities, so math will always stay in the loop.
Also, math might not be useful just for humanity, but also for AI, so the system might actively benefit from new math results itself. (Not sure if any of the recent proofs qualify, but future work might.)
I definitely think that this is marketing, just "with good side effects". My doubt is when they will be able to move to "marketing with better side effects", that is, research with more concrete outcomes (health, materials etc.).
Problem is, that type of research is much harder. Some doubt that progress in such areas will be quick (https://www.noahpinion.blog/p/wheres-the-intelligence-explos...).
While top AI labs no longer focus on chess, the community build way better chess engines.
Stockfish is probably stronger, than everything the top labs build.
Wouldn't we expect the same thing for math? That slowly the broader math community would engineer a harness/program... That will surpass the current labs, and be a community ran project
In the case of chess this seems fine, there isn't much value to society in creating an AI capable of beating top humans with a 4 pawn handicap rather than a 2 pawn one, but for maths where there are actual applications it is more complicated.
It is also fascinating, because I don't think there is any solution within our existing system, at least not any I know of. Theorem ownership is not a good solution (and neither are patents in general). Probably the most achievable (or rather the least unachievable) solution is a kind of communist utopia, where people can dedicate their time to a pursuit of any endeavor they see fit, as resources for a decent life are abundant and excessive power capture impossible. (The other option, somewhat dystopian, and which would not require humanity to change too much in its current mode of conduct, would be a totalitarian or caste-like capture of society by the scientific community.)
Incidentally, if AI proves as powerful as some expect it to become, it could bring about another solution of that issue by making all human science and mathematics obsolete, pushing its true market value to zero.
(With apologies for rambling.)
I wonder if it makes conceptualization simpler for models too, given that they're trained already on human-speak. And I'm also curious as to whether humans currently have an innate advantage into simplifying and contextualizing proofs, or will the machines get good at that as well?
I was bitter about that back in the day as I hoped for more answers, more matches, more "truth" about chess being shown. Soon after that community project Leela Chess Zero was started and not only surpassed original AlphaZero but added few hundred ELO points over it. Then the combination of NN and classical engines happened with NNUE and current Stockfish is again a few hundred ELO points stronger.
Today we pretty much know the truth in chess for all practical purposes. Human analysts/preparation experts focus on finding interesting path and opponent profiling (what is the most unpleasant for the opponent to face). They don't look for truth anymore. The game is doing great, it's more popular than it ever was.
https://en.chessbase.com/post/the-full-alphazero-paper-is-pu...
The result was that Stockfish lost many games by walking into known bad lines and lost way more games than it otherwise would.
> We also played a match that started from the set of opening positions used in the 2016 TCEC world championship, along with a series of additional matches against the most recent development version of Stockfish, and a variant of Stockfish that uses a strong opening book. In all matches, AlphaZero won.
https://deepmind.google/blog/alphazero-shedding-new-light-on...
I am not claiming AlphaZero wasn't stronger. It wasn't as strong as the PR piece suggested though and we have never seen the games being published. In chess this is extraordinary because basically all games in chess are publicly available - both human and computer games. Claiming "we have created a strong engine that has beaten Stockfish with opening book" while not showing those games (or details about opening book used) is akin to "we solved this math conjecture" without showing any kind of proof or argument.
Publishing a few 1000 of games costs nothing. Tens/hundreds of thousands of games are published every day.
But, I do think you are right that there will be some level of moving on. The spotlight is currently on maths and that won't last. It will move to some other area where there is more impact to be had. So while they might shift gears and put less focus on math, it will always be there as part of the portfolio.
Being able to present useful novel ideas would likely generate a lot of press, for a while. I don’t know how this would look since I’m useless at math, but Im sure there are plenty of unknown problems with massive implications, that once formulated can be solved.
One is that AI will continue hallucinating in a manner that is not easy to verify, second is that AI will not be enhanced to produced more simplified amd robust outputs, and third that a human will be required to do that. What humans in the loop are doing now is verify the process, propose shortcuts and add legitimacy, through the verification process, if that ends up being succesful its highly likely a lot less mathematicians will be required in the future.
The conclusion that this is not productive focuses on the mathematicians, but it is very productive in terms of hundreds of proofs being produced that had previously consumed uncountable hours of the brightest minds. Unless it ends up being the greatest hallucination ever ofcourse
This means when writing documentation, tutorials or commit messages, their output is often a garbled jumble. Assuming shared context, using invented terminology without explaining, leaking conversational states due to improper epistemic boundaries and failing to model the reader. This all usually leads to their freely generated explanations being terrible. Getting good explanations requires chaining questions that force them to line things up properly, which is not easy the less you know. These failures as something LLMs naturally struggle with make sense, given the nature of attention and RL with weak signals from human data.
Math is not merely a collection of proofs, it's a way of understanding. A proof presented in a manner that cannot be incorporated remains useless. It does not make it's way to physics like Riemannian geometry and matrix math did. This is no less true when done by humans too.
Your hallucination conclusion, checking if a proof is one, is exactly the counterproductive cost.
Most of us cannot verify that the claims in the OpenAI lore dump are in fact all correct. It will take tons of work from experts to do this. It took subject expert mathematicians to identify the discrepancy and disconnect in the Navier Stokes proofs, for example. LLMs will struggle to make use of their own proofs or turn them into knowledge that accumulates over time.
The act of proving is often more valuable than the proof itself. Human constraints and limitations force us to invent tools and abstractions that a 100,000 x 1M context swarm can bypass. The tradeoff from that AI swarm advantage is work that doesn't usually lend itself to being built upon. It's like doing all the side quests and reading all the books of an RPG versus min maxing a straight path with a guide. We might try to identify new abstractions, but the fact that we don't get access to CoT and that much of it will be illegible means mining LLM traces for what human mathematicians produce naturally will be a tedious chore.
This is a feature, and a huge step forward.
If you expect AI to do serious work, you can’t have it guessing what you “really meant”. Every sufficiently advanced task depends on very subtle details in the problem statement, and the correct default behavior for advanced AIs is to solve the task exactly as stated, unless a system prompt or other constraint tells it to do otherwise.
It is an old saw at this point, but what an LLM does still cannot be divided into hallucination and non-hallucination. This is literally an anthropomorphism trap.
Layers and layers of application-specific verification can reduce the risks inherent to LLMs, to a really remarkable degree, but nothing about what these tools are suggests that this problem will go away; it will just bubble up again somewhere else.
To an arbitrary degree.
Just like all of science. Reduce the error to the desired margin.
For all that I saw over the last few hundred hours with AI on software engineering, hallucinations are no longer a problem at all.
Not once have I seen a task fail due to what would have been a "hallucination". If they still occur, they can apparently be detected and corrected automatically, or are subtle enough to escape notice with presumably no significant impact on the results.
Why would this not also be the case for mathematics?
"Test suite passed" when it actually errored? Obvious hallucination, unless it ran a command that returned the wrong error code.
But is running a malformed command that does not achieve the expected effect itself a hallucination?
I'd say a hallucination (very misleading word) or confabulation or "making shit up" happens when an LLM uses factual / evidential language purely based on local statistical expectations of the text, instead of it drawing from actual evidence in its context pointing to it.
This is murkier in the case of general knowledge questions, like when and where was some famous person born. It may then be a spectrum from fully making something up based on how the name sounds, all the way to confidently retrieving it from its weights correctly. In between, we can get hallucinations. But newer models are taught to use Web Search when unsure, and it works pretty well, though not perfectly. I don't see any fundamental limit here. It's just not perfect. Trying to solve "the hallucination problem" is basically like saying "our dog vs. cat classifier is pretty good already with its 99% accuracy, now all we need to do is the tiny little task of eliminating the 1% error, and we will be golden". Like, no shit, there is some error yes. People are working to reduce it. It will never be absolutely 100%. It's not an insight to say we should remove hallucinations.
An old saw unless something that's widely accepted, but sadly it seems that many people don't recognize this, even many people working in the field.
Hold up, that's an even bigger assumption in the opposite direction, and I don't see anything to support it.
At least in terms LLMs getting all the "AI" hype these days, there is no structural/mathematical reason to believe they won't continue to have the same problem they've always had of generating plausible text over rational text, and I don't think anybody even has a clear idea how it could eventually be accomplished.
I've seen "then the magic singularity occurs and somehow it solves the problem for itself", but I would classify that more as mysticism than engineering.
Two years ago, hallucinating that the code worked or that a task was accomplished was a common occurrence.
We have seen that now agent swarms across thousands of agents can coordinate to achieve a result.
Clearly hallucinations are no longer the problem they once were, since now we can get working results for long horizon tasks that require massive compute.
Consequently it would seem unwise to assume that current limitations will remain as they are and prevent LLMs from coming up with solutions that they can explain to humans.
It happens in more subtle ways, but it still happens often enough for me to notice. For example I have had hallucinated checksums show up in lock files as recently as yesterday using a SOTA model.
This is not surprising, since the whole basis of LLM training is to produce output that humans will accept _as a proxy for actual training goals_. In a sense, the training process of an LLM “wants” to produce output that is statistically plausible much more than it “wants” to produce correct output. It’s always going to be a struggle to drive that system towards other goals (and we see this bourne out in practice by the amount of effort that is required to be spent on RL).
I think there will be some threshold of correctness (something like 99.999% of the time) that if the model surpasses it, I can stop needing to check it, but I think we’re still at 99% or something which sounds good, but when you are producing a ton of output you hit that 1% frequently.
> Consequently it would seem unwise to assume that current limitations will remain as they are and prevent LLMs from coming up with solutions that they can explain to humans.
I 100% agree with this. In fact explaining things to humans is something LLMs are particularly well suited for.
This is just a fact. I'm sorry if it messes with your narrative.
https://arxiv.org/abs/2401.11817
And yeah I get hallucinations all the time still. Maybe it's because I'm working on harder/more niche problems (like a compiler with an unusual type system), but it happens quite a lot. I don't record all of them.
Although the most common one you can find is them misattributing the source of changes from themselves and also other agents (Fable, Opus 5.5, deepseek, whatever). They'll say "your changes" or "you changed" or "your ruling." I didn't decide anything and it's in their own chat log, and yet...
You’re making the following assumptions:
1. the exercise of struggling to find proofs was not productive, but this is precisely how new techniques in math were produced. Brute forcing solutions doesn’t lend itself to the creation of much new mathematics (except maybe the exercise of developing verifiable proofs)
2. the point of doing mathematics is to be “productive” in the first place. This is silly. Many people get into mathematics because of the beauty of understanding, for example.
Are they independently wealthy? Or do they have a deal with their local supermarket that they can take food for free?
If your point is that capitalism fucks up the incentive structure and makes it all about maximizing productivity then I wholeheartedly agree with you.
There is literally not a single shred of evidence to indicate either of your supposed eventualities. The core technology of an LLM is sampling from a distribution so there is literally no way to make it deterministically robust (only probabilistically).
You might have a point if the goal was to have LLMs that spit out a correct proof without chain of thought or tool use. LLMs + agent harnesses are more than capable of self verification and course correction.
What has been demonstrated is a process that outputs lean proofs based on those probabilities. This happened after decades markov chain producing garbled texts and very shortly after gpt2 producing stories about unicorns.
An LLM mostly deterministically (except parallel processing nondeterminism that can be mitigated) produces a probability distribution that can be sampled deterministically: just take the highest probability token or use beam search.
[1] https://dev.to/natcher/researchers-develop-method-to-train-l...
Maybe if we start with giving simple AI generated analysis of those clumsy humans with their measly 2700 elo moves
With Waymo and Tesla increasingly doing what they said they would do, and a small number of early adopters happily paying money for their services, that do work.
So what's the critique? That the timelines are not correct? Sure. And how about the timeline of the people who said "research level math, never in my lifetime" and the people inside the ai companies who are apparently increasingly spooked by how quick the progress is? How about the various levels of code/programming jobs that AI was supposedly never going to be able to do, but, in reality, now just does?
We are engaging in some very one-sided discrediting, and I am not sure, why.
https://www.wsj.com/articles/self-driving-cars-dont-do-snow-...
More telling: Waymo just rolled out in Denver (1 month ago or so), apparently fairly confident they got this handled given the upcoming winter.
Progress on the obvious stuff keeps happening (which kind of brings me my to the first comment here).
But also: It does not have to do "snow" to be useful! A lot of cars/people don't drive when it snows heavily, and that's something we have always been okay with (at a societal level, YMMV of course). If was only useful 95% of the year that's still great. A lot of technology works like that.
Maybe the really hard thing isn't abstract cognitive capability, but perception+movement.
You just don’t notice that because evolution has given you 99% of what is needed to drive a car before you were even born.
Where does this "will probably not hold in the very near future" come from? People correctly warn about extrapolating current things onto the future, but then just throw some vague "probabilities" without providing any argument why their "probably" is somehow more grounded than others.
Here's a longer article which goes into a bit more of the details: https://terrytao.wordpress.com/2026/10/05/the-future-of-math...
In what world is OpenAI not “genuinely contributing to mathematics”?
I’m getting whiplash from the speed at which people are suddenly accusing them, and AI in general, of not doing enough.
But you are right, this is not OpenAI's "fault". The problem is - as others have said recently - that many people in mathematics want recognition for solving open questions more than they want the answers to the open questions. Everything about the economics and social environment of Mathematics will have to change.
I think that this is exactly the same split we see in software: there are those who mainly enjoy the craft aspect of building software, and are uninterested in the product or business they are supporting. Others are primarily interested in the production of useful software or building a platform or company.
I've always been in both camps myself. When it became obvious that AI was going to destroy the craft aspect - at least two years before it actually could do so - I became very discouraged, even depressed. But once it was actually good at building software, I became very excited about all the stuff I could now build. Sadly, I think a lot of people in our field have never had something they really wanted to build.
That community won't exist as many people simply won't even enter the field because it's reprehensible and contemptible, not to mention boring, just to read machine-generated proofs and verify them.
Collaborating on, or at least working on unsolved problems is what motivates most people.
AI is like a cheat code in a video game. You get to the end faster but fewer people want to play if the cheat code is always on. You can't turn it off either because the very challenge is to do something unique.
So effectively, stuff got proved, but people don't really understand how, so it's mostly fucking useless and done for OpenAI's marketing team, while also pissing off the maths world at large.
So you'd prefer if they kept their work secret? Or you don't want them working on these problems at all? Or they should be required to do the work the way you want them to?
I'm not clear what you see as a better option than dumping.
I tend to agree with this, but what is the alternative? Should OpenAI and Anthropic employ hundreds of mathematicians to do this work? Should they just not solve math problems within their reach?
It's unclear what more could be expected than releasing the presumably already verified results and write-ups for each problem. Should they run a mathematics school too?
Then the comment goes on to argue AI labs were not interested in actually advancing mathematics, and that investments into AI were manically excessive.
IMO none of this follows and demand is there to justify the investments.
The comment then goes further to argue that AI labs were putting too much effort into pretending there was exponential progress rather than actually making progress.
The factual basis for this claim seems to be that OpenAI released math results and write-ups, and it's not even clear what more they could do on that topic.
That's a very negative opinion.
So yes, that is exactly what they should do. Alternatively, if they are too lazy or incompetent to put in the effort themselves, do what AGMAI proposed and fund a third party to help out.
This is because have fundamentally different goals: being able to use results in calculation versus having a deeper understanding of the subject matter.
Peer reviewed and published insights are proven invalid all the time. This is the nature of research and how we learn.
We don't know the current system can work well enough at this scale, because that's un-knowable. We know it can find some of the problems. We don't know it can find all of them.
We do know it takes more effort - that's knowable. Increased data takes increased processing.
Whether the community has the required effort available, seems unlikely, considering the expertise required to be able to assess these things hasn't changed. Only the ability to generate them has increased.
I feel your concern stems from the risk that there is additional noise everyone needs to cut through.
In reality this isn't any tom, dick or Harry giving you their vibe code output. They have spent millions of dollars on this output, so there is a filter. The biggest filter of them all, funding.
Furthermore, LLM's have given us another gift semantic search, we can easily check your work against theirs, this is valuable insight so instead of researchers wasting decades and fortunes pursuing an avenue that shows no value (this includes methods), they can purse new avenues they know what to avoid, in the same breath they know what to work towards.
With convoluted and inelegant proofs, AI may fail to uncover those systems and patterns. As a most concrete example, it may fail to recognize some problems as isomorphic to other problems. Brute force solutions are a depth-first search.
To improve human mathematical understanding, AI is probably best used as a “copilot” (lol) rather than a black box oracle, like these AI companies appear to be doing.
If you're after new methods. Then new methods is the goal, the answer to the question is not the goal then. The animated response indicates the answer wasn't just a byproduct.
There is still something to glean from the answer. You have a further constraint. Otherwise whatever "new method" proposed may as well be hallucination, potentially taking you in the wrong direction away from the answer.
This line of thought is not unique, stonemasons made obsolete by uniform brickword suddenly were "worried about the art and preserving traditions".
What work do you think mathematicians do normally?
Like they sit whole day and have ideas? And where are the ideas?
The way I see it, _some_ mathematicians enjoy solving puzzles, and now AI is better at solving puzzles.
This does not affect people building new theories.
Also, it's quite prestigious to write a _book_ on some topic. And guess what writing a book entails? Refining and expanding. What you call grunt work.
Actually, the problem is, somehow the skill of building new theories in math is directly tied to slaving hard over a problem. It's the very experience of slaving away that actually somehow causes ideas to form. Pretty much all mathematicians understand this. Yes, senior mathematicians now can form some new theories, but what about junior ones who will have very little experience in working hard on a problem by hand?
Of course, they could work on the problem by hand anyway, but they won't because no one will pay them when a machine can do it.
Let's start with the fact that mathematicians aren't paid for results. Institution which gives them salary fundamentally doesn't give a fuck about theorems. They might care about having top-grade mathematicians for prestige, or because they believe that countries which are "good at math" also do better in science, engineering, etc.
> somehow the skill of building new theories in math is directly tied to slaving hard over a problem
We don't know if that's the only way. Perhaps collaboration with AI is just as good. Why reject it before it has been tried?
> what about junior ones
Junior guy with brilliant new ideas might benefit the most from AI as it can compensate for lacking technical chops and breadth of knowledge.
And it's not like this is something where we're loaned some top math genius for a limited amount of time and we have to make the most of it. Rather, this is a new high water mark. The accessibility of the results is no longer scarce. The scarcity has shifted, and that's where the focus of the math ecosystem should shift as well. And it doesn't help for a frontier community to saturate and take over messaging pipelines that were typically managed by the math ecosystem. It's not about "stay in your lane" but rather "we need coherence and be careful not to break the system."
Just two cents from someone who could screw up basic cashier math on any given day.
Many of the best startup ideas by the best product and engineering minds failed to gain attention and funding. Same with much of the best music - relegated to hard drives with derivative ideas only resurfaced decades later
I would expect much of the recent math dump will be leveraged by other LLM-driven research teams rather than read in depth by a human
The more information the better.
The entire purpose of published work is to remove noise (and perhaps incentivize work through attributing credit).
This information is now out there. You can choose to ignore it if you wish. You may just find yourself a century behind in research.
And on that point most of this research has been looked at by their mathematics panel and comes with lean certificates, it's not exactly noise.
This to me is more the old guard not willing to let go or change their ways.
I think a problem is that math seems like a deeply toxic, ego driven domain.
I think he argued that e.g because the navier stokes millennium problem ist considered solved now, you won't get any recognition for being the first human to solve.(How would you even proof you solved it yourself and not just regurgitated the ai proof?)
And since recognition is the main objective, noone would spend time on dissecting the proof, and perhaps finding some unique approach to solving the problem, that could be transferred to other open issues.
And therefore the problem is now "poisoned". Since it's assumed to be solved noone will research it, and the potential revelations won't be found
During your write up, I'd imagine you would check it's not already out there too. And once complete it's cheap and easy to run it through an LLM and ask is this covered by anything else out there. If its novel and not published it doesn't matter what others say.
Research is already messy as it stands. Something new can already be dismissed by incumbents as "not novel enough" especially in niche fields where they're likely to be the ones conducting peer review.
And if it's hard to understand (which seems to be the most common reaction) it's not exactly devoid of noise either
There was an opportunity for people to work with the AI to produce a proof, now it almost feels they're working against it.
Given sustained exponential growth is mathematically impossible to maintain with finite resources, it's funny to me they're using advanced mathematics to try and achieve this.
This is a transitive period. In a few years, verification and exchange between model instances will happen faster than humans can follow. Human input will be an ethical question, and not a productivity one, because it will be the bottleneck in any science.
If the focus in placed on maximizing some easily measurable output on a narrow perspective, situation is unlikely going to match a sweet spot of holistic equilibrium which is maximizing harmony and happiness through humanity as a whole.
Hm, kinda reminds me of my college days. "Proof trivial, left as home work." was a sentence my Profs loved to say.
For what?
If math is just about having a community of other mathematicians to hang out with, it still isn't a career. Nobody is paying money you need in order to to eat, just to hang out in a community.
Just like a software engineer, "Well AI can write all my projects now, but I have my local Rust Users Group to hang out with". Nobody is paying me to hang out and hand code Rust.
The amount of Confluence pages of "research" that is just a dump of LLM output someone passed to me to review is staggering
I hate this approach, it's unbelievably selfish
OAI obviously have a fiscal incentive here, but to presume a year from now we won't see improvements and more succinct work on the results coming from models?
If a random person is given a 60 page proof to digest and not the author, those hidden insights that _aren't_ in the paper might be completely inaccessible. Maybe the AI will "just" be able to provide the insights. Maybe. But pedagogy is tricky work, and despite these AIs being able to do all this fancy math we can't get them to write good cover letters yet, so....
Ultimately we might be left with just a bunch of intellectually unsatisfying proofs. This means way less drive to simplify the proofs or rework them.
End result: we generate a layer of "less efficient" mathematics, that won't get built upon. We will not actually have any shoulders upon which to stand.
The fact that they did not do so can only mean that either (1) their agents currently lack the capability to do it, or (2) OpenAI are completely indifferent and do not care in the slightest if the proofs are understood or not.
Come now, this is kind of unreasonable. When you're working on a new technology, you first get the ugly, inconvenient-to-use prototypes functioning with the core new thing you need; then you work on packaging it up into a format useable in production. I'm sure the very first digital camera sensors weren't very useful for photographers either; but it isn't really even possible to build the rest of the technology required to turn raw output of a digital sensor into something a professional photographer can use until you have the raw output itself.
The research is still on going on the raw output; getting things to the next stage, where the results are widely useable by professional mathematicians (and then on to engineers and scientists to whom the results would be practically useful), is a whole new research area.
Like, "Orr... maybe they care a lot, but haven't gotten to that part yet?"
The concern is being raised without evidence, because the evidence points to the gap simply being frontier models have just started to be able to get a raw proof out. Why, given existing progress, should we expect them to be unable to distill insights from those proofs?
Certainly this even more likely doesn't matter at all for applications: if I can send a radio signal further because my AIs design it a certain way, that's an unambiguous result. Which is really the next step here: turn a proof into a "mechanical" application.
OP didn’t suggest that.
The bar has been raised. Everyone has to meet it now. An inelegant solution squatted onto the internet doesn’t count as discovery per se, even if it’s impressive.
It’s fine that OpenAI posted their findings. It’s not fair to claim these problems have been solved. Not until someone can understand and verify the proof and then communicate the core, novel methodological element to someone else.
So yeah, probably we should stop saying "X has been solved", and instead say, "A Lean proof for X (or !X) has been generated". That doesn't change the fact that incentives are currently on finding the proof, and once the proof is generated by an AI, there's not currently a good mechanism / incentive structure to move that into the mathematical community. AI is here, so we need to find a new mechanism.
[1] https://arxiv.org/abs/2610.08144
Possibly yes such a representation is possible. But it doesn’t mean it’s certain there is a "best compressed representation". And even less one that encompass everything important and that is understandable by any human brain, even the most exceptionally brilliant ones sponsored by a whole society to reach their best possible achievable performance on that goal through full dedication on that sole task.
Your phrasing is illuminating that perhaps they aren't engaged in the creation, understanding, or integration of these proofs by humanity; they just have them. For them, this is a slidedeck they can pass to investors, creditors, the marketing department. Something they can add to the employee onboarding pamphlet.
What should they do? Hyperbolic maybe, but perhaps engage with humanity.
This isn't only bad for Math — it's bad for English too.
'Proof' is going to become the 2026 Most Misapplied Word of the Year.
I feel the same way about academia, the papers, the citations, the ego, the narcissism and the taxpayer codependency that got cut off and turns out wasn’t necessary at all thanks to a private sector entity running laps around them
I don’t feel that academics need to pursue the discipline and distributed brain-wracking that has sometimes resulted in the solved math problems, just because more times they find other nooks and crannies to explore along the way. I think the blueprint is enough. Standing on the shoulders of giants is good enough.
and if the concern is that they can’t figure out what to do with a proof, next year’s AI will
> genuinely contributing to mathematics
What's the difference between the two? Proofs are no longer the goalpost?
Those few thousand mathematicians are now getting a taste of their own medicine. After all, it was people with extraordinary mathematical talent who developed machine learning and large language models, leaving hundreds of millions of people who earn their living through speaking, writing, or teaching worried about their future job prospects.
Still, I believe almost everyone will be fine. Perhaps AI will also prove good at coming up with new conjectures, and some mathematicians may shift towards applied mathematics or other sciences.
Mathematics will revert to it's main practice, which is to study.
There are lots of weird panic reactions by some prominent problem solvers. See for example the ridiculous cease and desist like statement of AHM shared at Tao blog.
Put this Math 1-2.0 with that AHM statement together and you'll realize that this is a power struggle and that you see only one side of it.
Mathematicians have very diverse opinions about this. I, for example, am for as much as possible automatic harvesting of all these "low hanging" fruits. Should be disclosed as soon as possible, free of any bottleneck, and citable. The mathematics community may do whatever its various members desire to do with these results. Let them decide individually what to do with them. This AI tool is here to stay.
MVP 1.0 is always produced quickly to solve a particular business problem or conform to a spec, even if the solution has terrible code.
2.0 is when programmers refactor the code and internal APIs to make things nice and easily explainable.
>I think of mathematics as having a large component of psychology, because of its strong dependence on human minds. Dehumanized mathematics would be more like computer code, which is very different. Mathematical ideas, even simple ideas, are often hard to transplant from mind to mind....Translation in the direction conceptual -> concrete and symbolic is much easier than translation in the reverse direction, and symbolic forms often replaces the conceptual forms of understanding....
https://mathoverflow.net/questions/43690/whats-a-mathematici...
Also, this reads like you didn’t like math classes. That sucks, but it’s no basis for societal organization
The valuable part that Tao is feeling the absence of is the insight, and you can only get that from talking to the swarm of agents that developed the original proof with all of their context.
So to me it feels like we don't need Math 2.0, but Authorship 2.0. I want to "meet" the context that generated these proofs. I mean luckily these were not generated by faceless systems like a SAT solver, you can actually talk to it, but I'm not sure if we can step beyond our pride and grant the true authors of these proofs that recognition.
The idea would be that you should not fiddle with the minds who try to independently evaluate your works, so that is not a feasible approach to truth seeking.
While in organic chemistry, this way, valid progress was made, you can always avoid a perpetuum mobile inventor and get conned.
This is pretty much what a person that proivded patronage to a matematician used to be. API prompters are people who provide patronage for AI mathematicians.
You don't talk with them about the discoveries. About discoveries you should talk with who actually made them. Namely the LLMs.
Another analogy might by that you shouldn't expect to have interesting discussion about the essence of art with art producer.
Just thank them for the inference they covered and interact with the results instead.
"When the architect completes a fine building, he removes the scaffolding." - Carl Friedrich Gauss
Im not sure how that will work, but im convinced the current paradigm of just pushing agents into codebases for not much reason other than you can is going to make building software incredibly boring and push creative people away from the field and stagnate progress.
My prediction is software gets boring and building hardware projects will be the new frontier for creative engineers looking to push computing further. Which is probably a good thing.
I believe in the future, we're going to see a similar shift in "programmer" - instead of a human programming the computer, you'll give the ai an idea and it will spit out a program.
And just like how automating the act of computation revolutionized what we could compute, automating the act of writing code will change the act of programming - hopefully, as you described, allowing us to do things that simply were not practical in the past.
People wrote many books about software engineering, all from valuable experience from buildng expensive software systems. But in the age of AI, is there still anything learnable from generated code?
Personally I always ask AI to summarize its findings and lessons in a .md file. And I always learn something from it.
But could AI utilize some new patterns and paradigms I wasn't aware of? Very likely. Because we only learn from our personal grave mistakes, a summary from others gets neglected and forgotten
Sort of. An elegant proof is useful beyond what it shows. It hints at new mathematics, and can prompt discovery in applied fields. I don’t think I’ve heard of elegant code leading to discovery on its own.
Every software design pattern came from elegant code. People wrote code, summarize code, learnt from code, and taught code. That's discovery
Usually it's the opposite. "That's in prod? And it works? It shouldn't work and I thought it was doing something else. Why does it work?"
I’ve not heard of it either, but code is an abstraction of math, so I don’t see why this couldn’t theoretically happen.
Anecdotally, I’ve started spending time advancing my math skills beyond the early college level I stopped at and I’ve frequently found I already know concepts of more advanced math - I just didn’t know what they were called or how to apply them to an equation on paper, but I’ve been using them for years and intrinsically grasped the underlying academics.
The S fell short in actual reality for the most part, as it was merely a hiring requirement. A hiring requirement that didn't even make sense, because the skillset of academic CS only marginally overlaps with the skillset one wants to hire for.
Material engineering at least for the most part has actual real-world applications where one can push humanity further. CS (as practiced, not necessarily the idea of real CS but the CS we got due to it being used as a hiring filter) for the most part is just self-referential spinning with mostly unclear results.
There is real impressive work being done in that field, of course, but I'd argue that the majority of it over the last decade or so at least was just performative nonsense.
Maybe by again allocating new resources to other fields, what hides under the label CS can become more pure actual CS again. I think that would also be a much less miserable experience for everyone involved.
The will to implement the stuff we learned, instead of letting dark patterns and churn for the sake of churn get the better of that :P
Do we want these capabilities to be useful? It seems making them accessible so that the existing community can use it as a tool is the right approach. Automated proof methods are already used like that.
AI, however powerful, is a tool, only as important as the amount it helps mathematicians. Creating mathematics without human understanding is as sound as mass producing copies of Michelangelo David.
“Mathematics, rightly viewed, possesses not only truth, but supreme beauty — a beauty cold and austere, like that of sculpture [...] yet sublimely pure, and capable of a stern perfection such as only the greatest art can show.” - Bertrand Russell, "From The Study of Mathematics" (1902)
“A mathematician, like a painter or a poet, is a maker of patterns. [...] The mathematician's patterns, like the painter's or the poet's, must be beautiful; the ideas, like the colours or the words, must fit together in a harmonious way. Beauty is the first test: there is no permanent place in the world for ugly mathematics.” - G. H. Hardy, "A Mathematician's Apology" (1940)
This really expresses the heartburn you see across all fields, not exclusive to careerism. I certainly have friends in decomp and fan translation spaces that have been demotivated by the current rash of efforts happening there.
The rush to be "first" has always been over-celebrated, but it would be nice to believe there's a way to get beyond that thinking.
This I don't understand, seems like an obvious thing to automate, especially for byte-matching?
I don't think it's motivating to solve a black box by having AI generate another black box if what you want is to understand how the thing worked.
No? Decomp sources are often full of comments that claim that there's no clear reason as to why something is done a certain way or straight up full of question marks. Byte matching often is a result of bruteforcing a solution rather than understanding the original idea behind the code.
0: https://en.wikipedia.org/wiki/The_Unreasonable_Effectiveness...
I feel like we’re circling back to that 2010s energy of “everyone can be an entrepreneur.” Now it’s “everyone can build software”
But if you're consdiering a community, this falls apart. The maths community has universities, has professors who are paid, has students which are getting their degrees for varying reasons, it has conferences, has publications, papers, projects etc etc, all of which will get some negative impact some AI.
You know that line "when a measurement becomes a target it ceases to be a useful measurement". This line holds up to different degrees for various measurements and targets. For maths it holds up very well. The goal is "contribute to make the world better by increasing humanity's understanding of maths" and the measurement, which by evaluating an individual on it we're turning into a target, is "how much does the individual publish new findings". Measurement turned target holds up great. It's almost impossible to publish new findings and not contribute to humanity's understanding of maths. But with AI these two are being decoupled. You can produce lots of new findings, but the community is saturated and they don't get assimilated into humanity's understanding. Why do individuals use AI then? Because you've made the target "how much does the individual publish new findings" and they have to compete or lose.
Its different when your own job is on the line.
When human manual arts were being automated away it was supposed to be not only acceptable but any complain and you were told you were a progress blocking luddite.
Now that mental labor is getting automated, the response to automation is very different.
For example, a software engineer is like a car mechanic or a coal miner. None of those are anything like a mathematician.
The person you're rallying against isn't me, it's an imaginary hypocritical person which exists in your mind. I do mental labour, I welcome AI developments hard, and I still think TT is 100% correct.
Not the mathematicians I know. They’d happily drop the academic admin stuff, but they’d absolutely keep doing mathematics in much the same way.
1) Intellectually challenging, to such a degree that those wishing to enter the field need to have a certain level of intellectual prowess to do so. This creates some levels of mystique, with a sprinkle of elitism and gatekeeping.
2) Driven (among other things) by prestige. And the more pure the math is, the more prestigious it is.
3) So complex that people can spend their entire working careers chasing a handful of problems. The amount of time researchers spend on very specific problems is mind-boggling, if we think about the results.
4) Intensely captivating for the people deep in the weeds.
And the deeper you get, the longer you study, the more you start to value things like "mathematical beauty", and may start to view math as a form of art.
Like many similar fields, you end up with this ivory tower where people can dedicate their whole lives to thinking deeply about extremely niche and theoretical problems.
Quite a lot of people are not happy they aren't elite anymore, and many have spent years to decades to arrive here.
Simply put you invest years of your life to establish a kind of distinction over others, and that goes away. That hurts.
But its not something surprising. Most of these competitive programming problems were actually English languages puzzles, because you couldn't dial up the mathematical difficulty anymore making it a fields medal problem. And in most cases in simple language weren't even that hard to begin with, and you could look up solutions to these problems in an hour of Google searching.
Since the Greeks we've had the idea that "Being and thinking are one," or that Being (in the sense of all of existence as such) has some essential unity with thought, and therefore can be thought, and expressed or submitted to the logos or reason. Being is in some sense fundamentally intelligible, and mathematics is the most developed, exacting, and articulate expression of Being.
Logic was understood in this older sense up to roughly the the mid to late 19th century. This is why a work like Hegel's Science of Logic begins not with syllogisms or propositions but with Being and Nothing. But this was forgotten after logic was mathematized by the English around the time of Russell, and its connection to ontology was gradually overshadowed by a focus on epistemology (still, it should be remembered, originally as a means of getting back to ontology).
There may be truth in art, but it's always haunted by its own historicity or contingency, which is to say untruth. Mathematics seems on the contrary the only really timeless, absolute thing we have. Part of what makes it captivating is stumbling on a construction or concept or proposition or theorem that simply must be, independent of us.
The AIs are certainly now more than automatic theorem provers, mechanically traversing some space of true propositions. They are able to push things forward and connect seemingly disparate domains to get to a proof, but to my mind it remains to be seen how well they will be able to form new concepts and definitions.
Imagine the controversy surrounding Cantor, for example, but put an AI in the place of Cantor. If an AI proposed something like the (infinite) hierarchy of infinity, would we have accepted it? What would the intuitionism debates have looked like? Would they even have taken place? And aside from that, has it actually been shown conclusively that an AI could propose such a thing?
There are lots of attempts right now to recover a humanism for mathematics, or restore man's pride of place with respect to it, but maybe we don't need to worry about that. Tao's attempts to preserve the mathematical community, while allowing for practices to change through the crisis may look like a kind of rearguard action, but seems reasonable to me and not really dependent on any kind of humanism. It's a way to avoid the question for now while things play out (and not conservative/reactionary like Scholze and others), which may be exactly what we need, because after all, perhaps we still don't understand why we do mathematics, what it's really for, and what our relation is to it. Whether it's enough to preserve funding is another issue.
Academics and white collars now get to experience what blue collar workers experienced in the past.
Same as what developers in USA experienced who were and are getting replaced by Indians.
Disclaimer: I'm kinda familiar with the ideas of automatically provable software systems but haven't done anything with them myself. If I'm missing an obvious fact(s) here please let me know :)
I get that the proof is big and complicated, but "We messed up a +/- sign" kinda sounds like announcing that the next version of the Linux kernel is done ("Version 7.0.0 is awesome!") followed by realizing that it doesn't compile ("Turns out someone used 1 equals where they should have used 2. Stay tuned for V 7.0.01!").
I've got to be missing something here :)
Like given the choice, the vast majority of people would prefer one quality game like Minecraft, LoL, or Fortnite, vs. thousands of one-shot generated games, and looking at user playtime this is exactly what we see. If anything AI is just going to entrench these pre-AI franchises even more.
Most people would be happier with La Marzocco coffe machines which costs thousands of dollars, but if you get an OK shot with a 100 USD DeLongi, then the choice is clear for majority of the population.
I think people sway to low effort endeavours that still have a reward at the end (even if it lesser reward than high effort).
Yes, what a terrible thing to advance the field significantly and release the results publicly for everyone. Truly despicable.
And nothing stops mathematicians from solving it in a way that does advance the field. Claude's existence doesn't change that.
The funding can have been for advancing the understanding, by using a more measurable proxy and reasonable target that closely aligned with advancing understanding.
Perhaps another phrasing might be "we are paying people to go through the process of solving these problems" rather than "we are paying people for solutions". I might set a random task for my kids while on a hike to find X things, not because I want to find ten different leaves but the process of doing it means exploring and investigating in a certain kind of way a certain kind of area. If some sets up a leaf selling stand, they have advanced the field of "finding leaves" and kids can now very very easily get ten different leaves.
Now leaves here are frivolous and not useful. That example works better looking at, say, homework - clearly I don't care about having a list of words spelled correctly and I don't need the answer to 5x7, nor do we need more book reports on The Great Gatsby. We're doing them to teach, it's very explicitly about the result.
Research level maths however is a bit of both. The answers to some of these things are genuinely useful. Having the answer may be better than not having it. But having a lot of people working on solving it has other useful and beneficial outcomes.
We have structured large scale systems of huge numbers of people and institutes around how this works, and what top mathematicians are telling us is that open problems (particularly at new researcher level) are a key part of this process and are hard to find. Academia changes incredibly fucking slowly, just glacially slowly. Some aspects (most?) are barely changed across hundreds of years. And across an incredibly short space of time (less time than one paper can take to go from fully finished to actually published) we have gone from "this machine can solve school level work" to "this machine is solving major research level maths problems". The existing system will not work, the machines will not get dumber or slower, and some of the impacts are things you cannot undo.
Now that you can solve without understanding this setup is broken. Either the sources of funding will finally have to learn the difference and the value of the latter, or mathematics research ends.
3-5 years is the period of a grant, and grants have to make research progress, or you don’t get the next grant.
Everyone should have a portfolio of prompts that are indecipherable by other humans but when fed to a frontier model, produces shocking one paragraph english version of a 10000 line lean proof
Obfuscated prompt grant contest
I love smart people like this; even when there's a threat, instead of just being in denial or boycotting out of anger, they figure out a new path for their community
If I understand Tao correctly, he's saying that's going to have to be the focus going forward. I just default to thinking the models are going to be much better than us at that, too.
I wonder if this is true. The code produced by these models are not really getting any more elegant over time. On contrary, the models seem to be getting worse, often proposing really baroque architectures. You can use RL to optimize for correctness, optimizing the vibe seems much more difficult.
Einstein did not typically use the formal peer review system to "settle on published work." Almost all of his major papers (including his landmark 1905 Annus Mirabilis papers) were published directly by journal editors without formal peer review.
When Physical Review sent his 1936 draft to a referee, Einstein was so outraged that he withdrew the paper and vowed never to publish with the journal again. He corrected his math only after an informal, friendly discussion with colleague Howard Percy Robertson—who, unbeknownst to Einstein, was the anonymous reviewer.
In the beginning, calculus as invented by Newton used complex ruler and compass constructions. Newton had a high cognitive capacity, so it was understandable to Newton. It took mathematicians coming later, including Leibniz, to turn this technique into a body of work that fits more easily into the average human mind. Newton and Leibniz independently developed calculus, but Leibniz's notation and formalism in particular provided a much more compact way of expressing and manipulating the ideas of calculus.
Open up Spivak at any page and find a formula; you should probably find that it contains 7±2 'things'. Like an integral, say: the integral sign, lower and upper limits, the function inside, the variable of integration. Then the theory and the rules for transforming these expressions was created so that working with calculus becomes mostly a set of rote operations.
Now a 1st-year student can do more calculus in a week than Newton even could have done in a year.
Now imagine aliens land on the Earth which have 10x our cognitive capacity, and we ask them about their mathematics. It would probably be incomprehensible to us because it would not have gone through a cognitive bottleneck sufficiently small to force it to fit into our minds. They might be totally happy with a mathematical expression containing 700 'things.'
We now find ourselves in this situation, except the alien is an AI we created.
I believe a cognitive bottleneck needs to be maintained so that maths can still remain human maths.
EDIT: Basically, mathematical elegance is finding a representation which allows irrelevant detail to dissapear.
It’s more likely that instead of spending a 100K/year direct grant on two PhD students, PIs will hire 1 and have the student spend 50K on AI.
https://arxiv.org/abs/math/9404236
>On a more everyday level, it is common for people first starting to grapple with computers to make large-scale computations of things they might have done on a smaller scale by hand. They might print out a table of the first 10,000 primes, only to find that their printout isn’t something they really wanted after all. They discover by this kind of experience that what they really want is usually not some collection of “answers”—what they want is understanding.
The difference in reactions by the maths community should probably be spit up into those who have read and understood On Proof and Progress and actually thought deeply about why they do mathematics, and those who haven't.
Tao only changed the heading of his blog to that quote in the last month, which is part of my point. Somehow despite this framework for viewing the field of mathematics having been beautifully described by one of our greatest leaders over 30 years ago, Terence Tao who has spent the last 5 years thinking about this has only just come around to it. Perhaps it has something psychologically to do with his greatest skills being those most under attack (though of course I'm not actually implying anything about Tao, he is obviously an honest and good-natured contributor to the community).
The people who need to read On Proof and Progress are the undergraduates and PhD students who have ended up down an academic pathway without looking up because they've always been good at proof and understanding mathematics, without ever pausing to ask why.
Frontier Mathematicians will work on proving theorems, and Pure Mathematicians will work on taking proven results and making them understood.
My own prediction is the opposite. I believe that a lot more humans will have a much deeper understanding of math, as you have llms to teach us.
However, this will only unlock if we get a hold of our own attention.
Sounds like Terence Tao would have said the same about Ramanujan who basically just "solved" problems without much explanation / reasoning / communication other than it just arrived from god.
In the case of Ramanujan, others took on the responsibility of socializing and community building knowing that he wouldn't do it himself. Why can't the same approach happen here?
There will be people who want to just "solve" math problems now that they have a new tool that lets them express themselves this way. Maybe the don't want to participate in the broader math community, etc. Why discourage them, or add friction / a barrier to them participating in their own way? Why not take on the burden of socializing, making sense of, and community building yourself?
There may be valid reasons here I'm missing, but to me this seems a bit like wanting others to approach a field in a particular way even though the field can support many ways.
Translating LLM proofs to human-ese is instead grunt-work. One that is bound to dissapear in a few years anyways.
What would a more “responsible” approach have been?
So, now he can join the club. The AI gave him the proof, what more does he want?
1. Some present a unified line that the whole point of their craft is the human experience, and that automation is the antithesis of that. Marathon runners don't care that a car can get there faster, poets don't care that Poem Bot 2000 can write poems too. I think this is smart if you can credibly take this position. The difficulty is mostly convincing the buy side, which requires being very outspoken about your views.
2. Some appear to be undecided, with one faction taking the pro-human stance and another rushing to accelerate things with AI. A good example of this is mathematics, and I really wonder where they end up in the long haul. They have a very good claim on #1, because mathematics is pretty close to an art form and is robustly insulated from the pressures of the marketplace. But they can also choose option #3, below.
3. Some crafts prioritize results above all else, practitioners either rushing to extract as much money as possible before it all collapses, or believing that they can out-prompt everyone else forever and that their prompting skills are indispensable to their employers in the long haul. That's software engineering. I think this is going to be interesting to watch.
There's a reason there are maybe a few hundred professional marathon runners in the world vs tens of thousands of professional mathematicians. Bucket 1 is basically an "amusement for the upper classes" type of deal. Any field that goes in that direction would have to shrink down massively.
It also devalues the field in my opinion from something really profound with actual impact in the world to a somewhat vain leisure activity. (basically going back to gentleman scientists) But I know other people would see it exactly the opposite way.
I don't buy this. I think there are other reasons why "professional marathon running" is a niche thing; it's probably just that it's not all that interesting to most people to practice in function of what it demands of your body, and not that interesting to patronize / watch.
Take woodworkers. There's probably more professional woodworkers than professional mathematicians. In terms of utility, everything a woodworker does, a machine can do more cheaply and more quickly. The main reason the craft survives is just that we attach intrinsic value to furniture made by humans the old-fashioned way. This is shared by craftsmen and those who buy.
And you could argue the same thing you did for marathon runners: custom furniture is just amusement for well-off people. Sure, but there's enough people with money to keep it afloat.
You say marathons are just for fun, but prior to the wheel it was the only way to get around. (Other than horses in some places)
So it’s not that running is immune to automation, it’s that we’re already post-automation and that only people doing it for fun are left.
I'd end up with the buckets
a) Fully human
b) Hybrid human-AI
c) Fully AI
Math? Everyone is talking about math, with math-centric discussions and solutions. I think the obsession with math is simply to distract ourselves from the fact that it's coming for every occupation.
None of these essays even remotely consider it. Denial is a helluva drug.
This is similar to general solvability of the quintic equations - Abel provided a proof first but only with the advent of Gallois theory we could basically understand it in full and decide for any quinitic if it's solvable by radicals or no.
Not hating, it's just funny.
It is important that we stay focused - this is theatre. Incredibly impressive, but this doesn't yet show evidence of helping society, which is the whole reason we were doing this in the first place.
The industrial revolution did that to battles and wars and it inspired Tolkein's lores to a considerable degree. He loathed what mechanisation had done.
I feel something similar is happening to Mathematics. I shudder to think what would come of other human pursuit this mechanisation targets next.
I hope you simply lost you thoughttrain and did not meant that.
(and I don't even think that is exaggerated very much)
I didn't follow through to the PhD, but I spent few years building up to understanding of Fluid Dynamics and Functional Analysis to come close to NS. It's intriguing that it's "solved", but what interests me is then "what do we learn from it" and what lies beyond in non-linearity.
In the 10 years I've spent away from academia, I still cherish what Math taught me best: looking at equivalences and I still feel the kick that I surely wouldn't want an Agent to do on my behalf. NS was never the point. And who can't see that I can only feel that they missed out.
I loathe what's happened with agscience, all farms should be plowed by hand with donkeys and plows.
It's fascinating how people like to push any statement to its limits, because it circles around to absurdity and they believe they've made a point. Moderation seems to be chasing you, but you clearly are faster!
If your mental model breaks at the limits, it maybe means there's some truth in what you say, but there's a nuance that's clearly missing.
And that's the case here: nobody argues when automation comes for many types of other tasks, so why is mathematics special? Or even art? We should either find that line or accept that maybe they aren't as special as we assumed.
Which is a simple measurable goal requiring little bureaucracy. The mythical "all you need is a pen and paper and a lifetime of dedication"
> "Math 2.0" will need to ... value mathematical progress more holistically
which is directionally the opposite
> community building ... AI can contribute positively
what is this belief based on? Any other communities can illustrate?
Which should improve collaboration, Research and Clarity.
I would really appreciate if we come up with protocols for using ai in STEM field's it might be award at first but we could regulate properly using this method.
That is, if things go in the current trajectory. I don't see any reason why anything would change though.
On the other hand, it puts a premium on resources. AI is not cheap for mathematicians. Folks are fancy universities in rich countries with forward thinking ministries of science will have an advantage over the rest.
What is clearly in immediate crisis is the traditional model of doctoral education. Most of the problems that were "given" to ordinary doctoral students are solvable (quickly) even by something like Claude pro. Mathematicians need to adopt training models more like what is done in experimental and laboratory sciences - collaborative and structured.
Where Tao is wrong is in regards to exposition. AI already writes better lecture notes, problems, and exercises for mid level undergrad math classes than do most of my colleagues. It's exposition is generally well structured and clear and it can adjust level on request quite well. It writes research better than most professional mathematicians too.
Everything else is secondary (or the last of our priorities) and would be better automated?
This is a hard pill to swallow
Thats one of the timeless human debates.
We are now living in the perfect combo of low morality and general human automation. So i expect the next few decades dominated by people who think (and have a "proof") that doing something without an expected economic gain is useless.
It is interesting that AI is not yet superhuman at exposition, or at least exposition that can be understood by humans. But you haven't updated enough if you don't think that will happen soon. I'd also expect for AI to become superhuman at opening up new directions of study and theory building.
> Many fewer seminars, workshops, collaborations, or other activities are being generated from these results compared to traditional breakthroughs...the mere knowledge that a solution exists "contaminates" efforts by both humans and AI to find alternate routes to the problem that reveal additional insight
This is absurd. The mere knowledge contaminates...give me a break Tao! Of course having a (possible) solution changes how we're thinking about the problem. If that's what you mean by contaminate, fine. But if you're a person who's excited, curious, interested in mathematical knowledge for its own sake these AI results are a treasure trove. New approaches to old problems, some old approaches that we couldn't make work before. Why not whole seminars to take one of these results and dissect them, prompting the AIs to figure out where else we can use them, improving and simplifying, etc.
Look, I get Tao's anxiety. The ground is shifting and it's hard to solve for the equilibrium. How in the world do you write a grant proposal today when the person who will read it reads the headlines and thinks "math" is solved. That's something that the mathematical community will need to figure out over time. And it's possible that there'll less money for math research overall. When the marginal cost goes down, the market equilibrium changes (but don't forget Jevons paradox!). So I get the anxiety. I just expected better from some of the top people of the field.
AI only take us as far as our imagination thinks to ask it. This can be exhilarating when new models drop every month and we can continually reach a new threshold, basically for free. But it is only a one time gain and ultimately short-sighted. Where I find continuous value is using LLMs to help my understanding, full stop.
I use LLMs all day long as a SWE and I have tried many approaches, but the most satisfying and consistent approach is to lean heavily into understanding a problem space and a solution space. Yes, it whips up architecture and code, but I spend most of my time peppering it with questions about the design and how it handles certain situations, what about this edge case and that security concern and this future product need. I have it write a report breaking down the feature and how it integrates with existing code and if the report is too confusing I have it simplify either the report or the code until it makes sense to me, sometimes scaling back the work to a more manageable state. I do all of this before I look at any of the code it writes.
The difference from this approach is that I am not suffering reading through 3000 lines of AI slop but I am reviewing a PR that I fully understand. I can eyeball it quickly for anything that doesn't fit my mental model and dig deeper or quickly revise it. Only after I am happy with the bones do I consider the meat and skin of the code.
What I find most concerning is how frontier AI companies all seem to have this Math 1.0 perspective that they only want to type "solve Riemann" into the chat box and have the magic to happen. It is the same problem Google ran into, where a simple, no thinking solution serves most of the people best and most profitably, so you fully ignore or remove everything else (boolean operators, exact phrase search, verticals, filters, infinite pages of results, "nothing found" if there isn't, etc.) But that choice leads to the situation Google is in now, scrambling to stay relevant. In a different world, Google would have continuously augmented their search capabilities and eventually built a smooth, guidable AI interface.
But no, we must only have an input box and a Go button.
Everything looks like a nail when you build hammers, sell hammers, have infinite hammers to play with however you like and your company mission is to build a hammer starship to explore the hammerverse, whether or not that is even possible.
And when the job is done, I run retrospectives on old coding-agent sessions to find areas of friction and confusion. I also journal with pen and paper, as it's supposed to bring cognitive benefits, to help me stay on top of things.
In case you missed, Mathematics isn't about numbers and equations.
There will be no gap in understanding. Now there is because the models are discovering things at the edge of what they can do and so suck at explaining it. There's nothing particularly special about a newly solved problem in terms of learning it.
If we accept AI can explain all of existing math nicely, why shouldn't it be able to explain new proofs?
So much in AI is dependent on which of these two outcomes occur.
The job of professional mathematician might be the first to be completely eliminated by LLMs, save for those who can make money from a patron. I am hoping they are able to figure something out to save their profession, as other professions could use it as a blueprint as AI comes for them next.
Strong disagree.
Do you work in a math adjacent field? I do and I find having a mathematician around invaluable.
It's like a non-software person writing software. Yes, using a LLM will get you to a solution that works. But just talking with a software engineer will make the quality of that solution enormously better.
I find the same with math - I can get something to work using an LLM, but if I speak to a mathematician they'll say some magic words to try and I put that in the LLM and it is "oh yes this is a much better solution".
This is very different work to generating proofs though. Its things like "I'm trying to get my confidence intervals to properly deal with census like sampling but at small sample sizes" (yes, I know stats not pure math but still..)
You may say "hey, before AI people payed for math salaries even though they didnt understand the math or the economic outcome". But the issue is there are now "mathematitians" trying to convince not to fund.
In this new reality, you will get "mathematitians" trying to convince that only AI maths matter. And on the other side someone speaking about "understanding", "taste", "community". And the people deciding to fund dont have the skills to differentiate. So they will fund the AI boosters with a higher probability.
Thats how the job of professional mathematitian dissapears. By being replaced by something that on the surface looks similar, but its just an ugly copy.
Edit: If you'd like a better medicine based one, look to radiology, where AI is an omnipresent tool but claims that radiologists are no longer needed, based on an ignorant view that a radiologist's job is "classify images according to what diseases they indicate" have only contributed to a crippling worldwide shortage of radiologists.
Up until now the prize in pure (as opposed to applied) mathematics was the _understanding_ and the machine can't do that for you. What does it mean if we get "super powered alien maths" but humans can't do it? It's like inter univeral teichmuller theory but imagine if Mochizuki was right and it came with a lean proof?
In my view, over the last century, math has turned into an intellectual analogue of extreme bodybuilding competitions. A navel gazing runaway optimization in making useless stuff just to demonstrate cleverness. That's fine, why not. But society has no obligation to fund that, just as it doesn't fund other extreme hobbies. Ideally if we ever get something like UBI, math can be still their hobby.
Without human understanding you also might literally have no words for the thing you would otherwise want to ask for.
I think it'll be wildy useful but I also suspect human competence will still matter.
So make mathematicians proudly wear their flag colors and sing the anthem while lecturing to cheering spectators. Might work.
Maybe AI will take over some roles of doctors, but that's independent of curing diseases. Think about diseases that have a cure - do people with those diseases not go see a doctor?
I think most academic disciplines would benefit from such a re-evaluation. AI is still a scourge on the earth, but I suppose if it spurs such changes that's a modicum of a silver lining.
There is a world where we get to the edge of AI capabilities, and we build on top of that. As humans have always done with every new technology.
There is another more pessimistic view where LLMs just replace every human capability, and our economic overlords dont need us for anything and we just eat the small pieces of bread that are left.
This comes down to the fact of:
is human existence/intelligence just the simbolic representations we make in our brain? Or are they just a tool?
I tend to think of Godels incompleteness theorem as a proof that on the limit LLMs are useless. The real question for me is at what point approaching this limit becomes an issue, and if it has any practical consequences.
We don't build on top of that. No need for us to. AI does. That's sort of the whole point of this endeavor is it not? Humans need not apply.
In my experience, every new model release allows me to go further, although every time i see every time the limitations, and i identify where i add value. And this value gets bigger every time.
But on the other hand, every model release reduces the amount of people that are able to value this "added value", because it requires more skills.
So we have the paradox that the added value i can bring on top gets bigger and bigger, but the perception of the economic value for the majority of the population gets smaller.
AHM Statement on OpenAI's October 6 Release of Mathematical Documents
https://news.ycombinator.com/item?id=50000421 / https://news.ycombinator.com/item?id=49999159
Lastly, deep down I don't really get what mathematicians are so upset about. All open problems, once solved, are not solved by 99.9999% of mathematicians, because it's solved by one or a handful of others, and the others just learn of the solution/proof. Mathematicians can now still organize conferences about these proofs, discuss them, digest them, think of new avenues of research, etc. They don't even have to invite OpenAI, in 6 months whatever model is available on chatgpt.com will be this smart anyway, and they can use it in the workshops for explanations, etc.
[1] I was going to write "I'm a bit disappointed by the response of the math community..", but then I remembered, whatever T. Tao writes is not the position of the math community, it's his position. Then I was going to write "I'm a bit disappointed by the response of T. Tao..", but then I remembered, I don't know Tao personally, so why am I disappointed?
[2] Steve Ballmer of Microsoft, I believe
Grigori Perelman warned about this when he refused the Millenium problem prize. He understood mathematics should be a journey, not a destination.
Arbitrary conclusion. This is the corporate take on "mathematics"
Mathematics were meant to further our understanding of nature and solve people's problem. Not to serve corporate delusional CEOs for their psychopathic purposes.
but can someone please try to set aside their knee-jerk reactions for a while to give a good reason:
× You don't know how to farm — That doesn't prevent you from having food or cooking good meals.× You don't know how to mine raw materials — That doesn't prevent you from using computers/phones made with aluminum, copper, glass etc.
× You don't know how to fell trees and shape lumber — That doesn't prevent you from sitting in that comfy chair.
× You don't know assembly language or how to write operating systems — That doesn't prevent you from using Windows or macOS or Linux.
—
EVERYDAY you use hundreds of things made from THOUSANDS of technologies you don't understand, because other people already MASTERED them.
so YOU can go on to go do GREATER things.
(but you CAN still go do farming, mining, logging, writing your own OS, if you ENJOY it — nothing's stopping you — you just won't be as good as the technology that has been specialized for that over centuries, and almost certainly you won't be bringing anything new to those fields, and it'll take time away from doing other things.)
—
Maybe we shouldn't be wasting time on "oshit how do we uninvent or slow down this new technology because it makes things easier than what we grew up on"
and focus more on "what other greater things can we move on to?"
There's a whole freakin universe out there and we haven't even stepped off our home planet yet.
The second thing is to just ask what is there left to do. What are these "greater things" that people can dedicate time to, when clearly even classically cerebral activities like mathematics can be automated away. The industrial revolution already wrecked physical production of goods and made artisan workers obsolete outside of extremely niche scenarios -- that's why we call things artisanal, after all -- but there was still mental work. But now, mental work is also experiencing the same thing, and it's not clear what one should do as a human anymore.
And some people seem outright gleeful about these developments, which can be seen even in this thread. What happens when humanity becomes obsolete? And what happens when the machines that cause this obsolescence are controlled by a tiny amount of people, who suddenly don't need the rest of us? I can only hope that this turns out well for us and that with the development of these machines, humanity will get better, but the omnipresent existential dread is giving me doubts.
Their applications?
For example I'm not a mathematician but I love thinking about weird "useless" shit like how math might be like for aliens? Are numbers as fundamental as we assume? i.e. humans developed math for "arithmetic" first, then latched geometry etc on top of that. We took ages to admit zero and negative numbers.. what if an alien species develops math for "navigation" first, and starts out with complex numbers right away!?
> What happens when humanity becomes obsolete?
There's an infinity out there to explore.
> And what happens when the machines that cause this obsolescence are controlled by a tiny amount of people, who suddenly don't need the rest of us?
That's a social problem we needed to tackle more than 100 years before AI or even computers appeared.
Apparently the minutes hand was added to clock to keep time in factories, for the benefit of the factory owners, not the workers — something I learned from this 1991 show from the BBC with Terry Jones: "So This Is Progress" https://www.youtube.com/watch?v=-Em96NVxO9Q
At least, that's how I view most discussions on AI adoption. Technological advancements are great for humanity, but that doesn't mean it comes without costs. The luddites are a famous example that's very often mentioned in this forum.
And no, "reskilling" isn't an option for many people. If you are poor, if you have a family or have people dependent on you, you cannot put your life on pause to learn something new, especially if you have no guarantees it won't end up like last time.
Yes, but that's a social issue, external to technology but exacerbated by every new technology, AI or not:
UBI should be a thing: let people work on what they find fulfilling, instead of having to work to survive.
AI could help design a system for UBI that everyone agrees with, since it's so good at maths and shit now.
This problem HAS to be tackled. Removing/slowing AI will only kick it further down the road, not eliminate it.
It seems to me that all we are doing is a wild goose chase; Progress above all, to hell with any environmental/social impacts, the end justifies the means.
> This problem HAS to be tackled. Removing/slowing AI will only kick it further down the road, not eliminate it.
True, but people are generally selfish. They will put their own prosperity above that of the future generations, and I cannot blame them.
The actual fear underpinning this isn't even about being "poor", it's that being poor means starving, freezing, not having a bed to sleep on, not being able to get basic healthcare in emergencies..
It's possible to provide all those things without "giving away free money" to everybody, but..that's probably more complicated for now.
In any case, whether one "deserves" to live in basic comfort shouldn't depend on one's ability to do jobs that depend on holding back technological progress.
It shouldn't, but it does. So if we can't change this fact, we have to find a middle ground so that the people alive right now aren't thrown under the bus.
I don't disagree with what you are saying. Progress is inevitable in the end. I just wonder if we have to be destructive in our road to achieve it. Environmental and social damage are also problems that we need to tackle. Keep in mind that unstable societies, where people are fearful of the future, are prone to revolutions, and an unstable political climate is detrimental to technological progress.
I see very few people at the top speaking up about this, and that only makes me more skeptical of the usefulness of AI. If it's only going to be used against me, why would I ever support it?
This is what underpins the fear, I think.
You don't know how to farm but mentally you rest easy knowing a lot of other people do know. You also know there are books you could read to learn, if you needed to. Most things are like this, you could bootstrap your way to casting metals and probably even electric lights with only books and raw materials. Computer chips don't have this property.
Personally this is why I'd like to see libraries survive, even though I actually mostly read on my e-reader. I guess I took Anathem to heart.
Mathematics is an academy, and academies are human assemblages for producing truth (and the tools therein); they will stop producing if we forget to repair and refine them. It's just undeniable in the abstract.
(Sorry for the length, cut it as much as I could; mod(s) remove if you'd like. Talking to myself in the shadow of giants is how I'm coping with the ennui, I think.) That said, four philosophy nits on paradigms, scope, motivation, and pride:
1. Paradigms | The 'Math 1.0' rhetoric is undeniably powerful, but it makes it seem like he's unaware of his standpoint[1] by lumping all of "traditional mathematics" together. At the very least we've gone through four methodological revolutions in math, each one changing how the field is done on a fundamental level: ??? => Euclidean Certainty => Aristotlean Computation (~800s) => ~Newtonian Calculation (1600s) => ~Gaussian Systems (~1850s), and perhaps one in the 20th c. I lack the expertise to even gesture at. We also have clear analogues from parts of the other two acadamies in the 20th century alone: physics becoming an arcane, inelegant group effort in the ~1920s, and mainstream philosophy adopting a cloud of Kiki ideas vaguely revolving around Wittgeinstein & Chomsky in the ~1960s.
I totally understand this being distressing, especially when it's happening quickly. They, too, had people decrying the future of their fields. But we wouldn't obviously wouldn't change it, in hindsight; much of modern physics would be completely intractable without those strange, boring, unnerving methods, for example. More than intractable: unthinkable.
2. Scope | This all seems overly focused on Autumn 2026. Most egregiously, this is all built on the premise that RSI never happens, and we never acheive ASI. If we do, mathematics is almost assuredly A) the first academy to be completely outmoded, and B) the least of our problems. I cut a long thing about the caveats and effects here; at this point... if you know, you know.
3. Motivation | Ultimately this thread is focusing on human motivation throughout, a fact that would be more forgivable if acknowledged as an intentional tradeoff. Speculating that it'll be harder to have interest in math is just not worth withholding truth; for one thing, knowing that computers could solve a problem but it's banned to try would ruin motivation anyway, and worse. It's up to us to be motivated, and if I know us, we'll have no problem doing so as long as there's any utility there at all.
In more stark terms: trading progress in the fundamental academy for the sake of its current methods of recruitment and motivation seems like something posterity will almost definitely frown upon.
4. Pride | This is the common thread that weaves through all three preceeding points, I think, and is even stated in pretty blatant terms (that's Tao -- always a clear writer!):
Sure, his thesis acknowledges that some changes are welcome, but not radical ones; his tone implies tweaks to conference schedules and authorship norms rather than fundamental restructuring of what these professions are, and what it's like to dedicate one's life to the demos through them.Doing science (mathematic or otherwise) in this competitive, individualistic way is just clearly counterintuitive to me, even if it weren't a recent development. Imagine taking it to its conclusion and applying some kind of patent system to mathematics -- or even worse, copyright to combinations of symbols! Perhaps more riches would motivate some mathematicians, but it would so obviously eat away at the democratic principles that have brought us unimaginably far over the past 406 years.
---
TL;DR: What worked well for the past ~century is not particularly relevant, and I think Tao is missing the forest here, despite one of the best sylvan trailblazers around. On his side practically-speaking for heuristic and contingent reasons, regardless.
Of course it's good to have the discussion... So maybe, we listen to the nay-sayers, but defer judgement on the matter... That's wisdom.
Edit, to be clear, I consider Tao to be the wisdom provider, not an early nay-sayer!
The less said about Gary Marcus the better.
This is a technology not like prior technologies. Are we okay if the technology discourages a whole generation of Mathematicians? If the technology leads to 10x fewer mathematicians -- what impact does that have on the field? These are the questions Tao is asking. And I don't think he himself claims to have all the answers, he just doesn't wanna see the math _community_ die.
Take a look at this interview from two days ago: https://m.youtube.com/watch?v=oQypVVv1u1o
The interviewee is worried about the future of math research. He is not strictly worried about being replaced, instead he is worried that he will no longer be able to launder math-as-a-hobby through math-as-something-useful as is the case today. He lays out very clearly that grant proposals claim to have useful outcomes while the proposers know those claims are nonsense.
Business as usual in math, and frankly in all the other sciences, is to do research that furthers the researchers careers or personal interests and pretend that it’s somehow useful. This would be absolutely fine if it were privately funded, but it’s not, this is public money.
In every other endeavour, lying in order to get money is considered fraud.
We have collectively wasted a huge amount of taxpayer money and human time, entire careers, on things not likely to ever matter to anyone.
I look forward to science becoming automated so that we can have real progress instead of the current broken system.
People thought, back in the 17 century, that imaginary were useless (except as a trick for some calculations). Turns out the research into these numbers back then is amazingly useful today, 300 years later, in electronics and such.
Publicly funded maths research should continue, even if some taxpayers feel it's a waste of money.
However, with AI it actually may become so cheap that the scattergun random approach becomes more viable rather than less. It’s when human time and resources are scarce that you need to optimise. The hobbyist approach may therefore ironically continue, but without the hobbyists.
Note, I think the debate is mainly over what research should be funded, not whether any research should be funded.
In fact, the taxpayer should be funding fundamental research because it's so hard to justify profit from it - but that funding would benefit all in the long future.
So leaving the applied research that have commercial value be funded by private, commercial interests, would make more sense.
Thats great! So in other words it should be perfectly fine for AI to solve all of the supposedly useless math problems so that it might possibly be useful later. No need to worry about the mathematicians hobbies here.
of course not.
Hobbies can still be done even if AI gets it done more quickly. Just like today, where knives are made much faster/cheaper in presses and CNC machines, vs a blacksmith hammering. But still, there are hobby blacksmiths.
Do you really think a society with zero human mathematicians or scientists will outperform one with both human and AI ones?
So as measured by utility, I absolutely believe we don’t need humans doing science into the future. I’m sure people will continue doing it, but not for utility, for enjoyment - as a hobby. Probably we’ll all end up as dedicated hobbyists.
How much..?
> things that I do and that my colleagues do, this kind of like curiosity-driven, you know, applied math, computational physics type research has always been justified by, I would argue, intentionally blurring the line between what I would call, you know, science as product versus science as process
> science as product is very kind of clear-cut. It’s, you know, things like, you know, cure cancer, solve nuclear fusion, generate, you know, clean energy.
> And then there’s science as process, which is kind of the curiosity-driven stuff about, you know, like, “I want to understand protein folding,” or, “I want to understand, you know, turbulence,” or, “I want to understand quantum gravity,” or something. And broadly speaking, we have tended to justify the latter by kind of laundering it through the former
And the examples he gives are actually the more defensible ones, he talks about a friend of his working on some abstract algebra under the false guise of cryptography research later.
And it’s not just him saying it, this is simply true. He should be lauded for admitting it publicly, this is the only way any progress is made. At least, it used to be. Now it’ll be AI instead.
Because the other solutions are to a) quite literally become inhuman, with cyborg integrated TPUs running local models and networked interfaces to propierary models run in data centers, or b) assert dominance of human ignorance by burning civilization down, which doesn't sound pleasant.
Almost certainly not. It's just going to jump to a higher level of abstraction.