HN in RSCserver-reason-react
top.mdnew.mdbest.mdask.mdshow.mdjobs.md
← Back to stories

“Math 2.0” will need to value mathematical progress more holistically

536 pointsby ent101 12 hours ago535 comments

Discussion

Loading discussion
  • adrianN · 11 hours ago

    I wonder how we could formalize the notion of „interesting“ problems in a way that would allow us to automatically generate new interesting questions from the existing corpus of mathematics.

    • youoy · 11 hours ago

      For me this path means that our role as humans is just the understanding of intelligence? Everything else is secondary (or the last of our priorities) and would be better automated? This is a hard pill to swallow

      • latentsea · 10 hours ago

        We should think about rejecting this AI future.

    • sankhao · 11 hours ago

      Or maybe we should not let pure mathematicians decide which problems are interesting, but reward the practical applications instead.

      • adrianN · 11 hours ago

        If you only judge by practical applications you eliminate vast swaths of human endeavor.

        • etdznots · 10 hours ago

          Human endeavor and achievement does not make number go up for next quarter’s earnings

      • zorgmonkey · 10 hours ago

        This is wildly shortsighted view of mathematics. Historically many of the subfields of math which are presently most valuable were considered useless for decades or centuries. Number theory, non-Euclidean geometry, group theory, and Boolean algebra, to name a few.

        • sankhao · 10 hours ago

          This is false dichotomy. Many mathematical achievements were produced while looking at practical problems: Fourier is a good one. For every mathematician that accidentally made a practical contribution, there are many more that produced nothing of value, and would have with a better tuned reward function.

          • youoy · 10 hours ago

            I guess this AI wave will divide the population more between people that think that only economic value exists, and people that dont. Thats one of the timeless human debates. We are now living in the perfect combo of low morality and general human automation. So i expect the next few decades dominated by people who think (and have a "proof") that doing something without an expected economic gain is useless.

    • charcircuit · 10 hours ago

      A simple way is to just measure how long it takes to solve. If it takes more than 5 minutes then we don't have a good understanding of that area and its worth potentially investing into.

  • patternMachine · 11 hours ago

    "Taste"

  • cs_throwaway · 11 hours ago

    Let us know when it is clear that UCLA does not hire the candidate with the most top-tier journal papers. It’s more likely that instead of spending a 100K/year direct grant on two PhD students, PIs will hire 1 and have the student spend 50K on AI.

    • prodigycorp · 11 hours ago

      More likely is that frontier labs give huge discounts of EDUs if they consent to allowing training.

      • leni536 · 7 hours ago

        Plus requiring attribution in papers.

        • sebzim4500 · 6 hours ago

          Maybe, but I suspect that AI in maths will soon become so prevalent that it's use will be assumed by readers whether there is attribution or not.

    • rao-v · 9 hours ago

      My understanding is that top tier math programs don't hire based on journal pubs and haven't for a bit. Math journals are pretty slow and hiring is really via reputation building based on one or two big wow ideas / proofs propogated via arxiv and talks.

      • cs_throwaway · 7 hours ago

        It’s just as unlikely that this will cease. hiring a professor at a top school because they were an excellent math teacher… It won’t happen.

  • vasco · 11 hours ago

    He is right if model intelligence stalls. If model intelligence continues to improve soon there's no need for the prompter to understand anything or for any workshop as a mathematician will just be able to ask the model to explain how the proof works and models will do a good job at walking them through it step by step. There will be no gap in understanding. Now there is because the models are discovering things at the edge of what they can do and so suck at explaining it. There's nothing particularly special about a newly solved problem in terms of learning it. If we accept AI can explain all of existing math nicely, why shouldn't it be able to explain new proofs?

    • winterbourne · 11 hours ago

      >if model intelligence stalls. If model intelligence continues to improve So much in AI is dependent on which of these two outcomes occur.

      • doctoboggan · 11 hours ago

        I am not so sure that's true. Even if AI intelligence were to plateau at today's levels, there are still many gains to be had in speeding up today's intelligences. ASICs with burned in weights could become economical to invest in as they would retain usefulness longer than 18 months, and I feel we have only scratched the surface on possible usecases of local AI and what it means for almost any technology product or interface.

    • myaccountonhn · 11 hours ago

      If the mathematician doesn't understand anything, are they even a mathematician?

    • lancebeet · 10 hours ago

      That doesn't make sense to me. Terence Tao is great at explaining things, but he could spend months explaining some of his proofs to me without me understanding it. Even if the AI had a superhuman ability to explain things it's no guarantee that it could make a human understand.

      • vasco · 10 hours ago

        It couldn't make every human understand. But it could make Tao and other mathematicians understand. You're not going to become knowledgeable at everything suddenly, even though I doubt what you say. Many months of 1-1 tutoring and hard work from a student with a top mathematician would make you understand a lot. Just not sit down and read it first day.

  • ChrisArchitect · 11 hours ago

    Related: AHM Statement on OpenAI's October 6 Release of Mathematical Documents https://news.ycombinator.com/item?id=50000421 / https://news.ycombinator.com/item?id=49999159

  • teekert · 11 hours ago

    I think we can say by now that we should not listen to the early nay-sayers and just wait a bit. With every trend, not just "AI". They still have some points (the ethics and environment etc), but the we don't hear from the Stochastic Parrot folks anymore. Of course it's good to have the discussion... So maybe, we listen to the nay-sayers, but defer judgement on the matter... That's wisdom. Edit, to be clear, I consider Tao to be the wisdom provider, not an early nay-sayer!

    • tmule · 11 hours ago

      Yann LeCun and Gary Marcus are still at it - the latter having shifted goalposts.

      • zwaps · 11 hours ago

        Gary Marcus has has a long career now of shifting goalposts

        • vidarh · 11 hours ago

          Does he have time to do anything but shift goalposts, given how fast his are moving?

          • Chance-Device · 11 hours ago

            I believe he now just leaves them in his truck and drives it a bit further each day.

      • nl · 11 hours ago

        Yann LeCun's criticisms are at least balanced with an alternative approach he has teams actively working on and showing good progress in areas LLMs are weak. The less said about Gary Marcus the better.

    • computerfriend · 11 hours ago

      Tao was not an early nay-sayer.

      • teekert · 11 hours ago

        Certainly not implying that! He's the one with the deferred judgement and the wisdom.

        • computerfriend · 11 hours ago

          Oh, sorry, misinterpreted you! Although I am not sure I agree with your dismissal of nay-sayers. We can equally dismiss the opinions of early proponents. Perhaps maybe we should be hedging on all early opinions regardless.

          • teekert · 11 hours ago

            Yeah agree with you, and perhaps both are even important in our public opinion forming... Recently I've been hearing people like Grant Sanderson on AI, and they are all very "wise" and informed, neither dismissive nor mindlessly (p/b)ro, but really thinking implications of this technology through. Like Tao does here. I love it, these people provide real direction.

      • jatins · 10 hours ago

        If anything he is very pro-AI and far from a naysayer. Labelling anyone who doesn't isn't Linkedin style "stoked" "excited" about AI is not helpful. This is a technology not like prior technologies. Are we okay if the technology discourages a whole generation of Mathematicians? If the technology leads to 10x fewer mathematicians -- what impact does that have on the field? These are the questions Tao is asking. And I don't think he himself claims to have all the answers, he just doesn't wanna see the math _community_ die.

  • underdeserver · 11 hours ago

    Even when a problem got solved, there has always been value in publishing simpler proofs and corollaries that give better intuition into the broader field. If I understand Tao correctly, he's saying that's going to have to be the focus going forward. I just default to thinking the models are going to be much better than us at that, too.

    • Yamata · 10 hours ago

      There’s value but it is not rewarded commensurately with the contribution. Same for doing peer review.

    • matusp · 10 hours ago

      > the models are going to be much better than us at that, too. I wonder if this is true. The code produced by these models are not really getting any more elegant over time. On contrary, the models seem to be getting worse, often proposing really baroque architectures. You can use RL to optimize for correctness, optimizing the vibe seems much more difficult.

      • squidbeak · 5 hours ago

        For deeper complexity, elegance may turn out to be a hindrance. Better or worse may turn out to be descriptive of taste and nothing else. On the other hand, elegant alternatives might be brought into reach by this first messy contact with new ideas. Who can say at this stage?

  • yedhukrishnan · 11 hours ago

    This is a distilled version of what people say about the tech industry in the past year or so. Replace math with any field, and the statement is still relevant.

    • socializer · 11 hours ago

      I don't think it's that simple. I divide AI-impacted fields into three buckets: 1. Some present a unified line that the whole point of their craft is the human experience, and that automation is the antithesis of that. Marathon runners don't care that a car can get there faster, poets don't care that Poem Bot 2000 can write poems too. I think this is smart if you can credibly take this position. The difficulty is mostly convincing the buy side, which requires being very outspoken about your views. 2. Some appear to be undecided, with one faction taking the pro-human stance and another rushing to accelerate things with AI. A good example of this is mathematics, and I really wonder where they end up in the long haul. They have a very good claim on #1, because mathematics is pretty close to an art form and is robustly insulated from the pressures of the marketplace. But they can also choose option #3, below. 3. Some crafts prioritize results above all else, practitioners either rushing to extract as much money as possible before it all collapses, or believing that they can out-prompt everyone else forever and that their prompting skills are indispensable to their employers in the long haul. That's software engineering. I think this is going to be interesting to watch.

      • faeyanpiraat · 10 hours ago

        Seems like neat buckets but you may want to find one example for the first which can actually be done by an llm as running is not its strong suit afaik

        • socializer · 10 hours ago

          I gave two examples for #1 and one of them is something done by LLMs. Other examples include painting / drawing, songwriting & composing, etc. In all of these, gen AI output is pretty widely stigmatized.

      • autuni · 10 hours ago

        I think your second bucket would need to be changed, as it's not really undecided for some and Tao in fact argues (in a way) that there is no need for the field to decide. There is a need for results, but there is also a need for humans to fully understand why the results are the way they are and how they were obtained. This doesn't replace the human in the loop it's more of a cooperative effort. I'd end up with the buckets a) Fully human b) Hybrid human-AI c) Fully AI

        • lioeters · 2 hours ago

          Extrapolating from the list, the "Fully AI" camp is going to push the pedal to the metal, investing time and effort into maximizing their fusion with this technology. Mentally enhanced and amplified 24/7, and likely physically too, fusing their biology with hardware and robotics. Within a few years, not even a generation, they will become unrecognizable to the "Fully Human" category.

      • MisterMunchkin · 9 hours ago

        Ah but your buckets are based on your assumptions about the topic. You say marathons are just for fun, but prior to the wheel it was the only way to get around. (Other than horses in some places) So it’s not that running is immune to automation, it’s that we’re already post-automation and that only people doing it for fun are left.

      • magimas · 5 hours ago

        I don't think it's smart at all to position yourself in bucket 1. There's a reason there are maybe a few hundred professional marathon runners in the world vs tens of thousands of professional mathematicians. Bucket 1 is basically an "amusement for the upper classes" type of deal. Any field that goes in that direction would have to shrink down massively. It also devalues the field in my opinion from something really profound with actual impact in the world to a somewhat vain leisure activity. (basically going back to gentleman scientists) But I know other people would see it exactly the opposite way.

        • socializer · 3 hours ago

          > There's a reason there are maybe a few hundred professional marathon runners in the world vs tens of thousands of professional mathematicians. I don't buy this. I think there are other reasons why "professional marathon running" is a niche thing; it's probably just that it's not all that interesting to most people to practice in function of what it demands of your body, and not that interesting to patronize / watch. Take woodworkers. There's probably more professional woodworkers than professional mathematicians. In terms of utility, everything a woodworker does, a machine can do more cheaply and more quickly. The main reason the craft survives is just that we attach intrinsic value to furniture made by humans the old-fashioned way. This is shared by craftsmen and those who buy. And you could argue the same thing you did for marathon runners: custom furniture is just amusement for well-off people. Sure, but there's enough people with money to keep it afloat.

  • shubhamjain · 11 hours ago

    A very balanced perspective, and the concerns he raises are reasonable. He acknowledges that AI is going to transform mathematics, but simply dumping proofs on the math community and expecting others to do the grunt work of verifying, refining, and expanding on them is hardly a productive way to advance the field. There seems to be more interest in hitting some arbitrary benchmark (we proved X unsolved problems) than in genuinely contributing to mathematics. But what else is to be expected? It's become a maniacal race with too much money. Too much effort is being invested in proving that the exponential curve is still holding.

    • XorNot · 11 hours ago

      That seems short sighted though. A few years ago models couldn't do this at all, I'm not sure there's any evidence to suggest exploring and refining results is outside their capabilities or will remain so. OAI obviously have a fiscal incentive here, but to presume a year from now we won't see improvements and more succinct work on the results coming from models?

      • JumpCrisscross · 11 hours ago

        > I'm not sure there's any evidence to suggest exploring and refining results is outside their capabilities OP didn’t suggest that. The bar has been raised. Everyone has to meet it now. An inelegant solution squatted onto the internet doesn’t count as discovery per se , even if it’s impressive.

        • auggierose · 9 hours ago

          A correct solution verified in Lean will count in perpetuum. It is fine if you want more, but an achievement is an achievement, even if it is by AI.

          • oliculipolicula · 6 hours ago

            Hmmmm. It feels right that discovery is much more meaningful than achievement. "Bullshit lean proof" or "bullshit achievement" smells like it. "bullshit discovery" smells like a front-handed insult

      • Maxion · 11 hours ago

        But what should they do? They got all these proofs, should they just have sat on them?

        • JumpCrisscross · 11 hours ago

          > should they just have sat on them? It’s fine that OpenAI posted their findings. It’s not fair to claim these problems have been solved. Not until someone can understand and verify the proof and then communicate the core, novel methodological element to someone else.

          • CrimsonRain · 9 hours ago

            Solving a problem is not solving anymore. Up is down, pleasure is pain, darkness is light, slavery is freedom, madness is sanity.

          • gwd · 9 hours ago

            But this is Tao's point: Before, the mechanism by which a proof was verified and communicated and digested by the community was for the person who came up with the grotty, ugly first draft to engage with the community. Now there's nobody to really engage with, so the pipeline from "grotty, ugly draft" to "integrated into humanity's mathematical knowledge" has been broken. So yeah, probably we should stop saying "X has been solved", and instead say, "A Lean proof for X (or !X) has been generated". That doesn't change the fact that incentives are currently on finding the proof, and once the proof is generated by an AI, there's not currently a good mechanism / incentive structure to move that into the mathematical community. AI is here, so we need to find a new mechanism.

            • lern_too_spel · 8 hours ago

              It is not obvious to me that a single canonical human language write-up of a proof is the best output in this new world where write-ups are cheap. A human reader can query an LLM and get explanations of key points tailored to the reader's own background in mathematics.

          • Turneyboy · 9 hours ago

            Many of these are lean formalized. Arguably a much higher bar than whatever peer review provides in terms of verification.

            • psychoslave · 8 hours ago

              On some consideration, surely. On the other hand, but at some point this is borderline like saying "universe already solved every physical problems, including possibility to represent deep important point of its own structure in compressed intelligent ways" and then tell that reaching it in an actual grabbable artifact is left as an exercise. Possibly yes such a representation is possible. But it doesn’t mean it’s certain there is a "best compressed representation". And even less one that encompass everything important and that is understandable by any human brain, even the most exceptionally brilliant ones sponsored by a whole society to reach their best possible achievable performance on that goal through full dedication on that sole task.

            • cmceanga · 6 hours ago

              There is no guarantee that the lean proof is 1:1 with the natural language equivalent. The lean proof can be lesser. This happened in the Navier-Stokes proof, e.g. see [1] in example 3.1. Having the certificate doesn't necessarily imply correctness. [1] https://arxiv.org/abs/2610.08144

        • p_hoep · 10 hours ago

          No need. In the end the mathematicians that don't like this can just not look at the proofs or use them. They have that choice. Just like they didn't "ask" for them, they don't have to even acknowledge they exist.

        • usernomdeguerre · 10 hours ago

          >They got all these proofs... Your phrasing is illuminating that perhaps they aren't engaged in the creation, understanding, or integration of these proofs by humanity; they just have them. For them, this is a slidedeck they can pass to investors, creditors, the marketing department. Something they can add to the employee onboarding pamphlet. What should they do? Hyperbolic maybe, but perhaps engage with humanity.

          • nearbuy · 9 hours ago

            ...they did.

        • sans_souse · 10 hours ago

          BlackBox: The proof is in the pudding This isn't only bad for Math — it's bad for English too. 'Proof' is going to become the 2026 Most Misapplied Word of the Year.

        • pks016 · 7 hours ago

          At least check them properly. They have already withdrawn some of them.

          • pfdietz · 24 minutes ago

            That were not formalized.

      • rtpg · 10 hours ago

        The problem is that one a person writes a 60 page proof in theory that person has spent an inordinate amount of time on the proof and can answer questions, describe some insight, etc etc. If a random person is given a 60 page proof to digest and not the author, those hidden insights that _aren't_ in the paper might be completely inaccessible. Maybe the AI will "just" be able to provide the insights. Maybe. But pedagogy is tricky work, and despite these AIs being able to do all this fancy math we can't get them to write good cover letters yet, so.... Ultimately we might be left with just a bunch of intellectually unsatisfying proofs. This means way less drive to simplify the proofs or rework them. End result: we generate a layer of "less efficient" mathematics, that won't get built upon. We will not actually have any shoulders upon which to stand.

        • XorNot · 10 hours ago

          But why should process of discovering mathematical insights be any less attainable to AI models? The concern is being raised without evidence, because the evidence points to the gap simply being frontier models have just started to be able to get a raw proof out. Why, given existing progress, should we expect them to be unable to distill insights from those proofs? Certainly this even more likely doesn't matter at all for applications: if I can send a radio signal further because my AIs design it a certain way, that's an unambiguous result. Which is really the next step here: turn a proof into a "mechanical" application.

        • charcircuit · 10 hours ago

          AI can simplify and rework proofs too.

          • i_cannot_hack · 9 hours ago

            According to Scott Aaronsson, OpenAI set their agents on 8000 different problems, and got 372 final proofs. Even spending twice the original effort on simplifying and reworking those proofs so that they do not "feel like something written by someone who’s on psychedelics" would only increase the compute by less than 10% (assuming all the agents had a similar token budget). The fact that they did not do so can only mean that either (1) their agents currently lack the capability to do it, or (2) OpenAI are completely indifferent and do not care in the slightest if the proofs are understood or not.

            • gwd · 9 hours ago

              > The fact that they did not do so can only mean that either (1) their agents currently lack the capability to do it, or (2) OpenAI are completely indifferent and do not care in the slightest if the proofs are understood or not. Come now, this is kind of unreasonable. When you're working on a new technology, you first get the ugly, inconvenient-to-use prototypes functioning with the core new thing you need; then you work on packaging it up into a format useable in production. I'm sure the very first digital camera sensors weren't very useful for photographers either; but it isn't really even possible to build the rest of the technology required to turn raw output of a digital sensor into something a professional photographer can use until you have the raw output itself. The research is still on going on the raw output; getting things to the next stage, where the results are widely useable by professional mathematicians (and then on to engineers and scientists to whom the results would be practically useful), is a whole new research area.

              • i_cannot_hack · 9 hours ago

                Seems like you are just subscribing to the first option I gave, "their agents currently lack the capability to do it", but saying you think they will be more capable in the future if they can move away from the inconvenient-to-use prototypes after more research. Thinking it might be possible in the future is not in disagreement with anything I said, so I am not sure what you thought was unreasonable about my description.

                • gwd · 6 hours ago

                  Imagine someone looking at digital camera researchers showcasing a ground-breaking new sensor, and reacting by saying "Well obviously they're completely indifferent and do not care in the slightest if their work is used by real photographers or not." Like, "Orr... maybe they care a lot, but haven't gotten to that part yet?"

                  • i_cannot_hack · 5 hours ago

                    You are conflating the two options with each other (lack of capability vs indifference). One of them is true, not necessarily both ("or", not "and"). I did not claim a lack of capability was the same as indifference.

            • lern_too_spel · 8 hours ago

              Or (3) they feel a need to publish first, and that goal takes precedence over (2).

              • i_cannot_hack · 8 hours ago

                They claim each result used three hours of compute on average. Even spending significantly more on simplifying and reworking would delay the release with a single day at most. If avoiding such a minor delay took precedence over (2), I think "indifference" is the correct term. It has also been a while since the release now, so there is ample opportunity to post follow ups if time pressure was the only concern.

                • lern_too_spel · 7 hours ago

                  Why do that when they can spend more hours extending QRH to a proof of the full Riemann Hypothesis? The opportunity cost of digging up small potatoes is the whole enchilada.

    • mrheosuper · 10 hours ago

      Why not asking another chatbot to verify, like what we are doing with coding ?

      • kadoban · 10 hours ago

        That's not the hard part. The hard part still requires humans (and/or _maybe_ a bunch of tokens) and a lot of work.

      • iterance · 10 hours ago

        And then what?

    • gexla · 10 hours ago

      Right, easy comparison to make the the open source community for software. And it's not like this is something where we're loaned some top math genius for a limited amount of time and we have to make the most of it. Rather, this is a new high water mark. The accessibility of the results is no longer scarce. The scarcity has shifted, and that's where the focus of the math ecosystem should shift as well. And it doesn't help for a frontier community to saturate and take over messaging pipelines that were typically managed by the math ecosystem. It's not about "stay in your lane" but rather "we need coherence and be careful not to break the system." Just two cents from someone who could screw up basic cashier math on any given day.

      • pfdietz · 59 minutes ago

        OpenAI has kicked things over. Dealing with that is the math community's problem. OpenAI had absolutely no obligation to defer to that community's deficiencies. If this has upset their reward system and historical traditions, tough luck.

    • flexie · 10 hours ago

      Yes, it is indeed a balanced view. Those few thousand mathematicians are now getting a taste of their own medicine. After all, it was people with extraordinary mathematical talent who developed machine learning and large language models, leaving hundreds of millions of people who earn their living through speaking, writing, or teaching worried about their future job prospects. Still, I believe almost everyone will be fine. Perhaps AI will also prove good at coming up with new conjectures, and some mathematicians may shift towards applied mathematics or other sciences.

      • FrustratedMonky · 5 hours ago

        Pretty tenuous connection to blame mathematicians for everything bad created in the world, that happened to use math.

    • caaqil · 10 hours ago

      > There seems to be more interest in hitting some arbitrary benchmark (we proved X unsolved problems) > genuinely contributing to mathematics What's the difference between the two? Proofs are no longer the goalpost?

      • shubhamjain · 10 hours ago

        Not an expert, but explaining the proof and it being independently verifiable as important as just putting a paper of it. Reminds me of 1000 page proof of Goldbach’s conjecture that a Mathematician reached sometime back. He was told plainly that no one is going to invest time in verifying the proof because there’s a good chance there’s an error somewhere in between.

      • tene80i · 9 hours ago

        Proofs are valuable but I believe the mathematical community values understanding more. Proofs were previously a great way to develop understanding. Now, less so.

      • pegasus · 8 hours ago

        I recommend RTFA, it explains exactly that: why proofs should not be considered the goalpost, and why dumping all these AI-generated proofs might be an overall negative for mathematics as a whole.

      • desterothx · 8 hours ago

        Proofs of open problems are valuable because we are assuming that proving the problem requires some new method or infrastructure in math to prove it. Basically proving open problems isn't actually useful if it doesn't develop new tooling for mathematics, which can help us create new open problems, solve other ones etc.

    • mastermage · 10 hours ago

      Its kinda like the arms race in the cold war. There came out some truly marvelous technologies but the actualy goals were frankly terrifying.

    • antman · 10 hours ago

      This argument implicitly makes a few assumptions which will probably not hold in the very near future. One is that AI will continue hallucinating in a manner that is not easy to verify, second is that AI will not be enhanced to produced more simplified amd robust outputs, and third that a human will be required to do that. What humans in the loop are doing now is verify the process, propose shortcuts and add legitimacy, through the verification process, if that ends up being succesful its highly likely a lot less mathematicians will be required in the future. The conclusion that this is not productive focuses on the mathematicians, but it is very productive in terms of hundreds of proofs being produced that had previously consumed uncountable hours of the brightest minds. Unless it ends up being the greatest hallucination ever ofcourse

      • thereitgoes456 · 10 hours ago

        AI cannot explain chess moves it comes up with in an elegant way. What makes you think it will be able to do so for math?

        • jstummbillig · 9 hours ago

          It's kind of astonishing that after all we have seen in the last years people still find the position that AI will not be able to do an obviously valuable thing likely and it requiring an explanation (instead of the other way around).

          • hodgehog11 · 9 hours ago

            No, this is different, and this is coming from someone who has been studying deep learning for the last decade. We are talking about the difference between RLHF and RLVR strategies. The former benefits clarity and explanation, while the latter concerns only correctness. AI was moving in a particularly damaging direction by pushing on the first path, so it was natural to move to the second. But the second will come at the cost of clarity of explanation. It will likely get better at its explanations, but not fast enough to render its most advanced accomplishments readily understandable to the user. The chess example is a pretty good one (that is an RLVR approach).

            • red75prime · 9 hours ago

              The problem is that people strongly believe that this is an insurmountable problem that will persist indefinitely (or for a long time) and plan accordingly, while this, most likely, will be fixed soon by adding RLCAF (RL on conversational agent feedback) or something like that.

            • __s · 4 hours ago

              tbf GM explaining their 2700 elo moves are only understandable when vague, as elo goes up explanation becomes closer to "in this specific position there's these dpecific lines", why should 3500 elo moves have simple reasoning? Maybe if we start with giving simple AI generated analysis of those clumsy humans with their measly 2700 elo moves

          • n6242 · 8 hours ago

            Some of us still remember 2016, when we had a couple of cars sorta half-driving themselves, and Tesla, Uber and others promised we were only one year or two away from three million people in the US working as drivers being out of a job. And here we are, a decade later. AI is pretty amazing, but companies have a tendency to severely and comically overestimate and oversell it's capabilities, and underestimate the challenges.

            • jstummbillig · 7 hours ago

              > And here we are, a decade later. With Waymo and Tesla increasingly doing what they said they would do, and a small number of early adopters happily paying money for their services, that do work. So what's the critique? That the timelines are not correct? Sure. And how about the timeline of the people who said "research level math, never in my lifetime" and the people inside the ai companies who are apparently increasingly spooked by how quick the progress is? How about the various levels of code/programming jobs that AI was supposedly never going to be able to do, but, in reality, now just does? We are engaging in some very one-sided discrediting, and I am not sure, why.

              • jazzypants · 4 hours ago

                And, those companies are notably avoiding wet climates because they still struggle with self-driving in inclement weather. It's probably going to be another decade before we get to the point where these things can handle every situation. Just like all other engineering, the first 90% is the easy part. https://www.wsj.com/articles/self-driving-cars-dont-do-snow-...

                • jstummbillig · 3 hours ago

                  Alas, you will notice that article is 1.5 years old. Update on this issue (albeit soon from a year ago too): https://waymo.com/blog/2025/10/creating-an-all-weather-drive... More telling: Waymo just rolled out in Denver (1 month ago or so), apparently fairly confident they got this handled given the upcoming winter. Progress on the obvious stuff keeps happening (which kind of brings me my to the first comment here). But also: It does not have to do "snow" to be useful! A lot of cars/people don't drive when it snows heavily, and that's something we have always been okay with (at a societal level, YMMV of course). If was only useful 95% of the year that's still great. A lot of technology works like that.

                  • jazzypants · 2 hours ago

                    Great response, and thank you for the article! I'm skeptical that they're actually ready for real-world conditions outside of a test track, but that's a whole lot of training data and their lawyers must be convinced. We'll see.

            • p-e-w · 7 hours ago

              Driving a car is unimaginably more difficult than proving the Riemann hypothesis. You just don’t notice that because evolution has given you 99% of what is needed to drive a car before you were even born.

            • inglor_cz · 2 hours ago

              I remember a biologist wryly commenting that it took a lot longer to evolve good senses and appendages in nature than to evolve human intelligence, compared to the ape niveau. Maybe the really hard thing isn't abstract cognitive capability, but perception+movement.

              • pfdietz · 1 hour ago

                Moravec's argument strikes again.

        • ApolloFortyNine · 4 hours ago

          Interesting claim. That's true for old models that simply have no way to explain, LLMs however can. [1] [1] https://dev.to/natcher/researchers-develop-method-to-train-l...

      • mathisfun123 · 9 hours ago

        > One is that AI will continue hallucinating in a manner that is not easy to verify, second is that AI will not be enhanced to produced more simplified amd robust outputs, and third that a human will be required to do that. There is literally not a single shred of evidence to indicate either of your supposed eventualities. The core technology of an LLM is sampling from a distribution so there is literally no way to make it deterministically robust (only probabilistically).

        • antman · 9 hours ago

          The direction and pace of capability improvement has already been demonstrated by all models. The latest breakthroughs make that pretty evident, but there have been production systems that are based on probability since the beginning of computing. What has been demonstrated is a process that outputs lean proofs based on those probabilities. This happened after decades markov chain producing garbled texts and very shortly after gpt2 producing stories about unicorns.

        • red75prime · 8 hours ago

          > The core technology of an LLM is sampling from a distribution so there is literally no way to make it deterministically robust (only probabilistically). An LLM mostly deterministically (except parallel processing nondeterminism that can be mitigated) produces a probability distribution that can be sampled deterministically: just take the highest probability token or use beam search.

          • Vetch · 8 hours ago

            Deterministic yes, robust deterministic no. The most likely conjunction is not always the best nor representative of what the model is considering unless its certainty is high.

            • red75prime · 6 hours ago

              I think determinism has nothing to do with it. If you mean sensitivity to word ordering and such, it's a generalization failure.

          • gottheUIblues · 8 hours ago

            I think people on here tend to somewhat fixate on the determinism issue. Even with a deterministic LLM - stabilising the floating point arithmetic, and choosing from the distribution by a fixed method, or just save the random seeds - there is still a kind of a chaotic unpredictability that can exist between its inputs and outputs. However maybe that is a price that needs to be paid to get creativity.

        • bonoboTP · 4 hours ago

          Is a human deterministically robust? Or is a human also incapable of doing what you claim LLMs will never be able to do?

        • emtel · 3 hours ago

          No evidence other than the fact that this has been happening steadily in all areas for many years? You might have a point if the goal was to have LLMs that spit out a correct proof without chain of thought or tool use. LLMs + agent harnesses are more than capable of self verification and course correction.

      • Terr_ · 9 hours ago

        > assumptions which will probably not hold in the very near future [...] One is that AI will continue hallucinating in a manner that is not easy to verify Hold up, that's an even bigger assumption in the opposite direction, and I don't see anything to support it. At least in terms LLMs getting all the "AI" hype these days, there is no structural/mathematical reason to believe they won't continue to have the same problem they've always had of generating plausible text over rational text, and I don't think anybody even has a clear idea how it could eventually be accomplished. I've seen "then the magic singularity occurs and somehow it solves the problem for itself", but I would classify that more as mysticism than engineering.

        • hodgehog11 · 9 hours ago

          "Plausible" text was preferred over rational text when we trained LLMs using RLHF. It's rapidly shifting the other way now with RLVR, which enforces correctness by default.

        • user43928 · 7 hours ago

          Hallucinations are no longer much of a practical problem in software engineering. Two years ago, hallucinating that the code worked or that a task was accomplished was a common occurrence. We have seen that now agent swarms across thousands of agents can coordinate to achieve a result. Clearly hallucinations are no longer the problem they once were, since now we can get working results for long horizon tasks that require massive compute. Consequently it would seem unwise to assume that current limitations will remain as they are and prevent LLMs from coming up with solutions that they can explain to humans.

          • catlifeonmars · 5 hours ago

            It’s still a common occurrence. It happens in more subtle ways, but it still happens often enough for me to notice. For example I have had hallucinated checksums show up in lock files as recently as yesterday using a SOTA model. This is not surprising, since the whole basis of LLM training is to produce output that humans will accept _as a proxy for actual training goals_. In a sense, the training process of an LLM “wants” to produce output that is statistically plausible much more than it “wants” to produce correct output. It’s always going to be a struggle to drive that system towards other goals (and we see this bourne out in practice by the amount of effort that is required to be spent on RL). I think there will be some threshold of correctness (something like 99.999% of the time) that if the model surpasses it, I can stop needing to check it, but I think we’re still at 99% or something which sounds good, but when you are producing a ton of output you hit that 1% frequently. > Consequently it would seem unwise to assume that current limitations will remain as they are and prevent LLMs from coming up with solutions that they can explain to humans. I 100% agree with this. In fact explaining things to humans is something LLMs are particularly well suited for.

          • ashkankiani · 4 hours ago

            The confidence with which you, anonymous user, keep commenting that "hallucination is not much of a practical problem in software engineering anymore" based solely on your own anecdotal evidence is really remarkable, in not a good way.

            • user43928 · 4 hours ago

              You're free to substantiate your comment by telling us about your apparently different experience.

              • jazzypants · 4 hours ago

                I find hallucinations in my (mostly perfect) AI output every single day. If you're not finding them, you're just not looking hard enough. It's not surprising when everyone is screaming about how they don't read code these days. This is just a fact. I'm sorry if it messes with your narrative. https://arxiv.org/abs/2401.11817

                • user43928 · 4 hours ago

                  Thanks for the link. What kind of hallucination are you seeing, and does it affect the end result?

              • ashkankiani · 4 hours ago

                The sum total of all human observations is still not proof of the lack of hallucinations as a problem (even if their observations were perfect, which they aren't considering the volume produced vs reviewed carefully). That's why you can use a counter example only to disprove and not prove anything. And yeah I get hallucinations all the time still. Maybe it's because I'm working on harder/more niche problems (like a compiler with an unusual type system), but it happens quite a lot. I don't record all of them. Although the most common one you can find is them misattributing the source of changes from themselves and also other agents (Fable, Opus 5.5, deepseek, whatever). They'll say "your changes" or "you changed" or "your ruling." I didn't decide anything and it's in their own chat log, and yet...

      • Retr0id · 9 hours ago

        Even if you somehow have a 100% correct AI, it's not useful unless we can understand and internalise (and communicate) its results.

        • colordrops · 9 hours ago

          Who is this "we" you speak of? The professional mathematician community? Were pre-AI results useful outside of this community of people who could understand them?

          • dofm · 9 hours ago

            Something about this sentence makes me think about that Rob Auton bit, that before there were mobile phones, nobody had any reason to tell someone else that they were on a bus.

            • zer00eyz · 5 hours ago

              Before mobile phones you never called someone and asked "where are you" because a phone was tied to a location...

          • hodgehog11 · 9 hours ago

            Yes. Most probably do not understand the notation involved in, and the statement of, the Lindeberg-Levy Central Limit Theorem. But every scientist uses this theorem in one way or another. These ideas have a way of trickling down because to people who work thanklessly to do so.

      • dofm · 9 hours ago

        > One is that AI will continue hallucinating in a manner that is not easy to verify It is an old saw at this point, but what an LLM does still cannot be divided into hallucination and non-hallucination. This is literally an anthropomorphism trap. Layers and layers of application-specific verification can reduce the risks inherent to LLMs, to a really remarkable degree, but nothing about what these tools are suggests that this problem will go away; it will just bubble up again somewhere else.

        • antonvs · 9 hours ago

          > It is an old saw at this point An old saw unless something that's widely accepted, but sadly it seems that many people don't recognize this, even many people working in the field.

        • user43928 · 7 hours ago

          And why not? For all that I saw over the last few hundred hours with AI on software engineering, hallucinations are no longer a problem at all. Not once have I seen a task fail due to what would have been a "hallucination". If they still occur, they can apparently be detected and corrected automatically, or are subtle enough to escape notice with presumably no significant impact on the results. Why would this not also be the case for mathematics?

          • catlifeonmars · 5 hours ago

            I think OP is saying that hallucination or not is just semantics. There is nothing qualitatively different about hallucinated vs non-hallucinated output.

            • bonoboTP · 4 hours ago

              That's true in the same sense as "There is nothing qualitatively different about erroneous vs non-erroneous output" for a dog vs. cat image classifier.

              • user43928 · 4 hours ago

                To be fair, I guess the line is blurry between what could be labelled a regular mistake compared to a hallucination. "Test suite passed" when it actually errored? Obvious hallucination, unless it ran a command that returned the wrong error code. But is running a malformed command that does not achieve the expected effect itself a hallucination?

                • bonoboTP · 2 hours ago

                  If it makes a false claim, then it's an error. If it says the test was passed or a class was implemented but it was not, then it makes a false factual statement. I'd say a hallucination (very misleading word) or confabulation or "making shit up" happens when an LLM uses factual / evidential language purely based on local statistical expectations of the text, instead of it drawing from actual evidence in its context pointing to it. This is murkier in the case of general knowledge questions, like when and where was some famous person born. It may then be a spectrum from fully making something up based on how the name sounds, all the way to confidently retrieving it from its weights correctly. In between, we can get hallucinations. But newer models are taught to use Web Search when unsure, and it works pretty well, though not perfectly. I don't see any fundamental limit here. It's just not perfect. Trying to solve "the hallucination problem" is basically like saying "our dog vs. cat classifier is pretty good already with its 99% accuracy, now all we need to do is the tiny little task of eliminating the 1% error, and we will be golden". Like, no shit, there is some error yes. People are working to reduce it. It will never be absolutely 100%. It's not an insight to say we should remove hallucinations.

        • qarl · 2 hours ago

          > can reduce the risks inherent to LLMs, to a really remarkable degree To an arbitrary degree. Just like all of science. Reduce the error to the desired margin.

      • wolvesechoes · 9 hours ago

        Expression of tech-faith is not intellectually honest argument. Where does this "will probably not hold in the very near future" come from? People correctly warn about extrapolating current things onto the future, but then just throw some vague "probabilities" without providing any argument why their "probably" is somehow more grounded than others.

        • antman · 8 hours ago

          Markov chains garbled text to gpt took decades, gpt stories about unicorns to gpt production systems took a few years, gpt production system to gpt astra producing deterministic lean proofs of longstanding mathematical problems happened even faster. Scepticism to the point of requiring proof appears like an academic pursuit while production systems have already been built and are in the process of being enhanced

      • Vetch · 8 hours ago

        Putting hallucination aside, LLM "theory of mind" has gotten worse over time. I feel it peaked in Opus 3 and Sonnet 3.5, GPT 4 and then GPT 4.5 for OpenAI. Since then, even with Opus 5.5, phrasing has needed careful crafting, in order that it not be taken too literally. OpenAI models suffer from this much more than Anthropic models but Claudes have backslid over time too. This means when writing documentation, tutorials or commit messages, their output is often a garbled jumble. Assuming shared context, using invented terminology without explaining, leaking conversational states due to improper epistemic boundaries and failing to model the reader. This all usually leads to their freely generated explanations being terrible. Getting good explanations requires chaining questions that force them to line things up properly, which is not easy the less you know. These failures as something LLMs naturally struggle with make sense, given the nature of attention and RL with weak signals from human data. Math is not merely a collection of proofs, it's a way of understanding. A proof presented in a manner that cannot be incorporated remains useless. It does not make it's way to physics like Riemannian geometry and matrix math did. This is no less true when done by humans too. Your hallucination conclusion, checking if a proof is one, is exactly the counterproductive cost. Most of us cannot verify that the claims in the OpenAI lore dump are in fact all correct. It will take tons of work from experts to do this. It took subject expert mathematicians to identify the discrepancy and disconnect in the Navier Stokes proofs, for example. LLMs will struggle to make use of their own proofs or turn them into knowledge that accumulates over time. The act of proving is often more valuable than the proof itself. Human constraints and limitations force us to invent tools and abstractions that a 100,000 x 1M context swarm can bypass. The tradeoff from that AI swarm advantage is work that doesn't usually lend itself to being built upon. It's like doing all the side quests and reading all the books of an RPG versus min maxing a straight path with a guide. We might try to identify new abstractions, but the fact that we don't get access to CoT and that much of it will be illegible means mining LLM traces for what human mathematicians produce naturally will be a tedious chore.

        • p-e-w · 7 hours ago

          > Since then, even with Opus 5.5, phrasing has needed careful crafting, in order that it not be taken too literally. This is a feature, and a huge step forward. If you expect AI to do serious work, you can’t have it guessing what you “really meant”. Every sufficiently advanced task depends on very subtle details in the problem statement, and the correct default behavior for advanced AIs is to solve the task exactly as stated, unless a system prompt or other constraint tells it to do otherwise.

      • pegasus · 8 hours ago

        Did you even RTFA? His argument absolutely doesn't make any assumptions about hallucinations, implicit or not. It's you who assumes Tao must have surely been complaining about hallucinations or some such. You've not addressed any of his arguments and moreover ask questions his post answers. Here's a longer article which goes into a bit more of the details: https://terrytao.wordpress.com/2026/10/05/the-future-of-math...

      • catlifeonmars · 5 hours ago

        > but it is very productive in terms of hundreds of proofs being produced that had previously consumed uncountable hours of the brightest minds You’re making the following assumptions: 1. the exercise of struggling to find proofs was not productive, but this is precisely how new techniques in math were produced. Brute forcing solutions doesn’t lend itself to the creation of much new mathematics (except maybe the exercise of developing verifiable proofs) 2. the point of doing mathematics is to be “productive” in the first place. This is silly. Many people get into mathematics because of the beauty of understanding, for example.

        • bonoboTP · 4 hours ago

          > 2. the point of doing mathematics is to be “productive” in the first place. This is silly. Many people get into mathematics because of the beauty of understanding, for example. Are they independently wealthy? Or do they have a deal with their local supermarket that they can take food for free?

          • catlifeonmars · 2 hours ago

            So maybe we should just pay them more to do math and let them set the direction of research? If your point is that capitalism fucks up the incentive structure and makes it all about maximizing productivity then I wholeheartedly agree with you.

    • schleck8 · 9 hours ago

      > but simply dumping proofs on the math community and expecting others to do the grunt work of verifying, refining, and expanding on them is hardly a productive way to advance the field This is a transitive period. In a few years, verification and exchange between model instances will happen faster than humans can follow. Human input will be an ethical question, and not a productivity one, because it will be the bottleneck in any science.

      • Muromec · 9 hours ago

        And what use of that?

      • psychoslave · 8 hours ago

        Etic is a factor of productivity, and the larger the contextual window is the weightier it becomes. That’s even integrated within the paperclip parabola. If the focus in placed on maximizing some easily measurable output on a narrow perspective, situation is unlikely going to match a sweet spot of holistic equilibrium which is maximizing harmony and happiness through humanity as a whole.

      • RandomLensman · 7 hours ago

        How to make sure it doesn't evolve into some sort of Library of Babel of science?

        • schleck8 · 4 hours ago

          That's the billion dollar question. Nvidia currently wants to deploy chips with the sole purpose of monitoring agentic workloads and that doesn't seem farfetched but they obviously have a financial motive to sell more products. It's like a cat and mouse game, like cybersecurity in general.

    • kuboble · 9 hours ago

      I wonder if top labs will soon abandon math progress like they did go and chess. In example of go where I'm more familiar Google deep mind poured large resources to get a super human performance first, establish superiority and abandon it. The community then built their own tools starting from reproducing their papers. I think similar thing might happen to math. Nobody outside of math cares too much about Hamiltonian cycles in some bizarre graphs or proving lower bounds on complexity of some problem. Once those results stop being worthy of mainstream media attention, they will abandon math and the progress will be done by mathemicians guiding the models and the community will likely establish some new rules about what makes a valuable contribution. Merely solving not yet solved problem might not be it anymore.

      • diamondage · 8 hours ago

        This misses the raw advantage of a good proof. It makes conceptualization simpler . In some ways math is like a hash list of of theorems. This list makes it simpler to prove other calculations, and will always be useful, to both humans and AI models. I can see two new directions 1 - the creation of specialist theorem models; that can answer questions efficiently about one topic and 2 - we probably need to incentivize and codify ownership of theorems; charging a proportion of the compute saved by using them. Ultimately enabling mathematicians to be paid our true market value!

        • edot · 8 hours ago

          Oh boy, please not 2. What if this was a thing already and, since neither Newton nor Liebnitz had kids, we all had to pay some investors who bought the rights to calculus every time we took a derivative.

          • Dylan16807 · 8 hours ago

            You know patents only last 20 years right?

            • cnr · 8 hours ago

              As long as somebody with big $ decides: "let's make it 40!"

              • noworld · 6 hours ago

                Aaaaasaand now Disney owns the rights to room temperature superconductivity.

        • lh712 · 7 hours ago

          Considering how essential math and science is for the prosperity of mankind (not even speaking about the cultural value) the question of how to reward people working and contributing in these fields effectively and appropriately is of extreme importance. (And I think the current decline in our societies is to no small degree caused also by our utter failure to address that issue.) It is also fascinating, because I don't think there is any solution within our existing system, at least not any I know of. Theorem ownership is not a good solution (and neither are patents in general). Probably the most achievable (or rather the least unachievable) solution is a kind of communist utopia, where people can dedicate their time to a pursuit of any endeavor they see fit, as resources for a decent life are abundant and excessive power capture impossible. (The other option, somewhat dystopian, and which would not require humanity to change too much in its current mode of conduct, would be a totalitarian or caste-like capture of society by the scientific community.) Incidentally, if AI proves as powerful as some expect it to become, it could bring about another solution of that issue by making all human science and mathematics obsolete, pushing its true market value to zero. (With apologies for rambling.)

        • petesergeant · 6 hours ago

          > It makes conceptualization simpler I wonder if it makes conceptualization simpler for models too, given that they're trained already on human-speak. And I'm also curious as to whether humans currently have an innate advantage into simplifying and contextualizing proofs, or will the machines get good at that as well?

      • bluecalm · 8 hours ago

        Yeah chess is a good example. DeepMind came for publicity with AlphaZero. Arranged a match with Stockfish with rigged rules to make AlphaZero look better than it really was (it was amazing but the match wasn't fair) and then just published some games and went home. I was bitter about that back in the day as I hoped for more answers, more matches, more "truth" about chess being shown. Soon after that community project Leela Chess Zero was started and not only surpassed original AlphaZero but added few hundred ELO points over it. Then the combination of NN and classical engines happened with NNUE and current Stockfish is again a few hundred ELO points stronger. Today we pretty much know the truth in chess for all practical purposes. Human analysts/preparation experts focus on finding interesting path and opponent profiling (what is the most unpleasant for the opponent to face). They don't look for truth anymore. The game is doing great, it's more popular than it ever was.

        • squidbeak · 6 hours ago

          Why don't you mention the second match here, with its adjustments to meet Stockfish's quibbles - and the same result? https://en.chessbase.com/post/the-full-alphazero-paper-is-pu...

          • bluecalm · 5 hours ago

            Stockfish and other classical engines were never intended to be run in matches without opening books (of which there were plenty). Development assumed the presence of opening book and authors made 0 effort to make engines play well in openings because of it. This is also the reason classical Stockfish was a very small binary. A little effort to make it even by including even a very small opening book (like 50MB or something that would result in still smaller binary than NN engine with its net) would make it much more interesting. The result was that Stockfish lost many games by walking into known bad lines and lost way more games than it otherwise would.

            • squidbeak · 5 hours ago

              You aren't correct. > We also played a match that started from the set of opening positions used in the 2016 TCEC world championship, along with a series of additional matches against the most recent development version of Stockfish, and a variant of Stockfish that uses a strong opening book. In all matches, AlphaZero won. https://deepmind.google/blog/alphazero-shedding-new-light-on...

              • bluecalm · 2 hours ago

                Ok I remember it vaguely but the match that got publicity and the one published results were derived from was 1000 games match from starting position. Deepmind claimed AlphaZero also won from TCEC positions and vs Stockfish with good opening book but at least back in the day I don't think I could find those games being published or specifics about books/positions they have used. Can you? I am not claiming AlphaZero wasn't stronger. It wasn't as strong as the PR piece suggested though and we have never seen the games being published. In chess this is extraordinary because basically all games in chess are publicly available - both human and computer games. Claiming "we have created a strong engine that has beaten Stockfish with opening book" while not showing those games (or details about opening book used) is akin to "we solved this math conjecture" without showing any kind of proof or argument. Publishing a few 1000 of games costs nothing. Tens/hundreds of thousands of games are published every day.

      • im3w1l · 8 hours ago

        I disagree. Firstly, people in AI likely care about math on a personal level. Secondly math is useful . Playing go or chess is basically a party trick. Being useful gives it staying power. But, I do think you are right that there will be some level of moving on. The spotlight is currently on maths and that won't last. It will move to some other area where there is more impact to be had. So while they might shift gears and put less focus on math, it will always be there as part of the portfolio.

      • renyicircle · 8 hours ago

        I had the same idea recently. You've solved all the famous conjectures (all formulated by humans because humans found them interesting), what next? I doubt "AI formulated a math conjecture that nobody else cares about and immediately solved it" will produce that much hype. The actually interesting thing is indeed how mathematicians themselves will use these AI models going forward and how that will shape mathematics of the future.

      • pizza234 · 8 hours ago

        > I wonder if top labs will soon abandon math progress like they did go and chess. I definitely think that this is marketing, just "with good side effects". My doubt is when they will be able to move to "marketing with better side effects", that is, research with more concrete outcomes (health, materials etc.). Problem is, that type of research is much harder. Some doubt that progress in such areas will be quick ( https://www.noahpinion.blog/p/wheres-the-intelligence-explos... ).

        • td6 · 7 hours ago

          Would that be a bad thing? While top AI labs no longer focus on chess, the community build way better chess engines. Stockfish is probably stronger, than everything the top labs build. Wouldn't we expect the same thing for math? That slowly the broader math community would engineer a harness/program... That will surpass the current labs, and be a community ran project

          • sebzim4500 · 6 hours ago

            Yeah its certainly true now that Stockfish is much stronger than alphazero, but it's probably also true that had Deepmind spent another few years working on alphazero it would be enormously stronger than either. In the case of chess this seems fine, there isn't much value to society in creating an AI capable of beating top humans with a 4 pawn handicap rather than a 2 pawn one, but for maths where there are actual applications it is more complicated.

            • kuboble · 4 hours ago

              But arguably - it's better for chess community that the best engines are opensource than having superior GoogleChessBot.

        • simonh · 6 hours ago

          They aren't trying to 'solve' chess, go, or mathematical proofs as an end in themselves, but mainly in order to learn more about how to build better systems overall. The goal of AlphaZero was ultimately as a stepping stone towards AGI, and it's the same with LLMs.

          • catlifeonmars · 5 hours ago

            The assumption being that 1. these things are all stepping stones, not diversions 2. that ai labs have a singular goal of producing agi

            • simonh · 2 hours ago

              It's not an assumption, many of the people behind these projects explicitly say this what they are doing and why.

      • Xmd5a · 7 hours ago

        I don't think it will be the case, maths have real utility. I found something interesting at the intersection of combinatorics and information geometry. To be quite frank I don't understand what I'm doing. And yet, when I ask ChatGPT to use the framework we're developing to write an algorithm, it turns out it has quasi-parity with the state of the art. I have to measure absolute perfs to decide which one is better – theirs, not mine. Ok. Time to keep improving on what I have. And this implies dropping the code and going back to the blackboard doing more super abstract math that are way out of my league.

        • lh712 · 7 hours ago

          Good luck! We are at the point where the way in which humans do math and science changes significantly, and I have no good idea at all in what state is it going to settle down. But you are one of (many, I suppose) people exploring the new wilderness, so I wish you best.

          • Xirdus · 4 hours ago

            If we get AI singularity, then humans will stop doing scientific progress altogether. But if we don't, then it's pretty predictable what's going to happen - things will be much the same as now, except everyone will be using AI for proofs, data analysis, theoretical models and designing experiments, so important discoveries will happen more often. It's also possible that after the AI craze dies down, we'll have enough computational capacity to solve protein folding.

        • kuboble · 6 hours ago

          It has utility so some people will pursue it, but it has no immediate business value so I don't believe ai labs will keep spending millions on it. Unless they decide that trying p!=np is worth any money.

          • TheOtherHobbes · 6 hours ago

            Some math has almost incalculable business value, because math is the biggest driver of game-changer technology. We'd be nowhere without Laplace and Fourier transforms, Maxwell's equations, elliptic curve cryptography, and many more. Most math doesn't, but often these techniques are invented first and the applications come later. And the criticism of the current round of proofs is that while they may be true - likely for some, questionable for others - they're not adding new techniques or insights.

            • xeonmc · 5 hours ago

              The deluge of maybe-proofs have the same problem as the Library of Babel.

            • nbaksalyar · 5 hours ago

              > often these techniques are invented first and the applications come later There's a great paper from Abraham Flexner on this topic: https://worrydream.com/refs/Flexner_1939_-_The_Usefulness_of... It argues exactly that we should be allowed to pursue the seemingly "useless" knowledge. Previously discussed on HN: https://hn.algolia.com/?q=usefulness+of+useless+knowledge

          • IanCal · 6 hours ago

            There is a risk of this particularly if it's seen as advertising - at some point "ai model solves hard to explain problem" isn't going to be news and that benefit goes. However, there's some of this that's a proxy - the compute to solve these problems was very low (they claim a few hours of thinking time on a regular subscription). The large cost would have been the training and if training the models to be better at these things makes them smarter for useful tasks that's beneficial. I believe there was work done earlier on around showing that training the models on code made them better at broader reasoning tasks (not just writing the code itself). Another side is that if one goal is to improve the models themselves, their ability to work on mathsy problems must be high. That has very direct business value, and ideological value depending on what you think the motivations of the people running the companies are.

          • dannyw · 5 hours ago

            Why do you think this has no business value? It would be absolutely wasteful for OpenAI to not be doing this as part of a post-training RL rollout. There are architectural advancements yes, but lots of progress from LLMs really come from (1) better pre-training [generally through more cleaned data, and ofc more data], and (2) lots and lots of post-training. It's how we get more and more intelligent models for the same param sizes. The 'marketing' is just a useful side effect they get from their RL rollouts on maths and LEAN.

      • techpression · 6 hours ago

        I think there’s a venue where they start focusing on introducing hypotheses where the model currently can’t solve it, or maybe this is already happening? Being able to present useful novel ideas would likely generate a lot of press, for a while. I don’t know how this would look since I’m useless at math, but Im sure there are plenty of unknown problems with massive implications, that once formulated can be solved.

      • squidbeak · 6 hours ago

        AlphaGo and AlphaZero weren't generalized models. Math capability will presumably keep improving along with the other general capabilities, even if there wasn't a special RL focus for math itself.

      • curt15 · 5 hours ago

        The amount of money they're lately ploughing into proving math theorems is inconsistent with how societies and markets have priced pure mathematics. The entire US federal budget for math research is something like $100M annually. A single college football coach can already earn 10 percent of that. Pretty much the only enterprise that historically pays some mathematicians handsomely is quant finance, but those people are actually compensated not for proving theorems but rather for statistical modeling and programming skills. And even that industry is so technologically driven these days that pure research mathematicians no longer hold a clear edge over strong programmers with undergrad level probability and statistics at their fingertips.

        • jeremyjh · 5 hours ago

          As long as they continue making headlines they will continue spending. This is just marketing at this point.

        • dannyw · 5 hours ago

          Math is one of the most verifiable domains, esp thanks to LEAN, which also build coding skills. The $$$ they're pouring isn't just for marketing. Think of these papers/results more as "useful side effects" from large-scale RL rollouts and post-training. Every token being generated contributes to post-training in some way. There isn't a hard boundary between "training" or "inference", modern post-training is arguably inference-bound :)

          • bonoboTP · 4 hours ago

            It was shown quite some time ago that training LLMs on programming tasks improves their logical reasoning skills also in other natural language domains. So I could see math also being a training gym for AI even if the final use case is not directly math-related. Having to solve math problems efficiently can build in skills that come handy in all kinds of more everyday tasks or science and engineering.

            • FuckButtons · 2 hours ago

              It’s also very useful signal that the reasoning trace is leading to solving open problems - you can be certain that you’re not landing somewhere inside the training data.

          • sanderjd · 3 hours ago

            > There isn't a hard boundary between "training" or "inference", modern post-training is arguably inference-bound :) Ah this is an enlightening point. 8 hadn't thought about it this way, but you're right.

        • ogogmad · 5 hours ago

          Don't hire a straight-A student, unless it's to take exams; or a professor, unless it's to write papers. -- Nassim Taleb How interesting that Anthropic and OpenAI are full of professors and straight-A students!

          • jebarker · 4 hours ago

            It’s unclear to me what point you’re making here - can you elaborate?

            • conception · 2 hours ago

              Solving test questions well doesn’t necessarily translate into productive outcomes in the real world.

              • jimbokun · 1 hour ago

                So education is useless?

                • skydhash · 36 minutes ago

                  No, more like the average can be more important than some outlier's impact, at least for steady progress. So while you do need the Mozart and Einstein, a more productive effort would be to raise the general population's education level.

            • WarmWash · 1 hour ago

              If you track the best students futures and look the best workers pasts, there is not nearly as much overlap as society generally believes.

            • ogogmad · 1 hour ago

              A lot of people have been surprised by the interest that AI companies have in solving maths problems without obvious applications, as well as in the financially precarious state of these companies. I imagine that they thought that these companies were helmed by mere bean-counters, like at Boeing. But no, I reckon that they're still academics at heart, and so they work on Navier-Stokes and L=BPL, instead of on how to maximise value for their (soon to be) shareholders.

          • 01284a7e · 3 hours ago

            Don't quote Nassim Taleb, unless it's to be an arrogant dick. -- Me

            • sanderjd · 3 hours ago

              -- Michael Scott

      • mohamedkoubaa · 4 hours ago

        They already proved the point

      • singularity2001 · 3 hours ago

        Interesting path forwards, and probably partly true, but there are some important distinctions: Go was a specialized application. All the math results come as a side effect of reading the whole internet, and it will keep reading the whole internet. It will keep practicing thinking questions. Actually, math might be one of the best ways to keep them contemplating and measure their contemplation abilities, so math will always stay in the loop. Also, math might not be useful just for humanity, but also for AI, so the system might actively benefit from new math results itself. (Not sure if any of the recent proofs qualify, but future work might.)

      • sanderjd · 3 hours ago

        I think this is an interesting and good theory. They've probably eked out the large majority of the PR benefit at this point, so whether they continue in this vein will tell us a lot about their motivations for this work. To take this to the next step, what happened after deep mind pretty much solved Go is that they started looking for the next set of things that hadn't been done yet. It does strike me as very likely that this will follow that same path.