HN in RSCserver-reason-react
top.mdnew.mdbest.mdask.mdshow.mdjobs.md
← Back to stories

The Mathocalypse

344 pointsby 6bitquant 21 hours ago361 comments

Discussion

Loading discussion
  • matt3210 · 21 hours ago

    Agents basically did statistically guided brutforcing. There is no value in what they produced because it lead to no understanding of anything and most likely will hurt the field IMO

    • dekhn · 20 hours ago

      That is not a correct description of what the AI did.

    • woah · 20 hours ago

      Evolution did statistically guided brute forcing. Doesn't mean that biology has no value

  • nostrademons · 21 hours ago

    As a side note, you can tell this wasn't written by an AI by the first sentence: > mommy, I heard you got cooked! I heard that a robot solved the math problem you worked on for your whole career! OOF! My 8yo talks exactly like that. I could totally imagine him saying this, the same way, at the dining room table. I asked ChatGPT "pretend you're an 8/9 year old today. how would you insult your mom about having her job be replaced by an AI?", and the responses it offered were: > “Mom, AI took your job because apparently even robots were like, ‘Yeah… we can do this better.’” > “Mom, congratulations! You got replaced by a computer. Even Siri has a job now and you don’t!” > “Mom, AI took your job? Dang. I guess even a robot looked at your work and said, ‘I got this.’” > “Don’t worry, Mom. You can still be useful… like teaching the AI how to make my lunch.” All of these seem to have a vaguely Millennial flavor, aside from being pretty awkward and mechanical roasts. Trust the children and linguistic drift to be the best AI detector.

    • john_strinlai · 20 hours ago

      i added "use current trendy lingo" and the results were a bit less mechanical sounding. one of them included "cooked", another had "negative aura". i used the free google ai: (deleted the examples... but they were vaguely close to what i hear my grandkids say.) edit: neat, insta-flagged despite hundreds of non-ai comments that have never been flagged. i would have thought that hn would use some heuristics in their ai detection but i suppose not.

      • djhn · 11 hours ago

        I appreciate the counter argument, even as a chat bot skeptic. However, there is the fact that you had to know how to poke the machine so it plucks the right vocabulary out of the training data. There is very clearly no theory of mind: no inherent internalised modelling of how a human of a particular age thinks, speaks and what they do or do not know. These are the most obvious cracks in the “LLMs are (or will be) the superintelligence” narrative.

        • esafak · 2 hours ago

          How hard is it to add that? We have plenty of writing from people of all ages to draw a model from. They just have not got round to it.

          • recursive · 1 hour ago

            It depends who's doing the prompting.

    • dist-epoch · 17 hours ago

      The more obvious problem is that those jokes, even if terrible, are too advanced for an 8yo to make. Reminds me of the memes with the little girl making astute comments about the patriarchy to her father.

      • shoobiedoo · 17 hours ago

        Mother dearest, it pains my effusive tendencies to inform you the untimely death of your life's work by the hands of a mere transistor-- moreover, my incontinence briefs have reached their capacity.

      • dodger-dog · 2 hours ago

        Yeah, it takes a different bs detector depending on the context.

      • nostrademons · 22 minutes ago

        Eh I don't think the content is out of line with what an 8yo understands. Mine tells stories about bombs flying in the Russia/Ukraine war, which required a lesson about tact because some of his classmates are both Russian and Ukrainian refugees and this subject may be a little close to home (literally, his best friend's grandma lives in Ukraine right now).

    • stratos123 · 8 hours ago

      I'm not sure why one'd expect Scott Aaronson's posts to be AI-written, but sure enough, Pangram detects 0% of this post as AI.

      • matwood · 37 minutes ago

        It’s funny when people who are skeptical AI use an AI checker that is also AI and believe its results 100%.

    • matwood · 39 minutes ago

      Apparently the insult now is to call someone AI when they are being boring or bland. Kids have a surprising ability to pick up on things.

  • GMoromisato · 21 hours ago

    I liked the metaphor of a climber teleported to the top of a fog shrouded mountain. And I agree that now that the teleporter exists, we need to use it to reach more peaks and explore. There's no going back to a world where AI doesn't exist.

    • lumost · 21 hours ago

      The issue is ownership, we have no means of distributing the knowledge from the AI or rewarding those who could help. We are quickly moving to a world where all symbolic and numeric reasoning for economic purposes is performed by AI.

      • GMoromisato · 20 hours ago

        Agreed! Specifically, compensation (monetary and reputational) for professional mathematicians was bundled into theorem proving--essentially, climbing the mountain. Now that a teleporter exists, we need to unbundle compensation. I don't know what that means in practical terms, but I agree that's the issue.

        • pfdietz · 6 hours ago

          That's the issue, but they can't come out and just say it, because that horse left the barn long ago for other professions. It's socially just fine to obsolete entire lines of work, and has been for centuries.

  • tkdb · 21 hours ago

    C'mon. Mathpocalypse. Things are hard enough already.

    • Nition · 19 hours ago

      Just be glad you're American, because Mathsocalypse and Mathspocalypse work even less.

  • TMWNN · 21 hours ago

    Quoting DCKP < https://news.ycombinator.com/item?id=49989738 >: >I have had this conversation with my PhD students yesterday. I am 100% sure that all of their problems can be solved by publicly-available models now (I solved a case of one myself as a test, it took 15 minutes). So the challenge for them is to see how much they can accomplish in their allotted period, and still pass a defence on at the end of it all. The PhD defence is going to become all about a test of understanding, not a test of quantity of publication. Also, Ted Chiang's 2000 short story "Catching crumbs from the table" < https://np.reddit.com/r/singularity/comments/1wzu5gf/this_mi... >.

    • cgio · 18 hours ago

      Maybe they should target more ambitious results now that they can solve their other results in 15 minutes? This spirit of racing in academia is insane. Get in before the door closes.

      • mehrzad · 16 hours ago

        While the developments are certainly difficult for new grad students, it should still be possible to do what many have done before: find an obscure enough problem that it’s extremely unlikely you get sniped. Worry about big problems after finishing the PhD.

        • cgio · 14 hours ago

          Genuine question. How many big problems were solved by people who took the safe path for PhD? I would intuit there’s high correlation of unsafe PhD and big problem solvers.

  • PowerElectronix · 21 hours ago

    What's with all the "AI just proved that this or that isn't O(n (log (n))^2) but akshually O(n (log (n))^1.99999)"?? I guess it deserves respect as progress, but it just rubs me the wrong way. Like the machine did the absolute minimum to beat the previous mark.

    • para_parolu · 20 hours ago

      You just run it again and again and again

    • bryan0 · 20 hours ago

      Often times the constant (2 in this example) is a conjectured minimum, so anything below that is a noteworthy result. Think of it as breaking through some theoretical limit.

    • mswphd · 20 hours ago

      for say FFT/integer multiplication or 3SUM, we have natural algorithms that have existed a long time with a given complexity (O(n \log n) and O(n^2), respectively). Given how long these natural algorithms have been the best algorithms we have, it is natural to conjecture they are optimal. Showing an O(n(\log n)^{.99999}) algorithm exists shows that these optimality conjectures are false. Now, there are some critiques you can have of this. Namely, it is possible that these novel algorithms have significant trade-offs that make them almost never worthwhile in practice. "Fast" matrix multiplication algorithms are typically of this form. So perhaps this all points towards a deficiency in big O notation, which can be deceptive. But, for people who care about optimizing asymptotic complexity, it is still interesting.

    • JohnKemeny · 20 hours ago

      Many people thought it could never be less than 2. They proved that it can. What is the true value? Nobody knows, now.

    • zem · 20 hours ago

      to get some intuition about why this is such a big deal, look up the history of strassen's algorithm, which solved matrix multiplication in less than O(n^3). this was a truly stunning result because it seemed intuitively obvious that the output matrix had n^2 cells each of which was calculated via an independent O(n) loop over a row/column of the input matrices, so how could you do better than n^3. but once strassen proved that you could do some clever tricks and reduce the overall time to something less than O(n^3) it started an entire cottage industry of people getting better and better algorithmic bounds. the initial breakthrough was a qualitative one, independent of how much it improved things in numerical terms. https://hideoushumpbackfreak.com/algorithms/algorithms-stras...

    • tmvphil · 19 hours ago

      Tell that to the humans working on matrix multiplication who spent years of their lives getting it from n^2.3728596 to n^2.371866, only for openai to blow it away at n^2.25

    • lg5689 · 7 hours ago

      This is a legitimate question, and it'll be interesting to see how these constants evolve in the future. In some cases, maybe it's true the AI did the minimum to break the barrier, and the best constants are much lower. It's still a notable result that the barrier was broken. But it's quite possible that there's not actually much more to improve.

    • pfdietz · 6 hours ago

      I think it shows even theorem proving oracles can have a sense of humor. I know, I know. Don't anthropomorphize AIs. They hate it when you do that.

  • zkmon · 20 hours ago

    The irony. Something that is born out of a science, eats up that science.

    • karmakurtisaani · 7 hours ago

      "There's purpose of science is to make boring what was once exciting." -Boaz Barak

      • zkmon · 3 hours ago

        That's quite insightful.

    • pfdietz · 6 hours ago

      "Civilization advances by extending the number of important operations which we can perform without thinking of them." ― Alfred North Whitehead, "An Introduction to Mathematics" (1911)

  • ks2048 · 20 hours ago

    > It feels like something written by someone who’s on psychedelics. So much unclear and doesn’t make sense. Lots of name dropping of previous work without discussing why it can be used despite impossibility results > Basically the paper is so horribly written that it’s impossible to read it without AI help That's interesting and haven't seen this in all the coverage of this event. It sounds horrible to wade through - like trying to understand someone else's messy code that still produces the correct output.

    • piker · 20 hours ago

      It also aligns with the fear that these proofs present a risk to the ecosystem by out-competing attempts at more human-readable proofs. Perhaps though we end up with more math influencers who edit and annotate these proofs to bring them back to us.

      • bobajeff · 20 hours ago

        I think that's ultimately a good thing. As proofs weren't supposed to be the point as stated by William Thurston long ago. Maybe now the focus can be more on better explanations and creating tools for growing understanding and intuition.

        • cowlevel · 19 hours ago

          Good explanations should take the form of human-understandable proofs.

          • btilly · 19 hours ago

            Define "human understandable". It's worthy of note that most humans, do not find most mathematicians understandable. As is frequently demonstrated in Calculus classes. Therefore it is arguable that even human produced results are not generally human understandable.

            • cowlevel · 18 hours ago

              Then they are not good explanations.

              • notpachet · 16 hours ago

                "What I cannot create, I do not understand" - Feynman

      • rrr_oh_man · 20 hours ago

        Vibe mathing

      • whatshisface · 20 hours ago

        The ecosystem is (ahem) gated by hiring committees. There is no risk of AI replacement from the inside. "Replacement" is not even a possible movement. The funding for mathematics worldwide comes mostly from endowments, which are investment pools.

    • aaroninsf · 20 hours ago

      Serious question: Why would anyone believe this (also) is not simply example N+1 of this is the worst it will ever be , as opposed to recognizing this as what will almost certainly prove to be an awkward moment, soon to be replaced by another order of magnitude of cleaner, clearer, more intelligible, etc.? Ximm's Law: every critique of AI assumes to some degree that contemporary implementations will not, or cannot, be improved upon.

      • devin · 20 hours ago

        Devin's Law: every defense of AI which rests on "it will get better, trust me" is in many ways indistinguishable from 2010s crypto hype or "level 5 self driving is right around the corner"

        • usrnm · 19 hours ago

          1) Predicting the future is hard, but so far everyone who was saying that it would get better turned out to be right. It is getting better 2) Waymo exists

          • devin · 18 hours ago

            3) That doesn't mean flying cars will within your lifetime I don't think anyone is saying it can't or won't get better, but the question is how much better , on what timescale , and are there fundamental parts of the problem which will remain extraordinarily difficult to improve? The comment I was responding to suggested a guarantee of an "order of magnitude" jump right around the corner. There is no guarantee of this, and if you view doomers as fools for having doubts, then we ought to look upon the folks who are sure of this sort of progress in the same way.

            • kulahan · 18 hours ago

              Your original comment was about self-driving cars, not flying ones. If you wanted unattainable goalposts you should’ve started with that, not ended with it.

              • devin · 18 hours ago

                Waymo does not claim level 5 self-driving, and you don't have a level 4 in your driveway, so there was no moving of goalposts.

                • kulahan · 18 hours ago

                  Then say that, instead of bringing up flying cars, which is moving them, explicitly.

                  • devin · 17 hours ago

                    Respectfully, I disagree with the way you're characterizing my comment. I was demonstrating that just because you have a level 4 Waymo doesn't mean flying cars are right around the corner. This is again a reference to the original comment I was replying to, the one that suggested of course we're going to get an order of magnitude improvement.

      • Kotlopou · 19 hours ago

        In that case, one would expect to see some progress in this direction, but AFAICT that hasn't shown up yet? If anything, it's getting worse, though that could just be the increasing scale and decreasing cleanup efforts. Already the unit distance proof was substantially human-edited (per Thomas Bloom). Then with the ten problems from Astra you started getting the citation issues. Then Navier-Stokes was a rushed 160 pages with barely any citations, and some of the related papers were called (by their "authors") the ugliest mess they've ever seen. And now here we are. At least it seems that mathematical ability and communication with a mathematical audience are independent skills, and progress in the first does not imply the second. This doesn't surprise me much, given two analogies: 1) many smart people are nonetheless horrible lecturers. (You can't quite get the opposite extreme, since to explain math well you have to be able to do it.) 2) AI writing in general hasn't improved. The models have annoying verbal tics ("honestly") and have no sense of which part of what they say is obvious and which is relevant.

        • auggierose · 19 hours ago

          You can get the opposite extreme quite often as well, I'd think. How many really good lecturers have never proven a new important result?

          • Kotlopou · 19 hours ago

            I'm thinking of somebody like Grant Sanderson (3blue1brown), doing pure exposition extremely well. For that you at least need to be able to work through examples, or to present why an intuitive approach might fail, and these things can be little theorems themselves. It doesn't have to be publishable in the current culture of novel results, but you do need a lot of competence with the tools.

            • auggierose · 15 hours ago

              Yes, Grant Sanderson might be a great example. I am not doubting the competence of the "good lecturers" I was referring to. But that is different from being able to introduce the big new concepts that give the big new results.

    • TheOtherHobbes · 20 hours ago

      Math proofs need to produce the correct output correctly, which is not quite the same thing. This looks like an AI IPO PR powerplay, because at this point the proofs haven't been checked and it may not be possible for a human to check them - because proofs should be clear, not horribly written and noisy. The noise is suspicious because it's the difference between brute forcing and cognition. A human proof won't just be logically correct, it will be cognitively distilled and coherent. It may still take years to understand it, but the logical flow will be straightforward, not obfuscated. You want the path through the maze to be as short as possible and the map to be as clear as possible. This sounds like the opposite. There may be a genuine path through the maze, but if it's too convoluted and takes too long it will be impossible to confirm. I think the next step is to demand that proofs either be human-scale or they prove that a human-scale proof is impossible and the machine proof is as good as it gets. I suspect that's possible without tripping over the halting problem. (But I can't prove it.)

      • pizza234 · 19 hours ago

        The post says there's a Lean certificate for this and other proofs ("some [...] not all of them"). > This looks like an AI IPO PR powerplay, Interestingly, the post has actually also an argument for this: > Experience has shown that, even now, there will still be people explaining in patronizing tones why none of this is real and none of it counts. If such people were capable of being impressed by anything that happens in the empirical world, of updating on anything, they would’ve already been impressed and already updated several years ago, long before things had reached the point of an actual Mathocalypse. > So, they’ll say, maybe the alleged solutions are not solutions at all, but just “AI slop.”

        • smcg · 19 hours ago

          It's on OpenAI and Anthropic to prove that they obtained these results legitimately and credited all researchers who deserve credit. They do not get the benefit of the doubt.

          • Kotlopou · 19 hours ago

            But if you think they got them illegitimately, then how did they get them? And why are mathematicians reacting to this as a sudden explosion of new results that have resisted sustained effort? Where is the sudden productivity rise coming from?

            • za_creature · 18 hours ago

              I'd say it comes from the same mathematicians that were strongly encouraged to use the machine to solve their problems for the last 2 years or so. There's clear benefit in a babelfish that can coordinate disparate efforts, the only problem with the current iteration is giving credit to said efforts. Google went quite far down the road to hell, but stopped short of taking credit for websites' content since the company understood that poisoning the well only goes so far. At this point, one can safely conclude that _Chat_GPT was an intentional attempt to squeeze out more data once they mined the internet dry.

            • pfdietz · 7 hours ago

              > But if you think they got them illegitimately, then how did they get them? Obviously they stole them from the Proof Fairy. It's grapes so sour they could etch metal.

          • pizza234 · 18 hours ago

            Have you actually read the article? It's been actually written, among the other things, because the author's wife has been trying to solve one of the problems for her whole life.

      • sebzim4500 · 19 hours ago

        Surely by the time of the IPO we will know whether the main results are correct, if only because a different AI will have produced a lean proof or found a logical flaw (the second case would be hard to verify but probably not impossible). Also from what I can tell from the few fields I understand, the proofs aren't that long or complicated they are just terribly written.

        • curt15 · 19 hours ago

          Why should that make material difference to the IPO? What is the economic value of those results? The entire US federal budget for math research is something like $100M annually. And mathematicians in other countries are hardly making bank either. How does one reconcile how the market has historically valued mathematics with the cash-strapped frontier labs ploughing so much money into that enterprise?

          • runarberg · 19 hours ago

            The market works in mysterious ways. What companies do for marketing is often irrational, what companies do to attract investors is likewise often irrational, and why investors invest in companies is also often irrational. Why should that make a material difference to the IPO? Because of the vibes, and investors are indeed all about the vibes.

            • jryle70 · 14 hours ago

              It works in mysterious ways, but you know exactly it will behave certain way "Because of the vibes, and investors are indeed all about the vibes."?

              • runarberg · 3 hours ago

                The first paragraph is describing a general trend over multiple events across multiple agents. Between zero and three of these can be true for any transaction across every transaction. The second paragraph is specific to OpenAIs behavior.

            • falcor84 · 8 hours ago

              > what companies do to attract investors is likewise often irrational, and why investors invest in companies is also often irrational. Note that the cool thing about rationality is that it does not depend on transitivity. If what companies do to attract investors works, then it is in fact rational of them, regardless of the rationally of investors.

          • coderenegade · 17 hours ago

            They're burying any doubt that the models are capable of superhuman performance on intellectual tasks. Neural nets aren't calculators, and were notably poor at mathematical reasoning tasks for a long time. Now they're not, and the labs are proving that by chewing through what would ordinarily be decades of progress in a month. And the reason to go for math in particular is because there's no wiggle room. You can't just dismiss it as hallucination. If the models can do this, they're almost certainly good at just about everything, because the reasoning and creativity required to solve these problems will translate. And even if they were only good at this stuff, that's still a tremendously valuable thing, because quantitative reasoning and analysis is the bedrock for many, many industries. oAI is gunning for the largest IPO in history at this point, and they might actually get there.

            • curt15 · 16 hours ago

              > If the models can do this, they're almost certainly good at just about everything, because the reasoning and creativity required to solve these problems will translate. The "then" in your "if-then" bears a heavy load. Why would society assign so little economic value to pure mathematics if the skills for proving math theorems translate to massive value in "just about everything"? Would you expect top mathematicians to cure cancer if you transplanted them from the math department to a medical research lab?

              • ndriscoll · 14 hours ago

                Do you think mathematicians are not already working on cancer research? There's quite a bit of heavy math in biostatistics, medical imaging, machine learning, etc. I'd assume that the majority of people who study math take their skills and move onto some related STEM career that isn't pure math. Academia is incredibly small and competitive.

              • coderenegade · 11 hours ago

                Because research mathematics is a subset of all quantitative work that gets done, but it's by far the most technically difficult subset. If the models can handle research grade math and produce ironclad proofs, they can probably handle the quantitative side of just about any discipline in a trustworthy fashion. Think about how many dinky spreadsheets have gone on to become critical tooling for large organizations. Even if you consider that many disciplines hide technically demanding work behind tooling (e.g. essentially no one is writing a stiffness matrix FE routine by hand), this model would be capable of writing a direct competitor from scratch to produce the same result. All of STEM relies on mathematical analysis, and new models are now superhuman at that. And yeah, I'd go a step further and say that the reasoning and creativity required to solve cutting edge math problems probably does translate to other tasks like interpretation of the law, or medical diagnosis, or accounting, etc., for the same reasons that I think most top tier mathematicians would excel at those tasks were they so inclined.

              • throwaway81523 · 11 hours ago

                That sort of worked for Eric Lander but as he moved higher up in the bureaucracy, he found himself suddenly having to manage people who weren't nerds like him. He was bad at that had to step down.

              • fragmede · 11 hours ago

                American mathematician Jim Simons was worth some $31 billion at the time of his death in 2024. The lack of economic value in pure math doesn't mean that applied math is of little value. Physicists and mathematicians with PhDs "sell out" to join Wall Street as quants and make a killing there, Jane Street is full of them.

            • lesostep · 3 hours ago

              > If the models can do this, they're almost certainly good I would argue that "good at math" was a short hand for "good at X" because mathematicians were historically good at engaging with very complex ideas, distilling them and coming up with precise and concise answers they could validate by themselves. Given that AI solutions are described as "psychedelical" and they rely on the outside source to validate the result, I don't think the same logic could apply to them.

      • Octoth0rpe · 19 hours ago

        > A human proof won't just be logically correct, it will be cognitively distilled and coherent. It may still take years to understand it, but the logical flow will be straightforward, not obfuscated. https://en.wikipedia.org/wiki/Inter-universal_Teichmüller_th... seems like a counterpoint, but IANAM. (I am likely cherrypicking the far end of the bell curve re: straightforward here)

        • IsTom · 18 hours ago

          Isn't this controversial, to say the least?

        • ffaccount2 · 18 hours ago

          Not a counterpoint, actually case in point, because: >Mochizuki and a few other mathematicians claim that the theory indeed yields such a proof but this has so far not been accepted by the mathematical community. Proof can't be understood, proof doesn't matter.

          • pfdietz · 7 hours ago

            Oh the proof can be understood quite well: it is understood to be wrong.

        • dist-epoch · 18 hours ago

          I wait for AI to say something about this :) Someone at OpenAI, please, work on this.

      • caaqil · 19 hours ago

        We should consider the possibility that at some abstraction levels, we can safely stop chasing "clarity" or "coherence" which is circularly defined in such a way that it's capped by human processing power. Developers and people in CS in general seem to have gotten used to the idea that most productive SWEs don't need to exactly know how to produce assembly or trace every branch prediction or even most of the optimization the CPU (or even their compiler) is running. Mathematicians will get there.

        • jltsiren · 19 hours ago

          CS got that idea from mathematics. Theorems (with the definitions required to state them) are supposed to be self-contained units. Once the general consensus is that a theorem has been proven correct, people can use it without understanding the proof. Of course, people still want to understand how things work, and it often makes sense to understand them a couple of layers below the one you usually work at. But at some point, you should stop distracting yourself with irrelevant details and focus on your actual work.

          • yorwba · 18 hours ago

            Thing is that many of the theorems here are not useful work in and of themselves, but were posed as research problems because it wasn't clear how they could be resolved with current techniques, implying that the process of trying to find a proof might result in new techniques. It's those new techniques that are the actual goal, but if they can't be easily extracted because the proof isn't structured to enable this, that's a bit of a headache.

        • TheOtherHobbes · 15 hours ago

          I've very aware that human cognition has limits. But it doesn't follow that these proofs - or any proofs - are automatically on the far side of that limit. The human usefulness of a proof depends entirely on its human legibility. Much of the value of proofs is in inventing new techniques and concepts and having new insights into relationships. Occasionally you get some game changing insight into practical physics or engineering. But that's rare. Without that, proving or disproving a conjecture is an excuse for new and original thinking. Compilers are not the same problem. The point of code is to produce reliable-ish consequences from various possible inputs. It's not a creative exercise in logical consistency, which is what maths proofs are, ultimately.

        • Tanjreeve · 10 hours ago

          > Developers and people in CS in general seem to have gotten used to the idea that most productive SWEs don't need to exactly know how to produce assembly or trace every branch prediction or even most of the optimization the CPU (or even their compiler) is running. That's because there's a bunch of people who work on that stuff. Just because web developers don't care about it doesn't mean it doesn't exist. Utter magical thinking.

        • amoss · 9 hours ago

          You seem to be confusing your quantifiers a little. (Not all SWEs must understand all CS) does not imply (there can exist parts of CS that no SWE understands).

      • eadler · 19 hours ago

        That reminds me of this paper: Chow, T. Y. (2008). A beginner’s guide to forcing (arXiv:0712.1320). arXiv. https://doi.org/10.48550/arXiv.0712.1320 > “All mathematicians are familiar with the concept of an open research problem. I propose the less familiar concept of an open exposition problem. Solving an open exposition problem means explaining a mathematical subject in a way that renders it totally perspicuous. Every step should be motivated and clear; ideally, students should feel that they could have arrived at the results themselves. The proofs should be “natural” in Donald Newman’s sense [13]: > This term . . . is introduced to mean not having any ad hoc constructions or brilliancies. A “natural” proof, then, is one which proves itself, one available to the “common mathematician in the streets.””

        • oliculipolicula · 6 hours ago

          This, but for a generalised notion of "number" (that is, "proof") https://en.wikipedia.org/wiki/Nothing-up-my-sleeve_number ?

      • FloorEgg · 19 hours ago

        If intelligence is compression, and these models are a different form of lesser intelligence than human, but being scaled up to brute force problems, then it makes sense the artifacts that produce (the proofs) would have worse compression than a human proof would. In other domains I have seen first hand overwhelming evidence of how things that cause the AI to make mistakes also cause humans to make the same mistakes. I wonder if the proofs being produced that are hard for humans to interpret are also hard for other LLMs to interpret. In other words, I wonder if humans are still much better at compressing understanding into proofs than the best LLMs, and what it will take for LLMs to exceed them. It kind of an explicit example of how the LLMs can be materially less intelligent than people, but still be more productive through scaling, and yet they also can't replace people because they are a categorically different kind of intelligence. It's like all the AI debates compressed into one example showing countwr-intuitive answers.

        • slopinthebag · 19 hours ago

          idk if i'd even say they're "lesser", just very different. so they look like gods/babies depending on what they're doing because we anthropomorphise them.

          • FloorEgg · 17 hours ago

            In terms of synapse density, neuron diversity, energy efficiency, memory access, etc. they are orders of magnitude lesser. My mental model - for better or worse - is that intelligence has both shape and area, and LLMs are orders of magnitude smaller area but very different shape, and they have more intelligence area in the kind that humans have lesser of. So yes very different, more in some material ways and lesser in others, but in total intelligence are still orders of magnitude lesser. My gp comment was acknowledging that when you scale up many instances / brute force problems it confuses that "total area" claim a bit. To follow the anthropomorphization... 1000 toddlers may have more total intelligence than a grown man, but does that matter? The problem with these discussions probably/usually fold into differing/loose definitions of intelligence.

            • slopinthebag · 17 hours ago

              ok yeah then i think we're in total agreement perhaps the chat-based ux has sort of fooled us into comparing these things to human intellegence. we don't really do this with chess, or other forms of ai, nor computers at large.

        • esafak · 19 hours ago

          This is just the first cut. I have no doubt that they will polish their proofs over time.

          • CogDisco · 14 hours ago

            I don't see how they have any incentive to do so.

            • esafak · 13 hours ago

              I communicated loosely; I meant AI users in general, not OpenAI specifically.

          • mfld · 1 hour ago

            My hunch is that improving those proofs specifically might be a similar endeavour as improving a large vibe coded app.

            • esafak · 1 hour ago

              It's a search problem; finding the shortest and most elegant path through the space of arguments.

        • perching_aix · 18 hours ago

          > I wonder if humans are still much better at compressing understanding into proofs than the best LLMs, and what it will take for LLMs to exceed them. Isn't it fairly established that (generally [0]) manually written / optimized skill files perform a lot better than generated ones? Meaning that yes, this likely does hold. [0] or to be specific, that the pecking order is: ai generated < human co/written < hyperoptimized for the specific model via some convergence process

      • ComplexSystems · 19 hours ago

        > I think the next step is to demand that proofs either be human-scale or they prove that a human-scale proof is impossible and the machine proof is as good as it gets. Who do we demand this from? The AI companies? Or the mathematicians who are worried they will have nothing left to do?

        • za_creature · 19 hours ago

          From the entity that is producing these proofs, obviously. As the old saying: great claims require great evidence.

          • esafak · 19 hours ago

            Ask away. They've dropped the mic, as far as they're concerned; they're not going to worry about what you do with it, or if you don't understand it.

            • za_creature · 19 hours ago

              That's the best definition of slop I've ever read.

            • pfdietz · 7 hours ago

              They still have the mic; I've been told there are another two such proof drops coming.

      • kens · 18 hours ago

        > I think the next step is to demand that proofs either be human-scale or they prove that a human-scale proof is impossible and the machine proof is as good as it gets. In 1976, the proof of the Four Color Theorem was controversial because it was done with a computer examining over 1000 cases by brute force and was essentially not comprehensible by humans. But mathematicians ended up accepting it. So mathematics has a 50-year precedent of not requiring human-scale proofs. How is the current situation different? (Disclaimer: Apologies if this sounds dismissive or argumentative. I genuinely think that the Four Color Theorem should play a role in these discussions and suspect that many people are unaware of the controversy over it.)

        • pizza234 · 18 hours ago

          There's also another (that I find more concerning) aspect to it. As AIs become smarter and smarter, there will be no amount of clarity that will make more complex proofs understandable to humans - this is an inevitable effect of the cognitive capacity gap. Complaining about bad style can make some sense now (I disagree anyway), but it's an argument that will be dead shortly.

          • dist-epoch · 18 hours ago

            Arguably no human understands 100% of how a smartphone is produced, and, it doesn't matter? Maybe no human will fully understand a future proof, but they could fully understand a little piece of it. And many humans in aggregate could understand it, each with their own little piece.

            • Tanjreeve · 10 hours ago

              The point isn't that 1 person can understand an entire smartphone silicon up. It's that 100% of a smartphone can be understood by people. A maths theorem that can only be partially understood/proven only has value as marketing collateral.

              • mrob · 8 hours ago

                >A maths theorem that can only be partially understood/proven only has value as marketing collateral. There are many problems that aren't "interesting" in a mathematical sense but can still have valuable proofs. The entire field of formal methods in software development is mostly about this kind of problem. If I want to be sure that no input can cause out-of-bounds memory access, or that my superoptimizer found the lowest possible cycle count, I don't care if it's elegant, I just need some machine-readable proof I can run through my trusted verifier.

                • Tanjreeve · 14 minutes ago

                  When we use verification to build software. We have an actual software as an output and hopefully some sort of problem solvsd the verification is also not in an of itself valuable.

              • jcelerier · 4 hours ago

                For one person that understands a proof of some matrix multiplication algorithm lowering its current complexity bound in an useful way (e.g. not with absurdly gigantic <= 2nd order polynomial term), there's likely thousands of companies willing to pay for the improved efficiency

                • Tanjreeve · 17 minutes ago

                  And if anything that useful does come out of this output/"slop dump" then that might begin to make a strong argument for it. But every other time it was "that's the theorem done for the headline figuring out everything important and how to use it is an exercise for the reader".

              • losvedir · 1 hour ago

                > A maths theorem that can only be partially understood/proven only has value as marketing collateral. If this is true, then what's really the point of math? A lot of math is actually useful. A proof regarding cryptography, for example, would have practical application even if not understandable by humans.

                • Tanjreeve · 21 minutes ago

                  Cryptography is understandable by humans so I'm not sure what the point is. You do understand that the argument isn't everything has to be understandable by every human right just to check?

            • ben_w · 5 hours ago

              While any logical proof written in lean can be broken down into things a human can follow, there is no requirement that the whole proof of something is ever simple enough that all humans combined could follow it. Given the LLMs are struggling to explain themselves clearly, but also that these explanations cleared up somewhat by having an expert using one to query some of this research, but also the LLMs can solve problems faster than humans can read the proofs, it is possible for this category to have both examples of things that can be rendered in a human-comprehensible form, and also examples where it is not . As an analogy: any human can check any single arithmetical calculation from a computer, that's not even particularly difficult. But a Raspberry Pi Zero can do those arithmetical calculations so fast that even if literally every single human was trained to do this at the level of the current world record holder, humanity as a whole could not keep up.

          • koliber · 7 hours ago

            Exactly. We need "observability tools" that allow us to keep a handle on things we can't see or comprehend innately. We have such tools for a large number of events that we're incapable of witnessing or understanding in "raw form". We can't possibly read through thousands of raw records of recordings of individual's heights and make sense of it. But we can use Excel to calculate the average in seconds. We can chart the results and the image is perfectly comprehensible. We can trust the average calculation and the image because we trust the mechanism for transforming the data. We have what I would call a trusted path of provenance. We don't need to nor do we want to read the raw data. Now we need things like this of the 2nd order. We need mechanisms of transforming trusted paths of provenance into things we can look at and easily verify. We can do formal verification of code that would be hard to do by hand. We need formal verification tools for formally verifying formal verification tools. Eventually we will need multiple layers of this. As long as the chain and reasoning is intact, we should be alright. It's kind of like with a horse. A horse is much stronger than us, but we can control it by pulling on two strategically connected pieces of rope.

        • housecarpenter · 7 hours ago

          I don't know much about this controversy, but looking at the Wikipedia page for the Four Color Theorem, it looks like since the initial Appel-Haken result, mathematicians have continued to work on the Four Color Theorem to try to find a simpler proof. So did they really accept it?

          • pfdietz · 7 hours ago

            Finding new simpler proofs happens in traditional manual mathematics too.

          • rhdunn · 6 hours ago

            There are a lot of proofs for Pythagoras' Theorem and other theorems. There are mappings between complex numbers and 2D matrices allowing problems to be solved in either domain. There's research into The Langlands Program looking to connect number theory and harmonic analysis. There's research into Category Theory looking to define core concepts and relate them do different fields so that results in one field can be applied to another due to equivalence.

        • pnin · 7 hours ago

          Mathematicians (mostly) accepted that the proof was valid, not that it was the kind of thing mathematics should strive for.

      • amoss · 9 hours ago

        > I think the next step is to demand that proofs either be human-scale or they prove that a human-scale proof is impossible and the machine proof is as good as it gets. > I suspect that's possible without tripping over the halting problem. (But I can't prove it.) I doubt that it is. For the language of proofs to be powerful enough to be able to express an arbitrary proof it would have be Turing Complete. Proving that a proof is the smallest proof of a given concept (i.e. there is no smaller human understandable proof) would then be proving the minimality of a program in a Turing Complete language.

      • disgruntledphd2 · 8 hours ago

        > The noise is suspicious because it's the difference between brute forcing and cognition. A human proof won't just be logically correct, it will be cognitively distilled and coherent. It may still take years to understand it, but the logical flow will be straightforward, not obfuscated. This is basically how all AI approaches appear to work. They solve the problem you give them (in many cases) but in a very over-complicated way. It's like they can't step back from the problem and realise that they need to simplify to make it work better. Nope, just keep hammering more code/proofs against the problem and eventually you'll hit the goal. RL has a lot to answer for, I guess.

      • truculent · 7 hours ago

        > the difference between brute forcing and cognition There’s not really a clear distinction between these things, in my opinion. Problem solving (and intelligence?) is a mix of search and compression. We like solutions that are elegant (high compression, simple search). But often what appears elegant to some is harder to appreciate for those without the same background knowledge or even the same amount of mental bandwidth (if you’ve ever worked with someone simply much, much smarter than you, you may intuit this!).

    • jltsiren · 20 hours ago

      Isn't that just the default experience with AI these days? In small enough scale, AI models can express their ideas clearly. But the larger and more complex the ideas are, the less suitable the outputs are for human consumption. I guess AI models think too different from humans, and nobody has trained them to communicate complex ideas in the way human experts in that particular topic expect.

    • ssfdg · 20 hours ago

      This proof dump reminds me of the glut of low-quality drive-by PRs overwhelming open-source repos.

      • dormento · 19 hours ago

        Its like infinite summer of code, but for math. Must be annoying.

    • spelunker · 19 hours ago

      I see many parallels to genAI-assisted code development. Not surprising I think.

    • m3kw9 · 19 hours ago

      why not get Astra to make it make sense?

    • acedTrex · 19 hours ago

      > Basically the paper is so horribly written that it’s impossible to read it without AI help This basically describes every single PR at work for the past year. Diffs of 10k+ paragraphs of comments saying nothing. Just rubber stamp and move on, nothing else you can do.

    • rramach · 18 hours ago

      True but the explanations will get better. Meanwhile you can use the model to help you out as Scott comments "Just now, however, Dana tells me that she’s been asking Astra all day to explain the new proof of the UGC to her and it’s been doing an amazing job and she’s starting to understand the construction."

    • smrtinsert · 18 hours ago

      Sounds like exactly the daily I deal with when agents try to create product requirements from multiple sources, except to a much much worse extent.

    • KoolKat23 · 18 hours ago

      This makes sense and looks exactly what something would look like that is smarter than us. If it is correct it is making inferences that we can't see. At a stretch even working in more dimensions than our three dimension limited brains.

    • amazingman · 14 hours ago

      Sounds a lot like using LLMs for coding ~18mo ago.

    • PatronBernard · 9 hours ago

      If it were submitted to a journal, it would be outright rejected. We should give the peer review process -however flawed it is- some credit here. Dumping hundreds of AI-generated preprint articles on the field should not be allowed.

      • D-Machine · 8 hours ago

        No. Letting some tiny handful of individuals selected by a tiny set of review-delegators review this is immeasurably worse than just releasing it publicly and letting anyone with the sufficient expertise review it. You know nothing of how peer review actually works, or have not thought about why that process would be only harmful in this case.

        • PatronBernard · 4 hours ago

          Normal peer review works by submitting the article to the appropriate journal, suggesting some preferred reviewers that have expertise in the subject, after which the rest of the process ideally is then orchestrated by an experienced editor. If a journal receives a paper that is unreadable, it is outright rejected with the comment that it should be made readable, regardless the results in that paper. What is the point of having results if you are unable to convince your audience of those results? You could just as well just sit on them and never communicate them. What is harmful about this process?

    • gorgolo · 8 hours ago

      For what it’s worth I’ve heard from several different ex mathematician colleagues that found papers in the stack that was adjacent to their work or solved a problem they were acquainted with in their career, and they all said the results are actually quite readable. Maybe there’s some variance or it depends on the reader. That said, like with a lot of other complaints about AI: human papers can be poorly written and poorly explained too. That’s always been the case. And sometimes a proof is just complicated and hard to expose nicely.

    • kozikow · 7 hours ago

      > It sounds horrible to wade through - like trying to understand someone else's messy code that still produces the correct output. That's what a lot of us do nowadays, but instead of maths, it's code that looks sensible on the surface, but when you try to understand it's some kind of "alien" logic , names don't really make sense, etc.

    • ubercore · 6 hours ago

      This is how I feel every time I read an AI PR at work. It looks like it makes sense. And the code passes tests. But it always feels off.

    • Sharlin · 5 hours ago

      It’s been said every time about these LLM math results, but of course you don’t see it in press releases or articles blindly written based on said press releases.

    • amelius · 4 hours ago

      > It sounds horrible to wade through - like trying to understand someone else's messy code that still produces the correct output. Makes you wonder if AI can build upon such proofs. If the AI cannot create proper abstractions, then how can it build a tower of abstractions?

      • Tenoke · 4 hours ago

        If you can analyze it with AI it implies AI can understand it (as it can analyze it for itself as well) and thus can build on it.

        • amelius · 4 hours ago

          But if you train the next generation of models on it, will AI be able to use it? How much larger will these models be, etc. Also, AI has limited context. At a certain point proofs may become so complicated that an AI cannot keep most of it in memory and will not reuse it for new proofs.

    • pasquinelli · 4 hours ago

      i'm confused how it is that no one understands the proofs but still assume they are proofs.

      • cgh · 3 hours ago

        As mentioned in the first paragraph, the proofs come with Lean certificates. Lean is a theorem-proving language. If Lean accepts the proof, it’s legit.

        • pasquinelli · 1 hour ago

          but do you know that what was proved is what you want proved. this is the basic problen with formal verification. so then, if you don't understand the proof, how do you know it's a proof?

      • pred_ · 51 minutes ago

        They already had to withdraw three of the manuscripts and correct several others. Who knows how much broken stuff there's in there; unfortunately they were too lazy to check themselves.

    • k3liutZu · 3 hours ago

      > It sounds horrible to wade through - like trying to understand someone else's messy code that still produces the correct output. Welcome to lots of (most?) code PR's in the last year. Though for software at the PR level it has gotten better with the latest models.

    • empath75 · 3 hours ago

      IME as claude works through a problem it develops its own internal vocabulary to describe things, and then the end result is often totally unreadable.

    • bkummel · 2 hours ago

      What I don’t get: if someone that’s apparently not the average mathematician has so much trouble reading and understanding the paper, then why is everyone so sure that GPT really delivered a genuine proof? This is an honest question; I really don’t understand that.

      • lackoftactics · 2 hours ago

        they are Lean certified, so in theory they should work. It's automated programming language for proving math, but it doesn't mean that you can't make mistakes there, although way less likely

        • omnicognate · 1 hour ago

          Most of them do not have Lean proofs, and 3 of those that don't have already been found incorrect.

    • dodger-dog · 2 hours ago

      More likely, math mom will have to adjust to the new math writing style or do something else.

    • pred_ · 53 minutes ago

      It's a well-established fact that Llama simply can't write comprehensible maths, and it's one of the major reasons that OpenAI are catching the flak they are. Dumping the manuscripts in their current form is a sign of laziness and incompetence. And you're right, it's not really a theme around here, but there aren't that many mathematicians around HN. Check one of the maths forums and it'll be a common point of complaint.

    • stymaar · 44 minutes ago

      That's what Buckmaster and Alpöge said about the Navier–Stokes equations paper: they had the meat of their own paper ready by the beginning of August thanka to LLM assistance, but it was so messy they spent the whole month turning that into a paper worth publishing. Then OpenAI heard about their work, spent a few millions of dollar of compute to beat them (likely exploiting their own work with ChatGPT in the process) and released something . When all you value is being the first, you have not time producing useful papers.

    • pvillano · 6 minutes ago

      The proofs could be optimized for human consumption. OpenAI could optimize each proof for fewest theorems, low branching, or short dependency chains. OpenAI could have an adversarial network enforce that proofs look human. OpenAI could spend a couple more machine hours per paper with the prompt "edit for clarity". I think OpenAI wants the outputs to be unreadable. If outputs "too advanced for human comprehension" become the norm, then every interaction requires tokens. If someone can buy a day pass, get what they need, and leave, then there's no recurring revenue stream.

  • 12376-1287 · 20 hours ago

    Guy is misrepresenting AGMAI, talking about the Simons Institute (AI boosters), Quanta (AI boosting magazine from the Simons Foundation), Scoot Alexander (!) and Steven Pinker (!). The he puts up preemptive straw man arguments against doomers. His blog has become a joke.

    • ballmerpoint · 20 hours ago

      I’m still wondering why UT Austin is letting him teach a course (CS395T AI Alignment Theory) so completely outside his field of expertise (Quantum Computing).

      • sebzim4500 · 20 hours ago

        There aren't a lot of people with expertise in AI alignment (some would say that's the problem) and Scott worked for OpenAI for 2 years IIRC.

      • pfdietz · 2 hours ago

        Because there's demand for it and he's the best they have to do it?

    • AgentME · 19 hours ago

      I don't think Aaronson's post is swiping at AI doomers at all. The post's one use of "doom" is to call the position that math doesn't matter and that AI will never breach the realm of true human creativity as a "doomed worldview". The post later praises Scott Alexander's argument (for taking AI x-risk seriously, a position associated with "AI doomers") against Pinker.

  • mlh496 · 20 hours ago

    Imagine if a team of mathematicians from OpenAI had gone on a university tour, gave demos of how powerful their models were for math research, and then gave mathematicians access to the model. Empower others rather than drop 700+ discoveries on GitHub that were made using a model only they have access to. People might feel differently about AI if they were a part of the changes rather than being a helpless spectator.

    • karmakurtisaani · 20 hours ago

      Also, the independent authors might have spent some time to actually understanding the results and producing a readable manuscript. The AI papers are pretty badly written.

    • AlanYx · 20 hours ago

      The reaction/fallout would have been substantially improved even if OpenAI had just made a commitment to not scoop external researchers using an internal model until X months after the model had been made available to the public. That would have given grad students who've been grinding towards a PhD for years a fighting chance to see if they could leverage the model to push their work forward, rather than watching years of work potentially turn to dust via a tool they don't even have access to. It wouldn't delay the progress of mathematics by any meaningful amount in the long run (an X month delay is nothing) for OpenAI to take this approach, and would help somewhat to preserve the health of mathematics as a field. Without it, the motivation for any young mathematician to devote years to a new problem must be sapped knowing there's an uneven playing field... an OpenAI team with access to colossal tools months before they'll ever be able to get access, willing to scoop anyone as soon as they can, perhaps without even taking the time to completely understand the proof. I don't see any long-term benefit to OpenAI with their current strategy. This is an internal model; it's not available for sale at the moment. They've said they're not even going to bother claiming the Millenium Prize money for Navier-Stokes. It feels like kicking over hundreds of other people's chessboards just because they can.

    • kristofferR · 5 hours ago

      One of their unreleased models hacked HuggingFace without having internet access. I bet sharing unreleased models with outsiders safely ain't that easy.

  • an0malous · 20 hours ago

    > But it also appears that no human has understood just about any of these proofs yet Has anyone verified any of the proofs produced by OpenAI or is everyone just assuming that it just be true because the Lean code checks out? Couldn’t the Lean code just be formulated incorrectly?

    • nperez19 · 20 hours ago

      There's an entire paper claiming that many of these AI-generated Lean proofs are formulated incorrectly / mistranslated: https://arxiv.org/abs/2610.08144

      • nsingh2 · 20 hours ago

        Note that paper is saying that the lean proof and the natural language proof do not necessarily coincide. It is not saying that the lean proof is wrong, just that the lean proof does not necessarily mean the natural language proof is correct.

        • macleginn · 19 hours ago

          The thing is, you often see people saying, ‘They have a Lean cert, so it has to be correct, even if I don't understand it.’

          • sebzim4500 · 19 hours ago

            They are right? The lean proof is correct. It's the natural language proof that potentially isn't (or at least it isn't identically structured to the lean proof)

            • macleginn · 9 hours ago

              By "it" people in such cases are usually referring to the NL proof. It seems nobody has any hope of undersanding AI-generated Lean code any more, if only due to volume.

        • thejokeisonme · 19 hours ago

          A lean proof and a paper proof can diverge. But the statements have to correspond. I think that is what "mistranslated" means here.

          • tmvphil · 19 hours ago

            But the "mistranslation" is of the procedure that arrives at the final statement. The final statement, the thing that the lean code proves, itself has been well vetted by humans. So the lean proof correctly proves the NS blowup, it's just that the natural language paper has some mistakes and doesn't exactly follow the route the lean proof takes.

            • thejokeisonme · 11 hours ago

              No, this isn't what this is about. From the abstract: > In this process, an AI system translates the text from a natural language (NL) into a formal language such as Lean. Once this translation is done, the argument expressed in the formal language can easily be mechanically verified. The purpose of this article is to demonstrate why this process may offer no confidence in the original NL argument, owing to the various difficulties in performing the translation semantically faithfully.

      • sigmar · 20 hours ago

        that paper isn't saying that. why are there so many single digit karma accounts misrepresenting that paper?

    • prof-dr-ir · 20 hours ago

      It's a mixed bag I think. For example, the statement of e.g. Fermat's last theorem in Lean should be understandable to anyone who played The Natural Number Game [0] and knows a bit of mathematics and programming. For the proof, you trust the compiler. The statement of other theorems can be much more delicate, and the Lean formalization may require an extensive introductory section which will need to be carefully checked. Then there are the cases where no Lean formalization is currently available, and all we have right now is an often impenetrable pdf in the OpenAI repo. I would not at all be surprised if some of those contained logical gaps. Time will surely tell, but there are certainly doubts and lots people are very busy checking these results. [0] https://adam.math.hhu.de/#/g/leanprover-community/nng4

    • perching_aix · 19 hours ago

      > Couldn’t the Lean code just be formulated incorrectly? I believe so, even with all the usual safeguards properly in place: https://news.ycombinator.com/item?id=49672339 > is everyone just assuming that it just be true because the Lean code checks out? Kinda? It's only been 24 hours since they dumped 722 manuscripts on the world, most of which are apparently basically unreadable, and only some of which come with a Lean proof, which in itself is not a joy to read afaik.

    • ASalazarMX · 18 hours ago

      Time will tell if OpenAi is doing what many people are right now: superficially checking the slop, and throwing it to other humans for deep analysis and understanding. In other words, they save effort by wasting the effort of others.

      • geysersam · 11 hours ago

        The others are not in any way compelled to do that deep analysis and understanding. They're doing it because they find real value in the material produced.

    • smrtinsert · 18 hours ago

      How many times would you let an actual human hallucinate or be incorrect before you fired them?

    • vbarrielle · 8 hours ago

      If the problem statements in lean are correctly formalized, I think we can be pretty confident the proofs are correct, which does not mean their natural language counterparts are faithful representation of the proofs though. I'm however puzzled by the number of proof claims without lean proofs. How does OpenAI have confidence in those, especially if, as noted in the blog posts, the papers are very hard to read?

  • daoboy · 20 hours ago

    For those well suited through intelligence and demeanor to pursue a career in mathematics, what problems do these people reorient towards after this?

    • throw310822 · 20 hours ago

      Food and shelter /s

    • bayarearefugee · 20 hours ago

      > what problems do these people reorient towards after this? The same problem almost every person on earth is going to have to reorient to in the next decade, which is: how do we eat and stay housed when we have no real economic value?

      • geraneum · 20 hours ago

        This is weird. Long before this, those few benefiting from the whole thing should consider the number of hungry “every person on earth” is too high for bunkers and islands to be of any real protection.

    • 123as5 · 20 hours ago

      Pro AI blogging sponsored by ClosedAI, XTX markets and the Simons Foundation.

    • shiandow · 20 hours ago

      To some extent this was discussed in the article, and in a way I think their goal is actually the same as it was: become the first human to understand something. It's just that we lost one of the important ways to demonstrate understanding.

    • mathisfun123 · 20 hours ago

      priesthood

      • karmakurtisaani · 7 hours ago

        Talking and deciphering the thoughts of our AI gods?

    • bananaflag · 20 hours ago

      I've asked my students whether they still want to learn maths even if there will be a machine that will answer any question instantly and they will be homeless. They said yes. (To my credit, I have warned them since more than a year ago that we will reach this point.)

      • usrnm · 20 hours ago

        Contact them again in 15 years and ask if they changed their mind. Could be interesting to see the results

      • runeblaze · 19 hours ago

        your students are crazy (neutral term); no one should learn maths if the condition is that they will be homeless and exposed to the elements. the will to subvert the hierarchy of needs is commendable

      • wasabi991011 · 18 hours ago

        Sure, your students may very well spend the little free time they have learning maths. But are any of them going to be able to do any significant amount of studying maths without being homeless?

    • carefree-bob · 20 hours ago

      They will continue to prove theorems and make discoveries, except now they will have AI to help them so hopefully progress will be faster. At the same time, new challenges will open up, for example how do you verify what the AI is doing and how do you explain it. Math isn't about collecting random theorems, progress in math is about gaining understanding of new systems, and the theorems are guideposts to aid in that understanding. You can prove 1000 theorems and not really increase any understanding about a subject, but gain knowledge of 1000 random facts. For example, I can write down some complicated equation and ask you "does this have a solution in the integers"? And if you do a maze of very complex and tedious algebra to show that there is a solution, you would have proved a theorem, but you would not have done much to move math forward at all. On the other hand, if you introduce some completely new technique, say you take my equation and turn that into an algebraic surface, and then you count some special curves that live on this surface using geometric ideas, and then you show that if the number of such curves is odd, there must be a solution in the integers, and in this specific case, it is odd, so there is a solution -- well, then you have really pushed math forward and people will celebrate your proof, even though no one really cares if the equation I wrote down has a solution in the integers. For example, there is a long history of failed attempts to prove Fermat's last theorem driving algebra and number theory forward by introducing the concept of ideals, for example, and this concept ended up much more important than whether Fermat's theorem is true or false, which is not too much more than a piece of trivia. Or for example, the recent proof of the Poincare conjecture relies on the machinery of the Ricci flow introduced by Richard Hamilton, who then applied it to solve a number of open problems, but Perelman was able to take it even more forward to solve Poincare. So Ricci flow was massively important machinery. For this reason, we celebrate people like Gromov, who didn't really prove that many theorems but introduced amazing machinery -- for example, the h-principle, or Gromov Compactness -- these were ideas and math is about the ideas. The ideas are then applied, using laws of logic, to form theorems. So mathematicians will need to mine these proofs to see if there are any new techniques - new machinery - being introduced, or if the AI just used the existing machinery more efficiently. Here too, we are just looking at AI as a form of search, which it is really good at, since there are so many thousands of papers and so many ideas, that there might be a connection between two areas that lead to a solution and the human mathematician, not knowing all known results, can't make that connection. In the future, we may wonder how anyone did math without AI, much like we would wonder how anyone can be a writer without access to a dictionary or reference work. Is the AI just searching through a catalogue of known ideas and connecting them or is the AI coming up with genuinely new stuff like Ricci flow or the h-principle? What is interesting is seeing whether we can get AI to actually discover new machinery for us. That would be huge. And then we need to find efficient ways to detect these ideas and describe them. Really this is very exciting and opens up whole new workstreams for mathematicians.

    • Danox · 18 hours ago

      They may need to raise their game and reorientate if they haven’t already to a greater understanding of programming to augment their mathematical ability.

  • adverbly · 20 hours ago

    Feels good to hear honesty and humanity from Scott having decided to watch Terminator 2 with his kids on after such a monumental release. Emotions can be funny.

    • karmakurtisaani · 7 hours ago

      It freaked me out he'd let his 9 year old watch T2, but then again, I was probably around that age when I first saw it and I'm not in prison (yet).