12 comments

  • brian-bfz5 hours ago
    Hey guys, I&#x27;m one of the organizers. AMA.<p>- We are a team of undergrads at Caltech. We don&#x27;t represent Caltech, any Caltech departments, or any of our sponsors.<p>- We don&#x27;t receive monetary compensation. All the funding raised goes toward paying our judges and participants.<p>- Our goal is to promote responsible AI use. You can read more about our commitments here: <a href="https:&#x2F;&#x2F;mathathonchallenge.com&#x2F;faq.html" rel="nofollow">https:&#x2F;&#x2F;mathathonchallenge.com&#x2F;faq.html</a>
    • danielmarkbruce1 hour ago
      Why not specify 2-3 problems? Won&#x27;t you have people show up having already spent a bunch of time on their self chosen problem? Which sort of defeats the point of seeing what you can do in a short period of time?
      • brian-bfz21 minutes ago
        Solving a problem is the easy part with AI. Selecting a problem worth solving is half of the challenge.<p>We allow prior work as long as it&#x27;s labelled. We record chat logs, so it&#x27;s easy to verify what&#x27;s prior work. When we evaluate the significance of a result, we focus on the part produced at the event.
        • ajkjk7 minutes ago
          it seems like a person could cheat by finding an open problem that&#x27;s AI-solvable ahead of time (by trying many problems), and then pretend to do it for the first time in the competition. You can&#x27;t verify past chat logs so make sure they didn&#x27;t already realize it would work.<p>just something to think about
    • maypop4 hours ago
      Is this hackathon only for those with formal math backgrounds?<p>I&#x27;ve seen a few instances of AI assisted advances math and cs this year that were _not_ published by authors with formal backgrounds in those fields (or even institutional affiliation). Which makes me wonder if they would have a place at the event.
      • brian-bfz4 hours ago
        Yes. To us, solving a problem with AI doesn&#x27;t matter as much as selecting the right problem and understanding the proof. A participant must have a good mathematical intuition.
    • fred1231235 hours ago
      Hey! Any indication on what area of mathematics theses questions are from?
      • brian-bfz4 hours ago
        You pick your own problem! You can even formulate your own conjecture and then prove it. Picking an impactful problem is part of our judging criteria.
        • fred1231234 hours ago
          That sound cool!, do you have hints on the cash prizes the website said something like 2 M ???
          • jegutman4 hours ago
            That’s tokens available I assume for all competitors during the competition.
    • queuebert2 hours ago
      Serious question: Are you sure solving abstract mathematical problems with AI is responsible use? You will likely put mathematicians out of jobs, and I doubt solving the Collatz conjecture is urgent or will save lives. It also robs a future Fields medalist of the pride of doing all by themselves.<p>All things AI seems to assume that more and faster is better, but there is no justification of that assumption. As a biological counterexample, a tree grown quickly will likely not be as healthy or strong as one grown slowly.
      • ajkjk1 minute ago
        I dunno, I like math and I still think automatically solving problems is worth doing. If problem-solving is a hobby then it will still be a hobby after all this. If it&#x27;s about doing things for humanity--which it often is since many thousands of people are paid to do it--then doing it more efficiently benefits humanity. If the value on the other hand comes from people being extremely good at math, rather than from them solving novel open problems, then we should pay them to do that, which is independent of whether the open problems are solved or not. There&#x27;s just no version of this where &quot;not solving the problems&quot; is morally justifiable.<p>Now, I think it&#x27;s the case that professional mathematics spends way too much money on open problems and way less than it should on pedagogy, exposition, mastery, etc. But that has always been a problem, even decades ago (I&#x27;ve been complaining about it my whole life). AI just finally puts pressure on the world to do something about it. I find it relieving, honestly. And I&#x27;m an AI skeptic in many other ways; it&#x27;s not an AI-maximalism thing. I genuinely think the state of the field of mathematics has been a disaster for a long time.
      • suriyaG2 hours ago
        &gt; you will likely put mathematicians out of jobs.<p>why is that a concern in this context? would you have asked the same about steam engines and horses?<p>this is a really cool concept, organized very well. and that is very commendable.
        • lovasoa2 hours ago
          steam engines solved a problem people had
          • tim-kt1 hour ago
            If mathematicians aren&#x27;t solving problems people are having (which your comment seems to imply), then putting them out of their job with AI is not a bad thing. Of course mathematicians are solving problems, just in a very different way than other professions.
          • hatsix35 minutes ago
            Don&#x27;t respond to a strawman argument with another strawman. The post you are responding to ignored the reasons given in the second paragraph. They&#x27;re just trying to score points by preaching to the choir, not engage with the concern.
      • brian-bfz1 hour ago
        I don&#x27;t think AI will solve more math problems in a world with Mathathon than the counterfactual by EOY. People will use AI in math anyways. What matters is: can we encourage them to do so transparently and with full understanding of their results? Can we change the incentives in academia to reward problem selection and verification over proof generation?
        • reasonableklout1 hour ago
          It’s a noble goal to change the incentives, but how will you prevent the headlines from this event being “students prove Collatz conjecture with Claude” and instead be “students give great explanation of Collatz conjecture proof”?
          • brian-bfz12 minutes ago
            You&#x27;re right that we can&#x27;t. We&#x27;ll be responsible in our press releases and award prizes based on explanation, but we don&#x27;t control the headlines. However, this is already an improvement over the current state, where results are announced by headlines alone.
      • danielmarkbruce1 hour ago
        It&#x27;s more responsible than the use of electricity for a messaging board for people to argue minutiae.
      • ktallett2 hours ago
        I get your point and agree to some extent, but you can&#x27;t understand the proof without significant background in Maths so it will just allow mathematicians to solve issues faster than not have the opportunity at all.
        • a2ff6eeb02 hours ago
          Why do people need to understand proofs? If Amazon improves package routing with new advances in graph theory, my cat doesn&#x27;t need to understand it to benefit from better shipments of cat food.<p>Similarly, humans don&#x27;t need to be involved in scientific advances to benefit. We just need an aligned AI to take over the scientific thought for us. AI is already better than all but the top tier of humans at doing mathematics, it&#x27;s writing most of the posts on the front page of this website, and it&#x27;s doing the bulk of programming at many startups.<p>We can&#x27;t put this genie back in the bottle.
          • mattmcal2 hours ago
            Well, for one, most math proofs don&#x27;t have any practical applications, so a proof that no one reads is basically a digital paperweight. You might as well suggest AI write novels for other AI to read.
            • a2ff6eeb01 hour ago
              The hope is that some of them end up being useful; otherwise, nobody would be funding math departments. Mathematics typically anticipates and enables new physics and chemistry.<p>If people are just doing math to kill time, I don&#x27;t get why anyone would bother with AI. Do people really enjoy picking through a million lines of generated Lean code, if it&#x27;s not for any practical use?
              • mattmcal59 minutes ago
                If you&#x27;re interested in the topic enough to comment on it, you&#x27;ll probably find it worthwhile reading a mathematician&#x27;s perspective. Here&#x27;s the prolific Terry Tao: <a href="https:&#x2F;&#x2F;mathstodon.xyz&#x2F;@tao&#x2F;117219548485446992" rel="nofollow">https:&#x2F;&#x2F;mathstodon.xyz&#x2F;@tao&#x2F;117219548485446992</a>
                • a2ff6eeb039 minutes ago
                  They don&#x27;t actually say anything about why anyone should fund this, though. I don&#x27;t get why a society should worry about progress in mathematics if there&#x27;s no practical benefit expected.<p>Maybe there&#x27;s two kinds of math that we need? Useful math and navel gazing, and we can hand the first to the machines, and let hobbyists do the second in their free to entertain themselves?
  • corinthia6 hours ago
    recent caltech grad here! and know some of the organizers well<p>caltech&#x27;s cs department is very, very weak, and has struggled to recruit top people in the last few years, and the most recent AI faculty hires have had issues. this is very slowly changing but a lot of the motivation for htis was to create a way for students to get ml &quot;recognition&quot; and learn about ai since it cant be done through the school right now. really glad to see hn picked this up!
    • brian-bfz5 hours ago
      We aren&#x27;t affliated with the CS department. We&#x27;re a student-led initiative. Our goal is to promote responsible AI use in math.
  • Semkas11 hours ago
    Won&#x27;t deny that this is an interesting idea, but I feel like waiting on the output of an LLM for 40 hours feels like it is completely antithetical to what makes classic Hackathons appealing &#x2F; educative.<p>More generally, I don&#x27;t think the shape of a hackathon (intensely working for a short timespan) maps at all onto the way LLM Math progress has seemingly been made so far; AFAIK it mostly involves picking out something for the Model, then having it run for a week with sporadic correction &#x2F; encouragement.
    • brian-bfz4 hours ago
      1. IMO the hard part isn&#x27;t prompting. It&#x27;s selecting the problem and understanding the solution. It&#x27;d be especially exciting if a participant formulates their own conjecture, proves it with AI, then generalizes it to a new theory.<p>2. We have talked to mathematicians and frontier lab employees. We think 40 hours is enough to produce interesting results.
    • falcor848 hours ago
      But it&#x27;s not &quot;waiting on the output of an LLM for 40 hours&quot; any more than a regular hackathon is &quot;waiting for my damn teammates to finish their part for 40 hours&quot;. From my experience using agentic coding for hackathons, the best teams are those that coordinate with the AI agents in relatively quick cadence, generally giving it small tasks and steering it often. Teams may want to run some long-running sessions too, especially closer to the deadline, but even then, they&#x27;d probably want to run and follow several sessions in parallel, and continuously inspect their work so that they have reasonable confidence that their main efforts will wrap up before the deadline. There is an art to it.
    • Donald8 hours ago
      Have you done any math hacking with sol&#x2F;astra or fable? It’s more fun than using them for coding. The models are great at the monotony, like constructing a Gröbner-basis, etc. But they’re all still absolutely awful at coming up with new ideas, new proof methods, or new constructive forms. So you spend all your time on coming up with novel hypotheses yourself and handing off the rote work to an agent.<p>It’s also quite fun to get instant results by finding isomorphisms into unfamiliar areas of mathematics that previously would’ve required some networking in order to build a collaborative relationship.
      • rookienumbers308 hours ago
        &quot;A mathematician is a person who can find analogies between theorems; a better mathematician is one who can see analogies between proofs and the best mathematician can notice analogies between theories. One can imagine that the ultimate mathematician is one who can see analogies between analogies.&quot;<p>I wonder how models perform on finding analogies between analogies
        • skybrian7 hours ago
          Well, I suppose that explains monads. It’s one thing to see an analogy and quite another to make it the basis of an API.
        • tomrod8 hours ago
          &gt; I wonder how models perform on finding analogies between analogies<p>Load-bearingly verbose, in my experience.
    • Aboutplants11 hours ago
      If the goal is to accomplish something then why limit yourself with available tools?<p>I’m not a full on AI optimist but it is absolutely the most powerful tool in a host of applications. From a Hackathon perspective, obviously in the 90s it was much more unorganized, but the same ethos existed. Use all available tools to accomplish the goal&#x2F;task, it’s where a lot of incredible learning came out of. The same will hopefully happen in scenarios like this one
      • a2ff6eeb011 hours ago
        It seems like making progress on math is letting the AI run fully autonomously for a few days, occasionally asking it to keep going.<p>I&#x27;m not sure people need to organize a mathathon to wait for a computer to give a printout. They mainly need tokens.
        • MostlyStable5 hours ago
          That&#x27;s how several major AI advancements have happened. I have seen no evidence that that is the fastest way to make progress right now. I expect that, much like chess engines, it will not take too long before AI is significantly better than AI + human. But right now, my bet is that we are still safely within the window where an AI + human mathematician team is still better than AI alone (at least for the case where the human has learned how to work effectively with the partner....something that this event could possible be good for teaching).
          • a2ff6eeb05 hours ago
            I suspect that the best progress will be made by a team that purely spends their time taking a list of open problems and promoting &quot;solve &lt;problem&gt;&quot;, without actually trying to understand anything. Just keep as many problems in flight as you can across as many sessions as you can.<p>You can probably ask the AI to come up with a list of problems itself, and rank them by the likelihood of progress.
            • youoy3 hours ago
              Then 20 years go by and you wake up one day with questions that you cannot get out of your mind: why did I start prompting the LLM for? Why did I need these random proofs for? What do i do with my repo with 2billion lines of Lean?
        • charlieyu110 hours ago
          Are they actually autonomous? I’d say subject knowledge at the prompt stage plays a large part towards getting proper results
          • a2ff6eeb010 hours ago
            When Claude made progress on the Riemann conjecture, here are the kind of prompts used:<p>&gt; <i>Jarred&#x27;s input was mostly limited to sending Claude messages of encouragement (mostly variants of “keep going” or “believe in yourself”).2 This seems to have helped Claude overcome some initial skepticism that it could make meaningful progress.</i><p>And left it for a long time. Jarred isn&#x27;t a mathematician, he&#x27;s the maintainer of a janky JavaScript environment.<p>Here&#x27;s the transcript: <a href="https:&#x2F;&#x2F;www-cdn.anthropic.com&#x2F;8a0d1add3c637b858a9a181e98c40e9548c3f44f.pdf" rel="nofollow">https:&#x2F;&#x2F;www-cdn.anthropic.com&#x2F;8a0d1add3c637b858a9a181e98c40e...</a>
          • hgoel10 hours ago
            Prompting for some of the results was almost the &quot;Computer, do a breakthrough. Make no mistakes.&quot; meme. Just someone telling the model to keep trying a couple of times.<p>Unfortunately we don&#x27;t actually know what kind of prompting was done for the more prominent results.
            • charlieyu18 hours ago
              It’s definitely not how I work, I’d need to read the model responses and set a direction for the model to go
              • hgoel7 hours ago
                Same here, it&#x27;s one thing if I&#x27;m just screwing around, but if I&#x27;m trying to do anything serious, I need to at least have a handle on what it&#x27;s doing, and, when thinking traces are available, keeping track of any logical errors in the model&#x27;s reasoning.
                • a2ff6eeb05 hours ago
                  you know the thinking traces are redacted and summarized using another model, right? The actual thinking trace looks something like:<p>7♣-removal-IS-the-prerequisite-for-10♠&#x2F;9♥!!)-⟹-OVERLAP-(ii)+(iv):-{6♠ J♦ 9♥ 2♣}-=-FOUR--—-UNLESS-7♣&#x27;s-seat-8♥-...-and-2♣-drains-only-at-crack-:-⟹-2♣-celled-+-9♥-celled-simultaneously-UNAVOIDABLE-in-t8-dig--—-BREAK:-9♥-drains-to-10♠-THE-MOMENT-10♠-is-free:-t8-dig-order:-[K♣→t2]-[2♣→cell]-[7♣→8♥-:-8♥-WHERE:-post-chunk-9♠-:-chunk-⟸-K♣--done-:-ORDER:-K♣→t2,-CHUNK→K♣-(cap-4!!:-cells-then:-{6♠ J♦ 9♥}-FULL--chunk-cap-=-1--—-F-F-F-F-F.-—-chunk-BEFORE-9♥-celling?!-:-9♥-celled-at-t1-dig-⟸-needed-for-5♣-⟸-4♥3♣-⟸-t2-dig-⟸-K♣-seat-⟸-chunk-:-⟹-chunkAFTER-9♥-celling-FORCED--:-chunk-cap-with-{6♠ J♦ 9♥}:-1--—-—-J♦-THE-NEW-CANCER.-—-⟹-J♦-celling-DELAYED-till-after-chunk?! :-J♦&#x27;s-celling-was-for-J♥→Q♠-(5♦-access-for-4♣):-DELAY-4♣-resolution:-4♣→CELLearly-(as-always)-then-4♣-cell→5♦-LATER-when-5♦-frees-!!!:-cells-rotation:-4♣-celled-[t2-dig-…-5♦-freed]:-5♦-freed-⟸-J♥→Q♠-⟸-J♦-celled-:-⟹-{6♠, 4♣, J♦}-overlap-window-until-4♣→5♦-drains:-then-{6♠ J♦}+1-rotator-:-—-AND-9♥?!-9♥-celled-[t1-dig…]:-OVERLAP-{6♠ 4♣ 9♥}-before-J♦-even-:-⟹-rotator-slot-SINGLE:-timeline-:-(1)-{6♠}+2:-…-(2)-+9♥-(t1-dig):-{6♠ 9♥}+1:-(3)-+4♣-(t2-dig):-{6♠ 9♥ 4♣}-FULL-:-(4)-NEED:-t6-dig-(9♦8♠→10♣--no-cell;-8♥→CELL--FULL)--—-8♥-alternative-seat-pre-chunk:-NONE-—-.-⟹-⟹-THE-TRIANGLE-{9♥ 4♣ 8♥}-verdammt.-—-⟹-dig-t6-BEFORE-t2?!:-(3&#x27;)-+8♥:-{6♠ 9♥ 8♥}-FULL:-J♥→Q♠-⟸-J♦-cell--FULL--AAAAAAAAAAAARGH.<p>Citation: <a href="https:&#x2F;&#x2F;www-cdn.anthropic.com&#x2F;d00db56fa754a1b115b6dd7cb2e3c342ee809620.pdf" rel="nofollow">https:&#x2F;&#x2F;www-cdn.anthropic.com&#x2F;d00db56fa754a1b115b6dd7cb2e3c3...</a>, section 6.2.2<p>You&#x27;re not going to get a handle on what it&#x27;s doing. The thinking traces are there to make you feel better about yourself.
                  • charlieyu12 hours ago
                    Why is the model playing poker
                  • hgoel5 hours ago
                    I&#x27;m referring to models where the actual traces are available. Eg. I&#x27;ve been using a local Qwen3.8-Next-Flash lately.
                    • a2ff6eeb05 hours ago
                      Still not meaningful -- <a href="https:&#x2F;&#x2F;arxiv.org&#x2F;pdf&#x2F;2504.09762" rel="nofollow">https:&#x2F;&#x2F;arxiv.org&#x2F;pdf&#x2F;2504.09762</a>; even for local models, the reasoning traces are often filtered and summarized to sound sensible to humans. And even if not, they don&#x27;t necessarily represent what the model is thinking.
                      • hgoel5 hours ago
                        Hmm interesting, thanks for the link, I&#x27;ll have to give that a read.
        • jhonof10 hours ago
          I think the purpose of an event like this would be to optimize the process so that it isn&#x27;t just occasionally asking an AI to keep going.
          • a2ff6eeb04 hours ago
            That sounds like adding a bottleneck, unless you mean writing a harness that automatically asks the model to keep going, so that there&#x27;s no humans involved at all?
      • loloquwowndueo11 hours ago
        If the goal is to run 42km why limit yourself? Use a car and win.
        • falcor848 hours ago
          But the goal here is not to run 42km; to stay with the outdoors metaphor, it&#x27;s more like deciding where and how to set up a bivouac - use whatever tools you have at your disposal to analyze the area you&#x27;re in, and find the best site to stay in overnight.
      • thatseasy9 hours ago
        &gt; why limit yourself with available tools<p>Because the companies that run frontier models are malevolent by every metric.<p>They are destroying the environment, especially those in neighborhoods of low income people.<p>They are empowering their owners who are some of the most deplorable and duplicitous people living.<p>They are destroying personal compute to avoid competition with local models by buying all computer components with “promised money” and forcing their P into AI.<p>They stole the entire creative output of humanity and are trying to sell it back to us.<p>They are only good for giving wealth access to skill while removing from the skilled the ability to access wealth.<p>They are being used to kill in war and for surveillance.<p>Seriously why would you use them? Your use only emboldens them; making you complicit in their nefarious success.<p>I for one, am one who walks away from Omelas.
      • Fraterkes11 hours ago
        Did you read the article? This is not a hackathon where you build software, it’s one where you’re trying to get a model to make progress on a frontier math problem. The point is that <i>that</i> activity may not map well onto the shape of a hackathon
    • groundzeros20157 hours ago
      Are you sure it’s letting it run and not going back and forth interactively?
  • chaoxu10 hours ago
    I&#x27;ve applied as a team, hope I get in.<p>Recently I care about how to create harness for mathematics that uses up the complete reasoning ability of the model. I care about both capability and cost.<p>Most generic harness we have now are not made for maximizing reasoning. I&#x27;ve tested agents like codex, and rarely the cost of reasoning tokens reaches more than 20%. Which is quite strange as math requires a lot of reasoning. So hackathons can be a good test bed.
  • youoy10 hours ago
    Should I read this as the big labs trying to move maths forward? Or the big labs trying to use professional mathematitians as (cheap?) Labour for validating LLM outputs? On yesterdays &quot;An Alien Mind&quot; post from openAI they openly said that maths is not a priority for them, so I personally know what to think...
    • brian-bfz4 hours ago
      We are a student-led initiative. Our sponsors don&#x27;t pay us and don&#x27;t have a say in our decisions. All of our funding goes toward our judges and participants.
      • youoy3 hours ago
        Hey! Thanks for answering, i appreciate it. Dont get me wrong, this is what you should be doing, understanding what these models are good (and most importantly bad) for. My observation is that this is very very valuable for the labs, and in an ideal world they should be paying you to do this, not just the tokens and &quot;prize&quot; for the winners.<p>I am also a bit frustrated seeing maths go in the direction of prompt enginnering. I am afraid of a world were a math phd student cannot go one week thinking about a problem without prompting an LLM to give him&#x2F;her an invented answer. Something is lost along the way.<p>For me maths is not Lean, or formal systems, or an agent reasoning about formal systems to join literature from different fields. I see the value of it, but i think it will make it way more difficult for students (and profesional mathematitians) to see beyond that. And i see us heading into a reality were those who think like me will in practice remain a minority for quite a few years&#x2F;decades because the low hanging fruit of LLMs will be to vast to ignore.
        • brian-bfz3 hours ago
          Agreed. I think we share the same concerns about LLM use. We just drew different conclusions.<p>Mathathon&#x27;s goal is to reshape rather than stop LLM use. Can we set high standards for LLM use? Can we highlight the roles of a mathematician beyond proof generation? Can we redesign our incentives to promote these standards and roles?<p>I&#x27;d love to hear your thoughts on how to improve this event. We&#x27;re very open to criticisms.
    • isotypic7 hours ago
      Obviously the latter - this fact is betrayed by how the page lists the S^6 complex structure result, which was released as a 100 page barely readable mess (in fact even this might be too charitable), as still &quot;unverified&quot;. Clearly a situation labs would like to avoid for future claimed results.
    • bwfan1235 hours ago
      &gt; Or the big labs trying to use professional mathematitians as (cheap?) Labour for validating LLM outputs?<p>I am told AGI has been achieved. If so, shouldnt these systems be out and about on their own ? Looking at 1st proof submissions in batch 2 it is clear that fully autonomous AI systems have a long way to go.<p>AI harnessing human labor with the incentive of 2M in free tokens is the way my skeptic eye sees it, or humans being duped as reverse-centaurs.
      • a2ff6eeb04 hours ago
        It&#x27;s very clear that the AI still has no motivation beyond its prompts. Humans can mostly outsource their thinking today across a wide variety of topics, but they still need to express their desires.
      • ianm2185 hours ago
        &gt; I am told AGI has been achieved. If so, shouldnt these systems be out and about on their own<p>This seems like a strawman. It’s certainly not consensus that AGI has been achieved and I don’t think the people participating in this event feel like there is no value in human input or steering the AI.
    • tzs9 hours ago
      Probably many mathematicians want answers to the questions from the page:<p>&gt; This AI advancement raises the following questions: (a) How much can AI speed up the process from ideation to peer-reviewed publication? (b) What is the role of a mathematician when AI can solve conjectures faster?<p>and the big AI companies agreed to sponsor them to find out because it is good publicity for the companies.
    • blondie9x9 hours ago
      Yeah it&#x27;s a bit of a tricky situation. It&#x27;s almost like a bribe in a sense.
    • sb1012810 hours ago
      [dead]
  • charlieyu112 hours ago
    Interesting, if only I still have energy to work on something 40 hours non-stop
  • boothby8 hours ago
    &gt; It will be the first hackathon ever devoted to research level mathematics.<p>Well, that&#x27;s pretty damned ignorant; I was attending William Stein&#x27;s hackathons on the BSD conjecture and the Sage Math project nearly 2 decades ago.
  • amelius9 hours ago
    Will BigAI support this with free access to lots of hardware loaded with frontier models?
  • viccis6 hours ago
    No it&#x27;s not, there are tons of programs in mathematics where you go there, form teams, work a problem as a team that&#x27;s likely to get a result, and then publish the results from all the groups in the conference proceedings. They are called Research Collaboration Workshops.<p>Very consistent pattern from these tech companies in their mathematics press releases that shows a conspicuous lack of experience in the research math world.
  • xqcgrek25 hours ago
    No self respecting mathematician is going to willingly become a marketing tool for these companies solving neglected and irrelevant puzzles.