18 comments

  • joegibbs21 minutes ago
    There's one of mine in there where I predicted in 2023 that it would be 20 years until AI would be reliably able to entirely build and deploy arbitrary applications from a prompt. I was off by about 18 years on that one!
  • bmenrigh5 hours ago
    At least 1/3rd of these predictions aren't clear enough to determine exactly what is being claimed/predicted. Even after reading the full comment multiple times, on a lot of them I couldn't tell where the author had set the goalposts well enough to say whether we've crossed it or not.
    • jerf5 hours ago
      Well, I can answer this one: <a href="https:&#x2F;&#x2F;stoppels.ch&#x2F;goalposts&#x2F;?c=40662140" rel="nofollow">https:&#x2F;&#x2F;stoppels.ch&#x2F;goalposts&#x2F;?c=40662140</a><p>jerf, 2024: &quot;If it could be solved with a Math Overflow-post level of effort, even from Terence Tao, it isn&#x27;t what I was talking about as &quot;high level math&quot;.<p>&quot;I also am not surprised by &quot;Consider a generation function&quot; coming out of an LLM. I am talking about a system that could solve that problem, entirely, as doing high level math. A system that can emit &quot;have you considered using wood?&quot; is not a system that can build a house autonomously.<p>&quot;It especially won&#x27;t seem all that useful next to the generation of AIs I anticipate to be coming which use LLMs as a component to understand the world but are not just big LLMs.&quot;<p>The voting gloss: &quot;An AI fully solves a research-level math problem on its own, not just suggesting an approach.&quot;<p>Yes, I&#x27;m satisfied. I don&#x27;t even feel bad in hindsight. Coding assistants had a nice, gradual rise up the utility curve. Math went from &quot;lol, can&#x27;t add two six-digit numbers&quot; to research-math level almost overnight in comparison.
      • 4884485853 minutes ago
        But it still makes mistakes when adding numbers
        • shiandow5 minutes ago
          Honestly that makes me <i>more</i> convinced it&#x27;s actually doing mathematics.
    • happytoexplain5 hours ago
      Right - people on HN are <i>generally</i> reasonable about objective things. The vast majority of comments (outside those chosen for this website) are not &quot;AI will never ...&quot; but rather, &quot;AI does not currently ...&quot;. Of course the further you go back (I&#x27;m seeing a lot of comments from ten years ago!) the more skeptical they get, obviously. That&#x27;s a funny thing to go back and see with modern context, but it doesn&#x27;t really call for snideness&#x2F;mockery (something I think is sadly increasing on HN).
    • pitched5 hours ago
      &gt; cannot do precise things like coding software since humans will never be able to use natural language to specify their requirements.<p>To your point, this example. The issue expressed here is with humans, not AI. We are still pretty terrible at writing specs. TBF, the AIs are too but that wasn’t being voted on.
  • Retr0id5 hours ago
    Heh, there&#x27;s one of mine: <a href="https:&#x2F;&#x2F;stoppels.ch&#x2F;goalposts&#x2F;?c=39727943" rel="nofollow">https:&#x2F;&#x2F;stoppels.ch&#x2F;goalposts&#x2F;?c=39727943</a><p>&quot;GPT-4 looks at original ASCII art of a foot, not copied from the web, and says it is a foot.&quot;<p>The vote is currently 64% yes, 18% no.<p>Just now I asked Opus 5.5 to generate an ASCII art foot, and it did a passable job. It&#x27;s not great, but it&#x27;s a foot. Then I pasted it into ChatGPT (whatever they&#x27;re serving to the free tier by default, which seems to be 5.6 Luna), and it said it was a &quot;train&#x2F;locomotive&quot;: <a href="https:&#x2F;&#x2F;chatgpt.com&#x2F;share&#x2F;6abeaa39-cc80-83ed-851f-29370db08964" rel="nofollow">https:&#x2F;&#x2F;chatgpt.com&#x2F;share&#x2F;6abeaa39-cc80-83ed-851f-29370db089...</a><p>Maybe it&#x27;s Opus&#x27;s fault for drawing a bad foot but I think it&#x27;s fair to say LLMs are still pretty bad at ASCII art (without additional tool calling etc).
    • tedsanders4 hours ago
      For me, 6.1 Sol nailed it immediately:<p>&gt; A bare foot and ankle, pointing right, with three little toes.<p>I wonder how much of the wide variation in perceptions of LLM capabilities is driven by the gulf between free models and frontier models. Luna getting something wrong is not always great evidence for LLMs be unable to do that thing.<p>Edit: for curious skeptics without access to 6.1 Sol, I tried 3 times and it got it all 3 times. Convo share link: <a href="https:&#x2F;&#x2F;chatgpt.com&#x2F;share&#x2F;e&#x2F;6abeb955-7614-832e-a5e1-b1bd134f1971" rel="nofollow">https:&#x2F;&#x2F;chatgpt.com&#x2F;share&#x2F;e&#x2F;6abeb955-7614-832e-a5e1-b1bd134f...</a>
      • ben_w4 hours ago
        Case in point, this nonsense came out of 5.6 Luna: <a href="https:&#x2F;&#x2F;chatgpt.com&#x2F;share&#x2F;6abeae1e-42b4-83ed-8966-7e82ae0bef5c" rel="nofollow">https:&#x2F;&#x2F;chatgpt.com&#x2F;share&#x2F;6abeae1e-42b4-83ed-8966-7e82ae0bef...</a><p>Like, is this an ice-cream? A tooth?
      • tomalbrc4 hours ago
        Please share a link to the conversation, otherwise I am not buying it.<p>Because &quot;for me DeepSeek Flash 4.1 nailed it immediately&quot;, trust me bro.
    • ben_w4 hours ago
      Mm, I kina agree with the AI on this one:<p><pre><code> (_)(_)(_) represents the wheels </code></pre> They do look rather wheel-like; I have to assume you see them as toes though?<p>It&#x27;s like the duck-bunny picture to me. If I focus on the &quot;wheels&quot;, I see a steam train locomotive (but perhaps I&#x27;m only seeing that because I read your comment?); if I look at the ankle I see a foot.
    • adam_rb1 hour ago
      I think the problem is that you&#x27;re using basing your conclusion from the cheap&#x2F;dumb models available on the free tier of services. I just asked GPT6-Astra in Codex and it replied:<p>&quot;It’s ASCII art of a bare foot and lower leg, with the toes pointing to the right.&quot;<p>No tool calling, just an immediate reply with the correct answer.
      • brudgers1 hour ago
        Sure, but it’s already read the HN thread.
        • akavi30 minutes ago
          ...that&#x27;s not how LLM training works.
    • asdfasgasdgasdg46 minutes ago
      Opus 5.5 was able to parse and understand an ASCII art foot when I pasted one in.
    • howunfortunate4 hours ago
      Readers: before you vote or comment, look at that foot.<p>I think I would have failed this test!
    • hydrolox4 hours ago
      To be fair if a human was given a linear sequence representing ascii art you couldn&#x27;t tell either
    • nonameiguess4 hours ago
      This is a Rorschach test, not a foot. If you&#x27;d shown this to me without telling me what it was meant to be first, I&#x27;d have guessed a crematorium.
      • dec0dedab0de4 hours ago
        Has anyone done Roschach tests for AI? That would be an interesting study to see how different models responded.
  • outlore20 minutes ago
    These questions could benefit from being rephrased to make it clear what is being voted for
  • ben_w5 hours ago
    Very pleased one of my predictions was totally wrong: <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=23252711">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=23252711</a><p>Sure, sure, what LLMs make still isn&#x27;t &quot;efficient bug-free code&quot;: my prediction is falsified because while LLMs can write and train new models with machine learning, ML is fundamentally not advanced enough to throw arbitraty new tasks at like this.
    • FabCH5 hours ago
      Somewhat appropriate the site the OP links to is called „goalposts“ because as far as I can see, people keep shifting theirs.<p>In your case, the comment you link to says „business tasks“ and you expanded it now to „arbitrary new tasks“. Those are not the same. An LLM today sure can do many many many business-speak conversion tasks.
      • tripleee5 hours ago
        &gt; An LLM today sure can do many many many business-speak conversion tasks<p>Not reliably, and not without supervision. That&#x27;s the main point. I&#x27;m trying really hard to figure out a workflow that doesn&#x27;t require me to review the code and I just don&#x27;t see how it&#x27;s possible (yet)<p>You either need a comprehensive test suite (which requires understanding the code in order to create) or you need to review the actual implementation code to make sure it does the right thing
        • FabCH3 hours ago
          Code is a tiny part of &quot;business&quot;.<p>Most business is correspondence with people who want money from you and people you want money from.
      • ben_w5 hours ago
        I&#x27;m not always precise with my language, but business tasks can be pretty broad, I think &quot;arbitrary new tasks&quot; is not an unreasonable rephrasing on my part?<p>Consider I was replying to this:<p>&gt; So are we all going to be out of a job?<p>While your boss now has the capacity to ask Claude to train a new AI model to auto-balance a tower defence game&#x27;s mob, cost, and tower parameters (I know because I&#x27;ve done it), this only matters if you and your boss are working in a video games company.<p>If you and your boss are actually florists, you care if your boss can get Claude to automate a rose pruning, dead-heading, and fertilising robot.<p>People are trying, but I don&#x27;t think they&#x27;d be happy with 91.5% success rate: <a href="https:&#x2F;&#x2F;www.emerald.com&#x2F;ir&#x2F;article-abstract&#x2F;doi&#x2F;10.1108&#x2F;IR-04-2026-0198&#x2F;1398287&#x2F;Design-and-experimental-evaluation-of-an?redirectedFrom=fulltext" rel="nofollow">https:&#x2F;&#x2F;www.emerald.com&#x2F;ir&#x2F;article-abstract&#x2F;doi&#x2F;10.1108&#x2F;IR-0...</a>
        • FabCH3 hours ago
          Don&#x27;t get me wrong, we are all guilty of this.<p>It&#x27;s just amazing how quickly we accept that models are good at something.<p>My florist boss can&#x27;t get Claude to automate rose pruning. But she sure as hell doesn&#x27;t need to wait until Jacques is back in the shop to respond to that French supplier anymore. There is a lot of &quot;business tasks&quot; that are just paper being shuffled around no matter if you are a florist, baker, workshop owner, custom CNC shop, student offering lessons in extra time or whatever. And LLMs are already scary good at those.
          • ben_w3 hours ago
            &gt; There is a lot of &quot;business tasks&quot; that are just paper being shuffled around no matter if you are a florist, baker, workshop owner, custom CNC shop, student offering lessons in extra time or whatever. And LLMs are already scary good at those.<p>Yes indeed, but I was responding to &quot;So are we all going to be out of a job?&quot;, not &quot;Will AI radically change the jobs market?&quot;<p>We got the thing I thought would make everyone unemployed (AI which can make AI), but it turned out the AI good enough to make AI, happened before we figured out the general problem of few-shot learning that would mean the AI made by AI puts us all out of jobs.
      • Dylan168075 hours ago
        You can&#x27;t ignore the rest of the sentence. &quot;every other task their business does&quot; &quot;everyone will be out of a job&quot;<p>This means it has to handle basically all business tasks, so &quot;arbitrary&quot;. I&#x27;m not sure what percent you have in mind by &quot;many many many&quot; but I would say it can&#x27;t code half the things you need in an efficient and minimally buggy way.
        • FabCH3 hours ago
          What code does a village vet clinic need? In all seriousness.<p>Even IF they need code, they need at best a CRUD app to track patients, that&#x27;s it. There is no way Fable or Opus 5.5 can&#x27;t one-shot a village vet clinic app in 30 minutes, and only with &quot;I need a village vet clinic app&quot; as a prompt, and whatever questions it decides to ask along the way with it&#x27;s &quot;ask user&quot; tool.<p>Or a florist, to use the example from a sibling comment.<p>Code is tiny part of &quot;business&quot;.
          • Dylan168073 hours ago
            Anything you can&#x27;t solve with code just means the AI is doing worse on the benchmark isn&#x27;t it? That&#x27;s why I didn&#x27;t go into detail on that aspect.<p>And that one shot app is not going to be bug free.
          • ben_w3 hours ago
            &gt; What code does a village vet clinic need? In all seriousness.<p>Automated diagnostics, pharmacist, surgical robot, something to express anal glands without harming the patient.<p>Dog-English machine translation.
    • tripleee5 hours ago
      &gt; reliably convert business-speak into efficient bug-free code<p>I actually think this would take AGI to solve, which makes me optimistic about the future of software development.<p>All the benchmarks are currently testing against automated tests the AI can use as an oracle
    • vlyan5 hours ago
      so the conditions for your prediction simply haven&#x27;t been met yet.<p>if&#x2F;when you can tell a model to do a thing and be confident that it did the thing, it&#x27;s joever for 90% of knowledge workers.
      • ben_w5 hours ago
        The relevant condition was met; my misjudgement was that meeting it would require ML to be advanced enough to be able to train on arbitraty tasks from realistic (ie small) numbers of examples.
  • ErrantX5 hours ago
    What is interesting to me is in 2016 people were like; pass Turing test, write code, order me a coffee.<p>And even in 2024 the themes are similar, generally more complex or specific about the coding&#x2F;turing&#x2F;action test.<p>But in 2026 a huge shift, we have things like; can open a physical door, emulates human pettiness convincingly, makes novel scientific breakthroughs.<p>That alone tells you a lot IMO
    • ianjbutler4 hours ago
      Sigh, the whole &quot;obviously the turing test is solved&quot; meme is annoying.<p>Like, if we meant that it convincingly masquerades as a shitposter, ok. But everyone still bitches about AI slop, and everyone knows the writing is still bad. How does that even work if the turing test is obviously solved?<p>More to the point though, if you grill SOTA models on counterfactuals, causal world-models etc, you&#x27;ll trip them up in a way that actually <i>will not work</i> on ESL students and children. Certainly there&#x27;s no way to find a person that struggles with that <i>and</i> is also capable of cheerful fluent erudite discussion about astrophysics with perfect grammar. Yes, it&#x27;s getting harder obviously.. but detecting machines with determined, focused and intelligent interrogation remains pretty easy. If nothing else, the models are cooperative where people wouldn&#x27;t be and that&#x27;s a signal too.<p>The best progress we&#x27;ve made is that most people do agree that this <i>doesn&#x27;t practically matter</i> very much, i.e. we generally recognize the stakes were always overstated. But the constant vague appeals to common-sense that &quot;of course it&#x27;s a solved problem!&quot; always feels naive or fake.
      • johnsmith184018 minutes ago
        I was just thinking how anyone still thought AI didn&#x27;t pass turing already. There&#x27;s been literal papers proving average people cannot tell reliably.
  • 6thbit4 hours ago
    Not sure why this thread got flagged ?<p>Its fun. Can you add a sort by controversial? I&#x27;d like to know where people disagree the most between yes and no.
    • simonw4 hours ago
      Yeah this shouldn&#x27;t be flagged, it&#x27;s a neat project.
      • stabbles3 hours ago
        (OP here) It was fun as long as it lasted ;) I&#x27;ll leave it open for a few more days, but already it has enough votes for an interesting results page.
        • 6thbit19 minutes ago
          seems unflagged now :) would love to see a detailed results page
  • mrweasel5 hours ago
    The Turing test is interesting, because I believe that the current LLMs are perfectly capable of parsing the it in many situations. On the other hand we also have people are sound like they aren&#x27;t real.<p>Looking back, was the Turing test flawed perhaps? It failed to take into account that humans can be rather bad at telling actual people from a &quot;parrot&quot;. Turing was perhaps a little to optimistic about people.
    • bluefirebrand4 hours ago
      Of course the Turing test was flawed. We already knew that based on the Chinese Room argument.
    • suopspaces4 hours ago
      [dead]
  • johnsmith184011 minutes ago
    All I learned from this is that 40% of hackernews are AI haters which maps pretty well from the overtly negative sentiment on it constantly.
  • delichon5 hours ago
    If for each mistaken prediction there was some mild accountability, like someone shows up and slaps you with a trout, it would improve the site. But it should be added to the terms of service first.
    • pennomi4 hours ago
      Is it really AGI if it can’t come to my address and slap me with a trout? Clearly AI is all hype &#x2F;s
    • Retr0id5 hours ago
      Alternatively, you can bet on your predictions. If you&#x27;re wrong, you lose money.
      • jayGlow39 minutes ago
        you know that&#x27;s not a bad idea there are a lot of people who are very confident on both sides of the argument. I wonder how many would actually be willing to put their money where their mouth is.
  • travisgriggs5 hours ago
    How was this assembled? From a meta point of view, how much AI was used to curate and highlite the goals; how much was used to assemble the site itself? Or deploy it?
  • eternal_braid4 hours ago
    A chess scoresheet sometimes contains mistakes but chess players can figure out in many cases what was meant by thinking of what moves make sense and considering the level of play so far. Popular AIs tools fail at that.
    • dllu4 hours ago
      Chess is an interesting case. I remember in 2023, GPT 3.5 or something used to be surprisingly good at chess. There was even a &quot;stochastic parrot chess&quot; website [1]. I recall it was playing decently at around a 1800 level. Even as a fairly okay player myself (2100 bullet on lichess), I struggled to beat it. However, modern LLMs are a lot worse at chess. I guess having too much chess data in the training set probably regressed performance on stuff that actually matters, like coding.<p>[1] parrotchess.com, no longer available. Previous discussions: <a href="https:&#x2F;&#x2F;hn.algolia.com&#x2F;?q=parrotchess.com" rel="nofollow">https:&#x2F;&#x2F;hn.algolia.com&#x2F;?q=parrotchess.com</a>
  • simianwords5 hours ago
    I made a bet with a guy on HN that the market value of OpenAI + Anthropic would get to at least 2.5T by 2027. I think I&#x27;m on track to winning.<p><a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=48517353">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=48517353</a><p>I also made a bet that API inference margins are greater than 10% for OpenAI and Anthropic<p><a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=48500827">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=48500827</a><p>I can make another prediction about Agentic Commerce and I think it will get big. Muse + Grok Bot + Dots.
    • ngruhn4 hours ago
      I love it! Kinda wholesome how that heated discussion ended with that bet.
    • mrweasel5 hours ago
      You&#x27;re very lucky that market value and actual value isn&#x27;t the same thing.
    • kridsdale14 hours ago
      But there is no market value pre ipo
    • mcphage4 hours ago
      &gt; API inference margins are greater than 10% for OpenAI and Anthropic<p>How do you measure that?
      • 4884485849 minutes ago
        With Amodei&#x27;s special accounting ofc
  • JBits4 hours ago
    Quite a few of the challenges revolve around asking for LLMs to complete tasks reliably and aren&#x27;t about whether an instance of an LLM completing the task exists. Quite a few of the goalposts are consequently completely changed without the surrounding context, are not the same as what the HN commenter requested and hence seem disingenuous to me.
  • tamimio3 hours ago
    Well I said that before AI will soon make the pcb and electronics just like code, it seems some hw engineers didn’t like it, months later there are few products about the same idea :)
  • simianwords5 hours ago
    [flagged]
    • mcphage4 hours ago
      &gt; that never believed that AI could solve Millennium problems (or same in spirit)<p>How did that situation end up? Did it solve it on its own, or did it rip off another mathematician&#x27;s work?
      • JBits4 hours ago
        I have to say, it&#x27;s hilarious to me that solving a Millennium problem has given mathematicians a reason to doubt the mathematical abilities of LLMs.
        • mcphage3 hours ago
          I don&#x27;t think it was the LLM solving a Millennium problem—it was the LLM solving a Millennium problem followed immediately by a mathematician claiming that their work had been ripped off.
          • JBits1 hour ago
            I agree. The idea that mathematical achievements by LLMs could involve plagiarism didn&#x27;t seem common before but now the question can be asked of any new novel proof of construction generated by LLMs.<p>It&#x27;s also notable that the Open AI proof may not even be interesting to mathematicians.<p>Even if people already had an idea that LLMs were training on user inputs, it&#x27;s the first time it&#x27;s actually caused an issue. Mathematicians, and plenty of researchers, working in ambitious or competitive field now have a very good reason to avoid LLMs.