5 comments

  • Johnny_Bonk51 minutes ago
    I haven't joined your chats in a while but glad to see you put this together, I truly feel as though opus 5 is not much of an improvement. The only time i ever felt a wow factor was opus 4, 4.6 and fable pre trump admin lobotimizing
    • dhorthy46 minutes ago
      yeah this was just a start - the fastest cheapest thing we could try for a brand new model.<p>I&#x27;m hoping to do some more work with sol&#x2F;fable in the mix as well as exploring more languages and curating the problem set to include more of the benchmark<p>I also kinda felt like opus4.5 was dumber than 4.1 personally, maybe a little biased since 4.5 was 2.5x faster and 2.5x cheaper seems to indicate its a smaller model
    • dan_gee29 minutes ago
      As someone who doesn&#x27;t use AI, this is totally incomprehensible to me.<p>You might as well be talking about the difference between smoking Maui Wowie and Grandaddy Purp.
      • scrollaway24 minutes ago
        How is this useful or insightful?<p>You ever go to forums full of entomology specialists and tell them you don’t understand their fancy terms?
        • dan_gee22 minutes ago
          My point is that the differences between these models are so minor that obsessively benchmarking them comes across as navel-gazing.
          • joatmon-snoo0 minutes ago
            The evidence that proves a model is actually a step function change is these benchmarks.<p>If a model isn’t a step function change? Welcome to research.
          • Johnny_Bonk4 minutes ago
            like all good science, measure everything
  • killingtime7431 minutes ago
    Did you not benchmark latest GPT 5.6 or GLM 5.1&#x2F;Kimi K3 because of cost? I can run them if you share how you ran them
    • dhorthy25 minutes ago
      no i&#x27;m spinning those up at some point this week. here&#x27;s the first few prompts I used (claude opus 5 as the research orchestrator), (these were interspersed with lots of tools and assistant messages but it should get you kicked off.<p>&gt; fetch this article for slopcodebench and help me run an eval on a subset of problems with opus 5 <a href="https:&#x2F;&#x2F;arxiv.org&#x2F;html&#x2F;2603.24755v1" rel="nofollow">https:&#x2F;&#x2F;arxiv.org&#x2F;html&#x2F;2603.24755v1</a> &gt; Get all the context, fetch any mentioned repos, and then propose a plan to me.<p>&gt; i have an anthropic API key in .... &gt; Let&#x27;s do the three challenges with Opus 4.8 and Opus 5 and Fable please. I like your minimal set. Let&#x27;s try it. What do you need from me?<p>&gt; Actually I changed my mind. I want to do two of the easy ones you picked and then I want you to pick the one with more checkpoints, maybe one of the harder ones, not the very hardest one but one with a higher number of checkpoints.<p>&gt; Actually let&#x27;s do one easy, one medium, and one hard problem please. If we have a hard problem I&#x27;d like to see that.<p>&gt; lets rock - i think lets just do opus 4.8 and sonnet 5 and opus 5 since we our ZDR will block fable
  • knighthacker5 minutes ago
    This is where Opus 5 shines
  • dcl53 minutes ago
    finally the benchmark for me
    • dhorthy44 minutes ago
      i hope that is because you hate slop and not because you write it
  • cute_boi13 minutes ago
    Opus 5 is an overconfident stupid model. It tries to generate too much slop, tries to act like everything will fall. I have reversed back to fable and codex sol.
    • dhorthy10 minutes ago
      yes sol is still my daily driver for most coding tasks<p>I did find opus 5 quite handy for general knowledge work and visual design, without the cost of fable (e.g. the graphics in this post are made by opus 5)<p>but its not noticeably better than opus 4.8 in those regards, and I would not miss it if forced to go back to 4.8