9 comments

  • Gareth3214 hours ago
    OpenAI usage limits have been severely cut, and intelligence appears to be markedly declining, so I'm going to start trying these Chinese models seriously now. I don't mind if it takes longer. I just need the intelligence to predictably work the same way from day to day.
    • unsupp0rted4 hours ago
      Same- I pay $200&#x2F;mo for Codex but whereas I used to get a week&#x27;s work out of a weekly limit, now I get roughly 1~2 days.<p>I&#x27;ve stopped using Astra entirely and remain on Sol orchestrating Luna Xhigh, but it&#x27;s still not nearly a week&#x27;s usage for a week&#x27;s allotment.<p>And even then, whenever a new model is about to come out, it feels like the model I&#x27;m using is being dumbed down substantially.<p>I have no evidence for this and can have no evidence for this, but I can vote with my wallet regardless.
      • dangoodmanUT23 minutes ago
        see that&#x27;s interesting, because I&#x27;m prompting all day and I usually end up with 15-25% by the end of the week. I&#x27;m using Astra xhigh exclusively
      • Gareth3213 hours ago
        I strongly agree. Check out the Codex subreddit. Many empirical examples of Astra silently downgrading the models. One found Astra was silently using Luna Max (but still billing for Astra).<p>Even when I try to stick with Sol X&#x2F;High, my limits are at best half of what they were before Astra launched, and the intelligence has declined markedly.<p>I cancelled my $100 plan. This is absolutely absurd and frankly unusable now.
        • Muromec2 hours ago
          It feels bizarre reading about the amounts spent on it here and paying like 10 eurobucks a week for DS
          • christophilus53 minutes ago
            My spending with Deepseek was more than a Codex subscription, though. How are you keeping it to $10&#x2F;month? Super light usage?
          • f6v1 hour ago
            For me, DS Pro is still behind Sol. But I do think many people hitting the Codex limits in 1 or 2 days are doing something wrong.
            • sjbzbeiks18 minutes ago
              Entirely depends on the work asked and the repo involved.<p>When I’m doing work on a repo where I’m implementing a standard and the agents have to read the standard to keep from hallucinating my usage skyrockets.<p>Hell this changes depending on which language I’m working with.
          • Gareth3212 hours ago
            These tools were pretty great if you could afford them, but now they are expensive <i>and</i> shit, and that combination doesn&#x27;t work.
    • loloisi1 hour ago
      Sol 5.6 xhigh had been a very reliable workhorse for coding for me via the 200 bucks sub.<p>But this week they seem to have tweaked the system to a point at which all models (Astra, Sol, Luna) hit rate limits all_the_time without me being anywhere close to the weekly limit.<p>Early results with MiMo 2.6pro are quite encouraging for anything that&#x27;s non-UI work so likely switching spend for the time being
    • phoghed2 hours ago
      &gt; and intelligence appears to be markedly declining<p>Serious question: does anyone have evidence of this?<p>It’s something that’s constantly asserted, and has been since 2023. Every time someone posts a site that tries to track this though, I look at it and it’s just a flat line.
  • deanc11 minutes ago
    Why is it labelled as #1&#x2F;114 for artificial intelligence (at the top of the page) but if you scroll down and look at the graphs it&#x27;s obviously not?
    • coder5434 minutes ago
      [delayed]
    • MallocVoidstar1 minute ago
      Hover the #1&#x2F;114 and it shows only open weights, that&#x27;s probably why
      • deanc1 minute ago
        Feels a little pointless and disingenuous to present it this way (as default) unless you specifically filter it as such
  • egeres5 hours ago
    It feels suspicious that MiMo-V2.6 Pro gets 46 in de index while DeepSeek-V4.1 (<a href="https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;models&#x2F;deepseek-v4-1-flash" rel="nofollow">https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;models&#x2F;deepseek-v4-1-flash</a>) gets 39. According to the appendix at the bottom of <a href="https:&#x2F;&#x2F;mimo.xiaomi.com&#x2F;mimo-v2-6" rel="nofollow">https:&#x2F;&#x2F;mimo.xiaomi.com&#x2F;mimo-v2-6</a> the deepseek model sometimes surpasses mimo and it&#x27;s not so far behind in capabilities. A week ago opus 5 appeared 1 points ahead of fable 5 despite fable being a much smarter model (this has been corrected already)
    • SyneRyder3 hours ago
      The main AA benchmark keeps changing, and had to be radically changed when Astra came out and showed zero improvement over GPT 5.6 Sol in their benchmark. Opus 5 is still 1 point ahead of Fable 5.0 on the index, if you manually add Fable 5.0 back into the list, so it hasn&#x27;t actually been &quot;corrected&quot;. It&#x27;s only Fable 5.1 that is shown as ahead of Opus 5.<p>The AA benchmark is a weighted average of other benchmarks and some internal ones. I think the difficult part is finding benchmarks that reflect your own use of the models.
      • seahorseemoji1 hour ago
        The way Artificial Analysis keeps changing their weights feels kind of like deciding who the winner should be and making the weights reflect that. They’ve been changing their weights to add more weight to improved long-running agentic capabilities, but doing so means they’re reducing the relative importance of world knowledge and of writing ability.<p>I’ll grant that maybe world knowledge isn’t that important for these models. But writing ability is important for human understanding, and I think the weird turns of phrase and word choices reflect the labs’ underweighting of the importance of human understanding.
      • yt199825 minutes ago
        [dead]
    • GodelNumbering3 hours ago
      &gt; It feels suspicious that MiMo-V2.6 Pro gets 46 in de index while DeepSeek-V4.1 gets 39.<p>Why?
  • segmondy39 minutes ago
    Mimo2.5 is really good, but tended to loop too much for my taste. Locally, Pro2.5 wasn&#x27;t much better. I would reach for it for one shots, hopefully they sorted it out with v2.6, it&#x27;s a model that&#x27;s slept on by many. I found that most people that used it did so because it was free. It&#x27;s a top model worth exploring if you have never given it a go.
  • dom965 hours ago
    It is an impressive model. Agreed on most that is written on this page, with the exception of it being fast. I ran it on my own LLM benchmark suite[1] and it is faster than DeepSeek but still much slower than leading models. But it&#x27;s pricing is where it really shines.<p>KillSwitch-Bench 1.0<p><pre><code> Claude Opus 5 66.9 GPT-6 Astra 57.9 Claude Fable 5.1 46.7 MiMo-V2.6-Pro 38.8 Muse Spark 1.3 36.5 </code></pre> 1 - <a href="https:&#x2F;&#x2F;bench.killswitch-lang.org&#x2F;" rel="nofollow">https:&#x2F;&#x2F;bench.killswitch-lang.org&#x2F;</a>
    • drittich6 minutes ago
      Interesting - have you done MiMo-v2.6-flash?
    • ricardobeat1 hour ago
      Speed seems to vary a lot with demand. Last night it was reaching 80+ tok&#x2F;s
  • tensegrist1 hour ago
    where&#x27;s the flash model? it&#x27;s out already isn&#x27;t it
  • jampekka6 hours ago
    &quot;When evaluating the Intelligence Index, it generated 140M tokens, which is somewhat verbose in comparison to the median of 140M.&quot;
    • throwa3562626 hours ago
      Nowadays these error can be a good thing :)<p>Human error means this wasn&#x27;t just stopped together by some bot.
      • jampekka6 hours ago
        My bet is that it&#x27;s a bot error, but of a rule based one.
        • theanonymousone5 hours ago
          Yes, it seems they have a template that they fill with numbers. Similar issues spotted on Grok&#x27;s performance page: <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49789558">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49789558</a>
  • ignoramous3 hours ago
    Per Xiaomi, MiMo v2.6 training run cost $3.47m. A far cry from the estimated costs ($100m+) for the Big 5 (MSL, xAI, GDM, OAI, Ant). I wouldn&#x27;t be surprised if salaries and R&amp;D costs have similar drastic disparities.<p>For a model that matches <i>Muse Spark 1.3</i> in benchmarks, <i>MiMo v2.6 Pro</i> is incredibly cheap, given its cache rates will remain $0.0036 per million.
    • imjonse3 hours ago
      That is the RL training cost only. Their announcement blog mentions this: <a href="https:&#x2F;&#x2F;mimo.xiaomi.com&#x2F;mimo-v2-6#scaling-rl-fully-open-sourced" rel="nofollow">https:&#x2F;&#x2F;mimo.xiaomi.com&#x2F;mimo-v2-6#scaling-rl-fully-open-sour...</a>
    • drbscl3 hours ago
      My understanding of tech salaries in China is that they are pretty decent, but not as high as in SF; closer to typical European salaries.<p>Mostly due to lower cost of living; Shenzhen is way cheaper than SV
      • f6v1 hour ago
        I seriously doubt salaries are included. It must be just the electricity and GPU costs.
    • NortySpock3 hours ago
      I sorta got the impression that the $3.47 million only covered post-training , given that few of the graphs start at zero. Is a barely-trained model going to score 48 on DeepSWE v1.1 ?<p><a href="https:&#x2F;&#x2F;mimo.xiaomi.com&#x2F;rl&#x2F;" rel="nofollow">https:&#x2F;&#x2F;mimo.xiaomi.com&#x2F;rl&#x2F;</a>
  • kosolam5 hours ago
    Why sol is not in the comparison?
    • guelo3 hours ago
      the graph has a dropdown for selecting models
      • kosolam2 hours ago
        Not on mobile unfortunately
        • guelo2 hours ago
          Not the graph at the top, the one further down.