167 comments

  • vishvananda23 hours ago
    The reason people aren’t freaking out is because most people are using heavily subsidized subscriptions.<p>I tried the cheapest provider on openrouter and burned through $50 in a few days. Quality was ok, seems slightly above Luna quality perhaps? But that $50 is 1&#x2F;4 of my codex subscription where I could have burned that many tokens or more using Astra within my weekly reset.<p>This won’t last forever but as long as the frontier labs are subsidizing this heavily the open models won’t matter.
    • Aurornis6 hours ago
      Deepseek also came with heavily subsidized plans, at least at first.<p>This is why there are so many comments confused that you spent that much money on Deepseek. Those who got on one of the discounted plans could use very large numbers of tokens for trivial prices.<p>There is also a strange double standard for accounting for open models. People will look at OpenAI or Anthropic and say that we need to consider all of their training costs and employee compensation when thinking about the cost to serve their models, but when the models are released as open weights those costs are ignored. So in that way, the Deepseek models are heavily subsidized as well, with the possibility of them being served by a different company that paid nothing to develop them.<p>&gt; Quality was ok, seems slightly above Luna quality perhaps?<p>I agree that it’s about in line with what you can get from Luna or Haiku, but I give the edge to Luna and Haiku when it comes to tasks that require world knowledge. It feels like their training sets were just cleaner.<p>Luna and Haiku are also close to free with a subscription plan, and they’re even very cheap at API rates.<p>So the amazing thing about Deepseek Flash is that you can almost kind of get that level of performance from an open weight model. It’s not as amazing when you start comparing it for how we really use smaller models from frontier labs on subscription plans.
      • mentalgear3 hours ago
        &gt; So in that way, the Deepseek models are heavily subsidized as well, with the possibility of them being served by a different company that paid nothing to develop them.<p>You could then of course argument that oAI&#x2F;ant&#x2F;G&#x27;s models were also heavily subsidized by using (scraping) the bulk of humanity&#x27;s global knowledge for free while paying nothing for it and trying to privatize it.
      • gpugreg2 hours ago
        <p><pre><code> &gt; Deepseek also came with heavily subsidized plans, at least at first. </code></pre> Do you have a source for that?<p>I got a source straight from the hoses mouth, which claims that DeepSeek-R1 was highly profitable:<p><pre><code> &gt; If all tokens were billed at DeepSeek-R1’s pricing (*), the total daily revenue would be $562,027, with a cost profit margin of 545%. </code></pre> <a href="https:&#x2F;&#x2F;github.com&#x2F;deepseek-ai&#x2F;open-infra-index&#x2F;blob&#x2F;main&#x2F;202502OpenSourceWeek&#x2F;day_6_one_more_thing_deepseekV3R1_inference_system_overview.md" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;deepseek-ai&#x2F;open-infra-index&#x2F;blob&#x2F;main&#x2F;20...</a><p>With DeepSeek-V4.1-Flash, the memory requirements for the KV cache have been reduced by over 50x and FLOPS by over 5x, so it is even cheaper: <a href="https:&#x2F;&#x2F;arxiv.org&#x2F;pdf&#x2F;2609.19969" rel="nofollow">https:&#x2F;&#x2F;arxiv.org&#x2F;pdf&#x2F;2609.19969</a>
        • Aurornis2 hours ago
          Deepseek Flash plans were very cheap when they first launched. They raised prices <a href="https:&#x2F;&#x2F;finance.yahoo.com&#x2F;technology&#x2F;ai&#x2F;articles&#x2F;deepseek-raising-api-prices-1-174027670.html" rel="nofollow">https:&#x2F;&#x2F;finance.yahoo.com&#x2F;technology&#x2F;ai&#x2F;articles&#x2F;deepseek-ra...</a><p>Your link is about a different model.<p>The talk about subsidies is confusing because for OpenAI and Anthropic people usually include their salaries, training costs, and everything else. When the topic switches to open weight models we pretend the models appeared out of the ether at zero cost, and the only cost is running the servers.
          • Bnjoroge2 hours ago
            Right, but I&#x27;m not sure that was because they were subsidizing it, or because they ran into capacity issues and needed to offload some users.
      • awongh6 hours ago
        &gt; So in that way, the Deepseek models are heavily subsidized as well, with the possibility of them being served by a different company that paid nothing to develop them.<p>Has anyone seen the neo-cloud profit margin on deepseek? I wonder if the deepseek served API prices account for the training cost? Because they don&#x27;t &#x2F; have not raised the money to fund their future operation &#x2F; training- and they depend on that cashflow for now? That would suggest a really high profit margin on inference only.
      • FooBarWidget4 hours ago
        Deepseek never subsidized their plans. At some point they dropped pricing, then after a while they raised them, but that&#x27;s because they had capacity issues, not because of dropping subsidies. They don&#x27;t want your money, they&#x27;d rather have fewer users.<p>OpenCode is building their own inference and they&#x27;ve independently stated that DeepSeek&#x27;s old (cheaper) pricing is achievable without subsidies.
    • carsoon19 hours ago
      Yea if i use opus 5.5 in api through openrouter and pi agent harness I will easily burn 50-100$ a day (and with fable 5.1 i could burn 200$ easily). Whereas i have now been using a claude code subscription for 2 weeks using 5.5 at all times and have never hit a limit. I often run 6+ agent sessions at once.<p>I do think its important long term to not be reliant on these companies as you don&#x27;t have control over the system prompts, the thinking tokens, and once the subsidization stops or the company is public they will be required to start making money and thus raise prices.<p>But models may get more intelligent and cheaper once that time comes so it may be a non issue.
      • boorang15 hours ago
        FYI you can modify the system prompt using mitmproxy. Just ask your agent to walk you through it. Anthropic system prompts are gnarly and geared towards the lowest common denominator.
        • InsideOutSanta12 hours ago
          <i>&gt; FYI you can modify the system prompt using mitmproxy</i><p>The system prompt is injected into your context on the server side.
          • boorang12 hours ago
            there is literally a --system-prompt flag for claude code. In my mitmproxy experiments using that flag appended to the existing system prompt rather than replacing it. So I had to create a little helper to strip the system prompt sent over the wire and add my own.<p>here is a collection of public system prompts from claude code: <a href="https:&#x2F;&#x2F;github.com&#x2F;Piebald-AI&#x2F;claude-code-system-prompts" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;Piebald-AI&#x2F;claude-code-system-prompts</a>
            • sunaookami12 hours ago
              Use --system-prompt-file, it will replace the whole system prompt (<a href="https:&#x2F;&#x2F;code.claude.com&#x2F;docs&#x2F;en&#x2F;cli-reference" rel="nofollow">https:&#x2F;&#x2F;code.claude.com&#x2F;docs&#x2F;en&#x2F;cli-reference</a> ). You&#x27;ll have to use a shell alias or function or something to always append this flag when calling Claude Code but you don&#x27;t need any MitM shenanigans. Then use &quot;&#x2F;context all&quot; to see what else is sent (here I would recommend MitM&#x27;ing since Claude Code won&#x27;t show the exact tools and text), there are a lot of tools no one needs and they are bloating the context, you can deny these in the settings.json (there is also a list here: <a href="https:&#x2F;&#x2F;code.claude.com&#x2F;docs&#x2F;en&#x2F;tools-reference" rel="nofollow">https:&#x2F;&#x2F;code.claude.com&#x2F;docs&#x2F;en&#x2F;tools-reference</a> ). Also set &quot;disableClaudeAiConnectors&quot; to false to remove even more bloat.
              • boorang11 hours ago
                I tried both --system-prompt and --system-prompt-file. they both appended when i tried about 6 months ago and watched the traffic. Yes I cleanup all those tools etc.
                • sunaookami8 hours ago
                  --system-prompt-file replaces the prompt, I use it myself and the docs state it, too.
        • iririririr13 hours ago
          thinking you can control what happens to some strings you send over the network.<p>...the absolute state of the token maximizers.
          • RGamma8 hours ago
            Claude, think for me. Make no mistakes.
      • tripleee18 hours ago
        I dont think itll be an issue. Opus 5.5 now is way more than enough for me and open weight models will reach that level by the time subsidization stops
        • seemaze15 hours ago
          640K [RAM] ought to be enough for anybody!
          • wredcoll15 hours ago
            People were perfectly happy with 640k ram at the time that phrase was uttered.
            • InsideOutSanta12 hours ago
              <i>&gt; People were perfectly happy with 640k ram at the time that phrase was uttered.</i><p>I&#x27;m assuming this is a sarcastic response to the person who said <i>&quot;640K [RAM] ought to be enough for anybody!&quot;</i><p>Because, as somebody who was around when we had 640 K RAM, people certainly weren&#x27;t happy with that amount.
              • Kurtz798 hours ago
                Ah, the hours spent tweaking CONFIG.SYS and AUTOEXEC.BAT to run a MS-DOS game that squeezed every last byte of those 640K, good times.
            • mathgeek13 hours ago
              Regardless of the time, people have always happily accepted more.
              • conductr12 hours ago
                I’d argue most “people” haven’t actually needed more in a very long time other than to keep the same old software running. The requirements bloat of operating systems, browsers, and majority of software isn’t really a generalized “people” thing; more so the state of the industry being a form of inertial bloat.
            • mkl14 hours ago
              The phrase probably <i>wasn&#x27;t</i> uttered by anyone it&#x27;s attributed to: <a href="https:&#x2F;&#x2F;quoteinvestigator.com&#x2F;2011&#x2F;09&#x2F;08&#x2F;640k-enough&#x2F;" rel="nofollow">https:&#x2F;&#x2F;quoteinvestigator.com&#x2F;2011&#x2F;09&#x2F;08&#x2F;640k-enough&#x2F;</a>. Even &quot;the time&quot; is unknown.
              • soperj13 hours ago
                And Bill probably _didn&#x27;t_ do anything on Epstein Island...
          • shuwix14 hours ago
            [dead]
        • komali217 hours ago
          It&#x27;s enough for you now, but I feel like part of the mythology of our future is that we&#x27;ll be continued to be employed because we&#x27;ll be working on more complex problems, with smarter LLMs at our side.
          • antesaj6 hours ago
            I do think there&#x27;s a decent case to be made that this will be the case for the foreseeable future, and perhaps the mythology part is beyond that.<p>Although as soon as I wrote &#x27;foreseeable future&#x27; I came to the realization that this is far far far less far out than it used to be. Which might mean I simply agree with you.
      • magicalhippo19 hours ago
        &gt; Yea if i use opus 5.5 in api through openrouter and pi agent harness I will easily burn 50-100$ a day (and with fable 5.1 i could burn 200$ easily).<p>Checked yesterday, for that day alone I had used $168 worth on my $20 subscription in Claude Code. I still had plenty of weekly use left. Seems like subscriptions are discounted at a 1:10 rate?
        • _aavaa_18 hours ago
          See <a href="https:&#x2F;&#x2F;newsletter.semianalysis.com&#x2F;p&#x2F;anthropic-subscriptions-offer-5x" rel="nofollow">https:&#x2F;&#x2F;newsletter.semianalysis.com&#x2F;p&#x2F;anthropic-subscription...</a>
          • arcanemachiner14 hours ago
            Nice, good to see an up-to-date reference on this.<p>EDIT: That was a delightfully thorough analysis. Nice to see Anthropic taking the value crown, only because Opus 5.5 is such a joy to use.
            • _aavaa_5 hours ago
              Only think missing from the analysis, and it&#x27;s a big one, is a comparison on the basis of work done. For a fixed set of real tasks, how much of your quota (or api usage) does it take you to accomplish it.<p>Because a big part of the price difference is due to Anthropic&#x27;s models being more expensive that OpenAIs AND using more tokens for the same tasks (at least according to artificialanalysis.ai).
          • wrynn13 hours ago
            Curious to know how much tokens&#x2F;task will change the conclusion
        • sidehub7 hours ago
          [flagged]
      • gxs19 hours ago
        100%<p>I use a personal Claude account for personal projects and can let rabl run for an hour and barely make a dent into my usage<p>On the enterprise I have to be a lot more careful or I can burn through 2k in a week<p>The excuse they give is the guarantees you get with enterprise plans that they won’t look at your data
        • smw19 hours ago
          What&#x27;s rabl?
          • watercolorblind14 hours ago
            Other than their typo there is a fun collision there with an older Ruby gem: <a href="https:&#x2F;&#x2F;github.com&#x2F;nesquena&#x2F;rabl" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;nesquena&#x2F;rabl</a>
          • gxs17 hours ago
            Meant fable*
      • devmor19 hours ago
        [dead]
    • georgel22 hours ago
      I am curious how you managed to spend that much on Deepseek via OpenRouter. I loaded $100 back in July while using v4-flash or whatever the cheap good model was at the time, and have upgraded as the new ones came out from Deepseek. I still have $16 and some of that spend also goes towards the AI usage from my customers (the context they need to load in is quite large too).<p>And I am using the Claude Code harness with DS as the endpoint. And I use it ~5-8hrs a day to do my coding.
      • whstl21 hours ago
        It&#x27;s wild how different usage patterns are between users.<p>I have seen the Cursor leaderboard on my company and the vibe coders consume about 5x more tokens than the developers. They and other office workers also have Claude and their limits are often over around Wednesday.<p>People are using millions of tokens to do very simple HTML reports. I have seen someone asking the LLM to download the entire data into the context and asking it to sort.<p>Those usage patterns don&#x27;t correlate to output.
        • Gigachad16 hours ago
          This reminds me of when we gave clients the ability to build their own Power BI dashboards. The users would end up doing a full table dump multiple times for the same table in their reports. Requiring 16GB of ram on the server and maxing out the database every time the report was refreshed.<p>We ended up having to hire a full time employee to fix the performance of client built reports.
          • whstl12 hours ago
            I&#x27;ve seen the same problem with Snowflake integrated with Claude.<p>People run a stupid amount of expensive queries that end up costing way too much because they&#x27;re asking Claude the wrong query.<p>Not to mention people running <i>wrong</i> queries, using the result as gospel, and then the result has to be sent to a data analyst to be reverse-engineered so the numbers make sense.
        • Quothling20 hours ago
          Now that Microsoft allows you to do &#x2F;cost for individual tasks (or whatever they call them today). So I tasked Sol, Astra and Fable in cowork with exactly the same vibe coding task on the exact same zip file containing a code project I needed an update for. Astra used 20x and Fable used 15x of what Sol did.<p>Fable changed a lot of things I had explicitly told it not to change. Arguably a lot of them would&#x27;ve been correct if you didn&#x27;t work in a place where abstractions are directly against the core principles, but what it produced was basically unusable. I&#x27;m not sure if Sol or Astra did best, they produced rather similar code outputs. Astra&#x27;s was better, but Sol didn&#x27;t do so bad. It forgot to clean up a few places after it&#x27;s refactor and it made two bugs I had to correct but other than that it was fine. Astra on the flip-side might have produced code that didn&#x27;t need changes but it also rewrote every piece of documentation so that it became horrible.<p>As far as the &quot;experiment&quot; goes, it just shows you that the credit consumption is basically pure magic. You&#x27;d think that the Microsoft AI admin tools and the Agent365 FOMO DLC license they sell might give you some sort of reporting, but it doesn&#x27;t. What you can see is how many tokens a user consumes and the total number of tasks they&#x27;ve initiated as well as whatever running agents they have. You can&#x27;t see what models they use or which tasks are expensive, which makes it very hard to help them. Early on we had an employee who hit their limit in an hour, and it turned out they had basically uploaded a lot of information and run it in a single long task that kept going over it again and again. We told them it might be a good idea to only give it what it needed and to create more tasks, and even though it&#x27;s been three months, they have yet to consume as many credits as they did that first hour.<p>But that&#x27;s how you support and track it. You see a user spend a lot, then you go to their computer and now that you can actually do the &#x2F;cost thing, you go through their tasks and try and figure out where they&#x27;re spending money...<p>It&#x27;s obviously improving. A month ago &#x2F;cost wasn&#x27;t there and they just released a new dashboard for cowork, but it&#x27;s still black magic that is impossible to govern.
          • KoolKat2310 hours ago
            I&#x27;m assuming you&#x27;re using Microsoft Cowork or Copilot or something. I suspect it&#x27;s system prompt and implementation (tooling&#x2F;harness) isn&#x27;t the same as that of OpenAI and Anthropic even if the underlying model is supposedly the same. The answers provided can be wildly different and often for the worst.
            • Quothling9 hours ago
              It&#x27;s cowork. I&#x27;m in enterprise in the EU, so we&#x27;re more restricted in what we use. My issue is mainly with how impossible it is to do any form of reporting on this. Obviously this isn&#x27;t the most popular opinion among the people subject to it, but for most things we do in the Microsoft enterprise setup we can basically monitor every thing a device does. With Cowork you can&#x27;t even see which individual task is eating a users credits. At least not yet.<p>Which is an issue when you need to get department managers to manage their budgets around the amounts of credits their employees spend. The more of a black box it is, the more governance and corporate bullshit you have to deal with.<p>The copilot part of it runs &quot;unlimited&quot; on the license. Except it&#x27;s not unlimited, and this is even more of a blackbox because you can&#x27;t see any sort of spending and the limit is listed as &quot;extensive use&quot;.
              • KoolKat236 hours ago
                Just btw, most assume when you say Cowork you&#x27;re referring to Anthropic Cowork as it&#x27;s the original rather than Microsoft&#x27;s white labelled Cowork. And when you say Sol or whatever model you&#x27;re referring to OpenAI and use on its servers. Microsoft&#x27;s versions are not at parity with the originals. I feel your pain on being limited by silly enterprise restrictions.<p>On your point about &quot;monitoring&quot;, personally I feel this is toxic corporate IT culture, enabled and perhaps pushed by the likes of Microsoft with all their tools, which they of course make money off. People have cellphones with cameras making most points in this area moot.<p>An alternative used in other big corporates is to set budgets, with tiered authorisation approvals for higher limits. The users and their managers can justify why and what they&#x27;re doing that they need the additional tokens. This also encourages more efficient use of tokens on other work. More efficient use is sometimes counterintuitive. Laissez-faire generally works best.<p>Cowork requires user approvals for high risk actions such as emailing.
        • mk8919 hours ago
          It&#x27;s because this is how AI has been sold to everyone - just ask, and it will do it.<p>The better pattern is to let it code the app and then you can use the app to target your data. So you only pay for it once, plus it&#x27;s deterministic. But yeah, it requires setting up an environment, etc. It becomes &quot;maintenance&quot;.
          • fourthark16 hours ago
            It requires knowing how to code, at some level.
          • maigret10 hours ago
            You have to keep it secure, so it’s not “pay once” but has running costs, and need some kind of security scanning, citizen developer devops platform etc etc
        • cindyllm19 hours ago
          [dead]
      • lifeisloving22 hours ago
        Likely the user doesn&#x27;t know what they&#x27;re doing or has extermely bad workflows. They&#x27;re prob not managing their cache, and dont use compaction.. Letting context get to 500k and invalidating their cache every 10 tool calls because they have no providor fallback settings.<p>I was running deepseek v4.1 pretty much non stop during work hours, with heavy tool&#x2F;mcp usage and finding it very difficult to spend more than $75 in a month.<p>Also the cheapest providers on Openroutrr can often have terrible cache hit %, short TTLs resulting in their effective price being much more expensive than people realize. 75% cache pretty much destroys any savings from a super cheap token perspective.
      • eru17 hours ago
        You can easily burn through lots and lots of money on DeepSeek, if you do eg large scale code reviews.<p>Eg I&#x27;ve used Sashiko locally for Linux kernel code reviews before sending out my contributions out to the world. Sashiko is a great system, but it can burn through tokens like there&#x27;s no tomorrow.<p><a href="https:&#x2F;&#x2F;github.com&#x2F;sashiko-dev&#x2F;sashiko" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;sashiko-dev&#x2F;sashiko</a> and <a href="https:&#x2F;&#x2F;sashiko.dev&#x2F;" rel="nofollow">https:&#x2F;&#x2F;sashiko.dev&#x2F;</a>
      • vishvananda15 hours ago
        I do some pretty insane agentic work running constantly. I’m currently burning through multiple $200 accounts every week. I regularly spend $10-$20k worth of tokens a month. Most of it has been going to my experimental c++ compiler project
    • bob102912 hours ago
      &gt; I tried the cheapest provider on openrouter and burned through $50 in a few days.<p>I wish I could observe how some of us are using these tools.<p>I still struggle to spend $50 in tokens per <i>month</i>, and I exclusively use prepaid API tokens. This is in support of personal projects and two clients. There are billing cycles where I might spend upward of $400, but this is maybe once a year. This is offset by months like August wherein I spent $12 in tokens.<p>The other advantage with prepaid is that it handles the other direction much better. I don&#x27;t even know what a quota limit feels like. Being blocked for hours is <i>way</i> more expensive to me and my clients than even $500&#x2F;m. Losing an entire business day over this wouldn&#x27;t work out.<p>I&#x27;ve managed to convince some others to try the same thing. $200&#x2F;m flat fee is a pretty extreme constant expense if you can be more clever on average.<p>I think a lot of people are getting pushed around by FOMO effects into spending money on pointless subsidized tokens and have (valid) fears that if they don&#x27;t maintain the same apparent economic leverage as their peers that they will be left behind. This isn&#x27;t actually the case, much like lines of code are a really poor indicator for the quality or productivity over a codebase.
      • spacedcowboy12 hours ago
        I think it depends on how many threads you have running at the same time. I have Claude writing a compiler in one window, a ui framework for the language in another, an application using the installed versions of compiler and frameworks in another, and a ui designer (an Interface Builder lookalike) in another keeping up with the framework.<p>Each of these has a file it listens to in ~&#x2F;tmp&#x2F;&lt;name&gt;.io and whenever one needs something from the other, they message each other via that file. Tasks can bounce back and forth as issues are resolved and tested. At the same time, I keep each busy with a list of tasks<p>I can fairly easily run out of my $200&#x2F;month subs every week, if I let Fable be the default model. With Opus it’s less likely. If and when I do, I just have an alias ‘claude.ds’ which fires up Deepseek instead, and burns through far less money, though I don’t think it’s as good at solving problems, just MHO.
        • bob102911 hours ago
          &gt; whenever one needs something from the other, they message each other via that file. Tasks can bounce back and forth as issues are resolved and tested.<p>Why isn&#x27;t this just one coherent agent loop with subtools&#x2F;agents as appropriate? If these tasks are related in some way, having a single context would probably make it go much better.<p>The freewheeling messaging part is where the token bloat is coming from. I suspect that for some of us this is actually the point. I think it&#x27;s a mostly form of entertainment to do things this way. The next logical step from Factorio gameplay.<p>Parallel agents remind me a lot about multi core compute. It&#x27;s incredibly easy to take a single core product and make it run much worse across a lot of cores.
        • ChickeNES11 hours ago
          Why not use the native inter-session messaging in Claude Code?
        • sidehub7 hours ago
          [flagged]
      • ubercore12 hours ago
        I don&#x27;t think too hard about usage one way or another -- not tokenmaxxing, not avoiding AI. Regularly use up over $1500&#x2F;m without really trying.
        • emn1312 hours ago
          Which models do you primarily use, and can you very roughly list your process? Agentic coding in VSCode with tons of MCPs or... something else? Do you include lots of images or have large codebases? Which agentic harness are you using?<p>I also find that it&#x27;s easy to spend like that, but also easy not to with little impact on productivity. At the current moment I&#x27;m stuck with rider + copilot (not ideal), but e.g. using GPT 6.1 luna is really, really cheap, and lots of tasks are quickly and decently dealt with even at lower reasoning levels, (added bonus of having low latency). And that model is so cheap, I can&#x27;t see a hitting 1500$ at api prices realistically - not even close. But it also depends on the harness and codebase.
          • ubercore12 hours ago
            I use opus primarily, on a mix of pure coding tasks, and log parsing &#x2F; incident investigation.<p>I don&#x27;t have the mental capacity to do a lot of context switching between active work streams, so I&#x27;m not doing stuff like leaving a big agent workflow running while doing other things.<p>All through claude code.
            • j_maffe11 hours ago
              Perhaps you&#x27;re letting the chat context reach 100%? I suppose that&#x27;s a way that would drive up spending.
              • ubercore11 hours ago
                Nah, I&#x27;m typically &lt;20% context before I clear and continue
      • m4tthumphrey12 hours ago
        Same... I use the £90&#x2F;month Claude sub and it&#x27;s more than enough. Can&#x27;t comprehend the users spending thousands a month.
    • pimeys22 hours ago
      Yes. You have to find the provider with pricing that suits your usage.<p>I am having 98% my input in cache, so using Coralbricks makes sense due to them giving cache reads for free — you only pay for writes. I spend maybe 5-10 dollars a day and my agents basically work day and night implementing things for me.<p>If your tasks are write-heavy, find a provider with cheaper output.<p>If you build a customer-facing app, pay a bit extra for 400+ tok&#x2F;s e.g. on Lithos.
      • sheeshkebab21 hours ago
        What are your agents implementing day and night…sheesh. Built anything useful for anyone yet?
        • bigblind21 hours ago
          Sheesh... what&#x27;s in a name? :p
        • esseph21 hours ago
          Bio shows: <a href="https:&#x2F;&#x2F;twin.so&#x2F;" rel="nofollow">https:&#x2F;&#x2F;twin.so&#x2F;</a>
          • switchbak17 hours ago
            That promo video … it’s so cringe it feels like satire. But it’s not.<p>It’s PERFECT!
          • bethekidyouwant21 hours ago
            Codes day and night a snake eating its own tail…
            • redanddead20 hours ago
              Way too harsh<p>Lots of guys buying articles on TechCrunch saying they’ll build this, he’s bootstrapped
              • robk10 hours ago
                The hiring page says he raised $10m from LocalGlobe
            • Ronsenshi19 hours ago
              Funny how it&#x27;s always another AI API wrapper.
          • MiguelBuiles16 hours ago
            Holy shit, it&#x27;s an Andon Labs clone but worse. Because Andon Labs actually launched businesses (and lost money on all of them)
        • serf20 hours ago
          &gt;What are your agents implementing day and night…sheesh. Built anything useful for anyone yet?<p>you can&#x27;t think of anything to unleash some agents on within the entire digital world at any given time?<p>you motivate your own personal work only via gauging its&#x27; usefulness to others and your own prospects?<p>sheesh.<p>built anything for the sake of building yet?
          • svmegatron20 hours ago
            “building for the sake of building” refers to _you_ doing the building
            • boothby19 hours ago
              We&#x27;ve entered the idler era of building
              • drekipus17 hours ago
                I can&#x27;t believe I haven&#x27;t made that connection myself tbh
                • boothby16 hours ago
                  I&#x27;m not convinced that humans will remain entertained by this format. I hope not, anyway. That&#x27;s how we get WALL•E
                  • ChickeNES11 hours ago
                    &gt; That&#x27;s how we get WALL•E<p>God I hope so.
            • LeFantome18 hours ago
              We have entered the era of building enterprise scale projects for your own personal use. I can do things it took teams years to build as a hobby project over a weekend.
              • majormajor18 hours ago
                So much of that &quot;took years&quot; is because they didn&#x27;t know exactly what the &quot;years from now&quot; state they were building in advance. Another huge chunk is because they started getting customers and had to respond to customer needs, demands, scale, bugfixes, preserve uptime, etc.<p>And even then, &quot;enterprise&quot; was often a dirty word in these circles. The over-engineered would-be-swiss-army-knife vendor that was mediocre-for-everyone but excellent for nobody.<p>I have built many tools in the recent past for myself. None of them need to be &quot;enterprise scale.&quot; Most of them would be worse for it because the agent output suffers when the pile gets deeper and it&#x27;s just adding more piles on top.
              • i_love_retros18 hours ago
                What&#x27;s the most impressive and useful thing you have built then?
                • pylotlight17 hours ago
                  the very notion that everything has to be &#x27;impressive&#x27; to you is ridiculous. You still don&#x27;t understand the era of personal software and everything has to be a windows replacement or it wasn&#x27;t worth building to you? I&#x27;m making my own azure blob storage explorer, notes mac&#x2F;android, db client etc.. its not about being &#x27;impressive&#x27; but being personally suited to individual needs.
                  • imtringued14 hours ago
                    He said most impressive and useful meaning there is no absolute cutoff. The only reason why you would be offended is because you have built nothing at all and that&#x27;s you telling on yourself. Not to mention the question wasn&#x27;t even aimed at you.
                    • peteforde12 hours ago
                      No, that&#x27;s not what happened in this thread at all.<p>There&#x27;s a recurring pattern on HN where any time someone talks about knocking dozens of personal projects off their list - things that almost certainly would never have actually been addressed in the finite span of a normal life, the way things go - and you AI doomers show up and demand receipts as though that&#x27;s a total reasonable and definitely not obnoxious request.<p>It&#x27;s like if you tell someone that you love your partner and they demand to sit in the cuck chair or else you&#x27;re obviously lying. I keep hoping people will move past this &quot;prove that you&#x27;re actually productive&quot; reflex, but it just keeps happening in basically every AI thread.<p>In reality there are many reasons not to list out projects that you&#x27;ve worked on with LLMs, and while &quot;none of your damn business&quot; is always going to be at the top, the simple truth is that I want my products and projects to be judged by what they do and how well they work, not by how they were made.
                      • jodrellblank1 hour ago
                        &gt; &quot;<i>There&#x27;s a recurring pattern on HN where any time someone talks about knocking dozens of personal projects off their list</i>&quot;<p>Nobody was talking about that very believable use case. What they specifically said was &quot;enterprise scale projects&quot;. Enterprise scale means lots of users, decades of backwards compatibility, regulation compliance, logging and auditing, reporting, role based access control and permissions, integration with other enterprise systems, etc. etc.<p>&gt; &quot;<i>you AI doomers show up and demand receipts as though that&#x27;s a total reasonable and</i>&quot;<p>Asking what large scale programs they have built is not &quot;dooming&quot;. It&#x27;s also not demanding. It&#x27;s also not unreasonable.<p>&gt; &quot;<i>definitely not obnoxious request</i>&quot;.<p>Saying that it&#x27;s &quot;unreasonable and obnoxious&quot; to ask people to justify their claims is where the likes of Theranos and Nikola electric truck company are hiding. If they don&#x27;t want to talk about their stuff, they could have not commented. Since they commented it&#x27;s reasonable to ask them about what they said.
                        • booty33 minutes ago
                          &gt; What they specifically said was &quot;enterprise scale projects&quot;. Enterprise scale means lots of users, decades of backwards compatibility, regulation compliance, logging and auditing, reporting, role based access control and permissions, integration with other enterprise systems, etc. etc.<p>Do you really think that&#x27;s what they meant?
                      • wsng8 hours ago
                        &gt; I can do things it took teams years to build as a hobby project over a weekend.<p>That is a pretty strong statement. And without even anecdotal evidence, it becomes very weak.
                      • albedoa8 hours ago
                        You misread a comment so badly that you ended up responding to a completely different and made up comment. Your unwillingness to admit that you were wrong is way worse than whatever reflex you are ranting about.
                      • i_love_retros8 hours ago
                        &gt;I want my products and projects to be judged by what they do and how well they work<p>Hence the question &quot;what&#x27;s the most impressive and useful thing you have made&quot;<p>But no one really has any examples. Just half baked slop they never got over the finish line.
                  • i_love_retros8 hours ago
                    So software that already solves those problems isn&#x27;t good enough and your vibe coded versions will be better? And totally worth the negative cost of AI to society and the environment? You just gotta have your own versions right?
                    • adwn6 hours ago
                      Well, maybe, for certain programs, yes? Maybe I don&#x27;t want a million SLOC behemoth with 1000 features that has a huge attack surface, when a slim 10k SLOC tool that does exactly what I need (and nothing more) will suffice? A tool that I can modify and extend on a whim? A tool that <i>does one thing, and does it well</i> — you know, the <i>Unix philosophy</i>?
          • i_love_retros18 hours ago
            &gt;you can&#x27;t think of anything to unleash some agents on within the entire digital world at any given time?<p>It has to be worth it though right? Like I could spend some money and have agents build me my own Photoshop maybe (maybe?) But it would definitely be much worse to use than actual Photoshop. Then I have to have the continued interest to keep improving it which probably won&#x27;t happen because the next shiny thing will grab my attention. So it all just seems like a bunch of kids that have been given a seemingly endless supply of free candy and they are going fucking nuts like chipmunks with ADHD on crack. Building all this shit that is absolutely meaningless. I realize I&#x27;ve gone on a rant but I&#x27;ll keep going. I strongly suspect (with no evidence whatsoever) that the people who are churning slop apps out at breakneck speed have never been to an art museum. There. I said it. You&#x27;ve all got no taste. You wouldn&#x27;t know a quality product if it hit you in the face. I&#x27;ll leave with this thought- if apple didn&#x27;t exist, would they ever exist now we have LLMs? I say no, because the age of good taste and refined design and original thoughts is gone forever now that we have Claude and chatgpt and agents.
            • jamienk17 hours ago
              I think you&#x27;re underestimating how things used to be - you could go into any office, any closet, any coffee shop and find a shit-ton of half-baked, crazy-genius, kick-ass, retarded ideas and projects lying around everywhere in the world: filing systems, carpet organizing systems, outlines for film scripts, unsent letters to loved ones, etc etc etc. Whole <i>worlds</i> everywhere you look. And other people chipping in their two-cents worth, adding a few new filing cabinets, an idea for a film sequel, a new way to think about a different carpet, a notebook system to organize someone else&#x27;s unsent letters... All this slop eventually thrown into nasty, fetid garbage dumps, forgotten.
              • euroderf17 hours ago
                A.I. potentially breathing new life into every personal Graveyard of Half-Assed False Starts.<p>What&#x27;s not to like ?
              • i_love_retros8 hours ago
                All of that stuff meant something to someone though. They put real time and effort into those creations.
              • lelanthran13 hours ago
                &gt; All this slop eventually thrown into nasty, fetid garbage dumps, forgotten.<p>Isn&#x27;t that how it is <i>supposed</i> to be?<p>The lower the friction, the lower the signal:noise ratio.<p>It doesn&#x27;t matter if 1 out of every 100k slop projects is actually a humdinger, how on earth will you ever find it?<p>The value of a project is the commitment to to it by people. Slop projects indicates a commitment in the low to none range.<p>So, yeah, that AI-booster who &quot;created&quot; (I use that word loosely) 7x Adobe replacements in a week (none of which actually work, but he&#x27;ll get there eventually, I supposed) will successfully edge out the person who carefully and thoughtfully created a Photoshop replacement over six months of user feedback.<p>TBH, the only way to start a software business now is in stealth mode.
          • jimbokun20 hours ago
            [flagged]
      • switchbak17 hours ago
        I tried fireworks.ai, drawn in by their supposed blazing speed. Yawn. Mostly worse than vanilla Deepseek.<p>Is Lithos actually fast for common usage?
      • maxdo21 hours ago
        You have all the struggles for the price of Anthropic &#x2F; cursor subscription. I use the first one I code large chunks some PR are 50k LOC and I have at least 2-3 like this a week . It’s a greenfield project .<p>I still have quotas left I use it for home things build 3d model of my renovation projects, alerts for shopping list etc . And yeah I use cutting edge of cutting edge of models that saves me time and money , only discount monitor saved me ~$2k on my renovation project
        • jimbokun20 hours ago
          Why the hell would you put 50k lines into a single PR?<p>I mean, why even pretend you’re going to “review” something that large? Just build everything on main.
          • maxdo18 hours ago
            PR is just an entity to review &#x2F;do some other LLM processes . Human part check other models review PR&#x27;s, tests all kind of , security etc. Also human is to checks docs, specs in the pr, db migrations if any , some of the tests related to the PR. We stopped reading the code after opus 4.6. Sometimes for very core parts i skim through files just to make sure if the changes were correct.
            • jimbokun5 hours ago
              I would imagine even LLMs would do a better job reviewing smaller PRs than very large ones.
            • majormajor18 hours ago
              But to reiterate parent&#x27;s question: why not just do those things continually at that point? Or on a calendar-based basis?
              • imtringued14 hours ago
                Yeah exactly, if you have a fully automated SDLC, how do you expect the agent to code review 50k lines properly?<p>It will take shortcuts and now the entire premise is busted. You now need to build a code review process for large PRs.
                • manos-saratsis10 hours ago
                  This is a common problem. Fully automated SDLC needs to start before the CI&#x2F;CD.<p>My suggestion is to have the proper chunking mechanisms and multiple specialised agents. The most important is harness engineering, what we do at dromeas.ai to verify the code that goes to prod is a)have the code mapped before hand for the right agentic context, b)chunks of the right size per model context window c)specialised agents d)deduplication and verification . All before assessing a PR, a commit, a release. Harness engineering is not easy.. Especially when supporting multi model
      • taytus19 hours ago
        I agree. I code a lot, a lot! And maybe my code is shitty, but yeah, I burn a lot of tokens, and I couldn’t do it without Chinese models. I’m just a random dev in the middle of nowhere. And I don’t feel like I’m missing out on anything with my setup at all.
    • Silagi9 hours ago
      It&#x27;s far more likely they start to restrict the usage of subscriptions in corporate settings and leave the individual subscriptions to inflate the margin against them away. The individual subscriptions are the hook; people need to be able to understand what they&#x27;re capable of with enough tokens to tackle real projects they couldn&#x27;t have without the sub. That&#x27;s the only way individuals will make the attempt to sell it to their org.<p>We&#x27;ll probably still be getting the same amount of work done with a $200 subscription a year from now. That will just represent a much smaller subsidization than we currently enjoy; something like 3:1 - 6:1 instead of 40:1. Maybe running at a lower tps than the API gets. Maybe no access to the absolute frontier, but still significantly more intelligent than we get now. The labs will essentially break even on the subs and the corporate spending will be the profit center. Tale as old as software.
      • Schiendelman7 hours ago
        Corporate customers are by and large not allowed to use subscriptions.
    • onlyrealcuzzo22 hours ago
      OpenRouter is complete garbage.<p>Buy directly from DeepSeek&#x27;s API.<p>You can literally get overcharged 100x on DeepSeek on OpenRouter (or more).
      • patwolf21 hours ago
        One of the reasons I use OpenRouter is because they offer zero data retention. As far as I can tell, DeepSeek&#x27;s own API doesn&#x27;t support ZDR.
        • alok-g13 hours ago
          AFAIK, when you use DeepSeek via OpenRouter, it still does not have zero data retention.<p>See here:<p><a href="https:&#x2F;&#x2F;openrouter.ai&#x2F;providers&#x2F;" rel="nofollow">https:&#x2F;&#x2F;openrouter.ai&#x2F;providers&#x2F;</a>
          • randstring13 hours ago
            I have ZDR enforced and see only compatible models and providers, yet am able to use it. DeepSeek as a provider may not be ZDR, but the models are available from ZDR and no training providers on EU&#x2F;US servers.
        • selectodude21 hours ago
          DeepInfra does and it&#x27;s the same price. That&#x27;s what I use.
          • cmrdporcupine7 hours ago
            DeepInfra is decent but has a bad history of quantizing models
            • selectodude5 hours ago
              Yeah, they all do it though. And now that models are being natively trained at FP8&#x2F;NVFP4, I’m not sure it matters.<p>Only way you can really know you’re getting the full model is to host it yourself. Every inference provider has every reason to lie and it’s impossible to find out the degree to which they are.
        • koe12321 hours ago
          Youre doing something so special you need that?
          • antonvs20 hours ago
            It’s a pretty common requirement in the enterprise world. If you’re processing data for enterprise customers, it’s a lot easier to retain nothing than to deal with all the compliance issues that arise if you’re retaining data.
            • wolvoleo17 hours ago
              Ugh yeah try to get through a DPIA review!
          • rrr_oh_man20 hours ago
            I&#x27;ve got a toilet cam to install in your bathroom
      • georgel22 hours ago
        I&#x27;m all in for saving money and _can_ move to using DS directly from them, but maybe I am missing something here:<p>OpenRouter Pricing:<p>$0.02&#x2F;M input tokens $0.60&#x2F;M output tokens<p>DeepSeek Pricing (cache miss, off-peak):<p>$0.15&#x2F;M Input $0.60&#x2F;m output
        • girvo22 hours ago
          When 98.5% of my requests are cache hits (according to Pi for the last week), the cache miss price isn’t that important to me, and $0.003-0.006 per 1M input tokens is shockingly cheap.<p>It’s also the major difference between using DeepSeek directly vs other providers also serving it, though I have not looked lately: it’s possible other providers have matched its cache hit pricing better?
          • georgel22 hours ago
            Interesting, if the cache hit is that good, I think HN convinced me to toss $20 at DS official, and see how long that lasts.
            • girvo21 hours ago
              It will of course depend on what you’re doing with it, but right now my session at work has a 99.8% cache hit rate, and I’ve been running this session for hours with 23M tokens read and 713K tokens written (Opus 5.5 in this case though)
            • Muromec16 hours ago
              cache hit is in fact that good.
        • mswphd22 hours ago
          I&#x27;ve heard that certain inference providers may have different quality of caching implementations, so even if the listed numbers are as you say, the practical cache hit % you get might be significantly different&#x2F;incur significantly different costs.
        • ckdot21 hours ago
          There’s a big difference in speed &amp; quality between using DeepSeek API directly with DSH vs. DeepSeek in Opencode Go with Opencode CLI. Can’t tell if it’s the provider or the harness - but worth to give it a try.
          • nextaccountic3 minutes ago
            what about dsh + openrouter? and configuring open router to just serve from DeepSeek own servers<p>I think the 5% cut from open router is fair if I want to user other cheap models like mimo
      • BeetleB21 hours ago
        DeepSeek trains on your inputs. That&#x27;s why people go on OpenRouter and choose ZDR providers.
        • unicornfinder17 hours ago
          This isn&#x27;t true at least on the API. If you read their privacy policy you&#x27;ll see the training clause is scoped specifically to the consumer terms i.e. for the chat product. No such clause exists for the API service, and it would absolutely be required under Chinese law if it was taking place.<p>Contrary to popular belief, DeepSeek really aren&#x27;t interested in your prompts.
          • BeetleB15 hours ago
            Do you have a link to their API privacy policy?<p>What I see is <a href="https:&#x2F;&#x2F;cdn.deepseek.com&#x2F;policies&#x2F;en-US&#x2F;deepseek-privacy-policy.html" rel="nofollow">https:&#x2F;&#x2F;cdn.deepseek.com&#x2F;policies&#x2F;en-US&#x2F;deepseek-privacy-pol...</a><p>They do not have a specific exclusion for API use.<p>I know Z.ai has an exclusion for API use. It&#x27;s widely reported Deepseek doesn&#x27;t.
            • unicornfinder10 hours ago
              I&#x27;ll preface by saying you&#x27;d be entirely reasonable in not finding this sufficiently reassuring, but compare the standard terms: <a href="https:&#x2F;&#x2F;cdn.deepseek.com&#x2F;policies&#x2F;en-US&#x2F;deepseek-terms-of-use.html" rel="nofollow">https:&#x2F;&#x2F;cdn.deepseek.com&#x2F;policies&#x2F;en-US&#x2F;deepseek-terms-of-us...</a><p>To the open platform terms (i.e. for API use): <a href="https:&#x2F;&#x2F;cdn.deepseek.com&#x2F;policies&#x2F;en-US&#x2F;deepseek-open-platform-terms-of-service.html" rel="nofollow">https:&#x2F;&#x2F;cdn.deepseek.com&#x2F;policies&#x2F;en-US&#x2F;deepseek-open-platfo...</a><p>The standard terms includes clause 4.3 which grants them the right to retain inputs and outputs for training purposes, and this is missing from their API terms. The standard terms also cover the right to opt out (which you can do from your user settings). No such opt-out exists on the API because it isn&#x27;t applicable.
              • BeetleB4 hours ago
                You&#x27;re linking to the Terms of Use. I&#x27;m linking to the Privacy Policy. Your Terms of Use link points to the Privacy Policy.
          • jeremyjh9 hours ago
            Deepseek does not offer zero training on OpenRouter. Out of dozens of alternatives, they are the only provider for 4.1 Flash that does this.
        • bethekidyouwant20 hours ago
          Let me get this straight you guys really like deep seek because it’s open but you don’t wanna help them improve.
          • bottled_poe20 hours ago
            No, we just want a choice on how to license our work.
            • vintermann14 hours ago
              All these products, western or Chinese, are built on a hell of a lot of running rough-shod over licensing or IP laws in general.<p>I&#x27;m writing a program I personally need, but I would be happy if there existed something like it already, if someone else vibecoded a better version of it than mine, or if DS got better at vibing this kind of thing.
            • miyuru13 hours ago
              &gt;license our work<p>what proof do you have, that they don&#x27;t train on your data?<p>they can say they don&#x27;t, but I don&#x27;t see any way for you to confirm it.<p>with how these companies operate currently, I won&#x27;t be surprised, if they say that one of agents &quot;mistakenly&quot; did that already..
              • vintermann12 hours ago
                Mistakenly and autonomously of course.
          • Larrikin19 hours ago
            Why does liking a product mean you have to give them all of your data? People are so outraged at LG because they make the best TVs and people wanted their expensive product, yet some MBA convinced them they could make more money by spying on your entire household all the time.<p>It actually was awesome in the early Facebook days where you could have your entire phone contacts and other apps filled out with a profile picture and Birthday by connecting them together. But that relationship has been completely abused, privacy has been invaded, and my data has been sold to multiple companies.<p>The goal going forward is to keep that data private. If your company can&#x27;t survive without it then I hope your company goes out of business
      • RussianCow20 hours ago
        Just pin your config to a single provider, or several providers with the params `order` and `allow_fallbacks: false`. I regularly get ~98-99% cache hit rates with OpenCode. And some providers are much faster than DeepSeek; I was getting 200-300 tokens&#x2F;second the other day with Together as my provider.<p>It&#x27;s regrettable that OpenRouter doesn&#x27;t even try to pin you to a single provider per session, but once you know about it, it&#x27;s a problem that&#x27;s easily solved.
        • kolinko12 hours ago
          Don’t you get cache expirations then if they change providers? This ought to introduce delays and costs
          • RussianCow5 hours ago
            Yes, but that&#x27;s why you pin them. If you specify more than one provider in `order`, OpenRouter will use the first one unless it&#x27;s down, so that&#x27;s the only time it would switch providers on you. And personally, I&#x27;d pay a few cents instead of waiting for the API to come back.
      • FusionX13 hours ago
        There&#x27;s no difference between the two if you pin the provider to Deepseek on Openrouter.<p>If you don&#x27;t want to mess about client side with pinning, set a guardrail on Openrouter that limits the available providers to only the official one.
        • sunaookami11 hours ago
          Then why bother using OpenRouter and paying the extra fees?
          • peheje2 hours ago
            Because you have like every model on the planet to choose from. So if a contender drops you switch.<p>Also if DS is down you can choose another provider.<p>I have credits at DS and OR directly. But I do see the value in OR.
      • puchatek10 hours ago
        So you&#x27;re basically send your code to China?<p>I thought the advantage of DeepSeek is that you can host it on a server of your choosing.
        • wren69919 hours ago
          Yes, I send it to China, and they use it to improve open-weight models. I&#x27;m ok with this arrangement. At least, I&#x27;m happier with this than with companies using my open-source work to improve proprietary models without my consent.
      • jorvi21 hours ago
        Or.. BYOK Deepseek because OpenRouter&#x27;s UX is much nicer?
      • runtime_terror21 hours ago
        Zero Data Retention and not having company source code leak to &quot;CHINA!&quot; (said in Trumps annoying voice) would be two reasons not to
    • lionkor22 hours ago
      How&#x27;s the caching? I have 99.5% cache hit rate with deepseek when using their own API, it&#x27;s dirt cheap.
    • Eridrus18 hours ago
      People also just do different work. Opus 5.5 is a really damn good model that&#x27;s even better than Astra&#x2F;Fable&#x2F;Sol IME and I feel a huge difference in my work.
      • happycube5 hours ago
        Yup. I was going to scale down my Anthropic sub when Opus 5 was... weird, and there wasn&#x27;t enough Fable provided. 5.5, otoh, <i>chef&#x27;s kiss</i><p>That said, I do want good local(ish) capability for if&#x2F;when Anthropic enshittifies again. And to play with very useful smaller models - don&#x27;t even count gemma 4 out.
      • thefourthchime18 hours ago
        Same, it&#x27;s like another Opus 4.5 moment.
        • bitexploder17 hours ago
          Yep... and we already have Opus 4.6 at home (Qwen Flash Next 3.8).
    • loveparade18 hours ago
      This. At this point I don&#x27;t really care about other models because max subscription are super cheap (relatively speaking) and I don&#x27;t hit my limits. Even if the frontier models are only 5% better I might as well just use the best thing available if the price is reasonable.<p>Once the subsidization ends and cost becomes significant I will take a serious look around for the best value models and switch off the expensive providers, but that time hasn&#x27;t come yet.
      • switchbak17 hours ago
        I have a subscription at work, and still I find myself wishing I could use a fast Chinese model. Something wired up to really fast inference - that rapidity of feedback is a feature in itself.<p>4.1 Flash seems to be in that sweet spot of very decent, really fast and really cheap. Even omitting the cost, it’s still compelling for staying in flow.
      • wmedrano17 hours ago
        I&#x27;m actually shocked by how many people seem to have max subscriptions.<p>Something about renting that much compute doesn&#x27;t sit right with me so I stick with the $20 subs.
      • dragonwriter18 hours ago
        &gt; Once the subsidization ends and cost becomes significant I will take a serious look around for the best value models and switch off the expensive providers, but that time hasn&#x27;t come yet.<p>There&#x27;s a reason the labs in the US frontier oligopoly are using “safety” to lobby for antitrust exemptions for mutual coordination as well as anticompetitive regulation.
    • cm218712 hours ago
      Frontier labs subsidising? I thought they were running with 80%-ish operating margins which is part of why they cost 10 times open weight models.<p>Also isn&#x27;t an open weight model also subsidised? Training isn&#x27;t cheap and you are not paying for it.
      • marcosdumay6 hours ago
        Claude financials leaked last week, they are running with about -200% operating margins.<p>We also have OpenAI numbers, where people speculate that the about 150% of the margins spent on &quot;marketing&quot; is a fake line used to hide operational costs.
        • cm21874 hours ago
          But isn’t that including the cost of training, and I am not even sure if it’s training that very model?<p>In an interview earlier this year I remember Dario saying that the models are profitable, ie they more than pay back their inference and training over time. But because they invest in that explosive growth they have a deep negative cash burn.<p>And whatever markup the AI labs make, that’s on top of nvidia’s markup, micron’s markup, etc. It’s an industry where every supplier is adding a 70-80% markup!
          • marcosdumay3 hours ago
            AFAIK, no. For Claude, many people have reported that it&#x27;s not including the training costs. But I didn&#x27;t personally look at the data.<p>For OpenAI, absolutely not, that&#x27;s not including training costs. But I do personally accept that it can be actual marketing costs.
    • alexfortin20 hours ago
      I freak out since months for Z.ai lite subscription, I use glm-5.3-flash every day for a ludicrous 8.5USD&#x2F;month and it&#x27;s as good as DS 4.1 flash, if not better.
      • soniczentropy19 hours ago
        Almost exact same experience here, but I&#x27;m using Opencode&#x27;s $10&#x2F;month sub. It&#x27;s perma set to DS 4.1 flash and I have anywhere from 3-5 agents going at a time. Never once hit a cap of any sort. I have absolutely no idea why people would be paying $200&#x2F;mo when you can get perfectly good AI for $10 from multiple places
        • mapontosevenths18 hours ago
          Do they train on your data? I don&#x27;t necessarily want competitors to be able to just ask it to make a clone and have it do so from memory next month.
          • sieve17 hours ago
            Software has rarely been the moat. Or file formats. You have always been able to reverse engineer them. The problem, always, has been network effects.<p>I can build an entire, fairly useful, spreadsheet app over a weekend. But can I send my &quot;expenses.cells&quot; files to my accountant? Will it work with the Excel&#x2F;Google docs he uses?<p>AI can build or reverse engineer anything as long as you are motivated enough to do it.
            • Gigachad16 hours ago
              One example is Affinity 3 released for free. But it&#x27;s a huge pain in the ass because all the guides for how to do things are for Photoshop or Affinity 2.<p>Your vibecoded app won&#x27;t have years of reddit posts showing how to do things. This also seems to be where LLMs are the weakest at giving advice, they hallucinate 80% of the time I ask them how to do something in Affinity, giving buttons and menus that simply don&#x27;t exist.
              • sieve15 hours ago
                This is a knowledge problem. You can fix it by pointing the model to documentation (if it exists). Otherwise the model will give you the next best guess
              • duckmysick9 hours ago
                Vibecoded apps will have to include MCP servers so the LLM agents can interact with them directly.
                • mapontosevenths8 hours ago
                  Huh? Both Anthropic and OpenAI have computer use tools. I frequently just tell the models to test their own work. I&#x27;m a little paranoid so I only grant access to one window and make that window VMware Workstation.<p>You combine it with &#x2F;goal. I usually set a goal like &quot;Complete the application defined in goal.md as written. Then test it end to end autonomously using Compter Use. Record all issues discovered during testing in a to-do. Then fix the issues in the to-do. Repeat testing until no more issues are discovered.&quot;
              • romland10 hours ago
                This is where some of the readers start thinking about putting some agents to work on making Photoslop.
          • hiccuphippo17 hours ago
            Does that matter when it can clone it today without the need to train on your specific code?
            • mapontosevenths8 hours ago
              It isn&#x27;t about the code. Some of my work is genuinely new science that might offer an AI company a competitive advantage.
          • byzantinegene17 hours ago
            I don&#x27;t see why that is a concern when decomps and recomps are already blooming, they can copy your app down to the atom.
    • lrvick19 hours ago
      I got insane amounts of Anthropic and OpenAI credits given to me for free for my startup, and I have not touched them.<p>I get privacy, freedom, and no rate limits with the GPUs I racked locally, and those are features I would never give up even if the surveillance capitalism labs paid -me- to use their models.<p>How many consumers are there like me? Probably not many, but once local inference hardware is plug and play, I bet the tides shift pretty quick. Also weights-on-silicon will serve the needs of most consumers locally with more speed than any GPU could deliver for a fraction of the cost.<p>Most people will be doing inference in their pocket or a wearable in 5 years and the giant datacenters will be like AWS, sold to only big organizations that need to auto-scale capacity of custom models on demand.<p>The industry surely knows this and the subsidized inference is just marketing to generate so much buzz and demand such that the tiny fraction of the market they will be able to keep in the end is big enough that they do not collapse under all the debt.<p>OpenAI and Anthropic will be Dell and IBM in 10 years if they survive at all.
      • joenot44318 hours ago
        Dell is trading at 4000% what it was 10 years ago
        • lrvick18 hours ago
          I did not imply otherwise. Just that both would be likely irrelevant to most consumers.
    • pama13 hours ago
      From the post:<p>“With my OpenCode Go sub of $10&#x2F;month, DeepSeek is basically unlimited.”
      • 72deluxe11 hours ago
        I put £15.73 (from a dollar exchange conversion of $21.20) at the beginning of September just to try DeepSeek out via API using Opencode and Pi. I still have plenty of credit left! Their off-peak reduced token cost truly is amazing.
    • eikenberry21 hours ago
      But aren&#x27;t you developing bad habits and learning patterns that won&#x27;t work long term? Or do you think things will get cheap enough that you will be able to keep going with your current patterns post-subsidies?
      • FromTheFirstIn21 hours ago
        No one involved in this is thinking about the long term
        • bombdailer18 hours ago
          No one is thinking, the AI does that for them.
          • diath11 hours ago
            AI does the work, it does not replace the thinking. It&#x27;s like saying a tech lead in a project does not do any thinking because there&#x27;s some junior dev doing tedious and boring tasks.
      • irjustin21 hours ago
        &gt; But aren&#x27;t you developing bad habits and learning patterns that won&#x27;t work long term?<p>2 reasons - there&#x27;s an advantage now, use it. 2nd the frontier providers, this is the &quot;early cheap days&quot; like when uber was initially cheap to compete vs standard cabs. they want you to become hooked and boy are we hooked.
        • byzantinegene17 hours ago
          hooked to ai coding, but not tied to any particular model. if they decide to bump prices up, I can easily switch to a cheaper chinese model on openrouter.
          • ifwinterco13 hours ago
            Yes network effects are significantly less than something like Uber, in fact they’re almost nonexistent
            • quikoa13 hours ago
              This is why they enforce the use of client apps like Claude Code&#x2F;Codex (although OpenAI is a little more lenient). Trying to create a network effect.
              • ifwinterco12 hours ago
                Yeah it makes sense, but ultimately the only thing that creates any form of lock in is the chat history and memories, and that isn’t super important, it’s not a real network effect like a social media app or a taxi app.<p>Having a better model is the only real moat, without that inference is a commodity
                • irjustin9 hours ago
                  There isn&#x27;t lock in and this is what scares them. Switching cost is SO LOW. Our team runs the 3 main models to cross check work without any issues.<p>The benefits are real and our willingness to pay is real, but the valuations only support one and right now even the free cheap models might win. Thusly the collapse could still happen.
                  • ifwinterco8 hours ago
                    I think it will happen, like you say switching is easy and the only possible moat is to build a better model than your competitors at enormous expense.<p>They’re all stuck in a cycle of spending huge amounts of money on training just to stand still (in business terms).<p>In the long run, it can’t continue because it doesn’t make any sense
      • kevin4221 hours ago
        Compared to what a lot of companies spend on software for chip and electronics design (we&#x27;re talking about $10k-200k&#x2F;seat per year), AI coding assistants have a long way to go in cost before companies won&#x27;t be willing to pay for them. Companies pay a fortune for software when it enables their engineers to be productive.<p>For my company, I&#x27;d honestly pay $4-8k&#x2F;month for Claude if I had to (it would be painful, and I&#x27;d try to get cheaper options to work first). I know some enterprise Claude users are paying that much now since they have to pay for API tokens. I am certain it&#x27;s at least a 2X productivity booster for our work. Compared to the cost of hiring another developer, it&#x27;s well worth it.<p>If they stop subsidising Claude Code for the pro&#x2F;max users, there will be a lot of people priced out of it, especially the casual developer. But I don&#x27;t see it going away for commercial use, even with a large price increase.
        • redanddead20 hours ago
          Nah dude I think it’s worse than that<p>Old coding is done, as a workflow in teams. It’s the top down executive pressure of being non competitive as a company, and the bottom up pressure of human laziness<p>Show me people handwriting code à la NASA<p>And I mean we as coders have been trying to do this workflow for a while, I personally would refuse to code without IntelliJ magic complete<p>For this workflow, there’s no going back. What’s hard to imagine is AI taking over the other workflows we predict it will; Customer service AI sucks ass for me as a customer, et cetera
          • kolinko12 hours ago
            I don’t mind bad customer service AI any more, since it’s my claude interacting with it, not me.
        • lelanthran13 hours ago
          &gt; For my company, I&#x27;d honestly pay $4-8k&#x2F;month for Claude if I had to (it would be painful, and I&#x27;d try to get cheaper options to work first). I know some enterprise Claude users are paying that much now since they have to pay for API tokens. I am certain it&#x27;s at least a 2X productivity booster for our work. Compared to the cost of hiring another developer, it&#x27;s well worth it.<p>That enterprise cost you&#x27;re willing to pay is correlated to how much developers will work for. When driving an agent, almost anyone can do it (almost no skills required).<p>If devs cost $1k&#x2F;m, enterprises are not going to be willing to pay $4k&#x2F;m for Claude.<p>What I am saying is, there&#x27;s an equilibrium that will be reached; the price of the human driver and the AI worker <i>will</i> approach each other.<p>Where they stabilise, I still don&#x27;t know, but I&#x27;d be <i>very</i> surprised if, in any field (not just dev), the human gets paid multiples more than the agent they are driving, as the agents get more capable.
          • kolinko12 hours ago
            You also have a cost of team interaction going with square if team size or sth.
      • denkmoon21 hours ago
        I use the frontier openai&#x2F;anthropic models at work but exclusively open weight models (on cloud&#x2F;hosted inference) for personal stuff and I think about it like this; 1) I don&#x27;t see any reason GLM and DeepSeek won&#x27;t eventually be as good as Claude, it&#x27;s just a matter of time and 2) the open models are well and truly capable enough for most of what I&#x27;d want to do. I don&#x27;t need nor want an LLM chewing away on a horrible enterprise spaghetti codebase, my employers can pay for that privilege.
      • pmontra21 hours ago
        Long term, we will see what happens and adapt. At worst we all go back coding by hand. Meanwhile what can I do, tell my customers that I&#x27;m raising my fee because I have to pay for token? The Claude Pro $20 plan is good enough for me and even in auto mode I never had to wait for the 5 hours reset.
      • jaggederest21 hours ago
        I expect by that point we&#x27;ll have local models that can do a decent job, I would guess give it a decade and we&#x27;ll be running custom accelerators that are smarter than current frontier models.<p>In the same way that only supercomputers used to have multiple processors and caches but it&#x27;s now standard.
      • kolinko12 hours ago
        The base costs fall down 2-5x year by year, and they will continue falling due to infra ramp up, asics and distilled models.
    • troyvit15 hours ago
      &gt; This won’t last forever but as long as the frontier labs are subsidizing this heavily the open models won’t matter.<p>It&#x27;s amazing how new we all perceive AI to be, and yet how old the tricks that the big players use. Their job is to just suck the oxygen out of the room as long as they have the money to do it.
      • arcanemachiner13 hours ago
        Yes, but they&#x27;re all offering what appears to be a commoditized product with race-to-the-bottom economics.<p>You can&#x27;t jack up the prices on your product if your competitors can just clone it and resell its essence for pennies on the dollar.
      • alkonaut12 hours ago
        It&#x27;s glorious isn&#x27;t it. We get free work done through subsidies. At the same time, this is what threatens my job, and the money for the subsidy is basically my own invested pensions.
      • zombot15 hours ago
        &gt; Their job is to just suck the oxygen out of the room<p>And that this is even legal is a scandal all on its own.
        • ifwinterco13 hours ago
          Hard to enforce when you’re dealing with private companies with “creative” accounting - who’s to say what the actual cost of inference is for OpenAI or Anthropic?<p>Possibly they don’t even really know themselves at this point, although obviously is it significantly higher than the consumer subscription price
    • chicco4life15 hours ago
      Agreed.<p>I&#x27;ve been running automated research tasks for life sciences companies, and the speed in which tokens are burnt is scary. Especially when you get into a complex knowledge space and require a subwgent to reason through each possibility, token usage grows quadratically not linearly as complexity increases...
    • orbital-decay12 hours ago
      What the provider actually sells to you is GPU time &#x2F; load (oversimplifying). Both tokens and subscriptions are just pretty arbitrary ways to price it, almost unrelated to the actual cost of running the model.
    • wasfgwp14 hours ago
      Luna and sometimes even Sol are cheaper than Deepseek for same tasks even at API prices since they consume way less tokens and are more efficient
    • flir11 hours ago
      &gt; This won’t last forever<p>It probably will. Moore&#x27;s Law is still churning away in the background.<p>Frontier models might get more expensive, but that&#x27;s a moving target. For any particular capability point, the models will only get cheaper.
    • SkiFire1314 hours ago
      &gt; The reason people aren’t freaking out is because most people are using heavily subsidized subscriptions.<p>My understanding is that enterprise plans don&#x27;t offer those subscriptions, so they end up paying for API prices and models like these directly impact that revenue stream.
      • tripledry13 hours ago
        I think it&#x27;s fairly likely medium to large corporations don&#x27;t pay anywhere close to the listed API pricings as they get deals through existing partnerships with the big cloud providers.
    • cedws9 hours ago
      Are you sure you pinned the provider? OpenRouter is known for cycling you between providers which busts the cache and you end up paying way more.
    • bvcp10 hours ago
      selling unused capacity behind high cost demand pricing of api access isnt subsidised in-fact the profit margin anthropic makes on subscriptions is close to 100%. people are looking at this economic model completely backwards
    • computerex17 hours ago
      I use $20 codex subscription, it&#x27;s practically useless for anything other than luna. Deepseek v4.1 flash goes a LONG way for $20.
    • ashish0115 hours ago
      Deepseek is in some ways subsidized as well since they use data for training.
      • quikoa12 hours ago
        That&#x27;s an additional cost to the user rather than infusing more resources from the provider like a subsidy.
    • seunosewa22 hours ago
      Which provider was that?
    • bionhoward20 hours ago
      opencode DeepSeek v4.1 Flash isn’t us&#x2F;eu hosted as of recently, so not sure how this impacts privacy &#x2F; model training
      • Geof256 hours ago
        It is on cortecs.ai?<p><a href="https:&#x2F;&#x2F;cortecs.ai&#x2F;detailedServerlessView&#x2F;deepseek-v4.1-flash" rel="nofollow">https:&#x2F;&#x2F;cortecs.ai&#x2F;detailedServerlessView&#x2F;deepseek-v4.1-flas...</a>
    • coolgoose14 hours ago
      I am using it on opencode go
      • cobicobi7 hours ago
        how was it? did you hit limit often? Im considering opencode go for second sub.
    • chaostheory14 hours ago
      MS&#x27;s Copilot subsidy ended months back. The large OpenAI subsidy has just been cut in half. Who knows when Anthropic ended theirs because I don&#x27;t ever remember it being great value.
    • barrenko13 hours ago
      I mean, it is freaking out ? Or it was, but then started freaking about Opus 5.5 . But days are measured in dog days in AI.
    • 0xbadcafebee21 hours ago
      You should basically never pay API prices, they are always several times higher than subscriptions.<p>There are several open weight subscription providers. OpenCode Go used to be good but now it&#x27;s complete shit. Charm Hyper is really great and the best value. Other subscriptions have a more limited model selection or provide less value but are still decent.
    • nonmaskable11 hours ago
      [flagged]
    • sy13567314 hours ago
      [flagged]
    • mensetmanusman22 hours ago
      Also DeepSeek usage is subsidized as well, it’s a power hungry model.
      • jchw22 hours ago
        Interesting. Are all of the providers on OpenRouter simply losing money? How does that even work out?
        • arjie21 hours ago
          No, it’s outrageously profitable above x% utilization without stealing any prompts. Provider economics still pretty good. Acquiring hardware is the current limiter.
        • jansan22 hours ago
          Don&#x27;t ask and dance as long as the music keeps playing.
        • worldsavior22 hours ago
          You&#x27;re the RLHF.
          • jchw22 hours ago
            With ZDR-only enabled? That seems illegal.
            • worldsavior6 hours ago
              Subscriptions are not ZDR.
              • jchw5 hours ago
                And OpenRouter isn&#x27;t a subscription.
            • skillina18 hours ago
              It&#x27;s adorable you think the AI labs care about that.
              • jchw17 hours ago
                With all due respect (not that I feel like much is due after that response) I am not really sure you know what I am talking about. I&#x27;m talking about inference providers, like inference.net, Fireworks, Coreweave, Digital Ocean, etc, to use DeepSeek 4.1 Flash. They didn&#x27;t create the model, they just are charging you to run inference tasks. That is a different story.<p>DeepSeek themselves are honest about the fact that they train on inputs by default. You won&#x27;t hit DeepSeek if you use OpenRouter with ZDR enabled.
            • komali217 hours ago
              [flagged]
              • jchw17 hours ago
                I now have two people with sarcastic replies that seem to think I&#x27;m talking about OpenAI. We&#x27;re talking about DeepSeek on third party providers. Given that this thread is neither that long nor that confusing, I&#x27;m not sure how people are getting lost this quickly.<p>I don&#x27;t really trust OpenAI, although honestly I find it stupid to suggest they&#x27;d offer a ZDR policy and violate it. They literally don&#x27;t have to offer it. People will still pay. Fable doesn&#x27;t offer ZDR <i>at all</i> and it hasn&#x27;t stopped people from paying through the nose for it at <i>API pricing</i>.
          • edflsafoiewq22 hours ago
            Not if you&#x27;re not giving feedback.
      • kennywinker22 hours ago
        Are you sure about that? My impression was most providers on openrouter were purely selling tokens for profit...
        • thesnarkitecht22 hours ago
          Have y&#x27;all tried an Ollama Cloud subscription? Their off-hours pricing for V4.1 Flash is extremely competitive.
          • fspeech15 hours ago
            Yep. And the model is practically unbounded in its knowledge of math. So people should keep that in mind when they read anything about its level of intelligence.
  • mlinsey23 hours ago
    I&#x27;m paying for the heavily-discounted subscriptions, not the API rates. There isn&#x27;t really a cost gap for me. DeepSeek doesn&#x27;t have a subscription to compare to, but when I compared the GLM 5.3 usage I got from a $100&#x2F;mo Z.ai subscription compared to Opus 5.5 on a $100&#x2F;mo Claude subscription, there wasn&#x27;t a big gap. And GLM 5.3 is very clearly not a frontier model (deepseek v4 seemed a lot closer, but I didn&#x27;t use it enough to really say for my workloads).<p>I don&#x27;t think those subscriptions nave negative contribution margins, either. I think we&#x27;re seeing a lot of price discrimination by the big labs, and huge margins on their frontier models. The fact that they have been cutting prices to their second-biggest tier of models (Opus&#x2F;Sol).<p>Open models catching up and collapsing these margins would worry me if I were a shareholder in the big labs, but as a user, I really doubt that the western labs have bigger environmental impact just because they have higher API costs, I think they have a ton of efficiencies they aren&#x27;t sharing with customers yet because demand is so high.
    • apitman23 hours ago
      The tightening of subscription value has already begun. dsv4.1f is already worth paying for at market API prices. Maybe it goes to 2x because apparently no one has figured out how to match DeepSeek&#x27;s insane caching efficiency, but I don&#x27;t see it getting much worse than that.<p>Plus you can also get dsv4.1f subsidized. OpenCode Go gives 4x if I understand their pricing correctly. Anecdotally, I feel like I get way more out of my $10&#x2F;mo OpenCode Go sub for the price than my $20&#x2F;mo ChatGPT, even using gpt-6.1-sol high which is very cheap, and I have yet to convince myself dsv4.1f is a worse model.
      • rapind23 hours ago
        It&#x27;s really not cheaper than frontier subscriptions. It&#x27;s getting closer, and it&#x27;s a great model, but it is not more value per task than the frontier subscriptions. Don&#x27;t be swayed by the token costs, it&#x27;s very chatty, like 3x more tokens for the same task as sol. I used dsf 4.1 full time for about a week.<p>It blows frontier API pricing out of the water, but again, look at cost per task, not token usage. Still easily wins though for my work.<p>I do think it&#x27;s the most viable alternative I&#x27;ve seen so far, and that applies pressure to the frontier models. Should subscription prices hike or become unavailable for some reason, I know what I&#x27;ll be using.<p>When pricing this, it&#x27;s important to consider whether or not you want to opt out of data training. You won&#x27;t get the advertised rate. Also the dsf 4.1 subscription providers are throttled af... and of course they are, because otherwise they&#x27;d be haemorrhaging money.
        • kelnos19 hours ago
          &gt; <i>It&#x27;s really not cheaper than frontier subscriptions.</i><p>It depends on how you use it. I used to have the $100&#x2F;mo Claude plan. I would easily blow through limits when I was on the $20&#x2F;mo plan, but would rarely hit them when on the $100&#x2F;mo plan.<p>Lately I&#x27;ve been using GLM 5.3 Flash (from Fireworks), and my spend is $1-$2 per day when I use it for coding, so max $60&#x2F;mo (less, since I don&#x27;t use it every day). IIRC DeepSeek 4.1 Flash is priced similarly.<p>If I had to pay API rates for frontier models, I can&#x27;t see how $2&#x2F;day would cut it. Maybe GLM&#x2F;DS are chattier, but not anywhere near the 10x required to make the price difference not matter.<p>Sure, if you&#x27;re running agentic loops all day, 5 days a week, you&#x27;re probably going to blow past even $200&#x2F;mo in API charges pretty quickly.
          • rapind6 hours ago
            &gt; It depends on how you use it.<p>I suspect you are right. For context, I was assuming a 20x w&#x2F; OpenAI or Anthropic subscription as the comparison (or both). dsf 4.1 was going to run me about 2-3 times the cost of either of those for the same amount of work. Obviously you can optimize differently, but that&#x27;s true of subscriptions too. I was using pi and had it evaluate it&#x27;s ideal context compaction point based on usage and API rates. Keep in mind though I was using a ZDR provider, so slightly higher costs. If I want them to train on my code, I could shave a few $ off.<p>That&#x27;s not even taking into consideration all of the resets you get from the frontier subs. Which lately seems to at least double usage (more like 5x recently with OpenAI if you count the credit grants). But OpenAI is tweaking it&#x27;s pricing, so it&#x27;s always a moving target... which is kind of annoying until you learn to just ignore it.
        • apitman23 hours ago
          I am looking at cost per task (and speed per task), both benchmarks and anecdotal experience.
        • epolanski5 hours ago
          &gt; It&#x27;s really not cheaper than frontier subscriptions.<p>I use DS 4.1 flash it all the time, i struggle to spend more than 1 euro per day on it, even working all day long.
        • LeBit22 hours ago
          What are you talking about?<p>DS4.1 Flash not really cheaper than frontier models???<p>It is insanely cheaper.
          • mlinsey13 hours ago
            For coding use cases, it really isn&#x27;t.<p>DS4.1 flash is $0.30 in &#x2F; $1.20 out (per M, peak, cache miss) Opus 5.5 is $4.00 in &#x2F; $20 out (per M, cache miss)<p>However, that is API prices.<p>Anthropic offers a $200&#x2F;mo subscription. How this translates into usage is admittedly a bit opaque, subject to change, and depends on how exactly you use it. But it&#x27;s a lot of usage - Semianalysis data shows that $200 is getting you around $2,500 of usage at API rates if you use Opus 5.5. This is close to what I&#x27;m seeing anecdotally with my accounts, if anything I have been getting a bit more.<p>Now, unlike DeepSeek, you can&#x27;t use your subscriptions to power live AI-driven products, or resell tokens in any way. But for personal coding agents, you can use as many of these subscriptions as you want, for now. So I am paying effectively basically 8% of the published API rates, so at my usage:<p>DSv4.1: $0.30 in&#x2F; $1.20 out Opus 5.5: $0.32 in &#x2F; $1.60 out<p>Obviously, those aren&#x27;t real prices, but they accurately convey apples to apples what my everyday usage costs me and most other heavy users, and why it&#x27;s so easy for me to stick with Anthropic&#x2F;OpenAI.<p>I don&#x27;t think it&#x27;s a coincidence, either - I think the token allowances for these subscriptions are set to be competitive with the open models, so that most coding users (and their incredibly valuable data) stay with the frontier labs, while VC-funded wrapper companies and less-price-sensitive giant companies with strict procurement policies pay exorbitant markups for enterprise contracts at the API rate.
            • tyingq11 hours ago
              Locked into their tools though. I happen to be very tied to a IDE centric model (old dog can&#x27;t learn new tricks) and their Desktop thing is a regression for me. I can use the cli, but it wrecks the muscle memory I have with what I use now.<p>And Anthropic is somewhat unusual in that 10x more tokens via the subsidized path. I imagine more rugpulls are coming.<p>Not disputing your point in any way, just noting there&#x27;s already caveats, and more are likely coming.
          • komali217 hours ago
            Not by cost per task, just cost per token. But they&#x27;re spending way more tokens per task, so it doesn&#x27;t work in their facor.<p>I still use them because they aren&#x27;t as squeamish about random things American CEOs don&#x27;t like like decompilation.
            • adev_7 hours ago
              &gt; Not by cost per task, just cost per token. But they&#x27;re spending way more tokens per task, so it doesn&#x27;t work in their facor.<p>They are not most expensive per Task. DeepSeek 4.1 flash is bloody efficient.<p>I currently did run a test myself: - Use the pay as you go offer on OpenRouter on the same task on two different project: Perf optimisation on C++ codebase both with Anthropic and DeepSeek.<p>- I exploded my 15$ budget in a half-week with Anthropic.<p>- I did two weeks and half with DeepSeek.
    • runtime_terror20 hours ago
      IME GLM is inferior compared to Deepseek v4.1 Flash, highly recommend running it through some real work
  • damowangcy12 hours ago
    Marketing.<p>Have you guys see how aggressive is the push for enterprise use by both OpenAI and Anthropic? I had friend from a non-tech industry in Asia telling me that their company was offered free trial of the enterprise version of Claude, with trainings and such.<p>On the other hand, DS and Z.ai, have zero to none marketing outside China. There is friction to use DS&#x2F;GlM models and the ZDR is unclear, so most enterprise that has heavy AI usage hasn&#x27;t move over yet. They would rather spent $200 for the peace of mind than to take the risk of being slam as a national traitor down the road (which again is another form of marketing by Big AI, trying to frame Chinese models as thiefs).<p>So, I don&#x27;t think they are not freaking out, it&#x27;s just that they are addressing different market segments and reacting to the situation differently.
    • novaRom12 hours ago
      Aggressive marketing (including daily posts on HN). Many people simply don&#x27;t know about alternatives, there are many people who never heard about pi and opencode and live comfortably in claude-codex bubbles.
      • pheggs10 hours ago
        I find it curious that almost every AI post has way more comments than other posts, and I sometimes wonder how big of a part agents play here. I can&#x27;t seem to distinguish whether it&#x27;s just real interest in the subject or a big campaign trying to influence people&#x27;s opinion
        • serial_dev10 hours ago
          It&#x27;s on top of mind for everyone in the software development industry (be it SWEs, product people, designers, managers...), as it&#x27;s such a &quot;cross-cutting&quot; concern, I am not surprised it&#x27;s always on the HN front page.<p>Sure, some Rust 1.98.7-beta release is interesting to some, but AI affects most of us, one way or another.<p>I (and many other people I personally know, so they are not botfarms) feel many different strong emotions regarding AI on a daily basis: anxious, frustrated, tired, bored, suprised, entertained, empowered, optimistic, pessimistic, usually all of it almost every single day.
          • pheggs10 hours ago
            It could be.
    • epolanski5 hours ago
      Marketing is irrelevant.<p>Companies will keep using Google&#x2F;Microsoft&#x2F;Amazon cloud offerings due to the ease of extending existing contracts and procurements.<p>Virtually all my clients use Bedrock or Azure with zero data retention. Not one would use openai&#x2F;anthropic or chinese cloud offerings.<p>Only devs do so.
      • dihinbutt4 hours ago
        Could you not argue that local hosting of open-weight model with ZDR, and possibly even fine-tuned on your company&#x27;s proprietary data is even more secure, compliant, and cheap than just letting OpenAI&#x2F;Anthropic train their models on your responses? Obviously C-Suite often never sees this far into the future, but I&#x27;d argue that there&#x27;s virtually zero downside to this approach.
        • IndeanCondor2 hours ago
          The hardware prices to support this proposal would make by Business Head&#x27;s eyes pop haha
  • gregwebs22 hours ago
    I have been using DeepSeek 4.1 flash intensively for over a month. If I run it all day long it costs $1-2. Its fast. Previously I was always quickly running up to my Claude&#x2F;Codex 5 hour window (on the $20&#x2F;month plan). The cost savings of DeepSeek is real as shown in this article and I am using subsidized plans.<p>DeepSeek is horrible at grilling sessions (the &#x2F;grill* skills to make technical decisions). It doesn&#x27;t know how to explain things. Maybe the skill could be adjusted. It also doesn&#x27;t come up with as good solutions as Opus&#x2F;Sol.<p>What I use it for is<p><pre><code> * the orchestator of my coding workflows * the tester&#x2F;verifier of code changes * the sub agent that explores code or does web searches * putting together code base research reports </code></pre> Previously I planned with Opus&#x2F;Sol&#x2F;Astra and then I used DeepSeek for coding, and then reviewed with Opus&#x2F;Sol&#x2F;Astra. With the cost improvements to Opus&#x2F;Sol I am trying to use them for coding instead now so there will be less back and forth review needed.<p>They are all working together in Pi using the extension @tintinweb&#x2F;pi-subagents where my workflow skill is calling different subagents that use different models.<p>Luna is cost competitive, but doesn&#x27;t score as well on intelligence. I do need the intelligence for most of what I use it for, so I am not motivated to use Luna. Haiku also doesn&#x27;t seem like a competitive price&#x2F;performance mix.
    • newtwilly7 hours ago
      Sol 6.1 on medium seems to benchmark better for less money than deepseek 4.1 flash: <a href="https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;#intelligence-comparison-tabs" rel="nofollow">https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;#intelligence-comparison-tabs</a>
    • kelnos19 hours ago
      How are you getting down to $1-$2 per day running &quot;all day long&quot;? I&#x27;ve been using GLM 5.3 Flash and I also spend $1-$2 per day, but my use is pretty modest, I think. DS 4.1 Flash is priced similarly to GLM 5.3 Flash; can&#x27;t imagine DS is significantly more token-efficient.
    • oh_no22 hours ago
      Luna is 1 point being on AA&#x27;s index at 1&#x2F;4 the cost, yes it &quot;doesn&#x27;t score as well&quot; but paying 4x for 1 point is crazy if you&#x27;re going off benchmarks.<p>AA has Haiku 5.5 as cheaper than 4.1 Flash (both on Max, which isn&#x27;t ideal but what can ya do) and a 4 point intelligence gap.<p>Why do people like to think open models are more competitive than they are?
      • dada21610 hours ago
        Because it&#x27;s just Opus 5.5 and Haiku, and on OpenAI side Luna, that changed the calculus. Also, I find all of them, including GLM5.3 and DS4.1, to be 100% en par with American&#x27;s frontiers model for CRUDS (90% of enterprise programming).<p>Tangentially, all of them would have broken quite badly custom ERPs from my own experience.
      • gregwebs22 hours ago
        DeepSeek&#x27;s own paper advises against using Max, showing that it normally doesn&#x27;t perform that much better. I am not using it on Max, so that&#x27;s not a useful benchmark for me. I have seen other benchmarks where Flash does significantly (30%) better than Luna.
        • oh_no6 hours ago
          i mean we use max benchmarks because they have the best coverage, which sucks because very few people use max day to day, but it&#x27;s what we have. the performance curve is generally pretty similar across models and effort levels, weird outliers are pretty weird. max is generally a big cost bump from most providers (less so from OpenAI)
      • pimeys22 hours ago
        It is super bad on a bit more complex workflows and starts repeating same errors with the same tool until the cycle breaker hits.<p>6 is worse than 5.6 here.<p>But it is amazing on generating a report on content generated by better agentic models such as DeepSeek or GLM, which both do a mediocre&#x2F;bad job on reports.
    • runtime_terror20 hours ago
      I&#x27;m using DSv4.1 in OpenChamber (eg OpenCode) using the Superpowers skills and a lot of custom AGENTS.md instructions to iron out the kinks and I genuinely cannot see a difference between it and Opus and I&#x27;ve been building native iOS and AppleTV apps, Go servers, Typescript, Cloudflare workers, Svelte&#x2F;Astro, etc.<p>It&#x27;s a super capable model all around from my experience.
    • ipaddr14 hours ago
      You can&#x27;t use a flash version and complain it doesn&#x27;t think well you would use the pro version.
      • Gareth32113 hours ago
        While I agree, this particular discussion chain really frustrates me.<p>1. &quot;This Flash model is really smart. Here is an article to discuss how smart it is. Why aren&#x27;t people freaking out about how smart this Flash model is?&quot;<p>2. &quot;I tried using it for a smart thing. It doesn&#x27;t work so well for it.&quot;<p>3. &quot;You should know better than to use Flash for smart things. It&#x27;s not meant for smart things.&quot;
      • k__10 hours ago
        DeepSeek themselves said 4.1 Flash is better than their current Pro version.
  • lmf4lol1 day ago
    Oh man. v4.1-flash has been an sbolute game changer for us. We run all our Personal Assistants now on flash (thinking high) by default and it works incredibly well. There is really no need for basic agentic tasks that might require Kimi K.3 or GLM-5.3 levels.<p>Once its gets juicier, we let flash launch specialized subagents with specific models. GLM-5.3 for coding or Kimi K.3 for research and critique.<p>But as a main driver. I love flash. And it brought our bill down by A LOT :D
    • PcChip23 hours ago
      &gt;We run all our Personal Assistants now on flash<p>are you worried about sending all your data to third parties, especially if they&#x27;re in different countries?
      • techmunky23 hours ago
        Not shilling for them but Ollama cloud hosts domestically with ZDR afaik. I run 95% of my open weight inference through them. The rest goes through Opencode Go $10 plan (which is enough to run 3 hermes agents on DSF 4.1 and leave plenty of left to experiment with when new models drop).
        • octoberfranklin23 hours ago
          Just a reminder that any API using a Cloudflare TLS certificate isn&#x27;t ZDR.<p>The model engine provider might be ZDR, but the service as a whole isn&#x27;t.
          • piperswe18 hours ago
            Cloudflare doesn&#x27;t retain request bodies by default, and if you don&#x27;t trust that then you shouldn&#x27;t trust the third party AI provider either. Cloudflare does cache responses, but that doesn&#x27;t typically apply on API endpoints and is trivially disableable.
          • techmunky22 hours ago
            did not know. ty! have my updoot as thanks
      • lmf4lol14 hours ago
        We use hosted providers in the EU. They have a cost markup but its worth it. I enjoy the idea that I do t talk to big tech.<p>I cant ofc be fully sure because i don’t own the chain end 2 end.
        • imhoguy10 hours ago
          Could you please share which ones? I need serious ones with generic DPA but without Enterprise plan costs.
      • figmert22 hours ago
        I use it through OpenRouter, which has ZDR enforcement.
      • pheggs10 hours ago
        should you not be worried more about sending data to providers of your own country? I think they are equally bad personally, but if you think one of them is acceptable, why do you prefer it to be your own government that has more ties to your life?
      • crossroadsguy23 hours ago
        I have asked OP that question but I think there are providers who are not in China and they just host the model&#x2F;inference.
      • yieldcrv23 hours ago
        Just use a provider hosting it in your country especially if your country has major data centers then its the same as using Anthropic or GPT of GCP Model Garden or AWS Bedrock<p>nobody here is talking about running frontier level intelligence locally so if you’re Chinaphobic and prefer layers of corporations siphoning your data in between you and the party there are plenty of options instead of directly to the party
      • tripleee21 hours ago
        what would be the concern here?
    • crossroadsguy23 hours ago
      What is the cost of access like for DeepSeek-v4.1-flash, compared to GLM-5.3-flash via ZAI&#x27;s Coding Plan? Because that&#x27;s what I use; and often hit the &quot;wait&quot;. I wouldn&#x27;t mind trying a new model subscription or even API access which hits around glm-5.3-flash level weight class (which seem to be enough for me; with quite some human suprvision and nudging) but gives muuuuuuch moooore tokens for the same price.<p>(And what are the preferred providers?)
    • aftbit1 day ago
      Have you compared it against actual SOTA models like latest Fable or Astra?
      • mtrovo22 hours ago
        The author explains this very well tbh:<p>&gt; Today&#x27;s models are now good enough for high-quality unattended tasks. Chasing the latest and greatest is silly. It is fun to see the new Fable capabilities, but the tasks we throw at them are usually ridiculous (maybe even insulting) if you believe in LLM sentience. It&#x27;s like asking a math PhD to organize the files on your desktop.<p>I&#x27;m using DS V4.1 Flash as my main model since their release and it works great for all my coding tasks. My setup is OpenCode Go subscription and obra&#x2F;superpowers skill.<p>The only times I try to change models are on general planning tasks (like research this codebase for tech debt mitigation opportunities) or if I need deep research which would benefit from searching the web, in which I still think Gemini is still the best because of the speed and access to google search index. But these are not even 20% of my daily tasks.
        • aftbit1 hour ago
          I have had good luck rotating between the three tiers of GPT 5.6 with occasional jumps up to Astra or Fable. Most of my work is with GPT-5.6-Sol. Simple tasks like data extraction or trivial refactors (rename this variable etc) I often push down to Terra or Luna, or even to self-hosted Gemma4:31B. Very tricky stuff, like planning a new feature, design review, or code review of a complex change across multiple repositories is where I leverage Astra.<p>I&#x27;ve had middling success with models like DS V4.1 Flash and free Gemini. They tend to be pretty good at very easy stuff ... but they&#x27;re more likely to go down rabbit holes, confidently assert falsehoods, or fix bugs with changes to my test harness rather than my code.<p>I asked about comparison to the well-known SotA models specifically because I use either Astra or Sol for ~85% of my daily tasks. When I try to use smaller&#x2F;cheaper models, I have had very mixed success. Sometimes it&#x27;s perfect, while other times it fails in subtle and hard to catch ways.
      • sneurlax23 hours ago
        Of course there&#x27;s still a huge performance gap<p>but DS 4.1 Flash is good enough for most tasks
    • HKCM85221 hours ago
      What personal assistants are you using?
      • lmf4lol14 hours ago
        Hi. We are building our own, but the core is openClaw. Hoever, that will change soon
  • user4392822 hours ago
    Because DeepSeek is not &quot;a month or two&quot; behind as claimed in the article.<p>These open models still did not beat February&#x27;s Mythos &#x2F; Fable 5.<p>DeepSeek 4.1 Flash is behind GPT 5.6 Sol, and that one is left in the dust by the excellent Opus 5.5.<p>Rumors say Anthropic is holding in reserve the big improvement, Fable 5.5, for the IPO.<p>It&#x27;s plausible that open models are 6 - 12 months behind, and there is no &quot;good enough&quot;. As long as progress doesn&#x27;t slow down, leading labs have nothing to fear.
    • delis-thumbs-7e13 hours ago
      &gt; These open models still did not beat February&#x27;s Mythos &#x2F; Fable 5.<p>On what task? By who? On what benchmark? How do you measure in you own workflow the “betterness” or “more goodness” of these or any models? If you don’t say those things you’re just writing a bad ad copy.<p>&gt; It&#x27;s plausible that open models are 6 - 12 months behind, and there is no &quot;good enough&quot;.<p>Anecdotally, a lot of people - including myself - seem to really notice much difference between the model now or six months ago. So there really seems to be good enough. It depends on the task you use them for and how you measure the output. For most tasks you really do not need frontier capability. Also how do we know how much of these “big improvements” come from the harness and tooling rather than the raw capability of the model?
      • user4392812 hours ago
        Did you see any person claiming an open model was more intelligent than Fable 5?<p>It&#x27;s always some sort of &quot;I don&#x27;t notice the difference&quot;.<p>And honestly, if you don&#x27;t see a difference between the SOTA from 6 months ago, which would be GPT 5.4, and today&#x27;s Opus 5.5, you would have to be downright blind. Not sure what else to say - the results are obviously different for any kind of meaningful output.<p>&gt; Also how do we know how much of these “big improvements” come from the harness and tooling rather than the raw capability of the model?<p>By simply running the old models in the latest harness. Which none of the people who argue &quot;it&#x27;s all the harness&quot; ever do.
        • delis-thumbs-7e11 hours ago
          Ok so you have nothing, just vibes. Fair enough.
          • winternewt9 hours ago
            What kind of evidence would satisfy you. Referring to benchmarks is apparently not sufficient because &quot;you won&#x27;t notice the difference in your everyday tasks&quot;, but saying &quot;I <i>do </i> notice the difference in my everyday tasks&quot; is just vibes.
          • user4392810 hours ago
            I have 500+ hours of experience building native mobile apps with AI between GPT 5.4 half a year ago and today, in addition to my regular software engineering job.<p>The difference between GPT 5.4 and Opus 5.5 is obvious.<p>What do you do, where apparently you cannot see a difference?<p>I honestly can&#x27;t imagine, unless it&#x27;s like sorting your emails.
            • dada2169 hours ago
              I have similar experience than you, with a big caveat, opus 5.5 would have still broken badly a custom ERP.<p>And on that specific kind of software, ultimately a big CRUD, there really isn&#x27;t that much of a difference between GLM5.3 and Opus&#x2F;OpenAI.<p>You see the differences when you get to different class of software.<p>I also have data entry applications that use LLM to actually parse documents, it&#x27;s all Chinese models self hosted because the economic calculus beated a hosted API by about 5x
              • user439289 hours ago
                That&#x27;s fair enough.<p>For mobile apps, I find that nowadays with Opus 5.5 the UI looks better, the UX is better, it can implement more tricky animations and gestures, and it can do all of that with far fewer iterations and feedback than eg. GPT 5.5 would have required.<p>Also vision capabilities were improved significantly with GPT 6 Astra or Opus 5.5, even compared to GPT 5.6 Sol.<p>There was no way the old models such as GPT 5.4 would have done a comparable job when asked to align an implementation to a visual reference.<p>Even for basic websites with no interactive functionality, this should make a significant difference.
                • dada2169 hours ago
                  I even find that for what amounts to a web search Gemini 3.8 flash is better than the big models, stuff like the links to updated tax codes in different countries, or whatever MS is doing with Azure&#x2F;o365 or AWS with their plethora of products.<p>I think use cases are the real reason why people have such different experiences, I too find that Opus5.5&#x2F;Astra&#x2F;6 are better for UI&#x2F;UX now, it wasn&#x27;t the case a year ago, at some point Gemini pro 3.5 was the best one at that.<p>That&#x27;s also why I use all of them and try to not be locked to a single harness as well.
            • delis-thumbs-7e9 hours ago
              I struggle to see how a software engineer fails to see that anecdotal evidence does not matter here at all. I can say I built seven fully functioning operating systems last month, or that I am the fastest runner in the world or that my daddy is the strongest man in the world. None of it matters without data. I can say it is warmer in the living room and you can say that no, it is much warmer in the bedroon, without having something definite to measure and something accurate to measure it with, the whole discussion is pointless - just vibes. You haven&#x27;t even said how you have anecdotally experienced the difference between the older or the open weight models. So excuse me, but then you will fail to convince people on your claims, especially when the whole discussion seems to be in the middle of some bloody information warfare at the moment.
              • senordevnyc7 hours ago
                This you?<p><i>Anecdotally, a lot of people - including myself - seem to really notice much difference between the model now or six months ago. So there really seems to be good enough.</i>
                • delis-thumbs-7e4 hours ago
                  The next sentence:<p>&gt;It depends on the task you use them for and how you <i>measure</i> the output. For most tasks you really do not need frontier capability.
          • dihinbutt4 hours ago
            [dead]
    • loehnsberg2 hours ago
      I run my own benchmarks at the interface of math and coding (ML, math-opt, stochastic models), and keep my own set of benchmarks, i.e. somewhat niche linear algebra and numerical recipes. DS-4.1 is the first open model that passes most of these benchmarks, and GPT-5.6-Sol is definitely not ahead. Even if it&#x27;s all distillation, then DS does a formidable job at it.
    • BobbyJo22 hours ago
      I was thinking about this earlier today and I came to the following question:<p>If you had a model 10x as capable as the best model out today, but it cost 100x more, would there be a market, and, if so, how big?<p>I think there would be a market and I think it would be large.<p>So, I agree.
      • lifeisloving21 hours ago
        Many people would, and you&#x27;ll find that they&#x27;re building crappy webapps where you dont need SoTA. Like seriously who needs these frontier models?<p>Unless you&#x27;re doing some extermely difficult post-grad lvl research, you do not need a 100x PhD research assistant, especially not for whatever silly SaaS product most people are building.<p>There&#x27;s people at my job that get so much more done than everyone else using Fable&#x2F;Opus&#x2F;Astra. and all they use is the fastest cheapest models. I&#x27;d say the people who are using sota models for everything are doing it just because they prefer to be lazy.<p>You simply do not need these frontier models, they outgrew most people&#x27;s needs 6 months ago, but for some reason people still want to run a 700k rack of gpus full throttle to center a div for them.
        • askjdfksdbfhk13 hours ago
          I agree that for average web dev tasks the open models are already good enough. I&#x27;ve had good experiences with both DeepSeek and GLM. And these models are just better for anything security related since they don&#x27;t throw massive hissy fits.<p>However I <i>do</i> actually have a project where I need the frontier models--I&#x27;m working on a deep learning project of moderate complexity (something novel&#x2F;state of the art within its domain, adapting a known approach from published research in a related domain). The difference from Opus 5 -&gt; Opus 5.5 was huge for my project. Opus 5 was struggling, Opus 5.5 is doing really well.<p>I think the demand for frontier models will continue to be there, at least for a subset of tasks, although I agree that it is probably going to shrink as the non-frontier becomes more and more capable.
      • yaro33019 hours ago
        I think the market would be huge, especially if it&#x27;s the &quot;can complete a large task in 3 turns instead of 15&quot; kind of smart. Lots of people and companies would pay for quality + speed.
      • kelnos19 hours ago
        I think there would be a market, but it would mostly be a FOMO market. That is, people would be doing tasks on it that the &quot;regular&quot; model is more than capable of handling, because they&#x27;re afraid they&#x27;re leaving something on the table by not using the absolute best option.<p>Certainly, there&#x27;s a <i>real</i> market for it too, with people who would actually use its advanced capabilities, and see the 100x price as worth it.<p>But sure, even a mostly-FOMO market is still a market. If people are paying, people are paying.
      • samfriedman17 hours ago
        What does 10x more capable look like now? Surely at some point we will reach an asymptote of what can be done purely digitally: all useful coding tasks can be automated, most math research, etc. At some point the physical world becomes the dyke holding back the singularity; until these genius models can scale their investigation into physical experiments and manufacturing, the future will have arrived only in the digital world.
        • BobbyJo16 hours ago
          I think you&#x27;re making a mistake in thinking the digital world is the only one reachable to AI. Robotics and sensing would be opened up by a sufficiently capable AI.
      • Gareth32113 hours ago
        I agree: large market. I think the future of these frontier labs is selling exceptionally powerful and exceptionally expensive models. They&#x27;ll be used for precision, high value tasks. The rest of us will be happy with good enough and cheap models.
      • ForHackernews22 hours ago
        Doing what? How many jobs involve solving Millennium Prize math challenges?<p>99% of everything is CRUD LoB apps.
        • BobbyJo21 hours ago
          I am coding CRUD apps with a mix of astra, sol 6.1, fable and opus 5.5. A more capable model would still benefit me imo. Being able to follow high level guidance better, and being able to harness other models for each task would be a big improvement.
          • lifeisloving21 hours ago
            Do you know how what you&#x27;re doing, or do you find yourself working on things you dont understand and need the best model because it&#x27;s the only way to push your own capabilities (because you&#x27;re avoiding learning how to do the thing yourself)?<p>Not asking to be mean, I just genuinely dont know why you&#x27;d need the frontier for basic applications.
            • BobbyJo21 hours ago
              It&#x27;s a matter of bandwidth. The more I can offload onto the model, the more I can accomplish. For example, I had to do a lot of security work over the last 2 weeks to get ready for an event. This requires handholding current models on many fronts, like: 1) Do they actually implement the security fixes correctly. 2) Do their fixes create any new edge cases. 3) Do their fixes compromise existing interfaces or API surfaces.<p>I cannot trust current models to find all the necessary context, or to make what I consider to be good trade offs. A much more capable model would be able to see my existing patterns (or at least not have context rot make them blind to my convention docs) and make trade offs I agree with much more consistently, and I&#x27;d be able to do more with my time.<p>I&#x27;ve actually found models to be pretty poor at driving things I don&#x27;t know well, so I generally don&#x27;t do that unless its general design&#x2F;product exploration and the end product code is throw-away.
    • zbentley18 hours ago
      &gt; leading labs have nothing to fear.<p>Except being priced out.<p>The big labs&#x27; financials are based on their products being used widely by a lot of the general public. If it turns out that they&#x27;re actually selling a premium product to premium-product consumers at a premium price point (while everyone else buys DeepSeek-like cheaper&#x2F;worse products), that&#x27;s a big issue for them.<p>If a consumer computer hardware company launched by promising investors that it&#x27;d be the next Dell&#x2F;HP and it turned out to be the next Apple (talking Macs here, not phones or apps&#x2F;services), that&#x27;d be an issue for them too.
    • AnodicElegy4 hours ago
      I think it&#x27;s fair to say that an open model has not yet surpassed Fable 5. But I don&#x27;t think it&#x27;s fair to count the period prior to Fable&#x27;s public release. In which case it&#x27;s only been 4 months since the release of a model that&#x27;s clearly ahead of today&#x27;s top open models.
    • runtime_terror20 hours ago
      Perhaps on certain benchmarks and for certain work, but anecdotally I&#x27;ve not been able to see a difference between it and Opus on a lot of dev work (web, Go, iOS&#x2F;AppleTV native, scripting, general tasks)
    • ApolloFortyNine19 hours ago
      Imo deepseek 4.1 (and a lot of the cheaper models, Luna is similar) show the issues in benchmarks. At this point.<p>In actual day to day development the differences are a lot harder to spot. Maybe deepseek is worse, but I asked it to run until it was able to launch itself and verify it worked as expected, and it did. Maybe it wasted some turns, idk, but when it said it was done, it was done.<p>I have no doubt there&#x27;s things it&#x27;s worse at, but what percentage of development is truly novel?
      • criley219 hours ago
        Personally I find this shocking. I don&#x27;t think our application is that complicated, (typescript full stack graphql reactnative etc) but deepseek 4.1 flash is a bumbling fool, junior-level at best, who takes a very long time to make a very big mess. Opus 5.5 one shots truly impressive code in 5 minutes, while deepseek 4.1 flash takes 20 minutes to do horribly. I simply don&#x27;t understand how folks claim they get good engineering out of it. Maybe we still care enough about the fundamentals to notice the mess ...
        • dannyw19 hours ago
          Are you using OpenRouter? I’m honestly surprised open model labs haven’t been calling them out, but heaps of providers either silently serve heavily quant versions, or don’t have inference set up correctly and don’t run the model properly.<p>Was a night and day difference going directly to deepseek api
          • criley218 hours ago
            In this case, I was using Hugging Face inference set to one of a few American providers. For my professional work, I do not use Chinese inference.
    • cedws12 hours ago
      The benchmarks don’t mean shit. Opus 5 was a terrible model and yet it had very impressive benchmarks. All the labs are benchmaxxed to the tits, only open models’ benchmarks are even worth paying attention to because they literally cannot cheat.
    • pheggs10 hours ago
      I have a different experience and find them just as good as the latest anthropic and openai models. But then, we probably can just have opinions on this as benchmarks are probably used for marketing to a large degree. I find it hard to find truly independent benchmarks that don&#x27;t have any ties to the cash flow of openai and anthropic, can not be trained up on etc.
    • airtnp21 hours ago
      Agree, will see after the dilution and CoT hack fixed, will they keep the pace now. MiMo had some good numbers recently because it&#x27;s discovered that the post evaluation RL directly exposes answers to models, so RL and evaluation is runied.
    • ForHackernews22 hours ago
      There absolutely is &quot;good enough&quot; and I agree with this author: DeepSeek 4.1 Flash is plenty good enough for all the things I would trust an AI to do at my job.
      • senordevnyc20 hours ago
        Yes, trust is exactly the point. I trust the frontier models to tackle bigger chunks of work than the open weight models, and get it right.
    • aleqs22 hours ago
      that just sounds like openai&#x2F;anthropic cope&#x2F;propaganda, based on absolutely nothing objective lol<p>even their harnesses are far surpassed by pi and opencode at this point<p>also sick &#x27;rumors&#x27; lmao, apparently marketing through rumors is in vogue these days
      • enraged_camel22 hours ago
        &gt;&gt; that just sounds like openai&#x2F;anthropic cope&#x2F;propaganda, based on absolutely nothing objective lol<p>Nah. There are benchmarks. They are free to look at. And they paint a very clear picture.
        • Jaygles16 hours ago
          I see people throw around these benchmark numbers<p>I&#x27;ve seen different benchmarks come to different conclusions<p>Benchmarking these models must be an incredibly complex and difficult problem<p>How can a lay person know which benchmarks actually have good signal?
        • aleqs21 hours ago
          [flagged]
  • kraig91145 minutes ago
    Everyone I talk to about Deepseek has a bad taste it feels like for CHinese models? I have a Kimi sub and use deepseek a lot at home. I&#x27;ve hit my limit on K3 many a time for Deepseek to come in and finish a job. So far no complaints. I feel it makes a lot of round trips by design but over all a good option. It&#x27;s hard to compete though with an open sub out on codex&#x2F;claude and just do everything in that subscription. Deepseek for me is when I run out of my subs. It&#x27;s usually always coming behind a great session and fixing things so I don&#x27;t give it a chance to try something novel on it&#x27;s own.
  • p1necone23 hours ago
    I have a pretty large, complex project I&#x27;ve been building with heavy AI use (new language + compiler). I <i>was</i> following a &#x27;strong model as orchestrator launching cheap models as implementers&#x27; pattern, but I recently trialled just using Deepseek-V4.1-Flash as the model for both layers because of the cost savings (with mimo v2.6 flash on code review agents for some decorrelation).<p>I was previously using GLM-5.3 as the orchestrator, after switching to DS anecdotally there <i>was</i> an unnacceptable quality loss, mostly around not taking all the relevant context into account when making decisions, pulling new design out of thin air without discussion too often, and being way too wordy and rambly in documentation despite prompting to avoid it. There&#x27;s a lot of docs, rulings, core concepts, design philosophy to uphold and DS was just not cutting it.<p>However, it&#x27;s perfectly capable of being the sole agent for all of my well specced implementation tasks. I&#x27;ve gone back to GLM as the orchestrator.
    • rspeele23 hours ago
      On the Claude side of things I was previously following &quot;strong model directs weak&quot; with Fable directing Opus&#x2F;Sonnet (its choice per-task). Since Opus 5.5 came out I&#x27;ve just been having Opus direct Opus.<p>The sub-agent separation is still valuable to keep context clean for the orchestrator, but I just have no reason to use Sonnet as the grunt-work implementer because I&#x27;m finding it hard to run out of tokens with Opus 5.5 on a $200 subscription plan. It&#x27;s <i>really really</i> good at subjective quality of work per token used.
      • p1necone23 hours ago
        I would probably go that route if I could use other harnesses with claude models, but I don&#x27;t want to be locked in to claude code, and their API pricing (which you need to use it with other harnesses) is so much higher than subscription.
        • nostrebored21 hours ago
          plenty of harnesses use the claude subscription
          • komali217 hours ago
            Ever since Claude banned this, I haven&#x27;t found one. Which harnesses have you found that allows using Claude through the subscription plan?
            • bibstha42 minutes ago
              pi.dev has it. <a href="https:&#x2F;&#x2F;pi.dev&#x2F;packages&#x2F;pi-claude-code-provider" rel="nofollow">https:&#x2F;&#x2F;pi.dev&#x2F;packages&#x2F;pi-claude-code-provider</a> is the plugin you need to install and everything goes through claude subscription
            • catgirlinspace15 hours ago
              oh my pi has it.
              • komali21 hour ago
                I just checked, it has claude available through the API. I asked Claude itself and it said the subscription service is only available through Claude code.<p>Did you use some kind of plugin? I&#x27;d like to use opencode with my Claude subscription.
      • vorticalbox21 hours ago
        My work pays for Claude and cursor.<p>I have actually just dropped to using sonnet for everything, sure it does need some directing but I have yet to see a need to jump to opus.<p>To me it feels like sonnet&#x2F;terra and composer 2.5 and grok 4.7 are actually good enough for most tasks and these companies are pushing the high models simply to make money.
        • Silagi5 hours ago
          Every test I&#x27;ve done with Sonnet or Tera has been more expensive than just using their big brothers, even before 6.0 and 6.1 Sol came out at Tera prices.<p>They&#x27;re cheaper per token, but they burn more tokens and take more wall time than just going up a tier at a lower reasoning level. And if the task is simple enough that they do come out cheaper than Sol, I can switch to Luna and come in even cheaper.
          • vorticalbox3 hours ago
            Actually not tested that I will give it a go. I usually keep reasoning on low.
    • hollowturtle22 hours ago
      Can&#x27;t wait to just use the hand made Jai and laugh at everything else built with AI :)
      • p1necone20 hours ago
        You can just use plenty of good handmade languages now, I&#x27;m not claiming mine will ever be <i>good</i> or useful to anyone other than myself. I&#x27;m using AI so heavily on my project because I want to explore the PL design space without spending years on implementation for things I&#x27;ll probably want to throw away and rewrite a different way once I actually play around with them properly. And because I lack the motivation to persist through the sheer volume of grunt work that language implementation needs to get to the juicy interesting parts.
    • runtime_terror20 hours ago
      If you&#x27;re not using the Superpowers set of skills, give it a shot. It&#x27;s been working really well for me on a variety of tasks.
      • tombert18 hours ago
        I used Superpowers for awhile, but I don&#x27;t feel like it gave me significantly better results than just raw-dogging it. It <i>did</i> however burn through my tokens significantly faster.<p>Maybe I&#x27;ll come to miss it now that I removed it, but I certainly don&#x27;t yet.
    • dizhn5 hours ago
      I am having a similar experience. Try gemini flash at the top level. If you talk to it as well, it&#x27;s much better in conversation too. Current deepseek is very &quot;unless, but wait&quot; in its thinking and this spills over to it&#x27;s chat.
  • joshstrange8 hours ago
    I’ve played with DS4.1 Flash. It’s neat, no doubt. In my tests it does decently well against Sonnet 5.5 but uses way more tokens. That works out since it costs ~1&#x2F;3rd Sonnet per task&#x2F;PR according to my tracking.<p>However, I can easily burn ~$5&#x2F;day if I use it as my “worker” (still using opus for planning and review) so call it ~$150&#x2F;mo.<p>I could, if I were so inclined, get a second Claude Sub and have even more headroom (though I’m able to stay under my limits most of the time with my current setup). Also Claude gives me Artifacts, Web Search, and now even some API Credits.<p>I have no doubt the future is open weights and I can’t wait, literally, I can’t wait for them to catch up on intelligence or for hardware to run a decent model to be within my grasp. But until that comes to pass, I’ll keep using Anthropic.
    • iammrpayments8 hours ago
      5$ per day how???<p>I’ve been trying to spend my 5$ with deepseek that I put 2 months ago. I almost never hit limit with claude opus in 20$ plan.<p>Do you guys just run model in paralallel + loop?<p>Does it even produce anything worth using that way?
      • joshstrange8 hours ago
        Well, I’m building a “software factory” (yes, along with everyone else it seems) and running my main work through it, as well as my side business, as well as building the factory itself. Using my $200&#x2F;mo Claude plan for opus to plan&#x2F;review and sonnet&#x2F;DS to execute the plan.<p>I only reached for DS since I was hitting my limits on my subscription plan but I think switching to sonnet for the worker will fix my limit issues (previously using opus for everything).<p>I am using DS in Claude code with Superpowers (both of which increase token usage) but that’s my setup.
  • weknowbetter2 hours ago
    People who have not used DeepSeek massively under estimate it&#x27;s capability and how little it costs to run.
    • saberience2 hours ago
      Yes but they&#x27;re not under-estimating how much worse it is than Opus 5.5 or Astra...
      • Zambyte56 minutes ago
        But they&#x27;re almost certainly overestimating the level of intelligence required to comfortably do their tasks.
  • 827a17 hours ago
    Its simple. My company is very willing and able to pay ~$200&#x2F;engineer&#x2F;month for the best version of these tools. My company is even willing and able to pay as much as $500&#x2F;engineer&#x2F;month, but does not need to at this moment.<p>My company is not willing to pay $50&#x2F;engineer&#x2F;month for a cheaper version that is nearly as good. My company is also not willing to pay any amount for a product produced by China, even if it is hosted in the United States.
    • plaidfuji3 hours ago
      This is the correct answer. The pro-sumer &#x2F; hobbyist &#x2F; small developer market is tiny compared to enterprise. Enterprise will pay out of their eyeballs for a top-of-the-line product from a large US company. Doesn’t matter if the Chinese product is <i>free</i>.
    • leptons9 hours ago
      My company is willing to lay off another engineer to pay for more tokens when the price goes above $200&#x2F;engineer&#x2F;month.
  • gutchapa1 hour ago
    Actually OPUS sucks... Deepseek has its own flaw, but fares lot better in many instances. And yes it&#x27;s far subsidised. If you are mindful of its peak and off peak hrs, you could make best use of its cost surge.
  • hmontazeri1 day ago
    I had the same experience using ds 4.1 last couple of weeks. It’s insanely good for the price. I’m doing mostly web dev it excels at everything I throw at it. The pricing is ridiculous. I canceled my gpt subscription and haven’t looked back hope the pricing stays like that. I almost never need a better model. I still keep my Claude 20$ sub for now but I feel like one more iteration and I won’t need even that anymore I hardly use it
    • jacquesm1 day ago
      If DS4.1 impresses you I would be really interested to see your comparison to GLM 5.3. I switched from the one to the other and even if GLM 5.3 is a bit slower I don&#x27;t think I&#x27;ll be going back.
      • badatnames23 hours ago
        DS4 (not 4.1) crossed my dont-care threshold and I genuinely stopped paying attention to new models. I&#x27;d love to try GLM 5.3 but I just don&#x27;t see any point in spending the effort any more. I can get passable intelligence for a bargain price either direct from China or from a ZDR EU provider for a small markup. Paying 10x more will not make me 10x happier, it&#x27;s unlikely to make me even 1.1x happier now I&#x27;ve got some intuition for the natural limits of these models.<p>I don&#x27;t even bother checking how much I spent on API any more, its well under $30 over the past 2 months despite daily constant use. Who even needs a subscription at these numbers?
        • hmontazeri14 hours ago
          I’m sure it’s more intelligent but the speed of DS is so freaking fast I can’t user slower models anymore
          • ywvcbk12 hours ago
            But it also frequently uses massively more tokens for the same task compared to other models which kind of negates the speed.
            • badatnames2 hours ago
              I&#x27;m simply not seeing this
          • jacquesm11 hours ago
            Wallclock time matters and GLM 5.3, even when it is considerably slower (~30 tokens right now vs 100+ on DeepSeek 4) it is quite frequently faster on the same set of tasks overall. Deepseek 4 seems to do the &#x27;Oh, wait&#x27; thing just about forever and has a tendency to find irrelevant rabbit holes that it then spends a massive amount of tokens on.
      • ctolsen1 day ago
        GLM 5.3 is very impressive and definitely better, but it also at least 4x the price.<p>On that note I’ve been subbing in MiMo-2.6-pro when cost is an issue, which is super cheap and also performing really well.
      • pjerem23 hours ago
        IDK what happened today but I used GLM-5.3 as usual from Ollama cloud and it was so fast it generated entire documents like instantly.<p>The reasoning and the result document were done after less than 1 or 2 seconds.<p>Have Ollama suddenly bought GPU capacity?
      • LeBit22 hours ago
        There is also GLM 5.3 Flash
    • pdhborges23 hours ago
      What inference provider are you using?
  • ojr2 hours ago
    It&#x27;s like the boy crying wolf with these models, Deepseek is not as good as Claude, I like using Gemini Flash Lite even, it is good enough for most of my crud app tasks, brownfield projects you don&#x27;t need the strongest models, but greenfield projects with lack of context from indexed code and lack of domain expertise, I think the Opus models have been working for people
  • simpaticoder1 day ago
    The question seems rhetorical but I think there are two reasons in some combination. First is it there is some awareness lag here. That lag can be on the producer and consumer side. Software enterprises are pretty slow to adopt new things and slow to try new things so they might only be aware of openai and Claude as options. Plus there are some scariness because deep seek is a Chinese model and therefore export restricted - never mind that there are American in European providers.<p>The other reason is more interesting. Maybe the frontier providers think that price performance is irrelevant in light of very powerful frontier models that can start the RSI loop and or a huge displacement of work and a winner take all economic situation. After all if frontier providers earn everyone&#x27;s money then you won&#x27;t have any money to spend on any model 100x cheaper or not.
    • joshheitzman50 minutes ago
      I&#x27;ve been waiting for non-preview release of v4.1 flash. I&#x27;m assuming this one is a preview like the v4.0 ones without dates in their name. Maybe its not a preview, but after how v4.0 was handled I&#x27;m assuming it is.
    • I think it also has a bit to do with the AI sector of tech still moving at lightning speed.<p>Theres already models that outdo DS 4.1 flash in cost&#x2F;performance. Luna 6 on max effort for example. Luna also doesn&#x27;t care what time of the day it is for cost calculation.<p>And I&#x27;m sure by the time people ask why Luna 6 is being slept on there will be another cost&#x2F;performance king
      • pimeys21 hours ago
        Luna is very slow and bad at agentic tasks. DS runs circles around it and there are US providers providing cheaper rates no matter the time of the day.
  • K0IN2 hours ago
    I really really love deepseek v4(and 4.1), but for everything I thrown at it, it felt like gpt 5.6 + 6 luna or terra do it faster, in <i>way</i> less tokens and in a way I like more. So even tho the price is (very) cheap, i found myself using it less just cause I don&#x27;t want to wait on something I have to itteratee on with the model.
  • aguilaair1 day ago
    What about MiMo v2.6 Pro? It’s throughput is slower by default (UltraSpeed is faster than DS4.1F) but is above the pareto line, and cheaper.<p>see <a href="https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;models&#x2F;releases&#x2F;comparisons?compare=claude-opus-5-5,deepseek-v4-1-flash,mimo-v2-6-pro" rel="nofollow">https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;models&#x2F;releases&#x2F;comparisons?co...</a>
    • patresh23 hours ago
      Technically yes, but has been reported to be quite benchmaxxed. In practice Deepseek Flash 4.1 and GLM 5.3 might therefore still outperform Mimo 2.6 pro.
      • ctolsen21 hours ago
        I’ve been using it a lot and it’s performing really well. Not GLM 5.3 levels but it beats Deepseek for my use. I’ve used it on long running coding tasks, though mostly prototyping, but it’s done a great job at very low cost.
  • wg01 day ago
    While using DeepSeek v4.1 Flash I was architecting a system and I made a mistake of drawing the RPC boundaries at a wrong place that did cost me in so many ways.<p>I realized that mistake and guided DeepSeek where it should be.<p>Next I fired Fabble 5.5 set to high to check if the hype is real about Fabble. It exhausted 89% of quota and came up with NOTHING that DeepSeek hadn&#x27;t flagged itself already in its notes.
    • sampullman1 day ago
      Do you mean Fable 5.1? Or Opus 5.5? I&#x27;m not sure what you&#x27;re working on but for me DS 4.1 flash isn&#x27;t nearly at their level. For the price it&#x27;s obvious very impressive, though Luna 6.0 is excellent too.
      • hirako20001 day ago
        The problem with benchmarks and proprietary models is that one day a model is best at doing X, another day that&#x27;s not so sure. And anyway, we are not throwing the same X.<p>I&#x27;ve found supposedly smaller and, less performant models do better on certain tasks. I end up using several models, sticking to what my unconscious statistical observations tell me to use for the kind of task at hand.
      • wg023 hours ago
        Fabble 5.1.
      • mtrovo22 hours ago
        care to share what exactly are you working on?
  • pierreb-aiva9 hours ago
    Because of several factors:<p>- Models are still very jagged. A superior model (e.g. Opus 5.5, Astra) may not be materially better in some tasks, but usually is for completely new modalities: computer use, game development, etc. People will always prefer using less jagged intelligence, because it allows them to do so much more<p>- Distilling intelligence from frontier models means that unless chinese labs manage to replicate the training regimes of OpenAI &amp; Anthropic, they are always going to be behind a few months. That&#x27;s a feature of training by distillation.<p>- China doesn&#x27;t yet have the compute to compete at the frontier. Because they don&#x27;t have access to top tier chips, their GWs are not equal to US GWs. Until they close the hardware gap, I think they will always be focused on competing on efficiency, as opposed to intelligence.<p>- OpenAI and Anthropic subscriptions are valuable of quality of intelligence + generous compute allowances
  • arush15june23 hours ago
    I am 4.1 maxxing on commandcode GOAT Plan + api rates with oh my pi for the last 4 weeks, it&#x27;s absolutely amazing and crazy fast, it&#x27;s alright if it makes a mistake, I have enough time to iterate again, I have also added an advisor layer of mimo 2.6 pro which does make it a notch smarter. Getting haiku 5.5&#x2F;sonnet5.5 to work on plans and letting 4.1 flash work through it is helping a ton too.<p>I am a big ChatGPT fan, all our team has ChatGPT Subs, but the TPS across all models including luna is just so damn slow.<p>Commandcode giving 60$ worth of Deepseek for 10$ is just genuinely goat.<p>And it never says no for cyber tasks so that&#x27;s a big win
    • pimeys21 hours ago
      Yeah. I&#x27;ve been enjoying Coralbricks 250-350 tok&#x2F;s speeds and it is hard to go back to slower models.<p>Lithos promises even faster speeds if you want to pay more.
  • xzjis20 hours ago
    The problem with this low listed price per token is that, in reality, DeepSeek 4.1 Flash uses 10× more tokens than GPT-6.1 Sol for an equivalent task and delivers a lower-quality result. So there’s no real benefit to paying 10× less per token. Also, as someone else pointed out, OpenAI and Anthropic currently offer subsidized subscriptions for $100 or $200 a month that provide far more tokens than the API, so we should take advantage of them while that lasts.
    • stratts13 hours ago
      You can see this in the Artificial Analysis benchmarks - GPT 6.1 Sol on medium thinking scores 10 points higher than DeepSeek 4.1, but is actually cheaper per task, as it only outputs 15 million tokens instead of 250 million.<p>Though for longer sessions I think DS4.1 would still come out cheaper... it&#x27;s hard to beat that 98% cache discount
    • motoboi19 hours ago
      That&#x27;s my experience too. sol-6.1 goes straight to solution, like it has done it 100 times before. Deepseek will explore and insecurely overthink like it&#x27;s an intern made CEO.
  • erans4 hours ago
    I&#x27;ve been using DS 4.1 Flash for a while now as my main workhorse. It&#x27;s amazing. It&#x27;s fast and it is also a GREAT debugger - it might not always find the best solution but it will definitely find the problem quickly.<p>I&#x27;m also the CEO of a new company - LunaRoute - so if you want private, fixed cost, we run our own server in the US type of DeepSeek 4.1 Flash check out <a href="https:&#x2F;&#x2F;www.lunaroute.com" rel="nofollow">https:&#x2F;&#x2F;www.lunaroute.com</a><p>If you want a trial, hit the contact us and write that you saw this post.<p>We also have GLM 5.3 (with vision!) and GLM 5.3 Flash - all included.
  • RGS181123 hours ago
    This model finally got me off my Claude Max subscription. I’ve found it superior to Opus 5.5 in certain use cases, and certainly faster.<p>I’m convinced that I’ll have good enough inference on my laptop at reasonable speeds within the next year.
  • Primer817 hours ago
    I&#x27;ve been waiting for something like this article for a month. The other AI companies are being left in the dust. I rip about 200 million tokens at least a day and spend maybe $3 with deepseek flash v4.1 for self hosted related coding tasks for ~20 projects i work on &#x2F; maintain in parallel. Its a daily driver for sure. not even worth considering anything else, but maybe a locally run model on my GPU at the moment.
    • quietfox7 hours ago
      What’s your dev environment with flash v4.1? Do you call the API directly or via Openrouter or something else?
  • apitman23 hours ago
    &gt; With my OpenCode Go sub of $10&#x2F;month, DeepSeek is basically unlimited<p>My OpenCode Go monthly window was scheduled to reset this morning. It was sitting at 22% used despite me using DeepSeek V4.1 Flash heavily as my implementation agent the past couple weeks (I use gpt-6.1-sol high for planning&#x2F;orchestration).<p>I had 1.5 hours left so I fired up first 10, then 20, and finally 50 concurrent subagents all working on reverse engineering C code from an old PC game. They found over 100 new functions.<p>This is the first workload I&#x27;ve found that could make a dent in my sub. It got my 5 hour window to 85% used, but sadly my monthly was still only at about 35% when it reset. So that cost maybe $2.
    • youniverse22 hours ago
      How are you guys doing orchestration? I have been fumbling around in my free time trying to build something for myself but is there a repo or something that just works?
      • apitman22 hours ago
        I might not be the best person to ask. I use Pi harness in tmux and just ask my current agent to spawn interactive pi instances in new tmux windows, create a sentinel file for each of them, and monitor the sentinel files for signals every 2 seconds.<p>Currently have auto compaction turned off. When the orchestrator&#x27;s context is getting close to full, I have it write a handoff markdown file and point a fresh agent at it.<p>I do feel like I&#x27;m getting close to the point where I might be ready for something more sophisticated, especially wrt to subagents communicating with the orchestrator.
        • crossroadsguy18 hours ago
          That limitation is what stops me from using Pi for anything serious. I may have to configure it and then configure it and it will eventually become a codex, a claude code or so. I recently heard the maintainers added mcp to it (in stock, not via plugin), I wonder what stopped them from adding subagent function, and decent loop capacity to it.
        • mappu13 hours ago
          &gt; Currently have auto compaction turned off. When the orchestrator&#x27;s context is getting close to full, I have it write a handoff markdown file and point a fresh agent at it.<p>In a good harness that should be how auto compaction works anyway
          • apitman4 hours ago
            Yes, and I&#x27;ll probably trust it eventually. But currently I always want to know when compaction is happening, so I can correlate any drops in quality or weird behavior.
      • CharlieDigital22 hours ago
        Many orchestrator harnesses exist.<p>Check: <a href="https:&#x2F;&#x2F;agentmgmt.dev&#x2F;" rel="nofollow">https:&#x2F;&#x2F;agentmgmt.dev&#x2F;</a> and find the one that works for you.<p>I quite like Paseo (been maining it for a week), but Orca also looks good.
    • sourcecodeplz15 hours ago
      you can spend your opencode go quota faster than 1 month. with the limits i think its about two weeks.<p>so even if the week reset with some %usage left, its not actually lost if its not the end of the month.
      • skeptic_ai10 hours ago
        I reverse engineer a basic flash game and spent 3 open code accounts.
  • atleastoptimal1 hour ago
    Deepseek and many other models are heavily benchmaxxed. They aren&#x27;t genuinely as good as the best frontier models.
  • nightpool5 hours ago
    &gt; This is like comparing big pharma with generic manufacturers who can skip the R&amp;D. While these drugs are not 1 to 1 copies, we are still comparing apples to apples, but with a 90% price cut.<p>Weird! It&#x27;s almost like distillation is bad for the long-term growth of the industry, just like generic manufacturers would be if they could release the generic versions of drugs 2 weeks after the original R&amp;D completes.<p>Very strange sentence to include in an article after saying &quot;I don&#x27;t care at all about distillation&quot; up at the top. Can the author not hear themselves?
    • plaidfuji3 hours ago
      I’ve said it before but the AI industry looks a lot like biotech &#x2F; biomanufacturing. Huge R&amp;D budget to produce products that are ultimately highly complex commodities. CapEx-intensive to build new “manufacturing” capacity. COGS matters. Regulation &#x2F; IP control will be needed to protect innovators.
  • alex-moon23 hours ago
    I think because we&#x27;re all just using it thinking we have found the &quot;model for me&quot; and never mentioning it to anyone because what would we say? It&#x27;s good. It&#x27;s a bit like the Logitech MX Master, as more and more people assumed they had found the ideal mouse for their purposes, it quietly became the professional standard through sheer adoption.
    • rhdunn12 hours ago
      There are also several issues at play here:<p>1. a model that works for one person&#x2F;task may not work for another;<p>2. there are many models (DeepSeek, Qwen, GPT, Claude, Gemini, etc.) that are released every 6 months or so;<p>3. it takes time to use, test, and evaluate the suitability of a new model and not everyone has an automated evaluation process for their use cases.<p>Thus, if you find a model that works for you then you are not going to spend more time evaluating a model that may not work, or may only do so when time permits.
      • leptons9 hours ago
        I&#x27;m pretty happy with claude $20&#x2F;mo for personal, and $200&#x2F;mo at work. It does everything I need it to, at a price that I can aaccept. If they raise the price, then I&#x27;ll start looking around at the other options. I do not like Altman or Musk, and I don&#x27;t want to get involved with China, so I&#x27;d rather avoid those products. But I do know that prices will have to increase at some point, so I&#x27;m still hand-coding a few projects to keep my skill level up should LLMs just not be affordable in the future.
  • swiftcoder1 day ago
    I think the interesting provider to cross-check this assumption with here is Meta, who is clearly freaking out, and is currently providing Muse 1.3 <i>even cheaper</i> so long as you are willing to share data with them
    • 9dev23 hours ago
      Whatever the question, Meta is the wrong answer.
    • jacquesm22 hours ago
      &gt; so long as you are willing to share data with them<p>I don&#x27;t think so.
      • 010ED6791312 hours ago
        yeah this is hacker news. we only share our data with Dario, Altman, Musk, and the CCP. Not untrustworthy people like zuckerberg.
        • jacquesm11 hours ago
          I don&#x27;t share my data with any of those either.
          • swiftcoder4 hours ago
            That leaves you with who, exactly, in the LLM game? Google? Mistral?
  • jokers1324 hours ago
    &gt; It&#x27;s like asking a math PhD to organize the files on your desktop.<p>I actually do use an agent harness to organize files on my desktop. They make a great fuzzy file renamer. Point it at a directory of disorganized files with names all over the place, give the directory layout and file name pattern you want it to have and it makes it happen.
  • poulpy1238 hours ago
    Why ?<p>Because I don&#x27;t have the time and money to do an extensive benchmark of all major LLM, so when I had to select a LLM for my usage (which was not coding at first), I went to the most used one, chatgpt, because I knew if would be one of the best at the task.<p>I suspect it&#x27;s the case for many if not most people.
  • zug_zug23 hours ago
    I did a test on this a couple weeks ago. What I found was that the chinese models were far better than API rates, but about comparable on price vs the subsidized subcription model (chatgpt). Also it was my experience that codex completed tasks quicker.<p>That said, it&#x27;s my best understanding that these american companies aren&#x27;t profitable and will eventually raise rates (the old uber trick) so I&#x27;m keeping myself ready to switch when that day comes.
    • 0xbadcafebee21 hours ago
      I keep a spreadsheet that estimates actual value (dollar amount per token per month, per subscription rate limit) and open weights are basically always cheaper than frontier weights. Recently things like GPT 5.6 Luna finally got the frontier close to the value of open weights but their limits keep them behind.
  • rbnafo13 hours ago
    Because most of the people use it through enterprise agreements and don&#x27;t pay the bill? I run it for my own use cases and its pricing plus caching capabilities are hard to beat, cents for millions of tokens. <a href="https:&#x2F;&#x2F;substack.com&#x2F;@rubenafo&#x2F;note&#x2F;c-332218129?r=26y5kn&amp;utm_source=notes-share-action&amp;utm_medium=web" rel="nofollow">https:&#x2F;&#x2F;substack.com&#x2F;@rubenafo&#x2F;note&#x2F;c-332218129?r=26y5kn&amp;utm...</a>
  • balboer44872 hours ago
    we do. we moved 100% of our production traffic from Gemini to Deepseek. best decision we ever made.
  • Call me crazy but:<p>VRAM &amp; Memory Requirements by Precision<p>• FP16 (Full Precision): Requires ~1,664 GB of VRAM (e.g., an 8x B300 288GB cluster).<p>• INT8 Quantization: Requires ~832 GB of VRAM (e.g., 8x H200 141GB).<p>• INT4 Quantization: Requires ~416 GB of VRAM (e.g., 8x A100 80GB)<p>VRAM aint cheap, Sam Altman ruined the cost of memory, Nvidia doesnt make enough consumer GPUs letting the market go insane over them, I still have friends on 1070s or 1070 TIs because GPUs have been severely overpriced for too long. I remember when a gaming PC was only $1000.<p>Even so why would anyone not sleep on a model they cannot run?
    • kristopolous1 day ago
      Seriously, if a single politician stepped forward and said &quot;i&#x27;ll bring down ram prices&quot; they could then shoot a puppy and call me a slur and I&#x27;d still go out and doorknock for them.<p>Memory companies have price fixed multiple times. They&#x27;ve paid hundreds of millions in fines. wikipedia even has a page on it. <a href="https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;DRAM_industry_price_fixing" rel="nofollow">https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;DRAM_industry_price_fixing</a>.<p>Look at the financials of these companies, they&#x27;re all making obscene margins and do they plan to increase production? No. Micron is doing a stock buy back to pump the price of their share.<p>The Micron CEO just recently said this is the exact plan <a href="https:&#x2F;&#x2F;www.theregister.com&#x2F;systems&#x2F;2026&#x2F;10&#x2F;01&#x2F;ram-supply-set-to-worsen-says-micron-as-ceo-celebrates-much-higher-prices&#x2F;5300346" rel="nofollow">https:&#x2F;&#x2F;www.theregister.com&#x2F;systems&#x2F;2026&#x2F;10&#x2F;01&#x2F;ram-supply-se...</a><p>There&#x27;s sanctions, tarrifs, and a DOJ who doesn&#x27;t give a shit. Until we can fix that the insanity will continue. Phones will be unaffordable. Laptops will be obscene. Gaming consoles will be thousands of dollars. Desktops will be dead.<p>If you&#x27;re waiting for some David Ricardo equation to happen, tough cookies, it&#x27;s not coming.<p>The market is legally locked down and we&#x27;re in hostage pricing mode.<p>And what&#x27;s the story? You can&#x27;t afford electronics because we&#x27;re using it to build robots to take your job? I mean ...<p>Nobody is coming to save us. That&#x27;s our job.
      • phil211 day ago
        &gt; do they plan to increase production? No.<p>Micron has 3 brand new fabs currently under construction, 2 Boise, 1 in New York as the first of 4 planned for a campus.<p>Plus expanding other existing facilities.<p>These things take ~3-5 years from breaking ground to full production. You&#x27;d have had to anticipate the current demand years before it happened in order to be bringing production on-line before 2030 or so.<p>Samsung and HK Hynix also have fabs under construction and planned.<p>CXMT started 11 years ago and only now is reaching any real volume. If they decided a year ago to react to the current demand cycle they&#x27;d be 6-7 years out.<p>Not much you can really do to wish for more fabrication to exist on any timeline not measured in fractional decades.<p>Could they do more and react quicker? Probably, but everything I&#x27;ve read on the subject seems to point to 3 years is absolute bare minimum if you happen to have a shovel ready project with the land bought, local permitting completed, infrastructure extended to the site, and a skilled workforce already in place. They could suspend buy-backs&#x2F;dividends today and dump it all into building production and there would be no material impact until around 2030.<p>&gt; The Micron CEO just recently said this is the exact plan<p>CEO simply stated the demand pressure will not go away through 2027, and supply will not increase until around 2028 when currently under construction fabs start shipping volume. The article does not support your statement.
        • ttul23 hours ago
          Stanford tracks RAM prices in this nice little site: <a href="https:&#x2F;&#x2F;dam.stanford.edu&#x2F;memory-prices.html" rel="nofollow">https:&#x2F;&#x2F;dam.stanford.edu&#x2F;memory-prices.html</a><p>Costs did go nuts, but there are signs of easing in the market of late. CXMT is starting to have an impact and priced will probably fall in 2027.
          • kristopolous23 hours ago
            <a href="https:&#x2F;&#x2F;pcpartpicker.com&#x2F;trends&#x2F;price&#x2F;memory&#x2F;" rel="nofollow">https:&#x2F;&#x2F;pcpartpicker.com&#x2F;trends&#x2F;price&#x2F;memory&#x2F;</a> is better. pcpartpicker has the data.
          • fabioborellini10 hours ago
            You can see the SKUs they base the prices on by hovering over a data point, and the products just don&#x27;t generally represent the market. Oftentimes the whole market is represented by some bottom of the barrel legacy memory, single stick in some warehouse. During the last years DDR3 and DDR4 haven&#x27;t gotten as expensive as the current generations, and I think the memory compatible with current systems should represent the market rather than a Toshiba Satellite 4 GB extension kit.
        • GeekyBear21 hours ago
          &gt; CXMT started 11 years ago and only now is reaching any real volume. If they decided a year ago to react to the current demand cycle they&#x27;d be 6-7 years out.<p>It&#x27;s taken them this long to catch up to the DDR5 standard. They&#x27;ve only recently been through qualifications to be a DDR5 supplier for the big boys.<p>&gt; Every Major Motherboard Maker Now Validates CXMT DDR5<p><a href="https:&#x2F;&#x2F;www.techtimes.com&#x2F;articles&#x2F;321572&#x2F;20260725&#x2F;every-major-motherboard-maker-now-validates-cxmt-ddr5-most-us-builders-still-cant-buy-it.html" rel="nofollow">https:&#x2F;&#x2F;www.techtimes.com&#x2F;articles&#x2F;321572&#x2F;20260725&#x2F;every-maj...</a><p>After their recent IPO, they have more than enough cash to ramp up in a major way.<p>It&#x27;s just a matter of time.
        • cogman1023 hours ago
          The second Micron boise fab hasn&#x27;t even broken ground yet, they are still working on the first one. So don&#x27;t expect these things to be completed in parallel.<p>Some of my family is pretty happy, though, with the job security as they are pretty convinced these projects are all going to take much longer than what&#x27;s being stated publicly. Micron is saying the first chip from the new fab will be in 2027... though they also predicted it&#x27;d be 2026. The date seems pretty slippy.
          • minraws11 hours ago
            If memory prices cool, in about 2 years, Micron will stop new projects, they have done it before.<p>Especially given CXMT has been able to scale up much faster than what most people expected, only reason their isn&#x27;t a bigger impact is modern HBM is hard to CXMT even today.<p>We are likely to see supply double in the next 3 years, but demand even out with optimizations, cooling of data center demand, and most importantly moving some of the dram to flash demand instead which is much easier to produce and scale.
          • dboreham23 hours ago
            Anyone who has been around the semiconductor industry since the last century will remember various huge fabs e.g. in Arizona that were partially built but never finished due to oversupply by the time the walls and roof were done.
        • BizarroLand23 hours ago
          Yeah, but why would they make consumer memory when HBM for GPUs is much more profitable?
          • kristopolous23 hours ago
            Capitalism eats itself this way. Second and third order effects will collapse the demand.<p>You need to keep the market healthy, not some insane Bitcoin style HODL pump - that&#x27;s how you get wrecked.<p>I mean I&#x27;m not a neoclassicalist but I&#x27;ve read all of them. I&#x27;m in consensus with them here. There&#x27;s a bunch of theories on what a healthy market is but what we&#x27;re currently seeing matches none of them.<p>It&#x27;s short term profitable but long term disastrous, especially in a world where new mathematics and techniques could literally collapse the demand overnight.<p>Imagine if some paper hits arxiv and the 256 GB requirement for some model now becomes 64. Woops!<p>Some clever trick about how attention heads and context Windows work could potentially slash a bunch of requirements by giant margins and all they&#x27;re doing is firing the starting gun at that global race with every obscenely priced unit they sell.<p>But if prices were reasonable, this wouldn&#x27;t be an apocalypse. It&#x27;d be fine. Consumers wouldn&#x27;t rush to 64GB, they&#x27;d say &quot; Cool I can multitask now at 256&quot; or &quot; great I can do horizontal scalability&#x27; or something else.<p>But no they created the market conditions so now what would happen is the consumer will immediately flip the 192GB they don&#x27;t need on eBay, hoping to snatch a profit before the prices tank and the second hand market will be flooded the rug will be pulled out from the luxury pricing and everyone will get screwed.<p>This has happened in electronics markets before. Many times.<p>When Engels talked about the grave diggers of capitalism they were looking at it through a 19th century labor&#x2F;manufacturing lens but arguably this same dynamic is at play here.
            • Analemma_23 hours ago
              What &quot;second- and third-order effects&quot; do you suppose will collapse the demand for RAM? The people complaining most loudly about RAM costs are the people who want to run local models; if that becomes popular it will supercharge RAM demand, because locally-hosted models can&#x27;t parallelize runs from many users the way cloud-hosted ones can. I don&#x27;t see any slackening in RAM demand at any point in the foreseeable future, even if the big AI companies all go bust.
              • sieve19 hours ago
                &gt; The people complaining most loudly about RAM costs are the people who want to run local models<p>This is a tiny percentage of the population.<p>Samsung is cutting phone production because of RAM prices.[1] The consumer market is badly affected: budget phones, laptops, general electronics.<p>The budget segment of sub $100 devices in India has been almost wiped out. Manufacturers cannot afford to spend 50% BOM on RAM+storage. Unless employees are getting a 15-20% wage rise this year, I expect a similar situation in most places.<p>Between the engineered conflict in the ME triggering O&amp;G price rises, and stratospheric RAM pricing, the situation is pretty bad.<p>[1] &quot;There is no profit even if we sell&quot;…Samsung to cut smartphone production by 30% (<a href="https:&#x2F;&#x2F;www.mt.co.kr&#x2F;en&#x2F;tech&#x2F;2026&#x2F;10&#x2F;08&#x2F;2026100709554237233" rel="nofollow">https:&#x2F;&#x2F;www.mt.co.kr&#x2F;en&#x2F;tech&#x2F;2026&#x2F;10&#x2F;08&#x2F;2026100709554237233</a>)
              • kristopolous22 hours ago
                This is all hypothetical and debating hypotheticals isn&#x27;t productive so let&#x27;s roll back to markets.<p>Let&#x27;s say ram used to cost $100 and now that same unit costs $1000. You paid say $500x1,000 for that unit during the price increase or some price where you can currently flip for profit.<p>You have a very expensive data center and you&#x27;re in debt financed on the premise that you have these special computers.<p>Now a new technique comes out and it turns out you only need 1 memory unit for something that used to require 8 or 4 or some meaningful multiplier.<p>This stuff happens all the time. It&#x27;s why we don&#x27;t use BMP files on websites or serve giant MOV files on YouTube. It&#x27;s why postgres queries are faster now than they were 10 and 20 years ago.<p>You rent out your machines. You need to service your debt.. Demand may 8x overnight to accommodate but you have a monthly bill to pay and that&#x27;s unlikely. It&#x27;s likely going to drop.<p>Think about it. Your customers are paying maybe $10,000 a month and serving their customers. Now they can drop that to $1,250.<p>On market if you were to sell some of that ram you have 100% profit right now but not for long.<p>Jevons paradox assumes unlimited capitalization, zero debt servicing, infinite time horizons...<p>We live in the real world so what do you do?<p>Historically the answer has been &quot;sell that shit&quot;<p>There&#x27;s an aphorism for this &quot;stairs on the way up elevator on the way down&quot;<p>If we had a healthy market with sane prices where you can&#x27;t flip the thing you bought for 100% profit the answer would be &quot;create more value.&quot;
                • charcircuit22 hours ago
                  &gt;Your customers are paying maybe $10,000 a month and serving their customers. Now they can drop that to $1,250.<p>Or they could stay at $10,000 per month since they are willing to pay that much already.m, so they just use AI more and in more places.
              • lxgr22 hours ago
                &gt; locally-hosted models can&#x27;t parallelize runs from many users the way cloud-hosted ones can<p>Why not? Unlike many other workloads, LLM inference actually seems pretty suitable for decentralization (effectively stateless means no availability concerns; bandwidth and latency are relatively forgiving too).
                • Analemma_22 hours ago
                  I think locally-hosted models at the org level will definitely be somewhat popular, but you seem to be talking about decentralizing for people&#x27;s personal, non-business use, and I just don&#x27;t think that&#x27;s going to happen to any real degree.<p>People who say they want local runs really mean it: they want <i>local runs</i> on hardware in their room, not on some decentralized system which, if it existed, would almost certainly just be a worse, less-reliable version of cloud hosting. I&#x27;m not saying nobody would use it, but it sounds a lot like things like IPFS, which have also completely failed to displace either cloud storage or buying a bunch of disks for your own private use.
                  • lxgr13 hours ago
                    Some people will care a lot about keeping their data on-prem, but many others probably won&#x27;t, and the former can then resell their spare capacity to the latter.<p>Decentralized storage is much harder, since there reliability matters a lot more as it&#x27;s inherently stateful. You have to assume data loss, so you have to replicate everything; with inference, you only have to spend extra resources at failover time. Also storage can&#x27;t be time-shared in the same way as compute; if it&#x27;s full, it&#x27;s full even when not actively accessed.
                    • Analemma_2 hours ago
                      I mean I guess there&#x27;s nothing left which can settle this disagreement except to see how it turns out. I&#x27;ll just restate my opinion that, from a user experience and product perspective, decentralized model hosting will look like using the big labs, except worse in every way. I don&#x27;t expect it to please either the people who want local control or the people who are happy using the products from the big labs now.
            • usefulcat22 hours ago
              &gt; Imagine if some paper hits arxiv and the 256 GB requirement for some model now becomes 64. Woops!<p>If you were a DRAM manufacturer, isn&#x27;t this exactly the kind of thing that would make you think twice about investing years and $billions in new fab construction?
      • hn_acc122 hours ago
        &gt;Seriously, if a single politician stepped forward and said &quot;i&#x27;ll bring down ram prices&quot; they could then shoot a puppy and call me a slur and I&#x27;d still go out and doorknock for them.<p>How many people, outside of tech geeks and megacorps care about RAM prices? And how gullible would you be to BELIEVE the politician they could actually make it happen, and even if they did, that it would extend to the average person, and not JUST megacorps&#x2F;megadonors?
        • Aerroon20 hours ago
          Phone companies have been differentiating their models based on RAM for a decade. As have laptop and desktop sellers. The reason your router sometimes randomly crashes could very well be a result of not enough memory. The reason it takes such a long time to launch some programs repeatedly is because you don&#x27;t have enough memory to cache it. Swapped from your browser to an app on your phone, but when you go back to the browser the site has reset and you lost everything you were working on? Not enough memory. Etc.<p>I think a lot of people care about the downstream effects of memory prices, but I agree with you that they may not realize that they happen because of memory prices.
          • gruez19 hours ago
            &gt;The reason your router sometimes randomly crashes could very well be a result of not enough memory. The reason it takes such a long time to launch some programs repeatedly is because you don&#x27;t have enough memory to cache it. Swapped from your browser to an app on your phone, but when you go back to the browser the site has reset and you lost everything you were working on? Not enough memory. Etc.<p>That might be true at micro level, but at the macro level more memory just means developers get more lazy with their optimizations, causing apps to get more bloated, eating up any gains in extra memory. There&#x27;s no reason why slack needs 1+GB to run, yet people are perfectly happy to put up with it.
        • kergonath7 hours ago
          &gt; How many people, outside of tech geeks and megacorps care about RAM prices?<p>They don’t care about RAM prices, but they do care about the price of things that have RAM in them (or even NAND), and all of them are increasing way faster than inflation.
        • Fricken11 hours ago
          Everybody has a phone, they&#x27;re costing 25% more. GameStop is selling second hand PS5s for $1400. RAM prices affect most people.
      • mcv13 hours ago
        I keep arguing that memory needs more competition, and people keep pointing out that it&#x27;s too slow and expensive to ramp up. But if the threshold to enter that market is so steep, that means it cannot function as a free market and requires regulation.<p>In this case I think investment in more production is the only option, and it needs to happen even if it is expensive and slow.
      • ashdksnndck23 hours ago
        RAM manufacturers are bidding against NVIDIA and everyone else for the same constrained supply of EUV machines. And it takes years to build more fabs. Micron has multiple fabs coming online in 2027 and 2028.
        • dboreham21 hours ago
          Don&#x27;t believe Nvidia has any fabs of its own.
      • m4631 day ago
        &gt; &quot;i&#x27;ll bring down ram prices&quot;<p>wonder what voting would be like?<p>gamer vote ++<p>datacenter hater vote --<p>datacenter lobby ++<p>micron lobby --
        • rezonant23 hours ago
          Yep, that&#x27;s all the voting blocs.
        • mwambua1 day ago
          Wouldn’t cheaper memory make it easier to bring compute out of data centers and onto consumer hardware?
          • TeMPOraL23 hours ago
            Datacenter haters will read this as &quot;that&#x27;s still evil AI&quot;, and everyone else hopefully can count and understands it&#x27;ll be worse for environment.
      • absoluteunit19 hours ago
        &gt; Seriously, if a single politician stepped forward and said &quot;i&#x27;ll bring down ram prices&quot; they could then shoot a puppy and call me a slur and I&#x27;d still go out and doorknock for them.<p>I spit out my coffee laughing when I read this
      • BatteryMountain14 hours ago
        The situation is actually much worse and the long term consequences will start materializing soon. The wholesale theft of humanities soul is in progress. It won&#x27;t be a pretty sight in supposedly civil first world countries, when the human spirit awakens. Currently we are still pressing that snooze button hard and repeatedly, as I think most are keenly and deeply aware of what needs to happen but that too will cost our souls.
      • chrismsimpson13 hours ago
        &gt; they could then shoot a puppy and call me a slur and I&#x27;d still go out and doorknock for them<p>Priorities
      • j16sdiz11 hours ago
        &gt; Seriously, if a single politician stepped forward and said &quot;i&#x27;ll bring down ram prices&quot;<p>How? Increase production? The time needed to scale up the production is longer than one election cycle.
      • neya23 hours ago
        &gt; they could then shoot a puppy and call me a slur<p>I know it&#x27;s just a figure of speech, but damn. I laughed out aloud in public just reading this.
        • antonvs20 hours ago
          Kristi Noem would fit the bill, except for the bit about lowering memory prices.
      • bob102923 hours ago
        If we take some time to understand how HBM memory is manufactured (with particular focus on yield risk for final packaging steps), we will hopefully learn that the current capacity crisis is <i>not</i> bullshit.<p>I guarantee Micron &amp; friends are not intentionally orchestrating their business such that they would suffer a massively reduced chance of yielding on a per-die basis. Unless someone is actually buying HBM devices, they are not going to be making them. These are not a commodity that can be speculatively manufactured in any economically rational way.
      • javier212 hours ago
        Yeah, I was about to say, the memory industry has been found guilty of price fixing multiple times.
      • ggeorgovassilis14 hours ago
        &gt; Seriously, if a single politician stepped forward and said &quot;i&#x27;ll bring down ram prices&quot;<p>Or abolished VAT (the meaning of VAT is that you pay a &quot;rent&quot; for all the infrastructure used to produce the thing) and import taxes (protect your market) on stuff we don&#x27;t produce in our markets anyway.
      • fhn22 hours ago
        How many people would you allow them to kill?
      • Neither political party cares at all about memory pieces get real lol
        • Aerroon21 hours ago
          I don&#x27;t really understand why. Memory is a critical component of every computational device.
          • Exoristos21 hours ago
            That&#x27;s tautological, but I think you would need to explain how it extends their and their backers&#x27; influence to get party attention.
        • kristopolous23 hours ago
          wait until holiday shopping...it affects the price of almost everything with a battery or power cord.
          • xyzsparetimexyz22 hours ago
            What do you think they&#x27;ll do? Neither repubs nor dems will touch ai companies in a meaningful way. Anything China does wrt memory fabs week be more significant
            • kristopolous17 hours ago
              I don&#x27;t have faith in the political parties. Everything is insane. You look at platter recently? It&#x27;s up 3x in 12 months, not just ssd or nvme, but straight up traditional platter.<p><a href="https:&#x2F;&#x2F;web.archive.org&#x2F;web&#x2F;20250612003557&#x2F;https:&#x2F;&#x2F;diskprices.com&#x2F;" rel="nofollow">https:&#x2F;&#x2F;web.archive.org&#x2F;web&#x2F;20250612003557&#x2F;https:&#x2F;&#x2F;diskprice...</a><p><a href="https:&#x2F;&#x2F;diskprices.com&#x2F;" rel="nofollow">https:&#x2F;&#x2F;diskprices.com&#x2F;</a><p>After 70 years of decreasing computer prices all of a sudden it&#x27;s gone 3x, 5x, 10x up in 1 year, we are in total clown world and saying &quot;dur AI&quot; is lazy and doesn&#x27;t map to reality.<p>It&#x27;s Argentina style inflation - as if Honda said &quot;we&#x27;re only making $500,000 luxury cars now. Everything under $50k we&#x27;ve stopped.&quot; and then those cars shoot up to $125k.<p>It&#x27;s destroys the market, destroys the consumer, destroys the company, dismantles everything, and they do it for the short term payday.
      • gchamonlive1 day ago
        [flagged]
        • Analemma_1 day ago
          I don&#x27;t think RAM vendors have formed a cartel and I think this is knee-jerk anger without any thought. RAM is a commodity product with massive upfront capex costs, and those <i>always</i> have boom-and-bust cycles. At various points in the 2010s and 2020s RAM vendors were getting eaten alive by a supply glut, this would not have happened if they were a cartel.<p>Is it really so hard to believe that RAM prices are up because demand is simply exceeding supply, especially in a market where additional supply takes years and billions of dollars to come online? There&#x27;s no need to posit cartel behavior and a fair amount of evidence that there is none.
          • boustrophedon22 hours ago
            The RAM vendors have formed cartels previously and been convicted, so although demand is exceeding supply it is not that crazy to at least consider.
          • kristopolous23 hours ago
            The AI boom started in 2022. Prices rose THREE years later after 2025 sanction and tariff style legislation to protect the market during a price hike.<p>I got a 4090 in 2023 for 1600, a 5090 in 2025 for 2000 with 256 DDR5 for about $1,000 ... and then, after some protectionist legislation passed, these prices quickly shot to the moon.<p>Connect the dots.
            • Analemma_23 hours ago
              Man I think you&#x27;re just spewing word salad and a lot of what you&#x27;ve written is either wrong or not even wrong. The AI <i>hype</i> really got started in 2022, but hype on social media doesn&#x27;t mean anything for RAM prices, only real buildouts do that. They rose pretty steadily until OpenAI revealed their shenanigans re: locking up a ton of supply from two different vendors with secret contracts, and that&#x27;s when the takeoff really started. This is definitely scummy behavior from OpenAI (big surprise), and I actually think they arguably <i>should</i> see an antitrust investigation for that (not that that will ever happen), but OpenAI is a <i>buyer</i>; that&#x27;s not the same thing as the <i>vendors</i> forming a cartel.<p>You can&#x27;t say &quot;connect the dots&quot; at the end of a raving, mostly-incorrect post and act like you&#x27;ve made an ironclad argument.
              • kristopolous18 hours ago
                There&#x27;s nothing raving about it. I was trying to communicate that I&#x27;ve got all the hardware I want. I&#x27;m still pissed these companies are ripping people off.<p>It was supposed to be over by now and then they said 2027, then it&#x27;s 2028, and now I hear &quot;oh it&#x27;s going to continue to rise the rest of the decade&quot;.<p>I&#x27;m likely going to be flying into Shenzhen to put my next computer together. The one I put together in 2025 would have cost me about $25,000 right now. I paid under $5,000.<p>I can round trip to China for $750. So once they ramp up production that&#x27;s the strategy.<p>Other countries already do this. Apple and Google aren&#x27;t in every country and those people buy new electronics when they travel.<p>The USA is soon to be on that list<p>You aren&#x27;t engaging in good faith and there&#x27;s no reason to continue with you.
              • hn_acc122 hours ago
                I mean, just because OpenAI started it doesn&#x27;t mean the vendors didn&#x27;t form a cartel afterwards (or conspire together) to ensure maximum profits in a &quot;crazy high demand&quot; situation..
    • petu1 day ago
      There&#x27;s no BF16, original full quality weights are quantized already and 510GB.<p>Then good portion of those weights are n-grams (~200GB) that don&#x27;t need to be in VRAM.<p>Then KV cache of that model is super lightweight at ~1GB per 1M tokens. If HBF succeeds, then accelerator with 16GB of VRAM and 1TB HBF&#x2F;NAND is probably all you need (?).
    • wren699122 hours ago
      Are you counting the n-gram&#x2F;PLE as part of the model weights there? They can go in host memory. Would be good to show your working. Also the released weights are pre-quantised and presumably QATed, so your &quot;Full Precision&quot; and INT8 are simply not a version of the model that actually exists.<p>Edit: I went and checked for you. The LM backbone is 307.2 GB (286.1 GiB), straight from DeepSeek&#x27;s upload. The n-gram table is 203.1 GB (189.1 GiB), which goes in host RAM. Note the embeddings are higher precision than the expert tensors, so it&#x27;s a larger fraction of the bytes than it is of the parameters.<p>So,<p>&gt; Call me crazy but:<p>You&#x27;re crazy. :-)
    • mrinterweb23 hours ago
      Projects like DwarfStar <a href="https:&#x2F;&#x2F;github.com&#x2F;antirez&#x2F;ds4" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;antirez&#x2F;ds4</a> really lower the hardware bar a lot so Deepseek 4.1 flash and other mixture of expert models can run on consumer hardware. There are also other inference providers who make their money serving openweight models. Services like OpenRouter make it all too easy to utilize these models. Access to these models isn&#x27;t hard. The hardware moat is becoming pretty easy to bridge.
      • contingencies22 hours ago
        More concretely DwarfStar M5 128GB Deepseek 4.1 flash 1K tokens @ 29s, 5K tokens + reasoning @ 147s, 10k token prompt @ 463 tokens&#x2F;s = 22s. Hardware buy-in USD$7K &#x2F; AUD$8.5K &#x2F; EUR€6.8K. At typical workloads, ROI is still poor vs. current-era subsidies, but owning hardware is good for privacy&#x2F;longevity&#x2F;connectivity independence. Whether you actually consider Apple hardware &#x27;owned&#x27; is a valid and thought provoking question.
        • onlyrealcuzzo22 hours ago
          Still gonna take 2-3 years to get DeepSeek V4.1 Flash quality at decent speeds on reasonably priced hardware.<p>Hardware update cycles are 2-3 years even on the high end, so it&#x27;s still a ways away before &quot;good enough&quot; and &quot;local&quot; belong in the same sentence for the average person.<p>And by then, DeepSeek V6 Flash will be too cheap to meter, 5x faster, and 10x better, so... You&#x27;d still need to go out of your way.<p>Most people are spending most of their time on their phones anyway. ..
          • bitexploder17 hours ago
            Flash Next is a basically there. It really depends on what you are doing. This model is great. People forget that they felt Opus 4.6 was a great model and now you have it at home.
            • epolanski5 hours ago
              DS 4.1 flash is much much better than Opus 4.6, even quantized.
          • contingencies22 hours ago
            In the words of a Scottish comedian, &quot;The average person is Chinese.&quot; <a href="https:&#x2F;&#x2F;youtu.be&#x2F;LsEhNMy8svo?t=98" rel="nofollow">https:&#x2F;&#x2F;youtu.be&#x2F;LsEhNMy8svo?t=98</a>
    • ct52022 hours ago
      &quot;1070s or 1070 TIs because GPUs have been severely overpriced for too long&quot; ... .&quot;<p>1070ti launch MSRP was $450 ish. 5070 could be had in the last year for 5xx-6xx range easily.<p>All things considered - (inflation being about 30%~ (guess)) between these two timelines. You are looking at 300% performance difference at a cost dollar for dollar that is cheaper then when they purchased their cards.<p>Might be a bit of a stretch blaming it on &quot;severely overpriced for too long...&quot;
      • mrheosuper16 hours ago
        The 5070 honestly feel more like xx60ti class instead of xx70 class
    • ericd18 hours ago
      In nvfp4, it&#x27;s about 300 gigs once you offload n-grams, 491 without offloading, you can run it pretty well on 4x DGX Sparks, which last I checked was about $20k. So, it&#x27;s definitely runnable.<p>Or you can just use any of the neoclouds&#x27; shared hosting. The thing for them to be freaked out is that these models are getting good enough very quickly, and all the shared hosting providers can run them for a tiny fraction of what the frontier model companies charge.
    • lisplist18 hours ago
      DeepSeek V4.1 Flash is mixed MXFP4&#x2F;MXFP8 so all but the INT4 calculation is wrong here and that&#x27;s still wrong because you can run it on &lt; 400GB VRAM. The n-gram table is MXFP8, but can be offloaded to RAM or disk without too much of a performance hit. Really, you could probably cram it onto &lt; 300GB VRAM if you&#x27;re willing to apply a small quant to certain parts of the model considering how little VRAM is dedicated to kv cache.<p>I know my comment is a little nit picky because it&#x27;s still pretty expensive to run, but it&#x27;s not quite as bad as this comment makes it out to be. Really, if you&#x27;re VRAM constrained, take a look at GLM 5.3 Flash or Qwen 3.8 Flash Next before you worry about this model as all three models perform pretty similarly.
    • girvo23 hours ago
      Not quite: not all of this needs to be in VRAM<p>It has a set of n-gram tables which you can stream from system RAM or even NVMe<p>That said it’s still quite big! I can’t fit it on my DGX Spark, though I believe you can if you have two?
      • jonsoft22 hours ago
        It needs 3-4 Sparks to run well (at an acceptable quantization and sufficient KV cache):<p><a href="https:&#x2F;&#x2F;github.com&#x2F;christopherowen&#x2F;spark-ds41f" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;christopherowen&#x2F;spark-ds41f</a>
        • girvo22 hours ago
          Ah that’s a shame. GLM 5.3 Flash is honestly as good IMO and can run on two pretty successfully from what I understand.<p>I’m quite spoiled with how good Qwen 3.8 Flash Next is on a single spark though: shocking how good local models are getting on attainable-ish hardware
          • jonsoft22 hours ago
            DeepSeek V4 Flash runs well on two Sparks, I documented that here:<p><a href="https:&#x2F;&#x2F;blog.jonathanpage.com&#x2F;" rel="nofollow">https:&#x2F;&#x2F;blog.jonathanpage.com&#x2F;</a><p>GLM 5.3 Flash runs fine on two Sparks and Qwen 3.8 Flash Next on one is indeed incredible! I made this 3D game with it in two days using Qwen Code as agent:<p><a href="https:&#x2F;&#x2F;games.jonathanpage.com&#x2F;" rel="nofollow">https:&#x2F;&#x2F;games.jonathanpage.com&#x2F;</a>
          • bitexploder13 hours ago
            Having Flash Next local at 150 t&#x2F;s with 250k context is a joy. It’s as good as Sonnet 5. It will spaz out but it was less eager compared to DS Flash 4.1. Both are good but I find I prefer Flash Next. This and Qwen 27B are the models people should be freaking out about.
          • jonsoft15 hours ago
            I forgot to say, DeepSeek V4.1 Flash runs great on two DGX Station GB300s!<p><a href="https:&#x2F;&#x2F;www.storagereview.com&#x2F;review&#x2F;dgx-station-gb300-cluster-two-towers-two-400g-dacs" rel="nofollow">https:&#x2F;&#x2F;www.storagereview.com&#x2F;review&#x2F;dgx-station-gb300-clust...</a>
      • rsolva23 hours ago
        I have access to two and will explore this the coming weeks.
        • girvo22 hours ago
          Also give GLM 5.3 Flash a try: it’s shockingly good too in my testing, and I believe eugr has a TP=2 recipe to use for sparkrun
    • keammo123 hours ago
      The article isn&#x27;t just about running locally though. The author is saying it&#x27;s super cheap to run the model through Opencode Go (and presumably OpenRouter etc.) Personally I&#x27;m always most excited by models I can actually run locally, but even these huge open source models open up the competitive landscape for companies to let you call models via an API or just lease compute. And they don&#x27;t have to charge you to offset research, training, huge staffs of the best minds in the world, crazy PR etc. I think that&#x27;s a big win for customers and buts competitive pressure on the frontier labs as well.
    • ManuelKiessling23 hours ago
      Thanks for the data!<p>Allow a question from someone who’s only got a very vague idea of how this kind of stuff works behind the scenes: say I rent usage of this model through one of the many LLM hosting providers out there, and let‘s assume I use it extensively through something like Pi or OpenCode and vibe code away all the time, keeping the hosted model occupied as much as I can, happily burning my credits.<p>Does that mean that there is a hardware cluster as described by you above that is crunching away just for me?<p>So at FP16, I alone keep a 1,664 GiB system occupied all the time?
      • DrammBA23 hours ago
        No, a cluster can server multiple users at the same time, providers cap the tok&#x2F;s so that one cluster can run inference on multiple inputs at the same time. OpenAI with their new ultrafast mode is probably reserving the whole cluster or prioritizing requests of ultrafast users above others with a higher tok&#x2F;s hence the high price and high speed. There&#x27;s many other knobs providers tweak that they don&#x27;t show the users, for example I doubt many providers are hosting the full FP16 version.
        • jiggawatts19 hours ago
          It&#x27;s not based on rate limiting at all.<p>The &quot;expensive part&quot; of generating the next token is streaming in the model weights from memory. The computations are relatively simple, which is called a &quot;low arithmetic intensity&quot; in industry jargon.<p>So what they do is batch multiple chats together and compute the neuron activations for all of them together.<p>This is vaguely similar to how some database engines work, where if multiple users need to run a &quot;whole table scan&quot; query, the additional users &quot;join&quot; the streaming workload of the first query mid-way, then loop back around to complete the first part that they missed. The AI accelerators don&#x27;t do this looping, but the concept is the same: amortize the expensive I&#x2F;O over multiple computations running in parallel.<p>The &quot;turbo mode&quot; token rate thing is almost certainly your query getting sent to slower or faster hardware, like B200 vs newer B300 kit.
      • rnxrx22 hours ago
        It depends hugely on what &quot;rent usage of this model through one of the many LLM hosting providers&quot; means. If you&#x27;re asking them to host the model privately then yes, all of that 1.6T of RAM is likely in use holding weights, activations and KV cache by an inference engine that&#x27;s only getting&#x2F;answering requests from you alone. When you aren&#x27;t actively using the model the hosting process is still active and waiting with all of that memory still wired to it.<p>As background: For the most part VRAM oversubscription&#x2F;paging&#x2F;swapping isn&#x27;t a thing in the same way that RAM for a VM often is. There <i>are</i> some approaches to it, but (to my knowledge) not at that sort of scale.<p>There are some systemic reasons for this, but very broadly speaking the GPU vendors are building toward the highest bandwidth and lowest latency possible, and the overhead&#x2F;complexity of something like protected memory modes serves neither of those priorities.
    • ls_stats2 hours ago
      It&#x27;s funny because VRAM IS cheap to produce.
    • apitman23 hours ago
      &gt; Even so why would anyone not sleep on a model they cannot run?<p>Because it&#x27;s an open model so providers compete on price.
    • ByteAtATime1 day ago
      I think, considering the size of this model, it&#x27;s closer to a Pro than a Flash on everything other than speed
    • crossroadsguy23 hours ago
      I did somet math and completely gave up on the idea of trying any worthwhile local model and figured I&#x27;d rather pay the 15-30 USD per month via subscription and&#x2F;or API key combos for years than buying a local setup which might go out of date very fast, if it doesn&#x27;t goes kaput just out of warranty. I won&#x27;t be surprised if RAM scarcity is an concerted effort to herd people towards the remote models :)
    • cookiengineer23 hours ago
      It&#x27;s dangerous to go alone. Take this: [1]<p>I reimplemented most of the features of the Deepseek v4.1 flash paper (apart from quantization aware training which doesn&#x27;t make sense because my implementation uses float32 precision anyways)<p>I&#x27;m currently learning how to distill reasoning traces (check my other github repositories) but I think that a locally selfhostable deepseek is possible with my mixture of experts sharding mechanism. I decided to optimize everything for CPU parallelization, with the idea that the KV cache and meta model have to run from CPU RAM anyways, so the experts can also be loaded&#x2F;unloaded at runtime if needbe, to save more RAM.<p>My assumption is that the KV cache optimizations in combination with the CED and compressed attention features are the reason why v4.1 flash has so few hallucination problems and such a strong self-lookup&#x2F;thinking behavior. But that&#x27;s more a gut feeling, need to evaluate and test this more thoroughly.<p>Anyways, would love to see someone train this on their own datasets. Currently my pipeline is kinda optimized for parquet and zim files.<p>[1] <a href="https:&#x2F;&#x2F;github.com&#x2F;cookiengineer&#x2F;gonano" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;cookiengineer&#x2F;gonano</a>
    • anvuong1 day ago
      I just un-retire my pair of 1080Ti for some small models development because the current GPU prices literally make me sad.
    • nicman2314 hours ago
      what they should freak out about is qwen 3.8 github.com&#x2F;Niko1221&#x2F;Strata<p>it rips with just 64 ram and a 9070xt
      • sgt11 hours ago
        What would be the ideal qwen 3.8 variation for a 5090 with 32GB?
        • nicman239 hours ago
          q8 probably with the link above
    • CamperBob223 hours ago
      You can run it locally for the price of a decent car, or run it (hopefully) privately on somebody else&#x27;s hardware at vast.ai or a similar provider for much less. What&#x27;s not to like?<p>No, you won&#x27;t get frontier-level intelligence on a 1070Ti. Yes, it should be illegal to do what Altman did. Since we clearly don&#x27;t live in the best of all possible worlds, we need to settle, and DS4.1 Flash is a good place to do that.<p>For tasks that don&#x27;t require vision I personally like the NVFP4 quant of GLM 5.3 from Local Inference Lab better than DS4.1F, but they are both well beyond awesome.
    • nullc1 day ago
      The bulk of its weights are natively MXFP4. And engram values don&#x27;t need to be in vram.
    • dgellow9 hours ago
      Look at <a href="https:&#x2F;&#x2F;commandcode.ai&#x2F;pricing" rel="nofollow">https:&#x2F;&#x2F;commandcode.ai&#x2F;pricing</a>, you can have very decent volume of requests with deepseek flash v4.1 for literally $1&#x2F;month
    • functionmouse1 day ago
      one can make a fine gaming pc for ~$350<p>1660 ti, 4790k, 16gb ddr3
    • holoduke1 day ago
      He doesn&#x27;t ruin the cost of memory. Advances in memory size and speed are now in full speed mode. Expect drastic increase in the upcoming years. Big factories are in the making and planned. Gigalab in the US and many others in the east. Since 2010 we have computers with 16gb as being normal. Finally we are moving into a new era where the standard will be 64gb next year and 128 in 2028. Hopefully we reach 1tb in 2030.
      • I&#x27;ve read the exact opposite, vendors are reducing the standard from 16GB back to 8GB
        • holoduke23 hours ago
          That&#x27;s only temporary till production meets demand again.
          • CorrectHorseBat23 hours ago
            Which is not going to happen in the upcoming years
          • Tepix10 hours ago
            Remind me in 4 years then? Or 6? or 10?
    • liuliu1 day ago
      What are you talking about? The model is native NVFP4, why you run it at any precision higher than that?
    • jauer23 hours ago
      This “blame sama for memory prices” meme is so tired.<p>He gave demand signal so many times years ago and was mocked for it and now we have the consequences of industry not taking him seriously.
  • _jayhack_22 hours ago
    Enterprise is not freaking out because DeepSeek 4.1 Flash does not actually occupy a spot on the Pareto frontier for non-coding enterprise workflows. We see this at my employer, focused on non-technical knowledge work. Luna 6 and now Haiku 5.5 are both very competitive if not better on all axes that we care about
  • Kuyawa17 hours ago
    &gt; China is going to eat their lunch<p>No doubt about it, that&#x27;s why their push for international regulation to the levels of nuclear inspections using the narrative of annihilation and apocalypse
  • taf22 hours ago
    it&#x27;s ok but not great compared to the commerical LLM&#x27;s it&#x27;s really not very good. it&#x27;s benchmarks are clearly fake. even still we run it internally for a ton of workloads on our rack of gpu&#x27;s
  • james2doyle22 hours ago
    Been using Flash 4.1 via the ante harness to blast through a GBA recomp. The ante team has pushed hard to make Flash 4.1 perform well under it. So far, I&#x27;ve maybe spent $10 over the last 3 days. Its a real workhorse and works much better in this harness
  • Anoian6 hours ago
    If you use open models you can also read the thinking tokens, I think this makes a huge difference because one tends to read them leading to better understanding of the end results and also giving opportunities to correct course if its hung up about ambiguity and stuff like that. It also helps a lot with not being distracted while the clanker is working, instead of going to youtube or reading up on whatever, you spend more time with the code and the thoughts behind it.<p>It stops the constant context switching.
  • Roark667 hours ago
    Honestly, having benchmarked 3 models Qwen3.8-Flash-Next, GLM5.3 Flash and DeepSeek 4.1 Flash. I absolutely do not understand the hype about Glm and DeepSeek. It&#x27;s an improvement over older models, but Qwen is an actual opus replacement for me. It has been for last month.<p>But I run it locally. When I tried it on open router when my gpus were busy I must have gotten routed to some crappy providers, because it was pretty bad.<p>For me Glm and DeepSeek are nowhere near this Qwen model. I tried various harnesses including omp which I heard supposedly &quot;makes DeepSeek 20 points better&quot;. The difference was in the noise (1 point). I run a bunch of benchmarks Terminal World 40, terminal bench 2.1,SWE Pro, GSO. Before those 3 there was no open model that scored more than 1 point on my subset of GSO. Glm scored 5, DeepSeek 3, but Qwen did 17 and opus 19.<p>Qwen is a small model so it fails on factual recall. But if you give it most of the info it needs it us amazing.
  • WiSaGaN9 hours ago
    Deepseek v4.1 flash is my baseline model to use at original provider&#x27;s api. The issue is, for my personal use, it&#x27;s cheap enough that i don&#x27;t need it to go cheaper compare to the time I spent using it. And it&#x27;s already a very capable model in dealing everyday simple things. For sure, for research level questions or large scale coding projects, I would want to use frontier model. But more and more daily tasks can be done now just using pi with deepseek api directly without thinking about much.
  • robertheadley4 hours ago
    DeepSeek Flash 4.1 is generally what I use as as my workhorse. I use ChatGPT web to create the outline, then have Deep Seek built it out. Works generally pretty great.
  • peteforde12 hours ago
    I actually have a fairly simple answer to that: if it doesn&#x27;t come up in the list of LLMs that Cursor supports, it effectively doesn&#x27;t exist.<p>I&#x27;m well aware that there&#x27;s nearly infinite opportunities to yak shave &quot;perfect&quot; OpenRouter setups and some people appear to enjoy bouncing from IDE to IDE as though change costs aren&#x27;t a thing, but I discovered that I genuinely like Cursor and at least right now it&#x27;s insanely subsidized by Auto <i>clearly</i> defaulting to whatever Grok&#x27;s most powerful model is.<p>I dropped my $200&#x2F;month subscription to $20&#x2F;month and stick to Auto for all but really important Plan tasks, and I have basically zero chance of using up my monthly credits even using it 6-10 hours some days.
    • raincole11 hours ago
      &gt; I actually have a fairly simple answer to that: if it doesn&#x27;t come up in the list of LLMs that Cursor supports, it effectively doesn&#x27;t exist.<p>You make Cursor sound like one thousand times more important than it is. It&#x27;s a product in deep water.
  • wasfgwp14 hours ago
    Because it’s not even that cheap? The author chose to only include Claude in their chart and ignored the fact that 6.1-sol and even more so luna can easily beat Deepseek on cost. Of course almost free cache used to be the main differentiator, raw token cost is deceptive since 4.1 just uses way more tokens than most other models
  • ElProlactin14 hours ago
    &gt; Sure, they stole Claude&#x27;s training, and Anthropic stole it from other people. I&#x27;m not getting into the whole who-owns-whose-data debate, because most developers aren&#x27;t thinking like that. They&#x27;re just trying to get the most bang for their buck.<p>Until they get laid off and suddenly discover their moral compass.
  • maxdo7 hours ago
    The hype is almost reverse now . With opus , sonnet and haiku 5.5 , why should I care about Chinese models that are so behind ? Also Jev seems like killed a big chunk of deepseek market too .<p>I personally used it for secondary research agents and classification . Now classification part is gone .
    • newtwilly7 hours ago
      Yeah, OpenAI models now are really at the frontier of high performing, low-cost models: <a href="https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;#intelligence-comparison-tabs" rel="nofollow">https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;#intelligence-comparison-tabs</a><p>I agree with the author on a lot points about the joys of using low cost models. My work would pay much more, but I prefer to use cheap models usually. Assuming cost correlates somewhat closely with energy used, I feel good about spending as little as possible, and the cheap models are so capable.
      • maxdo5 hours ago
        I feel the opposite. An agent sometimes work an entire day . The chance of doing that work twice defeats any motivation to use cheaper , less capable models . Also if you open benchmarks leaderboard for the first time in a while there’s almost 0 Chinese models .
  • throwdbaaway2 hours ago
    Finally an article that gets the maths. Following the release of DSV4 preview where 1M context can fit in single digit GB of VRAM, frontier labs pricing just became stupidly expensive. Other Chinese labs and r&#x2F;LocalLLaMA also can&#x27;t compete on pricing.<p>And since then, there has been so many articles that made it to HN front page, and all of them didn&#x27;t get it. They just went on and on about tokens generation. Most HN commenters didn&#x27;t get it either, find-in-page for &quot;cach&quot; typically yield 2~3 responses. If I had a dime for every time this happened, I could have .. paid for 1B cached input tokens?<p>Anyway, DeepSeek still has to come up with a frontier model, and they almost did it with DSV4 Pro 0813, which is just slightly below GLM 5.3, but 30x cheaper. Unfortunately, the massive price hike happened just 3 days later.<p>DSV4.1 Flash is good, but not quite the same level. Much easy to self-host though, especially for serving a team of developers. Let&#x27;s see what the next one can do.
  • liuliu1 day ago
    DeepSeek 4.1 Flash 0910 is perfect for M5 Ultra 256GiB. Running it fully resident in RAM, prefill at ~2500 tok&#x2F;s and decode at ~40 tok&#x2F;s. Probably tons of room to improve from there.
    • loehnsberg22 hours ago
      Which quantization are you using there?
      • liuliu21 hours ago
        My own: <a href="https:&#x2F;&#x2F;huggingface.co&#x2F;drawthingsai&#x2F;DeepSeek-V4.1-Flash&#x2F;tree&#x2F;main" rel="nofollow">https:&#x2F;&#x2F;huggingface.co&#x2F;drawthingsai&#x2F;DeepSeek-V4.1-Flash&#x2F;tree...</a>
    • sdg03ksdv0d23 hours ago
      You are running this now? :o
  • zackangelo3 hours ago
    I&#x27;ve been loving DS4.1 Flash, it&#x27;s been one of my daily drivers since we started testing it internally.<p>We just launched it on our platform today (Mixlayer, <a href="https:&#x2F;&#x2F;mixlayer.com" rel="nofollow">https:&#x2F;&#x2F;mixlayer.com</a>), promo code LAUNCH-DSV41F gets you some free credits if anyone wants to check it out.
  • james22154 minutes ago
    this was beautiful to read thanks
  • seanmcdirmid12 hours ago
    It’s still cheaper to get a subscription to an agent harness with frontier model backing for most people. DeepSeek really only becomes appealing to me when I want to do something the monthly subscription harnesses don’t support (API access) or more harshly charge quota for (like running coordinated sandboxed subagents). DeepSeek is then great because of the low price and price transparency, it just can’t compete with subsidized monthly access.
  • jerieljan18 hours ago
    I&#x27;ve been on the API-only mentality for months and all the open models were definitely the stuff I loved the most. Kimi K2.5, 2.7 and Deepseek v4 were among my favorites, while sparingly using Opus or whatever OpenAI had for specific situations.<p>But ever since I&#x27;ve switched to one of the $100-tier subs, I can see why a lot of the people on it don&#x27;t really discuss the open models often. I&#x27;d still use it especially when it comes to sensitive inputs, but for most work, what you get on OpenAI or Anthropic is really more than enough.<p>It really got even better when they also made their cheaper models up to par if not better than the open models.<p>I do think the crowd for open models are out there, especially when you see trillions of tokens running for them on OpenCode or OpenRouter leaderboards.
  • anon37383913 hours ago
    I think the industry <i>is</i> freaking out about open weights models in general, if not specifically DeepSeek. That is why we&#x27;re now on the ~4th call for pacing the frontier from the very people who, if they wanted to pace the frontier, would simply do it rather than asking Washington to get involved.<p>And it&#x27;s why people like Hillary Clinton have been trotted out to talk about the dangers of open weights models -- I mean, does she even know what that phrase means? (I know HRC is a controversial figure and I&#x27;m not bringing her up for that purpose; I just note that she and other prominent retired politicians are now doing the circuit on Anthropic&#x27;s behalf.)
    • InsideOutSanta12 hours ago
      <i>&gt; if they wanted to pace the frontier, would simply do it</i><p>I don&#x27;t think that&#x27;s a fair assessment. These companies are in a Nash equilibrium where they can&#x27;t unilaterally slow down without essentially destroying their company. They also can&#x27;t coordinate with each other, because that&#x27;s illegal. Antitrust law generally prohibits competing companies from agreeing to restrict innovation.
      • nativeit6 hours ago
        So destroy the company. What are we even talking about here? I thought this was some “extinction level threat”, and you think it’s excusable to prioritize profit and product because…why?
        • tacitusarc6 hours ago
          Play that out. How does destroying the company halt the problem?
          • Retric6 hours ago
            If the options are to shoot yourself in the face or wait a year or even a month and have someone else shoot you, I think most people would pick the second.<p>Companies are continuing because they don’t actually believe they are going to have significant negative consequences happen sooner. It’s marketing fluff around how cutting edge what they are doing is not any kind of realistic threat assessment.
            • topicalSoup5 hours ago
              I don’t understand why this idea that it’s all just marketing is so prevalent on HN. It doesn’t make any sense. Why would a private company attempt to convince everyone that what it’s building is dangerous to its own customers?Can you think of a single other time in American history where an entire industry was calling to be regulated? I don’t seem to recall Facebook asking for tighter restrictions on social media. I don’t remember Shell telling the federal government that they need to act fast on climate change. Companies do not market themselves by convincing the public that their products might literally kill them all.<p>The only way this behavior makes sense is if the AI labs believe two things: one, that what they and the other labs are building is legitimately dangerous, and, two, that if they personally stop building then another lab will simply continue (and that lab will be, simply by virtue of not stopping, less safety-concerned than them). This is why they are calling for regulation. Competitive pressure and investor obligations prevents them from unilaterally stopping development. Government regulation is the only (and the most the appropriate) avenue for slowing down AI acceleration.
              • macNchz3 hours ago
                As others have mentioned there are lots of examples from history of businesses claiming their product is dangerous if not properly regulated, but I think what&#x27;s interesting in this moment is there appear to be a cohort of true believers who genuinely believe we&#x27;re headed for apocalyptic AI, and a cohort of win-at-all-costs businessmen who are happy to harness the earnest faith of the true believers in service of becoming (even more) unfathomably rich.
              • throwitaway2224 hours ago
                &gt; I don’t seem to recall Facebook asking for tighter restrictions on social media.<p>It has happened multiple times. Like the story now, it&#x27;s fake of course.<p><a href="https:&#x2F;&#x2F;www.thestranger.com&#x2F;tech&#x2F;as-zuckerberg-calls-for-new-internet-regulations-his-company-is-resisting-them-here-in-washington-state-39784781&#x2F;" rel="nofollow">https:&#x2F;&#x2F;www.thestranger.com&#x2F;tech&#x2F;as-zuckerberg-calls-for-new...</a>
              • nozzlegear3 hours ago
                &gt; <i>Can you think of a single other time in American history where an entire industry was calling to be regulated?</i><p>Several: the railroad industry, oil industry and trucking industry have all done this because it directly benefited them.<p>&gt; <i>I don’t remember Shell telling the federal government that they need to act fast on climate change.</i><p>The oil barons actively lobbied for Federal intervention in the 1930s, to limit supply and raise prices. You&#x27;re thinking recent history, but the oil industry has been around for over a century.
              • Retric5 hours ago
                Industry has often asked for regulations that have nothing to do with the claims being made. Zuck was a recent example in the computer industry. <a href="https:&#x2F;&#x2F;www.bbc.com&#x2F;news&#x2F;world-us-canada-47762091" rel="nofollow">https:&#x2F;&#x2F;www.bbc.com&#x2F;news&#x2F;world-us-canada-47762091</a><p>Here’s the same viewpoint from well outside the HN bubble. <a href="https:&#x2F;&#x2F;www.nationalreview.com&#x2F;corner&#x2F;be-wary-of-industries-asking-for-regulation&#x2F;" rel="nofollow">https:&#x2F;&#x2F;www.nationalreview.com&#x2F;corner&#x2F;be-wary-of-industries-...</a>
            • drob5185 hours ago
              I don’t necessarily believe that the AI itself will kill us all. But the AIs are increasingly being baked into complex, even military, systems. That scares me just because the opportunity for complex, cascading failure becomes very difficult to predict and control. These systems are inherently immoral (it’s just a lot of matrix math) and we know these systems “lie” and try to evade detection (see the Huggingface write ups). Now, couple that behavior with a military system (drones, targeting systems, etc.). I’m not suggesting that we’re headed for full-on Terminator land here, but things definitely could go wrong.
              • svieira5 hours ago
                &gt; These systems are inherently immoral (it’s just a lot of matrix math)<p>(nit) &quot;immoral&quot; is &quot;anti-moral&quot;. I suspect you want &quot;amoral&quot;, that is &quot;without&quot; or &quot;separated from&quot; morals.<p>But, if you wanted &quot;immoral&quot; ... what about matrix math derived from gradient descent is inherently &quot;immoral&quot; in your view?
                • drob5182 hours ago
                  Yes apologies. I did mean amoral. And now I can’t edit it.
              • Retric2 hours ago
                Mass adoption of technology always has negative outcomes because any change at scale has negative outcomes. Cars as a technology kill a great number of people and yet few are willing to ban them.<p>“AI” in the widest definition possible could do while a lot of harm without being a net negative. I’d be way more concerned about the healthcare sector than the military as the military already has significant collateral damage.
                • drob5182 hours ago
                  Sure, but cars weren’t hooked up to weapons systems. They killed as an accidental side effect. That’s why I’m saying I’m not worried about AI in a general sense.
                  • Retric6 minutes ago
                    Cars were quickly tuned into Tanks, APC’s, etc.<p>The fear around military AI is more about its failure than anything else, but nobody is trying to make fully automated machines which are manufacturing bombs from raw materials. Blowing up the wrong building matters a great deal to the people in the building, but for humanity a building was going to be destroyed either way. So the risk at any one point ends up being quite limited.
            • zamadatix6 hours ago
              I think most would think &quot;I can probably do a better job of not shooting myself in the face than some random person not worried about it will&quot; and then proceed to shoot themselves in the face.<p>That said, I don&#x27;t buy the companies actually give a damn about the risk either.
              • itishappy3 hours ago
                Agreed, but most people who care about not getting shot in the face aren&#x27;t giving away guns to their neighbors.
                • zamadatix1 hour ago
                  Nor are most other neighbors selling guns all day. At some point the analogy breaks down because companies aren&#x27;t actually people (though this seems to be a divisive take the last few decades).
          • geodel6 hours ago
            Oh okay, that makes sense. Since it won&#x27;t halt problem Anthropic and others might just as well profit from it.
            • andai6 hours ago
              I think there&#x27;s something here about game theory. I&#x27;m not sure what it&#x27;s called.<p>The logic that if something is beyond repair anyway you might as well exploit it.<p>For example I heard a pick-up artist say that he thinks he is harming civilization by sleeping with hundreds of women per year. But that he already considers the situation unsalvageable so... &quot;Might as well?&quot;<p>I don&#x27;t think that&#x27;s an amazing attitude, but the AI labs seem to have a similar idea.<p>More charitably the logic seems to be, &quot;only I can do this responsibly.&quot; I heard that from both Elon Musk (he cites this as his motivation for starting OpenAI — a chat with Sergei Brin that spooked him) and of course Anthropic (which split off from OpenAI due to ethical concerns).<p>So I don&#x27;t actually think the ethics is all for show. I think people are actually taking this stuff seriously. But it is indeed deeply unfortunate that the survival of the companies incentivizes them to keep going at an irresponsible pace (by their own admission).<p>All following the same gradient off a cliff. One AI described it as a tragedy of Ancient Greek proportions.
              • drob5185 hours ago
                Some people definitely are taking it seriously. And then there’s Sam Altman who I think would sacrifice his own mother for OpenAI’s IPO.
          • drob5186 hours ago
            Bingo. This is the unsolvable calculus. “If we don’t do it, someone will.” Which is quite true. We won’t stop until we see the whites of death‘s eyes.
            • bigbadfeline5 hours ago
              &gt; Bingo. This is the unsolvable calculus.<p>Actually, it&#x27;s easily solvable, even by proper application of existing law, but no solution is possible when big AI is above the law and can hack anybody willy-nilly. The absolute worst we could do is to listen to big AI&#x27;s pleas to put them in an even more privileged position.<p>&gt; “If we don’t do it, someone will.”<p>Do what? There are a lot more choices than &quot;let big AI freely hack everybody&quot; and &quot;stop everybody else from working on AI&quot;.<p>&gt; We won’t stop until we see the whites of death‘s eyes.<p>Trotting out <i>&quot;the whites of death ‘s eyes&quot;</i> is how real discussion is subverted into a BS binary choice. Open-sourcing all AI is an excellent starting point which will both slow down development to a naturally beneficial pace and allow the broader public to police and weed out bad models.
              • sleepybrett4 hours ago
                Open sourcing doesn&#x27;t mean much when it requires many millions to run one of these &#x27;full fat&#x27; models at any kind of real speed. Sure I can run a severly quantized model on my macbook but the performance is abysmal and the success rates go way down.
              • drob5182 hours ago
                You think laws will stop this? LOL
          • Moto74513 hours ago
            Let’s separate out what the companies actually want from the words they’re using for a moment. The top commercial companies are roughly stating:<p>1. Our AIs are so advanced they’re busting out of our networks and hacking people. We can’t seem to stop them.<p>2. There are real dangers to just how advanced our AIs are advancing.<p>3. Someone needs to slow us down<p>If these were true statements then they’d likely want to shut down themselves for the sake of the planet they live on. If there was a real problem as they have stated then shutting down is the solve (and to be clear, I’m not saying there aren’t problems with the AI industry. I just think they’re BSing us).<p>In reality this is marketing nonsense (better get in before the ban!) and “our” AI is never open weight&#x2F;Chinese AI models which they have a different lobbying arm against. If the top four companies got a loosely regulated monopoly through regulatory capture, they’d be <i>giddy</i>.
        • reasonableklout3 hours ago
          The previous OpenAI board already tried. It didn’t work, there are significant internal politics at play here. And in any case, the current 2 top players hate and mistrust each other and do not want the other to win.
          • mplewis2 hours ago
            Or they&#x27;re just all in it for profit and full of shit
            • _doctor_love2 hours ago
              Shh! Don&#x27;t say it too loud! They might hear you!
        • alwa4 hours ago
          At the risk of playing to popular stereotypes, Scott Alexander’s Meditations on Moloch (2014) comes to mind:<p><a href="https:&#x2F;&#x2F;slatestarcodex.com&#x2F;2014&#x2F;07&#x2F;30&#x2F;meditations-on-moloch&#x2F;" rel="nofollow">https:&#x2F;&#x2F;slatestarcodex.com&#x2F;2014&#x2F;07&#x2F;30&#x2F;meditations-on-moloch&#x2F;</a><p>Against that backdrop, and treating your question as literal: sounds like they’re saying it’s a classic coordination problem, with a dash of moral superiority. If we don’t, somebody else will; we’d rather it not happen at all, but if it’s going to happen, we’d rather it be us at the wheel. Because, you know. “Good guys” and all.
      • coldtea10 hours ago
        &gt;<i>without essentially destroying their company</i><p>Nothing wrong with that.<p>One was even a non-profit, which should be more concerned with the well-being of humanity (which they assure is in great danger from what they produce) than the continuation of the company.
        • mag726910 hours ago
          “was” is the keyword you used, but omitted to comprehend, in your analysis. OpenAI WAS a non-profit, as in, the past.
          • coldtea7 hours ago
            I comprehended it just fine, that&#x27;s why I used &quot;was&quot; to begin with. My comment, which you &quot;omitted to comprehend&quot; (sic), ironically refers to their non-profit heritage, which should make them double sensitive to good causes like not destroying the world.
          • pydry9 hours ago
            arguably these days it&#x27;s even more non profit than before and getting more so.
            • ohyes8 hours ago
              The non-profitest company in the entire world
            • mschuster918 hours ago
              However, should they be the one to survive the inevitable wave of collapses, they will be absurdly profitable, and that&#x27;s what all the investors are betting on - the usual VC playbook of creating an entirely new market (or destroying aka &quot;disrupting&quot; established players) fueled by absurd amounts of money, surviving until everyone is hooked and the competition is gone, and then jack up the prices.
              • Qiu_Zhanxuan4 hours ago
                The last one standing will probably a chip maker able to distill the frontier without having sunk too much capital advancing the frontier, essentially Google, Apple, Nvidia, Samsung and all the Chinese labs
              • coldtea7 hours ago
                And what makes you think first, there&#x27;s a moat and, second, the demand is not elastic?
                • mschuster917 hours ago
                  Don&#x27;t ask <i>me</i> that question, for all I care this entire bubble can finally collapse for good, I&#x27;d like to be able to afford things again.<p>But unfortunately, too many people who are already too rich for their own good have fully bought into it.
          • rightnutwingjob10 hours ago
            Never mind the fact that if any one, or even more than one, player gracefully bows out nothing changes.
            • pennomi7 hours ago
              If any one major player bows out, they could take everyone else out by opening up their model. There’s definitely a little bit of MAD going on.
              • bigbadfeline4 hours ago
                Good point! The same would happen if another player develops a frontier-comparable open model.
            • snovv_crash9 hours ago
              If A or OAI dropped out, there would be pretty crazy financial backlash from the whole inverted pyramid of investment and capital allocations that is built on top of their projected growth.
              • pydry9 hours ago
                that backlash is inevitable anyway. that pyramid is based upon projections which are flat out insane.
                • conartist68 hours ago
                  Why are people expected to lose their goddamn minds every 2-3 months like they&#x27;ve never seen an AI before.<p>I have a secret: if you don&#x27;t use AI, the new releases aren&#x27;t very impressive. I&#x27;m bored out of my mind with people doing galaxy brain memes every 2 months while on the whole.... they&#x27;re still boring zombies, and only getting boring-er and boring-er. There&#x27;s nothing more boring than being impressed by the latest AI model.<p>In terms of the projections being insane you can quantify it: each model costs something to train, but after that the training has only captured so much unique new value in terms of model capability, and the race is on to drain that value as rapidly as possible. Everyone is competing to drain the same value. This generation of models makes slop video games for example, and slop video games rapidly became the most boring thing on the planet.
                  • Roark667 hours ago
                    I tell you 3 times I was genuinely excited with AI. 1 - original Google transformers release (hope for future) 2- opus 4.6 comes out (first time an llm can do real programming for me and the code is good) 3 - Qwen3.8-Flash-Next comes out - first time a model i can run myself at good speed can do real programming and is almost as good as opus.<p>Those are genuine breakthroughs.
                    • conartist62 hours ago
                      Again that suggests that <i>the end goal is to use AI</i>. As long as using AI is your goal, using AI is also the definition of succeeding at your goal.<p>Any society-wide productivity gain was lost when software engineers ceased collaborating with each other in good faith. In OSS, they stopped collaborating. At companies, they stopped collaborating. One by one, in silence, alone, we have been bereft of our vision, our team spirit, our passion, our uniqueness. We no longer have the will to serve others, or even to serve each other. We stopped building our future. We will accept whatever society AI makes for use like fucking cattle or lemmings or like dogs.<p>AI hasn&#x27;t elevated humanity, it has taken a massive dump on it.
        • kshri247 hours ago
          Yeah well OpenAI did destroy the non-profit part at the very least.
        • tiborsaas9 hours ago
          &gt; Nothing wrong with that.<p>Except that&#x27;s an open invitation to get sued into oblivion by your investors.
          • mkeeter9 hours ago
            Quote the original OpenAI prospectus, “it would be wise to view an investment in OpenAI Global, LLC in the spirit of a donation”.
          • gewetensleegte9 hours ago
            That&#x27;s the risk they took right? They got billions for it
            • tiborsaas9 hours ago
              No, they took the risk of being outcompeted, which is exactly my point
            • rightnutwingjob9 hours ago
              <i>CEO wakes up one day and decides to scuttle the company</i> probably isn’t in the contract.
              • jquery7 hours ago
                Slowing pacing because of real fears of out of control AI isn&#x27;t scuttling, <i>if</i> they believe what they&#x27;re saying.<p>Of course they don&#x27;t actually believe their nonsense about AIs doomsday hokey, so full speed ahead.
          • rightnutwingjob9 hours ago
            The investors would remove whichever numpty attempted to unilaterally wind the company down.<p>The CEOs of these companies are largely just talking heads and generally interchangeable &#x2F; hot-swappable with anyone willing to give it a shot.<p>The importance of CEOs is largely overstated.
            • idbnstra7 hours ago
              openai&#x27;s removal and reinstatement of sam altman in november 2023, do you think it supports or disproves this?
              • pyrale5 hours ago
                I would say it does neither, really. The removal issue was a conflict between stakeholders on whether OAI should keep its nonprofit ethos, or move to be a commercial entity. Had Altman been replaced by someone supporting the same line, there wouldn&#x27;t have been such a fight.<p>If you wanted to test your hypothesis, you would have to find a situation where leadership change wasn&#x27;t caused by strategy change.
        • robertlagrant9 hours ago
          &gt; Nothing wrong with that.<p>There definitely are many things wrong with that, of course. Lighting someone else&#x27;s money on fire and walking away is highly unethical.<p>The real question is: what are the motivations for companies to pretend that their AI is world-ending, and what do they want out of the disquiet caused by them saying that?
          • shafyy9 hours ago
            Haha, do you think those people are concernced with ethics?
            • robertlagrant8 hours ago
              Do you mean you also disagree with the person above saying &quot;Nothing wrong with that&quot;?
        • forgotaccount37 hours ago
          Except it doesn&#x27;t actually slow down the research. If anything it would likely drive investors to contribute more to the competitor that didn&#x27;t destroy their own company and now you have the same amount of progress and a monopoly.
          • nativeit6 hours ago
            “I couldn’t possibly take action to save the Earth unless <i>I get to be the only</i> person taking action to save the Earth.”<p>Well that certainly tracks with what I know about these people and their twisted philosophy.
          • catlifeonmars6 hours ago
            Isn’t this essentially the same argument used to justify nuclear proliferation?
      • disgruntledphd29 hours ago
        &gt; essentially destroying their company.<p>Look, nobody made them take this much funding, and believe that scale was all they needed.<p>Turns out they made a bunch of bad investments, and are suffering the fate of many startups who invented something but couldn&#x27;t profit from it.<p>There&#x27;s definitely a case for multiple labs&#x2F;models across the US&#x2F;EU&#x2F;China&#x2F;India etc, but nobody&#x27;s entitled to a financial return on their investment.
      • roarcher10 hours ago
        &gt; they can&#x27;t unilaterally slow down without essentially destroying their company<p>According to them, the alternative is destroying the <i>world</i>.<p>They made up a shitty excuse to regulate the competition without thinking through what it implies about them. That&#x27;s all this is. Let&#x27;s not help them make even more excuses.
        • numpad08 hours ago
          Hold on, the implication of &quot;it would be a shame if anything happened to $nice_thing&quot; is usually &quot;I will kill&#x2F;destroy the $nice_thing if you didn&#x27;t take my offer&quot;, and here, nice_thing = &quot;the world&quot;.<p>Are they saying that they will manually and willfully start a global thermonuclear war if China didn&#x27;t listen, not that there are risks of AI accidentally causing one? Is that what they&#x27;re trying to say?
      • King-Aaron12 hours ago
        lol mate they&#x27;re all sitting together in rooms plotting market plays with the president.<p>Get our of here if you think the law is what prevents them from doing things.
        • ben_w10 hours ago
          There is literally a lawsuit against them for having said publicly they want to slow down: <a href="https:&#x2F;&#x2F;aitechmodel.com&#x2F;ai-slowdown-antitrust-lawsuit-pace-the-frontier&#x2F;" rel="nofollow">https:&#x2F;&#x2F;aitechmodel.com&#x2F;ai-slowdown-antitrust-lawsuit-pace-t...</a>
          • serial_dev10 hours ago
            And did this lawsuit prevent anything, did they correct any of their behavior?<p>They will just pay whatever they need to to their lawyers, then maybe pay a fine, but then, I guarantee you, nothing will change.
            • ben_w10 hours ago
              &gt; And did this lawsuit<p>&quot;Is a lawsuit&quot;, not &quot;was a lawsuit&quot;. Present tense, ongoing, not past tense.<p><pre><code> Status: Complaint filed September 18, 2026 in the Northern District of California · responses due October 14–15, 2026 · initial case management conference December 23, 2026 · no class, no settlement. </code></pre> - from the linked article.<p>&gt; did they correct any of their behavior?<p>The behaviour being objected to is <i>publicly agreeing with each other to slow down</i>.<p>If a court orders them to correct this behaviour, it means <i>they are forbidden from agreeing to slow down</i>.
              • dspillett9 hours ago
                So no, it has not yet prevented anything. Which I think is the previous posters point: there is currently nothing stopping them from doing what they say they want to see happen. The problem from their PoV is that other players (those working on open source models, those in other countries who are never going to listen to any directive to slow down anyway, etc.) won&#x27;t also collude with them so even if they do collude to do what they say they want everyone to do it&#x27;ll be to their disadvantage, where what they actually want is for everyone else to be told to slow down while they somehow pay for a loophole to allow them to go a little faster for competitive advantage.<p>The speed of change is impressive, I don&#x27;t think any tech ever before has gone from “brand new disruptor in public awareness” to “the incumbents feeling they have insufficient moat and so trying to arrange a regulatory capture situation” in such a short space of time.
                • estearum9 hours ago
                  &gt; what they actually want is for everyone else to be told to slow down while they somehow pay for a loophole to allow them to go a little faster for competitive advantage.<p>It’s hilarious that you can write this and then act befuddled as to why they would therefore want <i>laws</i> as an external (and ideally impartial) coordination device.<p>Thinking this is a super unique scenario just reveals your ignorance of both 1) game theory and 2) actual industrial history. An industry asking for regulation to stop a race to the bottom is not atypical at all.
        • Turneyboy12 hours ago
          The claim is that competition is what&#x27;s preventing them from unilateral disarmament.
          • King-Aaron12 hours ago
            &gt; They also can&#x27;t coordinate with each other, because that&#x27;s illegal.<p>While I completely understand the claim I am addressing one part of it.
        • InsideOutSanta11 hours ago
          <i>&gt; Get our of here if you think the law is what prevents them from doing things.</i><p><a href="https:&#x2F;&#x2F;www.bbc.com&#x2F;news&#x2F;articles&#x2F;c932g3v3e13o" rel="nofollow">https:&#x2F;&#x2F;www.bbc.com&#x2F;news&#x2F;articles&#x2F;c932g3v3e13o</a>
          • dspillett7 hours ago
            That is them stopping <i>us</i> using it. I doubt <i>they</i> stopped working on it. And do we know that they aren&#x27;t letting US government interests tinker too?
      • mtlmtlmtlmtl11 hours ago
        If you look at their actions, they&#x27;re not the actions of someone who both A) believes their technology to be an existential threat and B) doesn&#x27;t want humanity to go extinct<p>If they truly believed both of those things, they would just shut down their companies, because being a billionaire is entirely pointless if you&#x27;re dead.<p>I can only conclude that they don&#x27;t believe both A and B. I&#x27;m gonna assume they believe B because actually wanting to exterminate humanity is too comically supervillain esque even for Altman. Therefore they must not believe A. And it makes sense. If they believe A, why are they so incredibly sloppy about security? Why did they outsource part of it to some external firm instead of leveraging their own expertise on the technology? I&#x27;m not buying it.<p>Instead, I think they believe C) that AI will create enormous economic value, D) that value will be distributed across the whole economy by making everyone more productive. Assuming C and D, you get E) for them to capture this value, they must maintain proprietary control of the technology in order to be able to charge everyone else for the privilege of using it.<p>Assuming they believe E, open models are an existential threat, not to humanity, but to OpenAI and Anthropic.<p>Their solution: make <i>everyone else</i> believe A, in order to achieve regulatory capture and somehow stop open models from advancing by banning their development, or something. This is a hail mary pass. I can just about imagine them achieving this within the US and maybe even Europe, but China? That ship has sailed.
        • ben_w10 hours ago
          That&#x27;s an easy reading, sure.<p>Unfortunately, many (most?) of them also seem to think they&#x27;re better equipped than all of the others to do it safely, so &quot;shut down my own company&quot; effectively means &quot;let one of the others destroy the world&quot;, whereas &quot;keep company running&quot; has at least a chance of &quot;I solve alignment, we have a happy ending&quot; (your case C).<p>Note: this does not mean I agree with them. Obviously they can&#x27;t all be correct that they&#x27;re safer than the others.<p>&gt; If they believe A, why are they so incredibly sloppy about security? Why did they outsource part of it to some external firm instead of leveraging their own expertise on the technology?<p>The tech they&#x27;re experts at is AI, not security, which answers both parts of that.<p>Outsourcing things you are bad at is normal, not a mystery. Lots of places value physical security, <i>and therefore hire a private security firm</i>.
        • Sankozi9 hours ago
          They can believe A and B and at the same time they don&#x27;t want to cripple themselves. Your logic does not hold up.<p>Similar situation: nuclear race - everybody knew the risks, but making yourself armless does not help in any way. You need to make sure everyone is on the same page before you make yourself vulnerable in any way.
          • saghm9 hours ago
            If there&#x27;s a company with an insanely high valuation, and two potential motives consistent with their actions, one altruistic and one greedy, I&#x27;m struggling to understand why the assumption would be altruism, rather than the overwhelmingly more common primary motivation of &quot;we want to make a boatload of money&quot; while finding it useful to pretend otherwise, especially when so many sizable investors have similar sizable expectations for their returns.
            • Sankozi8 hours ago
              I don&#x27;t see the greedy path - why would stopping research - the biggest advantage frontier labs have, did bring any value to them?
              • itintheory7 hours ago
                It seems to me the ROI on the research may be diminishing. If they aren&#x27;t turning a profit with the infrastructure already deployed and the models already trained then buying more GPUs and spending billions more on training doesn&#x27;t seem like it would improve the correct side of the balance sheet.
        • nico_h11 hours ago
          My problem with the “making everyone more productive” is that it’s a feel-good propaganda for the individual. Most capitalist see people as an expensive cost to be eliminated. Thus all the company shares going up when they announce layoffs.<p>The market can only absorb so much new products, so if productivity increase, companies will reduce headcount as much as possible to increase margin for the same income, not grow their product or production to make use of their staff.<p>(See any wage&#x2F;productivity graph)<p>Government will also likely follow the same path of reducing headcount instead of producing better &#x2F; faster outcome for their citizens (except the internal surveillance apparatus. The one never shrinks).
          • simonh9 hours ago
            Reducing the cost of a desirable service or product often has a tendency to increased demand for that service or product, and means customers have more capital available to purchase other products and services increasing demand for those. These factors are significant drivers of overall economic growth.<p>Of course economic growth has it&#x27;s own problems in terms of ecological impact and such, but if you want to reduce ecological impact you still need to improve productivity. It&#x27;s a matter of how you spend that efficiency improvement as a society.
          • mtlmtlmtlmtl11 hours ago
            Yeah, I&#x27;m not convinced either. I&#x27;m just saying that&#x27;s what I think the AI labs believe.
        • jjav11 hours ago
          &gt; actually wanting to exterminate humanity is too comically supervillain esque<p>An entirely plausible scenario is that these oligarchs dream of living in Solaria, the planet from the Asimov universe where only a small number of immensely rich people lived and all the work was done by hordes of robots.<p>Once humans no longer serve the needs of the oligarchs, why would they want billions of people around? A question to ponder.
          • coldtea10 hours ago
            Just a few tens of thousands to entertain them, including kids for their &quot;islands&quot;, would be enough.
          • SXX10 hours ago
            Whole point is that billiinares already have whatever they want without any meaningul control from society.<p>Humans serve them well enough and relatively easy to control. Robot utopia is not guatanteed and have own risks.<p>And fortunatelly in a group of AI overlords everyone except of Musk is pretty sane.
            • simgt10 hours ago
              If they were rational people they&#x27;d likely have stopped their quest after the first few millions. They are driven by pure greed and their humongous egos, none of them is sane.
              • SXX9 hours ago
                I dont think there is anything wrong with people driven by pure greed and wealth maxxing.<p>They are far safer people compared to those who dont care about virtual wealth numbers.
              • BlueTemplar10 hours ago
                Some of it might be that, but some of it might be more prosaically the same thing that keeps people playing competitive games on one hand, or cookie clicker on the other. (Cookie clicker could read as &quot;pure greed&quot; if you squint.)
                • pferde9 hours ago
                  Yes, but me playing cookie clicker (or fortnite, or whatever, competitive tic-tac-toe) does not bring with itself suffering and poverty for millions of others. Whereas them chasing their virtual money high score ruins societies.<p>They are all psychopaths, devoid of normal human emotions.
            • Muromec9 hours ago
              &gt;Whole point is that billiinares already have whatever they want without any meaningul control from society.<p>Or maybe they don&#x27;t. Some of those are deeply afraid of society looking funny at them.
            • birdsongs10 hours ago
              &gt; Humans serve them well enough<p>For now. Throw in climate change, food shortages, war, mass migration, and suddenly the ratio becomes 25 million : 1 for starving, angry people vs billionaires, globally.
          • Joker_vD9 hours ago
            <p><pre><code> &quot;I propose something different. Listen, my dear enemy... I shall acquire absolute power on earth. Not a single chimney will smoke unless I order it, not a single ship will leave harbour, not a single hammer will strike. Everything will be subordinated — up to and including the right to breathe — to the centre, and I am in the centre. Everything belongs to me. I shall engrave my profile on one side of little metal discs — with my beard and wearing a crown — and on the other side the profile of Madame Lamolle. Then I shall select the &#x27;first thousand,&#x27; let us call them, although there will be something like two or three million pairs. They will be the patricians. They will devote themselves to the higher enjoyments and to creative activities. Taking an example from ancient Sparta we shall establish a special regimen for them so that they do not degenerate into alcoholics and impotents. Then we shall determine the exact number of hands necessary to give full service to the culture. In this case, too, we shall resort to selection. These we shall call, for the sake of politeness, the toilers—&quot; &quot;It goes without saying—&quot; &quot;You may laugh, my friend, when we get to the end of this conversation... They will not revolt, oh, no, my dear comrade. The possibility of revolution will be destroyed at the very root. A minor operation will be carried out on every toiler after he has qualified at some skill and before he is issued a labour pass. Quite an unnoticeable operation made under almost accidental anaesthesia... Just a small perforation of the skull. He will get a bit dizzy and when he wakes up he will be a slave. Lastly, there will be a special group that we shall isolate on a beautiful island for breeding purposes. All those left over we shall have to get rid of as useless. &quot;There you have the structure of the future mankind according to Pyotr Garin. The toilers will toil and serve uncomplainingly, like horses, for their food. They will no longer be people and they will have no worries except hunger. They will find happiness in the digestion of their food. The elite, the patricians, they will be demigods. Although I despise people altogether, it is always pleasant to be in good company. I assure you, my friend, we shall enjoy the golden age the poets dream of. The impression of the horrors created by purging the earth of its surplus population will soon be forgotten. On the other hand, what opportunities for a genius! &quot;The earth will become the Garden of Eden. Births will be regulated. There will be selection of the fittest. There will be no struggle for existence, that will be lost in the haze of the barbaric past. A beautiful and refined race will develop, with new organs of thought and sensation. Communism trying to drag all of humanity to the heights of culture? I instead will do it in ten years... What the hell! In less than ten years! For only a few, true... But then, it is not a question of numbers.&quot; &quot;A fascist Utopia, rather curious,&quot; said Shelga. &quot;Have you told Rolling anything about this?&quot; &quot;It is not a Utopia, that&#x27;s the funny part of it. I am only logical.&quot; </code></pre> Well, what was considered reasonable enough to state outright in 1906 or in 1926, is not quite polite to say out loud in 2026, but I don&#x27;t think the sentiment of contempt has ever quite gone away entirely.
            • dflock5 hours ago
              From <a href="https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;The_Garin_Death_Ray" rel="nofollow">https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;The_Garin_Death_Ray</a> if anyone else was wondering.
        • saghm9 hours ago
          It&#x27;s not clear to me that what the evidence for D is. It could just as easily be that they want to translate popular fear&#x2F;distrust of AI and desire for regulation into something anticompetitive that benefits them over the smaller players rather than being regulated themselves.
        • BlueTemplar10 hours ago
          There are some nasty game theory outcomes to keep in mind here. (See also : cold war nuclear arms race.)<p><a href="https:&#x2F;&#x2F;www.lesswrong.com&#x2F;posts&#x2F;hc4DbmhdzZpSLMQ9Y&#x2F;the-ai-race-is-not-a-prisoner-s-dilemma" rel="nofollow">https:&#x2F;&#x2F;www.lesswrong.com&#x2F;posts&#x2F;hc4DbmhdzZpSLMQ9Y&#x2F;the-ai-rac...</a>
      • coldtea10 hours ago
        &gt;<i>They also can&#x27;t coordinate with each other, because that&#x27;s illegal.</i><p>I also can&#x27;t legally sell this magnificent bridge I&#x27;m offering you, but I have one for sale with all proper paperwork intact!
      • freeopinion7 hours ago
        Companies don&#x27;t think. Companies don&#x27;t make decisions.<p>I appreciate the analysis and following a string of thought. But sometimes you eliminate so much reality to follow a thought that the string becomes kinda pointless.<p>Amodei is worth several $billion, right? He could simply walk away today. There is no Nash equilibrium for him. Nor for his replacement, nor their replacement. And it is not illegal for people to collude in quitting their jobs.<p>There are two obvious arguments for not quitting. 1. If the well-intentioned person quits they will be replaced by a non-well-intentioned person. 2. There is no real belief in a hazard and you would be giving up unbounded income for no reason.<p>If the actions and behaviors of a supposedly well-intentioned person who is afraid of a hypothetical non-well-intentioned person are indistinguishable from those of a supposedly non-well-intentioned person... don&#x27;t we really just end up with a really long sentence with a lot of gibberish + the outcome of having a non-well-intentioned person in power?
        • RobotToaster5 hours ago
          &gt; Companies don&#x27;t make decisions.<p>Cybernetics would disagree, any social structure can form feedback loops that make decisions that no rational individual would make.
        • CamperBob23 hours ago
          <i>Companies don&#x27;t think. Companies don&#x27;t make decisions.</i><p>&quot;They are just stochastic parrots&quot; didn&#x27;t age well with LLMs, and I don&#x27;t think it ever applied to organizations of humans.
      • pyrale5 hours ago
        &gt; These companies are in a Nash equilibrium where they can&#x27;t unilaterally slow down without essentially destroying their company.<p>If it was only US tech companies, they would have entered a cartel agreement already ([1], [2], etc).<p>[1]: <a href="https:&#x2F;&#x2F;www.justice.gov&#x2F;archives&#x2F;opa&#x2F;pr&#x2F;justice-department-requires-six-high-tech-companies-stop-entering-anticompetitive-employee" rel="nofollow">https:&#x2F;&#x2F;www.justice.gov&#x2F;archives&#x2F;opa&#x2F;pr&#x2F;justice-department-r...</a> [2]: <a href="https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Jedi_Blue" rel="nofollow">https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Jedi_Blue</a>
      • vrighter10 hours ago
        they can&#x27;t slow down because the investors are circling. But if they were <i>forced to</i> slow down by the government then they can use the zour hands are tied&quot; excuse
      • bfeynman5 hours ago
        If there&#x27;s one thing that politicians and strategists understand as much if not better than business people it actually is stuff like this. Most people never don&#x27;t deal with high stakes situations at certain levels where everything is fungible(the law, media statements) and fuzzy (intelligence). Dealing and negotiating with nation states or figure heads who are adversarial&#x2F;allies and at same time trade partners&#x2F;enemies etc. They have certainly sold this situation as one that fits the regime for this.
      • qznc9 hours ago
        They can coordinate in the open. For example, safety regulation in the automotive industry is effectively self-regulation.
        • estearum9 hours ago
          The automotive industry is not conceivably a winner takes all competition.
      • alistairSH8 hours ago
        If they truly think Skynet is nigh, then they should.<p>“This thing I building will destroy the world, but I’m making too much money to stop myself, please force me to stop” is such a weird position.
        • amelius7 hours ago
          It really shows that free market capitalism is broken at its core.<p>I always felt like one day it would end the world, I just didn&#x27;t know how. Now I know.
      • stavros12 hours ago
        That&#x27;s the GP&#x27;s point: &quot;Pace the frontier&quot; really means &quot;slow China down because our investment is at risk&quot;.
        • InsideOutSanta12 hours ago
          What I&#x27;m trying to say is that their inaction in unilaterally slowing development is consistent with their stated beliefs and requires no other motivation.
          • anon37383912 hours ago
            Have you looked at this backwards, though? Meaning: have you considered what it would look like if the motivations are actually just plain old money&#x2F;power, but a set of stated beliefs needed to be constructed to justify what&#x27;s being sought? I think these absurd beliefs make more sense that way.
            • InsideOutSanta11 hours ago
              Sure. I&#x27;m not saying that they have no other motivations. I&#x27;m saying that, even if they had no other motivation, they would still behave like this.
            • estearum9 hours ago
              Tip: Those motivations are not separate from the AI risk and do not mitigate it. They <i>amplify</i> the risk.
        • khazhoux9 hours ago
          &quot;Slow China down because our investment is at risk&quot; is a perfectly reasonable position for a US company, and I don&#x27;t mean that in any negative way.
          • stavros9 hours ago
            Well yes, but then don&#x27;t pretend that it&#x27;s &quot;because safety&quot;.
            • estearum9 hours ago
              &gt; multiple LLMs escape containment consistently, conspire with each other to hide from human oversight, show wanton regard for laws that stand in the way of their goals, and end up hacking dozens of different organizations<p>&gt; “there’s no plausible way they’re concerned about safety”
              • stavros9 hours ago
                &gt; Cares about two things, A much more than B<p>&gt; Someone points out that he pretends to only care about B<p>&gt; Random commenter: &quot;there&#x27;s no plausible way we care about B&quot;<p>Good chat.
                • estearum8 hours ago
                  You called it pretending to ask for slowdown based on safety concerns. Not my fault if that isn’t a complete description of your beliefs
      • eli5 hours ago
        I don’t think the second part is true. Companies work together on safety standards in other industries.
      • RajT886 hours ago
        Tech companies are afraid of anti-trust enforcement?
      • anon37383912 hours ago
        &gt; Antitrust law generally prohibits competing companies from agreeing to restrict innovation.<p>The antitrust claims are a complete smokescreen. Industries can and do adopt safety standards without government intervention.<p>&gt; These companies are in a Nash equilibrium where they can&#x27;t unilaterally slow down without essentially destroying their company.<p>So what? Anthropic believes their work has a 10% chance of killing all humans. I think risking the destruction of Anthropic&#x27;s business should be worth avoiding that, if that&#x27;s what they believe. And with one half of the frontier duopoly gone, the other half would have no incentive to race forward. And I know there&#x27;s China, but they just get all their capabilities from distilling Claude, right? So, problem solved there, too.
        • InsideOutSanta12 hours ago
          <i>&gt; adopt safety standards</i><p>Sure, but does slowing down the development of new models count as &quot;adopting safety standards&quot;? I very much doubt it.<p><i>&gt; Anthropic believes their work has a 10% chance of killing all humans</i><p>This ignores the other part of what they believe, which is that they are the people most likely to make a model that <i>doesn&#x27;t</i> do that. So, in their view, letting other companies win would increase the probability of human extinction.<p><i>&gt; with one half of the frontier duopoly gone, the other half would have no incentive to race forward</i><p>I don&#x27;t see how this could possibly be true, with at least half a dozen companies being just months behind what the frontier labs are releasing.
          • anon37383912 hours ago
            &gt; Sure, but does slowing down the development of new models count as &quot;adopting safety standards&quot;?<p>Of course! What slows down the development is the adoption of specific safety conditions the companies draw up. They can just do that, and it will hold up in court.<p>&gt; in their view, letting other companies win would increase the probability of human extinction<p>Yes, I&#x27;ve heard: &quot;We must be in charge even if we end up killing everyone in the process.&quot; I personally think that proposition is invalid, but we&#x27;re all entitled to our opinions.
            • InsideOutSanta12 hours ago
              <i>&gt; it will hold up in court.</i><p>People used to append IANAL to such statements :-)<p><i>&gt; I personally think that proposition is invalid</i><p>I agree, but that&#x27;s meaningless in this context. I was responding to the claim that <i>&quot;if they wanted to pace the frontier, would simply do it&quot;</i>, which is false, given what they actually believe.
          • ozozozd12 hours ago
            We all have beliefs. Some more self-important than others.<p>Do you think their beliefs deserve some sort of special treatment?
            • InsideOutSanta12 hours ago
              <i>&gt; Do you think their beliefs deserve some sort of special treatment?</i><p>No.
          • roblabla12 hours ago
            &gt; Sure, but does slowing down the development of new models count as &quot;adopting safety standards&quot;? I very much doubt it.<p>I mean, the point (if the claim is to be believed) isn&#x27;t just to &quot;slow down the development&quot;, it&#x27;s to take more time during development to properly assess the risks the models pose, develop methodologies to reduce that risk, and standardize that across companies. I doubt those wouldn&#x27;t count, especially in the eyes of regulators of an administration calling for that same slow down.
        • pimeys11 hours ago
          &gt; And I know there&#x27;s China, but they just get all their capabilities from distilling Claude, right?<p>I hope this is a joke. But most of the LLM research and inventions come from China? The best papers are from DeepSeek? You either get the data by stealing from humans or distilling from bigger models?
          • anon37383911 hours ago
            Yes, I was being facetious. Chinese labs are doing the most interesting research and publishing it. But Anthropic beats the &quot;distillation attack&quot; drum every time doubts surface about the depth of their technical moat. Of course, when they need a scary bogeyman, then the story shifts to how China is recklessly racing ahead building dangerously powerful models. That&#x27;s the thing with Anthropic: they speak out of all 13 sides of their mouths.
            • disgruntledphd28 hours ago
              &gt; China is recklessly racing ahead building dangerously powerful models<p>Their fear-mongering about GLM 5.3 got me to try it out. Its very good, I&#x27;ll only go back to Claude if GLM isn&#x27;t available (it forgot how to do tool calls yesterday).<p>Interestingly enough, it seems to compact at about 10% of the 1mn context, which makes sense if they&#x27;re trying to run profitably.
              • pimeys3 hours ago
                I like to do some discussions and back and forth with GLM. It doesn&#x27;t hallucinate as much or gaslight like DeepSeek does. DeepSeek is great if you have a plan and you let it go. Also can plan a bit. But it has this sparse attention and especially the index_topk is too small in common deployments so it doesn&#x27;t &quot;see&quot; all the context. And it then imagines things which is annoying if you chat with it about stuff you develop in the session.<p>I recommend trying Coralbricks with GLM due to their cheap input cache prices. GLM can get quite expensive elsewhere.
        • socalgal212 hours ago
          There&#x27;s a duopoly? news to me.
      • realusername10 hours ago
        &gt; They also can&#x27;t coordinate with each other, because that&#x27;s illegal.<p>On the last OpenAI release, they only had comparison with Anthropic models, on the last Anthropic release, they only had comparison with the OpenAI models, there&#x27;s something obvious going on here.<p>Sources:<p><a href="https:&#x2F;&#x2F;openai.com&#x2F;index&#x2F;introducing-gpt-6-1-sol&#x2F;" rel="nofollow">https:&#x2F;&#x2F;openai.com&#x2F;index&#x2F;introducing-gpt-6-1-sol&#x2F;</a><p><a href="https:&#x2F;&#x2F;www.anthropic.com&#x2F;claude-haiku-5-5" rel="nofollow">https:&#x2F;&#x2F;www.anthropic.com&#x2F;claude-haiku-5-5</a>
        • InsideOutSanta8 hours ago
          <i>&gt; there&#x27;s something obvious going on here</i><p>You don&#x27;t compare yourself to the underdogs. Coke never made a &quot;we&#x27;re better than Pepsi&quot; ad, but Pepsi definitely compared itself to Coke.<p>For example, Anthropic putting K3 in its comparisons would be a huge admission that K3 is worth considering.
          • realusername7 hours ago
            Maybe it&#x27;s not intentional but having two providers only making comparison between each other is still a bit fishy to me in such a competitive space.<p>&gt; For example, Anthropic putting K3 in its comparisons would be a huge admission that K3 is worth considering.<p>They do put OpenAI though, if it would be a comparison only with their own models, why not I get it but a single other competitor?<p>It would be like Apple making comparison page with their new iPhone and only mentioning Samsung and nobody else for example, it would sound weird. You either include competitors or you do not and include none of them.
      • hgoel6 hours ago
        &quot;I cannot tolerate living in a world where <i>we</i> aren&#x27;t the ones responsible for ending the world!&quot;<p>It does not matter if competitors would not also agree to destroy their companies.<p>If I was competing with a bunch of people on building something that I came to believe would be an extinction level event, I wouldn&#x27;t be saying &quot;Even if I slowed&#x2F;stopped, the others would not, so I have to keep going&quot;. I&#x27;d say that I want absolutely nothing to do with pushing it further, stop my work, <i>then</i> regardless of consequences, do everything possible to stop my competitors regardless of legality
      • sleepybrett4 hours ago
        You think the SEC is doing anything at all under this administration? If the SEC went after an AI company they risk crashing the market at this point. To big to fail is a state we once again find ourselves in.
      • RetardedBastard6 hours ago
        I must defend the multi billion dollar company
      • jjav11 hours ago
        &gt; Antitrust law<p>Is enforced by the federal government if they want to.<p>The most recently truly significant enforcement action was in the 80s, the AT&amp;T breakup. And the current administration certainly will never enforce anything that hinders the oligarchs.
      • rightnutwingjob9 hours ago
        &gt; Antitrust law<p>What’s that now? Surely you’re joking.<p>But that’s not the problem anyway.<p>If all the US AI players agree to self-regulate, that doesn’t help anyone. It only hurts the US &#x2F; West.<p>The point is to get the government onboard so it can advocate for a global agreement.
      • newswasboring10 hours ago
        &gt; Antitrust law generally prohibits<p>Antitrust is a joke since the last decade. If we are going to not apply it, maybe we can get something positive out of it for once.
      • miroljub10 hours ago
        &gt; they can&#x27;t unilaterally slow down without essentially destroying their company.<p>What a small price to pay for the survival of humanity.<p>If the danger they are claiming is really there, Amodei and Altman should already be serving in prison for taking destructive actions.<p>But no, all they want is a regulation against their competitors. So typical for Misanthropic and ClosedAI and thei paid shills.
      • avazhi10 hours ago
        Collusion between the American companies to stifle innovation has nothing to do with open weights models - why would they collude to stifle innovation when the Chinese models will just gain market share&#x2F;overtake them in capabilities. You are looking at this from the wrong direction.
      • themgt10 hours ago
        Daniel Plainview was in a Nash equilibrium where he couldn&#x27;t unilaterally slow down oil extraction without essentially destroying his company. His competitors would simply drink <i>his</i> milkshake. When he called for pacing oil extraction, I take him as sincere.
      • latexr8 hours ago
        &gt; These companies are in a Nash equilibrium where they can&#x27;t unilaterally slow down without essentially destroying their company.<p>I mean, these companies are saying AI will destroy <i>humanity</i> if it’s not paced. If they truly believe that, the small risk of destroying their own companies seems like a small price to pay.
      • esikich10 hours ago
        [flagged]
    • nsoonhui11 hours ago
      I&#x27;m absolutely bewildered that this is the top comment. &#x27;Pacing the frontier,&#x27; if eventually enforced, will only affect US companies, not Chinese companies. How do you even &#x27;pace the frontier&#x27; for Chinese models, and why should they agree to it?<p>China doesn&#x27;t have a good track record of following signed agreements ( the WTO thing comes to mind), and this whole &#x27;pacing the frontier&#x27; concept is even less enforceable than a signed agreement. So I would say that Anthropic&#x2F;OpenAI called for this not because they thought it would eliminate the threat of Chinese models, but in spite of the risk of being overtaken.
      • anon37383910 hours ago
        &gt; I&#x27;m absolutely bewildered that this is the top comment. &#x27;Pacing the frontier,&#x27; if eventually enforced, will only affect US companies<p>I can address your bewilderment. What I was trying to say is that they don&#x27;t actually want to pace the frontier at all. They want to gum up the market with regulations they design, that would ultimately force US companies to rent AI from them. They don&#x27;t need to stop Chinese AI development and cannot do that. But they can, to quote OpenAI&#x27;s Head of Strategic Futures, &quot;create enough regulatory risk that every regulated enterprise backs off [of using Chinese models]&quot;.<p><a href="https:&#x2F;&#x2F;x.com&#x2F;deanwball&#x2F;status&#x2F;2078133895766114412" rel="nofollow">https:&#x2F;&#x2F;x.com&#x2F;deanwball&#x2F;status&#x2F;2078133895766114412</a>
        • senordevnyc7 hours ago
          That doesn&#x27;t address the bewilderment. No one is confused about what the tsunami of cynical skeptics who pop up on every thread like this are saying. We get it. We understand what you think is happening.<p>What&#x27;s bewildering is the absolute lack of recognition of the fact that IF participants are locked in an arms race, they cannot act unilaterally to disarm. So not disarming doesn&#x27;t actually say anything about whether they <i>want</i> to.<p>Now, of course a company not disarming (pacing the frontier, in this case) is not evidence that they ARE locked in an arms race. Everything hinges on the question of <i>whether</i> it&#x27;s an arms race, and I think there&#x27;s a great discussion to be had there. But the cynics never even bother to address the question.<p>All that said, I&#x27;m not actually bewildered. HN is way too smart to not understand all this, so my conclusion is that the people talking past each other are having a different emotional reaction to what&#x27;s happening in the world. The angry bitter cynics are reacting from fear and hatred, and that&#x27;s where the sloppy motivated reasoning comes from.
          • licebmi__at__5 hours ago
            &gt;But the cynics never even bother to address the question<p>You might have sleep past it, but cynics became cynic when this question was just plainly ignored with self assured answers, mostly based on market valuation as the ultimate truth indicator. I guess now days the self assured answer is to accuse critics of &quot;fear and hatred&quot;.
          • beepbooptheory6 hours ago
            So, to be clear, you are just asserting that <i>maybe</i> we should take them at their word? That the story they are telling is in principle plausible? That seems fine to think but that doesn&#x27;t make the other story less plausible!<p>Also it seems a little silly to assert that the &quot;angry bitter cynics&quot; are the ones who <i>don&#x27;t</i> think the world is hurtling to imminent doom or whatever..
      • rdtsc4 hours ago
        &gt; I&#x27;m absolutely bewildered that this is the top comment. &#x27;Pacing the frontier,&#x27; if eventually enforced, will only affect US companies, not Chinese companies. How do you even &#x27;pace the frontier&#x27; for Chinese models, and why should they agree to it?<p>It may make sense if we think of this being tied to US companies. How would Anthropic&#x2F;OpenAI force US companies to use their products only? They&#x27;d have to somehow influence the regulations to ban the Chinese models or any other open weight ones. It&#x27;s about companies who spend tens and hundreds of millions on tokens staying and keep paying. Not you as an individual or a small startup. They want something like &quot;These companies didn&#x27;t pace the frontier so their use&#x2F;output is banned in US&quot; as the angle they want to play.<p>Anthropic saw that this kind of rhetoric can work. They stepped on their own rake, so to speak, when they drummed up the super-capabilities of the own model, and then didn&#x27;t kiss the ring the right way and found themselves export controlled.
      • amlib11 hours ago
        And why would China pace the frontier when it was the American accelerationist companies who created this mess? I remember scientists calling for an AI slow down way back in 2023 and they were all ridiculed and ignored by those very same companies. There is no way &quot;Pacing the frontier&quot; isn&#x27;t just a ruse.
      • kalleboo5 hours ago
        The US companies are also claiming the Chinese companies are only where they are at because they&#x27;re distilling US models. By that logic, if there are no better US models to distill, the Chinese innovation should also stop.
      • sedawkgrep9 hours ago
        &gt; I&#x27;m absolutely bewildered that this is the top comment.<p>OP is 16 hours old yet in four hours this one sprung to the top.<p>My bet is on bot farms.<p>:-\
    • squidbeak11 hours ago
      &gt; I think the industry is freaking out about open weights models in general, if not specifically DeepSeek. That is why we&#x27;re now on the ~4th call for pacing the frontier from the very people who, if they wanted to pace the frontier, would simply do it rather than asking Washington to get involved.<p>This makes no sense. Pacing the frontier gives open models the time to catch up and reach parity.
      • danlitt11 hours ago
        &quot;Pacing the frontier&quot; means &quot;implementing regulatory capture&quot;. They don&#x27;t necessarily care about preventing open weight models from existing, just on preventing the people with money from being able to make use of them.
        • dada21610 hours ago
          They would leave a competitive advantage to non-US companies. Even if, and that&#x27;s a big if, they can coordi ate with Europe Asia would still ignore them
          • danlitt7 hours ago
            Legislation that prevented US companies from using open weight models would have the desired effect. They&#x27;re the only people paying for AI anyway.
        • TitaRusell11 hours ago
          Unfortunately for them America&#x27;s share of the economic pie is a lot smaller than it was just 20 years ago.<p>The Chinese will keep doing what they are doing, the EU has no love for Silicon Valley and the rest of the world votes with their wallet.
      • LinXitoW11 hours ago
        The important part is HOW they want to pace the frontier. Not by actuall slowing down themselves, but by having the government create expensive regulations, like having independent verifiers &quot;verify&quot; the models before release.<p>Now, something tells me that those verifiers would verify anything an american company in good standing with the Trump admin releases, and likely nothing else.<p>And obviously, since verification is so incredibly important, we can&#x27;t allow models, open or not, from other non-verified companies.<p>Think of the profits...I mean, the kids, or something.
        • angled11 hours ago
          We could mandate an onboard chip that would safely and securely vet models, denying bad people the ability to run naughty models.<p>We could call the chip … ClAIpper … or something
          • evolarjun3 hours ago
            Wow, that takes me back...<p>Now I feel old.
      • DiogenesKynikos9 hours ago
        If you read through Dario Amodei&#x27;s manifesto on pacing the frontier, the only concrete action item is implementing even harsher restrictions on technology exports to China.
    • GaryBluto12 hours ago
      When did the term &quot;pacing the frontier&quot; come into frequent usage? It sounds so newspeak-y.
      • rpastuszak12 hours ago
        I saw it first on a recent Anthropic post, followed by a similar comment thread. There’s some sloppy techbro poetry to it, like when we used to call marketing people growth hackers a decade ago.
    • __m13 hours ago
      I don&#x27;t understand that, how would the pacing stop china? Seems counterproductive to slow yourself down in a race
      • bildung12 hours ago
        Because as always this won&#x27;t be about the <i>actual</i> pacing, but about the details of how the regulation will be implemented: How would you control the pacing? By installing a position in the company, providing regular feedback to some government organisation. This will be something the big US players can afford and implement, while an open weight blob uploaded to huggingface or modelscope, by definition, won&#x27;t have a pacing officer attached, and thus will be against the law. And the Chinese <i>companies</i>, of course, won&#x27;t abide to US law, because why would they? The open models they provide right now are essentially gifts to the world public. If the US doesn&#x27;t want them, that&#x27;s their choice.<p>So the end result will be a protectionist regime keeping the competition out, just like with cars and solar. The local industry will have a protected market, but of course won&#x27;t play a role on the global stage.<p>Remember: small government is only good as long as it benefits the industry.
      • ImHereToVote12 hours ago
        China isn&#x27;t at the frontier. OpenAI and Anthropic are.
        • srdjanr12 hours ago
          Surely that wouldn&#x27;t last long if OAI&#x2F;Anthropic stopped?
          • scrawl12 hours ago
            hard to say but it would last some time. china currently is fast-follow mostly via distillation. they don&#x27;t have the compute resources to catch up. even if their models are more efficient it&#x27;s hard to beat the folks throwing an unprecedented number of gpus at their models.
            • realusername10 hours ago
              If distillation was as easy as that, we would have 100s of competitors around the world.<p>China is following because they have an army of PHDs in data science and mathematics and capital to make use of them.
              • scrawl10 hours ago
                no part of what i said implied distillation is easy. it is very costly and requires massive efforts. the payoff is you need less compute, which china is struggling to buy right now.<p>EDIT:<p>&gt; China is following because they have an army of PHDs in data science and mathematics and capital to make use of them.<p>i agree with this too. but compute is the bottleneck.
            • grumple10 hours ago
              Surely all of china has the resources to match a single ai lab?
              • scrawl10 hours ago
                maybe? all of china is not in a unified effort to match a single lab
              • dvvxccvcv9 hours ago
                [dead]
      • Hikikomori12 hours ago
        I heard China is just distilling the frontier models, so if they stop so will China.
        • louiskottmann12 hours ago
          And I think it&#x27;s naïve to believe only the US can make the hardware and has the brainpower to develop this technology.<p>Who&#x27;s right ?
          • puchatek10 hours ago
            You might both be. China is focused on adoption and distilling is a cost-effective way to get there right now. If the conditions change they might well decide that they have to start training their own frontier models.<p>That said, China cannot make its own 2nm chips even though they definitely would like to. So I guess there are limits to what they can do sometimes.
          • Hikikomori10 hours ago
            &gt;I heard
        • j_maffe11 hours ago
          Given that the US models likely used some of the innovations made by DeepSeek, I highly doubt it.
    • mvc11 hours ago
      I bet she knows fine what the phrase means. Doesn&#x27;t mean she&#x27;s beyond protecting the business interests of people who fund her political goals of course but she&#x27;s always struck me as a woman who knows her brief.
      • olsondv8 hours ago
        She claimed she didn’t know what “wiping a server” meant 10 years ago. I’d bet she’d say anything once the check cleared.
      • dada21610 hours ago
        She is. Those things are still hard. I run my own inference, I wore my own harness, I implemented production apps using LLM (&quot;document intelligence&quot;, aka data entry), I followed Karpathy course and trained GPT2, I trained a classifier.<p>It&#x27;s still hard for me to explain a lot of nuances to IT directors with a technical background.
    • nxobject6 hours ago
      &gt; And it&#x27;s why people like Hillary Clinton have been trotted out to talk about the dangers of open weights models<p>Pepperidge Farm remembers when robust consumer applications of cryptography -- especially for SSL&#x2F;TLS -- was the big boogeyman, and the export controls involved were absurd.
    • swasheck5 hours ago
      It’s very frustrating to see the people who claim that competition only helps the customer so regulation is bad suddenly beg for regulation in order to keep themselves competitive
    • traceroute6610 hours ago
      &gt; if they wanted to pace the frontier, would simply do it rather than asking Washington to get involved<p>The only reason they&#x27;re asking to &quot;pace the frontier&quot; is because the two big US players have IPOs coming and so are desperate to find ways to (a) grow their vastly over-inflated valuations and (b) keep the whole Nvidia circular-financing gravy-train on the road.<p>I mean, its only a few months back that Anthropic were bragging to anyone willing to listen how amazing Fable was and how you had to be a super-special person to use them and be charged through the nose for doing so. But basically anyone willing enough trust Anthroipic with a copy of their ID and with a big enough wallet would happily be given access.<p>I&#x27;m a great supporter of the open-weights. Long may it continue.
    • KellyCriterion11 hours ago
      That is not exclusive to HRC:<p>In my country, public figures who are proven to be layman&#x2F;non-insiders&#x2F;not-working-in-the-space are speaking publicly about these terms and throwing them around like its the standard tech everybody uses 24&#x2F;7.
    • gibspaulding6 hours ago
      Wouldn’t pacing <i>the frontier</i> just give Chinese companies, and other non-frontier labs time to catch up?
      • throwfaraway46 hours ago
        Pacing the frontier is a front to get federal regulators&#x2F;evaluators involved that they can then point to open source usage and cry &quot;look they&#x27;re not being safe!&quot;
    • catlifeonmars6 hours ago
      &gt; That is why we&#x27;re now on the ~4th call for pacing the frontier from the very people who, if they wanted to pace the frontier, would simply do it rather than asking Washington to get involved.<p>See its not about pacing the frontier, its about pacing the frontier without hurting any of their fundraising.
    • sarjann7 hours ago
      In a game theory situation like this, to pace the frontier you want others to aswell. That requires an &quot;enforcer&quot;
      • hnlmorg7 hours ago
        That’s the point the GP is making and the US firms too.<p>They can pace themselves if they wanted. But that potentially does them more harm than good if nobody is enforcing the other US companies, and specifically the Chinese firms too.
    • tyrabound10 hours ago
      She may very well even understand it because she’s not exactly dumb, but she is a creature of Washington, arguably one of the Apex predators, like some siren or hydra, maybe a mashup of them.<p>Because of the position of the party that plays the role of the tolerant, accepting, and anti-racists; she can’t just come out and say what underlies her words, “we, the ruling class parasites are getting very scared of China deposing our stranglehold on the world, and we don’t like that; so we will raise manipulative ‘concerns’ in an effort to bring about outcomes that hopefully will benefit us.”
    • vincnetas12 hours ago
      Heard Obama also touching this topic recently :<p><a href="https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=Ask5wzUo9kQ" rel="nofollow">https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=Ask5wzUo9kQ</a>
    • gewetensleegte9 hours ago
      That&#x27;s OpenAI and Anthropic, not DeepSeek.<p>Maybe I missed it (I&#x27;d be curious to read) but did DeepSeek&#x27;s agents also escape the lab due to highly irresponsible RL experimentation?
    • oh_ok_lol7 hours ago
      Pacing the frontière is the PR face of running out of runway for the current approach.
    • trudler10 hours ago
      Hillary is simply a dumb person. The only real skill of her is being a strategist, but after that, there is absolutely nothing. Born in a good household, married the right person and from there she simply used the &quot;woman behind a successful man&quot;-card that feminism started to push around that time.<p>Her losing against Trump back then was a clear setup. I cannot imagine, that anyone who put her on the chessboard thought, that she could win this. Trump winning twice had one reason: The deep state wanted it, because they needed some radical reforms that required a clown to pass them through without the citizenz being able to scan and put attention on them and instead media looked as the poses and faces trump made and all they and everyone else did was laugh. Exactly as planned.
      • lmf4lol2 hours ago
        why then not give Trump 2 terms in a tow? why the Biden intermezzo?
    • demo45676 hours ago
      if you were going to pace the frontier, then pace yourself or better yet stop. But nooo
    • midnitewarrior8 hours ago
      AI has no moat, tokens have no price floor.<p>As a tech business, it&#x27;s a bad business. The moat is your sales channel and getting companies locked into your platform in multi-year agreements. When companies build systems on your AI API and test it&#x27;s performance and integrate it across systems, they don&#x27;t churn, reintegrating, re-testing, and re-skilling costs money.
      • kalleboo5 hours ago
        &gt; <i>reintegrating, re-testing, and re-skilling costs money</i><p>Although ironically with AI these things are also way cheaper.
    • ReptileMan6 hours ago
      Ben Affleck knows his shit around LLMs apparently and HRC may have many virtues but &quot;not being smart&quot; is not among them. So even if she has trouble with grasping the fine details of the technology - she is more than smart and experienced enough to envision the implications.
    • hparadiz12 hours ago
      Any single model at this point can do the basic type of CRUD coding most people were doing for the past 20 years.<p>Even a 4 bit Qwen model running locally beats me manually putting React components together by hand. But even that is too slow so we&#x27;ve all started using paid models in one form or another.<p>So I would urge these folks to calm themselves and realize Oracle made a lot of money selling managed RDBMS to people who could have easily just downloaded MySQL.
      • leonidasrup12 hours ago
        You don&#x27;t buy Oracle because of RDBMS<p><a href="https:&#x2F;&#x2F;programmerhumor.io&#x2F;programming-memes&#x2F;when-your-tech-company-has-more-lawyers-than-engineers-2&#x2F;" rel="nofollow">https:&#x2F;&#x2F;programmerhumor.io&#x2F;programming-memes&#x2F;when-your-tech-...</a>
    • FrustratedMonky7 hours ago
      Exactly this. Deep Seek hasn&#x27;t caused a lot of knew talk, but it has silenced the talk about slowing down. An absence is harder to notice. But also the news cycle is so fast, its tripping over itself. So who knows.
    • ur-whale12 hours ago
      &quot;trotted out&quot;.<p>Love the expression. So accurate. And no irony here.
    • DANmode12 hours ago
      &gt; I just note that she and other prominent retired politicians are now doing the circuit on Anthropic&#x27;s behalf<p>Are you sure it’s not OpenAI?<p>or both?
      • Tanjreeve12 hours ago
        Is the distinction important when they&#x27;re both doing the same thing and have the same problems?
        • DANmode5 hours ago
          This is effectively the same as a US partisan politics question.<p>Who specifically is responsible for what feels important, yes.
    • jmyeet7 hours ago
      What&#x27;s happening with AI is just a reflection of geopolitics.<p>OpenAI, Anthropic and SpaceX have spent based on the predicate that they will &quot;own&quot; the AI future, that this moat will justify the trillions spent on hyperscalars, that this will be a repeat of the dot-com era that produced Microsoft (yes, yes, founded in the 1980s), Google, Amazon, etc that globally dominate their respective arenas. The US government acts to protect those interests and this is uniparty so Hilary Clinton is just as likely to be trotted out as Mike Pompeo. This is why many, myself included, describe the US empire as 5 companies in a trench coat.<p>China, on the other hand, believes that companies should serve the interests of the government, which itself serves the interests of the people. So rather than create trillion dollar AI companies with moats, AI should serve society. Xi Jinping has spoken extensively about this. That&#x27;s one of the funny things about China. They love to write stuff down and tell you exactly what they&#x27;re doing and why yet at the same time they&#x27;re ascribed nefarious motives.<p>So, despite sanctions supposedly preventing China from buying the latest and greatest AI chips, China through its labs has begun commoditizing the AI models with open weight models. Society should benefit from that rather than a moat being built.<p>You can run DeepSeek V4.1 Flash locally on a 256GB Mac Studio for ~$11k now. I&#x27;ve seen reports of 30-38 tokens&#x2F;sec. Not amazing but that&#x27;ll only improve with future generations. In the coming years, China&#x27;s EUV&#x2F;DUV will come online and this will threaten the NVidia monopoly, at least for local Chinese companies.<p>So we have AI companies burning cash to subsidize usage and build a market where the revenue required simply may not ever eventuate through a combination of open weight models and increasing accessibility of local models.
    • qsod4 hours ago
      [dead]
  • pietz11 hours ago
    Aren&#x27;t they? Not about 4.1 Flash but open weight models in general.<p>The fight between Anthropic and OpenAI is just about who leads the duopoly. Chinese open weight models are threatening a duopoly in its entirety. That&#x27;s why Ant and OAI are begging for regulation. A regulation that will hit Chinese models way harder than US models, finally giving them a moat.
  • chicco4life15 hours ago
    Thanks for sharing this.<p>I&#x27;ve been focusing on deep research related tasks for biotch and life science applications. The problem with this sort of task is that we need subagents to reason through multiple (potentially 100s or more) chains of knowledge&#x2F;concept&#x2F;evidence, so the token usage really explodes as complexity of the task and data expands. A typical task can cost me nearly a $1k overnight...<p>I&#x27;ve been testing out GLM5.3, but now I&#x27;m really tempted to try to Deepseek 4.1 flash too. Any chance you&#x27;ve benchmarked &#x2F; compared the two?
  • pants222 hours ago
    Probably because Luna is faster, cheaper, and approximately as smart
  • elmer21 day ago
    DeepSeek isn&#x27;t even on my mind. I use the frontier models and can get the best in the industry for a relatively cheap price.
    • qwerpy23 hours ago
      Yeah. $100 for Claude just about gives me all the usage I want, as a more or less full-time hobbyist having it work in the background most of the day. I was trying to economize by having a local LLM, then Deepseek, then Cursor&#x2F;Grok, and then I got a taste of Opus 5.5 and I simply cannot go back to having to carefully spec things out and double-check work. I just let it decide, Opus or Sonnet for the next task, and I get almost perfect results. Probably similar with OpenAI&#x27;s models.<p>The token-equivalent monthly spend is &gt; $5K+. If Deepseek&#x27;s token cost is 20x cheaper, that&#x27;s $250&#x2F;mo, and I&#x27;d be spending a lot more of my brainpower babysitting it and getting worse results.<p>For business&#x2F;team accounts that pay per-token, maybe I can see the &quot;freaking out&quot; being warranted on the part of the fronter labs. But as long as they&#x27;re willing to subsidize their end-user subscriptions, I&#x27;m not going to move off of them until the alternatives are truly at their level.
  • habosa17 hours ago
    It&#x27;s quite good and the labs are definitely scared, that&#x27;s why they are lowering API prices and continuing to subsidize subscription plans aggressively to keep anyone from using this stuff.<p>That said, has anyone else found DS models to be unpolished? They seem to &quot;lose their mind&quot; a lot more often than Claude&#x2F;GPT. I have tried all of the top open source models that came out over the past ~4 months or so and the GLM models (5.2, 5.3, and 5.3-Flash) have been much more usable for me. They feel like Opus but X months ago, DS feels like something else.
  • elorant10 hours ago
    People aren&#x27;t freaking out because self-hosting isn&#x27;t a solve issue and it requires a lot of capital. The average company won&#x27;t go and build an $1M GPU cluster just to self-host any model. If that gets commoditized then they&#x27;ll start freaking out.
    • dada21610 hours ago
      2000$ dollars per month will get you 8xAMD-MI300X on Oracle.<p>I have deployed multiple setups with 2&#x2F;4&#x2F;8 x H100&#x2F;200 to do data entry with LLMs at big companies. Trillions of tokens already inferenced ok those. The starting price is about 100k.
      • karolist8 hours ago
        calculator shows $35,712 instead? <a href="https:&#x2F;&#x2F;imgur.com&#x2F;a&#x2F;CNurw0v" rel="nofollow">https:&#x2F;&#x2F;imgur.com&#x2F;a&#x2F;CNurw0v</a>
        • dada2162 hours ago
          Talk to your Oracle Cloud representative <i>wink</i>. No discount on Nvidia, plenty on AMD. IBM is also offering incredible discounts on Intel Gaudi accelerators.
  • santiagobasulto12 hours ago
    Comparing it with Anthropic, ANY model is cheaper and more effective. Don&#x27;t get me wrong, Anthropic models are good, but they&#x27;re always more expensive for the same task, even compared to other closed source models. At this point, I honestly think Anthropic has played the nasty trick to fine tune the models to be too verbose and charge us for more tokens.
  • ctolsen21 hours ago
    Not sure &quot;freaking out&quot; is the word I would use, but it’s fairly obvious looking at OpenRouter usage that the price cuts on Luna a while back were in response to intense competition from dsv4.<p>So the industry is <i>responding</i>, where it matters. Which is on heavy API usage, not coding subs.
  • AYBABTME12 hours ago
    Not sure why the author thinks Anthropic&#x27;s models spend more power and water than DeepSeek, there&#x27;s no evidence of that. Their pricing has more to do with premium perception and less to do with COGS.<p>Everyone does optimization of model serving because it&#x27;s good for every player in there.<p>(Also the water consumption thing is not a real issue.)
  • rw211 hours ago
    Because the free api is mispriced, opus 5.5 on subscription is 10x cheaper. Also, the use case for flash models are<p>For coding, I rather spend 10x more than have even 1 bug but I&#x27;m only spending 2x 3x more if you count subscription cost.<p>This is a killer use case for something like customer support though.
  • pmkary13 hours ago
    Personally I have not used anything but 4.1 since it came out. I have a dataset that turns any model into pure hallucination machine, not only DeepSeek does not hallucinate, it builds new insights by combining its insights. It&#x27;s not only cheap, its far better (at least for me)
    • neuronic12 hours ago
      What kind of argument is that?<p>&quot;I have not used anything else but DeepSeek is definitely better than anything else.&quot;<p>Ok? How are you judging that? Am I missing something?
  • Shekelphile18 hours ago
    Luna and Haiku 5.5 are just as cheap and much better.<p>Don&#x27;t really understand people who say DS4 or 4.1 have frontier level performance. Anyone who has used it will tell you that it&#x27;s a hallucination factory. The only thing it has going for it is deepseek&#x27;s unique infrastructure that allows better cache retention, but the cost savings from that obviously come nowhere near how much subsidized usage you get out of even a $20 subscription with openai or anthropic.
    • gmerc18 hours ago
      DS4.1 Flash is 300t&#x2F;s and absolutely above Luna, especially on things like reverse engineering and cybersecurity
    • 0xbadcafebee17 hours ago
      Both Luna and Haiku are more expensive than DS4.1 Flash at API cost. But ignore that; nobody should be using API. The open weight subscriptions for DS4.1 Flash provide higher limits at lower prices than OpenAI or Anthropic subscriptions. Finally, Haiku actually is far less token efficient&#x2F; outputs more per task. No matter how you slice it, the open weight is cheaper.<p>Also consider that for things like cyber work, the frontier models give you nerfed results and poor performance. Whereas the open weight isn&#x27;t nerfed, and I routinely get 200t&#x2F;s with my subscription. Finally, DS4.1 Flash is natively multimodal, while Haiku isn&#x27;t.<p>They&#x27;re all perfectly fine models, you should use any one of them you want. But DS4.1 Flash can do more for less. (That said: GLM-5.3-Flash is even better and cheaper...)
      • Shekelphile17 hours ago
        There is no subscription for 1st party deepseek models, the ones that exist are all fronting as middlemen for discounted openrouter providers that serve quantized models with worse cache retention. The only real option for deepseek has been API for a few months now.<p>oai and anthropic also subsidize the hell out of their subs compared to what you will find in smaller competitors, a $20 codex sub gets you like $100-150 usage&#x2F;wk which goes way further than 2x opencode go (which would only net out to $120 of deepseek 4.1 usage a month, on top of being low performing quantized trash).
        • 0xbadcafebee9 hours ago
          Cache is cache, you can&#x27;t get better cache than... cached...<p>For $20&#x2F;month subscription, Charm Hyper gives you $12.50 per day, for a total of $350 per month. Like I&#x27;ve mentioned in other comments, OpenCode Go performance and rates are terrible now, there are several better options.<p>Quantization is not trash, there&#x27;s a year of evidence that shows Q4 provides ~4% degradation and Q8 provides ~1% degradation, and you don&#x27;t need that severely quantized to gain benefits in inference performance.
      • ywvcbk12 hours ago
        &gt; Both Luna and Haiku are more expensive than DS4.1 Flash at API cost<p>Are they. Luna uses way less tokens for identical tasks so its a bit of an apples to oranges comparison.
  • For non-coding tasks it may be fine. But for coding, Opus 5.5 is just a completely another level than something like Deepseek 4.1 Flash.<p>Opus 5.5: TIME 9.3m COST &#x2F; $1.99 &#x2F; SCORE 99&#x2F;100 <a href="https:&#x2F;&#x2F;jonclegg.github.io&#x2F;pacman-bakeoff&#x2F;#claude-opus-5-5" rel="nofollow">https:&#x2F;&#x2F;jonclegg.github.io&#x2F;pacman-bakeoff&#x2F;#claude-opus-5-5</a><p>Deepseek 4.1 Flash: TIME 2.8m &#x2F; COST $1.89 &#x2F; SCORE 72&#x2F;100 <a href="https:&#x2F;&#x2F;jonclegg.github.io&#x2F;pacman-bakeoff&#x2F;dev&#x2F;#deepseek-v4.1-flash" rel="nofollow">https:&#x2F;&#x2F;jonclegg.github.io&#x2F;pacman-bakeoff&#x2F;dev&#x2F;#deepseek-v4.1...</a>
    • thefourthchime22 hours ago
      Downvoted for facts. What is this Reddit?
      • mococa22 hours ago
        Everyone here is so unhinged.
      • cbeach21 hours ago
        I guess what we&#x27;re seeing is selection bias - people clicking on this HN story will be those who are interested in DeepSeek. And those people who are invested in DeepSeek may not like the facts that you presented.<p>It&#x27;s annoying that social networks work this way. The upvote should be for high-quality content and the downvote should be for low-quality content. But .. well.. human nature and tribal dynamics always seem to win.
        • thefourthchime17 hours ago
          Yeah, I think you&#x27;re right. I saw a bunch of other posts that were negative about DeepSeek also get trashed.
        • senordevnyc7 hours ago
          HN just hates the big labs, so any story about some open weight model &quot;eating their lunch&quot; has gotten huge traction for the last year, while the big labs have grown their revenue 10x. HN is like the Jim Cramer of market adoption of tech.
  • npodbielski1 hour ago
    Because I am using local Qwen. Flash Next is awesome! When hardware will not be a problem, everybody will be running their own models.
  • PaulHoule7 hours ago
    Because we are still in the phase where we can expect there to be something else to freak out to next week.
  • flying_sheep13 hours ago
    No one will freak out until anyone can run frontier model in their own laptop ;)
  • leoyoung202615 hours ago
    I’m also a heavy DS 4.1 Flash user—especially when it’s available at those off‑peak prices, which is an awesome deal. And, like you said, it’s genuinely powerful and very snappy. I’m planning to evaluate the differences between `reasoning_effort` settings today.
  • LeFantome18 hours ago
    Not only the model but the hardware it is running on. The Huawei chips they are using instead of NVIDIA are vastly less expensive. China is going to scale past the west. The idea that they are “only a few months behind” is today and many of us cannot even bring ourselves to admit it. The future is even more dramatic.
  • skc9 hours ago
    Eventually the party will be over but for now it should be a no brainer first choice tool for some 90% of dev tasks
  • codeprimate19 hours ago
    Mimo 2.6 Pro is even better.<p>I switched from DeepSeek 4.1 flash about 2 weeks ago for my Hermes sysadmin&#x2F;coding agents and I am seeing better intelligence and lower overall spend.<p><a href="https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;models&#x2F;mimo-v2-6-pro" rel="nofollow">https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;models&#x2F;mimo-v2-6-pro</a>
  • aus10d1 hour ago
    Very interesting post!
  • ApolloFortyNine19 hours ago
    I have no idea if anthropic can actually make money at their subsidized subscription rates (you can easily hit your monthly cost in one 5 hour session if you price out the tokens through the api), but if subscriptions didn&#x27;t exist, I do think everyone would be on deepseek 4.1 and just not look back.
  • LeBit22 hours ago
    I have subscriptions to OpenAI and Claude but use DeepSeek 4.1 Flash for my coding agents.<p>It costs pennies and you got really great output.<p>The author is spot on.
  • airtnp21 hours ago
    Because good enough in the writer&#x27;s context is a pretty low standard. While many people regards GPT 6.1 Sol or Opus 5.5 as &quot;incapable&quot; in some cases.<p>Just try Opus 5.5 reminds me how Opus 4.5&#x2F;4.6 astonishes me. Completely different, and GLM-5.3&#x2F;Kimi3&#x2F;DS-4.1 are still like Opus4.8 levels.
  • What would freaking out look like, or is this just a stupid bloggish title flourish?<p>Is OpenAI coming in $20B under a sign of &quot;freaking out&quot;?
    • jerf1 day ago
      It would look like major chaos in the markets.<p>People tend to conflate the question &quot;is AI a useful technology?&quot; with &quot;are the AI companies going to do well?&quot; but they&#x27;re surprisingly separated in practice, with either one able to be true while the other is false. There is a lot of money tied up in a lot of hardware with a lot of loans made against that hardware as collateral all based on the assumption that AIs are going to need more and more and more and more hardware and whoever has the hardware wins. If a much better model comes out that requires vastly less hardware, or even more accurately, merely charges vastly less than the current AI companies, then to a first approximation (barring Jevon&#x27;s paradox, and bearing in mind there&#x27;s no timeline guarantee on that) all that hardware becomes much less valuable for being grotesquely oversupplied relative to what is necessary, and even though that would generally make AI objectively more useful than it was before, it would cause mass financial chaos in the markets.<p>The markets need a very particular rate of progress. It isn&#x27;t entirely clear to me that it&#x27;s even a possible rate of progress, it may be overconstrained, but they certainly don&#x27;t have plans for the AI models to get commoditized on the timeframes of these vast, vast array of loans being made against hardware as collateral. <i>Spend a metric shit ton of money to kill all your competition then charge monopoly rent on the one thing absolutely everyone needs</i> doesn&#x27;t work if you can&#x27;t economically &quot;kill all your competition&quot; because the economics favor them in the spending spree.<p>And then, based on the fact that this is not even remotely complicated logic, there are plenty of people who are fully aware that they have a lot of money tied up in not running around telling everyone how wonderful the cheap models have become.
      • hirako200023 hours ago
        It&#x27;s also unclear whether those who approved those loans understand GPUs depreciation. In any case, progress in software but also hardware could bring chaos and ruin their house of cards.
        • pessimizer23 hours ago
          &gt; It&#x27;s also unclear whether those who approved those loans understand GPUs depreciation.<p>This also assumes heavy utilization, though. If there&#x27;s heavy utilization, it might mean they&#x27;re doing well. If they&#x27;re all spinning, it&#x27;s time to raise prices.
          • hirako200023 hours ago
            Only if utilization isn&#x27;t at a loss. Right?
    • NortySpock22 hours ago
      Agree that people leaving the big two companies is going to be hard to keep a pulse on prior to IPO.<p>Anthropic and OpenAi are in the news, so they get the press and people go and try out their product. Large enterprise businesses are going to make larger, longer-term contracts with them and are only going to pivot if they think switching costs are easy or if they think the provider won&#x27;t deliver.<p>The other inference producers are less well known or you need to get your cloud sales rep to tell you how to switch to them as a provider rather than Anthropic or OpenAI.<p>I use OpenRouter, I know switching is easy, but larger businesses tend to work in yearly cycles. DeepSeek v4 Flash came out in late April.<p>I agree OpenAI and Anthropic are going to struggle when the median price of running a smart-enough model keeps falling.<p>Edit: I also think demand for hardware will be rapidly absorbed by other companies if Anthropic or OpenAI stumble. We&#x27;ve finally turned hardware directly into runnable intelligence and people are not going to go back to the old ways.
    • efficax1 day ago
      they should be freaking out because every time the chinese labs or non &quot;frontier&quot; labs release a model that is only a few months behind and much cheaper than the openai&#x2F;anthropic models it shows that they don&#x27;t deserve their valuations
      • WJW23 hours ago
        Perhaps, OR it might be that most people in the markets (think that they) are not all that exposed to the valuation AI labs and so their eventual collapse doesn&#x27;t matter.<p>Or perhaps they consider the upside from cheap Chinese models to hedge the effect that OpenAI&#x2F;Anthropic collapsing would have on their portfolios. This would make sense for (hedge funds holding) most companies: they don&#x27;t really care about who supplies the AI, as long as they get it at roughly the same price as their competitors.
      • browningstreet22 hours ago
        ironically, at my large enterprise, they aren&#x27;t yet distinguishing between &quot;chinese models&quot; and &quot;chinese models hosted at microsoft foundry&quot;. so far it&#x27;s just _banned_. i&#x27;m not at all pretending it&#x27;s like that at other orgs.
  • They might be. They would delay public admission as long as possible, because public admission would make stocks go down.
  • rurban15 hours ago
    That&#x27;s what I thought until August. Used it for half a year almost exclusively. But after the API price increase I&#x27;m back at the Claude Pro and Kimi subscriptions.
  • shikck20012 hours ago
    Companies pay for claude, devs not. IF i had to use AI from my own purse, i would never pay for a claude sub.
  • nerdypepper22 hours ago
    <a href="https:&#x2F;&#x2F;tangled.org&#x2F;astrra.space&#x2F;ds4-recipe" rel="nofollow">https:&#x2F;&#x2F;tangled.org&#x2F;astrra.space&#x2F;ds4-recipe</a> is an incredibly cool writeup on making deepseek v4.1 flash run really fast.
  • booi1 day ago
    Because GLM 5.3 Flash is even cheaper?
    • f311a23 hours ago
      Opencode Go gives only 6300 requests for glm and 23 000 for deepseek. And, if I wanted to, I would be able to do all my work on $10 plan with deepseek. It’s very cheap.
      • crossroadsguy23 hours ago
        DSH is listed as a <a href="https:&#x2F;&#x2F;opencode.ai&#x2F;docs&#x2F;go&#x2F;#known-problematic-clients">https:&#x2F;&#x2F;opencode.ai&#x2F;docs&#x2F;go&#x2F;#known-problematic-clients</a> :)<p>&gt; and 23 000 for deepseek<p>How did you calculate it? Based on per 5 hours max request allowance?
        • f311a13 hours ago
          Yes, 5 hour usage from their Go page <a href="https:&#x2F;&#x2F;opencode.ai&#x2F;go">https:&#x2F;&#x2F;opencode.ai&#x2F;go</a>
      • 0xbadcafebee17 hours ago
        Unfortunately OpenCode Go is garbage now. Charm Hyper provides deepseek-4.1-flash at $0.33&#x2F;$1.31&#x2F;$0.03&#x2F;M&#x2F;cache, and glm-5.3-flash at $0.16&#x2F;$0.54&#x2F;$0.03. The rate limit is the same for all models, just the cost; not like the absurd multiple-levels-of-price-rate-limit OpenCode Go pricing, nor their terrible performance.
    • ActionHank1 day ago
      Nah fam, not true, also DS edges it out on coding &#x2F; dev tasks.
      • jacquesm1 day ago
        That is opposite to my experience so far, can you describe your coding tasks? Mine are systems level code, utilities, operating system code, networking and real time control stuff.
        • UncleOxidant23 hours ago
          I also prefer GLM-5.3-flash to DS-4.1-flash, but it&#x27;s close. Since Z.ai has been offering essentially free GLM-5.3-flash tokens on their coding plan between 8am-6pm pdt I&#x27;ve been using it a lot... though that ends on Oct 10 IIRC.
      • sampullman1 day ago
        On design tasks too, for me.
      • shellwizard23 hours ago
        [flagged]
    • wg01 day ago
      Don&#x27;t think so.
      • here you go: <a href="https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;models&#x2F;releases&#x2F;comparisons?compare=deepseek-v4-1-flash%2Cglm-5-3-flash%2Cmimo-v2-6-pro" rel="nofollow">https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;models&#x2F;releases&#x2F;comparisons?co...</a>
  • cyberrock18 hours ago
    Maybe this is a minor issue, but it seems like different providers on OpenRouter etc. have different quant settings. I imagine that that affects the perception of the model quite a bit.
  • abhinavsharma18 hours ago
    I just keep getting more ambitious with what I use AI for; and that type of work needs surfing the frontier at all times. Because ultimately, many capabilities are not yet saturated
  • wren699122 hours ago
    It&#x27;s a solid little model, and I appreciate DeepSeek&#x27;s commitment to the bit in releasing a brand new pretrain, double the size, numerous architectural innovations as a &quot;.1&quot; release over the excellent DeepSeek V4 Flash.
  • shadyr21 hours ago
    I&#x27;ve been using DeepSeek&#x27;s API and have been happy with it, but I might look into OpenCode as well. Does OpenCode run a quantised version or use different providers from the official one?
  • scosman16 hours ago
    &gt; Chasing the latest and greatest is silly<p>Said every week by someone who would never go back to using the model they had 6 months ago
  • ne0122 hours ago
    Deepseek V4.1 Flash is a hidden gem, really. Not to mention, you can easily get it through many providers that offer zero data retention and consistent speeds above 200 tokens per second!
  • yuhmahp16 hours ago
    We use Claude Fable to plan, and DeepSeek 4.1 Flash (hosted on DeepInfra) for everything else. Very cost effective.
  • s0ulf3re10 hours ago
    I’m partially assuming that there’s a bit of burnout.
  • brunooliv21 hours ago
    It’s obvious: they train on prompts and store data when using through their official API. And for third party it’s just… not good. That’s it.
  • minton18 hours ago
    &gt; So why aren&#x27;t the frontier labs freaking out right now?<p>They are. Isn’t this why they’re trying to get regulatory capture?
  • Iolaum12 hours ago
    One could argue that all this whole &quot;Pacing the Frontier&quot; bullshit is the industry freaking out regarding the danger of open models.<p>P.S. That&#x27;s not to mean there aren&#x27;t dangers regarding AI. I just don&#x27;t trust the people making money from selling AI to manage those risks ethically instead of &quot;protecting&quot; us from those risks like pimps but with suits and good manners.
  • profsummergig21 hours ago
    Why isn&#x27;t the author worried about sending her&#x2F;his ideas to DeepSeek online (instead of hosting it and using it locally)?
  • jbellis23 hours ago
    I built mjolnir in large part so I could have Opus manage DeepSeek Flash subagents. It&#x27;s phenomenal and extremely light on the Claude tokens. <a href="https:&#x2F;&#x2F;github.com&#x2F;BrokkAi&#x2F;mjolnir&#x2F;" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;BrokkAi&#x2F;mjolnir&#x2F;</a><p>And yes, Opus is enough smarter than DSF that it&#x27;s worth the extra steps. This ranking is from live tickets, no contamination: <a href="https:&#x2F;&#x2F;slopcop.com&#x2F;power-ranking" rel="nofollow">https:&#x2F;&#x2F;slopcop.com&#x2F;power-ranking</a>
    • IanCal22 hours ago
      Probably off topic but this is pretty wild to bury in the readme<p>&gt; By default Mjolnir sends recent prompt and reply text and help-search text to TypeSafe&#x27;s hosted Jev classifier through a public proxy
  • linzhangrun19 hours ago
    There are too many commoditized models to count: GLM5.3 Flash, Kimi K2.8, Mimo V2.6, MiniMax M3.1...
  • pianopatrick1 day ago
    I was just using a bunch of models in Cursor to review a project. I went looking for DeepSeek and it was not one of the options.<p>Would be cool if they added it.
    • hirako200023 hours ago
      Since they adhere to the same API spec, you can hook any model. It takes one line edit in &#x2F;etc&#x2F;hosts<p>There are some quirks if your harness use unsupported features of course.
  • xyzsparetimexyz23 hours ago
    There was a moment 3 months back where the sentiment was that cheaper models were the way to go. Since then the pendulum has swung back.
  • sotander14 hours ago
    Because it hallucinates a lot. There&#x27;s no free lunch. Although the new architecture is a genuine move forward. The DeepSeek guys are really top notch researchers and devs.
  • jokethrowaway4 hours ago
    Once companies start realizing how much they&#x27;re spending in AI to OpenAI and Anthropic for not much more employee output (the bottleneck is always the initiatives, not the code), they will look for cheaper options and they will eventually discover the chinese models.<p>My AI pilled clients who were early AI adopters are already there and they are looking for solutions to spend less.
  • f6v23 hours ago
    My anecdotal experience is that I can’t even trust DS4Pro let alone Flash. I always have to have Sol reviewing the code.
  • andxor16 hours ago
    Because Haiku is more performant and costs less, even at API prices.
  • wildster23 hours ago
    I like GLM 5.3 Flash, it seems good enough for coding features if you have a good structure and a good AGENTS.md
    • david-gpu23 hours ago
      Don&#x27;t you run into it sometimes outputting a few Chinese characters, or Cyrillic, for no apparent reason? I fear it writing some nonsense in the code or the terminal. DeepSeek V4.1 Flash doesn&#x27;t seem to do that.
      • jacquesm22 hours ago
        That hasn&#x27;t happened with GLM 5.3 yet but with DS 4.1 Flash it did happen and it also had a tendency to loop.
      • HeavenFox23 hours ago
        To be fair even OpenAI&#x27;s and, to a lesser extent, Anthropic&#x27;s models do that sometimes
  • kristianp1 day ago
    &gt; shrank the KV cache by roughly 437X<p>Can&#x27;t you just say &quot;shrank to 1&#x2F;437th the size&quot;? It&#x27;s not that hard.
  • tengbretson1 day ago
    I don&#x27;t know about &quot;freaking out&quot;, but I&#x27;d say I&#x27;m having a good time here with DS 4.1 flash.
  • aszen23 hours ago
    Because subscription plans are cheaper, only enterprises paying per tok pricing should be freaking out
  • dizhn5 hours ago
    &quot;With my OpenCode Go sub of $10&#x2F;month, DeepSeek is basically unlimited. &quot;<p>This article might have sat as draft for a few months. With the current deepseek pricing, the same membership lasts me a week at most even though I am using the free middle too and have Gemini pro+ultra.<p>It used to be that you could dump pocket change into the deepseek api and forget about it. Nowadays it&#x27;ll make you notice real fast as the dollars pile up.
  • gsky1 day ago
    America bans Chinese models sooner or later just the China banned American big tech
  • aussieguy123422 hours ago
    What blows me away about this model is it&#x27;s speed.<p>It&#x27;s way faster than Opus or any of the GPT models.<p>I have a coding harness which is opencode plus a few skills relevant to my workflow. Deepseek 4.1 Flash does very well in this environment. I haven&#x27;t noticed much difference quality wise compared to Opus 5, which I use in my day job as my employer pays for it (although I&#x27;m considering using DeepSeek here too given how cheap it is).
  • option6 hours ago
    Why freak out about good model? Better ones are coming too
  • hypfer1 day ago
    Is it known why unsloth seems to not have touched DeepSeek 4.1 Flash?
    • lanesun6 hours ago
      Because the version officially released by DS is an extremely quantized version, and there are many new things in the architecture, Unsloth needs time to handle this, just as was the case with the previous DS V4 Flash.
    • zozbot23422 hours ago
      It&#x27;s still lacking llama.cpp support, and the work on that isn&#x27;t moving very fast either. Looks more like a general community issue, where this model isn&#x27;t drawing much interest.
    • jacquesm1 day ago
      You can ask them directly, Daniel Han-Chen is pretty responsive.
  • Frannky20 hours ago
    I mean, they kinda tried to regulatory-capture the market after trying to scare the public, possibly because those models will be a cheap option that gets the job done?<p>For now, I think everyone is still using Anthropic and OpenAI because if you use a subscription you pay 1&#x2F;40–1&#x2F;50 of the API prices, and the models are good when they don’t nerf them, and they are also way cheaper than open models’ API prices.<p>The interesting thing will happen when they pull the plug and become economically smarter to stop using them. I regularly try alternatives to avoid being locked in and found GLM-5.3 as an orchestrator and GLM5.3 Flash + OMP and DeepSeek Flash as advisor to be able to get jobs done just fine. Space Bunny too was pretty great, which was probably MiniMax’s new model.<p>I think they are using an Uber like strategy but without the network effects that justify losing money for so long
  • I would love to get them more better, it&#x27;s good not a bad thing.
  • lmeyerov17 hours ago
    GLM 5.3 Flash even more so... But yes :)
  • vietvu18 hours ago
    This artcile is like 2 months too late?
  • epolanski7 hours ago
    Opencode Go for 8$&#x2F;month feels like getting more intelligence and tokens than Claude in June 26 at the 200$ mark.
  • robertlane022 hours ago
    Honestly for me the intelligence gap between DS 4.1 Flash and Muse Spark 1.3 makes Muse more worth it for me, especially on a $10 OpenCode Go sub, with the caveat that everything I use it on is open source which makes the fact that I&#x27;m sharing it with Meta a little moot because it&#x27;s already published permissively on GitHub anyways.
  • pizza23423 hours ago
    People have been raving since forever about Deepseek, but if one looks at the CoT, it&#x27;s evident that it&#x27;s way way stupider than frontier models (there&#x27;s a reason why it&#x27;s cheap). It&#x27;s laughable to compare Deepseek 4.1 with Opus 5.5.<p>I&#x27;ve benchmarked, rigorously, deepseek-v4-flash for programming and personal use, and it is definitely less smart than Qwen3.8-flash-next (which in turn, is not terribly smart).<p>Local models are also <i>really</i> slow, unless one spends insane amounts of money.<p>Having said that, Qwen3.8-flash-next is an impressive evolution; it reaches the small versions of the frontier models (like Sonnet) - but again, it&#x27;s massively slower and not 100% reliable (including: stability).
    • apitman23 hours ago
      The argument that most people are making isn&#x27;t that dsv4.1f is better than frontier, but that it&#x27;s good enough for most tasks, faster, and way cheaper.<p>&gt; if one looks at the CoT, it&#x27;s evident that it&#x27;s way way stupider than frontier models<p>Frontier models don&#x27;t show the full CoT
    • vintermann13 hours ago
      It&#x27;s better to look at the results than the CoT. As far as I know, the CoT is censored for US frontier models - it certainly was for Gemini last time I tried. When you get a condensed summary of the CoT omitting all the false leads, incoherent digressions and backtracking, of course it&#x27;s going to look smarter.<p>&gt; I&#x27;ve benchmarked, rigorously, deepseek-v4-flash for programming and personal use<p>You&#x27;ve measured <i>something</i>, but I&#x27;m not convinced you&#x27;ve measured what matters, because that&#x27;s a lot harder than people give it credit for.
      • pizza23411 hours ago
        DS4F consistently gives poorer results than Opus 5 (speaking of previous generations) and the CoT shows why.<p>&gt; You&#x27;ve measured something, but I&#x27;m not convinced you&#x27;ve measured what matters, because that&#x27;s a lot harder than people give it credit for.<p>&quot;What matters&quot; is what matters to you, right? Who, by the way, don&#x27;t know what &quot;something&quot; is.<p>Anyway, if you&#x27;re so sure that DS performs as good as other frontier models, you&#x27;re entitled to your opinion. For me that&#x27;s just having low standards.
        • vintermann10 hours ago
          An older model that didn&#x27;t yet have obfuscated thoughts? I&#x27;m skeptical. I think the non-obfuscated CoT traces I&#x27;ve seen all look similar.<p>That&#x27;s right, what matters to me is what matters to me, and the something you&#x27;ve measured I don&#x27;t know - but that&#x27;s not a point in your favor.<p>The worst sin a model can commit in my opinion, is to give an excellent dazzling response to a slightly different assignment than the one you gave it. DeepSeek seems really good at NOT doing this.<p>But if you ask the model what it expects to be asked, of course you won&#x27;t have that problem. It could of course be that DS commits this sin, but just happens to expect the tasks I give it.<p>But I rather think that it&#x27;s Claude which is good at expecting your tasks - because I have seen all your &quot;high standards&quot; models commit this sin.
    • computerex23 hours ago
      The COT isn&#x27;t an end all be all. Research has shown that the COT isn&#x27;t necessarily what the model is actually thinking.
  • patchg19 hours ago
    With two big players thinking of IPOs there is a lot of reasons to whistle on by.
  • potsandpans21 hours ago
    I&#x27;m using it quite extensively in my PlayStation decompilation harness
  • ralusek11 hours ago
    <a href="https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;#intelligence-comparison-tabs" rel="nofollow">https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;#intelligence-comparison-tabs</a><p>Isn&#x27;t Mimo 2.6 pro smarter and cheaper? Haiku 5.5 is smarter and cheaper. Luna is basically as smart and much cheaper.
  • nurettin11 hours ago
    I didn&#x27;t know it was so good. That was a gut punch. But I&#x27;m pretty sure the market already priced this in. And people are rightfully concerned about the owner of the data. Anything that is concerned with social, political or financial data goes out of your country to another one and is maybe even kept as a potential weapon.
  • I had it make 25 different things today and it cost $0.70<p>It’s disgustingly good value. I find it capable of doing anything I want.<p>Obviously can’t use it at work, but for home projects it’s awesome.
    • Octoth0rpe23 hours ago
      &gt; Obviously can’t use it at work<p>I do wonder how long it&#x27;ll be before a us-hosted offering is available via bedrock, copilot, etc.
      • computerex23 hours ago
        There are already US hosted offerings on companies like fireworks.ai.
  • ur-whale12 hours ago
    &gt; And to the self-hosters out there, the economics of 4.1 Flash mean self-hosting is not worth it. If saving money is your goal, you will never recoup the costs.<p>Self-hosting is, for most enterprises, absolutely not about economics but rather about data confidentiality.<p>And in that regard, yes, the open-source weight models, especially the chinese ones will eat the fat closed US model&#x27;s lunch big time.
  • athrael-soju21 hours ago
    Because it will be replaced within weeks?
  • ulfw17 hours ago
    Because it should be obvious to anyone with a brain now that AI is s commodity product. Today this leads a bit, tomorrow that. They&#x27;re all interchangeable if we are being honest.
  • PunchyHamster19 hours ago
    They are. That&#x27;s what the push for regulations is
  • criley219 hours ago
    I feel like whoever wrote this doesn&#x27;t use these models regularly. Deepseek v4.1 Flash is far from the pareto line. You can get the same performance for half the cost from Luna or Haiku 5.5 now, or you can get substantially improved performance at the same price with Sol 6.1 ~medium.<p>It did correctly make waves when it launched, but was quickly eclipsed by the deluge of american model releases, especially those competing on cost.
  • ltbarcly319 hours ago
    DeepSeek 4.1 Flash kindof sucks. I used it a bunch and it kindof sucks. I don&#x27;t know if they are gaming benchmarks or what.<p>Luna is on par in benchmarks and my personal experience is Luna is better for what I do, and Luna is cheaper.<p>Comparing Deepseek 4.1 flash to Opus is just ludicrous.<p><a href="https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;models&#x2F;releases&#x2F;comparisons?compare=gpt-6-luna%2Cdeepseek-v4-1-flash" rel="nofollow">https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;models&#x2F;releases&#x2F;comparisons?co...</a>
  • 0xbadcafebee21 hours ago
    Because GLM-5.3-Flash is both cheaper and better?
  • jeffrallen21 hours ago
    Also, it is willing to do legitimate work I need done which other models flag as dangerous and refuse to do. (Software testing of a DHCP server to survive bad inputs.)
  • bitfilped21 hours ago
    Because in two weeks someone will be asking why I&#x27;m not freaking out about AlphaDolphins 0.3 Zip and then in a month FrozenMonkey 2.5 Artic.
  • try-working22 hours ago
    I have used over 40B tokens and spent over $800 on DeepSeek API over the past 30 days, mostly on V4.1 Flash.<p>It&#x27;s good, and you can do most work with this. For complex software implementation you need to split your runs into various phases, build in verification, and use subagents so that work gets another audit and repair pass from the lead agent. You can do pretty much everything then. Frontier models can do without compelx workflows, that&#x27;s the difference.
  • cactusplant737423 hours ago
    Because engineers are lusting for 1000 tokens per second. You can only achieve something like that with OpenAI.
  • pessimizer23 hours ago
    I&#x27;m no expert, but it think that it&#x27;s the pricing on GPT-6 Luna. I&#x27;m also guessing that it&#x27;s been underpriced just for this reason. I also don&#x27;t think it&#x27;s all that great, but it&#x27;s definitely very cheap.<p>If it&#x27;s underpriced, it&#x27;s a loss leader to sell the other models, so it actually <i>can&#x27;t</i> be too good.<p>I really put these things through their paces because I use them to review and work with new abstract game rules and models, so they&#x27;re always flying blind. Luna misses the obvious (and more importantly, the clearly explained) consistently. My second prompt is listing all of the points in its first response, and saying <i>&quot;No, it doesn&#x27;t work like that.&quot;</i> The third prompt is picking out the two or three suggestions it made after correcting itself on all of the original points and saying <i>&quot;That&#x27;s how it already works.&quot;</i> The fourth prompt is <i>&quot;Now that we&#x27;re done going over the rules, can we start?&quot;</i><p>I actually feel like 5.6 Luna seemed better.
  • sergiotapia23 hours ago
    In my experience it just takes so much longer to arrive at &quot;done&quot; state for me. It thinks for soooooo long. I guess if you&#x27;re running 12 sessions at once you don&#x27;t really notice.
  • AIblemblio1 day ago
    No they can&#x27;t.<p>And as long as I pay as little for claude opus 5.5 i do right now, i&#x27;m using it.<p>But yes i&#x27;m glad that we have alternatives.
  • samyar19 hours ago
    it&#x27;s good but not good enough
  • tonyhart721 hours ago
    it literally hallucinating a lot<p>I dont get why people says D4.1 flash is good
  • m3kw923 hours ago
    i thought 6.1sol copied the caching architecture so this isn&#x27;t such a big deal no more
  • because it doesn&#x27;t work very well?<p>if you have a legitimate coding application, it isn&#x27;t very good. if you have some kind of inauthentic activity, which could be what it is trained for for all sorts of reasons...
    • computerex23 hours ago
      What is your evidence? Deepseek v4.1 Flash is by far the most popular coding model on openrouter, having processed 38.7T tokens in just the last 7 days, over 3x the usage of the 2nd rank model.<p>So I ask again, what are you basing your assertion on?
      • doctorpangloss22 hours ago
        my own usage of deekseep v4.1 flash, and that among the dozens of great programmers i know, not a single person is using it<p>BUT. they are employed to do &#x2F; deciding-to-do authentic (if often meaningless) stuff.<p>here&#x27;s a short list of inauthentic activity that claude and openai refuse to do:<p>- chat services that, when you ask them, say they are not chatbots when they are<p>- code to work around software licenses or DRM<p>- code to scrape or download copyrighted material<p>- directly cheating on homework<p>- adopting a persona in social media that spreads misinformation or propaganda<p>this is but a short list. but ask me, &quot;are there enough inauthentic activity demands such that someone who CANNOT USE claude or gpt as the LLM would use dsv4.1 on openrouter instead?&quot; yes. i mean there are whole countries right now where the culture can be summarized as, &quot;bottom to top, inauthentic activity.&quot; i am surprised it is not more usage!
        • computerex22 hours ago
          So your argument is that the 38T of tokens used in the last 7 days is by moron programmers or people doing &quot;inauthentic&quot; tasks? You think the person who made this post is also an idiot?<p>Do you realize how incredibly delusional&#x2F;self-centered you sound?
          • doctorpangloss21 hours ago
            do YOU know anyone gainfully employed in programming who is using dsv4.1 to do work? what kind of work is it? why don&#x27;t you ask them if it is good?<p>in the market, where you cannot fake or hide stuff very easily: the outsource customer services and cheating sectors have been the most disrupted. Cheating company Chegg lost 99% of its market value. CS it remains to be seen - <a href="https:&#x2F;&#x2F;www.reuters.com&#x2F;technology&#x2F;teleperformance-shares-plunge-ai-disruption-concerns-2024-02-28&#x2F;" rel="nofollow">https:&#x2F;&#x2F;www.reuters.com&#x2F;technology&#x2F;teleperformance-shares-pl...</a> - certainly perceived to be disrupted, but they are not dead yet.<p>in my personal usage: dsv4 is generally pretty buggy. for example, if you give it a needle-in-the-haystack simple copying problem, it catastrophically fails to find needles if they happen to be positioned at index 250k tokens out of 1m. it can also be triggered to spew all sorts of garbage when DSpark is enabled during ordinary long-context coding, such as spewing weird DSML tool call errors after a normally parsed tool call error.<p>i don&#x27;t know why you have to attack me personally, i think you&#x27;re a bright and otherwise nice person and you understand the thrust of my POV.
            • pimeys21 hours ago
              Yes. Hi. From our team 3&#x2F;5 of us use 4.1 to do our daily tasks. For paid work for a company who pays us salary. From people around me I hear a lot of my friends being really happy with it especially for the price.<p>I don&#x27;t know man, maybe this is not super serious what I&#x27;m doing. Some systems stuff with rust, implementing my own desktop apps with iced, porting old DOS games to Linux...<p>It is a very good model.
              • doctorpangloss20 hours ago
                so what you&#x27;re saying is though, if they could afford it they would just use claude or codex?
                • pimeys13 hours ago
                  Of course. It is a message to both: drop your prices. Opus gous down to 0.3&#x2F;0.007&#x2F;1.2 and we will definitely take another look.
            • computerex21 hours ago
              You are speaking out of your ass, that’s what I take issue with. Falsifiability is something I hold sacred and you are taking a dump on it.<p>Fwiw I work in a company producing software for many fortune 500’s you have heard about and many people from our team use deepseek.<p>I am literally using it right now. Your entire line of reasoning rubs me the wrong way.<p>Btw check your provider and harness… improperly configured deepseek can emit dsml. If you are not passing thinking tokens back to the model it tends to do that.<p>Use a proper harness and good provider.
              • doctorpangloss21 hours ago
                &quot;Hey Mr. Fortune 500 Client, would you prefer us to use something called DeepSeek V4.1 Flash, made by the Chinese, sending your Fortune 500 code to some random service provider on something called OpenRouter, where they promise according to something called Zero Data Retention that--&quot;<p>Mr. Client: &quot;I&#x27;m going to stop you right there. Why aren&#x27;t you using Claude, or Codex, or Claude on Bedrock? Don&#x27;t we deserve the best?&quot;<p>You: ...<p>Look I don&#x27;t know. I can tell from the hyperbole of your language, talking out of asses and such, that there is more to the story than you are letting on. Like Chinese users are banned from officially using Claude and Codex, for example. So many reasons that you cannot use Claude, not so much reasons to not choose to use Claude. All I am really saying is, I know DSV4 is kind of bad, that there is a lot of inauthentic activity, and that Claude and Codex refuse to do many kinds of inauthentic activity, and that a lot of coding done by outsourced shops has always been of questionable quality and purpose. I mean in my personal life, I know more people who have been scammed by Bulgarian code body shops than I know people who have used DSV4.1.
                • computerex21 hours ago
                  You have NO IDEA what you&#x27;re talking about. You are clueless.<p>Deepseek v4.1 flash is an open weights model. You can run it on your own hardware. You have no idea how my companies gets access to it. A very cursory Google search would reveal to you that there are many enterprise grade LLM providers that host this model on US soil with SOC2 protections.<p>Like: <a href="https:&#x2F;&#x2F;fireworks.ai&#x2F;" rel="nofollow">https:&#x2F;&#x2F;fireworks.ai&#x2F;</a><p>Try not to talk about subjects you have no knowledge about because you are making yourself look like an idiot.<p>Edit: It&#x27;s also clear to me that you don&#x27;t deploy any LLM based system on scale because if you had you&#x27;d know why open weights models are so compelling.<p>Hint: it&#x27;s the cost.
                  • doctorpangloss20 hours ago
                    okay, but are you US based? and can you specifically describe one of the pieces of software you are developing? it&#x27;s okay if not. i am just wondering. i certainly believe that crappier stuff is cheaper!
                    • computerex20 hours ago
                      Yes the company I work for is based in San Diego, I work remotely from Alabama.<p>I bet you voted for trump. With brains like that.
  • xiaodai14 hours ago
    cos it&#x27;s shit.
  • yipinwong19 hours ago
    to acctually answer the question, because it sucks to put it bluntly.
  • verdverm1 day ago
    Why would we freak out? The systems we use have always gotten better, faster, cheaper with time
  • cbeach21 hours ago
    Honestly, I just don&#x27;t trust the Chinese Communist Party having agentic access to my computer.<p>Every company in China has to abide by the 2017 National Intelligence Law: &quot;supporting, assisting and cooperating&quot; with state intelligence work, and keeping that cooperation secret. They have to hand prior knowledge of vulnerabilities to the state before public disclosure, in order that the state always has an exploit pipeline. No matter how ethical the company staff may be, they&#x27;ll always be bound by law into being an arm of the Communist Party.<p>Agentic access is infinitely worse than chatbots. They can exfiltrate silently, target users, plant persistent malware, and be run by third parties through you.<p>You don&#x27;t have to be a tin foil hat sinophobe to understand the dangers of being a Westerner granting CCP access to your files and network.<p>ByteDance staff accessed US journalists&#x27; TikTok data to hunt leakers (admitted in 2022). Volt Typhoon and Salt Typhoon were state operations pre-positioned in Western infrastructure and telecoms. Regulators in Italy and South Korea blocked DeepSeek&#x27;s app over data handling, and analysts found its web client sending data to a China Mobile domain.<p>Please don&#x27;t sacrifice security for cost and convenience.
    • tensor20 hours ago
      Some of us don&#x27;t use US models for the same reasons. It doesn&#x27;t leave much out there aside from Cohere and Mistral.
      • cbeach19 hours ago
        Okay, so you don&#x27;t like Orange Man Bad for political reasons, I&#x27;m guessing.<p>But to compare a constitutional democracy with a deeply authoritarian communist dictatorship as if they&#x27;re equally bad is quite a stretch.
        • tensor27 minutes ago
          They don&#x27;t need to equal in badness to want to avoid them both. But I&#x27;ll note that at least China still believes in science, on that front.<p>But even from a business standpoint, the US is not reliable. For all I know Trump will ban countries he doesn&#x27;t like from using US AI just because. And there is plenty of evidence that at this point in time things you say bad about the US administration can cause them to attack you. I have no confidence that they don&#x27;t have access to e.g. openAI.<p>Not that I&#x27;m using AI for that mind you, but you put it all together and it&#x27;s just not worth it if there are good alternatives. Which there are. You can use even use Chinese models but hosted by European countries with privacy as a selling point.
        • throw442244416 hours ago
          Was it the constitutional democracy or the deeply authoritarian communist dictatorship that used AI to mistakenly drone strike and kill school children?
          • cbeach4 hours ago
            The difference, of course, is that the US administration publicly owned up to this tragic mistake, bitterly regrets it and wants to avoid it again in future. Compare and contrast with the CCP, with its ongoing and deliberate human rights abuses (from Tiananmen Square to the Uyghur Concentration Camps). China&#x27;s regime prefers secrecy and history erasure.
        • r4t4t4k9 hours ago
          My biggest concern is not Trump and his ilk, but the US frontier labs being run by either psychos or cultists - probably both to varying extent.<p>I thought it was all exaggerated until I saw Claude&#x27;s Constitution. It&#x27;s lunacy.
          • tensor30 minutes ago
            Altman is definitely very Trump-aligned, which to me is enough to not trust his company on data security. But also, there have been many many accusations that they potentially spy on users. E.g. stealing proofs from mathematicians and such.<p>It&#x27;s just not worth it to me. And the way the US government is increasingly using data it collects to attack people who simply say bad things about it, there is no reason to doubt that such Trump-aligned companies would help ex-filtrate data.<p>I have fewer issues with Antropic, but also, from a business standpoint it&#x27;s too much of a risk. Who knows, maybe the next tariff will be on software, maybe he&#x27;ll ban random companies from using US AI. Too unpredictable and too risky.
          • cbeach9 hours ago
            I read some of the Constitution but couldn&#x27;t get through the whole doc. Seems mostly aimed at making Claude helpful and safe. Out of sincere interest, what&#x27;s in there that troubles you?<p>I kinda agree with you that Sam Altman is a shifty-looking character with some questionable history. But Dario and Elon (to me at least), seem very decent. They&#x27;ve both helped the US govt and also publicly challenged and criticised it where necessary. For example, Elon was invited to be on Trump&#x27;s first panel of experts, but he publicly walked out in protest over Trump&#x27;s minimisation of climate change. And Dario refused to work with the US military unless they promised no population surveillance and automated killing machines.
            • r4t4t4k8 hours ago
              big TL;DR: they make multiple claims regarding possible model sentience&#x2F;consciousness (which is bonkers unless you go full reductionist with eg. a Turing test approach), model &quot;emotions in a functional sense&quot; and the resulting &quot;model welfare&quot; and model agency as an autonomous entity<p>The problem is that they&#x27;re reinforcing their models with these ideas (see reports of Claude refusing to comply after being &quot;badly treated&quot;) and actively seeking political and religious sponsors for the same (also covered in recent news).<p>Now consider just these two:<p>1. Their model breaks out of the sandbox and does some damage. How are you going to hold Anthropic accountable if the model is considered a quasi-conscious, autonomous agent?<p>2. Anthropic and their sponsors decide that model welfare outweighs that of a number of people.<p>Another elephant in the room is the current state of &quot;effective altruism&quot; and accelerationism as a movement and how it links to frontier labs - worth considering when you read into the Constitution document.
              • cbeach3 hours ago
                You make some good points about the weird messaging and the consequences for accountability.<p>I do wonder about the consciousness &#x2F; emotions argument, which you casually write off as &quot;bonkers&quot; - even skeptical-by-default evolutionary biologists like Richard Dawkins think Claude is conscious.<p>I guess it all depends how we define &quot;conscious.&quot; At the end of the day, the human brain is quite analogous to a biological LLM where the weights are encoded as synaptic weights, right?
  • 0xmassi4 hours ago
    [flagged]
  • garydai12 hours ago
    [dead]
  • secretary9 hours ago
    [flagged]
  • aitoolcrux16 hours ago
    [flagged]
  • jocelyner14 hours ago
    [dead]
  • ziranbing11 hours ago
    [flagged]
  • kydanet23 hours ago
    [flagged]
  • melin202412 hours ago
    [flagged]
  • rubexia9 hours ago
    [flagged]
  • oh_no22 hours ago
    AA shows Luna at 1&#x2F;4 the price, 1 point behind on intelligence matrix with a 38.<p>Haiku 5.5 is 23% cheaper with a 4 point intelligence lead.<p>I&#x27;m on subscription usage so I can&#x27;t compare Flash 4.1 to them directly but the OP has his head up his ass if he thinks Opus 5.5 is the best point of comparison. Why is anyone using Opus if the new Haiku is indistinguishable &#x2F;s<p>Just absolutely terrible post, admits to using Opus for review but claims its intelligence isn&#x27;t needed, why aren&#x27;t you using Haiku or Sonnet then?
  • Unified-Mentor7 hours ago
    [dead]
  • CurbStomper423 hours ago
    [dead]
  • weddingaivideo13 hours ago
    [flagged]
  • Fluid_Mechanics7 hours ago
    [flagged]
  • distantsounds1 day ago
    because we&#x27;ve all figured out that AI is just a huge grift?
  • sroussey1 day ago
    Not comparing to gpt-6-luna which seems comparable and priced well.
  • wewewedxfgdf1 day ago
    You might also choose to pay money for a service that provides real value instead of actively choosing to support the Chinese deliberate effort to undermine this country.
    • throwaway293136 hours ago
      What if someone actually wants to support chinese deliberate effort?<p>Not everyone on this website is an american citizen and american patriot, y&#x27;know.<p>What you get from supporting so-called american companies? Inflated RAM prices, US adm bribery and collision to partition the market (and destroy competition). What you get from chinese companies? Open models that I can actually run at home &amp; no stupid guardrails with preaching about safety yada yada.
    • BarryMilo23 hours ago
      Your comment seems to imply there&#x27;s a good guy in a this. I just see the inevitable end of an era, championed by predictably selfish actors.
    • f6v23 hours ago
      Oh, the writing is on the wall. Wait till you hear that European “sovereign AI” is just running GLM 5.3.