20 comments

  • GodelNumbering1 hour ago
    Back of the envelope calculation (could be off, correct me if I am)<p>If you are a large enough company that spends million+ on inference a month, it makes sense to buy a GB300 rack ($6M on top range from what I could find) which has 20.7 TB. Since the model is mixed trained (MXFP4), you would need less than 10% of the rack&#x27;s memory to serve the full model. Aggregate HBM bandwidth: 576 TB&#x2F;s. You can run over 6000 parallel agentic workflows (each with ~100k context on average) at ~30 tok&#x2F;s.<p>Assuming the annual amortization+electricity at $1.5M&#x2F;year and about 50% average annual utilization, you get less than 60 cents (USD) per million output token, for a frontier model with plenty of capacity to share, all your data never leaving premises and well over an order of magnitude cheaper!<p>As long as a company believes that the openweight models will continue to get more capable and &#x27;AI is here to stay&#x27;, this model provides the first solid footing for a decision to just buy a rack.
    • 9999000009991 hour ago
      And hire 2 or 3 dev ops to keep it running ?<p>That another 400 to 700k.<p>It becomes your problem and not someone else’s. However, I don’t trust hosted LLMs for anything that needs to be private.
      • wongarsu1 hour ago
        Where do I sign up to get 200k&#x2F;yr to keep one rack running? Sounds like an incredibly chill job
        • arjie1 hour ago
          Apparently it’s going to take the 3 of us to do this, mate. Going to get so much reading done.
        • LeonM52 minutes ago
          What you get is not what you cost.<p>40% overhead is quite typical, so you&#x27;d be looking at $120k&#x2F;year. In the USA I&#x27;d consider that a competitive salary for an admin capable or keeping a $6M rack of specialized hardware running 24&#x2F;7.
          • margalabargala12 minutes ago
            Yeah but you don&#x27;t need <i>two</i> such people, or even one, dedicated to this single rack.<p>A company of the size that this is worthwhile for, probably has dedicated devops on staff already and can add this rack to the inventory with no additional staff.
            • 9999000009997 minutes ago
              Who is going to upgrade the models ?<p>Who is going to fix it when the api does something weird ?<p>Who is going to proactively make sure it’s not overheating?<p>Chat GPT has enterprise contracts for a reason.
      • russell_h1 hour ago
        &gt; However, I don’t trust hosted LLMs for anything that needs to be private.<p>Why not? Do you trust AWS with things that need to be private?
        • Kevcmk1 hour ago
          More than I trust frontier labs. AWS doesn&#x27;t need to recoup 9 digits USD of capex
      • lumost1 hour ago
        There will be cloud&#x2F;SaaS vendors who have lower cost of labor&#x2F;capital due to automation and financing terms.<p>Having these models in the open caps the inference margin.
      • GodelNumbering1 hour ago
        &gt; And hire 2 or 3 dev ops to keep it running<p>Not a devops but I&#x27;d say one full time is already too many.
        • dboreham1 hour ago
          Yes but zero is not enough and where do you get a fraction of a competent dev op from?
          • layer857 minutes ago
            From the other dev-op work you’re doing.
            • ayewo22 minutes ago
              Understood but sharing your existing devops resources with this will soon become a bottleneck especially when any major downtime will keep several engineers (and long-running agents) blocked from any meaningful work until availability improves.
              • layer84 minutes ago
                Not my experience, from an SMB that maintains its own hardware and services. You have a certain contingent of competent engineers who distribute their work across projects, and it generally works out fine. Or course you plan with some redundancy and fall-back plans in your systems.
      • slicktux1 hour ago
        Just like that new jobs created by AI! Localized model maintainer&#x2F;technician.
      • toomuchtodo1 hour ago
        You’ll slap some training on existing technologists&#x2F;infra&#x2F;sysadmin folks and perhaps have a support contract for the edge cases (hardware troubleshooting and advanced replacement).<p>(managed an entire data center building with thousands of servers a lifetime ago with ~2-3 other people, it’s only gotten easier over the last two decades imho)
      • clint1 hour ago
        Just let it manage itself, what could go wrong! :)
    • reckless1 hour ago
      I think the licensing that would likely apply to a company that&#x27;s able to afford ~$6M rack and the associated infrastructure muddies this somewhat
      • petu1 hour ago
        I think internal use is allowed at any scale in the license?<p>&gt; 4. The requirements set forth in Sections 2 and 3 do not apply to: (a) internal use of the Software, defined as any use that does not make the Software, its outputs, or its underlying capabilities available to third parties; [...]
        • jsnell41 minutes ago
          It&#x27;s very hard to make sure none of the outputs are ever made available to third parties. Source code can end up widely distributed (e.g. client-side js, open source). Prose will frequently get shared across organization boundaries (e.g. emails, websites, documents).
      • vidarh1 hour ago
        As far as I can tell, license fees are only applicable if you have more than 20m USD&#x2F;year revenue from services provided using the model, or serve more than 100m users.
  • m_ke2 hours ago
    Also open sourced a bunch of infra to go with it.<p>Anyone who claims open source and open weights models are &quot;decel&quot; needs to get their head checked<p><a href="https:&#x2F;&#x2F;github.com&#x2F;MoonshotAI&#x2F;MoonEP" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;MoonshotAI&#x2F;MoonEP</a><p><a href="https:&#x2F;&#x2F;github.com&#x2F;kvcache-ai&#x2F;AgentEnv" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;kvcache-ai&#x2F;AgentEnv</a><p><a href="https:&#x2F;&#x2F;github.com&#x2F;MoonshotAI&#x2F;FlashKDA" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;MoonshotAI&#x2F;FlashKDA</a>
    • jvanderbot2 hours ago
      This comment would be much better without the second line
      • ike_a2 hours ago
        I&#x27;m not sure I understand the case for open-source models being decelerationist, is this it?<p>Decel:<p>- Potentially reduces investor appetite for funding big labs.<p>- More risk of powerful AI getting in bad hands -&gt; more regulation.<p>Accel:<p>- More competition so big labs can&#x27;t rest on laurels.<p>- More research in open, so all labs can accrete advancements faster.<p>I feel like open-source = acceleration has a much more clear argument. (and how bad would deceleration be in any case?)
        • m_ke0 minutes ago
          the argument is that we should all fold and let sam altman burn trillions of dollars on naive scaling and pay monopoly prices for their closed APIs until the models are good enough to be closed off for &quot;safety&quot; reasons so that they can take an even larger cut by competing directly with us
        • sosodev1 hour ago
          I think the argument is that decentralization leads to deceleration because it means less centralized funding and data. Those are the two primary ingredients for accel.<p>The problem with the decel&#x2F;accel rhetoric is that it lacks nuance.
        • Smaug1231 hour ago
          If your worldview is “most of the progress is made by closed labs, then open labs fast-follow” (which isn’t implausible given the documented distillation of Fable), and further that open labs cannot make make meaningful progress vs the closed labs <i>except</i> by fast-following and that they won’t pick up the ability to make progress after the closed labs are gone, then driving closed labs out of business slows down overall progress.
          • reissbaker1 hour ago
            I think it&#x27;s pretty hard to hold that worldview: Anthropic couldn&#x27;t ship a reasoning model until they copied DeepSeek R1&#x27;s homework, and they&#x27;ve all copied DS-style super-sparse MoEs at this point too.
            • chorizo53 minutes ago
              That’s a really good point. Folks really need to read the papers coming out of these Chinese labs. Every paper from the DeepSeek team has been a step change.
        • StevenWaterman2 hours ago
          I think it&#x27;s basically open weights =&gt; more inference competition =&gt; less profit from inference =&gt; less training competition
          • f311a1 hour ago
            &gt; less training competition<p>I think you meant less research and experiments in big labs because they don&#x27;t get all the AI money.<p>Training is expensive, but they also have more than 10 000 of employees combined and they cost a lot of money.
        • Iolaum2 hours ago
          Open Source models decelerate growth of closed AI. For people who think (or want) AI = closed_AI then that argument has weight. Good luck getting them to update their priors.
        • zozbot2342 hours ago
          Open source AI is actually a lot less &quot;powerful&quot; than genuine frontier models, i.e. it has a much tighter inherent capability ceiling. This is &quot;decelerationist&quot; from a purely AGI-pilled point of view but it&#x27;s actually great if you&#x27;re worried about a capabilities arms race putting AI Safety at severe risk.<p>Kimi K3 is plausibly a lot <i>less</i> dangerous than a totally jailbroken ChatGPT&#x2F;Gemini&#x2F;Claude Sonnet (let alone Opus or Fable!) and it&#x27;s quite deeply weird how no one seems to be calling for <i>those</i> models to be banned or restrained by further regulation. Why the double standard against the less concerning (but more efficient!) open weight models?
          • ike_a1 hour ago
            Do you think they are inherently less powerful? I&#x27;d imagined that closed labs have a head start &#x2F; more funding so the open labs are playing catch-up.<p>Is there a world where open source models end up at the frontier, or do you think there are structural&#x2F;first-principles reasons why this won&#x27;t happen?
            • zozbot2341 hour ago
              If you&#x27;re targeting widespread local&#x2F;on prem deployment which is what many open weight models are doing, that inherently limits your scale in terms of total model weights&#x2F;inference-time compute compared to running in a few centralized datacenters. A centralized model will always be able to leverage a larger scale of deployment, placing it much closer to the genuine &quot;frontier&quot;.
      • janalsncm16 minutes ago
        It’s a different topic but it’s correct imo. Open sourcing things is the best way to accelerate development.<p>Put another way, if you want to slow things down, put it behind a paywall, tag ideas ans “intellectual property” (meaning you’re the only one who can use it) and get the lawyers involved (injecting our slow legal system).<p>None of the above is a judgement call on whether development should be accelerated.
      • SirLordBoss1 hour ago
        Why? Absolutely correct, especially considering the position of the person they&#x27;re referring to
      • acedTrex2 hours ago
        Why?
      • Der_Einzige2 hours ago
        [flagged]
      • viccis1 hour ago
        I know it&#x27;s a hard ask on this site, but I need you to start parsing content and not tone. It was a helpful bit of context, even if it was a bit vitriolic.
        • BeetleB11 minutes ago
          Some battles are simply not going to be won.<p>I, for example, dislike reading comments complaining the submission (or another comment) is LLM generated. Focus on the content, not the style.<p>I&#x27;m not going to have my way, and nor shall you.
        • Almondsetat1 hour ago
          Why should I waste time parsing content and not tone? Why can&#x27;t the commenter just avoid the tone? It even saves time since you can write less!
          • viccis1 hour ago
            Because discourse around tone isn&#x27;t productive. You could wipe this whole comment thread, starting with the parent of mine, and lose exactly zero information.
          • m_ke1 hour ago
            [flagged]
            • oxzidized1 hour ago
              &gt; because it gets people like you going<p>So you&#x27;re admiting you were trolling?
              • m_ke1 hour ago
                no, the goal was to spark a conversation about the value of *open AI* and it looks like it worked
        • SubiculumCode1 hour ago
          The content was &#x27;you need to get your head checked&#x27;. That isn&#x27;t tone, that directly implying that if you hold that position, there is something wrong with you. It&#x27;s rude and unnecessary.
        • jvanderbot36 minutes ago
          It depends on if we expect posters to be informative and include nuance.
    • vanuatu1 hour ago
      To me its clear that it is decel<p>the only reason other labs can catch up is because the frontier labs can be distilled, and they siphon a % of the labs&#x27; revenue to reinvest into the next iteration<p>full accel would mean nationalizing the big 2 labs and locking in manhattan project style until RSI<p>(Edit: some great counterpoints in the replies. my view has definitely been changed!)
      • m_ke1 hour ago
        only if you only get your news from main stream business press and Big Lab propaganda channels<p>There&#x27;s no chance K3 is a distill of Fable, it came out way too soon after the limited fable release to be feasbile.<p>If you look at all of the top ML conferences, chinese labs contribute way more to advances in ML than &quot;Open&quot;AI and Anthropic: <a href="https:&#x2F;&#x2F;www.reddit.com&#x2F;r&#x2F;TheMachineGod&#x2F;comments&#x2F;1pi4q7f&#x2F;papers_at_neurips_2025&#x2F;" rel="nofollow">https:&#x2F;&#x2F;www.reddit.com&#x2F;r&#x2F;TheMachineGod&#x2F;comments&#x2F;1pi4q7f&#x2F;pape...</a><p>This K3 release just helped every other lab on the planet stay in the race by making it possible for them to build on top of it, placing them at the frontier starting line instead of having to spend billions of their own dollars and risking it all to attempt to catch up.<p>The open source contributions I linked to above will move the whole field forward and reduce the costs of training and inference for everyone.<p>Open science compounds on it self, every new advancement pushes the field forwards and opens up new grounds for future improvements.<p>It is impossible for a single closed lab to consistently stay ahead of the rest of the field, especially in a huge growing research area like machine learning. The only only advantage the big labs have is money, but the naive scaling game is not sustainable long term when you have to pay 10-100x more then the fast followers and we start getting more and more open models or use case specific models that can handle 90% of high volume use cases.<p>Research is a high variance, low expected value activity, meaning that the few large concentrated labs have to be conservative with their bets and double down on proven things when scaling up. The rest of the field is like a diversified portfolio, with thousands of players making smaller riskier bets that only require a few of them to succeed (like K3 did here, and DeepSeek a year ago)<p>EDIT: also if you look at most of the work from OpenAI, it&#x27;s mostly taking existing promising open research work and scaling it up. (except for things like CLIP and etc from Alec Radford)
        • chorizo59 minutes ago
          Appreciate this point. Back in grad school, I published a peer reviewed paper with all the source code and datasets used. Got heckled at a conference talk by staff from a commercial lab. They shouted, we figured this out five years ago, lol. Also our approach is still better. But they don’t release their source code or publish much, so no one knows what this approach is or if it’s actually better.<p>And years down the line, lots of other research labs used my code and cited my paper.
          • m_ke50 minutes ago
            Oh this definitely happens all the time. I was an early employee at Clarifai, which won imagenet a year after alexnet and we were able to stay on the frontier for about 2 years before a bunch of open source models were matching our results. It was always some random PhD research project spinout or some random kid in Boston named Alec Radford.<p>We had a bunch of things that we never published that ended up being major research findings years later at top conferences.
      • calebkaiser40 minutes ago
        OpenAI&#x27;s head of strategic futures publicly stated that you can&#x27;t explain the quality of the newest Kimi via distillation.<p>Further, you can just read the papers released alongside most open models. Plenty of hugely influential research results published that drive the frontier forward. It&#x27;s not like these models are just existing architectures downloaded from Huggingface and trained on frontier lab APIs.
      • ZeroGravitas1 hour ago
        Distilling can be done by fast follower closed models too, so this argument against open models doesn&#x27;t hold up.
      • pluto_modadic1 hour ago
        being able to replicate it in the open means that there&#x27;s nothing special about frontier models.<p>Frontier models would have to do something extraordinary or unique, or unreplicatable, because clearly there is no moat, and US companies are sitting on huge nvidia valuations and get surprised when competitors beat them.
  • fahrradflucht2 hours ago
    License: <a href="https:&#x2F;&#x2F;huggingface.co&#x2F;moonshotai&#x2F;Kimi-K3&#x2F;blob&#x2F;main&#x2F;LICENSE" rel="nofollow">https:&#x2F;&#x2F;huggingface.co&#x2F;moonshotai&#x2F;Kimi-K3&#x2F;blob&#x2F;main&#x2F;LICENSE</a><p>&gt; If the Licensee or any of its affiliates operates a Model as a Service business, and the aggregate revenue of the Licensee and its affiliates exceeds 20 million US dollars (or the equivalent in other currencies) in total over any consecutive 12 months, the Licensee must enter into a separate agreement with Moonshot AI before using the Software or its derivative works for any commercial purpose.<p>+ the existing 100 million monthly active users, or more than 20 million US dollars for commercial products have to name Kimi clause
    • simonw1 hour ago
      Kimi 2.6 had the same janky license: <a href="https:&#x2F;&#x2F;huggingface.co&#x2F;moonshotai&#x2F;Kimi-K2.6&#x2F;blob&#x2F;main&#x2F;LICENSE" rel="nofollow">https:&#x2F;&#x2F;huggingface.co&#x2F;moonshotai&#x2F;Kimi-K2.6&#x2F;blob&#x2F;main&#x2F;LICENS...</a> - looks like they&#x27;ve been doing that at least as far back as K2.<p>The models they released in 2025 - <a href="https:&#x2F;&#x2F;huggingface.co&#x2F;moonshotai&#x2F;models" rel="nofollow">https:&#x2F;&#x2F;huggingface.co&#x2F;moonshotai&#x2F;models</a> - were clean MIT. They started doing the &quot;modified MIT&quot; thing in January 2026 with moonshotai&#x2F;Kimi-K2-Thinking
      • zozbot2341 hour ago
        To be clear, the restriction on &quot;Model as a Service&quot; past $20M yearly revenue is new to K3. K2.x had the attribution requirement for any commercial use with more than $20M monthly revenue.
        • simonw1 hour ago
          Good catch, thanks.
      • lossolo1 hour ago
        In what way was the license &quot;janky&quot;? Moonshot isn&#x27;t a trillion dollar US behemoth. If these terms allow it to release near frontier models with open weights, so startups and anyone with enough hardware can use them, while charging only companies with more than $20 million in revenue or 100 million users, that seems like a reasonable trade off.
        • simonw1 hour ago
          I use the term &quot;janky&quot; for any time someone releases something under a supposedly open source license that doesn&#x27;t comply with the OSI definition.<p>&quot;Modified MIT&quot; is the perfect example of that.<p>I&#x27;m not saying it&#x27;s unreasonable, or that you <i>can&#x27;t</i> release under such a license - it&#x27;s your software, use whatever license you like!<p>I&#x27;ll call it &quot;janky&quot; when you do.
          • lossolo58 minutes ago
            So it&#x27;s more like your personal definition, not the normal meaning of the word. Ok, because &quot;janky&quot; is informal slang also meaning poor quality, unreliable etc. I understand that you&#x27;re using janky as shorthand for not OSI compliant. I just think that can blur two separate issues: whether a license qualifies as open source under the OSI definition, and whether the license itself is unreasonable&#x2F;badly designed.
            • simonw54 minutes ago
              I&#x27;m trying to make fetch happen.
        • bilbo0s1 hour ago
          Reasonable for people at the bottom of the tech pyramid. Certainly for startup guys this is heaven sent.<p>For people at the top of the tech pyramid, I could totally see why they&#x27;d want to put an end to all this &quot;Open Source LLM&quot; stuff.
    • cosmojg28 minutes ago
      It&#x27;s worth noting that the weights of machine learning models are not subject to copyright in the United States as they are the product of an automated optimization process (e.g., stochastic gradient descent, expectation maximization, genetic algorithms) rather than human authorship. Granted, this has yet to be fully tested in court and going to court is expensive, so it&#x27;s likely that your employer would prefer to err on the side of caution and respect such attempts at model licensing anyway. Nonetheless, this was partially tested last year in <i>Thaler v. Perlmutter</i> which affirmed that copyright requires human authorship, reading the Copyright Act&#x27;s provisions on ownership, duration, and transferability as presupposing a human author[1].<p>If you want to assess the position of the U.S. Copyright Office for yourself, the relevant text can be found in the <i>Compendium of U.S. Copyright Office Practices</i> § 313.2, &quot;Works That Lack Human Authorship&quot;[2], which states:<p>&gt; […] the Copyright Act protects “original works of <i>authorship</i>.” 17 U.S.C. § 102(a) (emphasis added). To qualify as a work of “authorship” a work must be created by a human being. See Burrow-Giles Lithographic Co., 111 U.S. at 58. Works that do not satisfy this requirement are not copyrightable.<p>&gt; […] the Office will not register works produced by a machine or mere mechanical process that operates randomly or automatically without any creative input or intervention from a human author. The crucial question is “whether the ‘work’ is basically one of human authorship, with the computer [or other device] merely being an assisting instrument, or whether the traditional elements of authorship in the work (literary, artistic, or musical expression or elements of selection, arrangement, etc.) were actually conceived and executed not by man but by a machine.” U.S. COPYRIGHT OFFICE, REPORT TO THE LIBRARIAN OF CONGRESS BY THE REGISTER OF COPYRIGHTS 5 (1965).<p>[1] <a href="https:&#x2F;&#x2F;media.cadc.uscourts.gov&#x2F;opinions&#x2F;docs&#x2F;2025&#x2F;03&#x2F;23-5233.pdf" rel="nofollow">https:&#x2F;&#x2F;media.cadc.uscourts.gov&#x2F;opinions&#x2F;docs&#x2F;2025&#x2F;03&#x2F;23-523...</a><p>[2] <a href="https:&#x2F;&#x2F;www.copyright.gov&#x2F;comp3&#x2F;chap300&#x2F;ch300-copyrightable-authorship.pdf" rel="nofollow">https:&#x2F;&#x2F;www.copyright.gov&#x2F;comp3&#x2F;chap300&#x2F;ch300-copyrightable-...</a>
    • throwaway274482 hours ago
      I wonder how they&#x27;ll figure out who to target for litigation when this license is violated.
      • paxys2 hours ago
        How many companies host and serve models via API and have a $20M+ revenue? Going to be pretty straightforward to catch offenders.
      • bityard1 hour ago
        Software licenses aren&#x27;t enforced through litigation as much as they are enforced through the _threat_ of litigation and legal risk. In other words, pretty much every company pays lawyers to minimize legal risk. Those lawyers inevitably look at all the contracts, agreements, and software licenses, and tell the C-suite what to do in order to keep their legal exposure as low as possible. &quot;Don&#x27;t violate other companies&#x27; IP,&quot; is pretty low-hanging fruit in those conversations.<p>It is a very rare (and ballsy, and perhaps incompetent) company that ignores their lawyers&#x27; recommendations to adhere to the letter of all of the software licenses they are bound to.
      • mike_hearn38 minutes ago
        They have certainly trained in some secret call&#x2F;response pairs that would uniquely identify Kimi serving.
      • ffsm82 hours ago
        They haven&#x27;t litigated the last public non-compliance... Despite that one being extremely public. So probably not at all for now.
        • JumpCrisscross2 hours ago
          &gt; <i>They haven&#x27;t litigated the last public non-compliance</i><p>What was it?
          • drawnwren2 hours ago
            Cursor was thought to be but they were later found to be using an authorized provider [1]<p>1 - <a href="https:&#x2F;&#x2F;x.com&#x2F;Kimi_Moonshot&#x2F;status&#x2F;2035074972943831491?lang=en" rel="nofollow">https:&#x2F;&#x2F;x.com&#x2F;Kimi_Moonshot&#x2F;status&#x2F;2035074972943831491?lang=...</a>
        • Iolaum2 hours ago
          Maybe because a non public agreement was in place?
    • ericpauley1 hour ago
      What is it a license for, though? Copyright? Are model weights even copyrightable?
    • embedding-shape1 hour ago
      Hah, so much for &quot;open weights&quot; :D Fair enough, they&#x27;re at least downloadable, shame they didn&#x27;t end up being actually open, nor open source, the community had really high hopes for this. But again, still available for download, so better than nothing else I suppose.
      • jhonof1 hour ago
        Open source projects frequently have separate commercial licenses no? This isn&#x27;t abnormal.
        • embedding-shape1 hour ago
          No, I don&#x27;t know a single project I&#x27;d call &quot;open source&quot; (as understood by the FOSS community) that restricts what you can do with it, that&#x27;d make it very much &quot;not open source&quot; as you&#x27;re discriminating against specific persons&#x2F;groups&#x2F;fields of endeavor.
          • parodysbird1 hour ago
            GPL licenses also restrict what you can do with it...
            • embedding-shape1 hour ago
              Beyond reciprocity which is the entire point of that license, what restrictions does it come with? I guess you could say that it has a restriction of adding new restrictions, but surely that&#x27;s not what you&#x27;re talking about?
  • eamag2 hours ago
    &gt; we build a self-evolving, hierarchically organized knowledge graph that agents continuously expand through web-scale exploration across knowledge-intensive and coding domains<p>That&#x27;s interesting!
  • whimsicalism2 hours ago
    It&#x27;s funny that we&#x27;ve finally returned to tanh activation functions, time is a circle.
    • rhdunn46 minutes ago
      From the paper (page 6 with a comparison to GLU and SwiGLU) they are not using tanh directly (i.e. f(x) = tanh(x)) but:<p><pre><code> f_gate(b,x) = b * tanh(x &#x2F; b) * sigmoid(x) f_up(b,x) = b * tanh(x &#x2F; b) </code></pre> Looking at the graph I wonder if this is to try and get the best of both GLU (better representation at higher values of x &gt;~ 5) and SwiGLU (the value bump just before 0).
      • whimsicalism43 minutes ago
        Still more of a tanh than I&#x27;ve seen in years
    • pinkmuffinere1 hour ago
      Wow this is fascinating enough that I’m actually going to read tfa lol
  • a-dub2 hours ago
    knowledge graph guided task synthesis. very cool! i have long wondered about the &quot;how do you get good coverage of all the tasks&quot; problem.<p>maybe some interesting theoretical work there around the rate of production of new knowledge itself and various mechanisms (human approaches, mechanistic approaches, etc).
  • storus2 hours ago
    What would be the current best method to fine-tune it for my own specific agentic tasks? LoRA + DPO? GRPO? Something else?
    • whimsicalism2 hours ago
      LoRA + SFT, but it&#x27;ll be big - better to wait for a finetuning API from one of the providers, I wouldn&#x27;t jump straight to RL or off-policy pseudo-RL like DPO.
  • eamag2 hours ago
    Can someone explain what are teachers in Multi-Teacher On-Policy Distillation? I can imagine math, coding and other verifiable domains, but they also have biology? Is it where distillation from bigger models come in?
    • sosodev1 hour ago
      They reference <a href="https:&#x2F;&#x2F;thinkingmachines.ai&#x2F;blog&#x2F;on-policy-distillation&#x2F;" rel="nofollow">https:&#x2F;&#x2F;thinkingmachines.ai&#x2F;blog&#x2F;on-policy-distillation&#x2F;</a><p>If I understand correctly, it&#x27;s distillation via having a teacher model score each of the student&#x27;s tokens for a problem based on their own probabilities of generating that token at each step in the sequence. The reward&#x2F;loss is then applied as RL.<p>The multi-teacher bit seems to imply they&#x27;re distilling from multiple models. It&#x27;s light on the details, but it seems like it could be part of distilling from frontier&#x2F;closed models. Provided they calculate the logprobs, which OpenAI seems to allow via API but not Anthropic. Maybe they have a way of estimating the logprobs externally?<p>This method can be used to learn any domain from the teacher. Biology included.
      • mike_hearn37 minutes ago
        It&#x27;s not that light on the details. I read the paper and they train several different models in parallel over a few different domains and then they distill from their own models to get the final model.
    • porridgeraisin1 hour ago
      They&#x27;re other models yes.
  • Tejush53 minutes ago
    Any guess on the pre-training tokens&#x2F;flops they consumed?
  • weberer1 hour ago
    Does anyone know if a torrent is available? I think it would take quite a while to download 1.5tb from their servers.
  • colesantiago2 hours ago
    This is amazing to witness. Moonshot open sourcing Kimi K3, a frontier AI and other components really means we are getting abundant AI for all of humanity.<p>Kudos to Moonshot for truly being what OpenAI should have been.<p>Fable-level and frontier AI should be open source and available to everyone for free.
    • embedding-shape1 hour ago
      &gt; Kudos to Moonshot for truly being what OpenAI should have been.<p>Kudos to Moonshot for making these weights available for download. Lets not fool ourselves and claim these are &quot;open source&quot; by any understanding of the concept though, there are usage restrictions (even if you download them) and also training data isn&#x27;t clearly broken down either, nor it it actually using a FOSS license.
      • lossolo1 hour ago
        &gt; Kudos to Moonshot for making these weights available for download. Lets not fool ourselves and claim these are &quot;open source&quot; by any understanding of the concept though, there are usage restrictions (even if you download them) and also training data isn&#x27;t clearly broken down either, nor it it actually using a FOSS license.<p>They will probably never release the training data because that represents a large part of their competitive moat. The same is true of US companies (Google, OpenAI, Meta etc) none of which has released the full training data for its open models.<p>They use private datasets that cost a lot to acquire, synthetic datasets and a lot copyrighted material for which they don&#x27;t have licensing.<p>What matters most is that, with the necessary hardware, I can download a near frontier model, run it and modify it however I want. The other concerns you mentioned are just noise. And if my company is generating $20 million in revenue or serving 100 million users, it can probably afford a relatively inexpensive commercial license.
        • embedding-shape1 hour ago
          &gt; They use private datasets that cost a lot to acquire, synthetic datasets and a lot copyrighted material for which they don&#x27;t have licensing.<p>Yeaaah, and this, of course, is worth it, because it leads to you being able to download a near frontier model. Don&#x27;t get me wrong, long-term humanity is probably better of with science with little regards to pesky things like ethics and provenance, but we also have a tendency to not fully realize the downstream or wider effects until way too late.<p>And sure, there is a lot of reasons to go with keeping your software proprietary too, I&#x27;m not trying to claim otherwise, same with training data. It&#x27;s just that usually we don&#x27;t call proprietary software &quot;open source&quot; unless it is open source, regardless of the reasons someone keep it proprietary or not, could be for whatever reason really.
        • bilbo0s1 hour ago
          &gt;<i>They will probably never release the training data because that represents a large part of their competitive moat</i><p>Not that what you&#x27;ve written isn&#x27;t the case. However, in <i>addition</i> to what you&#x27;ve written, (or probably even before what you&#x27;ve written), there&#x27;s the fact that everything they&#x27;re training on is stolen IP. Same with US LLM labs.<p>Let&#x27;s not kid ourselve&#x27;s about where the training data is coming from. They are not asking artists, writers, coders, content creators, etc etc etc for permission to use their creations.<p>Anthropic, Moonshot et al are doing incredible things, but we shouldn&#x27;t gloss over the costs. Both present and future costs are kind of enormous.
        • dboreham1 hour ago
          Exactly. I wish people would just stop saying &quot;open source&quot; regarding models. Even &quot;open weight&quot; is disingenuous. &quot;self-hostable&quot; would be more honest.
          • embedding-shape50 minutes ago
            It&#x27;s a gradient almost, with some steps. So far, I think you could categorize every single released so far as one of:<p>- Proprietary - No access beyond remote endpoints<p>- Downloadable - You can run it, but there are restrictions and training data&#x2F;code isn&#x27;t public and&#x2F;or under FOSS license, nor are the weights under a FOSS license<p>- Open weights - The weights are under a FOSS license and downloadable without restrictions, but not all of training code&#x2F;data is public or under a FOSS license<p>- Open source - The model architecture, weights, training code and data are all under a FOSS license, and you can freely download and use them without any restrictions<p>Few models hit the mark for &quot;open source model&quot;, many models people call &quot;open weights&quot; would go under &quot;downloadable&quot; here, which personally I think would be accurate.
  • lenerdenator2 hours ago
    What would it take to get an American open model to compete with this?
    • embedding-shape1 hour ago
      Latest &quot;big&quot; release from any of the bigger American lab must have been GPT-OSS-120b I think? Released ~summer 2025, so pretty much one years ago. Doesn&#x27;t seem like it&#x27;ll happen by itself, so something either forcing their hand figuratively, or something forcing their hand literally.<p>Personally I was wishing&#x2F;hoping for one of the recent Gemma releases to be in the ~100B class at least, but sadly Google is keeping that all for themselves.
      • Stagnant1 hour ago
        NVIDIA-Nemotron-3-Ultra-550B-A55B was released in June 4th 2026 and I think it was the largest open US model until thinking machines&#x27; Inkling (975B) was released a couple of weeks ago.
      • layer81 hour ago
        August 2025.
      • lossolo1 hour ago
        Yeah, besides that, the only big other open source US model that was worth looking at was Inkling (975B params), Jul 15, 2026.<p><a href="https:&#x2F;&#x2F;thinkingmachines.ai&#x2F;news&#x2F;introducing-inkling&#x2F;" rel="nofollow">https:&#x2F;&#x2F;thinkingmachines.ai&#x2F;news&#x2F;introducing-inkling&#x2F;</a>
      • whimsicalism1 hour ago
        the longer i read this comment the wronger it gets
        • embedding-shape1 hour ago
          Don&#x27;t make me add more of my thoughts to the comment, I can still edit it.<p>Deeply fun &quot;care about improving people&#x27;s lives&quot; quote on your user page :p
          • whimsicalism47 minutes ago
            how is 2025 two years ago? GPT-120b is not even close to the biggest recent American release, etc. etc.
    • 0x000xca0xfe5 minutes ago
      A rogue employee with a USB stick.
    • segmondy1 hour ago
      if it&#x27;s to be believed, it&#x27;s so easy. you distill it, and you have a copy in 2 weeks, isn&#x27;t that what China is doing? so 2 weeks from now, we should have one.
    • boomskats1 hour ago
      An act of G̶o̶d̶ Congress?
    • chrsw1 hour ago
      Something beyond my imagination
  • abratabia1 hour ago
    [flagged]
  • histiq2 hours ago
    [flagged]
  • tudou5272 hours ago
    [flagged]
  • kiaansaraiya2 hours ago
    [dead]
  • samxli2 hours ago
    [flagged]
  • m00dy2 hours ago
    I would want to see three things before drawing strong conclusions:<p>End-to-end tokens&#x2F;sec and cost on realistic coding agent trajectories, including tool outputs and retries, not isolated decode benchmarks.<p>Cache hit rates and prefill cost for branching, multi-turn sessions.<p>Router-load distributions after post-training, where expert collapse or specialization problems often show up.
  • brcmthrowaway1 hour ago
    Is this p-hacking?