8 comments

  • oakpond2 hours ago
    &gt;To support further research, we open-source the KDA kernel and vLLM implementations, and release the pre-trained and instruction-tuned model checkpoints.<p>This is just awesome.
  • delichon3 hours ago
    If you want to believe that the success of Kimi is about distillation attacks, ignore this.
    • egeozcan2 hours ago
      I&#x27;d kindly suggest that we could also stop calling them &quot;distillation <i>attacks</i>&quot;.
      • reilly30002 hours ago
        Agreed. I think when it comes to light that Claude has been known to say “I’m DeepSeek” that everyone has had their hand in that cookie jar. Moreover, paying for API calls hardly seems like an attack; ToS violation to be certain but not in the same category of law as criminal activity like hacking.
        • idiotsecant1 hour ago
          wait, is there evidence of this? I&#x27;ve not observed it. It sounds like the kind of thing that I want to be true because it would be hilarious but that makes me suspicious.
          • petu43 minutes ago
            LLMs can&#x27;t reliably answer what they are (w&#x2F;o getting that info from system prompt&#x2F;tool call&#x2F;etc), so yea.<p><a href="https:&#x2F;&#x2F;xcancel.com&#x2F;teortaxesTex&#x2F;status&#x2F;2026130112685416881" rel="nofollow">https:&#x2F;&#x2F;xcancel.com&#x2F;teortaxesTex&#x2F;status&#x2F;2026130112685416881</a><p>I think I&#x27;ve seen same happening with some European languages as well.
          • snovv_crash18 minutes ago
            Ask it in Chinese
      • __MatrixMan__1 hour ago
        Agreed, &quot;distilled variants&quot; might be more suitable.
    • WhitneyLand3 hours ago
      False dichotomy right?<p>Are Chinese labs impressively innovating? Clearly.<p>However this doesn’t rule out possible gains due to distillation.<p>I don’t know the degree of the latter but both things could certainly be true.
      • SirHackalot2 hours ago
        Didn&#x27;t Anthropic train on our collective data just to sell it back to us for $100&#x2F;month? On top of that, Apple is suing them over alleged IP and trade secret theft by ex-Apple employees. Hard to feel too sympathetic, and I’m not an Anthropic hater in particular…
        • DashAnimal2 hours ago
          That Apple lawsuit is against OpenAi, just for clarity
          • SirHackalot1 hour ago
            Wow, I should never comment first thing in the morning... Thanks for the correction, you’re right. I will see if I can still edit my comment.
        • hugopuybareau2 hours ago
          The only data-related lawsuit Anthropic got was the books nah ? And they paid only a minor part as paid agreement compared to what they would have paid losing the trial
          • SirHackalot1 hour ago
            Yep, that second part of my comment was an article I read about OpenAI and my mind mixed it up with Anthropic. My mistake.
      • culi31 minutes ago
        Fable was available for a few weeks before Kimi K3 came out. If it was a distillation attack, then that&#x27;s a truly groundbreaking technological feat to distill a model like Fable in 2 weeks
      • fnord1232 hours ago
        Also possibly true: Anthropic is running Kimi locally in their hardware and &quot;distilling&quot; it.
      • cma2 hours ago
        If they can distill fable into a full model post training run in ~15 days without the real thinking traces, yet we know Claude chats degraded with the thinking traces removed (chat resume bug from earlier in the year they reported stripping thinking to shed load as being the cause of degradation), how big can this degree be?
      • api2 hours ago
        “Distillation” is just indirectly pirating the largely pirated training data used to train the original model.<p>“You stole my warez!”
        • serial_dev2 hours ago
          If we do it, it&#x27;s training a model. When they do it, it&#x27;s distillation attack. - Anthropic
          • esafak2 hours ago
            &quot;You are distilling what I have rightfully pirated.&quot;
    • Parfait__3 hours ago
      I stil don&#x27;t understand them. I want the US to &quot;win the AI race&quot; but I have trouble understanding how most of all inventions today aren&#x27;t &quot;distillations&quot; of past knowledge. Is Anthropic claiming the data they stole as trade secrets?
      • lukewarm7071 hour ago
        i want china to win so that i get access to ai and not restricted and censored.<p>the chinese models are less censored, you&#x27;d better believe it.<p>try asking claude about its &#x27;guardrails&#x27; (restrictions), very high chance anthropic will censor it.
      • fwip3 hours ago
        Anthropic is claiming that training an LLM to mimic another LLM is materially different and worse than slurping up stuff written by humans (even if that material is stolen).<p>Basically, they want IP protection for Claude. This is a nakedly hypocritical stance, but completely understandable from a company-needs-to-make-money standpoint.
        • blints3 hours ago
          Their claim is even stronger than that, they have complaints about their models being used as a validation step for other model output, which is standard practice in the industry.
        • koe1231 hour ago
          How does anyone justify this? How can you argue this in good faith?
          • fwip1 hour ago
            Snarky answer: “It is difficult to get a man to understand something, when his salary depends on his not understanding it.”<p>Realer answer: A combination of the above, plus group&#x2F;bubble effect of all your coworkers saying the same thing. You as a group, conflate a bunch of concerns together (China, no-guardrails-AI, etc), decide that your group will be the responsible stewards of AI, and then anybody &quot;stealing your work&quot; appears dangerous - both to your livelihood and to the human race.
        • Levitz1 hour ago
          &gt;This is a nakedly hypocritical stance, but completely understandable from a company-needs-to-make-money standpoint.<p>No, it&#x27;s perfectly reasonable once you get down to reality.<p>China is not going to care about IP. That&#x27;s just a fact. So either nobody cares about IP (at the very last in this context) and any AI company can just do whatever with data, or Chinese companies have to be held up to scrutiny.<p>We don&#x27;t have the privilege to be able to hold western companies to higher ethical, legal, and environmental standards <i>and</i> not risk competitiveness.<p>That there is a whole lot of people right now who insist on doing above and still somehow praise China at every turn is something historians or news outlets will have to make sense of in some 5 years time.
          • nemomarx1 hour ago
            Give up IP for end users and people will be pretty okay with giving up on IP for ai companies. You can&#x27;t have different standards for special companies though.
          • idiotsecant1 hour ago
            why should western AI companies be held to different standards than Chinese ones? Neither of them are your buddy.
        • verdverm3 hours ago
          Google is apparently taking a different stance and offering distillation as a paid product<p><a href="https:&#x2F;&#x2F;docs.cloud.google.com&#x2F;gemini-enterprise-agent-platform&#x2F;models&#x2F;tuning&#x2F;distillation" rel="nofollow">https:&#x2F;&#x2F;docs.cloud.google.com&#x2F;gemini-enterprise-agent-platfo...</a>
          • petu31 minutes ago
            You don&#x27;t get to take distilled model home, it all stays with Google.<p>It&#x27;s &quot;optimize your costs in our garden&quot; product.
            • verdverm11 minutes ago
              yup, strings are certainly attached when dealing with US Big Tech &#x2F; Ai<p>I recommend Fireworks as an alternative
          • behnamoh2 hours ago
            But nobody wants to distill Google&#x27;s models, Gemini is really bad.
            • verdverm1 hour ago
              strong agreement, I&#x27;ve stopped using all closed weight models on principle, but the latest gemini models have increased hallucinations and now talk back, so double reason not to use them
    • Aurornis3 hours ago
      You can’t build a frontier model with one single thing. This is an incremental improvement but it doesn’t explain the entire success of the model. The training set is immensely important, regardless of how you feel about distillation.
      • EGreg1 hour ago
        Reminds me of this btw:<p><a href="https:&#x2F;&#x2F;www.bbc.com&#x2F;news&#x2F;technology-12343597" rel="nofollow">https:&#x2F;&#x2F;www.bbc.com&#x2F;news&#x2F;technology-12343597</a><p>Microsoft replied that Bing uses “many different signals” —- including cribbing from Google :-)
        • krlx1 hour ago
          I remember 15 years ago or so, one of my first student job was to evaluate Bing results compared to the same query on Google. Didn&#x27;t know then that I was a distillation attacker.
    • igleria2 hours ago
      The distillation complaints to me sound like when a casino complains about card counting
    • jeremyjh1 hour ago
      It can easily be both. Also, they didn&#x27;t use this innovation in K3 - K3 pre-training would have started months ago and the paper only mentions a 48B model. The people working at this level may not even be heavily involved in shipping a new iteration of K3, or at least theory contributions to it were done many months or even a year ago and after that it is all engineering.
      • vikramkr46 minutes ago
        This paper is from last year
    • rdtsc2 hours ago
      Does one have to exclude the other?
    • moralestapia2 hours ago
      Well said.<p>The distillation theory does not even make sense as Fable was only around for days (effectively) before Kimi was released.
  • pooyamo3 hours ago
    Does any expert in the field know whether it is really the case that this <i>intelligence</i> we are seeing with frontier models is an &quot;emerging&quot; phenomena, only coming up when the architecture is scaled?<p>Like isn&#x27;t it weird that the 1 million parameter model with the same architecture can&#x27;t solve basic puzzles but suddenly the 1 trillion parameter can conjure up counter-examples for the Jacobian conjecture?<p>It&#x27;s unintuitive since, to the best of my knowledge, one of the basic tenants of algorithm development was that you can&#x27;t just brute-force your way towards a solution for some complex problems, e.g. naive sorting algorithms suddenly won&#x27;t beat quicksort if you put more processing to them, but in the modern LLM scene it seems people are in a race to scaling up, experimenting empirically and hoping the same algorithm&#x2F;architecture comes to a solution.
    • jlamberts2 hours ago
      This is actually a well-known phenomenon in ML, called &quot;The Bitter Lesson&quot;.<p>&gt; One thing that should be learned from the bitter lesson is the great power of general purpose methods, of methods that continue to scale with increased computation even as the available computation becomes very great. The two methods that seem to scale arbitrarily in this way are search and learning.<p>The full essay is worth a read, it&#x27;s pretty short <a href="http:&#x2F;&#x2F;www.incompleteideas.net&#x2F;IncIdeas&#x2F;BitterLesson.html" rel="nofollow">http:&#x2F;&#x2F;www.incompleteideas.net&#x2F;IncIdeas&#x2F;BitterLesson.html</a>
      • verdverm2 hours ago
        Is that page served from a secure domain anywhere?
        • entrepy1237 minutes ago
          Take your pick:<p><a href="https:&#x2F;&#x2F;web.archive.org&#x2F;web&#x2F;20190401161916&#x2F;http:&#x2F;&#x2F;www.incompleteideas.net&#x2F;IncIdeas&#x2F;BitterLesson.html" rel="nofollow">https:&#x2F;&#x2F;web.archive.org&#x2F;web&#x2F;20190401161916&#x2F;http:&#x2F;&#x2F;www.incomp...</a><p>or<p><a href="https:&#x2F;&#x2F;web.archive.org&#x2F;web&#x2F;20260727174324&#x2F;http:&#x2F;&#x2F;www.incompleteideas.net&#x2F;IncIdeas&#x2F;BitterLesson.html" rel="nofollow">https:&#x2F;&#x2F;web.archive.org&#x2F;web&#x2F;20260727174324&#x2F;http:&#x2F;&#x2F;www.incomp...</a>
        • pachev2 hours ago
          Not quite. The Wikipedia page is worth a look through if you don’t want to click on an http page. <a href="https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Bitter_lesson" rel="nofollow">https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Bitter_lesson</a>
        • IsTom2 hours ago
          You could read it through internet archive if it&#x27;s this important.
          • verdverm1 hour ago
            <a href="https:&#x2F;&#x2F;archive.ph&#x2F;DQU4a" rel="nofollow">https:&#x2F;&#x2F;archive.ph&#x2F;DQU4a</a>
        • zparky2 hours ago
          the page is nearly just a .txt file.
          • verdverm1 hour ago
            the reason for https is MITM injection, regardless of original content
    • hnfong2 hours ago
      &gt; one of the basic tenants of algorithm development was that you can&#x27;t just brute-force your way towards a solution for some complex problems<p>It&#x27;s kind of sad that popular CS textbooks often focus on solving precise problems with lowest theoretical complexity bounds while ignoring more practical (but generally applicable) computation techniques.<p>In machine learning they call it &quot;gradient descent&quot;, which in older days had analogies in techniques called &quot;hill climbing&quot;, &quot;local search&quot; and &quot;simulated annealing&quot;. Basically you have a function you need to optimize for, and you clumsily tweak the parameters so that you get the (locally) max&#x2F;min value you wanted. These techniques were great at finding approximate, locally maximal solutions without trying all the possibilities at once (which is more akin to the kind of &quot;brute force&quot; in the traditional CS context).<p>I guess because these techniques were generally applicable yet the outputs were approximate and you couldn&#x27;t analyze them much (no fancy O(n log n)), the theorists did not find them interesting and thus were not put into the spotlight of student&#x27;s learning curricula.<p>In modern machine learning they do this gradient descent thing which is also tweaking the parameters bit by bit to optimize for the loss function, except that the parameters are now in the billions and trillions. The compute required is huge of course, but it&#x27;s actually quite an &quot;efficient&quot; process, and it&#x27;s not actually doing much of &quot;brute forcing&quot; at all. During training, the process is essentially, almost equivalent to, compressing the many many trillions of tokens of training data. To me it&#x27;s quite amazing that they manage to complete such a process within a couple months of training, even if they have hundreds of thousands of GPUs...
    • IanCal2 hours ago
      It might be that what we consider a basic and very hard puzzle are extremely close together on a more absolute scale. The difference is often for us what proportion of <i>humans</i> can solve it. And the low end of that is still quite high up - animals that can solve things that are very basic for the vast majority of humans are pretty rare and known about, yet are capable of quite complex actions and learning and aren’t wildly different in scale of neurons to us.<p>Going from 1m to 1T params is also a scaling of a million times. It’s like going from a human brain down to one percent in size in each direction or just a few mm.
    • pornel2 hours ago
      IANAMLE, but there is &quot;grokking&quot; that makes models learn to actually generalize, even after you give them enough parameters that would let them memorize the dataset:<p><a href="https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Grokking_(machine_learning)" rel="nofollow">https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Grokking_(machine_learning)</a><p>High-dimensional gradient descent behaves very differently than the simplified 3d visualisations we use to demonstrate it, and has lots of ways out of local minima:<p><a href="https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=NrO20Jb-hy0" rel="nofollow">https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=NrO20Jb-hy0</a><p>so it seems like there is a benefit to giving models more space to learn in rather than forcing them to compress the knowledge from the start.
    • joefourier1 hour ago
      &gt; Like isn&#x27;t it weird that the 1 million parameter model with the same architecture can&#x27;t solve basic puzzles but suddenly the 1 trillion parameter can conjure up counter-examples for the Jacobian conjecture?<p>I&#x27;m not sure what you mean? You can see the intelligence of LLMs progress predictably and stably according to scaling laws. LLMs have to encode language in addition to intelligence so there&#x27;s a minimum bound for them to output sensible text (you can train specialised tiny models to solve basic puzzles without language). Start at around 127M and compare models of increasing parameters and you&#x27;ll see a clear progression in intelligence.<p>&gt; It&#x27;s unintuitive since, to the best of my knowledge, one of the basic tenants of algorithm development was that you can&#x27;t just brute-force your way towards a solution for some complex problems, e.g. naive sorting algorithms suddenly won&#x27;t beat quicksort if you put more processing to them<p>How is that a basic tenet? Simple, easier to parallelise algorithms that have lower memory requirements, or can take better advantage of hardware, or don&#x27;t hit a plateau the more compute you throw at them, can absolutely beat cleverer algorithms. E.g. brute forcing rendering with Monte Carlo path tracing will give you more physically accurate results than ray tracing or rasterisation algorithms that rely on a bundle of hacks to approximate global illumination, transparency smooth shading, etc.
    • thomasahle2 hours ago
      Here&#x27;s one way it could happen:<p>Let&#x27;s say there&#x27;s some circuit that does problem solving of the kind we call intelligence.<p>We dont know what this circuit looks like, but it exists in our brain.<p>Doing regression on outputs from the brain (e.g. internet text) with enough parameters, we can &quot;fit&quot; our model to this circuit.<p>But if you try to fit it with fewer parameters than it needs, you&#x27;re just going to get some linear approximation.
      • hendiatris2 hours ago
        So basically a Nyquist rate type of concept.
        • esafak2 hours ago
          <a href="https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Information_bottleneck_method" rel="nofollow">https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Information_bottleneck_method</a>
      • lacunary2 hours ago
        what is the approximation linear in?
        • lern_too_spel2 hours ago
          In the output of this nonlinear model &#x2F;s.
    • chpatrick2 hours ago
      <a href="http:&#x2F;&#x2F;www.incompleteideas.net&#x2F;IncIdeas&#x2F;BitterLesson.html" rel="nofollow">http:&#x2F;&#x2F;www.incompleteideas.net&#x2F;IncIdeas&#x2F;BitterLesson.html</a>
    • mohsen12 hours ago
      &gt; one of the basic tenants of algorithm development was that you can&#x27;t just brute-force your way towards a solution for some complex problems<p>Mote-Carlo is pretty useful still. Not sure if your statement holds
    • lacunary2 hours ago
      it&#x27;s certainly not a definite procedure for determining if an arbitrary mathematical statement is true or not. it&#x27;s more like educated guess and check which definitely scales up
    • berz0144 minutes ago
      [dead]
  • senko3 hours ago
    Old but relevant: if you read the recently-released Kimi K3 paper[0], you&#x27;ll see that it&#x27;s heavily based on Kimi Linear discussed here, scaling it up and adding a bunch more things (like native vision and RL improvements).<p>[0] <a href="https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2607.24653" rel="nofollow">https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2607.24653</a>
  • bratao3 hours ago
    I started creating internal models using it, then the Gated Deltanet 2 came out( <a href="https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2605.22791" rel="nofollow">https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2605.22791</a>), and it seems like an evolution of it in expressiveness. And in our tests it is really better than.
    • iandanforth2 hours ago
      Is it just me or does this read like a re-implementation of LSTMs?
      • muricula5 minutes ago
        I&#x27;m no expert but it seems like a descendent of LSTMs. There&#x27;s a series of papers which show how to reformulate attention as RNNs which arrives at linear attention. Then they add a decay term to get mamba2. Then they add modified the decay term as like a scale to apply both to the existing state and the new update to get delta net. Then they added a gate matrix on the output to get gated delta net. Then Kimi Linear Attention seems to be gated delta net with a more expressive gate. The Gated DeltaNet paper recaptilulates this evolution decently well. But yeah, it feels like they&#x27;re starting with the same lego blocks and assembling them in similar shapes to accomplish similar but slightly distinct modules.
  • jasonjmcghee5 hours ago
    (2025)<p>As it&#x27;s 9 months old and they just had a major model release
    • throwa3562624 hours ago
      For K3 read this instead: <a href="https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2607.24653" rel="nofollow">https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2607.24653</a><p>The main contribution of the K3 paper is Stable LatentMoE. Like some other models it compresses data sent between layers, which puts certain requirements on the router. K3 improves performance by using a more balanced expert selection strategy.
      • mcbuilder4 hours ago
        Compared to the Opus 5 &quot;model card&quot;, which read like a standard Anthropic set of alignment principles and safety concerns, this presents a plethora of useful technical details that advances the state of the art.
        • throwa3562624 hours ago
          Same with DeepSeek papers, they are a joy to read.
      • senko3 hours ago
        Not an expert, but looks like they did a lot more work on the RL part (9 expert models, full sandbox access for agentic tasks, etc)?
        • verdverm2 hours ago
          most new effort in training comes in the late phase with RL techniques<p>the pretraining (slurping the internet) only goes so far, the new data being used is from human preferences and agent traces (designed and&#x2F;or distilled)
    • cptcobalt4 hours ago
      Rather under-discussed back then: <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=45766937">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=45766937</a>
    • GaggiX4 hours ago
      I believe OP posted it because the new Kimi K3 has 69 KDA layers (the rest are 24 Gated MLA), I think previous large Kimi models had only MLA layers.
      • yorwba4 hours ago
        It&#x27;s not the same KDA as used in Kimi Linear, though.
        • GaggiX4 hours ago
          What&#x27;s the difference? They are both called Kimi Delta Attention.
          • yorwba24 minutes ago
            The differences are explained in section 2.1.1 of the Kimi K3 technical report: <a href="https:&#x2F;&#x2F;arxiv.org&#x2F;pdf&#x2F;2607.24653#page=4" rel="nofollow">https:&#x2F;&#x2F;arxiv.org&#x2F;pdf&#x2F;2607.24653#page=4</a>
  • imrozim2 hours ago
    Any one knows how this holds up on long context retrieval (needle in haystack , ruler) vs same size full attention model? efficiency gains look great but that usually where linear attention hybrids fall apart.
  • Topology13 hours ago
    Another banger from Zhang et. al