8 comments

  • aabdi1 hour ago
    I don’t think it would be surprising that people want to write their own kernels.<p>A big problem with the existing engines like llama or sd is that they don’t support optimal graph compilation. Usually this means about a real 2 or 3x multiplier loss relative to optimal. Cuda graphs do okay but they still leave a lot on the floor<p>It’s usually worth it to optimize in that context if you are willing to peer into the mechanics.<p>Of course that’s expensive. You need to know how to appropriately pipeline and merge your kernels.
  • stephbook6 hours ago
    Should have started with writing your own blog posts.
    • lelanthran4 hours ago
      &gt; Should have started with writing your own blog posts.<p>While the page looks vibe-coded[1], the content itself does not have any AI tells. What are the tells you are seeing?<p>[1] Too many sites I find on HN frontpage these days slow my PC to a crawl. I assume they are all using the same autogenerated HTML, Javascrip and CSS to make animated backgrounds :-( On this specific site scrolling is laggy.
      • interpol_p2 hours ago
        I stopped reading almost immediately. The stylistic choices in the writing just felt like LLM to me. Examples:<p>&quot;depth estimation that beats PyTorch on CPU in half the memory&quot; — &quot;…beats X in Y…&quot;<p>&quot;Most LocalAI backends wrap somebody else’s engine, and that is the right default.&quot; — &quot;…and that is the right&quot;<p>&quot;MLX and the rest are maintained by people who are better at those models than we are&quot; — &quot;better at those models than we are&quot; — it&#x27;s this thing that LLMs do where they are kind of weirdly confident but overly deferential<p>&quot;This post is about what those ports buy&quot; — &quot;…buy&quot; used in this context<p>&quot;Same model, 1.31x the speed&quot; — &quot;Same X, something Y&quot; — it&#x27;s this overconfident yet deferential writing style<p>The further I read, the more tells there are. I find it incredibly tiring to read LLM generated prose and I&#x27;m not sure why. Is it because I&#x27;m aware it&#x27;s not human written and have an unconscious bias? Or is it because the style is just full-on, &quot;Not X but Y. Those performance gains are bought, not earned. This stops, that starts. Read on, or don&#x27;t, that&#x27;s the follow-up&quot;
        • MattPalmer10862 hours ago
          For me, it&#x27;s that AI writing is always trying to be clever for every single point it makes (and constantly uses the same language patterns when doing so).<p>Its like listening to an insufferable clever dick, who is not as bright as they think they are. You would also find it incredibly irritating if a human talked like that
        • lelanthran2 hours ago
          Now that you point it out, there are quite a few tells, still not as many as most of the slop that gets posted here.<p>I think it&#x27;s because of the laggy scrolling that I didn&#x27;t read the whole thing anyway, just the first few screens.
        • layer82 hours ago
          Also, “honest reading” — without any context explaining why one would contemplate a dishonest reading.
        • nnevatie1 hour ago
          [dead]
        • PatronBernard1 hour ago
          Why are you using an em-dash though?
      • wonnage4 hours ago
        [flagged]
    • winter_blue4 hours ago
      I found the post insightful and interesting. I&#x27;m not sure it was written with AI assistance, but even if it was, I don&#x27;t see that as a reason to dismiss it. For what it&#x27;s worth, I spend hours everyday reading AI output and summaries.
    • pjmlp3 hours ago
      Same could be said for all that talk about having Claude do their work.
      • xienze2 hours ago
        I&#x27;ve had this debate before on HN. The excuses are generally &quot;well it can write better code than most developers, but an LLM can&#x27;t write better prose than most people&quot; (I strongly disagree with this) and, what I think is at the heart of the matter, &quot;text is for the reader to read directly, code is hidden.&quot; Or in other words, &quot;as long as I can&#x27;t tell it&#x27;s AI, it&#x27;s fine.&quot;
        • nnevatie1 hour ago
          I would rather read faulty English, succinct sentences and getting to the point, than the generic filler LLMs produce.
          • pjmlp43 minutes ago
            I would also code review code that people actually put some effort learning on how to write it, even if it had one bug or two.
    • nnevatie6 hours ago
      Came here to say the same. Really tiring to read these slop-infested posts, where everything has the “right shape”.
      • polotics3 hours ago
        The thing is... although the writing is unmistakably full of LLMisms, I can&#x27;t fault the `author` for having produced a slop readme. The content earns its keep, it only grates because of the robotic personality. We need another word than &quot;slop&quot; for this.<p>&quot;blland&quot;, &quot;llame&quot;,... ?
    • altmanaltman5 hours ago
      I went through the post because of your comment but it really doesn&#x27;t look like AI slop. Can you please share why you feel like its slop and not written by a human? I can also say &quot;should have started writing your own comments&quot; to you and its unfalsifiable. Blanket accusations with no proof is not a good move really.
      • nnevatie4 hours ago
        The post is full of signs. Here&#x27;s only a couple of examples:<p>&gt; The method, the measurements, and what it costs us.<p>&gt; That is the general shape of these wins.<p>&gt; Parity is the gate, speed is the follow-up<p>I could go on and on, but you probably get the point. If you don&#x27;t find anything funny with the above, you might have not been enough-exposed to slop.
        • altmanaltman3 hours ago
          What do you mean you could go on and on? Why do you think those sentences are AI written.<p>And okay, your second argument is that I just don&#x27;t know slop because I am not exposed to it? But you don&#x27;t know anything about me or what I am exposed.<p>You&#x27;re just making random claims and stating they are correct without any evidence or arguments.
          • bendmorris2 hours ago
            What kind of evidence do you expect beyond &quot;random claims&quot; here?<p>This post is incredibly obviously AI generated, to the extent that I doubt a human author edited it at all. Not &quot;written with AI assistance&quot; but full on &quot;give Claude some bullets and hit publish.&quot; It contains tons of tropes that show up in all AI writing and which people are highlighting here.<p>What would convince you of that?
          • skavi1 hour ago
            <a href="https:&#x2F;&#x2F;www.pangram.com&#x2F;history&#x2F;125358dc-9f80-4fe2-8540-85a9c314aa6c" rel="nofollow">https:&#x2F;&#x2F;www.pangram.com&#x2F;history&#x2F;125358dc-9f80-4fe2-8540-85a9...</a>
          • rcarmo3 hours ago
            They follow the tropes I get when I ask AI to do docs or summaries. Very Opus style, this one.
        • wannabe444 hours ago
          It&#x27;s always hyping up something and throwing punch lines in every sentence. Normies love this shit.
          • nnevatie4 hours ago
            Yes, it’s basically business-as-usual but on speed.
  • dennis163846 hours ago
    I had a similar success with Model2Vec static embedder and NER inference (both GGUF, compiled for WASM), ported to plain C from ONNX Runtime.<p>Wasm size from 30Mb to 300kb and 1.5x speedup. It&#x27;s definitely worth it for performance or distribution size.
  • scottcodie5 hours ago
    I did took a native c++ approach when writing a relational transformers engine (RelativeDB). My journey was pytorch -&gt; c++ -&gt; Triton (lang). While C++ was more performant than Triton, I couldn&#x27;t afford to optimize on every gpu. I just accepted the ~15% throughput loss for my cloud service, which honestly wasn&#x27;t bad for the amount of flexibility I got out of it.<p>But the cpp port of vllm looks great, that&#x27;d be great if you&#x27;ll maintain that. I hit the same limitations with vllm.
  • piterrro3 hours ago
    Could this vllm port be faster to install? Im starting gpu machine multiple times a day and it takes 5 minutes to set vllm up. If Inise this port that time is minimized?
  • adithyassekhar5 hours ago
    What you get: X is the A, Y is the B.
  • federicoTXTS1 day ago
    [flagged]
  • openrockets2 hours ago
    [dead]