Why we write our own C and C++ inference engines

(localai.io)

37 points | by eatonphil 2 days ago

6 comments

  • dennis16384 3 hours ago
    I had a similar success with Model2Vec static embedder and NER inference (both GGUF, compiled for WASM), ported to plain C from ONNX Runtime.

    Wasm size from 30Mb to 300kb and 1.5x speedup. It's definitely worth it for performance or distribution size.

  • stephbook 3 hours ago
    Should have started with writing your own blog posts.
    • lelanthran 1 hour ago
      > Should have started with writing your own blog posts.

      While the page looks vibe-coded[1], the content itself does not have any AI tells. What are the tells you are seeing?

      [1] Too many sites I find on HN frontpage these days slow my PC to a crawl. I assume they are all using the same autogenerated HTML, Javascrip and CSS to make animated backgrounds :-( On this specific site scrolling is laggy.

    • pjmlp 39 minutes ago
      Same could be said for all that talk about having Claude do their work.
    • winter_blue 1 hour ago
      I found the post insightful and interesting. I'm not sure it was written with AI assistance, but even if it was, I don't see that as a reason to dismiss it. For what it's worth, I spend hours everyday reading AI output and summaries.
    • nnevatie 3 hours ago
      Came here to say the same. Really tiring to read these slop-infested posts, where everything has the “right shape”.
      • polotics 6 minutes ago
        The thing is... although the writing is unmistakably full of LLMisms, I can't fault the `author` for having produced a slop readme. The content earns its keep, it only grates because of the robotic personality. We need another word than "slop" for this.

        "blland", "llame",... ?

    • altmanaltman 2 hours ago
      I went through the post because of your comment but it really doesn't look like AI slop. Can you please share why you feel like its slop and not written by a human? I can also say "should have started writing your own comments" to you and its unfalsifiable. Blanket accusations with no proof is not a good move really.
      • nnevatie 1 hour ago
        The post is full of signs. Here's only a couple of examples:

        > The method, the measurements, and what it costs us.

        > That is the general shape of these wins.

        > Parity is the gate, speed is the follow-up

        I could go on and on, but you probably get the point. If you don't find anything funny with the above, you might have not been enough-exposed to slop.

        • wannabe44 1 hour ago
          It's always hyping up something and throwing punch lines in every sentence. Normies love this shit.
          • nnevatie 1 hour ago
            Yes, it’s basically business-as-usual but on speed.
        • altmanaltman 15 minutes ago
          What do you mean you could go on and on? Why do you think those sentences are AI written.

          And okay, your second argument is that I just don't know slop because I am not exposed to it? But you don't know anything about me or what I am exposed.

          You're just making random claims and stating they are correct without any evidence or arguments.

          • rcarmo 5 minutes ago
            They follow the tropes I get when I ask AI to do docs or summaries. Very Opus style, this one.
  • scottcodie 2 hours ago
    I did took a native c++ approach when writing a relational transformers engine (RelativeDB). My journey was pytorch -> c++ -> Triton (lang). While C++ was more performant than Triton, I couldn't afford to optimize on every gpu. I just accepted the ~15% throughput loss for my cloud service, which honestly wasn't bad for the amount of flexibility I got out of it.

    But the cpp port of vllm looks great, that'd be great if you'll maintain that. I hit the same limitations with vllm.

  • piterrro 52 minutes ago
    Could this vllm port be faster to install? Im starting gpu machine multiple times a day and it takes 5 minutes to set vllm up. If Inise this port that time is minimized?
  • adithyassekhar 2 hours ago
    What you get: X is the A, Y is the B.
  • federicoTXTS 1 day ago
    [flagged]