4 comments

  • dghlsakjg 34 minutes ago
    I know everyone wants to crap all over these setups that are impractical, but this is how progress happens.

    People will keep plugging away at this and figure out how to avoid wearing the hard drive, how to make it run faster, custom hardware buses etc.

    Keep going! I personally can't wait for the day when a 1t param model runs off a $200 SSD instead of a $50k rack of Nvidia chips.

    • arjie 18 minutes ago
      Haha 1T on $50k might be a bit hopeful, mate, even at FP8. But I too am hopeful.
      • hedora 12 minutes ago
        AMD already demonstrated 1T on strix halo clusters. << $10K at original MSRP.
  • AHASIC 31 minutes ago
    I read a comment on here a few months back I wanna restate. Basically, there is a good chance that Apple is betting that the LLMs in the future will be so efficient that those that consumers will use everyday will be easily computed by the iPhone or even bigger ones on Macs. Honestly makes the most sense that we are heading that way in a few years latest.
    • Mistletoe 19 minutes ago
      What hardware advances would we need to see for that to happen? It feels like everything in that arena has kind of plateaued.
      • sudo_cowsay 4 minutes ago
        It could be on software side too. OpenAI has certainly not plateaued.
  • jbird99 44 minutes ago
    At what, 10 tokens per hour? These disk swapping methods all have the same drawbacks - kill your drive early, and slow as hell.
    • kennywinker 34 minutes ago
      It says very prominently in the post: 4.5-5t/s for 80b on an M5
    • wat10000 34 minutes ago
      Isn’t it only writes that kill drives?
      • sudo_cowsay 3 minutes ago
        Yeah, that's why most of these comments seem weird to me.
      • Alpha3031 7 minutes ago
        Yes for NAND, and I suppose nobody is using mechanical hard drives for this.
  • brrrrrm 53 minutes ago
    this is cool but like, are we just vibe coding NAND burners at this point? these decode times don't really tell the whole story, because prefill becomes the bottleneck.

    half an hour to process 10k tokens on an M5 seems... not great

    • kennywinker 31 minutes ago
      Not great for coding, or realtime agent interactions. But for background processing tasks overnight? Seems like it’d work pretty well