DeepSeek V4 Flash on a Single AMD MI300X

(github.com)

57 points | by zhoutong 1 hour ago

2 comments

  • majke 53 minutes ago
    I don't think you can buy a single "MI300X" unit, right? Only the box with x8 of these at a cost of ~250K EUR.
    • Lwerewolf 6 minutes ago
      The MI350p exists and should run a decent quant (say, the ~96GB antirez mix) well, but you can get two rtx pro 6000s for one of these, or 8x (actually more) r9700 + probably the gear to run them, etc.

      Otherwise, you can probably buy one of these second hand from somewhere (SXM A100s are available that way) and run it in an adapter board.

    • zhoutong 42 minutes ago
      It’s available on demand from a few cloud providers. Seems like the cheapest is AMD Developer Cloud (https://www.amd.com/en/developer/resources/cloud-access/amd-...) powered by Digital Ocean at $1.99/hour.

      Edit: Now I think about it, this might be the cheapest way to run the DeepSeek V4 Flash 0731 on a dedicated inference server at original weights. I haven’t run mixed load benchmarks but I guess it’s possible to generate $3-$4 worth of tokens per hour and still maintain a usable per-user throughput.

      • langs 4 minutes ago
        You need to optimize the KVCache part(save to disk to save compute) to achieve this goal.
      • WASDx 20 minutes ago
        At 830tok/s * 1 hour that's almost 3M tokens which is just $0.54 worth of tokens at Deepseeks current output price.
    • baalimago 47 minutes ago
      Give it an AI-bubble pop and these will be flooding the market.
      • _factor 41 minutes ago
        They will be instantly bought out by companies, not individuals. The consumer bubble won’t pop for quite a while yet. Production also won’t ramp up while lack of real competition keeps the demand high.
  • jkwang 33 minutes ago
    [flagged]