17 comments

  • SomeHacker44 2 minutes ago
    > DeepSeek V4 Flash 0731 (Reasoning, Max Effort) is amongst the leading models in intelligence and well priced when comparing to other models of similar price.

    Similar price? Doesn't make sense. Maybe they meant power, capability or speed?

  • throwaw12 1 hour ago
    If deepseek v4 flash is beating DeepSeek V4 Pro, can we expect new V4 Pro which is on par with Opus 5 in couple weeks (even better if it beats Opus)?
    • websap 11 minutes ago
      I’m salivating at the thought of this!

      Deepseek v4 Pro prices with Opus 5 perf would be freaking unbelievable!!

      This is probably a dream.

    • jmathai 28 minutes ago
      I’ve been using v4 flash for an app I’m building [1] and it’s amazing how cost effective and good it is coming from having always used gpt, opus and sonnet models.

      It’s so cost effective I can offer a generous free tier since my goal isn’t to make money with it.

      [1] https://trysojourn.app

  • monooso 2 hours ago
    404. I believe this is the correct URL:

    https://artificialanalysis.ai/models/deepseek-v4-flash

    • theanonymousone 1 hour ago
      Yes, sorry, I went into anti-procrastination mode after I posted. I hope someone fixes it.
  • WithinReason 2 hours ago
    Already beat Luna on price/task, by about 2x:

    https://artificialanalysis.ai/models/deepseek-v4-flash?intel...

    • spwa4 1 hour ago
      Maybe I'm reading that incorrectly, but it seems to me the cost is on the X-axis.

      First, your direct comparison, Deepseek V4 Flash 0731 (max effort) $0.03 (rounded up) per task @ index 50.

      OpenAI Luna:

      * high effort $0.03 (rounded down) @ index 46

      * xhigh effort $0.04 @ index 49

      * max effort $0.07 @ index 51

      So I would say a fair statement would be "OpenAI Luna between 2x and 3x the price of Deepseek Flash, what you get is 2 to 5 times faster inference"

      The cheapest OpenAI model that beats it is OpenAI Luna (max effort) $0.07 @ index 51 (if you take the rounding out it summarizes to triple the price for similar performance), but still close to 3x faster.

      And can SOMEONE please tell artificialanalysis that using dark blue for both Deepseek AND OpenAI is an especially unfortunate choice of colors, especially today?

      • andai 48 minutes ago
        For anything substantial, you'd want a bigger model anyway.

        For simple tasks, they're already saturated, and you'd prefer the faster model, so that you can have a realtime/interactive-ish experience.

        Or to put it bluntly, it's cheaper if you don't value your time. That goes for smaller models in general -- need more handholding, more correcting -- but the Chinese ones are slower on top of that.

        As for speed, Sol on Low is faster than Luna on most settings.

  • baalimago 51 minutes ago
    New Deepseek models are like Christmas for me. Really big fan of low cost API models, noone does it better than DS. Until VRAM price is low enough to run models locally, this is the way to go.

    The subsidized subscription model won't last, API pricing "feels" closer to a true sustainable business model.

  • freakynit 3 minutes ago
    I will get downvoted, but fck it.

    The ban on these open models is coming within weeks, if not days. As usual, the excuse will be "national security".

  • epolanski 8 minutes ago
    I was writing a benchmark for my own harness, and DS4 flash answers as well as Fable 5 on any query.

    The specific agent is focused on getting precise and on point answers about a codebase.

    The starting point was nowhere near. E.g. asked why was X implemented in a certain way it would give bogus answers when the real answer was that there was no reason at all.

    The benchmark included more than 50 questions or different difficulty.

    But when the agent was improved in its prompt and rooting it was impossible to have it perform worse than closed source sota.

    Just to say that the quality of the harness is as important as agents intelligence.

  • embedding-shape 2 hours ago
    Is the "Output Tokens per Intelligence Index Task" data actually correct or am I reading it wrong? It says there that "Kimi K3 (Max)" would think/reason less than than deepseek-v4-flash, and a whole bunch of other models, like less than hy3 and even gpt-oss-120b, but in my experience, K3 is probably the model that thinks/reasons the longest of all of these.

    Am I just using it on tasks that makes it go on forever vs these benchmarks that are short&sweet, or something like that? I've been throwing bunch of identical prompts at different models at the same time, and when comparing hy3 and K3 I've never once had K3 reason less than hy3, as just one anecdotal data point.

  • paoliniluis 39 minutes ago
    Would be awesome to see a new ds4 release. Having so much in something that can be run locally is mind blowing
  • NooneAtAll3 13 minutes ago
    what a horribly heavy and resource-consuming website...
  • k1e 54 minutes ago
    No speed (tokens/s) benchmarks?
    • k__ 48 minutes ago
      On OpenRouter it's 93 TPS.
  • qtalen 1 hour ago
    Unfortunately, DeepSeek Flash still doesn’t support multimodal; otherwise, it would offer better value than GPT 5.6 SOL.
  • try-working 23 minutes ago
    Now let's see Dario's price cut.
  • lostmsu 2 hours ago
  • madikz 1 hour ago
    [flagged]
  • BedVibe_Studios 1 hour ago
    [flagged]
  • yonisto 1 hour ago
    Does it already know the answer to what happen at Tiananmen Square? Or still avoiding it?
    • xbmcuser 52 minutes ago
      Who cares if it is programming correctly I would be more worried about it not doing things like find security bugs because US or Chinese government does not want to. Which LLM is more likely to do that?
    • edot 36 minutes ago
      It’s open weight, you can (or you can wait for someone else to) uncensor it. We shouldn’t be upset at the researchers making this for the mandates their government puts on them.
    • avazhi 20 minutes ago
      Western models censor just as much shit as the Chinese models do, big guy, it’s just different material. While we should be pushing for universal fully uncensored models, this comment is lazy and trite at this point.

      But you already know that.

    • hiherer 16 minutes ago
      [dead]