Qwen 3.8 follows GPT-5.5 Pro reasoning prefills

(gist.github.com)

63 points | by wsxiaoys 1 hour ago

5 comments

  • wongarsu 1 hour ago
    That writing style might be a tad too tense

    If I got it correct (appending B from https://stolen-thoughts.com/paper.pdf is essential) they are the authors of the well-known exploit to recover readable CoT from OpenAI and Anthropic models. They use that to find hints of distillation, by running a benchmark with a SotA model, recovering the CoT, then taking the first 1% of the CoT and running the open-source model as if that was the start of its own CoT. In the paper they found that Kimi-K3 gets a lot closer to Claude 4.8 answers when prefilled with the start of Claude 4.8 reasoning, suggesting that Claude 4.8 was used in its post-training. This blog post is the follow-up with results that suggest that Qwen3.8 was post-trained with the help of GPT-5.5 Pro (or some similarly responding GPT model, it's unclear how many models they tested)

  • 7734128 1 hour ago
    The problem with this is obviously that the only GPT 5.5 thoughts that we have access to are from stolen thought.

    Qwen 3.8 0902 was trained after the release of the paper on August 10, so it should have seen those specific thoughts.

    • usernomdeguerre 49 minutes ago
      seems like only the companies in question could run this sort analysis long-term; since they have full access to their CoTs not in public datasets.
      • verdverm 20 minutes ago
        and we have to "trust them bro" to be fair and accurate, something I am very unlikely to do given their other false / misleading statements to date
        • sureMan6 13 minutes ago
          And the only end result would be that the Chinese trained on their data just like OAI and Anthropic trained on our data so who cares
          • verdverm 3 minutes ago
            capitalism ensures I get high marx on my Ai bill
    • refulgentis 48 minutes ago
      The thoughts trick was known before their paper / August.

      I "independently" "invented" it for the first Anthropic reasoning models because the API required you have thoughts for each assistant message. My app lets you switch AIs within a chat, and their API used to require thinking for all messages if thinking was enabled, so I needed to get a valid thinking stub to insert.

      Time has flew by for me the last 3 years, but, I'd guess it's been at least 18 months. And IMHO it wasn't very complicated to work through how to do once you were dead set on making it happen. I expect it was well-known to distillers before the paper.

  • brcmthrowaway 2 minutes ago
    This makes me very sad

    If Qwen and other Chinese labs are just copying reasoning traces, then those labs are more than a year behind the frontier.

  • jari_mustonen 49 minutes ago
    > Qwen barely moved toward Opus 4.8 in the earlier experiment, but moved by +20.58 points toward GPT-5.5 Pro here, including a large effect on the private synthetic puzzles. The data suggest that Qwen may have learned from GPT-5.5 Pro, or from a closely related GPT model, rather than from Opus.

    How does this suggest anyting of the sorts?

  • CamperBob2 45 minutes ago
    News flash: people who scraped the Internet without permission to build their product complain when something vaguely similar is done to them. Water still wet, sky still blue. Film at 11.
    • RivieraKid 2 minutes ago
      For some reason, what China is doing seems worse. Part of it is that I want the US to stay ahead of China.
    • verdverm 18 minutes ago
      I would not be surprised in the slightest if we later find out they are running those same open models to find useful traces or bits to incorporate into their own training. Lots of rules for thee but not for me from Big Ai

      I look forward to a day when open models are so dominant that we stop considering traces to be some form of intellectual property that must be hidden from / manipulated for paying users.

      It's that manipulation of inputs and outputs that really rubs me the wrong way

    • noir_lord 27 minutes ago
      I'd send them the worlds smallest violin but Rufus is getting in the way of me finding it.
    • vlyan 15 minutes ago
      not just the internet, but every commercially published written work in existence, and I doubt their highly publicized destructive scanning thing had managed to legitimize even a fraction of a percent.

      this what is permissible for Jupiter is not permissible for a cow bullshit alone should tell people all they need to know about what kind of greasy sociopaths run "open"ai and (mis)anthropic, and how seriously you should take their purported stances on "safety" and other self-serving shit.