7 comments

  • melvinroest 41 minutes ago
    Wow, this announcement is good content marketing.

    Don't get me wrong, it's interesting. But there is no technical discussion as to how they did it. It's simply: we did it and Mythos and Codex didn't.

    It's good to know that it's possible, but I'd have already expected it. Put a base model versus a base model + harness + whatever else, and yea, if you do it right then you have a better system to find vulnerabilities.

    > We then ran AISLE's autonomous AI system against curl.

    They don't even mention what models the use under the hood. It wouldn't surprise me if they are from Anthropic and OpenAI.

    • tux3 26 minutes ago
      The homepage says something about about AI guided fuzzing based on libfuzzer or AFL. Looks like they have the LLMs identify a bunch of interesting functions to test, generate some test harnesses, and then sort through the fuzzer findings at a high level, which sounds like a pretty good idea.
      • bch 5 minutes ago
        Also sounds incredibly compute intensive.
      • melvinroest 23 minutes ago
        Thanks for figuring that out. Sort of sounds like AI programming programs to find vulnerabilities, of which fuzzing is one of the proven techniques to do it.
    • wky 15 minutes ago
      It wouldn’t surprise me if AISLE uses many different providers’ models, and what’s holding back OpenAI and Anthropic is only using first-party models. Just because OpenAI and Anthropic have arguably the strongest models overall doesn’t mean their models are the strongest at finding any given class of vulnerability or lead to follow.
    • whizzter 22 minutes ago
      Their system can run with various models, they go into more details in this article.

      https://aisle.com/blog/system-over-model-zero-day-discovery-...

    • drdrd 37 minutes ago
      > what models the use under the hood

      Presumably their own, wouldn’t they?

      • melvinroest 34 minutes ago
        You mean their own trained models, or do you think it's an open source model that they fine-tuned? If they use their own, I'd guess it's the latter.
  • rwmj 36 minutes ago
    We had a few AISLE-generated security reports, and the signal to noise was reasonably good.

    The most notable bug/exploit their scanner found was: https://gitlab.com/nbdkit/libnbd/-/commit/e50bbd2681117c2dd8...

    The tool basically had to chain two exploits together to reach this. It also came up with a patch to fix which was fairly sensible (but I ended up editing it further for clarity).

  • bluGill 19 minutes ago
    OpenAI and Anthropic have both been studying CURL for a while though. Anything they found was already fixed.

    If you want to compare you need to start with something that none of studied. Somebody please take the source to a 2023 release of CURL (It shouldn't be hard to find one) - before all the current AI craze, and run all the tools on them to see what they find. Only then can we compare numbers. (and even then severity may come into place - all 6 are rated low impact)

    • thih9 4 minutes ago
      I guess this would also require models trained on pre-2023 data - or not trained on later curl code, changelogs, blog posts discussing curl security fixes, etc.
  • guptadagger 6 minutes ago
    This is an ad. I didn't learn anything from reading it.
  • TechTechTech 35 minutes ago
    Good marketing and definitive proof that local (read: on-prem & air-gapped) models with correct context and tools are good enough to perform on par and above SOTA cloud hosted solutions.

    We have seen this point many times before with different technologies. The first computers at university were big and expensive, same as this machine. Give it a few years and this functionality will be a commodity.

  • anilgulecha 42 minutes ago
    That's bragging rights correctly earned, i think! As marketing-y as this post is, definitely something to keep an eye on.
  • Surac 20 minutes ago
    Marketing Slop