How well do agents use test/verification techniques?

(danluu.com)

29 points | by vinhnx 2 hours ago

5 comments

  • gz09 2 minutes ago
    Results seem somewhat reasonable given that the amount of verus/TLA/Creusot/Lean code out there is tiny compared to all the other non-formal code.

    So it's understandable that the agents wont be able to go beyond proving trival things, given how much more difficult it is to write such code.

    A (more) interesting experiment (to me) would be to write a high level spec manually for a non-trivial system (liveness etc.) and see if the agent can produce an implementation using guided refinements that satisfies this specification.

  • vetronauta 18 minutes ago
    If I understood correctly, the author had the agent perform manual mutation testing (write code, write passing tests, manually change the code to see some tests fail, then revert the change), rather than automated mutation testing. Why not use a mutation-testing framework and consider the build failed if a certain percentage of mutations are not killed?

    The article claims that the agents didn’t actually use TDD or mutation testing. While it’s possible for an agent to ignore the TDD procedure even when instructed to follow it, it can’t ignore a build failure.

  • siscia 6 minutes ago
    It is still early, but I find that this experiment makes little to no sense and it is barely useful.

    The way you test code cannot (and should not) be decoupled by the way in which you architect the code itself.

    80%+ of effective testing is not in the testing framework but in the code architecture.

    The author doesn't mention how the code is being architected and managed.

    For what it is worth, I found that forcing agents on DI/hexagonal architecture and forcing a trivial coverage check is quite useful and produce overall good enough code with relative little effort

  • tomrod 24 minutes ago
    Timely. I'm also looking at this now, actually! We tend to throw benchmark after benchmark at systems, but miss that models are one part of the system. Harnesses are more than models and need tuning too, and in doing so there can be gains or loss of prior tested function as well.
  • kestrelquant 1 hour ago
    [flagged]