7 comments

  • _matthew_ 6 minutes ago
    I don't think it makes sense to have the prompt be that short. This is basically a bench.ark of how models interpret an overly vague prompt. It should at least be "Create a pacman clone in a single html page. Make it faithful to the original" if that's what we're scoring it on.
  • hoistway 11 minutes ago
    Always assumed Pac-Man was an easy solve for modern AI. Guess those ghost patterns are trickier than they look for one-shot learning.
  • nedo_var 16 minutes ago
    Curious if models grasp the ghost patterns or just react. Pac-Man's more complex than it seems for one-shot learning.
  • strataspace 4 hours ago
    I tried this with DOOM. Fable 5 did a pretty shit job. Astra made pretty crazy animated sprites and was pretty good considering.

    The fact that these are at all playable and 100x my programming skill level is pretty depressing from a certain pov. The ThreeJS dude posted ab how demotivated he was to continue his work, and while I was never a dev that did much with webgl, I commiserate.

    • acomjean 50 minutes ago
      I wonder if we need better programming abstractions/ languages that can make programming easier for people.

      It seems like a lot of the programming is cookie cutter type stuff (where these ai programs shine) and should be easier.

      They can be amazing, but promoting “make a pac man game” seems like an exercise in which model has the best training code that was Pac-Man.

  • blindflag 6 minutes ago
    I'm curious; can you explain why you picked Pac-Man, in particular?
  • dang 58 minutes ago
  • Computer0 6 hours ago
    Opus 5-5 seemed like a perfect clone, with others displaying flaws in an initial look. Astra notably created a bunch of surrounding ugly crap to look at.
    • vunderba 6 hours ago
      I was just coming here to say this. I looked at all of them, and Opus 5.5 at high effort seems significantly better than every other clone.

      Buffered controls, different pathfinding AI for each ghost, level transitions, etc.

      Gpt-5.6 sol high technically completed the assignment as well, but the gaps within the pipes used to construct the maze, the lack of pause when eating a ghost, etc made it feel far less polished.