12 comments

  • dmurray 1 hour ago
    Oh no! Stratego had been on my mind as something we just hadn't tried hard enough to make a winning bot for, including the DeepMind effort from 2022. I was planning to make the first one.

    I thought this was slightly less crank-coded than trying to prove the Riemann Hypothesis, but maybe these days you just ask Claude to do that and it tells you there's a counterexample at 1 + πi that no one ever noticed before.

    • bananaflag 14 minutes ago
      I wonder why people act as if Riemann hypothesis is somehow already solved, when it is still highly probable that it is impossible for humans and slightly less impossible for AI overminds.

      (Of course, tomorrow Google might announce that it has solved it.)

    • NooneAtAll3 39 minutes ago
      recently we got Starcraft 1 RL bot and Advance Wars bot, both on upper human level
      • root_axis 9 minutes ago
        The frontier is moving, but there's still no human-level starcraft bot that plays entirely through vision, like a human.

        IMO it also needs to use a real mouse before I think it's a true comparison, even a casual player would have a massive advantage if they could issue selections and unit commands via query.

      • gregdeon 38 minutes ago
        We got an Advance Wars bot? Say more...
  • hnedeotes 59 minutes ago
    I think that what makes these games beatable repeatedly is that they're static. Not saying an algorithm properly trained won't play better than the average player a game like MtG, or my own https://aethersummon.com (specially now while it has under 90 possible scrolls only) but if you have a regular release cadence (say weekly or bi-weekly) of relevant new "cards", then I think the playing field is much more even for humans.

    Those new additions can invalidate the whole training data by a single new "card" that changes completely the dynamics and would be easy for a player to understand and incorporate but not for an algorithm (perhaps with enough compute to re-train it regularly it could) - that along with the decision trees being orders of magnitude deeper, wider and with more conditionalities than go, chess or stratego - even through the same turn with the same cards available and same table state - would probably pose much harder problems for a compute bound algo.

    • Arainach 37 minutes ago
      > Those new additions can invalidate the whole training data by a single new "card" that changes completely the dynamics

      This doesn't follow. You're basically proposing that new combo decks be added all the time, and it's far simpler for an agent to scan the new cards for potential interactions with the thousands of other cards in circulation than for a human to remember all of them.

      Your analogy is akin to saying that all you have to do is keep landing new code all the time, and since the agents weren't trained on the code they won't be able to identify and respond to security vulnerabilities in it as fast as humans, which hasn't turned out to be correct

    • qsort 33 minutes ago
      There are very few missing pieces for a game like MTG. The main reasons we don't have a Stockfish for MTG is that it's a PITA to implement the rules and that nobody cares (or at least not enough to make it happen.)

      There is nothing that, in principle, makes MTG different from poker or bridge, and we have superhuman engines for both.

      • hnedeotes 10 minutes ago
        MTG is also severely constrained (small hand, mana -> possible moves) although I don't think it's anywhere near the same. In my opinion the rules are effectively what change the whole dynamics. You can't plan as efficiently without knowing what your opponent holds and having to take into account all possibilities (with infinite energy/compute time perhaps)... I don't doubt you can train a model to play well, I just think it should be much more level to the human player. In MtG you also have the randomness which is not easy to model nor account for - the perfect play by an LLM can be the worse once the opponet draws next.

        In my own game you don't have shuffle/draw randomness but the pool of options is statistically tending to infinite (if I would have 500 or 1000 scrolls designed and MtG depending on the format has that depth) when compared to something like chess, or this game. On the other hand in my own game you have to account for much more depth on the possible options your opponent has.

  • rovr138 58 minutes ago
    > Now, a team of researchers from Carnegie Mellon, MIT, New York University, and Stanford University has done it. Their AI, called Ataraxos, beat Pim Niemeijer, arguably the best Stratego player of all time, 15 games to one, with four draws. And it took just 16 GPUs and a few thousand dollars to train it.

    Just 16 GPUs, and a few thousand dollars?

    What about “researchers from Carnegie Mellon, MIT, New York University, and Stanford University” this wasn’t just anyone.

  • SwellJoe 25 minutes ago
    I loved Stratego so much as a kid. But, I eventually couldn't find anyone to play with me because I crushed everyone, including my dad who was much better than me at chess. But, I never would have thought it'd be a game that models would have a hard time with. It feels relatively simple. And, inexplicably, I never even thought that there might be serious players...I've kept the board game, one of the very few things I have from young childhood, but it's been twenty years since I played. I guess it's time to find an online Stratego. Surely someone in the whole world can beat me.

    Too bad I never played against an AI before they cracked it.

  • smokel 1 hour ago
    This approach also works for Hanabi, which is a very interesting game. You can't see your own cards, but the other players can. I bought the game because someone on a reinforcement learning podcast [2] mentioned it, and actually played it multiple times.

    [1] https://en.wikipedia.org/wiki/Hanabi_(card_game)

    [2] https://www.talkrl.com/episodes/jakob-foerster

  • gritzko 1 hour ago
    I recall playing this game as a preschooler. It was mostly psychology and bluff. Very interesting.
    • changoplatanero 1 hour ago
      yes back then the game was hard partly because you couldn't remember all of your opponents pieces that you had seen. An AI would never forget though.
      • cyanydeez 1 hour ago
        eventually though, resimulations will never to revolve around trying to inject bad context into the the stream and try to distract them.
  • smokel 1 hour ago
    This puts the earlier "Mastering the Game of Stratego with Model-Free Multiagent Reinforcement Learning", 2022 [1] in some perspective. Apparently the "mastering" in 2022 wasn't quite there yet. Four years later, the new approach seems to actually be better than humans.

    [1] https://arxiv.org/abs/2206.15378

  • wakamoleguy 55 minutes ago
    Looks like a game archive exists here: https://ataraxosai.github.io/
  • osti 1 hour ago
    Wait, wasn't there that strong stratego bot that came out from deepmind in 2022?
  • gavinlilly 1 hour ago
    I wonder how capable current AIs are with the "silent defense" variant of Stratego [1]? The article states the high level of uncertainty presents a challenge. With silent defense the uncertainty is even higher.

    [1] https://www.hasbro.com/common/instruct/Stratego.PDF "When an attack is made, the attacker is the only player who has to declare the number of his or her piece. The defender does not reveal the number of his or her piece, but resolves the attack by removing whatever piece has a lower number from the gameboard. Players keep their own captured pieces. Exception: when a Scout attacks, the defender must reveal the number of his or her piece.

    • janzer 45 minutes ago
      As a kid, a friend of mine had "Electronic Stratego"[1], the biggest gameplay change was that you could carry out fights without revealing the strength of either piece to the other side. I found this made for a much more interesting game and we had quite a bit of fun playing it.

      1. https://boardgamegeek.com/boardgame/3513/electronic-stratego (We generally banned the use of the 'probing' feature)