Work for Undaunted Normandy has been continuing for the last few months. At the moment most of the work has been going into the AI.
One of the most important components in any strategy game is the AI strength. The graphics can be the most beautiful ever created but if the AI opponent plays badly then, outside of multiplayer, there’s very little challenge for the player.
The original plan was to make a decision engine solution. Each turn the AI would consider the board position and make a move according to a set of hard and fast rules defined for it. For example, if the opponent is on a tile where it could win, either move on to that tile or shoot at it. In the most part this works ok, and this is currently the fully implemented AI. You can find a video of me playing against the AI here (https://www.youtube.com/watch?v=CyIyBfLqQ0o) and you’ll see it makes fairly sensible decisions all of the time.
I’ll admit now that in online playtesting I’ve discovered that I’m not the strongest player of Undaunted Normandy. I’ve played a lot of games in every scenario, and I lose as many games as I win. Probably more. However, against the decision engine AI I’m winning most of the time, so although the engine plays a good game, it could still improve.
Having given this a lot of thought, and with a few false starts I’ve decided that the solution to this is potentially to use a machine learning (ML) AI solution. This means setting up a training environment where the AI plays against a closely matched version of itself, gradually improving as it progresses and learning from what it’s doing. The best analogy I read was that it’s a bit like a tennis player. If you play tennis against Roger Federer, then you’ll lose heavily and won’t have learnt anything. Similarly if you play tennis against a 5 year old child you’ll win easily but also won’t have learnt anything. The best way to improve is to play against someone just a bit better than you.
You then have to consider how you’re going to configure the ML environment. The best approach is to give the AI one point for winning, and one point for losing, and then let it just go and figure the rest out for itself. In order to encourage it to get a bit of a move on you deduct a very small fraction of a point each turn. You also need to be able to provide the ML solution with enough information about what is on the board at any one point in time in order to make a decision. You provide it with a list of valid orders for a card and then it’ll pick one and perform it.
Having set this up, the next step is to provide gradual training. All of this has taken place in Scenario one in the game (La Raye) since I still want to prove that it’ll work, and it’s the simplest scenario – two types of units, standard victory conditions.
In the first instance, I just filled up all the tiles with scouted markers and victory points. The AI learnt pretty quickly that it could win by picking up the victory points, but it gave it the right idea. The second step was then to remove the victory points so that they are in random points on the board (in more or less every other space). This means that the AI has to go looking for the victory tiles. The final step was to remove the scout markers that weren’t already defined at the scenario, so that the ML solution could learn how to move scouts.
I did then make one final step, which was to make the AI play the defined scenario over and over again. However, this caused it to focus entirely on the tile in the bottom left hand corner of the map. The reason for this is (I think) that in the first scenario this is the tile that most often decides who wins the game. However, it started to create quite inflexible behaviour, so I’ve gone back to the previous step to try and keep the AI as flexible as possible (which should benefit subsequent scenarios).
At this point the final challenge with the ML solution is to get it to play thousands of games. At the moment it’s running at 20x normal speed, and can play one game every 20-30 seconds. The eventual intention is that you the player will then be able to play against an opponent who has played thousands, or maybe even hundreds of thousands of games, and this will create a proper challenge.
The risk is that we spend weeks and weeks trying to get a ML solution working and it just doesn’t learn properly, but I think the reward is definitely worth it if it does in that the player will be properly challenged when playing in single player mode.
