Hello again. As I mentioned in my last blog, the majority of work on the game right now is going into the AI. The goal is to make the best possible AI solution that provides the player with a real and repeatable challenge.
A decision based AI has already been written and tested. By and large the goal with it is to not do anything stupid. For example, if it’s standing on a victory location where capturing would win the game then it should do this. The AI is made up of lots of little rules like this to help it make the decision. In this sense it’s not too far from Undaunted Reinforcements. In fact I did consider basing the AI on the Undaunted Reinforcements one but too many of the game elements are changed by this expansion, and I also wanted an AI that played according to the basic game rules.
Moving on to the Machine Learning (ML) AI then. It uses a branch of ML called reinforcement learning. The idea is that it constantly plays against a version of itself that is just a little bit better or worse than the current iteration it’s testing. When it wins a game it gets a points reward, and through this it aims to achieve a version that achieve this as often as possible.
In order to allow the AI to understand how it archived the reward, it needs to understand the game state at each decision point. This means that you have to provide it with the information that it’s going to use in order to make this. If you’re playing Undaunted Normandy at home, you’ll look at your cards, look at the board, and make a decision about your next move. What you’re actually doing is considering a number of variables relating to the game state and then making your decision based on this information.
The ML AI solution tries to achieve the exact same thing. The key question is, “what is the minimum set of information I need to make a decision about my next move”. In ML parlance this is called collecting observations. And you want to use 1s and 0s wherever possible to simplify the computation required. The answer I arrived at for this question was the following: 5 indicators, one for each possible card (this is just the first scenario remember); the number of points needed for the AI player to win; the number of points needed for the opposition player to win; the number of AI riflemen cards in the deck (for pinning rules); the number of enemy riflemen cards in the deck (ditto); and for each tile on the board information relating to its x and y position, defence value, tile victory points, whether it’s scouted or controlled, and whether each of the four unit types in the scenario is in the tile.
That’s a total of 213 observations! Think about that the next time you’re taking a turn in Undaunted Normandy. It’s remarkable how much information the human brain can take in without realising it. Subsequent scenarios will be more complicated - more units, more tiles, etc. Oh, and throw the random numbers that are generated by combat into this equation.
The AI ML then starts up and running. At time of writing it’s currently made over 2.6 million decisions and counting. It plays one game of Undaunted Normandy every 20 seconds. If I had to guess I’d say that so far it’s played around 85,000 games (unfortunately it crashes occasionally when the AI doesn’t get a response, so I can’t leave it running constantly).
And after this it’s still taking some pretty stupid moves. There could be two reasons for this: either (a) it’s just a slow process and with that much information is just going to take time, or (b) the ML configuration isn’t quite right somewhere and needs tweaking. I suppose a third option could be that the random factor of combat is messing with things a bit as well. My hope is that this evens itself out over time.
In a typical data science ML environment (the type Netflix might use for their recommendation engine) a data scientist will build a model, test it, change it, test it, etc until it’s just right. But that is a more scalable working environment, so it’s easier to iterate. Unfortunately I can only have one version of the game running at one time on my dev laptop, so it’s not a scalable environment, so it just takes longer.
For now then I’m sticking with the current solution, and hopefully brute force will get the job done.
In the meantime of course the release date is getting nearer and nearer. I think what this means is that the game will go up in Early Access with the decision based AI available on all scenarios, the ML AI available as an experimental option for the first scenario (expect it to do some stupid things). And multiplayer of course. More on that next time.
That’s all for this time. If you’ve learnt something in all of this then my work here is done for now!
