No, we don't literally melt servers, but remember yesterday we promised to explain why we are doing the Melt-the-Servers Event? If you missed it https://www.worldsadrift.com/blog/melting-servers/
Today we have Tristan explain what and why it's required! And why YOU should join us on Thursday!!
So MMOs run on servers! Who knew? Well I guess most people do, but what you might not know is how we go about determining how many of those servers we need to run the game.
Given an infinite amount of money we could just buy a supercomputer (maybe IBM’s Watson?) and call it a day right? Unfortunately not, well how about multiple Watsons? Still nope. The reasons for this are many and until we stress test the game and get a good understanding of its performance characteristics, we would just be guessing what hardware would be the most suitable and not actually going to cost us an arm and a leg.
Stress testing is the process of putting servers under load, to be able to determine where the game is likely to be bottlenecked first. Stress testing can be done in many different ways, such as through scripts designed to query the servers over and over again, through bots written to simulate players playing the game or through lots of actual players logging in and playing the game. For Worlds Adrift we need to do test using all these methods as none of them in isolation gives us a full view over how the game performs across a range of key metrics.
It’s using these metrics that we are able to determine how our game scales and how best to provision hardware to run it the most efficiently.
When running the tests we are looking at a ton of different metrics, from CPU, memory, hard disk usage to bandwidth, latency and queued messages in the system. These metrics give us an idea of where our bottlenecks may lie and what we need to do in terms of scaling hardware, re-engineer systems or limiting concurrent players to alleviate these bottlenecks. Going back to our Watson analogy, just throwing any old hardware at the problem is probably not going to help, as it might be bottlenecked by the raw bandwidth a single server can handle or even the overhead of managing multiple supercomputers. Not everything scales well together and its understanding what doesn’t scale well that helps us determine what hardware we need and what systems need rethinking.
While the tests are ongoing we are watching with eagle-eyes lots of graphs, logs and player reports to keep tabs on what metrics are scaling well with the amount of load and which aren’t (hence we are doing this on a Thursday evening). And if all is going well we are adapting and finding the perfect combination of software engineering and raw hardware to ensure that on the day of release we know to the best of our ability what hardware we need to keep up with the requirements of many, many people playing the game all at once.
There is a lot involved in stress testing a game, finding the perfect balance of hardware, software and internet connection to ensure that we can handle all the load you guys could possibly throw at us, and make sure the servers aren't going to crash or you are suffering from severe lag when playing the game. I have only really touched the surface here but when it really comes down to it, stress testing is about running the servers as close to capacity as possible simulating the type of behaviour real players would have on the server in our worst case scenarios (ie. everyone ever all wanting to play the game all at once!). If we are successful in our testing, not even an armada of ships chasing down a herd Thuntomites and Mantas can stop us!
Well that’s all for now!
Tristan - Lead Developer on Worlds Adrift
See you in the skies!