All good

4 playing

Everything on record

Europe

Scheduled restart

Europe shut down and came back. Not one of the daily restarts.

Twenty-three people were playing on Europe when the server stopped dead for nine seconds, and it was restarted a minute after that. Both were us, deliberately, while looking for a memory fault. Two people were dropped by the freeze and everyone was dropped by the restart. Nothing anybody had done was lost.

Started
29 August, 20:24:30 UTC
Ended
29 August, 20:24:40 UTC
How long
10s

The tick readings around this were not kept. The monitor holds its raw samples for about two days, and the record keeps the readings for its most recent entries only - so this far back, what happened is on record and what it looked like is not.

About this restart · written 29 August

How it went

  1. 20:23:17 UTC

    We ask the server for a complete copy of what it is holding in memory. Writing that copy stops the game while it happens - there is no version of it that does not.

  2. 20:23:28 UTC

    The server records the pause itself: one moment of the game took 9.4 seconds, where a normal one takes a twentieth of a second. To anyone playing, the world simply stopped.

  3. 20:23:34 UTC

    The copy is finished and the game runs on. Two of the twenty-three players had dropped out in the meantime; the rest carried on. The server never actually went down.

  4. 20:24:18 UTC

    The server is shut down by hand, saving every world and every profile first.

  5. 20:24:38 UTC

    Back up and taking players, 9.2 seconds later.

Why we did it

Europe had been slowly filling up with things it should have thrown away. Every time somebody finished a fight against a practice bot, the bot was cleaned up in the world but one leftover of it stayed in memory - exactly one, every single time, all day, a few thousand of them by the evening. Left alone that ends one way: the server gets slower, and then it needs a restart it did not choose.

We could count them. We could not see what was keeping them, and that is the part you have to know before you can fix anything - otherwise you are guessing at code that looks reasonable. The only thing that shows it is a full copy of the server’s memory, and taking one means stopping the server for as long as it takes to write.

Why on a live evening

Because it does not happen anywhere else. We ran the same fights on a copy of the server we keep for testing and the leftovers did not appear - so the evidence only exists on the real one, with real players on it.

We picked the moment rather than the time of day: nine seconds is short enough that most people read it as a hitch, and it was over before the server’s own watchdog had anything to say. We would rather spend nine seconds now than an unplanned restart later, in the middle of something that matters to you.

The restart

That was us too, a minute later, and it is the one part of this that fixed something: the pile-up lives in memory, so starting fresh empties it. The shutdown was an orderly one - every world and every profile written to disk first - which is why the restart cost nine seconds and nothing else.

The fault behind it is not fixed yet. The copy is what tells us where to look, and that work is going on now. When it lands, this stops being something we have to restart our way out of.

About your progress

Nothing was lost. During the freeze the game was not playing, so there was nothing happening to save: no hits landed, no blocks placed, no items picked up. Everything before it had already been written in the normal way, and the restart wrote everything again before it closed.

If you were dropped by either, rejoining was the whole of the fix. Nobody needed to do anything else.