Europe
Unplanned restart
Europe shut down and came back. Not one of the daily restarts.
Anyone trying to join sat on the loading screen until it timed out, and anyone already playing was disconnected. It ran through the middle of the afternoon and took a restart to clear.
- Started
- 28 August, 14:22:28 UTC
- Ended
- 28 August, 14:22:48 UTC
- How long
- 20s
The tick readings around this were not kept. The monitor holds its raw samples for about two days, and the record keeps the readings for its most recent entries only - so this far back, what happened is on record and what it looked like is not.
Incident report · written 28 August
How it went
13:03:50 UTC
The last reading arrives. The server is running normally, at 20 TPS.
13:04:09 UTC
The server warns itself that it has stopped responding, and goes on warning every five seconds for the next hour. Nothing after this point is played or logged in: joins reach the loading screen and sit there.
14:22:32 UTC
The server is restarted by hand, and the world starts loading.
14:22:41 UTC
The world finishes loading in 8.7 seconds. Players can join again.
14:22:51 UTC
Readings resume, at 20 TPS.
What went wrong
A player joining is the moment the server does the most work about one person at once, and part of that work - working out what a player is allowed to do - was being done on two threads at the same time. Two threads writing to the same list left that list in a state that cannot be read to the end, and the loop reading it never finished.
That loop runs on the one thread that runs everything: the world, the mobs, every hit and block placed, twenty times a second. It stopped, and the whole game stopped with it. The server was still running - accepting connections, using a full core, answering pings - and it never played another tick.
Why it went on so long
Minecraft servers ship with a watchdog for exactly this, and it fired within seconds. It could not end it. The watchdog is written to stop the server when the game loop stops answering, and stopping the server is itself work for the game loop - so it queued its own shutdown behind the loop it was trying to rescue and waited there with everybody else.
Our monitoring did not catch it either, for a duller reason: it was watching for a server that had gone quiet or gone away, and this one had done neither. It was up, it was busy, and the numbers it had last sent were healthy. The page said the server was offline, which was true and read to a player as "back in a minute".
What we changed
The fault itself is fixed: the permission work now happens on the same thread as the rest of a join, in the order the rest of the server expects.
The monitoring now watches for this shape specifically - a server that has stopped reporting while its machine is still burning CPU, which is the difference between a server that has crashed and one that is stuck. It alerts about a minute in, at any hour.
This page changed too. An outage that is still going now appears in the record while it is happening instead of only once it ends, and when a server is stuck rather than down it says so, in the words that matter to somebody deciding whether to wait: waiting will not fix this one.
About your progress
Nothing was lost to this. The loop that plays the game had stopped, so from 13:04 onwards nothing happened on the server to save: no blocks placed, no kills, no items picked up. Everything up to that point had already been written to disk in the normal way, and it was all there when the server came back.