All good

4 playing

Everything on record

Europe

Offline

Europe could not stay up: it went down and came back 4 times. Nobody could rely on it for that stretch.

Twenty-six people were playing when Europe stopped responding, and it then went down and came back four times over eleven minutes. Everyone on it was dropped each time, and for one stretch of just under six minutes nobody could get in at all. It has been steady since 13:47.

Started
31 August, 13:35:47 UTC
Ended
31 August, 13:47:39 UTC
How long
11m 52s

Around it

13:32 to 13:40 UTC

The slowest 5% of updates, 10 seconds to a point, against a ceiling of 50ms. Typically 2.9ms through this window. The dashed line is when it started; point at the chart to read any part of it.

Incident report · written 31 August

How it went

  1. 13:35:47 UTC

    Twenty-six people playing, and every reading normal.

  2. 13:36:24 UTC

    The server stops responding and says so about itself. Anyone playing is dropped at this point.

  3. 13:36:48 UTC

    It has restarted itself and is taking players again, about seventy seconds after it went. Five of the twenty-six are back within the minute.

  4. 13:40:19 UTC

    It falls over again. Then again at 13:46, and again at 13:47. The longest stretch with nobody able to join is five minutes and forty-five seconds.

  5. 13:47:32 UTC

    Back up, and this time it stays up.

  6. 13:54:31 UTC

    Eighteen people playing again, seven minutes later.

What went wrong

The server had been running right at the edge of the memory it is allowed, for at least half an hour before it fell over, through the busiest part of the day. Past that edge it spends more and more of its time tidying up after itself and less of it running the game, and there is a point where it stops answering at all. That is what happened at 13:36.

It is worth saying what did not happen, because a run of restarts usually means somebody pushed something. Nothing was being changed or deployed. The server was running exactly what it had been running all morning.

Why four times and not once

Restarting is automatic, which is why each gap was about a minute rather than however long it takes somebody to notice and act. But a restart does not change what caused it. Each time the server came back it came back to the same busy afternoon, filled up again, and went the same way.

The fourth one held, and the uncomfortable part is why: by then far fewer people were on it. It steadied because it had emptied out, not because anything had been fixed.

Where this stands

We know what it ran out of, and what has been using it was already being looked at before today. That work is not finished, and we are not going to pretend a restart was a fix. Until it is done, a busy afternoon can do this again.

What has changed is that we can see it. This page recorded all four restarts as they happened, with the times, which is the reason there is anything to write up at all.

About your progress

Your progress is written to disk about once a minute. These were hard stops rather than orderly ones, so the most that can have been lost is whatever happened in the last minute before each fall - a kill or two, not a session. Everything older than that was already saved.

If you were dropped, rejoining was the whole of it. Nothing needed doing on your side.