Timeline of the incident on Saturday, 19th April 2025
at 10:16 number of connections between internal relay services dropped to zero (0), causing the internal relay egress rate also drop to zero (0);
at 10:18 memory of all internal relay servers was at maximum capacity, causing swapping and server lock-up;
at 10:22 message dispatching halted;
at 10:27 internal relay idling alert was triggered;
at 10:34 investigation into alert was started and incident created;
at 10:53 first internal relay server was restarted, resuming delayed message delivery;
at 11:17 last of deferred bulk messages were delivered;
at 11:23 the issue was considered resolved.
What caused the incident?
The incident was caused by our internal relay, which is configured to consume up to 75% of server's physical memory. Under normal conditions this wouldn't be an issue, but as the number of connections between relays dropped to zero (0) the relay service started to consume multiple times more memory before it could start shedding the load and shrinking the queue. This caused the server to start swapping, increasing CPU load and in turn locked up the server.
How we address the issue?
We configured relay services to start load shedding when 60% of server's physical memory is consumed. We believe it gives enough headroom to services to start shrinking the queue before the server's memory is fully consumed without sacrificing performance.