Our infrastructure provider carried out emergency Linux security maintenance overnight, which forcibly restarted several production servers. This caused four separate incidents and woke Hampus, our on-call engineer, four times throughout the night. The final incident was the most severe, making our entire Kubernetes cluster unreachable and significantly affecting Fluxer's availability.
Recovery was made harder when several services failed to restart cleanly after the forced reboots. Hampus worked through the night to restore the cluster, recover the affected services, and remove connection limits exposed during the restart process.
Everything is now operational. We have improved automatic service recovery and increased connection capacity. We are sorry for the disruption, and grateful for everyone's patience while a very tired human worked to put out a series of infrastructure fires.