On 18 July 2026 we had a significant partial degradation. One Kubernetes worker node appeared healthy from the outside but lost access to in-cluster service addresses, including cluster DNS, so every workload on it could no longer reach internal services or datastores. Because the node hosted many workloads, critical features were badly degraded while all of our monitored signals stayed green.
The cause was an incorrect permission for the Kubernetes service proxy, which needs to list and watch node records to build internal routing rules. It was introduced on 27 May 2026 during a bulk re-apply of base cluster permissions from an out-of-date manifest set, which also disabled the automatic reconciliation that would have corrected it. The fault lay dormant for about seven weeks and triggered at roughly 06:01 UTC on 18 July, when that node's proxy restarted and could no longer rebuild its rules.
We restored the correct permission and its reconciliation, and full service returned by about 11:45 UTC. The degradation lasted just under six hours. Because everything we monitored stayed green while critical paths were impaired, we are improving our monitoring to page on-call for partial degradations, including per-node reachability of cluster DNS and services, node-level workload health, and drift in critical permissions.