On July 14, 2026, at 10:57 UTC, Akamai identified an increase in 502 errors and latency affecting customers using the Linode API, CLI, and Cloud Manager. This disruption resulted in moderate service impact, with customers reporting elevated error rates. Our initial investigation traced the issue to latency with IAM services, which was resolved, but elevated errors persisted.
Further analysis by relevant subject matter experts determined that the incident was triggered by a manual failback to the Cloud IAM primary load balancer from the secondary load balancer. This action was prompted by a warning alert indicating that the secondary load balancer was acting as the keepalived master. The manual process of starting and stopping services to initiate the failback differed from the automated process and led to a cascade of stale GRPC connections, causing increased latency and API errors. Restarting the API servers cleared the stale connections and restored normal operations. Customer impact was mitigated by approximately 13:10 UTC on July 14, 2026.
To prevent recurrence, Akamai is investigating why the manual failback caused this behavior. The team is considering implementing a drain command to clear GRPC connections during failover and setting up alerts to detect stale connections for proactive intervention. However, the immediate focus remains on understanding the root cause, with alerting and automation planned for later phases.
Several customers have confirmed resolution across their deployments. Akamai will continue to monitor system health and await additional customer feedback before declaring full recovery.
This summary provides an overview of our current understanding of the incident given the information available. Our investigation is ongoing and any information herein is subject to change.