Incident with GitHub.com

Severity: Critical
Category: Misconfiguration
Service: GitHub

This summary is created by Generative AI and may differ from the actual content.

Overview

On August 17, 2026, from 13:28–21:15 UTC (7h 47m), GitHub.com experienced widespread elevated errors and latency across Issues, Pull Requests, APIs, Actions, and Copilot services. Web/API error rates reached approximately 20% at peak, while archive and raw-content downloads reached approximately 50%. SAML/OIDC authentication, SCIM, and Team Sync were also affected. Most services recovered by 16:36 UTC as the Central US datacenter recovered; Actions recovered by 18:03 UTC; and Copilot Token Service fully recovered by 21:02 UTC. The incident was caused by network saturation on load balancers in Central US due to an Istio sidecar pod reaching concurrency limits with a misconfigured autoscaling policy. This cascaded to four HAProxy nodes exhausting their flow limits, degrading authentication and causing widespread failures. A latent retry bug in VS Code amplified Copilot Token Service traffic by approximately 10x, delaying recovery. Recovery involved pausing HAProxy nodes, reducing gateway retry logic, and blocking inbound Copilot Token Service requests at load balancers before gradually ramping traffic back up.

Impact

Widespread service degradation affecting multiple critical GitHub services for 7 hours 47 minutes. Web/API error rates reached approximately 20% at peak, with archive and raw-content downloads experiencing approximately 50% error rates. Authentication services (SAML/OIDC, SCIM, Team Sync) were impacted. Copilot Token Service traffic increased from normal 7–9K RPS to 70–100K RPS. The incident affected Git Operations, Webhooks, API Requests, Issues, Pull Requests, Actions, Pages, and Copilot services.

Trigger

A new peak in traffic caused network saturation on load balancers in Central US. An Istio sidecar pod reached its concurrency limits and failed to auto scale correctly due to a misconfigured autoscaling policy that watched host service limits but not sidecar limits. This initial failure cascaded, causing four HAProxy nodes to exhaust their flow limits and degrade the gateway authentication path.

Detection

Monitoring systems detected high error rates and latency across multiple services. The incident was identified through elevated error rates around 20% for web experiences and API traffic, and approximately 50% error rates for archive downloads and raw repository content downloads. SAML/OIDC authentication, SCIM, and Team Sync failures were also detected.

Resolution

Immediate mitigation involved pausing HAProxy on the affected nodes, which produced immediate broad recovery. Additional mitigations included: 1) temporarily reducing gateway retry logic via a PR, 2) blocking inbound Copilot Token Service token requests at load balancers with a 403 response, and 3) gradually ramping back up traffic per-site to allow callers to succeed. Traffic was moved from Central US to Northern Virginia during the network failure investigation and resolution. Recovery was completed by reducing gateway authentication retries and blocking retry-triggering responses to stabilize the Copilot Token Service.

Root Cause

Cascading failure caused by multiple factors: 1) Istio sidecar pod reaching concurrency limits with a misconfigured autoscaling policy that watched host service limits but not sidecar limits, 2) resulting network saturation on Central US load balancers, 3) four HAProxy nodes exhausting their flow limits, 4) a latent retry bug in VS Code that amplified Copilot Token Service traffic by approximately 10x, and 5) optimistic retry logic that overloaded internal load balancers. Complicating factors included scraping attacks on codeload endpoints and client retry behavior that amplified load during token operation failures.