Incident with Actions
This summary is created by Generative AI and may differ from the actual content.
Overview
On August 6, 2026, between 15:05 UTC and 00:14 UTC on August 7, GitHub Actions experienced degraded availability lasting approximately 9 hours. At peak, 71% of workflow runs experienced infrastructure failures and 75% of remaining runs were delayed by more than 5 minutes. The incident was triggered by a routine deployment to an internal Actions service that exposed existing capacity and concurrency weaknesses. As pods were replaced, remaining capacity became saturated, causing services to crash and cascading failures across multiple clusters. A latent bug in job assignment services caused runners to be assigned invalid jobs and get stuck retrying them, preventing pickup of valid work. Some Actions Runner Controller runners remained stuck offline after the incident due to a mitigation deployed during the incident that inadvertently affected them. Some jobs created during the incident were left stuck unable to be retried or canceled.
Impact
At peak, 71% of workflow runs experienced infrastructure failures and 75% of remaining workflow runs were delayed by more than 5 minutes. Both GitHub-hosted and self-hosted runners were affected. Downstream services including GitHub Pages, Copilot code review, Copilot coding agent, and GitHub Enterprise Importer migrations were also impacted. Some workflow-triggering events were not processed during the incident and cannot be replayed automatically.
Trigger
A routine deployment to an internal Actions service responsible for processing events and generating Actions jobs. The deployment exposed an existing capacity and concurrency weakness in the system.
Detection
Monitoring systems detected high levels of workflow run failures and delays. Engineers identified the source of the disruption and narrowed the impact to runners stuck retrying invalid jobs through investigation of job completion rates and queue status.
Resolution
Services recovered at 17:00 UTC after expanding capacity, throttling incoming webhook-triggered work to allow system recovery, and increasing processing capacity for accumulated backlogs. Deployed fixes to prevent runners from repeatedly attempting to acquire invalid jobs. Deployed changes to address runners being assigned jobs that are no longer valid. Rolled back the mitigation that inadvertently affected ARC runners. Provided CLI and UI solutions for customers to address stuck jobs. Planned automatic recovery mechanisms in upcoming Runner and Actions Runner Controller releases.
Root Cause
Cascading failure caused by a combination of factors: (1) A routine deployment that exposed existing capacity and concurrency weaknesses in an internal Actions service; (2) As pods were replaced during deployment, remaining capacity became saturated, causing services to crash; (3) A latent bug in job assignment services that caused runners to be assigned invalid jobs, which then got stuck retrying those jobs instead of picking up valid work, preventing recovery of the system.
