Actions delays in starting runs
This summary is created by Generative AI and may differ from the actual content.
Overview
On August 24, 2026, between 13:33 UTC and 14:04 UTC, GitHub Actions experienced degraded performance with 3.8% of runs experiencing start delays over 5 minutes and 1.25% of runs failing outright. A disk failure on a node hosting a service instance responsible for processing runner assignment events caused the incident. Although the node was severely degraded and unable to perform disk operations, it continued sending healthy signals, preventing automatic remediation. Events accumulated in the queue until an automatic rebalance at 13:54 UTC redirected processing to healthy components. The queue backlog was cleared by 14:00 UTC and normal processing resumed by 14:04 UTC.
Impact
3.8% of Actions runs experienced start delays exceeding 5 minutes, and 1.25% of Actions runs failed outright during the 31-minute incident window.
Trigger
Disk failure on a node hosting one of the service instances responsible for processing runner assignment events, combined with the node continuing to send healthy signals despite being severely degraded and unable to perform disk operations.
Detection
Degraded performance was detected through monitoring systems, which identified failures while queuing and running Actions jobs for a subset of customers. The system detected high levels of processing delays and failures, prompting investigation at 13:56 UTC.
Resolution
An automatic rebalance at 13:54 UTC redirected event processing from the affected component to healthy components. The queue backlog was cleared by 14:00 UTC, and normal processing resumed by 14:04 UTC. Future prevention includes improving detection and automated remediation for unhealthy nodes that aren't fully offline, and strengthening application-level resiliency to automatically remove stalled consumers and reassign their work without waiting for node recovery.
Root Cause
Disk failure on a node that continued sending healthy signals despite being severely degraded and unable to perform disk operations. The system's health detection mechanism failed to identify the node as unhealthy, preventing automatic pod removal and work reassignment. This caused runner assignment events to accumulate in the queue until the automatic rebalance mechanism eventually redirected processing to healthy components.
