Intermittent failures in runner group and runner-related permissions pages
This summary is created by Generative AI and may differ from the actual content.
Overview
On August 18, 2026, between 05:02 UTC and 11:30 UTC, customers were unable to view or manage Actions Runners and Runner Groups through the GitHub UI and API. The issue was caused by failures in backend requests reading runner and runner group data due to an expired authentication certificate unique to this service. The certificate had been rotated in KeyVault, but a step to enable use at runtime had been paused to prevent recurrence of previous incidents triggered by this operation. The impact was mitigated by completing the enablement of the new certificate in the backend system.
Impact
Customers were unable to view or manage Actions Runners and Runner Groups through the GitHub UI and API for approximately 6 hours and 28 minutes. The failure affected runner-related permissions and the ability to load runner groups, particularly impacting customers using Larger Runners.
Trigger
The trigger was an expired authentication certificate unique to the runner service. The certificate had been rotated in KeyVault, but the runtime enablement step was paused to prevent recurrence of previous incidents triggered by this operation.
Detection
Customers reported failures to load runner groups and runner-related permissions. Monitoring detected impacted performance for GitHub services, leading to investigation reports starting at 07:40 UTC on August 18, 2026.
Resolution
The mitigation was completed by enabling the new certificate in the backend system at runtime. This restored communication between Actions services and resolved the failures in backend requests reading runner and runner group data. Additional monitoring was added for this and other certificates.
Root Cause
The root cause was an expired authentication certificate unique to the runner service. The certificate had been rotated in KeyVault, but the step to enable use at runtime had been paused as a preventive measure against previous incidents triggered by this operation, leaving the service unable to use the new certificate.
