Incident with Pages - Deployment Lag

Severity: Major
Category: Misconfiguration
Service: GitHub

This summary is created by Generative AI and may differ from the actual content.

Overview

On August 6, 2026, at 07:00 UTC, a configuration change inadvertently reduced the capacity of the GitHub Pages deployment processing service. As traffic increased over the following hours, latency in the deployment pipeline progressively increased. At 12:09 UTC, latency crossed the alerting threshold and the team began investigating. The team reverted the invalid configuration and applied additional mitigations, including reducing status deployment processing to lower the load on the Redis cluster. Latency returned to normal levels at 15:40 UTC. Customer impact occurred from 11:34 to 15:32 UTC, during which approximately 128,000 deployments failed to process. The team updated alerts to detect elevated processing latency sooner and to notify immediately when latency causes deployment processing failures. Future work includes updating how GitHub Pages availability is measured to ensure incidents like this are accurately reflected in availability metrics.

Impact

Approximately 128,000 deployments failed to process during the 3 hour 58 minute customer impact window from 11:34 to 15:32 UTC on August 6, 2026. The incident was not fully captured by existing availability metrics.

Trigger

A configuration change that inadvertently reduced the capacity of the service processing GitHub Pages deployments, combined with increasing traffic that caused latency in the deployment pipeline to progressively increase.

Detection

Latency crossed the alerting threshold at 12:09 UTC on August 6, 2026, and the on-call team began investigating at that time.

Resolution

The team reverted the invalid configuration change and applied additional mitigations, including reducing status deployment processing to lower the load on the Redis cluster. Latency returned to normal levels at 15:40 UTC. Updated alerts were implemented to detect elevated processing latency sooner and to notify immediately when latency causes deployment processing failures. Future work includes updating how GitHub Pages availability is measured to accurately reflect incidents like this.

Root Cause

A configuration change that inadvertently reduced the capacity of the GitHub Pages deployment processing service, which when combined with increasing traffic, caused progressive latency increases in the deployment pipeline.