Incident with GraphQL API Requests

Severity: Minor
Category: Scalability
Service: GitHub

This summary is created by Generative AI and may differ from the actual content.

Overview

GraphQL API service experienced degradation between 14:00 UTC and 16:00 UTC on August 11, 2026. Timeout rates averaged 0.06% and peaked at 0.14% of requests. The incident was caused by increased utilization at one site leading to resource contention across dependencies. Capacity was increased to mitigate the bottleneck and error rates returned to normal levels by 16:49 UTC.

Impact

Timeout rate averaged 0.06% and peaked at 0.14% of GraphQL API requests, affecting customers with higher than normal timeouts during the 2-hour window.

Trigger

Increased utilization at one of the sites caused resource contention across dependencies, leading to increased timeouts for GraphQL requests.

Detection

Investigation of reports began at 14:50 UTC when degraded performance for API Requests was identified. Error rates were confirmed and monitoring showed the degradation affecting API Requests.

Resolution

Capacity was increased to alleviate the capacity bottleneck. A fix to increase service capacity was deployed, and error rates returned to normal levels by 16:49 UTC. Continued monitoring was implemented to ensure stability.

Root Cause

Increased utilization at one site caused resource contention across dependencies, resulting in a capacity bottleneck that led to increased timeouts for GraphQL API requests.