PagerDuty Advance Experiencing Delays and Timeouts

Severity: Major
Category: Dependencies
Service: PagerDuty

This summary is created by Generative AI and may differ from the actual content.

Overview

On June 24, 2026, between 21:59 UTC and 22:43 UTC, PagerDuty customers in the US service region experienced degraded performance for PagerDuty Advance features. The slowdown was caused by increased latency and intermittent errors from a large‑language‑model (LLM) provider used for inference. Core platform components such as event ingestion, notification delivery, and the web console remained fully operational; only request processing for certain Advance features was affected. The issue resolved automatically as the provider restored normal service by 22:43 UTC.

Impact

The impact was limited to slower request processing, delays, and occasional timeouts for PagerDuty Advance users in the US region. No outage of core PagerDuty services occurred, and there was no effect on event ingestion, incident notification delivery, or the PagerDuty SRE web console. Consequently, the incident caused a performance degradation rather than a functional outage.

Trigger

The incident was triggered by increased latency and intermittent errors originating from one of the large‑language‑model providers that PagerDuty Advance relies on for inference on specific features. The provider experienced a region‑wide disruption that affected all of its service regions concurrently.

Detection

PagerDuty’s enhanced monitoring and observability for LLM provider health detected the increased latency and error rates shortly after they began at 21:59 UTC. Internal service health metrics, combined with provider status updates, alerted the on‑call team to the degradation, enabling rapid diagnosis.

Resolution

Recovery began at approximately 22:35 UTC when the LLM provider started to return to normal performance. By 22:43 UTC the provider’s service was fully restored, and PagerDuty Advance returned to normal operation without any manual intervention from the PagerDuty team.

Root Cause

The root cause was a provider‑side disruption affecting the LLM service across all its regions, leading to increased latency and intermittent errors for requests routed to that provider’s models. This provider outage limited the effectiveness of existing cross‑region failover mechanisms, exposing PagerDuty Advance to a single‑provider dependency.