Intermittent failures creating agent tasks
This summary is created by Generative AI and may differ from the actual content.
Overview
Between 13:57 UTC on August 20 and 00:37 UTC on August 21, 2026, some users of the Copilot Cloud Agent experienced delays of up to 60 to 90 minutes in seeing the status and results of their agent tasks. The agent tasks themselves continued to run and complete during this time; only the visibility of their status was delayed. The cause was a regional outage in a third-party cloud database service that Copilot uses to store agent task status.
Impact
Users experienced delays of up to 60 to 90 minutes in seeing the status and results of their agent tasks. No task data was lost and the agent tasks themselves continued to run and complete normally during the incident.
Trigger
Regional outage in a third-party cloud database service that Copilot uses to store agent task status.
Detection
Users reported delays when starting tasks using Copilot Cloud Agent and inability to see the status of these tasks. The issue was identified through user reports and subsequent investigation.
Resolution
Failed over the affected database to a healthy region, added processing capacity to work through the backlog, and restored normal operation once the underlying service recovered. Removed the database configuration that made the system vulnerable to regional outages and improved database failover procedures.
Root Cause
Regional outage in a third-party cloud database service that stores agent task status, combined with a database configuration that made the system vulnerable to regional outages and insufficient failover procedures.
