Resilimap

Multi-cloud reliability monitoring

Site generated: 2026-08-06 12:04:58 UTC 8/6/2026, 12:04:58 PM local
Home / Github / Incident Details

Incident with Actions

Resolved Scheduled Maintenance
Started
Sat, Jul 25, 2026, 09:25 AM UTC
Last Updated
Sat, Jul 25, 2026, 09:25 AM UTC
Resolved
N/A
This maintenance window has been completed (Total duration: 12d 2h 40m)

Incident History

Jul 25, 09:25 UTC
Resolved - On July 25, 2026, GitHub Actions experienced two related periods of degradation that caused some workflow runs to be delayed by more than 5 minutes or end with infrastructure failures.

First period (08:45 – 09:13 UTC): During planned maintenance on a critical-path Redis cluster for Actions, one participating region was left in a degraded state. Separately, an independent capacity operation temporarily removed another region from the cluster and redirected its traffic to the degraded region. This created cross-region inconsistencies in job-assignment state, causing workflow runs to be delayed, exhaust retries, or fail outright. At peak, about 7% of runs were delayed by more than 5 minutes, and 25% of runs failed with an infrastructure error during the course of the incident. We mitigated the incident at 09:13 UTC by returning traffic to its normal distribution.

Second period (12:08 – 12:48 UTC): As part of mitigating the first incident, traffic was returned to the regional instance that was still undergoing its capacity increase. Multiple Redis nodes in the scaling region experienced failures, increasing traffic to healthy nodes and causing connection limits to be reached on many nodes. At peak, 30% of runs were delayed by more than 5 minutes, and 60% of runs failed with an infrastructure error during the course of the incident. We mitigated the incident at 12:48 UTC by redirecting workflow traffic away from the scaling region.

We are adding stronger regional health and capacity checks before maintenance and requiring a stable observation period before restoring traffic. We are also improving automated connection resiliency, and partnering with our platform dependency to automatically detect and remediate unhealthy cluster members and shard imbalance. More generally, we already had work underway to improve the resiliency and scale of this piece of Actions infrastructure.

Jul 25, 09:20 UTC
Update - We identified an issue causing delays in GitHub Actions run starts. Some users may have experienced longer than expected wait times when triggering workflow runs. We have applied mitigations and have recovered. Our team continues to monitor and investigate the root cause.

Jul 25, 09:13 UTC
Monitoring - The degradation affecting Actions has been mitigated. We are monitoring to ensure stability.

Jul 25, 08:59 UTC
Investigating - We are investigating reports of degraded performance for Actions

External Resources