Resilimap

Multi-cloud reliability monitoring

Site generated: 2026-08-06 12:04:58 UTC 8/6/2026, 12:04:58 PM local
Back to all providers

Datadog

Last updated: 7/30/2026, 10:03:12 AM

✓ All Systems Operational
Active Incidents
0
Resolved
25
Scheduled Maint.
0
Total Incidents
25
Total Maint.
0
Critical
0

Recently Resolved Incidents

Jul 30, 04:27 EDT
Resolved - This incident has been resolved.

Jul 30, 04:16 EDT
Monitoring - A fix has been implemented and we are monitoring the results.

Jul 30, 03:49 EDT
Identified - The issue has been identified and a fix is being implemented.

Jul 30, 03:44 EDT
Investigating - We are investigating delays in Service Check monitors evaluation, which began at 6:35AM UTC.

AI Analysis
Impact: major
Categories: monitoring
Users: all-users
Root Cause: Delayed evaluation of Service Check monitors
Started: 7/30/2026, 8:27:01 AM Resolved: N/A Duration: 171h 38m

Jul 29, 13:28 EDT
Resolved - This incident has been resolved.

Jul 29, 13:12 EDT
Monitoring - A fix has been implemented and we are monitoring the results.

Jul 29, 12:53 EDT
Update - The RUM delays have been resolved. We're continuing to mitigate the issues on distribution metrics.

Jul 29, 12:35 EDT
Identified - The issue has been identified and a fix is being implemented.

Jul 29, 11:42 EDT
Update - We are continuing to investigate this issue.

Jul 29, 11:41 EDT
Investigating - We are currently investigating an issue delaying the ingestion of Real User Monitoring (RUM) measure metrics, custom process metrics, and a subset of distribution metrics for some customers.

As a result, some users may see delays or gaps in RUM, custom process, and distribution metrics. To prevent false alerts caused by delayed data, affected monitors will not notify and will automatically resume once current data is available. All other monitors will operate normally. No data has been lost.

We will provide further updates as the situation progresses.

Started: 7/29/2026, 5:28:08 PM Resolved: N/A Duration: 186h 37m

Jul 27, 15:52 EDT
Resolved - This incident has been resolved.

Jul 27, 15:46 EDT
Monitoring - A fix has been implemented and we're monitoring the results.

Jul 27, 15:36 EDT
Identified - We've identified the issue publishing Synthetic results. Synthetic tests are continuing to run with delayed results.

Jul 27, 15:25 EDT
Update - We're continuing to investigate an issue publishing Synthetic results. Synthetic tests are running and results are delayed.

Jul 27, 15:05 EDT
Investigating - We are currenting investigating an issue running Synthetic tests.

Started: 7/27/2026, 7:52:05 PM Resolved: N/A Duration: 232h 13m

Jul 23, 14:13 EDT
Resolved - This incident has been resolved.

Jul 23, 14:02 EDT
Monitoring - A fix has been implemented and we are monitoring the results.

Jul 23, 13:31 EDT
Identified - The issue has been identified and a fix is being implemented.

Jul 23, 13:08 EDT
Investigating - We are investigating increased latency with Pages and Monitor state changes. As a result of this issue, some users may see delays.

Started: 7/23/2026, 6:13:36 PM Resolved: N/A Duration: 329h 52m

Jul 10, 11:59 EDT
Resolved - This incident has been resolved.

Jul 10, 11:18 EDT
Monitoring - A fix has been implemented and we are monitoring the results. A small percentage of data sent between 6:23 UTC and 14:43 UTC is still being backfilled.

Jul 10, 10:35 EDT
Update - We are continuing to investigate the issue.

Jul 10, 10:09 EDT
Investigating - We are currently investigating an issue affecting Data Streams Monitoring (DSM) in our US1 datacenter. As a result, data stream latency metrics may be delayed or not updating for some users. If you have monitors configured on those metrics, they may fail to alert as expected during this incident. We are not aware of any data loss at this time. Our engineering team is actively working on a fix, and we'll follow up with more information shortly.

Started: 7/10/2026, 3:59:28 PM Resolved: N/A Duration: 644h 6m

Jul 2, 21:29 EDT
Resolved - This incident has been resolved.

Jul 2, 21:21 EDT
Monitoring - A fix has been implemented and we are monitoring the results.

Jul 2, 20:55 EDT
Identified - The issue has been identified and a fix is being implemented.

Jul 2, 20:25 EDT
Investigating - We are investigating errors pulling Datadog container images from registry.datadoghq.com. As a result of this issue, some users may be unable to download or deploy the Datadog Agent, Cluster Agent, and other Datadog container images. Containers that are already running are not affected.

Started: 7/3/2026, 1:29:11 AM Resolved: N/A Duration: 826h 36m

Elevated Error Rates

Resolved Unknown

Jul 1, 16:37 EDT
Resolved - This incident has been resolved.

Jul 1, 15:43 EDT
Update - We are continuing to monitor the situation and are taking additional preventive measures.

Jul 1, 15:19 EDT
Monitoring - We have deployed mitigation actions and are monitoring the recovery.

Jul 1, 15:10 EDT
Identified - We have identified the issue of elevated error rates and are working through mitigation actions

Jul 1, 14:57 EDT
Update - We are investigating elevated error rates across multiple products including: Fleet Automation, Security Products, Cloudcraft, Cloud Network Monitoring, Serverless, Kubernetes Autoscaling, and Database Monitoring.

Jul 1, 14:31 EDT
Investigating - We are actively investigating elevated error rates across multiple Datadog products.

Started: 7/1/2026, 8:37:08 PM Resolved: N/A Duration: 855h 28m

Jun 30, 16:32 EDT
Resolved - This incident has been resolved.

Jun 30, 15:36 EDT
Monitoring - A fix has been implemented and we are monitoring the results.

Jun 30, 15:06 EDT
Update - We are continuing to work on a fix for this issue.

Jun 30, 14:57 EDT
Identified - The issue has been identified and a fix is being implemented.

Jun 30, 14:36 EDT
Update - We are continuing to investigate this issue.

Jun 30, 14:28 EDT
Update - We are continuing to investigate this issue.

Jun 30, 13:54 EDT
Update - We are continuing to investigate this issue.

Jun 30, 13:51 EDT
Investigating - We are currently investigating an issue affecting CI Visibility, Test Optimization, Code Coverage, Software Delivery, Preemptive Alerting, Session Replay, RUM Explorer, Agent Observability, Database Monitoring, Event Correlation, Error Tracking, and all Security products in our US1 datacenter. As a result, some users may experience delays in data processing and errors when using affected products. Our engineering teams have identified the issue and are actively working on a resolution.

Our engineering teams are actively working to resolve the issue, and we'll follow up with more information shortly. Please let us know if you have any questions or need anything else in the meantime.

Started: 6/30/2026, 8:32:44 PM Resolved: N/A Duration: 879h 33m

Jun 30, 07:18 EDT
Resolved - This incident has been resolved.

Jun 30, 07:11 EDT
Monitoring - We have deployed a fix and we are monitoring the results. We will provide another update once the issue is fully resolved.

Jun 30, 06:53 EDT
Identified - We have identified the underlying issue and are working on a fix.
It is important to note that no data has been lost, and notifications will be caught up once the service is operational again.

Jun 30, 06:35 EDT
Investigating - We are investigating delays in Monitors Notifications, which began at 09:52 UTC.

Started: 6/30/2026, 11:18:50 AM Resolved: N/A Duration: 888h 47m

Jun 26, 09:45 EDT
Resolved - This incident has been resolved.

Jun 26, 09:41 EDT
Monitoring - A fix has been implemented and we are monitoring the results.

Jun 26, 09:38 EDT
Identified - We have identified increased latency processing APM Trace Metrics and are working on a fix.
As a result of this issue, some users may see delayed APM Trace Metrics since 13:07 UTC.
To prevent false monitor alerts due to delayed data, monitors affected by the delay will not notify and will automatically resume once current data is available. All other monitors will operate normally..

Started: 6/26/2026, 1:45:37 PM Resolved: N/A Duration: 982h 20m

Showing 10 of 25 resolved incidents