On July 14, 2026, at 19:55 UTC, customers in our us7: Cloud - US West region experienced delays and failures in CloudWorks™ integration job processing. Integration jobs that were scheduled to run in this region did not complete as expected, resulting in a processing backlog that was cleared once service was fully restored.
Root cause
A back-end component supporting the CloudWorks™ service temporarily exhausted its available memory, leading to a brief connectivity disruption. Although the database recovered immediately, the CloudWorks™ processing components retained inactive connections and were unable to automatically reconnect. This prevented integration jobs from completing and caused a queue of scheduled tasks to accumulate.
Recovery
Our engineering team identified the issue and took immediate action. The team restarted the CloudWorks™ processing components to clear all stale and inactive connections. To quickly clear the backlog of pending integration jobs, we temporarily increased the processing capacity of the components. After verifying that integration jobs were completed normally and the backlog had been fully processed, we returned the system to its original capacity. By 21:14 UTC, the issue was fully resolved.
Corrective and preventative actions
We are implementing the following actions to prevent recurrence:
Memory allocation increase: We are increasing the memory allocation for the back-end component in us7. This reduces the likelihood of a similar memory pressure event occurring in the future.
Background process optimization: We've disabled certain background metric collection processes that were contributing unnecessary load to the back-end component. This further lowers the risk of memory pressure building over time.
Resilience testing: Our engineering teams are running controlled tests in non-production environments to better understand how the CloudWorks™ service behaves during brief connectivity disruptions. This work will help us improve the service's ability to recover automatically — without requiring manual intervention — should a similar event occur in the future.
We apologize for the impact this issue has had on your operations. We are committed to the improvements outlined above to prevent similar disruptions. If you have questions or concerns, please contact Support.