[{"id":"3BvH3LVGcupoYqV6F4Nw","number":"11738310674864064645","begin":"2026-07-15T23:57:00+00:00","created":"2026-07-16T03:30:59+00:00","end":"2026-07-16T12:25:00+00:00","modified":"2026-07-25T13:16:55+00:00","external_desc":"Google Cloud VMware Engine (GCVE), Google Cloud NetApp Volumes, and Bare Metal Solutions (BMS) services are experiencing a service outage in europe-west4-a due to a cooling failure.","updates":[{"created":"2026-07-25T13:16:55+00:00","modified":"2026-07-25T13:16:55+00:00","when":"2026-07-25T13:16:55+00:00","text":"## \\# Incident Report\n## \\#\\# Summary\nOn Wednesday, 15 July 2026, Google Cloud VMware Engine (GCVE), Bare Metal Solution (BMS), and Google Cloud NetApp Volumes (GCNV) experienced service interruptions for a total duration of 14 hours, 55 minutes.\nThe root cause for this outage is a 3ms voltage drop in the power feed from the utility provider and subsequent utility breaker protective action. This incident was mitigated by the data center provider rectifying the failed systems followed by the Google engineering team restoring the services.\n## \\#\\# Root Cause\nA regional data center hosting services in europe-west4-a experienced an upstream voltage transient affecting both utility power feeds A and B. During this event, the utility breakers on both A and B feeds tripped and initiated the transfer to the back-up power source DRUPS (Diesel Rotary Uninterruptible Power Supply). The transition of side B to DRUPS system was successful without any power interruption. The DRUPS back-up power system for side A failed to take over the facility load due to electrical component failures.\nThe DRUPS failure to take over load initiated automatic transfer of the affected 3 rows from feed A to redundant power feed B. Rows 1 and 2 transferred to power feed B successfully. Row 3 failed to transfer to power feed B and experienced complete power loss due to an overload protection breaker trip. The root cause of the failed transfer on the affected row was identified as load deployment discrepancy which is under further investigation by Google engineering team. This resulted in loss of both redundant power feeds to the single row.\nDuring the voltage transient event, the server data hall experienced an increased temperature due to a cooling system failure. The chiller controller dropped offline during the voltage transient event, failing to signal the chilled water distribution pumps to restart and ultimately causing the chiller system A to shut down. The redundant source was not available due to known ongoing construction work at the facility. This resulted in the data hall temperatures reaching 44°C and subsequent shutdown of the affected data hall machines. Google engineering teams initiated machine shutdown procedures for the remainder of reachable devices as part of the cooling emergency shutdown process.\nA timeline of events during the incident is provided below.\n### DataCenter Events\n| Timestamp(PST) | Event Description |\n| :---- | :---- |\n| 07-15-2026 16:24 | An electrical fault occurred on the utility grid upstream of the data center, disrupting the electrical distribution. |\n| 07-15-2026 16:24 | DRUPS back-up power system for side A failed to take over the data hall load. |\n| 07-15-2026 16:24 | The chiller controller dropped offline during the voltage transient event causing distribution pumps to stop and unable to restart. |\n| 07-15-2026 16:37 | Rows 1 and 2 transferred to power feed B successfully. Row 3 failed to transfer to power feed B and lost power. |\n| 07-15-2026 18:29 | Data hall temperatures reached 44C and crossed the safe machine operating threshold. |\n| 07-15-2026 19:55 | Machine shutdown procedures implemented. |\n| 07-15-2026 21:05 | Cooling system fully recovered and temperatures in the data hall returned to normal operating range. |\n| 07-15-2026 21:24 | Notification about cooling recovery sent by provider |\n| 07-15-2026 21:46 | Notification about power recovery sent by provider |\n### GCVE Timelines\n| Timestamp(PST) | Event Description |\n| :---- | :---- |\n| 07-15-2026 16:30 | GCVE service power monitoring alert received |\n| 07-15-2026 16:41 | First symptom detected, servers reporting redundancy power feed failure. |\n| 07-15-2026 17:05 | Switch temperature \\\u003e60°C alerts triggered. |\n| 07-15-2026 17:09 | Start of user impact, first prober alert failure for north south traffic for customers. |\n| 07-15-2026 19:00 | Server / Network devices shutdown initiated. |\n| 07-15-2026 19:55 | All reachable server/network devices were shut down. |\n| 07-15-2026 21:24 | Cooling system fully recovered and temperatures in the datahall were returning to normal operating range. |\n| 07-15-2026 21:56 | Network restoration started with reachable hydra devices. |\n| 07-15-2026 22:40 | Onsite technician arrived, to recover the console servers. |\n| 07-15-2026 23:10 | Console server reboot/recovery complete. |\n| 07-16-2026 00:33 | Network recovery for all Placement Groups complete. |\n| 07-16-2026 02:36 | Incident mitigated, after recovering all the customer PCs are fully healthy. |\nThis outage impacted 24 private clouds of 20 distinct customers in europe-west4-a region.\n### Bare Metal Solutions (BMS) Timelines\n| Timestamp(PST) | Event Description |\n| :---- | :---- |\n| 07-15-2026 16:41 | First symptom detected, servers and storage reporting redundancy power feed failure |\n| 07-15-2026 17:52 | Automated monitoring detected rising temperatures |\n| 07-15-2026 18:37 | Four Netapp SAN storage nodes failed and were shut down due to overheating, resulting in storage availability issue |\n| 07-15-2026 18:37 | Customer impact started, first server shutdown detected |\n| 07-15-2026 18:59 | Reserve and buffer servers started to get powered off |\n| 07-15-2026 20:43 | Drop in the temperature has been detected |\n| 07-15-2026 20:52 | Four Netapp SAN storage nodes recovered, mitigating storage availability issue |\n| 07-15-2026 21:05 | The cooling system fully recovered and temperatures in the datahall were returning to normal operating range. |\n| 07-15-2026 23:23 | First customer server reboot started |\n| 07-16-2026 06:32 | Last customer server reboot started |\n| 07-16-2026 08:42 | Incident mitigated. All impacted customer servers were rebooted and confirmed in healthy state. |\nThis outage impacted 9 distinct BMS customers in europe-west4\n### NetApp Timelines\n| Timestamp(PST) | Event Description |\n| :---- | :---- |\n| 07-15-2026 16:32 | Netapp facility remote monitoring alert on high temp and network switches switches failure received |\n| 07-15-2026 17:32 | All 6 clusters failed and were shut down automatically due to high temperature |\n| 07-15-2026 21:05 | Cooling capacity restored |\n| 07-15-2026 22:08 | Cluster recovery process initiated for all 6 clusters. 1st cluster recovered |\n| 07-15-2026 23:55 | All 6 Netapp Clusters and all customers recovered and validated as healthy \u0026 operational |\n## \\#\\# Remediation and Prevention\nGoogle engineers were alerted to the infrastructure power loss on Wednesday, 15 July 2026 at 16:30 US/Pacific and subsequent increase in data hall temperatures at 17:44 US/Pacific via our automated monitoring and telemetry, and immediately began investigating.\nDuring the voltage transient event, the chillers maintained power but shut down because the distribution pumps failed to restart, resulting in no water flow. The pumps were manually switched from auto to hand mode to restore circulation and cooling. The facility team deployed an interim portable UPS to support the chiller controllers and prevent localized controller power loss.\nUpon confirmation of utility power restoration, the facility returned to utility power via automatic switching with the exception of two Remote Power Panels (RPPs), which required manual verification before restoration of power to affected server row. To provide full resiliency the facility team aligned a swing DRUPS unit to restore the redundant configuration on the A feed and is working on repairing the faulty DRUPS.\nTo remediate the power loss at Row 3, once the electrical fault was confirmed to be fully isolated, teams reset the tripped breakers, successfully restoring both redundant power to the impacted server racks.\nGoogle is committed to preventing a repeat of this issue in the future and is completing the following actions:\n* Detailed investigation by the engineering team of the sequence of events that prevented transfer of load to DRUPS system and identified system improvements **\\[ETA Aug 2026\\].**\n* Determine and resolve the cause of overload conditions that caused power loss to Row 3 **\\[ETA Aug 2026\\].**\n* Perform investigation of chiller pump control system redundancy setup to ensure cooling system resiliency **\\[ETA Sep 2026\\].**\n* Implement improvements to facility system monitoring alerts for power events and configure early alert thresholds for cooling excursion detection. **\\[ETA Aug 2026\\].**\n* The GCVE engineering team is working to improve the automated triggering of shutdown during quickly evolving thermal runoff situations **\\[ETA Oct 2026\\]**.\n* The BMS engineering team is working with partner teams to update the incident classification and SLO definitions to improve personnel availability during major datacenter events **\\[ETA Sep 2026\\]**.\n* The GCNV engineering team is collaborate with the data center team to review incident handling SLA and runbooks, identify potential future occurrences and create playbooks to address/mitigate them, establish monthly joint emergency drill, and review if the Netapp cluster recovery process can be expedited **\\[ETA Aug 2026\\].**\n## \\#\\# Detailed Description of Impact\nOn 15 July 2026 from 16:39 to 16 July 2026 07:34 US/Pacific, customers experienced service disruptions across several products in europe-west4-a zone:\n* **Google Cloud VMware Engine**: Customers lost connectivity to their private clouds and workloads as foundational hosts and network switches were powered off for thermal protection.\n* **Bare Metal Solution**: Customers experienced connectivity loss to their database servers and storage appliances as the underlying hardware was powered down.\n* **Google Cloud NetApp Volumes**: Customers in the STANDARD, PREMIUM, and EXTREME service levels were unable to access their storage volumes. Control plane operations, including the creation of new storage pools, volumes, and backups, experienced failures in the affected region.","status":"AVAILABLE","affected_locations":[]},{"created":"2026-07-18T00:22:37+00:00","modified":"2026-07-25T13:16:55+00:00","when":"2026-07-18T00:22:37+00:00","text":"# Preliminary Incident Report\nWe sincerely apologize for the disruption this incident caused to your business. We know how much you rely on Google Cloud, and we regret the impact on your productivity.\nPlease note, this information is based on our best knowledge at the time of posting and is subject to change as our investigation continues. A final Incident Report with preventative actions will be posted once our investigation is complete.\nIf you have experienced impact outside of what is listed below, please reach out to Google Cloud Support using https://cloud.google.com/support.\n## Date/Time of the Issue (All time US/Pacific)\n**Google Cloud VMware Engine Impact:**\n- Start: 15 July 2026 17:09\n- End: 16 July 2026 2:33\n- Duration: 9 Hours, 24 Minutes\n**Google Cloud NetApp Volumes Impact:**\n- Start: 15 July 2026 16:39\n- End: 16 July 2026 01:10\n- Duration: 8 Hours 31 Minutes\n**Bare Metal Solution Impact:**\n- Start: 15 July 2026 18:37\n- End: 16 July 2026 07:34\n- Duration: 12 Hours 57 Minutes\n## Summary\nOn Wednesday, 15 July 2026, Google Cloud VMware Engine, Bare Metal Solution, and Google Cloud NetApp Volumes experienced service interruptions for a total duration of 14 hours, 55 minutes. We are taking immediate steps to ensure this doesn’t happen again.\n## Preliminary Root Cause\nA regional data center hosting services in europe-west4-a experienced a loss of utility power and subsequent cooling capacity.\nThe sequence of events leading to customer impact was as follows:\n- An electrical fault occurred on the utility grid upstream of the data center, disrupting the electrical distribution gear and cooling equipment.\n- The loss of cooling infrastructure resulted in ambient temperatures rising rapidly within the affected data halls.\nHost servers, storage clusters, and network switches were shut down to prevent equipment damages due to the extreme heat.\n- The high temperatures in the data hall resulted in disruption to customer workloads and control plane operations across the affected services.\nGoogle engineers have begun a full root cause analysis and we will provide additional information once it is available.\n## Remediation\nGoogle engineering teams were alerted to the issue via automated temperature and hardware unreachability telemetry starting at 16:39 US/Pacific.\nEngineers collaborated with the third-party facility provider to safely restore primary utility power and cooling systems, returning ambient room temperatures to safe operational levels. With the environment stabilized, field support engineers and remote teams systematically booted and verified the server hosts, storage nodes, and network fabric in a controlled sequence. Automated recovery playbooks were executed to restore underlying network switches and routing infrastructure, allowing storage appliances and compute clusters to be brought back online and verified for health. The underlying infrastructure has been fully recovered, and normal operations have been successfully restored.\n## Description of Impact\nOn 15 July 2026 from 16:39 to 16 July 2026 07:34 US/Pacific, customers experienced service disruptions across several products in europe-west4-a zone:\n- **Google Cloud VMware Engine:** Customers lost connectivity to their private clouds and workloads as foundational hosts and network switches were powered off for thermal protection.\n- **Bare Metal Solution:** Customers experienced connectivity loss to their database servers and storage appliances as the underlying hardware was powered down.\n- **Google Cloud NetApp Volumes:** Customers in the STANDARD, PREMIUM, and EXTREME service levels were unable to access their storage volumes. Control plane operations, including the creation of new storage pools, volumes, and backups, experienced failures in the affected region.","status":"AVAILABLE","affected_locations":[]},{"created":"2026-07-16T12:25:47+00:00","modified":"2026-07-18T00:22:37+00:00","when":"2026-07-16T12:25:47+00:00","text":"**Summary**\nGoogle Cloud VMware Engine (GCVE), Google Cloud NetApp Volumes, and Bare Metal Solution (BMS) services experienced a service outage in europe-west4-a due to a cooling failure.\n**Description**\nThe datacenter serving europe-west4-a for GCVE, BMS, and NetApp has experienced a power failure, which subsequently caused a cooling failure. Google proactively turned down workloads in order to protect customer data from any risks posed by running infrastructure in a high temperature environment.\nCooling and power have been restored at the site. The issue with GCVE and Google Cloud NetApp Volumes has been resolved for all affected users as of Thursday, 2026-07-16 02:31 PDT and 01:00 PDT respectively.\nThe issue with BMS is largely resolved and is believed to be affecting a very small number of customers. Our Engineering will continue to recover this residual impact directly with the impacted customers.\nIf you have questions or are still impacted, please open a case with the Support Team and we will work with you until this issue is resolved.\n**GCVE:** Service restoration is complete across all the Private Clouds. Customers have been informed about this restoration and requested to verify that their VMs are fully operational.\n**Google Cloud NetApp Volumes:** Service restoration is complete, and all affected servers are now operational.\n**BMS:** Service restoration is largely complete. We continue working with a very small number of remaining affected customers to achieve full recovery.\nWe thank you for your patience while we continue to fully resolve the issue.\n**Customer Symptoms**\nThe issue is now resolved for GCVE and Google Cloud NetApp Volumes. The issue with BMS is largely resolved and our engineering will continue to recover this residual impact directly with the impacted customers.\n**GCVE:** Customers would have observed a loss of connectivity to their private cloud.\n**Google Cloud NetApp Volumes:** Customers would have lost connectivity to ONTAP clusters, as they were proactively shut down to protect these clusters from any damage. Only Standard, Premium, and Extreme service levels were affected.\n**BMS:** Customers may have lost connectivity to their servers. As a protective measure for hardware integrity and recovery, we had shut down some servers in this zone.\n**Workaround**\nNone at this time.","status":"SERVICE_DISRUPTION","affected_locations":[{"title":"Netherlands (europe-west4)","id":"europe-west4"}]},{"created":"2026-07-16T11:39:12+00:00","modified":"2026-07-16T12:25:47+00:00","when":"2026-07-16T11:39:12+00:00","text":"**Summary**\nGoogle Cloud VMware Engine (GCVE), Google Cloud NetApp Volumes, and Bare Metal Solution (BMS) services experienced a service outage in europe-west4-a due to a cooling failure.\n**Description**\nThe datacenter serving europe-west4-a for GCVE, BMS, and NetApp has experienced a power failure, which subsequently caused a cooling failure. Google proactively turned down workloads in order to protect customer data from any risks posed by running infrastructure in a high temperature environment.\nCooling and power have been restored at the site. The issue with GCVE and Google Cloud NetApp Volumes has been resolved for all affected users. We continue to recover the impact for BMS at the moment.\n**GCVE [Mitigated]:** Service restoration is complete across all the Private Clouds. Customers have been informed about this restoration and requested to verify that their VMs are fully operational.\n**Google Cloud NetApp Volumes [Mitigated]:** Service restoration is complete, and all affected servers are now operational. We continue to monitor these for full service recovery.\n**BMS:** Server restoration is largely complete, with most systems now online and reporting a healthy status. We are actively working through the remaining validation efforts to achieve full recovery.\nFor Google Cloud NetApp Volumes and GCVE, if customers are still experiencing impact from this issue, please contact Support, and we will work with you to resolve any residual impact.\nWe will provide the BMS restoration status update by Thursday, 2026-07-16 05:30 PDT with current details.\n**Customer Symptoms**\nThe issue is now resolved for GCVE and Google Cloud NetApp Volumes. We continue to recover the impact for BMS at the moment.\n**GCVE [Mitigated]:** Customers would have observed a loss of connectivity to their private cloud.\n**Google Cloud NetApp Volumes [Mitigated]:** Customers would have lost connectivity to ONTAP clusters, as they were proactively shut down to protect these clusters from any damage. Only Standard, Premium, and Extreme service levels were affected. Customers should now be able to access their volumes again.\n**BMS:** Customers may have lost connectivity to their servers. As a protective measure for hardware integrity and recovery, we had shut down some servers in this zone. We are continuing to restore servers and monitor their health.\n**Workaround**\nNone at this time.","status":"SERVICE_DISRUPTION","affected_locations":[{"title":"Netherlands (europe-west4)","id":"europe-west4"}]},{"created":"2026-07-16T10:14:07+00:00","modified":"2026-07-16T11:39:12+00:00","when":"2026-07-16T10:14:07+00:00","text":"**Summary**\nGoogle Cloud VMware Engine (GCVE), Google Cloud NetApp Volumes, and Bare Metal Solution (BMS) services experienced a service outage in europe-west4-a due to a cooling failure.\n**Description**\nThe datacenter serving europe-west4-a for GCVE, BMS, and NetApp has experienced a power failure, which subsequently caused a cooling failure. Google proactively turned down workloads in order to protect customer data from any risks posed by running infrastructure in a high temperature environment.\nCooling and power have been restored at the site. The issue with GCVE and Google Cloud NetApp Volumes has been resolved for all affected users. We continue to recover the impact for BMS at the moment.\n**GCVE [Mitigated]:** Service restoration is complete across all the Private Clouds. Customers have been informed about this restoration and requested to verify that their VMs are fully operational.\n**Google Cloud NetApp Volumes [Mitigated]:** Service restoration is complete, and all affected servers are now operational. We continue to monitor these for full service recovery.\n**BMS:** Server restoration is largely complete, with most systems now online and reporting a healthy status. We are actively working through the remaining validation efforts to achieve full recovery.\nFor Google Cloud NetApp Volumes and GCVE, if customers are still experiencing impact from this issue, please contact Support, and we will work with you to resolve any residual impact.\nWe will provide the BMS restoration status update by Thursday, 2026-07-16 04:30 PDT with current details.\n**Customer Symptoms**\nThe issue is now resolved for GCVE and Google Cloud NetApp Volumes. We continue to recover the impact for BMS at the moment.\n**GCVE [Mitigated]:** Customers would have observed a loss of connectivity to their private cloud.\n**Google Cloud NetApp Volumes [Mitigated]:** Customers would have lost connectivity to ONTAP clusters, as they were proactively shut down to protect these clusters from any damage. Only Standard, Premium, and Extreme service levels were affected. Customers should now be able to access their volumes again.\n**BMS:** Customers may have lost connectivity to their servers. As a protective measure for hardware integrity and recovery, we had shut down some servers in this zone. We are continuing to restore servers and monitor their health.\n**Workaround**\nNone at this time.","status":"SERVICE_DISRUPTION","affected_locations":[{"title":"Netherlands (europe-west4)","id":"europe-west4"}]},{"created":"2026-07-16T09:21:21+00:00","modified":"2026-07-16T10:14:07+00:00","when":"2026-07-16T09:21:21+00:00","text":"**Summary**\nGoogle Cloud VMware Engine (GCVE), Google Cloud NetApp Volumes, and Bare Metal Solution (BMS) services experienced a service outage in europe-west4-a due to a cooling failure.\n**Description**\nThe datacenter serving europe-west4-a for GCVE, BMS, and NetApp has experienced a power failure, which subsequently caused a cooling failure. Google proactively turned down workloads in order to protect customer data from any risks posed by running infrastructure in a high temperature environment.\nCooling and power have been restored at the site. The datacenter operator team is now validating site stability.\n**GCVE:** Network fabric has been restored, and Private Cloud restoration remains underway. We are notifying the affected customers as their Private Clouds are restored to allow for VM operational verification.\n**BMS:** Server restoration is largely complete, with most systems now online and reporting a healthy status. We are actively working through the remaining validation efforts to achieve full recovery.\n**Google Cloud NetApp Volumes:** Service restoration is complete, and all affected servers are now operational. We continue to monitor these for full service recovery. If customers are still experiencing impact from this issue, please contact Support, and we will work with you to resolve any residual impact.\nWe will provide an update by Thursday, 2026-07-16 03:30 PDT with current details.\n**Customer Symptoms**\n**GCVE:** Customers would have observed a loss of connectivity to their private cloud. As a protective measure for hardware integrity and recovery, we had shut down all Private Clouds in this zone. We are in the process of bringing all restored Private Clouds online.\n**BMS:** Customers may have lost connectivity to their servers. As a protective measure for hardware integrity and recovery, we had shut down some servers in this zone. We are continuing to restore servers and monitor their health.\n**Google Cloud NetApp Volumes:** Customers would have lost connectivity to ONTAP clusters, as they were proactively shut down to protect these clusters from any damage. Only Standard, Premium, and Extreme service levels were affected. Customers should now be able to access their volumes again.\n**Workaround**\nNone at this time.","status":"SERVICE_DISRUPTION","affected_locations":[{"title":"Netherlands (europe-west4)","id":"europe-west4"}]},{"created":"2026-07-16T08:32:22+00:00","modified":"2026-07-16T09:21:21+00:00","when":"2026-07-16T08:32:22+00:00","text":"**Summary**\nGoogle Cloud VMware Engine (GCVE), Google Cloud NetApp Volumes, and Bare Metal Solution (BMS) services experienced a service outage in europe-west4-a due to a cooling failure.\n**Description**\nThe datacenter serving europe-west4-a for GCVE, BMS, and NetApp has experienced a power failure, which subsequently caused a cooling failure. Google proactively turned down workloads in order to protect customer data from any risks posed by running infrastructure in a high temperature environment.\nCooling and power have been restored at the site. The datacenter operator team is now validating site stability.\n**GCVE:** Network fabric has been restored, and Private Cloud restoration remains underway. We are notifying the affected customers as their Private Clouds are restored to allow for VM operational verification.\n**BMS:** Server restoration is largely complete, with most systems now online and reporting a healthy status. We are actively working through the remaining validation efforts to achieve full recovery.\n**Google Cloud NetApp Volumes:** Service restoration is complete, and all affected servers are now operational. We continue to monitor these for full service recovery. If customers are still experiencing impact from this issue, please contact Support, and we will work with you to resolve any residual impact.\nWe will provide an update by Thursday, 2026-07-16 02:30 PDT with current details.\n**Customer Symptoms**\n**GCVE:** Customers would have observed a loss of connectivity to their private cloud. As a protective measure for hardware integrity and recovery, we had shut down all Private Clouds in this zone. We are in the process of bringing all restored Private Clouds online.\n**BMS:** Customers may have lost connectivity to their servers. As a protective measure for hardware integrity and recovery, we had shut down some servers in this zone. We are continuing to restore servers and monitor their health.\n**Google Cloud NetApp Volumes:** Customers would have lost connectivity to ONTAP clusters, as they were proactively shut down to protect these clusters from any damage. Only Standard, Premium, and Extreme service levels were affected. Customers should now be able to access their volumes again.\n**Workaround**\nThere are no workarounds available for this issue at this time. However, customers with multi-regional deployments are advised to route traffic to an alternate site.","status":"SERVICE_DISRUPTION","affected_locations":[{"title":"Netherlands (europe-west4)","id":"europe-west4"}]},{"created":"2026-07-16T07:39:01+00:00","modified":"2026-07-16T08:32:22+00:00","when":"2026-07-16T07:39:01+00:00","text":"**Summary**\nGoogle Cloud VMware Engine (GCVE), Google Cloud NetApp Volumes, and Bare Metal Solution (BMS) services are experiencing a service outage in europe-west4-a due to a cooling failure.\n**Description**\nThe datacenter serving europe-west4-a for GCVE, BMS, and NetApp has experienced a power failure, which subsequently caused a cooling failure. Google proactively turned down workloads in order to protect customer data from any risks posed by running infrastructure in a high temperature environment.\nCooling and power have been restored at the site. The datacenter operator team is now working to safely restart affected workloads.\n**GCVE:** As a precautionary measure, our engineering teams had shut down private clouds to protect these clouds from any damage. Network fabric has been restored, and Private Cloud restoration remains underway. We will notify affected customers once their Private Clouds are restored to allow for VM operational verification.\n**BMS:** We are continuing to restore servers and monitor their health. Customers may be unable to access these systems until service restoration is complete.\n**Google Cloud NetApp Volumes:** Our engineering team had proactively shut down ONTAP clusters to protect the data stored on them. Recovery is currently in progress. Customers may experience unavailability of Cloud NetApp Volumes until this issue is resolved.\nWe will provide an update by Thursday, 2026-07-16 01:30 PDT with current details.\n**Customer Symptoms**\n**GCVE:** Customers would have already observed a loss of connectivity to their private cloud. As a protective measure for hardware integrity and recovery, we have shut down all private clouds in this zone.\n**BMS:** Customers may have lost connectivity to their servers. As a protective measure for hardware integrity and recovery, we shut down some servers in this zone. We are continuing to restore servers and monitor their health. Customers may be unable to access these systems until service restoration is complete.\n**Google Cloud NetApp Volumes:** Customers would have already lost connectivity to ONTAP clusters, as we proactively shut them down to protect these clusters from any damage. Only Standard, Premium, and Extreme service levels are affected.\n**Workaround**\nThere are no workarounds available for this issue at this time. However, customers with multi-regional deployments are advised to route traffic to an alternate site.","status":"SERVICE_DISRUPTION","affected_locations":[{"title":"Netherlands (europe-west4)","id":"europe-west4"}]},{"created":"2026-07-16T06:29:45+00:00","modified":"2026-07-16T07:39:01+00:00","when":"2026-07-16T06:29:45+00:00","text":"**Summary**\nGoogle Cloud VMware Engine (GCVE), Google Cloud NetApp Volumes, and Bare Metal Solution (BMS) services are experiencing a service outage in europe-west4-a due to a cooling failure.\n**Description**\nThe datacenter serving europe-west4-a for GCVE, BMS, and NetApp has experienced a power failure, which subsequently caused a cooling failure. Google proactively turned down workloads in order to protect customer data from any risks posed by running infrastructure in a high temperature environment.\nCooling and power have been restored at the site. The datacenter operator team is now working to safely restart affected workloads.\n**GCVE:** As a precautionary measure, our engineering teams had shut down private clouds to protect these clouds from any damage. Service restoration is still in progress. We will notify customers as soon as their private cloud has been powered back up and it is safe to restart the VMs.\n**BMS:** We are continuing to restore servers and monitor their health. Customers may be unable to access these systems until service restoration is complete.\n**Google Cloud NetApp Volumes:** Our engineering team had proactively shut down ONTAP clusters to protect the data stored on them. Customers may experience unavailability of Cloud NetApp Volumes until this issue is resolved.\nWe will provide an update by Thursday, 2026-07-16 00:30 PDT with current details.\n**Customer Symptoms**\n**GCVE:** Customers would have already observed a loss of connectivity to their private cloud. As a protective measure for hardware integrity and recovery, we have shut down all private clouds in this zone.\n**BMS:** Customers may have lost connectivity to their servers. As a protective measure for hardware integrity and recovery, we shut down some servers in this zone. We are continuing to restore servers and monitor their health. Customers may be unable to access these systems until service restoration is complete.\n**Google Cloud NetApp Volumes:** Customers would have already lost connectivity to ONTAP clusters, as we proactively shut them down to protect these clusters from any damage. Only Standard, Premium, and Extreme service levels are affected.\n**Workaround**\nThere are no workarounds available for this issue at this time. However, customers with multi-regional deployments are advised to route traffic to an alternate site.","status":"SERVICE_DISRUPTION","affected_locations":[{"title":"Netherlands (europe-west4)","id":"europe-west4"}]},{"created":"2026-07-16T04:41:44+00:00","modified":"2026-07-16T06:31:30+00:00","when":"2026-07-16T04:41:44+00:00","text":"**Summary**\nGoogle Cloud VMware Engine (GCVE), Google Cloud NetApp Volumes, and Bare Metal Solution (BMS) services are experiencing a service outage in europe-west4-a due to a cooling failure\n**Description**\nThe datacenter serving europe-west4-a for GCVE, BMS, and NetApp has experienced a power failure, which subsequently caused a cooling failure. Google has proactively turned down workloads in order to protect customer data from any risks posed by running infrastructure in a high temperature environment. The datacenter operator team is working diligently to bring cooling and power back online so we can safely restart affected workloads. Some cooling has been restored, and we are working to restore affected workloads as available capacity allows.\n**GCVE:** Cooling has been restored at the site. As a precautionary measure, our engineering teams had shut down private clouds to protect these clouds from any damage. We are beginning service restoration and will notify customers as soon as their private cloud has been powered back up and it is safe to restart the VMs.\n**BMS:** Some Bare Metal Solution servers experienced high temperatures and we proactively shut them down. We are starting to bring the servers back online, and will provide an update once they are ready. Customers may be unable to access these systems until service is restored.\n**Google Cloud NetApp Volumes:** Our engineering team has proactively shut down ONTAP clusters to protect the data stored on them. Customers may experience unavailability of Cloud NetApp Volumes until this issue is resolved.\nThere is no estimated time for resolution. We will provide an update by Wednesday, 2026-07-15 23:30 PDT with current details.\nWe apologize to all who are affected by the disruption.\n**Customer Symptoms**\n**GCVE:** Customers would have already observed a loss of connectivity to their private cloud. As a protective measure for hardware integrity and recovery, we have shut down all private clouds in this zone.\n**BMS:** Customers may have lost connectivity to their BMS machines. As a protective measure for hardware integrity and recovery, we shut down some servers in this zone, and are actively restoring them. Customers may be unable to access these systems until restoration is complete.\n**Google Cloud NetApp Volumes:** Customers would have already lost connectivity to ONTAP clusters, as we proactively shut them down to protect these clusters from any damage. Only Standard, Premium, and Extreme service levels are affected.\n**Workaround**\nThere are no workarounds available for this issue at this time. However, customers with multi-regional deployments are advised to route traffic to an alternate site.","status":"SERVICE_DISRUPTION","affected_locations":[{"title":"Netherlands (europe-west4)","id":"europe-west4"}]},{"created":"2026-07-16T04:17:44+00:00","modified":"2026-07-16T06:31:21+00:00","when":"2026-07-16T04:17:44+00:00","text":"**Summary**\nGoogle Cloud VMware Engine (GCVE), Google Cloud NetApp Volumes, and Bare Metal Solution (BMS) services are experiencing temperature alerts in europe-west4-a\n**Description**\nThe datacenter serving europe-west4-a for GCVE, BMS, and NetApp has experienced a power failure, which subsequently caused a cooling failure. Google has proactively turned down workloads in order to protect customer data from any risks posed by running infrastructure in a high temperature environment. The datacenter operator team is working diligently to bring cooling and power back online so we can safely restart affected workloads.\n**GCVE:** Our engineering teams have shut down private clouds, and we expect that others will shut down as well. We will notify customers as soon as their private cloud has been powered back up and it is safe to restart the VMs.\n**BMS:** Some Bare Metal Solution servers are experiencing high temperatures and we are in the process of shutting them down. Customers may be unable to access their systems until we resolve the issue.\n**Google Cloud NetApp Volumes** Our engineering team has proactively shut down ONTAP clusters to protect the data stored on them. Customers may experience unavailability of Cloud NetApp Volumes until this issue is resolved. There is no estimated time for resolution. We will provide an update by Wednesday, 2026-07-15 22:30 PDT with current details. We apologize to all who are affected by the disruption.\n**Customer Symptoms**\n**GCVE:** Customers would have already observed a loss of connectivity to their private cloud. As a protective measure for hardware integrity and a recovery, we have shut down all private clouds in this zone.\n**BMS:** Customers may have already lost connectivity to their BMS machines. As a protective measure for hardware integrity and a recovery, we are shutting down some servers in this zone.\n**Google Cloud NetApp Volumes** Customers would have already lost connectivity to ONTAP clusters, as we proactively shut them down for their protection. Only Standard, Premium, and Extreme service levels are affected.\n**Workaround**\nThere are no workarounds available for this issue at this time. Customers with multi-regional deployments are advised to route traffic to an alternate site.","status":"SERVICE_DISRUPTION","affected_locations":[{"title":"Netherlands (europe-west4)","id":"europe-west4"}]},{"created":"2026-07-16T03:30:59+00:00","modified":"2026-07-16T06:30:59+00:00","when":"2026-07-16T03:30:59+00:00","text":"**Summary**\nGoogle Cloud VMware Engine (GCVE), Google Cloud NetApp Volumes, and Bare Metal Solution (BMS) services are experiencing temperature alerts in europe-west4-a.\n**Description**\n**GCVE:** Our engineering teams are currently working with data center facility teams to address the rising temperatures. A shutdown has already occurred for some private clouds, and we expect that others will shut down as well. **We recommend shutting down your VMware workload VMs for precaution so that recovery is safer. A shutdown of all systems started at 19:00 PDT \\[4:00 Local Time (GMT+2)\\]** We will notify you as soon as your Private Cloud has been powered back up and it is safe to restart the VMs.\n**BMS:** Some Bare Metal Solution servers are experiencing high temperatures and we are in the process of shutting them down. Customers may be unable to access their systems until we resolve the issue.\n**Google Cloud NetApp Volumes** Our engineering team has proactively shut down ONTAP clusters to protect the data stored on them. Customers may experience unavailability of Cloud NetApp Volumes until this issue is resolved. There is no estimated time for resolution. We will provide an update by Wednesday, 2026-07-15 21:30 PDT with current details. We apologize to all who are affected by the disruption.\n**Customer Symptoms**\n**GCVE:** Customers may have already observed a loss of connectivity to their private cloud. As a protective measure for hardware integrity and a recovery, we are shutting down all Private Clouds in this zone.\n**BMS:** Customers may have already lost connectivity to their BMS machines. As a protective measure for hardware integrity and a recovery, we are shutting down servers in this zone.\n**Google Cloud NetApp Volumes** Customers may have already lost connectivity to ONTAP clusters, or will lose connectivity to them shortly, as we proactively shut them down for their protection. Only Standard, Premium, and Extreme service levels are affected.\n**Workaround**\nThere are no workarounds available for this issue at this time. Customers with multi-regional deployments are advised to route traffic to an alternate site.","status":"SERVICE_DISRUPTION","affected_locations":[{"title":"Netherlands (europe-west4)","id":"europe-west4"}]}],"most_recent_update":{"created":"2026-07-25T13:16:55+00:00","modified":"2026-07-25T13:16:55+00:00","when":"2026-07-25T13:16:55+00:00","text":"## \\# Incident Report\n## \\#\\# Summary\nOn Wednesday, 15 July 2026, Google Cloud VMware Engine (GCVE), Bare Metal Solution (BMS), and Google Cloud NetApp Volumes (GCNV) experienced service interruptions for a total duration of 14 hours, 55 minutes.\nThe root cause for this outage is a 3ms voltage drop in the power feed from the utility provider and subsequent utility breaker protective action. This incident was mitigated by the data center provider rectifying the failed systems followed by the Google engineering team restoring the services.\n## \\#\\# Root Cause\nA regional data center hosting services in europe-west4-a experienced an upstream voltage transient affecting both utility power feeds A and B. During this event, the utility breakers on both A and B feeds tripped and initiated the transfer to the back-up power source DRUPS (Diesel Rotary Uninterruptible Power Supply). The transition of side B to DRUPS system was successful without any power interruption. The DRUPS back-up power system for side A failed to take over the facility load due to electrical component failures.\nThe DRUPS failure to take over load initiated automatic transfer of the affected 3 rows from feed A to redundant power feed B. Rows 1 and 2 transferred to power feed B successfully. Row 3 failed to transfer to power feed B and experienced complete power loss due to an overload protection breaker trip. The root cause of the failed transfer on the affected row was identified as load deployment discrepancy which is under further investigation by Google engineering team. This resulted in loss of both redundant power feeds to the single row.\nDuring the voltage transient event, the server data hall experienced an increased temperature due to a cooling system failure. The chiller controller dropped offline during the voltage transient event, failing to signal the chilled water distribution pumps to restart and ultimately causing the chiller system A to shut down. The redundant source was not available due to known ongoing construction work at the facility. This resulted in the data hall temperatures reaching 44°C and subsequent shutdown of the affected data hall machines. Google engineering teams initiated machine shutdown procedures for the remainder of reachable devices as part of the cooling emergency shutdown process.\nA timeline of events during the incident is provided below.\n### DataCenter Events\n| Timestamp(PST) | Event Description |\n| :---- | :---- |\n| 07-15-2026 16:24 | An electrical fault occurred on the utility grid upstream of the data center, disrupting the electrical distribution. |\n| 07-15-2026 16:24 | DRUPS back-up power system for side A failed to take over the data hall load. |\n| 07-15-2026 16:24 | The chiller controller dropped offline during the voltage transient event causing distribution pumps to stop and unable to restart. |\n| 07-15-2026 16:37 | Rows 1 and 2 transferred to power feed B successfully. Row 3 failed to transfer to power feed B and lost power. |\n| 07-15-2026 18:29 | Data hall temperatures reached 44C and crossed the safe machine operating threshold. |\n| 07-15-2026 19:55 | Machine shutdown procedures implemented. |\n| 07-15-2026 21:05 | Cooling system fully recovered and temperatures in the data hall returned to normal operating range. |\n| 07-15-2026 21:24 | Notification about cooling recovery sent by provider |\n| 07-15-2026 21:46 | Notification about power recovery sent by provider |\n### GCVE Timelines\n| Timestamp(PST) | Event Description |\n| :---- | :---- |\n| 07-15-2026 16:30 | GCVE service power monitoring alert received |\n| 07-15-2026 16:41 | First symptom detected, servers reporting redundancy power feed failure. |\n| 07-15-2026 17:05 | Switch temperature \\\u003e60°C alerts triggered. |\n| 07-15-2026 17:09 | Start of user impact, first prober alert failure for north south traffic for customers. |\n| 07-15-2026 19:00 | Server / Network devices shutdown initiated. |\n| 07-15-2026 19:55 | All reachable server/network devices were shut down. |\n| 07-15-2026 21:24 | Cooling system fully recovered and temperatures in the datahall were returning to normal operating range. |\n| 07-15-2026 21:56 | Network restoration started with reachable hydra devices. |\n| 07-15-2026 22:40 | Onsite technician arrived, to recover the console servers. |\n| 07-15-2026 23:10 | Console server reboot/recovery complete. |\n| 07-16-2026 00:33 | Network recovery for all Placement Groups complete. |\n| 07-16-2026 02:36 | Incident mitigated, after recovering all the customer PCs are fully healthy. |\nThis outage impacted 24 private clouds of 20 distinct customers in europe-west4-a region.\n### Bare Metal Solutions (BMS) Timelines\n| Timestamp(PST) | Event Description |\n| :---- | :---- |\n| 07-15-2026 16:41 | First symptom detected, servers and storage reporting redundancy power feed failure |\n| 07-15-2026 17:52 | Automated monitoring detected rising temperatures |\n| 07-15-2026 18:37 | Four Netapp SAN storage nodes failed and were shut down due to overheating, resulting in storage availability issue |\n| 07-15-2026 18:37 | Customer impact started, first server shutdown detected |\n| 07-15-2026 18:59 | Reserve and buffer servers started to get powered off |\n| 07-15-2026 20:43 | Drop in the temperature has been detected |\n| 07-15-2026 20:52 | Four Netapp SAN storage nodes recovered, mitigating storage availability issue |\n| 07-15-2026 21:05 | The cooling system fully recovered and temperatures in the datahall were returning to normal operating range. |\n| 07-15-2026 23:23 | First customer server reboot started |\n| 07-16-2026 06:32 | Last customer server reboot started |\n| 07-16-2026 08:42 | Incident mitigated. All impacted customer servers were rebooted and confirmed in healthy state. |\nThis outage impacted 9 distinct BMS customers in europe-west4\n### NetApp Timelines\n| Timestamp(PST) | Event Description |\n| :---- | :---- |\n| 07-15-2026 16:32 | Netapp facility remote monitoring alert on high temp and network switches switches failure received |\n| 07-15-2026 17:32 | All 6 clusters failed and were shut down automatically due to high temperature |\n| 07-15-2026 21:05 | Cooling capacity restored |\n| 07-15-2026 22:08 | Cluster recovery process initiated for all 6 clusters. 1st cluster recovered |\n| 07-15-2026 23:55 | All 6 Netapp Clusters and all customers recovered and validated as healthy \u0026 operational |\n## \\#\\# Remediation and Prevention\nGoogle engineers were alerted to the infrastructure power loss on Wednesday, 15 July 2026 at 16:30 US/Pacific and subsequent increase in data hall temperatures at 17:44 US/Pacific via our automated monitoring and telemetry, and immediately began investigating.\nDuring the voltage transient event, the chillers maintained power but shut down because the distribution pumps failed to restart, resulting in no water flow. The pumps were manually switched from auto to hand mode to restore circulation and cooling. The facility team deployed an interim portable UPS to support the chiller controllers and prevent localized controller power loss.\nUpon confirmation of utility power restoration, the facility returned to utility power via automatic switching with the exception of two Remote Power Panels (RPPs), which required manual verification before restoration of power to affected server row. To provide full resiliency the facility team aligned a swing DRUPS unit to restore the redundant configuration on the A feed and is working on repairing the faulty DRUPS.\nTo remediate the power loss at Row 3, once the electrical fault was confirmed to be fully isolated, teams reset the tripped breakers, successfully restoring both redundant power to the impacted server racks.\nGoogle is committed to preventing a repeat of this issue in the future and is completing the following actions:\n* Detailed investigation by the engineering team of the sequence of events that prevented transfer of load to DRUPS system and identified system improvements **\\[ETA Aug 2026\\].**\n* Determine and resolve the cause of overload conditions that caused power loss to Row 3 **\\[ETA Aug 2026\\].**\n* Perform investigation of chiller pump control system redundancy setup to ensure cooling system resiliency **\\[ETA Sep 2026\\].**\n* Implement improvements to facility system monitoring alerts for power events and configure early alert thresholds for cooling excursion detection. **\\[ETA Aug 2026\\].**\n* The GCVE engineering team is working to improve the automated triggering of shutdown during quickly evolving thermal runoff situations **\\[ETA Oct 2026\\]**.\n* The BMS engineering team is working with partner teams to update the incident classification and SLO definitions to improve personnel availability during major datacenter events **\\[ETA Sep 2026\\]**.\n* The GCNV engineering team is collaborate with the data center team to review incident handling SLA and runbooks, identify potential future occurrences and create playbooks to address/mitigate them, establish monthly joint emergency drill, and review if the Netapp cluster recovery process can be expedited **\\[ETA Aug 2026\\].**\n## \\#\\# Detailed Description of Impact\nOn 15 July 2026 from 16:39 to 16 July 2026 07:34 US/Pacific, customers experienced service disruptions across several products in europe-west4-a zone:\n* **Google Cloud VMware Engine**: Customers lost connectivity to their private clouds and workloads as foundational hosts and network switches were powered off for thermal protection.\n* **Bare Metal Solution**: Customers experienced connectivity loss to their database servers and storage appliances as the underlying hardware was powered down.\n* **Google Cloud NetApp Volumes**: Customers in the STANDARD, PREMIUM, and EXTREME service levels were unable to access their storage volumes. Control plane operations, including the creation of new storage pools, volumes, and backups, experienced failures in the affected region.","status":"AVAILABLE","affected_locations":[]},"status_impact":"SERVICE_DISRUPTION","severity":"medium","service_key":"zall","service_name":"Multiple Products","affected_products":[{"title":"Bare Metal Solution","id":"5gQF7Qk9XfQ32pStX6Hb"},{"title":"Google Cloud NetApp Volumes","id":"vtzMUyQ4z9CbA1x6z85s"},{"title":"VMWare engine","id":"9H6gWUHvb2ZubeoxzQ1Y"}],"uri":"incidents/3BvH3LVGcupoYqV6F4Nw","currently_affected_locations":[],"previously_affected_locations":[{"title":"Netherlands (europe-west4)","id":"europe-west4"}]},{"id":"T8gmtofFSTGT5tbhyciF","number":"5362718329524538692","begin":"2026-07-14T17:00:00+00:00","created":"2026-07-14T20:24:18+00:00","end":"2026-07-15T03:40:00+00:00","modified":"2026-07-24T07:14:47+00:00","external_desc":"Google Cloud VMware Engine (GCVE) Stretched Cluster customers are experiencing zonal outages impacting network connectivity across multiple regions.","updates":[{"created":"2026-07-24T07:10:18+00:00","modified":"2026-07-24T07:14:47+00:00","when":"2026-07-24T07:10:18+00:00","text":"## Incident Report\n## Summary\nOn Tuesday, 14 July 2026 10:00 PT, Google Cloud VMware Engine (GCVE) Stretched Cluster customers in the australia-southeast2 and europe-west3 zones experienced inter-site communication failures.\nThe disruption was traced to a network configuration update that introduced a conflict, causing inter-site communication failures, which triggered VMware Stretched Cluster failover events. Google engineers successfully mitigated the impact by rolling back the issue-causing configuration change and resetting the network state.\nWe sincerely apologize for the disruption this incident caused to your business. We know how much you rely on Google Cloud, and we regret the impact on your productivity. We are working to address the root cause and prevent this from occurring in the future.\n## Root Cause\nA network configuration update intended to prepare the cloud network infrastructure for new capabilities was deployed to the foundational network control plane. While the configuration payload itself was structurally valid, it exposed an implementation gap within the control plane's routing logic.\nThis logical gap caused the underlying network hosts to misconfigure program routing tables. As a result, traffic destined for the private IP address space used by GCVE Stretched Clusters was dropped. Because Stretched Clusters rely on this private address space to establish routing sessions for inter-zonal connectivity, the traffic drops severed communications between the active zones of the clusters, triggering VMware High Availability (HA) failovers.\nStandard routing health-checking protocols, such as Border Gateway Protocol (BGP) and Bidirectional Forwarding Detection (BFD), remained fully functional because their control plane sessions run on separate, unaffected address spaces. Because the control plane remained healthy, standard failover mechanisms failed to detect that the selective private IP range used for inter-zonal data tunneling was being dropped. Automated safeguards did not block the deployment because the configuration passed initial payload validations. The impact only manifested once the update began routing data traffic through the specific affected IP range.\n## Remediation and Prevention\nTo stabilize the environment, engineers identified the working network paths and deployed configuration changes to reroute traffic and restore connectivity. GCVE Stretched Cluster inter-zonal connectivity was completely restored for all supported locations on Tuesday, 14 July 2026 at 20:40 US/Pacific.\nGoogle is committed preventing a repeat of this issue in the future and is completing the following actions:\n* **Expanded Testing:** We are adding more detailed GCVE network setups to our existing testing environments. This allows us to automatically test future network updates against these configurations before they go live.\n* **Service-Level Data Path Failover:** We are implementing additional service-level data path failover mechanisms that actively probe the specific data-tunneling traffic space. This will ensure a path failover is triggered if the data plane itself is degraded even when the BGP control plane remains functional.\n* **Detailed Alerting**: We are adding faster, more specific alerts for connection issues between zones. This builds on our current platform monitoring to catch minor disruptions early and speed up our response.\n* **Improved Cluster Resilience:** We are fine-tuning the cluster's high-availability and storage settings. This makes virtual machines more resilient to short network drops, preventing them from restarting unnecessarily if the main site is still healthy.\n* **Workload Resilience Alignment (Shared Responsibility):** We are proactively reaching out to customers utilizing non-vSAN replicated virtual machine configurations within Stretched Clusters. Because these workloads are pinned to a single zone without active cross-site replication, they cannot survive inter-site network disruptions. We are ready to assist customers in auditing their storage policies, adjusting Stretched Cluster configurations, and planning the secondary zone capacity required to enable robust high-availability failovers.\n## Detailed Description of Impact\nOn Tuesday, 14 July 2026 from 10:00 to 20:40 US/Pacific, customers utilizing GCVE Stretched Clusters in the affected zones (australia-southeast2, and europe-west3) experienced inter-site communication failures. For some customers, depending on their architecture, this disruption led to a VMWare HA event causing VM restarts/movement across zones as designed, host disconnects, VSAN alarms and intermittent access to VMware Management (vCenter/NSX Manager) components.","status":"AVAILABLE","affected_locations":[]},{"created":"2026-07-20T18:31:26+00:00","modified":"2026-07-24T07:10:18+00:00","when":"2026-07-20T18:31:26+00:00","text":"## Preliminary Incident Report\nWe sincerely apologize for the disruption this incident caused to your business. We know how much you rely on Google Cloud, and we regret the impact on your productivity. We are working to address the root cause and prevent this from occurring in the future.\nPlease note, this information is based on our best knowledge at the time of posting and is subject to change as our investigation continues. A final Incident Report with preventative actions will be posted once our investigation is complete.\nIf you have experienced impact outside of what is listed below, please reach out to Google Cloud Support using [**https://cloud.google.com/support**](https://cloud.google.com/support).\n## Date/Time of the Issue (All time US/Pacific)\nIncident Start: 14 July 2026 10:00\nIncident End: 14 July 2026 20:40\nDuration: 10 hours, 40 minutes\n## Summary\nOn Tuesday, 14 July 2026, Google Cloud VMware Engine (GCVE) Stretched Cluster customers in the australia-southeast2 and europe-west3 zones experienced inter-site communication failures.\nThe disruption was traced to a network configuration update that introduced a conflict, causing inter-site communication failures, which triggered VMware Stretch Cluster failover events. Google engineers successfully mitigated the disruption by rolling back the configuration change.\n## Preliminary Root Cause\nA configuration change deployed within the Google Cloud network affected traffic between the VMware Engine stretched cluster zones. This resulted in a loss of connectivity between zones in the stretched cluster deployment.\nGoogle engineers are doing a full root cause analysis and will provide additional information once it is available.\n## Remediation\nTo stabilize the environment, engineers identified the working network paths and deployed configuration changes to reroute traffic and restore connectivity. GCVE stretch cluster inter-zonal connectivity was completely restored for all supported locations on Tuesday, 14 July 2026 at 20:40 US/Pacific.\n## Description of Impact\nOn Tuesday, 14 July 2026 from 10:00 to 20:40 US/Pacific, customers utilizing GCVE stretch clusters in the affected zones (australia-southeast2, and europe-west3) experienced inter-site communication failures. For some customers, depending on their architecture, this disruption led to VMWare HA event causing VM restarts/movement across zones as designed, Host disconnects, VSAN alarms and Intermittent access to VMware Management (vCenter/NSX Manager) components.\n---","status":"AVAILABLE","affected_locations":[]},{"created":"2026-07-15T05:34:26+00:00","modified":"2026-07-20T18:31:26+00:00","when":"2026-07-15T05:34:26+00:00","text":"**Summary**\nGoogle Cloud VMware Engine (GCVE) Stretched Cluster customers may have experienced zonal outages impacting network connectivity across multiple regions.\n**Description**\nWe experienced an inter-site communication issue with Google Cloud VMware Engine (GCVE) Stretched Cluster customers beginning Tuesday, 2026-07-14 at 10:00 PDT. We identified the affected customers, and have worked with them to fully mitigate the issue by Tuesday, 2026-07-14 at 21:46 PDT.\nPreliminary analysis indicates that a network configuration update was the cause of the inter-zone network disruption. Our engineering team mitigated the issue by rolling back the faulty configuration to its last-known good value.\nIf customers are still experiencing impact from this issue, please contact Support, and we will work with you to resolve any residual impact.\nWe thank you for your patience while we worked to resolve this issue.\n**Customer Symptoms**\nSome GCVE customers using Stretched Cluster may have experienced inter-site communication failures to their GCVE environments within the affected zones.\n**Workaround**\nThe issue is now mitigated.","status":"AVAILABLE","affected_locations":[]},{"created":"2026-07-15T03:27:54+00:00","modified":"2026-07-15T05:34:26+00:00","when":"2026-07-15T03:27:54+00:00","text":"**Summary**\nGoogle Cloud VMware Engine (GCVE) Stretched Cluster customers are experiencing zonal outages impacting network connectivity across multiple regions.\n**Description**\nWe are experiencing an inter-site communication issue with Google Cloud VMware Engine (GCVE) stretched cluster customers beginning Tuesday, 2026-07-14 at 10:00 PDT. We have identified the customers that are affected by this issue and we are working with them on mitigation.\nA remediation rollout is currently in progress to address the underlying network issue.\n**Failover / VM Migration:** Migrating affected VMs to the healthy, secondary zone of your stretch cluster remains the primary mitigation strategy. Because of the complexities surrounding failover risks and secondary zone health, we highly encourage you to open a ticket with Google Cloud Support if you are severely impacted.\nWe do not currently have an ETA for resolution. We will provide another update by Tuesday, 2026-07-14 22:00 PDT with current details.\n**Customer Symptoms**\nSome GCVE customers using Stretched Cluster may experience inter-site communication failures to their GCVE environments within the affected zones.\n**Workaround**\nWhile we work on restoring full connectivity, we recommend the following to restore access to your workloads:\n* VM Migration (recommended): Where possible, migrate your affected VMs to the healthy and unaffected side of the stretch cluster. We strongly recommend consulting with Google Cloud Support before proceeding.","status":"SERVICE_DISRUPTION","affected_locations":[{"title":"Sydney (australia-southeast1)","id":"australia-southeast1"},{"title":"Melbourne (australia-southeast2)","id":"australia-southeast2"},{"title":"Frankfurt (europe-west3)","id":"europe-west3"}]},{"created":"2026-07-15T01:55:24+00:00","modified":"2026-07-15T03:27:54+00:00","when":"2026-07-15T01:55:24+00:00","text":"**Summary**\nGoogle Cloud VMware Engine (GCVE) Stretched Cluster customers are experiencing zonal outages impacting network connectivity across multiple regions.\n**Description**\nWe are experiencing an inter-site communication issue with Google Cloud VMware Engine (GCVE) stretched cluster customers beginning Tuesday, 2026-07-14 at 10:00 PDT. We have identified the customers that are affected by this issue and we are working with them on mitigation.\nA remediation rollout is currently in progress to address the underlying network issue.\n**Failover / VM Migration:** Migrating affected VMs to the healthy, secondary zone of your stretch cluster remains the primary mitigation strategy. Because of the complexities surrounding failover risks and secondary zone health, we highly encourage you to open a ticket with Google Cloud Support if you are severely impacted.\nWe do not currently have an ETA for resolution. We will provide another update by Tuesday, 2026-07-14 20:30 PDT with current details.\n**Customer Symptoms**\nSome GCVE customers using Stretched Cluster may experience inter-site communication failures to their GCVE environments within the affected zones.\n**Workaround**\nWhile we work on restoring full connectivity, we recommend the following workarounds to restore access to your workloads:\n- VM Migration (recommended): Where possible, migrate your affected VMs to the healthy and unaffected side of the stretch cluster. We strongly recommend consulting with Google Support before proceeding.","status":"SERVICE_DISRUPTION","affected_locations":[{"title":"Sydney (australia-southeast1)","id":"australia-southeast1"},{"title":"Melbourne (australia-southeast2)","id":"australia-southeast2"},{"title":"Frankfurt (europe-west3)","id":"europe-west3"}]},{"created":"2026-07-15T01:15:28+00:00","modified":"2026-07-15T01:55:24+00:00","when":"2026-07-15T01:15:28+00:00","text":"**Summary**\nGoogle Cloud VMware Engine (GCVE) Stretched Cluster customers are experiencing zonal outages impacting network connectivity across multiple regions.\n**Description**\nWe are experiencing an inter-site communication issue with Google Cloud VMware Engine (GCVE) stretched cluster customers beginning Tuesday, 2026-07-14 at 10:00 PDT. We have identified the customers that are affected by this issue and we are working with them on mitigation.\nA remediation rollout is currently in progress to address the underlying network issue.\n**Failover / VM Migration:** Migrating affected VMs to the healthy, secondary zone of your stretch cluster remains the primary mitigation strategy. Because of the complexities surrounding failover risks and secondary zone health, we highly encourage you to open a ticket with Google Cloud Support if you are severely impacted.\nWe do not currently have an ETA for resolution. We will provide another update by Tuesday, 2026-07-14 19:00 PDT with current details.\n**Customer Symptoms**\nSome GCVE customers using Stretched Cluster may experience inter-site communication failures to their GCVE environments within the affected zones.\n**Workaround**\nWhile we work on restoring full connectivity, we recommend the following workarounds to restore access to your workloads:\n- VM Migration (recommended): Where possible, migrate your affected VMs to the healthy and unaffected side of the stretch cluster. We strongly recommend consulting with Google Support before proceeding.","status":"SERVICE_DISRUPTION","affected_locations":[{"title":"Sydney (australia-southeast1)","id":"australia-southeast1"},{"title":"Melbourne (australia-southeast2)","id":"australia-southeast2"},{"title":"Frankfurt (europe-west3)","id":"europe-west3"}]},{"created":"2026-07-15T00:13:18+00:00","modified":"2026-07-15T01:15:28+00:00","when":"2026-07-15T00:13:18+00:00","text":"**Summary**\nGoogle Cloud VMware Engine (GCVE) Stretched Cluster customers are experiencing zonal outages impacting network connectivity across multiple regions.\n**Description**\nWe are experiencing an inter-site communication issue with Google Cloud VMware Engine (GCVE) stretched cluster customers beginning Tuesday, 2026-07-14 at 10:00 PDT. We have identified the customers that are affected by this issue and we are working with them on mitigation.\nA remediation rollout is currently in progress to address the underlying network issue.\n**Failover / VM Migration:** Migrating affected VMs to the healthy, secondary zone of your stretch cluster remains the primary mitigation strategy. Because of the complexities surrounding failover risks and secondary zone health, we highly encourage you to open a ticket with Google Cloud Support if you are severely impacted.\nWe do not currently have an ETA for resolution. We will provide another update by Tuesday, 2026-07-14 18:00 PDT with current details.\n**Customer Symptoms**\nSome GCVE customers using Stretched Cluster may experience inter-site communication failures to their GCVE environments within the affected zones.\nUpon further investigation we identified that the northamerica-northeast2 region was not impacted.\n**Workaround**\nWhile we work on restoring full connectivity, we recommend the following workarounds to restore access to your workloads:\n- VM Migration (recommended): Where possible, migrate your affected VMs to the healthy and unaffected side of the stretch cluster. We strongly recommend consulting with Google Support before proceeding.","status":"SERVICE_DISRUPTION","affected_locations":[{"title":"Sydney (australia-southeast1)","id":"australia-southeast1"},{"title":"Melbourne (australia-southeast2)","id":"australia-southeast2"},{"title":"Frankfurt (europe-west3)","id":"europe-west3"},{"title":"Toronto (northamerica-northeast2)","id":"northamerica-northeast2"}]},{"created":"2026-07-14T23:05:33+00:00","modified":"2026-07-15T00:13:18+00:00","when":"2026-07-14T23:05:33+00:00","text":"**Summary**\nGoogle Cloud VMware Engine (GCVE) Stretched Cluster customers are experiencing zonal outages impacting network connectivity across multiple regions.\n**Description**\nWe are experiencing an inter-site communication issue with Google Cloud VMware Engine (GCVE) stretched cluster customers beginning Tuesday, 2026-07-14 at 10:00 PDT. We have identified the customers that are affected by this issue and we are working with them on mitigation.\nOur investigation has identified a recent configuration update that is the likely cause of the inter-zone network disruption. Teams are working on remediation.\n**Failover / VM Migration:** Migrating affected VMs to the healthy, secondary zone of your stretch cluster remains the primary mitigation strategy. Because of the complexities surrounding failover risks and secondary zone health, we highly encourage you to open a ticket with Google Cloud Support if you are severely impacted.\nWe do not currently have an ETA for resolution. We will provide another update by Tuesday, 2026-07-14 17:00 PDT with current details.\n**Customer Symptoms**\nSome GCVE customers using Stretched Cluster may experience inter-site communication failures to their GCVE environments within the affected zones.\n**Workaround**\nWhile we work on restoring full connectivity, we recommend the following workarounds to restore access to your workloads:\n- VM Migration (recommended): Where possible, migrate your affected VMs to the healthy and unaffected side of the stretch cluster. We strongly recommend consulting with Google Support before proceeding.","status":"SERVICE_DISRUPTION","affected_locations":[{"title":"Sydney (australia-southeast1)","id":"australia-southeast1"},{"title":"Melbourne (australia-southeast2)","id":"australia-southeast2"},{"title":"Frankfurt (europe-west3)","id":"europe-west3"},{"title":"Toronto (northamerica-northeast2)","id":"northamerica-northeast2"}]},{"created":"2026-07-14T22:16:55+00:00","modified":"2026-07-14T23:05:33+00:00","when":"2026-07-14T22:16:55+00:00","text":"**Summary**\nGoogle Cloud VMware Engine (GCVE) Stretched Cluster customers are experiencing zonal outages impacting network connectivity across multiple regions.\n**Description**\nWe are experiencing an inter-site communication issue with Google Cloud VMware Engine (GCVE) stretched cluster customers beginning Tuesday, 2026-07-14 at 10:00 PDT. We have identified the customers that are affected by this issue and we are working with them on mitigation.\nOur investigation has identified underlying inter-zone communication failures and Border Gateway Protocol (BGP) session flapping between cluster zones. Specifically, network connectivity has been lost between the affected zones and the witness appliance. Because the witness appliance is currently unreachable, the cluster zones are unable to safely synchronize state. As a result, VMs on the affected sites are becoming isolated and may be left without writable data.\n**Failover / VM Migration:** Migrating affected VMs to the healthy, secondary zone of your stretch cluster remains the primary mitigation strategy. Because of the complexities surrounding failover risks and secondary zone health, we highly encourage you to open a ticket with Google Cloud Support if you are severely impacted.\nWe do not currently have an ETA for resolution. We will provide another update by Tuesday, 2026-07-14 16:00 PDT with current details.\n**Customer Symptoms**\nSome GCVE customers using Stretched Cluster may experience inter-site communication failures to their GCVE environments within the affected zones.\n**Workaround**\nWhile we work on restoring full connectivity, we recommend the following workarounds to restore access to your workloads:\n- VM Migration (recommended): Where possible, migrate your affected VMs to the healthy and unaffected side of the stretch cluster. We strongly recommend consulting with Google Support before proceeding.","status":"SERVICE_DISRUPTION","affected_locations":[{"title":"Sydney (australia-southeast1)","id":"australia-southeast1"},{"title":"Melbourne (australia-southeast2)","id":"australia-southeast2"},{"title":"Frankfurt (europe-west3)","id":"europe-west3"},{"title":"Toronto (northamerica-northeast2)","id":"northamerica-northeast2"}]},{"created":"2026-07-14T21:31:52+00:00","modified":"2026-07-14T22:16:55+00:00","when":"2026-07-14T21:31:52+00:00","text":"**Summary**\nGoogle Cloud VMware Engine (GCVE) Stretched Cluster customers are experiencing zonal outages impacting network connectivity across multiple regions.\n**Description**\nWe are experiencing an inter-site communication issue with Google Cloud VMware Engine (GCVE) stretched cluster customers beginning Tuesday, 2026-07-14 at 10:00 PDT. We have identified the customers that are affected by this issue and we are working with them on mitigation.\nOur preliminary investigation indicates this is stemming from an underlying network connectivity issue affecting the infrastructure that links the zones within a stretch cluster. This disruption is causing synchronization issues between the affected zones. We believe Storage and Compute services remain unaffected, and VMs are running as expected, though connectivity to them may be degraded.\nWe will provide another update by Tuesday, 2026-07-14 15:00 PDT with current details.\n**Customer Symptoms**\nSome GCVE customers using Stretched Cluster may experience inter-site communication failures to their GCVE environments within the affected zones.\n**Workaround**\nWhile we work on restoring full connectivity, we recommend the following workarounds to restore access to your workloads:\n- VM Migration (recommended): Where possible, migrate your affected VMs to the healthy and unaffected side of the stretch cluster. We strongly recommend consulting with Google Support before proceeding.","status":"SERVICE_DISRUPTION","affected_locations":[{"title":"Melbourne (australia-southeast2)","id":"australia-southeast2"},{"title":"Frankfurt (europe-west3)","id":"europe-west3"},{"title":"Toronto (northamerica-northeast2)","id":"northamerica-northeast2"}]},{"created":"2026-07-14T20:24:18+00:00","modified":"2026-07-14T21:31:52+00:00","when":"2026-07-14T20:24:18+00:00","text":"**Summary**\nGoogle Cloud VMware Engine (GCVE) Stretched Cluster customers are experiencing zonal outages impacting network connectivity across multiple regions.\n**Description**\nWe are experiencing an issue with Google Cloud VMware Engine (GCVE) beginning Tuesday, 2026-07-14 at 10:00 PDT.\nWe are continuing to investigate and mitigate the issue. We believe this disruption is isolated to network connectivity issues for stretch clusters, while Storage and Compute services appear to be unaffected. GCVE VMs are running as expected but customers may experience connectivity issues to the VMs.\nWe will provide another update by Tuesday, 2026-07-14 14:00 PDT with current details.\n**Customer Symptoms**\nSome GCVE customers may experience inter-site communication failures to their GCVE environments within the affected zones.\n**Workaround**\nNone at this time.","status":"SERVICE_DISRUPTION","affected_locations":[{"title":"Melbourne (australia-southeast2)","id":"australia-southeast2"},{"title":"Frankfurt (europe-west3)","id":"europe-west3"},{"title":"Toronto (northamerica-northeast2)","id":"northamerica-northeast2"}]}],"most_recent_update":{"created":"2026-07-24T07:10:18+00:00","modified":"2026-07-24T07:14:47+00:00","when":"2026-07-24T07:10:18+00:00","text":"## Incident Report\n## Summary\nOn Tuesday, 14 July 2026 10:00 PT, Google Cloud VMware Engine (GCVE) Stretched Cluster customers in the australia-southeast2 and europe-west3 zones experienced inter-site communication failures.\nThe disruption was traced to a network configuration update that introduced a conflict, causing inter-site communication failures, which triggered VMware Stretched Cluster failover events. Google engineers successfully mitigated the impact by rolling back the issue-causing configuration change and resetting the network state.\nWe sincerely apologize for the disruption this incident caused to your business. We know how much you rely on Google Cloud, and we regret the impact on your productivity. We are working to address the root cause and prevent this from occurring in the future.\n## Root Cause\nA network configuration update intended to prepare the cloud network infrastructure for new capabilities was deployed to the foundational network control plane. While the configuration payload itself was structurally valid, it exposed an implementation gap within the control plane's routing logic.\nThis logical gap caused the underlying network hosts to misconfigure program routing tables. As a result, traffic destined for the private IP address space used by GCVE Stretched Clusters was dropped. Because Stretched Clusters rely on this private address space to establish routing sessions for inter-zonal connectivity, the traffic drops severed communications between the active zones of the clusters, triggering VMware High Availability (HA) failovers.\nStandard routing health-checking protocols, such as Border Gateway Protocol (BGP) and Bidirectional Forwarding Detection (BFD), remained fully functional because their control plane sessions run on separate, unaffected address spaces. Because the control plane remained healthy, standard failover mechanisms failed to detect that the selective private IP range used for inter-zonal data tunneling was being dropped. Automated safeguards did not block the deployment because the configuration passed initial payload validations. The impact only manifested once the update began routing data traffic through the specific affected IP range.\n## Remediation and Prevention\nTo stabilize the environment, engineers identified the working network paths and deployed configuration changes to reroute traffic and restore connectivity. GCVE Stretched Cluster inter-zonal connectivity was completely restored for all supported locations on Tuesday, 14 July 2026 at 20:40 US/Pacific.\nGoogle is committed preventing a repeat of this issue in the future and is completing the following actions:\n* **Expanded Testing:** We are adding more detailed GCVE network setups to our existing testing environments. This allows us to automatically test future network updates against these configurations before they go live.\n* **Service-Level Data Path Failover:** We are implementing additional service-level data path failover mechanisms that actively probe the specific data-tunneling traffic space. This will ensure a path failover is triggered if the data plane itself is degraded even when the BGP control plane remains functional.\n* **Detailed Alerting**: We are adding faster, more specific alerts for connection issues between zones. This builds on our current platform monitoring to catch minor disruptions early and speed up our response.\n* **Improved Cluster Resilience:** We are fine-tuning the cluster's high-availability and storage settings. This makes virtual machines more resilient to short network drops, preventing them from restarting unnecessarily if the main site is still healthy.\n* **Workload Resilience Alignment (Shared Responsibility):** We are proactively reaching out to customers utilizing non-vSAN replicated virtual machine configurations within Stretched Clusters. Because these workloads are pinned to a single zone without active cross-site replication, they cannot survive inter-site network disruptions. We are ready to assist customers in auditing their storage policies, adjusting Stretched Cluster configurations, and planning the secondary zone capacity required to enable robust high-availability failovers.\n## Detailed Description of Impact\nOn Tuesday, 14 July 2026 from 10:00 to 20:40 US/Pacific, customers utilizing GCVE Stretched Clusters in the affected zones (australia-southeast2, and europe-west3) experienced inter-site communication failures. For some customers, depending on their architecture, this disruption led to a VMWare HA event causing VM restarts/movement across zones as designed, host disconnects, VSAN alarms and intermittent access to VMware Management (vCenter/NSX Manager) components.","status":"AVAILABLE","affected_locations":[]},"status_impact":"SERVICE_DISRUPTION","severity":"medium","service_key":"9H6gWUHvb2ZubeoxzQ1Y","service_name":"VMWare engine","affected_products":[{"title":"VMWare engine","id":"9H6gWUHvb2ZubeoxzQ1Y"}],"uri":"incidents/T8gmtofFSTGT5tbhyciF","currently_affected_locations":[],"previously_affected_locations":[{"title":"Melbourne (australia-southeast2)","id":"australia-southeast2"},{"title":"Frankfurt (europe-west3)","id":"europe-west3"}]},{"id":"5fGQt4VbkDnr3Yp8PXPr","number":"1464522124090870782","begin":"2026-06-05T07:00:00+00:00","created":"2026-06-10T01:13:50+00:00","end":"2026-06-26T19:00:00+00:00","modified":"2026-06-29T23:06:10+00:00","external_desc":"Network traffic to Google Cloud originating from Delhi, Chennai, Mumbai and surrounding areas is experiencing intermittent periods of elevated latency and possible packet loss.","updates":[{"created":"2026-06-29T23:06:10+00:00","modified":"2026-06-29T23:06:10+00:00","when":"2026-06-29T23:06:10+00:00","text":"**Summary**\nNetwork traffic to Google Cloud originating from Delhi, Chennai, Mumbai and surrounding areas experienced intermittent periods of elevated latency and possible packet loss.\n**Description**\nTraffic rerouting from the impacted Delhi facility caused a subset of Hybrid Connectivity, Virtual Private Cloud (VPC) and Media CDN customers to experience intermittent latency spikes as demand exceeded regional capacity.\nWe completed the augmentation of out-of-region Internet Edge regional peering capacity in Chennai to provide additional load-balancing and redundancy to large ISPs in India. Service to a large portion of Internet Edge peering capacity has been restored to reduce latency in the local Delhi metropolitan area. We are now recovered and returned to normal service as of Friday, 2026-06-26 PDT.\nFollowing safety clearance, our team restored all lost capacity. We have restored capacity between Delhi-Chennai and Delhi-Mumbai and will continue to closely monitor latency deviations and packet drops.\nWe thank you for your patience during the resolution of this issue.\n**Symptoms**\nThe impacted customers may have experienced slightly elevated latency and non-optimal network routing into Google Cloud.\n**Workaround**\nNone","status":"AVAILABLE","affected_locations":[]},{"created":"2026-06-23T22:52:49+00:00","modified":"2026-06-29T23:06:10+00:00","when":"2026-06-23T22:52:49+00:00","text":"**Summary**\nNetwork traffic to Google Cloud originating from Delhi, Chennai, Mumbai and surrounding areas is experiencing intermittent periods of elevated latency and possible packet loss.\n**Description**\nTraffic rerouting from the impacted Delhi facility has caused a subset of Hybrid Connectivity and Virtual Private Cloud (VPC) customers to experience intermittent latency spikes as demand exceeds regional capacity. Media CDN customers may also notice increased latencies.\nInitial traffic mitigations have yielded positive results for some Cloud customers. We have restored a large portion of Internet Edge peering capacity to reduce latency in the local Delhi metropolitan area. Since the last update, we have completed the augmentation of out-of-region Internet Edge regional peering capacity in Chennai to provide additional load-balancing and redundancy to large ISPs in India.\nFollowing safety clearance, our team has accessed the damaged site and is continuing to restore additional capacity throughout this week. We have optimized capacity across network backbones to increase available headroom and augmented our Delhi user-facing backbone capacity. We will continue to closely monitor latency deviations and packet drops.\nWe will provide our next update by Monday, 2026-06-29 at 17:00 PDT.\n**Symptoms**\nCustomers may experience slightly elevated latency and non-optimal network routing into Google Cloud until the affected facility is fully restored.\n**Workaround**\nThere is no workaround at this time.","status":"SERVICE_DISRUPTION","affected_locations":[{"title":"Delhi (asia-south2)","id":"asia-south2"},{"title":"Global","id":"global"}]},{"created":"2026-06-22T23:51:23+00:00","modified":"2026-06-23T22:52:49+00:00","when":"2026-06-22T23:51:23+00:00","text":"**Summary**\nNetwork traffic to Google Cloud originating from Delhi, Chennai, Mumbai and surrounding areas is experiencing intermittent periods of elevated latency and possible packet loss.\n**Description**\nTraffic rerouting from the impacted Delhi facility has caused a subset of Hybrid Connectivity and Virtual Private Cloud (VPC) customers to experience intermittent latency spikes as demand exceeds regional capacity. Media CDN customers may also notice increased latencies.\nInitial traffic mitigations have yielded positive results for some Cloud customers. We have restored a large portion of Internet Edge peering capacity to reduce latency in the local Delhi metropolitan area. Since the last update, we have completed the augmentation of out-of-region Internet Edge regional peering capacity in Chennai to provide additional load-balancing and redundancy to large ISPs in India.\nFollowing safety clearance, teams obtained access to the damaged site and will further restore additional incremental user-facing backbone capacity on Tuesday, 2026-06-23. We have optimized capacity across network backbones to increase available headroom, and augmented our Delhi user-facing backbone capacity. We will continue to closely monitor latency deviations and packet drops.\nWe will provide our next update by Tuesday, 2026-06-23 at 17:00 PDT.\n**Symptoms**\nCustomers may experience slightly elevated latency and non-optimal network routing into Google Cloud until the affected facility is fully restored.\n**Workaround**\nThere is no workaround at this time.","status":"SERVICE_DISRUPTION","affected_locations":[{"title":"Delhi (asia-south2)","id":"asia-south2"},{"title":"Global","id":"global"}]},{"created":"2026-06-17T22:36:07+00:00","modified":"2026-06-22T23:51:23+00:00","when":"2026-06-17T22:36:07+00:00","text":"**Summary**\nNetwork traffic to Google Cloud originating from Delhi, Chennai, Mumbai and surrounding areas is experiencing intermittent periods of elevated latency and possible packet loss.\n**Description**\nTraffic rerouting from the impacted Delhi facility has caused a subset of Hybrid Connectivity and Virtual Private Cloud (VPC) customers to experience intermittent latency spikes as demand exceeds regional capacity. Media CDN customers may also notice increased latencies.\nInitial traffic mitigations have yielded positive results for some Cloud customers. We have restored a portion of Internet Edge peering capacity to reduce latency in the local Delhi metropolitan area. Further, we are nearly finished with the augmentation of out-of-region Internet Edge regional peering capacity in Chennai to provide additional load-balancing and redundancy to large ISPs in India.\nWe have optimized capacity across network backbones to increase available headroom. Additionally, we have augmented our Delhi backbone capacity. We have also restored additional capacity between Delhi-Chennai and Delhi-Mumbai. We will continue to closely monitor latency deviations and packet drops.\nWe will provide our next update by Monday, 2026-06-22 17:00 PDT.\n**Symptoms**\nCustomers may experience slightly elevated latency and non-optimal network routing into Google Cloud until the affected facility is fully restored.\n**Workaround**\nThere is no workaround at this time.","status":"SERVICE_DISRUPTION","affected_locations":[{"title":"Delhi (asia-south2)","id":"asia-south2"},{"title":"Global","id":"global"}]},{"created":"2026-06-15T23:11:38+00:00","modified":"2026-06-17T22:36:07+00:00","when":"2026-06-15T23:11:38+00:00","text":"**Summary**\nNetwork traffic to Google Cloud originating from Delhi, Chennai, Mumbai and surrounding areas is experiencing intermittent periods of elevated latency and possible packet loss.\n**Description**\nA fire at a third-party data center facility required an emergency power shutdown of networking equipment, isolating a non-compute local Point of Presence (POP) in Delhi and reducing available network capacity in the metro area.\nWe rerouted significant traffic from the impacted facility in Delhi to address reduced local serving capabilities. As a result, a subset of Hybrid Connectivity and Virtual Private Cloud (VPC) customers may be impacted by the routing changes made to address reduced local, latency-optimized serving capabilities in Delhi. Affected customers may experience intermittent latency spikes due to demand exceeding capacity across Indian metros and regional ISPs. Media CDN customers may have experienced higher latencies than normal.\nInitial traffic mitigations have yielded positive results for some Cloud customers. We have restored a portion of Internet Edge peering capacity to reduce latency in the local Delhi metropolitan area. Further, we are augmenting out-of-region Internet Edge regional peering capacity in Chennai to provide additional load-balancing and redundancy to large ISPs in India (expected to be done by Wednesday, 2026-06-17 PDT). We have optimized capacity across network backbones to increase available headroom. Additionally, we have augmented our Delhi backbone capacity over the weekend. We will continue to closely monitor latency deviations and packet drops.\nWe will provide our next update by Wednesday, 2026-06-17 at 17:00 PDT.\n**Symptoms**\nCustomers may experience slightly elevated latency and non-optimal network routing into Google Cloud until the affected facility is fully restored.\n**Workaround**\nThere is no workaround at this time.","status":"SERVICE_DISRUPTION","affected_locations":[{"title":"Delhi (asia-south2)","id":"asia-south2"},{"title":"Global","id":"global"}]},{"created":"2026-06-12T22:50:20+00:00","modified":"2026-06-15T23:11:38+00:00","when":"2026-06-12T22:50:20+00:00","text":"**Summary**\nNetwork traffic to Google Cloud originating from Delhi, Chennai, Mumbai and surrounding areas is experiencing intermittent periods of elevated latency and possible packet loss.\n**Description**\nA fire at a third-party data center facility required an emergency power shutdown of networking equipment, isolating a non-compute local Point of Presence (POP) in Delhi and reducing available network capacity in the metro area.\nWe rerouted significant traffic from the impacted facility in Delhi to address reduced local serving capabilities. As a result, a subset of Hybrid Connectivity and Virtual Private Cloud (VPC) customers may be impacted by the routing changes made to address reduced local, latency-optimized serving capabilities in Delhi. Affected customers may experience intermittent latency spikes due to demand exceeding capacity across Indian metros and regional ISPs.\nInitial traffic mitigations have yielded positive results for some Cloud customers. In parallel, we are pursuing additional Internet Edge peering capacity to reduce latency in the local Delhi metropolitan area. Further, we are augmenting out-of-region Internet Edge regional peering capacity in Chennai to provide additional load-balancing and redundancy to large ISPs in India (expected to be done by Wednesday, 2026-06-17 PDT). We have optimized capacity across network backbones to increase available headroom. Additionally, we are further augmenting our Delhi backbone capacity (expected to be complete by Monday, 2026-06-15 PDT). We will continue to closely monitor latency deviations and packet drops.\nWe will provide our next update by Monday, 2026-06-15 17:00 PDT.\n**Symptoms**\nCustomers may experience slightly elevated latency and non-optimal network routing into Google Cloud until the affected facility is fully restored.\n**Workaround**\nThere is no workaround at this time.","status":"SERVICE_DISRUPTION","affected_locations":[{"title":"Delhi (asia-south2)","id":"asia-south2"}]},{"created":"2026-06-11T23:25:15+00:00","modified":"2026-06-12T22:50:20+00:00","when":"2026-06-11T23:25:15+00:00","text":"**Summary**\nNetwork traffic to Google Cloud originating from Delhi, Chennai, Mumbai and surrounding areas is experiencing intermittent periods of elevated latency and possible packet loss.\n**Description**\nA fire at a third-party data center facility required an emergency power shutdown of networking equipment, isolating a non-compute local Point of Presence (POP) in Delhi and reducing available network capacity in the metro area.\nWe rerouted significant traffic from the impacted facility in Delhi to address reduced local serving capabilities. As a result, a subset of Hybrid Connectivity and Virtual Private Cloud (VPC) customers may be impacted by the routing changes made to address reduced local, latency-optimized serving capabilities in Delhi. Affected customers may experience intermittent latency spikes due to demand exceeding capacity across Indian metros and regional ISPs.\nInitial traffic mitigations have yielded positive results for some Cloud customers. In parallel, we are pursuing additional Internet Edge peering capacity to reduce existing fragility in the metro. We are also optimizing capacity across network backbones to increase additional headroom for VPC customers in the affected region. Additionally, we are planning to augment our local Delhi POP and migrating select peering partners to further increase regional capacity. We will continue to closely monitor latency deviations and packet drops.\nWe will provide our next update by Friday, 2026-06-12 at 17:00 PDT.\n**Symptoms**\nCustomers may experience slightly elevated latency and non-optimal network routing into Google Cloud until the affected facility is fully restored.\n**Workaround**\nThere is no workaround at this time.","status":"SERVICE_DISRUPTION","affected_locations":[{"title":"Delhi (asia-south2)","id":"asia-south2"}]},{"created":"2026-06-10T22:25:38+00:00","modified":"2026-06-11T23:25:15+00:00","when":"2026-06-10T22:25:38+00:00","text":"**Summary**\nNetwork traffic to Google Cloud originating from Delhi, Chennai, Mumbai and surrounding areas is experiencing intermittent periods of elevated latency and possible packet loss.\n**Description**\nA fire at a third-party data center facility required an emergency power shutdown of networking equipment, isolating a non-compute local Point of Presence (POP) in Delhi and reducing available network capacity in the metro area.\nWe rerouted significant traffic from the impacted facility in Delhi to address reduced local serving capabilities. As a result, a subset of Hybrid Connectivity and Virtual Private Cloud (VPC) customers may be impacted by the routing changes made to address reduced local, latency-optimized serving capabilities in Delhi. Affected customers may experience intermittent latency spikes due to demand exceeding capacity across Indian metros and regional ISPs.\nWe are investigating additional traffic mitigations and Internet Edge peering augmentation to alleviate the latency issues affecting our customers. We are continuing to work with the ISP partners in the region to mitigate any additional impact from unplanned failures.\nWe will provide our next update by Thursday, 2026-06-11 17:00 PDT.\n**Symptoms**\nCustomers may experience slightly elevated latency and non-optimal network routing into Google Cloud until the affected facility is fully restored.\n**Workaround**\nThere is no workaround at this time.","status":"SERVICE_DISRUPTION","affected_locations":[{"title":"Delhi (asia-south2)","id":"asia-south2"}]},{"created":"2026-06-10T01:13:50+00:00","modified":"2026-06-10T22:25:38+00:00","when":"2026-06-10T01:13:50+00:00","text":"**Summary**\nNetwork traffic to Google Cloud originating from Delhi, Chennai, Mumbai and surrounding areas is experiencing intermittent periods of elevated latency and possible packet loss.\n**Description**\nA fire at a third-party data center facility required an emergency power shutdown of networking equipment, isolating a non-compute local Point of Presence (POP) in Delhi and reducing available network capacity in the metro area.\nWe rerouted significant traffic from the impacted facility in Delhi to address reduced local serving capabilities. As a result, a subset of Hybrid Connectivity and Virtual Private Cloud (VPC) customers may be impacted by the routing changes made to address reduced local, latency-optimized serving capabilities in Delhi. Affected customers may experience intermittent latency spikes due to demand exceeding capacity across Indian metros and regional ISPs.\nWe are investigating additional traffic mitigations and Internet Edge peering augmentation to alleviate the latency issues affecting our customers.\nWe will provide our next update by Wednesday, 2026-06-10 17:00 PDT.\n**Symptoms**\nCustomers may experience slightly elevated latency and non-optimal network routing into Google Cloud until the affected facility is fully restored.\n**Workaround**\nThere is no workaround at this time.","status":"SERVICE_DISRUPTION","affected_locations":[{"title":"Delhi (asia-south2)","id":"asia-south2"}]}],"most_recent_update":{"created":"2026-06-29T23:06:10+00:00","modified":"2026-06-29T23:06:10+00:00","when":"2026-06-29T23:06:10+00:00","text":"**Summary**\nNetwork traffic to Google Cloud originating from Delhi, Chennai, Mumbai and surrounding areas experienced intermittent periods of elevated latency and possible packet loss.\n**Description**\nTraffic rerouting from the impacted Delhi facility caused a subset of Hybrid Connectivity, Virtual Private Cloud (VPC) and Media CDN customers to experience intermittent latency spikes as demand exceeded regional capacity.\nWe completed the augmentation of out-of-region Internet Edge regional peering capacity in Chennai to provide additional load-balancing and redundancy to large ISPs in India. Service to a large portion of Internet Edge peering capacity has been restored to reduce latency in the local Delhi metropolitan area. We are now recovered and returned to normal service as of Friday, 2026-06-26 PDT.\nFollowing safety clearance, our team restored all lost capacity. We have restored capacity between Delhi-Chennai and Delhi-Mumbai and will continue to closely monitor latency deviations and packet drops.\nWe thank you for your patience during the resolution of this issue.\n**Symptoms**\nThe impacted customers may have experienced slightly elevated latency and non-optimal network routing into Google Cloud.\n**Workaround**\nNone","status":"AVAILABLE","affected_locations":[]},"status_impact":"SERVICE_DISRUPTION","severity":"medium","service_key":"zall","service_name":"Multiple Products","affected_products":[{"title":"Hybrid Connectivity","id":"5x6CGnZvSHQZ26KtxpK1"},{"title":"Media CDN","id":"FK8WX6iZ3FuQL6qUwski"},{"title":"Virtual Private Cloud (VPC)","id":"BSGtCUnz6ZmyajsjgTKv"}],"uri":"incidents/5fGQt4VbkDnr3Yp8PXPr","currently_affected_locations":[],"previously_affected_locations":[{"title":"Delhi (asia-south2)","id":"asia-south2"},{"title":"Global","id":"global"}]},{"id":"41E5S3mkTGDfkZuJZH5k","number":"6876619551109882402","begin":"2026-02-27T12:37:00+00:00","created":"2026-02-27T16:12:30+00:00","end":"2026-02-27T14:35:00+00:00","modified":"2026-03-09T05:25:43+00:00","external_desc":"Vertex AI Gemini API customers experienced increased error rates when accessing the global endpoint.","updates":[{"created":"2026-03-09T05:25:43+00:00","modified":"2026-03-09T05:25:43+00:00","when":"2026-03-09T05:25:43+00:00","text":"# Incident Report\n## Summary\nOn Friday, 27 February 2026 at 04:37 US/Pacific, customers using Vertex AI Gemini API models experienced increased error rates. Impacted services included Google Cloud Support, Agent Assist, Vertex Gemini API and Dialogflow CX in US regions and the global endpoint. The issue persisted for a duration of 1 hour and 58 minutes.\nThis is not the level of quality and reliability we strive to offer you, and we have taken immediate steps to improve the platform’s performance and availability.\n## Root Cause\nThis incident was caused by a configuration change to a safety filtering service that supports all Gemini models. For some specific requests, this created code paths that eventually led to service disruptions and capacity loss for the safety filtering service. Consequently, customers encountered overload (429 and 503) errors for their queries, with some users reporting elevated error rates for specific models in US regions.\n## Remediation and Prevention\nGoogle engineers were alerted to the issue via our automated monitoring system on Friday, 27 February 2026 04:54 US/Pacific and immediately started an investigation.\nEngineers identified the faulty configuration change and initiated a rollback to restore the previous stable configuration. Engineers also added more capacity to the service to stabilize it. Full service restoration was confirmed by 06:35 US/Pacific as the rollback propagated and servers became healthy. \\ \\\nGoogle is committed to preventing a repeat of this issue and is taking the following actions:\n* Reinforcing rollout processes to include mandatory validation checkpoints.\n* Improving alerting systems to monitor critical dependencies more closely.\n## ## Detailed Description of Impact\nOn Friday, 27 February 2026 between 04:37 and 06:35 US/Pacific, customers accessing Vertex Gemini APIs may have experienced the following:\n* **Affected Models:** All Vertex AI Gemini API models were affected, including gemini-2.5-flash, gemini-2.5-flash-lite, gemini-2.5-pro, gemini-3.0-flash-preview, gemini-3.0-pro-preview, gemini-2.0-flash, gemini-2.0-flash-lite.\n* **Error Experience:** * **PayGo Customers:** Experienced primarily 429 Resource Exhausted errors. * **Provisioned Throughput (PT) Customers:** Received 503 Service Unavailable errors. * For PT customers, most errors stopped at **06:00**. For PayGo customers, most errors stopped at **06:20**.\n* **Geographic Scope:** Global endpoint, us-central1, us-east4, and other US regions were impacted.","status":"AVAILABLE","affected_locations":[]},{"created":"2026-03-04T23:23:18+00:00","modified":"2026-03-09T05:25:43+00:00","when":"2026-03-04T23:23:18+00:00","text":"# Preliminary Incident Report\nWe apologize for the inconvenience this service disruption may have caused. We would like to provide some information about this incident below. Please note, this information is based on our best knowledge at the time of posting and is subject to change as our investigation continues. A final Incident Report with preventative actions will be posted once our investigation is complete. If you have experienced impact outside of what is listed below, please reach out to Google Cloud Support using https://cloud.google.com/support.\n## Date/Time of the Issue (All time US/Pacific)\nIncident Start: 27 February 2026 04:37\nIncident End: 27 February 2026 06:35\nDuration: 1 hour, 58 minutes\n## Summary\nOn Friday, 27 February 2026 at 04:37 US/Pacific, customers using Vertex AI Gemini API models (including Gemini 2.0, 2.5, and 3.0 previews) experienced increased error rates. Impacted services included Google Cloud Support, Agent Assist, the Vertex Gemini API and Dialogflow CX in US regions and the global endpoint for a duration of 1 hour and 58 minutes.\nThis is not the level of quality and reliability we strive to offer you, and we are taking immediate steps to improve the platform’s performance and availability.\n## Preliminary Root Cause\nThis incident was caused by a configuration change to a safety filtering service that supports all Gemini models. This configuration change enabled a code path that interacted poorly with specific requests, leading to service disruption for the safety filtering service. This in turn led to customers seeing overload (429 and 503) errors for their queries.\nGoogle engineers have begun a full root cause analysis and will provide additional information once it is available.\n## Remediation\nGoogle engineers were alerted to the service disruption via automated alert on Friday, 27 February 2026 04:54 US/Pacific and immediately started an investigation.\nEngineers identified the faulty configuration change for a safety filtering service and initiated a rollback to restore the previous stable configuration. Additionally, engineers added more capacity to the service. Full service restoration was confirmed by 06:35 US/Pacific as the rollback propagated and servers became healthy.\n## Description of Impact\nOn Friday, 27 February 2026 between 04:37 and 06:35 US/Pacific, customers accessing Vertex Gemini APIs may have experienced the following:\n- Affected Models: All Gemini versions were affected, including gemini-2.5-flash, gemini-2.5-flash-lite, gemini-2.5-pro, gemini-3.0-flash-preview, gemini-3.0-pro-preview, gemini-2.0-flash, gemini-2.0-flash-lite.\n- Error Experience: - PayGo Customers: Experienced primarily 429 Resource Exhausted errors. - Provisioned Throughput (PT) Customers: Received 503 Service Unavailable errors. - For PT customers, most errors stopped at 06:00. For PayGo customers, most errors stopped at 06:20.\n- Geographic Scope: Global endpoint, us-central1, us-east4, and other US regions were impacted.","status":"AVAILABLE","affected_locations":[]},{"created":"2026-02-27T16:12:30+00:00","modified":"2026-03-04T23:23:18+00:00","when":"2026-02-27T16:12:30+00:00","text":"**Description** \\\nBetween Friday, 2026-02-27, 04:36 and 06:45 PST, customers experienced increased error rates when accessing the Vertex Gemini API Global endpoint. The issue impacted API requests to multiple Gemini models.\nThe incident also caused downstream impact to Dialogflow CX, Agent Assist, Google Cloud Support AI agent, and Customer Experience Agent Studio, which rely on Gemini APIs.\nPreliminary analysis indicates the issue was triggered by a recent configuration change. Service was fully restored after the configuration change was rolled back.\nWe thank you for your patience while we worked on resolving the issue.\n**Symptom**\n\\\nCustomers experienced increased error rates when sending API requests to impacted multiple Gemini models through the global endpoint.","status":"SERVICE_INFORMATION","affected_locations":[{"title":"Montréal (northamerica-northeast1)","id":"northamerica-northeast1"},{"title":"São Paulo (southamerica-east1)","id":"southamerica-east1"},{"title":"Iowa (us-central1)","id":"us-central1"},{"title":"South Carolina (us-east1)","id":"us-east1"},{"title":"Northern Virginia (us-east4)","id":"us-east4"},{"title":"Columbus (us-east5)","id":"us-east5"},{"title":"Oregon (us-west1)","id":"us-west1"}]}],"most_recent_update":{"created":"2026-03-09T05:25:43+00:00","modified":"2026-03-09T05:25:43+00:00","when":"2026-03-09T05:25:43+00:00","text":"# Incident Report\n## Summary\nOn Friday, 27 February 2026 at 04:37 US/Pacific, customers using Vertex AI Gemini API models experienced increased error rates. Impacted services included Google Cloud Support, Agent Assist, Vertex Gemini API and Dialogflow CX in US regions and the global endpoint. The issue persisted for a duration of 1 hour and 58 minutes.\nThis is not the level of quality and reliability we strive to offer you, and we have taken immediate steps to improve the platform’s performance and availability.\n## Root Cause\nThis incident was caused by a configuration change to a safety filtering service that supports all Gemini models. For some specific requests, this created code paths that eventually led to service disruptions and capacity loss for the safety filtering service. Consequently, customers encountered overload (429 and 503) errors for their queries, with some users reporting elevated error rates for specific models in US regions.\n## Remediation and Prevention\nGoogle engineers were alerted to the issue via our automated monitoring system on Friday, 27 February 2026 04:54 US/Pacific and immediately started an investigation.\nEngineers identified the faulty configuration change and initiated a rollback to restore the previous stable configuration. Engineers also added more capacity to the service to stabilize it. Full service restoration was confirmed by 06:35 US/Pacific as the rollback propagated and servers became healthy. \\ \\\nGoogle is committed to preventing a repeat of this issue and is taking the following actions:\n* Reinforcing rollout processes to include mandatory validation checkpoints.\n* Improving alerting systems to monitor critical dependencies more closely.\n## ## Detailed Description of Impact\nOn Friday, 27 February 2026 between 04:37 and 06:35 US/Pacific, customers accessing Vertex Gemini APIs may have experienced the following:\n* **Affected Models:** All Vertex AI Gemini API models were affected, including gemini-2.5-flash, gemini-2.5-flash-lite, gemini-2.5-pro, gemini-3.0-flash-preview, gemini-3.0-pro-preview, gemini-2.0-flash, gemini-2.0-flash-lite.\n* **Error Experience:** * **PayGo Customers:** Experienced primarily 429 Resource Exhausted errors. * **Provisioned Throughput (PT) Customers:** Received 503 Service Unavailable errors. * For PT customers, most errors stopped at **06:00**. For PayGo customers, most errors stopped at **06:20**.\n* **Geographic Scope:** Global endpoint, us-central1, us-east4, and other US regions were impacted.","status":"AVAILABLE","affected_locations":[]},"status_impact":"SERVICE_INFORMATION","severity":"low","service_key":"zall","service_name":"Multiple Products","affected_products":[{"title":"Agent Assist","id":"eUntUKqUrHdbBLNcVVXq"},{"title":"Dialogflow CX","id":"BnCicQdHSdxaCv8Ya6Vm"},{"title":"Google Cloud Support","id":"bGThzF7oEGP5jcuDdMuk"},{"title":"Vertex Gemini API","id":"Z0FZJAMvEB4j3NbCJs6B"}],"uri":"incidents/41E5S3mkTGDfkZuJZH5k","currently_affected_locations":[],"previously_affected_locations":[{"title":"Global","id":"global"},{"title":"Montréal (northamerica-northeast1)","id":"northamerica-northeast1"},{"title":"São Paulo (southamerica-east1)","id":"southamerica-east1"},{"title":"Iowa (us-central1)","id":"us-central1"},{"title":"South Carolina (us-east1)","id":"us-east1"},{"title":"Northern Virginia (us-east4)","id":"us-east4"},{"title":"Columbus (us-east5)","id":"us-east5"},{"title":"Oregon (us-west1)","id":"us-west1"}]}]