Resolved – 02:17 CEST All services have been verified healthy and this incident is resolved. The applications that were still restarting have completed, storage has finished resynchronising, and every backup job that was delayed during the outage has since run successfully.
Impact: all mdapi.ch services were unavailable from about 18:48 to 19:37 CEST; core services (mail, sign-in, DNS, NTP) were back by 19:50 and the remaining applications returned over the following 15 minutes.
Cause: a storage I/O stall locked up two of the three cluster nodes; the hardware watchdog did not reset them and they were reset manually via out-of-band management. The stall and the failed automatic reset are being reviewed.
Posted Sep 04, 2026 - 02:17 CEST
Update
Update – 20:05 CEST All core services are restored: mail, sign-in, GitLab and platform tooling, DNS, NTP and cloud/notes/documents are operational. A handful of applications are still restarting after the outage and are expected back within the next 15 minutes. Backup jobs are catching up. We will resolve this incident once everything is confirmed healthy.
Posted Sep 03, 2026 - 20:04 CEST
Monitoring
Update – 19:50 CEST Both locked nodes have been reset and rejoined the cluster; the control plane has been fully available since 19:37 CEST. Core services (DNS, authentication gateway, remote access) are back. Remaining applications, including GitLab, are restarting and should return over the next 15 to 30 minutes as storage finishes resynchronising. We are monitoring until everything is confirmed healthy.
Posted Sep 03, 2026 - 19:51 CEST
Identified
Update – 19:30 CEST Two of the three cluster nodes locked up at 18:35 and 18:37 CEST after a storage I/O stall. The hardware watchdog fired but the nodes did not reset themselves. The third node crashed and rebooted on its own at 18:46 and is healthy, but the cluster cannot serve with a single node.
We are hard-resetting the two locked nodes one at a time via out-of-band management. The first reset was issued at 19:27 CEST. Services should start returning once that node has rebooted (roughly 10 minutes), with full capacity after the second reset.
No data loss is expected. Next update by 20:00 CEST.
Posted Sep 03, 2026 - 19:31 CEST
Investigating
Since 18:48 CEST (16:48 UTC) on 3 September 2026, all services hosted on the mdapi.ch platform are unavailable. This includes web applications, GitLab, single sign-on, DNS-backed services and the remote-access gateway.
The cluster hosting these services lost its control plane after a storage I/O stall. One node crashed and rebooted on its own; the other two became unresponsive at the same time and have not recovered yet. The platform cannot serve traffic while fewer than two nodes are healthy.
We have out-of-band access to the hardware and are working to bring the two unresponsive nodes back. There is no indication of data loss at this stage. Internet connectivity and edge networking are not affected.
Next update within 60 minutes or as soon as the nodes are back online.
Posted Sep 03, 2026 - 19:18 CEST
This incident affected: Mail, Sign-in (SSO), Platform, Files & Documents, Backups, Websites, DNS, Smart Home, Internet, and NTP.