Web App and MCP Unavailable
Resolved
Oct 1, 2026 at 8:30pm UTC
Background
Early this week, Windmill rolled out enhanced monitoring for the core infrastructure which processes all requests coming into all our services (Windy, web app, MCP, and integrations).
We noticed a small number (>0.01%) of requests were failing due to a networking misconfiguration in how our services talk to each other. Although there was no customer impact since these failing requests get retried, we proactively rolled out a fix the evening of 9/30/2026.
Windmill runs its services in a high-availability configuration, to ensure our services are always available. To maintain high-availability during upgrades, we use rolling deploys where old versions of services stay online until the new version is ready.
Once the new version is ready, the old version enters a draining state where it will finish servicing all pending requests before it goes offline. The configuration change we rolled out on 9/30/2026 caused a latent bug in how we handle draining to block the old versions from shutting down. We balance load evenly across all instances, meaning the 50% of requests going to the stuck old version would fail.
Timeline
- 9/30 @ 11:52 PM EST - Fix deployed for networking misconfiguration, exposing latent bug in service shutdown
- 10/1 @ 7:00 AM EST - A small number of MCP servers get shut down during load balancing, but fail to drain, causing minor degradation in MCP service.
- 10/1 @ 8:30 AM EST - Error rates increase, reported internally and incident declared.
- 10/1 @ 9:00 AM EST - Root cause identified.
- 10/1 @ 9:46 AM EST - Fix is identified, applied, and roll out begins.
- 10/1 @ 9:52 AM EST - Roll out reaches our API servers, error rates begin to rise as the old version fails to drain properly.
- 10/1 @ 10:01 AM EST - First user-facing errors accessing app.gowindmill.com reported and incident escalated.
- 10/1 @ 10:09 AM EST - Roll out of new API servers completes, but old servers continue to fail to drain. 50% error rate observed.
- 10/1 @ 10:21 AM EST - Infrastructure vendor involved to forcefully shutdown the erroring API servers.
- 10/1 @ 10:25 AM EST - Force shutdown triggered by vendor, error rates remain elevated
- 10/1 @ 10:30 AM EST - Identified another service was blocking the force shutdown, restart on the other service triggered
- 10/1 @ 10:36 AM EST - Restart completes, all the erroring servers terminate, error rate drops to 0%.
- 10/1 @ 10:43 AM EST - Continued monitoring reveals no further errors, incident resolved.
Remediation
We understand the impact this downtime had - our customers rely on Windmill for critical work and we apologize greatly for interrupting it.
Internally, we are implementing changes to prevent similar incidents from occurring in the future. This includes:
Enhanced testing in staging prior to promotion to production
Comprehensive audits for similar latent risks and remediation
Continuing to improve monitoring and shorten incident response time
More clear error pages
For continuous updates on Windmill’s real-time status, we encourage customers to subscribe to status page here: https://status.gowindmill.com/
Affected services
Updated
Oct 1, 2026 at 2:43pm UTC
All of the problematic instances have been fully replaced and the system has stabilized. We apologize for any inconvenience and are continuing to monitor the situation.
Affected services
Updated
Oct 1, 2026 at 2:39pm UTC
Our deployment to resolve the MCP connection issues triggered the same underlying issue on our API. Those instances have drained and the error rates are decreasing. We are continuing to monitor the situation.
Affected services
Updated
Oct 1, 2026 at 2:02pm UTC
The connection issues have spread to Windmill's web app; a fix is still actively being deployed
Affected services
Updated
Oct 1, 2026 at 1:46pm UTC
A fix is deploying now to resolve the issues
Affected services
Updated
Oct 1, 2026 at 1:00pm UTC
We have identified the root cause and are working on a solution.
Affected services
Created
Oct 1, 2026 at 12:30pm UTC
We are experiencing intermittent issues connecting to our MCP server. We are investigating the situation.
Affected services