Root Cause Analysis — SIP Trunking Calls Failing (July 31, 2026)
Duration:
EU and IN regions: July 31, 2026, widespread outage between 00:45 – 11:12 UTC with degradation persisting for some customers until 13:22 UTC
US region: July 31, 2026, full outage 00:45 – 06:14 and 09:22 - 10:30 UTC with the outage persisting for some customers until 19:00 UTC
Description of the Issue:
During the affected windows, calls made using the ElevenAgents platform over SIP telephony will have failed.
After an initial outage period, functionality was restored for the majority of customers at 10:30 UTC in the US and 11:12 UTC in EU/IN regions.
For any customers that had made edits to their SIP configurations during the initial outage, additional remediation was required for these configurations. Functionality was restored gradually to those configurations until full remediation for all customers and SIP configurations was completed at 13:22 UTC in the EU/IN region and 19:00 UTC for the US.
Following this incident, full functionality was restored and no additional action is necessary by those affected.
Root Cause:
As part of a planned cleanup of legacy storage behavior for SIP trunk configuration, a configuration change was rolled out ahead of the corresponding service update. Another related update hadn't rolled out yet, so the still-running version defaulted to the old storage location, and temporarily read and wrote trunk configuration there instead of the desired storage path.
A fix for this issue was deployed at 04:56 UTC in the US. This fix ensured that all reads and writes were coming exclusively from the appropriate storage location, and operations that used the other storage location were ignored. At this point all SIP configs that hadn’t been edited during the initial impact were restored and being read correctly. Unfortunately, as we were still seeing issues from configs that had been written during the incident and were now out of sync between the two storage layers, our investigating team did not identify this as the correct fix and so continued with alternative hypotheses.
Another new version of our SIP service was released at 09:22 UTC which reverted the change in the US and introduced other candidate fixes. This again meant that all SIP configurations were failing to read and made the root cause clear to our team.
The true fix was redeployed at 10:30 UTC in the US and 11:05 UTC in the EU.
For any SIP configurations that had been edited during the initial downtime, additional backfilling of data was required to ensure the data was stored consistently. This backfill took a number of hours, and progressively restoring service to customers. It took longer in the US due to a larger number of configurations which made the targeted backfill process more complex.
Timeline & Actions Taken:
July 31, 2026, 00:55 UTC: A misconfiguration applied and calls started failing, as SIP information could no longer be read from its storage location.
July 31, 2026, 06:14 UTC: A fix was deployed to the US environment that restored reads and writes to the appropriate storage location. This fixed the issue for all customers that hadn’t made changes to their configurations during the impacted period.
July 31, 2026, 09:22 UTC: The deployed fix was rolled back in the US after we mistakenly attributed the remediation to a different action. This resulted in calls failing in the US again.
July 31, 2026, 10:30 UTC: The original fix was reapplied to the US region, restoring calls for all SIP configurations that hadn’t been edited during the incident thus far.
July 31, 2026, 11:12 UTC: A fix was rolled out to the EU/IN region, restoring calls for all SIP configurations that hadn’t been edited during the incident thus far.
July 31, 2026, 13:22 UTC: The complete data storage remediation in the EU/IN region was complete and all calls now succeed.
July 31, 2026, 18:45 UTC: The complete data storage remediation in the US region was complete and all calls now succeed.
Preventative Actions & Learnings:
We have confirmed that persistent data, including SIP trunk configuration, is now stored in the sole and appropriate data path.
In addition, we are:
Improving detection and alerting for failed calls – Our alerting failed to notify us in this incident as a result of a change to our structured log format from an unrelated change. We will be running an audit of all services to ensure that this is not impacting any other service and remediating any location where we see this is the case.
Investigating our review and release process for configuration and service updates to ensure configuration changes are not deployed to invalid service versions
Improving our investigation response procedures to involve Account Executives and CSM teams earlier to facilitate proactive communication to affected customers
We apologize for the disruption this caused to your workstreams and are committed to preventing similar issues moving forward.