Incident Overview
Calls routed through a subset of our SIP infrastructure experienced connection timeouts and call failures. The issue was caused by network port exhaustion on specific SIP media processing servers.
Date: Oct 6th 2026
Duration: 2h 13mins (9:53 UTC to 12:06 UTC)
Impact: ~66% of inbound and 100% of outbound calls using EU SIP servers with static IPs crashed or timed out
Root Cause Details
Capacity Context: Our SIP servers are configured with network port ranges designed to support at least 30 times our peak call volume.
Failure Mechanism: A software bug during call teardown prevented network ports from being properly closed and returned to the available pool when calls ended.
Impact Timeline: Because the SIP servers ran continuously for a couple of weeks without restarts, the unreleased ports accumulated over time. Once all available ports on a server were exhausted, any new calls assigned to that instance failed to establish a connection and timed out.
Action Items
Automated Server Restarts: Added scheduled periodic restarts for the SIP servers to routinely flush stale connections and clear the port pool.
Port Usage Monitoring & Alerting: Building on our existing infrastructure monitoring suite, we have enhanced our telemetry to specifically track used versus available network ports on every server instance. This adds a new layer of proactive alerting to notify engineering long before port usage reaches critical levels.
Port Leak Code Fix: Actively auditing the call termination routines to isolate the exact code path leaving ports open and deploy a permanent software patch.
We apologize for the impact this had on your workstreams and are committed to preventing similar issues moving forward.