Bouncing Trunks

Status
Not open for further replies.

dkirk-ads

Customer
Joined
May 6, 2019
Messages
74
Reaction score
25
This is an OLD problem that started nearly a year ago. It has not been helped by newer versions of the 3CX software and we are currently running 16.0.504 on an internally hosted Debian server. Firewall is a Sonicwall TZxxx and the 3CX firewall check always passes with a green check mark. I have two SIP trunks both going to SIP.US and using the 3CX templates for those connections. SIP.US provides connections at gw1.sip.us and gw2.sip.us and 90% of the time this works great.

The problem comes seemingly randomly where the 3CX server can only see one of the SIP servers for a period of time and then it suddenly snaps back to working again. Most of the time the secondary is unreachable but not always.

Registration at SIP.US Secondary has failed. Destination (sip:74.81.71.18:5060;lr) is not reachable, DNS error resolving FQDN, or service is not available.

03/31/2020 8:51:06 AM - [CM504005]: Registration failed for: Lc:10001(@SIP.US Secondary[<sip:[email protected]:0/UDP>]); Cause: Cause: 408 Request Timeout/REGISTER from local

This morning I rebooted the Debian server and this time the secondary came online while the primary was locked out

Registration at SIP.US Primary has failed. Destination (sip:65.254.44.194:5060;transport=TCP;lr) is not reachable, DNS error resolving FQDN, or service is not available.

To be consistent I rebooted the Debian server yet again, a mere 2 minutes later and this time the Primary came online again, with the secondary down. This will continue for a day or so and then suddenly start working again. Since it shows the correct IP it would appear DNS is not the problem. When we do a trace on SIP.US server side their server never receives a connection from my static IP here at the office.

I have removed both SIP trunks and added them back in more times than I can remember, and nothing changes in a few hours or a day. This time, however, it has been failing for just over a week now, solid.

This is very annoying and I would like to get this resolved once and for all. Does anybody have any suggestions? I have a WireShark packet capture of a SIP registration attempt, what should I be looking for?

GREATLY appreciate any help.
 
Looking at the packet capture this struck me as problematic, thoughts? The x.x.15.35 is an internal DNS server while x.x.15.19 is the 3CX server.

44 1.971991 192.168.15.35 192.168.15.19 DNS 140 Standard query response 0x0011 No such name SRV _sip._udp.gw2.sip.us SOA ns10.dnsmadeeasy.com
 
A more complete picture of the registration attempt:

31 1.905525 192.168.15.186 192.168.15.19 SIP 881 Request: REGISTER sip:192.168.15.19:5060 (1 binding) |

32 1.905689 192.168.15.19 192.168.15.35 DNS 70 Standard query 0x0010 NAPTR gw2.sip.us

41 1.944477 192.168.15.35 192.168.15.19 DNS 130 Standard query response 0x0010 NAPTR gw2.sip.us SOA ns10.dnsmadeeasy.com

42 1.944714 192.168.15.19 192.168.15.35 DNS 80 Standard query 0x0011 SRV _sip._udp.gw2.sip.us

43 1.944791 192.168.15.19 192.168.15.186 SIP 446 Status: 200 OK (1 binding) |

44 1.971991 192.168.15.35 192.168.15.19 DNS 140 Standard query response 0x0011 No such name SRV _sip._udp.gw2.sip.us SOA ns10.dnsmadeeasy.com

45 2.173371 192.168.15.19 74.81.71.18 SIP 574 Request: REGISTER sip:gw2.sip.us (1 binding) |
 
From the Debian 3cx server both SIP servers are pingable by name.

Where is _sip._udp.gw2.sip.us coming from, shown in the trace above?

2020-03-31 10_09_06-Window.png
 
Using dig on the Debian 3CX server it shows the following DNS SRV record, so it's as if the Debian side of things is working, but 3CX is faltering:

2020-03-31 11_17_49-Window.png
 
Have you tried using the Google DNS servers... 8.8.8.8 and/or 8.8.4.4
 
I was using them, switched to CloudFlare, no change. Just switched back to Google again, rebooted the server, and sure enough .gw2 is alive and .gw1 is dead. Rebooted again, same. Rebooted again same. Rebooted again and this time they swapped back to .gw1 is alive and .gw2 is dead.
 
Last edited:
Packet capture is showing that I am not sending the credentials. Why, sometimes it does, sometimes it doesn't, sure looks like a 3CX issue somewhere deep? Am I wrong?
 
When you say "a 3CX" issue, it may very well be a 3CX configuration issue. If it were a bug, affecting everyone with the same installation, I'm certain that there would be more posts about it. It could be corruption, or a hardware issue on your particular install, or something on your network, or ISP. If possible (not easily done), you might consider a process of elimination.
Do you have a second provider to test with and see if the same thing happens with them? Can the PBX be attached to a different ISP, temporarily, as a test? Have you gone to the (extreme) measure of a complete new install?
 
Reconfiguring 3CX to have one SIP trunk (gw1.) and setting the alternative proxy to gw2. seems to have cured the condition. This 3cx installation simply would not reliably send two registrations. Test the failover of this new configuration shows it failing over properly. Strange, but I guess case closed. Tried to purchase a support ticket from 3cx but their web site won't do anything but show rotating circles; I tried that route as well.
 
Status
Not open for further replies.

Forum statistics

Threads
111,940
Messages
589,852
Members
164,832
Latest member
Boblatino