Multiple 3CX Servers hanging at random since 16.0 Release

Status
Not open for further replies.

BrenttG

Platinum Partner
Advanced Certified
Joined
Nov 17, 2017
Messages
912
Reaction score
586
OK, ill break down the technical side of what we are seeing in our cloud environment of nearly 100 3CX Servers.

We have numerous 3CX servers this has occurred on now, there does not seem to be any warning...

  1. Customer will call us and report that they are not getting inbound calls, but can still call out.
  2. We enter the 3CX Management Console, and about 50% of the time everything looks green, the other half, the Services indicator will be Red.
  3. Upon entering the services list, we will see one of two possible things:
a: All services are running, however the SIP Server has 1.2-1.5 GB of RAM usage
b: IVR Service will be in Stopped state, and the SIP Server has 1.2-1.5 GB of RAM usage.

Our hypothesis is that the SIP Server is depleting the Server of RAM, which causes the IVR Service and/or another service to Halt or Exit to stopped state.

In troubleshooting, the IVR Service does not easily restart if we attempt to just do it, however if we restart the SIP Server which has high memory usage, and then do the IVR Service they both begin to respond normally again fairly quickly.

This has occurred on systems with only 6 phones, and a system that has ~45 phones, it does not seem to matter how many extensions are on the system. This has taken place exclusively on systems that were upgraded from 15.5 to 16.0, we have not seen it occur on systems that are new as of V16.0, but that does not mean it wont, it just has not yet.

All the affected PBXs were of the following specs, running on VMware vCenter/ESXi 6.5
2 x CPU Cores (Xeon E5 2640 v2s, Hosts are all dual socket)
2GB RAM
50 GB HDD
Debian 9 built from 3CX Linux ISO, No GUI Running or Installed, we use console exclusively to reduce resource usage.
No CPU throttling is in place
All VMware Hosts are ~25%-35% Memory Load, none are even close to being peaked.
Our hosts are all dual Socket Dell with 256 GB RAM per host.

All of our Linux 3CX Servers average 20-30% memory usage during normal operation, but when the problem occurs, the memory usage skyrockets, and it is the SIP Server, that is causing it.

I will add any further technical details as i learn of them or am notified of them from the rest of my team. We have encountered this 5-7 times now, and have just put together during a meeting that this same symptoms, and resolution and risen over and over now, but so far never on the same system more than once yet, we are going to start a tally to track the impact pattern. The first incident of this as best we can tell, was about 3 days ago.
 
Last edited:
@Gunnar Brynjólfsson @Azhar

Can you both confirm if you are using the Instance Manager connected to a Reseller account on the affected machines ?
Hi, I was using "Instance Manager" connected to the Reseller account on the affected machine, those not using instance manager were not affected.
 
Thank you for your feedback, the issue reported is under investigation.

For the time being I would suggest to disable the instance manager on the affected machines, and then reboot to clear any stuck processes while we look further into the matter.
 
@Gunnar Brynjólfsson @Azhar

Can you both confirm if you are using the Instance Manager connected to a Reseller account on the affected machines ?

We do not use the Instance Manager, but all of our PBXs are registered to our reseller account, and i believe the Instance Manager is turned on, however, i must note, the last several times the problem has occurred for us, including last weekend wheni t hit 30 servers at once, it was within minutes of the automatic update schedule hitting. Im not 100% convinced its the instance manager due to that fact that 2 minutes after the system tries to update, BOOM! And there wasn't even any 3CX updates pending, they were all already up to date, the only update showing was the BETA, which we don't do on production systems.
 
I just logged into the two biggest systems that were affected on sunday night, and both do not have instance manager turned on.... it is off. +1 to the scheduled updater having something to do with it.

Note: the 30 servers that went down, were all set to update automatically in 3CX, at 8:00 PM Arizona,USA timezone, on sunday. They all went offline at varying times between 8:02-8:05 PM. The times they went offline are a matter of record because we have active monitoring running against them.
 
Hi everyone,

Just an update on this. We made a few adjustments to the scripts that are ran in the background by the Instance Manager which should prevent this kind of behavior.
To get the adjustments, all you should need to do is either:
  • Stop the Instance Manager, then start it again
or
  • Restart the whole machine 3CX is running on
(whatever is easier for you)

I think we haven't missed something, but looking forward to any feedback you may have.
 
Status
Not open for further replies.

Forum statistics

Threads
112,089
Messages
590,696
Members
165,053
Latest member
jmc-tim