Receptionist Friendly Disaster Recovery

Status
Not open for further replies.

as2563

Forum User
Joined
Apr 23, 2020
Messages
3
Reaction score
0
We have been running on 3CX for 6-7 weeks now. 65 extensions using Yealink T46S phones and 3CX version 16.0.4 is installed onsite as a Hyper-V virtual machine, Debian OS, 4 cores, 8 GB RAM, 16 call Enterprise license. We use Bandwidth.com sip trunks resold to us by Cloudco.

This morning I was in a meeting and as I was coming out our receptionist flagged me down and informed me that "the phones are down". All of our extensions said "No Service" on the display. Calls in to our main number were dead quiet with no ring or busy signal. The 3CX web client looked normal. The 3CX console looked normal, except that it was showing a steady 25% CPU utilization which is much higher than normal.

I logged into the Hyper-V host and initiated a shutdown from the hypervisor management console. As I observed the 3CX console from within HyperV, I noticed it hung up waiting on the SIP server to terminate for much longer than it did the other services. After shutdown completed, I started the VM and everything returned to normal.

Here is my problem. We are a busy company, we get lots of calls and handling our calls well is a big part of our company culture. I am a busy IT manager with a lot on my plate. Some times I go on vacation (imagine that). What are we supposed to do the next time 3CX goes bonkers and needs a timeout, and I'm not here to make it happen? Having to log into the hypervisor to restart the phone system is not something I would even attempt to train a receptionist on.

I am considering getting a mini PC, installing 3CX on it, and tucking it under the reception desk. Then I can show her the power button and say "next time things go nuts, just shut off this computer, wait a few seconds, and turn it back on."

Tell me why this is a bad idea, and how could it be better. Final solution must be receptionist friendly.
 
We have been running on 3CX for 6-7 weeks now. 65 extensions using Yealink T46S phones and 3CX version 16.0.4 is installed onsite as a Hyper-V virtual machine, Debian OS, 4 cores, 8 GB RAM, 16 call Enterprise license. We use Bandwidth.com sip trunks resold to us by Cloudco.

This morning I was in a meeting and as I was coming out our receptionist flagged me down and informed me that "the phones are down". All of our extensions said "No Service" on the display. Calls in to our main number were dead quiet with no ring or busy signal. The 3CX web client looked normal. The 3CX console looked normal, except that it was showing a steady 25% CPU utilization which is much higher than normal.

I logged into the Hyper-V host and initiated a shutdown from the hypervisor management console. As I observed the 3CX console from within HyperV, I noticed it hung up waiting on the SIP server to terminate for much longer than it did the other services. After shutdown completed, I started the VM and everything returned to normal.

Here is my problem. We are a busy company, we get lots of calls and handling our calls well is a big part of our company culture. I am a busy IT manager with a lot on my plate. Some times I go on vacation (imagine that). What are we supposed to do the next time 3CX goes bonkers and needs a timeout, and I'm not here to make it happen? Having to log into the hypervisor to restart the phone system is not something I would even attempt to train a receptionist on.

I am considering getting a mini PC, installing 3CX on it, and tucking it under the reception desk. Then I can show her the power button and say "next time things go nuts, just shut off this computer, wait a few seconds, and turn it back on."

Tell me why this is a bad idea, and how could it be better. Final solution must be receptionist friendly.
Move to the Cloud instead and use an SBC? Then, if something is wrong, calls will still work because the system is cloud and mobile apps can work.

You problem oddly sounds like a networking issue. What's the network setup? Also, have you followed the 3CX Hyper-V docs?
 
I don't see how moving to cloud helps in a situation where the PBX itself is clearly hung up and maxing out a processor core and has to be rebooted. If anything it would have been worse as I would need to get logged into a cloud account in order to reboot the VM.

We have a Ubiquiti Unifi LAN with a dedicated VLAN for VOIP and that VLAN is set as the voice network on all ports. I did make some tweaks last night to fully align with 3CX Hyper-V docs.

However, in the scenario outlined above, with a hard reset of the PBX, is 3CX prone to database corruption? For instance, the original Unifi cloud key devices were prone to corrupt the internal database when hit with a power outage.

Obviously this is not an ideal way to reboot a PBX but I need something simple that a receptionist can do if I am not available.
 
We don't know what happened and under what circumstances in that case, so it's best not to speculate. The CPU usage could

I would encourage you to share more details as per my previous post
https://www.3cx.com/community/threads/information-to-provide-when-requesting-help.67558/
and in addition include the following information

- Host Server CPU type
- If other VMs are running and sharing the CPU
- If you have anything else installed in the VM's OS other than 3CX (including monitoring software)
- How the phones are networked in relation to the PBX
- How many virtual NICs the VM has
 
I would enable all the email notifications you can and have them come to your mailbox, maybe make a rule to file them under a folder within your inbox (Settings > Email > Notifications tab) you might get alerted to an issue before you need to take action.

Rather than give a user access the turn on/off the phone system, you could make up a powershell script that they could run from their workstation, you'll find the commands you need here https://docs.microsoft.com/en-us/powershell/module/hyper-v/restart-vm?view=win10-ps you might want to add something in to your powershell script to log when it has been run or even alert you somehow (email?) or you might find out users have just been rebooting it for the past 3 months every week without letting you know.

I've haven't come across any issues with 3cx so far when turning off VMs, not like the original cloudkey - I basically prepared to restore from backup every time I rebooted/upgraded one of them!
 
Sounds like you need to hire some help! or outsource your phone system support.

You could schedule the VM to reboot every so often or get a debian install and show receptionist how to log in to MGMT console and reboot via terminal.
 
Move to the Cloud instead and use an SBC? Then, if something is wrong, calls will still work because the system is cloud and mobile apps can work.

How does moving to the cloud help a hung system?
 
You could schedule the VM to reboot every so often or get a debian install and show receptionist how to log in to MGMT console and reboot via terminal.

It is a Debian...
 
So as you said you made some tweaks to the VM. I'm betting one of those had to do with time synchronization. But ultimately I don't understand your question. Do you not have other servers there? Do you expect the receptionist to reboot those as well when they go down? 3CX is no more or less robust than any other software. If you are concerned about things happening when you are away, then you do what most other companies do. You hire someone or an outsourced firm to cover when you are on vacation. But ultimately if configured in a supported environment what you described shouldn't happen and I doubt it will happen again.
 
  • Like
Reactions: accentlogic
Status
Not open for further replies.

Latest Posts

Forum statistics

Threads
111,945
Messages
589,866
Members
164,835
Latest member
Firefox Technologies