Post v18 Upgrade Lockup

Status
Not open for further replies.

Stoopler

Customer
Joined
Mar 25, 2021
Messages
10
Reaction score
3
We have 3CX installed on an Amazon Lightsail instance that has been running spectacularly until yesterday, about an hour after upgrade.

I took a snapshot and backup and upgraded 16->18 which seems to have worked without any issues. Upgraded our two SBC's to 18 as well, which worked.

The instance now seems to run for about an hour and then goes completely non-responsive.

Is there anything I could look into to try and figure out what is going on? This started when v18 was installed.
 
The instance now seems to run for about an hour and then goes completely non-responsive.
To first clarify a few things, when you say "non-responsive" do you mean the machine is completely non-responsive or the 3CX PBX?
• What tests have you made to determine either
• If it's the machine, can you not even SSH? Is it actually running if you check via Lightsail's console?
• You say it runs for about an hour and then this happens, what do you do to bring it back up?
• How was this 3CX Instance deployed, was it via the PBX Express, Marketplace, etc or was it a manual 3CX install onto a Debian 9 machine?


Quick things to check:

• Check resources usage of the machine while it's running. Make sure the VM is compliant with our hardware specs: https://www.3cx.com/docs/recommended-hardware-specifications-for-3cx/
• Check /var/log/syslog for the last entries before the issue occurred. The info there might shed some light on the matter.
• If it's the 3CX PBX that's going unresponsive, check if all services are running. You can use:
systemctl list-units 3CX* nginx* postgresq*
What do you see when you run this?
 
In every crash that we're seeing the last lines in syslog are this:

Dec 20 09:17:01 ip-172-26-2-255 CRON[32455]: (root) CMD ( cd / && run-parts --report /etc/cron.hourly)

Lightsail metrics report both a status check failure and instance check failure at the same time, which from my continued research indicates that there is an AWS hardware level issue that can be resolved with stopping/starting the instance (rather than rebooting).

The interesting thing here is that it's happening about every hour exactly.

The next time it happens, I will try to run the command mentioned. We're on a 16 GB RAM, 4 vCPUs, 320 GB SSD instance, so resources shouldn't be an issue (if anything, it's overkill for our use).

Instance was deployed via PBX Express around June 2021 on v16 and then upgraded yesterday to v18.
 
As an update, it just locked up again, VM completely inaccessible even through the Lightsail SSH console. I was logged in through the Lightsail web SSH console and running a continuous ping out and when the machine went non-responsive the console was also dead in the water.
 
As an update, it just locked up again, VM completely inaccessible even through the Lightsail SSH console. I was logged in through the Lightsail web SSH console and running a continuous ping out and when the machine went non-responsive the console was also dead in the water.
Sounds like it would be best to restore from a 3CX backup and move the IP across in Lightsail.
 
Sounds like it would be best to restore from a 3CX backup and move the IP across in Lightsail.

Is there a defined method for reinstalling somewhere? I'd like to have as little (which is neat, considering it's dying every hour) interruption to our PBX as possible. If I can reinstall and then restore the backup, followed by moving the IP over that would be great.
 
Continuing to troubleshoot before I reinstall, wondering if it isn't related/similar to this reddit post:


--

We're dropping every hour on the dot when watching it, getting these results when running the commands mentioned in that thread (see attached).
 
What you references does seem to be exactly what you are describing too.
I don't think it is something widespread as I think the internet would be blowing up about it, but if you want to try the re-install on a new machine, all you would need to do from a 3CX perspective is:

Take a Full Backup including all options and epsecially the License key and FQDN option, then:
  1. Power off the current VM
  2. Log into your Customer Portal and in Subscriptions, find your License Key
  3. You should see a "Reinstall" option, click it
  4. Follow the wizard and restore the backup when asked to.
I am not sure how you move the IP across from one instance to another tbh, but I think it's doable from the Lightsail interface.
 
  • Like
Reactions: ChrisC_3CX
Solution on this one was pretty dumb, it was because we had IPv6 disabled on v16 to resolve some issues it was causing. Networking service wouldn't come up because it failed to get a v6 lease.

Re-enabled IPv6 on the VM and rebooted and everything has been peachy ever since.
 
  • Like
Reactions: NickD_3CX
Solution on this one was pretty dumb, it was because we had IPv6 disabled on v16 to resolve some issues it was causing. Networking service wouldn't come up because it failed to get a v6 lease.

Re-enabled IPv6 on the VM and rebooted and everything has been peachy ever since.
Glad to hear you it's up and running again now!
 
Status
Not open for further replies.

Forum statistics

Threads
111,974
Messages
590,083
Members
164,901
Latest member
Silent_Guru