Hardware for local AI

GWinchester

Customer
Joined
Oct 23, 2025
Messages
20
Reaction score
17
I am looking forward to be able to run the AI/transcription locally. however there still seems to be limited information on hardware.
in the transcription engine server instructions https://www.3cx.com/docs/transcription-engine-server/ wich by the way i can only find a link to in one of the bogs (not in the help menu yet) it has GPU: Minimum Recommended (Fast): Nvidia 24GB RTX GPU. in theory an L4 would work?
we have dell PowerEdge R660 so a 5090 just isn't going to fit.
what have people tried and how did it go?

Another suggestion in one of the bogs was a GB10 based unit, however there seems to be no instructions for this type of setup.
Is there any more information available for GB10 or that option is not going to be supported?
 
Hello,

Anything running Debian/Ubuntu with an Nvidia RTX GPU (can use RTX drivers) with at least 24Gb of VRAM will do.

A Google L4 has max 22GB so its asking for trouble. You will be on the edge..

I have used an ADA 6000 in tests myself, and ofcourse 5090s work, and 40xx even if the VRAM is enough.

GB10 is to be supported, and would be a very good choice. ARM support will come in the near future and will be announced.
 
Last edited:
Adding to what @KyriacosS_3CX said,

Nice choice with the DGX Spark. It’s a great device.
At the moment, the onboard transcriber won’t work on it because the transcriber isn’t yet compatible with ARM-based devices. This is something we’re actively working on right now.
Once ARM support is completed, the onboard transcriber will run normally on the DGX Spark.
 
So, I guess the next question then is:
Will the Update 8 support the GB10 at the time of release or will it come later?
 
  • Like
Reactions: nikolascx
Hi Dears
It's always fun to discuss hardware :)

No not update 8.
But our order arrived just last week ;) and we are very eager to see the transcriber run on these devices.

This device is pure fun.
And it is going to be a baseline as all the big manufacturers are rebranding this so they can say they have their own DGX Spark. I compiled a small list of the big guys and their DGX Spark models:
  • Acer Veriton GN100 AI Mini Workstation
  • ASUS Ascent GX10
  • Dell Pro Max with GB10
  • GIGABYTE AI Top Atom
  • HP ZGX Nano AI Station G1n
  • Lenovo ThinkStation PGX
  • MSI EdgeXpert MS-C931
Fun times ahead..

Team - remember this.
This development initiative is completely separate from the PBX. Which means that to release support for ARM is not bound to a release of the server.

Rest assured that when we have this, we will independently launch it without delays. We are working on it as we speak.

God bless
 
Looks like we will most likely go with a Dell Pro Max with GB10.

Just I understood correctly, update 8 will be released and sometime later, hopefully not to long, transcription for GB10 will be release. which means there will be a period of time with no transcription unless we use another option in the meantime.
 
If you choose to hold until the GB10 support is rolled out, then a number of options are still available - renting a GPU from a cloud provider, using an ENT/PLUS licence that includes a certain transcription volume by 3CX or using a 3rd party transcription service, from the two build into the PBX.
If you are already using one of these options, you can just extend it for a while longer.
 
Very excited for GB10 and ability to use local transcription for multiple 3CX instances. Thanks for discussion and all the hard work!
 
Looks like we can get a Dell Pro Max with GB10 on 60 day trial.
It takes about a week for delivery.
Any time frame on when there might be something to test?
 
  • Like
Reactions: Evolute IT
No, you misunderstood. U8 is released for transcription with the hardware mentioned in the notes. No compatibility for GB10 has been announced at all.
 
  • Like
Reactions: Evolute IT
Apologies for the quick reply above - I only meant that I'm excited for future developments and possibilites of GB10 and understand that there is no release or timeline.
 
No problem at all! Just we can not commit on the fly to these types of things. A lot of work goes into it it has to be tested etc . Thank you for your understanding.
 
DGX Spark + Multi-PBX Transcription may be a game changer for anyone hosting multiple customers... Now if they could keep the price stable on memory and storage...
 
We are also investigating our options for deploying resources and hosting our own transcription servers. Our existing private cloud environment has server chassis which I've been advised wouldn't support the recommend GPUs. A couple of questions I was hoping you could provide some feedback on.

  • Would NVIDIA L4 cards work with the transcription server? They meet the 24GB memory requirement
  • Do we need a transcription server per 3CX instance or can this be shared by multiple instances? (reading the last comment from cmp1)
Thanks
 
  • Like
Reactions: ElenaF_3CX
Hello dear @mattcFuse2
If you want you can send the list of what your private cloud supports and I will be more than happy to help you choose.
We can do a shortlist exercise together.It's actually fun.
L4 is an excellent option yes. You can go ahead.
Re shared transcription server - Be a little bit patient but yes.. currently prototyping this. ;)
 
Great news about the shared transcription being on the roadmap.

We are using Dell PowerEdge so looking for SFF GPUs with no fans. The shortlist I made with my infra team was:

  • L4
  • A10
  • A30
  • A40
  • A2 - only has 16GB
Appreciate your thoughts on these options. Thanks for your response
 
  • Like
Reactions: ElenaF_3CX
Hello dear @mattcFuse2

L4 (2023) 24Gb - this is the newest card built on Ada - highly efficient. Send this to our final candidate list and then base your decision on price. (Would have been interesting if you posted the prices too)

A2 will be too small. Cannot use this. Drop it.

A10 and A30 OK - released together same year. 2021. A30 uses faster memory than A10. So Send both to shortlist but put A30 higher preference and judge by price.

A40 this is the oldest from the list - 2020 but still strong with 48GB. Normally used for video or large batch inference. Im curious on the prices. But I would drop it.

So for now, the candidate list looks like this - (top down with the highest = the most preferred.)

=====

#1: L4 Most modern architecture and incredibly efficient. 24GB And does this by drawing only 72 watts of power.
#2: A30 This guy is based on the GA100 chip which is the same heavy compute that the flagship A100 is built on. Its memory is faster which means it can put data in the GPU faster. But this one draws 165W of power.
#3: A10 Lacks Ada support of the L4 and the fast HBM2 memory of the A30. But its a great card but I would just use it if difference of 1 and 2 were too great. Draws 150W of power.

Hope this helps. When you have questions like this, send me prices so I can complete the analysis :cool: – now you left me hanging xD

Great work and a pretty nice mix your infra gave you. Well done!
 
In regards to the GB10 development is it planned to only port the transcription module?
or the entire phone system so the entire thing including transcription can run on GB10?
 
Hi
Definitely not the entire phone system.
Maybe you are missing the point of the GB10.
The GB10 is a desktop developer tool. The point of the GB10 is to bring data center-class capabilities to a desktop form factor.

Which means that developers can code and run their full AI solution on the $4k GB10. And if your code works on the GB10, it is architecturally "future-proofed" for a large-scale $40k H100 cluster.

Also we are not even sure about the onboard AI either. If we have to create a separate repository for the GB10, we will not do it.
 
  • Like
Reactions: Evolute IT
Hello dear @mattcFuse2

L4 (2023) 24Gb - this is the newest card built on Ada - highly efficient. Send this to our final candidate list and then base your decision on price. (Would have been interesting if you posted the prices too)

A2 will be too small. Cannot use this. Drop it.

A10 and A30 OK - released together same year. 2021. A30 uses faster memory than A10. So Send both to shortlist but put A30 higher preference and judge by price.

A40 this is the oldest from the list - 2020 but still strong with 48GB. Normally used for video or large batch inference. Im curious on the prices. But I would drop it.

So for now, the candidate list looks like this - (top down with the highest = the most preferred.)

=====

#1: L4 Most modern architecture and incredibly efficient. 24GB And does this by drawing only 72 watts of power.
#2: A30 This guy is based on the GA100 chip which is the same heavy compute that the flagship A100 is built on. Its memory is faster which means it can put data in the GPU faster. But this one draws 165W of power.
#3: A10 Lacks Ada support of the L4 and the fast HBM2 memory of the A30. But its a great card but I would just use it if difference of 1 and 2 were too great. Draws 150W of power.

Hope this helps. When you have questions like this, send me prices so I can complete the analysis :cool: – now you left me hanging xD

Great work and a pretty nice mix your infra gave you. Well done!
Thanks for this! I'll let you know how we get on and the final decision made
 

Members Online Now

No members online now.

Forum statistics

Threads
111,831
Messages
589,277
Members
164,660
Latest member
RJenkinsROCK