No GPU, No Downside: Flagship LLMs On A GPU-less Teenaged Server
There are a lot of causes to run LLMs domestically, from privateness considerations to only eager to futz about with the know-how, however if you wish to run the large fashions, the usual logic is that you just want huge cash for heaps and many VRAM. [MattMo] is asking that into query with his recent video — embedded under, naturally — by which he will get GLM 5.3 Flash, Qwen 3.8 Flash and Qwen 3.8 27B all operating on a 14-year-old server with out a single GPU.
The secret, for those who can name it that, is that they aren’t operating very quick: 4 tokens per second was in regards to the max. Those 4 tokens are excreted from the dual Xeon processors of the classic Dell PowerEdge R720 server, with the fashions residing in its 348 GB of DDR3 system reminiscence. That’s sufficient even for the biggest flagship fashions, however as you’ll be able to see by the pace, issues are a bit bottlenecked by having solely 20 threads obtainable between the 2 processors. Said processors are additionally sufficiently old to lack sure directions that may have helped pace issues up. Still, [MattMo] argues within the video that that is greater than only a dancing bear: there are workloads the place batch-processing at 4tps would possibly make sense, and there are individuals who have already got servers of this class laying round of their homelabs. The intersection of that Venn diagram might be fairly lonely, but when that’s you — hey! [MattMo] says it’ll work, so give it a shot.
If you had to purchase the {hardware}, effectively, it’s additionally fairly cheap on the second-hand market, with [MattMo] estimating about $600 given prevailing costs. He additionally factors out that the newer fashions should get some optimization to enhance speeds, however don’t anticipate real-time conversations with Hal 9000. Still, when it comes to native LLMs, it actually beats the pants off toy fashions operating through Llama on the PSP or the C64.


