AMD's new workstation pairs a 96-core Threadripper with 576GB of GPU memory

Skye Jacobs

Posts: 2,162   +62
Staff
First look: AMD used its IFA 2026 keynote to introduce the Threadripper Halo Station, a liquid-cooled workstation designed for running AI models with more than one trillion parameters locally rather than through a cloud service. AMD has not said when the Halo Station will be available. But the system points to a growing market for workstation-class AI hardware that gives developers more direct control over large-model workloads, including where data is processed and how models are deployed.

The system pairs the 96-core Ryzen Threadripper PRO 9995WX processor with up to four Instinct MI350P accelerators. AMD demonstrated a two-card version at IFA, while a fully configured model with four accelerators would provide 576 GB of HBM3E memory – enough, the company said, to hold a trillion-parameter model on the machine itself.

AMD has not announced pricing, availability, or full system specifications. It also did not detail the Halo Station's storage, networking, or chassis options, or whether it plans to sell the system directly or through workstation manufacturers and system integrators.

Each Instinct MI350P card includes 144 GB of HBM3E memory and delivers 4 TB/s of memory bandwidth. The two-card system demonstrated at IFA had 288 GB of HBM3E memory. A four-card version would provide 576 GB.

That memory capacity is central to the Halo Station's purpose. Large models require substantial high-bandwidth memory to store weights and process data efficiently. AMD said the MI350P's 4 TB/s of bandwidth is 14 times higher than that of any LPDDR5X variant.

The system also has significant power and cooling needs. Each MI350P accelerator can draw up to 600 watts of total board power. AMD said the cards are independently liquid-cooled, and the Threadripper processor has its own liquid-cooling setup.

The Threadripper PRO 9995WX, code-named Shimada Peak, is based on AMD's Zen 5 architecture. It has 96 cores, 192 threads and a boost clock of up to 5.4 GHz. The processor supports up to 2 TB of DDR5 system memory.

Jack Huynh, AMD's senior vice president and general manager of Computing and Graphics, described the product as a new workstation category intended to bring "supercomputer-class compute" to individual users and developers.

Permalink to story:

 
What's the price of such thing? And "what" kind of models can you run on these things?

Like I have a Pro GPT subscription, 229 EU a month or 2500+ a year. I figured if a thing like that would be faster, would be better, or have the latest or high tech models that are available, and it can work on whatever I throw at it, the expense would be made quite easy.
 
What's the price of such thing? And "what" kind of models can you run on these things?

Like I have a Pro GPT subscription, 229 EU a month or 2500+ a year. I figured if a thing like that would be faster, would be better, or have the latest or high tech models that are available, and it can work on whatever I throw at it, the expense would be made quite easy.
The local AI models have been growing a lot recently and you can run models that require about 60-120GB VRAM on a single card like that with enough space for context (like Qwen 3.6 27B FP16 or Devstral-2 123B Q8)
 
What's the price of such thing? And "what" kind of models can you run on these things?

Like I have a Pro GPT subscription, 229 EU a month or 2500+ a year. I figured if a thing like that would be faster, would be better, or have the latest or high tech models that are available, and it can work on whatever I throw at it, the expense would be made quite easy.
Well, a fully specced out 9995x machine tends to run 25k+… without the cards or exorbitant RAM… so… ballpark 100k maybe?
 
Separately, Microsoft announced that Project Zenith would not be available for these machines as as they are not Windows Copilot certified. (I jest.)

Local hardware is going to be very appealing for certain cases, especially where compliance demands the data not leave the premises, and for customized models and adjacent services. Particularly when compared against higher per-token cloud API costs vs. the subsidized subscription plans.

For the person whose needs fit a $200 or less monthly subscription, that value is going to be tough to beat with dedicated hardware. The hardware will be much more expensive, and may also not efficiently provide the same burst where say you may be running 5 parallel sessions with multiple agents each during the work day while doing zero for half the hours in a week or more.
 
Who are all these people coming out of the woodwork to run these LLM's and for what purpose? Every single hardware release is just about running AI.
 
Who are all these people coming out of the woodwork to run these LLM's and for what purpose? Every single hardware release is just about running AI.

As I used to be skeptic just like everyone else, it is amazing what AI these days is capable of, and I'm not talking about your average chatbot.

I mean if you have some degree of coding, or some degree of good idea's, I suggest give it a try. We are completely moving to writing prompts now.

I'm managing 4000 websites by now through AI and through a single prompt. Where a client used to ask, hey, the number of Whatsapp has expired, can you change it? I'd simply type a prompt and it changes it for me in a minute including a rollback copy. Either through the website interface, FTP or even root access on the complete server.

I managed to code a **** ton in less then 2 weeks where I normally would take 6 months over considering the research that was needed or the expertise I needed to hire elsewhere. I no longer need it and I can work all on my own here. For my case it is worth the 229 EU a month easily. And I have yet to max it even out considering I run heavy loads on a daily basis.

Yes it still is not perfect, but you as a human are required to (bug) test things and fix these. I came up with so many weird idea's that I would never had any idea on even how to work it out. Simple example was, "Design me a PCB or Slotket for not Slot 1 but Slot A in where I can start using a Athlon XP CPU". And it pretty much printed all the schematics for me including everything else required.

d21d0ba0-9492-4994-bcc4-0039c6aaf7ff.png


Not saying it worked out of the box, but with the help of AI you can do really crazy idea's, things, that you as a unskilled individual have no sense of how to even do it, can get along very well. The options with Ai are endless.

Got a weird engine noise? Make a video and ask Ai to determine what's going on with the engine. You be surprised how accurate it is. I have yet to catch it hallucinate.

You run cloudflare on your websites? And your having 600 zones active in an account? Instead of manually applying any change into or over 600 zones you now prompt it and it runs everything in less then 3 minutes where normally I would be pushing for over 45 minutes and be exhausted lol.

Blind server installation - application management - monitoring - you name it. I do everything from my phone now. So yeah if there's a alternative for running one locally and where models are on par with current or latest Astra then ill buy one.
 
Last edited:
Back