AMD has introduced a desktop AI workstation capable of holding up to 576 GB of HBM3e memory and delivering 16 TB/s of memory bandwidth, according to an announcement covered by The Register. The system, dubbed the Threadripper Halo, was revealed at IFA 2026 and represents the first time AMD's Instinct accelerators have appeared in a workstation configuration. The company is targeting machine learning researchers who want to run large AI models locally, with a launch planned for next year.
The Threadripper Halo is built around a 96-core Threadripper PRO 9995WX processor, which first appeared in 2025, and supports up to 2 TB of DDR5 memory for a total of up to 2.6 TB of combined system memory, according to the report. The system can be equipped with as many as four PCIe-based Instinct GPUs, though the units on display at IFA contained only two. Each MI350P accelerator—essentially half of an MI350X adapted to PCIe form—packs 144 GB of HBM3e memory capable of 4 TB/s bandwidth and up to 4.6 petaFLOPS of FP4 compute in its 600 W configuration, though AMD is expected to run the cards at a more modest 450 W setting in this workstation. With four accelerators installed, the system promises 576 GB of HBM3e and 16 TB/s of memory bandwidth. AMD has not disclosed pricing, but the system is expected to cost between $100,000 and $150,000 as configured.
The report notes that the combined memory enables the system to run models exceeding a trillion parameters in size at four-bit precision entirely in GPU memory. By offloading portions of the model to system memory, the workstation should be able to handle the largest open-weights models, including Moonshot.AI's 2.8 trillion-parameter Kimi K3. AMD claims the system offers up to 3.4 times the total system memory and more than twice the memory bandwidth of Nvidia's DGX Station, which features a 252 GB B300 GPU, a 72-core Grace CPU, 496 GB of LPDDR5x memory, and retails for around $100,000 when available.
The massive memory capacity addresses a key bottleneck in running frontier-class AI models: keeping model weights accessible without constant shuttling between slower storage and faster compute, according to the report. This advantage should translate to faster LLM inference, provided tensor parallel operations don't hit bottlenecks on the CPU's PCIe bus. But deploying a quad MI350P configuration will push the limits of standard North American power outlets unless AMD either underclocks the cards from 600 W to 300 W or requires a 20-amp circuit, which may explain why the demonstration units at IFA contained only two accelerators. The ability to run trillion-parameter models on a desk-based system represents a shift from cloud-dependent workflows to local control, though the price point and power requirements limit the market to well-funded research labs and institutions. AMD says the Threadripper Halo will ship starting next year, though availability may vary by market depending on electrical infrastructure. Organizations weighing this investment face a trade-off between model control and operational complexity, particularly as power and thermal management become gating factors even before budget constraints enter the conversation.

