Two AMD Ryzen AI Halo PCs Run 397B Quen 3.5 in Local Test
TL;DR
A local test claims that two PCs running AMD's Ryzen AI Halo can execute Qwen 3.5 at 397 billion parameters. Each unit carries 128 GB of shared memory, and clustered together the capacity is said to cover models up to 400 billion parameters. For developers this points at a new class of compact on premise systems able to run large models without a cloud connection. The figures come from a single Geeky Gadgets report, with speed, quantization, power draw and cost still unspecified.
Nauti's Take
The first test should use a real workflow and measure latency, tokens per second, power draw, and the quantization actually required. Two devices with 256 GB of shared memory sound promising for local prototypes, yet without solid benchmarks it remains unclear whether the cluster improves speed or mainly adds capacity.
Small teams should verify availability and total cost before comparing the setup with cloud spending.