LocalLLM: Running Qwen3.8-Flash-Next on Strix Halo (Asus ProArt PX13)

1. tl;dr

See the final command at the end.

Laptop Asus ProArt PX13 (Strix Halo)
CPU AMD Ryzen AI MAX+ 395, 16 cores
GPU Radeon 8060S iGPU
RAM 128 GB unified memory
OS CachyOS, Linux 7.2.7
Model Qwen3.8-Flash-Next-125B-A6B
Quantization UD-Q4_K_XL
Context 256k
Prompt processing ~900 tokens/second (AC) / ~500 (battery)
Token generation ~35 tokens/second (AC) / ~20 (battery)

In my previous post, I ran a 35B Mixture-of-Experts model on a 12 GB GPU, carefully moving experts to the CPU until the VRAM was full. This time, there is no VRAM juggling at all: a 125B model just fits into 128 GB of unified memory, and it is fast enough to actually enjoy working with it.

2. Why Strix Halo

The story so far: in 2023 I experimented with local AI, in June 2026 I ran Qwen3.6-35B-A3B on an RTX 4070 Ti, and both times the bottleneck was memory: either VRAM or bandwidth.

AMD's Strix Halo (Ryzen AI Max) is the first mainstream chip that removes this bottleneck: an iGPU with up to 128 GB of unified memory. No more splitting experts between CPU and GPU, no more quantizing down because the model doesn't fit. You just load the model and it's there.

So I got one: an Asus ProArt PX13 with 128 GB RAM, running CachyOS on Linux 7.2.7. I paid 3.000 EUR for it (open box), regular price was 3.600 EUR and now it is not even available anymore.

I also considered getting a Macbook Pro with an M5 Ultra and 128 GB, but Apple just raised the prices from ~6.000 EUR to 7.800 EUR, so that was no option anymore.

I was also eyeing the GPD Win 5, but you could not order it in Germany with 128 GB. And there was the Bosgame M5 mini PC, available for 2.500 EUR, but for the small price difference I found a laptop the more attractive option.

3. The model: Qwen3.8-Flash-Next-125B-A6B

The interesting part is not the hardware alone, it is the model generation. Qwen3.8-Flash-Next is a Mixture-of-Experts model with 125B total parameters, but only ~6B active parameters per token.

This combination suits Strix Halo well:

I use the Unsloth Dynamic quantization UD-Q4_K_XL, which comes as four GGUF shards:

Qwen3.8-Flash-Next-UD-Q4_K_XL-00001-of-00004.gguf

It fits nicely into the 128 GB together with the 256k context, the MTP model and the vision projector, and in my usage the quality is good enough to not think about quantization at all.

4. Serving with gufo in Docker

I serve the model with gufo, a local inference engine specifically optimized for Strix Halo, in Docker. It provides an OpenAI-compatible API that I use from opencode. The container talks to the AMD GPU via /dev/kfd and /dev/dri:

Two features are worth noting:

I switched engines more than once to get here, and every switch was about one number: prompt processing. In the order I tried them:

Engine License Technology Prefill (tok/s) Token generation (tok/s) Comment
llama.cpp MIT Vulkan, ROCm – – did not support the model at all when it launched
apepojken MIT Vulkan Plugged-in: 200
Battery: 100
Plugged-in: 40 the llama.cpp fork that made the model runnable at all, with RDNA 3.5 kernel work; stable, but prefill stayed slow
pwilkin MIT ROCm (custom ROCr and HIP) README: 1.180 README: 27 the first stack that was open source and fast at prefill; but you bring the ROCm yourself, on a pinned llama.cpp branch, so the setup is a project of its own
halogen proprietary ROCm (own kernels) README: 1.424 README: 46 blistering, every kernel written for this one GPU and this one model family; but a binary container under an EULA
gufo MIT ROCm (HIP) README: 1.600
Plugged-in: 900
Battery: 500
README: 32
Plugged-in: 35
Battery: 20
fast like halogen, open source, and it carries its ROCm inside a single Docker image: this is where I settled

My numbers are all at 64k context, plugged in, on the Balanced power profile at 50 W TDP. The README rows are the authors' own benchmarks at whatever prompt length they chose, so read them as a ballpark next to my 64k figures rather than as a shoot-out. apepojken's fork is still a little ahead on token generation, but for agents the prefill is the half you wait for: 4-5x there is what made me stop looking. The benchmarks go into more detail than the numbers in this post, with results per context length, and they match well what I see day to day.

This matters a lot for agentic coding: every tool call sends the whole conversation back to the model, so a large part of your wall-clock time is spent re-processing the context. At 200 tokens/second, a 100k context means an 8-minute stall. At 900 tokens/second, the same re-evaluation takes under two minutes, and with caching, only the incremental part has to be processed on most turns.

With 256k context and 8k max output tokens, there is plenty of room for real coding sessions without constant compaction.

5. Experience in opencode

This is where the "models got MUCH stronger" part comes in.

Last time, my honest conclusion was: cloud models are still far superior, the model loops from time to time, give it small tasks only.

This time, it is just fun.

I handed it a mid-sized Maven project, and it built it autonomously: reading build errors, narrowing down the cause, fixing, rebuilding, repeat. It found the critical points on its own, without me explaining Maven's usual traps.

What changed compared to my June setup:

6. Not perfect yet: amdgpu crashes

To be honest: from time to time, the amdgpu driver crashed on me. I don't know yet what triggers it. I have not found a reliable reproduction. But Linux managed to recover the driver every time, so I lost a running generation, not the machine.

This is the kind of rough edge you should expect when running an iGPU at sustained full load on Linux. It did not stop me from enjoying the setup, but if you want a turn-key experience, be aware of it.

7. Power profiles matter

Performance depends heavily on two things:

  1. Whether the laptop is plugged in.
  2. Which power profile is active.

So if you benchmark a Strix Halo laptop and get disappointing numbers, check the power profile first. Mine is a laptop, not a desktop replacement that runs at full tilt all the time, and that is fine, but you should know it when comparing tok/s figures from the internet.

8. Not only for AI: a do-it-yourself Steam Machine

Strix Halo makes a fine gaming machine, too. In my experience many games run well on the 8060S at 1080p, and it competes well with Valve's new Steam Machine. With one machine you can switch between a gaming session and serving a local AI model: that beats two boxes under my desk.

9. Outlook: audio, image, and video

What also appeals to me about gufo is the direction: it wants to be a one-stop shop for local AI on this hardware. Its model table already includes speech recognition (Qwen3-ASR) and text-to-speech (Qwen3-TTS), while image generation (Qwen-Image) and video generation (MiniMax H3) are listed as in progress.

I recently played around with MiniMax H3 on my desktop and was impressed by what it produced. If it runs on the PX13 through the same tool one day, I would have text, audio, image, and video generation in one place, on my own hardware.

As far as I can tell, the video support today is text-to-video; image-to-video driven by an additional text description does not seem to be there yet. Something to watch.

10. Comparison to the cloud

Cloud models are still ahead, but honestly: for daily coding work, the gap no longer feels like a gap. The model builds my Maven project, finds the real issues, and does not need me to babysit it through every step.

And it runs entirely on my machine, on my desk, on battery, with no prompt ever leaving the laptop.

Three years after my first local experiments, that used to be a compromise. Now it is just a preference.

11. Final command

docker run --rm \
  --user 1000:1000 \
  --group-add render \
  --device /dev/kfd \
  --device /dev/dri \
  --ulimit memlock=-1 \
  -p 8080:8080 \
  -v /srv/llama/models/Qwen3.8-Flash-Next-125B-A6B:/models/:ro \
  ghcr.io/gufo-org/toolboxes/gufo-runtime:20260924T104050 \
  gufo serve llm \
  --model /models/UD-Q4_K_XL/Qwen3.8-Flash-Next-UD-Q4_K_XL-00001-of-00004.gguf \
  --speculative mtp \
  --mtp-model /models/mtp-Qwen3.8-Flash-Next-shared-Q8_0.gguf \
  --mmproj /models/mmproj-BF16.gguf \
  --context 262144 \
  --max-tokens 8192 \
  --host 0.0.0.0 \
  --port 8080