A 35B MoE went from 42.88 to 111.89 tok/s_gen on one RX 7900 XTX in a single day, and the gains came from removing per-element scalar dequant loads, a serial top-8 router, and about a third of the kernel launches per token, not from any new algorithm.
At the start of the session on 2026-09-15, RegesCore-35B decoded at 42.88 tokens per second on one RX 7900 XTX, less than half of llama.cpp's 109.4 on the same GGUF. By the end of that day it read 111.89, ahead of llama.cpp's same-GGUF number of 109.92 by a hair, and landed there through a round that missed its own preregistered performance bar and shipped anyway, on an explicit override. This is what happened between those two numbers, and it's worth writing down precisely because none of the individual changes are interesting engineering. They're all the same lesson: measure per-kernel device time before guessing, because the actual bottleneck is rarely the one that looks clever to fix.
The first three rounds of the day (bench/moe-perf-protocol.md) came from one rocprofv3 kernel trace read before any code changed. Two things dominated device time: the expert weight kernels were dequantizing Q4_K and Q8_0 blocks with per-element scalar loads, one value at a time out of packed 4-bit or 8-bit storage, and the top-8 router was running its eight argmax passes serially inside a single kernel launch rather than overlapping them with anything.
R1 (8130f65) vectorized the Q4_K expert dot product's dequant loads. R2 (d49bfc3) did the same for the Q8_0 row-dot path. R3 rewrote the router as a one-wave kernel instead of the serial version. None of these are new numerical methods; they're the same arithmetic reading memory in wider, better-shaped chunks. The 20-prompt medians moved 42.88 to 55.98 to 71.89 to 93.46, agreement against the reference held at 53.15 out of 64 against a 53.20 starting point (both inside the accepted band), and llama.cpp's own number on the same GGUF, 109.4, went from a target more than double this engine's speed to one within 15%.
That the fix was "load memory correctly" rather than "invent a better kernel" is not a small point. A model that looks compute-bound from the outside can be memory-access-bound in a way that a naive read of the algorithm hides completely; the actual dequant math per element is trivial, and trivial math executed one scalar at a time is where the 42.88 came from.
The next lever wasn't inside a kernel at all. It was how many kernels got launched per token. bench/moe-launch-count.sh, backed by an actual rocprofv3 kernel trace rather than a guess, found 1217.0 launches per token on the profile at the time. R4 (7145b71) traced two of them to real bugs: 60 memsets zeroing SSM gate partial planes that a different kernel already writes into rather than accumulates into, so the zeroing was pure waste, and 40 host-buffer copies giving the shared expert an index of zero that was also silently clobbering the router's own idx[0]. Removing exactly 100 unnecessary launches, verified by the trace count itself (1217.0 to 1117.0, not an estimate), moved the 20-prompt median from 94.68 to 99.92, a 1.055x gain, with identity 20/20 bit-identical to the previous engine throughout.
R6.0 (1e270b5) went after five more launch folds, bringing the count to 857.0 from 1117.0 and the median to 107.27 from 94.79, a 1.132x gain, identity 20/20 held again. That round also carried a persistent-kernel design for the whole MoE token, re-estimated at the time as an extra-large piece of work, and Mario's call was to defer it rather than build it that day.
R6.0b (4f9929d) folded further, from 857 to 727 launches, and measured 111.89 tok/s_gen against 107.28 for R6.0, a ratio of 1.043x. The round's own preregistered kill line, set before the round ran, was +5%. It landed at 4.3%, below the line it had set for itself, and the commit was kept only on Mario's recorded override rather than the round's own passing grade. Identity held at 20/20 bit-identical throughout.
The same-stint comparison against llama.cpp is worth stating precisely rather than rounding: llama.cpp's own run in that stint measured 109.92, while the R6.0 build measured 106.95 in the identical session, a ratio of 0.973x, meaning the earlier build was behind llama.cpp in that specific stint even though it beat llama.cpp's number from an earlier session. Both llama.cpp figures, 109.4 and 109.92, are real and both are correct; they're two different runs, and a MoE engine's throughput on this hardware moves enough between sessions that the comparison arm has to be measured freshly rather than reused.
A kill line set before the round exists to be obeyed, and an override that keeps a round below its own line has to say so out loud in the same place the number is reported, not just in a commit message nobody reads afterward.
What counts:
- A kernel trace naming an exact launch count, 1217.0, 1117.0, 857.0, 727.0, from rocprofv3, not an estimate from reading the source.
- A 20-prompt median with identity checked bit-exact at every step, because a launch-count reduction that also changed a result silently would be worthless no matter how fast it ran.
- Naming which round missed its own bar and by how much, 4.3% against a 5% line, rather than rounding the whole day's story up to a clean success.
What doesn't: - A single-prompt speedup. Every figure in this post is a 20-prompt median; a MoE model's per-token compute depends on which experts a given prompt routes to, so a lucky prompt tells you almost nothing about the model's typical cost. - Treating "faster than llama.cpp" as one fixed comparison. The two llama.cpp numbers in this post, from two different stints, differ by more than half a percent from each other; the comparison arm needs re-measuring in the same session, every time.
The most expensive thing this day cost wasn't GPU time, it was the temptation to call 107.27 a finished number and move on. R6.0b's own preregistered threshold said it hadn't earned its place, and the round shipped anyway on a human call rather than a passing gate. That's now the visible fact next to the number in docs/BASELINE.md, rather than a detail buried in a commit log: 111.89 tok/s_gen, landed BELOW its own preregistered +5% kill line, kept only on Mario's recorded override. A reader who takes the headline number without that sentence is missing the actual state of the work.
Every commit named above is real and buildable: 8130f65, d49bfc3, the R3 router change, 7145b71, 1e270b5, 4f9929d. bench/moe-launch-count.sh reproduces the launch counts from a live rocprofv3 trace rather than from this post's word, and bench/moe-perf-protocol.md and bench/moe-persist-protocol.md carry the per-kernel device-time receipts the rounds were decided on. Run the 20-prompt suite at each commit and check the medians against the numbers above; if the launch counts or medians don't reproduce within the reported spread, this post's central claim is wrong and should be corrected.
Every number in this post comes from one RX 7900 XTX. It says nothing about whether the same scalar-load and launch-count fixes matter on a card with more compute relative to its memory bandwidth, where the original bottleneck might not exist in the same shape, or none at all.
AMD RX 7900 XTX (gfx1100, RDNA3), 24 GB, one card. ROCm 7.2. Mojo 1.0.0 via max[all]==26.5.0. rocprofv3 for kernel-launch and device-time traces. Every figure above traces to the commit named beside it or to bench/moe-perf-protocol.md / bench/moe-persist-protocol.md / docs/BASELINE.md in amarbaro/mojo-baro.
Kommentare
Noch keine Kommentare.
Anmelden um zu kommentieren.