← Blog
✎article de blog

Experts in Host RAM

A close-up of a green graphics card PCB with a cooling fan and heatsink fins, teal ribbon cables in the foreground and blurred orange lights behind, the AMARBARO mark and wordmark centered over the board.

Routing a 35B MoE's expert weights out of VRAM into host RAM cuts the pack's VRAM footprint from 21 GB to 2.68 GB with bit-identical output on 20 of 20 prompts, and turning that from a working demo into a usable engine took three separate rounds, not one.

RegesCore-35B is a 256-expert mixture-of-experts model, top-8 routing, and its full weight pack is 21 gigabytes. An RX 7900 XTX has 24. That arithmetic works, barely, and it stops working the moment anything else on the card wants memory: a KV cache past a few thousand tokens, a second model, a desktop compositor. On 2026-09-15 this project built a version of the engine that loads 2.68 gigabytes of trunk weights into VRAM and leaves the other 18.3 gigabytes of expert weights sitting in host RAM, fetched on demand. It produced bit-identical output to the full-pack engine on 20 of 20 prompts. It also ran at 39 tokens per second, against 107 for the full-pack engine in the same stint.

That gap is the actual subject of this post. The identity result is the interesting claim, but a MoE engine that answers correctly at a third of the speed of the thing it's replacing isn't a result you'd ship, and the two rounds after it are what turned "correct and slow" into "correct and only a third slower, not two-thirds."

The mechanism: a cache in front of a cache miss

The two kernels that touch expert weights, amar_moe_gate_up_q4k and amar_moe_down_q4k, don't know how many experts exist. Each one reads expert e at e * stride from a base pointer it's handed, with the stride as a runtime argument. That's the fact that made the whole design cheap: a VRAM cache holding cap experts, indexed not by expert id but by cache slot, runs both kernels completely unmodified. The only new code is the thing that decides which slot holds which expert, and fetches the ones that are missing.

prepare() reads the router's top-8 back after it runs, checks which of those eight experts are already resident in the per-layer LRU cache, and issues a fetch for the rest before the gate/up kernel launches. The first version of that fetch was a blocking read through the host page cache, no pinning. It worked. It was also 40 host round trips per token, one for every layer's set of misses, and on the first prompt of the 20-prompt set, fetch_s came back at 1.74 seconds out of a 1.75-second decode. Nearly the whole cost of generating a token was waiting on the fetches, not computing anything.

The engine's own trace made this legible rather than mysterious: a 4.5 GB, 64-slot cache covers 76.7% of the requests this model actually makes for its top-8 experts, measured against an offline replay of 51,200 rows across 20 prompts before a line of the live tier existed. The live cache landed at 0.7743, inside the replay's predicted band. So the cache's policy was never the bottleneck. Nothing about what to fetch was slow; how it was fetched was.

A second instance: pinning changed what the cost actually was

The fix for the round-trip cost wasn't a smarter cache. It was pinning the host memory the experts live in, so the transfer path is a direct PCIe DMA instead of a page-cache read that the kernel driver has to fault in. BARO_TIER_PINNED=1 (which is the default since 2026-09-17) took the same design from 48.57 tok/s_gen page-cached to 67.12 tok/s_gen pinned, identity 20/20 unchanged. Nothing about the cache's contents changed. The only thing that changed was whether the host side of the transfer was a syscall into the page cache or a direct DMA source, and that difference alone was a 1.38x speedup.

This is worth sitting with, because it repeats a shape that shows up elsewhere in this project: two engines can execute the identical algorithm, on the identical data, and differ by 38% because of what the memory underneath them actually is. Pinning host RAM has a real cost of its own, 18.3 gigabytes locked out of the pageable pool on a 64 GB box that's also running other agents, and the project's own stage-1 measurement found pinned and pageable host-to-device transfer differ by only 0.6% in raw bandwidth. The 38% wasn't bandwidth. It was avoiding the page-fault path entirely.

A third instance: reading a miss in-kernel instead of copying it first

The next round asked a sharper question: if a gate/up or down kernel is going to need an expert's weights anyway, does it need those weights copied into VRAM first, or can it read them directly from pinned host memory while it runs? BARO_TIER_ZC (zero-copy, opt-in, requires pinning) answers that for two of the three kernel families: a missed q4_k expert is read straight from the pinned host store by the kernel doing the math, with the fetched slot filled in for next time as a side effect of the same pass. The three q6_k down layers still copy first; they weren't converted.

Measured: 67.71 to 72.02 tok/s_gen, a 1.064x gain, identity 20/20 held, and kernels/test_moe_block.mojo's zero-copy arm checked bit-exact against the copy-first arm. A standalone probe (bench/tier_zerocopy_probe.mojo) explains the shape of the win and its limit: in-kernel pinned-host reads hit 24.6 to 27.2 GB/s when 8 to 24 pieces are in flight at once, similar to the 21.9 GB/s of a per-piece DMA copy, but drop to 14 to 15 GB/s at only 2 pieces in flight. Zero-copy wins when a kernel naturally has enough independent loads outstanding to hide the latency of reading over PCIe; it doesn't automatically beat a copy when it doesn't.

The rule

A cache that fetches correctly is not the same claim as a cache that fetches fast; measure separately what you fetch and how.

What counts as evidence here, and what doesn't

What counts: - A 20-prompt identity check against the full-pack engine, not a handful of hand-picked prompts, because the whole claim is that pulling 18.3 GB of weights out of VRAM changes nothing about what the model says. - fetch_s read back from the engine's own per-request counters, because "the cache should be fast" is a prediction, not a measurement, and this project's own protocol rules exist specifically because a passed flag is not evidence it took effect. - A same-arm comparison of pinned against pageable at identical cache contents, which is what isolated the 1.38x to the transfer path rather than to the cache design.

What doesn't: - The offline replay's hit rate alone. It predicted this design would work; it didn't and couldn't predict the 40-round-trip cost, because a replay against a trace doesn't have a host round trip to measure. - A single-prompt tok/s figure. The 39.06, 67.12, and 72.02 numbers above are all 20-prompt medians; anything less would hide exactly the kind of stall a single unlucky prompt full of cold experts can produce.

What it cost

Two commits worth of engineering time went into a design whose headline speed, even after both fixes, is still 72.02 tok/s against 111.89 for the full-pack engine in a comparable stint. The tier is not a replacement for having enough VRAM; it's a way to run a model that doesn't fit at all, at roughly two-thirds the speed of one that does. Whether that trade is worth it depends entirely on whether the alternative is "don't run the model," and this project has not measured what the same design does on a model whose experts genuinely don't fit in 24 GB, because tracing that model's routing locality is unfinished work.

Pinning 18.3 GB of host RAM is also a standing cost on a shared machine. It's memory nothing else on the box can use for the duration the engine runs, on a box that at various points in this project has three other agents active.

How to prove this wrong

The identity claim is checkable directly: build the engine, load qwen35moe with BARO_TIER=64, run the same 20 prompts against the full-pack engine, and diff every token. A single divergence anywhere falsifies the "bit-identical" claim in this post. The speed claims are checkable by reading fetch_s and decode_s back from the engine's own request log rather than trusting the tok/s number in isolation; if pinning or zero-copy don't move that ratio the way described here, the mechanism explanation in this post is wrong, not just the number.

This runs on one RX 7900 XTX. It says nothing about whether the same tier design helps or hurts on a card with a faster or slower PCIe link, a different host memory bandwidth, or a NUMA topology that puts host RAM further from the GPU's root complex than this one desktop does.

Provenance

AMD RX 7900 XTX (gfx1100, RDNA3), 24 GB, one card. ROCm 7.2. Mojo 1.0.0 via max[all]==26.5.0. Tier module serve/expert_tier.mojo, commit 8f8c8c3; wiring f798ed4; pinned default 5ad2e2d; zero-copy 404ac04. Reports: exchange/lane-B4-stage2b-report.md, exchange/lane-MOE3-report.md, bench/moe-tier-protocol.md. Figures cited from docs/BASELINE.md and the reports above, in amarbaro/mojo-baro.

Commentaires

Pas encore de commentaires.

Se connecter pour commenter.