# Tag

MoE

Posts with this tag: 4

← Alle berichten
Practice

Keeping the Failures In

BARO_SPEC defaults to on. Setting it on the MoE model changes nothing, because the actual gate in the code is BARO_SPEC and a model flag that is always false for that architecture, and the document says so in the same line rather than in a footnote.

Read
GPU kernels A 35B MoE on one consumer card · part 3

Prefill as a Batch

A chunked prefill path for a 35B MoE passes identity on 23 of 23 prompts and runs 3.7 to 7.9 times faster than replaying the decode kernel one token at a time. Two variants that would have been faster still, WMMA attention and a dense SSM scan, reorder floating-point sums and stay off by default because they fail the same identity check.

Read
Systems design A 35B MoE on one consumer card · part 1

Experts in Host RAM

The routed experts of a 35B MoE moved to host RAM, VRAM footprint down to 2.68 GB from 21, output identical on 20 of 20 prompts. The first working version ran at 39 tok/s. Pinning the host store and reading misses in-kernel got it to 72.

Read