# Tag

WMMA

Posts with this tag: 2

← All posts
GPU kernels A 35B MoE on one consumer card · part 3

Prefill as a Batch

A chunked prefill path for a 35B MoE passes identity on 23 of 23 prompts and runs 3.7 to 7.9 times faster than replaying the decode kernel one token at a time. Two variants that would have been faster still, WMMA attention and a dense SSM scan, reorder floating-point sums and stay off by default because they fail the same identity check.

Read