GPU kernels
A 35B MoE on one consumer card · part 3
Prefill as a Batch
A chunked prefill path for a 35B MoE passes identity on 23 of 23 prompts and runs 3.7 to 7.9 times faster than replaying the decode kernel one token at a time. Two variants that would have been faster still, WMMA attention and a dense SSM scan, reorder floating-point sums and stay off by default because they fail the same identity check.
Read