GPU kernelsKernels on a gaming GPU · part 4
Beating llama.cpp at Decode, Without Speculating
The trunk decode path had been sitting a few percent behind llama.cpp for two days running. Quantizing to Q4_0 and fusing the whole forward pass into one persistent GPU launch closed that gap and passed it — 130.7 tok/s against llama.cpp's 110.0, no speculative decoding on either side.
Read