int8 WMMA Is Not Faster
v_wmma_i32_16x16x16_iu8 issues at 1.007x the rate of the bf16 WMMA instruction on gfx1100, both at 133 T-ops/s. An int8 MMQ prefill kernel built on the expected 2x lost to the project's own bf16 kernel, 0.91x at n=1024, and the lane closed negative.
Read