← Blog
✎article de blog

int8 WMMA Is Not Faster

Two identical machined aluminium engine blocks side by side on a steel bench under matched cold teal lights, amber light inside their bores, the AMARBARO mark centered between them.

On gfx1100, the int8 WMMA instruction issues at the same rate as the bf16 WMMA instruction, so an int8 GEMM built to exploit a spec-sheet 2x cannot beat a well-scheduled bf16 GEMM on this card.

The RDNA3 spec sheet advertises int8 matrix-core throughput at twice the rate of bf16. That number is real, for some instruction shapes. It is not real for the one this project's prefill GEMM needed, and a full lane's worth of int8 MMQ kernel work was built and measured before that stopped being an assumption and became a receipt.

The measurement that mattered

bench/bench_wmma_peak_i8.mojo isolates the instruction itself: a register-only loop, no memory traffic to confound the result, interleaved issue, seven repeats. It measures v_wmma_i32_16x16x16_iu8 against v_wmma_f32_16x16x16_bf16 on the same 7900 XTX.

Ratio: 1.007. Both instructions issue at 133 T-ops/s, which works out to 512 operations per clock per compute unit for each. Not close to 2x. Effectively identical.

That single number decides the rest of this post. An int8 GEMM on this hardware cannot beat the same tile schedule run in bf16, because the instruction it would need to be faster on issues at the same rate as the one it is competing against. Whatever gain int8 offers on this card has to come from somewhere other than raw matrix-core throughput, most plausibly narrower memory traffic from packing weights at half the byte width. On a kernel that already keeps its weight traffic close to the wire, that gain has nowhere to hide.

What was built on top of the wrong assumption anyway

The kernel work did not stop at the peak measurement, because a slower matrix core does not automatically mean a slower end-to-end GEMM if the memory traffic saved elsewhere makes up the difference. kernels/matmul_mmq.mojo implements amar_quant_q8 (converting bf16 activations to int8 per 32-element block, carrying an f32 scale and an accumulator-order bias term) and amar_matmul_lds_q4, an LDS-pipelined schedule instantiated for both int8 and bf16: eight waves in a 4x2 arrangement, a 2x4 wave tile, a block-K of 32 matching one q4 weight block, double-buffered shared memory with an XOR swizzle, one barrier per step. The bf16 instantiation of this same schedule is bit-exact with the project's existing amar_matmul_prefill_q4 kernel, which is what makes the comparison between the two instantiations fair: same schedule, same everything except the operand dtype.

The result closed the lane negative. Against a gate of at most 2092 microseconds at n=1024, the int8 MMQ kernel measured 2486 microseconds, a fail, and 0.91x of the bf16 version of the identical schedule, which itself stood at 1.86x over the project's earlier R4 prefill kernel. bf16-lds, not int8, is the keeper.

The ablation that explains why: the weight (B operand) load occupies 46 to 49 percent of the inner K-loop on both the int8 and bf16 arms, not the roughly 10 percent that had been predicted going in. Int8 halves the byte width of the A operand, the activations, which was never the expensive side of this particular loop. Halving the cheap operand while the matrix core itself runs no faster leaves nothing to spend the saved bytes on. A follow-up attempt at a uniform loader, meant to rebalance where the loop spent its time, produced no measurable gain either.

Why this happens

RDNA3's WMMA int8 path is built for a shape and issue pattern that does not automatically transfer to every GEMM tiling. The spec-sheet 2x is a property of the instruction under favorable conditions, not a guarantee that any kernel using it inherits the multiplier. Here, the actual bottleneck inside the K-loop was the weight load, which int8 does not touch, and the matrix core itself, measured directly and in isolation, showed no headroom to exploit even if the loop had been rebalanced. Two separate places to look for a win, and both came back empty on this card, for this schedule.

The rule

A vendor's per-instruction throughput number is a claim about that instruction, at that shape, under conditions the vendor chose. Measure the instruction your kernel actually issues, in isolation, before rescheduling a kernel around it.

What counts and what does not

What counts as having tested the assumption: an isolated, register-only peak measurement of the exact instruction the kernel will issue, repeated enough times to rule out one-shot noise, run on the same card the kernel will ship on. That is what bench_wmma_peak_i8.mojo did, and it is what turned "int8 should be 2x" from a spec-sheet number into a falsifiable, falsified claim before a full kernel was built around it.

What does not count: reading a peak-throughput figure off a vendor spec sheet and treating it as the ceiling a new kernel can reach. It did not hold here, on this instruction shape, and there was no way to know that without measuring the instruction directly.

What it cost

A full kernel, amar_quant_q8 plus an int8 instantiation of the LDS-pipelined schedule, plus a parity test and an ablation harness, plus a follow-up uniform-loader attempt, all closed negative. None of it shipped. What it bought is a receipt that will stop the same idea from being re-proposed on this card: BARO_DOT, the int8-dot path this lane also touched, defaults off for exactly this reason, and the standing note against reproposing int8 WMMA for speed on gfx1100 is now written down rather than left to be rediscovered by the next person who reads the spec sheet.

How to prove this wrong

Run bench/bench_wmma_peak_i8.mojo on your own gfx1100 card. If v_wmma_i32_16x16x16_iu8 measures meaningfully above 1.0x relative to v_wmma_f32_16x16x16_bf16, the central claim here is card- or driver-specific and does not generalize the way this post assumes. A ratio at or near parity, matching the 1.007 measured here, would support it. Either way, that is one command and does not require building the full MMQ kernel to check.

Provenance

AMD RX 7900 XTX (gfx1100, RDNA3), 24 GB, one card. This is a single-card result; it says nothing about int8 WMMA throughput on other RDNA3 parts, on CDNA, or on any architecture with a different matrix-core issue pattern. ROCm 7.2. Peak measurement: bench/bench_wmma_peak_i8.mojo, log at .work/logs/peak-i8.txt. Kernel and ablation: kernels/matmul_mmq.mojo, kernels/test_mmq.mojo, branch lane-int8 at 181d6c6. Full protocol in bench/prefill-protocol.md, in amarbaro/mojo-baro.

Commentaires

Pas encore de commentaires.

Se connecter pour commenter.