An inference stack for the RX 7900 XTX with its GPU kernels written in Mojo, benchmarked against hipBLASLt under a frozen protocol. Every figure on this page is read from a machine-generated receipt.
A benchmark without its hardware is a rumour. This block is read from results/report-gfx1100-…-61eec10.json, written by bench/report.sh at commit 61eec10 on a clean tree.
Receipt self-check: valid — no problems recorded.
A pipelined WMMA kernel against the vendor library through the same C++ shim, with a 10-second clock warm-up before every measurement and 200 timed iterations per point. Ahead at 10 of 10 sizes measured here, best 1.38x at 1536³.
mojo-baro hipBLASLtbars scaled to 106,095 GFLOP/s
| Size | mojo-baro | hipBLASLt | Ratio | Tile · warps · grid | Vendor algo | Max error |
|---|---|---|---|---|---|---|
| 256³ | 6,444 | 6,230 | 1.03x | 64×64×32 · 2×2 · 4×4 | #25/32 | 1.7e-6 |
| 512³ | 31,807 | 26,107 | 1.22x | 64×64×32 · 2×2 · 8×8 | #21/32 | 1.4e-5 |
| 768³ | 64,239 | 54,931 | 1.17x | 64×64×32 · 2×2 · 12×12 | #21/32 | 2.1e-5 |
| 1024³ | 75,142 | 64,462 | 1.17x | 64×128×32 · 2×4 · 8×16 | #13/32 | 0.0e+0 |
| 1536³ | 93,840 | 68,067 | 1.38x | 128×128×32 · 4×2 · 12×12 | #13/32 | 0.0e+0 |
| 2048³ | 93,180 | 81,955 | 1.14x | 128×128×32 · 4×2 · 16×16 | #10/32 | 4.8e-7 |
| 2560³ | 98,333 | 83,983 | 1.17x | 128×128×32 · 4×2 · 20×20 | #0/32 | 9.5e-7 |
| 3072³ | 100,248 | 98,206 | 1.02x | 128×128×32 · 4×2 · 24×24 | #0/32 | 0.0e+0 |
| 3584³ | 106,095 | 85,027 | 1.25x | 128×128×32 · 4×2 · 28×28 | #0/32 | 0.0e+0 |
| 4096³ | 89,484 | 82,405 | 1.09x | 128×128×32 · 4×2 · 32×32 | #0/32 | 2.2e-4 |
GFLOP/s on this card, this ROCm, these sizes. Not a general claim about RDNA3 versus CDNA, and nothing here is shown to transfer to another card.
The other kernel result — single-token decode, where the weights stream out of HBM — took four rounds, and three of them falsified their own prediction. Each round was frozen in bench/coldcache-protocol.md before it ran. M=1, K=4096, N=12288 skinny GEMM. 8 rotating device buffers (working set >> the 96 MB Infinity Cache, so every launch streams from HBM), 1 s clock warm, 200 timed launches, whole measurement repeated 10x in-process.
A speedup is claimed only if the two ranges over the 10 repeats do not overlap and the stability prediction held.
The weight layout is coalescing-bound at ~250 GB/s (26% of peak), not byte-bound — so halving the bytes only added dequant ALU work and q8 came out slower. A post-hoc arm that transposes the weights at load (B-layout) ran 194–195 µs; disclosed as not preregistered.
q8b is faster than bf16 but far less than modeled: dequant ALU plus the per-32 scale reload eats most of the byte win, leaving 344 GB/s on its own byte stream. The vendor reaches ~812 GB/s, 85% of peak. Beating it cold needs bandwidth efficiency, not fewer bytes. No beat-vendor claim made.
CPT=8 was predicted to regress and was instead the best arm. 723 GB/s, 75% of peak — but the vendor is still 1.10x ahead at this point.
12.3% over v2, and the m1 range sits strictly below the vendor's with no overlap. Margin ~1%. The mechanism is occupancy relief, not fewer bytes: dropping the 8-row LDS staging shrinks the accumulator to 64 lanes, so more waves stay resident. Byte traffic is identical between the two. One 817 µs vendor outlier (repeat 8) was excluded and disclosed.
The long form on each result, one post per kernel.
Square GEMM is the benchmark everyone runs; decode is M=1 streaming weights out of HBM. Four frozen rounds, three wrong predictions, and a one-percent win that came from occupancy relief rather than fewer bytes.
ReadA pipelined fp16 WMMA GEMM beats hipBLASLt at ten square sizes on RDNA3 — but only once tile shape stops being fixed. Plus a retracted 2x claim and why clock warm-up was worth more than any kernel change.
Read