⌁project / evidence

mojo-baro

An LLM inference engine for AMD RDNA3 with its GPU kernels written in Mojo. This page shows what the current machine-generated receipt measures, and keeps the rest of the engine's limits visible.

A working engine, with the limits attached

mojo-baro serves an OpenAI-compatible API on AMD RDNA3. This page separates what the repository ships from what this receipt actually measures.

receipt
results/report-gfx1100-AMD-Radeon-RX-7900-XTX-61eec10.json
self-check
valid
sampled sizes
10
best ratio
1.38×

01 / the machine

A benchmark without its hardware is a rumour

This block is read from results/report-gfx1100-AMD-Radeon-RX-7900-XTX-61eec10.json, written by bench/report.sh at commit 61eec10 on a clean tree.

GPU
AMD Radeon RX 7900 XTX
Architecture
gfx1100 · 96 CUs · wave32
Max clock
2,371 MHz
ROCm
7.2.4
Vendor library
libhipblaslt.so.1.2
Compiler
Mojo 1.0.0 (ed45d567)
Clocks before
44.0 °C · 30 MHz · 15.0 W
Clocks after
56.0 °C · 29 MHz · 48.0 W

Receipt self-check: valid, no problems recorded.

02 / the measured kernel

Square fp16 GEMM against hipBLASLt

A pipelined WMMA kernel against the vendor library through the same C++ shim, with a 10.0-second clock warm-up and 200 timed iterations per point. It leads at 10 of 10 sampled sizes, with a best ratio of 1.38× at 1536³.

  1. 256³ 1.03×
  2. 512³ 1.22×
  3. 768³ 1.17×
  4. 1024³ 1.17×
  5. 1536³ 1.38×
  6. 2048³ 1.14×
  7. 2560³ 1.17×
  8. 3072³ 1.02×
  9. 3584³ 1.25×
  10. 4096³ 1.09×

mojo-baro hipBLASLt bars scaled to 106,095 GFLOP/s

Every column below is a field of the receipt: throughput, launcher geometry, vendor algorithm selection and the correctness error each kernel was gated on.
Tamaño mojo-baro hipBLASLt Proporción Azulejo · warps · cuadrícula Algoritmo del proveedor Error máximo
256³ 6,444 6,230 1.03× 64×64×32 · 2×2 · 4×4 #25/32 1.7e-06
512³ 31,807 26,107 1.22× 64×64×32 · 2×2 · 8×8 #21/32 1.4e-05
768³ 64,239 54,931 1.17× 64×64×32 · 2×2 · 12×12 #21/32 2.1e-05
1024³ 75,142 64,462 1.17× 64×128×32 · 2×4 · 8×16 #13/32 0.0e+00
1536³ 93,840 68,067 1.38× 128×128×32 · 4×2 · 12×12 #13/32 0.0e+00
2048³ 93,180 81,955 1.14× 128×128×32 · 4×2 · 16×16 #10/32 4.8e-07
2560³ 98,333 83,983 1.17× 128×128×32 · 4×2 · 20×20 #0/32 9.5e-07
3072³ 100,248 98,206 1.02× 128×128×32 · 4×2 · 24×24 #0/32 0.0e+00
3584³ 106,095 85,027 1.25× 128×128×32 · 4×2 · 28×28 #0/32 0.0e+00
4096³ 89,484 82,405 1.09× 128×128×32 · 4×2 · 32×32 #0/32 2.2e-04

GFLOP/s on this card, this ROCm, these sizes. Not a general claim about RDNA3 versus CDNA, and not shown to transfer to another card.

03 / engine results

The engine runs beyond the square kernel

The same repository carries a profile-driven dense path, matched-weight decode comparisons and speculative sampling. Each figure below stays beside the protocol that froze its question and its falsifier.

Dense-family targets verified against llama.cpp on the same GGUF, 20 prompts, teacher-forced agreement.
Model Arch Agreement tok/s_gen Status
Llama-3.2-1B-Instruct Q4_K_M llama 95.3-100% 451-458 servable
Qwen2.5-7B-Instruct Q4_K_M qwen2 96.9-100% 96.8-97.8 servable
granite-4.2-3b BF16 granite 98.4-100% 164.5-169.0 servable
lily-cybersecurity-7b Q6_K llama (Mistral-arch, SPM) 96.9-100% 94.1-95.0 servable

All four targets are servable through the profile-driven Spark path. Source: bench/dense-protocol.md.

decode arms

Same model, different weights, different verdicts

Same box, same GGUF where noted, 20-prompt medians unless the protocol says otherwise.

Model and weights llama.cpp mojo-baro Ratio
Qwythos-9B, Q4_0 110.0 130.7 1.19x
Qwythos-9B, q8 74.1 68.8 0.94x
Ornith-1.5-9B, Q4_K_M 88.8 80.8 0.91x

The Q4_0 side has since moved to 136.1-136.8 tok/s_gen under bench/attn-protocol.md. llama.cpp was not re-measured in that stint, so no newer ratio is claimed.

Receipt: bench/attn-latency-protocol.md

speculative decode

MTP clears the greedy baseline

median ratio
1.1042x
mojo-baro
150.96 tok/s_gen
greedy baseline
136.72 tok/s_gen

20-prompt median, q4 pack, k=2. Output stayed identical to plain greedy decode on every prompt tested. Receipt: bench/mtp-protocol.md.

device sampling

Sampling composes with speculation

temperature / top-p
0.7 / 0.9
with speculation
147.15 tok/s_gen
without speculation
109.19 tok/s_gen
greedy, no speculation
134.97 tok/s_gen

One dense q4 stint. The sampler runs inside the decode loop. Receipt: bench/spec-sample-protocol.md.

MoE path

RegesCore-35B decodes and serves

Mean teacher-forced agreement: 53.20/64 over 20 prompts. Decode: 111.89 tok/s_gen against llama.cpp at 109.92 on the same GGUF.

The two decode numbers come from different stints. That is stated here because the comparison is useful, but it is not a same-stint ratio. Receipt: bench/moe-persist-protocol.md.

04 / scope and next

What exists, what is measured, what comes next

ships now

Inference engine on AMD RDNA3

Mojo kernels sit under an OpenAI-compatible API. The repository also carries dense and MoE engine paths, sampling, tool calls, checkpoints and state transfer.

measured here

Square fp16 GEMM

This page publishes the receipt-backed WMMA comparison only. Hardware, ROCm, tile geometry, vendor algorithm choice and correctness error stay attached to each row.

measured engine

Dense, q4 and MoE paths

Four profile-driven dense targets are servable. Q4 decode and MoE results stay paired with their protocol caveats, including losing arms and different stints.

next design round

Batching

One request decodes at a time today. Batching is the next design round, not a shipped capability.