A Mojo-ported arm64 kernel set can match llama.cpp bit for bit on two real phones, and matching on one phone is not evidence about the other: the OnePlus fails a steadiness gate the Xiaomi passes clean.
The claim in this repo's README is narrow and it is worth reading narrowly: a Mojo inference engine produces the same greedy tokens as llama.cpp, bit for bit, on two Android phones. Not "Mojo runs on Android." Not "the app runs a local model." The kernels, ported by hand from a KleidiAI binding into Mojo, agree with the reference implementation down to the byte, on a Xiaomi and a OnePlus, for Qwen2.5-0.5B at Q4_0. That is the whole of what is proven, and it took three layers of gates to get there, one of which still fails.
The chain runs from vendor kernel to ported kernel to full engine. First, the Q4_0 matmul
through a KleidiAI binding is checked byte-identical to ggml's own MUL_MAT on both phones.
Second, the same kernel ported into Mojo is checked byte-identical to that binding. Third, the
full engine, wired together from the ported kernels, produces greedy tokens bit-exact against
llama.cpp on both phones for the 0.5B model. Each of the three steps is a gate that could have
failed independently, and none of them tells you the next one will pass: a kernel can be
byte-exact in isolation and an engine built from correct kernels can still diverge, because
engines add rounding order, accumulation order, and libm calls that a single matmul test never
exercises.
None of this proves the engine is fast, or that it generalizes past Qwen2.5-0.5B, or that it works inside the app. It proves one thing precisely: for this model, on these two devices, the arithmetic path taken by a from-scratch Mojo port produces the same numbers as llama.cpp's C++ kernels, to the bit. That is a correctness claim, not a performance claim, and the two are kept separate on purpose because a project that reports "faster and correct" in one sentence tends to have measured only one of the two.
IQ1_S is byte-identical to ggml's own implementation of the same format, 2.1x to 3.6x faster
per matmul than the reference path on these phones. The name suggests 1-bit weights; the format
IQ1_S is documented elsewhere as averaging 1.56 bits per weight. Measured on the actual packed
tensors used in this repo's gate, the real figure is 2.31 bits per weight, not 1.56.
The gap is not a bug in the kernel and not a lie in the format name. It is the difference between a format's theoretical average across an idealized weight distribution and what a real model's weight distribution actually packs into once codebooks, block scales, and the super-block metadata are counted. A quantization format's headline bit count describes a distribution, and a real checkpoint is one sample from a much messier distribution than the one the format was tuned against. Anyone reporting an "N bits per weight" figure without measuring the packed tensor size of the specific model in hand is reporting the format's marketing number, not its cost on their model.
The correctness side of this kernel is solid: byte-identical to ggml, gated the same way as Q4_0. The bits-per-weight side is a second, independent measurement, and the two do not imply each other. A kernel can be perfectly correct and still cost more memory than its name promises.
Thread scaling was gated on both phones with a frozen rule set before the runs: pass if every run of twenty stays above 0.7x of that run's own median. The Xiaomi passed clean: prefill 1.142x and decode 1.066x of llama.cpp, zero stalls over the full run. The OnePlus is faster on every comparison that reports a median: prefill 2.16x, decode 1.58x of llama.cpp. It also failed its own gate, because three of the twenty runs landed below 0.7x of the median, with spread as wide as 59 percent on one measurement and 32 percent on the other.
This is the same failure mode as a synthetic benchmark that looks great on average and hides its own worst case: a median is exactly the statistic that a small number of bad runs cannot move much, which is why the gate does not check the median, it checks every run against it. A phone that is faster on average and unsteady is not simply "the same, plus noise." Whatever mechanism drops the OnePlus below 0.7x on three of twenty runs is unresolved in this repo as of this writing; the megaplan names core parking between parallel regions as the suspect and hands the question to a separate lane, not yet closed.
A phone that beats the reference on every median and fails its own steadiness gate on three runs of twenty has not passed threading; it has a threading result and an open bug, and the two get reported together or not at all.
The OnePlus is the project's stated constraint phone: 2.5 GB of free RAM against the Xiaomi's larger headroom. The Xiaomi's clean pass says the engine's threading model can run steady on this class of hardware. It says nothing about the OnePlus, and the OnePlus's own result says so directly: the same code, the same gate, one phone clean and one phone not. A single-device result, run once, generalizes to nothing beyond that device; two devices generalize only to the extent the two devices agree, and here they do not.
The engine's own on-device tab in the app is a separate case worth naming for contrast, because it fails a different gate for a different reason. It runs llama.cpp directly, not this project's kernels, and reproduces the desktop reference exactly on only 1 of 5 prompts after one repair round. The report attributes the divergence to near-ties on a flat probability distribution rather than to a bug in the app's llama.cpp binding, which is an explanation for why the divergence happens, not a pass. It stays a fail in the capability table either way: an explanation does not move a number across a gate.
What counts as a bit-exact claim in this repo:
- Greedy tokens identical, prompt by prompt, against a desktop llama.cpp built from the exact
pinned commit, ca3d5a3e1, on the same phone.
- A kernel checked byte-identical against ggml's own kernel on real device hardware, not a
simulator.
- A steadiness claim checked against a frozen rule (here, 0.7x of the run's own median) stated
before the runs, not chosen afterward to fit whatever the runs produced.
What does not count: - A single fast run. The OnePlus's 2.16x and 1.58x medians are real, and they are not the whole result; the three slow runs are part of the same result. - A format's published bits-per-weight figure standing in for a measurement of the actual packed tensor. 1.56 is IQ1_S's average; 2.31 is what this model's weights actually cost, packed. Both numbers are true, of different things. - One phone's pass as evidence about a second phone. The Xiaomi and the OnePlus disagree on steadiness under identical code, which is the entire reason two phones were used instead of one.
Running every gate on two physical devices over adb, with a desktop llama.cpp built from a
pinned commit as the reference for both, is slower than running once. It is the only arrangement
that could have caught the OnePlus's failure at all: a project that gated on one phone would have
shipped the number 1.58x and never seen the 32 percent spread sitting underneath it. The kaipack
pre-packed loader is the other place this discipline forced an accepted failure rather than a
convenient pass: it hits bit-exact tokens 15 of 15 on both phones and clears peak RSS at 0.50x of
the GGUF path, but its first-token latency lands at 0.53x on the OnePlus and 0.65x on the Xiaomi
against a 0.5x bar set before the run. The feature ships anyway, because it frees the OnePlus's
anonymous memory from 276 MB down to 5 MB, which the constraint phone needs regardless of the
missed latency bar. The bar stays where it was frozen. The feature stays in on its own merits,
labeled with the bar it did not clear.
The steadiness claim is the one worth attacking, because it is the one this repo has not closed.
Build the engine for a OnePlus 12 (or the closest available OnePlus with comparable free RAM),
run mojo-engine's gate 3 threading check twenty times against the pinned llama.cpp commit, and
look at every run against that run's own median, not just the median across the set. Twenty runs
above 0.7x contradicts this post's second half outright. A different distribution of failures,
say all twenty low but consistently low, would say the mechanism is not the intermittent one this
repo currently suspects, and that is worth reporting too.
The IQ1_S bits-per-weight claim is narrower and easier to check independently: pack the same model into IQ1_S with any standard tool, measure the resulting file's tensor bytes divided by parameter count, and compare against 2.31. A different model, a different quantization calibration, or a different codebook could reasonably land somewhere else; the point is that 1.56 is not that number for any specific model without measuring it.
Xiaomi and OnePlus, Android arm64, gates run over adb against a desktop reference. llama.cpp
fetched at the pinned commit ca3d5a3e1, never vendored; a comparison against any other commit
of llama.cpp is not this comparison. Mojo 1.0 toolchain, cross-compiled for arm64-v8a. Every
number above is cited in amarbaro/baro.apk's README and
its exchange/ lane reports (lane-M2-report.md, lane-M4-report.md,
lane-MOJOQ4-report.md), against gates at commit 34e827c for the app tabs and the megaplan
document for the threading and IQ1_S receipts.
Comentarios
Aún no hay comentarios.
Iniciar sesión para comentar.