← Blog
✎post del blog

An Arm64 Engine, Bit-Exact on Two Phones, and One Steady on One

Two phones lying face up on a dark stone-colored surface, connected by teal charging cables, a laptop visible out of focus in the background, with the AMARBARO seed-root-boot mark centered over the scene.

A Mojo-ported arm64 kernel set can match llama.cpp bit for bit on two real phones, and matching on one phone is not evidence about the other: the OnePlus fails a steadiness gate the Xiaomi passes clean.

The claim in this repo's README is narrow and it is worth reading narrowly: a Mojo inference engine produces the same greedy tokens as llama.cpp, bit for bit, on two Android phones. Not "Mojo runs on Android." Not "the app runs a local model." The kernels, ported by hand from a KleidiAI binding into Mojo, agree with the reference implementation down to the byte, on a Xiaomi and a OnePlus, for Qwen2.5-0.5B at Q4_0. That is the whole of what is proven, and it took three layers of gates to get there, one of which still fails.

What "bit-exact" actually checked

The chain runs from vendor kernel to ported kernel to full engine. First, the Q4_0 matmul through a KleidiAI binding is checked byte-identical to ggml's own MUL_MAT on both phones. Second, the same kernel ported into Mojo is checked byte-identical to that binding. Third, the full engine, wired together from the ported kernels, produces greedy tokens bit-exact against llama.cpp on both phones for the 0.5B model. Each of the three steps is a gate that could have failed independently, and none of them tells you the next one will pass: a kernel can be byte-exact in isolation and an engine built from correct kernels can still diverge, because engines add rounding order, accumulation order, and libm calls that a single matmul test never exercises.

None of this proves the engine is fast, or that it generalizes past Qwen2.5-0.5B, or that it works inside the app. It proves one thing precisely: for this model, on these two devices, the arithmetic path taken by a from-scratch Mojo port produces the same numbers as llama.cpp's C++ kernels, to the bit. That is a correctness claim, not a performance claim, and the two are kept separate on purpose because a project that reports "faster and correct" in one sentence tends to have measured only one of the two.

The 1-bit kernel is correct and heavier than its own name

IQ1_S is byte-identical to ggml's own implementation of the same format, 2.1x to 3.6x faster per matmul than the reference path on these phones. The name suggests 1-bit weights; the format IQ1_S is documented elsewhere as averaging 1.56 bits per weight. Measured on the actual packed tensors used in this repo's gate, the real figure is 2.31 bits per weight, not 1.56.

The gap is not a bug in the kernel and not a lie in the format name. It is the difference between a format's theoretical average across an idealized weight distribution and what a real model's weight distribution actually packs into once codebooks, block scales, and the super-block metadata are counted. A quantization format's headline bit count describes a distribution, and a real checkpoint is one sample from a much messier distribution than the one the format was tuned against. Anyone reporting an "N bits per weight" figure without measuring the packed tensor size of the specific model in hand is reporting the format's marketing number, not its cost on their model.

The correctness side of this kernel is solid: byte-identical to ggml, gated the same way as Q4_0. The bits-per-weight side is a second, independent measurement, and the two do not imply each other. A kernel can be perfectly correct and still cost more memory than its name promises.

The Xiaomi is steady; the OnePlus is faster and is not

Thread scaling was gated on both phones with a frozen rule set before the runs: pass if every run of twenty stays above 0.7x of that run's own median. The Xiaomi passed clean: prefill 1.142x and decode 1.066x of llama.cpp, zero stalls over the full run. The OnePlus is faster on every comparison that reports a median: prefill 2.16x, decode 1.58x of llama.cpp. It also failed its own gate, because three of the twenty runs landed below 0.7x of the median, with spread as wide as 59 percent on one measurement and 32 percent on the other.

This is the same failure mode as a synthetic benchmark that looks great on average and hides its own worst case: a median is exactly the statistic that a small number of bad runs cannot move much, which is why the gate does not check the median, it checks every run against it. A phone that is faster on average and unsteady is not simply "the same, plus noise." Whatever mechanism drops the OnePlus below 0.7x on three of twenty runs is unresolved in this repo as of this writing; the megaplan names core parking between parallel regions as the suspect and hands the question to a separate lane, not yet closed.

A phone that beats the reference on every median and fails its own steadiness gate on three runs of twenty has not passed threading; it has a threading result and an open bug, and the two get reported together or not at all.

What one phone's pass proves about the other's

The OnePlus is the project's stated constraint phone: 2.5 GB of free RAM against the Xiaomi's larger headroom. The Xiaomi's clean pass says the engine's threading model can run steady on this class of hardware. It says nothing about the OnePlus, and the OnePlus's own result says so directly: the same code, the same gate, one phone clean and one phone not. A single-device result, run once, generalizes to nothing beyond that device; two devices generalize only to the extent the two devices agree, and here they do not.

The engine's own on-device tab in the app is a separate case worth naming for contrast, because it fails a different gate for a different reason. It runs llama.cpp directly, not this project's kernels, and reproduces the desktop reference exactly on only 1 of 5 prompts after one repair round. The report attributes the divergence to near-ties on a flat probability distribution rather than to a bug in the app's llama.cpp binding, which is an explanation for why the divergence happens, not a pass. It stays a fail in the capability table either way: an explanation does not move a number across a gate.

What counts here and what does not

What counts as a bit-exact claim in this repo: - Greedy tokens identical, prompt by prompt, against a desktop llama.cpp built from the exact pinned commit, ca3d5a3e1, on the same phone. - A kernel checked byte-identical against ggml's own kernel on real device hardware, not a simulator. - A steadiness claim checked against a frozen rule (here, 0.7x of the run's own median) stated before the runs, not chosen afterward to fit whatever the runs produced.

What does not count: - A single fast run. The OnePlus's 2.16x and 1.58x medians are real, and they are not the whole result; the three slow runs are part of the same result. - A format's published bits-per-weight figure standing in for a measurement of the actual packed tensor. 1.56 is IQ1_S's average; 2.31 is what this model's weights actually cost, packed. Both numbers are true, of different things. - One phone's pass as evidence about a second phone. The Xiaomi and the OnePlus disagree on steadiness under identical code, which is the entire reason two phones were used instead of one.

What this costs

Running every gate on two physical devices over adb, with a desktop llama.cpp built from a pinned commit as the reference for both, is slower than running once. It is the only arrangement that could have caught the OnePlus's failure at all: a project that gated on one phone would have shipped the number 1.58x and never seen the 32 percent spread sitting underneath it. The kaipack pre-packed loader is the other place this discipline forced an accepted failure rather than a convenient pass: it hits bit-exact tokens 15 of 15 on both phones and clears peak RSS at 0.50x of the GGUF path, but its first-token latency lands at 0.53x on the OnePlus and 0.65x on the Xiaomi against a 0.5x bar set before the run. The feature ships anyway, because it frees the OnePlus's anonymous memory from 276 MB down to 5 MB, which the constraint phone needs regardless of the missed latency bar. The bar stays where it was frozen. The feature stays in on its own merits, labeled with the bar it did not clear.

How to prove this wrong

The steadiness claim is the one worth attacking, because it is the one this repo has not closed. Build the engine for a OnePlus 12 (or the closest available OnePlus with comparable free RAM), run mojo-engine's gate 3 threading check twenty times against the pinned llama.cpp commit, and look at every run against that run's own median, not just the median across the set. Twenty runs above 0.7x contradicts this post's second half outright. A different distribution of failures, say all twenty low but consistently low, would say the mechanism is not the intermittent one this repo currently suspects, and that is worth reporting too.

The IQ1_S bits-per-weight claim is narrower and easier to check independently: pack the same model into IQ1_S with any standard tool, measure the resulting file's tensor bytes divided by parameter count, and compare against 2.31. A different model, a different quantization calibration, or a different codebook could reasonably land somewhere else; the point is that 1.56 is not that number for any specific model without measuring it.

Provenance

Xiaomi and OnePlus, Android arm64, gates run over adb against a desktop reference. llama.cpp fetched at the pinned commit ca3d5a3e1, never vendored; a comparison against any other commit of llama.cpp is not this comparison. Mojo 1.0 toolchain, cross-compiled for arm64-v8a. Every number above is cited in amarbaro/baro.apk's README and its exchange/ lane reports (lane-M2-report.md, lane-M4-report.md, lane-MOJOQ4-report.md), against gates at commit 34e827c for the app tabs and the megaplan document for the threading and IQ1_S receipts.

Commenti

Ancora nessun commento.

Accedi per commentare.