An identity gate built on generated token ids cannot certify that a transferred model state is correct, because a deliberately corrupted state can still produce the right ids on a share of prompts.
The falsifier was supposed to be blunt: take a saved conversation state, swap its K and V tensors on purpose, hand the corrupted state to a node, let it continue generating, and check whether the output tokens are wrong. If they are not wrong, the gate cannot tell a good state from a bad one, and the whole exercise of moving state between two engines cannot be trusted on ids alone.
Five states were swapped this way during the lane's gate-2 identity run on 2026-09-17. Three produced wrong token ids, as expected. Two did not. Both the f32 arm and the int8 arm reproduced this exactly: 5 of 5 swapped states were rejected as "flip refused" by the transport layer's own consistency check, but only 3 of 5 actually diverged in their generated ids once accepted and continued.
| control S | 100 Mbit | 1 Gbit | 10 Gbit | rate-dependent prompts | flip refused | swapkv restored / wrong ids | |
|---|---|---|---|---|---|---|---|
| f32 | 20 | 20 | 20 | 20 | 0 | 5 of 5 | 5 of 5 / 3 of 5 |
| int8 (the plan's format) | 20 | 16 | 16 | 16 | 0 | 5 of 5 | 5 of 5 / 3 of 5 |
The frozen falsifier for this gate required at least 4 of 5 K/V-swapped states to give wrong ids. Two strongly determined short prompts gave the right 32 ids anyway. The verdict, as the frozen scorer printed it, was FAIL: the harness had not proven itself. No threshold was moved after seeing that number, and the f32 row above is not reported as a pass despite every rate reaching 20/20 on unswapped states.
The mechanism is not exotic. On a short, strongly determined prompt, the model's next-token distribution can be so confidently peaked that a corrupted K/V still produces the same argmax at several positions in a row. A wrong internal state and a right output token are not the same fact, and an ids-only gate conflates them whenever the prompt happens to be easy enough.
Since the ids could not certify a state, the state was compared byte for byte instead (bench/fork-bytes-check.sh): one node re-exports exactly what it imported, and a second arm sends a deliberately K/V-swapped state so that only a node which actually kept the imported bytes, rather than silently regenerating something plausible, can produce the expected re-export.
Its first run found a real, pre-existing bug in serve/engine.mojo: save_state chose which checkpoint to export by position alone. Two different prompts that both happened to be at pos 14, p05-math and p02-python-fib, collided on that key, and p05-math's export silently carried p02-python-fib's conversation and SSM state, 52 of the 61 MB exported, under a valid sha and passing every identity check that existed at the time. One conversation's state leaking into another's export, and any receiver that trusted the sha would have accepted it without complaint.
This was fixed in 511cdb4. The fix was reproduced first at 6 of 7, then verified at 7 of 7 bit-identical exports, including deliberately swapped states, with ci-checks exiting 0 and cargo nextest passing 75 of 75.
The transport itself checked out cleanly on every measure that did not depend on generated tokens: node B's sha256 matched on every import, a corrupted state came back through the 409 path relayed by the sending node, and node B kept and placed every imported byte of conversation, SSM, and K/V state. What the ids alone established, and established weakly, was narrower: that 32 tokens continue the same as a reference on most prompts. Control S, an unmoved engine restoring its own state, reached 20 of 20 on the same prompts, so this engine's restored run does reproduce its own cold run where llama.cpp's equivalent restore, measured in a related gate the same week, does not.
The same class of gap showed up one gate earlier in the same lane. Gate 4 asked whether llama.cpp, handed this project's exported state, would continue it the same way llama.cpp continues its own restored bytes. K and V were swapped in the written slot file for that comparison too: n_restored and cache_n still came back correct on 20 of 20 swapped files, and llama.cpp accepted and reused every one of them, with generated ids wrong at index 0 on all 20. The reuse check, which only confirms a state was not silently recomputed, passed every time on data that was wrong from the first token. Only the ids caught that particular failure, which is the mirror image of this post's finding: reuse checks and id checks each catch a different kind of corruption, and neither one alone is a gate.
A wrong state and a wrong output are different facts. A gate needs a check that can see each one, because either check alone will pass some kinds of corruption the other one catches.
What counts as a load-bearing correctness check, based on what actually caught something here:
n_restored, cache_n against the prompt length) that fires independently of whether the continuation is correct. This is what caught llama.cpp silently reusing swapped K/V without recomputing, on data an ids-only check missed until generation actually diverged.What does not count as having verified a state transfer:
The checkpoint collision had been in serve/engine.mojo for long enough to pass every identity check the earlier gates ran, because those checks were built on ids and on shas, neither of which this bug touched until the byte-level check specifically went looking for cross-prompt contamination. Writing and running bench/fork-bytes-check.sh cost a dedicated pass that produced no useful signal on 3 of its 5 swapped states, by design, since the point of running it was to find what the ids-based gate could not, not to replace that gate.
Take any engine that saves and restores conversation state, generate five conversations of varying length, save each state, swap K and V inside the saved file, restore, and generate 32 tokens from each. If fewer than 2 of 5 swapped states produce the same ids as an unswapped continuation, this project's specific 2-of-5 finding does not reproduce on that engine, though the general argument that ids cannot certify a state would still hold as long as at least one swapped state slips through. The stronger claim, that a byte-level re-export check is required in addition to any ids check, is falsified if a pure ids gate ever catches every state corruption a byte check catches on the same test set; that has not been observed here.
AMD RX 7900 XTX (gfx1100, RDNA3), 24 GB, one card, for the f32 and int8 rows above; the byte check itself is CPU-side comparison of saved state files and does not depend on the GPU. serve/engine.mojo fix at 511cdb4; full gate results and receipts under .work/fork/g2-gate-fixed and .work/fork/bytes-check-fixed, protocol bench/p1-fork-protocol.md, in amarbaro/mojo-baro.
Commenti
Ancora nessun commento.
Accedi per commentare.