A forced-agreement floor is only meaningful once it has been measured against a same-arm control, because a floor that sits between two achievable scores is unreachable by any arm, patched or not.
On 2026-09-17 a LoRA write-back on mlp.down_proj of the Qwythos champion's top eight layers was run through the project's forced-agreement gate against llama.cpp: 20 prompts, generation length 64, a frozen floor of 90% minimum-over-20. The patched model scored 96.89% aggregate and, on its single worst prompt, p09-explain-gpu, 89.06%. Under the floor. The write-back looked broken by 0.94 points.
It was not broken. The floor was unreachable by anything, patched or not, and the gate had never been asked whether that was true.
The fix was to run the same gate, unchanged, against the unpatched base model, which had never been measured against this particular floor before. Every earlier receipt for that prompt was ours-versus-ours spec identity, not the base against its own llama.cpp reference.
The unpatched control also failed. 89.06%, on a different prompt (p03-story instead of p09-explain-gpu), aggregate 97.31%.
| control (unpatched) | patched | |
|---|---|---|
| minimum | p03-story 89.06% | p09-explain-gpu 89.06% |
| aggregate | 1230/1264 = 97.31% | 1216/1255 = 96.89% |
Both minima land on exactly the same number: 57 correct out of 64, which is 89.06%. That is not a coincidence, and it is not noise. It is the shape of the scoring grid. At 64 generated tokens the achievable per-prompt scores step from 57/64 (89.06%) straight to 58/64 (90.63%). There is no value between them. A floor set at 90% sits in a gap nothing can land in. Any prompt that comes in one token short of perfect will read 89.06%, and no amount of the model being genuinely correct raises that specific prompt's score past the gap, because there is nothing to raise it to short of 90.63%.
The gate is deterministic within an arm, so this was not run once and left as a guess. The control was run twice, on separate servers and ports, and reproduced all 20 prompts exactly, 1230/1264 both times, zero differing prompts. The prompt that sits worst is not stable across arms, though: p03 is the worst prompt unpatched (89.06%) and the third-best patched (96.88%); p09 goes the other way. Min-over-20 reports whichever single prompt happens to land worst in that specific arm, and which prompt that is moves around while the aggregate barely does, -0.42 percentage points between control and patch.
The floor was amended at commit c24dd32 ("amend P5b's forced-agreement bar to aggregate-vs-control"), and the merge at 41f60f8 records the LoRA write-back as passing on the amended aggregate-versus-control bar. The rule this leaves behind: if you re-run this or any similar forced-agreement gate and the unpatched base has never been measured against the same floor, do not trust a PASS or a FAIL from the patched arm alone.
This is not the first time a bar in this project turned out to be a property of the scoring instrument rather than of the thing being scored. bench/PROTOCOL-RULES.md P14 exists because of an earlier, larger instance: a gate required 20 prompts at 64/64 teacher-forced agreement against llama.cpp, the project's MoE path reached 53.20 against that gate and was treated as failing for a whole session, and the dense path that actually ships, verified, on Qwythos reached 51.90 against the identical reference on a quant-matched arm the same day. The MoE path was already above the ceiling of the known-good path while being called broken. P14's statement of the rule: before a pass threshold is frozen, measure what the mature path on the same hardware actually reaches against the same reference on a quant-matched arm.
The same document records why greedy 64-token equality is never trusted past roughly 256 generated ids in this project at all: llama.cpp's own f16-KV configuration fails its own f32 reference at 5 of 7 tested lengths. A reference implementation checked against itself, in a different but supported configuration, still misses. That is the reason the standing identity check here is teacher-forced agreement (BARO_FORCE), which re-anchors every position to the reference regardless of what came before, rather than free-running greedy equality, which lets one early divergence compound forward through the rest of the generation.
A floor is not a fact about the model. Measure it against a same-arm control before you trust what it reports about anything else.
What counts as having validated a floor, in this project:
n_predict=64 the grid is in sixty-fourths; a floor placed between two adjacent fractions is unreachable by anything.What does not count:
A whole repair round was nearly spent on the LoRA write-back before the control was run: retraining, or adjusting the patched layers, against a bar that the unpatched model could not clear either. The coordinator's kill/repair call on the write-back was left open specifically because that control had not yet been run, which is the right response to an unclear signal, but it is also lost time that a same-arm control run first would have avoided entirely.
What it bought is the amended rule itself, and a repeatable instance of the P14 principle at a different generation length and on a different kind of gate: a forced-agreement floor and a teacher-forced equality gate both turned out to be describing the instrument, not the arm, once someone ran the reference through its own test.
Run bench/p1-bridge-protocol.md's or the P5b forced-agreement harness on your own hardware at n_predict=64 against an unpatched model of your choosing, same 90% min-over-20 style floor. If the per-prompt minimum clears 90% with no prompt landing on exactly 57/64, the grid-gap argument here does not hold for your setup, and the specific claim that the 90% floor is unreachable by construction at this generation length should be treated as local to this repo's prompt set and model family rather than a general property of forced agreement at 64 tokens.
Separately, the llama.cpp f16-KV-versus-f32 comparison at 5 of 7 lengths is a receipt anyone running llama.cpp with both configurations can attempt to reproduce; a report that reaches 7 of 7 agreement would contradict the standing reason this project avoids greedy equality gates past 256 ids and should be taken seriously.
AMD RX 7900 XTX (gfx1100, RDNA3), 24 GB, one card. The forced-agreement figures above are per-prompt and aggregate scores from a single gate run; nothing here compares throughput, and none of it says anything about how this bar behaves on a different card or a different generation length. Control receipts: .work/team-B/sonnet/p5b/control/forced/results.txt; patched receipts: .work/team-B/sonnet/p5b/forced/results.txt, both in the lane-team-b worktree. Floor amendment c24dd32; merge recording the amended result 41f60f8; P14 in bench/PROTOCOL-RULES.md in amarbaro/mojo-baro.
Comentarii
Niciun comentariu încă.
Conectare pentru a comenta.