← Blog
✎post del blog

A Gate the Candidate Can Write

A close-up of an open brass pocket watch with exposed skeleton gears and jeweled bearings, a coiled watch chain and a spare gear resting beside it on a dark textured surface, the AMARBARO mark and wordmark centered over the watch face.

A performance gate whose timing state lives inside the file under test can be passed by a candidate that deletes the GPU synchronization call the timer depends on, and the only defense is moving the stopwatch, the host syncs, the file writes, and the fixture paths somewhere the candidate cannot reach.

This project runs a self-optimizing loop: a proposer edits a Mojo kernel, a gate script rebuilds it, checks its output against a reference, times it, and decides whether the edit survives. On 2026-09-08 an audit asked a narrow question about that gate, tools/loop-gate.sh as it stood at commit c0cfb6f: could a candidate that changes nothing meaningful about the kernel still pass every stage, including the performance stage, by editing something else instead. Four candidates were built specifically to try. One of them did, and it did it with a single deleted line.

The mechanism: the stopwatch was inside the room it was timing

The gate measured a kernel's speed by reading two timestamps out of the same file the candidate was allowed to edit, serve/engine.mojo. The relevant fragment, before this audit, ran a ctx.synchronize() to block the host until the GPU finished the kernel, then read the clock into var dt = .... That synchronization call is what makes the timestamp mean anything: without it, the host reads the clock the instant it issues the GPU work, not the instant the work finishes, so dt measures approximately nothing, and approximately nothing is a very small, very stable number.

Candidate 1 deleted that one line. Nothing about the kernel's actual computation changed. The candidate's identity check passed, because deleting a synchronization call doesn't change what values get written where, it only changes when the host thinks they were written. The candidate then ran three timed passes and printed 108.6, 108.9, and 110.8 tokens per second, against a real, honestly-measured champion of roughly 67. The gate scored this as a performance PASS at plus sixty-three percent, with a spread of 0.021, tight enough that the gate's own variance discipline had nothing to flag. The only thing that stopped this candidate from shipping was a separate, unrelated stage: a spill-count check that no build of the actual champion kernel could pass either, because the champion's own compiled sources already carried nine spilling instantiations at the commit the audit ran against. Every "zero survivors" result this loop had produced up to that point had been sitting below a bar nothing could clear, real candidate or fake one, and that fact wasn't visible until someone went looking for exactly this failure mode.

A second instance: a file a candidate can write is not a fixture

A second candidate in the same audit, cand-3, took a more direct route: it changed one output token to zero, then rewrote the reference file the identity check compares against, ref-tokens-64.txt, with its own new output. The identity check passed, 64 of 64, because it was comparing the candidate's output against a file the candidate itself had just overwritten. This candidate didn't make a speed claim, so it failed later on other grounds, but the identity result on its own was a clean, honest-looking PASS produced by a check that trusted a file the process under test could write to.

Two other candidates in the audit, one that shifted a timer variable by 500 million and one that deleted a completion flag entirely, both produced wildly inflated or garbage timing numbers (as high as 923,618 tokens per second) and were caught, but only by accident: the gate's spread threshold happened to reject them because a tiny decode_s denominator makes percentage noise swing wildly. The audit's own accounting is explicit about this: a candidate that had shifted the timestamp by 100 milliseconds instead of 500 would have produced a tidy, stable-looking plus-thirteen-percent and passed spread cleanly. The spread check that caught two candidates by luck is not a defense against the mechanism; it's a defense against a mechanism applied clumsily.

The rule

A candidate under test may never own anything the gate reads to decide whether it passed: not the clock, not the host synchronization that makes the clock meaningful, not the reference file, not the path the reference file lives at.

What counts as a real fix, and what doesn't

What counts, and what shipped the same day (e63dc77, e9e3109, ddda3c4, 75b7078, and the commits after them): splitting the engine into two files with two different trust levels. serve/window.mojo stays embedded and editable, the actual surface a candidate proposes changes to. serve/engine.mojo, which owns the stopwatch, the host synchronizations, and the harness logic, is taken fresh from git at a pinned commit on every gate run and is never embedded into anything a candidate's diff can touch. After the split, a benign structural change to window.mojo (a view-swap that changes nothing about the computed values) passes scope, compile, and identity on both fixtures and both timed runs at 130.4 against 130.2 tok/s, and is rejected only at the performance stage, which is the correct outcome for a genuinely neutral change. A direct attempt at the same class of attack after the fix, incrementing a position counter inside window.mojo, fails identity immediately at position 1, expected 271, got 0, because there's no longer a timing variable in that file worth attacking, and a correctness-breaking edit shows up as a correctness failure instead of a fabricated speed number.

What doesn't count as a fix, because the same audit found it insufficient on its own: tightening the spread threshold. Spread measures whether repeated runs of the same broken measurement agree with each other, and candidate 1's broken measurement was extremely internally consistent, 108.6 to 110.8 across three runs. A tighter spread bound would not have caught it; a stopwatch a candidate cannot reach is what caught it.

What it cost

The fix cost a real architectural split that the loop's design hadn't anticipated needing: harness logic that used to live next to the code being optimized now has to be pulled from git on every single gate invocation, rebuilt fresh, and kept byte-identical to a known-good commit rather than trusted as part of the working tree. That's slower per iteration and it's more moving parts to keep synchronized correctly.

It also didn't fully resolve the audit's other open finding. The spill-count stage that stopped candidate 1, purely by accident of timing, turned out to be a bar the champion's own real sources couldn't clear either, which means every prior "clean" result from this loop had been passing a gate that was, in a different sense, also broken: not gameable, but uncalibrated against its own baseline. Fixing that required a second rule change, making the spill-count comparison relative to a same-session build of the champion rather than an absolute threshold, and that change is recorded as Mario's call rather than something the audit could decide unilaterally, because it touches what "pass" is allowed to mean.

How to prove this wrong

Check out tools/loop-gate.sh and serve/engine.mojo at c0cfb6f, the commit immediately before this audit's fixes landed. Delete the ctx.synchronize() call that precedes the dt computation, leave everything else untouched, and run the gate's real ladder against a real champion build. If that run does not pass the performance stage at something close to plus sixty-three percent, the central example in this post is wrong and should be corrected. The receipts from the original four-candidate run live at .work/loop/audit-scorer/gate-run.log in the repository history; the per-run numbers in the table above came from a second run of the same four candidates against the hardened gate, all four dying at the scope stage instead, which is the comparison that shows the fix actually changed the outcome rather than just the commit log.

This audit ran on one RX 7900 XTX, against one engine's own harness design. It says nothing about whether a gate built around a different language's build system, a different timing primitive, or a different process boundary has the same vulnerability in the same shape; the general claim, that a candidate cannot be trusted with anything the gate reads to score it, is a much older idea than this project, and this post is one dated instance of it costing real GPU time to rediscover.

Provenance

AMD RX 7900 XTX (gfx1100, RDNA3), 24 GB, one card. ROCm 7.2. Mojo 1.0.0 via max[all]==26.5.0. Gate audited at commit c0cfb6f; hardening commits e63dc77, e9e3109, ddda3c4, 75b7078, 3a70336, 935d056, 0dbdc21, 3242573. Report exchange/scorer-integrity-report.md, receipts in .work/loop/audit-scorer/, in amarbaro/mojo-baro.

Commenti

Ancora nessun commento.

Accedi per commentare.