← Blog
✎article de blog

The GPU Is a Queue

A row of round teal and amber lamps mounted on a dark ribbed metal wall panel, receding toward a vanishing point, the AMARBARO seed-root-boot mark centered in the foreground.

Almost all GPU job failures on a single shared card are discoverable without ever touching the GPU, so a CPU preflight before every GPU launch is not caution, it is the majority of the fix.

One RX 7900 XTX drives this desktop's compositor, its terminal emulator, and every AI workload run on this machine. It is one card, and every one of those consumers has a claim on it at any given moment. On 2026-09-16, a single day's count of GPU jobs across ongoing kernel and model work came to 135, of which 34 failed. The number that matters is not 34 of 135. It is that every one of the 34 was discoverable on the CPU, before the GPU was ever asked to do anything.

What "discoverable on the CPU" actually means

A GPU job fails for one of roughly two reasons: something about the job itself was wrong before it ever reached the card (a build error, a flag that silently did nothing, a config file with a typo, an engine given empty stdin), or something about the card's actual state at the moment of launch was wrong (out of memory, another job already holding the VRAM, a hold in effect that the queue could not see). The first category never needed a GPU cycle to catch. A binary that fails to build fails identically whether or not a GPU is attached. A flag that is accepted and silently ignored, this project's own recurring failure mode, produces the same silent no-op on a CPU dry run as it does mid-job on the card. An engine handed empty stdin prints ready and exits 0 whether or not anything is listening on the other end. None of these needed to burn a queue slot, hold VRAM, or wait behind another job to be caught.

The second category, actual GPU state, is real and does need the card, but it is a much smaller slice than a first glance suggests, because most of what looks like a resource conflict is actually a bookkeeping failure that a CPU-side check would have caught: a hold recorded only on a project whiteboard, invisible to the queue itself, or a script whose cleanup path never actually reaped the process it launched, so the job shows running for nine minutes after its own log already says it failed.

A second instance: the queue has no state for the wrong kind of hold

On 2026-09-11, a session wrote "GPU HELD for Mario since 17:41" into a project whiteboard's resume card. gpu-wait's own subcommand set is submit run status list logs cancel wait version gpu stats; there is no hold subcommand, so the daemon's queue was empty and admitted the next job without complaint. That job ran for roughly 36 minutes against a hold that existed only as prose in a markdown file the queue never reads. The session that launched it had read the whiteboard's head hours earlier, at session start, and never again before the GPU launch.

The mechanism this exposes is not "the tooling is incomplete," though it is; it is that a hold recorded in a place the queue cannot see is not a hold, it is a note to a human who has to remember to re-read it before every single launch, not once per session. gpu-wait list returning empty is not permission when a hold note exists elsewhere; treating an empty queue as sufficient evidence of a free card is exactly the CPU-side check that was skipped.

A third and smaller instance from the same week: gpu-wait run does not forward the caller's shell environment or stdin into the job it launches internally. env VAR=val placed before gpu-wait run sets that variable on the wrapper process, not on the job; a knob meant to cap a sweep at 50 items instead read its variable as unset, defaulted to unlimited, and ran a full 4,799-item sweep silently, with both output files simply growing past their expected size with no error at any point. The fix is placing the environment assignment inside the queued command itself, and the failure mode, like the others, produced no GPU-side symptom at all: the job ran to completion, on the GPU, successfully, doing four orders of magnitude more work than intended.

The rule

Preflight on CPU first: build the binary, run the gate once on its smallest input, read back every arm-defining flag from the running process. A GPU job should never be the place a build error, a missing flag, or a bad config is discovered.

What counts as GPU-worthy and what does not

What belongs on the card: - The actual timed run, once the binary is known to build, the gate is known to pass on a small input, and every flag the run depends on has been read back from a live process rather than assumed from the command line typed. - A resident server, held only for as long as GPU work is actually queued behind it, reusing one process across a sequence of jobs rather than restarting per job. - The deterministic reference arm, cached by model hash, binary commit, parameters, and input ids, so it runs on the GPU exactly once per unique configuration rather than once per comparison.

What does not: - Anything a CPU build or a dry run would catch: compilation, a flag's effect on a config dump, a stdin/env wiring check, a settings-template lint. - CPU-side scoring, report generation, or log parsing, none of which needs the card and all of which can run while the next GPU job is already queued. - A hold that exists only as prose. If it matters enough to block a launch, it needs to live where the launching mechanism reads it, not where a human might remember to look.

What it cost, measured against not doing this

The 2026-09-19 incident is the clearest single price tag: a gpu-wait run --vram 22 job on the same 24 GB card that also drives the desktop compositor ran the GPU near its ceiling while ghostty and every herdr pane inside it were open. At 04:16 the kernel logged "Not enough memory for command submission," ghostty's own GPU context was rejected, and it exited, taking every pane running inside it down with it. Recovery was a claude --resume per session in fresh workspaces, not a lost result, but it is a cost that a five-second check of gpu-wait gpu's live VRAM figure against the job's declared footprint would have avoided entirely, on the CPU, before the job ever launched.

The --timeout flag's own failure mode, from 2026-09-09, is a smaller version of the same lesson applied to a different axis: it is a maximum runtime, not a queue-wait budget, and a long-lived ComfyUI render was killed mid-step because a batch-job flag was applied to a server. The daemon's own log said timeout plainly; the client just reported a bare connection refusal, which is what made the diagnosis take longer than the five-second read of a log line that already had the answer in it.

How to prove this wrong

Track your own GPU job failures for a week, sorted by whether the cause was visible without the GPU: a build failure, a flag with no effect, a stdin or environment variable that never reached the job, a config value that a CPU-side dump would have shown was wrong, versus an actual runtime resource conflict that only manifests once the job is on the card. If your own ratio comes back with the majority in the second category, this post's central claim, that CPU preflight is the majority of the fix rather than a nice-to-have on top of it, does not hold for your workload, and that would be worth knowing. The specific figure this post opens with, 34 of 135 on 2026-09-16, is one day on one machine with one card shared between a desktop and AI work; a dedicated headless GPU box with no compositor to starve and no human editing config files by hand would reasonably see a different split.

Provenance

Single AMD RX 7900 XTX (gfx1100, RDNA3), 24 GB, shared between a KDE Wayland desktop and all AI GPU work on the machine, managed by the gpu-waitd user daemon (cooperative scheduling, not enforced at the driver level). The 34-of-135 figure and the preflight rule it produced are stated in the operating rules at ~/.claude/CLAUDE.md section 17, dated 2026-09-16. The whiteboard-hold incident (2026-09-11), the environment-forwarding incident, the timeout incident (2026-09-09), and the ghostty crash (2026-09-19) are each their own dated entry in ~/Brain/m.ledger/gpuwaitingroom.md. gpu-waitd's own source and README are at ~/Projects/gpu-media/gpuwaitingroom/.

Commentaires

Pas encore de commentaires.

Se connecter pour commenter.