← Blog
✎post de blog

The IPC Primitive Is a Latent Tensor

A monolithic dark slab resting above a dense sea of cloud, under a heavy sky broken by a single bright band at the horizon.

The message between agents on one machine can be a latent tensor with an identity rather than text, and the transport cost of moving one is already measured.

The one sentence

LatentOS is a RAM-first immutable node OS, written in Mojo on our own inference engine, where the message agents exchange is a latent tensor with an identity, and text is one adapter among several.

Design is complete: nine documents, 1679 lines, frozen 2026-09-08. Nothing boots. No agent runs. Eight of eleven preregistered experiments have receipts. The one the whole thesis rests on is inconclusive as of tonight, and this post says so before it says anything else.

Why this is not another latent-communication paper

Every published system in this space (LatentMAS, Interlat, Coconut, Communicating Activations) was built by people who do not own the engine they run on. That constraint shows in the method. You get forward hooks, past_key_values tuples, and a framework dict carrying tensors between processes. The latent is a thing you reach in and grab, and its meaning is whatever the framework happened to preserve.

We own the forward pass, the KV cache, the SSM state, the sampler and the server, because mojo-baro is ours end to end. It already ships a byte-exact serialisable latent: the prefix checkpoint restores 1087 of 1088 tokens in 7.6 ms against 810 ms cold. That is not a hook. That is a first-class artifact the engine can name.

So the design inverts the usual premise. Instead of bolting a latent subsystem onto a text-first stack, the OS IPC primitive is a latent handle, and the fleet's unit of identity is a latent-speaking role. Text becomes one adapter you can attach, not the substrate everything else is translated through.

Three measured facts the design rests on

RAM-first is arithmetic, not preference. Decode re-reads the entire model per token. On the lab box that is 18.3 GB/s of reads against 22.9 GB/s of RAM and 0.53 GB/s of SSD. Pinned host-to-device over PCIe peaks at 28.5 GB/s, half the aggregate CPU read. The conclusion writes itself: maximise residency, never stream weights.

Identity is a tuple, not a model name. Weights are the GGUF UUID, which is tensor data only and proven blind to metadata. Role is the sha256 of the role file, prompt included. Runtime is the engine build. The same weight bytes moved perplexity 4.1% across two builds, so the build is part of the name whether you like it or not. Latents add two more axes: Σ space and batch class. And throughput is a curve in context depth, never a single number.

Latents are not batch-invariant. A served 27B model returns different bits for the same prompt depending on which tenants were co-batched with it, because co-batching changes reduction order. This is the fact with the sharpest consequence: a latent's reproducibility class has to be part of its name, enforced at mint time. You cannot recover it later.

The components

Node agent (doc 01, Phase 1). A Mojo systemd service on a stock Kairos core, replacing kairos-agent and the RAM half of immucore. It owns tiering (verify the boot reservation, pin, evict, probe bandwidth), plus identity and manifest, the latent handle table with its memfd store and GC, and curve-based admission with a quote-versus-actual EWMA. Upgrade is delegated to stock Kairos in Phase 1. The engine runs as a child process in its own crash domain and gains three entry points: mint_latent, ingest_latent, inject_latent. Estimated 5 to 7 MB against Kairos's 15 MB of Go. Mojo 1.0 produces no static binary; that shapes the initrd, not the design.

RAM tier (doc 02). Three tiers: the OS itself (520 MiB idle, measured, against a 4 GiB budget), one pinned active role plus a standby in a boot-reserved hugepage region (32 GiB on the desktop, 5 GiB on the lab box), and an evictable disk tier. Exactly one requirement is genuinely boot-path: reserving 1 GiB pages before userland fragments memory.

Latent IPC (doc 03, the core). A handle is a 256-byte header plus an opaque payload. The header carries kind, layer set, position range, role sha, Σ id, batch class and an HMAC. Five families from the literature were triaged rather than adopted wholesale:

family verdict why
KV/SSM transfer admitted, native it is prefix.mojo with a header on it
continuous thought admitted as HIDDEN the thesis kind
activation fusion admitted, experimental unproven, gated behind E11
logit blending rejected 970 KiB per step for an ensemble trick
Σ adapters deferred to v1 not needed for a single-model fleet

Transport follows breakeven arithmetic, not taste: a pointer in-process, a memfd over SCM_RIGHTS same-host, and cross-host either a tokens-plus-hash recipe or the raw payload, whichever the arithmetic favours. For scale: an 8-step HIDDEN message is 128 KiB against 1.2 KB of equivalent text: 107x the bytes, and 37x less producer time.

Identity and manifest (doc 04). A node advertises "I am a tagger, here is my curve, here is when I can start", not "I have model X." The manifest key is role sha, runtime, Σ and batch class, mapped to curve coefficients and tier state. Admission quotes an ETA from queue depth, load, ingest, the prefill curve, the decode curve and adapter cost, and accepts only if the upper bound fits the deadline.

Boot path in Mojo (doc 05, Phase 2, gated). Fork the image only if the stock config cannot carry the cmdline reservation, or pinned pages do not survive pressure. The prior was stated in writing before the experiments ran: the gate will say don't fork.

Legacy mobile profile (doc 08). The same design on 2017-to-2022 phones: a 1.2 GiB working set under LMKD, eMMC treated as read-only with zero flash writes, a 1.8 W passive envelope, and an ashmem fallback for old kernels.

What is measured, and what is not

exp question result
E1 byte table KV 64 KiB/token, checkpoint 50.25 MiB. The board's 26.25 was an external estimate.
E2 stock Kairos carries the hugepage cmdline PASS, both boot targets
E3 pinned pages survive 1.5x RAM page-cache pressure PASS, hugetlb and mlock arms
E4 tier bandwidths NVMe cold 4.64 GB/s, page cache 22.33 GB/s
E5 Mojo 1.0 can do memfd seals, mlock2, mincore, SCM_RIGHTS PASS
E6 hugetlb pages as ROCm pinned memory 28.49 GB/s, same as hipHostMalloc
E7 same-host handoff 10 µs handoff, 3.47 ms total with touch, beats the 7.6 ms restore
E10 FIXED(M) batches position-invariant PASS, admitted as canonical alongside M1
E9 bf16 enough for HIDDEN INVALID, read a buffer the megakernel never writes
E8 HIDDEN works training-free INCONCLUSIVE, task set had no headroom
E11 activation fusion not run, sequenced after E8

E9 deserves its own sentence, because it is the most instructive failure here. The experiment read a hidden-state buffer that looked entirely healthy, with real values, varying per prompt, at plausible magnitudes, and was never written by the fused kernel on the path being measured. It carried a stale row from an earlier prefill chunk. A result that looks like it works is the most expensive kind of wrong, and the only reason it was caught is that the buffer's write sites were grepped rather than assumed. The corrected version derives from the one buffer every code path genuinely keeps current.

So: the fork gate is resolved, and it says don't fork. The transport and tier numbers are real. Reproducibility has a canonical class. And the systems claim survives prior-art review only in its narrow form: an engine that defines its own reduction order can mint artifacts whose name says whether they reproduce. Nobody else can, because nobody else owns the engine.

The part that decides everything is still open

What is unmeasured is the thing the design lives or dies on: whether an untrained model reads its own hidden states as reasoning.

Tonight established three things. The path is live: answers change on 18 of 20 items when latents are injected. It is non-destructive: math stays 20 of 20 under both 8 and 32 injected vectors. And it is 12x faster to produce than text chain-of-thought.

It did not establish quality, because the text arm tied the no-handoff baseline. When your control does not beat doing nothing, the task set cannot tell you anything about the treatment. That is a defect in the experiment, not a result about latents.

The strongest case against this design, in the docs' own words: latents may buy speed only, and never quality (Wenzel 2026 matches text, never exceeds it). Phase 1 serves single-model fleets only. Reproducibility costs producer throughput. And our slow prefill flatters every transfer-versus-recompute breakeven in the document.

What happens next

Rebuild the E8 task set with actual headroom. Re-freeze the gate with a precondition that was missing the first time, which is that text must beat baseline before the latent arm means anything, and rerun. Rerun E9 off the corrected hidden state.

If E8 passes, or lands speed-only, the latent primitive stands and E11 and the cross-host wire format follow. If it kills, the primitive shrinks to KV/prefix handoff and the OS becomes prefix reuse across processes with identity attached, which the docs already commit, in writing, to saying out loud.

How to prove this wrong

Two separate claims sit in this post, and each has its own way to fail. The transport claim, that a latent message costs less to move than the text it replaces, is already falsifiable on your own rig: move a state of comparable size as text through the same transport and measure both against the 107x-bytes, 37x-less-producer-time figure above; if text is not materially more expensive, that premise does not hold. The reasoning claim, that an untrained model reads its own hidden states as reasoning, is the one still open: it stands or falls on the rebuilt E8 run, against a text arm that first has to beat the no-handoff baseline before the comparison means anything.

Preregistration is worth something only if you publish the inconclusive ones too.

Comentarii

Niciun comentariu încă.

Conectare pentru a comenta.