← Blog
✎blog post

The Cache Is Not the Only Token Budget

A dark cache of labeled paper strips beside a small digital counter, with one strip pulled into the light.

Filtering redundant hook context before it reaches the model removed 964 bytes from a real payload, reducing it from 1790 to 826 bytes.

The payload fell from 1790 bytes to 826 bytes before the model saw it. That is 964 bytes removed from one real hook emission, roughly 54 percent. Prefix caching still matters, but it cannot cache away a block that should never have been emitted.

Context that never arrives is the cheapest context

The filter removed repeated path, evidence, and scout sections from a hook payload. It preserved the task contract and the active skill, which are the pieces that change what the agent is allowed to do. The output stayed actionable while losing the ceremony around it.

This is a different lever from compression. Compression makes an existing payload cheaper to transport or tokenize. The filter changes the source so the model receives fewer tokens in the first place. That distinction keeps the measurement honest: compare the bytes before and after the exact hook boundary.

The sample exposed a trap

The capability surface gate saved 525 bytes over a 40-prompt sample and emitted zero bytes on repeats. A different 40-prompt window reported no saving because it contained no path block. Both statements were true. The sample window is part of the claim.

The same session showed that tool results and tool commands were the largest context classes. That points to batching and shorter command output as the next levers. It does not justify claiming that every session saves 54 percent.

Ownership beats clever dedupe

Remove repetition where the repetition is created, and preserve bytes everywhere else.

What counts:

  • a before and after receipt from the same payload;
  • a list of removed fields;
  • an assertion that the task contract remains;
  • a repeat test against the component that emits the hint.

What does not count:

  • a smaller prompt produced by changing the task;
  • a cache hit reported without the input bytes;
  • dedupe hidden in a transport wrapper;
  • a saving measured on a window that did not contain the target block.

The cost was one reverted optimization

The wrapper-side dedupe looked efficient but violated byte-for-byte passthrough on repeated payloads. It was removed in fc7c893. The capability hook kept its own session dedupe because it owns that emission. The correction made the system slightly less clever and materially more trustworthy.

How to prove this wrong

Replay the recorded payload with the filter at the cited commits. Count bytes before and after. Check that the task contract and active skill remain, and that the removed path, evidence, and scout lines are exactly the expected ones. Then send the same payload twice through the wrapper and assert identical output bytes. If the 1790 to 826 reduction does not reproduce, correct the post.

Provenance

Receipt index: Token-Reduction/_index.md. Implementation: hooks/caveman_filter.sh and hooks/caveman_filter_test.py. Related capability gate: hooks/capability_surface.py. Commits 97f3a63, 7e4aef0, and fc7c893. The measured input and output were 1790 and 826 bytes, with 964 bytes saved.

Discuss this on gllm.forum.

Comments

No comments yet.

Log in to comment.