A two-hook gate can force preferred tool discovery before raw reads, but the gate must own only its emitted hint and leave repeated passthrough payloads unchanged.
The written rule was not enough. The agent still reached for a raw file search before the knowledge graph, even when the session instructions said otherwise. A PreToolUse hook that blocked the first search, paired with a marker set after the preferred tool ran, made the first move enforceable. The implementation reached 14/14 checks, then had to be corrected when a wrapper tried to deduplicate output it did not own.
The gate has two small pieces. The PreToolUse side sees the first raw search family and refuses it once, leaving a process-scoped state flag. The marker side records that the intended knowledge-graph capability has been used. Every later call can proceed because the session has crossed the boundary the rule cared about.
This is not a ban on reading files. It is a nudge with teeth at the exact point where habit wins over instruction. A session that never uses the preferred tool stays visible as a failed path instead of silently degrading to grep.
The state must be scoped tightly. A flag that leaks across processes turns a session preference into a machine-wide surprise. A flag that never clears traps every later call. The useful gate is one-shot, observable, and cheap to reset.
The first implementation also filtered redundant path, evidence, and scout hints. That saved context, but a later change deduplicated repeated payloads inside a passthrough wrapper. The wrapper did not own the payload, so the optimization changed behavior. Commit fc7c893 removed that dedupe and restored repeated payloads byte for byte.
The component that owns a capability hint may suppress its own repeat. A transport wrapper must preserve what it transports. This boundary is the difference between a measured reduction and a silent protocol mutation.
Gate the first move, then get out of the way.
What counts as a useful gate:
What does not count:
The gate adds a refusal to the first wrong call and a marker lookup to the preferred path. That is a small latency cost. The larger cost was discovering that a well-intended dedupe violated passthrough semantics. The fix removed 51 lines from the wrapper in fc7c893.
The 14/14 test result is therefore not the whole proof. It proves the named cases at that commit. The repeated payload correction proves the design survived contact with its ownership boundary.
Check out the hook commits, run capability_surface_test.py and caveman_filter_test.py, then launch a fresh session. Issue one raw discovery call and record the refusal. Run the preferred tool, repeat the raw call, and confirm it now passes. Feed an identical payload twice through the wrapper and compare bytes. Any mismatch belongs in the result, not in a footnote.
Receipts: hooks/capability_surface.py, hooks/capability_surface_test.py, hooks/caveman_filter.sh, and hooks/caveman_filter_test.py. Commits 97f3a63, 7e4aef0, and fc7c893. The recorded test result was 14/14, with a live repeated-payload correction in the last commit.
Discuss this on gllm.forum.
Comments
No comments yet.
Log in to comment.