Files
hyperguild/brain/sessions/analysis/agent-intent-column.md
mathiasandClaude Opus 4.8 e8dbcf6eef
CI / Lint / Test / Vet (push) Successful in 14s
CI / Mirror to GitHub (push) Successful in 4s
docs(brain): add agent-consumer brain-MCP intent analysis
Agent-consumer column of the two-part brain intent↔interface study.
Reconstructs the knowledge-act behind every brain MCP call in the
Claude Code agent transcripts on koala (the only corpus with brain
calls; brain/sessions and agentsquad eval logs carry none).

46 distinct knowledge-acts, 37% interface mismatch, 0% intent_unclear.
Headline gap: no update/supersede verb → agents blind re-write same
slug (5x); no read-after-write → lexical re-query of own note (4x);
lexical-only reads → semantic-as-keyword-stuffing chains (3x);
brain_answer hedged with parallel brain_query.

Canonical schema brain-intent-extraction.md absent on host; vocab
reconstructed, every row tagged schema_source=reconstructed.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-15 19:44:59 +02:00

8.8 KiB
Raw Permalink Blame History

Agent-Consumer Brain Intent Analysis — koala column

Consumer: autonomous_agent (all rows). Host: koala. Date: 2026-06-15. Raw rows: agent-intent-column.jsonl (46 real knowledge-acts + 1 excluded false-positive).

Caveat — canonical schema not on this host

The shared closed-vocabulary file brain-intent-extraction.md does not exist on koala — the only reference to it is inside this task's own prompt. The intent vocabulary below was reconstructed from the prompt's framing + brain/schema.md. Every row carries schema_source: reconstructed. If the canonical vocab differs, re-map the intent field; the intent_tool_match / workaround / observed_friction evidence stands regardless of label names.

Reconstructed closed vocab: semantic_retrieval, lexical_lookup, check_prior_art, synthesized_answer, store_new_knowledge, update_or_supersede, ingest_raw_source, verify_write_landed, discover_capability, intent_unclear.

Corpus

  • ~/.claude/projects/*/*.jsonl — Claude Code agent transcripts (98 files). The only source with brain calls.
  • brain/sessions/*.jsonl — empty (only .gitkeep).
  • agentsquad docs/eval/*.jsonl — code-review eval results, not brain calls.
  • No separate Crush session logs on this host.
  • Zero brain calls were human-typed. Every brain act is agent-initiated (the CLAUDE.md "query as reflex / close-the-loop write" behaviour), so all qualify as autonomous_agent. The human gave the top-level task; the agent chose every brain act.

1. Intent histogram, split by intent_tool_match

intent match mismatch partial total
check_prior_art 10 0 0 10
store_new_knowledge 14 1 0 15
update_or_supersede 0 5 0 5
synthesized_answer 2 2 1 5
verify_write_landed 0 4 0 4
semantic_retrieval 0 3 0 3
ingest_raw_source 2 0 0 2
discover_capability 0 2 0 2
total 28 17 1 46

Batch note: several rows collapse a same-second fan-out of identical-intent calls (a10=3, a11=4, a14=5, a33=3, a34=3, a35=3). Call-level the corpus is ~62 brain calls; the table counts the 46 distinct knowledge-acts. Frequency is deliberately not the point — the mismatch column is.

37% of agent knowledge-acts (17/46) are interface mismatches. Every mismatch falls into one of four intents: update_or_supersede, verify_write_landed, semantic_retrieval, discover_capability — plus one store that had to use HTTP.

2. Mismatch list, grouped by intent (primary deliverable)

update_or_supersede → re-write same slug (5/5 mismatch) — HIGHEST VALUE

There is no update / patch / append / supersede verb. When an agent improves a note it already wrote, the only move is to call brain_write/brain_ingest again with the same filename/source and hope the index replaces rather than duplicates. Observed:

slug 1st write 2nd write gap what changed
tapir-rls-identity-bootstrapping (ingest) 14:00:59 14:02:08 69s condensed body
homelab-security-chains-not-bugs.md 09:28:22 10:32:49 64m new worked example
webfetch-readme-when-image-or-flag-uncertain.md 18:31:23 18:34:23 3m near-identical (retry)
gitea-mcp-per-repo-tools-404-and-rest-fallback 05:57:37 05:57:55 18s content tweak
cannot-move-ingress-host-across-namespaces-flux-dryrun 13:59:50 14:00:13 23s content tweak

Sub-30s gaps (3 of 5) read as "I wanted to edit but can only overwrite." The agent has no way to know whether the second write deduped or created a contradictory v2 — opacity the brain's own design principle (mcp-tool-design-get-needs-list-partner.md, written by one of these very agents) would flag: every _write needs a _get/_update partner.

verify_write_landed → lexical re-query (4/4 mismatch)

No read-after-write / get-by-id confirmation. After every substantive write, agents re-query lexically to check the note is retrievable — and reformulate the keywords because they can't predict what BM25 indexed:

  • exit-255 lesson: query exit 255 unknown reason restart loop diagnosis → write → query exit 255 unknown SIGKILL containerd.
  • extension-version-lags: query extension build pinned version...pgvector → write → query pgvector postgres extension version compile error bump.
  • webfetch-readme: write → write → query stuffed with the note's own section headings.

semantic_retrieval → BM25 keyword-stuffing (3/3 mismatch)

Agent knows the meaning but not the indexed words, so it brute-forces phrasings of one need against a lexical tool:

  • 4-step chain on one Go bug: bytes.Buffer Bytes Reset aliasing slice sharinggo buffer reuse map backing array bug → (write) → go map values all show same content after loopbytes.Buffer Bytes returns same data every iteration. Symptom and cause both known; no single lexical query bridges them.
  • Natural-language questions (how to enable prometheus metrics on litellm proxy callback, moved compose stack to new directory volumes disappeared empty) shoved into brain_query.

synthesized_answer → fall back to / hedge with brain_query (2 mismatch + 1 partial)

brain_answer is frequently not trusted as terminal:

  • Strongest signal: 3× brain_answer at 15:08 (compose volumes / pi rebuild time / litellm ModuleNotFound) → 3× brain_query at 15:09 on the same three topics. The agent asked the synthesizer, was unsatisfied, and immediately re-ran raw search.
  • Dex→Authentik: brain_query and brain_answer fired 8s apart on one question — the agent pays both interfaces because it can't predict which returns usable knowledge.
  • (Counter-examples exist: brain_answer for episodic recall — "what happened with the litellm migration on 2026-05-16?" — was consumed and not re-queried. So answer works for episodic/temporal recall, fails for how-do-I / does-X-hold reasoning.)

discover_capability + store-via-HTTP (3 mismatch)

brain tools are not ambient — they are deferred and must be ToolSearch-loaded each session. Agents fumble the discovery (ToolSearch 'brain knowledge memory''brain''brain ingestion knowledge wiki') and, when MCP load/auth fails, drop to raw curl against brain-mcp / hand-built /tmp/brain_entry.json. One agent even ToolSearch-ed an authenticate tool, then abandoned MCP for the HTTP bodge — matching the known "brain/op MCP auth lapses too often" footgun.

Write-interface / layer schema confusion (cross-cutting)

brain_write was called with three different param shapes in the same corpus: {filename, type:"lesson", content}, {filename, content} (no type), and {wing:"tapir", hall:"failures", filename, content} — plus brain_ingest {source, content}. Agents are unsure which verb and which layer (flat slug vs wing/hall knowledge routing vs raw ingest) a given knowledge-act maps to. This is the knowledge/ vs wiki/ confusion expressed at the parameter level.

3. intent_unclear rate

0 / 46 genuine brain acts (0%). Agent intent is unusually legible because these are Claude Code transcripts: the surrounding task, the query/filename strings, and the write content all disambiguate. One row (a47) was tagged intent_unclear and excluded — its brain_ substring was inside a git commit message, not a brain call. Example of the only ambiguity that arose: a brain_answer on "gitea MCP not working... how to authenticate" followed by an unrelated query — can't tell if the answer satisfied or was abandoned (partial, row a28).

4. The single biggest intent↔interface gap

The brain offers one write verb and one lexical read verb, but autonomous agents perform four distinct knowledge-acts against them — and three of the four have no fitting interface. The deepest gap is the missing update/supersede path: agents close every task by writing a lesson (the CLAUDE.md ritual), routinely improve it minutes-to-an-hour later, and — having no edit primitive — re-write the same slug blind, unable to tell whether they corrected the entry or forked a contradiction into the index. This compounds with the lexical-only read side: because there is no get-by-id or semantic retrieval, agents can't even reliably find their own just-written note to check it, so they keyword-stuff reformulated queries and hedge brain_answer with parallel brain_query. The interface is built for append-and-keyword-search; the agents are trying to curate a living, deduplicated knowledge base, and the seam between those two shows up as the 5 blind re-writes, 4 read-after-write re-queries, and 3 semantic-as-lexical chains that dominate the mismatch column.


Evidence-only per task scope — no redesign proposed.