Agent-consumer column of the two-part brain intent↔interface study. Reconstructs the knowledge-act behind every brain MCP call in the Claude Code agent transcripts on koala (the only corpus with brain calls; brain/sessions and agentsquad eval logs carry none). 46 distinct knowledge-acts, 37% interface mismatch, 0% intent_unclear. Headline gap: no update/supersede verb → agents blind re-write same slug (5x); no read-after-write → lexical re-query of own note (4x); lexical-only reads → semantic-as-keyword-stuffing chains (3x); brain_answer hedged with parallel brain_query. Canonical schema brain-intent-extraction.md absent on host; vocab reconstructed, every row tagged schema_source=reconstructed. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
8.8 KiB
Agent-Consumer Brain Intent Analysis — koala column
Consumer: autonomous_agent (all rows). Host: koala. Date: 2026-06-15.
Raw rows: agent-intent-column.jsonl (46 real knowledge-acts + 1 excluded false-positive).
Caveat — canonical schema not on this host
The shared closed-vocabulary file brain-intent-extraction.md does not exist on
koala — the only reference to it is inside this task's own prompt. The intent
vocabulary below was reconstructed from the prompt's framing + brain/schema.md.
Every row carries schema_source: reconstructed. If the canonical vocab differs,
re-map the intent field; the intent_tool_match / workaround / observed_friction
evidence stands regardless of label names.
Reconstructed closed vocab: semantic_retrieval, lexical_lookup,
check_prior_art, synthesized_answer, store_new_knowledge, update_or_supersede,
ingest_raw_source, verify_write_landed, discover_capability, intent_unclear.
Corpus
~/.claude/projects/*/*.jsonl— Claude Code agent transcripts (98 files). The only source with brain calls.brain/sessions/*.jsonl— empty (only.gitkeep).agentsquad docs/eval/*.jsonl— code-review eval results, not brain calls.- No separate Crush session logs on this host.
- Zero brain calls were human-typed. Every brain act is agent-initiated (the
CLAUDE.md "query as reflex / close-the-loop write" behaviour), so all qualify as
autonomous_agent. The human gave the top-level task; the agent chose every brain act.
1. Intent histogram, split by intent_tool_match
| intent | match | mismatch | partial | total |
|---|---|---|---|---|
| check_prior_art | 10 | 0 | 0 | 10 |
| store_new_knowledge | 14 | 1 | 0 | 15 |
| update_or_supersede | 0 | 5 | 0 | 5 |
| synthesized_answer | 2 | 2 | 1 | 5 |
| verify_write_landed | 0 | 4 | 0 | 4 |
| semantic_retrieval | 0 | 3 | 0 | 3 |
| ingest_raw_source | 2 | 0 | 0 | 2 |
| discover_capability | 0 | 2 | 0 | 2 |
| total | 28 | 17 | 1 | 46 |
Batch note: several rows collapse a same-second fan-out of identical-intent calls (a10=3, a11=4, a14=5, a33=3, a34=3, a35=3). Call-level the corpus is ~62 brain calls; the table counts the 46 distinct knowledge-acts. Frequency is deliberately not the point — the mismatch column is.
37% of agent knowledge-acts (17/46) are interface mismatches. Every mismatch falls
into one of four intents: update_or_supersede, verify_write_landed,
semantic_retrieval, discover_capability — plus one store that had to use HTTP.
2. Mismatch list, grouped by intent (primary deliverable)
update_or_supersede → re-write same slug (5/5 mismatch) — HIGHEST VALUE
There is no update / patch / append / supersede verb. When an agent improves a note
it already wrote, the only move is to call brain_write/brain_ingest again with the
same filename/source and hope the index replaces rather than duplicates. Observed:
| slug | 1st write | 2nd write | gap | what changed |
|---|---|---|---|---|
tapir-rls-identity-bootstrapping (ingest) |
14:00:59 | 14:02:08 | 69s | condensed body |
homelab-security-chains-not-bugs.md |
09:28:22 | 10:32:49 | 64m | new worked example |
webfetch-readme-when-image-or-flag-uncertain.md |
18:31:23 | 18:34:23 | 3m | near-identical (retry) |
gitea-mcp-per-repo-tools-404-and-rest-fallback |
05:57:37 | 05:57:55 | 18s | content tweak |
cannot-move-ingress-host-across-namespaces-flux-dryrun |
13:59:50 | 14:00:13 | 23s | content tweak |
Sub-30s gaps (3 of 5) read as "I wanted to edit but can only overwrite." The agent has
no way to know whether the second write deduped or created a contradictory v2 — opacity
the brain's own design principle (mcp-tool-design-get-needs-list-partner.md, written
by one of these very agents) would flag: every _write needs a _get/_update partner.
verify_write_landed → lexical re-query (4/4 mismatch)
No read-after-write / get-by-id confirmation. After every substantive write, agents re-query lexically to check the note is retrievable — and reformulate the keywords because they can't predict what BM25 indexed:
exit-255lesson: queryexit 255 unknown reason restart loop diagnosis→ write → queryexit 255 unknown SIGKILL containerd.extension-version-lags: queryextension build pinned version...pgvector→ write → querypgvector postgres extension version compile error bump.webfetch-readme: write → write → query stuffed with the note's own section headings.
semantic_retrieval → BM25 keyword-stuffing (3/3 mismatch)
Agent knows the meaning but not the indexed words, so it brute-forces phrasings of one need against a lexical tool:
- 4-step chain on one Go bug:
bytes.Buffer Bytes Reset aliasing slice sharing→go buffer reuse map backing array bug→ (write) →go map values all show same content after loop→bytes.Buffer Bytes returns same data every iteration. Symptom and cause both known; no single lexical query bridges them. - Natural-language questions (
how to enable prometheus metrics on litellm proxy callback,moved compose stack to new directory volumes disappeared empty) shoved intobrain_query.
synthesized_answer → fall back to / hedge with brain_query (2 mismatch + 1 partial)
brain_answer is frequently not trusted as terminal:
- Strongest signal: 3×
brain_answerat 15:08 (compose volumes / pi rebuild time / litellm ModuleNotFound) → 3×brain_queryat 15:09 on the same three topics. The agent asked the synthesizer, was unsatisfied, and immediately re-ran raw search. - Dex→Authentik:
brain_queryandbrain_answerfired 8s apart on one question — the agent pays both interfaces because it can't predict which returns usable knowledge. - (Counter-examples exist:
brain_answerfor episodic recall — "what happened with the litellm migration on 2026-05-16?" — was consumed and not re-queried. Soanswerworks for episodic/temporal recall, fails for how-do-I / does-X-hold reasoning.)
discover_capability + store-via-HTTP (3 mismatch)
brain tools are not ambient — they are deferred and must be ToolSearch-loaded each
session. Agents fumble the discovery (ToolSearch 'brain knowledge memory' →
'brain' → 'brain ingestion knowledge wiki') and, when MCP load/auth fails, drop to
raw curl against brain-mcp / hand-built /tmp/brain_entry.json. One agent even
ToolSearch-ed an authenticate tool, then abandoned MCP for the HTTP bodge — matching
the known "brain/op MCP auth lapses too often" footgun.
Write-interface / layer schema confusion (cross-cutting)
brain_write was called with three different param shapes in the same corpus:
{filename, type:"lesson", content}, {filename, content} (no type), and
{wing:"tapir", hall:"failures", filename, content} — plus brain_ingest {source, content}. Agents are unsure which verb and which layer (flat slug vs wing/hall
knowledge routing vs raw ingest) a given knowledge-act maps to. This is the
knowledge/ vs wiki/ confusion expressed at the parameter level.
3. intent_unclear rate
0 / 46 genuine brain acts (0%). Agent intent is unusually legible because these are
Claude Code transcripts: the surrounding task, the query/filename strings, and the
write content all disambiguate. One row (a47) was tagged intent_unclear and
excluded — its brain_ substring was inside a git commit message, not a brain call.
Example of the only ambiguity that arose: a brain_answer on "gitea MCP not working...
how to authenticate" followed by an unrelated query — can't tell if the answer satisfied
or was abandoned (partial, row a28).
4. The single biggest intent↔interface gap
The brain offers one write verb and one lexical read verb, but autonomous agents
perform four distinct knowledge-acts against them — and three of the four have no fitting
interface. The deepest gap is the missing update/supersede path: agents close every
task by writing a lesson (the CLAUDE.md ritual), routinely improve it minutes-to-an-hour
later, and — having no edit primitive — re-write the same slug blind, unable to tell
whether they corrected the entry or forked a contradiction into the index. This compounds
with the lexical-only read side: because there is no get-by-id or semantic retrieval,
agents can't even reliably find their own just-written note to check it, so they
keyword-stuff reformulated queries and hedge brain_answer with parallel brain_query.
The interface is built for append-and-keyword-search; the agents are trying to
curate a living, deduplicated knowledge base, and the seam between those two shows up
as the 5 blind re-writes, 4 read-after-write re-queries, and 3 semantic-as-lexical chains
that dominate the mismatch column.
Evidence-only per task scope — no redesign proposed.