# Agent-Consumer Brain Intent Analysis — koala column **Consumer:** `autonomous_agent` (all rows). **Host:** koala. **Date:** 2026-06-15. **Raw rows:** `agent-intent-column.jsonl` (46 real knowledge-acts + 1 excluded false-positive). ## Caveat — canonical schema not on this host The shared closed-vocabulary file `brain-intent-extraction.md` **does not exist on koala** — the only reference to it is inside *this task's own prompt*. The intent vocabulary below was **reconstructed** from the prompt's framing + `brain/schema.md`. Every row carries `schema_source: reconstructed`. If the canonical vocab differs, re-map the `intent` field; the `intent_tool_match` / `workaround` / `observed_friction` evidence stands regardless of label names. **Reconstructed closed vocab:** `semantic_retrieval`, `lexical_lookup`, `check_prior_art`, `synthesized_answer`, `store_new_knowledge`, `update_or_supersede`, `ingest_raw_source`, `verify_write_landed`, `discover_capability`, `intent_unclear`. ## Corpus - `~/.claude/projects/*/*.jsonl` — Claude Code agent transcripts (98 files). **The only source with brain calls.** - `brain/sessions/*.jsonl` — empty (only `.gitkeep`). - `agentsquad docs/eval/*.jsonl` — code-review eval results, **not** brain calls. - No separate Crush session logs on this host. - **Zero** brain calls were human-typed. Every brain act is agent-initiated (the CLAUDE.md "query as reflex / close-the-loop write" behaviour), so all qualify as `autonomous_agent`. The human gave the top-level task; the agent chose every brain act. ## 1. Intent histogram, split by `intent_tool_match` | intent | match | mismatch | partial | total | |---|---|---|---|---| | check_prior_art | 10 | 0 | 0 | 10 | | store_new_knowledge | 14 | 1 | 0 | 15 | | update_or_supersede | 0 | 5 | 0 | 5 | | synthesized_answer | 2 | 2 | 1 | 5 | | verify_write_landed | 0 | 4 | 0 | 4 | | semantic_retrieval | 0 | 3 | 0 | 3 | | ingest_raw_source | 2 | 0 | 0 | 2 | | discover_capability | 0 | 2 | 0 | 2 | | **total** | **28** | **17** | **1** | **46** | > Batch note: several rows collapse a same-second fan-out of identical-intent calls > (a10=3, a11=4, a14=5, a33=3, a34=3, a35=3). Call-level the corpus is ~62 brain calls; > the table counts the 46 distinct knowledge-acts. Frequency is deliberately *not* the > point — the mismatch column is. **37% of agent knowledge-acts (17/46) are interface mismatches.** Every mismatch falls into one of four intents: `update_or_supersede`, `verify_write_landed`, `semantic_retrieval`, `discover_capability` — plus one `store` that had to use HTTP. ## 2. Mismatch list, grouped by intent (primary deliverable) ### update_or_supersede → re-write same slug (5/5 mismatch) — HIGHEST VALUE There is **no update / patch / append / supersede verb**. When an agent improves a note it already wrote, the only move is to call `brain_write`/`brain_ingest` **again with the same filename/source** and hope the index replaces rather than duplicates. Observed: | slug | 1st write | 2nd write | gap | what changed | |---|---|---|---|---| | `tapir-rls-identity-bootstrapping` (ingest) | 14:00:59 | 14:02:08 | 69s | condensed body | | `homelab-security-chains-not-bugs.md` | 09:28:22 | 10:32:49 | 64m | new worked example | | `webfetch-readme-when-image-or-flag-uncertain.md` | 18:31:23 | 18:34:23 | 3m | near-identical (retry) | | `gitea-mcp-per-repo-tools-404-and-rest-fallback` | 05:57:37 | 05:57:55 | 18s | content tweak | | `cannot-move-ingress-host-across-namespaces-flux-dryrun` | 13:59:50 | 14:00:13 | 23s | content tweak | Sub-30s gaps (3 of 5) read as "I wanted to edit but can only overwrite." The agent has no way to know whether the second write deduped or created a contradictory v2 — opacity the brain's own design principle (`mcp-tool-design-get-needs-list-partner.md`, written *by one of these very agents*) would flag: every `_write` needs a `_get`/`_update` partner. ### verify_write_landed → lexical re-query (4/4 mismatch) No read-after-write / get-by-id confirmation. After every substantive write, agents re-query lexically to check the note is retrievable — and *reformulate the keywords* because they can't predict what BM25 indexed: - `exit-255` lesson: query `exit 255 unknown reason restart loop diagnosis` → write → query `exit 255 unknown SIGKILL containerd`. - `extension-version-lags`: query `extension build pinned version...pgvector` → write → query `pgvector postgres extension version compile error bump`. - `webfetch-readme`: write → write → query stuffed with the note's own section headings. ### semantic_retrieval → BM25 keyword-stuffing (3/3 mismatch) Agent knows the *meaning* but not the *indexed words*, so it brute-forces phrasings of one need against a lexical tool: - **4-step chain on one Go bug:** `bytes.Buffer Bytes Reset aliasing slice sharing` → `go buffer reuse map backing array bug` → (write) → `go map values all show same content after loop` → `bytes.Buffer Bytes returns same data every iteration`. Symptom and cause both known; no single lexical query bridges them. - Natural-language questions (`how to enable prometheus metrics on litellm proxy callback`, `moved compose stack to new directory volumes disappeared empty`) shoved into `brain_query`. ### synthesized_answer → fall back to / hedge with brain_query (2 mismatch + 1 partial) `brain_answer` is frequently **not trusted as terminal**: - **Strongest signal:** 3× `brain_answer` at 15:08 (compose volumes / pi rebuild time / litellm ModuleNotFound) → 3× `brain_query` at 15:09 on the *same three topics*. The agent asked the synthesizer, was unsatisfied, and immediately re-ran raw search. - Dex→Authentik: `brain_query` and `brain_answer` fired **8s apart on one question** — the agent pays both interfaces because it can't predict which returns usable knowledge. - (Counter-examples exist: `brain_answer` for episodic recall — "what happened with the litellm migration on 2026-05-16?" — was consumed and not re-queried. So `answer` works for *episodic/temporal* recall, fails for *how-do-I / does-X-hold* reasoning.) ### discover_capability + store-via-HTTP (3 mismatch) brain tools are **not ambient** — they are deferred and must be `ToolSearch`-loaded each session. Agents fumble the discovery (`ToolSearch 'brain knowledge memory'` → `'brain'` → `'brain ingestion knowledge wiki'`) and, when MCP load/auth fails, drop to **raw `curl` against `brain-mcp` / hand-built `/tmp/brain_entry.json`**. One agent even `ToolSearch`-ed an `authenticate` tool, then abandoned MCP for the HTTP bodge — matching the known "brain/op MCP auth lapses too often" footgun. ### Write-interface / layer schema confusion (cross-cutting) `brain_write` was called with **three different param shapes** in the same corpus: `{filename, type:"lesson", content}`, `{filename, content}` (no type), and `{wing:"tapir", hall:"failures", filename, content}` — plus `brain_ingest {source, content}`. Agents are unsure which verb and which layer (flat slug vs `wing`/`hall` knowledge routing vs raw ingest) a given knowledge-act maps to. This is the `knowledge/ vs wiki/` confusion expressed at the parameter level. ## 3. `intent_unclear` rate **0 / 46 genuine brain acts (0%).** Agent intent is unusually legible because these are Claude Code transcripts: the surrounding task, the query/filename strings, and the write content all disambiguate. One row (`a47`) was tagged `intent_unclear` and **excluded** — its `brain_` substring was inside a git commit message, not a brain call. Example of the only ambiguity that arose: a `brain_answer` on "gitea MCP not working... how to authenticate" followed by an unrelated query — can't tell if the answer satisfied or was abandoned (`partial`, row a28). ## 4. The single biggest intent↔interface gap **The brain offers one write verb and one lexical read verb, but autonomous agents perform four distinct knowledge-acts against them — and three of the four have no fitting interface.** The deepest gap is the **missing update/supersede path**: agents close every task by writing a lesson (the CLAUDE.md ritual), routinely improve it minutes-to-an-hour later, and — having no edit primitive — re-write the same slug blind, unable to tell whether they corrected the entry or forked a contradiction into the index. This compounds with the lexical-only read side: because there is no `get-by-id` or semantic retrieval, agents can't even reliably *find their own just-written note* to check it, so they keyword-stuff reformulated queries and hedge `brain_answer` with parallel `brain_query`. The interface is built for *append-and-keyword-search*; the agents are trying to *curate a living, deduplicated knowledge base*, and the seam between those two shows up as the 5 blind re-writes, 4 read-after-write re-queries, and 3 semantic-as-lexical chains that dominate the mismatch column. --- *Evidence-only per task scope — no redesign proposed.*