docs(brain): add agent-consumer brain-MCP intent analysis
Agent-consumer column of the two-part brain intent↔interface study. Reconstructs the knowledge-act behind every brain MCP call in the Claude Code agent transcripts on koala (the only corpus with brain calls; brain/sessions and agentsquad eval logs carry none). 46 distinct knowledge-acts, 37% interface mismatch, 0% intent_unclear. Headline gap: no update/supersede verb → agents blind re-write same slug (5x); no read-after-write → lexical re-query of own note (4x); lexical-only reads → semantic-as-keyword-stuffing chains (3x); brain_answer hedged with parallel brain_query. Canonical schema brain-intent-extraction.md absent on host; vocab reconstructed, every row tagged schema_source=reconstructed. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -0,0 +1,148 @@
|
||||
# Agent-Consumer Brain Intent Analysis — koala column
|
||||
|
||||
**Consumer:** `autonomous_agent` (all rows). **Host:** koala. **Date:** 2026-06-15.
|
||||
**Raw rows:** `agent-intent-column.jsonl` (46 real knowledge-acts + 1 excluded false-positive).
|
||||
|
||||
## Caveat — canonical schema not on this host
|
||||
|
||||
The shared closed-vocabulary file `brain-intent-extraction.md` **does not exist on
|
||||
koala** — the only reference to it is inside *this task's own prompt*. The intent
|
||||
vocabulary below was **reconstructed** from the prompt's framing + `brain/schema.md`.
|
||||
Every row carries `schema_source: reconstructed`. If the canonical vocab differs,
|
||||
re-map the `intent` field; the `intent_tool_match` / `workaround` / `observed_friction`
|
||||
evidence stands regardless of label names.
|
||||
|
||||
**Reconstructed closed vocab:** `semantic_retrieval`, `lexical_lookup`,
|
||||
`check_prior_art`, `synthesized_answer`, `store_new_knowledge`, `update_or_supersede`,
|
||||
`ingest_raw_source`, `verify_write_landed`, `discover_capability`, `intent_unclear`.
|
||||
|
||||
## Corpus
|
||||
|
||||
- `~/.claude/projects/*/*.jsonl` — Claude Code agent transcripts (98 files). **The only
|
||||
source with brain calls.**
|
||||
- `brain/sessions/*.jsonl` — empty (only `.gitkeep`).
|
||||
- `agentsquad docs/eval/*.jsonl` — code-review eval results, **not** brain calls.
|
||||
- No separate Crush session logs on this host.
|
||||
- **Zero** brain calls were human-typed. Every brain act is agent-initiated (the
|
||||
CLAUDE.md "query as reflex / close-the-loop write" behaviour), so all qualify as
|
||||
`autonomous_agent`. The human gave the top-level task; the agent chose every brain act.
|
||||
|
||||
## 1. Intent histogram, split by `intent_tool_match`
|
||||
|
||||
| intent | match | mismatch | partial | total |
|
||||
|---|---|---|---|---|
|
||||
| check_prior_art | 10 | 0 | 0 | 10 |
|
||||
| store_new_knowledge | 14 | 1 | 0 | 15 |
|
||||
| update_or_supersede | 0 | 5 | 0 | 5 |
|
||||
| synthesized_answer | 2 | 2 | 1 | 5 |
|
||||
| verify_write_landed | 0 | 4 | 0 | 4 |
|
||||
| semantic_retrieval | 0 | 3 | 0 | 3 |
|
||||
| ingest_raw_source | 2 | 0 | 0 | 2 |
|
||||
| discover_capability | 0 | 2 | 0 | 2 |
|
||||
| **total** | **28** | **17** | **1** | **46** |
|
||||
|
||||
> Batch note: several rows collapse a same-second fan-out of identical-intent calls
|
||||
> (a10=3, a11=4, a14=5, a33=3, a34=3, a35=3). Call-level the corpus is ~62 brain calls;
|
||||
> the table counts the 46 distinct knowledge-acts. Frequency is deliberately *not* the
|
||||
> point — the mismatch column is.
|
||||
|
||||
**37% of agent knowledge-acts (17/46) are interface mismatches.** Every mismatch falls
|
||||
into one of four intents: `update_or_supersede`, `verify_write_landed`,
|
||||
`semantic_retrieval`, `discover_capability` — plus one `store` that had to use HTTP.
|
||||
|
||||
## 2. Mismatch list, grouped by intent (primary deliverable)
|
||||
|
||||
### update_or_supersede → re-write same slug (5/5 mismatch) — HIGHEST VALUE
|
||||
There is **no update / patch / append / supersede verb**. When an agent improves a note
|
||||
it already wrote, the only move is to call `brain_write`/`brain_ingest` **again with the
|
||||
same filename/source** and hope the index replaces rather than duplicates. Observed:
|
||||
|
||||
| slug | 1st write | 2nd write | gap | what changed |
|
||||
|---|---|---|---|---|
|
||||
| `tapir-rls-identity-bootstrapping` (ingest) | 14:00:59 | 14:02:08 | 69s | condensed body |
|
||||
| `homelab-security-chains-not-bugs.md` | 09:28:22 | 10:32:49 | 64m | new worked example |
|
||||
| `webfetch-readme-when-image-or-flag-uncertain.md` | 18:31:23 | 18:34:23 | 3m | near-identical (retry) |
|
||||
| `gitea-mcp-per-repo-tools-404-and-rest-fallback` | 05:57:37 | 05:57:55 | 18s | content tweak |
|
||||
| `cannot-move-ingress-host-across-namespaces-flux-dryrun` | 13:59:50 | 14:00:13 | 23s | content tweak |
|
||||
|
||||
Sub-30s gaps (3 of 5) read as "I wanted to edit but can only overwrite." The agent has
|
||||
no way to know whether the second write deduped or created a contradictory v2 — opacity
|
||||
the brain's own design principle (`mcp-tool-design-get-needs-list-partner.md`, written
|
||||
*by one of these very agents*) would flag: every `_write` needs a `_get`/`_update` partner.
|
||||
|
||||
### verify_write_landed → lexical re-query (4/4 mismatch)
|
||||
No read-after-write / get-by-id confirmation. After every substantive write, agents
|
||||
re-query lexically to check the note is retrievable — and *reformulate the keywords*
|
||||
because they can't predict what BM25 indexed:
|
||||
- `exit-255` lesson: query `exit 255 unknown reason restart loop diagnosis` → write →
|
||||
query `exit 255 unknown SIGKILL containerd`.
|
||||
- `extension-version-lags`: query `extension build pinned version...pgvector` → write →
|
||||
query `pgvector postgres extension version compile error bump`.
|
||||
- `webfetch-readme`: write → write → query stuffed with the note's own section headings.
|
||||
|
||||
### semantic_retrieval → BM25 keyword-stuffing (3/3 mismatch)
|
||||
Agent knows the *meaning* but not the *indexed words*, so it brute-forces phrasings of
|
||||
one need against a lexical tool:
|
||||
- **4-step chain on one Go bug:** `bytes.Buffer Bytes Reset aliasing slice sharing` →
|
||||
`go buffer reuse map backing array bug` → (write) → `go map values all show same
|
||||
content after loop` → `bytes.Buffer Bytes returns same data every iteration`. Symptom
|
||||
and cause both known; no single lexical query bridges them.
|
||||
- Natural-language questions (`how to enable prometheus metrics on litellm proxy
|
||||
callback`, `moved compose stack to new directory volumes disappeared empty`) shoved
|
||||
into `brain_query`.
|
||||
|
||||
### synthesized_answer → fall back to / hedge with brain_query (2 mismatch + 1 partial)
|
||||
`brain_answer` is frequently **not trusted as terminal**:
|
||||
- **Strongest signal:** 3× `brain_answer` at 15:08 (compose volumes / pi rebuild time /
|
||||
litellm ModuleNotFound) → 3× `brain_query` at 15:09 on the *same three topics*. The
|
||||
agent asked the synthesizer, was unsatisfied, and immediately re-ran raw search.
|
||||
- Dex→Authentik: `brain_query` and `brain_answer` fired **8s apart on one question** —
|
||||
the agent pays both interfaces because it can't predict which returns usable knowledge.
|
||||
- (Counter-examples exist: `brain_answer` for episodic recall — "what happened with the
|
||||
litellm migration on 2026-05-16?" — was consumed and not re-queried. So `answer`
|
||||
works for *episodic/temporal* recall, fails for *how-do-I / does-X-hold* reasoning.)
|
||||
|
||||
### discover_capability + store-via-HTTP (3 mismatch)
|
||||
brain tools are **not ambient** — they are deferred and must be `ToolSearch`-loaded each
|
||||
session. Agents fumble the discovery (`ToolSearch 'brain knowledge memory'` →
|
||||
`'brain'` → `'brain ingestion knowledge wiki'`) and, when MCP load/auth fails, drop to
|
||||
**raw `curl` against `brain-mcp` / hand-built `/tmp/brain_entry.json`**. One agent even
|
||||
`ToolSearch`-ed an `authenticate` tool, then abandoned MCP for the HTTP bodge — matching
|
||||
the known "brain/op MCP auth lapses too often" footgun.
|
||||
|
||||
### Write-interface / layer schema confusion (cross-cutting)
|
||||
`brain_write` was called with **three different param shapes** in the same corpus:
|
||||
`{filename, type:"lesson", content}`, `{filename, content}` (no type), and
|
||||
`{wing:"tapir", hall:"failures", filename, content}` — plus `brain_ingest {source,
|
||||
content}`. Agents are unsure which verb and which layer (flat slug vs `wing`/`hall`
|
||||
knowledge routing vs raw ingest) a given knowledge-act maps to. This is the
|
||||
`knowledge/ vs wiki/` confusion expressed at the parameter level.
|
||||
|
||||
## 3. `intent_unclear` rate
|
||||
|
||||
**0 / 46 genuine brain acts (0%).** Agent intent is unusually legible because these are
|
||||
Claude Code transcripts: the surrounding task, the query/filename strings, and the
|
||||
write content all disambiguate. One row (`a47`) was tagged `intent_unclear` and
|
||||
**excluded** — its `brain_` substring was inside a git commit message, not a brain call.
|
||||
Example of the only ambiguity that arose: a `brain_answer` on "gitea MCP not working...
|
||||
how to authenticate" followed by an unrelated query — can't tell if the answer satisfied
|
||||
or was abandoned (`partial`, row a28).
|
||||
|
||||
## 4. The single biggest intent↔interface gap
|
||||
|
||||
**The brain offers one write verb and one lexical read verb, but autonomous agents
|
||||
perform four distinct knowledge-acts against them — and three of the four have no fitting
|
||||
interface.** The deepest gap is the **missing update/supersede path**: agents close every
|
||||
task by writing a lesson (the CLAUDE.md ritual), routinely improve it minutes-to-an-hour
|
||||
later, and — having no edit primitive — re-write the same slug blind, unable to tell
|
||||
whether they corrected the entry or forked a contradiction into the index. This compounds
|
||||
with the lexical-only read side: because there is no `get-by-id` or semantic retrieval,
|
||||
agents can't even reliably *find their own just-written note* to check it, so they
|
||||
keyword-stuff reformulated queries and hedge `brain_answer` with parallel `brain_query`.
|
||||
The interface is built for *append-and-keyword-search*; the agents are trying to
|
||||
*curate a living, deduplicated knowledge base*, and the seam between those two shows up
|
||||
as the 5 blind re-writes, 4 read-after-write re-queries, and 3 semantic-as-lexical chains
|
||||
that dominate the mismatch column.
|
||||
|
||||
---
|
||||
*Evidence-only per task scope — no redesign proposed.*
|
||||
Reference in New Issue
Block a user