docs(brain): add two-column intent merge scaffold
CI / Lint / Test / Vet (push) Successful in 14s
CI / Mirror to GitHub (push) Successful in 4s

Drop-in unified findings doc for the brain-MCP intent study. Locks the
shared row schema + closed intent vocab both columns must conform to.
Agent column filled from agent-intent-column.jsonl (46 acts, 37%
mismatch); human column left as PENDING cells + <<SYNTH>> blocks so the
Claude.ai-history analysis merges in without re-deriving structure.

Pre-seeds the cross-consumer divergence questions: agent mismatch is
write-side-heavy (supersede + verify-landed); hypothesis is human
mismatch is read-side-heavy (semantic + answer) — interface may fail the
two consumers at opposite ends.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-06-16 20:45:24 +02:00
co-authored by Claude Opus 4.8
parent e8dbcf6eef
commit aa918388b9
@@ -0,0 +1,140 @@
# Brain-MCP Intent↔Interface Findings — Unified (two-column merge)
**Status — 2026-06-16**
-**Agent column** filled from `agent-intent-column.jsonl` (46 acts, koala).
-**Human column** = `PENDING`. Drop the Claude.ai-history analysis into
`human-intent-column.jsonl` (same dir, schema below), then fill the `PENDING`
cells and the synthesis blocks marked `<<SYNTH>>`.
- ⚠️ Canonical `brain-intent-extraction.md` still absent on koala. Vocab below is
the **reconstructed** lock both columns must share. If the real file surfaces,
re-map `intent` labels in *both* columns identically before merging.
---
## Shared schema (LOCKED — both columns conform)
Per-call row, JSONL:
| field | values / form | notes |
|---|---|---|
| `id` | `a01..` (agent) / `h01..` (human) | column prefix kept distinct |
| `session` | string | source session/conversation id |
| `ts` | ISO-8601 | best-effort |
| `tool` | brain tool name (+ `(HTTP-curl)` / `(HTTP-staged)` suffix for bodges) | |
| `intent` | closed vocab ↓ | the knowledge-act WANTED |
| `intent_tool_match` | `match` \| `mismatch` \| `partial` | does the called tool fit the want |
| `consumer_type` | `autonomous_agent` \| `human_interactive` | fixed per column |
| `workaround` | string \| null | the bodge when mismatch — **primary signal** |
| `observed_friction` | string \| null | reformulation chains, discovery tax, hedging |
| `evidence` | string | excerpt anchoring the classification |
| `schema_source` | `reconstructed` | flip to `canonical` if real vocab lands |
### Closed intent vocab (LOCKED)
`semantic_retrieval`, `lexical_lookup`, `check_prior_art`, `synthesized_answer`,
`store_new_knowledge`, `update_or_supersede`, `ingest_raw_source`,
`verify_write_landed`, `discover_capability`, `intent_unclear`.
---
## Master comparison — by intent
| intent | agent acts | agent mismatch | human acts | human mismatch | shared gap |
|---|---|---|---|---|---|
| check_prior_art | 10 | 0% | `PENDING` | `PENDING` | — |
| store_new_knowledge | 15 | 7% (1/15) | `PENDING` | `PENDING` | `<<SYNTH>>` |
| update_or_supersede | 5 | **100%** (5/5) | `PENDING` | `PENDING` | `<<SYNTH>>` no edit verb |
| synthesized_answer | 5 | 40% (2/5)+1 partial | `PENDING` | `PENDING` | `<<SYNTH>>` |
| verify_write_landed | 4 | **100%** (4/4) | `PENDING` | `PENDING` | `<<SYNTH>>` no read-after-write |
| semantic_retrieval | 3 | **100%** (3/3) | `PENDING` | `PENDING` | `<<SYNTH>>` lexical-only read |
| ingest_raw_source | 2 | 0% | `PENDING` | `PENDING` | — |
| discover_capability | 2 | **100%** (2/2) | `PENDING` | `PENDING` | agent-specific (ToolSearch/auth)? |
| intent_unclear | 0 | — | `PENDING` | `PENDING` | divergence expected ↓ |
| **TOTAL** | **46** | **37% (17)** | `PENDING` | `PENDING` | |
---
## Per-intent merged findings
### update_or_supersede — agent: 5/5 mismatch (highest value)
**Agent:** no edit/patch/append verb. Agents re-write same slug blind:
`homelab-security-chains-not-bugs.md` (+64m), `tapir-rls-identity-bootstrapping`,
`webfetch-readme...`, `gitea-mcp-per-repo-tools-404...`,
`cannot-move-ingress-host...` — 3 of 5 sub-30s ("wanted edit, got overwrite").
Cannot tell if write deduped or forked a contradiction.
**Human:** `PENDING` — *look for: user editing a prior note, asking "update what I
saved about X", or expressing frustration that an old fact is stale/duplicated.*
**<<SYNTH>>** shared verdict once both filled.
### verify_write_landed — agent: 4/4 mismatch
**Agent:** no `get-by-id`/read-after-write. Agents lexically re-query their own
fresh note with reformulated keywords (`exit 255 unknown reason``...SIGKILL
containerd`; `extension build pinned...``pgvector ...compile error bump`).
**Human:** `PENDING` — *humans may not exhibit this (they trust the write UI
confirmation). If absent in human column, it's an agent-specific gap → flag.*
**<<SYNTH>>**.
### semantic_retrieval — agent: 3/3 mismatch
**Agent:** meaning known, indexed words unknown → BM25 keyword-stuffing. 4-step
chain on one Go `bytes.Buffer` bug; NL questions shoved into `brain_query`.
**Human:** `PENDING` — *humans likely hit this HARDER (they phrase conversationally).
Compare reformulation-chain length agent vs human.*
**<<SYNTH>>** — likely the strongest cross-consumer overlap.
### synthesized_answer — agent: 2 mismatch + 1 partial
**Agent:** `brain_answer` not trusted terminal — 3 answers → 3 same-topic queries
1min later; query+answer fired 8s apart hedging one need. Works for *episodic*
recall, fails for *how-do-I / does-X-hold*.
**Human:** `PENDING` — *humans may prefer `brain_answer` as primary (chat-native).
If human match-rate >> agent, the tool fits humans not agents → key divergence.*
**<<SYNTH>>**.
### store_new_knowledge — agent: 14/15 match
**Agent:** healthy, except 1 HTTP-staged bodge when MCP auth lapsed. Also surfaced
write-schema confusion: 3 param shapes (`{filename,type}` / `{filename}` /
`{wing,hall,filename}`) + `ingest{source}`.
**Human:** `PENDING`*humans rarely write directly; expect low volume.*
**<<SYNTH>>**.
### check_prior_art / ingest_raw_source — agent: 0% mismatch
Lexical fits named-entity recall and raw-source capture. **Human:** `PENDING`.
### discover_capability — agent: 2/2 mismatch (agent-specific)
Brain tools deferred → `ToolSearch`-load each session; auth lapse → `curl` bodge.
**Likely has NO human analog** (humans get ambient connectors). Candidate for
"agent-only gap" bucket. **Human:** `PENDING` to confirm absent.
---
## Cross-consumer divergence — questions to resolve at merge
1. **intent_unclear rate.** Agent = 0% (transcripts self-document). Human expected
higher (conversational, implicit). Big delta = the columns measure legibility
differently, not just intent.
2. **Where does each consumer's mismatch concentrate?** Agent mismatch is
write-side-heavy (supersede + verify-landed = 9/17). Hypothesis: human mismatch
is read-side-heavy (semantic + answer). If true → **the interface fails the two
consumers at opposite ends.**
3. **Agent-only gaps** (`discover_capability`, `verify_write_landed`) vs
**shared gaps** (`semantic_retrieval`, `update_or_supersede`). Shared gaps =
highest-priority evidence; agent-only = harness/auth issues.
---
## Combined headline — `<<SYNTH>>` (fill when human column lands)
> Agent-side draft (to be reconciled with human-side):
> Brain = append + keyword-search; agents want a curated, dedup'd, self-verifying KB.
> Missing update/supersede path + lexical-only reads are the seam. **Open question
> for the merge: do humans hit the same read-side wall, making semantic-retrieval the
> universal gap — or do agents uniquely suffer the write-side (supersede / verify)
> wall that humans sidestep via the chat UI?**
---
## Drop-in checklist (when human column arrives)
1. Place `human-intent-column.jsonl` in this dir; conform to LOCKED schema.
2. Fill every `PENDING` cell in master table + per-intent blocks.
3. Resolve the 3 divergence questions with evidence.
4. Replace each `<<SYNTH>>` with the reconciled verdict; write the combined headline.
5. If canonical vocab surfaced: re-map both columns' `intent`, flip `schema_source`.
6. Commit as `docs(brain): merge human+agent intent columns`.