Drop-in unified findings doc for the brain-MCP intent study. Locks the shared row schema + closed intent vocab both columns must conform to. Agent column filled from agent-intent-column.jsonl (46 acts, 37% mismatch); human column left as PENDING cells + <<SYNTH>> blocks so the Claude.ai-history analysis merges in without re-deriving structure. Pre-seeds the cross-consumer divergence questions: agent mismatch is write-side-heavy (supersede + verify-landed); hypothesis is human mismatch is read-side-heavy (semantic + answer) — interface may fail the two consumers at opposite ends. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
7.1 KiB
Brain-MCP Intent↔Interface Findings — Unified (two-column merge)
Status — 2026-06-16
- ✅ Agent column filled from
agent-intent-column.jsonl(46 acts, koala). - ⏳ Human column =
PENDING. Drop the Claude.ai-history analysis intohuman-intent-column.jsonl(same dir, schema below), then fill thePENDINGcells and the synthesis blocks marked<<SYNTH>>. - ⚠️ Canonical
brain-intent-extraction.mdstill absent on koala. Vocab below is the reconstructed lock both columns must share. If the real file surfaces, re-mapintentlabels in both columns identically before merging.
Shared schema (LOCKED — both columns conform)
Per-call row, JSONL:
| field | values / form | notes |
|---|---|---|
id |
a01.. (agent) / h01.. (human) |
column prefix kept distinct |
session |
string | source session/conversation id |
ts |
ISO-8601 | best-effort |
tool |
brain tool name (+ (HTTP-curl) / (HTTP-staged) suffix for bodges) |
|
intent |
closed vocab ↓ | the knowledge-act WANTED |
intent_tool_match |
match | mismatch | partial |
does the called tool fit the want |
consumer_type |
autonomous_agent | human_interactive |
fixed per column |
workaround |
string | null | the bodge when mismatch — primary signal |
observed_friction |
string | null | reformulation chains, discovery tax, hedging |
evidence |
string | excerpt anchoring the classification |
schema_source |
reconstructed |
flip to canonical if real vocab lands |
Closed intent vocab (LOCKED)
semantic_retrieval, lexical_lookup, check_prior_art, synthesized_answer,
store_new_knowledge, update_or_supersede, ingest_raw_source,
verify_write_landed, discover_capability, intent_unclear.
Master comparison — by intent
| intent | agent acts | agent mismatch | human acts | human mismatch | shared gap |
|---|---|---|---|---|---|
| check_prior_art | 10 | 0% | PENDING |
PENDING |
— |
| store_new_knowledge | 15 | 7% (1/15) | PENDING |
PENDING |
<<SYNTH>> |
| update_or_supersede | 5 | 100% (5/5) | PENDING |
PENDING |
<<SYNTH>> no edit verb |
| synthesized_answer | 5 | 40% (2/5)+1 partial | PENDING |
PENDING |
<<SYNTH>> |
| verify_write_landed | 4 | 100% (4/4) | PENDING |
PENDING |
<<SYNTH>> no read-after-write |
| semantic_retrieval | 3 | 100% (3/3) | PENDING |
PENDING |
<<SYNTH>> lexical-only read |
| ingest_raw_source | 2 | 0% | PENDING |
PENDING |
— |
| discover_capability | 2 | 100% (2/2) | PENDING |
PENDING |
agent-specific (ToolSearch/auth)? |
| intent_unclear | 0 | — | PENDING |
PENDING |
divergence expected ↓ |
| TOTAL | 46 | 37% (17) | PENDING |
PENDING |
Per-intent merged findings
update_or_supersede — agent: 5/5 mismatch (highest value)
Agent: no edit/patch/append verb. Agents re-write same slug blind:
homelab-security-chains-not-bugs.md (+64m), tapir-rls-identity-bootstrapping,
webfetch-readme..., gitea-mcp-per-repo-tools-404...,
cannot-move-ingress-host... — 3 of 5 sub-30s ("wanted edit, got overwrite").
Cannot tell if write deduped or forked a contradiction.
Human: PENDING — look for: user editing a prior note, asking "update what I
saved about X", or expressing frustration that an old fact is stale/duplicated.
<> shared verdict once both filled.
verify_write_landed — agent: 4/4 mismatch
Agent: no get-by-id/read-after-write. Agents lexically re-query their own
fresh note with reformulated keywords (exit 255 unknown reason → ...SIGKILL containerd; extension build pinned... → pgvector ...compile error bump).
Human: PENDING — humans may not exhibit this (they trust the write UI
confirmation). If absent in human column, it's an agent-specific gap → flag.
<>.
semantic_retrieval — agent: 3/3 mismatch
Agent: meaning known, indexed words unknown → BM25 keyword-stuffing. 4-step
chain on one Go bytes.Buffer bug; NL questions shoved into brain_query.
Human: PENDING — humans likely hit this HARDER (they phrase conversationally).
Compare reformulation-chain length agent vs human.
<> — likely the strongest cross-consumer overlap.
synthesized_answer — agent: 2 mismatch + 1 partial
Agent: brain_answer not trusted terminal — 3 answers → 3 same-topic queries
1min later; query+answer fired 8s apart hedging one need. Works for episodic
recall, fails for how-do-I / does-X-hold.
Human: PENDING — humans may prefer brain_answer as primary (chat-native).
If human match-rate >> agent, the tool fits humans not agents → key divergence.
<>.
store_new_knowledge — agent: 14/15 match
Agent: healthy, except 1 HTTP-staged bodge when MCP auth lapsed. Also surfaced
write-schema confusion: 3 param shapes ({filename,type} / {filename} /
{wing,hall,filename}) + ingest{source}.
Human: PENDING — humans rarely write directly; expect low volume.
<>.
check_prior_art / ingest_raw_source — agent: 0% mismatch
Lexical fits named-entity recall and raw-source capture. Human: PENDING.
discover_capability — agent: 2/2 mismatch (agent-specific)
Brain tools deferred → ToolSearch-load each session; auth lapse → curl bodge.
Likely has NO human analog (humans get ambient connectors). Candidate for
"agent-only gap" bucket. Human: PENDING to confirm absent.
Cross-consumer divergence — questions to resolve at merge
- intent_unclear rate. Agent = 0% (transcripts self-document). Human expected higher (conversational, implicit). Big delta = the columns measure legibility differently, not just intent.
- Where does each consumer's mismatch concentrate? Agent mismatch is write-side-heavy (supersede + verify-landed = 9/17). Hypothesis: human mismatch is read-side-heavy (semantic + answer). If true → the interface fails the two consumers at opposite ends.
- Agent-only gaps (
discover_capability,verify_write_landed) vs shared gaps (semantic_retrieval,update_or_supersede). Shared gaps = highest-priority evidence; agent-only = harness/auth issues.
Combined headline — <<SYNTH>> (fill when human column lands)
Agent-side draft (to be reconciled with human-side): Brain = append + keyword-search; agents want a curated, dedup'd, self-verifying KB. Missing update/supersede path + lexical-only reads are the seam. Open question for the merge: do humans hit the same read-side wall, making semantic-retrieval the universal gap — or do agents uniquely suffer the write-side (supersede / verify) wall that humans sidestep via the chat UI?
Drop-in checklist (when human column arrives)
- Place
human-intent-column.jsonlin this dir; conform to LOCKED schema. - Fill every
PENDINGcell in master table + per-intent blocks. - Resolve the 3 divergence questions with evidence.
- Replace each
<<SYNTH>>with the reconciled verdict; write the combined headline. - If canonical vocab surfaced: re-map both columns'
intent, flipschema_source. - Commit as
docs(brain): merge human+agent intent columns.