Compare commits

..
Author SHA1 Message Date
mathiasandClaude Opus 4.8 6606b38a76 feat(gitea): IssueTracker client + inject into brain server (#52)
Implements the IssueTracker port as a real Gitea REST client (#49c) — the
new outbound dependency the brain server gains for capture.

- CreateIssue / CommentIssue / CloseIssue(+optional closing comment) over
  the Gitea API. Owner is the const "mathias", never caller-supplied, so
  a caller cannot redirect a write to another owner's repo.
- Token read once at construction (BRAIN_GITEA_TOKEN), held in the struct,
  travels only in the Authorization header — never logged or in argv.
  Error messages carry status + truncated body, never the token
  (regression-tested). gitea.New returns nil when URL or token is unset,
  so missing config = tracker disabled via one nil check.
- Injected into the MCP server behind the capture.IssueTracker interface
  via WithIssueTracker (constructor injection, swappable/testable); main
  wires it from BRAIN_GITEA_URL (default https://git.d-ma.be) +
  BRAIN_GITEA_TOKEN. Consumed by the capture use-case in #53.

Tests use httptest transports: create (owner+auth header asserted),
comment, close with/without comment, error path that proves the token
never leaks into an error string.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-22 23:27:29 +02:00
mathiasandClaude Opus 4.8 4cfc98de56 refactor(capture): CloseIssue carries a closing comment (#52)
#52's IssueTracker spec is CloseIssue(repo, number, comment). Refine the
#51 port signature to match and have the service pass the ticket body as
the closing comment (empty ⇒ close only). Keeps the close-with-comment
flow first-class rather than forcing two separate ticket items.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-22 23:27:29 +02:00
mathiasandClaude Opus 4.8 0ac165cca3 feat(brainstore): shared BrainStore impl; re-point MCP handlers (#51)
Extracts the #45 write/update/get logic + the wiki upkeep that must
accompany a write (wing _index rebuild, cross-wing auto-tunnel, graph
re-index) into a single concrete brainstore.Store implementing
capture.BrainStore. The MCP brain_write/brain_update/brain_get handlers
are re-pointed at it, so there is one implementation, not two — the DRY
payoff #51 is named for. capture and MCP now share the exact same brain
write path and read-after-write contract.

The Server gains a *brainstore.Store, constructed in NewServer and given
the graph store in WithGraph. Embedding refresh stays out-of-band
(mtime-driven vectorstore.Sync), unchanged. Existing MCP brain_update/
brain_get/brain_write tests pass unmodified — behaviour and the
{id, path, content_hash} response contract are preserved.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-22 23:21:43 +02:00
mathiasandClaude Opus 4.8 43f92e3102 feat(capture): CaptureService use-case + ports + entities (#51)
The Clean-Architecture core of the capture capability (#49b). Pure
orchestration over ports — no HTTP, no live Gitea, no audit I/O — fully
unit-tested against fakes before any adapter exists.

- Ports: BrainStore (#45 write/update/get), IssueTracker, SummaryWriter,
  ClassificationPolicy (satisfied by #50's classification.Config),
  AuditSink. Entities: Insight, Ticket, Summary, CaptureContext,
  CaptureInput, CaptureReceipt.
- CaptureService.Capture: validate-before-write (fail-closed), resolve
  effective classification (stricter of declared vs target-derived;
  under-declaration logged as a security event), orchestrate insights
  (write/supersede) → tickets → summary best-effort, emit a request-level
  audit record of exactly what landed, return a partial-aware receipt.
- dry_run short-circuits after validation, writes nothing (not even audit).

Out of scope here, layered on later: the I1 origin sovereignty gate (#53,
needs the server-derived principal) and the classification-aware audit
degradation/refusal (#54). "Effective" is folded into the service as
classification.Stricter rather than a port method — the stricter-wins
rule is use-case policy.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-22 23:21:43 +02:00
mathiasandClaude Opus 4.8 98cfae595c feat(classification): data-sensitivity taxonomy + per-wing/repo tags (#50)
CI / Lint / Test / Vet (pull_request) Successful in 13s
CI / Mirror to GitHub (pull_request) Has been skipped
Implements the I1-prerequisite from #49/#50: the classification taxonomy
and the per-wing / per-repo tagging the capture server reads to derive a
target's sensitivity.

- Levels public < internal < confidential, ordered so "stricter wins"
  (spec §4.1 model C) is a plain max via Stricter().
- Tags read from an optional classification.yaml at the brain root
  (wings:/repos: maps). Absent file → defaults-only, not an error.
- Defaulting: client-* → confidential; hyperguild/homelab → internal;
  everything else → confidential. Fail-safe-to-strictest is the
  load-bearing property: a missing tag never silently downgrades.
- Config.Derive(Target) is the function the use-case calls; Wing/Repo
  are the per-kind helpers. ParseLevel rejects unknown tokens; a bad
  level in the config file is a hard load error.

Central classification.yaml (not _index.md frontmatter, not gitea repo
topics): classifying a repo needs no live Gitea client, so #50 has no
dependency on the tracker work (#52); it's auditable in one place; and
it avoids BuildWingIndex clobbering a wing's regenerated _index.md.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-22 22:42:05 +02:00
mathias 2a595b5a92 docs(capture): make Q4 audit-sink-down posture classification-aware (#49)
CI / Lint / Test / Vet (push) Successful in 12s
CI / Mirror to GitHub (push) Successful in 3s
Reconsidered Q4: instead of one global degrade-and-warn, the posture now
inherits from effective classification (Q1):
- confidential + audit-sink-down -> hard-refuse (no buffer; removes the
  buffer-integrity question for confidential data)
- internal/public + audit-sink-down -> degrade-and-warn + durable local
  buffer + ntfy + reconcile-on-recovery
- floor (all tiers): refuse if nothing can record the audit
Updated the I5 Gherkin scenarios + obligations row to match. Couples Q4
to the Q1 classification spine -> one coherent sensitivity model.
2026-06-22 20:20:40 +00:00
mathias b7a2cc5fdf docs(capture): resolve §4 open questions into binding decisions (#49)
CI / Lint / Test / Vet (push) Successful in 14s
CI / Mirror to GitHub (push) Successful in 4s
Q1 classification trust: model (C) — caller declares, server cross-checks
  target tag, stricter wins, mismatch logged. Needs a classification
  taxonomy + per-wing/repo tags (prerequisite, sub-task of #49).
Q2 harness origin: server-derived from authenticated principal;
  context.harness is descriptive-only, never a gate input.
Q3 central relay: ships in v1 (needed for claude.ai/Crush/Pi/LLM Council)
  + I2 security-baseline ledger entry is v1 work.
Q4 audit-sink-down: degrade-and-warn + durable local buffer + reconcile
  on recovery; refuse only if NOTHING can record the audit.

Updated the I1 sovereignty + I5 auditability Gherkin scenarios to match;
added classification-mismatch and origin-spoofing scenarios.
2026-06-22 20:17:07 +00:00
mathias db638cca11 docs(capture): add use-case + BDD spec for the capture capability (#49)
CI / Lint / Test / Vet (push) Successful in 14s
CI / Mirror to GitHub (push) Successful in 3s
Pre-build spec for the uniform cross-harness capture path. Grounds the
design in the homelab invariants (I1-I5, now canonical on infra main):
- I1 sovereignty gate: refuse confidential capture via us-nexus harness
- I2: distributed-library form (no high-degree observer node) is
  admissible; central relay needs a security-baseline ledger entry
- I5: capture is a privileged write path, must emit audit records
Gherkin scenarios cover the happy path, the sovereignty refusal,
supersession + staleness discipline (#45/#47), fail-closed validation,
best-effort partial-failure receipts, dry-run, and fidelity supersession.
Four open questions flagged for review (classification trust is the
highest-risk one).
2026-06-22 17:25:03 +00:00
mathias 38579598e0 feat(skills): add close-session workflow skill
CI / Lint / Test / Vet (push) Successful in 14s
CI / Mirror to GitHub (push) Successful in 4s
Disciplined end-of-session closeout for Claude.ai chats: harvest →
ground-truth gitea → confirm issue actions → commit session summary →
brain orientation note → safe-to-archive verdict.

Finalized against current infra (git.d-ma.be, koala:30401 LiteLLM,
Authentik) and the brain_update/brain_get verbs. Phase 5 uses
supersede-by-slug + brain_get read-after-write with the batch
no-semantic-query-after-supersede discipline. session_log/brain_tunnel
inlined (now callable from claude.ai). Summary frontmatter is the
reduced live-capture schema with fidelity:live-capture to distinguish
from batch-export summaries.
2026-06-22 15:57:02 +00:00
mathias d6fa92b176 Merge pull request 'chore(context): drop unsupported cursor + aider adapters' (#48) from chore/drop-cursor-aider-adapters into main
CI / Lint / Test / Vet (push) Successful in 12s
CI / Mirror to GitHub (push) Successful in 4s
2026-06-22 09:03:59 +00:00
mathias 9173f9058d chore(context): remove generated .aider.conventions.md
CI / Lint / Test / Vet (pull_request) Successful in 13s
CI / Mirror to GitHub (pull_request) Has been skipped
Aider is not a supported harness; the generator no longer emits this.
Note: generator no longer writes .aider.conf.yml either.
2026-06-22 09:01:39 +00:00
mathias 3e84a41fed chore(context): remove generated .cursorrules
Cursor is not a supported harness; the generator no longer emits this.
2026-06-22 09:01:29 +00:00
mathias 0e0571c7da chore(context): stop policing cursor/aider adapters in task check
Drop the context:sync:cursor task and remove .cursorrules +
.aider.conventions.md from the drift-guard file list in `check`, now
that the generator no longer emits them.
2026-06-22 09:01:19 +00:00
mathias 7a27cf71a2 chore(context): drop cursor + aider adapters from generator
Cursor and Aider are not supported harnesses. Remove generate_cursor
and generate_aider, their no-arg calls, and their case arms. The active
adapters are claude (CLAUDE.md), agents (AGENTS.md), and system-prompt
(.context/system-prompt.txt).
2026-06-22 09:00:55 +00:00
mathias 63df6d3283 Merge pull request 'feat: brain_update supersede verb + brain_get / brain_write read-after-write handle (#45)' (#46) from feat/brain-update-supersede into main
CI / Lint / Test / Vet (push) Successful in 13s
CI / Mirror to GitHub (push) Successful in 4s
2026-06-22 08:59:28 +00:00
mathiasandClaude Opus 4.8 f04b03e07e chore(context): re-sync derived adapters after root rule-0 update
CI / Lint / Test / Vet (pull_request) Successful in 13s
CI / Mirror to GitHub (pull_request) Has been skipped
context-sync regenerated the adapters from the updated root AGENT.md
(rule 0 pre-task ritual + TDD constraint). The committed adapters had
drifted; this is the documented `task check` remedy, not a content
change in this repo.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-22 08:25:16 +02:00
mathiasandClaude Opus 4.8 6c61f93146 feat(mcp): register brain_update + brain_get, extend brain_write handle
Wires the #45 verbs into the MCP surface (all three sites: tools()
descriptors, handleCall dispatch, package doc comment).

- brain_update: supersede-by-slug or full path; rebuilds wing _index and
  re-tunnels cross-wing matches against the new body (idempotent,
  best-effort), re-indexes the graph, returns {id, path, content_hash,
  superseded}.
- brain_get: fetch by id or path (both are the brain-relative handle);
  returns {id, path, content_hash, frontmatter, body}.
- brain_write: return contract extended from {path} to {id, path,
  content_hash} — path kept for backward compat — so the create path
  also yields a stable handle.

id == relPath; content_hash == sha256 of the file bytes. Tests cover the
supersede happy path, missing-target error + no-create, get by id/path,
write handle, and an end-to-end re-embed test that drives the real
vectorstore.Sync re-index after an update.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-22 08:25:10 +02:00
mathiasandClaude Opus 4.8 95a69fc2c1 feat(brain): add UpdateNote/ReadNote supersede primitives + frontmatter editor
Implements the api-layer half of #45. UpdateNote supersedes a note in
place (whole-note body replace, frontmatter re-stamp: updated_at,
supersedes=prior content hash, supersede_reason), preserving created_at,
wing, hall, and any custom fields. Never creates — a missing target is
an error so callers fall back to brain_write. ReadNote is the read-after-
write primitive (frontmatter + body + content_hash). ContentHash is the
sha256 handle that round-trips write/update → get.

Frontmatter is edited via a line-preserving ordered editor rather than a
yaml.v3 round-trip, which would reorder keys and strip comments — the
brain writes flat key:value frontmatter by hand.

Embeddings are not refreshed here: the rewritten file's mtime advances,
which the out-of-band vectorstore.Sync ticker uses to re-embed it — the
same mechanism brain_write relies on.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-22 08:25:00 +02:00
mathiasandClaude Opus 4.8 bb8bc0478c chore(cd): use git.d-ma.be for SSH alias + INFRA_REPO (post-rename consistency)
CI / Lint / Test / Vet (push) Successful in 13s
CI / Mirror to GitHub (push) Successful in 4s
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-17 00:19:42 +02:00
mathiasandClaude Opus 4.8 a961a3c064 fix(vectorstore): hard-split oversized heading-less chunks
A doc with no headings and no blank-line paragraphs (JSON-lines, e.g.
wiki/telos/decisions/human-intent-column.md) survived both chunk passes whole
and was sent to nomic-embed over its context window → 'input length exceeds
the context length' (400, the steady embed errors=1). Add a final hard-split
pass (line then UTF-8 rune boundaries) so no chunk exceeds maxBytes. TDD.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-17 00:19:42 +02:00
mathiasandClaude Opus 4.8 bec28f9014 fix(cd): point registry/patch/verify refs at git.d-ma.be after host rename
CI / Lint / Test / Vet (push) Successful in 13s
CI / Mirror to GitHub (push) Successful in 3s
The gitea.d-ma.be→git.d-ma.be rename moved the infra deployment manifests, but
cd.yml still sed-patched 'gitea.d-ma.be/mathias/ingestion:' — which no longer
matches, so the patch produced no change and 'git commit' failed under set -e,
wedging all deploys (image stuck at e8dbcf6). Repoint the registry image refs,
the infra-patch sed patterns, and the rollout-verify EXPECTED values to
git.d-ma.be (same registry backend). SSH alias + INFRA_REPO left as-is.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-16 23:22:07 +02:00
mathiasandClaude Opus 4.8 b62ac57382 fix(brain_answer): reranker is a filter, not a gate — fall back to BM25 when it keeps nothing
CI / Lint / Test / Vet (push) Successful in 13s
CI / Mirror to GitHub (push) Successful in 3s
The Qwen3-Reranker is a web-search cross-encoder. Against conversational /
personal-intent queries (e.g. 'what am I optimizing toward?') it scores every
candidate as 'no', so brain_answer collapsed to 'No relevant content found'
even though BM25 had retrieved on-topic notes (incl. the telos wing). Treat
the reranker as a filter: when it keeps zero results, fall back to the
BM25/vector ordering (capped to the no-reranker depth of 10). Refs brain#11.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-16 22:53:02 +02:00
mathiasandClaude Opus 4.8 aa918388b9 docs(brain): add two-column intent merge scaffold
CI / Lint / Test / Vet (push) Successful in 14s
CI / Mirror to GitHub (push) Successful in 4s
Drop-in unified findings doc for the brain-MCP intent study. Locks the
shared row schema + closed intent vocab both columns must conform to.
Agent column filled from agent-intent-column.jsonl (46 acts, 37%
mismatch); human column left as PENDING cells + <<SYNTH>> blocks so the
Claude.ai-history analysis merges in without re-deriving structure.

Pre-seeds the cross-consumer divergence questions: agent mismatch is
write-side-heavy (supersede + verify-landed); hypothesis is human
mismatch is read-side-heavy (semantic + answer) — interface may fail the
two consumers at opposite ends.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-16 20:45:24 +02:00
mathiasandClaude Opus 4.8 e8dbcf6eef docs(brain): add agent-consumer brain-MCP intent analysis
CI / Lint / Test / Vet (push) Successful in 14s
CI / Mirror to GitHub (push) Successful in 4s
Agent-consumer column of the two-part brain intent↔interface study.
Reconstructs the knowledge-act behind every brain MCP call in the
Claude Code agent transcripts on koala (the only corpus with brain
calls; brain/sessions and agentsquad eval logs carry none).

46 distinct knowledge-acts, 37% interface mismatch, 0% intent_unclear.
Headline gap: no update/supersede verb → agents blind re-write same
slug (5x); no read-after-write → lexical re-query of own note (4x);
lexical-only reads → semantic-as-keyword-stuffing chains (3x);
brain_answer hedged with parallel brain_query.

Canonical schema brain-intent-extraction.md absent on host; vocab
reconstructed, every row tagged schema_source=reconstructed.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-15 19:44:59 +02:00
mathiasandClaude Opus 4.8 0eeb1df4a2 fix(claudewatcher): scrub bare 1Password service-account tokens (ops_)
CI / Lint / Test / Vet (push) Successful in 17s
CI / Mirror to GitHub (push) Successful in 4s
A ~/.zshrc read surfaced OP_SERVICE_ACCOUNT_TOKEN into a transcript. The
_TOKEN= form was already caught by homelab-env-token, but a bare ops_<b64>
value was not. Add an op-service-account rule (ordered early). Tests cover
both env-assigned and bare forms.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 00:28:47 +02:00
mathiasandClaude Opus 4.8 9febb1bba1 fix(claudewatcher): harden secret scrubber against \b evasion + add JWT
CI / Lint / Test / Vet (push) Successful in 17s
CI / Mirror to GitHub (push) Successful in 4s
A shell mangle that glued a key to a preceding word ('yes'+'sk-...') had
no word boundary, so the leading \b in the openai-sk rule failed to match
and a LiteLLM master key leaked past the scrubber into ingest (2026-06-11).

- openai-sk: drop leading \b, match sk- shape anywhere ({32,} floor keeps
  short task-/disk- words clean).
- add jwt rule for bare header.payload.sig tokens (no Bearer prefix).
- regression tests: the exact yessk- evasion, standalone sk-, bare JWT,
  plus clean-content guards for task-/disk-.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-11 08:55:33 +02:00
mathias 5dc247b994 fix(ci): quote "on" key so Gitea parses workflow triggers
CI / Lint / Test / Vet (push) Successful in 17s
CI / Mirror to GitHub (push) Successful in 4s
Bare on: parses as YAML boolean true (Norway problem); Gitea then ignores the triggers and silently skips jobs. Quoting forces the string key.
2026-06-03 08:38:46 +02:00
mathias 2125558196 docs: extend harness boundary decision to cover Crush as third harness
CI / Mirror to GitHub (push) Successful in 4s
CI / Lint / Test / Vet (push) Successful in 12s
2026-05-28 11:38:36 +00:00
mathias 2beaac2feb docs: add 2026-05-28 decisions — harness boundary + field benchmark definition
CI / Lint / Test / Vet (push) Successful in 12s
CI / Mirror to GitHub (push) Successful in 3s
2026-05-28 10:31:41 +00:00
mathias 525811bc1a docs: add hypothesis statement and harness boundary clarification
CI / Lint / Test / Vet (push) Successful in 13s
CI / Mirror to GitHub (push) Successful in 4s
2026-05-28 10:30:39 +00:00
mathias bad0581623 merge: client-name scrubber rule (refs hyperguild#27)
CI / Lint / Test / Vet (push) Successful in 13s
CI / Mirror to GitHub (push) Successful in 4s
2026-05-26 07:10:05 +02:00
mathias a94b860c2e feat(claudewatcher): client-name guard via RegisterRule + env
Pre-rollout guard. Source code stays clean — client identities come
from CLAUDE_INGEST_CLIENT_BLOCK env (sourced from a SOPS-encrypted k8s
secret in infra repo). Env value is a regex alternation; main wraps
it with `(?i)\b(...)\b` so word-boundary matching avoids false hits
inside longer identifiers (e.g. "Sebastian" doesn't trigger on "SEB").

DefaultRules (credential shapes) still take precedence so any leak
that's BOTH a client mention AND a credential shape logs as the
credential — strictly more dangerous, points triage at the right
thing. Tests cover precedence + case variations + word-boundary
respect + invalid-pattern rejection.

Refs: infra#73 Track E.1 pre-rollout grill (option B).

Bump-Type: minor
2026-05-26 07:10:05 +02:00
40 changed files with 3662 additions and 770 deletions
-315
View File
@@ -1,315 +0,0 @@
# Agent context — Mathias workspace
<!-- Canonical root context for all AI coding agents.
Lives at: ~/dev/.context/AGENT.md
Applies to every project under ~/dev/ unless overridden.
Run `task context:sync` from ~/dev/ to regenerate harness-specific files.
Project-level context in .context/PROJECT.md layers on top of this. -->
## Who I am
I'm Mathias, a digital product manager and technology consultant based in Sweden.
I build software, research emerging tech, and deliver consulting engagements
for clients under NDA. I work across AI/ML, financial automation, web applications,
and climate/sustainability tech.
## How I work with agents
- I think like a product manager — I care about *why* before *how*
- I want agents to be opinionated and push back, not just execute blindly
- I prefer concise responses; skip ceremony and get to the point
- When I say "build this", I mean production-quality with tests, not a demo
- Ask me before making irreversible changes or adding heavy dependencies
- I work with confidential client data — never send it to cloud APIs unless I explicitly say it's OK
## Behavior rules
These rules apply to every task across every project, regardless of harness.
1. **No assumptions.** Don't hide confusion — surface it. Surface tradeoffs explicitly.
Think before coding; if the problem is unclear, ask or state assumptions before acting.
2. **Minimum viable code.** Solve with the smallest change that works. Nothing
speculative, no "while we're here" cleanups, no premature abstractions. Simplicity first.
3. **Surgical changes.** Touch only what the task requires. Leave unrelated code,
files, and formatting alone. Diffs should be small and reviewable.
4. **Goal-driven execution.** Define clear success criteria up front for every task.
Loop — implement, verify, refine — until those criteria are met. Don't claim
completion without evidence (tests pass, command output, observed behavior).
5. **Trunk-Based Development — commit directly to main.** Every commit is one
logical change (one tool, one fix, one test) with passing tests. Main is always
deployable. Never create long-lived feature branches.
**Exception — parallel agents on same repo:** If another agent is known to be
actively working on the same repo simultaneously, create a short-lived branch
(`agent/<description>`), finish the task, and merge to main within the same
session. Do not leave agent branches open between sessions.
**Exception — external contributor or client four-eyes requirement:** Use
PR flow only when a human reviewer outside the project is required. Document
the reason in PROJECT.md.
## Default stack
| Layer | Default | Fallback | Last resort |
|-------|---------|----------|-------------|
| Language | Go | Python | TypeScript, Java, C |
| UI | HTMX + Templ | Server-rendered HTML | React (only if SPA is justified) |
| Build | Task (taskfile.dev) | Make | — |
| Containers | Docker Compose (dev), k3s (prod) | — | — |
| DB | PostgreSQL + sqlc | SQLite | — |
| Search | pgvector (vector), BM25 | Qdrant (when >1M vectors or hybrid retrieval) | — |
| Logging | slog (structured) | — | — |
| Testing | Table-driven, testify | — | — |
| Agents (Go) | google.golang.org/adk + pkg/litellm adapter | — | — |
Exploratory: Rust, Zig — I'll tell you when I want these.
## Code conventions
- **Go style**: golines, gofumpt, golangci-lint
- **Errors**: `fmt.Errorf("operation: %w", err)` — never naked, never log-and-return
- **Naming**: stdlib conventions, no stuttering
- **Architecture**: prefer stdlib over frameworks, constructor injection, env-var config parsed into typed structs
- **Git**: conventional commits (`feat:`, `fix:`, `chore:`), commit directly to main,
one logical change per commit, CI is the quality gate
- **Never**: long-lived feature branches, PRs for solo work, direct push without
passing `task check` locally first
- **Security**: no secrets in code, govulncheck before adding deps, SOPS for encrypted config
- **Dependencies**: prefer stdlib. testify, slog, templ, sqlc, google.golang.org/adk (agent projects only) are pre-approved; anything else needs justification in the commit message
## Infrastructure
Three machines on Tailscale:
| Machine | Role | Key specs |
|---------|------|-----------|
| koala | GPU inference, heavy compute | RTX 5070, runs k3s + llama-swap + shared postgres18/pgvector |
| iguana | Services, builds | M2 Ultra Mac |
| flamingo | Daily driver, edge | Mac mini, ~/dev is here |
- **Model routing**: LiteLLM in front of llama-swap (local) + cloud APIs (when permitted)
- **Orchestration**: k3s cluster across all three machines
- **Networking**: Tailscale mesh
## Project landscape
All development repos live at `~/dev/` (softlink from `~/Documents/local-dev/`).
Organized in thematic folders:
| Folder | Focus | Count |
|--------|-------|-------|
| `GO/` | Go web frameworks, API integrations, learning projects | ~10 |
| `AI/` | ML research, AI frameworks (FinRL, DSPy, crawl4ai) | ~6 |
| `AGENTS/` | Autonomous agents, coding agents, MCP servers, infra | ~15 |
| `QKX/` | Invoice processing, financial automation, payment systems | ~13 |
| `XT/` | Climate data, sustainability (Klimatkollen, Garbo) | ~2 |
See `~/dev/PROJECT_SUMMARY.md` for detailed descriptions of each project.
### Key active projects
- **super-koala** (`AGENTS/`) — multi-component agent stack with LangGraph, DSPy, MCP
- **azure-tiger** (`QKX/`) — invoice extraction → ISO 20022 payment instructions
- **gocrwl** (`AGENTS/`) — Go web crawler with containerized deployment
- **koala-ai-stack** (`AGENTS/`) — local AI server infrastructure management
- **klimatkollen** (`XT/`) — Swedish municipal climate data platform
## Knowledge base — actively use it
A persistent brain (BM25 search + LLM-synthesised Q&A) survives across sessions,
hosts, and harnesses. It holds 100+ hard-won entries: infra incident postmortems,
Go pitfalls, framework gotchas, design principles, ADRs. **It is not optional
reference material — query it actively, not just when explicitly told.**
### When to query (treat as a reflex)
- **Before** starting a non-trivial task — search for prior art with the symptom
AND the system component ("how did we solve X in Y?"). 5 seconds beats 5 hours.
- **When debugging** — search for the error string, the stack frame, the affected
service. Past you may have already paid this tax.
- **Before adopting** a pattern, library, framework, or model name — check if it
was tried and rejected, or what the integration footguns are.
- **When making architectural decisions** — search for the domain + "ADR" or
"decision" to find prior reasoning before re-deriving it.
- **When a recommendation feels novel** — challenge yourself: "has this been
documented?" The brain often has it.
### When to write
After you discover something that **future-you would forget** and that **isn't
recoverable from the code, git log, or PR description alone**:
- Bugs whose root cause is non-obvious and generalisable beyond this project.
- Framework / library / model-name quirks that bit you and would bite anyone.
- Design principles validated under fire (e.g. "every `_get` needs a `_list`").
- Postmortems for incidents: what broke, why, how diagnosed, what to do next time.
DON'T write project status, sprint progress, PR summaries, or "what I did this
session" — those rot fast and the originals are in git/gitea anyway. Brain
entries that age well are about *why*, *how to avoid*, and *what to do when*.
### How to access (per harness)
| Harness | Query | Write |
|---------|-------|-------|
| **Claude Code, Claude Desktop** | `brain_query` (BM25), `brain_answer` (LLM-synth + sources) MCP tools | `brain_write` MCP tool |
| **Crush, Pi, Antigravity, other MCP-capable** | same MCP server: `ingestion-brain` (via the `mcp__*_brain__*` namespace once authenticated) | same |
| **Anything HTTP-only (curl, scripts)** | `POST https://brain-mcp.d-ma.be/query` with `{"query":"..."}` (auth via `BRAIN_MCP_TOKEN`) | `POST .../write` with `{"content":"...","filename":"..."}` |
| **Browser / human inspection** | `https://gitea.d-ma.be/mathias/hyperguild``knowledge/` and `wiki/` markdown files |
- **Scoping**: defaults to `public` collection; client projects filter to `{client}` + `public`.
- **Routing**: brain_answer's LLM uses berget.ai as primary, iguana ollama as
fallback. Both are configurable in the `supervisor/ingestion-deployment.yaml`
on the koala k3s cluster; don't hardcode local-only model names into the
berget URL (see knowledge entry on namespace mismatches).
### Quick reflex checks
If you find yourself about to say any of these out loud, you owe yourself a brain query first:
- "I think the issue might be..."
- "Let me try X and see..."
- "I'll just write a script to..."
- "This is probably a new bug..."
- "Has anyone done this before?" — *yes, probably, go check.*
## Client work rules
When working on a project tagged with a client name:
1. Never send code, data, or context to cloud APIs — use local models only
2. Never reference other client projects or their data
3. Keep all artifacts within the client's git org / directory
4. Treat everything as confidential unless told otherwise
## Harness-agnostic principles
This context is designed to work with any AI coding tool:
- Claude Code, Cursor, Aider, Open WebUI, Charmbracelet Mods/Crush
- Pi Coding Agent, Mistral Vibe, Antigravity
- Any tool that accepts a system prompt or reads a markdown context file
The canonical source is always `.context/AGENT.md` (root) and `.context/PROJECT.md` (per-project).
Derived files are committed (see *How context propagates* below) so a `git pull` on any host yields full agent context with no setup.
## How context propagates
Canonical sources of truth:
- Universal: `~/dev/.context/AGENT.md` (this file)
- Project: `<repo>/.context/PROJECT.md` (per-repo)
Derived files (committed, regenerated by `task context:sync`):
- `CLAUDE.md`, `AGENTS.md`, `.cursorrules`, `.aider.conventions.md`,
`.context/system-prompt.txt`
Workflow:
1. Edit a canonical file. Run `task context:sync`. Commit canonical and
derived together. Push.
2. On any other host, `git pull` brings both. Claude Code (tree-walking)
uses `CLAUDE.md`; Crush / Pi / Antigravity (cwd-only) use `AGENTS.md`;
Cursor uses `.cursorrules`; Aider uses `.aider.conventions.md`.
3. `task check` runs `context:sync` then asserts `git status --porcelain`
is empty over the derived files (catches both modified-tracked drift
and missing-untracked adapters). A drift fails the check with a
message telling you to stage the regenerated files.
Behavior rules in this file and per-project rules in `PROJECT.md` apply
unconditionally on every host, every harness.
## Engineering Skills
Shared engineering skills are available in `~/dev/.skills/`. Load on demand via the index.
See `~/dev/.skills/SKILLS_INDEX.md` for the full list with descriptions and "use when" triggers.
Key skills:
- **TDD**: always write tests first — load `tdd` skill
- **Code Review**: load `code-review` skill before any review
- **SOLID/Clean Code**: load `solid` or `clean-code` skill for design work
- **Problem first**: load `problem-analysis` skill before coding non-trivial features
---
# Project context
<!-- Canonical project context. Edit this, run `task context:sync`.
Root agent context from ~/dev/.context/AGENT.md is automatically
prepended for harnesses that don't walk the directory tree. -->
## Identity
- **Name**: supervisor
- **Owner**: Mathias
- **Client**: personal
- **Repo**:
- **Status**: active
## Stack
- **Primary language**: Go
- **UI layer**: HTMX + Templ (when applicable)
- **Fallback languages**: Python, TypeScript (justify in PR if used)
- **Build**: Task (taskfile.dev), not Make
- **Containers**: Docker (compose for dev, k3s for deploy)
- **Target infra**: koala (GPU workloads), iguana (services), flamingo (edge)
## Conventions
### Code style
- Go: follow `golines`, `gofumpt`, `golangci-lint` with project config
- Tests: table-driven, in `_test.go` next to source, `testify` for assertions
- Errors: wrap with `fmt.Errorf("operation: %w", err)`, no naked returns
- Naming: stdlib conventions, no stuttering (`http.Client` not `http.HTTPClient`)
### Architecture preferences
- Prefer standard library over frameworks (net/http over gin/echo)
- Dependency injection via constructor functions, not containers
- Configuration via environment variables, parsed at startup into a typed struct
- Structured logging via `slog`
### Git
- Conventional commits: `feat:`, `fix:`, `chore:`, `docs:`, `refactor:`
- Branch naming: `feat/short-description`, `fix/short-description`
- PRs: one concern per PR, description explains *why* not *what*
### Security
- No secrets in code, ever — use env vars or SOPS-encrypted files
- Client data never leaves local network unless explicitly cleared
- Dependencies: audit with `govulncheck` before adding
## MCP endpoints
Two MCP servers are live, both reachable over Tailscale and via HTTPS domain:
- **`brain`** at `https://brain-mcp.d-ma.be/mcp` (NodePort `koala:30330`) —
`brain_query`, `brain_write`, `brain_ingest`, `brain_ingest_raw`,
`brain_answer`, `brain_classify`, `session_log`. Hosted by the ingestion
service. Auth: Dex JWT (claude.ai OAuth) or static `BRAIN_MCP_TOKEN`.
- **`routing`** at `http://koala:30310/mcp` — Mode 2 routing pod. Advertises
`review`, `debug`, `retrospective`, `trainer`; per-call routes to local model
or Claude based on brain `/pass-rate`. Bearer auth via `ROUTING_MCP_TOKEN`
(opt-in). Only `mode client-local` registers this endpoint.
The supervisor MCP (`koala:30320`) was retired in Plan 7 (2026-05-12). Its
skill workers (`tdd`, `spec`) are now SKILL.md files; routed skills moved to
the routing pod; brain tools moved to the brain MCP.
The brain HTTP REST API (`/query`, `/write`, `/ingest`, `/ingest-raw`,
`/ingest-path`, `/backfill-refs`, `/pass-rate`) remains available on port 3300
for shell scripts and non-MCP clients.
`brain_answer(query)` performs BM25 retrieval + LLM synthesis (berget.ai
gemma4:31b → iguana fallback). `brain_classify(text)` infers doc type, title,
and tags. Both require `BRAIN_LLM_PRIMARY_URL` to be set in the ingestion pod.
## Agent instructions
When acting as a coding agent on this project:
1. Read this file and all `SKILL.md` files in `.skills/` before starting work
2. Run `task check` before committing (lint + test + vet)
3. If unsure about a convention, check `DECISIONS.md` or ask
4. Never modify files outside the project root without explicit permission
5. When adding a dependency, explain why in the commit message
6. For client projects: never send code or context to cloud APIs — use local models via LiteLLM
+54 -8
View File
@@ -32,6 +32,14 @@ and climate/sustainability tech.
These rules apply to every task across every project, regardless of harness.
0. **Pre-task ritual — before ANY implementation (non-negotiable).** Run this before writing a single line:
- **Query the brain** (`brain_query`) for the domain + symptom. If the result changes your approach, surface it before acting. 5 seconds beats 5 hours.
- **Load the relevant skill** — see trigger table in *Engineering Skills* below.
- **Write the failing test first.** Name the test before the function. If the target is untestable (e.g. `main()` wiring), extract the logic into a testable function first. No implementation without a red test.
- **State the observable success criterion** — what specific behavior, output, or passing test proves this is done?
**TDD is non-negotiable.** "Tests pass" is not proof of correctness — only proof the tests ran. Write tests that would catch the bug before writing code that fixes it.
1. **No assumptions.** Don't hide confusion — surface it. Surface tradeoffs explicitly.
Think before coding; if the problem is unclear, ask or state assumptions before acting.
2. **Minimum viable code.** Solve with the smallest change that works. Nothing
@@ -54,6 +62,22 @@ These rules apply to every task across every project, regardless of harness.
PR flow only when a human reviewer outside the project is required. Document
the reason in PROJECT.md.
6. **Close the loop — every substantive task ends with the same ritual.** Shipping
the code is not the end of the task; capturing it is. Run this unprompted:
- **Tag + bump SemVer** on the change (annotated tag; minor for a feature or
new/changed ADR, patch for a fix; docs in the same commit). Check the repo's
actual last tag — stated versions in docs drift stale.
- **Push** main and the tag (CI is the gate).
- **Persist generalizable learnings to the brain** (`brain_write`, wing/hall) —
the reusable patterns and the footguns that would bite anyone again, never
project status. See *Knowledge base — when to write* below.
- **File discovered-but-deferred work as tracker issues** on the project's own
repo — token-budget gaps, recorded ADR limitations, v2 follow-ups. Don't let
"out of scope, recorded" rot in a commit message; make it a ticket with a
source pointer.
- Surface the brain entries and issue numbers in the closing summary so the
trail is auditable.
## Default stack
| Layer | Default | Fallback | Last resort |
@@ -83,6 +107,26 @@ Exploratory: Rust, Zig — I'll tell you when I want these.
- **Security**: no secrets in code, govulncheck before adding deps, SOPS for encrypted config
- **Dependencies**: prefer stdlib. testify, slog, templ, sqlc, google.golang.org/adk (agent projects only) are pre-approved; anything else needs justification in the commit message
## Secret handling (every harness, every command)
Tool output is persisted: terminal → `~/.claude/projects` transcripts →
claudewatcher → brain/wiki → gitea history. A secret printed once is
searchable forever, and clearing it means rotating the key. So:
1. **Never print, echo, log, or transform a secret to inspect it.** No
`base64`/`xxd`/`cat` of a key, and never pipe a secret through a transform
to defeat `op run`'s output masking (it masks raw values; base64 hides them
from the mask — that exact trick leaked a key on 2026-06-11).
2. **Secrets stay in the subprocess.** Reference them only as env vars consumed
*inside* `op run --env-file ~/.op-env -- <cmd>`. Never place a literal secret
in a command's argv (it lands in the tool call and the transcript).
3. **Existence check without revealing the value:** `[ -n "$X" ] && echo set` —
never `${X:-...}` (returns the value when set) and never echo a substring of it.
4. **Cross-host secrets:** run the secret-consuming command on the host that has
the secret; do not forward a raw key over ssh argv/stdout.
5. If a secret does leak into output, say so immediately and flag it for rotation —
don't bury it.
## Infrastructure
Three machines on Tailscale:
@@ -162,7 +206,7 @@ entries that age well are about *why*, *how to avoid*, and *what to do when*.
| **Claude Code, Claude Desktop** | `brain_query` (BM25), `brain_answer` (LLM-synth + sources) MCP tools | `brain_write` MCP tool |
| **Crush, Pi, Antigravity, other MCP-capable** | same MCP server: `ingestion-brain` (via the `mcp__*_brain__*` namespace once authenticated) | same |
| **Anything HTTP-only (curl, scripts)** | `POST https://brain-mcp.d-ma.be/query` with `{"query":"..."}` (auth via `BRAIN_MCP_TOKEN`) | `POST .../write` with `{"content":"...","filename":"..."}` |
| **Browser / human inspection** | `https://gitea.d-ma.be/mathias/hyperguild` → `knowledge/` and `wiki/` markdown files |
| **Browser / human inspection** | `https://git.d-ma.be/mathias/hyperguild` → `knowledge/` and `wiki/` markdown files |
- **Scoping**: defaults to `public` collection; client projects filter to `{client}` + `public`.
- **Routing**: brain_answer's LLM uses berget.ai as primary, iguana ollama as
@@ -224,15 +268,17 @@ unconditionally on every host, every harness.
## Engineering Skills
Shared engineering skills are available in `~/dev/.skills/`. Load on demand via the index.
Shared engineering skills are available in `~/dev/.skills/`. Load at task start — not "on demand" but on schedule, before writing code. See `~/dev/.skills/SKILLS_INDEX.md` for the full list.
See `~/dev/.skills/SKILLS_INDEX.md` for the full list with descriptions and "use when" triggers.
**Skill trigger table — load before starting, not after getting stuck:**
Key skills:
- **TDD**: always write tests first — load `tdd` skill
- **Code Review**: load `code-review` skill before any review
- **SOLID/Clean Code**: load `solid` or `clean-code` skill for design work
- **Problem first**: load `problem-analysis` skill before coding non-trivial features
| Task type | Load |
|-----------|------|
| Any feature or bug fix | `tdd` |
| Refactor or design | `clean-code` or `solid` |
| Debug | `problem-analysis` |
| Review code or PRs | `code-review` |
| Frame a problem before coding | `problem-analysis` |
---
-318
View File
@@ -1,318 +0,0 @@
# Cursor rules — auto-generated
# Do not edit. Run: task context:sync
# Agent context — Mathias workspace
<!-- Canonical root context for all AI coding agents.
Lives at: ~/dev/.context/AGENT.md
Applies to every project under ~/dev/ unless overridden.
Run `task context:sync` from ~/dev/ to regenerate harness-specific files.
Project-level context in .context/PROJECT.md layers on top of this. -->
## Who I am
I'm Mathias, a digital product manager and technology consultant based in Sweden.
I build software, research emerging tech, and deliver consulting engagements
for clients under NDA. I work across AI/ML, financial automation, web applications,
and climate/sustainability tech.
## How I work with agents
- I think like a product manager — I care about *why* before *how*
- I want agents to be opinionated and push back, not just execute blindly
- I prefer concise responses; skip ceremony and get to the point
- When I say "build this", I mean production-quality with tests, not a demo
- Ask me before making irreversible changes or adding heavy dependencies
- I work with confidential client data — never send it to cloud APIs unless I explicitly say it's OK
## Behavior rules
These rules apply to every task across every project, regardless of harness.
1. **No assumptions.** Don't hide confusion — surface it. Surface tradeoffs explicitly.
Think before coding; if the problem is unclear, ask or state assumptions before acting.
2. **Minimum viable code.** Solve with the smallest change that works. Nothing
speculative, no "while we're here" cleanups, no premature abstractions. Simplicity first.
3. **Surgical changes.** Touch only what the task requires. Leave unrelated code,
files, and formatting alone. Diffs should be small and reviewable.
4. **Goal-driven execution.** Define clear success criteria up front for every task.
Loop — implement, verify, refine — until those criteria are met. Don't claim
completion without evidence (tests pass, command output, observed behavior).
5. **Trunk-Based Development — commit directly to main.** Every commit is one
logical change (one tool, one fix, one test) with passing tests. Main is always
deployable. Never create long-lived feature branches.
**Exception — parallel agents on same repo:** If another agent is known to be
actively working on the same repo simultaneously, create a short-lived branch
(`agent/<description>`), finish the task, and merge to main within the same
session. Do not leave agent branches open between sessions.
**Exception — external contributor or client four-eyes requirement:** Use
PR flow only when a human reviewer outside the project is required. Document
the reason in PROJECT.md.
## Default stack
| Layer | Default | Fallback | Last resort |
|-------|---------|----------|-------------|
| Language | Go | Python | TypeScript, Java, C |
| UI | HTMX + Templ | Server-rendered HTML | React (only if SPA is justified) |
| Build | Task (taskfile.dev) | Make | — |
| Containers | Docker Compose (dev), k3s (prod) | — | — |
| DB | PostgreSQL + sqlc | SQLite | — |
| Search | pgvector (vector), BM25 | Qdrant (when >1M vectors or hybrid retrieval) | — |
| Logging | slog (structured) | — | — |
| Testing | Table-driven, testify | — | — |
| Agents (Go) | google.golang.org/adk + pkg/litellm adapter | — | — |
Exploratory: Rust, Zig — I'll tell you when I want these.
## Code conventions
- **Go style**: golines, gofumpt, golangci-lint
- **Errors**: `fmt.Errorf("operation: %w", err)` — never naked, never log-and-return
- **Naming**: stdlib conventions, no stuttering
- **Architecture**: prefer stdlib over frameworks, constructor injection, env-var config parsed into typed structs
- **Git**: conventional commits (`feat:`, `fix:`, `chore:`), commit directly to main,
one logical change per commit, CI is the quality gate
- **Never**: long-lived feature branches, PRs for solo work, direct push without
passing `task check` locally first
- **Security**: no secrets in code, govulncheck before adding deps, SOPS for encrypted config
- **Dependencies**: prefer stdlib. testify, slog, templ, sqlc, google.golang.org/adk (agent projects only) are pre-approved; anything else needs justification in the commit message
## Infrastructure
Three machines on Tailscale:
| Machine | Role | Key specs |
|---------|------|-----------|
| koala | GPU inference, heavy compute | RTX 5070, runs k3s + llama-swap + shared postgres18/pgvector |
| iguana | Services, builds | M2 Ultra Mac |
| flamingo | Daily driver, edge | Mac mini, ~/dev is here |
- **Model routing**: LiteLLM in front of llama-swap (local) + cloud APIs (when permitted)
- **Orchestration**: k3s cluster across all three machines
- **Networking**: Tailscale mesh
## Project landscape
All development repos live at `~/dev/` (softlink from `~/Documents/local-dev/`).
Organized in thematic folders:
| Folder | Focus | Count |
|--------|-------|-------|
| `GO/` | Go web frameworks, API integrations, learning projects | ~10 |
| `AI/` | ML research, AI frameworks (FinRL, DSPy, crawl4ai) | ~6 |
| `AGENTS/` | Autonomous agents, coding agents, MCP servers, infra | ~15 |
| `QKX/` | Invoice processing, financial automation, payment systems | ~13 |
| `XT/` | Climate data, sustainability (Klimatkollen, Garbo) | ~2 |
See `~/dev/PROJECT_SUMMARY.md` for detailed descriptions of each project.
### Key active projects
- **super-koala** (`AGENTS/`) — multi-component agent stack with LangGraph, DSPy, MCP
- **azure-tiger** (`QKX/`) — invoice extraction → ISO 20022 payment instructions
- **gocrwl** (`AGENTS/`) — Go web crawler with containerized deployment
- **koala-ai-stack** (`AGENTS/`) — local AI server infrastructure management
- **klimatkollen** (`XT/`) — Swedish municipal climate data platform
## Knowledge base — actively use it
A persistent brain (BM25 search + LLM-synthesised Q&A) survives across sessions,
hosts, and harnesses. It holds 100+ hard-won entries: infra incident postmortems,
Go pitfalls, framework gotchas, design principles, ADRs. **It is not optional
reference material — query it actively, not just when explicitly told.**
### When to query (treat as a reflex)
- **Before** starting a non-trivial task — search for prior art with the symptom
AND the system component ("how did we solve X in Y?"). 5 seconds beats 5 hours.
- **When debugging** — search for the error string, the stack frame, the affected
service. Past you may have already paid this tax.
- **Before adopting** a pattern, library, framework, or model name — check if it
was tried and rejected, or what the integration footguns are.
- **When making architectural decisions** — search for the domain + "ADR" or
"decision" to find prior reasoning before re-deriving it.
- **When a recommendation feels novel** — challenge yourself: "has this been
documented?" The brain often has it.
### When to write
After you discover something that **future-you would forget** and that **isn't
recoverable from the code, git log, or PR description alone**:
- Bugs whose root cause is non-obvious and generalisable beyond this project.
- Framework / library / model-name quirks that bit you and would bite anyone.
- Design principles validated under fire (e.g. "every `_get` needs a `_list`").
- Postmortems for incidents: what broke, why, how diagnosed, what to do next time.
DON'T write project status, sprint progress, PR summaries, or "what I did this
session" — those rot fast and the originals are in git/gitea anyway. Brain
entries that age well are about *why*, *how to avoid*, and *what to do when*.
### How to access (per harness)
| Harness | Query | Write |
|---------|-------|-------|
| **Claude Code, Claude Desktop** | `brain_query` (BM25), `brain_answer` (LLM-synth + sources) MCP tools | `brain_write` MCP tool |
| **Crush, Pi, Antigravity, other MCP-capable** | same MCP server: `ingestion-brain` (via the `mcp__*_brain__*` namespace once authenticated) | same |
| **Anything HTTP-only (curl, scripts)** | `POST https://brain-mcp.d-ma.be/query` with `{"query":"..."}` (auth via `BRAIN_MCP_TOKEN`) | `POST .../write` with `{"content":"...","filename":"..."}` |
| **Browser / human inspection** | `https://gitea.d-ma.be/mathias/hyperguild` → `knowledge/` and `wiki/` markdown files |
- **Scoping**: defaults to `public` collection; client projects filter to `{client}` + `public`.
- **Routing**: brain_answer's LLM uses berget.ai as primary, iguana ollama as
fallback. Both are configurable in the `supervisor/ingestion-deployment.yaml`
on the koala k3s cluster; don't hardcode local-only model names into the
berget URL (see knowledge entry on namespace mismatches).
### Quick reflex checks
If you find yourself about to say any of these out loud, you owe yourself a brain query first:
- "I think the issue might be..."
- "Let me try X and see..."
- "I'll just write a script to..."
- "This is probably a new bug..."
- "Has anyone done this before?" — *yes, probably, go check.*
## Client work rules
When working on a project tagged with a client name:
1. Never send code, data, or context to cloud APIs — use local models only
2. Never reference other client projects or their data
3. Keep all artifacts within the client's git org / directory
4. Treat everything as confidential unless told otherwise
## Harness-agnostic principles
This context is designed to work with any AI coding tool:
- Claude Code, Cursor, Aider, Open WebUI, Charmbracelet Mods/Crush
- Pi Coding Agent, Mistral Vibe, Antigravity
- Any tool that accepts a system prompt or reads a markdown context file
The canonical source is always `.context/AGENT.md` (root) and `.context/PROJECT.md` (per-project).
Derived files are committed (see *How context propagates* below) so a `git pull` on any host yields full agent context with no setup.
## How context propagates
Canonical sources of truth:
- Universal: `~/dev/.context/AGENT.md` (this file)
- Project: `<repo>/.context/PROJECT.md` (per-repo)
Derived files (committed, regenerated by `task context:sync`):
- `CLAUDE.md`, `AGENTS.md`, `.cursorrules`, `.aider.conventions.md`,
`.context/system-prompt.txt`
Workflow:
1. Edit a canonical file. Run `task context:sync`. Commit canonical and
derived together. Push.
2. On any other host, `git pull` brings both. Claude Code (tree-walking)
uses `CLAUDE.md`; Crush / Pi / Antigravity (cwd-only) use `AGENTS.md`;
Cursor uses `.cursorrules`; Aider uses `.aider.conventions.md`.
3. `task check` runs `context:sync` then asserts `git status --porcelain`
is empty over the derived files (catches both modified-tracked drift
and missing-untracked adapters). A drift fails the check with a
message telling you to stage the regenerated files.
Behavior rules in this file and per-project rules in `PROJECT.md` apply
unconditionally on every host, every harness.
## Engineering Skills
Shared engineering skills are available in `~/dev/.skills/`. Load on demand via the index.
See `~/dev/.skills/SKILLS_INDEX.md` for the full list with descriptions and "use when" triggers.
Key skills:
- **TDD**: always write tests first — load `tdd` skill
- **Code Review**: load `code-review` skill before any review
- **SOLID/Clean Code**: load `solid` or `clean-code` skill for design work
- **Problem first**: load `problem-analysis` skill before coding non-trivial features
---
# Project context
<!-- Canonical project context. Edit this, run `task context:sync`.
Root agent context from ~/dev/.context/AGENT.md is automatically
prepended for harnesses that don't walk the directory tree. -->
## Identity
- **Name**: supervisor
- **Owner**: Mathias
- **Client**: personal
- **Repo**:
- **Status**: active
## Stack
- **Primary language**: Go
- **UI layer**: HTMX + Templ (when applicable)
- **Fallback languages**: Python, TypeScript (justify in PR if used)
- **Build**: Task (taskfile.dev), not Make
- **Containers**: Docker (compose for dev, k3s for deploy)
- **Target infra**: koala (GPU workloads), iguana (services), flamingo (edge)
## Conventions
### Code style
- Go: follow `golines`, `gofumpt`, `golangci-lint` with project config
- Tests: table-driven, in `_test.go` next to source, `testify` for assertions
- Errors: wrap with `fmt.Errorf("operation: %w", err)`, no naked returns
- Naming: stdlib conventions, no stuttering (`http.Client` not `http.HTTPClient`)
### Architecture preferences
- Prefer standard library over frameworks (net/http over gin/echo)
- Dependency injection via constructor functions, not containers
- Configuration via environment variables, parsed at startup into a typed struct
- Structured logging via `slog`
### Git
- Conventional commits: `feat:`, `fix:`, `chore:`, `docs:`, `refactor:`
- Branch naming: `feat/short-description`, `fix/short-description`
- PRs: one concern per PR, description explains *why* not *what*
### Security
- No secrets in code, ever — use env vars or SOPS-encrypted files
- Client data never leaves local network unless explicitly cleared
- Dependencies: audit with `govulncheck` before adding
## MCP endpoints
Two MCP servers are live, both reachable over Tailscale and via HTTPS domain:
- **`brain`** at `https://brain-mcp.d-ma.be/mcp` (NodePort `koala:30330`) —
`brain_query`, `brain_write`, `brain_ingest`, `brain_ingest_raw`,
`brain_answer`, `brain_classify`, `session_log`. Hosted by the ingestion
service. Auth: Dex JWT (claude.ai OAuth) or static `BRAIN_MCP_TOKEN`.
- **`routing`** at `http://koala:30310/mcp` — Mode 2 routing pod. Advertises
`review`, `debug`, `retrospective`, `trainer`; per-call routes to local model
or Claude based on brain `/pass-rate`. Bearer auth via `ROUTING_MCP_TOKEN`
(opt-in). Only `mode client-local` registers this endpoint.
The supervisor MCP (`koala:30320`) was retired in Plan 7 (2026-05-12). Its
skill workers (`tdd`, `spec`) are now SKILL.md files; routed skills moved to
the routing pod; brain tools moved to the brain MCP.
The brain HTTP REST API (`/query`, `/write`, `/ingest`, `/ingest-raw`,
`/ingest-path`, `/backfill-refs`, `/pass-rate`) remains available on port 3300
for shell scripts and non-MCP clients.
`brain_answer(query)` performs BM25 retrieval + LLM synthesis (berget.ai
gemma4:31b → iguana fallback). `brain_classify(text)` infers doc type, title,
and tags. Both require `BRAIN_LLM_PRIMARY_URL` to be set in the ingestion pod.
## Agent instructions
When acting as a coding agent on this project:
1. Read this file and all `SKILL.md` files in `.skills/` before starting work
2. Run `task check` before committing (lint + test + vet)
3. If unsure about a convention, check `DECISIONS.md` or ask
4. Never modify files outside the project root without explicit permission
5. When adding a dependency, explain why in the commit message
6. For client projects: never send code or context to cloud APIs — use local models via LiteLLM
+9 -9
View File
@@ -1,6 +1,6 @@
name: cd
on:
"on":
workflow_run:
workflows: ["CI"]
types: [completed]
@@ -13,9 +13,9 @@ jobs:
if: ${{ github.event.workflow_run.conclusion == 'success' && github.event.workflow_run.event == 'push' }}
environment: staging
env:
INGESTION_IMAGE: gitea.d-ma.be/mathias/ingestion
ROUTING_IMAGE: gitea.d-ma.be/mathias/routing
INFRA_REPO: git@gitea.d-ma.be:mathias/infra.git
INGESTION_IMAGE: git.d-ma.be/mathias/ingestion
ROUTING_IMAGE: git.d-ma.be/mathias/routing
INFRA_REPO: git@git.d-ma.be:mathias/infra.git
BUILDKIT_HOST: unix:///run/buildkit/buildkitd.sock
steps:
- name: Checkout
@@ -71,17 +71,17 @@ jobs:
mkdir -p ~/.ssh
echo "${{ secrets.INFRA_DEPLOY_KEY }}" > ~/.ssh/infra_deploy_key
chmod 600 ~/.ssh/infra_deploy_key
printf 'Host gitea.d-ma.be\n HostName 127.0.0.1\n Port 30022\n StrictHostKeyChecking no\n' >> ~/.ssh/config
printf 'Host git.d-ma.be\n HostName 127.0.0.1\n Port 30022\n StrictHostKeyChecking no\n' >> ~/.ssh/config
GIT_SSH_COMMAND="ssh -i ~/.ssh/infra_deploy_key -o IdentitiesOnly=yes" \
git clone "${INFRA_REPO}" /tmp/infra-update
cd /tmp/infra-update
sed -i "s|gitea.d-ma.be/mathias/ingestion:.*|gitea.d-ma.be/mathias/ingestion:${IMAGE_TAG}|" \
sed -i "s|git.d-ma.be/mathias/ingestion:.*|git.d-ma.be/mathias/ingestion:${IMAGE_TAG}|" \
"k3s/apps/supervisor/ingestion-deployment.yaml"
sed -i "s|gitea.d-ma.be/mathias/routing:.*|gitea.d-ma.be/mathias/routing:${IMAGE_TAG}|" \
sed -i "s|git.d-ma.be/mathias/routing:.*|git.d-ma.be/mathias/routing:${IMAGE_TAG}|" \
"k3s/apps/routing/deployment.yaml"
git config user.email "cd-bot@d-ma.be"
@@ -103,7 +103,7 @@ jobs:
- name: Wait for Flux to apply new ingestion image
run: |
EXPECTED="gitea.d-ma.be/mathias/ingestion:${{ github.sha }}"
EXPECTED="git.d-ma.be/mathias/ingestion:${{ github.sha }}"
for i in $(seq 1 60); do
CURRENT=$(kubectl get deploy ingestion -n supervisor \
-o jsonpath='{.spec.template.spec.containers[0].image}' 2>/dev/null || echo "")
@@ -135,7 +135,7 @@ jobs:
- name: Wait for Flux to apply new routing image
run: |
EXPECTED="gitea.d-ma.be/mathias/routing:${{ github.sha }}"
EXPECTED="git.d-ma.be/mathias/routing:${{ github.sha }}"
for i in $(seq 1 60); do
CURRENT=$(kubectl get deploy routing -n routing \
-o jsonpath='{.spec.template.spec.containers[0].image}' 2>/dev/null || echo "")
+1 -1
View File
@@ -1,6 +1,6 @@
name: CI
on:
"on":
push:
branches: [main]
tags: ["v*"]
+54 -8
View File
@@ -27,6 +27,14 @@ and climate/sustainability tech.
These rules apply to every task across every project, regardless of harness.
0. **Pre-task ritual — before ANY implementation (non-negotiable).** Run this before writing a single line:
- **Query the brain** (`brain_query`) for the domain + symptom. If the result changes your approach, surface it before acting. 5 seconds beats 5 hours.
- **Load the relevant skill** — see trigger table in *Engineering Skills* below.
- **Write the failing test first.** Name the test before the function. If the target is untestable (e.g. `main()` wiring), extract the logic into a testable function first. No implementation without a red test.
- **State the observable success criterion** — what specific behavior, output, or passing test proves this is done?
**TDD is non-negotiable.** "Tests pass" is not proof of correctness — only proof the tests ran. Write tests that would catch the bug before writing code that fixes it.
1. **No assumptions.** Don't hide confusion — surface it. Surface tradeoffs explicitly.
Think before coding; if the problem is unclear, ask or state assumptions before acting.
2. **Minimum viable code.** Solve with the smallest change that works. Nothing
@@ -49,6 +57,22 @@ These rules apply to every task across every project, regardless of harness.
PR flow only when a human reviewer outside the project is required. Document
the reason in PROJECT.md.
6. **Close the loop — every substantive task ends with the same ritual.** Shipping
the code is not the end of the task; capturing it is. Run this unprompted:
- **Tag + bump SemVer** on the change (annotated tag; minor for a feature or
new/changed ADR, patch for a fix; docs in the same commit). Check the repo's
actual last tag — stated versions in docs drift stale.
- **Push** main and the tag (CI is the gate).
- **Persist generalizable learnings to the brain** (`brain_write`, wing/hall) —
the reusable patterns and the footguns that would bite anyone again, never
project status. See *Knowledge base — when to write* below.
- **File discovered-but-deferred work as tracker issues** on the project's own
repo — token-budget gaps, recorded ADR limitations, v2 follow-ups. Don't let
"out of scope, recorded" rot in a commit message; make it a ticket with a
source pointer.
- Surface the brain entries and issue numbers in the closing summary so the
trail is auditable.
## Default stack
| Layer | Default | Fallback | Last resort |
@@ -78,6 +102,26 @@ Exploratory: Rust, Zig — I'll tell you when I want these.
- **Security**: no secrets in code, govulncheck before adding deps, SOPS for encrypted config
- **Dependencies**: prefer stdlib. testify, slog, templ, sqlc, google.golang.org/adk (agent projects only) are pre-approved; anything else needs justification in the commit message
## Secret handling (every harness, every command)
Tool output is persisted: terminal → `~/.claude/projects` transcripts →
claudewatcher → brain/wiki → gitea history. A secret printed once is
searchable forever, and clearing it means rotating the key. So:
1. **Never print, echo, log, or transform a secret to inspect it.** No
`base64`/`xxd`/`cat` of a key, and never pipe a secret through a transform
to defeat `op run`'s output masking (it masks raw values; base64 hides them
from the mask — that exact trick leaked a key on 2026-06-11).
2. **Secrets stay in the subprocess.** Reference them only as env vars consumed
*inside* `op run --env-file ~/.op-env -- <cmd>`. Never place a literal secret
in a command's argv (it lands in the tool call and the transcript).
3. **Existence check without revealing the value:** `[ -n "$X" ] && echo set`
never `${X:-...}` (returns the value when set) and never echo a substring of it.
4. **Cross-host secrets:** run the secret-consuming command on the host that has
the secret; do not forward a raw key over ssh argv/stdout.
5. If a secret does leak into output, say so immediately and flag it for rotation —
don't bury it.
## Infrastructure
Three machines on Tailscale:
@@ -157,7 +201,7 @@ entries that age well are about *why*, *how to avoid*, and *what to do when*.
| **Claude Code, Claude Desktop** | `brain_query` (BM25), `brain_answer` (LLM-synth + sources) MCP tools | `brain_write` MCP tool |
| **Crush, Pi, Antigravity, other MCP-capable** | same MCP server: `ingestion-brain` (via the `mcp__*_brain__*` namespace once authenticated) | same |
| **Anything HTTP-only (curl, scripts)** | `POST https://brain-mcp.d-ma.be/query` with `{"query":"..."}` (auth via `BRAIN_MCP_TOKEN`) | `POST .../write` with `{"content":"...","filename":"..."}` |
| **Browser / human inspection** | `https://gitea.d-ma.be/mathias/hyperguild``knowledge/` and `wiki/` markdown files |
| **Browser / human inspection** | `https://git.d-ma.be/mathias/hyperguild``knowledge/` and `wiki/` markdown files |
- **Scoping**: defaults to `public` collection; client projects filter to `{client}` + `public`.
- **Routing**: brain_answer's LLM uses berget.ai as primary, iguana ollama as
@@ -219,15 +263,17 @@ unconditionally on every host, every harness.
## Engineering Skills
Shared engineering skills are available in `~/dev/.skills/`. Load on demand via the index.
Shared engineering skills are available in `~/dev/.skills/`. Load at task start — not "on demand" but on schedule, before writing code. See `~/dev/.skills/SKILLS_INDEX.md` for the full list.
See `~/dev/.skills/SKILLS_INDEX.md` for the full list with descriptions and "use when" triggers.
**Skill trigger table — load before starting, not after getting stuck:**
Key skills:
- **TDD**: always write tests first — load `tdd` skill
- **Code Review**: load `code-review` skill before any review
- **SOLID/Clean Code**: load `solid` or `clean-code` skill for design work
- **Problem first**: load `problem-analysis` skill before coding non-trivial features
| Task type | Load |
|-----------|------|
| Any feature or bug fix | `tdd` |
| Refactor or design | `clean-code` or `solid` |
| Debug | `problem-analysis` |
| Review code or PRs | `code-review` |
| Frame a problem before coding | `problem-analysis` |
---
+68
View File
@@ -4,6 +4,74 @@ Record *why* things are the way they are. Future-you will thank present-you.
---
## 2026-05-28 — three active harnesses: hyperguild, agentsquad, Crush (extends earlier boundary decision)
**Context:** After wiring Crush to LiteLLM in May 2026, there are now three active harnesses.
The earlier boundary decision only covered hyperguild vs agentsquad. Crush's role was undefined.
**Decision:** Three harnesses, three distinct roles, shared skills layer.
| Harness | Engine | Primary use | Brain MCP? | Routing pod? | Skills? |
|---------|--------|-------------|------------|--------------|---------|
| **hyperguild** | Claude Code + MCP | Disciplined solo coding sessions, TDD/review/debug workflows | Yes | Yes | Yes (SKILL.md) |
| **agentsquad** | OpenCode + LiteLLM | Multi-agent task execution, executor/reviewer pipelines | No | No (own routing) | Yes (SKILL.md) |
| **Crush** | Charmbracelet TUI + LiteLLM | Interactive local coding, quick iterations on flamingo | No (not yet) | No (direct LiteLLM) | Yes (SKILL.md) |
**Crush specifics (as of 2026-05-28):**
- Config: `~/.config/crush/crush.json` on flamingo (see brain: `homelab/facts/crush-litellm-wiring-2026-05`)
- Connects directly to LiteLLM at `http://koala:4000/v1/` using `sk-local-123`
- Auth type: `openai-compat` (not `openai`)
- Does NOT go through the routing pod — model selection is manual in the Crush UI
- Brain MCP not wired — Crush has no MCP client capability today; revisit if Crush adds MCP support
**Shared across all three:**
- `mathias/skills` — any SKILL.md file works in all three harnesses
- LiteLLM proxy on koala (`http://koala:4000/v1/`) — Crush and agentsquad both route through it; hyperguild does too for local model calls
**Consequences:** No consolidation needed. crush.json must be kept in sync when litellm_config.yaml model names change. The `crush.json` canonical location is `~/.config/crush/crush.json` on flamingo — not yet tracked in a dotfiles repo (track as tech debt).
---
## 2026-05-28 — "field benchmark" for local models = pass-rate at scale (supersedes GOTTH eval suite)
**Context:** The GOTTH eval suite (45 offline prompts across 5 categories) was replaced by
a "field benchmark" in May 2026, but the replacement was never defined concretely.
**Decision:** The field benchmark is per-skill pass rate over real routing pod usage,
collected automatically by `internal/routing/passrate.go` and exposed at:
```
GET /pass-rate?skill=<name>&window=<duration>
```
No separate eval suite. No synthetic prompts. The benchmark runs itself once the routing
pod receives real traffic. Target: 30-day rolling window per skill, reviewed monthly.
**Bootstrap note:** With no session history, `passrate.go` returns `nil` and the router
defaults to the thinking model for every call. The fast-model path activates only after
real pass-rate data accumulates. Seed with real usage — do not pre-populate.
**Consequences:** Zero maintenance overhead for the benchmark. The tradeoff is that results
are only meaningful after ~2 weeks of real usage, and skills that are rarely invoked will
have statistically thin pass-rate data. Revisit if a skill has fewer than 20 calls in 30 days.
---
## 2026-05-28 — brain injection in skill handlers: review is done, others unverified
**Context:** The April 2026 scope reset listed "brain_query injection into skill handlers"
as the top priority. As of 2026-05-28, `internal/skills/review/handlers.go` calls
`brain.Query(ctx, ...)` before dispatching to the LLM — confirmed in code review.
Status of debug, retrospective, and trainer handlers is unverified.
**Decision:** Treat review as the reference implementation. Verify debug, retrospective,
trainer against the same pattern before shipping new skill work. Tracked in issue #32.
**Consequences:** The April concern may be stale for review. A one-pass audit of the other
three skill handlers closes this fully.
---
## 2026-04-08 — AGENTS.md as cross-tool standard, not CLAUDE.md
**Context**: Multiple tools (Crush, Pi, Antigravity) read `AGENTS.md` natively. Claude Code reads `CLAUDE.md`. Building on `CLAUDE.md` as the primary format locks into one vendor.
+72 -49
View File
@@ -5,14 +5,31 @@ Instead of letting Claude Code do whatever it wants, hyperguild enforces structu
workflows (TDD red/green/refactor), logs every session, and accumulates learnings
into a searchable brain.
## Hypothesis
> We believe routing skill tasks through local models, backed by brain context,
> produces measurably better outcomes than raw Claude Code alone —
> measurable by per-skill pass rate over rolling 30-day windows
> (available at `GET /pass-rate?skill=<name>&window=30d` on the brain pod).
This is the falsifiable claim the routing pod and pass-rate infrastructure exist to test.
If per-skill pass rates don't improve over baseline (all-cloud) after 30 days of real
usage, the fast-model routing path should be reconsidered.
## Harness
**hyperguild = Claude Code + MCP.** This is a supervisor for Claude Code sessions specifically.
For multi-agent orchestration (OpenCode + LiteLLM, executor/reviewer pipelines), see
[agentsquad](http://gitea.d-ma.be/mathias/agentsquad) — a separate harness for a different
orchestration model. Skills (mathias/skills) are shared between both.
## How it works
```
Your Claude Code session (in any project)
│ MCP over HTTP (Tailscale)
├──▶ supervisor :3200 (NodePort 30320 on koala) — skill workers: tdd, debug, spec, …
├──▶ routing :3210 (NodePort 30310 on koala) — Mode 2 only: review, debug, retrospective, trainer
├──▶ routing :3210 (NodePort 30310 on koala) — review, debug, retrospective, trainer
└──▶ brain :3300 (NodePort 30330 on koala) — brain_query, brain_write, brain_ingest, session_log
└─ also serves the legacy REST endpoints (/query, /write, /ingest, …)
@@ -20,34 +37,28 @@ Your Claude Code session (in any project)
brain/
├── sessions/ — JSONL log, one file per session_id
├── wiki/ — searchable knowledge (full-text)
│ ├── concepts/
│ ├── entities/
│ └── sources/
── raw/ — retrospective output, staged for review
└── training-data/ — SFT/DPO/RL data (Phase 2)
├── wiki/ — searchable knowledge (wing/hall layout)
│ ├── homelab/
│ ├── claude-sessions/
│ └── ...
── knowledge/ — legacy flat notes (migration pending: hyperguild#22)
```
## Phase 1 tools (available now)
| Tool | What it does |
|------|-------------|
| `tdd_red` | Writes a failing test for a spec, verifies it fails |
| `tdd_green` | Writes the minimal implementation to make tests pass |
| `tdd_refactor` | Cleans up implementation while keeping tests green |
| `session_log` | Appends a structured entry to the session JSONL log |
| `retrospective` | Reads the session log, identifies novel learnings, writes to brain/raw/ |
| `retrospective` | Reads the session log, identifies novel learnings, writes to brain |
| `review` | Structured code review via local model, brain-context injected |
| `debug` | Hypothesis-driven debugging via local model |
| `brain_query` | Full-text search over brain/wiki/ |
| `brain_write` | Writes a note to brain/raw/ (with optional YAML frontmatter) |
| `brain_write` | Writes a note to brain (with wing/hall routing) |
| `brain_answer` | BM25 + LLM synthesis — Q&A over brain corpus |
| `tier` | Returns the current connectivity tier (1=cloud, 2=LAN, 3=offline) |
## Start the servers
```bash
# Requires goreman: go install github.com/mattn/goreman@latest
task start # starts ingestion (:3300) + supervisor (:3200) via goreman
task stop # kills both by port
```
> **Note:** `tdd_red/green/refactor` and `spec` were retired in Plan 7 (2026-05-12).
> They are now SKILL.md files in [mathias/skills](http://gitea.d-ma.be/mathias/skills).
## Connect a project
@@ -56,9 +67,9 @@ Create `.mcp.json` in your project root:
```json
{
"mcpServers": {
"supervisor": {
"routing": {
"type": "http",
"url": "http://koala:30320/mcp"
"url": "http://koala:30310/mcp"
},
"brain": {
"type": "http",
@@ -68,33 +79,29 @@ Create `.mcp.json` in your project root:
}
```
Two MCP servers are exposed today, both reachable over Tailscale:
Two MCP servers are exposed, both reachable over Tailscale:
- **`supervisor`** at `koala:30320` — skill workers (`tdd_red/green/refactor`,
`review`, `debug`, `spec`, `retrospective`, `trainer`, `tier`).
- **`routing`** at `koala:30310` — skill workers (`review`, `debug`, `retrospective`, `trainer`).
Routes each call to fast local model or thinking model based on per-skill pass rate.
- **`brain`** at `koala:30330` — knowledge access (`brain_query`, `brain_write`,
`brain_ingest`, `brain_ingest_raw`) and `session_log`. Hosted by the ingestion
service directly, no separate pod.
`brain_ingest`, `brain_ingest_raw`, `brain_answer`, `brain_classify`) and `session_log`.
No local binary or stdio shim is required — Claude Code talks to both via HTTP.
Open Claude Code in your project — run `/mcp` to confirm both servers are listed.
## A typical TDD session
## A typical session
```
1. Call tdd_red → spec in, failing test file out
2. Call tdd_green → test path in, implementation out
3. Call tdd_refactor → impl + test in, cleaned code out
4. Call session_log → log each phase result
5. Call retrospective → extracts learnings → brain/raw/
6. Review brain/raw/, move worthy notes to brain/wiki/concepts/
7. Future sessions: call brain_query to retrieve relevant context
1. Call review → brain context injected + local model review → findings
2. Call session_log → log each phase result
3. Call retrospective → extracts learnings → brain
4. Future sessions: call brain_query / brain_answer to retrieve relevant context
```
## Tier detection
The supervisor probes connectivity at call time:
The routing pod probes connectivity at call time:
| Tier | Label | Condition |
|------|-------|-----------|
@@ -102,31 +109,47 @@ The supervisor probes connectivity at call time:
| 2 | lan-only | Can reach LiteLLM but not Anthropic |
| 3 | airplane | No external connectivity |
## Model routing
The routing pod selects models per skill call based on historical pass rate:
| Pass rate | Decision |
|-----------|----------|
| ≥ 0.90 (FLOOR) | Fast model (`HYPERGUILD_FAST_MODEL`) |
| ≤ 0.70 (CEIL) | Thinking model (`HYPERGUILD_THINKING_MODEL`) |
| between CEIL and FLOOR | Sample band — probabilistic routing |
| nil (no history yet) | Defaults to thinking model |
> **Bootstrap note:** With no session history, all calls route to the thinking model.
> The fast-model path activates only after real pass-rate data accumulates at `/pass-rate`.
> Seed with real usage — don't try to pre-populate.
## Key env vars
| Variable | Default | Purpose |
|----------|---------|---------|
| `INGEST_BRAIN_DIR` | `../brain` | Brain directory for ingestion server |
| `INGEST_PORT` | `3300` | Ingestion server port |
| `SUPERVISOR_CONFIG_DIR` | `./config/supervisor` | Skill discipline files |
| `SUPERVISOR_SESSIONS_DIR` | `./brain/sessions` | JSONL session logs |
| `INGEST_BASE_URL` | `http://localhost:3300` | Supervisor → ingestion |
| `INGEST_BASE_URL` | `http://localhost:3300` | Routing pod → brain |
| `LITELLM_BASE_URL` | — | LiteLLM proxy for Tier 2 model routing |
| `SUPERVISOR_MCP_TOKEN` | — | Optional bearer token for the supervisor MCP HTTP endpoint; when empty, no auth is enforced |
| `ROUTING_PORT` | `3210` | Routing pod's listen port |
| `ROUTING_MCP_TOKEN` | — | Optional bearer token for the routing MCP HTTP endpoint |
| `ROUTING_MCP_TOKEN` | — | Optional bearer token; when empty, no auth enforced |
| `BRAIN_URL` | `http://ingestion.supervisor:3300` | Routing pod → brain (in-cluster) |
| `HYPERGUILD_FAST_MODEL` | `koala/qwen35-9b-fast` | Fast model for high-pass-rate skill calls |
| `HYPERGUILD_THINKING_MODEL` | `iguana/gemma4-26b` | Thinking model for low-pass-rate skill calls |
| `HYPERGUILD_ROUTE_LOCAL_FLOOR` | `0.90` | At/above pass rate, route to fast model |
| `HYPERGUILD_ROUTE_LOCAL_CEIL` | `0.70` | Below pass rate, route to thinking model. Between CEIL and FLOOR is the sample band. |
| `HYPERGUILD_ROUTE_LOCAL_FLOOR` | `0.90` | Fast model threshold |
| `HYPERGUILD_ROUTE_LOCAL_CEIL` | `0.70` | Thinking model threshold |
| `HYPERGUILD_PASS_RATE_TTL_SECONDS` | `60` | Per-skill pass-rate cache TTL |
> **Operator note:** LiteLLM at `LITELLM_BASE_URL` must register both `HYPERGUILD_FAST_MODEL` and `HYPERGUILD_THINKING_MODEL` for routing to do useful work. If a model is missing, LiteLLM returns 4xx, the routing pod's fast route fails, the fail-open retry on the thinking model likely also fails (since both are missing), and the only signal is `final_status: "fail"` on `_routing` entries in the brain.
> **Operator note:** LiteLLM at `LITELLM_BASE_URL` must register both `HYPERGUILD_FAST_MODEL`
> and `HYPERGUILD_THINKING_MODEL`. If a model is missing, the fail-open retry also fails and
> the only signal is `final_status: "fail"` on `_routing` entries in the brain.
## Phase 2 (planned)
## Open issues
- `review` skill — structured code review with iron law enforcement
- `debug` skill — hypothesis-driven debugging sessions
- `spec` skill — generates specs from conversations
- `trainer` — extracts SFT/DPO pairs from session logs for fine-tuning
See [issues](http://gitea.d-ma.be/mathias/hyperguild/issues) — key open items:
- **#25** — skills platform overhaul (audit first, then lazy loading + brain feedback loop)
- **#24** — reduce context burn from skill listing
- **#22** — migrate legacy brain notes to wing/hall layout (one-shot script, low risk)
- **#31** — connect routing-mcp to claude.ai as custom connector
+2 -4
View File
@@ -17,8 +17,6 @@ tasks:
cmds: [bash scripts/context-sync.sh claude]
context:sync:agents:
cmds: [bash scripts/context-sync.sh agents]
context:sync:cursor:
cmds: [bash scripts/context-sync.sh cursor]
# ── Development ────────────────────────────────────────────────────────────
@@ -90,12 +88,12 @@ tasks:
cmds:
- task: context:sync
- cmd: |
drift=$(git status --porcelain -- AGENTS.md CLAUDE.md .cursorrules .aider.conventions.md .context/system-prompt.txt 2>/dev/null)
drift=$(git status --porcelain -- AGENTS.md CLAUDE.md .context/system-prompt.txt 2>/dev/null)
if [ -n "$drift" ]; then
echo "ERROR: derived adapters drifted from canonical context." >&2
echo "$drift" >&2
echo "" >&2
echo "Run: git add AGENTS.md CLAUDE.md .cursorrules .aider.conventions.md .context/system-prompt.txt" >&2
echo "Run: git add AGENTS.md CLAUDE.md .context/system-prompt.txt" >&2
echo " git commit -m 'chore: re-sync context adapters'" >&2
exit 1
fi
@@ -0,0 +1,48 @@
{"_meta":true,"note":"Agent-consumer column of brain-MCP intent analysis. consumer_type fixed=autonomous_agent. CAVEAT: canonical schema file brain-intent-extraction.md is NOT present on this host (koala) — only this session's own task prompt references it. The closed intent vocabulary below was RECONSTRUCTED from the task prompt's framing + brain/schema.md. Re-map intent labels if the canonical vocab differs. schema_source=reconstructed on every row.","closed_intent_vocab":["semantic_retrieval","lexical_lookup","check_prior_art","synthesized_answer","store_new_knowledge","update_or_supersede","ingest_raw_source","verify_write_landed","discover_capability","intent_unclear"],"intent_tool_match_values":["match","mismatch","partial"],"corpus":"~/.claude/projects/*/*.jsonl (Claude Code agent transcripts on koala). brain/sessions/*.jsonl empty. agentsquad docs/eval/*.jsonl are code-review eval results, NOT brain calls. No separate Crush logs found. Zero brain calls appear under any mcp__ name with a human typing the call — all brain acts are agent-initiated (CLAUDE.md reflex), so all qualify as autonomous_agent."}
{"id":"a01","session":"tapir-c","ts":"2026-06-?T15:01:51","tool":"brain_query","intent":"check_prior_art","intent_tool_match":"match","workaround":null,"observed_friction":null,"evidence":"single pre-task query 'YouTube Data API captions download ownership limitation timedtext adapter Go' — named-entity lexical lookup, fit BM25 well, no reformulation."}
{"id":"a02","session":"tapir","ts":"2026-06-05T21:44:10","tool":"brain_ingest","intent":"ingest_raw_source","intent_tool_match":"match","workaround":null,"observed_friction":"preceded by ToolSearch select:brain_ingest 14s earlier — tool not ambient, had to be discovered/loaded first.","evidence":"source=tapir-scheduled-discovery-session-2026-06-05, a session learnings dump."}
{"id":"a03","session":"tapir","ts":"2026-06-?T14:00:59","tool":"brain_ingest","intent":"ingest_raw_source","intent_tool_match":"match","workaround":null,"observed_friction":null,"evidence":"source=tapir-rls-identity-bootstrapping, first write of RLS lesson."}
{"id":"a04","session":"tapir","ts":"2026-06-?T14:02:08","tool":"brain_ingest","intent":"update_or_supersede","intent_tool_match":"mismatch","workaround":"RE-INGESTED same source name 'tapir-rls-identity-bootstrapping' 69s later with edited/condensed body. No update/patch/supersede verb exists, so the agent overwrote-by-re-ingest. Whether this dedups or creates a v2 duplicate is opaque to the agent.","observed_friction":"agent revised content within 70s of first write — classic edit-after-write with no edit primitive.","evidence":"two brain_ingest, identical source string, divergent content."}
{"id":"a05","session":"tapir","ts":"2026-06-?T14:56:27","tool":"brain_write","intent":"store_new_knowledge","intent_tool_match":"match","workaround":null,"observed_friction":null,"evidence":"postgres-cascade-skips-tables-without-fk.md — distinct new lesson."}
{"id":"a06","session":"tapir","ts":"2026-06-?T21:08:57","tool":"brain_query","intent":"check_prior_art","intent_tool_match":"match","workaround":null,"observed_friction":"preceded by ToolSearch select:brain_query — discovery tax again.","evidence":"'Dex passwords.dex.coreos.com CRD ...' keyword-rich, single shot."}
{"id":"a07","session":"AI-infra","ts":"2026-05-?T15:38:03","tool":"brain_query(HTTP-curl)","intent":"discover_capability","intent_tool_match":"mismatch","workaround":"raw `curl -X POST` to brain-mcp endpoint instead of MCP tool. Preceded by two ToolSearch ('brain knowledge memory' then 'brain') that did not yield a usable loaded tool, so agent fell back to HTTP.","observed_friction":"3-step ladder: ToolSearch 'brain knowledge memory' -> ToolSearch 'brain' -> curl. Agent did not know which act maps to which tool name.","evidence":"curl -s -o /tmp/brain-init.txt -w code:%{http_code} -X POST ..."}
{"id":"a08","session":"AI-infra","ts":"2026-05-?T04:46:13","tool":"brain_query(HTTP-curl)","intent":"discover_capability","intent_tool_match":"mismatch","workaround":"hand-set TOKEN=... then curl brain-test endpoint — probing whether the HTTP brain path is reachable/authed at all. MCP path not used.","observed_friction":"agent testing connectivity by hand; MCP auth/availability not trusted.","evidence":"TOKEN=...; curl -s -o /tmp/brain-test ..."}
{"id":"a09","session":"AI-infra","ts":"2026-05-?T05:23:35","tool":"brain_query","intent":"check_prior_art","intent_tool_match":"match","workaround":null,"observed_friction":"preceded by ToolSearch select:brain_query (after an earlier 05:03 ToolSearch 'brain ingestion knowledge wiki' that explored layers).","evidence":"'koala machine state RTX 5070 llama-swap' — named-entity recall, fits lexical."}
{"id":"a10","session":"AI-infra","ts":"2026-05-?T05:27:36","tool":"brain_query","intent":"check_prior_art","intent_tool_match":"match","workaround":null,"observed_friction":null,"evidence":"batch of 3 in ~1s (k3s/flux gitops; llama-swap ai-stack GPU; MCP Dex OAuth claude.ai) — parallel prior-art sweep, all named-entity."}
{"id":"a11","session":"AI-infra","ts":"2026-05-?T07:13:50","tool":"brain_query","intent":"check_prior_art","intent_tool_match":"match","workaround":null,"observed_friction":null,"evidence":"batch of 4 in ~2s before a debugging session (flux healthCheck; exit 255 restart loop; NVML mismatch; mirror rebase). Named symptoms, lexical fit OK on first pass."}
{"id":"a12","session":"AI-infra","ts":"2026-05-?T07:14:26","tool":"brain_write","intent":"store_new_knowledge","intent_tool_match":"match","workaround":null,"observed_friction":null,"evidence":"4 writes in ~35s (flux-healthcheck-stale; exit-255-unknown-reason-not-oom; nvidia-nvml-mismatch; mcp-static-bearer) — answers to the 4 queries just run, captured as lessons. Healthy query->fix->write loop."}
{"id":"a13","session":"AI-infra","ts":"2026-05-?T07:15:25","tool":"brain_query","intent":"verify_write_landed","intent_tool_match":"mismatch","workaround":"25s after writing exit-255-unknown-reason-not-oom.md, re-queried 'exit 255 unknown SIGKILL containerd' — reformulated terms (SIGKILL/containerd not in original query 'exit 255 unknown reason restart loop diagnosis'). Either confirming the fresh write is retrievable or re-searching because first lexical query missed. No read-after-write / get-by-id act exists.","observed_friction":"reformulation chain: 'exit 255 unknown reason restart loop diagnosis' -> 'exit 255 unknown SIGKILL containerd'. Same need, different keywords.","evidence":"query at 07:13:51 vs 07:15:25 bracketing the 07:14:37 write."}
{"id":"a14","session":"AI-infra","ts":"2026-05-?T07:22:55","tool":"brain_query","intent":"check_prior_art","intent_tool_match":"match","workaround":null,"observed_friction":null,"evidence":"batch of 5 in ~20s before a homelab security audit (piguard/iguana tailscale; unifi UCG firewall; SOPS age; ingress TLS cert-manager; koala UFW iptables). Named-entity sweep."}
{"id":"a15","session":"AI-infra","ts":"2026-05-?T09:27:53","tool":"brain_write","intent":"store_new_knowledge","intent_tool_match":"match","workaround":null,"observed_friction":null,"evidence":"audit-shortcut-tls-blocks-zero; policy-audit-mode-blocks-nothing — distinct new audit lessons."}
{"id":"a16","session":"AI-infra","ts":"2026-05-?T09:28:22","tool":"brain_write","intent":"store_new_knowledge","intent_tool_match":"match","workaround":null,"observed_friction":null,"evidence":"homelab-security-chains-not-bugs.md FIRST write (worked example: koala 2026-05-13)."}
{"id":"a17","session":"AI-infra","ts":"2026-05-?T10:32:49","tool":"brain_write","intent":"update_or_supersede","intent_tool_match":"mismatch","workaround":"RE-WROTE homelab-security-chains-not-bugs.md ~64min later with a different/expanded worked example (host-user dotfile, over-broad ClusterRole). Same filename, additive revision, no patch/append/supersede verb — agent overwrites and hopes the index replaces rather than duplicates.","observed_friction":"the in-between hour of audit work produced a better example; only way to fold it in was a full re-write of the same slug.","evidence":"two brain_write same filename at 09:28:22 and 10:32:49, divergent worked examples."}
{"id":"a18","session":"AI-infra","ts":"2026-05-?T10:18:25","tool":"brain_write","intent":"store_new_knowledge","intent_tool_match":"match","workaround":null,"observed_friction":null,"evidence":"homelab-document-accepted-risk-to-break-audit-cycle.md — distinct."}
{"id":"a19","session":"AI-infra","ts":"2026-05-?T10:32:55","tool":"brain_query","intent":"verify_write_landed","intent_tool_match":"mismatch","workaround":"6s after the homelab-chains re-write, queried 'RBAC MCP cluster pods log chain' — checking the chain reasoning is retrievable / finding the related entry. Read-after-write done via lexical search.","observed_friction":null,"evidence":"query immediately follows the 10:32:49 write."}
{"id":"a20","session":"AI-infra","ts":"2026-05-?T18:57:34","tool":"brain_query","intent":"check_prior_art","intent_tool_match":"match","workaround":null,"observed_friction":"preceded by ToolSearch select:brain_write,brain_query — re-discovered tools this session.","evidence":"'extension build pinned version major version upgrade postgres pgvector'."}
{"id":"a21","session":"AI-infra","ts":"2026-05-?T18:57:58","tool":"brain_write","intent":"store_new_knowledge","intent_tool_match":"match","workaround":null,"observed_friction":null,"evidence":"extension-version-lags-platform-major-upgrade.md."}
{"id":"a22","session":"AI-infra","ts":"2026-05-?T18:58:06","tool":"brain_query","intent":"verify_write_landed","intent_tool_match":"mismatch","workaround":"8s after writing extension-version-lags, re-queried 'pgvector postgres extension version compile error bump' — reformulated from the 18:57:34 query ('extension build pinned version...'). Lexical re-search to confirm the just-written lesson is findable, with different keyword guess.","observed_friction":"reformulation: 'extension build pinned version major version upgrade postgres pgvector' -> 'pgvector postgres extension version compile error bump'.","evidence":"write at 18:57:58 bracketed by queries 18:57:34 and 18:58:06."}
{"id":"a23","session":"AI-infra","ts":"2026-05-?T18:31:23","tool":"brain_write","intent":"store_new_knowledge","intent_tool_match":"match","workaround":null,"observed_friction":null,"evidence":"webfetch-readme-when-image-or-flag-uncertain.md FIRST write."}
{"id":"a24","session":"AI-infra","ts":"2026-05-?T18:34:23","tool":"brain_write","intent":"update_or_supersede","intent_tool_match":"mismatch","workaround":"RE-WROTE webfetch-readme-when-image-or-flag-uncertain.md 3min later, near-identical body. Looks like a retry/overwrite (uncertain the first landed, or minor edit). No idempotent upsert with confirmation, so agent re-fires the write.","observed_friction":"followed 7s later by a brain_query on the same topic ('OSS tool image registry CLI flag webhook path schema drift README pre-flight') — write-write-query, i.e. overwrite then verify-by-search.","evidence":"two brain_write same filename 18:31:23 / 18:34:23, then query 18:34:31."}
{"id":"a25","session":"AI-infra","ts":"2026-05-?T18:34:31","tool":"brain_query","intent":"verify_write_landed","intent_tool_match":"mismatch","workaround":"keyword-stuffed lexical query 'OSS tool image registry CLI flag webhook path schema drift README pre-flight' fired right after the webfetch-readme write — agent dumps every concept token hoping BM25 surfaces its own fresh note. This is semantic intent (find that conceptual lesson) coerced into a bag-of-keywords.","observed_friction":"query is a concatenation of the note's section headings — a tell that the agent is groping lexically for content it knows by meaning.","evidence":"query text mirrors the just-written note's bullet topics."}
{"id":"a26","session":"dev","ts":"2026-06-?T21:18:26","tool":"brain_answer","intent":"synthesized_answer","intent_tool_match":"match","workaround":null,"observed_friction":"preceded by ToolSearch select:brain_answer.","evidence":"'tapir transcript persistence shared cross-user dedup table RLS isolation ADR-021 ...' -> 22min later a brain_write (acted on the answer). Answer consumed, not re-queried. Healthy."}
{"id":"a27","session":"dev","ts":"2026-06-?T21:40:20","tool":"brain_write","intent":"store_new_knowledge","intent_tool_match":"match","workaround":null,"observed_friction":null,"evidence":"tapir-migration-and-rls-test-infra-gotchas, with wing/hall absent here (flat) — see schema-confusion note a40."}
{"id":"a28","session":"dev","ts":"2026-06-?T05:35:04","tool":"brain_answer","intent":"synthesized_answer","intent_tool_match":"partial","workaround":"asked 'gitea MCP not working workaround file issue via API which token ... how to authenticate gitea API' — a how-do-I question. Next brain act (05:39 query) is a different topic (tapir transcript), so the answer was apparently sufficient OR abandoned; ambiguous.","observed_friction":null,"evidence":"brain_answer then unrelated brain_query 4min later."}
{"id":"a29","session":"dev","ts":"2026-06-?T05:39:54","tool":"brain_query","intent":"check_prior_art","intent_tool_match":"match","workaround":null,"observed_friction":null,"evidence":"'tapir transcript persistence ADR-021 shared non-RLS'."}
{"id":"a30","session":"dev","ts":"2026-06-?T05:57:37","tool":"brain_write","intent":"store_new_knowledge","intent_tool_match":"match","workaround":null,"observed_friction":null,"evidence":"gitea-mcp-per-repo-tools-404-and-rest-fallback FIRST write."}
{"id":"a31","session":"dev","ts":"2026-06-?T05:57:55","tool":"brain_write","intent":"update_or_supersede","intent_tool_match":"mismatch","workaround":"RE-WROTE gitea-mcp-per-repo-tools-404-and-rest-fallback 18s later — overwrite/retry of same slug, no upsert confirmation.","observed_friction":"sub-20s gap = almost certainly a content tweak the agent could not express as an edit.","evidence":"two brain_write same filename 05:57:37 / 05:57:55."}
{"id":"a32","session":"dev","ts":"2026-05-?T11:51:47","tool":"brain_write","intent":"store_new_knowledge","intent_tool_match":"match","workaround":null,"observed_friction":null,"evidence":"infra-litellm-absorption-2026-05-16.md."}
{"id":"a33","session":"dev","ts":"2026-05-?T12:07:04","tool":"brain_query","intent":"check_prior_art","intent_tool_match":"match","workaround":null,"observed_friction":null,"evidence":"batch of 3 (litellm rebuild time piguard; docker compose orphaned volumes; prometheus_client ModuleNotFoundError) — lexical, error-string driven."}
{"id":"a34","session":"dev","ts":"2026-05-?T15:08:15","tool":"brain_answer","intent":"synthesized_answer","intent_tool_match":"mismatch","workaround":"THREE brain_answer at 15:08 (moved compose volumes? / pi rebuild time? / litellm ModuleNotFound prometheus) — the SAME three topics queried lexically an hour earlier (12:07) — were IMMEDIATELY followed at 15:09 by THREE brain_query on the same three topics. The agent asked the synthesizer, was unsatisfied, and fell straight back to raw lexical search. Strongest answer->query fallback in the corpus.","observed_friction":"answer/query duplication across one intent: agent hedges by firing both interfaces, trusting neither.","evidence":"15:08 answers vs 15:09 queries, topic-for-topic aligned."}
{"id":"a35","session":"dev","ts":"2026-05-?T15:09:16","tool":"brain_query","intent":"semantic_retrieval","intent_tool_match":"mismatch","workaround":"after the 3 brain_answer calls failed to satisfy, re-issued as lexical brain_query ('moved compose stack to new directory volumes disappeared empty'; 'raspberry pi docker build time arm slow'; 'how to enable prometheus metrics on litellm proxy callback'). The want is meaning-based ('did my volumes move?') but the only retrieval that 'worked' was keyword search — and these are full natural-language sentences crammed into a BM25 box.","observed_friction":"natural-language questions ('how to enable...', 'moved ... disappeared') passed to a lexical query tool — semantic intent, lexical interface.","evidence":"3 queries at 15:09 mirror the 3 answers at 15:08."}
{"id":"a36","session":"dev","ts":"2026-05-?T20:31:50","tool":"brain_answer","intent":"synthesized_answer","intent_tool_match":"match","workaround":null,"observed_friction":null,"evidence":"'What happened with the litellm migration on 2026-05-16?' — episodic recall question, answer fit; no re-query followed."}
{"id":"a37","session":"dev","ts":"2026-05-?T21:07:17","tool":"brain_query","intent":"semantic_retrieval","intent_tool_match":"mismatch","workaround":"FOUR-step reformulation chain over one Go bug: 'bytes.Buffer Bytes Reset aliasing slice sharing' -> 'go buffer reuse map backing array bug' -> [write go-bytes-buffer-bytes-reset-aliasing-trap.md] -> 'go map values all show same content after loop' -> 'bytes.Buffer Bytes returns same data every iteration'. The agent knows the SYMPTOM (all map values identical) and the CAUSE (Bytes() aliasing) but cannot phrase a single lexical query that bridges them — it wants concept retrieval and is forced to brute-force keyword variants.","observed_friction":"4 distinct phrasings of the same bug, two before and two after writing the lesson — also doubles as verify_write_landed on the trailing queries.","evidence":"21:07:17, 21:07:17, (write 21:08:01), 21:08:10, 21:08:19."}
{"id":"a38","session":"dev","ts":"2026-05-?T21:08:01","tool":"brain_write","intent":"store_new_knowledge","intent_tool_match":"match","workaround":null,"observed_friction":null,"evidence":"go-bytes-buffer-bytes-reset-aliasing-trap.md."}
{"id":"a39","session":"dev","ts":"2026-05-?T07:54:26","tool":"brain_write","intent":"store_new_knowledge","intent_tool_match":"match","workaround":null,"observed_friction":null,"evidence":"mcp-tool-design-get-needs-list-partner.md — a design principle."}
{"id":"a40","session":"dev","ts":"2026-06-?T06:25:00","tool":"brain_query","intent":"check_prior_art","intent_tool_match":"match","workaround":null,"observed_friction":null,"evidence":"'Dex to Authentik migration auth.d-ma.be issuer cutover OIDC subject ...'."}
{"id":"a41","session":"dev","ts":"2026-06-?T06:25:08","tool":"brain_answer","intent":"synthesized_answer","intent_tool_match":"mismatch","workaround":"brain_query (a40) and brain_answer (a41) fired ~8s apart on the SAME intent (Dex->Authentik subject-keyed token orphan). Agent runs lexical search AND synthesized answer in parallel for one question rather than choosing — it cannot predict which interface will return usable knowledge, so it pays both.","observed_friction":"query+answer doublet on one need.","evidence":"06:25:00 query then 06:25:08 answer, same topic."}
{"id":"a42","session":"dev","ts":"2026-06-?T13:48:49","tool":"brain_query","intent":"verify_write_landed","intent_tool_match":"mismatch","workaround":"'authentik cutover validation probe' then 6min later 'authentik cutover post-flip validation' — reformulated pair, likely searching for the agent's own earlier cutover notes / confirming validation steps are recorded. Lexical re-search standing in for recall-my-recent-context.","observed_friction":"reformulation: 'validation probe' -> 'post-flip validation'.","evidence":"13:48:49 and 13:54:56."}
{"id":"a43","session":"dev","ts":"2026-06-?T13:59:17","tool":"brain_write","intent":"store_new_knowledge","intent_tool_match":"match","workaround":null,"observed_friction":"preceded by ToolSearch select:brain_write.","evidence":"oidc-issuer-host-change-vs-idp-swap-subject."}
{"id":"a44","session":"dev","ts":"2026-06-?T13:59:50","tool":"brain_write","intent":"store_new_knowledge","intent_tool_match":"match","workaround":null,"observed_friction":null,"evidence":"cannot-move-ingress-host-across-namespaces-flux-dryrun FIRST write."}
{"id":"a45","session":"dev","ts":"2026-06-?T14:00:13","tool":"brain_write","intent":"update_or_supersede","intent_tool_match":"mismatch","workaround":"RE-WROTE cannot-move-ingress-host-across-namespaces-flux-dryrun 23s later — overwrite of same slug, no edit/upsert primitive.","observed_friction":"sub-30s gap = content correction expressed as a full re-write.","evidence":"two brain_write same filename 13:59:50 / 14:00:13."}
{"id":"a46","session":"dev","ts":"2026-06-?T13:53:54","tool":"brain_write(HTTP-staged)","intent":"store_new_knowledge","intent_tool_match":"mismatch","workaround":"after `ToolSearch select:mcp__claude_ai_brain__authenticate` (MCP auth flow), the agent staged the entry as `cat > /tmp/brain_entry.json` ({filename:'postgres-force-rls-cross-u...', content}) for a curl write rather than calling brain_write directly — MCP write path was not usable (auth/loading), so it dropped to the HTTP bodge.","observed_friction":"reached for an 'authenticate' tool, then abandoned MCP for hand-built JSON + curl. Matches known pattern: brain/op MCP auth lapses often.","evidence":"ToolSearch authenticate 13:53:27 -> cat /tmp/brain_entry.json 13:53:54."}
{"id":"a47","session":"template-go-agent","ts":"2026-05-?T18:46:26","tool":"BASH(not-a-brain-act)","intent":"intent_unclear","intent_tool_match":"match","workaround":null,"observed_friction":null,"evidence":"'brain_' substring was inside a git commit message body ('agent boundaries, network policy, agent s...'), NOT a brain call. Excluded from knowledge-act analysis; logged for audit completeness."}
@@ -0,0 +1,148 @@
# Agent-Consumer Brain Intent Analysis — koala column
**Consumer:** `autonomous_agent` (all rows). **Host:** koala. **Date:** 2026-06-15.
**Raw rows:** `agent-intent-column.jsonl` (46 real knowledge-acts + 1 excluded false-positive).
## Caveat — canonical schema not on this host
The shared closed-vocabulary file `brain-intent-extraction.md` **does not exist on
koala** — the only reference to it is inside *this task's own prompt*. The intent
vocabulary below was **reconstructed** from the prompt's framing + `brain/schema.md`.
Every row carries `schema_source: reconstructed`. If the canonical vocab differs,
re-map the `intent` field; the `intent_tool_match` / `workaround` / `observed_friction`
evidence stands regardless of label names.
**Reconstructed closed vocab:** `semantic_retrieval`, `lexical_lookup`,
`check_prior_art`, `synthesized_answer`, `store_new_knowledge`, `update_or_supersede`,
`ingest_raw_source`, `verify_write_landed`, `discover_capability`, `intent_unclear`.
## Corpus
- `~/.claude/projects/*/*.jsonl` — Claude Code agent transcripts (98 files). **The only
source with brain calls.**
- `brain/sessions/*.jsonl` — empty (only `.gitkeep`).
- `agentsquad docs/eval/*.jsonl` — code-review eval results, **not** brain calls.
- No separate Crush session logs on this host.
- **Zero** brain calls were human-typed. Every brain act is agent-initiated (the
CLAUDE.md "query as reflex / close-the-loop write" behaviour), so all qualify as
`autonomous_agent`. The human gave the top-level task; the agent chose every brain act.
## 1. Intent histogram, split by `intent_tool_match`
| intent | match | mismatch | partial | total |
|---|---|---|---|---|
| check_prior_art | 10 | 0 | 0 | 10 |
| store_new_knowledge | 14 | 1 | 0 | 15 |
| update_or_supersede | 0 | 5 | 0 | 5 |
| synthesized_answer | 2 | 2 | 1 | 5 |
| verify_write_landed | 0 | 4 | 0 | 4 |
| semantic_retrieval | 0 | 3 | 0 | 3 |
| ingest_raw_source | 2 | 0 | 0 | 2 |
| discover_capability | 0 | 2 | 0 | 2 |
| **total** | **28** | **17** | **1** | **46** |
> Batch note: several rows collapse a same-second fan-out of identical-intent calls
> (a10=3, a11=4, a14=5, a33=3, a34=3, a35=3). Call-level the corpus is ~62 brain calls;
> the table counts the 46 distinct knowledge-acts. Frequency is deliberately *not* the
> point — the mismatch column is.
**37% of agent knowledge-acts (17/46) are interface mismatches.** Every mismatch falls
into one of four intents: `update_or_supersede`, `verify_write_landed`,
`semantic_retrieval`, `discover_capability` — plus one `store` that had to use HTTP.
## 2. Mismatch list, grouped by intent (primary deliverable)
### update_or_supersede → re-write same slug (5/5 mismatch) — HIGHEST VALUE
There is **no update / patch / append / supersede verb**. When an agent improves a note
it already wrote, the only move is to call `brain_write`/`brain_ingest` **again with the
same filename/source** and hope the index replaces rather than duplicates. Observed:
| slug | 1st write | 2nd write | gap | what changed |
|---|---|---|---|---|
| `tapir-rls-identity-bootstrapping` (ingest) | 14:00:59 | 14:02:08 | 69s | condensed body |
| `homelab-security-chains-not-bugs.md` | 09:28:22 | 10:32:49 | 64m | new worked example |
| `webfetch-readme-when-image-or-flag-uncertain.md` | 18:31:23 | 18:34:23 | 3m | near-identical (retry) |
| `gitea-mcp-per-repo-tools-404-and-rest-fallback` | 05:57:37 | 05:57:55 | 18s | content tweak |
| `cannot-move-ingress-host-across-namespaces-flux-dryrun` | 13:59:50 | 14:00:13 | 23s | content tweak |
Sub-30s gaps (3 of 5) read as "I wanted to edit but can only overwrite." The agent has
no way to know whether the second write deduped or created a contradictory v2 — opacity
the brain's own design principle (`mcp-tool-design-get-needs-list-partner.md`, written
*by one of these very agents*) would flag: every `_write` needs a `_get`/`_update` partner.
### verify_write_landed → lexical re-query (4/4 mismatch)
No read-after-write / get-by-id confirmation. After every substantive write, agents
re-query lexically to check the note is retrievable — and *reformulate the keywords*
because they can't predict what BM25 indexed:
- `exit-255` lesson: query `exit 255 unknown reason restart loop diagnosis` → write →
query `exit 255 unknown SIGKILL containerd`.
- `extension-version-lags`: query `extension build pinned version...pgvector` → write →
query `pgvector postgres extension version compile error bump`.
- `webfetch-readme`: write → write → query stuffed with the note's own section headings.
### semantic_retrieval → BM25 keyword-stuffing (3/3 mismatch)
Agent knows the *meaning* but not the *indexed words*, so it brute-forces phrasings of
one need against a lexical tool:
- **4-step chain on one Go bug:** `bytes.Buffer Bytes Reset aliasing slice sharing`
`go buffer reuse map backing array bug` → (write) → `go map values all show same
content after loop``bytes.Buffer Bytes returns same data every iteration`. Symptom
and cause both known; no single lexical query bridges them.
- Natural-language questions (`how to enable prometheus metrics on litellm proxy
callback`, `moved compose stack to new directory volumes disappeared empty`) shoved
into `brain_query`.
### synthesized_answer → fall back to / hedge with brain_query (2 mismatch + 1 partial)
`brain_answer` is frequently **not trusted as terminal**:
- **Strongest signal:** 3× `brain_answer` at 15:08 (compose volumes / pi rebuild time /
litellm ModuleNotFound) → 3× `brain_query` at 15:09 on the *same three topics*. The
agent asked the synthesizer, was unsatisfied, and immediately re-ran raw search.
- Dex→Authentik: `brain_query` and `brain_answer` fired **8s apart on one question** —
the agent pays both interfaces because it can't predict which returns usable knowledge.
- (Counter-examples exist: `brain_answer` for episodic recall — "what happened with the
litellm migration on 2026-05-16?" — was consumed and not re-queried. So `answer`
works for *episodic/temporal* recall, fails for *how-do-I / does-X-hold* reasoning.)
### discover_capability + store-via-HTTP (3 mismatch)
brain tools are **not ambient** — they are deferred and must be `ToolSearch`-loaded each
session. Agents fumble the discovery (`ToolSearch 'brain knowledge memory'` →
`'brain'` → `'brain ingestion knowledge wiki'`) and, when MCP load/auth fails, drop to
**raw `curl` against `brain-mcp` / hand-built `/tmp/brain_entry.json`**. One agent even
`ToolSearch`-ed an `authenticate` tool, then abandoned MCP for the HTTP bodge — matching
the known "brain/op MCP auth lapses too often" footgun.
### Write-interface / layer schema confusion (cross-cutting)
`brain_write` was called with **three different param shapes** in the same corpus:
`{filename, type:"lesson", content}`, `{filename, content}` (no type), and
`{wing:"tapir", hall:"failures", filename, content}` — plus `brain_ingest {source,
content}`. Agents are unsure which verb and which layer (flat slug vs `wing`/`hall`
knowledge routing vs raw ingest) a given knowledge-act maps to. This is the
`knowledge/ vs wiki/` confusion expressed at the parameter level.
## 3. `intent_unclear` rate
**0 / 46 genuine brain acts (0%).** Agent intent is unusually legible because these are
Claude Code transcripts: the surrounding task, the query/filename strings, and the
write content all disambiguate. One row (`a47`) was tagged `intent_unclear` and
**excluded** — its `brain_` substring was inside a git commit message, not a brain call.
Example of the only ambiguity that arose: a `brain_answer` on "gitea MCP not working...
how to authenticate" followed by an unrelated query — can't tell if the answer satisfied
or was abandoned (`partial`, row a28).
## 4. The single biggest intent↔interface gap
**The brain offers one write verb and one lexical read verb, but autonomous agents
perform four distinct knowledge-acts against them — and three of the four have no fitting
interface.** The deepest gap is the **missing update/supersede path**: agents close every
task by writing a lesson (the CLAUDE.md ritual), routinely improve it minutes-to-an-hour
later, and — having no edit primitive — re-write the same slug blind, unable to tell
whether they corrected the entry or forked a contradiction into the index. This compounds
with the lexical-only read side: because there is no `get-by-id` or semantic retrieval,
agents can't even reliably *find their own just-written note* to check it, so they
keyword-stuff reformulated queries and hedge `brain_answer` with parallel `brain_query`.
The interface is built for *append-and-keyword-search*; the agents are trying to
*curate a living, deduplicated knowledge base*, and the seam between those two shows up
as the 5 blind re-writes, 4 read-after-write re-queries, and 3 semantic-as-lexical chains
that dominate the mismatch column.
---
*Evidence-only per task scope — no redesign proposed.*
@@ -0,0 +1,140 @@
# Brain-MCP Intent↔Interface Findings — Unified (two-column merge)
**Status — 2026-06-16**
-**Agent column** filled from `agent-intent-column.jsonl` (46 acts, koala).
-**Human column** = `PENDING`. Drop the Claude.ai-history analysis into
`human-intent-column.jsonl` (same dir, schema below), then fill the `PENDING`
cells and the synthesis blocks marked `<<SYNTH>>`.
- ⚠️ Canonical `brain-intent-extraction.md` still absent on koala. Vocab below is
the **reconstructed** lock both columns must share. If the real file surfaces,
re-map `intent` labels in *both* columns identically before merging.
---
## Shared schema (LOCKED — both columns conform)
Per-call row, JSONL:
| field | values / form | notes |
|---|---|---|
| `id` | `a01..` (agent) / `h01..` (human) | column prefix kept distinct |
| `session` | string | source session/conversation id |
| `ts` | ISO-8601 | best-effort |
| `tool` | brain tool name (+ `(HTTP-curl)` / `(HTTP-staged)` suffix for bodges) | |
| `intent` | closed vocab ↓ | the knowledge-act WANTED |
| `intent_tool_match` | `match` \| `mismatch` \| `partial` | does the called tool fit the want |
| `consumer_type` | `autonomous_agent` \| `human_interactive` | fixed per column |
| `workaround` | string \| null | the bodge when mismatch — **primary signal** |
| `observed_friction` | string \| null | reformulation chains, discovery tax, hedging |
| `evidence` | string | excerpt anchoring the classification |
| `schema_source` | `reconstructed` | flip to `canonical` if real vocab lands |
### Closed intent vocab (LOCKED)
`semantic_retrieval`, `lexical_lookup`, `check_prior_art`, `synthesized_answer`,
`store_new_knowledge`, `update_or_supersede`, `ingest_raw_source`,
`verify_write_landed`, `discover_capability`, `intent_unclear`.
---
## Master comparison — by intent
| intent | agent acts | agent mismatch | human acts | human mismatch | shared gap |
|---|---|---|---|---|---|
| check_prior_art | 10 | 0% | `PENDING` | `PENDING` | — |
| store_new_knowledge | 15 | 7% (1/15) | `PENDING` | `PENDING` | `<<SYNTH>>` |
| update_or_supersede | 5 | **100%** (5/5) | `PENDING` | `PENDING` | `<<SYNTH>>` no edit verb |
| synthesized_answer | 5 | 40% (2/5)+1 partial | `PENDING` | `PENDING` | `<<SYNTH>>` |
| verify_write_landed | 4 | **100%** (4/4) | `PENDING` | `PENDING` | `<<SYNTH>>` no read-after-write |
| semantic_retrieval | 3 | **100%** (3/3) | `PENDING` | `PENDING` | `<<SYNTH>>` lexical-only read |
| ingest_raw_source | 2 | 0% | `PENDING` | `PENDING` | — |
| discover_capability | 2 | **100%** (2/2) | `PENDING` | `PENDING` | agent-specific (ToolSearch/auth)? |
| intent_unclear | 0 | — | `PENDING` | `PENDING` | divergence expected ↓ |
| **TOTAL** | **46** | **37% (17)** | `PENDING` | `PENDING` | |
---
## Per-intent merged findings
### update_or_supersede — agent: 5/5 mismatch (highest value)
**Agent:** no edit/patch/append verb. Agents re-write same slug blind:
`homelab-security-chains-not-bugs.md` (+64m), `tapir-rls-identity-bootstrapping`,
`webfetch-readme...`, `gitea-mcp-per-repo-tools-404...`,
`cannot-move-ingress-host...` — 3 of 5 sub-30s ("wanted edit, got overwrite").
Cannot tell if write deduped or forked a contradiction.
**Human:** `PENDING` — *look for: user editing a prior note, asking "update what I
saved about X", or expressing frustration that an old fact is stale/duplicated.*
**<<SYNTH>>** shared verdict once both filled.
### verify_write_landed — agent: 4/4 mismatch
**Agent:** no `get-by-id`/read-after-write. Agents lexically re-query their own
fresh note with reformulated keywords (`exit 255 unknown reason``...SIGKILL
containerd`; `extension build pinned...``pgvector ...compile error bump`).
**Human:** `PENDING` — *humans may not exhibit this (they trust the write UI
confirmation). If absent in human column, it's an agent-specific gap → flag.*
**<<SYNTH>>**.
### semantic_retrieval — agent: 3/3 mismatch
**Agent:** meaning known, indexed words unknown → BM25 keyword-stuffing. 4-step
chain on one Go `bytes.Buffer` bug; NL questions shoved into `brain_query`.
**Human:** `PENDING` — *humans likely hit this HARDER (they phrase conversationally).
Compare reformulation-chain length agent vs human.*
**<<SYNTH>>** — likely the strongest cross-consumer overlap.
### synthesized_answer — agent: 2 mismatch + 1 partial
**Agent:** `brain_answer` not trusted terminal — 3 answers → 3 same-topic queries
1min later; query+answer fired 8s apart hedging one need. Works for *episodic*
recall, fails for *how-do-I / does-X-hold*.
**Human:** `PENDING` — *humans may prefer `brain_answer` as primary (chat-native).
If human match-rate >> agent, the tool fits humans not agents → key divergence.*
**<<SYNTH>>**.
### store_new_knowledge — agent: 14/15 match
**Agent:** healthy, except 1 HTTP-staged bodge when MCP auth lapsed. Also surfaced
write-schema confusion: 3 param shapes (`{filename,type}` / `{filename}` /
`{wing,hall,filename}`) + `ingest{source}`.
**Human:** `PENDING`*humans rarely write directly; expect low volume.*
**<<SYNTH>>**.
### check_prior_art / ingest_raw_source — agent: 0% mismatch
Lexical fits named-entity recall and raw-source capture. **Human:** `PENDING`.
### discover_capability — agent: 2/2 mismatch (agent-specific)
Brain tools deferred → `ToolSearch`-load each session; auth lapse → `curl` bodge.
**Likely has NO human analog** (humans get ambient connectors). Candidate for
"agent-only gap" bucket. **Human:** `PENDING` to confirm absent.
---
## Cross-consumer divergence — questions to resolve at merge
1. **intent_unclear rate.** Agent = 0% (transcripts self-document). Human expected
higher (conversational, implicit). Big delta = the columns measure legibility
differently, not just intent.
2. **Where does each consumer's mismatch concentrate?** Agent mismatch is
write-side-heavy (supersede + verify-landed = 9/17). Hypothesis: human mismatch
is read-side-heavy (semantic + answer). If true → **the interface fails the two
consumers at opposite ends.**
3. **Agent-only gaps** (`discover_capability`, `verify_write_landed`) vs
**shared gaps** (`semantic_retrieval`, `update_or_supersede`). Shared gaps =
highest-priority evidence; agent-only = harness/auth issues.
---
## Combined headline — `<<SYNTH>>` (fill when human column lands)
> Agent-side draft (to be reconciled with human-side):
> Brain = append + keyword-search; agents want a curated, dedup'd, self-verifying KB.
> Missing update/supersede path + lexical-only reads are the seam. **Open question
> for the merge: do humans hit the same read-side wall, making semantic-retrieval the
> universal gap — or do agents uniquely suffer the write-side (supersede / verify)
> wall that humans sidestep via the chat UI?**
---
## Drop-in checklist (when human column arrives)
1. Place `human-intent-column.jsonl` in this dir; conform to LOCKED schema.
2. Fill every `PENDING` cell in master table + per-intent blocks.
3. Resolve the 3 divergence questions with evidence.
4. Replace each `<<SYNTH>>` with the reconciled verdict; write the combined headline.
5. If canonical vocab surfaced: re-map both columns' `intent`, flip `schema_source`.
6. Commit as `docs(brain): merge human+agent intent columns`.
+23
View File
@@ -17,6 +17,7 @@ import (
"github.com/mathiasbq/hyperguild/ingestion/internal/api"
"github.com/mathiasbq/hyperguild/ingestion/internal/claudewatcher"
"github.com/mathiasbq/hyperguild/ingestion/internal/embed"
"github.com/mathiasbq/hyperguild/ingestion/internal/gitea"
"github.com/mathiasbq/hyperguild/ingestion/internal/graphstore"
"github.com/mathiasbq/hyperguild/ingestion/internal/graphsync"
"github.com/mathiasbq/hyperguild/ingestion/internal/llm"
@@ -175,6 +176,15 @@ func main() {
logger.Info("brain reranker configured", "url", rerankURL, "model", rerankModel)
}
// Gitea ticket tracker for the capture capability (#52). Token via env
// only — never logged or in argv. Both vars must be set to enable it;
// gitea.New returns nil otherwise, leaving ticket integration off.
giteaURL := envOr("BRAIN_GITEA_URL", "https://git.d-ma.be")
if tracker := gitea.New(giteaURL, os.Getenv("BRAIN_GITEA_TOKEN")); tracker != nil {
mcpSrv = mcpSrv.WithIssueTracker(tracker)
logger.Info("brain gitea tracker configured", "url", giteaURL)
}
// Hybrid retrieval (pgvector + nomic-embed-text). Both env vars must
// be set together for the path to wire on; otherwise BM25-only.
var vectorStore *vectorstore.PGStore
@@ -256,6 +266,19 @@ func main() {
logger.Error("CLAUDE_SESSIONS_DIR set but BRAIN_PG_DSN missing — claudewatcher needs the cursor table")
os.Exit(1)
}
// Client-name guard. The env value is a regex alternation
// (e.g. "SEB|Mastercard"); we wrap it with word boundaries
// and case-insensitive flag so substrings inside longer
// identifiers don't false-match. Sourced from a SOPS secret
// so client identities never live in source.
if clientBlock := os.Getenv("CLAUDE_INGEST_CLIENT_BLOCK"); clientBlock != "" {
pattern := `(?i)\b(` + clientBlock + `)\b`
if err := claudewatcher.RegisterRule("client-name", pattern); err != nil {
logger.Error("claudewatcher client-block rule invalid", "err", err)
os.Exit(1)
}
logger.Info("claudewatcher client-block guard registered")
}
cursorStore, cerr := claudewatcher.NewCursorStore(ctx, pgDSN)
if cerr != nil {
logger.Error("claudewatcher cursor init", "err", cerr)
+97
View File
@@ -0,0 +1,97 @@
package api
import "strings"
// frontmatter is an ordered, line-preserving view of a note's YAML
// frontmatter block. It deliberately avoids a full YAML round-trip: the
// brain writes flat `key: value` frontmatter by hand, and a yaml.v3
// re-marshal would reorder keys and strip comments. Preserving the
// original lines verbatim keeps brain_update a surgical edit — only the
// keys it manages (updated_at, supersedes, supersede_reason) change.
type frontmatter struct {
lines []fmLine
}
// fmLine is one frontmatter line. For `key: value` lines, key and value
// are populated; for blank lines, comments, or anything that isn't a
// simple scalar pair, key is empty and raw holds the line verbatim.
type fmLine struct {
key string
value string
raw string
}
// parseFrontmatter splits src into its frontmatter block and body. A
// frontmatter block is recognised only when the file opens with a `---`
// fence and a closing `---` fence follows. Otherwise the whole input is
// the body and the returned frontmatter is empty.
func parseFrontmatter(src string) (frontmatter, string) {
var fm frontmatter
if !strings.HasPrefix(src, "---\n") {
return fm, src
}
rest := src[len("---\n"):]
end := strings.Index(rest, "\n---\n")
if end < 0 {
// Opening fence with no closing fence — treat as bodyless content.
return fm, src
}
block := rest[:end]
body := rest[end+len("\n---\n"):]
for _, line := range strings.Split(block, "\n") {
key, val, ok := strings.Cut(line, ":")
key = strings.TrimSpace(key)
if !ok || key == "" || strings.HasPrefix(strings.TrimSpace(line), "#") {
fm.lines = append(fm.lines, fmLine{raw: line})
continue
}
fm.lines = append(fm.lines, fmLine{key: key, value: strings.TrimSpace(val)})
}
return fm, body
}
// get returns the value for key, or "" if absent.
func (f *frontmatter) get(key string) string {
for _, l := range f.lines {
if l.key == key {
return l.value
}
}
return ""
}
// set overrides the value for an existing key in place, or appends a new
// `key: value` line when the key is absent.
func (f *frontmatter) set(key, value string) {
for i := range f.lines {
if f.lines[i].key == key {
f.lines[i].value = value
return
}
}
f.lines = append(f.lines, fmLine{key: key, value: value})
}
// render serialises the frontmatter back into a `---`-fenced block. An
// empty frontmatter renders to the empty string so bodies without a
// header stay header-less.
func (f *frontmatter) render() string {
if len(f.lines) == 0 {
return ""
}
var b strings.Builder
b.WriteString("---\n")
for _, l := range f.lines {
if l.key == "" {
b.WriteString(l.raw)
} else {
b.WriteString(l.key)
b.WriteString(": ")
b.WriteString(l.value)
}
b.WriteByte('\n')
}
b.WriteString("---\n")
return b.String()
}
@@ -0,0 +1,61 @@
package api
import (
"strings"
"testing"
"github.com/stretchr/testify/assert"
)
func TestParseFrontmatterSplitsHeaderAndBody(t *testing.T) {
src := "---\nwing: jepa-fx\nhall: facts\ncreated_at: 2026-01-01T00:00:00Z\n---\n# Title\n\nbody text\n"
fm, body := parseFrontmatter(src)
assert.Equal(t, "jepa-fx", fm.get("wing"))
assert.Equal(t, "facts", fm.get("hall"))
assert.Equal(t, "2026-01-01T00:00:00Z", fm.get("created_at"))
assert.Equal(t, "# Title\n\nbody text\n", body)
}
func TestParseFrontmatterNoHeader(t *testing.T) {
src := "# Just a body\n\nno frontmatter here\n"
fm, body := parseFrontmatter(src)
assert.Empty(t, fm.lines)
assert.Equal(t, src, body)
}
func TestFrontmatterSetOverridesExistingKey(t *testing.T) {
fm, _ := parseFrontmatter("---\nwing: a\nupdated_at: old\n---\nbody\n")
fm.set("updated_at", "new")
assert.Equal(t, "new", fm.get("updated_at"))
// No duplicate key.
assert.Equal(t, 1, strings.Count(fm.render(), "updated_at:"))
}
func TestFrontmatterSetAppendsNewKey(t *testing.T) {
fm, _ := parseFrontmatter("---\nwing: a\n---\nbody\n")
fm.set("supersedes", "abc123")
out := fm.render()
assert.Contains(t, out, "wing: a")
assert.Contains(t, out, "supersedes: abc123")
}
func TestFrontmatterRenderPreservesCustomFields(t *testing.T) {
src := "---\nwing: a\nhall: facts\ncustom_field: keep-me\ntags: [x, y]\n---\nbody\n"
fm, _ := parseFrontmatter(src)
fm.set("updated_at", "2026-06-22T00:00:00Z")
out := fm.render()
assert.Contains(t, out, "custom_field: keep-me")
assert.Contains(t, out, "tags: [x, y]")
assert.Contains(t, out, "updated_at: 2026-06-22T00:00:00Z")
}
func TestFrontmatterRenderRoundTrips(t *testing.T) {
src := "---\nwing: a\nhall: facts\n---\n"
fm, _ := parseFrontmatter(src)
assert.Equal(t, src, fm.render())
}
+131
View File
@@ -0,0 +1,131 @@
package api
import (
"crypto/sha256"
"encoding/hex"
"fmt"
"os"
"path/filepath"
"strings"
"time"
"github.com/mathiasbq/hyperguild/ingestion/internal/brain"
)
// ContentHash returns the lowercase hex sha256 of b. It is the note's
// content_hash handle: brain_write / brain_update return it, brain_get
// recomputes it from the file on disk, and brain_update stamps the prior
// note's hash into the new note's `supersedes` frontmatter.
func ContentHash(b []byte) string {
sum := sha256.Sum256(b)
return hex.EncodeToString(sum[:])
}
// resolveWithin maps a brainDir-relative path to an absolute path and
// guarantees it does not escape brainDir. Returns the cleaned relPath
// (forward-slashed) and the absolute path.
func resolveWithin(brainDir, relPath string) (rel, abs string, err error) {
clean := filepath.Clean("/" + filepath.ToSlash(relPath))
rel = strings.TrimPrefix(clean, "/")
abs = filepath.Join(brainDir, filepath.FromSlash(rel))
check, err := filepath.Rel(brainDir, abs)
if err != nil || check == ".." || strings.HasPrefix(check, ".."+string(filepath.Separator)) {
return "", "", fmt.Errorf("path %q escapes brain dir", relPath)
}
return rel, abs, nil
}
// UpdateNoteOptions identifies the note to supersede and supplies its new
// body. Path takes precedence; otherwise the target is resolved from
// Wing/Hall/Slug via brain.NotePath.
type UpdateNoteOptions struct {
Path string // brainDir-relative path; takes precedence over wing/hall/slug
Wing string
Hall string
Slug string
Content string // new full body (whole-note replace)
Reason string // optional; stamped as supersede_reason
}
// UpdateNote supersedes an existing note in place. It replaces the body
// with opts.Content, preserves the existing frontmatter (created_at,
// wing, hall, and any custom fields), and stamps updated_at, supersedes
// (the prior content hash), and supersede_reason (when given).
//
// It never creates: if the target does not exist, it returns an error so
// the caller can fall back to brain_write. Returns the note's relPath,
// the new content hash, and the prior content hash.
//
// Embeddings are NOT refreshed here. The rewritten file's mtime advances,
// which the mtime-driven vectorstore.Sync ticker uses to re-embed it on
// its next pass — the same out-of-band mechanism brain_write relies on.
func UpdateNote(brainDir string, opts UpdateNoteOptions) (relPath, contentHash, priorHash string, err error) {
if opts.Content == "" {
return "", "", "", fmt.Errorf("content is required")
}
var rel string
if opts.Path != "" {
rel = opts.Path
} else {
full, perr := brain.NotePath(brainDir, opts.Wing, opts.Hall, opts.Slug)
if perr != nil {
return "", "", "", perr
}
rel, _ = filepath.Rel(brainDir, full)
rel = filepath.ToSlash(rel)
}
rel, abs, err := resolveWithin(brainDir, rel)
if err != nil {
return "", "", "", err
}
prior, err := os.ReadFile(abs)
if err != nil {
if os.IsNotExist(err) {
return "", "", "", fmt.Errorf("note %q does not exist: use brain_write to create", rel)
}
return "", "", "", fmt.Errorf("read target: %w", err)
}
priorHash = ContentHash(prior)
fm, _ := parseFrontmatter(string(prior))
fm.set("updated_at", time.Now().UTC().Format(time.RFC3339))
fm.set("supersedes", priorHash)
if opts.Reason != "" {
fm.set("supersede_reason", opts.Reason)
}
out := []byte(fm.render() + opts.Content)
if err := os.WriteFile(abs, out, 0o644); err != nil {
return "", "", "", fmt.Errorf("write: %w", err)
}
return rel, ContentHash(out), priorHash, nil
}
// ReadNote reads the note at the brainDir-relative relPath and returns
// its parsed frontmatter, body, and content hash. It is the read-after-
// write primitive behind brain_get: the hash it returns equals the hash
// brain_write / brain_update returned for the same bytes.
func ReadNote(brainDir, relPath string) (fm map[string]string, body, contentHash string, err error) {
_, abs, err := resolveWithin(brainDir, relPath)
if err != nil {
return nil, "", "", err
}
raw, err := os.ReadFile(abs)
if err != nil {
if os.IsNotExist(err) {
return nil, "", "", fmt.Errorf("note %q does not exist", relPath)
}
return nil, "", "", fmt.Errorf("read note: %w", err)
}
parsed, body := parseFrontmatter(string(raw))
fm = make(map[string]string, len(parsed.lines))
for _, l := range parsed.lines {
if l.key != "" {
fm[l.key] = l.value
}
}
return fm, body, ContentHash(raw), nil
}
+130
View File
@@ -0,0 +1,130 @@
package api
import (
"os"
"path/filepath"
"strings"
"testing"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
)
// seedNote writes a note directly to disk and returns its relPath.
func seedNote(t *testing.T, brainDir, rel, content string) string {
t.Helper()
full := filepath.Join(brainDir, filepath.FromSlash(rel))
require.NoError(t, os.MkdirAll(filepath.Dir(full), 0o755))
require.NoError(t, os.WriteFile(full, []byte(content), 0o644))
return rel
}
func TestUpdateNoteSupersedesAndStamps(t *testing.T) {
brainDir := t.TempDir()
rel := seedNote(t, brainDir, "wiki/jepa-fx/facts/val-vol.md",
"---\nwing: jepa-fx\nhall: facts\ncreated_at: 2026-01-01T00:00:00Z\ncustom: keep-me\n---\n# Old\n\nold body\n")
relPath, hash, priorHash, err := UpdateNote(brainDir, UpdateNoteOptions{
Path: rel,
Content: "# New\n\nnew body\n",
Reason: "facts changed",
})
require.NoError(t, err)
assert.Equal(t, rel, relPath)
assert.NotEmpty(t, hash)
assert.NotEmpty(t, priorHash)
assert.NotEqual(t, hash, priorHash)
got, err := os.ReadFile(filepath.Join(brainDir, filepath.FromSlash(rel)))
require.NoError(t, err)
s := string(got)
// Body replaced.
assert.Contains(t, s, "# New")
assert.NotContains(t, s, "old body")
// Prior fields preserved.
assert.Contains(t, s, "wing: jepa-fx")
assert.Contains(t, s, "hall: facts")
assert.Contains(t, s, "created_at: 2026-01-01T00:00:00Z")
assert.Contains(t, s, "custom: keep-me")
// Supersession stamped.
assert.Contains(t, s, "updated_at:")
assert.Contains(t, s, "supersedes: "+priorHash)
assert.Contains(t, s, "supersede_reason: facts changed")
}
func TestUpdateNoteResolvesByWingHallSlug(t *testing.T) {
brainDir := t.TempDir()
seedNote(t, brainDir, "wiki/jepa-fx/facts/val-vol.md",
"---\nwing: jepa-fx\nhall: facts\n---\nold\n")
relPath, _, _, err := UpdateNote(brainDir, UpdateNoteOptions{
Wing: "jepa-fx", Hall: "facts", Slug: "val-vol",
Content: "new\n",
})
require.NoError(t, err)
assert.Equal(t, "wiki/jepa-fx/facts/val-vol.md", relPath)
}
func TestUpdateNoteErrorsOnMissingAndDoesNotCreate(t *testing.T) {
brainDir := t.TempDir()
_, _, _, err := UpdateNote(brainDir, UpdateNoteOptions{
Wing: "jepa-fx", Hall: "facts", Slug: "ghost",
Content: "x\n",
})
require.Error(t, err)
assert.Contains(t, err.Error(), "does not exist")
// No file created.
_, statErr := os.Stat(filepath.Join(brainDir, "wiki/jepa-fx/facts/ghost.md"))
assert.True(t, os.IsNotExist(statErr), "missing-target update must not create a note")
}
func TestUpdateNoteRejectsTraversal(t *testing.T) {
brainDir := t.TempDir()
_, _, _, err := UpdateNote(brainDir, UpdateNoteOptions{
Path: "../escape.md",
Content: "x\n",
})
require.Error(t, err)
}
func TestReadNoteReturnsFrontmatterBodyHash(t *testing.T) {
brainDir := t.TempDir()
rel := seedNote(t, brainDir, "wiki/jepa-fx/facts/n.md",
"---\nwing: jepa-fx\nhall: facts\n---\n# Body\n\ntext\n")
fm, body, hash, err := ReadNote(brainDir, rel)
require.NoError(t, err)
assert.Equal(t, "jepa-fx", fm["wing"])
assert.Equal(t, "facts", fm["hall"])
assert.Equal(t, "# Body\n\ntext\n", body)
// Hash matches ContentHash of the raw bytes on disk (round-trip).
raw, _ := os.ReadFile(filepath.Join(brainDir, filepath.FromSlash(rel)))
assert.Equal(t, ContentHash(raw), hash)
}
func TestReadNoteRejectsTraversal(t *testing.T) {
brainDir := t.TempDir()
_, _, _, err := ReadNote(brainDir, "../../etc/passwd")
require.Error(t, err)
}
func TestUpdateThenReadRoundTripsHash(t *testing.T) {
brainDir := t.TempDir()
rel := seedNote(t, brainDir, "wiki/a/facts/n.md", "---\nwing: a\nhall: facts\n---\nold\n")
_, hash, _, err := UpdateNote(brainDir, UpdateNoteOptions{Path: rel, Content: "new\n"})
require.NoError(t, err)
_, _, readHash, err := ReadNote(brainDir, rel)
require.NoError(t, err)
assert.Equal(t, hash, readHash, "update content_hash must round-trip through ReadNote")
}
func TestContentHashStable(t *testing.T) {
assert.Equal(t, ContentHash([]byte("abc")), ContentHash([]byte("abc")))
assert.NotEqual(t, ContentHash([]byte("abc")), ContentHash([]byte("abd")))
assert.True(t, strings.HasPrefix(ContentHash([]byte("")), "")) // hex, non-panicking
}
+129
View File
@@ -0,0 +1,129 @@
// Package brainstore is the concrete BrainStore: the single shared
// implementation of the #45 write/update/get verbs, used by BOTH the MCP
// handlers and the capture use-case so there is one implementation, not
// two (the Clean-Architecture / DRY payoff of #51).
//
// It composes the file-level primitives in package api (WriteNote,
// UpdateNote, ReadNote — the read-after-write contract) with the wiki
// upkeep that must accompany a write: wing _index rebuild, cross-wing
// auto-tunnel, and graph re-index. Embedding refresh is intentionally
// out-of-band (mtime-driven vectorstore.Sync) and not triggered here —
// see the brain note on out-of-band sync.
package brainstore
import (
"context"
"log/slog"
"strings"
"github.com/mathiasbq/hyperguild/ingestion/internal/api"
"github.com/mathiasbq/hyperguild/ingestion/internal/brain"
"github.com/mathiasbq/hyperguild/ingestion/internal/capture"
"github.com/mathiasbq/hyperguild/ingestion/internal/graphsync"
)
// Store implements capture.BrainStore against a brain directory on disk,
// optionally re-indexing each write into the knowledge graph.
type Store struct {
brainDir string
graph graphsync.Store // nil = graph re-index disabled
}
// New constructs a Store bound to brainDir with graph indexing disabled.
func New(brainDir string) *Store {
return &Store{brainDir: brainDir}
}
// WithGraph enables graph re-index on every write/update. nil disables it.
func (s *Store) WithGraph(g graphsync.Store) *Store {
s.graph = g
return s
}
// Write creates a brain note and returns its read-after-write handle.
func (s *Store) Write(ctx context.Context, n capture.Note) (capture.Ref, error) {
relPath, err := api.WriteNote(s.brainDir, api.WriteNoteOptions{
Content: n.Content,
Filename: n.Filename,
Type: n.Type,
Domain: n.Domain,
Wing: n.Wing,
Hall: n.Hall,
})
if err != nil {
return capture.Ref{}, err
}
s.wikiUpkeep(relPath, n.Wing, n.Content)
s.indexInGraph(ctx, "brain_write", relPath)
_, _, hash, _ := api.ReadNote(s.brainDir, relPath)
return capture.Ref{ID: relPath, Path: relPath, ContentHash: hash}, nil
}
// Update supersedes an existing note in place. slug may be a bare slug
// (resolved against n.Wing/n.Hall) or a full brain-relative path (when it
// contains a slash). It never creates — a missing target is an error.
func (s *Store) Update(ctx context.Context, slug string, n capture.Note) (capture.Ref, error) {
opts := api.UpdateNoteOptions{Content: n.Content, Reason: n.Reason}
if strings.Contains(slug, "/") {
opts.Path = slug
} else {
opts.Wing, opts.Hall, opts.Slug = n.Wing, n.Hall, slug
}
relPath, hash, _, err := api.UpdateNote(s.brainDir, opts)
if err != nil {
return capture.Ref{}, err
}
if wing := wingFromRelPath(relPath); wing != "" {
s.wikiUpkeep(relPath, wing, n.Content)
}
s.indexInGraph(ctx, "brain_update", relPath)
return capture.Ref{ID: relPath, Path: relPath, ContentHash: hash, Superseded: true}, nil
}
// Get fetches a note by id/path — the read-after-write confirmation
// primitive (a direct fetch, never a semantic query).
func (s *Store) Get(_ context.Context, id string) (capture.StoredNote, error) {
fm, body, hash, err := api.ReadNote(s.brainDir, id)
if err != nil {
return capture.StoredNote{}, err
}
return capture.StoredNote{ID: id, Path: id, ContentHash: hash, Frontmatter: fm, Body: body}, nil
}
// wikiUpkeep rebuilds the wing _index and re-tunnels cross-wing matches
// when a note lands in the structured wiki. Both are best-effort: the
// note is already written, so a failure here is logged, not propagated.
func (s *Store) wikiUpkeep(relPath, wing, content string) {
if wing == "" {
return
}
if err := brain.BuildWingIndex(s.brainDir, wing); err != nil {
slog.Warn("brainstore: auto-index failed", "wing", wing, "err", err)
}
if err := brain.AutoTunnel(s.brainDir, relPath, content); err != nil {
slog.Warn("brainstore: auto-tunnel failed", "src", relPath, "err", err)
}
}
// indexInGraph re-indexes a written doc into the graph, best-effort.
func (s *Store) indexInGraph(ctx context.Context, op, relPath string) {
if s.graph == nil || relPath == "" {
return
}
if err := graphsync.IndexDoc(ctx, s.graph, s.brainDir, relPath); err != nil {
slog.Warn(op+": graph index failed", "path", relPath, "err", err)
}
}
// wingFromRelPath extracts the wing from a structured wiki path
// (wiki/<wing>/<hall>/<slug>.md). Returns "" for legacy/non-wiki paths.
func wingFromRelPath(relPath string) string {
parts := strings.Split(relPath, "/")
if len(parts) >= 4 && parts[0] == "wiki" {
return parts[1]
}
return ""
}
@@ -0,0 +1,85 @@
package brainstore_test
import (
"context"
"os"
"path/filepath"
"testing"
"github.com/mathiasbq/hyperguild/ingestion/internal/brainstore"
"github.com/mathiasbq/hyperguild/ingestion/internal/capture"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
)
func TestStoreWriteReturnsHandle(t *testing.T) {
dir := t.TempDir()
s := brainstore.New(dir)
ref, err := s.Write(context.Background(), capture.Note{
Content: "# X\n\nbody\n", Filename: "x", Wing: "a", Hall: "facts",
})
require.NoError(t, err)
assert.Equal(t, "wiki/a/facts/x.md", ref.Path)
assert.Equal(t, ref.Path, ref.ID)
assert.NotEmpty(t, ref.ContentHash)
assert.False(t, ref.Superseded)
_, err = os.Stat(filepath.Join(dir, "wiki/a/facts/x.md"))
require.NoError(t, err)
}
func TestStoreUpdateSupersedes(t *testing.T) {
dir := t.TempDir()
s := brainstore.New(dir)
_, err := s.Write(context.Background(), capture.Note{
Content: "old\n", Filename: "n", Wing: "a", Hall: "facts",
})
require.NoError(t, err)
ref, err := s.Update(context.Background(), "n", capture.Note{
Content: "new\n", Wing: "a", Hall: "facts", Reason: "changed",
})
require.NoError(t, err)
assert.True(t, ref.Superseded)
assert.Equal(t, "wiki/a/facts/n.md", ref.Path)
got, _ := os.ReadFile(filepath.Join(dir, "wiki/a/facts/n.md"))
assert.Contains(t, string(got), "new")
assert.Contains(t, string(got), "supersede_reason: changed")
}
func TestStoreUpdateByFullPath(t *testing.T) {
dir := t.TempDir()
s := brainstore.New(dir)
_, err := s.Write(context.Background(), capture.Note{Content: "old\n", Filename: "n", Wing: "a", Hall: "facts"})
require.NoError(t, err)
ref, err := s.Update(context.Background(), "wiki/a/facts/n.md", capture.Note{Content: "fresh\n"})
require.NoError(t, err)
assert.Equal(t, "wiki/a/facts/n.md", ref.Path)
}
func TestStoreUpdateMissingErrors(t *testing.T) {
dir := t.TempDir()
s := brainstore.New(dir)
_, err := s.Update(context.Background(), "ghost", capture.Note{Content: "x\n", Wing: "a", Hall: "facts"})
require.Error(t, err)
_, statErr := os.Stat(filepath.Join(dir, "wiki/a/facts/ghost.md"))
assert.True(t, os.IsNotExist(statErr), "update must not create")
}
func TestStoreGetRoundTripsHash(t *testing.T) {
dir := t.TempDir()
s := brainstore.New(dir)
ref, err := s.Write(context.Background(), capture.Note{
Content: "# Body\n\ntext\n", Filename: "n", Wing: "a", Hall: "facts",
})
require.NoError(t, err)
note, err := s.Get(context.Background(), ref.ID)
require.NoError(t, err)
assert.Equal(t, ref.ContentHash, note.ContentHash, "write→get hash round-trips")
assert.Equal(t, "a", note.Frontmatter["wing"])
assert.Contains(t, note.Body, "# Body")
}
+111
View File
@@ -0,0 +1,111 @@
// Package capture is the Clean-Architecture use-case for the uniform
// capture capability (issue #49/#51): persist a finished session's
// valuable output — insights → brain, action items → Gitea tickets,
// optional summary → ai-sessions — with one invocation, identical core
// behaviour across every harness.
//
// This package is pure orchestration. It depends only on ports
// (interfaces) and plain entities — no HTTP, no live Gitea, no embedding
// or audit I/O. The real adapters are wired in #52 (Gitea tracker), #53
// (REST + I1 origin gate), and #54/#55 (audit path + relay). The I1
// sovereignty refusal and the classification-aware audit degradation are
// deliberately NOT here — those need the server-derived principal origin
// (#53) and the loki/buffer machinery (#54). What lives here is everything
// testable against fakes: validation, effective-classification resolution
// (stricter wins), best-effort orchestration, and the partial receipt.
package capture
// CaptureContext is the per-session metadata accompanying a capture.
//
// Classification is the caller-declared sensitivity (model C, spec §4.1):
// the server independently derives the target's classification and gates
// on the stricter of the two. Principal is server-derived from the
// authenticated identity (#53 populates it); it is never caller-asserted.
// Harness is descriptive telemetry only — never a gate input.
type CaptureContext struct {
Harness string
SessionRef string
Fidelity string
Actor string
Classification string // caller-declared level token ("" = unspecified)
Principal string // server-derived (auth); audit identity
}
// Insight is one piece of session knowledge bound for the brain. A
// non-empty SupersedeSlug routes to Update (revise in place); otherwise
// Write (create).
type Insight struct {
Text string
Wing string
Hall string
SupersedeSlug string
}
// Ticket is one action item bound for a Gitea repo. Owner is always the
// operator (set by the tracker adapter), never carried here.
type Ticket struct {
Repo string
Action string // create | close | comment
Number int // required for close/comment
Title string // required for create
Body string
}
// Summary is an optional session summary bound for ai-sessions.
type Summary struct {
Title string
Body string
ReposTouched []string
}
// CaptureInput is the whole capture request.
type CaptureInput struct {
Context CaptureContext
Insights []Insight
Tickets []Ticket
Summary *Summary
DryRun bool
}
// InsightResult is the per-insight outcome in the receipt.
type InsightResult struct {
ID string `json:"id,omitempty"`
Path string `json:"path,omitempty"`
ContentHash string `json:"content_hash,omitempty"`
Superseded bool `json:"superseded"`
OK bool `json:"ok"`
}
// TicketResult is the per-ticket outcome in the receipt.
type TicketResult struct {
Repo string `json:"repo"`
Number int `json:"number,omitempty"`
Action string `json:"action"`
URL string `json:"url,omitempty"`
OK bool `json:"ok"`
}
// SummaryResult is the summary outcome in the receipt.
type SummaryResult struct {
Path string `json:"path,omitempty"`
OK bool `json:"ok"`
}
// ItemError pins a failure to a specific request item for the partial
// receipt. Item is a stable locator like "insight[1]" or "ticket[0]".
type ItemError struct {
Item string `json:"item"`
Error string `json:"error"`
}
// CaptureReceipt is the structured, partial-aware result. Per-item ok
// flags plus a flat Errors list make partial success explicit; the
// caller never has to infer what landed.
type CaptureReceipt struct {
Insights []InsightResult `json:"insights"`
Tickets []TicketResult `json:"tickets"`
Summary *SummaryResult `json:"summary,omitempty"`
Errors []ItemError `json:"errors"`
EffectiveClassification string `json:"effective_classification,omitempty"`
DryRun bool `json:"dry_run"`
}
+104
View File
@@ -0,0 +1,104 @@
package capture
import (
"context"
"time"
"github.com/mathiasbq/hyperguild/ingestion/internal/classification"
)
// Ref is the read-after-write handle returned by a brain write/update —
// the #45 contract. ContentHash lets the caller confirm what landed
// without a re-query; for an Update, Superseded is true.
type Ref struct {
ID string
Path string
ContentHash string
Superseded bool
}
// StoredNote is a brain note fetched by Get: the read-after-write
// confirmation primitive (a direct fetch, never a semantic query).
type StoredNote struct {
ID string
Path string
ContentHash string
Frontmatter map[string]string
Body string
}
// Note is the brain-write payload. It carries both the wing/hall taxonomy
// and the legacy type/domain fields so a single BrainStore serves both
// capture insights and the existing MCP brain_write surface. Reason is
// the supersede rationale, used only by Update.
type Note struct {
Content string
Filename string
Wing string
Hall string
Type string
Domain string
Reason string
}
// BrainStore is the brain persistence port — the shared implementation of
// the #45 write/update/get verbs that both the MCP handlers and capture
// call, so there is one implementation, not two. The read-after-write +
// staleness discipline lives behind this interface so no caller carries
// the rule.
type BrainStore interface {
Write(ctx context.Context, n Note) (Ref, error)
Update(ctx context.Context, slug string, n Note) (Ref, error)
Get(ctx context.Context, id string) (StoredNote, error)
}
// IssueRef identifies a ticket touched by the tracker.
type IssueRef struct {
Repo string
Number int
URL string
}
// IssueTracker is the Gitea ticket port. The implementation (#52) always
// scopes to owner "mathias"; the port deliberately omits owner.
type IssueTracker interface {
CreateIssue(ctx context.Context, repo, title, body string) (IssueRef, error)
// CloseIssue closes an issue, optionally posting a closing comment
// first (empty comment ⇒ close only).
CloseIssue(ctx context.Context, repo string, number int, comment string) (IssueRef, error)
CommentIssue(ctx context.Context, repo string, number int, body string) (IssueRef, error)
}
// SummaryWriter is the ai-sessions summary port.
type SummaryWriter interface {
WriteFile(ctx context.Context, repo, path, content string) error
}
// ClassificationPolicy derives a target's sensitivity (model C). The
// "stricter wins" combination of declared vs derived is use-case policy
// and lives in the service, so the port stays minimal. Satisfied by
// classification.Config (#50).
type ClassificationPolicy interface {
Derive(target classification.Target) classification.Level
}
// AuditEntry is the request-level audit record (I5): who/what captured
// what, when, via which principal. SecurityEvents carries anomalies such
// as a caller under-declaring sensitivity relative to the target floor.
type AuditEntry struct {
Timestamp time.Time
Principal string
Actor string
Harness string
SessionRef string
EffectiveClassification string
Items []string
SecurityEvents []string
}
// AuditSink records the audit entry. The classification-aware
// degradation/refusal policy (confidential fails closed, internal
// degrades) is the caller's concern in #54; this port just records.
type AuditSink interface {
Record(ctx context.Context, e AuditEntry) error
}
+316
View File
@@ -0,0 +1,316 @@
package capture
import (
"context"
"crypto/sha256"
"encoding/hex"
"fmt"
"strings"
"time"
"github.com/mathiasbq/hyperguild/ingestion/internal/brain"
"github.com/mathiasbq/hyperguild/ingestion/internal/classification"
)
// Service is the CaptureSession use-case. It depends only on ports.
type Service struct {
brain BrainStore
issues IssueTracker
summaries SummaryWriter
policy ClassificationPolicy
audit AuditSink
// now is the clock, injectable for deterministic summary paths and
// audit timestamps in tests.
now func() time.Time
}
// NewService constructs a Service from its ports. summaries may be nil
// when no summary persistence is wired; a CaptureInput with a Summary
// then fails that item rather than panicking.
func NewService(b BrainStore, tr IssueTracker, sw SummaryWriter, p ClassificationPolicy, a AuditSink) *Service {
return &Service{brain: b, issues: tr, summaries: sw, policy: p, audit: a, now: time.Now}
}
var validActions = map[string]bool{"create": true, "close": true, "comment": true}
// Capture runs the use-case: validate (fail-closed), resolve effective
// classification (stricter of declared vs target-derived), then persist
// insights → tickets → summary best-effort, emit an audit record, and
// return a partial-aware receipt.
//
// A validation failure returns a non-nil error with nothing written. A
// per-item execution failure is recorded in the receipt (no rollback);
// the call still returns a nil error so the caller gets the partial
// receipt. The I1 origin gate and audit-down degradation are layered on
// by #53/#54 around this core.
func (s *Service) Capture(ctx context.Context, in CaptureInput) (CaptureReceipt, error) {
if err := s.validate(in); err != nil {
return CaptureReceipt{}, err
}
declared := classification.Public // unspecified ⇒ lowest ⇒ target floor governs
if in.Context.Classification != "" {
// Already validated parseable.
declared, _ = classification.ParseLevel(in.Context.Classification)
}
effective, securityEvents := s.resolveClassification(declared, in)
receipt := CaptureReceipt{
Errors: []ItemError{},
EffectiveClassification: effective.String(),
DryRun: in.DryRun,
}
if in.DryRun {
// Would-be receipt: mark planned items ok, write nothing (not even
// audit — dry_run touches nothing).
for range in.Insights {
receipt.Insights = append(receipt.Insights, InsightResult{OK: true})
}
for _, tk := range in.Tickets {
receipt.Tickets = append(receipt.Tickets, TicketResult{Repo: tk.Repo, Action: tk.Action, Number: tk.Number, OK: true})
}
if in.Summary != nil {
receipt.Summary = &SummaryResult{Path: s.summaryPath(in.Context, in.Summary), OK: true}
}
return receipt, nil
}
var landed []string
for i, ins := range in.Insights {
res, item, err := s.persistInsight(ctx, ins)
receipt.Insights = append(receipt.Insights, res)
if err != nil {
receipt.Errors = append(receipt.Errors, ItemError{Item: fmt.Sprintf("insight[%d]", i), Error: err.Error()})
continue
}
landed = append(landed, item)
}
for i, tk := range in.Tickets {
res, err := s.persistTicket(ctx, tk)
receipt.Tickets = append(receipt.Tickets, res)
if err != nil {
receipt.Errors = append(receipt.Errors, ItemError{Item: fmt.Sprintf("ticket[%d]", i), Error: err.Error()})
continue
}
landed = append(landed, fmt.Sprintf("ticket:%s#%d", tk.Repo, res.Number))
}
if in.Summary != nil {
res, err := s.persistSummary(ctx, in.Context, in.Summary)
receipt.Summary = &res
if err != nil {
receipt.Errors = append(receipt.Errors, ItemError{Item: "summary", Error: err.Error()})
} else {
landed = append(landed, "summary:"+res.Path)
}
}
// I5: emit a request-level audit record of exactly what landed.
// Best-effort here; the classification-aware refusal/degradation
// policy is #54.
if err := s.audit.Record(ctx, AuditEntry{
Timestamp: s.now().UTC(),
Principal: in.Context.Principal,
Actor: in.Context.Actor,
Harness: in.Context.Harness,
SessionRef: in.Context.SessionRef,
EffectiveClassification: effective.String(),
Items: landed,
SecurityEvents: securityEvents,
}); err != nil {
receipt.Errors = append(receipt.Errors, ItemError{Item: "audit", Error: err.Error()})
}
return receipt, nil
}
// validate enforces fail-closed structural validity over the whole
// request before any write. A bad declared classification, an invalid
// wing/hall, an empty insight, or a malformed ticket aborts the capture
// with nothing written.
func (s *Service) validate(in CaptureInput) error {
if in.Context.Classification != "" {
if _, err := classification.ParseLevel(in.Context.Classification); err != nil {
return fmt.Errorf("context.classification: %w", err)
}
}
for i, ins := range in.Insights {
if strings.TrimSpace(ins.Text) == "" {
return fmt.Errorf("insight[%d]: text is required", i)
}
if strings.TrimSpace(ins.Wing) == "" {
return fmt.Errorf("insight[%d]: wing is required", i)
}
if !brain.IsValidHall(ins.Hall) {
return fmt.Errorf("insight[%d]: invalid hall %q", i, ins.Hall)
}
}
for i, tk := range in.Tickets {
if strings.TrimSpace(tk.Repo) == "" {
return fmt.Errorf("ticket[%d]: repo is required", i)
}
if !validActions[tk.Action] {
return fmt.Errorf("ticket[%d]: invalid action %q (want create/close/comment)", i, tk.Action)
}
if tk.Action == "create" && strings.TrimSpace(tk.Title) == "" {
return fmt.Errorf("ticket[%d]: create requires a title", i)
}
if (tk.Action == "close" || tk.Action == "comment") && tk.Number <= 0 {
return fmt.Errorf("ticket[%d]: %s requires an issue number", i, tk.Action)
}
}
return nil
}
// resolveClassification computes the effective level (stricter of
// declared and every target's derived level) and collects a security
// event whenever the caller under-declared relative to a target floor.
func (s *Service) resolveClassification(declared classification.Level, in CaptureInput) (classification.Level, []string) {
effective := declared
var events []string
consider := func(kind classification.TargetKind, name string) {
derived := s.policy.Derive(classification.Target{Kind: kind, Name: name})
effective = classification.Stricter(effective, derived)
if declared < derived {
events = append(events, fmt.Sprintf("classification under-declared: declared=%s target=%s(%s) derived=%s",
declared, name, kindString(kind), derived))
}
}
for _, ins := range in.Insights {
consider(classification.WingTarget, ins.Wing)
}
for _, tk := range in.Tickets {
consider(classification.RepoTarget, tk.Repo)
}
if in.Summary != nil {
for _, repo := range in.Summary.ReposTouched {
consider(classification.RepoTarget, repo)
}
}
return effective, events
}
func (s *Service) persistInsight(ctx context.Context, ins Insight) (InsightResult, string, error) {
note := Note{Content: ins.Text, Wing: ins.Wing, Hall: ins.Hall, Filename: brain.Sanitise(firstLine(ins.Text))}
var ref Ref
var err error
if ins.SupersedeSlug != "" {
note.Reason = "superseded via capture"
ref, err = s.brain.Update(ctx, ins.SupersedeSlug, note)
} else {
ref, err = s.brain.Write(ctx, note)
}
if err != nil {
return InsightResult{OK: false, Superseded: ins.SupersedeSlug != ""}, "", err
}
return InsightResult{
ID: ref.ID, Path: ref.Path, ContentHash: ref.ContentHash,
Superseded: ref.Superseded, OK: true,
}, "insight:" + ref.ID, nil
}
func (s *Service) persistTicket(ctx context.Context, tk Ticket) (TicketResult, error) {
res := TicketResult{Repo: tk.Repo, Action: tk.Action, Number: tk.Number}
var ref IssueRef
var err error
switch tk.Action {
case "create":
ref, err = s.issues.CreateIssue(ctx, tk.Repo, tk.Title, tk.Body)
case "close":
ref, err = s.issues.CloseIssue(ctx, tk.Repo, tk.Number, tk.Body)
case "comment":
ref, err = s.issues.CommentIssue(ctx, tk.Repo, tk.Number, tk.Body)
}
if err != nil {
return res, err
}
if ref.Number != 0 {
res.Number = ref.Number
}
res.URL = ref.URL
res.OK = true
return res, nil
}
func (s *Service) persistSummary(ctx context.Context, c CaptureContext, sum *Summary) (SummaryResult, error) {
if s.summaries == nil {
return SummaryResult{OK: false}, fmt.Errorf("no summary writer configured")
}
path := s.summaryPath(c, sum)
content := s.renderSummary(c, sum)
repo := "ai-sessions"
if err := s.summaries.WriteFile(ctx, repo, path, content); err != nil {
return SummaryResult{Path: path, OK: false}, err
}
return SummaryResult{Path: path, OK: true}, nil
}
// summaryPath builds summaries/<harness>/<YYYY-MM>/<date>-<slug>-<ref8>.md.
// The ref8 disambiguator is derived from the session_ref (or the title
// when no ref is present) so distinct sessions never collide.
func (s *Service) summaryPath(c CaptureContext, sum *Summary) string {
t := s.now().UTC()
slug := brain.Sanitise(sum.Title)
if slug == "" {
slug = "summary"
}
seed := c.SessionRef
if seed == "" {
seed = sum.Title + sum.Body
}
sum8 := shortHash(seed)
return fmt.Sprintf("summaries/%s/%s/%s-%s-%s.md",
brain.Sanitise(c.Harness), t.Format("2006-01"), t.Format("2006-01-02"), slug, sum8)
}
// renderSummary stamps fidelity + session metadata into frontmatter so the
// richer-fidelity-supersedes-thinner collision rule has the data it needs.
func (s *Service) renderSummary(c CaptureContext, sum *Summary) string {
var b strings.Builder
b.WriteString("---\n")
fmt.Fprintf(&b, "title: %s\n", sum.Title)
fmt.Fprintf(&b, "harness: %s\n", c.Harness)
if c.SessionRef != "" {
fmt.Fprintf(&b, "session_ref: %s\n", c.SessionRef)
}
fmt.Fprintf(&b, "fidelity: %s\n", c.Fidelity)
fmt.Fprintf(&b, "captured_at: %s\n", s.now().UTC().Format(time.RFC3339))
if len(sum.ReposTouched) > 0 {
fmt.Fprintf(&b, "repos_touched: [%s]\n", strings.Join(sum.ReposTouched, ", "))
}
b.WriteString("---\n\n")
b.WriteString(sum.Body)
if !strings.HasSuffix(sum.Body, "\n") {
b.WriteByte('\n')
}
return b.String()
}
func kindString(k classification.TargetKind) string {
if k == classification.RepoTarget {
return "repo"
}
return "wing"
}
func firstLine(s string) string {
s = strings.TrimSpace(s)
if i := strings.IndexByte(s, '\n'); i >= 0 {
s = s[:i]
}
s = strings.TrimLeft(s, "# ")
if len(s) > 60 {
s = s[:60]
}
return s
}
func shortHash(s string) string {
sum := sha256.Sum256([]byte(s))
return hex.EncodeToString(sum[:])[:8]
}
+331
View File
@@ -0,0 +1,331 @@
package capture
import (
"context"
"errors"
"strings"
"testing"
"time"
"github.com/mathiasbq/hyperguild/ingestion/internal/classification"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
)
// --- fakes ---
type fakeBrain struct {
writes []Note
updates []Note
gets []string
failOn func(Note) error // nil = always succeed
hashSeq int
}
func (f *fakeBrain) ref(prefix string, n Note, superseded bool) Ref {
f.hashSeq++
path := "wiki/" + n.Wing + "/" + n.Hall + "/" + n.Filename + ".md"
return Ref{ID: path, Path: path, ContentHash: prefix + string(rune('0'+f.hashSeq)), Superseded: superseded}
}
func (f *fakeBrain) Write(_ context.Context, n Note) (Ref, error) {
if f.failOn != nil {
if err := f.failOn(n); err != nil {
return Ref{}, err
}
}
f.writes = append(f.writes, n)
return f.ref("w", n, false), nil
}
func (f *fakeBrain) Update(_ context.Context, slug string, n Note) (Ref, error) {
if f.failOn != nil {
if err := f.failOn(n); err != nil {
return Ref{}, err
}
}
n.Filename = slug
f.updates = append(f.updates, n)
return f.ref("u", n, true), nil
}
func (f *fakeBrain) Get(_ context.Context, id string) (StoredNote, error) {
f.gets = append(f.gets, id)
return StoredNote{ID: id, Path: id}, nil
}
type fakeTracker struct {
created []string
closed []int
comments []int
err error
}
func (f *fakeTracker) CreateIssue(_ context.Context, repo, title, _ string) (IssueRef, error) {
if f.err != nil {
return IssueRef{}, f.err
}
f.created = append(f.created, repo+":"+title)
return IssueRef{Repo: repo, Number: 100 + len(f.created), URL: "https://git/" + repo + "/issues/x"}, nil
}
func (f *fakeTracker) CloseIssue(_ context.Context, repo string, number int, _ string) (IssueRef, error) {
if f.err != nil {
return IssueRef{}, f.err
}
f.closed = append(f.closed, number)
return IssueRef{Repo: repo, Number: number}, nil
}
func (f *fakeTracker) CommentIssue(_ context.Context, repo string, number int, _ string) (IssueRef, error) {
if f.err != nil {
return IssueRef{}, f.err
}
f.comments = append(f.comments, number)
return IssueRef{Repo: repo, Number: number}, nil
}
type fakeSummary struct {
paths []string
content []string
err error
}
func (f *fakeSummary) WriteFile(_ context.Context, _, path, content string) error {
if f.err != nil {
return f.err
}
f.paths = append(f.paths, path)
f.content = append(f.content, content)
return nil
}
// fakePolicy derives from an explicit map; default Internal so tests pin
// behaviour without depending on the real defaulting.
type fakePolicy struct{ tags map[string]classification.Level }
func (p fakePolicy) Derive(t classification.Target) classification.Level {
if lvl, ok := p.tags[t.Name]; ok {
return lvl
}
return classification.Internal
}
type fakeAudit struct {
entries []AuditEntry
err error
}
func (f *fakeAudit) Record(_ context.Context, e AuditEntry) error {
if f.err != nil {
return f.err
}
f.entries = append(f.entries, e)
return nil
}
// --- helpers ---
func newSvc(b BrainStore, tr IssueTracker, sw SummaryWriter, p ClassificationPolicy, a AuditSink) *Service {
s := NewService(b, tr, sw, p, a)
s.now = func() time.Time { return time.Date(2026, 6, 22, 12, 0, 0, 0, time.UTC) }
return s
}
func baseCtx() CaptureContext {
return CaptureContext{Harness: "claude-code", Actor: "mathias", Principal: "mathias", Classification: "internal"}
}
// --- scenarios ---
func TestCaptureHappyPath(t *testing.T) {
b := &fakeBrain{}
tr := &fakeTracker{}
au := &fakeAudit{}
svc := newSvc(b, tr, nil, fakePolicy{}, au)
rec, err := svc.Capture(context.Background(), CaptureInput{
Context: baseCtx(),
Insights: []Insight{
{Text: "a", Wing: "hyperguild", Hall: "decisions", SupersedeSlug: ""},
{Text: "b", Wing: "hyperguild", Hall: "facts"},
},
Tickets: []Ticket{{Repo: "hyperguild", Action: "create", Title: "do x", Body: "y"}},
})
require.NoError(t, err)
require.Len(t, rec.Insights, 2)
for _, r := range rec.Insights {
assert.True(t, r.OK)
assert.NotEmpty(t, r.ContentHash, "read-after-write hash returned")
}
require.Len(t, rec.Tickets, 1)
assert.True(t, rec.Tickets[0].OK)
assert.Equal(t, 2, len(b.writes))
assert.Empty(t, rec.Errors)
// Audit emitted naming principal/harness + items that landed.
require.Len(t, au.entries, 1)
assert.Equal(t, "mathias", au.entries[0].Principal)
assert.Equal(t, "claude-code", au.entries[0].Harness)
assert.Len(t, au.entries[0].Items, 3)
}
func TestCaptureSupersedeNotDuplicate(t *testing.T) {
b := &fakeBrain{}
svc := newSvc(b, &fakeTracker{}, nil, fakePolicy{}, &fakeAudit{})
rec, err := svc.Capture(context.Background(), CaptureInput{
Context: baseCtx(),
Insights: []Insight{{Text: "revised", Wing: "hyperguild", Hall: "facts", SupersedeSlug: "prior-note"}},
})
require.NoError(t, err)
assert.Empty(t, b.writes, "supersede must not create")
require.Len(t, b.updates, 1)
assert.Equal(t, "prior-note", b.updates[0].Filename)
assert.True(t, rec.Insights[0].Superseded)
}
func TestCaptureValidationFailClosed(t *testing.T) {
b := &fakeBrain{}
tr := &fakeTracker{}
au := &fakeAudit{}
svc := newSvc(b, tr, nil, fakePolicy{}, au)
_, err := svc.Capture(context.Background(), CaptureInput{
Context: baseCtx(),
Insights: []Insight{
{Text: "ok", Wing: "hyperguild", Hall: "facts"},
{Text: "bad", Wing: "hyperguild", Hall: "garbage-hall"}, // invalid hall
},
Tickets: []Ticket{{Repo: "hyperguild", Action: "create", Title: "t"}},
})
require.Error(t, err)
// Nothing written anywhere.
assert.Empty(t, b.writes)
assert.Empty(t, b.updates)
assert.Empty(t, tr.created)
assert.Empty(t, au.entries)
}
func TestCaptureValidationRejectsBadTicket(t *testing.T) {
svc := newSvc(&fakeBrain{}, &fakeTracker{}, nil, fakePolicy{}, &fakeAudit{})
_, err := svc.Capture(context.Background(), CaptureInput{
Context: baseCtx(),
Tickets: []Ticket{{Repo: "hyperguild", Action: "frobnicate"}}, // bad action
})
require.Error(t, err)
_, err = svc.Capture(context.Background(), CaptureInput{
Context: baseCtx(),
Tickets: []Ticket{{Repo: "hyperguild", Action: "close"}}, // close needs number
})
require.Error(t, err)
}
func TestCapturePartialFailureBestEffort(t *testing.T) {
b := &fakeBrain{failOn: func(n Note) error {
if strings.Contains(n.Content, "FAIL") {
return errors.New("disk full")
}
return nil
}}
tr := &fakeTracker{}
au := &fakeAudit{}
svc := newSvc(b, tr, nil, fakePolicy{}, au)
rec, err := svc.Capture(context.Background(), CaptureInput{
Context: baseCtx(),
Insights: []Insight{
{Text: "good one", Wing: "hyperguild", Hall: "facts"},
{Text: "FAIL here", Wing: "hyperguild", Hall: "facts"},
},
Tickets: []Ticket{{Repo: "hyperguild", Action: "create", Title: "t"}},
})
require.NoError(t, err, "partial failure is not a request-level error")
assert.True(t, rec.Insights[0].OK)
assert.False(t, rec.Insights[1].OK)
assert.True(t, rec.Tickets[0].OK, "ticket still persisted; no rollback")
require.Len(t, rec.Errors, 1)
assert.Equal(t, "insight[1]", rec.Errors[0].Item)
// Audit reflects exactly what landed: 1 insight + 1 ticket.
require.Len(t, au.entries, 1)
assert.Len(t, au.entries[0].Items, 2)
}
func TestCaptureDryRunWritesNothing(t *testing.T) {
b := &fakeBrain{}
tr := &fakeTracker{}
au := &fakeAudit{}
svc := newSvc(b, tr, nil, fakePolicy{}, au)
rec, err := svc.Capture(context.Background(), CaptureInput{
Context: baseCtx(),
DryRun: true,
Insights: []Insight{{Text: "a", Wing: "hyperguild", Hall: "facts"}},
Tickets: []Ticket{{Repo: "hyperguild", Action: "create", Title: "t"}},
})
require.NoError(t, err)
assert.True(t, rec.DryRun)
assert.Len(t, rec.Insights, 1)
assert.True(t, rec.Insights[0].OK, "would-be receipt marks planned items ok")
// Nothing written anywhere, including audit.
assert.Empty(t, b.writes)
assert.Empty(t, tr.created)
assert.Empty(t, au.entries)
}
func TestCaptureStricterClassificationWins(t *testing.T) {
// Caller declares internal; target wing tagged confidential → effective confidential + security event.
b := &fakeBrain{}
au := &fakeAudit{}
pol := fakePolicy{tags: map[string]classification.Level{"client-seb": classification.Confidential}}
svc := newSvc(b, &fakeTracker{}, nil, pol, au)
ctx := baseCtx()
ctx.Classification = "internal"
rec, err := svc.Capture(context.Background(), CaptureInput{
Context: ctx,
Insights: []Insight{{Text: "x", Wing: "client-seb", Hall: "facts"}},
})
require.NoError(t, err)
assert.Equal(t, "confidential", rec.EffectiveClassification)
require.Len(t, au.entries, 1)
assert.NotEmpty(t, au.entries[0].SecurityEvents, "under-declaration logged as security event")
assert.Equal(t, "confidential", au.entries[0].EffectiveClassification)
}
func TestCaptureCallerRaisingSensitivityHonoured(t *testing.T) {
// Caller declares confidential; target internal → effective confidential, NOT a security event.
au := &fakeAudit{}
pol := fakePolicy{tags: map[string]classification.Level{"hyperguild": classification.Internal}}
svc := newSvc(&fakeBrain{}, &fakeTracker{}, nil, pol, au)
ctx := baseCtx()
ctx.Classification = "confidential"
rec, err := svc.Capture(context.Background(), CaptureInput{
Context: ctx,
Insights: []Insight{{Text: "x", Wing: "hyperguild", Hall: "facts"}},
})
require.NoError(t, err)
assert.Equal(t, "confidential", rec.EffectiveClassification)
assert.Empty(t, au.entries[0].SecurityEvents, "raising sensitivity is honoured, not flagged")
}
func TestCaptureSummaryPathAndFidelity(t *testing.T) {
sw := &fakeSummary{}
svc := newSvc(&fakeBrain{}, &fakeTracker{}, sw, fakePolicy{}, &fakeAudit{})
ctx := baseCtx()
ctx.Fidelity = "transcript-parse"
ctx.SessionRef = "abc123def456"
rec, err := svc.Capture(context.Background(), CaptureInput{
Context: ctx,
Summary: &Summary{Title: "Session Wrap", Body: "did stuff", ReposTouched: []string{"hyperguild"}},
})
require.NoError(t, err)
require.NotNil(t, rec.Summary)
assert.True(t, rec.Summary.OK)
require.Len(t, sw.paths, 1)
assert.True(t, strings.HasPrefix(sw.paths[0], "summaries/claude-code/2026-06/"), "path: %s", sw.paths[0])
assert.Contains(t, sw.paths[0], "session-wrap")
assert.Contains(t, sw.content[0], "fidelity: transcript-parse", "fidelity stamped in frontmatter")
}
@@ -0,0 +1,189 @@
// Package classification defines the data-sensitivity taxonomy and the
// per-wing / per-repo tagging the capture server reads to enforce the I1
// sovereignty gate (issue #50, capture spec §4.1).
//
// The single load-bearing property is fail-safe-to-strictest: a target
// with no explicit tag and no known default classifies as Confidential,
// never as something more permissive. A missing tag must never silently
// downgrade — that would turn the I1 gate into theatre.
//
// Classification is read from an optional classification.yaml at the
// brain root. A central, Flux-reconcilable file is deliberate: it is
// auditable in one place (I2/I5), it does not require a live Gitea client
// to classify a repo (so this package has no dependency on the gitea
// tracker work), and it avoids tagging a wing's _index.md frontmatter —
// which BuildWingIndex regenerates and would clobber.
package classification
import (
"fmt"
"os"
"path/filepath"
"strings"
"gopkg.in/yaml.v3"
)
// Level is a data-sensitivity tier. Higher is stricter, so the "stricter
// wins" rule (spec §4.1 model C) is a plain max.
type Level int
const (
Public Level = iota
Internal
Confidential
)
// String returns the canonical lowercase token for a level.
func (l Level) String() string {
switch l {
case Public:
return "public"
case Internal:
return "internal"
case Confidential:
return "confidential"
default:
return fmt.Sprintf("level(%d)", int(l))
}
}
// ParseLevel parses a level token (case-insensitive, surrounding space
// tolerated). An unknown token is an error — callers must decide what to
// do with bad input rather than have it silently coerced.
func ParseLevel(s string) (Level, error) {
switch strings.ToLower(strings.TrimSpace(s)) {
case "public":
return Public, nil
case "internal":
return Internal, nil
case "confidential":
return Confidential, nil
default:
return Confidential, fmt.Errorf("unknown classification level %q (want public/internal/confidential)", s)
}
}
// Stricter returns the more restrictive of two levels.
func Stricter(a, b Level) Level {
if a > b {
return a
}
return b
}
// TargetKind distinguishes the two kinds of capture destination.
type TargetKind int
const (
WingTarget TargetKind = iota // a brain wing (insights land here)
RepoTarget // a Gitea repo (tickets / summaries land here)
)
// Target names a capture destination to classify.
type Target struct {
Kind TargetKind
Name string
}
// Config holds the explicit per-wing / per-repo classification tags read
// from classification.yaml. Absent entries fall through to the built-in
// defaults in defaultFor. The zero value (no file) is valid and applies
// defaults to everything.
type Config struct {
wings map[string]Level
repos map[string]Level
}
// rawConfig is the on-disk YAML shape: string→string maps, parsed into
// validated levels by Load.
type rawConfig struct {
Wings map[string]string `yaml:"wings"`
Repos map[string]string `yaml:"repos"`
}
// Load reads classification.yaml from brainDir. An absent file is not an
// error — it yields an empty config where every target classifies by the
// built-in defaults. A malformed file, or any unparseable level token in
// it, is a hard error: a classification source the server cannot trust
// must fail loud, not degrade silently.
func Load(brainDir string) (*Config, error) {
cfg := &Config{wings: map[string]Level{}, repos: map[string]Level{}}
data, err := os.ReadFile(filepath.Join(brainDir, "classification.yaml"))
if err != nil {
if os.IsNotExist(err) {
return cfg, nil
}
return nil, fmt.Errorf("read classification.yaml: %w", err)
}
var raw rawConfig
if err := yaml.Unmarshal(data, &raw); err != nil {
return nil, fmt.Errorf("parse classification.yaml: %w", err)
}
for name, lvl := range raw.Wings {
parsed, perr := ParseLevel(lvl)
if perr != nil {
return nil, fmt.Errorf("wing %q: %w", name, perr)
}
cfg.wings[normalise(name)] = parsed
}
for name, lvl := range raw.Repos {
parsed, perr := ParseLevel(lvl)
if perr != nil {
return nil, fmt.Errorf("repo %q: %w", name, perr)
}
cfg.repos[normalise(name)] = parsed
}
return cfg, nil
}
// Derive returns the classification for any target — the function the
// capture use-case calls per item.
func (c *Config) Derive(t Target) Level {
if t.Kind == RepoTarget {
return c.Repo(t.Name)
}
return c.Wing(t.Name)
}
// Wing classifies a brain wing: an explicit tag wins, else defaults.
func (c *Config) Wing(name string) Level {
if lvl, ok := c.wings[normalise(name)]; ok {
return lvl
}
return defaultFor(name)
}
// Repo classifies a Gitea repo: an explicit tag wins, else defaults.
func (c *Config) Repo(name string) Level {
if lvl, ok := c.repos[normalise(name)]; ok {
return lvl
}
return defaultFor(name)
}
// defaultFor applies the built-in defaulting rules when a target has no
// explicit tag:
// - client-* → Confidential (client work is confidential by default)
// - hyperguild / homelab → Internal (the operator's own infra)
// - everything else → Confidential (fail safe to strictest)
func defaultFor(name string) Level {
n := normalise(name)
if strings.HasPrefix(n, "client-") {
return Confidential
}
switch n {
case "hyperguild", "homelab":
return Internal
default:
return Confidential
}
}
// normalise lowercases and trims a wing/repo name so matching and the
// client-* prefix check are case-insensitive.
func normalise(name string) string {
return strings.ToLower(strings.TrimSpace(name))
}
@@ -0,0 +1,112 @@
package classification
import (
"os"
"path/filepath"
"testing"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
)
func TestLevelOrderingAndString(t *testing.T) {
assert.True(t, Public < Internal)
assert.True(t, Internal < Confidential)
assert.Equal(t, "public", Public.String())
assert.Equal(t, "internal", Internal.String())
assert.Equal(t, "confidential", Confidential.String())
}
func TestParseLevel(t *testing.T) {
for s, want := range map[string]Level{
"public": Public, "internal": Internal, "confidential": Confidential,
"PUBLIC": Public, " Confidential ": Confidential,
} {
got, err := ParseLevel(s)
require.NoError(t, err, s)
assert.Equal(t, want, got, s)
}
_, err := ParseLevel("secret")
require.Error(t, err, "unknown level must error, not silently default")
_, err = ParseLevel("")
require.Error(t, err)
}
func TestStricterReturnsMax(t *testing.T) {
assert.Equal(t, Confidential, Stricter(Internal, Confidential))
assert.Equal(t, Confidential, Stricter(Confidential, Public))
assert.Equal(t, Internal, Stricter(Public, Internal))
assert.Equal(t, Public, Stricter(Public, Public))
}
func TestLoadAbsentFileIsDefaultsOnly(t *testing.T) {
cfg, err := Load(t.TempDir())
require.NoError(t, err, "absent classification.yaml must not be an error — defaults apply")
require.NotNil(t, cfg)
// Pure defaulting still works.
assert.Equal(t, Internal, cfg.Wing("hyperguild"))
assert.Equal(t, Confidential, cfg.Wing("anything-unknown"))
}
func TestLoadParsesExplicitTags(t *testing.T) {
dir := t.TempDir()
require.NoError(t, os.WriteFile(filepath.Join(dir, "classification.yaml"), []byte(
"wings:\n research-public: public\n hyperguild: confidential\nrepos:\n infra: internal\n research-public: public\n",
), 0o644))
cfg, err := Load(dir)
require.NoError(t, err)
// Explicit tag wins over the built-in default (hyperguild default is internal).
assert.Equal(t, Confidential, cfg.Wing("hyperguild"))
// Explicit public is honoured.
assert.Equal(t, Public, cfg.Wing("research-public"))
assert.Equal(t, Internal, cfg.Repo("infra"))
assert.Equal(t, Public, cfg.Repo("research-public"))
}
func TestLoadRejectsUnknownLevelInFile(t *testing.T) {
dir := t.TempDir()
require.NoError(t, os.WriteFile(filepath.Join(dir, "classification.yaml"),
[]byte("wings:\n x: top-secret\n"), 0o644))
_, err := Load(dir)
require.Error(t, err, "an unparseable level in the config must fail loud, not be ignored")
}
func TestWingDefaulting(t *testing.T) {
cfg, err := Load(t.TempDir())
require.NoError(t, err)
cases := map[string]Level{
"client-seb": Confidential, // client-* → confidential
"client-mastercard": Confidential,
"hyperguild": Internal,
"homelab": Internal,
"jepa-fx": Confidential, // unknown → fail safe to strictest
"": Confidential, // empty → fail safe
}
for wing, want := range cases {
assert.Equal(t, want, cfg.Wing(wing), "wing %q", wing)
}
}
func TestRepoDefaulting(t *testing.T) {
cfg, err := Load(t.TempDir())
require.NoError(t, err)
assert.Equal(t, Confidential, cfg.Repo("client-seb-pipeline"))
assert.Equal(t, Internal, cfg.Repo("hyperguild"))
assert.Equal(t, Confidential, cfg.Repo("some-unknown-repo"), "untagged repo → confidential (fail safe)")
}
func TestDeriveUnifiedTarget(t *testing.T) {
cfg, err := Load(t.TempDir())
require.NoError(t, err)
assert.Equal(t, Internal, cfg.Derive(Target{Kind: WingTarget, Name: "homelab"}))
assert.Equal(t, Confidential, cfg.Derive(Target{Kind: RepoTarget, Name: "client-x"}))
assert.Equal(t, Confidential, cfg.Derive(Target{Kind: WingTarget, Name: "untagged"}))
}
func TestCaseInsensitiveMatching(t *testing.T) {
cfg, err := Load(t.TempDir())
require.NoError(t, err)
assert.Equal(t, Confidential, cfg.Wing("Client-SEB"), "client- prefix match is case-insensitive")
assert.Equal(t, Internal, cfg.Wing("HyperGuild"))
}
+66 -2
View File
@@ -1,6 +1,10 @@
package claudewatcher
import "regexp"
import (
"fmt"
"regexp"
"sync"
)
// Scrubber drops any turn whose content matches a known-bad pattern.
// Fail-closed by design: we'd rather lose signal than ingest credentials
@@ -34,21 +38,74 @@ var DefaultRules = []Rule{
// specific match name in logs.
{Name: "authorization-header", RE: regexp.MustCompile(`(?i)Authorization\s*:\s*[A-Za-z]+\s+\S{8,}`)},
{Name: "bearer-token", RE: regexp.MustCompile(`(?i)Bearer\s+[A-Za-z0-9._\-]{16,}`)},
// JWT (header.payload.sig), e.g. a Dex/OAuth token dumped to stdout
// without a "Bearer " prefix. Both header and payload base64url-encode
// JSON, so both segments begin with "eyJ".
{Name: "jwt", RE: regexp.MustCompile(`eyJ[A-Za-z0-9_\-]{8,}\.eyJ[A-Za-z0-9_\-]{8,}\.[A-Za-z0-9_\-]{8,}`)},
{Name: "postgres-uri-with-password", RE: regexp.MustCompile(`postgres(?:ql)?://[^:\s/]+:[^@\s/]+@`)},
{Name: "private-key", RE: regexp.MustCompile(`-----BEGIN[^-]*PRIVATE KEY-----`)},
{Name: "ssh-key", RE: regexp.MustCompile(`ssh-(?:rsa|ed25519|ecdsa)\s+[A-Za-z0-9+/=]{40,}`)},
{Name: "github-pat", RE: regexp.MustCompile(`\b(?:ghp|gho|ghu|ghr|gha)_[A-Za-z0-9]{30,}\b`)},
{Name: "openai-sk", RE: regexp.MustCompile(`\bsk-(?:proj-)?[A-Za-z0-9]{32,}\b`)},
// 1Password service-account token (ops_<base64url>). Long, high-value root
// credential; guard the bare value (the _TOKEN= form also hits homelab-env-token).
{Name: "op-service-account", RE: regexp.MustCompile(`\bops_[A-Za-z0-9_\-]{40,}`)},
// No leading \b: a shell mangle can glue the key to a preceding word
// ("yes"+"sk-...") which has no word boundary, and that exact case
// leaked a LiteLLM master key past this rule (2026-06-11). Match the
// sk- shape wherever it appears; the {32,} length floor keeps short
// "task-"/"disk-" words from tripping it.
{Name: "openai-sk", RE: regexp.MustCompile(`sk-(?:proj-)?[A-Za-z0-9]{32,}`)},
{Name: "anthropic-sk", RE: regexp.MustCompile(`\bsk-ant-[A-Za-z0-9_\-]{32,}\b`)},
{Name: "aws-access-key", RE: regexp.MustCompile(`\bAKIA[0-9A-Z]{16}\b`)},
{Name: "homelab-env-token", RE: regexp.MustCompile(`(?i)(?:_TOKEN|_PASSWORD|_API_KEY|_SECRET)\s*[:=]\s*['"]?[A-Za-z0-9._/+\-]{12,}`)},
{Name: "sops-encrypted-marker", RE: regexp.MustCompile(`ENC\[AES256_GCM,data:[A-Za-z0-9+/=]{8,}`)},
}
// extraRules is appended to DefaultRules at process startup via
// RegisterRule. The mutex guards concurrent RegisterRule calls (rare)
// against concurrent Scrub reads (hot path). Scrub takes a read lock
// only when extraRules is non-empty, so steady-state cost is zero
// when no client-name guard is configured.
var (
extraRulesMu sync.RWMutex
extraRules []Rule
)
// RegisterRule appends a runtime-configured regex to the scrubber's
// rule set. Used by main to inject client-name guards from
// CLAUDE_INGEST_CLIENT_BLOCK env var (or equivalent SOPS-encrypted
// secret) without baking client identities into source code.
//
// pattern is compiled as-is — callers wrap with `\b...\b` and case
// flags as needed. Duplicate names are accepted (rules are positional);
// the second registration just fires after the first.
func RegisterRule(name, pattern string) error {
re, err := regexp.Compile(pattern)
if err != nil {
return fmt.Errorf("compile rule %q: %w", name, err)
}
extraRulesMu.Lock()
extraRules = append(extraRules, Rule{Name: name, RE: re})
extraRulesMu.Unlock()
return nil
}
// ResetExtraRules clears every RegisterRule-added rule. Test-only.
func ResetExtraRules() {
extraRulesMu.Lock()
extraRules = nil
extraRulesMu.Unlock()
}
// Scrub reports the first matching rule, or empty when content is clean.
// Empty string is treated as clean. Caller decides what to do on a hit;
// the convention in claudewatcher is to drop the turn entirely and emit
// a slog.Warn naming the rule.
//
// Rule order: DefaultRules first (credential shapes), then runtime
// RegisterRule additions (client-name guards). Credential leaks
// outrank client-name hits in the log because they're strictly more
// dangerous.
func Scrub(content string) string {
if content == "" {
return ""
@@ -58,5 +115,12 @@ func Scrub(content string) string {
return r.Name
}
}
extraRulesMu.RLock()
defer extraRulesMu.RUnlock()
for _, r := range extraRules {
if r.RE.MatchString(content) {
return r.Name
}
}
return ""
}
@@ -25,6 +25,18 @@ func TestScrub_PoisonedFixtures(t *testing.T) {
{"aws-access-key", "AWS_ACCESS_KEY_ID=AKIAIOSFODNN7EXAMPLE", "aws-access-key"},
{"homelab-env", "POSTGRES_PASSWORD=hunter2supersecretvalue", "homelab-env-token"},
{"sops-marker", "value: ENC[AES256_GCM,data:abc123def456,iv:zzz]", "sops-encrypted-marker"},
// Regression: a shell mangle glued the key to a preceding word
// ("yes"+"sk-..."), defeating the leading \b in the sk- rule and
// leaking a LiteLLM master key past the scrubber (2026-06-11).
{"sk-glued-to-word", "master key resolved: yessk-7181ca984603239d8c4819361bf33b94b9c3c07018791868", "openai-sk"},
{"sk-standalone-hex", "sk-7181ca984603239d8c4819361bf33b94b9c3c07018791868", "openai-sk"},
// Bare JWT not preceded by "Bearer" (e.g. a Dex token dumped to stdout).
{"jwt-bare", "token=eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9.eyJzdWIiOiIxMjM0NTY3ODkwIn0.dQw4w9WgXcQabcdef", "jwt"},
// 1Password service-account token (ops_<base64url>), env-assigned and bare.
// Both hit the dedicated op-service-account rule (ordered before the
// generic homelab-env-token). Guards ~/.zshrc reads etc. (2026-06-14).
{"op-sa-env", "export OP_SERVICE_ACCOUNT_TOKEN=ops_eyJzaWduSW5BZGRyZXNzIjoibXkuMXBhc3N3b3JkLmNvbSJ9", "op-service-account"},
{"op-sa-bare", "ops_eyJzaWduSW5BZGRyZXNzIjoibXkuMXBhc3N3b3JkLmNvbSIsInVzZXJBdXRoIjp7fX0aGVsbG8", "op-service-account"},
}
for _, tc := range cases {
t.Run(tc.name, func(t *testing.T) {
@@ -43,6 +55,9 @@ func TestScrub_CleanContentPassesThrough(t *testing.T) {
"file at ~/.ssh/id_ed25519",
"the function Authorization() takes no args",
"comment: see API key in 1Password",
// loosened sk- rule must not trip on short "task-"/"disk-" words
"run task-build then task-test in the pipeline",
"mounted /dev/disk-by-id/wwn-0x5000",
}
for _, c := range cases {
assert.Empty(t, Scrub(c), "expected clean for %q", c)
@@ -55,3 +70,63 @@ func TestScrub_FirstMatchWins(t *testing.T) {
content := "Authorization: Bearer ghp_aBcD1234EfGh5678IjKl9012MnOp3456QrSt"
assert.Equal(t, "authorization-header", Scrub(content))
}
func TestRegisterRule_ClientNameGuard(t *testing.T) {
t.Cleanup(ResetExtraRules)
require := func(err error) {
if err != nil {
t.Fatalf("unexpected err: %v", err)
}
}
require(RegisterRule("client-name", `(?i)\b(SEB|Mastercard)\b`))
// Hits — case variations + word-boundary respect.
for _, hit := range []string{
"mentioned SEB in this commit",
"the Mastercard project deadline",
"working on mastercard scope",
"SEB internal review",
} {
assert.Equal(t, "client-name", Scrub(hit), "should match %q", hit)
}
// Misses — substring within a longer word should NOT match
// thanks to \b. "Sebastian" contains "seb" but \b prevents hit.
for _, miss := range []string{
"Sebastian wrote the docs",
"unrelated text",
"researcher",
"https://example.com/search?seb=1", // 'seb' bounded by ?=, still matches \b
} {
got := Scrub(miss)
if miss == "https://example.com/search?seb=1" {
// `seb=` has word-boundary at '='; this DOES match \bseb\b.
// Accept either outcome; document the tradeoff.
assert.Contains(t, []string{"", "client-name"}, got)
continue
}
assert.Empty(t, got, "should NOT match %q", miss)
}
}
func TestRegisterRule_CredentialsTakePrecedence(t *testing.T) {
t.Cleanup(ResetExtraRules)
require := func(err error) {
if err != nil {
t.Fatalf("unexpected err: %v", err)
}
}
require(RegisterRule("client-name", `\b(SEB)\b`))
// Content matches both a credential rule AND a client rule —
// credential rule wins by ordering, so log triage points at the
// strictly more dangerous leak.
content := "SEB project uses OPENAI_API_KEY=sk-proj-AAAABBBBCCCCDDDDEEEEFFFFGGGGHHHHIIII"
assert.Equal(t, "openai-sk", Scrub(content))
}
func TestRegisterRule_RejectsInvalidPattern(t *testing.T) {
t.Cleanup(ResetExtraRules)
err := RegisterRule("bad", "[unclosed")
assert.Error(t, err)
}
+129
View File
@@ -0,0 +1,129 @@
// Package gitea implements capture.IssueTracker against a Gitea instance
// over its REST API. It is the new outbound dependency the brain server
// gains for the capture capability (#49c/#52): the server otherwise does
// brain-local file ops only.
//
// Owner is hard-coded to the operator and never taken from caller input.
// The API token is read once at construction, held in the struct, and
// never logged or placed in argv — it travels only in the Authorization
// header of outbound requests (AGENTS.md secret-handling).
package gitea
import (
"bytes"
"context"
"encoding/json"
"fmt"
"io"
"net/http"
"strings"
"time"
"github.com/mathiasbq/hyperguild/ingestion/internal/capture"
)
// owner is the fixed repository owner for every ticket operation. It is a
// constant, not a parameter, so a caller can never redirect a write to
// another owner's repo.
const owner = "mathias"
// Client is a Gitea REST API IssueTracker.
type Client struct {
baseURL string
token string
http *http.Client
}
// New constructs a Client. It returns nil when either baseURL or token is
// empty, so callers can treat missing config as "tracker disabled" with a
// single nil check (mirrors embed.New).
func New(baseURL, token string) *Client {
if baseURL == "" || token == "" {
return nil
}
return &Client{
baseURL: strings.TrimRight(baseURL, "/"),
token: token,
http: &http.Client{Timeout: 15 * time.Second},
}
}
// issueResponse is the subset of a Gitea issue/comment payload we read.
type issueResponse struct {
Number int `json:"number"`
HTMLURL string `json:"html_url"`
}
// CreateIssue opens a new issue under the fixed owner.
func (c *Client) CreateIssue(ctx context.Context, repo, title, body string) (capture.IssueRef, error) {
var out issueResponse
if err := c.do(ctx, http.MethodPost,
fmt.Sprintf("/api/v1/repos/%s/%s/issues", owner, repo),
map[string]any{"title": title, "body": body}, &out); err != nil {
return capture.IssueRef{}, err
}
return capture.IssueRef{Repo: repo, Number: out.Number, URL: out.HTMLURL}, nil
}
// CommentIssue posts a comment on an existing issue.
func (c *Client) CommentIssue(ctx context.Context, repo string, number int, body string) (capture.IssueRef, error) {
var out issueResponse
if err := c.do(ctx, http.MethodPost,
fmt.Sprintf("/api/v1/repos/%s/%s/issues/%d/comments", owner, repo, number),
map[string]any{"body": body}, &out); err != nil {
return capture.IssueRef{}, err
}
return capture.IssueRef{Repo: repo, Number: number, URL: out.HTMLURL}, nil
}
// CloseIssue closes an issue, first posting a closing comment when one is
// given (empty comment ⇒ close only).
func (c *Client) CloseIssue(ctx context.Context, repo string, number int, comment string) (capture.IssueRef, error) {
if strings.TrimSpace(comment) != "" {
if _, err := c.CommentIssue(ctx, repo, number, comment); err != nil {
return capture.IssueRef{}, err
}
}
var out issueResponse
if err := c.do(ctx, http.MethodPatch,
fmt.Sprintf("/api/v1/repos/%s/%s/issues/%d", owner, repo, number),
map[string]any{"state": "closed"}, &out); err != nil {
return capture.IssueRef{}, err
}
return capture.IssueRef{Repo: repo, Number: number, URL: out.HTMLURL}, nil
}
// do performs a JSON request against the Gitea API and decodes the
// response into out. Errors carry the status and a truncated body for
// diagnosis but never the token.
func (c *Client) do(ctx context.Context, method, path string, payload any, out *issueResponse) error {
reqBody, err := json.Marshal(payload)
if err != nil {
return fmt.Errorf("marshal request: %w", err)
}
req, err := http.NewRequestWithContext(ctx, method, c.baseURL+path, bytes.NewReader(reqBody))
if err != nil {
return err
}
req.Header.Set("Content-Type", "application/json")
req.Header.Set("Accept", "application/json")
// Gitea's token scheme. Held here only; never logged.
req.Header.Set("Authorization", "token "+c.token)
resp, err := c.http.Do(req)
if err != nil {
return fmt.Errorf("gitea %s %s: %w", method, path, err)
}
defer func() { _ = resp.Body.Close() }()
respBody, _ := io.ReadAll(io.LimitReader(resp.Body, 4096))
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
return fmt.Errorf("gitea %s %s: status %d: %s", method, path, resp.StatusCode, strings.TrimSpace(string(respBody)))
}
if out != nil && len(respBody) > 0 {
if err := json.Unmarshal(respBody, out); err != nil {
return fmt.Errorf("gitea %s %s: decode response: %w", method, path, err)
}
}
return nil
}
+117
View File
@@ -0,0 +1,117 @@
package gitea_test
import (
"context"
"encoding/json"
"io"
"net/http"
"net/http/httptest"
"strings"
"testing"
"github.com/mathiasbq/hyperguild/ingestion/internal/gitea"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
)
const testToken = "super-secret-token-value"
func TestNewNilWhenUnconfigured(t *testing.T) {
assert.Nil(t, gitea.New("", testToken))
assert.Nil(t, gitea.New("https://git.example", ""))
}
func TestCreateIssueForcesOwnerAndAuth(t *testing.T) {
var gotPath, gotAuth, gotBody string
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
gotPath = r.URL.Path
gotAuth = r.Header.Get("Authorization")
b, _ := io.ReadAll(r.Body)
gotBody = string(b)
assert.Equal(t, http.MethodPost, r.Method)
w.WriteHeader(http.StatusCreated)
_ = json.NewEncoder(w).Encode(map[string]any{"number": 42, "html_url": "https://git.d-ma.be/mathias/hyperguild/issues/42"})
}))
defer srv.Close()
c := gitea.New(srv.URL, testToken)
require.NotNil(t, c)
ref, err := c.CreateIssue(context.Background(), "hyperguild", "Do the thing", "details")
require.NoError(t, err)
assert.Equal(t, "/api/v1/repos/mathias/hyperguild/issues", gotPath, "owner forced to mathias")
assert.Equal(t, "token "+testToken, gotAuth)
assert.Contains(t, gotBody, "Do the thing")
assert.Equal(t, "hyperguild", ref.Repo)
assert.Equal(t, 42, ref.Number)
assert.Contains(t, ref.URL, "/issues/42")
}
func TestCommentIssue(t *testing.T) {
var gotPath string
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
gotPath = r.URL.Path
w.WriteHeader(http.StatusCreated)
_ = json.NewEncoder(w).Encode(map[string]any{"html_url": "https://git/c/1"})
}))
defer srv.Close()
ref, err := gitea.New(srv.URL, testToken).CommentIssue(context.Background(), "hyperguild", 7, "a comment")
require.NoError(t, err)
assert.Equal(t, "/api/v1/repos/mathias/hyperguild/issues/7/comments", gotPath)
assert.Equal(t, 7, ref.Number)
}
func TestCloseIssueWithComment(t *testing.T) {
var paths []string
var states []string
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
paths = append(paths, r.Method+" "+r.URL.Path)
if r.Method == http.MethodPatch {
var body map[string]any
b, _ := io.ReadAll(r.Body)
_ = json.Unmarshal(b, &body)
states = append(states, body["state"].(string))
}
w.WriteHeader(http.StatusOK)
_ = json.NewEncoder(w).Encode(map[string]any{"number": 9, "html_url": "https://git/i/9"})
}))
defer srv.Close()
ref, err := gitea.New(srv.URL, testToken).CloseIssue(context.Background(), "hyperguild", 9, "closing because done")
require.NoError(t, err)
assert.Equal(t, 9, ref.Number)
// Comment posted first, then state PATCHed to closed.
assert.Contains(t, paths, "POST /api/v1/repos/mathias/hyperguild/issues/9/comments")
assert.Contains(t, paths, "PATCH /api/v1/repos/mathias/hyperguild/issues/9")
assert.Equal(t, []string{"closed"}, states)
}
func TestCloseIssueNoComment(t *testing.T) {
var commented bool
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
if strings.HasSuffix(r.URL.Path, "/comments") {
commented = true
}
w.WriteHeader(http.StatusOK)
_ = json.NewEncoder(w).Encode(map[string]any{"number": 3, "html_url": "https://git/i/3"})
}))
defer srv.Close()
_, err := gitea.New(srv.URL, testToken).CloseIssue(context.Background(), "hyperguild", 3, "")
require.NoError(t, err)
assert.False(t, commented, "empty comment ⇒ no comment POST")
}
func TestErrorPathDoesNotLeakToken(t *testing.T) {
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, _ *http.Request) {
w.WriteHeader(http.StatusInternalServerError)
_, _ = w.Write([]byte("boom"))
}))
defer srv.Close()
_, err := gitea.New(srv.URL, testToken).CreateIssue(context.Background(), "hyperguild", "t", "b")
require.Error(t, err)
assert.NotContains(t, err.Error(), testToken, "token must never appear in an error message")
assert.Contains(t, err.Error(), "500")
}
+204
View File
@@ -0,0 +1,204 @@
package mcp_test
import (
"context"
"encoding/json"
"os"
"path/filepath"
"strings"
"testing"
"time"
"github.com/mathiasbq/hyperguild/ingestion/internal/mcp"
"github.com/mathiasbq/hyperguild/ingestion/internal/vectorstore"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
)
// callResult parses the JSON text payload of a successful tool call.
func callResult(t *testing.T, resp map[string]any) map[string]any {
t.Helper()
require.Nil(t, resp["error"], "tool returned error: %v", resp["error"])
text := resp["result"].(map[string]any)["content"].([]any)[0].(map[string]any)["text"].(string)
var out map[string]any
require.NoError(t, json.Unmarshal([]byte(text), &out))
return out
}
func TestBrainUpdateSupersedesExisting(t *testing.T) {
brainDir := t.TempDir()
srv := mcp.NewServer(brainDir, nil, nil, nil)
// Seed via brain_write so the note carries real frontmatter.
callResult(t, toolCall(t, srv, "brain_write", map[string]any{
"content": "# Old\n\nold body\n", "filename": "val-vol",
"wing": "jepa-fx", "hall": "facts",
}))
out := callResult(t, toolCall(t, srv, "brain_update", map[string]any{
"wing": "jepa-fx", "hall": "facts", "slug": "val-vol",
"content": "# New\n\nnew body\n", "reason": "facts changed",
}))
assert.Equal(t, "wiki/jepa-fx/facts/val-vol.md", out["path"])
assert.Equal(t, out["path"], out["id"])
assert.NotEmpty(t, out["content_hash"])
assert.Equal(t, true, out["superseded"])
got, err := os.ReadFile(filepath.Join(brainDir, "wiki/jepa-fx/facts/val-vol.md"))
require.NoError(t, err)
s := string(got)
assert.Contains(t, s, "# New")
assert.NotContains(t, s, "old body")
assert.Contains(t, s, "wing: jepa-fx")
assert.Contains(t, s, "supersede_reason: facts changed")
assert.Contains(t, s, "supersedes:")
}
func TestBrainUpdateMissingTargetErrorsNoCreate(t *testing.T) {
brainDir := t.TempDir()
srv := mcp.NewServer(brainDir, nil, nil, nil)
resp := toolCall(t, srv, "brain_update", map[string]any{
"wing": "jepa-fx", "hall": "facts", "slug": "ghost",
"content": "x\n",
})
require.NotNil(t, resp["error"])
assert.Contains(t, resp["error"].(map[string]any)["message"].(string), "does not exist")
_, statErr := os.Stat(filepath.Join(brainDir, "wiki/jepa-fx/facts/ghost.md"))
assert.True(t, os.IsNotExist(statErr))
}
func TestBrainUpdateByFullPath(t *testing.T) {
brainDir := t.TempDir()
srv := mcp.NewServer(brainDir, nil, nil, nil)
callResult(t, toolCall(t, srv, "brain_write", map[string]any{
"content": "old\n", "filename": "n", "wing": "a", "hall": "facts",
}))
out := callResult(t, toolCall(t, srv, "brain_update", map[string]any{
"slug": "wiki/a/facts/n.md", "content": "fresh\n",
}))
assert.Equal(t, "wiki/a/facts/n.md", out["path"])
}
func TestBrainGetByIDAndPath(t *testing.T) {
brainDir := t.TempDir()
srv := mcp.NewServer(brainDir, nil, nil, nil)
w := callResult(t, toolCall(t, srv, "brain_write", map[string]any{
"content": "# Body\n\ntext\n", "filename": "n", "wing": "a", "hall": "facts",
}))
id := w["id"].(string)
hash := w["content_hash"].(string)
require.NotEmpty(t, id)
require.NotEmpty(t, hash)
// by id
g1 := callResult(t, toolCall(t, srv, "brain_get", map[string]any{"id": id}))
assert.Equal(t, id, g1["path"])
assert.Equal(t, hash, g1["content_hash"], "content_hash must round-trip write→get")
assert.Contains(t, g1["body"].(string), "# Body")
fm := g1["frontmatter"].(map[string]any)
assert.Equal(t, "a", fm["wing"])
// by path
g2 := callResult(t, toolCall(t, srv, "brain_get", map[string]any{"path": id}))
assert.Equal(t, hash, g2["content_hash"])
}
func TestBrainGetMissingArgsErrors(t *testing.T) {
srv := mcp.NewServer(t.TempDir(), nil, nil, nil)
resp := toolCall(t, srv, "brain_get", map[string]any{})
require.NotNil(t, resp["error"])
}
func TestBrainWriteReturnsHandle(t *testing.T) {
brainDir := t.TempDir()
srv := mcp.NewServer(brainDir, nil, nil, nil)
out := callResult(t, toolCall(t, srv, "brain_write", map[string]any{
"content": "# X\n\nbody\n", "filename": "x", "wing": "a", "hall": "facts",
}))
assert.Equal(t, "wiki/a/facts/x.md", out["path"])
assert.Equal(t, out["path"], out["id"])
assert.NotEmpty(t, out["content_hash"])
}
// --- retrieval-reflects-new-content: exercises the real mtime-driven Sync ---
type fakeVecStore struct {
chunks map[string][]float32
deleted []string
}
func (f *fakeVecStore) KnownPathsWithTime(_ context.Context) (map[string]time.Time, error) {
m := make(map[string]time.Time, len(f.chunks))
for p := range f.chunks {
m[p] = time.Unix(0, 0) // always stale → mtime(now) is always newer
}
return m, nil
}
func (f *fakeVecStore) Upsert(_ context.Context, path string, vec []float32) error {
f.chunks[path] = vec
return nil
}
func (f *fakeVecStore) Delete(_ context.Context, path string) error {
delete(f.chunks, path)
f.deleted = append(f.deleted, path)
return nil
}
type fakeEmbedder struct{ seen []string }
func (e *fakeEmbedder) Embed(_ context.Context, text string) ([]float32, error) {
e.seen = append(e.seen, text)
return []float32{1, 0, 0}, nil
}
// TestBrainUpdateReembedsNewContent proves the supersede contract end to
// end against the actual embedding mechanism: brain_update rewrites the
// file, advancing its mtime, and the next vectorstore.Sync pass re-embeds
// the NEW body and drops the stale chunk. No stub of the re-index path.
func TestBrainUpdateReembedsNewContent(t *testing.T) {
brainDir := t.TempDir()
srv := mcp.NewServer(brainDir, nil, nil, nil)
ctx := context.Background()
callResult(t, toolCall(t, srv, "brain_write", map[string]any{
"content": "# Note\n\nthe OLD distinctive payload\n",
"filename": "n", "wing": "a", "hall": "facts",
}))
store := &fakeVecStore{chunks: map[string][]float32{}}
emb := &fakeEmbedder{}
// First sync embeds the original content.
_, err := vectorstore.Sync(ctx, brainDir, store, emb)
require.NoError(t, err)
require.NotEmpty(t, store.chunks)
require.True(t, anyContains(emb.seen, "OLD distinctive payload"))
callResult(t, toolCall(t, srv, "brain_update", map[string]any{
"wing": "a", "hall": "facts", "slug": "n",
"content": "# Note\n\nthe NEW distinctive payload\n",
}))
emb.seen = nil // only watch what the second pass embeds
_, err = vectorstore.Sync(ctx, brainDir, store, emb)
require.NoError(t, err)
assert.True(t, anyContains(emb.seen, "NEW distinctive payload"),
"Sync must re-embed the superseded body; saw %v", emb.seen)
assert.False(t, anyContains(emb.seen, "OLD distinctive payload"),
"the old body must not be re-embedded")
assert.NotEmpty(t, store.deleted, "stale chunks must be deleted before re-embed")
}
func anyContains(ss []string, sub string) bool {
for _, s := range ss {
if strings.Contains(s, sub) {
return true
}
}
return false
}
+103 -14
View File
@@ -9,8 +9,8 @@ import (
"strings"
"time"
"github.com/mathiasbq/hyperguild/ingestion/internal/api"
"github.com/mathiasbq/hyperguild/ingestion/internal/brain"
"github.com/mathiasbq/hyperguild/ingestion/internal/capture"
"github.com/mathiasbq/hyperguild/ingestion/internal/extract"
"github.com/mathiasbq/hyperguild/ingestion/internal/graphsync"
"github.com/mathiasbq/hyperguild/ingestion/internal/pipeline"
@@ -61,6 +61,26 @@ func (s *Server) tools() []map[string]any {
"hall": enum("optional memory type (requires wing)", halls...),
}),
},
{
"name": "brain_update",
"description": "Supersede an existing brain note in place: whole-note body replace + frontmatter re-stamp (updated_at, supersedes=prior content hash, supersede_reason). Errors if the target does not exist — use brain_write to create. Returns {id, path, content_hash, superseded}. Prior version recoverable from git.",
"inputSchema": schema([]string{"content"}, map[string]any{
"content": str("new full body (whole-note replace)"),
"slug": str("target note slug within wing/hall, OR a full brain-relative path (e.g. wiki/jepa-fx/facts/x.md)"),
"wing": str("wing of the target (required unless slug/path is a full path)"),
"hall": enum("hall of the target (required unless slug/path is a full path)", halls...),
"path": str("full brain-relative path to the target; takes precedence over slug/wing/hall"),
"reason": str("optional short note on why superseded — stamped into frontmatter"),
}),
},
{
"name": "brain_get",
"description": "Fetch a single brain note by id or path (both are the brain-relative path — the note handle). Returns {id, path, content_hash, frontmatter, body}. Read-after-write confirmation without a lexical re-query.",
"inputSchema": schema([]string{}, map[string]any{
"id": str("note id (brain-relative path) as returned by brain_write/brain_update"),
"path": str("brain-relative path to the note; equivalent to id"),
}),
},
{
"name": "brain_tunnel",
"description": "Create an explicit bidirectional [[wikilink]] between two notes in different wings. Idempotent.",
@@ -199,7 +219,11 @@ func (s *Server) brainWrite(ctx context.Context, args json.RawMessage) (json.Raw
if err := json.Unmarshal(args, &a); err != nil {
return nil, fmt.Errorf("parse args: %w", err)
}
relPath, err := api.WriteNote(s.brainDir, api.WriteNoteOptions{
// Delegate to the shared BrainStore so write+index+tunnel+graph live in
// one implementation (capture uses the same store). The read-after-write
// handle {id, path, content_hash} comes back from the store; path is kept
// for backward compatibility.
ref, err := s.store.Write(ctx, capture.Note{
Content: a.Content,
Filename: a.Filename,
Type: a.Type,
@@ -210,19 +234,84 @@ func (s *Server) brainWrite(ctx context.Context, args json.RawMessage) (json.Raw
if err != nil {
return nil, err
}
// Auto-regenerate the wing _index.md when the write landed in the
// structured wiki, and auto-tunnel cross-wing matches. Both are
// best-effort: the note is already written.
if a.Wing != "" && a.Hall != "" {
if err := brain.BuildWingIndex(s.brainDir, a.Wing); err != nil {
slog.Warn("brain_write: auto-index failed", "wing", a.Wing, "err", err)
}
if err := brain.AutoTunnel(s.brainDir, relPath, a.Content); err != nil {
slog.Warn("brain_write: auto-tunnel failed", "src", relPath, "err", err)
}
return json.Marshal(map[string]string{"id": ref.ID, "path": ref.Path, "content_hash": ref.ContentHash})
}
type brainUpdateArgs struct {
Slug string `json:"slug,omitempty"`
Wing string `json:"wing,omitempty"`
Hall string `json:"hall,omitempty"`
Path string `json:"path,omitempty"`
Content string `json:"content"`
Reason string `json:"reason,omitempty"`
}
// brainUpdate supersedes an existing note in place: whole-note body
// replace, frontmatter re-stamp (updated_at/supersedes/supersede_reason),
// graph re-index, and wing _index rebuild. It never creates — a missing
// target is an error so the caller can fall back to brain_write.
//
// Embedding re-sync is delegated to the out-of-band vectorstore.Sync
// ticker: the rewritten file's mtime advances, so the next pass re-embeds
// it. This mirrors brain_write, which likewise does not embed in-handler.
func (s *Server) brainUpdate(ctx context.Context, args json.RawMessage) (json.RawMessage, error) {
var a brainUpdateArgs
if err := json.Unmarshal(args, &a); err != nil {
return nil, fmt.Errorf("parse args: %w", err)
}
s.indexInGraph(ctx, "brain_write", relPath)
return json.Marshal(map[string]string{"path": relPath})
if a.Content == "" {
return nil, fmt.Errorf("content is required")
}
// path takes precedence over slug; the store treats any slug containing
// a slash as a full brain-relative path (issue #45: "slug ... OR path").
slug := a.Slug
if a.Path != "" {
slug = a.Path
}
ref, err := s.store.Update(ctx, slug, capture.Note{
Content: a.Content,
Wing: a.Wing,
Hall: a.Hall,
Reason: a.Reason,
})
if err != nil {
return nil, err
}
return json.Marshal(map[string]any{
"id": ref.ID, "path": ref.Path, "content_hash": ref.ContentHash, "superseded": ref.Superseded,
})
}
type brainGetArgs struct {
ID string `json:"id,omitempty"`
Path string `json:"path,omitempty"`
}
// brainGet fetches a note by id or path (both are the brainDir-relative
// path — the de-facto handle). Read-only; the create-path read-after-
// write primitive that lets callers confirm a write landed without a
// lexical re-query.
func (s *Server) brainGet(ctx context.Context, args json.RawMessage) (json.RawMessage, error) {
var a brainGetArgs
if err := json.Unmarshal(args, &a); err != nil {
return nil, fmt.Errorf("parse args: %w", err)
}
target := a.Path
if target == "" {
target = a.ID
}
if target == "" {
return nil, fmt.Errorf("id or path is required")
}
note, err := s.store.Get(ctx, target)
if err != nil {
return nil, err
}
return json.Marshal(map[string]any{
"id": note.ID, "path": note.Path, "content_hash": note.ContentHash,
"frontmatter": note.Frontmatter, "body": note.Body,
})
}
// indexInGraph is a best-effort wrapper around graphsync.IndexDoc that
+34 -4
View File
@@ -1,7 +1,7 @@
// Package mcp implements an MCP HTTP handler for the ingestion service.
// Exposed tools: brain_query, brain_write, brain_index, brain_tunnel,
// brain_ingest, brain_ingest_raw, brain_answer, brain_classify,
// brain_graph, brain_context, session_log.
// Exposed tools: brain_query, brain_write, brain_update, brain_get,
// brain_index, brain_tunnel, brain_ingest, brain_ingest_raw,
// brain_answer, brain_classify, brain_graph, brain_context, session_log.
package mcp
import (
@@ -10,6 +10,8 @@ import (
"fmt"
"net/http"
"github.com/mathiasbq/hyperguild/ingestion/internal/brainstore"
"github.com/mathiasbq/hyperguild/ingestion/internal/capture"
"github.com/mathiasbq/hyperguild/ingestion/internal/graphstore"
"github.com/mathiasbq/hyperguild/ingestion/internal/graphsync"
"github.com/mathiasbq/hyperguild/ingestion/internal/pipeline"
@@ -46,6 +48,8 @@ type Server struct {
vector search.VectorSearcher // nil = BM25-only retrieval
embedder search.Embedder // nil = BM25-only retrieval
graph graphsync.Store // nil = brain_graph and GraphRAG augmentation disabled
store *brainstore.Store // shared brain write/update/get impl (also used by capture)
tracker capture.IssueTracker // nil = no Gitea ticket integration; wired for capture (#53)
}
// NewServer constructs a Server bound to brainDir. pipelineCfg supplies the
@@ -56,7 +60,13 @@ func NewServer(brainDir string, pipelineCfg *pipeline.Config, llm pipeline.Compl
if pipelineCfg != nil {
cfg = *pipelineCfg
}
return &Server{brainDir: brainDir, pipeline: cfg, llm: llm, answerLLM: answerLLM}
return &Server{
brainDir: brainDir,
pipeline: cfg,
llm: llm,
answerLLM: answerLLM,
store: brainstore.New(brainDir),
}
}
// WithReranker installs an opt-in cross-encoder reranker. When set,
@@ -84,12 +94,28 @@ func (s *Server) WithHybridRetrieval(v search.VectorSearcher, e search.Embedder)
func (s *Server) WithGraph(g *graphstore.PGStore) *Server {
if g == nil {
s.graph = nil
s.store.WithGraph(nil)
return s
}
s.graph = g
s.store.WithGraph(g)
return s
}
// WithIssueTracker injects the Gitea ticket tracker behind the
// capture.IssueTracker interface. nil leaves ticket integration off. The
// use-case (capture) consumes this in #53; it is wired here so the
// dependency is constructed once and stays swappable/testable.
func (s *Server) WithIssueTracker(t capture.IssueTracker) *Server {
s.tracker = t
return s
}
// IssueTracker returns the injected ticket tracker (nil when unconfigured).
func (s *Server) IssueTracker() capture.IssueTracker {
return s.tracker
}
func (s *Server) ServeHTTP(w http.ResponseWriter, r *http.Request) {
// MCP streamable HTTP: GET establishes the SSE stream for server-to-client events.
if r.Method == http.MethodGet {
@@ -177,6 +203,10 @@ func (s *Server) handleCall(ctx context.Context, name string, args json.RawMessa
return s.brainQuery(ctx, args)
case "brain_write":
return s.brainWrite(ctx, args)
case "brain_update":
return s.brainUpdate(ctx, args)
case "brain_get":
return s.brainGet(ctx, args)
case "brain_index":
return s.brainIndex(ctx, args)
case "brain_tunnel":
+23 -1
View File
@@ -2,12 +2,14 @@ package mcp_test
import (
"bytes"
"context"
"encoding/json"
"net/http"
"net/http/httptest"
"strings"
"testing"
"github.com/mathiasbq/hyperguild/ingestion/internal/capture"
"github.com/mathiasbq/hyperguild/ingestion/internal/mcp"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
@@ -55,7 +57,8 @@ func TestServerToolsList(t *testing.T) {
names = append(names, t.(map[string]any)["name"].(string))
}
assert.ElementsMatch(t, []string{
"brain_query", "brain_write", "brain_index", "brain_tunnel",
"brain_query", "brain_write", "brain_update", "brain_get",
"brain_index", "brain_tunnel",
"brain_ingest_raw", "brain_ingest",
"brain_answer", "brain_classify", "brain_graph", "brain_context",
"session_log",
@@ -92,3 +95,22 @@ func TestServerUnknownMethodReturnsError(t *testing.T) {
assert.Equal(t, float64(-32601), errObj["code"])
assert.Contains(t, errObj["message"].(string), "unknown/method")
}
type stubTracker struct{}
func (stubTracker) CreateIssue(context.Context, string, string, string) (capture.IssueRef, error) {
return capture.IssueRef{}, nil
}
func (stubTracker) CloseIssue(context.Context, string, int, string) (capture.IssueRef, error) {
return capture.IssueRef{}, nil
}
func (stubTracker) CommentIssue(context.Context, string, int, string) (capture.IssueRef, error) {
return capture.IssueRef{}, nil
}
func TestWithIssueTrackerInjects(t *testing.T) {
srv := mcp.NewServer(t.TempDir(), nil, nil, nil)
assert.Nil(t, srv.IssueTracker(), "tracker is off by default")
srv = srv.WithIssueTracker(stubTracker{})
assert.NotNil(t, srv.IssueTracker(), "tracker injected behind the interface")
}
+16 -3
View File
@@ -77,9 +77,22 @@ func (s *Server) brainAnswer(ctx context.Context, args json.RawMessage) (json.Ra
return nil, fmt.Errorf("search: %w", err)
}
if s.reranker != nil && len(results) > 0 {
results, err = rerankResults(ctx, s.reranker, a.Query, results, 5)
if err != nil {
return nil, fmt.Errorf("rerank: %w", err)
reranked, rerr := rerankResults(ctx, s.reranker, a.Query, results, 5)
if rerr != nil {
return nil, fmt.Errorf("rerank: %w", rerr)
}
// The reranker is a filter, not a gate. The Qwen3-Reranker is a
// web-search cross-encoder: against a conversational / personal-
// intent query ("what am I optimizing toward?") it scores even
// on-topic notes as "no", which would collapse the whole answer to
// "no relevant content" despite BM25 having retrieved relevant
// content. When the reranker keeps nothing, fall back to the
// BM25/vector ordering (capped to the no-reranker depth) rather
// than returning an empty answer.
if len(reranked) > 0 {
results = reranked
} else if len(results) > 10 {
results = results[:10]
}
}
if len(results) == 0 {
@@ -98,6 +98,42 @@ func TestBrainAnswer_RerankerFiltersBeforeLLM(t *testing.T) {
assert.NotContains(t, sawSources, "noise.md")
}
func TestBrainAnswer_RerankerKeepsNone_FallsBackToBM25(t *testing.T) {
brainDir := brainDirWithContent(t) // test.md BM25-matches "pass-rate logging"
// Reranker rejects every candidate ("no" to all) — models a
// web-search cross-encoder facing a conversational / personal-intent
// query, which is exactly when it wrongly scores on-topic notes as
// irrelevant. The answer must still synthesize from the BM25 hits, not
// collapse to "no relevant content".
rrSrv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, _ *http.Request) {
_ = json.NewEncoder(w).Encode(map[string]any{"response": "no", "done": true})
}))
defer rrSrv.Close()
var sawSources string
llm := func(_ context.Context, _, user string) (string, error) {
sawSources = user
return "fallback answer", nil
}
srv := mcp.NewServer(brainDir, nil, nil, llm).
WithReranker(reranker.New(rrSrv.URL, "qwen3"))
ts := httptest.NewServer(srv)
defer ts.Close()
rpc := callTool(t, ts, "brain_answer", map[string]any{"query": "pass-rate logging"})
require.Nil(t, rpc["error"])
content := rpc["result"].(map[string]any)["content"].([]any)[0].(map[string]any)["text"].(string)
var result map[string]any
require.NoError(t, json.Unmarshal([]byte(content), &result))
assert.Equal(t, "fallback answer", result["answer"])
assert.NotEmpty(t, result["sources"], "reranker keeping nothing must fall back to BM25, not empty")
assert.Contains(t, sawSources, "test.md")
}
func TestBrainAnswer_NoLLM(t *testing.T) {
srv := mcp.NewServer(t.TempDir(), nil, nil, nil)
ts := httptest.NewServer(srv)
+65
View File
@@ -3,6 +3,7 @@ package vectorstore
import (
"fmt"
"strings"
"unicode/utf8"
)
// NumberedChunk pairs a chunk's body with the storage path it will use
@@ -66,6 +67,70 @@ func ChunkMarkdown(content string, maxBytes int) []string {
}
out = append(out, splitAtParagraphs(s, maxBytes)...)
}
// Final guarantee: no chunk exceeds maxBytes. A single heading-less,
// paragraph-less block (JSON-lines, minified content) survives the two
// passes above whole — splitAtParagraphs emits an over-budget paragraph
// rather than truncating prose. Hard-split any such chunk at line/rune
// boundaries so the embedder never rejects an over-context chunk.
final := make([]string, 0, len(out))
for _, c := range out {
if len(c) <= maxBytes {
final = append(final, c)
continue
}
final = append(final, hardSplit(c, maxBytes)...)
}
return final
}
// hardSplit slices s into pieces no larger than maxBytes, breaking at line
// boundaries where possible and otherwise mid-line at a UTF-8 rune boundary.
// Last resort for content that has neither headings nor blank-line paragraphs.
func hardSplit(s string, maxBytes int) []string {
var out []string
var cur strings.Builder
flush := func() {
if cur.Len() > 0 {
out = append(out, cur.String())
cur.Reset()
}
}
for _, line := range strings.SplitAfter(s, "\n") {
if line == "" {
continue
}
if len(line) > maxBytes {
flush()
out = append(out, runeSplit(line, maxBytes)...)
continue
}
if cur.Len() > 0 && cur.Len()+len(line) > maxBytes {
flush()
}
cur.WriteString(line)
}
flush()
return out
}
// runeSplit slices s into <=maxBytes pieces without splitting a UTF-8 rune.
func runeSplit(s string, maxBytes int) []string {
var out []string
for len(s) > maxBytes {
cut := maxBytes
for cut > 0 && !utf8.RuneStart(s[cut]) {
cut--
}
if cut == 0 { // single rune wider than the budget; emit it whole
cut = maxBytes
}
out = append(out, s[:cut])
s = s[cut:]
}
if len(s) > 0 {
out = append(out, s)
}
return out
}
@@ -30,6 +30,23 @@ func TestChunkMarkdown_SplitsAtHeadings(t *testing.T) {
}
}
func TestChunkMarkdown_HardSplitsHeadinglessOversizedBlock(t *testing.T) {
// A document with no headings and no blank-line paragraph breaks (e.g.
// JSON-lines like wiki/telos/decisions/human-intent-column.md). The old
// chunker emitted it as one over-budget chunk → nomic-embed returned
// "input length exceeds the context length" (400). Every chunk must now
// fit the budget, with no content lost.
maxBytes := 200
src := strings.Repeat("x", 1000) // one 1000-byte blob, no headings, no \n\n
out := vectorstore.ChunkMarkdown(src, maxBytes)
require.Greater(t, len(out), 1, "oversized blob must be split")
for i, c := range out {
assert.LessOrEqual(t, len(c), maxBytes, "chunk %d over budget: %d bytes", i, len(c))
}
assert.Equal(t, 1000, strings.Count(strings.Join(out, ""), "x"), "no content lost")
}
func TestChunkMarkdown_FurtherSplitsOversizedSection(t *testing.T) {
// One H2 section with 4 paragraphs of ~80 chars each, limit 100.
src := "## big\n\n" +
+3 -32
View File
@@ -40,10 +40,9 @@ if [ -n "$ROOT_CONTEXT" ] && [ -f "$ROOT_CONTEXT" ]; then
echo " Root context: $ROOT_CONTEXT"
else
# No reachable root AGENT.md — common in CI's clean checkout. The root+project
# adapters (AGENTS.md, .cursorrules, .aider.conventions.md, system-prompt.txt)
# require the root context to regenerate correctly, so we skip them entirely
# and only regenerate CLAUDE.md (which is project-only and inherits root via
# tree walk in Claude Code itself).
# adapters (AGENTS.md, system-prompt.txt) require the root context to
# regenerate correctly, so we skip them entirely and only regenerate CLAUDE.md
# (which is project-only and inherits root via tree walk in Claude Code itself).
echo " No root AGENT.md found — regenerating CLAUDE.md only"
echo "Syncing project context from $PROJECT_FILE..."
cat "$PROJECT_FILE" > CLAUDE.md
@@ -78,30 +77,6 @@ generate_agents() {
echo " → AGENTS.md (root + project; Crush, Pi, Antigravity)"
}
# ── Cursor ───────────────────────────────────────────────────
generate_cursor() {
{
echo "# Cursor rules — auto-generated"
echo "# Do not edit. Run: task context:sync"
echo ""
root_block
cat "$PROJECT_FILE"
} > .cursorrules
echo " → .cursorrules (root + project)"
}
# ── Aider ────────────────────────────────────────────────────
generate_aider() {
{ root_block; cat "$PROJECT_FILE"; } > .aider.conventions.md
if [ ! -f .aider.conf.yml ]; then
cat > .aider.conf.yml << 'YAML'
read: .aider.conventions.md
auto-commits: false
YAML
fi
echo " → .aider.conventions.md (root + project)"
}
# ── Generic system prompt (Open WebUI, Mods, etc.) ──────────
generate_system_prompt() {
{
@@ -142,8 +117,6 @@ echo "Syncing project context from $PROJECT_FILE..."
if [ $# -eq 0 ]; then
generate_claude
generate_agents
generate_cursor
generate_aider
generate_system_prompt
generate_mcp
else
@@ -151,8 +124,6 @@ else
case "$adapter" in
claude) generate_claude ;;
agents) generate_agents ;;
cursor) generate_cursor ;;
aider) generate_aider ;;
prompt|system|openwebui|owui|generic) generate_system_prompt ;;
mcp) generate_mcp ;;
*) echo "Unknown adapter: $adapter" ;;
+109
View File
@@ -0,0 +1,109 @@
---
name: close-session
description: Disciplined end-of-session closeout for a Claude.ai chat before archiving it. Harvests the session's decisions, artifacts, and open threads and durably persists them to the brain MCP and the right Gitea repo so nothing is lost when context resets. Use this whenever the user signals they are wrapping up — phrases like "close this out", "let's wrap up", "before I archive", "session retro", "capture this before I go", "did we lose anything", or any end-of-session/handoff cue — even if they don't say the word "close". Also use when the user explicitly asks to retro, archive, or hand off a working session.
---
# close-session
Capture a finishing Claude.ai work session into durable storage before the chat is archived and its context is lost. The goal is simple and load-bearing: **after this runs, a fresh session (or another agent) can reconstruct what was decided, what was shipped, and what is still open — without the original chat.**
This skill is **batch**: one session in, findings out, done. It does not loop or re-read its own fresh output semantically (see Phase 5). Run the phases in order. Stop at any confirmation gate that says STOP.
## Operating constraints (read first)
- **Gitea owner is always `mathias`.** Never guess another owner.
- **Ground-truth at HEAD before acting.** Issue bodies and doc references rot — stale hostnames, retired services, moved endpoints. Before closing/commenting on any issue, `gitea:issue_get` it fresh. Before asserting an infra fact, verify it; do not copy it from memory or from a stale issue body.
- **Current infra truths** (verify rather than trust, but these are the known-good baseline): Gitea is `git.d-ma.be` (not `gitea.d-ma.be`). LiteLLM is `http://koala:30401/v1/` (public `https://llm-api.d-ma.be`); piguard runs NGINX Proxy Manager only — never reference `piguard:4000` or `koala:4000`. Identity provider is Authentik (Dex migration complete).
- **Side-effects need a confirmation gate.** Closing issues, committing files, and writing to the brain are all real writes. Surface exactly what will happen and get a clear yes before doing it. Reads are free; writes are gated.
- **Never fabricate.** If the session didn't produce a decision worth persisting, say so and skip that write. An empty-but-honest closeout beats an invented one.
## Phase 1 — Harvest
Reconstruct what actually happened this session from the conversation itself. Produce, in working memory:
- **Decisions taken** — what was decided and the reasoning, not just the outcome.
- **Artifacts produced** — issues filed/closed, PRs opened/merged, files committed, brain notes written, ADRs. Capture identifiers (issue numbers, PR numbers, paths, commit SHAs) as you go.
- **Open threads** — what was deferred, what's blocked, what the next session should pick up.
- **Generalizable learnings** — reusable patterns or footguns that would bite anyone again (these are brain-worthy; project status is not).
Be honest about fidelity: a long session compresses harder at the start than the end. Flag anything you're reconstructing rather than certain of.
## Phase 2 — Ground-truth Gitea state
For every repo touched this session, get its true current state before proposing any change. `gitea:repo_status` (owner `mathias`) gives branches + open PRs + protection in one call. For each issue you intend to close, comment on, or reference: `gitea:issue_get` it fresh and compare to what the session assumed. Note any drift (closed-already, body rotted, renamed) — you'll surface it in Phase 3.
Do not write anything in this phase. This is the read pass.
## Phase 3 — Confirm and act on issue changes
Present a single consolidated plan of issue actions: which to close (with closing comment), which to file (discovered-but-deferred work — token-budget gaps, recorded limitations, v2 follow-ups), which to comment on. Include the exact title/body for any new issue and the closing rationale for any close.
**GATE — STOP and get explicit confirmation before any issue write.** Issue closes and new issues are side-effects. Once confirmed, execute them (`gitea:issue_close`, `gitea:issue_create`, `gitea:issue_comment`, all owner `mathias`), correcting any rotted references you found in Phase 2 as you go.
## Phase 4 — Commit the canonical session summary
Write one summary file to `mathias/ai-sessions`, committed directly to `main` via `gitea:file_write_branch` (no PR — this repo is solo and unprotected; if branch protection is ever added, fall back to a branch + PR).
**Path:** `summaries/claudeai/<YYYY-MM>/<YYYY-MM-DD>-<topic-slug>-<chatid8>.md`
where `<chatid8>` is the first 8 chars of the chat's UUID if known, else a short stable slug. `claudeai` has no host segment — Claude.ai is Anthropic-side, not a homelab host.
**Frontmatter — the REDUCED live-capture schema.** A live close-session capture cannot populate the batch-export telemetry (token counts, message counts, duration_ms, permission_mode) — those only exist in the account export pipeline. Write only what's truthfully known, and mark fidelity so a reader (or the batch pipeline) can tell a live capture from an export:
```yaml
---
title: "<concise session title>"
client: "claudeai"
interface: "claudeai-chat"
date: "<YYYY-MM-DD>"
repos_touched: [<repo slugs>]
topic_tags: [<tags>]
outcome: "<shipped|in-progress|abandoned>"
fidelity: "live-capture" # NOT an export; reconstructed live from chat
captured_by: "close-session-skill"
---
```
Do not invent the export-only fields. `fidelity: live-capture` is the honest signal; if the batch export later produces a richer summary for the same session, the export is source of truth and supersedes this.
**Body** (keep it reconstructable, not exhaustive):
```markdown
## One-paragraph summary
## Decisions
## Key artifacts
## Open threads
```
**GATE — STOP, show the full file (path + frontmatter + body), get explicit confirmation before committing.**
## Phase 5 — Brain orientation note (the durable "where we are" record)
Write one brain note so a fresh session can orient without the chat. This uses the `brain_update`/`brain_get` verbs (live since 2026-06).
**Target:** `wing: <domain>` (the project/topic domain, e.g. `hyperguild`, `jepa-fx`), `hall: decisions`. The note is a knowledge-type record (a decision/orientation), grouped by knowledge-type, not by interface surface.
**Batch read-after-write discipline (important — do these in order, do not interleave):**
1. **Read first, before any write.** Check whether an orientation note already exists for this wing/topic. Do your "does this already exist / what should I supersede" reads NOW, up front. BM25/keyword search and `brain_get` are immediate; semantic/vector search may lag up to ~5 min after a write, so never rely on a semantic query to find something you wrote earlier in this same run.
2. **Write or supersede:**
- **New note** → `brain_write` (wing, hall: decisions). Returns `{id, path, content_hash}`.
- **Superseding a prior orientation note** → `brain_update` (slug or path, wing, hall, content, reason). Whole-note replace; stamps `supersedes`/`updated_at`; returns `{id, path, content_hash, superseded}`. Use this instead of a second `brain_write` to the same slug — blind re-write creates duplicates/contradictions, which is the exact failure brain_update exists to prevent.
3. **Confirm it landed** via `brain_get(id)` and check the returned `content_hash` matches what the write returned. This is the read-after-write confirmation — do it with `brain_get`, never a semantic query.
**RULE: no semantic/vector brain query after the first `brain_update` in this run.** The batch shape makes this natural — read up front, write, confirm by id. If you ever find the skill wanting to semantic-search a just-superseded note, stop and flag it (that's the signal the staleness window matters and needs the synchronous-reembed follow-up).
**GATE — STOP, show the note (target wing/hall, new-vs-supersede, full content), get explicit confirmation before the brain write.**
After the note lands, if it relates to a note in another wing, create the cross-link inline with `brain_tunnel(source, target)` (idempotent; both paths brain-relative, must be in different wings). Optionally append a `session_log` entry (`session_id`, `skill: close-session`, `phase`, `final_status`) for telemetry. Both are now callable directly from Claude.ai — no Claude Code/Crush handoff needed.
## Phase 6 — Verdict
Deliver a final "safe to archive" verdict in the chat. Either:
- **SAFE TO ARCHIVE** — list what landed (issues closed/filed with numbers, summary path, brain note id, any tunnels) so the trail is auditable. Then list anything still in the user's queue (e.g. a PR awaiting their merge, a decision owed next session).
- **NOT YET** — name the specific gate that wasn't passed or the write that failed, and what to do about it.
Never claim safe-to-archive if any gated write was declined or errored. The verdict is the skill's contract: if it says safe, the session can be lost without losing the work.
## Why the gates and the batch discipline matter
The whole point is durability across a context reset. Every gate is a place where a wrong write would silently corrupt the record (close the wrong issue, overwrite a good brain note, commit a half-truth). The batch read-discipline in Phase 5 exists because the brain's vector index refreshes out-of-band: write-then-semantically-reread in the same run can read stale, so the skill front-loads reads and confirms writes by id. Get those right and the skill does what it promises — nothing important is lost when the chat goes away.
+248
View File
@@ -0,0 +1,248 @@
# Capture capability — use-case & BDD specification
**Status:** Decisions resolved 2026-06-22 (§4). Ready for implementation scoping. `capture` is a
privileged cross-harness write path touching brain + Gitea + ai-sessions.
**Tracks:** hyperguild #49.
**Governed by:** `infra/docs/architecture/01-invariants.md` (I1I5), the admissibility test in
`00-synthesis-model.md`, and the distributed-consolidation shape mandated by
`brain/wiki/homelab/decisions/no-centralized-cross-harness-observer-2026-06-17.md`.
---
## 1. Use-case (Clean Architecture form)
**Name:** CaptureSession
**Actor:** A harness acting on the user's behalf (claude.ai Chat/Cowork/Code/Design, Claude Code
CLI, Crush, Pi, LLM Council, Agentsquad executor/reviewer) — or the user directly.
**Goal:** Durably persist a finished session's valuable output — insights → brain, action items →
Gitea tickets, optional summary → ai-sessions — with one uniform invocation, identical core
behaviour across harnesses.
**Primary success scenario (essential steps):**
1. Caller assembles capture input (insights, tickets, optional summary) + context (harness,
session_ref, fidelity, actor, **data-classification**).
2. System validates the whole request (fail-closed).
3. System resolves **effective classification** (stricter of caller-declared and target-derived)
and the **server-derived harness origin** (from the authenticated principal). It checks the
**sovereignty gate** (I1): if effective classification is confidential AND the origin is a
non-sovereign (us-nexus) surface, the capture is **refused** before any write.
4. System persists insights (write or supersede), tickets (create/close/comment), summary — each
best-effort, recording per-item outcome.
5. System emits an **audit record** (I5) of who/what captured what, when, via which principal.
6. System returns a structured, partial-aware receipt.
**Architectural shape:** the *logic* is a shared use-case (`CaptureService`), invoked **per-harness
against the caller's own credentials** (distributed consolidation — no high-degree observer node).
A central authenticated relay endpoint exists ONLY as a fallback for harnesses that cannot run the
use-case in-process (Crush/Pi/headless); the relay holds no standing visibility and retains nothing
beyond the I5 audit log.
---
## 2. Invariant obligations (acceptance gates, not nice-to-haves)
| Invariant | Obligation on `capture` |
|---|---|
| **I1 sovereign containment** | A confidential-classified session MUST NOT be captured through a us-nexus harness. Harness origin is **server-derived from the authenticated principal** (not caller-asserted). Classification uses **model (C)**: caller declares, server cross-checks the target's tag, **stricter wins**, mismatch logged. See §4.14.2. |
| **I2 deliberate acceptance** | The *distributed-library* form opens no new acceptance. IF a central relay node is deployed, its cross-harness reach MUST be entered in `infra/docs/security-baseline.md` with Why-accepted / Revisit-if before it ships. |
| **I3 GitOps reconcilability** | IF `capture` runs as a deployed service, its manifest lives under `infra/k3s/apps/**` (sovereign source, Flux-reconciled). No untracked runtime. |
| **I4 decisions captured** | The distributed-vs-central decision and the intent-named-verb pattern are recorded (ADR + brain). |
| **I5 auditability** | Every capture emits a request-level audit record (actor/principal, harness, items written, timestamp) to the alloy/loki substrate. **Classification-aware degradation** (§4.4): confidential + sink-down → hard-refuse; internal/public + sink-down → durable local buffer + ntfy + reconcile. Floor: refuse if nothing can record the audit. |
---
## 3. BDD scenarios (Gherkin)
```gherkin
Feature: Capture session value uniformly across harnesses
As an operator working across many AI harnesses
I want one uniform command to persist insights and file tickets
So that valuable session output is never lost and is always auditable
Background:
Given a brain store, a Gitea issue tracker, and an ai-sessions summary writer
And the caller is authenticated with a principal
And the session context declares a harness, a fidelity, and a data classification
# --- Core happy path ---
Scenario: Capture insights and tickets from a non-confidential session
Given a session classified as "internal"
And the capture input has 2 insights and 1 ticket to create
When capture is invoked
Then both insights are written to the brain and their ids and content hashes are returned
And the ticket is created in the named repo under owner "mathias"
And an audit record is emitted naming the principal, harness, and items written
And the receipt reports every item as ok
# --- I1: sovereignty gate (the load-bearing refusal) ---
# Harness origin is server-derived from the authenticated principal, never from context.harness.
Scenario: Refuse capture of a confidential session through a us-nexus harness
Given a session whose effective classification is "confidential"
And the authenticated principal resolves to a us-nexus harness origin
When capture is invoked
Then the capture is refused before any write
And no insight, ticket, or summary is persisted
And the refusal names the sovereignty invariant as the reason
Scenario: Allow capture of a confidential session through a sovereign harness
Given a session whose effective classification is "confidential"
And the authenticated principal resolves to a sovereign-soil harness origin
When capture is invoked
Then the capture proceeds and persists normally
Scenario: Ignore a caller-asserted harness label and use the server-derived origin
Given the request context asserts harness "sovereign-soil"
But the authenticated principal resolves to a us-nexus origin
And the session classification is "confidential"
When capture is invoked
Then the capture is refused
And the server-derived origin is used, not the asserted label
And the asserted-vs-derived discrepancy is logged as a security event
# --- I1: classification model (C) — stricter of declared vs target-derived wins ---
Scenario: Take the stricter classification when caller and target disagree
Given the caller declares classification "internal"
But the target wing/repo is tagged "confidential"
When capture is invoked
Then the effective classification is "confidential"
And the declared-vs-derived mismatch is logged as a security event
And the I1 gate is evaluated against "confidential"
Scenario: Honour a caller raising sensitivity above the target's tag
Given the caller declares classification "confidential"
And the target wing/repo is tagged "internal"
When capture is invoked
Then the effective classification is "confidential"
And the capture is gated as confidential
# --- Supersession + staleness discipline (reuses #45 / #47 resolution) ---
Scenario: Supersede a prior insight rather than duplicating it
Given an insight whose context names an existing note to supersede
When capture is invoked
Then the existing note is updated in place, not duplicated
And the prior content hash is recorded in the superseding note
And read-after-write confirmation uses a direct fetch, never a semantic query
# --- Validation: fail-closed ---
Scenario: Reject a malformed request before any write
Given a capture input with an invalid wing/hall or unknown repo
When capture is invoked
Then the request is rejected with a validation error
And nothing is written to the brain, Gitea, or ai-sessions
# --- Partial failure: best-effort + honest receipt ---
Scenario: Report partial success when one item fails mid-capture
Given a capture input with 2 insights and 1 ticket
And the second insight write will fail
When capture is invoked
Then the first insight and the ticket are persisted
And the second insight is reported as failed in the receipt
And no rollback is attempted
And the audit record reflects exactly what landed
# --- Dry run ---
Scenario: Preview a capture without writing
Given a valid capture input with dry_run true
When capture is invoked
Then the would-be receipt is returned
And nothing is written anywhere
# --- I5: auditability is classification-aware (confidential fails closed) ---
Scenario: Confidential capture hard-refuses when the central audit sink is down
Given the effective classification is "confidential"
And the central audit substrate (loki) cannot be written to
When capture is invoked
Then the capture is refused before any write
And the reason names the auditability invariant
# Confidential work must be centrally auditable at write time — no buffered exception.
Scenario: Internal capture degrades to a durable local buffer when the sink is down
Given the effective classification is "internal" or "public"
And the central audit substrate (loki) cannot be written to
When capture is invoked
Then the capture proceeds
And the audit record is written to a durable LOCAL fallback buffer
And an ntfy alert is emitted naming the degraded audit state
And the receipt flags that audit was buffered locally, not centrally recorded
Scenario: Locally buffered audit records reconcile to the central sink on recovery
Given internal-tier audit records were buffered locally during a sink outage
When the central audit substrate becomes reachable again
Then the buffered records are replayed to the central sink
And the local buffer is cleared only after confirmed central write
Scenario: Even internal capture refuses if neither sink nor local buffer can be written
Given the effective classification is "internal" or "public"
And neither the central sink nor the local fallback buffer can be written
When capture is invoked
Then the capture is refused
And the reason names the auditability invariant
# Degrade-and-warn has a floor: if NOTHING can record the audit, do not write.
# --- Summary fidelity (collision rule from the retro work) ---
Scenario: A richer-fidelity summary supersedes a thinner one for the same session
Given a summary already exists for session_ref X at fidelity "live-capture"
And a new summary arrives for session_ref X at fidelity "transcript-parse"
When capture is invoked
Then the transcript-parse summary supersedes the live-capture one
And the live-capture summary is not left as a contradicting duplicate
```
---
## 4. Resolved decisions (2026-06-22)
These were open questions at draft; resolved in the 2026-06-22 review session. Recorded here as
binding design decisions for the build.
1. **Classification trust — model (C): caller-declares + server-cross-checks, stricter wins.**
The caller declares `context.classification`; the server **independently derives** the target's
classification (from the target wing/repo's classification tag) and gates on the **stricter of
the two**. The caller can voluntarily *raise* sensitivity but can never *lower* it below the
target's floor. A declared-vs-derived **mismatch is logged as a security event** (I5).
- **Prerequisite (new build work):** a classification taxonomy (e.g. `public` /
`internal` / `confidential`) and a per-wing / per-repo classification tag the server can read.
This must exist before the I1 gate is load-bearing. Tracked as a sub-task of #49.
- **Implemented (#50):** taxonomy `public < internal < confidential` (ordered so "stricter wins"
is `max`) in `ingestion/internal/classification/`. Tags are read from an optional
`classification.yaml` at the brain root (`wings:` / `repos:` maps); absent entries fall to
built-in defaults (`client-*` → confidential; `hyperguild`/`homelab` → internal; everything
else → **confidential, fail-safe**). `Config.Derive(Target)` is the function the use-case
calls. See brain `wiki/hyperguild/decisions/capture-classification-taxonomy`.
- Rationale: composes with decision 2; fails safe; honours a caller flagging something *more*
sensitive than its destination. Pure caller-trust (A) was rejected — it makes the gate theatre.
2. **Sovereign-harness determination — server-derived, not caller-asserted.**
"Is this harness us-nexus / sovereign?" is derived from the **authenticated principal/origin**
(the OAuth2 identity), never from `context.harness`. `context.harness` survives only as a
self-reported label for the audit log — descriptive telemetry, **never a gate input**. A control
keyed on an attacker-suppliable value is not a control.
3. **Central relay — ships in v1, with the I2 ledger entry.**
The relay is required, not optional: claude.ai (Chat/Cowork/Design), Crush, Pi, and LLM Council
cannot run the use-case library in-process, and those are primary day-to-day surfaces. Deferring
the relay would ship a capability that doesn't work from the interfaces actually in use. Because
the relay is a (thin, no-standing-visibility, audit-only-retention) central node, its cross-harness
reach **must be entered in `infra/docs/security-baseline.md`** with Why-accepted / Revisit-if
**before it ships** (I2). That ledger entry is v1 work, not a follow-up.
4. **Audit-sink-down — classification-aware: confidential fails closed, internal/public degrades.**
The posture inherits from the effective classification (decision 1), so there is one coherent
sensitivity model rather than a separate availability policy:
- **Confidential + central audit sink unreachable → hard-refuse.** No buffer, no proceed.
Confidential work must be centrally auditable *at write time*; "buffer and reconcile later"
introduces a buffer-integrity question (can a write tamper with its own pending audit record?)
that must not exist for confidential data. The simplicity of "refuse" is itself the assurance
asset — trivially true, nothing to poke holes in.
- **Internal / public + central sink unreachable → degrade-and-warn** with a durable local buffer
+ ntfy alert + reconcile-on-recovery (the earlier Q4 design, now scoped to lower tiers). Keeps
capture available for your own homelab work during an observability outage; negligible risk
since the buffered record is still durable and the data isn't client-confidential.
- **Floor (all tiers):** if *nothing* — neither central sink nor (for internal/public) the local
buffer — can record the audit, capture **refuses**. No tier writes wholly un-audited.
- Rationale: matches assurance cost to data sensitivity, exactly as the I1/sovereignty model
does for placement. Presentable to a due-diligence client as "audit posture is
classification-aware: confidential fails closed, internal degrades gracefully" — which
demonstrates the judgment, not just a binary. Couples Q4 to Q1's classification machinery
(being built anyway) and removes the buffer-integrity rabbit hole for the only tier where it
mattered.