Compare commits
26
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
98cfae595c | ||
|
|
2a595b5a92 | ||
|
|
b7a2cc5fdf | ||
|
|
db638cca11 | ||
|
|
38579598e0 | ||
|
|
d6fa92b176 | ||
|
|
9173f9058d | ||
|
|
3e84a41fed | ||
|
|
0e0571c7da | ||
|
|
7a27cf71a2 | ||
|
|
63df6d3283 | ||
|
|
f04b03e07e | ||
|
|
6c61f93146 | ||
|
|
95a69fc2c1 | ||
|
|
bb8bc0478c | ||
|
|
a961a3c064 | ||
|
|
bec28f9014 | ||
|
|
b62ac57382 | ||
|
|
aa918388b9 | ||
|
|
e8dbcf6eef | ||
|
|
0eeb1df4a2 | ||
|
|
9febb1bba1 | ||
|
|
5dc247b994 | ||
|
|
2125558196 | ||
|
|
2beaac2feb | ||
|
|
525811bc1a |
@@ -1,315 +0,0 @@
|
||||
# Agent context — Mathias workspace
|
||||
|
||||
<!-- Canonical root context for all AI coding agents.
|
||||
Lives at: ~/dev/.context/AGENT.md
|
||||
Applies to every project under ~/dev/ unless overridden.
|
||||
|
||||
Run `task context:sync` from ~/dev/ to regenerate harness-specific files.
|
||||
Project-level context in .context/PROJECT.md layers on top of this. -->
|
||||
|
||||
## Who I am
|
||||
|
||||
I'm Mathias, a digital product manager and technology consultant based in Sweden.
|
||||
I build software, research emerging tech, and deliver consulting engagements
|
||||
for clients under NDA. I work across AI/ML, financial automation, web applications,
|
||||
and climate/sustainability tech.
|
||||
|
||||
## How I work with agents
|
||||
|
||||
- I think like a product manager — I care about *why* before *how*
|
||||
- I want agents to be opinionated and push back, not just execute blindly
|
||||
- I prefer concise responses; skip ceremony and get to the point
|
||||
- When I say "build this", I mean production-quality with tests, not a demo
|
||||
- Ask me before making irreversible changes or adding heavy dependencies
|
||||
- I work with confidential client data — never send it to cloud APIs unless I explicitly say it's OK
|
||||
|
||||
## Behavior rules
|
||||
|
||||
These rules apply to every task across every project, regardless of harness.
|
||||
|
||||
1. **No assumptions.** Don't hide confusion — surface it. Surface tradeoffs explicitly.
|
||||
Think before coding; if the problem is unclear, ask or state assumptions before acting.
|
||||
2. **Minimum viable code.** Solve with the smallest change that works. Nothing
|
||||
speculative, no "while we're here" cleanups, no premature abstractions. Simplicity first.
|
||||
3. **Surgical changes.** Touch only what the task requires. Leave unrelated code,
|
||||
files, and formatting alone. Diffs should be small and reviewable.
|
||||
4. **Goal-driven execution.** Define clear success criteria up front for every task.
|
||||
Loop — implement, verify, refine — until those criteria are met. Don't claim
|
||||
completion without evidence (tests pass, command output, observed behavior).
|
||||
5. **Trunk-Based Development — commit directly to main.** Every commit is one
|
||||
logical change (one tool, one fix, one test) with passing tests. Main is always
|
||||
deployable. Never create long-lived feature branches.
|
||||
|
||||
**Exception — parallel agents on same repo:** If another agent is known to be
|
||||
actively working on the same repo simultaneously, create a short-lived branch
|
||||
(`agent/<description>`), finish the task, and merge to main within the same
|
||||
session. Do not leave agent branches open between sessions.
|
||||
|
||||
**Exception — external contributor or client four-eyes requirement:** Use
|
||||
PR flow only when a human reviewer outside the project is required. Document
|
||||
the reason in PROJECT.md.
|
||||
|
||||
## Default stack
|
||||
|
||||
| Layer | Default | Fallback | Last resort |
|
||||
|-------|---------|----------|-------------|
|
||||
| Language | Go | Python | TypeScript, Java, C |
|
||||
| UI | HTMX + Templ | Server-rendered HTML | React (only if SPA is justified) |
|
||||
| Build | Task (taskfile.dev) | Make | — |
|
||||
| Containers | Docker Compose (dev), k3s (prod) | — | — |
|
||||
| DB | PostgreSQL + sqlc | SQLite | — |
|
||||
| Search | pgvector (vector), BM25 | Qdrant (when >1M vectors or hybrid retrieval) | — |
|
||||
| Logging | slog (structured) | — | — |
|
||||
| Testing | Table-driven, testify | — | — |
|
||||
| Agents (Go) | google.golang.org/adk + pkg/litellm adapter | — | — |
|
||||
|
||||
Exploratory: Rust, Zig — I'll tell you when I want these.
|
||||
|
||||
## Code conventions
|
||||
|
||||
- **Go style**: golines, gofumpt, golangci-lint
|
||||
- **Errors**: `fmt.Errorf("operation: %w", err)` — never naked, never log-and-return
|
||||
- **Naming**: stdlib conventions, no stuttering
|
||||
- **Architecture**: prefer stdlib over frameworks, constructor injection, env-var config parsed into typed structs
|
||||
- **Git**: conventional commits (`feat:`, `fix:`, `chore:`), commit directly to main,
|
||||
one logical change per commit, CI is the quality gate
|
||||
- **Never**: long-lived feature branches, PRs for solo work, direct push without
|
||||
passing `task check` locally first
|
||||
- **Security**: no secrets in code, govulncheck before adding deps, SOPS for encrypted config
|
||||
- **Dependencies**: prefer stdlib. testify, slog, templ, sqlc, google.golang.org/adk (agent projects only) are pre-approved; anything else needs justification in the commit message
|
||||
|
||||
## Infrastructure
|
||||
|
||||
Three machines on Tailscale:
|
||||
|
||||
| Machine | Role | Key specs |
|
||||
|---------|------|-----------|
|
||||
| koala | GPU inference, heavy compute | RTX 5070, runs k3s + llama-swap + shared postgres18/pgvector |
|
||||
| iguana | Services, builds | M2 Ultra Mac |
|
||||
| flamingo | Daily driver, edge | Mac mini, ~/dev is here |
|
||||
|
||||
- **Model routing**: LiteLLM in front of llama-swap (local) + cloud APIs (when permitted)
|
||||
- **Orchestration**: k3s cluster across all three machines
|
||||
- **Networking**: Tailscale mesh
|
||||
|
||||
## Project landscape
|
||||
|
||||
All development repos live at `~/dev/` (softlink from `~/Documents/local-dev/`).
|
||||
|
||||
Organized in thematic folders:
|
||||
|
||||
| Folder | Focus | Count |
|
||||
|--------|-------|-------|
|
||||
| `GO/` | Go web frameworks, API integrations, learning projects | ~10 |
|
||||
| `AI/` | ML research, AI frameworks (FinRL, DSPy, crawl4ai) | ~6 |
|
||||
| `AGENTS/` | Autonomous agents, coding agents, MCP servers, infra | ~15 |
|
||||
| `QKX/` | Invoice processing, financial automation, payment systems | ~13 |
|
||||
| `XT/` | Climate data, sustainability (Klimatkollen, Garbo) | ~2 |
|
||||
|
||||
See `~/dev/PROJECT_SUMMARY.md` for detailed descriptions of each project.
|
||||
|
||||
### Key active projects
|
||||
|
||||
- **super-koala** (`AGENTS/`) — multi-component agent stack with LangGraph, DSPy, MCP
|
||||
- **azure-tiger** (`QKX/`) — invoice extraction → ISO 20022 payment instructions
|
||||
- **gocrwl** (`AGENTS/`) — Go web crawler with containerized deployment
|
||||
- **koala-ai-stack** (`AGENTS/`) — local AI server infrastructure management
|
||||
- **klimatkollen** (`XT/`) — Swedish municipal climate data platform
|
||||
|
||||
## Knowledge base — actively use it
|
||||
|
||||
A persistent brain (BM25 search + LLM-synthesised Q&A) survives across sessions,
|
||||
hosts, and harnesses. It holds 100+ hard-won entries: infra incident postmortems,
|
||||
Go pitfalls, framework gotchas, design principles, ADRs. **It is not optional
|
||||
reference material — query it actively, not just when explicitly told.**
|
||||
|
||||
### When to query (treat as a reflex)
|
||||
|
||||
- **Before** starting a non-trivial task — search for prior art with the symptom
|
||||
AND the system component ("how did we solve X in Y?"). 5 seconds beats 5 hours.
|
||||
- **When debugging** — search for the error string, the stack frame, the affected
|
||||
service. Past you may have already paid this tax.
|
||||
- **Before adopting** a pattern, library, framework, or model name — check if it
|
||||
was tried and rejected, or what the integration footguns are.
|
||||
- **When making architectural decisions** — search for the domain + "ADR" or
|
||||
"decision" to find prior reasoning before re-deriving it.
|
||||
- **When a recommendation feels novel** — challenge yourself: "has this been
|
||||
documented?" The brain often has it.
|
||||
|
||||
### When to write
|
||||
|
||||
After you discover something that **future-you would forget** and that **isn't
|
||||
recoverable from the code, git log, or PR description alone**:
|
||||
|
||||
- Bugs whose root cause is non-obvious and generalisable beyond this project.
|
||||
- Framework / library / model-name quirks that bit you and would bite anyone.
|
||||
- Design principles validated under fire (e.g. "every `_get` needs a `_list`").
|
||||
- Postmortems for incidents: what broke, why, how diagnosed, what to do next time.
|
||||
|
||||
DON'T write project status, sprint progress, PR summaries, or "what I did this
|
||||
session" — those rot fast and the originals are in git/gitea anyway. Brain
|
||||
entries that age well are about *why*, *how to avoid*, and *what to do when*.
|
||||
|
||||
### How to access (per harness)
|
||||
|
||||
| Harness | Query | Write |
|
||||
|---------|-------|-------|
|
||||
| **Claude Code, Claude Desktop** | `brain_query` (BM25), `brain_answer` (LLM-synth + sources) MCP tools | `brain_write` MCP tool |
|
||||
| **Crush, Pi, Antigravity, other MCP-capable** | same MCP server: `ingestion-brain` (via the `mcp__*_brain__*` namespace once authenticated) | same |
|
||||
| **Anything HTTP-only (curl, scripts)** | `POST https://brain-mcp.d-ma.be/query` with `{"query":"..."}` (auth via `BRAIN_MCP_TOKEN`) | `POST .../write` with `{"content":"...","filename":"..."}` |
|
||||
| **Browser / human inspection** | `https://gitea.d-ma.be/mathias/hyperguild` → `knowledge/` and `wiki/` markdown files |
|
||||
|
||||
- **Scoping**: defaults to `public` collection; client projects filter to `{client}` + `public`.
|
||||
- **Routing**: brain_answer's LLM uses berget.ai as primary, iguana ollama as
|
||||
fallback. Both are configurable in the `supervisor/ingestion-deployment.yaml`
|
||||
on the koala k3s cluster; don't hardcode local-only model names into the
|
||||
berget URL (see knowledge entry on namespace mismatches).
|
||||
|
||||
### Quick reflex checks
|
||||
|
||||
If you find yourself about to say any of these out loud, you owe yourself a brain query first:
|
||||
|
||||
- "I think the issue might be..."
|
||||
- "Let me try X and see..."
|
||||
- "I'll just write a script to..."
|
||||
- "This is probably a new bug..."
|
||||
- "Has anyone done this before?" — *yes, probably, go check.*
|
||||
|
||||
## Client work rules
|
||||
|
||||
When working on a project tagged with a client name:
|
||||
1. Never send code, data, or context to cloud APIs — use local models only
|
||||
2. Never reference other client projects or their data
|
||||
3. Keep all artifacts within the client's git org / directory
|
||||
4. Treat everything as confidential unless told otherwise
|
||||
|
||||
## Harness-agnostic principles
|
||||
|
||||
This context is designed to work with any AI coding tool:
|
||||
- Claude Code, Cursor, Aider, Open WebUI, Charmbracelet Mods/Crush
|
||||
- Pi Coding Agent, Mistral Vibe, Antigravity
|
||||
- Any tool that accepts a system prompt or reads a markdown context file
|
||||
|
||||
The canonical source is always `.context/AGENT.md` (root) and `.context/PROJECT.md` (per-project).
|
||||
Derived files are committed (see *How context propagates* below) so a `git pull` on any host yields full agent context with no setup.
|
||||
|
||||
## How context propagates
|
||||
|
||||
Canonical sources of truth:
|
||||
- Universal: `~/dev/.context/AGENT.md` (this file)
|
||||
- Project: `<repo>/.context/PROJECT.md` (per-repo)
|
||||
|
||||
Derived files (committed, regenerated by `task context:sync`):
|
||||
- `CLAUDE.md`, `AGENTS.md`, `.cursorrules`, `.aider.conventions.md`,
|
||||
`.context/system-prompt.txt`
|
||||
|
||||
Workflow:
|
||||
1. Edit a canonical file. Run `task context:sync`. Commit canonical and
|
||||
derived together. Push.
|
||||
2. On any other host, `git pull` brings both. Claude Code (tree-walking)
|
||||
uses `CLAUDE.md`; Crush / Pi / Antigravity (cwd-only) use `AGENTS.md`;
|
||||
Cursor uses `.cursorrules`; Aider uses `.aider.conventions.md`.
|
||||
3. `task check` runs `context:sync` then asserts `git status --porcelain`
|
||||
is empty over the derived files (catches both modified-tracked drift
|
||||
and missing-untracked adapters). A drift fails the check with a
|
||||
message telling you to stage the regenerated files.
|
||||
|
||||
Behavior rules in this file and per-project rules in `PROJECT.md` apply
|
||||
unconditionally on every host, every harness.
|
||||
|
||||
## Engineering Skills
|
||||
|
||||
Shared engineering skills are available in `~/dev/.skills/`. Load on demand via the index.
|
||||
|
||||
See `~/dev/.skills/SKILLS_INDEX.md` for the full list with descriptions and "use when" triggers.
|
||||
|
||||
Key skills:
|
||||
- **TDD**: always write tests first — load `tdd` skill
|
||||
- **Code Review**: load `code-review` skill before any review
|
||||
- **SOLID/Clean Code**: load `solid` or `clean-code` skill for design work
|
||||
- **Problem first**: load `problem-analysis` skill before coding non-trivial features
|
||||
|
||||
---
|
||||
|
||||
# Project context
|
||||
|
||||
<!-- Canonical project context. Edit this, run `task context:sync`.
|
||||
Root agent context from ~/dev/.context/AGENT.md is automatically
|
||||
prepended for harnesses that don't walk the directory tree. -->
|
||||
|
||||
## Identity
|
||||
|
||||
- **Name**: supervisor
|
||||
- **Owner**: Mathias
|
||||
- **Client**: personal
|
||||
- **Repo**:
|
||||
- **Status**: active
|
||||
|
||||
## Stack
|
||||
|
||||
- **Primary language**: Go
|
||||
- **UI layer**: HTMX + Templ (when applicable)
|
||||
- **Fallback languages**: Python, TypeScript (justify in PR if used)
|
||||
- **Build**: Task (taskfile.dev), not Make
|
||||
- **Containers**: Docker (compose for dev, k3s for deploy)
|
||||
- **Target infra**: koala (GPU workloads), iguana (services), flamingo (edge)
|
||||
|
||||
## Conventions
|
||||
|
||||
### Code style
|
||||
- Go: follow `golines`, `gofumpt`, `golangci-lint` with project config
|
||||
- Tests: table-driven, in `_test.go` next to source, `testify` for assertions
|
||||
- Errors: wrap with `fmt.Errorf("operation: %w", err)`, no naked returns
|
||||
- Naming: stdlib conventions, no stuttering (`http.Client` not `http.HTTPClient`)
|
||||
|
||||
### Architecture preferences
|
||||
- Prefer standard library over frameworks (net/http over gin/echo)
|
||||
- Dependency injection via constructor functions, not containers
|
||||
- Configuration via environment variables, parsed at startup into a typed struct
|
||||
- Structured logging via `slog`
|
||||
|
||||
### Git
|
||||
- Conventional commits: `feat:`, `fix:`, `chore:`, `docs:`, `refactor:`
|
||||
- Branch naming: `feat/short-description`, `fix/short-description`
|
||||
- PRs: one concern per PR, description explains *why* not *what*
|
||||
|
||||
### Security
|
||||
- No secrets in code, ever — use env vars or SOPS-encrypted files
|
||||
- Client data never leaves local network unless explicitly cleared
|
||||
- Dependencies: audit with `govulncheck` before adding
|
||||
|
||||
## MCP endpoints
|
||||
|
||||
Two MCP servers are live, both reachable over Tailscale and via HTTPS domain:
|
||||
|
||||
- **`brain`** at `https://brain-mcp.d-ma.be/mcp` (NodePort `koala:30330`) —
|
||||
`brain_query`, `brain_write`, `brain_ingest`, `brain_ingest_raw`,
|
||||
`brain_answer`, `brain_classify`, `session_log`. Hosted by the ingestion
|
||||
service. Auth: Dex JWT (claude.ai OAuth) or static `BRAIN_MCP_TOKEN`.
|
||||
- **`routing`** at `http://koala:30310/mcp` — Mode 2 routing pod. Advertises
|
||||
`review`, `debug`, `retrospective`, `trainer`; per-call routes to local model
|
||||
or Claude based on brain `/pass-rate`. Bearer auth via `ROUTING_MCP_TOKEN`
|
||||
(opt-in). Only `mode client-local` registers this endpoint.
|
||||
|
||||
The supervisor MCP (`koala:30320`) was retired in Plan 7 (2026-05-12). Its
|
||||
skill workers (`tdd`, `spec`) are now SKILL.md files; routed skills moved to
|
||||
the routing pod; brain tools moved to the brain MCP.
|
||||
|
||||
The brain HTTP REST API (`/query`, `/write`, `/ingest`, `/ingest-raw`,
|
||||
`/ingest-path`, `/backfill-refs`, `/pass-rate`) remains available on port 3300
|
||||
for shell scripts and non-MCP clients.
|
||||
|
||||
`brain_answer(query)` performs BM25 retrieval + LLM synthesis (berget.ai
|
||||
gemma4:31b → iguana fallback). `brain_classify(text)` infers doc type, title,
|
||||
and tags. Both require `BRAIN_LLM_PRIMARY_URL` to be set in the ingestion pod.
|
||||
|
||||
## Agent instructions
|
||||
|
||||
When acting as a coding agent on this project:
|
||||
|
||||
1. Read this file and all `SKILL.md` files in `.skills/` before starting work
|
||||
2. Run `task check` before committing (lint + test + vet)
|
||||
3. If unsure about a convention, check `DECISIONS.md` or ask
|
||||
4. Never modify files outside the project root without explicit permission
|
||||
5. When adding a dependency, explain why in the commit message
|
||||
6. For client projects: never send code or context to cloud APIs — use local models via LiteLLM
|
||||
@@ -32,6 +32,14 @@ and climate/sustainability tech.
|
||||
|
||||
These rules apply to every task across every project, regardless of harness.
|
||||
|
||||
0. **Pre-task ritual — before ANY implementation (non-negotiable).** Run this before writing a single line:
|
||||
- **Query the brain** (`brain_query`) for the domain + symptom. If the result changes your approach, surface it before acting. 5 seconds beats 5 hours.
|
||||
- **Load the relevant skill** — see trigger table in *Engineering Skills* below.
|
||||
- **Write the failing test first.** Name the test before the function. If the target is untestable (e.g. `main()` wiring), extract the logic into a testable function first. No implementation without a red test.
|
||||
- **State the observable success criterion** — what specific behavior, output, or passing test proves this is done?
|
||||
|
||||
**TDD is non-negotiable.** "Tests pass" is not proof of correctness — only proof the tests ran. Write tests that would catch the bug before writing code that fixes it.
|
||||
|
||||
1. **No assumptions.** Don't hide confusion — surface it. Surface tradeoffs explicitly.
|
||||
Think before coding; if the problem is unclear, ask or state assumptions before acting.
|
||||
2. **Minimum viable code.** Solve with the smallest change that works. Nothing
|
||||
@@ -54,6 +62,22 @@ These rules apply to every task across every project, regardless of harness.
|
||||
PR flow only when a human reviewer outside the project is required. Document
|
||||
the reason in PROJECT.md.
|
||||
|
||||
6. **Close the loop — every substantive task ends with the same ritual.** Shipping
|
||||
the code is not the end of the task; capturing it is. Run this unprompted:
|
||||
- **Tag + bump SemVer** on the change (annotated tag; minor for a feature or
|
||||
new/changed ADR, patch for a fix; docs in the same commit). Check the repo's
|
||||
actual last tag — stated versions in docs drift stale.
|
||||
- **Push** main and the tag (CI is the gate).
|
||||
- **Persist generalizable learnings to the brain** (`brain_write`, wing/hall) —
|
||||
the reusable patterns and the footguns that would bite anyone again, never
|
||||
project status. See *Knowledge base — when to write* below.
|
||||
- **File discovered-but-deferred work as tracker issues** on the project's own
|
||||
repo — token-budget gaps, recorded ADR limitations, v2 follow-ups. Don't let
|
||||
"out of scope, recorded" rot in a commit message; make it a ticket with a
|
||||
source pointer.
|
||||
- Surface the brain entries and issue numbers in the closing summary so the
|
||||
trail is auditable.
|
||||
|
||||
## Default stack
|
||||
|
||||
| Layer | Default | Fallback | Last resort |
|
||||
@@ -83,6 +107,26 @@ Exploratory: Rust, Zig — I'll tell you when I want these.
|
||||
- **Security**: no secrets in code, govulncheck before adding deps, SOPS for encrypted config
|
||||
- **Dependencies**: prefer stdlib. testify, slog, templ, sqlc, google.golang.org/adk (agent projects only) are pre-approved; anything else needs justification in the commit message
|
||||
|
||||
## Secret handling (every harness, every command)
|
||||
|
||||
Tool output is persisted: terminal → `~/.claude/projects` transcripts →
|
||||
claudewatcher → brain/wiki → gitea history. A secret printed once is
|
||||
searchable forever, and clearing it means rotating the key. So:
|
||||
|
||||
1. **Never print, echo, log, or transform a secret to inspect it.** No
|
||||
`base64`/`xxd`/`cat` of a key, and never pipe a secret through a transform
|
||||
to defeat `op run`'s output masking (it masks raw values; base64 hides them
|
||||
from the mask — that exact trick leaked a key on 2026-06-11).
|
||||
2. **Secrets stay in the subprocess.** Reference them only as env vars consumed
|
||||
*inside* `op run --env-file ~/.op-env -- <cmd>`. Never place a literal secret
|
||||
in a command's argv (it lands in the tool call and the transcript).
|
||||
3. **Existence check without revealing the value:** `[ -n "$X" ] && echo set` —
|
||||
never `${X:-...}` (returns the value when set) and never echo a substring of it.
|
||||
4. **Cross-host secrets:** run the secret-consuming command on the host that has
|
||||
the secret; do not forward a raw key over ssh argv/stdout.
|
||||
5. If a secret does leak into output, say so immediately and flag it for rotation —
|
||||
don't bury it.
|
||||
|
||||
## Infrastructure
|
||||
|
||||
Three machines on Tailscale:
|
||||
@@ -162,7 +206,7 @@ entries that age well are about *why*, *how to avoid*, and *what to do when*.
|
||||
| **Claude Code, Claude Desktop** | `brain_query` (BM25), `brain_answer` (LLM-synth + sources) MCP tools | `brain_write` MCP tool |
|
||||
| **Crush, Pi, Antigravity, other MCP-capable** | same MCP server: `ingestion-brain` (via the `mcp__*_brain__*` namespace once authenticated) | same |
|
||||
| **Anything HTTP-only (curl, scripts)** | `POST https://brain-mcp.d-ma.be/query` with `{"query":"..."}` (auth via `BRAIN_MCP_TOKEN`) | `POST .../write` with `{"content":"...","filename":"..."}` |
|
||||
| **Browser / human inspection** | `https://gitea.d-ma.be/mathias/hyperguild` → `knowledge/` and `wiki/` markdown files |
|
||||
| **Browser / human inspection** | `https://git.d-ma.be/mathias/hyperguild` → `knowledge/` and `wiki/` markdown files |
|
||||
|
||||
- **Scoping**: defaults to `public` collection; client projects filter to `{client}` + `public`.
|
||||
- **Routing**: brain_answer's LLM uses berget.ai as primary, iguana ollama as
|
||||
@@ -224,15 +268,17 @@ unconditionally on every host, every harness.
|
||||
|
||||
## Engineering Skills
|
||||
|
||||
Shared engineering skills are available in `~/dev/.skills/`. Load on demand via the index.
|
||||
Shared engineering skills are available in `~/dev/.skills/`. Load at task start — not "on demand" but on schedule, before writing code. See `~/dev/.skills/SKILLS_INDEX.md` for the full list.
|
||||
|
||||
See `~/dev/.skills/SKILLS_INDEX.md` for the full list with descriptions and "use when" triggers.
|
||||
**Skill trigger table — load before starting, not after getting stuck:**
|
||||
|
||||
Key skills:
|
||||
- **TDD**: always write tests first — load `tdd` skill
|
||||
- **Code Review**: load `code-review` skill before any review
|
||||
- **SOLID/Clean Code**: load `solid` or `clean-code` skill for design work
|
||||
- **Problem first**: load `problem-analysis` skill before coding non-trivial features
|
||||
| Task type | Load |
|
||||
|-----------|------|
|
||||
| Any feature or bug fix | `tdd` |
|
||||
| Refactor or design | `clean-code` or `solid` |
|
||||
| Debug | `problem-analysis` |
|
||||
| Review code or PRs | `code-review` |
|
||||
| Frame a problem before coding | `problem-analysis` |
|
||||
|
||||
---
|
||||
|
||||
|
||||
-318
@@ -1,318 +0,0 @@
|
||||
# Cursor rules — auto-generated
|
||||
# Do not edit. Run: task context:sync
|
||||
|
||||
# Agent context — Mathias workspace
|
||||
|
||||
<!-- Canonical root context for all AI coding agents.
|
||||
Lives at: ~/dev/.context/AGENT.md
|
||||
Applies to every project under ~/dev/ unless overridden.
|
||||
|
||||
Run `task context:sync` from ~/dev/ to regenerate harness-specific files.
|
||||
Project-level context in .context/PROJECT.md layers on top of this. -->
|
||||
|
||||
## Who I am
|
||||
|
||||
I'm Mathias, a digital product manager and technology consultant based in Sweden.
|
||||
I build software, research emerging tech, and deliver consulting engagements
|
||||
for clients under NDA. I work across AI/ML, financial automation, web applications,
|
||||
and climate/sustainability tech.
|
||||
|
||||
## How I work with agents
|
||||
|
||||
- I think like a product manager — I care about *why* before *how*
|
||||
- I want agents to be opinionated and push back, not just execute blindly
|
||||
- I prefer concise responses; skip ceremony and get to the point
|
||||
- When I say "build this", I mean production-quality with tests, not a demo
|
||||
- Ask me before making irreversible changes or adding heavy dependencies
|
||||
- I work with confidential client data — never send it to cloud APIs unless I explicitly say it's OK
|
||||
|
||||
## Behavior rules
|
||||
|
||||
These rules apply to every task across every project, regardless of harness.
|
||||
|
||||
1. **No assumptions.** Don't hide confusion — surface it. Surface tradeoffs explicitly.
|
||||
Think before coding; if the problem is unclear, ask or state assumptions before acting.
|
||||
2. **Minimum viable code.** Solve with the smallest change that works. Nothing
|
||||
speculative, no "while we're here" cleanups, no premature abstractions. Simplicity first.
|
||||
3. **Surgical changes.** Touch only what the task requires. Leave unrelated code,
|
||||
files, and formatting alone. Diffs should be small and reviewable.
|
||||
4. **Goal-driven execution.** Define clear success criteria up front for every task.
|
||||
Loop — implement, verify, refine — until those criteria are met. Don't claim
|
||||
completion without evidence (tests pass, command output, observed behavior).
|
||||
5. **Trunk-Based Development — commit directly to main.** Every commit is one
|
||||
logical change (one tool, one fix, one test) with passing tests. Main is always
|
||||
deployable. Never create long-lived feature branches.
|
||||
|
||||
**Exception — parallel agents on same repo:** If another agent is known to be
|
||||
actively working on the same repo simultaneously, create a short-lived branch
|
||||
(`agent/<description>`), finish the task, and merge to main within the same
|
||||
session. Do not leave agent branches open between sessions.
|
||||
|
||||
**Exception — external contributor or client four-eyes requirement:** Use
|
||||
PR flow only when a human reviewer outside the project is required. Document
|
||||
the reason in PROJECT.md.
|
||||
|
||||
## Default stack
|
||||
|
||||
| Layer | Default | Fallback | Last resort |
|
||||
|-------|---------|----------|-------------|
|
||||
| Language | Go | Python | TypeScript, Java, C |
|
||||
| UI | HTMX + Templ | Server-rendered HTML | React (only if SPA is justified) |
|
||||
| Build | Task (taskfile.dev) | Make | — |
|
||||
| Containers | Docker Compose (dev), k3s (prod) | — | — |
|
||||
| DB | PostgreSQL + sqlc | SQLite | — |
|
||||
| Search | pgvector (vector), BM25 | Qdrant (when >1M vectors or hybrid retrieval) | — |
|
||||
| Logging | slog (structured) | — | — |
|
||||
| Testing | Table-driven, testify | — | — |
|
||||
| Agents (Go) | google.golang.org/adk + pkg/litellm adapter | — | — |
|
||||
|
||||
Exploratory: Rust, Zig — I'll tell you when I want these.
|
||||
|
||||
## Code conventions
|
||||
|
||||
- **Go style**: golines, gofumpt, golangci-lint
|
||||
- **Errors**: `fmt.Errorf("operation: %w", err)` — never naked, never log-and-return
|
||||
- **Naming**: stdlib conventions, no stuttering
|
||||
- **Architecture**: prefer stdlib over frameworks, constructor injection, env-var config parsed into typed structs
|
||||
- **Git**: conventional commits (`feat:`, `fix:`, `chore:`), commit directly to main,
|
||||
one logical change per commit, CI is the quality gate
|
||||
- **Never**: long-lived feature branches, PRs for solo work, direct push without
|
||||
passing `task check` locally first
|
||||
- **Security**: no secrets in code, govulncheck before adding deps, SOPS for encrypted config
|
||||
- **Dependencies**: prefer stdlib. testify, slog, templ, sqlc, google.golang.org/adk (agent projects only) are pre-approved; anything else needs justification in the commit message
|
||||
|
||||
## Infrastructure
|
||||
|
||||
Three machines on Tailscale:
|
||||
|
||||
| Machine | Role | Key specs |
|
||||
|---------|------|-----------|
|
||||
| koala | GPU inference, heavy compute | RTX 5070, runs k3s + llama-swap + shared postgres18/pgvector |
|
||||
| iguana | Services, builds | M2 Ultra Mac |
|
||||
| flamingo | Daily driver, edge | Mac mini, ~/dev is here |
|
||||
|
||||
- **Model routing**: LiteLLM in front of llama-swap (local) + cloud APIs (when permitted)
|
||||
- **Orchestration**: k3s cluster across all three machines
|
||||
- **Networking**: Tailscale mesh
|
||||
|
||||
## Project landscape
|
||||
|
||||
All development repos live at `~/dev/` (softlink from `~/Documents/local-dev/`).
|
||||
|
||||
Organized in thematic folders:
|
||||
|
||||
| Folder | Focus | Count |
|
||||
|--------|-------|-------|
|
||||
| `GO/` | Go web frameworks, API integrations, learning projects | ~10 |
|
||||
| `AI/` | ML research, AI frameworks (FinRL, DSPy, crawl4ai) | ~6 |
|
||||
| `AGENTS/` | Autonomous agents, coding agents, MCP servers, infra | ~15 |
|
||||
| `QKX/` | Invoice processing, financial automation, payment systems | ~13 |
|
||||
| `XT/` | Climate data, sustainability (Klimatkollen, Garbo) | ~2 |
|
||||
|
||||
See `~/dev/PROJECT_SUMMARY.md` for detailed descriptions of each project.
|
||||
|
||||
### Key active projects
|
||||
|
||||
- **super-koala** (`AGENTS/`) — multi-component agent stack with LangGraph, DSPy, MCP
|
||||
- **azure-tiger** (`QKX/`) — invoice extraction → ISO 20022 payment instructions
|
||||
- **gocrwl** (`AGENTS/`) — Go web crawler with containerized deployment
|
||||
- **koala-ai-stack** (`AGENTS/`) — local AI server infrastructure management
|
||||
- **klimatkollen** (`XT/`) — Swedish municipal climate data platform
|
||||
|
||||
## Knowledge base — actively use it
|
||||
|
||||
A persistent brain (BM25 search + LLM-synthesised Q&A) survives across sessions,
|
||||
hosts, and harnesses. It holds 100+ hard-won entries: infra incident postmortems,
|
||||
Go pitfalls, framework gotchas, design principles, ADRs. **It is not optional
|
||||
reference material — query it actively, not just when explicitly told.**
|
||||
|
||||
### When to query (treat as a reflex)
|
||||
|
||||
- **Before** starting a non-trivial task — search for prior art with the symptom
|
||||
AND the system component ("how did we solve X in Y?"). 5 seconds beats 5 hours.
|
||||
- **When debugging** — search for the error string, the stack frame, the affected
|
||||
service. Past you may have already paid this tax.
|
||||
- **Before adopting** a pattern, library, framework, or model name — check if it
|
||||
was tried and rejected, or what the integration footguns are.
|
||||
- **When making architectural decisions** — search for the domain + "ADR" or
|
||||
"decision" to find prior reasoning before re-deriving it.
|
||||
- **When a recommendation feels novel** — challenge yourself: "has this been
|
||||
documented?" The brain often has it.
|
||||
|
||||
### When to write
|
||||
|
||||
After you discover something that **future-you would forget** and that **isn't
|
||||
recoverable from the code, git log, or PR description alone**:
|
||||
|
||||
- Bugs whose root cause is non-obvious and generalisable beyond this project.
|
||||
- Framework / library / model-name quirks that bit you and would bite anyone.
|
||||
- Design principles validated under fire (e.g. "every `_get` needs a `_list`").
|
||||
- Postmortems for incidents: what broke, why, how diagnosed, what to do next time.
|
||||
|
||||
DON'T write project status, sprint progress, PR summaries, or "what I did this
|
||||
session" — those rot fast and the originals are in git/gitea anyway. Brain
|
||||
entries that age well are about *why*, *how to avoid*, and *what to do when*.
|
||||
|
||||
### How to access (per harness)
|
||||
|
||||
| Harness | Query | Write |
|
||||
|---------|-------|-------|
|
||||
| **Claude Code, Claude Desktop** | `brain_query` (BM25), `brain_answer` (LLM-synth + sources) MCP tools | `brain_write` MCP tool |
|
||||
| **Crush, Pi, Antigravity, other MCP-capable** | same MCP server: `ingestion-brain` (via the `mcp__*_brain__*` namespace once authenticated) | same |
|
||||
| **Anything HTTP-only (curl, scripts)** | `POST https://brain-mcp.d-ma.be/query` with `{"query":"..."}` (auth via `BRAIN_MCP_TOKEN`) | `POST .../write` with `{"content":"...","filename":"..."}` |
|
||||
| **Browser / human inspection** | `https://gitea.d-ma.be/mathias/hyperguild` → `knowledge/` and `wiki/` markdown files |
|
||||
|
||||
- **Scoping**: defaults to `public` collection; client projects filter to `{client}` + `public`.
|
||||
- **Routing**: brain_answer's LLM uses berget.ai as primary, iguana ollama as
|
||||
fallback. Both are configurable in the `supervisor/ingestion-deployment.yaml`
|
||||
on the koala k3s cluster; don't hardcode local-only model names into the
|
||||
berget URL (see knowledge entry on namespace mismatches).
|
||||
|
||||
### Quick reflex checks
|
||||
|
||||
If you find yourself about to say any of these out loud, you owe yourself a brain query first:
|
||||
|
||||
- "I think the issue might be..."
|
||||
- "Let me try X and see..."
|
||||
- "I'll just write a script to..."
|
||||
- "This is probably a new bug..."
|
||||
- "Has anyone done this before?" — *yes, probably, go check.*
|
||||
|
||||
## Client work rules
|
||||
|
||||
When working on a project tagged with a client name:
|
||||
1. Never send code, data, or context to cloud APIs — use local models only
|
||||
2. Never reference other client projects or their data
|
||||
3. Keep all artifacts within the client's git org / directory
|
||||
4. Treat everything as confidential unless told otherwise
|
||||
|
||||
## Harness-agnostic principles
|
||||
|
||||
This context is designed to work with any AI coding tool:
|
||||
- Claude Code, Cursor, Aider, Open WebUI, Charmbracelet Mods/Crush
|
||||
- Pi Coding Agent, Mistral Vibe, Antigravity
|
||||
- Any tool that accepts a system prompt or reads a markdown context file
|
||||
|
||||
The canonical source is always `.context/AGENT.md` (root) and `.context/PROJECT.md` (per-project).
|
||||
Derived files are committed (see *How context propagates* below) so a `git pull` on any host yields full agent context with no setup.
|
||||
|
||||
## How context propagates
|
||||
|
||||
Canonical sources of truth:
|
||||
- Universal: `~/dev/.context/AGENT.md` (this file)
|
||||
- Project: `<repo>/.context/PROJECT.md` (per-repo)
|
||||
|
||||
Derived files (committed, regenerated by `task context:sync`):
|
||||
- `CLAUDE.md`, `AGENTS.md`, `.cursorrules`, `.aider.conventions.md`,
|
||||
`.context/system-prompt.txt`
|
||||
|
||||
Workflow:
|
||||
1. Edit a canonical file. Run `task context:sync`. Commit canonical and
|
||||
derived together. Push.
|
||||
2. On any other host, `git pull` brings both. Claude Code (tree-walking)
|
||||
uses `CLAUDE.md`; Crush / Pi / Antigravity (cwd-only) use `AGENTS.md`;
|
||||
Cursor uses `.cursorrules`; Aider uses `.aider.conventions.md`.
|
||||
3. `task check` runs `context:sync` then asserts `git status --porcelain`
|
||||
is empty over the derived files (catches both modified-tracked drift
|
||||
and missing-untracked adapters). A drift fails the check with a
|
||||
message telling you to stage the regenerated files.
|
||||
|
||||
Behavior rules in this file and per-project rules in `PROJECT.md` apply
|
||||
unconditionally on every host, every harness.
|
||||
|
||||
## Engineering Skills
|
||||
|
||||
Shared engineering skills are available in `~/dev/.skills/`. Load on demand via the index.
|
||||
|
||||
See `~/dev/.skills/SKILLS_INDEX.md` for the full list with descriptions and "use when" triggers.
|
||||
|
||||
Key skills:
|
||||
- **TDD**: always write tests first — load `tdd` skill
|
||||
- **Code Review**: load `code-review` skill before any review
|
||||
- **SOLID/Clean Code**: load `solid` or `clean-code` skill for design work
|
||||
- **Problem first**: load `problem-analysis` skill before coding non-trivial features
|
||||
|
||||
---
|
||||
|
||||
# Project context
|
||||
|
||||
<!-- Canonical project context. Edit this, run `task context:sync`.
|
||||
Root agent context from ~/dev/.context/AGENT.md is automatically
|
||||
prepended for harnesses that don't walk the directory tree. -->
|
||||
|
||||
## Identity
|
||||
|
||||
- **Name**: supervisor
|
||||
- **Owner**: Mathias
|
||||
- **Client**: personal
|
||||
- **Repo**:
|
||||
- **Status**: active
|
||||
|
||||
## Stack
|
||||
|
||||
- **Primary language**: Go
|
||||
- **UI layer**: HTMX + Templ (when applicable)
|
||||
- **Fallback languages**: Python, TypeScript (justify in PR if used)
|
||||
- **Build**: Task (taskfile.dev), not Make
|
||||
- **Containers**: Docker (compose for dev, k3s for deploy)
|
||||
- **Target infra**: koala (GPU workloads), iguana (services), flamingo (edge)
|
||||
|
||||
## Conventions
|
||||
|
||||
### Code style
|
||||
- Go: follow `golines`, `gofumpt`, `golangci-lint` with project config
|
||||
- Tests: table-driven, in `_test.go` next to source, `testify` for assertions
|
||||
- Errors: wrap with `fmt.Errorf("operation: %w", err)`, no naked returns
|
||||
- Naming: stdlib conventions, no stuttering (`http.Client` not `http.HTTPClient`)
|
||||
|
||||
### Architecture preferences
|
||||
- Prefer standard library over frameworks (net/http over gin/echo)
|
||||
- Dependency injection via constructor functions, not containers
|
||||
- Configuration via environment variables, parsed at startup into a typed struct
|
||||
- Structured logging via `slog`
|
||||
|
||||
### Git
|
||||
- Conventional commits: `feat:`, `fix:`, `chore:`, `docs:`, `refactor:`
|
||||
- Branch naming: `feat/short-description`, `fix/short-description`
|
||||
- PRs: one concern per PR, description explains *why* not *what*
|
||||
|
||||
### Security
|
||||
- No secrets in code, ever — use env vars or SOPS-encrypted files
|
||||
- Client data never leaves local network unless explicitly cleared
|
||||
- Dependencies: audit with `govulncheck` before adding
|
||||
|
||||
## MCP endpoints
|
||||
|
||||
Two MCP servers are live, both reachable over Tailscale and via HTTPS domain:
|
||||
|
||||
- **`brain`** at `https://brain-mcp.d-ma.be/mcp` (NodePort `koala:30330`) —
|
||||
`brain_query`, `brain_write`, `brain_ingest`, `brain_ingest_raw`,
|
||||
`brain_answer`, `brain_classify`, `session_log`. Hosted by the ingestion
|
||||
service. Auth: Dex JWT (claude.ai OAuth) or static `BRAIN_MCP_TOKEN`.
|
||||
- **`routing`** at `http://koala:30310/mcp` — Mode 2 routing pod. Advertises
|
||||
`review`, `debug`, `retrospective`, `trainer`; per-call routes to local model
|
||||
or Claude based on brain `/pass-rate`. Bearer auth via `ROUTING_MCP_TOKEN`
|
||||
(opt-in). Only `mode client-local` registers this endpoint.
|
||||
|
||||
The supervisor MCP (`koala:30320`) was retired in Plan 7 (2026-05-12). Its
|
||||
skill workers (`tdd`, `spec`) are now SKILL.md files; routed skills moved to
|
||||
the routing pod; brain tools moved to the brain MCP.
|
||||
|
||||
The brain HTTP REST API (`/query`, `/write`, `/ingest`, `/ingest-raw`,
|
||||
`/ingest-path`, `/backfill-refs`, `/pass-rate`) remains available on port 3300
|
||||
for shell scripts and non-MCP clients.
|
||||
|
||||
`brain_answer(query)` performs BM25 retrieval + LLM synthesis (berget.ai
|
||||
gemma4:31b → iguana fallback). `brain_classify(text)` infers doc type, title,
|
||||
and tags. Both require `BRAIN_LLM_PRIMARY_URL` to be set in the ingestion pod.
|
||||
|
||||
## Agent instructions
|
||||
|
||||
When acting as a coding agent on this project:
|
||||
|
||||
1. Read this file and all `SKILL.md` files in `.skills/` before starting work
|
||||
2. Run `task check` before committing (lint + test + vet)
|
||||
3. If unsure about a convention, check `DECISIONS.md` or ask
|
||||
4. Never modify files outside the project root without explicit permission
|
||||
5. When adding a dependency, explain why in the commit message
|
||||
6. For client projects: never send code or context to cloud APIs — use local models via LiteLLM
|
||||
@@ -1,6 +1,6 @@
|
||||
name: cd
|
||||
|
||||
on:
|
||||
"on":
|
||||
workflow_run:
|
||||
workflows: ["CI"]
|
||||
types: [completed]
|
||||
@@ -13,9 +13,9 @@ jobs:
|
||||
if: ${{ github.event.workflow_run.conclusion == 'success' && github.event.workflow_run.event == 'push' }}
|
||||
environment: staging
|
||||
env:
|
||||
INGESTION_IMAGE: gitea.d-ma.be/mathias/ingestion
|
||||
ROUTING_IMAGE: gitea.d-ma.be/mathias/routing
|
||||
INFRA_REPO: git@gitea.d-ma.be:mathias/infra.git
|
||||
INGESTION_IMAGE: git.d-ma.be/mathias/ingestion
|
||||
ROUTING_IMAGE: git.d-ma.be/mathias/routing
|
||||
INFRA_REPO: git@git.d-ma.be:mathias/infra.git
|
||||
BUILDKIT_HOST: unix:///run/buildkit/buildkitd.sock
|
||||
steps:
|
||||
- name: Checkout
|
||||
@@ -71,17 +71,17 @@ jobs:
|
||||
mkdir -p ~/.ssh
|
||||
echo "${{ secrets.INFRA_DEPLOY_KEY }}" > ~/.ssh/infra_deploy_key
|
||||
chmod 600 ~/.ssh/infra_deploy_key
|
||||
printf 'Host gitea.d-ma.be\n HostName 127.0.0.1\n Port 30022\n StrictHostKeyChecking no\n' >> ~/.ssh/config
|
||||
printf 'Host git.d-ma.be\n HostName 127.0.0.1\n Port 30022\n StrictHostKeyChecking no\n' >> ~/.ssh/config
|
||||
|
||||
GIT_SSH_COMMAND="ssh -i ~/.ssh/infra_deploy_key -o IdentitiesOnly=yes" \
|
||||
git clone "${INFRA_REPO}" /tmp/infra-update
|
||||
|
||||
cd /tmp/infra-update
|
||||
|
||||
sed -i "s|gitea.d-ma.be/mathias/ingestion:.*|gitea.d-ma.be/mathias/ingestion:${IMAGE_TAG}|" \
|
||||
sed -i "s|git.d-ma.be/mathias/ingestion:.*|git.d-ma.be/mathias/ingestion:${IMAGE_TAG}|" \
|
||||
"k3s/apps/supervisor/ingestion-deployment.yaml"
|
||||
|
||||
sed -i "s|gitea.d-ma.be/mathias/routing:.*|gitea.d-ma.be/mathias/routing:${IMAGE_TAG}|" \
|
||||
sed -i "s|git.d-ma.be/mathias/routing:.*|git.d-ma.be/mathias/routing:${IMAGE_TAG}|" \
|
||||
"k3s/apps/routing/deployment.yaml"
|
||||
|
||||
git config user.email "cd-bot@d-ma.be"
|
||||
@@ -103,7 +103,7 @@ jobs:
|
||||
|
||||
- name: Wait for Flux to apply new ingestion image
|
||||
run: |
|
||||
EXPECTED="gitea.d-ma.be/mathias/ingestion:${{ github.sha }}"
|
||||
EXPECTED="git.d-ma.be/mathias/ingestion:${{ github.sha }}"
|
||||
for i in $(seq 1 60); do
|
||||
CURRENT=$(kubectl get deploy ingestion -n supervisor \
|
||||
-o jsonpath='{.spec.template.spec.containers[0].image}' 2>/dev/null || echo "")
|
||||
@@ -135,7 +135,7 @@ jobs:
|
||||
|
||||
- name: Wait for Flux to apply new routing image
|
||||
run: |
|
||||
EXPECTED="gitea.d-ma.be/mathias/routing:${{ github.sha }}"
|
||||
EXPECTED="git.d-ma.be/mathias/routing:${{ github.sha }}"
|
||||
for i in $(seq 1 60); do
|
||||
CURRENT=$(kubectl get deploy routing -n routing \
|
||||
-o jsonpath='{.spec.template.spec.containers[0].image}' 2>/dev/null || echo "")
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
name: CI
|
||||
|
||||
on:
|
||||
"on":
|
||||
push:
|
||||
branches: [main]
|
||||
tags: ["v*"]
|
||||
|
||||
@@ -27,6 +27,14 @@ and climate/sustainability tech.
|
||||
|
||||
These rules apply to every task across every project, regardless of harness.
|
||||
|
||||
0. **Pre-task ritual — before ANY implementation (non-negotiable).** Run this before writing a single line:
|
||||
- **Query the brain** (`brain_query`) for the domain + symptom. If the result changes your approach, surface it before acting. 5 seconds beats 5 hours.
|
||||
- **Load the relevant skill** — see trigger table in *Engineering Skills* below.
|
||||
- **Write the failing test first.** Name the test before the function. If the target is untestable (e.g. `main()` wiring), extract the logic into a testable function first. No implementation without a red test.
|
||||
- **State the observable success criterion** — what specific behavior, output, or passing test proves this is done?
|
||||
|
||||
**TDD is non-negotiable.** "Tests pass" is not proof of correctness — only proof the tests ran. Write tests that would catch the bug before writing code that fixes it.
|
||||
|
||||
1. **No assumptions.** Don't hide confusion — surface it. Surface tradeoffs explicitly.
|
||||
Think before coding; if the problem is unclear, ask or state assumptions before acting.
|
||||
2. **Minimum viable code.** Solve with the smallest change that works. Nothing
|
||||
@@ -49,6 +57,22 @@ These rules apply to every task across every project, regardless of harness.
|
||||
PR flow only when a human reviewer outside the project is required. Document
|
||||
the reason in PROJECT.md.
|
||||
|
||||
6. **Close the loop — every substantive task ends with the same ritual.** Shipping
|
||||
the code is not the end of the task; capturing it is. Run this unprompted:
|
||||
- **Tag + bump SemVer** on the change (annotated tag; minor for a feature or
|
||||
new/changed ADR, patch for a fix; docs in the same commit). Check the repo's
|
||||
actual last tag — stated versions in docs drift stale.
|
||||
- **Push** main and the tag (CI is the gate).
|
||||
- **Persist generalizable learnings to the brain** (`brain_write`, wing/hall) —
|
||||
the reusable patterns and the footguns that would bite anyone again, never
|
||||
project status. See *Knowledge base — when to write* below.
|
||||
- **File discovered-but-deferred work as tracker issues** on the project's own
|
||||
repo — token-budget gaps, recorded ADR limitations, v2 follow-ups. Don't let
|
||||
"out of scope, recorded" rot in a commit message; make it a ticket with a
|
||||
source pointer.
|
||||
- Surface the brain entries and issue numbers in the closing summary so the
|
||||
trail is auditable.
|
||||
|
||||
## Default stack
|
||||
|
||||
| Layer | Default | Fallback | Last resort |
|
||||
@@ -78,6 +102,26 @@ Exploratory: Rust, Zig — I'll tell you when I want these.
|
||||
- **Security**: no secrets in code, govulncheck before adding deps, SOPS for encrypted config
|
||||
- **Dependencies**: prefer stdlib. testify, slog, templ, sqlc, google.golang.org/adk (agent projects only) are pre-approved; anything else needs justification in the commit message
|
||||
|
||||
## Secret handling (every harness, every command)
|
||||
|
||||
Tool output is persisted: terminal → `~/.claude/projects` transcripts →
|
||||
claudewatcher → brain/wiki → gitea history. A secret printed once is
|
||||
searchable forever, and clearing it means rotating the key. So:
|
||||
|
||||
1. **Never print, echo, log, or transform a secret to inspect it.** No
|
||||
`base64`/`xxd`/`cat` of a key, and never pipe a secret through a transform
|
||||
to defeat `op run`'s output masking (it masks raw values; base64 hides them
|
||||
from the mask — that exact trick leaked a key on 2026-06-11).
|
||||
2. **Secrets stay in the subprocess.** Reference them only as env vars consumed
|
||||
*inside* `op run --env-file ~/.op-env -- <cmd>`. Never place a literal secret
|
||||
in a command's argv (it lands in the tool call and the transcript).
|
||||
3. **Existence check without revealing the value:** `[ -n "$X" ] && echo set` —
|
||||
never `${X:-...}` (returns the value when set) and never echo a substring of it.
|
||||
4. **Cross-host secrets:** run the secret-consuming command on the host that has
|
||||
the secret; do not forward a raw key over ssh argv/stdout.
|
||||
5. If a secret does leak into output, say so immediately and flag it for rotation —
|
||||
don't bury it.
|
||||
|
||||
## Infrastructure
|
||||
|
||||
Three machines on Tailscale:
|
||||
@@ -157,7 +201,7 @@ entries that age well are about *why*, *how to avoid*, and *what to do when*.
|
||||
| **Claude Code, Claude Desktop** | `brain_query` (BM25), `brain_answer` (LLM-synth + sources) MCP tools | `brain_write` MCP tool |
|
||||
| **Crush, Pi, Antigravity, other MCP-capable** | same MCP server: `ingestion-brain` (via the `mcp__*_brain__*` namespace once authenticated) | same |
|
||||
| **Anything HTTP-only (curl, scripts)** | `POST https://brain-mcp.d-ma.be/query` with `{"query":"..."}` (auth via `BRAIN_MCP_TOKEN`) | `POST .../write` with `{"content":"...","filename":"..."}` |
|
||||
| **Browser / human inspection** | `https://gitea.d-ma.be/mathias/hyperguild` → `knowledge/` and `wiki/` markdown files |
|
||||
| **Browser / human inspection** | `https://git.d-ma.be/mathias/hyperguild` → `knowledge/` and `wiki/` markdown files |
|
||||
|
||||
- **Scoping**: defaults to `public` collection; client projects filter to `{client}` + `public`.
|
||||
- **Routing**: brain_answer's LLM uses berget.ai as primary, iguana ollama as
|
||||
@@ -219,15 +263,17 @@ unconditionally on every host, every harness.
|
||||
|
||||
## Engineering Skills
|
||||
|
||||
Shared engineering skills are available in `~/dev/.skills/`. Load on demand via the index.
|
||||
Shared engineering skills are available in `~/dev/.skills/`. Load at task start — not "on demand" but on schedule, before writing code. See `~/dev/.skills/SKILLS_INDEX.md` for the full list.
|
||||
|
||||
See `~/dev/.skills/SKILLS_INDEX.md` for the full list with descriptions and "use when" triggers.
|
||||
**Skill trigger table — load before starting, not after getting stuck:**
|
||||
|
||||
Key skills:
|
||||
- **TDD**: always write tests first — load `tdd` skill
|
||||
- **Code Review**: load `code-review` skill before any review
|
||||
- **SOLID/Clean Code**: load `solid` or `clean-code` skill for design work
|
||||
- **Problem first**: load `problem-analysis` skill before coding non-trivial features
|
||||
| Task type | Load |
|
||||
|-----------|------|
|
||||
| Any feature or bug fix | `tdd` |
|
||||
| Refactor or design | `clean-code` or `solid` |
|
||||
| Debug | `problem-analysis` |
|
||||
| Review code or PRs | `code-review` |
|
||||
| Frame a problem before coding | `problem-analysis` |
|
||||
|
||||
---
|
||||
|
||||
|
||||
@@ -4,6 +4,74 @@ Record *why* things are the way they are. Future-you will thank present-you.
|
||||
|
||||
---
|
||||
|
||||
## 2026-05-28 — three active harnesses: hyperguild, agentsquad, Crush (extends earlier boundary decision)
|
||||
|
||||
**Context:** After wiring Crush to LiteLLM in May 2026, there are now three active harnesses.
|
||||
The earlier boundary decision only covered hyperguild vs agentsquad. Crush's role was undefined.
|
||||
|
||||
**Decision:** Three harnesses, three distinct roles, shared skills layer.
|
||||
|
||||
| Harness | Engine | Primary use | Brain MCP? | Routing pod? | Skills? |
|
||||
|---------|--------|-------------|------------|--------------|---------|
|
||||
| **hyperguild** | Claude Code + MCP | Disciplined solo coding sessions, TDD/review/debug workflows | Yes | Yes | Yes (SKILL.md) |
|
||||
| **agentsquad** | OpenCode + LiteLLM | Multi-agent task execution, executor/reviewer pipelines | No | No (own routing) | Yes (SKILL.md) |
|
||||
| **Crush** | Charmbracelet TUI + LiteLLM | Interactive local coding, quick iterations on flamingo | No (not yet) | No (direct LiteLLM) | Yes (SKILL.md) |
|
||||
|
||||
**Crush specifics (as of 2026-05-28):**
|
||||
- Config: `~/.config/crush/crush.json` on flamingo (see brain: `homelab/facts/crush-litellm-wiring-2026-05`)
|
||||
- Connects directly to LiteLLM at `http://koala:4000/v1/` using `sk-local-123`
|
||||
- Auth type: `openai-compat` (not `openai`)
|
||||
- Does NOT go through the routing pod — model selection is manual in the Crush UI
|
||||
- Brain MCP not wired — Crush has no MCP client capability today; revisit if Crush adds MCP support
|
||||
|
||||
**Shared across all three:**
|
||||
- `mathias/skills` — any SKILL.md file works in all three harnesses
|
||||
- LiteLLM proxy on koala (`http://koala:4000/v1/`) — Crush and agentsquad both route through it; hyperguild does too for local model calls
|
||||
|
||||
**Consequences:** No consolidation needed. crush.json must be kept in sync when litellm_config.yaml model names change. The `crush.json` canonical location is `~/.config/crush/crush.json` on flamingo — not yet tracked in a dotfiles repo (track as tech debt).
|
||||
|
||||
---
|
||||
|
||||
## 2026-05-28 — "field benchmark" for local models = pass-rate at scale (supersedes GOTTH eval suite)
|
||||
|
||||
**Context:** The GOTTH eval suite (45 offline prompts across 5 categories) was replaced by
|
||||
a "field benchmark" in May 2026, but the replacement was never defined concretely.
|
||||
|
||||
**Decision:** The field benchmark is per-skill pass rate over real routing pod usage,
|
||||
collected automatically by `internal/routing/passrate.go` and exposed at:
|
||||
|
||||
```
|
||||
GET /pass-rate?skill=<name>&window=<duration>
|
||||
```
|
||||
|
||||
No separate eval suite. No synthetic prompts. The benchmark runs itself once the routing
|
||||
pod receives real traffic. Target: 30-day rolling window per skill, reviewed monthly.
|
||||
|
||||
**Bootstrap note:** With no session history, `passrate.go` returns `nil` and the router
|
||||
defaults to the thinking model for every call. The fast-model path activates only after
|
||||
real pass-rate data accumulates. Seed with real usage — do not pre-populate.
|
||||
|
||||
**Consequences:** Zero maintenance overhead for the benchmark. The tradeoff is that results
|
||||
are only meaningful after ~2 weeks of real usage, and skills that are rarely invoked will
|
||||
have statistically thin pass-rate data. Revisit if a skill has fewer than 20 calls in 30 days.
|
||||
|
||||
---
|
||||
|
||||
## 2026-05-28 — brain injection in skill handlers: review is done, others unverified
|
||||
|
||||
**Context:** The April 2026 scope reset listed "brain_query injection into skill handlers"
|
||||
as the top priority. As of 2026-05-28, `internal/skills/review/handlers.go` calls
|
||||
`brain.Query(ctx, ...)` before dispatching to the LLM — confirmed in code review.
|
||||
Status of debug, retrospective, and trainer handlers is unverified.
|
||||
|
||||
**Decision:** Treat review as the reference implementation. Verify debug, retrospective,
|
||||
trainer against the same pattern before shipping new skill work. Tracked in issue #32.
|
||||
|
||||
**Consequences:** The April concern may be stale for review. A one-pass audit of the other
|
||||
three skill handlers closes this fully.
|
||||
|
||||
---
|
||||
|
||||
## 2026-04-08 — AGENTS.md as cross-tool standard, not CLAUDE.md
|
||||
|
||||
**Context**: Multiple tools (Crush, Pi, Antigravity) read `AGENTS.md` natively. Claude Code reads `CLAUDE.md`. Building on `CLAUDE.md` as the primary format locks into one vendor.
|
||||
|
||||
@@ -5,14 +5,31 @@ Instead of letting Claude Code do whatever it wants, hyperguild enforces structu
|
||||
workflows (TDD red/green/refactor), logs every session, and accumulates learnings
|
||||
into a searchable brain.
|
||||
|
||||
## Hypothesis
|
||||
|
||||
> We believe routing skill tasks through local models, backed by brain context,
|
||||
> produces measurably better outcomes than raw Claude Code alone —
|
||||
> measurable by per-skill pass rate over rolling 30-day windows
|
||||
> (available at `GET /pass-rate?skill=<name>&window=30d` on the brain pod).
|
||||
|
||||
This is the falsifiable claim the routing pod and pass-rate infrastructure exist to test.
|
||||
If per-skill pass rates don't improve over baseline (all-cloud) after 30 days of real
|
||||
usage, the fast-model routing path should be reconsidered.
|
||||
|
||||
## Harness
|
||||
|
||||
**hyperguild = Claude Code + MCP.** This is a supervisor for Claude Code sessions specifically.
|
||||
For multi-agent orchestration (OpenCode + LiteLLM, executor/reviewer pipelines), see
|
||||
[agentsquad](http://gitea.d-ma.be/mathias/agentsquad) — a separate harness for a different
|
||||
orchestration model. Skills (mathias/skills) are shared between both.
|
||||
|
||||
## How it works
|
||||
|
||||
```
|
||||
Your Claude Code session (in any project)
|
||||
│
|
||||
│ MCP over HTTP (Tailscale)
|
||||
├──▶ supervisor :3200 (NodePort 30320 on koala) — skill workers: tdd, debug, spec, …
|
||||
├──▶ routing :3210 (NodePort 30310 on koala) — Mode 2 only: review, debug, retrospective, trainer
|
||||
├──▶ routing :3210 (NodePort 30310 on koala) — review, debug, retrospective, trainer
|
||||
└──▶ brain :3300 (NodePort 30330 on koala) — brain_query, brain_write, brain_ingest, session_log
|
||||
│
|
||||
└─ also serves the legacy REST endpoints (/query, /write, /ingest, …)
|
||||
@@ -20,34 +37,28 @@ Your Claude Code session (in any project)
|
||||
▼
|
||||
brain/
|
||||
├── sessions/ — JSONL log, one file per session_id
|
||||
├── wiki/ — searchable knowledge (full-text)
|
||||
│ ├── concepts/
|
||||
│ ├── entities/
|
||||
│ └── sources/
|
||||
├── raw/ — retrospective output, staged for review
|
||||
└── training-data/ — SFT/DPO/RL data (Phase 2)
|
||||
├── wiki/ — searchable knowledge (wing/hall layout)
|
||||
│ ├── homelab/
|
||||
│ ├── claude-sessions/
|
||||
│ └── ...
|
||||
└── knowledge/ — legacy flat notes (migration pending: hyperguild#22)
|
||||
```
|
||||
|
||||
## Phase 1 tools (available now)
|
||||
|
||||
| Tool | What it does |
|
||||
|------|-------------|
|
||||
| `tdd_red` | Writes a failing test for a spec, verifies it fails |
|
||||
| `tdd_green` | Writes the minimal implementation to make tests pass |
|
||||
| `tdd_refactor` | Cleans up implementation while keeping tests green |
|
||||
| `session_log` | Appends a structured entry to the session JSONL log |
|
||||
| `retrospective` | Reads the session log, identifies novel learnings, writes to brain/raw/ |
|
||||
| `retrospective` | Reads the session log, identifies novel learnings, writes to brain |
|
||||
| `review` | Structured code review via local model, brain-context injected |
|
||||
| `debug` | Hypothesis-driven debugging via local model |
|
||||
| `brain_query` | Full-text search over brain/wiki/ |
|
||||
| `brain_write` | Writes a note to brain/raw/ (with optional YAML frontmatter) |
|
||||
| `brain_write` | Writes a note to brain (with wing/hall routing) |
|
||||
| `brain_answer` | BM25 + LLM synthesis — Q&A over brain corpus |
|
||||
| `tier` | Returns the current connectivity tier (1=cloud, 2=LAN, 3=offline) |
|
||||
|
||||
## Start the servers
|
||||
|
||||
```bash
|
||||
# Requires goreman: go install github.com/mattn/goreman@latest
|
||||
task start # starts ingestion (:3300) + supervisor (:3200) via goreman
|
||||
task stop # kills both by port
|
||||
```
|
||||
> **Note:** `tdd_red/green/refactor` and `spec` were retired in Plan 7 (2026-05-12).
|
||||
> They are now SKILL.md files in [mathias/skills](http://gitea.d-ma.be/mathias/skills).
|
||||
|
||||
## Connect a project
|
||||
|
||||
@@ -56,9 +67,9 @@ Create `.mcp.json` in your project root:
|
||||
```json
|
||||
{
|
||||
"mcpServers": {
|
||||
"supervisor": {
|
||||
"routing": {
|
||||
"type": "http",
|
||||
"url": "http://koala:30320/mcp"
|
||||
"url": "http://koala:30310/mcp"
|
||||
},
|
||||
"brain": {
|
||||
"type": "http",
|
||||
@@ -68,65 +79,77 @@ Create `.mcp.json` in your project root:
|
||||
}
|
||||
```
|
||||
|
||||
Two MCP servers are exposed today, both reachable over Tailscale:
|
||||
Two MCP servers are exposed, both reachable over Tailscale:
|
||||
|
||||
- **`supervisor`** at `koala:30320` — skill workers (`tdd_red/green/refactor`,
|
||||
`review`, `debug`, `spec`, `retrospective`, `trainer`, `tier`).
|
||||
- **`routing`** at `koala:30310` — skill workers (`review`, `debug`, `retrospective`, `trainer`).
|
||||
Routes each call to fast local model or thinking model based on per-skill pass rate.
|
||||
- **`brain`** at `koala:30330` — knowledge access (`brain_query`, `brain_write`,
|
||||
`brain_ingest`, `brain_ingest_raw`) and `session_log`. Hosted by the ingestion
|
||||
service directly, no separate pod.
|
||||
`brain_ingest`, `brain_ingest_raw`, `brain_answer`, `brain_classify`) and `session_log`.
|
||||
|
||||
No local binary or stdio shim is required — Claude Code talks to both via HTTP.
|
||||
|
||||
Open Claude Code in your project — run `/mcp` to confirm both servers are listed.
|
||||
|
||||
## A typical TDD session
|
||||
## A typical session
|
||||
|
||||
```
|
||||
1. Call tdd_red → spec in, failing test file out
|
||||
2. Call tdd_green → test path in, implementation out
|
||||
3. Call tdd_refactor → impl + test in, cleaned code out
|
||||
4. Call session_log → log each phase result
|
||||
5. Call retrospective → extracts learnings → brain/raw/
|
||||
6. Review brain/raw/, move worthy notes to brain/wiki/concepts/
|
||||
7. Future sessions: call brain_query to retrieve relevant context
|
||||
1. Call review → brain context injected + local model review → findings
|
||||
2. Call session_log → log each phase result
|
||||
3. Call retrospective → extracts learnings → brain
|
||||
4. Future sessions: call brain_query / brain_answer to retrieve relevant context
|
||||
```
|
||||
|
||||
## Tier detection
|
||||
|
||||
The supervisor probes connectivity at call time:
|
||||
The routing pod probes connectivity at call time:
|
||||
|
||||
| Tier | Label | Condition |
|
||||
|------|-------|-----------|
|
||||
|------|-------|-----------|
|
||||
| 1 | full-online | Can reach api.anthropic.com |
|
||||
| 2 | lan-only | Can reach LiteLLM but not Anthropic |
|
||||
| 3 | airplane | No external connectivity |
|
||||
|
||||
## Model routing
|
||||
|
||||
The routing pod selects models per skill call based on historical pass rate:
|
||||
|
||||
| Pass rate | Decision |
|
||||
|-----------|----------|
|
||||
| ≥ 0.90 (FLOOR) | Fast model (`HYPERGUILD_FAST_MODEL`) |
|
||||
| ≤ 0.70 (CEIL) | Thinking model (`HYPERGUILD_THINKING_MODEL`) |
|
||||
| between CEIL and FLOOR | Sample band — probabilistic routing |
|
||||
| nil (no history yet) | Defaults to thinking model |
|
||||
|
||||
> **Bootstrap note:** With no session history, all calls route to the thinking model.
|
||||
> The fast-model path activates only after real pass-rate data accumulates at `/pass-rate`.
|
||||
> Seed with real usage — don't try to pre-populate.
|
||||
|
||||
## Key env vars
|
||||
|
||||
| Variable | Default | Purpose |
|
||||
|----------|---------|---------|
|
||||
|----------|---------|---------|
|
||||
| `INGEST_BRAIN_DIR` | `../brain` | Brain directory for ingestion server |
|
||||
| `INGEST_PORT` | `3300` | Ingestion server port |
|
||||
| `SUPERVISOR_CONFIG_DIR` | `./config/supervisor` | Skill discipline files |
|
||||
| `SUPERVISOR_SESSIONS_DIR` | `./brain/sessions` | JSONL session logs |
|
||||
| `INGEST_BASE_URL` | `http://localhost:3300` | Supervisor → ingestion |
|
||||
| `INGEST_BASE_URL` | `http://localhost:3300` | Routing pod → brain |
|
||||
| `LITELLM_BASE_URL` | — | LiteLLM proxy for Tier 2 model routing |
|
||||
| `SUPERVISOR_MCP_TOKEN` | — | Optional bearer token for the supervisor MCP HTTP endpoint; when empty, no auth is enforced |
|
||||
| `ROUTING_PORT` | `3210` | Routing pod's listen port |
|
||||
| `ROUTING_MCP_TOKEN` | — | Optional bearer token for the routing MCP HTTP endpoint |
|
||||
| `ROUTING_MCP_TOKEN` | — | Optional bearer token; when empty, no auth enforced |
|
||||
| `BRAIN_URL` | `http://ingestion.supervisor:3300` | Routing pod → brain (in-cluster) |
|
||||
| `HYPERGUILD_FAST_MODEL` | `koala/qwen35-9b-fast` | Fast model for high-pass-rate skill calls |
|
||||
| `HYPERGUILD_THINKING_MODEL` | `iguana/gemma4-26b` | Thinking model for low-pass-rate skill calls |
|
||||
| `HYPERGUILD_ROUTE_LOCAL_FLOOR` | `0.90` | At/above pass rate, route to fast model |
|
||||
| `HYPERGUILD_ROUTE_LOCAL_CEIL` | `0.70` | Below pass rate, route to thinking model. Between CEIL and FLOOR is the sample band. |
|
||||
| `HYPERGUILD_ROUTE_LOCAL_FLOOR` | `0.90` | Fast model threshold |
|
||||
| `HYPERGUILD_ROUTE_LOCAL_CEIL` | `0.70` | Thinking model threshold |
|
||||
| `HYPERGUILD_PASS_RATE_TTL_SECONDS` | `60` | Per-skill pass-rate cache TTL |
|
||||
|
||||
> **Operator note:** LiteLLM at `LITELLM_BASE_URL` must register both `HYPERGUILD_FAST_MODEL` and `HYPERGUILD_THINKING_MODEL` for routing to do useful work. If a model is missing, LiteLLM returns 4xx, the routing pod's fast route fails, the fail-open retry on the thinking model likely also fails (since both are missing), and the only signal is `final_status: "fail"` on `_routing` entries in the brain.
|
||||
> **Operator note:** LiteLLM at `LITELLM_BASE_URL` must register both `HYPERGUILD_FAST_MODEL`
|
||||
> and `HYPERGUILD_THINKING_MODEL`. If a model is missing, the fail-open retry also fails and
|
||||
> the only signal is `final_status: "fail"` on `_routing` entries in the brain.
|
||||
|
||||
## Phase 2 (planned)
|
||||
## Open issues
|
||||
|
||||
- `review` skill — structured code review with iron law enforcement
|
||||
- `debug` skill — hypothesis-driven debugging sessions
|
||||
- `spec` skill — generates specs from conversations
|
||||
- `trainer` — extracts SFT/DPO pairs from session logs for fine-tuning
|
||||
See [issues](http://gitea.d-ma.be/mathias/hyperguild/issues) — key open items:
|
||||
|
||||
- **#25** — skills platform overhaul (audit first, then lazy loading + brain feedback loop)
|
||||
- **#24** — reduce context burn from skill listing
|
||||
- **#22** — migrate legacy brain notes to wing/hall layout (one-shot script, low risk)
|
||||
- **#31** — connect routing-mcp to claude.ai as custom connector
|
||||
|
||||
+2
-4
@@ -17,8 +17,6 @@ tasks:
|
||||
cmds: [bash scripts/context-sync.sh claude]
|
||||
context:sync:agents:
|
||||
cmds: [bash scripts/context-sync.sh agents]
|
||||
context:sync:cursor:
|
||||
cmds: [bash scripts/context-sync.sh cursor]
|
||||
|
||||
# ── Development ────────────────────────────────────────────────────────────
|
||||
|
||||
@@ -90,12 +88,12 @@ tasks:
|
||||
cmds:
|
||||
- task: context:sync
|
||||
- cmd: |
|
||||
drift=$(git status --porcelain -- AGENTS.md CLAUDE.md .cursorrules .aider.conventions.md .context/system-prompt.txt 2>/dev/null)
|
||||
drift=$(git status --porcelain -- AGENTS.md CLAUDE.md .context/system-prompt.txt 2>/dev/null)
|
||||
if [ -n "$drift" ]; then
|
||||
echo "ERROR: derived adapters drifted from canonical context." >&2
|
||||
echo "$drift" >&2
|
||||
echo "" >&2
|
||||
echo "Run: git add AGENTS.md CLAUDE.md .cursorrules .aider.conventions.md .context/system-prompt.txt" >&2
|
||||
echo "Run: git add AGENTS.md CLAUDE.md .context/system-prompt.txt" >&2
|
||||
echo " git commit -m 'chore: re-sync context adapters'" >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
@@ -0,0 +1,48 @@
|
||||
{"_meta":true,"note":"Agent-consumer column of brain-MCP intent analysis. consumer_type fixed=autonomous_agent. CAVEAT: canonical schema file brain-intent-extraction.md is NOT present on this host (koala) — only this session's own task prompt references it. The closed intent vocabulary below was RECONSTRUCTED from the task prompt's framing + brain/schema.md. Re-map intent labels if the canonical vocab differs. schema_source=reconstructed on every row.","closed_intent_vocab":["semantic_retrieval","lexical_lookup","check_prior_art","synthesized_answer","store_new_knowledge","update_or_supersede","ingest_raw_source","verify_write_landed","discover_capability","intent_unclear"],"intent_tool_match_values":["match","mismatch","partial"],"corpus":"~/.claude/projects/*/*.jsonl (Claude Code agent transcripts on koala). brain/sessions/*.jsonl empty. agentsquad docs/eval/*.jsonl are code-review eval results, NOT brain calls. No separate Crush logs found. Zero brain calls appear under any mcp__ name with a human typing the call — all brain acts are agent-initiated (CLAUDE.md reflex), so all qualify as autonomous_agent."}
|
||||
{"id":"a01","session":"tapir-c","ts":"2026-06-?T15:01:51","tool":"brain_query","intent":"check_prior_art","intent_tool_match":"match","workaround":null,"observed_friction":null,"evidence":"single pre-task query 'YouTube Data API captions download ownership limitation timedtext adapter Go' — named-entity lexical lookup, fit BM25 well, no reformulation."}
|
||||
{"id":"a02","session":"tapir","ts":"2026-06-05T21:44:10","tool":"brain_ingest","intent":"ingest_raw_source","intent_tool_match":"match","workaround":null,"observed_friction":"preceded by ToolSearch select:brain_ingest 14s earlier — tool not ambient, had to be discovered/loaded first.","evidence":"source=tapir-scheduled-discovery-session-2026-06-05, a session learnings dump."}
|
||||
{"id":"a03","session":"tapir","ts":"2026-06-?T14:00:59","tool":"brain_ingest","intent":"ingest_raw_source","intent_tool_match":"match","workaround":null,"observed_friction":null,"evidence":"source=tapir-rls-identity-bootstrapping, first write of RLS lesson."}
|
||||
{"id":"a04","session":"tapir","ts":"2026-06-?T14:02:08","tool":"brain_ingest","intent":"update_or_supersede","intent_tool_match":"mismatch","workaround":"RE-INGESTED same source name 'tapir-rls-identity-bootstrapping' 69s later with edited/condensed body. No update/patch/supersede verb exists, so the agent overwrote-by-re-ingest. Whether this dedups or creates a v2 duplicate is opaque to the agent.","observed_friction":"agent revised content within 70s of first write — classic edit-after-write with no edit primitive.","evidence":"two brain_ingest, identical source string, divergent content."}
|
||||
{"id":"a05","session":"tapir","ts":"2026-06-?T14:56:27","tool":"brain_write","intent":"store_new_knowledge","intent_tool_match":"match","workaround":null,"observed_friction":null,"evidence":"postgres-cascade-skips-tables-without-fk.md — distinct new lesson."}
|
||||
{"id":"a06","session":"tapir","ts":"2026-06-?T21:08:57","tool":"brain_query","intent":"check_prior_art","intent_tool_match":"match","workaround":null,"observed_friction":"preceded by ToolSearch select:brain_query — discovery tax again.","evidence":"'Dex passwords.dex.coreos.com CRD ...' keyword-rich, single shot."}
|
||||
{"id":"a07","session":"AI-infra","ts":"2026-05-?T15:38:03","tool":"brain_query(HTTP-curl)","intent":"discover_capability","intent_tool_match":"mismatch","workaround":"raw `curl -X POST` to brain-mcp endpoint instead of MCP tool. Preceded by two ToolSearch ('brain knowledge memory' then 'brain') that did not yield a usable loaded tool, so agent fell back to HTTP.","observed_friction":"3-step ladder: ToolSearch 'brain knowledge memory' -> ToolSearch 'brain' -> curl. Agent did not know which act maps to which tool name.","evidence":"curl -s -o /tmp/brain-init.txt -w code:%{http_code} -X POST ..."}
|
||||
{"id":"a08","session":"AI-infra","ts":"2026-05-?T04:46:13","tool":"brain_query(HTTP-curl)","intent":"discover_capability","intent_tool_match":"mismatch","workaround":"hand-set TOKEN=... then curl brain-test endpoint — probing whether the HTTP brain path is reachable/authed at all. MCP path not used.","observed_friction":"agent testing connectivity by hand; MCP auth/availability not trusted.","evidence":"TOKEN=...; curl -s -o /tmp/brain-test ..."}
|
||||
{"id":"a09","session":"AI-infra","ts":"2026-05-?T05:23:35","tool":"brain_query","intent":"check_prior_art","intent_tool_match":"match","workaround":null,"observed_friction":"preceded by ToolSearch select:brain_query (after an earlier 05:03 ToolSearch 'brain ingestion knowledge wiki' that explored layers).","evidence":"'koala machine state RTX 5070 llama-swap' — named-entity recall, fits lexical."}
|
||||
{"id":"a10","session":"AI-infra","ts":"2026-05-?T05:27:36","tool":"brain_query","intent":"check_prior_art","intent_tool_match":"match","workaround":null,"observed_friction":null,"evidence":"batch of 3 in ~1s (k3s/flux gitops; llama-swap ai-stack GPU; MCP Dex OAuth claude.ai) — parallel prior-art sweep, all named-entity."}
|
||||
{"id":"a11","session":"AI-infra","ts":"2026-05-?T07:13:50","tool":"brain_query","intent":"check_prior_art","intent_tool_match":"match","workaround":null,"observed_friction":null,"evidence":"batch of 4 in ~2s before a debugging session (flux healthCheck; exit 255 restart loop; NVML mismatch; mirror rebase). Named symptoms, lexical fit OK on first pass."}
|
||||
{"id":"a12","session":"AI-infra","ts":"2026-05-?T07:14:26","tool":"brain_write","intent":"store_new_knowledge","intent_tool_match":"match","workaround":null,"observed_friction":null,"evidence":"4 writes in ~35s (flux-healthcheck-stale; exit-255-unknown-reason-not-oom; nvidia-nvml-mismatch; mcp-static-bearer) — answers to the 4 queries just run, captured as lessons. Healthy query->fix->write loop."}
|
||||
{"id":"a13","session":"AI-infra","ts":"2026-05-?T07:15:25","tool":"brain_query","intent":"verify_write_landed","intent_tool_match":"mismatch","workaround":"25s after writing exit-255-unknown-reason-not-oom.md, re-queried 'exit 255 unknown SIGKILL containerd' — reformulated terms (SIGKILL/containerd not in original query 'exit 255 unknown reason restart loop diagnosis'). Either confirming the fresh write is retrievable or re-searching because first lexical query missed. No read-after-write / get-by-id act exists.","observed_friction":"reformulation chain: 'exit 255 unknown reason restart loop diagnosis' -> 'exit 255 unknown SIGKILL containerd'. Same need, different keywords.","evidence":"query at 07:13:51 vs 07:15:25 bracketing the 07:14:37 write."}
|
||||
{"id":"a14","session":"AI-infra","ts":"2026-05-?T07:22:55","tool":"brain_query","intent":"check_prior_art","intent_tool_match":"match","workaround":null,"observed_friction":null,"evidence":"batch of 5 in ~20s before a homelab security audit (piguard/iguana tailscale; unifi UCG firewall; SOPS age; ingress TLS cert-manager; koala UFW iptables). Named-entity sweep."}
|
||||
{"id":"a15","session":"AI-infra","ts":"2026-05-?T09:27:53","tool":"brain_write","intent":"store_new_knowledge","intent_tool_match":"match","workaround":null,"observed_friction":null,"evidence":"audit-shortcut-tls-blocks-zero; policy-audit-mode-blocks-nothing — distinct new audit lessons."}
|
||||
{"id":"a16","session":"AI-infra","ts":"2026-05-?T09:28:22","tool":"brain_write","intent":"store_new_knowledge","intent_tool_match":"match","workaround":null,"observed_friction":null,"evidence":"homelab-security-chains-not-bugs.md FIRST write (worked example: koala 2026-05-13)."}
|
||||
{"id":"a17","session":"AI-infra","ts":"2026-05-?T10:32:49","tool":"brain_write","intent":"update_or_supersede","intent_tool_match":"mismatch","workaround":"RE-WROTE homelab-security-chains-not-bugs.md ~64min later with a different/expanded worked example (host-user dotfile, over-broad ClusterRole). Same filename, additive revision, no patch/append/supersede verb — agent overwrites and hopes the index replaces rather than duplicates.","observed_friction":"the in-between hour of audit work produced a better example; only way to fold it in was a full re-write of the same slug.","evidence":"two brain_write same filename at 09:28:22 and 10:32:49, divergent worked examples."}
|
||||
{"id":"a18","session":"AI-infra","ts":"2026-05-?T10:18:25","tool":"brain_write","intent":"store_new_knowledge","intent_tool_match":"match","workaround":null,"observed_friction":null,"evidence":"homelab-document-accepted-risk-to-break-audit-cycle.md — distinct."}
|
||||
{"id":"a19","session":"AI-infra","ts":"2026-05-?T10:32:55","tool":"brain_query","intent":"verify_write_landed","intent_tool_match":"mismatch","workaround":"6s after the homelab-chains re-write, queried 'RBAC MCP cluster pods log chain' — checking the chain reasoning is retrievable / finding the related entry. Read-after-write done via lexical search.","observed_friction":null,"evidence":"query immediately follows the 10:32:49 write."}
|
||||
{"id":"a20","session":"AI-infra","ts":"2026-05-?T18:57:34","tool":"brain_query","intent":"check_prior_art","intent_tool_match":"match","workaround":null,"observed_friction":"preceded by ToolSearch select:brain_write,brain_query — re-discovered tools this session.","evidence":"'extension build pinned version major version upgrade postgres pgvector'."}
|
||||
{"id":"a21","session":"AI-infra","ts":"2026-05-?T18:57:58","tool":"brain_write","intent":"store_new_knowledge","intent_tool_match":"match","workaround":null,"observed_friction":null,"evidence":"extension-version-lags-platform-major-upgrade.md."}
|
||||
{"id":"a22","session":"AI-infra","ts":"2026-05-?T18:58:06","tool":"brain_query","intent":"verify_write_landed","intent_tool_match":"mismatch","workaround":"8s after writing extension-version-lags, re-queried 'pgvector postgres extension version compile error bump' — reformulated from the 18:57:34 query ('extension build pinned version...'). Lexical re-search to confirm the just-written lesson is findable, with different keyword guess.","observed_friction":"reformulation: 'extension build pinned version major version upgrade postgres pgvector' -> 'pgvector postgres extension version compile error bump'.","evidence":"write at 18:57:58 bracketed by queries 18:57:34 and 18:58:06."}
|
||||
{"id":"a23","session":"AI-infra","ts":"2026-05-?T18:31:23","tool":"brain_write","intent":"store_new_knowledge","intent_tool_match":"match","workaround":null,"observed_friction":null,"evidence":"webfetch-readme-when-image-or-flag-uncertain.md FIRST write."}
|
||||
{"id":"a24","session":"AI-infra","ts":"2026-05-?T18:34:23","tool":"brain_write","intent":"update_or_supersede","intent_tool_match":"mismatch","workaround":"RE-WROTE webfetch-readme-when-image-or-flag-uncertain.md 3min later, near-identical body. Looks like a retry/overwrite (uncertain the first landed, or minor edit). No idempotent upsert with confirmation, so agent re-fires the write.","observed_friction":"followed 7s later by a brain_query on the same topic ('OSS tool image registry CLI flag webhook path schema drift README pre-flight') — write-write-query, i.e. overwrite then verify-by-search.","evidence":"two brain_write same filename 18:31:23 / 18:34:23, then query 18:34:31."}
|
||||
{"id":"a25","session":"AI-infra","ts":"2026-05-?T18:34:31","tool":"brain_query","intent":"verify_write_landed","intent_tool_match":"mismatch","workaround":"keyword-stuffed lexical query 'OSS tool image registry CLI flag webhook path schema drift README pre-flight' fired right after the webfetch-readme write — agent dumps every concept token hoping BM25 surfaces its own fresh note. This is semantic intent (find that conceptual lesson) coerced into a bag-of-keywords.","observed_friction":"query is a concatenation of the note's section headings — a tell that the agent is groping lexically for content it knows by meaning.","evidence":"query text mirrors the just-written note's bullet topics."}
|
||||
{"id":"a26","session":"dev","ts":"2026-06-?T21:18:26","tool":"brain_answer","intent":"synthesized_answer","intent_tool_match":"match","workaround":null,"observed_friction":"preceded by ToolSearch select:brain_answer.","evidence":"'tapir transcript persistence shared cross-user dedup table RLS isolation ADR-021 ...' -> 22min later a brain_write (acted on the answer). Answer consumed, not re-queried. Healthy."}
|
||||
{"id":"a27","session":"dev","ts":"2026-06-?T21:40:20","tool":"brain_write","intent":"store_new_knowledge","intent_tool_match":"match","workaround":null,"observed_friction":null,"evidence":"tapir-migration-and-rls-test-infra-gotchas, with wing/hall absent here (flat) — see schema-confusion note a40."}
|
||||
{"id":"a28","session":"dev","ts":"2026-06-?T05:35:04","tool":"brain_answer","intent":"synthesized_answer","intent_tool_match":"partial","workaround":"asked 'gitea MCP not working workaround file issue via API which token ... how to authenticate gitea API' — a how-do-I question. Next brain act (05:39 query) is a different topic (tapir transcript), so the answer was apparently sufficient OR abandoned; ambiguous.","observed_friction":null,"evidence":"brain_answer then unrelated brain_query 4min later."}
|
||||
{"id":"a29","session":"dev","ts":"2026-06-?T05:39:54","tool":"brain_query","intent":"check_prior_art","intent_tool_match":"match","workaround":null,"observed_friction":null,"evidence":"'tapir transcript persistence ADR-021 shared non-RLS'."}
|
||||
{"id":"a30","session":"dev","ts":"2026-06-?T05:57:37","tool":"brain_write","intent":"store_new_knowledge","intent_tool_match":"match","workaround":null,"observed_friction":null,"evidence":"gitea-mcp-per-repo-tools-404-and-rest-fallback FIRST write."}
|
||||
{"id":"a31","session":"dev","ts":"2026-06-?T05:57:55","tool":"brain_write","intent":"update_or_supersede","intent_tool_match":"mismatch","workaround":"RE-WROTE gitea-mcp-per-repo-tools-404-and-rest-fallback 18s later — overwrite/retry of same slug, no upsert confirmation.","observed_friction":"sub-20s gap = almost certainly a content tweak the agent could not express as an edit.","evidence":"two brain_write same filename 05:57:37 / 05:57:55."}
|
||||
{"id":"a32","session":"dev","ts":"2026-05-?T11:51:47","tool":"brain_write","intent":"store_new_knowledge","intent_tool_match":"match","workaround":null,"observed_friction":null,"evidence":"infra-litellm-absorption-2026-05-16.md."}
|
||||
{"id":"a33","session":"dev","ts":"2026-05-?T12:07:04","tool":"brain_query","intent":"check_prior_art","intent_tool_match":"match","workaround":null,"observed_friction":null,"evidence":"batch of 3 (litellm rebuild time piguard; docker compose orphaned volumes; prometheus_client ModuleNotFoundError) — lexical, error-string driven."}
|
||||
{"id":"a34","session":"dev","ts":"2026-05-?T15:08:15","tool":"brain_answer","intent":"synthesized_answer","intent_tool_match":"mismatch","workaround":"THREE brain_answer at 15:08 (moved compose volumes? / pi rebuild time? / litellm ModuleNotFound prometheus) — the SAME three topics queried lexically an hour earlier (12:07) — were IMMEDIATELY followed at 15:09 by THREE brain_query on the same three topics. The agent asked the synthesizer, was unsatisfied, and fell straight back to raw lexical search. Strongest answer->query fallback in the corpus.","observed_friction":"answer/query duplication across one intent: agent hedges by firing both interfaces, trusting neither.","evidence":"15:08 answers vs 15:09 queries, topic-for-topic aligned."}
|
||||
{"id":"a35","session":"dev","ts":"2026-05-?T15:09:16","tool":"brain_query","intent":"semantic_retrieval","intent_tool_match":"mismatch","workaround":"after the 3 brain_answer calls failed to satisfy, re-issued as lexical brain_query ('moved compose stack to new directory volumes disappeared empty'; 'raspberry pi docker build time arm slow'; 'how to enable prometheus metrics on litellm proxy callback'). The want is meaning-based ('did my volumes move?') but the only retrieval that 'worked' was keyword search — and these are full natural-language sentences crammed into a BM25 box.","observed_friction":"natural-language questions ('how to enable...', 'moved ... disappeared') passed to a lexical query tool — semantic intent, lexical interface.","evidence":"3 queries at 15:09 mirror the 3 answers at 15:08."}
|
||||
{"id":"a36","session":"dev","ts":"2026-05-?T20:31:50","tool":"brain_answer","intent":"synthesized_answer","intent_tool_match":"match","workaround":null,"observed_friction":null,"evidence":"'What happened with the litellm migration on 2026-05-16?' — episodic recall question, answer fit; no re-query followed."}
|
||||
{"id":"a37","session":"dev","ts":"2026-05-?T21:07:17","tool":"brain_query","intent":"semantic_retrieval","intent_tool_match":"mismatch","workaround":"FOUR-step reformulation chain over one Go bug: 'bytes.Buffer Bytes Reset aliasing slice sharing' -> 'go buffer reuse map backing array bug' -> [write go-bytes-buffer-bytes-reset-aliasing-trap.md] -> 'go map values all show same content after loop' -> 'bytes.Buffer Bytes returns same data every iteration'. The agent knows the SYMPTOM (all map values identical) and the CAUSE (Bytes() aliasing) but cannot phrase a single lexical query that bridges them — it wants concept retrieval and is forced to brute-force keyword variants.","observed_friction":"4 distinct phrasings of the same bug, two before and two after writing the lesson — also doubles as verify_write_landed on the trailing queries.","evidence":"21:07:17, 21:07:17, (write 21:08:01), 21:08:10, 21:08:19."}
|
||||
{"id":"a38","session":"dev","ts":"2026-05-?T21:08:01","tool":"brain_write","intent":"store_new_knowledge","intent_tool_match":"match","workaround":null,"observed_friction":null,"evidence":"go-bytes-buffer-bytes-reset-aliasing-trap.md."}
|
||||
{"id":"a39","session":"dev","ts":"2026-05-?T07:54:26","tool":"brain_write","intent":"store_new_knowledge","intent_tool_match":"match","workaround":null,"observed_friction":null,"evidence":"mcp-tool-design-get-needs-list-partner.md — a design principle."}
|
||||
{"id":"a40","session":"dev","ts":"2026-06-?T06:25:00","tool":"brain_query","intent":"check_prior_art","intent_tool_match":"match","workaround":null,"observed_friction":null,"evidence":"'Dex to Authentik migration auth.d-ma.be issuer cutover OIDC subject ...'."}
|
||||
{"id":"a41","session":"dev","ts":"2026-06-?T06:25:08","tool":"brain_answer","intent":"synthesized_answer","intent_tool_match":"mismatch","workaround":"brain_query (a40) and brain_answer (a41) fired ~8s apart on the SAME intent (Dex->Authentik subject-keyed token orphan). Agent runs lexical search AND synthesized answer in parallel for one question rather than choosing — it cannot predict which interface will return usable knowledge, so it pays both.","observed_friction":"query+answer doublet on one need.","evidence":"06:25:00 query then 06:25:08 answer, same topic."}
|
||||
{"id":"a42","session":"dev","ts":"2026-06-?T13:48:49","tool":"brain_query","intent":"verify_write_landed","intent_tool_match":"mismatch","workaround":"'authentik cutover validation probe' then 6min later 'authentik cutover post-flip validation' — reformulated pair, likely searching for the agent's own earlier cutover notes / confirming validation steps are recorded. Lexical re-search standing in for recall-my-recent-context.","observed_friction":"reformulation: 'validation probe' -> 'post-flip validation'.","evidence":"13:48:49 and 13:54:56."}
|
||||
{"id":"a43","session":"dev","ts":"2026-06-?T13:59:17","tool":"brain_write","intent":"store_new_knowledge","intent_tool_match":"match","workaround":null,"observed_friction":"preceded by ToolSearch select:brain_write.","evidence":"oidc-issuer-host-change-vs-idp-swap-subject."}
|
||||
{"id":"a44","session":"dev","ts":"2026-06-?T13:59:50","tool":"brain_write","intent":"store_new_knowledge","intent_tool_match":"match","workaround":null,"observed_friction":null,"evidence":"cannot-move-ingress-host-across-namespaces-flux-dryrun FIRST write."}
|
||||
{"id":"a45","session":"dev","ts":"2026-06-?T14:00:13","tool":"brain_write","intent":"update_or_supersede","intent_tool_match":"mismatch","workaround":"RE-WROTE cannot-move-ingress-host-across-namespaces-flux-dryrun 23s later — overwrite of same slug, no edit/upsert primitive.","observed_friction":"sub-30s gap = content correction expressed as a full re-write.","evidence":"two brain_write same filename 13:59:50 / 14:00:13."}
|
||||
{"id":"a46","session":"dev","ts":"2026-06-?T13:53:54","tool":"brain_write(HTTP-staged)","intent":"store_new_knowledge","intent_tool_match":"mismatch","workaround":"after `ToolSearch select:mcp__claude_ai_brain__authenticate` (MCP auth flow), the agent staged the entry as `cat > /tmp/brain_entry.json` ({filename:'postgres-force-rls-cross-u...', content}) for a curl write rather than calling brain_write directly — MCP write path was not usable (auth/loading), so it dropped to the HTTP bodge.","observed_friction":"reached for an 'authenticate' tool, then abandoned MCP for hand-built JSON + curl. Matches known pattern: brain/op MCP auth lapses often.","evidence":"ToolSearch authenticate 13:53:27 -> cat /tmp/brain_entry.json 13:53:54."}
|
||||
{"id":"a47","session":"template-go-agent","ts":"2026-05-?T18:46:26","tool":"BASH(not-a-brain-act)","intent":"intent_unclear","intent_tool_match":"match","workaround":null,"observed_friction":null,"evidence":"'brain_' substring was inside a git commit message body ('agent boundaries, network policy, agent s...'), NOT a brain call. Excluded from knowledge-act analysis; logged for audit completeness."}
|
||||
@@ -0,0 +1,148 @@
|
||||
# Agent-Consumer Brain Intent Analysis — koala column
|
||||
|
||||
**Consumer:** `autonomous_agent` (all rows). **Host:** koala. **Date:** 2026-06-15.
|
||||
**Raw rows:** `agent-intent-column.jsonl` (46 real knowledge-acts + 1 excluded false-positive).
|
||||
|
||||
## Caveat — canonical schema not on this host
|
||||
|
||||
The shared closed-vocabulary file `brain-intent-extraction.md` **does not exist on
|
||||
koala** — the only reference to it is inside *this task's own prompt*. The intent
|
||||
vocabulary below was **reconstructed** from the prompt's framing + `brain/schema.md`.
|
||||
Every row carries `schema_source: reconstructed`. If the canonical vocab differs,
|
||||
re-map the `intent` field; the `intent_tool_match` / `workaround` / `observed_friction`
|
||||
evidence stands regardless of label names.
|
||||
|
||||
**Reconstructed closed vocab:** `semantic_retrieval`, `lexical_lookup`,
|
||||
`check_prior_art`, `synthesized_answer`, `store_new_knowledge`, `update_or_supersede`,
|
||||
`ingest_raw_source`, `verify_write_landed`, `discover_capability`, `intent_unclear`.
|
||||
|
||||
## Corpus
|
||||
|
||||
- `~/.claude/projects/*/*.jsonl` — Claude Code agent transcripts (98 files). **The only
|
||||
source with brain calls.**
|
||||
- `brain/sessions/*.jsonl` — empty (only `.gitkeep`).
|
||||
- `agentsquad docs/eval/*.jsonl` — code-review eval results, **not** brain calls.
|
||||
- No separate Crush session logs on this host.
|
||||
- **Zero** brain calls were human-typed. Every brain act is agent-initiated (the
|
||||
CLAUDE.md "query as reflex / close-the-loop write" behaviour), so all qualify as
|
||||
`autonomous_agent`. The human gave the top-level task; the agent chose every brain act.
|
||||
|
||||
## 1. Intent histogram, split by `intent_tool_match`
|
||||
|
||||
| intent | match | mismatch | partial | total |
|
||||
|---|---|---|---|---|
|
||||
| check_prior_art | 10 | 0 | 0 | 10 |
|
||||
| store_new_knowledge | 14 | 1 | 0 | 15 |
|
||||
| update_or_supersede | 0 | 5 | 0 | 5 |
|
||||
| synthesized_answer | 2 | 2 | 1 | 5 |
|
||||
| verify_write_landed | 0 | 4 | 0 | 4 |
|
||||
| semantic_retrieval | 0 | 3 | 0 | 3 |
|
||||
| ingest_raw_source | 2 | 0 | 0 | 2 |
|
||||
| discover_capability | 0 | 2 | 0 | 2 |
|
||||
| **total** | **28** | **17** | **1** | **46** |
|
||||
|
||||
> Batch note: several rows collapse a same-second fan-out of identical-intent calls
|
||||
> (a10=3, a11=4, a14=5, a33=3, a34=3, a35=3). Call-level the corpus is ~62 brain calls;
|
||||
> the table counts the 46 distinct knowledge-acts. Frequency is deliberately *not* the
|
||||
> point — the mismatch column is.
|
||||
|
||||
**37% of agent knowledge-acts (17/46) are interface mismatches.** Every mismatch falls
|
||||
into one of four intents: `update_or_supersede`, `verify_write_landed`,
|
||||
`semantic_retrieval`, `discover_capability` — plus one `store` that had to use HTTP.
|
||||
|
||||
## 2. Mismatch list, grouped by intent (primary deliverable)
|
||||
|
||||
### update_or_supersede → re-write same slug (5/5 mismatch) — HIGHEST VALUE
|
||||
There is **no update / patch / append / supersede verb**. When an agent improves a note
|
||||
it already wrote, the only move is to call `brain_write`/`brain_ingest` **again with the
|
||||
same filename/source** and hope the index replaces rather than duplicates. Observed:
|
||||
|
||||
| slug | 1st write | 2nd write | gap | what changed |
|
||||
|---|---|---|---|---|
|
||||
| `tapir-rls-identity-bootstrapping` (ingest) | 14:00:59 | 14:02:08 | 69s | condensed body |
|
||||
| `homelab-security-chains-not-bugs.md` | 09:28:22 | 10:32:49 | 64m | new worked example |
|
||||
| `webfetch-readme-when-image-or-flag-uncertain.md` | 18:31:23 | 18:34:23 | 3m | near-identical (retry) |
|
||||
| `gitea-mcp-per-repo-tools-404-and-rest-fallback` | 05:57:37 | 05:57:55 | 18s | content tweak |
|
||||
| `cannot-move-ingress-host-across-namespaces-flux-dryrun` | 13:59:50 | 14:00:13 | 23s | content tweak |
|
||||
|
||||
Sub-30s gaps (3 of 5) read as "I wanted to edit but can only overwrite." The agent has
|
||||
no way to know whether the second write deduped or created a contradictory v2 — opacity
|
||||
the brain's own design principle (`mcp-tool-design-get-needs-list-partner.md`, written
|
||||
*by one of these very agents*) would flag: every `_write` needs a `_get`/`_update` partner.
|
||||
|
||||
### verify_write_landed → lexical re-query (4/4 mismatch)
|
||||
No read-after-write / get-by-id confirmation. After every substantive write, agents
|
||||
re-query lexically to check the note is retrievable — and *reformulate the keywords*
|
||||
because they can't predict what BM25 indexed:
|
||||
- `exit-255` lesson: query `exit 255 unknown reason restart loop diagnosis` → write →
|
||||
query `exit 255 unknown SIGKILL containerd`.
|
||||
- `extension-version-lags`: query `extension build pinned version...pgvector` → write →
|
||||
query `pgvector postgres extension version compile error bump`.
|
||||
- `webfetch-readme`: write → write → query stuffed with the note's own section headings.
|
||||
|
||||
### semantic_retrieval → BM25 keyword-stuffing (3/3 mismatch)
|
||||
Agent knows the *meaning* but not the *indexed words*, so it brute-forces phrasings of
|
||||
one need against a lexical tool:
|
||||
- **4-step chain on one Go bug:** `bytes.Buffer Bytes Reset aliasing slice sharing` →
|
||||
`go buffer reuse map backing array bug` → (write) → `go map values all show same
|
||||
content after loop` → `bytes.Buffer Bytes returns same data every iteration`. Symptom
|
||||
and cause both known; no single lexical query bridges them.
|
||||
- Natural-language questions (`how to enable prometheus metrics on litellm proxy
|
||||
callback`, `moved compose stack to new directory volumes disappeared empty`) shoved
|
||||
into `brain_query`.
|
||||
|
||||
### synthesized_answer → fall back to / hedge with brain_query (2 mismatch + 1 partial)
|
||||
`brain_answer` is frequently **not trusted as terminal**:
|
||||
- **Strongest signal:** 3× `brain_answer` at 15:08 (compose volumes / pi rebuild time /
|
||||
litellm ModuleNotFound) → 3× `brain_query` at 15:09 on the *same three topics*. The
|
||||
agent asked the synthesizer, was unsatisfied, and immediately re-ran raw search.
|
||||
- Dex→Authentik: `brain_query` and `brain_answer` fired **8s apart on one question** —
|
||||
the agent pays both interfaces because it can't predict which returns usable knowledge.
|
||||
- (Counter-examples exist: `brain_answer` for episodic recall — "what happened with the
|
||||
litellm migration on 2026-05-16?" — was consumed and not re-queried. So `answer`
|
||||
works for *episodic/temporal* recall, fails for *how-do-I / does-X-hold* reasoning.)
|
||||
|
||||
### discover_capability + store-via-HTTP (3 mismatch)
|
||||
brain tools are **not ambient** — they are deferred and must be `ToolSearch`-loaded each
|
||||
session. Agents fumble the discovery (`ToolSearch 'brain knowledge memory'` →
|
||||
`'brain'` → `'brain ingestion knowledge wiki'`) and, when MCP load/auth fails, drop to
|
||||
**raw `curl` against `brain-mcp` / hand-built `/tmp/brain_entry.json`**. One agent even
|
||||
`ToolSearch`-ed an `authenticate` tool, then abandoned MCP for the HTTP bodge — matching
|
||||
the known "brain/op MCP auth lapses too often" footgun.
|
||||
|
||||
### Write-interface / layer schema confusion (cross-cutting)
|
||||
`brain_write` was called with **three different param shapes** in the same corpus:
|
||||
`{filename, type:"lesson", content}`, `{filename, content}` (no type), and
|
||||
`{wing:"tapir", hall:"failures", filename, content}` — plus `brain_ingest {source,
|
||||
content}`. Agents are unsure which verb and which layer (flat slug vs `wing`/`hall`
|
||||
knowledge routing vs raw ingest) a given knowledge-act maps to. This is the
|
||||
`knowledge/ vs wiki/` confusion expressed at the parameter level.
|
||||
|
||||
## 3. `intent_unclear` rate
|
||||
|
||||
**0 / 46 genuine brain acts (0%).** Agent intent is unusually legible because these are
|
||||
Claude Code transcripts: the surrounding task, the query/filename strings, and the
|
||||
write content all disambiguate. One row (`a47`) was tagged `intent_unclear` and
|
||||
**excluded** — its `brain_` substring was inside a git commit message, not a brain call.
|
||||
Example of the only ambiguity that arose: a `brain_answer` on "gitea MCP not working...
|
||||
how to authenticate" followed by an unrelated query — can't tell if the answer satisfied
|
||||
or was abandoned (`partial`, row a28).
|
||||
|
||||
## 4. The single biggest intent↔interface gap
|
||||
|
||||
**The brain offers one write verb and one lexical read verb, but autonomous agents
|
||||
perform four distinct knowledge-acts against them — and three of the four have no fitting
|
||||
interface.** The deepest gap is the **missing update/supersede path**: agents close every
|
||||
task by writing a lesson (the CLAUDE.md ritual), routinely improve it minutes-to-an-hour
|
||||
later, and — having no edit primitive — re-write the same slug blind, unable to tell
|
||||
whether they corrected the entry or forked a contradiction into the index. This compounds
|
||||
with the lexical-only read side: because there is no `get-by-id` or semantic retrieval,
|
||||
agents can't even reliably *find their own just-written note* to check it, so they
|
||||
keyword-stuff reformulated queries and hedge `brain_answer` with parallel `brain_query`.
|
||||
The interface is built for *append-and-keyword-search*; the agents are trying to
|
||||
*curate a living, deduplicated knowledge base*, and the seam between those two shows up
|
||||
as the 5 blind re-writes, 4 read-after-write re-queries, and 3 semantic-as-lexical chains
|
||||
that dominate the mismatch column.
|
||||
|
||||
---
|
||||
*Evidence-only per task scope — no redesign proposed.*
|
||||
@@ -0,0 +1,140 @@
|
||||
# Brain-MCP Intent↔Interface Findings — Unified (two-column merge)
|
||||
|
||||
**Status — 2026-06-16**
|
||||
- ✅ **Agent column** filled from `agent-intent-column.jsonl` (46 acts, koala).
|
||||
- ⏳ **Human column** = `PENDING`. Drop the Claude.ai-history analysis into
|
||||
`human-intent-column.jsonl` (same dir, schema below), then fill the `PENDING`
|
||||
cells and the synthesis blocks marked `<<SYNTH>>`.
|
||||
- ⚠️ Canonical `brain-intent-extraction.md` still absent on koala. Vocab below is
|
||||
the **reconstructed** lock both columns must share. If the real file surfaces,
|
||||
re-map `intent` labels in *both* columns identically before merging.
|
||||
|
||||
---
|
||||
|
||||
## Shared schema (LOCKED — both columns conform)
|
||||
|
||||
Per-call row, JSONL:
|
||||
|
||||
| field | values / form | notes |
|
||||
|---|---|---|
|
||||
| `id` | `a01..` (agent) / `h01..` (human) | column prefix kept distinct |
|
||||
| `session` | string | source session/conversation id |
|
||||
| `ts` | ISO-8601 | best-effort |
|
||||
| `tool` | brain tool name (+ `(HTTP-curl)` / `(HTTP-staged)` suffix for bodges) | |
|
||||
| `intent` | closed vocab ↓ | the knowledge-act WANTED |
|
||||
| `intent_tool_match` | `match` \| `mismatch` \| `partial` | does the called tool fit the want |
|
||||
| `consumer_type` | `autonomous_agent` \| `human_interactive` | fixed per column |
|
||||
| `workaround` | string \| null | the bodge when mismatch — **primary signal** |
|
||||
| `observed_friction` | string \| null | reformulation chains, discovery tax, hedging |
|
||||
| `evidence` | string | excerpt anchoring the classification |
|
||||
| `schema_source` | `reconstructed` | flip to `canonical` if real vocab lands |
|
||||
|
||||
### Closed intent vocab (LOCKED)
|
||||
`semantic_retrieval`, `lexical_lookup`, `check_prior_art`, `synthesized_answer`,
|
||||
`store_new_knowledge`, `update_or_supersede`, `ingest_raw_source`,
|
||||
`verify_write_landed`, `discover_capability`, `intent_unclear`.
|
||||
|
||||
---
|
||||
|
||||
## Master comparison — by intent
|
||||
|
||||
| intent | agent acts | agent mismatch | human acts | human mismatch | shared gap |
|
||||
|---|---|---|---|---|---|
|
||||
| check_prior_art | 10 | 0% | `PENDING` | `PENDING` | — |
|
||||
| store_new_knowledge | 15 | 7% (1/15) | `PENDING` | `PENDING` | `<<SYNTH>>` |
|
||||
| update_or_supersede | 5 | **100%** (5/5) | `PENDING` | `PENDING` | `<<SYNTH>>` no edit verb |
|
||||
| synthesized_answer | 5 | 40% (2/5)+1 partial | `PENDING` | `PENDING` | `<<SYNTH>>` |
|
||||
| verify_write_landed | 4 | **100%** (4/4) | `PENDING` | `PENDING` | `<<SYNTH>>` no read-after-write |
|
||||
| semantic_retrieval | 3 | **100%** (3/3) | `PENDING` | `PENDING` | `<<SYNTH>>` lexical-only read |
|
||||
| ingest_raw_source | 2 | 0% | `PENDING` | `PENDING` | — |
|
||||
| discover_capability | 2 | **100%** (2/2) | `PENDING` | `PENDING` | agent-specific (ToolSearch/auth)? |
|
||||
| intent_unclear | 0 | — | `PENDING` | `PENDING` | divergence expected ↓ |
|
||||
| **TOTAL** | **46** | **37% (17)** | `PENDING` | `PENDING` | |
|
||||
|
||||
---
|
||||
|
||||
## Per-intent merged findings
|
||||
|
||||
### update_or_supersede — agent: 5/5 mismatch (highest value)
|
||||
**Agent:** no edit/patch/append verb. Agents re-write same slug blind:
|
||||
`homelab-security-chains-not-bugs.md` (+64m), `tapir-rls-identity-bootstrapping`,
|
||||
`webfetch-readme...`, `gitea-mcp-per-repo-tools-404...`,
|
||||
`cannot-move-ingress-host...` — 3 of 5 sub-30s ("wanted edit, got overwrite").
|
||||
Cannot tell if write deduped or forked a contradiction.
|
||||
**Human:** `PENDING` — *look for: user editing a prior note, asking "update what I
|
||||
saved about X", or expressing frustration that an old fact is stale/duplicated.*
|
||||
**<<SYNTH>>** shared verdict once both filled.
|
||||
|
||||
### verify_write_landed — agent: 4/4 mismatch
|
||||
**Agent:** no `get-by-id`/read-after-write. Agents lexically re-query their own
|
||||
fresh note with reformulated keywords (`exit 255 unknown reason` → `...SIGKILL
|
||||
containerd`; `extension build pinned...` → `pgvector ...compile error bump`).
|
||||
**Human:** `PENDING` — *humans may not exhibit this (they trust the write UI
|
||||
confirmation). If absent in human column, it's an agent-specific gap → flag.*
|
||||
**<<SYNTH>>**.
|
||||
|
||||
### semantic_retrieval — agent: 3/3 mismatch
|
||||
**Agent:** meaning known, indexed words unknown → BM25 keyword-stuffing. 4-step
|
||||
chain on one Go `bytes.Buffer` bug; NL questions shoved into `brain_query`.
|
||||
**Human:** `PENDING` — *humans likely hit this HARDER (they phrase conversationally).
|
||||
Compare reformulation-chain length agent vs human.*
|
||||
**<<SYNTH>>** — likely the strongest cross-consumer overlap.
|
||||
|
||||
### synthesized_answer — agent: 2 mismatch + 1 partial
|
||||
**Agent:** `brain_answer` not trusted terminal — 3 answers → 3 same-topic queries
|
||||
1min later; query+answer fired 8s apart hedging one need. Works for *episodic*
|
||||
recall, fails for *how-do-I / does-X-hold*.
|
||||
**Human:** `PENDING` — *humans may prefer `brain_answer` as primary (chat-native).
|
||||
If human match-rate >> agent, the tool fits humans not agents → key divergence.*
|
||||
**<<SYNTH>>**.
|
||||
|
||||
### store_new_knowledge — agent: 14/15 match
|
||||
**Agent:** healthy, except 1 HTTP-staged bodge when MCP auth lapsed. Also surfaced
|
||||
write-schema confusion: 3 param shapes (`{filename,type}` / `{filename}` /
|
||||
`{wing,hall,filename}`) + `ingest{source}`.
|
||||
**Human:** `PENDING` — *humans rarely write directly; expect low volume.*
|
||||
**<<SYNTH>>**.
|
||||
|
||||
### check_prior_art / ingest_raw_source — agent: 0% mismatch
|
||||
Lexical fits named-entity recall and raw-source capture. **Human:** `PENDING`.
|
||||
|
||||
### discover_capability — agent: 2/2 mismatch (agent-specific)
|
||||
Brain tools deferred → `ToolSearch`-load each session; auth lapse → `curl` bodge.
|
||||
**Likely has NO human analog** (humans get ambient connectors). Candidate for
|
||||
"agent-only gap" bucket. **Human:** `PENDING` to confirm absent.
|
||||
|
||||
---
|
||||
|
||||
## Cross-consumer divergence — questions to resolve at merge
|
||||
|
||||
1. **intent_unclear rate.** Agent = 0% (transcripts self-document). Human expected
|
||||
higher (conversational, implicit). Big delta = the columns measure legibility
|
||||
differently, not just intent.
|
||||
2. **Where does each consumer's mismatch concentrate?** Agent mismatch is
|
||||
write-side-heavy (supersede + verify-landed = 9/17). Hypothesis: human mismatch
|
||||
is read-side-heavy (semantic + answer). If true → **the interface fails the two
|
||||
consumers at opposite ends.**
|
||||
3. **Agent-only gaps** (`discover_capability`, `verify_write_landed`) vs
|
||||
**shared gaps** (`semantic_retrieval`, `update_or_supersede`). Shared gaps =
|
||||
highest-priority evidence; agent-only = harness/auth issues.
|
||||
|
||||
---
|
||||
|
||||
## Combined headline — `<<SYNTH>>` (fill when human column lands)
|
||||
|
||||
> Agent-side draft (to be reconciled with human-side):
|
||||
> Brain = append + keyword-search; agents want a curated, dedup'd, self-verifying KB.
|
||||
> Missing update/supersede path + lexical-only reads are the seam. **Open question
|
||||
> for the merge: do humans hit the same read-side wall, making semantic-retrieval the
|
||||
> universal gap — or do agents uniquely suffer the write-side (supersede / verify)
|
||||
> wall that humans sidestep via the chat UI?**
|
||||
|
||||
---
|
||||
|
||||
## Drop-in checklist (when human column arrives)
|
||||
1. Place `human-intent-column.jsonl` in this dir; conform to LOCKED schema.
|
||||
2. Fill every `PENDING` cell in master table + per-intent blocks.
|
||||
3. Resolve the 3 divergence questions with evidence.
|
||||
4. Replace each `<<SYNTH>>` with the reconciled verdict; write the combined headline.
|
||||
5. If canonical vocab surfaced: re-map both columns' `intent`, flip `schema_source`.
|
||||
6. Commit as `docs(brain): merge human+agent intent columns`.
|
||||
@@ -0,0 +1,97 @@
|
||||
package api
|
||||
|
||||
import "strings"
|
||||
|
||||
// frontmatter is an ordered, line-preserving view of a note's YAML
|
||||
// frontmatter block. It deliberately avoids a full YAML round-trip: the
|
||||
// brain writes flat `key: value` frontmatter by hand, and a yaml.v3
|
||||
// re-marshal would reorder keys and strip comments. Preserving the
|
||||
// original lines verbatim keeps brain_update a surgical edit — only the
|
||||
// keys it manages (updated_at, supersedes, supersede_reason) change.
|
||||
type frontmatter struct {
|
||||
lines []fmLine
|
||||
}
|
||||
|
||||
// fmLine is one frontmatter line. For `key: value` lines, key and value
|
||||
// are populated; for blank lines, comments, or anything that isn't a
|
||||
// simple scalar pair, key is empty and raw holds the line verbatim.
|
||||
type fmLine struct {
|
||||
key string
|
||||
value string
|
||||
raw string
|
||||
}
|
||||
|
||||
// parseFrontmatter splits src into its frontmatter block and body. A
|
||||
// frontmatter block is recognised only when the file opens with a `---`
|
||||
// fence and a closing `---` fence follows. Otherwise the whole input is
|
||||
// the body and the returned frontmatter is empty.
|
||||
func parseFrontmatter(src string) (frontmatter, string) {
|
||||
var fm frontmatter
|
||||
if !strings.HasPrefix(src, "---\n") {
|
||||
return fm, src
|
||||
}
|
||||
rest := src[len("---\n"):]
|
||||
end := strings.Index(rest, "\n---\n")
|
||||
if end < 0 {
|
||||
// Opening fence with no closing fence — treat as bodyless content.
|
||||
return fm, src
|
||||
}
|
||||
block := rest[:end]
|
||||
body := rest[end+len("\n---\n"):]
|
||||
|
||||
for _, line := range strings.Split(block, "\n") {
|
||||
key, val, ok := strings.Cut(line, ":")
|
||||
key = strings.TrimSpace(key)
|
||||
if !ok || key == "" || strings.HasPrefix(strings.TrimSpace(line), "#") {
|
||||
fm.lines = append(fm.lines, fmLine{raw: line})
|
||||
continue
|
||||
}
|
||||
fm.lines = append(fm.lines, fmLine{key: key, value: strings.TrimSpace(val)})
|
||||
}
|
||||
return fm, body
|
||||
}
|
||||
|
||||
// get returns the value for key, or "" if absent.
|
||||
func (f *frontmatter) get(key string) string {
|
||||
for _, l := range f.lines {
|
||||
if l.key == key {
|
||||
return l.value
|
||||
}
|
||||
}
|
||||
return ""
|
||||
}
|
||||
|
||||
// set overrides the value for an existing key in place, or appends a new
|
||||
// `key: value` line when the key is absent.
|
||||
func (f *frontmatter) set(key, value string) {
|
||||
for i := range f.lines {
|
||||
if f.lines[i].key == key {
|
||||
f.lines[i].value = value
|
||||
return
|
||||
}
|
||||
}
|
||||
f.lines = append(f.lines, fmLine{key: key, value: value})
|
||||
}
|
||||
|
||||
// render serialises the frontmatter back into a `---`-fenced block. An
|
||||
// empty frontmatter renders to the empty string so bodies without a
|
||||
// header stay header-less.
|
||||
func (f *frontmatter) render() string {
|
||||
if len(f.lines) == 0 {
|
||||
return ""
|
||||
}
|
||||
var b strings.Builder
|
||||
b.WriteString("---\n")
|
||||
for _, l := range f.lines {
|
||||
if l.key == "" {
|
||||
b.WriteString(l.raw)
|
||||
} else {
|
||||
b.WriteString(l.key)
|
||||
b.WriteString(": ")
|
||||
b.WriteString(l.value)
|
||||
}
|
||||
b.WriteByte('\n')
|
||||
}
|
||||
b.WriteString("---\n")
|
||||
return b.String()
|
||||
}
|
||||
@@ -0,0 +1,61 @@
|
||||
package api
|
||||
|
||||
import (
|
||||
"strings"
|
||||
"testing"
|
||||
|
||||
"github.com/stretchr/testify/assert"
|
||||
)
|
||||
|
||||
func TestParseFrontmatterSplitsHeaderAndBody(t *testing.T) {
|
||||
src := "---\nwing: jepa-fx\nhall: facts\ncreated_at: 2026-01-01T00:00:00Z\n---\n# Title\n\nbody text\n"
|
||||
fm, body := parseFrontmatter(src)
|
||||
|
||||
assert.Equal(t, "jepa-fx", fm.get("wing"))
|
||||
assert.Equal(t, "facts", fm.get("hall"))
|
||||
assert.Equal(t, "2026-01-01T00:00:00Z", fm.get("created_at"))
|
||||
assert.Equal(t, "# Title\n\nbody text\n", body)
|
||||
}
|
||||
|
||||
func TestParseFrontmatterNoHeader(t *testing.T) {
|
||||
src := "# Just a body\n\nno frontmatter here\n"
|
||||
fm, body := parseFrontmatter(src)
|
||||
|
||||
assert.Empty(t, fm.lines)
|
||||
assert.Equal(t, src, body)
|
||||
}
|
||||
|
||||
func TestFrontmatterSetOverridesExistingKey(t *testing.T) {
|
||||
fm, _ := parseFrontmatter("---\nwing: a\nupdated_at: old\n---\nbody\n")
|
||||
fm.set("updated_at", "new")
|
||||
|
||||
assert.Equal(t, "new", fm.get("updated_at"))
|
||||
// No duplicate key.
|
||||
assert.Equal(t, 1, strings.Count(fm.render(), "updated_at:"))
|
||||
}
|
||||
|
||||
func TestFrontmatterSetAppendsNewKey(t *testing.T) {
|
||||
fm, _ := parseFrontmatter("---\nwing: a\n---\nbody\n")
|
||||
fm.set("supersedes", "abc123")
|
||||
|
||||
out := fm.render()
|
||||
assert.Contains(t, out, "wing: a")
|
||||
assert.Contains(t, out, "supersedes: abc123")
|
||||
}
|
||||
|
||||
func TestFrontmatterRenderPreservesCustomFields(t *testing.T) {
|
||||
src := "---\nwing: a\nhall: facts\ncustom_field: keep-me\ntags: [x, y]\n---\nbody\n"
|
||||
fm, _ := parseFrontmatter(src)
|
||||
fm.set("updated_at", "2026-06-22T00:00:00Z")
|
||||
|
||||
out := fm.render()
|
||||
assert.Contains(t, out, "custom_field: keep-me")
|
||||
assert.Contains(t, out, "tags: [x, y]")
|
||||
assert.Contains(t, out, "updated_at: 2026-06-22T00:00:00Z")
|
||||
}
|
||||
|
||||
func TestFrontmatterRenderRoundTrips(t *testing.T) {
|
||||
src := "---\nwing: a\nhall: facts\n---\n"
|
||||
fm, _ := parseFrontmatter(src)
|
||||
assert.Equal(t, src, fm.render())
|
||||
}
|
||||
@@ -0,0 +1,131 @@
|
||||
package api
|
||||
|
||||
import (
|
||||
"crypto/sha256"
|
||||
"encoding/hex"
|
||||
"fmt"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"time"
|
||||
|
||||
"github.com/mathiasbq/hyperguild/ingestion/internal/brain"
|
||||
)
|
||||
|
||||
// ContentHash returns the lowercase hex sha256 of b. It is the note's
|
||||
// content_hash handle: brain_write / brain_update return it, brain_get
|
||||
// recomputes it from the file on disk, and brain_update stamps the prior
|
||||
// note's hash into the new note's `supersedes` frontmatter.
|
||||
func ContentHash(b []byte) string {
|
||||
sum := sha256.Sum256(b)
|
||||
return hex.EncodeToString(sum[:])
|
||||
}
|
||||
|
||||
// resolveWithin maps a brainDir-relative path to an absolute path and
|
||||
// guarantees it does not escape brainDir. Returns the cleaned relPath
|
||||
// (forward-slashed) and the absolute path.
|
||||
func resolveWithin(brainDir, relPath string) (rel, abs string, err error) {
|
||||
clean := filepath.Clean("/" + filepath.ToSlash(relPath))
|
||||
rel = strings.TrimPrefix(clean, "/")
|
||||
abs = filepath.Join(brainDir, filepath.FromSlash(rel))
|
||||
check, err := filepath.Rel(brainDir, abs)
|
||||
if err != nil || check == ".." || strings.HasPrefix(check, ".."+string(filepath.Separator)) {
|
||||
return "", "", fmt.Errorf("path %q escapes brain dir", relPath)
|
||||
}
|
||||
return rel, abs, nil
|
||||
}
|
||||
|
||||
// UpdateNoteOptions identifies the note to supersede and supplies its new
|
||||
// body. Path takes precedence; otherwise the target is resolved from
|
||||
// Wing/Hall/Slug via brain.NotePath.
|
||||
type UpdateNoteOptions struct {
|
||||
Path string // brainDir-relative path; takes precedence over wing/hall/slug
|
||||
Wing string
|
||||
Hall string
|
||||
Slug string
|
||||
Content string // new full body (whole-note replace)
|
||||
Reason string // optional; stamped as supersede_reason
|
||||
}
|
||||
|
||||
// UpdateNote supersedes an existing note in place. It replaces the body
|
||||
// with opts.Content, preserves the existing frontmatter (created_at,
|
||||
// wing, hall, and any custom fields), and stamps updated_at, supersedes
|
||||
// (the prior content hash), and supersede_reason (when given).
|
||||
//
|
||||
// It never creates: if the target does not exist, it returns an error so
|
||||
// the caller can fall back to brain_write. Returns the note's relPath,
|
||||
// the new content hash, and the prior content hash.
|
||||
//
|
||||
// Embeddings are NOT refreshed here. The rewritten file's mtime advances,
|
||||
// which the mtime-driven vectorstore.Sync ticker uses to re-embed it on
|
||||
// its next pass — the same out-of-band mechanism brain_write relies on.
|
||||
func UpdateNote(brainDir string, opts UpdateNoteOptions) (relPath, contentHash, priorHash string, err error) {
|
||||
if opts.Content == "" {
|
||||
return "", "", "", fmt.Errorf("content is required")
|
||||
}
|
||||
|
||||
var rel string
|
||||
if opts.Path != "" {
|
||||
rel = opts.Path
|
||||
} else {
|
||||
full, perr := brain.NotePath(brainDir, opts.Wing, opts.Hall, opts.Slug)
|
||||
if perr != nil {
|
||||
return "", "", "", perr
|
||||
}
|
||||
rel, _ = filepath.Rel(brainDir, full)
|
||||
rel = filepath.ToSlash(rel)
|
||||
}
|
||||
|
||||
rel, abs, err := resolveWithin(brainDir, rel)
|
||||
if err != nil {
|
||||
return "", "", "", err
|
||||
}
|
||||
|
||||
prior, err := os.ReadFile(abs)
|
||||
if err != nil {
|
||||
if os.IsNotExist(err) {
|
||||
return "", "", "", fmt.Errorf("note %q does not exist: use brain_write to create", rel)
|
||||
}
|
||||
return "", "", "", fmt.Errorf("read target: %w", err)
|
||||
}
|
||||
priorHash = ContentHash(prior)
|
||||
|
||||
fm, _ := parseFrontmatter(string(prior))
|
||||
fm.set("updated_at", time.Now().UTC().Format(time.RFC3339))
|
||||
fm.set("supersedes", priorHash)
|
||||
if opts.Reason != "" {
|
||||
fm.set("supersede_reason", opts.Reason)
|
||||
}
|
||||
|
||||
out := []byte(fm.render() + opts.Content)
|
||||
if err := os.WriteFile(abs, out, 0o644); err != nil {
|
||||
return "", "", "", fmt.Errorf("write: %w", err)
|
||||
}
|
||||
return rel, ContentHash(out), priorHash, nil
|
||||
}
|
||||
|
||||
// ReadNote reads the note at the brainDir-relative relPath and returns
|
||||
// its parsed frontmatter, body, and content hash. It is the read-after-
|
||||
// write primitive behind brain_get: the hash it returns equals the hash
|
||||
// brain_write / brain_update returned for the same bytes.
|
||||
func ReadNote(brainDir, relPath string) (fm map[string]string, body, contentHash string, err error) {
|
||||
_, abs, err := resolveWithin(brainDir, relPath)
|
||||
if err != nil {
|
||||
return nil, "", "", err
|
||||
}
|
||||
raw, err := os.ReadFile(abs)
|
||||
if err != nil {
|
||||
if os.IsNotExist(err) {
|
||||
return nil, "", "", fmt.Errorf("note %q does not exist", relPath)
|
||||
}
|
||||
return nil, "", "", fmt.Errorf("read note: %w", err)
|
||||
}
|
||||
parsed, body := parseFrontmatter(string(raw))
|
||||
fm = make(map[string]string, len(parsed.lines))
|
||||
for _, l := range parsed.lines {
|
||||
if l.key != "" {
|
||||
fm[l.key] = l.value
|
||||
}
|
||||
}
|
||||
return fm, body, ContentHash(raw), nil
|
||||
}
|
||||
@@ -0,0 +1,130 @@
|
||||
package api
|
||||
|
||||
import (
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"testing"
|
||||
|
||||
"github.com/stretchr/testify/assert"
|
||||
"github.com/stretchr/testify/require"
|
||||
)
|
||||
|
||||
// seedNote writes a note directly to disk and returns its relPath.
|
||||
func seedNote(t *testing.T, brainDir, rel, content string) string {
|
||||
t.Helper()
|
||||
full := filepath.Join(brainDir, filepath.FromSlash(rel))
|
||||
require.NoError(t, os.MkdirAll(filepath.Dir(full), 0o755))
|
||||
require.NoError(t, os.WriteFile(full, []byte(content), 0o644))
|
||||
return rel
|
||||
}
|
||||
|
||||
func TestUpdateNoteSupersedesAndStamps(t *testing.T) {
|
||||
brainDir := t.TempDir()
|
||||
rel := seedNote(t, brainDir, "wiki/jepa-fx/facts/val-vol.md",
|
||||
"---\nwing: jepa-fx\nhall: facts\ncreated_at: 2026-01-01T00:00:00Z\ncustom: keep-me\n---\n# Old\n\nold body\n")
|
||||
|
||||
relPath, hash, priorHash, err := UpdateNote(brainDir, UpdateNoteOptions{
|
||||
Path: rel,
|
||||
Content: "# New\n\nnew body\n",
|
||||
Reason: "facts changed",
|
||||
})
|
||||
require.NoError(t, err)
|
||||
assert.Equal(t, rel, relPath)
|
||||
assert.NotEmpty(t, hash)
|
||||
assert.NotEmpty(t, priorHash)
|
||||
assert.NotEqual(t, hash, priorHash)
|
||||
|
||||
got, err := os.ReadFile(filepath.Join(brainDir, filepath.FromSlash(rel)))
|
||||
require.NoError(t, err)
|
||||
s := string(got)
|
||||
// Body replaced.
|
||||
assert.Contains(t, s, "# New")
|
||||
assert.NotContains(t, s, "old body")
|
||||
// Prior fields preserved.
|
||||
assert.Contains(t, s, "wing: jepa-fx")
|
||||
assert.Contains(t, s, "hall: facts")
|
||||
assert.Contains(t, s, "created_at: 2026-01-01T00:00:00Z")
|
||||
assert.Contains(t, s, "custom: keep-me")
|
||||
// Supersession stamped.
|
||||
assert.Contains(t, s, "updated_at:")
|
||||
assert.Contains(t, s, "supersedes: "+priorHash)
|
||||
assert.Contains(t, s, "supersede_reason: facts changed")
|
||||
}
|
||||
|
||||
func TestUpdateNoteResolvesByWingHallSlug(t *testing.T) {
|
||||
brainDir := t.TempDir()
|
||||
seedNote(t, brainDir, "wiki/jepa-fx/facts/val-vol.md",
|
||||
"---\nwing: jepa-fx\nhall: facts\n---\nold\n")
|
||||
|
||||
relPath, _, _, err := UpdateNote(brainDir, UpdateNoteOptions{
|
||||
Wing: "jepa-fx", Hall: "facts", Slug: "val-vol",
|
||||
Content: "new\n",
|
||||
})
|
||||
require.NoError(t, err)
|
||||
assert.Equal(t, "wiki/jepa-fx/facts/val-vol.md", relPath)
|
||||
}
|
||||
|
||||
func TestUpdateNoteErrorsOnMissingAndDoesNotCreate(t *testing.T) {
|
||||
brainDir := t.TempDir()
|
||||
|
||||
_, _, _, err := UpdateNote(brainDir, UpdateNoteOptions{
|
||||
Wing: "jepa-fx", Hall: "facts", Slug: "ghost",
|
||||
Content: "x\n",
|
||||
})
|
||||
require.Error(t, err)
|
||||
assert.Contains(t, err.Error(), "does not exist")
|
||||
|
||||
// No file created.
|
||||
_, statErr := os.Stat(filepath.Join(brainDir, "wiki/jepa-fx/facts/ghost.md"))
|
||||
assert.True(t, os.IsNotExist(statErr), "missing-target update must not create a note")
|
||||
}
|
||||
|
||||
func TestUpdateNoteRejectsTraversal(t *testing.T) {
|
||||
brainDir := t.TempDir()
|
||||
_, _, _, err := UpdateNote(brainDir, UpdateNoteOptions{
|
||||
Path: "../escape.md",
|
||||
Content: "x\n",
|
||||
})
|
||||
require.Error(t, err)
|
||||
}
|
||||
|
||||
func TestReadNoteReturnsFrontmatterBodyHash(t *testing.T) {
|
||||
brainDir := t.TempDir()
|
||||
rel := seedNote(t, brainDir, "wiki/jepa-fx/facts/n.md",
|
||||
"---\nwing: jepa-fx\nhall: facts\n---\n# Body\n\ntext\n")
|
||||
|
||||
fm, body, hash, err := ReadNote(brainDir, rel)
|
||||
require.NoError(t, err)
|
||||
assert.Equal(t, "jepa-fx", fm["wing"])
|
||||
assert.Equal(t, "facts", fm["hall"])
|
||||
assert.Equal(t, "# Body\n\ntext\n", body)
|
||||
|
||||
// Hash matches ContentHash of the raw bytes on disk (round-trip).
|
||||
raw, _ := os.ReadFile(filepath.Join(brainDir, filepath.FromSlash(rel)))
|
||||
assert.Equal(t, ContentHash(raw), hash)
|
||||
}
|
||||
|
||||
func TestReadNoteRejectsTraversal(t *testing.T) {
|
||||
brainDir := t.TempDir()
|
||||
_, _, _, err := ReadNote(brainDir, "../../etc/passwd")
|
||||
require.Error(t, err)
|
||||
}
|
||||
|
||||
func TestUpdateThenReadRoundTripsHash(t *testing.T) {
|
||||
brainDir := t.TempDir()
|
||||
rel := seedNote(t, brainDir, "wiki/a/facts/n.md", "---\nwing: a\nhall: facts\n---\nold\n")
|
||||
|
||||
_, hash, _, err := UpdateNote(brainDir, UpdateNoteOptions{Path: rel, Content: "new\n"})
|
||||
require.NoError(t, err)
|
||||
|
||||
_, _, readHash, err := ReadNote(brainDir, rel)
|
||||
require.NoError(t, err)
|
||||
assert.Equal(t, hash, readHash, "update content_hash must round-trip through ReadNote")
|
||||
}
|
||||
|
||||
func TestContentHashStable(t *testing.T) {
|
||||
assert.Equal(t, ContentHash([]byte("abc")), ContentHash([]byte("abc")))
|
||||
assert.NotEqual(t, ContentHash([]byte("abc")), ContentHash([]byte("abd")))
|
||||
assert.True(t, strings.HasPrefix(ContentHash([]byte("")), "")) // hex, non-panicking
|
||||
}
|
||||
@@ -0,0 +1,189 @@
|
||||
// Package classification defines the data-sensitivity taxonomy and the
|
||||
// per-wing / per-repo tagging the capture server reads to enforce the I1
|
||||
// sovereignty gate (issue #50, capture spec §4.1).
|
||||
//
|
||||
// The single load-bearing property is fail-safe-to-strictest: a target
|
||||
// with no explicit tag and no known default classifies as Confidential,
|
||||
// never as something more permissive. A missing tag must never silently
|
||||
// downgrade — that would turn the I1 gate into theatre.
|
||||
//
|
||||
// Classification is read from an optional classification.yaml at the
|
||||
// brain root. A central, Flux-reconcilable file is deliberate: it is
|
||||
// auditable in one place (I2/I5), it does not require a live Gitea client
|
||||
// to classify a repo (so this package has no dependency on the gitea
|
||||
// tracker work), and it avoids tagging a wing's _index.md frontmatter —
|
||||
// which BuildWingIndex regenerates and would clobber.
|
||||
package classification
|
||||
|
||||
import (
|
||||
"fmt"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
|
||||
"gopkg.in/yaml.v3"
|
||||
)
|
||||
|
||||
// Level is a data-sensitivity tier. Higher is stricter, so the "stricter
|
||||
// wins" rule (spec §4.1 model C) is a plain max.
|
||||
type Level int
|
||||
|
||||
const (
|
||||
Public Level = iota
|
||||
Internal
|
||||
Confidential
|
||||
)
|
||||
|
||||
// String returns the canonical lowercase token for a level.
|
||||
func (l Level) String() string {
|
||||
switch l {
|
||||
case Public:
|
||||
return "public"
|
||||
case Internal:
|
||||
return "internal"
|
||||
case Confidential:
|
||||
return "confidential"
|
||||
default:
|
||||
return fmt.Sprintf("level(%d)", int(l))
|
||||
}
|
||||
}
|
||||
|
||||
// ParseLevel parses a level token (case-insensitive, surrounding space
|
||||
// tolerated). An unknown token is an error — callers must decide what to
|
||||
// do with bad input rather than have it silently coerced.
|
||||
func ParseLevel(s string) (Level, error) {
|
||||
switch strings.ToLower(strings.TrimSpace(s)) {
|
||||
case "public":
|
||||
return Public, nil
|
||||
case "internal":
|
||||
return Internal, nil
|
||||
case "confidential":
|
||||
return Confidential, nil
|
||||
default:
|
||||
return Confidential, fmt.Errorf("unknown classification level %q (want public/internal/confidential)", s)
|
||||
}
|
||||
}
|
||||
|
||||
// Stricter returns the more restrictive of two levels.
|
||||
func Stricter(a, b Level) Level {
|
||||
if a > b {
|
||||
return a
|
||||
}
|
||||
return b
|
||||
}
|
||||
|
||||
// TargetKind distinguishes the two kinds of capture destination.
|
||||
type TargetKind int
|
||||
|
||||
const (
|
||||
WingTarget TargetKind = iota // a brain wing (insights land here)
|
||||
RepoTarget // a Gitea repo (tickets / summaries land here)
|
||||
)
|
||||
|
||||
// Target names a capture destination to classify.
|
||||
type Target struct {
|
||||
Kind TargetKind
|
||||
Name string
|
||||
}
|
||||
|
||||
// Config holds the explicit per-wing / per-repo classification tags read
|
||||
// from classification.yaml. Absent entries fall through to the built-in
|
||||
// defaults in defaultFor. The zero value (no file) is valid and applies
|
||||
// defaults to everything.
|
||||
type Config struct {
|
||||
wings map[string]Level
|
||||
repos map[string]Level
|
||||
}
|
||||
|
||||
// rawConfig is the on-disk YAML shape: string→string maps, parsed into
|
||||
// validated levels by Load.
|
||||
type rawConfig struct {
|
||||
Wings map[string]string `yaml:"wings"`
|
||||
Repos map[string]string `yaml:"repos"`
|
||||
}
|
||||
|
||||
// Load reads classification.yaml from brainDir. An absent file is not an
|
||||
// error — it yields an empty config where every target classifies by the
|
||||
// built-in defaults. A malformed file, or any unparseable level token in
|
||||
// it, is a hard error: a classification source the server cannot trust
|
||||
// must fail loud, not degrade silently.
|
||||
func Load(brainDir string) (*Config, error) {
|
||||
cfg := &Config{wings: map[string]Level{}, repos: map[string]Level{}}
|
||||
|
||||
data, err := os.ReadFile(filepath.Join(brainDir, "classification.yaml"))
|
||||
if err != nil {
|
||||
if os.IsNotExist(err) {
|
||||
return cfg, nil
|
||||
}
|
||||
return nil, fmt.Errorf("read classification.yaml: %w", err)
|
||||
}
|
||||
|
||||
var raw rawConfig
|
||||
if err := yaml.Unmarshal(data, &raw); err != nil {
|
||||
return nil, fmt.Errorf("parse classification.yaml: %w", err)
|
||||
}
|
||||
for name, lvl := range raw.Wings {
|
||||
parsed, perr := ParseLevel(lvl)
|
||||
if perr != nil {
|
||||
return nil, fmt.Errorf("wing %q: %w", name, perr)
|
||||
}
|
||||
cfg.wings[normalise(name)] = parsed
|
||||
}
|
||||
for name, lvl := range raw.Repos {
|
||||
parsed, perr := ParseLevel(lvl)
|
||||
if perr != nil {
|
||||
return nil, fmt.Errorf("repo %q: %w", name, perr)
|
||||
}
|
||||
cfg.repos[normalise(name)] = parsed
|
||||
}
|
||||
return cfg, nil
|
||||
}
|
||||
|
||||
// Derive returns the classification for any target — the function the
|
||||
// capture use-case calls per item.
|
||||
func (c *Config) Derive(t Target) Level {
|
||||
if t.Kind == RepoTarget {
|
||||
return c.Repo(t.Name)
|
||||
}
|
||||
return c.Wing(t.Name)
|
||||
}
|
||||
|
||||
// Wing classifies a brain wing: an explicit tag wins, else defaults.
|
||||
func (c *Config) Wing(name string) Level {
|
||||
if lvl, ok := c.wings[normalise(name)]; ok {
|
||||
return lvl
|
||||
}
|
||||
return defaultFor(name)
|
||||
}
|
||||
|
||||
// Repo classifies a Gitea repo: an explicit tag wins, else defaults.
|
||||
func (c *Config) Repo(name string) Level {
|
||||
if lvl, ok := c.repos[normalise(name)]; ok {
|
||||
return lvl
|
||||
}
|
||||
return defaultFor(name)
|
||||
}
|
||||
|
||||
// defaultFor applies the built-in defaulting rules when a target has no
|
||||
// explicit tag:
|
||||
// - client-* → Confidential (client work is confidential by default)
|
||||
// - hyperguild / homelab → Internal (the operator's own infra)
|
||||
// - everything else → Confidential (fail safe to strictest)
|
||||
func defaultFor(name string) Level {
|
||||
n := normalise(name)
|
||||
if strings.HasPrefix(n, "client-") {
|
||||
return Confidential
|
||||
}
|
||||
switch n {
|
||||
case "hyperguild", "homelab":
|
||||
return Internal
|
||||
default:
|
||||
return Confidential
|
||||
}
|
||||
}
|
||||
|
||||
// normalise lowercases and trims a wing/repo name so matching and the
|
||||
// client-* prefix check are case-insensitive.
|
||||
func normalise(name string) string {
|
||||
return strings.ToLower(strings.TrimSpace(name))
|
||||
}
|
||||
@@ -0,0 +1,112 @@
|
||||
package classification
|
||||
|
||||
import (
|
||||
"os"
|
||||
"path/filepath"
|
||||
"testing"
|
||||
|
||||
"github.com/stretchr/testify/assert"
|
||||
"github.com/stretchr/testify/require"
|
||||
)
|
||||
|
||||
func TestLevelOrderingAndString(t *testing.T) {
|
||||
assert.True(t, Public < Internal)
|
||||
assert.True(t, Internal < Confidential)
|
||||
assert.Equal(t, "public", Public.String())
|
||||
assert.Equal(t, "internal", Internal.String())
|
||||
assert.Equal(t, "confidential", Confidential.String())
|
||||
}
|
||||
|
||||
func TestParseLevel(t *testing.T) {
|
||||
for s, want := range map[string]Level{
|
||||
"public": Public, "internal": Internal, "confidential": Confidential,
|
||||
"PUBLIC": Public, " Confidential ": Confidential,
|
||||
} {
|
||||
got, err := ParseLevel(s)
|
||||
require.NoError(t, err, s)
|
||||
assert.Equal(t, want, got, s)
|
||||
}
|
||||
_, err := ParseLevel("secret")
|
||||
require.Error(t, err, "unknown level must error, not silently default")
|
||||
_, err = ParseLevel("")
|
||||
require.Error(t, err)
|
||||
}
|
||||
|
||||
func TestStricterReturnsMax(t *testing.T) {
|
||||
assert.Equal(t, Confidential, Stricter(Internal, Confidential))
|
||||
assert.Equal(t, Confidential, Stricter(Confidential, Public))
|
||||
assert.Equal(t, Internal, Stricter(Public, Internal))
|
||||
assert.Equal(t, Public, Stricter(Public, Public))
|
||||
}
|
||||
|
||||
func TestLoadAbsentFileIsDefaultsOnly(t *testing.T) {
|
||||
cfg, err := Load(t.TempDir())
|
||||
require.NoError(t, err, "absent classification.yaml must not be an error — defaults apply")
|
||||
require.NotNil(t, cfg)
|
||||
// Pure defaulting still works.
|
||||
assert.Equal(t, Internal, cfg.Wing("hyperguild"))
|
||||
assert.Equal(t, Confidential, cfg.Wing("anything-unknown"))
|
||||
}
|
||||
|
||||
func TestLoadParsesExplicitTags(t *testing.T) {
|
||||
dir := t.TempDir()
|
||||
require.NoError(t, os.WriteFile(filepath.Join(dir, "classification.yaml"), []byte(
|
||||
"wings:\n research-public: public\n hyperguild: confidential\nrepos:\n infra: internal\n research-public: public\n",
|
||||
), 0o644))
|
||||
|
||||
cfg, err := Load(dir)
|
||||
require.NoError(t, err)
|
||||
// Explicit tag wins over the built-in default (hyperguild default is internal).
|
||||
assert.Equal(t, Confidential, cfg.Wing("hyperguild"))
|
||||
// Explicit public is honoured.
|
||||
assert.Equal(t, Public, cfg.Wing("research-public"))
|
||||
assert.Equal(t, Internal, cfg.Repo("infra"))
|
||||
assert.Equal(t, Public, cfg.Repo("research-public"))
|
||||
}
|
||||
|
||||
func TestLoadRejectsUnknownLevelInFile(t *testing.T) {
|
||||
dir := t.TempDir()
|
||||
require.NoError(t, os.WriteFile(filepath.Join(dir, "classification.yaml"),
|
||||
[]byte("wings:\n x: top-secret\n"), 0o644))
|
||||
_, err := Load(dir)
|
||||
require.Error(t, err, "an unparseable level in the config must fail loud, not be ignored")
|
||||
}
|
||||
|
||||
func TestWingDefaulting(t *testing.T) {
|
||||
cfg, err := Load(t.TempDir())
|
||||
require.NoError(t, err)
|
||||
cases := map[string]Level{
|
||||
"client-seb": Confidential, // client-* → confidential
|
||||
"client-mastercard": Confidential,
|
||||
"hyperguild": Internal,
|
||||
"homelab": Internal,
|
||||
"jepa-fx": Confidential, // unknown → fail safe to strictest
|
||||
"": Confidential, // empty → fail safe
|
||||
}
|
||||
for wing, want := range cases {
|
||||
assert.Equal(t, want, cfg.Wing(wing), "wing %q", wing)
|
||||
}
|
||||
}
|
||||
|
||||
func TestRepoDefaulting(t *testing.T) {
|
||||
cfg, err := Load(t.TempDir())
|
||||
require.NoError(t, err)
|
||||
assert.Equal(t, Confidential, cfg.Repo("client-seb-pipeline"))
|
||||
assert.Equal(t, Internal, cfg.Repo("hyperguild"))
|
||||
assert.Equal(t, Confidential, cfg.Repo("some-unknown-repo"), "untagged repo → confidential (fail safe)")
|
||||
}
|
||||
|
||||
func TestDeriveUnifiedTarget(t *testing.T) {
|
||||
cfg, err := Load(t.TempDir())
|
||||
require.NoError(t, err)
|
||||
assert.Equal(t, Internal, cfg.Derive(Target{Kind: WingTarget, Name: "homelab"}))
|
||||
assert.Equal(t, Confidential, cfg.Derive(Target{Kind: RepoTarget, Name: "client-x"}))
|
||||
assert.Equal(t, Confidential, cfg.Derive(Target{Kind: WingTarget, Name: "untagged"}))
|
||||
}
|
||||
|
||||
func TestCaseInsensitiveMatching(t *testing.T) {
|
||||
cfg, err := Load(t.TempDir())
|
||||
require.NoError(t, err)
|
||||
assert.Equal(t, Confidential, cfg.Wing("Client-SEB"), "client- prefix match is case-insensitive")
|
||||
assert.Equal(t, Internal, cfg.Wing("HyperGuild"))
|
||||
}
|
||||
@@ -38,11 +38,23 @@ var DefaultRules = []Rule{
|
||||
// specific match name in logs.
|
||||
{Name: "authorization-header", RE: regexp.MustCompile(`(?i)Authorization\s*:\s*[A-Za-z]+\s+\S{8,}`)},
|
||||
{Name: "bearer-token", RE: regexp.MustCompile(`(?i)Bearer\s+[A-Za-z0-9._\-]{16,}`)},
|
||||
// JWT (header.payload.sig), e.g. a Dex/OAuth token dumped to stdout
|
||||
// without a "Bearer " prefix. Both header and payload base64url-encode
|
||||
// JSON, so both segments begin with "eyJ".
|
||||
{Name: "jwt", RE: regexp.MustCompile(`eyJ[A-Za-z0-9_\-]{8,}\.eyJ[A-Za-z0-9_\-]{8,}\.[A-Za-z0-9_\-]{8,}`)},
|
||||
{Name: "postgres-uri-with-password", RE: regexp.MustCompile(`postgres(?:ql)?://[^:\s/]+:[^@\s/]+@`)},
|
||||
{Name: "private-key", RE: regexp.MustCompile(`-----BEGIN[^-]*PRIVATE KEY-----`)},
|
||||
{Name: "ssh-key", RE: regexp.MustCompile(`ssh-(?:rsa|ed25519|ecdsa)\s+[A-Za-z0-9+/=]{40,}`)},
|
||||
{Name: "github-pat", RE: regexp.MustCompile(`\b(?:ghp|gho|ghu|ghr|gha)_[A-Za-z0-9]{30,}\b`)},
|
||||
{Name: "openai-sk", RE: regexp.MustCompile(`\bsk-(?:proj-)?[A-Za-z0-9]{32,}\b`)},
|
||||
// 1Password service-account token (ops_<base64url>). Long, high-value root
|
||||
// credential; guard the bare value (the _TOKEN= form also hits homelab-env-token).
|
||||
{Name: "op-service-account", RE: regexp.MustCompile(`\bops_[A-Za-z0-9_\-]{40,}`)},
|
||||
// No leading \b: a shell mangle can glue the key to a preceding word
|
||||
// ("yes"+"sk-...") which has no word boundary, and that exact case
|
||||
// leaked a LiteLLM master key past this rule (2026-06-11). Match the
|
||||
// sk- shape wherever it appears; the {32,} length floor keeps short
|
||||
// "task-"/"disk-" words from tripping it.
|
||||
{Name: "openai-sk", RE: regexp.MustCompile(`sk-(?:proj-)?[A-Za-z0-9]{32,}`)},
|
||||
{Name: "anthropic-sk", RE: regexp.MustCompile(`\bsk-ant-[A-Za-z0-9_\-]{32,}\b`)},
|
||||
{Name: "aws-access-key", RE: regexp.MustCompile(`\bAKIA[0-9A-Z]{16}\b`)},
|
||||
{Name: "homelab-env-token", RE: regexp.MustCompile(`(?i)(?:_TOKEN|_PASSWORD|_API_KEY|_SECRET)\s*[:=]\s*['"]?[A-Za-z0-9._/+\-]{12,}`)},
|
||||
|
||||
@@ -25,6 +25,18 @@ func TestScrub_PoisonedFixtures(t *testing.T) {
|
||||
{"aws-access-key", "AWS_ACCESS_KEY_ID=AKIAIOSFODNN7EXAMPLE", "aws-access-key"},
|
||||
{"homelab-env", "POSTGRES_PASSWORD=hunter2supersecretvalue", "homelab-env-token"},
|
||||
{"sops-marker", "value: ENC[AES256_GCM,data:abc123def456,iv:zzz]", "sops-encrypted-marker"},
|
||||
// Regression: a shell mangle glued the key to a preceding word
|
||||
// ("yes"+"sk-..."), defeating the leading \b in the sk- rule and
|
||||
// leaking a LiteLLM master key past the scrubber (2026-06-11).
|
||||
{"sk-glued-to-word", "master key resolved: yessk-7181ca984603239d8c4819361bf33b94b9c3c07018791868", "openai-sk"},
|
||||
{"sk-standalone-hex", "sk-7181ca984603239d8c4819361bf33b94b9c3c07018791868", "openai-sk"},
|
||||
// Bare JWT not preceded by "Bearer" (e.g. a Dex token dumped to stdout).
|
||||
{"jwt-bare", "token=eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9.eyJzdWIiOiIxMjM0NTY3ODkwIn0.dQw4w9WgXcQabcdef", "jwt"},
|
||||
// 1Password service-account token (ops_<base64url>), env-assigned and bare.
|
||||
// Both hit the dedicated op-service-account rule (ordered before the
|
||||
// generic homelab-env-token). Guards ~/.zshrc reads etc. (2026-06-14).
|
||||
{"op-sa-env", "export OP_SERVICE_ACCOUNT_TOKEN=ops_eyJzaWduSW5BZGRyZXNzIjoibXkuMXBhc3N3b3JkLmNvbSJ9", "op-service-account"},
|
||||
{"op-sa-bare", "ops_eyJzaWduSW5BZGRyZXNzIjoibXkuMXBhc3N3b3JkLmNvbSIsInVzZXJBdXRoIjp7fX0aGVsbG8", "op-service-account"},
|
||||
}
|
||||
for _, tc := range cases {
|
||||
t.Run(tc.name, func(t *testing.T) {
|
||||
@@ -43,6 +55,9 @@ func TestScrub_CleanContentPassesThrough(t *testing.T) {
|
||||
"file at ~/.ssh/id_ed25519",
|
||||
"the function Authorization() takes no args",
|
||||
"comment: see API key in 1Password",
|
||||
// loosened sk- rule must not trip on short "task-"/"disk-" words
|
||||
"run task-build then task-test in the pipeline",
|
||||
"mounted /dev/disk-by-id/wwn-0x5000",
|
||||
}
|
||||
for _, c := range cases {
|
||||
assert.Empty(t, Scrub(c), "expected clean for %q", c)
|
||||
|
||||
@@ -0,0 +1,204 @@
|
||||
package mcp_test
|
||||
|
||||
import (
|
||||
"context"
|
||||
"encoding/json"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"github.com/mathiasbq/hyperguild/ingestion/internal/mcp"
|
||||
"github.com/mathiasbq/hyperguild/ingestion/internal/vectorstore"
|
||||
"github.com/stretchr/testify/assert"
|
||||
"github.com/stretchr/testify/require"
|
||||
)
|
||||
|
||||
// callResult parses the JSON text payload of a successful tool call.
|
||||
func callResult(t *testing.T, resp map[string]any) map[string]any {
|
||||
t.Helper()
|
||||
require.Nil(t, resp["error"], "tool returned error: %v", resp["error"])
|
||||
text := resp["result"].(map[string]any)["content"].([]any)[0].(map[string]any)["text"].(string)
|
||||
var out map[string]any
|
||||
require.NoError(t, json.Unmarshal([]byte(text), &out))
|
||||
return out
|
||||
}
|
||||
|
||||
func TestBrainUpdateSupersedesExisting(t *testing.T) {
|
||||
brainDir := t.TempDir()
|
||||
srv := mcp.NewServer(brainDir, nil, nil, nil)
|
||||
|
||||
// Seed via brain_write so the note carries real frontmatter.
|
||||
callResult(t, toolCall(t, srv, "brain_write", map[string]any{
|
||||
"content": "# Old\n\nold body\n", "filename": "val-vol",
|
||||
"wing": "jepa-fx", "hall": "facts",
|
||||
}))
|
||||
|
||||
out := callResult(t, toolCall(t, srv, "brain_update", map[string]any{
|
||||
"wing": "jepa-fx", "hall": "facts", "slug": "val-vol",
|
||||
"content": "# New\n\nnew body\n", "reason": "facts changed",
|
||||
}))
|
||||
assert.Equal(t, "wiki/jepa-fx/facts/val-vol.md", out["path"])
|
||||
assert.Equal(t, out["path"], out["id"])
|
||||
assert.NotEmpty(t, out["content_hash"])
|
||||
assert.Equal(t, true, out["superseded"])
|
||||
|
||||
got, err := os.ReadFile(filepath.Join(brainDir, "wiki/jepa-fx/facts/val-vol.md"))
|
||||
require.NoError(t, err)
|
||||
s := string(got)
|
||||
assert.Contains(t, s, "# New")
|
||||
assert.NotContains(t, s, "old body")
|
||||
assert.Contains(t, s, "wing: jepa-fx")
|
||||
assert.Contains(t, s, "supersede_reason: facts changed")
|
||||
assert.Contains(t, s, "supersedes:")
|
||||
}
|
||||
|
||||
func TestBrainUpdateMissingTargetErrorsNoCreate(t *testing.T) {
|
||||
brainDir := t.TempDir()
|
||||
srv := mcp.NewServer(brainDir, nil, nil, nil)
|
||||
|
||||
resp := toolCall(t, srv, "brain_update", map[string]any{
|
||||
"wing": "jepa-fx", "hall": "facts", "slug": "ghost",
|
||||
"content": "x\n",
|
||||
})
|
||||
require.NotNil(t, resp["error"])
|
||||
assert.Contains(t, resp["error"].(map[string]any)["message"].(string), "does not exist")
|
||||
_, statErr := os.Stat(filepath.Join(brainDir, "wiki/jepa-fx/facts/ghost.md"))
|
||||
assert.True(t, os.IsNotExist(statErr))
|
||||
}
|
||||
|
||||
func TestBrainUpdateByFullPath(t *testing.T) {
|
||||
brainDir := t.TempDir()
|
||||
srv := mcp.NewServer(brainDir, nil, nil, nil)
|
||||
callResult(t, toolCall(t, srv, "brain_write", map[string]any{
|
||||
"content": "old\n", "filename": "n", "wing": "a", "hall": "facts",
|
||||
}))
|
||||
|
||||
out := callResult(t, toolCall(t, srv, "brain_update", map[string]any{
|
||||
"slug": "wiki/a/facts/n.md", "content": "fresh\n",
|
||||
}))
|
||||
assert.Equal(t, "wiki/a/facts/n.md", out["path"])
|
||||
}
|
||||
|
||||
func TestBrainGetByIDAndPath(t *testing.T) {
|
||||
brainDir := t.TempDir()
|
||||
srv := mcp.NewServer(brainDir, nil, nil, nil)
|
||||
w := callResult(t, toolCall(t, srv, "brain_write", map[string]any{
|
||||
"content": "# Body\n\ntext\n", "filename": "n", "wing": "a", "hall": "facts",
|
||||
}))
|
||||
id := w["id"].(string)
|
||||
hash := w["content_hash"].(string)
|
||||
require.NotEmpty(t, id)
|
||||
require.NotEmpty(t, hash)
|
||||
|
||||
// by id
|
||||
g1 := callResult(t, toolCall(t, srv, "brain_get", map[string]any{"id": id}))
|
||||
assert.Equal(t, id, g1["path"])
|
||||
assert.Equal(t, hash, g1["content_hash"], "content_hash must round-trip write→get")
|
||||
assert.Contains(t, g1["body"].(string), "# Body")
|
||||
fm := g1["frontmatter"].(map[string]any)
|
||||
assert.Equal(t, "a", fm["wing"])
|
||||
|
||||
// by path
|
||||
g2 := callResult(t, toolCall(t, srv, "brain_get", map[string]any{"path": id}))
|
||||
assert.Equal(t, hash, g2["content_hash"])
|
||||
}
|
||||
|
||||
func TestBrainGetMissingArgsErrors(t *testing.T) {
|
||||
srv := mcp.NewServer(t.TempDir(), nil, nil, nil)
|
||||
resp := toolCall(t, srv, "brain_get", map[string]any{})
|
||||
require.NotNil(t, resp["error"])
|
||||
}
|
||||
|
||||
func TestBrainWriteReturnsHandle(t *testing.T) {
|
||||
brainDir := t.TempDir()
|
||||
srv := mcp.NewServer(brainDir, nil, nil, nil)
|
||||
out := callResult(t, toolCall(t, srv, "brain_write", map[string]any{
|
||||
"content": "# X\n\nbody\n", "filename": "x", "wing": "a", "hall": "facts",
|
||||
}))
|
||||
assert.Equal(t, "wiki/a/facts/x.md", out["path"])
|
||||
assert.Equal(t, out["path"], out["id"])
|
||||
assert.NotEmpty(t, out["content_hash"])
|
||||
}
|
||||
|
||||
// --- retrieval-reflects-new-content: exercises the real mtime-driven Sync ---
|
||||
|
||||
type fakeVecStore struct {
|
||||
chunks map[string][]float32
|
||||
deleted []string
|
||||
}
|
||||
|
||||
func (f *fakeVecStore) KnownPathsWithTime(_ context.Context) (map[string]time.Time, error) {
|
||||
m := make(map[string]time.Time, len(f.chunks))
|
||||
for p := range f.chunks {
|
||||
m[p] = time.Unix(0, 0) // always stale → mtime(now) is always newer
|
||||
}
|
||||
return m, nil
|
||||
}
|
||||
|
||||
func (f *fakeVecStore) Upsert(_ context.Context, path string, vec []float32) error {
|
||||
f.chunks[path] = vec
|
||||
return nil
|
||||
}
|
||||
|
||||
func (f *fakeVecStore) Delete(_ context.Context, path string) error {
|
||||
delete(f.chunks, path)
|
||||
f.deleted = append(f.deleted, path)
|
||||
return nil
|
||||
}
|
||||
|
||||
type fakeEmbedder struct{ seen []string }
|
||||
|
||||
func (e *fakeEmbedder) Embed(_ context.Context, text string) ([]float32, error) {
|
||||
e.seen = append(e.seen, text)
|
||||
return []float32{1, 0, 0}, nil
|
||||
}
|
||||
|
||||
// TestBrainUpdateReembedsNewContent proves the supersede contract end to
|
||||
// end against the actual embedding mechanism: brain_update rewrites the
|
||||
// file, advancing its mtime, and the next vectorstore.Sync pass re-embeds
|
||||
// the NEW body and drops the stale chunk. No stub of the re-index path.
|
||||
func TestBrainUpdateReembedsNewContent(t *testing.T) {
|
||||
brainDir := t.TempDir()
|
||||
srv := mcp.NewServer(brainDir, nil, nil, nil)
|
||||
ctx := context.Background()
|
||||
|
||||
callResult(t, toolCall(t, srv, "brain_write", map[string]any{
|
||||
"content": "# Note\n\nthe OLD distinctive payload\n",
|
||||
"filename": "n", "wing": "a", "hall": "facts",
|
||||
}))
|
||||
|
||||
store := &fakeVecStore{chunks: map[string][]float32{}}
|
||||
emb := &fakeEmbedder{}
|
||||
|
||||
// First sync embeds the original content.
|
||||
_, err := vectorstore.Sync(ctx, brainDir, store, emb)
|
||||
require.NoError(t, err)
|
||||
require.NotEmpty(t, store.chunks)
|
||||
require.True(t, anyContains(emb.seen, "OLD distinctive payload"))
|
||||
|
||||
callResult(t, toolCall(t, srv, "brain_update", map[string]any{
|
||||
"wing": "a", "hall": "facts", "slug": "n",
|
||||
"content": "# Note\n\nthe NEW distinctive payload\n",
|
||||
}))
|
||||
|
||||
emb.seen = nil // only watch what the second pass embeds
|
||||
_, err = vectorstore.Sync(ctx, brainDir, store, emb)
|
||||
require.NoError(t, err)
|
||||
|
||||
assert.True(t, anyContains(emb.seen, "NEW distinctive payload"),
|
||||
"Sync must re-embed the superseded body; saw %v", emb.seen)
|
||||
assert.False(t, anyContains(emb.seen, "OLD distinctive payload"),
|
||||
"the old body must not be re-embedded")
|
||||
assert.NotEmpty(t, store.deleted, "stale chunks must be deleted before re-embed")
|
||||
}
|
||||
|
||||
func anyContains(ss []string, sub string) bool {
|
||||
for _, s := range ss {
|
||||
if strings.Contains(s, sub) {
|
||||
return true
|
||||
}
|
||||
}
|
||||
return false
|
||||
}
|
||||
@@ -61,6 +61,26 @@ func (s *Server) tools() []map[string]any {
|
||||
"hall": enum("optional memory type (requires wing)", halls...),
|
||||
}),
|
||||
},
|
||||
{
|
||||
"name": "brain_update",
|
||||
"description": "Supersede an existing brain note in place: whole-note body replace + frontmatter re-stamp (updated_at, supersedes=prior content hash, supersede_reason). Errors if the target does not exist — use brain_write to create. Returns {id, path, content_hash, superseded}. Prior version recoverable from git.",
|
||||
"inputSchema": schema([]string{"content"}, map[string]any{
|
||||
"content": str("new full body (whole-note replace)"),
|
||||
"slug": str("target note slug within wing/hall, OR a full brain-relative path (e.g. wiki/jepa-fx/facts/x.md)"),
|
||||
"wing": str("wing of the target (required unless slug/path is a full path)"),
|
||||
"hall": enum("hall of the target (required unless slug/path is a full path)", halls...),
|
||||
"path": str("full brain-relative path to the target; takes precedence over slug/wing/hall"),
|
||||
"reason": str("optional short note on why superseded — stamped into frontmatter"),
|
||||
}),
|
||||
},
|
||||
{
|
||||
"name": "brain_get",
|
||||
"description": "Fetch a single brain note by id or path (both are the brain-relative path — the note handle). Returns {id, path, content_hash, frontmatter, body}. Read-after-write confirmation without a lexical re-query.",
|
||||
"inputSchema": schema([]string{}, map[string]any{
|
||||
"id": str("note id (brain-relative path) as returned by brain_write/brain_update"),
|
||||
"path": str("brain-relative path to the note; equivalent to id"),
|
||||
}),
|
||||
},
|
||||
{
|
||||
"name": "brain_tunnel",
|
||||
"description": "Create an explicit bidirectional [[wikilink]] between two notes in different wings. Idempotent.",
|
||||
@@ -222,7 +242,111 @@ func (s *Server) brainWrite(ctx context.Context, args json.RawMessage) (json.Raw
|
||||
}
|
||||
}
|
||||
s.indexInGraph(ctx, "brain_write", relPath)
|
||||
return json.Marshal(map[string]string{"path": relPath})
|
||||
// Read-after-write handle: id == relPath, content_hash == sha256 of
|
||||
// the bytes just written. path is kept for backward compatibility.
|
||||
_, _, hash, _ := api.ReadNote(s.brainDir, relPath)
|
||||
return json.Marshal(map[string]string{"id": relPath, "path": relPath, "content_hash": hash})
|
||||
}
|
||||
|
||||
type brainUpdateArgs struct {
|
||||
Slug string `json:"slug,omitempty"`
|
||||
Wing string `json:"wing,omitempty"`
|
||||
Hall string `json:"hall,omitempty"`
|
||||
Path string `json:"path,omitempty"`
|
||||
Content string `json:"content"`
|
||||
Reason string `json:"reason,omitempty"`
|
||||
}
|
||||
|
||||
// brainUpdate supersedes an existing note in place: whole-note body
|
||||
// replace, frontmatter re-stamp (updated_at/supersedes/supersede_reason),
|
||||
// graph re-index, and wing _index rebuild. It never creates — a missing
|
||||
// target is an error so the caller can fall back to brain_write.
|
||||
//
|
||||
// Embedding re-sync is delegated to the out-of-band vectorstore.Sync
|
||||
// ticker: the rewritten file's mtime advances, so the next pass re-embeds
|
||||
// it. This mirrors brain_write, which likewise does not embed in-handler.
|
||||
func (s *Server) brainUpdate(ctx context.Context, args json.RawMessage) (json.RawMessage, error) {
|
||||
var a brainUpdateArgs
|
||||
if err := json.Unmarshal(args, &a); err != nil {
|
||||
return nil, fmt.Errorf("parse args: %w", err)
|
||||
}
|
||||
if a.Content == "" {
|
||||
return nil, fmt.Errorf("content is required")
|
||||
}
|
||||
|
||||
opts := api.UpdateNoteOptions{Content: a.Content, Reason: a.Reason}
|
||||
switch {
|
||||
case a.Path != "":
|
||||
opts.Path = a.Path
|
||||
case strings.Contains(a.Slug, "/"):
|
||||
// slug carries a full path (issue #45: "slug ... OR full path").
|
||||
opts.Path = a.Slug
|
||||
default:
|
||||
opts.Wing, opts.Hall, opts.Slug = a.Wing, a.Hall, a.Slug
|
||||
}
|
||||
|
||||
relPath, hash, _, err := api.UpdateNote(s.brainDir, opts)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
|
||||
// Best-effort wiki upkeep, mirroring brain_write: rebuild the wing
|
||||
// _index and re-tunnel cross-wing matches against the new body. Both
|
||||
// are idempotent and never block — the note is already superseded.
|
||||
if wing := wingFromRelPath(relPath); wing != "" {
|
||||
if err := brain.BuildWingIndex(s.brainDir, wing); err != nil {
|
||||
slog.Warn("brain_update: auto-index failed", "wing", wing, "err", err)
|
||||
}
|
||||
if err := brain.AutoTunnel(s.brainDir, relPath, a.Content); err != nil {
|
||||
slog.Warn("brain_update: auto-tunnel failed", "src", relPath, "err", err)
|
||||
}
|
||||
}
|
||||
s.indexInGraph(ctx, "brain_update", relPath)
|
||||
|
||||
return json.Marshal(map[string]any{
|
||||
"id": relPath, "path": relPath, "content_hash": hash, "superseded": true,
|
||||
})
|
||||
}
|
||||
|
||||
// wingFromRelPath extracts the wing segment from a structured wiki path
|
||||
// (wiki/<wing>/<hall>/<slug>.md). Returns "" for legacy/non-wiki paths.
|
||||
func wingFromRelPath(relPath string) string {
|
||||
parts := strings.Split(relPath, "/")
|
||||
if len(parts) >= 4 && parts[0] == "wiki" {
|
||||
return parts[1]
|
||||
}
|
||||
return ""
|
||||
}
|
||||
|
||||
type brainGetArgs struct {
|
||||
ID string `json:"id,omitempty"`
|
||||
Path string `json:"path,omitempty"`
|
||||
}
|
||||
|
||||
// brainGet fetches a note by id or path (both are the brainDir-relative
|
||||
// path — the de-facto handle). Read-only; the create-path read-after-
|
||||
// write primitive that lets callers confirm a write landed without a
|
||||
// lexical re-query.
|
||||
func (s *Server) brainGet(_ context.Context, args json.RawMessage) (json.RawMessage, error) {
|
||||
var a brainGetArgs
|
||||
if err := json.Unmarshal(args, &a); err != nil {
|
||||
return nil, fmt.Errorf("parse args: %w", err)
|
||||
}
|
||||
target := a.Path
|
||||
if target == "" {
|
||||
target = a.ID
|
||||
}
|
||||
if target == "" {
|
||||
return nil, fmt.Errorf("id or path is required")
|
||||
}
|
||||
fm, body, hash, err := api.ReadNote(s.brainDir, target)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
return json.Marshal(map[string]any{
|
||||
"id": target, "path": target, "content_hash": hash,
|
||||
"frontmatter": fm, "body": body,
|
||||
})
|
||||
}
|
||||
|
||||
// indexInGraph is a best-effort wrapper around graphsync.IndexDoc that
|
||||
|
||||
@@ -1,7 +1,7 @@
|
||||
// Package mcp implements an MCP HTTP handler for the ingestion service.
|
||||
// Exposed tools: brain_query, brain_write, brain_index, brain_tunnel,
|
||||
// brain_ingest, brain_ingest_raw, brain_answer, brain_classify,
|
||||
// brain_graph, brain_context, session_log.
|
||||
// Exposed tools: brain_query, brain_write, brain_update, brain_get,
|
||||
// brain_index, brain_tunnel, brain_ingest, brain_ingest_raw,
|
||||
// brain_answer, brain_classify, brain_graph, brain_context, session_log.
|
||||
package mcp
|
||||
|
||||
import (
|
||||
@@ -177,6 +177,10 @@ func (s *Server) handleCall(ctx context.Context, name string, args json.RawMessa
|
||||
return s.brainQuery(ctx, args)
|
||||
case "brain_write":
|
||||
return s.brainWrite(ctx, args)
|
||||
case "brain_update":
|
||||
return s.brainUpdate(ctx, args)
|
||||
case "brain_get":
|
||||
return s.brainGet(ctx, args)
|
||||
case "brain_index":
|
||||
return s.brainIndex(ctx, args)
|
||||
case "brain_tunnel":
|
||||
|
||||
@@ -55,7 +55,8 @@ func TestServerToolsList(t *testing.T) {
|
||||
names = append(names, t.(map[string]any)["name"].(string))
|
||||
}
|
||||
assert.ElementsMatch(t, []string{
|
||||
"brain_query", "brain_write", "brain_index", "brain_tunnel",
|
||||
"brain_query", "brain_write", "brain_update", "brain_get",
|
||||
"brain_index", "brain_tunnel",
|
||||
"brain_ingest_raw", "brain_ingest",
|
||||
"brain_answer", "brain_classify", "brain_graph", "brain_context",
|
||||
"session_log",
|
||||
|
||||
@@ -77,9 +77,22 @@ func (s *Server) brainAnswer(ctx context.Context, args json.RawMessage) (json.Ra
|
||||
return nil, fmt.Errorf("search: %w", err)
|
||||
}
|
||||
if s.reranker != nil && len(results) > 0 {
|
||||
results, err = rerankResults(ctx, s.reranker, a.Query, results, 5)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("rerank: %w", err)
|
||||
reranked, rerr := rerankResults(ctx, s.reranker, a.Query, results, 5)
|
||||
if rerr != nil {
|
||||
return nil, fmt.Errorf("rerank: %w", rerr)
|
||||
}
|
||||
// The reranker is a filter, not a gate. The Qwen3-Reranker is a
|
||||
// web-search cross-encoder: against a conversational / personal-
|
||||
// intent query ("what am I optimizing toward?") it scores even
|
||||
// on-topic notes as "no", which would collapse the whole answer to
|
||||
// "no relevant content" despite BM25 having retrieved relevant
|
||||
// content. When the reranker keeps nothing, fall back to the
|
||||
// BM25/vector ordering (capped to the no-reranker depth) rather
|
||||
// than returning an empty answer.
|
||||
if len(reranked) > 0 {
|
||||
results = reranked
|
||||
} else if len(results) > 10 {
|
||||
results = results[:10]
|
||||
}
|
||||
}
|
||||
if len(results) == 0 {
|
||||
|
||||
@@ -98,6 +98,42 @@ func TestBrainAnswer_RerankerFiltersBeforeLLM(t *testing.T) {
|
||||
assert.NotContains(t, sawSources, "noise.md")
|
||||
}
|
||||
|
||||
func TestBrainAnswer_RerankerKeepsNone_FallsBackToBM25(t *testing.T) {
|
||||
brainDir := brainDirWithContent(t) // test.md BM25-matches "pass-rate logging"
|
||||
|
||||
// Reranker rejects every candidate ("no" to all) — models a
|
||||
// web-search cross-encoder facing a conversational / personal-intent
|
||||
// query, which is exactly when it wrongly scores on-topic notes as
|
||||
// irrelevant. The answer must still synthesize from the BM25 hits, not
|
||||
// collapse to "no relevant content".
|
||||
rrSrv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, _ *http.Request) {
|
||||
_ = json.NewEncoder(w).Encode(map[string]any{"response": "no", "done": true})
|
||||
}))
|
||||
defer rrSrv.Close()
|
||||
|
||||
var sawSources string
|
||||
llm := func(_ context.Context, _, user string) (string, error) {
|
||||
sawSources = user
|
||||
return "fallback answer", nil
|
||||
}
|
||||
|
||||
srv := mcp.NewServer(brainDir, nil, nil, llm).
|
||||
WithReranker(reranker.New(rrSrv.URL, "qwen3"))
|
||||
ts := httptest.NewServer(srv)
|
||||
defer ts.Close()
|
||||
|
||||
rpc := callTool(t, ts, "brain_answer", map[string]any{"query": "pass-rate logging"})
|
||||
require.Nil(t, rpc["error"])
|
||||
|
||||
content := rpc["result"].(map[string]any)["content"].([]any)[0].(map[string]any)["text"].(string)
|
||||
var result map[string]any
|
||||
require.NoError(t, json.Unmarshal([]byte(content), &result))
|
||||
|
||||
assert.Equal(t, "fallback answer", result["answer"])
|
||||
assert.NotEmpty(t, result["sources"], "reranker keeping nothing must fall back to BM25, not empty")
|
||||
assert.Contains(t, sawSources, "test.md")
|
||||
}
|
||||
|
||||
func TestBrainAnswer_NoLLM(t *testing.T) {
|
||||
srv := mcp.NewServer(t.TempDir(), nil, nil, nil)
|
||||
ts := httptest.NewServer(srv)
|
||||
|
||||
@@ -3,6 +3,7 @@ package vectorstore
|
||||
import (
|
||||
"fmt"
|
||||
"strings"
|
||||
"unicode/utf8"
|
||||
)
|
||||
|
||||
// NumberedChunk pairs a chunk's body with the storage path it will use
|
||||
@@ -66,6 +67,70 @@ func ChunkMarkdown(content string, maxBytes int) []string {
|
||||
}
|
||||
out = append(out, splitAtParagraphs(s, maxBytes)...)
|
||||
}
|
||||
|
||||
// Final guarantee: no chunk exceeds maxBytes. A single heading-less,
|
||||
// paragraph-less block (JSON-lines, minified content) survives the two
|
||||
// passes above whole — splitAtParagraphs emits an over-budget paragraph
|
||||
// rather than truncating prose. Hard-split any such chunk at line/rune
|
||||
// boundaries so the embedder never rejects an over-context chunk.
|
||||
final := make([]string, 0, len(out))
|
||||
for _, c := range out {
|
||||
if len(c) <= maxBytes {
|
||||
final = append(final, c)
|
||||
continue
|
||||
}
|
||||
final = append(final, hardSplit(c, maxBytes)...)
|
||||
}
|
||||
return final
|
||||
}
|
||||
|
||||
// hardSplit slices s into pieces no larger than maxBytes, breaking at line
|
||||
// boundaries where possible and otherwise mid-line at a UTF-8 rune boundary.
|
||||
// Last resort for content that has neither headings nor blank-line paragraphs.
|
||||
func hardSplit(s string, maxBytes int) []string {
|
||||
var out []string
|
||||
var cur strings.Builder
|
||||
flush := func() {
|
||||
if cur.Len() > 0 {
|
||||
out = append(out, cur.String())
|
||||
cur.Reset()
|
||||
}
|
||||
}
|
||||
for _, line := range strings.SplitAfter(s, "\n") {
|
||||
if line == "" {
|
||||
continue
|
||||
}
|
||||
if len(line) > maxBytes {
|
||||
flush()
|
||||
out = append(out, runeSplit(line, maxBytes)...)
|
||||
continue
|
||||
}
|
||||
if cur.Len() > 0 && cur.Len()+len(line) > maxBytes {
|
||||
flush()
|
||||
}
|
||||
cur.WriteString(line)
|
||||
}
|
||||
flush()
|
||||
return out
|
||||
}
|
||||
|
||||
// runeSplit slices s into <=maxBytes pieces without splitting a UTF-8 rune.
|
||||
func runeSplit(s string, maxBytes int) []string {
|
||||
var out []string
|
||||
for len(s) > maxBytes {
|
||||
cut := maxBytes
|
||||
for cut > 0 && !utf8.RuneStart(s[cut]) {
|
||||
cut--
|
||||
}
|
||||
if cut == 0 { // single rune wider than the budget; emit it whole
|
||||
cut = maxBytes
|
||||
}
|
||||
out = append(out, s[:cut])
|
||||
s = s[cut:]
|
||||
}
|
||||
if len(s) > 0 {
|
||||
out = append(out, s)
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
|
||||
@@ -30,6 +30,23 @@ func TestChunkMarkdown_SplitsAtHeadings(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
func TestChunkMarkdown_HardSplitsHeadinglessOversizedBlock(t *testing.T) {
|
||||
// A document with no headings and no blank-line paragraph breaks (e.g.
|
||||
// JSON-lines like wiki/telos/decisions/human-intent-column.md). The old
|
||||
// chunker emitted it as one over-budget chunk → nomic-embed returned
|
||||
// "input length exceeds the context length" (400). Every chunk must now
|
||||
// fit the budget, with no content lost.
|
||||
maxBytes := 200
|
||||
src := strings.Repeat("x", 1000) // one 1000-byte blob, no headings, no \n\n
|
||||
out := vectorstore.ChunkMarkdown(src, maxBytes)
|
||||
|
||||
require.Greater(t, len(out), 1, "oversized blob must be split")
|
||||
for i, c := range out {
|
||||
assert.LessOrEqual(t, len(c), maxBytes, "chunk %d over budget: %d bytes", i, len(c))
|
||||
}
|
||||
assert.Equal(t, 1000, strings.Count(strings.Join(out, ""), "x"), "no content lost")
|
||||
}
|
||||
|
||||
func TestChunkMarkdown_FurtherSplitsOversizedSection(t *testing.T) {
|
||||
// One H2 section with 4 paragraphs of ~80 chars each, limit 100.
|
||||
src := "## big\n\n" +
|
||||
|
||||
+3
-32
@@ -40,10 +40,9 @@ if [ -n "$ROOT_CONTEXT" ] && [ -f "$ROOT_CONTEXT" ]; then
|
||||
echo " Root context: $ROOT_CONTEXT"
|
||||
else
|
||||
# No reachable root AGENT.md — common in CI's clean checkout. The root+project
|
||||
# adapters (AGENTS.md, .cursorrules, .aider.conventions.md, system-prompt.txt)
|
||||
# require the root context to regenerate correctly, so we skip them entirely
|
||||
# and only regenerate CLAUDE.md (which is project-only and inherits root via
|
||||
# tree walk in Claude Code itself).
|
||||
# adapters (AGENTS.md, system-prompt.txt) require the root context to
|
||||
# regenerate correctly, so we skip them entirely and only regenerate CLAUDE.md
|
||||
# (which is project-only and inherits root via tree walk in Claude Code itself).
|
||||
echo " No root AGENT.md found — regenerating CLAUDE.md only"
|
||||
echo "Syncing project context from $PROJECT_FILE..."
|
||||
cat "$PROJECT_FILE" > CLAUDE.md
|
||||
@@ -78,30 +77,6 @@ generate_agents() {
|
||||
echo " → AGENTS.md (root + project; Crush, Pi, Antigravity)"
|
||||
}
|
||||
|
||||
# ── Cursor ───────────────────────────────────────────────────
|
||||
generate_cursor() {
|
||||
{
|
||||
echo "# Cursor rules — auto-generated"
|
||||
echo "# Do not edit. Run: task context:sync"
|
||||
echo ""
|
||||
root_block
|
||||
cat "$PROJECT_FILE"
|
||||
} > .cursorrules
|
||||
echo " → .cursorrules (root + project)"
|
||||
}
|
||||
|
||||
# ── Aider ────────────────────────────────────────────────────
|
||||
generate_aider() {
|
||||
{ root_block; cat "$PROJECT_FILE"; } > .aider.conventions.md
|
||||
if [ ! -f .aider.conf.yml ]; then
|
||||
cat > .aider.conf.yml << 'YAML'
|
||||
read: .aider.conventions.md
|
||||
auto-commits: false
|
||||
YAML
|
||||
fi
|
||||
echo " → .aider.conventions.md (root + project)"
|
||||
}
|
||||
|
||||
# ── Generic system prompt (Open WebUI, Mods, etc.) ──────────
|
||||
generate_system_prompt() {
|
||||
{
|
||||
@@ -142,8 +117,6 @@ echo "Syncing project context from $PROJECT_FILE..."
|
||||
if [ $# -eq 0 ]; then
|
||||
generate_claude
|
||||
generate_agents
|
||||
generate_cursor
|
||||
generate_aider
|
||||
generate_system_prompt
|
||||
generate_mcp
|
||||
else
|
||||
@@ -151,8 +124,6 @@ else
|
||||
case "$adapter" in
|
||||
claude) generate_claude ;;
|
||||
agents) generate_agents ;;
|
||||
cursor) generate_cursor ;;
|
||||
aider) generate_aider ;;
|
||||
prompt|system|openwebui|owui|generic) generate_system_prompt ;;
|
||||
mcp) generate_mcp ;;
|
||||
*) echo "Unknown adapter: $adapter" ;;
|
||||
|
||||
@@ -0,0 +1,109 @@
|
||||
---
|
||||
name: close-session
|
||||
description: Disciplined end-of-session closeout for a Claude.ai chat before archiving it. Harvests the session's decisions, artifacts, and open threads and durably persists them to the brain MCP and the right Gitea repo so nothing is lost when context resets. Use this whenever the user signals they are wrapping up — phrases like "close this out", "let's wrap up", "before I archive", "session retro", "capture this before I go", "did we lose anything", or any end-of-session/handoff cue — even if they don't say the word "close". Also use when the user explicitly asks to retro, archive, or hand off a working session.
|
||||
---
|
||||
|
||||
# close-session
|
||||
|
||||
Capture a finishing Claude.ai work session into durable storage before the chat is archived and its context is lost. The goal is simple and load-bearing: **after this runs, a fresh session (or another agent) can reconstruct what was decided, what was shipped, and what is still open — without the original chat.**
|
||||
|
||||
This skill is **batch**: one session in, findings out, done. It does not loop or re-read its own fresh output semantically (see Phase 5). Run the phases in order. Stop at any confirmation gate that says STOP.
|
||||
|
||||
## Operating constraints (read first)
|
||||
|
||||
- **Gitea owner is always `mathias`.** Never guess another owner.
|
||||
- **Ground-truth at HEAD before acting.** Issue bodies and doc references rot — stale hostnames, retired services, moved endpoints. Before closing/commenting on any issue, `gitea:issue_get` it fresh. Before asserting an infra fact, verify it; do not copy it from memory or from a stale issue body.
|
||||
- **Current infra truths** (verify rather than trust, but these are the known-good baseline): Gitea is `git.d-ma.be` (not `gitea.d-ma.be`). LiteLLM is `http://koala:30401/v1/` (public `https://llm-api.d-ma.be`); piguard runs NGINX Proxy Manager only — never reference `piguard:4000` or `koala:4000`. Identity provider is Authentik (Dex migration complete).
|
||||
- **Side-effects need a confirmation gate.** Closing issues, committing files, and writing to the brain are all real writes. Surface exactly what will happen and get a clear yes before doing it. Reads are free; writes are gated.
|
||||
- **Never fabricate.** If the session didn't produce a decision worth persisting, say so and skip that write. An empty-but-honest closeout beats an invented one.
|
||||
|
||||
## Phase 1 — Harvest
|
||||
|
||||
Reconstruct what actually happened this session from the conversation itself. Produce, in working memory:
|
||||
|
||||
- **Decisions taken** — what was decided and the reasoning, not just the outcome.
|
||||
- **Artifacts produced** — issues filed/closed, PRs opened/merged, files committed, brain notes written, ADRs. Capture identifiers (issue numbers, PR numbers, paths, commit SHAs) as you go.
|
||||
- **Open threads** — what was deferred, what's blocked, what the next session should pick up.
|
||||
- **Generalizable learnings** — reusable patterns or footguns that would bite anyone again (these are brain-worthy; project status is not).
|
||||
|
||||
Be honest about fidelity: a long session compresses harder at the start than the end. Flag anything you're reconstructing rather than certain of.
|
||||
|
||||
## Phase 2 — Ground-truth Gitea state
|
||||
|
||||
For every repo touched this session, get its true current state before proposing any change. `gitea:repo_status` (owner `mathias`) gives branches + open PRs + protection in one call. For each issue you intend to close, comment on, or reference: `gitea:issue_get` it fresh and compare to what the session assumed. Note any drift (closed-already, body rotted, renamed) — you'll surface it in Phase 3.
|
||||
|
||||
Do not write anything in this phase. This is the read pass.
|
||||
|
||||
## Phase 3 — Confirm and act on issue changes
|
||||
|
||||
Present a single consolidated plan of issue actions: which to close (with closing comment), which to file (discovered-but-deferred work — token-budget gaps, recorded limitations, v2 follow-ups), which to comment on. Include the exact title/body for any new issue and the closing rationale for any close.
|
||||
|
||||
**GATE — STOP and get explicit confirmation before any issue write.** Issue closes and new issues are side-effects. Once confirmed, execute them (`gitea:issue_close`, `gitea:issue_create`, `gitea:issue_comment`, all owner `mathias`), correcting any rotted references you found in Phase 2 as you go.
|
||||
|
||||
## Phase 4 — Commit the canonical session summary
|
||||
|
||||
Write one summary file to `mathias/ai-sessions`, committed directly to `main` via `gitea:file_write_branch` (no PR — this repo is solo and unprotected; if branch protection is ever added, fall back to a branch + PR).
|
||||
|
||||
**Path:** `summaries/claudeai/<YYYY-MM>/<YYYY-MM-DD>-<topic-slug>-<chatid8>.md`
|
||||
where `<chatid8>` is the first 8 chars of the chat's UUID if known, else a short stable slug. `claudeai` has no host segment — Claude.ai is Anthropic-side, not a homelab host.
|
||||
|
||||
**Frontmatter — the REDUCED live-capture schema.** A live close-session capture cannot populate the batch-export telemetry (token counts, message counts, duration_ms, permission_mode) — those only exist in the account export pipeline. Write only what's truthfully known, and mark fidelity so a reader (or the batch pipeline) can tell a live capture from an export:
|
||||
|
||||
```yaml
|
||||
---
|
||||
title: "<concise session title>"
|
||||
client: "claudeai"
|
||||
interface: "claudeai-chat"
|
||||
date: "<YYYY-MM-DD>"
|
||||
repos_touched: [<repo slugs>]
|
||||
topic_tags: [<tags>]
|
||||
outcome: "<shipped|in-progress|abandoned>"
|
||||
fidelity: "live-capture" # NOT an export; reconstructed live from chat
|
||||
captured_by: "close-session-skill"
|
||||
---
|
||||
```
|
||||
|
||||
Do not invent the export-only fields. `fidelity: live-capture` is the honest signal; if the batch export later produces a richer summary for the same session, the export is source of truth and supersedes this.
|
||||
|
||||
**Body** (keep it reconstructable, not exhaustive):
|
||||
```markdown
|
||||
## One-paragraph summary
|
||||
## Decisions
|
||||
## Key artifacts
|
||||
## Open threads
|
||||
```
|
||||
|
||||
**GATE — STOP, show the full file (path + frontmatter + body), get explicit confirmation before committing.**
|
||||
|
||||
## Phase 5 — Brain orientation note (the durable "where we are" record)
|
||||
|
||||
Write one brain note so a fresh session can orient without the chat. This uses the `brain_update`/`brain_get` verbs (live since 2026-06).
|
||||
|
||||
**Target:** `wing: <domain>` (the project/topic domain, e.g. `hyperguild`, `jepa-fx`), `hall: decisions`. The note is a knowledge-type record (a decision/orientation), grouped by knowledge-type, not by interface surface.
|
||||
|
||||
**Batch read-after-write discipline (important — do these in order, do not interleave):**
|
||||
|
||||
1. **Read first, before any write.** Check whether an orientation note already exists for this wing/topic. Do your "does this already exist / what should I supersede" reads NOW, up front. BM25/keyword search and `brain_get` are immediate; semantic/vector search may lag up to ~5 min after a write, so never rely on a semantic query to find something you wrote earlier in this same run.
|
||||
2. **Write or supersede:**
|
||||
- **New note** → `brain_write` (wing, hall: decisions). Returns `{id, path, content_hash}`.
|
||||
- **Superseding a prior orientation note** → `brain_update` (slug or path, wing, hall, content, reason). Whole-note replace; stamps `supersedes`/`updated_at`; returns `{id, path, content_hash, superseded}`. Use this instead of a second `brain_write` to the same slug — blind re-write creates duplicates/contradictions, which is the exact failure brain_update exists to prevent.
|
||||
3. **Confirm it landed** via `brain_get(id)` and check the returned `content_hash` matches what the write returned. This is the read-after-write confirmation — do it with `brain_get`, never a semantic query.
|
||||
|
||||
**RULE: no semantic/vector brain query after the first `brain_update` in this run.** The batch shape makes this natural — read up front, write, confirm by id. If you ever find the skill wanting to semantic-search a just-superseded note, stop and flag it (that's the signal the staleness window matters and needs the synchronous-reembed follow-up).
|
||||
|
||||
**GATE — STOP, show the note (target wing/hall, new-vs-supersede, full content), get explicit confirmation before the brain write.**
|
||||
|
||||
After the note lands, if it relates to a note in another wing, create the cross-link inline with `brain_tunnel(source, target)` (idempotent; both paths brain-relative, must be in different wings). Optionally append a `session_log` entry (`session_id`, `skill: close-session`, `phase`, `final_status`) for telemetry. Both are now callable directly from Claude.ai — no Claude Code/Crush handoff needed.
|
||||
|
||||
## Phase 6 — Verdict
|
||||
|
||||
Deliver a final "safe to archive" verdict in the chat. Either:
|
||||
|
||||
- **SAFE TO ARCHIVE** — list what landed (issues closed/filed with numbers, summary path, brain note id, any tunnels) so the trail is auditable. Then list anything still in the user's queue (e.g. a PR awaiting their merge, a decision owed next session).
|
||||
- **NOT YET** — name the specific gate that wasn't passed or the write that failed, and what to do about it.
|
||||
|
||||
Never claim safe-to-archive if any gated write was declined or errored. The verdict is the skill's contract: if it says safe, the session can be lost without losing the work.
|
||||
|
||||
## Why the gates and the batch discipline matter
|
||||
|
||||
The whole point is durability across a context reset. Every gate is a place where a wrong write would silently corrupt the record (close the wrong issue, overwrite a good brain note, commit a half-truth). The batch read-discipline in Phase 5 exists because the brain's vector index refreshes out-of-band: write-then-semantically-reread in the same run can read stale, so the skill front-loads reads and confirms writes by id. Get those right and the skill does what it promises — nothing important is lost when the chat goes away.
|
||||
@@ -0,0 +1,248 @@
|
||||
# Capture capability — use-case & BDD specification
|
||||
|
||||
**Status:** Decisions resolved 2026-06-22 (§4). Ready for implementation scoping. `capture` is a
|
||||
privileged cross-harness write path touching brain + Gitea + ai-sessions.
|
||||
**Tracks:** hyperguild #49.
|
||||
**Governed by:** `infra/docs/architecture/01-invariants.md` (I1–I5), the admissibility test in
|
||||
`00-synthesis-model.md`, and the distributed-consolidation shape mandated by
|
||||
`brain/wiki/homelab/decisions/no-centralized-cross-harness-observer-2026-06-17.md`.
|
||||
|
||||
---
|
||||
|
||||
## 1. Use-case (Clean Architecture form)
|
||||
|
||||
**Name:** CaptureSession
|
||||
**Actor:** A harness acting on the user's behalf (claude.ai Chat/Cowork/Code/Design, Claude Code
|
||||
CLI, Crush, Pi, LLM Council, Agentsquad executor/reviewer) — or the user directly.
|
||||
**Goal:** Durably persist a finished session's valuable output — insights → brain, action items →
|
||||
Gitea tickets, optional summary → ai-sessions — with one uniform invocation, identical core
|
||||
behaviour across harnesses.
|
||||
|
||||
**Primary success scenario (essential steps):**
|
||||
1. Caller assembles capture input (insights, tickets, optional summary) + context (harness,
|
||||
session_ref, fidelity, actor, **data-classification**).
|
||||
2. System validates the whole request (fail-closed).
|
||||
3. System resolves **effective classification** (stricter of caller-declared and target-derived)
|
||||
and the **server-derived harness origin** (from the authenticated principal). It checks the
|
||||
**sovereignty gate** (I1): if effective classification is confidential AND the origin is a
|
||||
non-sovereign (us-nexus) surface, the capture is **refused** before any write.
|
||||
4. System persists insights (write or supersede), tickets (create/close/comment), summary — each
|
||||
best-effort, recording per-item outcome.
|
||||
5. System emits an **audit record** (I5) of who/what captured what, when, via which principal.
|
||||
6. System returns a structured, partial-aware receipt.
|
||||
|
||||
**Architectural shape:** the *logic* is a shared use-case (`CaptureService`), invoked **per-harness
|
||||
against the caller's own credentials** (distributed consolidation — no high-degree observer node).
|
||||
A central authenticated relay endpoint exists ONLY as a fallback for harnesses that cannot run the
|
||||
use-case in-process (Crush/Pi/headless); the relay holds no standing visibility and retains nothing
|
||||
beyond the I5 audit log.
|
||||
|
||||
---
|
||||
|
||||
## 2. Invariant obligations (acceptance gates, not nice-to-haves)
|
||||
|
||||
| Invariant | Obligation on `capture` |
|
||||
|---|---|
|
||||
| **I1 sovereign containment** | A confidential-classified session MUST NOT be captured through a us-nexus harness. Harness origin is **server-derived from the authenticated principal** (not caller-asserted). Classification uses **model (C)**: caller declares, server cross-checks the target's tag, **stricter wins**, mismatch logged. See §4.1–4.2. |
|
||||
| **I2 deliberate acceptance** | The *distributed-library* form opens no new acceptance. IF a central relay node is deployed, its cross-harness reach MUST be entered in `infra/docs/security-baseline.md` with Why-accepted / Revisit-if before it ships. |
|
||||
| **I3 GitOps reconcilability** | IF `capture` runs as a deployed service, its manifest lives under `infra/k3s/apps/**` (sovereign source, Flux-reconciled). No untracked runtime. |
|
||||
| **I4 decisions captured** | The distributed-vs-central decision and the intent-named-verb pattern are recorded (ADR + brain). |
|
||||
| **I5 auditability** | Every capture emits a request-level audit record (actor/principal, harness, items written, timestamp) to the alloy/loki substrate. **Classification-aware degradation** (§4.4): confidential + sink-down → hard-refuse; internal/public + sink-down → durable local buffer + ntfy + reconcile. Floor: refuse if nothing can record the audit. |
|
||||
|
||||
---
|
||||
|
||||
## 3. BDD scenarios (Gherkin)
|
||||
|
||||
```gherkin
|
||||
Feature: Capture session value uniformly across harnesses
|
||||
As an operator working across many AI harnesses
|
||||
I want one uniform command to persist insights and file tickets
|
||||
So that valuable session output is never lost and is always auditable
|
||||
|
||||
Background:
|
||||
Given a brain store, a Gitea issue tracker, and an ai-sessions summary writer
|
||||
And the caller is authenticated with a principal
|
||||
And the session context declares a harness, a fidelity, and a data classification
|
||||
|
||||
# --- Core happy path ---
|
||||
Scenario: Capture insights and tickets from a non-confidential session
|
||||
Given a session classified as "internal"
|
||||
And the capture input has 2 insights and 1 ticket to create
|
||||
When capture is invoked
|
||||
Then both insights are written to the brain and their ids and content hashes are returned
|
||||
And the ticket is created in the named repo under owner "mathias"
|
||||
And an audit record is emitted naming the principal, harness, and items written
|
||||
And the receipt reports every item as ok
|
||||
|
||||
# --- I1: sovereignty gate (the load-bearing refusal) ---
|
||||
# Harness origin is server-derived from the authenticated principal, never from context.harness.
|
||||
Scenario: Refuse capture of a confidential session through a us-nexus harness
|
||||
Given a session whose effective classification is "confidential"
|
||||
And the authenticated principal resolves to a us-nexus harness origin
|
||||
When capture is invoked
|
||||
Then the capture is refused before any write
|
||||
And no insight, ticket, or summary is persisted
|
||||
And the refusal names the sovereignty invariant as the reason
|
||||
|
||||
Scenario: Allow capture of a confidential session through a sovereign harness
|
||||
Given a session whose effective classification is "confidential"
|
||||
And the authenticated principal resolves to a sovereign-soil harness origin
|
||||
When capture is invoked
|
||||
Then the capture proceeds and persists normally
|
||||
|
||||
Scenario: Ignore a caller-asserted harness label and use the server-derived origin
|
||||
Given the request context asserts harness "sovereign-soil"
|
||||
But the authenticated principal resolves to a us-nexus origin
|
||||
And the session classification is "confidential"
|
||||
When capture is invoked
|
||||
Then the capture is refused
|
||||
And the server-derived origin is used, not the asserted label
|
||||
And the asserted-vs-derived discrepancy is logged as a security event
|
||||
|
||||
# --- I1: classification model (C) — stricter of declared vs target-derived wins ---
|
||||
Scenario: Take the stricter classification when caller and target disagree
|
||||
Given the caller declares classification "internal"
|
||||
But the target wing/repo is tagged "confidential"
|
||||
When capture is invoked
|
||||
Then the effective classification is "confidential"
|
||||
And the declared-vs-derived mismatch is logged as a security event
|
||||
And the I1 gate is evaluated against "confidential"
|
||||
|
||||
Scenario: Honour a caller raising sensitivity above the target's tag
|
||||
Given the caller declares classification "confidential"
|
||||
And the target wing/repo is tagged "internal"
|
||||
When capture is invoked
|
||||
Then the effective classification is "confidential"
|
||||
And the capture is gated as confidential
|
||||
|
||||
# --- Supersession + staleness discipline (reuses #45 / #47 resolution) ---
|
||||
Scenario: Supersede a prior insight rather than duplicating it
|
||||
Given an insight whose context names an existing note to supersede
|
||||
When capture is invoked
|
||||
Then the existing note is updated in place, not duplicated
|
||||
And the prior content hash is recorded in the superseding note
|
||||
And read-after-write confirmation uses a direct fetch, never a semantic query
|
||||
|
||||
# --- Validation: fail-closed ---
|
||||
Scenario: Reject a malformed request before any write
|
||||
Given a capture input with an invalid wing/hall or unknown repo
|
||||
When capture is invoked
|
||||
Then the request is rejected with a validation error
|
||||
And nothing is written to the brain, Gitea, or ai-sessions
|
||||
|
||||
# --- Partial failure: best-effort + honest receipt ---
|
||||
Scenario: Report partial success when one item fails mid-capture
|
||||
Given a capture input with 2 insights and 1 ticket
|
||||
And the second insight write will fail
|
||||
When capture is invoked
|
||||
Then the first insight and the ticket are persisted
|
||||
And the second insight is reported as failed in the receipt
|
||||
And no rollback is attempted
|
||||
And the audit record reflects exactly what landed
|
||||
|
||||
# --- Dry run ---
|
||||
Scenario: Preview a capture without writing
|
||||
Given a valid capture input with dry_run true
|
||||
When capture is invoked
|
||||
Then the would-be receipt is returned
|
||||
And nothing is written anywhere
|
||||
|
||||
# --- I5: auditability is classification-aware (confidential fails closed) ---
|
||||
Scenario: Confidential capture hard-refuses when the central audit sink is down
|
||||
Given the effective classification is "confidential"
|
||||
And the central audit substrate (loki) cannot be written to
|
||||
When capture is invoked
|
||||
Then the capture is refused before any write
|
||||
And the reason names the auditability invariant
|
||||
# Confidential work must be centrally auditable at write time — no buffered exception.
|
||||
|
||||
Scenario: Internal capture degrades to a durable local buffer when the sink is down
|
||||
Given the effective classification is "internal" or "public"
|
||||
And the central audit substrate (loki) cannot be written to
|
||||
When capture is invoked
|
||||
Then the capture proceeds
|
||||
And the audit record is written to a durable LOCAL fallback buffer
|
||||
And an ntfy alert is emitted naming the degraded audit state
|
||||
And the receipt flags that audit was buffered locally, not centrally recorded
|
||||
|
||||
Scenario: Locally buffered audit records reconcile to the central sink on recovery
|
||||
Given internal-tier audit records were buffered locally during a sink outage
|
||||
When the central audit substrate becomes reachable again
|
||||
Then the buffered records are replayed to the central sink
|
||||
And the local buffer is cleared only after confirmed central write
|
||||
|
||||
Scenario: Even internal capture refuses if neither sink nor local buffer can be written
|
||||
Given the effective classification is "internal" or "public"
|
||||
And neither the central sink nor the local fallback buffer can be written
|
||||
When capture is invoked
|
||||
Then the capture is refused
|
||||
And the reason names the auditability invariant
|
||||
# Degrade-and-warn has a floor: if NOTHING can record the audit, do not write.
|
||||
|
||||
# --- Summary fidelity (collision rule from the retro work) ---
|
||||
Scenario: A richer-fidelity summary supersedes a thinner one for the same session
|
||||
Given a summary already exists for session_ref X at fidelity "live-capture"
|
||||
And a new summary arrives for session_ref X at fidelity "transcript-parse"
|
||||
When capture is invoked
|
||||
Then the transcript-parse summary supersedes the live-capture one
|
||||
And the live-capture summary is not left as a contradicting duplicate
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 4. Resolved decisions (2026-06-22)
|
||||
|
||||
These were open questions at draft; resolved in the 2026-06-22 review session. Recorded here as
|
||||
binding design decisions for the build.
|
||||
|
||||
1. **Classification trust — model (C): caller-declares + server-cross-checks, stricter wins.**
|
||||
The caller declares `context.classification`; the server **independently derives** the target's
|
||||
classification (from the target wing/repo's classification tag) and gates on the **stricter of
|
||||
the two**. The caller can voluntarily *raise* sensitivity but can never *lower* it below the
|
||||
target's floor. A declared-vs-derived **mismatch is logged as a security event** (I5).
|
||||
- **Prerequisite (new build work):** a classification taxonomy (e.g. `public` /
|
||||
`internal` / `confidential`) and a per-wing / per-repo classification tag the server can read.
|
||||
This must exist before the I1 gate is load-bearing. Tracked as a sub-task of #49.
|
||||
- **Implemented (#50):** taxonomy `public < internal < confidential` (ordered so "stricter wins"
|
||||
is `max`) in `ingestion/internal/classification/`. Tags are read from an optional
|
||||
`classification.yaml` at the brain root (`wings:` / `repos:` maps); absent entries fall to
|
||||
built-in defaults (`client-*` → confidential; `hyperguild`/`homelab` → internal; everything
|
||||
else → **confidential, fail-safe**). `Config.Derive(Target)` is the function the use-case
|
||||
calls. See brain `wiki/hyperguild/decisions/capture-classification-taxonomy`.
|
||||
- Rationale: composes with decision 2; fails safe; honours a caller flagging something *more*
|
||||
sensitive than its destination. Pure caller-trust (A) was rejected — it makes the gate theatre.
|
||||
|
||||
2. **Sovereign-harness determination — server-derived, not caller-asserted.**
|
||||
"Is this harness us-nexus / sovereign?" is derived from the **authenticated principal/origin**
|
||||
(the OAuth2 identity), never from `context.harness`. `context.harness` survives only as a
|
||||
self-reported label for the audit log — descriptive telemetry, **never a gate input**. A control
|
||||
keyed on an attacker-suppliable value is not a control.
|
||||
|
||||
3. **Central relay — ships in v1, with the I2 ledger entry.**
|
||||
The relay is required, not optional: claude.ai (Chat/Cowork/Design), Crush, Pi, and LLM Council
|
||||
cannot run the use-case library in-process, and those are primary day-to-day surfaces. Deferring
|
||||
the relay would ship a capability that doesn't work from the interfaces actually in use. Because
|
||||
the relay is a (thin, no-standing-visibility, audit-only-retention) central node, its cross-harness
|
||||
reach **must be entered in `infra/docs/security-baseline.md`** with Why-accepted / Revisit-if
|
||||
**before it ships** (I2). That ledger entry is v1 work, not a follow-up.
|
||||
|
||||
4. **Audit-sink-down — classification-aware: confidential fails closed, internal/public degrades.**
|
||||
The posture inherits from the effective classification (decision 1), so there is one coherent
|
||||
sensitivity model rather than a separate availability policy:
|
||||
- **Confidential + central audit sink unreachable → hard-refuse.** No buffer, no proceed.
|
||||
Confidential work must be centrally auditable *at write time*; "buffer and reconcile later"
|
||||
introduces a buffer-integrity question (can a write tamper with its own pending audit record?)
|
||||
that must not exist for confidential data. The simplicity of "refuse" is itself the assurance
|
||||
asset — trivially true, nothing to poke holes in.
|
||||
- **Internal / public + central sink unreachable → degrade-and-warn** with a durable local buffer
|
||||
+ ntfy alert + reconcile-on-recovery (the earlier Q4 design, now scoped to lower tiers). Keeps
|
||||
capture available for your own homelab work during an observability outage; negligible risk
|
||||
since the buffered record is still durable and the data isn't client-confidential.
|
||||
- **Floor (all tiers):** if *nothing* — neither central sink nor (for internal/public) the local
|
||||
buffer — can record the audit, capture **refuses**. No tier writes wholly un-audited.
|
||||
- Rationale: matches assurance cost to data sensitivity, exactly as the I1/sovereignty model
|
||||
does for placement. Presentable to a due-diligence client as "audit posture is
|
||||
classification-aware: confidential fails closed, internal degrades gracefully" — which
|
||||
demonstrates the judgment, not just a binary. Couples Q4 to Q1's classification machinery
|
||||
(being built anyway) and removes the buffer-integrity rabbit hole for the only tier where it
|
||||
mattered.
|
||||
Reference in New Issue
Block a user