Compare commits
22
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
9173f9058d | ||
|
|
3e84a41fed | ||
|
|
0e0571c7da | ||
|
|
7a27cf71a2 | ||
|
|
63df6d3283 | ||
|
|
f04b03e07e | ||
|
|
6c61f93146 | ||
|
|
95a69fc2c1 | ||
|
|
bb8bc0478c | ||
|
|
a961a3c064 | ||
|
|
bec28f9014 | ||
|
|
b62ac57382 | ||
|
|
aa918388b9 | ||
|
|
e8dbcf6eef | ||
|
|
0eeb1df4a2 | ||
|
|
9febb1bba1 | ||
|
|
5dc247b994 | ||
|
|
2125558196 | ||
|
|
2beaac2feb | ||
|
|
525811bc1a | ||
|
|
bad0581623 | ||
|
|
a94b860c2e |
@@ -1,315 +0,0 @@
|
|||||||
# Agent context — Mathias workspace
|
|
||||||
|
|
||||||
<!-- Canonical root context for all AI coding agents.
|
|
||||||
Lives at: ~/dev/.context/AGENT.md
|
|
||||||
Applies to every project under ~/dev/ unless overridden.
|
|
||||||
|
|
||||||
Run `task context:sync` from ~/dev/ to regenerate harness-specific files.
|
|
||||||
Project-level context in .context/PROJECT.md layers on top of this. -->
|
|
||||||
|
|
||||||
## Who I am
|
|
||||||
|
|
||||||
I'm Mathias, a digital product manager and technology consultant based in Sweden.
|
|
||||||
I build software, research emerging tech, and deliver consulting engagements
|
|
||||||
for clients under NDA. I work across AI/ML, financial automation, web applications,
|
|
||||||
and climate/sustainability tech.
|
|
||||||
|
|
||||||
## How I work with agents
|
|
||||||
|
|
||||||
- I think like a product manager — I care about *why* before *how*
|
|
||||||
- I want agents to be opinionated and push back, not just execute blindly
|
|
||||||
- I prefer concise responses; skip ceremony and get to the point
|
|
||||||
- When I say "build this", I mean production-quality with tests, not a demo
|
|
||||||
- Ask me before making irreversible changes or adding heavy dependencies
|
|
||||||
- I work with confidential client data — never send it to cloud APIs unless I explicitly say it's OK
|
|
||||||
|
|
||||||
## Behavior rules
|
|
||||||
|
|
||||||
These rules apply to every task across every project, regardless of harness.
|
|
||||||
|
|
||||||
1. **No assumptions.** Don't hide confusion — surface it. Surface tradeoffs explicitly.
|
|
||||||
Think before coding; if the problem is unclear, ask or state assumptions before acting.
|
|
||||||
2. **Minimum viable code.** Solve with the smallest change that works. Nothing
|
|
||||||
speculative, no "while we're here" cleanups, no premature abstractions. Simplicity first.
|
|
||||||
3. **Surgical changes.** Touch only what the task requires. Leave unrelated code,
|
|
||||||
files, and formatting alone. Diffs should be small and reviewable.
|
|
||||||
4. **Goal-driven execution.** Define clear success criteria up front for every task.
|
|
||||||
Loop — implement, verify, refine — until those criteria are met. Don't claim
|
|
||||||
completion without evidence (tests pass, command output, observed behavior).
|
|
||||||
5. **Trunk-Based Development — commit directly to main.** Every commit is one
|
|
||||||
logical change (one tool, one fix, one test) with passing tests. Main is always
|
|
||||||
deployable. Never create long-lived feature branches.
|
|
||||||
|
|
||||||
**Exception — parallel agents on same repo:** If another agent is known to be
|
|
||||||
actively working on the same repo simultaneously, create a short-lived branch
|
|
||||||
(`agent/<description>`), finish the task, and merge to main within the same
|
|
||||||
session. Do not leave agent branches open between sessions.
|
|
||||||
|
|
||||||
**Exception — external contributor or client four-eyes requirement:** Use
|
|
||||||
PR flow only when a human reviewer outside the project is required. Document
|
|
||||||
the reason in PROJECT.md.
|
|
||||||
|
|
||||||
## Default stack
|
|
||||||
|
|
||||||
| Layer | Default | Fallback | Last resort |
|
|
||||||
|-------|---------|----------|-------------|
|
|
||||||
| Language | Go | Python | TypeScript, Java, C |
|
|
||||||
| UI | HTMX + Templ | Server-rendered HTML | React (only if SPA is justified) |
|
|
||||||
| Build | Task (taskfile.dev) | Make | — |
|
|
||||||
| Containers | Docker Compose (dev), k3s (prod) | — | — |
|
|
||||||
| DB | PostgreSQL + sqlc | SQLite | — |
|
|
||||||
| Search | pgvector (vector), BM25 | Qdrant (when >1M vectors or hybrid retrieval) | — |
|
|
||||||
| Logging | slog (structured) | — | — |
|
|
||||||
| Testing | Table-driven, testify | — | — |
|
|
||||||
| Agents (Go) | google.golang.org/adk + pkg/litellm adapter | — | — |
|
|
||||||
|
|
||||||
Exploratory: Rust, Zig — I'll tell you when I want these.
|
|
||||||
|
|
||||||
## Code conventions
|
|
||||||
|
|
||||||
- **Go style**: golines, gofumpt, golangci-lint
|
|
||||||
- **Errors**: `fmt.Errorf("operation: %w", err)` — never naked, never log-and-return
|
|
||||||
- **Naming**: stdlib conventions, no stuttering
|
|
||||||
- **Architecture**: prefer stdlib over frameworks, constructor injection, env-var config parsed into typed structs
|
|
||||||
- **Git**: conventional commits (`feat:`, `fix:`, `chore:`), commit directly to main,
|
|
||||||
one logical change per commit, CI is the quality gate
|
|
||||||
- **Never**: long-lived feature branches, PRs for solo work, direct push without
|
|
||||||
passing `task check` locally first
|
|
||||||
- **Security**: no secrets in code, govulncheck before adding deps, SOPS for encrypted config
|
|
||||||
- **Dependencies**: prefer stdlib. testify, slog, templ, sqlc, google.golang.org/adk (agent projects only) are pre-approved; anything else needs justification in the commit message
|
|
||||||
|
|
||||||
## Infrastructure
|
|
||||||
|
|
||||||
Three machines on Tailscale:
|
|
||||||
|
|
||||||
| Machine | Role | Key specs |
|
|
||||||
|---------|------|-----------|
|
|
||||||
| koala | GPU inference, heavy compute | RTX 5070, runs k3s + llama-swap + shared postgres18/pgvector |
|
|
||||||
| iguana | Services, builds | M2 Ultra Mac |
|
|
||||||
| flamingo | Daily driver, edge | Mac mini, ~/dev is here |
|
|
||||||
|
|
||||||
- **Model routing**: LiteLLM in front of llama-swap (local) + cloud APIs (when permitted)
|
|
||||||
- **Orchestration**: k3s cluster across all three machines
|
|
||||||
- **Networking**: Tailscale mesh
|
|
||||||
|
|
||||||
## Project landscape
|
|
||||||
|
|
||||||
All development repos live at `~/dev/` (softlink from `~/Documents/local-dev/`).
|
|
||||||
|
|
||||||
Organized in thematic folders:
|
|
||||||
|
|
||||||
| Folder | Focus | Count |
|
|
||||||
|--------|-------|-------|
|
|
||||||
| `GO/` | Go web frameworks, API integrations, learning projects | ~10 |
|
|
||||||
| `AI/` | ML research, AI frameworks (FinRL, DSPy, crawl4ai) | ~6 |
|
|
||||||
| `AGENTS/` | Autonomous agents, coding agents, MCP servers, infra | ~15 |
|
|
||||||
| `QKX/` | Invoice processing, financial automation, payment systems | ~13 |
|
|
||||||
| `XT/` | Climate data, sustainability (Klimatkollen, Garbo) | ~2 |
|
|
||||||
|
|
||||||
See `~/dev/PROJECT_SUMMARY.md` for detailed descriptions of each project.
|
|
||||||
|
|
||||||
### Key active projects
|
|
||||||
|
|
||||||
- **super-koala** (`AGENTS/`) — multi-component agent stack with LangGraph, DSPy, MCP
|
|
||||||
- **azure-tiger** (`QKX/`) — invoice extraction → ISO 20022 payment instructions
|
|
||||||
- **gocrwl** (`AGENTS/`) — Go web crawler with containerized deployment
|
|
||||||
- **koala-ai-stack** (`AGENTS/`) — local AI server infrastructure management
|
|
||||||
- **klimatkollen** (`XT/`) — Swedish municipal climate data platform
|
|
||||||
|
|
||||||
## Knowledge base — actively use it
|
|
||||||
|
|
||||||
A persistent brain (BM25 search + LLM-synthesised Q&A) survives across sessions,
|
|
||||||
hosts, and harnesses. It holds 100+ hard-won entries: infra incident postmortems,
|
|
||||||
Go pitfalls, framework gotchas, design principles, ADRs. **It is not optional
|
|
||||||
reference material — query it actively, not just when explicitly told.**
|
|
||||||
|
|
||||||
### When to query (treat as a reflex)
|
|
||||||
|
|
||||||
- **Before** starting a non-trivial task — search for prior art with the symptom
|
|
||||||
AND the system component ("how did we solve X in Y?"). 5 seconds beats 5 hours.
|
|
||||||
- **When debugging** — search for the error string, the stack frame, the affected
|
|
||||||
service. Past you may have already paid this tax.
|
|
||||||
- **Before adopting** a pattern, library, framework, or model name — check if it
|
|
||||||
was tried and rejected, or what the integration footguns are.
|
|
||||||
- **When making architectural decisions** — search for the domain + "ADR" or
|
|
||||||
"decision" to find prior reasoning before re-deriving it.
|
|
||||||
- **When a recommendation feels novel** — challenge yourself: "has this been
|
|
||||||
documented?" The brain often has it.
|
|
||||||
|
|
||||||
### When to write
|
|
||||||
|
|
||||||
After you discover something that **future-you would forget** and that **isn't
|
|
||||||
recoverable from the code, git log, or PR description alone**:
|
|
||||||
|
|
||||||
- Bugs whose root cause is non-obvious and generalisable beyond this project.
|
|
||||||
- Framework / library / model-name quirks that bit you and would bite anyone.
|
|
||||||
- Design principles validated under fire (e.g. "every `_get` needs a `_list`").
|
|
||||||
- Postmortems for incidents: what broke, why, how diagnosed, what to do next time.
|
|
||||||
|
|
||||||
DON'T write project status, sprint progress, PR summaries, or "what I did this
|
|
||||||
session" — those rot fast and the originals are in git/gitea anyway. Brain
|
|
||||||
entries that age well are about *why*, *how to avoid*, and *what to do when*.
|
|
||||||
|
|
||||||
### How to access (per harness)
|
|
||||||
|
|
||||||
| Harness | Query | Write |
|
|
||||||
|---------|-------|-------|
|
|
||||||
| **Claude Code, Claude Desktop** | `brain_query` (BM25), `brain_answer` (LLM-synth + sources) MCP tools | `brain_write` MCP tool |
|
|
||||||
| **Crush, Pi, Antigravity, other MCP-capable** | same MCP server: `ingestion-brain` (via the `mcp__*_brain__*` namespace once authenticated) | same |
|
|
||||||
| **Anything HTTP-only (curl, scripts)** | `POST https://brain-mcp.d-ma.be/query` with `{"query":"..."}` (auth via `BRAIN_MCP_TOKEN`) | `POST .../write` with `{"content":"...","filename":"..."}` |
|
|
||||||
| **Browser / human inspection** | `https://gitea.d-ma.be/mathias/hyperguild` → `knowledge/` and `wiki/` markdown files |
|
|
||||||
|
|
||||||
- **Scoping**: defaults to `public` collection; client projects filter to `{client}` + `public`.
|
|
||||||
- **Routing**: brain_answer's LLM uses berget.ai as primary, iguana ollama as
|
|
||||||
fallback. Both are configurable in the `supervisor/ingestion-deployment.yaml`
|
|
||||||
on the koala k3s cluster; don't hardcode local-only model names into the
|
|
||||||
berget URL (see knowledge entry on namespace mismatches).
|
|
||||||
|
|
||||||
### Quick reflex checks
|
|
||||||
|
|
||||||
If you find yourself about to say any of these out loud, you owe yourself a brain query first:
|
|
||||||
|
|
||||||
- "I think the issue might be..."
|
|
||||||
- "Let me try X and see..."
|
|
||||||
- "I'll just write a script to..."
|
|
||||||
- "This is probably a new bug..."
|
|
||||||
- "Has anyone done this before?" — *yes, probably, go check.*
|
|
||||||
|
|
||||||
## Client work rules
|
|
||||||
|
|
||||||
When working on a project tagged with a client name:
|
|
||||||
1. Never send code, data, or context to cloud APIs — use local models only
|
|
||||||
2. Never reference other client projects or their data
|
|
||||||
3. Keep all artifacts within the client's git org / directory
|
|
||||||
4. Treat everything as confidential unless told otherwise
|
|
||||||
|
|
||||||
## Harness-agnostic principles
|
|
||||||
|
|
||||||
This context is designed to work with any AI coding tool:
|
|
||||||
- Claude Code, Cursor, Aider, Open WebUI, Charmbracelet Mods/Crush
|
|
||||||
- Pi Coding Agent, Mistral Vibe, Antigravity
|
|
||||||
- Any tool that accepts a system prompt or reads a markdown context file
|
|
||||||
|
|
||||||
The canonical source is always `.context/AGENT.md` (root) and `.context/PROJECT.md` (per-project).
|
|
||||||
Derived files are committed (see *How context propagates* below) so a `git pull` on any host yields full agent context with no setup.
|
|
||||||
|
|
||||||
## How context propagates
|
|
||||||
|
|
||||||
Canonical sources of truth:
|
|
||||||
- Universal: `~/dev/.context/AGENT.md` (this file)
|
|
||||||
- Project: `<repo>/.context/PROJECT.md` (per-repo)
|
|
||||||
|
|
||||||
Derived files (committed, regenerated by `task context:sync`):
|
|
||||||
- `CLAUDE.md`, `AGENTS.md`, `.cursorrules`, `.aider.conventions.md`,
|
|
||||||
`.context/system-prompt.txt`
|
|
||||||
|
|
||||||
Workflow:
|
|
||||||
1. Edit a canonical file. Run `task context:sync`. Commit canonical and
|
|
||||||
derived together. Push.
|
|
||||||
2. On any other host, `git pull` brings both. Claude Code (tree-walking)
|
|
||||||
uses `CLAUDE.md`; Crush / Pi / Antigravity (cwd-only) use `AGENTS.md`;
|
|
||||||
Cursor uses `.cursorrules`; Aider uses `.aider.conventions.md`.
|
|
||||||
3. `task check` runs `context:sync` then asserts `git status --porcelain`
|
|
||||||
is empty over the derived files (catches both modified-tracked drift
|
|
||||||
and missing-untracked adapters). A drift fails the check with a
|
|
||||||
message telling you to stage the regenerated files.
|
|
||||||
|
|
||||||
Behavior rules in this file and per-project rules in `PROJECT.md` apply
|
|
||||||
unconditionally on every host, every harness.
|
|
||||||
|
|
||||||
## Engineering Skills
|
|
||||||
|
|
||||||
Shared engineering skills are available in `~/dev/.skills/`. Load on demand via the index.
|
|
||||||
|
|
||||||
See `~/dev/.skills/SKILLS_INDEX.md` for the full list with descriptions and "use when" triggers.
|
|
||||||
|
|
||||||
Key skills:
|
|
||||||
- **TDD**: always write tests first — load `tdd` skill
|
|
||||||
- **Code Review**: load `code-review` skill before any review
|
|
||||||
- **SOLID/Clean Code**: load `solid` or `clean-code` skill for design work
|
|
||||||
- **Problem first**: load `problem-analysis` skill before coding non-trivial features
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
# Project context
|
|
||||||
|
|
||||||
<!-- Canonical project context. Edit this, run `task context:sync`.
|
|
||||||
Root agent context from ~/dev/.context/AGENT.md is automatically
|
|
||||||
prepended for harnesses that don't walk the directory tree. -->
|
|
||||||
|
|
||||||
## Identity
|
|
||||||
|
|
||||||
- **Name**: supervisor
|
|
||||||
- **Owner**: Mathias
|
|
||||||
- **Client**: personal
|
|
||||||
- **Repo**:
|
|
||||||
- **Status**: active
|
|
||||||
|
|
||||||
## Stack
|
|
||||||
|
|
||||||
- **Primary language**: Go
|
|
||||||
- **UI layer**: HTMX + Templ (when applicable)
|
|
||||||
- **Fallback languages**: Python, TypeScript (justify in PR if used)
|
|
||||||
- **Build**: Task (taskfile.dev), not Make
|
|
||||||
- **Containers**: Docker (compose for dev, k3s for deploy)
|
|
||||||
- **Target infra**: koala (GPU workloads), iguana (services), flamingo (edge)
|
|
||||||
|
|
||||||
## Conventions
|
|
||||||
|
|
||||||
### Code style
|
|
||||||
- Go: follow `golines`, `gofumpt`, `golangci-lint` with project config
|
|
||||||
- Tests: table-driven, in `_test.go` next to source, `testify` for assertions
|
|
||||||
- Errors: wrap with `fmt.Errorf("operation: %w", err)`, no naked returns
|
|
||||||
- Naming: stdlib conventions, no stuttering (`http.Client` not `http.HTTPClient`)
|
|
||||||
|
|
||||||
### Architecture preferences
|
|
||||||
- Prefer standard library over frameworks (net/http over gin/echo)
|
|
||||||
- Dependency injection via constructor functions, not containers
|
|
||||||
- Configuration via environment variables, parsed at startup into a typed struct
|
|
||||||
- Structured logging via `slog`
|
|
||||||
|
|
||||||
### Git
|
|
||||||
- Conventional commits: `feat:`, `fix:`, `chore:`, `docs:`, `refactor:`
|
|
||||||
- Branch naming: `feat/short-description`, `fix/short-description`
|
|
||||||
- PRs: one concern per PR, description explains *why* not *what*
|
|
||||||
|
|
||||||
### Security
|
|
||||||
- No secrets in code, ever — use env vars or SOPS-encrypted files
|
|
||||||
- Client data never leaves local network unless explicitly cleared
|
|
||||||
- Dependencies: audit with `govulncheck` before adding
|
|
||||||
|
|
||||||
## MCP endpoints
|
|
||||||
|
|
||||||
Two MCP servers are live, both reachable over Tailscale and via HTTPS domain:
|
|
||||||
|
|
||||||
- **`brain`** at `https://brain-mcp.d-ma.be/mcp` (NodePort `koala:30330`) —
|
|
||||||
`brain_query`, `brain_write`, `brain_ingest`, `brain_ingest_raw`,
|
|
||||||
`brain_answer`, `brain_classify`, `session_log`. Hosted by the ingestion
|
|
||||||
service. Auth: Dex JWT (claude.ai OAuth) or static `BRAIN_MCP_TOKEN`.
|
|
||||||
- **`routing`** at `http://koala:30310/mcp` — Mode 2 routing pod. Advertises
|
|
||||||
`review`, `debug`, `retrospective`, `trainer`; per-call routes to local model
|
|
||||||
or Claude based on brain `/pass-rate`. Bearer auth via `ROUTING_MCP_TOKEN`
|
|
||||||
(opt-in). Only `mode client-local` registers this endpoint.
|
|
||||||
|
|
||||||
The supervisor MCP (`koala:30320`) was retired in Plan 7 (2026-05-12). Its
|
|
||||||
skill workers (`tdd`, `spec`) are now SKILL.md files; routed skills moved to
|
|
||||||
the routing pod; brain tools moved to the brain MCP.
|
|
||||||
|
|
||||||
The brain HTTP REST API (`/query`, `/write`, `/ingest`, `/ingest-raw`,
|
|
||||||
`/ingest-path`, `/backfill-refs`, `/pass-rate`) remains available on port 3300
|
|
||||||
for shell scripts and non-MCP clients.
|
|
||||||
|
|
||||||
`brain_answer(query)` performs BM25 retrieval + LLM synthesis (berget.ai
|
|
||||||
gemma4:31b → iguana fallback). `brain_classify(text)` infers doc type, title,
|
|
||||||
and tags. Both require `BRAIN_LLM_PRIMARY_URL` to be set in the ingestion pod.
|
|
||||||
|
|
||||||
## Agent instructions
|
|
||||||
|
|
||||||
When acting as a coding agent on this project:
|
|
||||||
|
|
||||||
1. Read this file and all `SKILL.md` files in `.skills/` before starting work
|
|
||||||
2. Run `task check` before committing (lint + test + vet)
|
|
||||||
3. If unsure about a convention, check `DECISIONS.md` or ask
|
|
||||||
4. Never modify files outside the project root without explicit permission
|
|
||||||
5. When adding a dependency, explain why in the commit message
|
|
||||||
6. For client projects: never send code or context to cloud APIs — use local models via LiteLLM
|
|
||||||
@@ -32,6 +32,14 @@ and climate/sustainability tech.
|
|||||||
|
|
||||||
These rules apply to every task across every project, regardless of harness.
|
These rules apply to every task across every project, regardless of harness.
|
||||||
|
|
||||||
|
0. **Pre-task ritual — before ANY implementation (non-negotiable).** Run this before writing a single line:
|
||||||
|
- **Query the brain** (`brain_query`) for the domain + symptom. If the result changes your approach, surface it before acting. 5 seconds beats 5 hours.
|
||||||
|
- **Load the relevant skill** — see trigger table in *Engineering Skills* below.
|
||||||
|
- **Write the failing test first.** Name the test before the function. If the target is untestable (e.g. `main()` wiring), extract the logic into a testable function first. No implementation without a red test.
|
||||||
|
- **State the observable success criterion** — what specific behavior, output, or passing test proves this is done?
|
||||||
|
|
||||||
|
**TDD is non-negotiable.** "Tests pass" is not proof of correctness — only proof the tests ran. Write tests that would catch the bug before writing code that fixes it.
|
||||||
|
|
||||||
1. **No assumptions.** Don't hide confusion — surface it. Surface tradeoffs explicitly.
|
1. **No assumptions.** Don't hide confusion — surface it. Surface tradeoffs explicitly.
|
||||||
Think before coding; if the problem is unclear, ask or state assumptions before acting.
|
Think before coding; if the problem is unclear, ask or state assumptions before acting.
|
||||||
2. **Minimum viable code.** Solve with the smallest change that works. Nothing
|
2. **Minimum viable code.** Solve with the smallest change that works. Nothing
|
||||||
@@ -54,6 +62,22 @@ These rules apply to every task across every project, regardless of harness.
|
|||||||
PR flow only when a human reviewer outside the project is required. Document
|
PR flow only when a human reviewer outside the project is required. Document
|
||||||
the reason in PROJECT.md.
|
the reason in PROJECT.md.
|
||||||
|
|
||||||
|
6. **Close the loop — every substantive task ends with the same ritual.** Shipping
|
||||||
|
the code is not the end of the task; capturing it is. Run this unprompted:
|
||||||
|
- **Tag + bump SemVer** on the change (annotated tag; minor for a feature or
|
||||||
|
new/changed ADR, patch for a fix; docs in the same commit). Check the repo's
|
||||||
|
actual last tag — stated versions in docs drift stale.
|
||||||
|
- **Push** main and the tag (CI is the gate).
|
||||||
|
- **Persist generalizable learnings to the brain** (`brain_write`, wing/hall) —
|
||||||
|
the reusable patterns and the footguns that would bite anyone again, never
|
||||||
|
project status. See *Knowledge base — when to write* below.
|
||||||
|
- **File discovered-but-deferred work as tracker issues** on the project's own
|
||||||
|
repo — token-budget gaps, recorded ADR limitations, v2 follow-ups. Don't let
|
||||||
|
"out of scope, recorded" rot in a commit message; make it a ticket with a
|
||||||
|
source pointer.
|
||||||
|
- Surface the brain entries and issue numbers in the closing summary so the
|
||||||
|
trail is auditable.
|
||||||
|
|
||||||
## Default stack
|
## Default stack
|
||||||
|
|
||||||
| Layer | Default | Fallback | Last resort |
|
| Layer | Default | Fallback | Last resort |
|
||||||
@@ -83,6 +107,26 @@ Exploratory: Rust, Zig — I'll tell you when I want these.
|
|||||||
- **Security**: no secrets in code, govulncheck before adding deps, SOPS for encrypted config
|
- **Security**: no secrets in code, govulncheck before adding deps, SOPS for encrypted config
|
||||||
- **Dependencies**: prefer stdlib. testify, slog, templ, sqlc, google.golang.org/adk (agent projects only) are pre-approved; anything else needs justification in the commit message
|
- **Dependencies**: prefer stdlib. testify, slog, templ, sqlc, google.golang.org/adk (agent projects only) are pre-approved; anything else needs justification in the commit message
|
||||||
|
|
||||||
|
## Secret handling (every harness, every command)
|
||||||
|
|
||||||
|
Tool output is persisted: terminal → `~/.claude/projects` transcripts →
|
||||||
|
claudewatcher → brain/wiki → gitea history. A secret printed once is
|
||||||
|
searchable forever, and clearing it means rotating the key. So:
|
||||||
|
|
||||||
|
1. **Never print, echo, log, or transform a secret to inspect it.** No
|
||||||
|
`base64`/`xxd`/`cat` of a key, and never pipe a secret through a transform
|
||||||
|
to defeat `op run`'s output masking (it masks raw values; base64 hides them
|
||||||
|
from the mask — that exact trick leaked a key on 2026-06-11).
|
||||||
|
2. **Secrets stay in the subprocess.** Reference them only as env vars consumed
|
||||||
|
*inside* `op run --env-file ~/.op-env -- <cmd>`. Never place a literal secret
|
||||||
|
in a command's argv (it lands in the tool call and the transcript).
|
||||||
|
3. **Existence check without revealing the value:** `[ -n "$X" ] && echo set` —
|
||||||
|
never `${X:-...}` (returns the value when set) and never echo a substring of it.
|
||||||
|
4. **Cross-host secrets:** run the secret-consuming command on the host that has
|
||||||
|
the secret; do not forward a raw key over ssh argv/stdout.
|
||||||
|
5. If a secret does leak into output, say so immediately and flag it for rotation —
|
||||||
|
don't bury it.
|
||||||
|
|
||||||
## Infrastructure
|
## Infrastructure
|
||||||
|
|
||||||
Three machines on Tailscale:
|
Three machines on Tailscale:
|
||||||
@@ -162,7 +206,7 @@ entries that age well are about *why*, *how to avoid*, and *what to do when*.
|
|||||||
| **Claude Code, Claude Desktop** | `brain_query` (BM25), `brain_answer` (LLM-synth + sources) MCP tools | `brain_write` MCP tool |
|
| **Claude Code, Claude Desktop** | `brain_query` (BM25), `brain_answer` (LLM-synth + sources) MCP tools | `brain_write` MCP tool |
|
||||||
| **Crush, Pi, Antigravity, other MCP-capable** | same MCP server: `ingestion-brain` (via the `mcp__*_brain__*` namespace once authenticated) | same |
|
| **Crush, Pi, Antigravity, other MCP-capable** | same MCP server: `ingestion-brain` (via the `mcp__*_brain__*` namespace once authenticated) | same |
|
||||||
| **Anything HTTP-only (curl, scripts)** | `POST https://brain-mcp.d-ma.be/query` with `{"query":"..."}` (auth via `BRAIN_MCP_TOKEN`) | `POST .../write` with `{"content":"...","filename":"..."}` |
|
| **Anything HTTP-only (curl, scripts)** | `POST https://brain-mcp.d-ma.be/query` with `{"query":"..."}` (auth via `BRAIN_MCP_TOKEN`) | `POST .../write` with `{"content":"...","filename":"..."}` |
|
||||||
| **Browser / human inspection** | `https://gitea.d-ma.be/mathias/hyperguild` → `knowledge/` and `wiki/` markdown files |
|
| **Browser / human inspection** | `https://git.d-ma.be/mathias/hyperguild` → `knowledge/` and `wiki/` markdown files |
|
||||||
|
|
||||||
- **Scoping**: defaults to `public` collection; client projects filter to `{client}` + `public`.
|
- **Scoping**: defaults to `public` collection; client projects filter to `{client}` + `public`.
|
||||||
- **Routing**: brain_answer's LLM uses berget.ai as primary, iguana ollama as
|
- **Routing**: brain_answer's LLM uses berget.ai as primary, iguana ollama as
|
||||||
@@ -224,15 +268,17 @@ unconditionally on every host, every harness.
|
|||||||
|
|
||||||
## Engineering Skills
|
## Engineering Skills
|
||||||
|
|
||||||
Shared engineering skills are available in `~/dev/.skills/`. Load on demand via the index.
|
Shared engineering skills are available in `~/dev/.skills/`. Load at task start — not "on demand" but on schedule, before writing code. See `~/dev/.skills/SKILLS_INDEX.md` for the full list.
|
||||||
|
|
||||||
See `~/dev/.skills/SKILLS_INDEX.md` for the full list with descriptions and "use when" triggers.
|
**Skill trigger table — load before starting, not after getting stuck:**
|
||||||
|
|
||||||
Key skills:
|
| Task type | Load |
|
||||||
- **TDD**: always write tests first — load `tdd` skill
|
|-----------|------|
|
||||||
- **Code Review**: load `code-review` skill before any review
|
| Any feature or bug fix | `tdd` |
|
||||||
- **SOLID/Clean Code**: load `solid` or `clean-code` skill for design work
|
| Refactor or design | `clean-code` or `solid` |
|
||||||
- **Problem first**: load `problem-analysis` skill before coding non-trivial features
|
| Debug | `problem-analysis` |
|
||||||
|
| Review code or PRs | `code-review` |
|
||||||
|
| Frame a problem before coding | `problem-analysis` |
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
|
|||||||
-318
@@ -1,318 +0,0 @@
|
|||||||
# Cursor rules — auto-generated
|
|
||||||
# Do not edit. Run: task context:sync
|
|
||||||
|
|
||||||
# Agent context — Mathias workspace
|
|
||||||
|
|
||||||
<!-- Canonical root context for all AI coding agents.
|
|
||||||
Lives at: ~/dev/.context/AGENT.md
|
|
||||||
Applies to every project under ~/dev/ unless overridden.
|
|
||||||
|
|
||||||
Run `task context:sync` from ~/dev/ to regenerate harness-specific files.
|
|
||||||
Project-level context in .context/PROJECT.md layers on top of this. -->
|
|
||||||
|
|
||||||
## Who I am
|
|
||||||
|
|
||||||
I'm Mathias, a digital product manager and technology consultant based in Sweden.
|
|
||||||
I build software, research emerging tech, and deliver consulting engagements
|
|
||||||
for clients under NDA. I work across AI/ML, financial automation, web applications,
|
|
||||||
and climate/sustainability tech.
|
|
||||||
|
|
||||||
## How I work with agents
|
|
||||||
|
|
||||||
- I think like a product manager — I care about *why* before *how*
|
|
||||||
- I want agents to be opinionated and push back, not just execute blindly
|
|
||||||
- I prefer concise responses; skip ceremony and get to the point
|
|
||||||
- When I say "build this", I mean production-quality with tests, not a demo
|
|
||||||
- Ask me before making irreversible changes or adding heavy dependencies
|
|
||||||
- I work with confidential client data — never send it to cloud APIs unless I explicitly say it's OK
|
|
||||||
|
|
||||||
## Behavior rules
|
|
||||||
|
|
||||||
These rules apply to every task across every project, regardless of harness.
|
|
||||||
|
|
||||||
1. **No assumptions.** Don't hide confusion — surface it. Surface tradeoffs explicitly.
|
|
||||||
Think before coding; if the problem is unclear, ask or state assumptions before acting.
|
|
||||||
2. **Minimum viable code.** Solve with the smallest change that works. Nothing
|
|
||||||
speculative, no "while we're here" cleanups, no premature abstractions. Simplicity first.
|
|
||||||
3. **Surgical changes.** Touch only what the task requires. Leave unrelated code,
|
|
||||||
files, and formatting alone. Diffs should be small and reviewable.
|
|
||||||
4. **Goal-driven execution.** Define clear success criteria up front for every task.
|
|
||||||
Loop — implement, verify, refine — until those criteria are met. Don't claim
|
|
||||||
completion without evidence (tests pass, command output, observed behavior).
|
|
||||||
5. **Trunk-Based Development — commit directly to main.** Every commit is one
|
|
||||||
logical change (one tool, one fix, one test) with passing tests. Main is always
|
|
||||||
deployable. Never create long-lived feature branches.
|
|
||||||
|
|
||||||
**Exception — parallel agents on same repo:** If another agent is known to be
|
|
||||||
actively working on the same repo simultaneously, create a short-lived branch
|
|
||||||
(`agent/<description>`), finish the task, and merge to main within the same
|
|
||||||
session. Do not leave agent branches open between sessions.
|
|
||||||
|
|
||||||
**Exception — external contributor or client four-eyes requirement:** Use
|
|
||||||
PR flow only when a human reviewer outside the project is required. Document
|
|
||||||
the reason in PROJECT.md.
|
|
||||||
|
|
||||||
## Default stack
|
|
||||||
|
|
||||||
| Layer | Default | Fallback | Last resort |
|
|
||||||
|-------|---------|----------|-------------|
|
|
||||||
| Language | Go | Python | TypeScript, Java, C |
|
|
||||||
| UI | HTMX + Templ | Server-rendered HTML | React (only if SPA is justified) |
|
|
||||||
| Build | Task (taskfile.dev) | Make | — |
|
|
||||||
| Containers | Docker Compose (dev), k3s (prod) | — | — |
|
|
||||||
| DB | PostgreSQL + sqlc | SQLite | — |
|
|
||||||
| Search | pgvector (vector), BM25 | Qdrant (when >1M vectors or hybrid retrieval) | — |
|
|
||||||
| Logging | slog (structured) | — | — |
|
|
||||||
| Testing | Table-driven, testify | — | — |
|
|
||||||
| Agents (Go) | google.golang.org/adk + pkg/litellm adapter | — | — |
|
|
||||||
|
|
||||||
Exploratory: Rust, Zig — I'll tell you when I want these.
|
|
||||||
|
|
||||||
## Code conventions
|
|
||||||
|
|
||||||
- **Go style**: golines, gofumpt, golangci-lint
|
|
||||||
- **Errors**: `fmt.Errorf("operation: %w", err)` — never naked, never log-and-return
|
|
||||||
- **Naming**: stdlib conventions, no stuttering
|
|
||||||
- **Architecture**: prefer stdlib over frameworks, constructor injection, env-var config parsed into typed structs
|
|
||||||
- **Git**: conventional commits (`feat:`, `fix:`, `chore:`), commit directly to main,
|
|
||||||
one logical change per commit, CI is the quality gate
|
|
||||||
- **Never**: long-lived feature branches, PRs for solo work, direct push without
|
|
||||||
passing `task check` locally first
|
|
||||||
- **Security**: no secrets in code, govulncheck before adding deps, SOPS for encrypted config
|
|
||||||
- **Dependencies**: prefer stdlib. testify, slog, templ, sqlc, google.golang.org/adk (agent projects only) are pre-approved; anything else needs justification in the commit message
|
|
||||||
|
|
||||||
## Infrastructure
|
|
||||||
|
|
||||||
Three machines on Tailscale:
|
|
||||||
|
|
||||||
| Machine | Role | Key specs |
|
|
||||||
|---------|------|-----------|
|
|
||||||
| koala | GPU inference, heavy compute | RTX 5070, runs k3s + llama-swap + shared postgres18/pgvector |
|
|
||||||
| iguana | Services, builds | M2 Ultra Mac |
|
|
||||||
| flamingo | Daily driver, edge | Mac mini, ~/dev is here |
|
|
||||||
|
|
||||||
- **Model routing**: LiteLLM in front of llama-swap (local) + cloud APIs (when permitted)
|
|
||||||
- **Orchestration**: k3s cluster across all three machines
|
|
||||||
- **Networking**: Tailscale mesh
|
|
||||||
|
|
||||||
## Project landscape
|
|
||||||
|
|
||||||
All development repos live at `~/dev/` (softlink from `~/Documents/local-dev/`).
|
|
||||||
|
|
||||||
Organized in thematic folders:
|
|
||||||
|
|
||||||
| Folder | Focus | Count |
|
|
||||||
|--------|-------|-------|
|
|
||||||
| `GO/` | Go web frameworks, API integrations, learning projects | ~10 |
|
|
||||||
| `AI/` | ML research, AI frameworks (FinRL, DSPy, crawl4ai) | ~6 |
|
|
||||||
| `AGENTS/` | Autonomous agents, coding agents, MCP servers, infra | ~15 |
|
|
||||||
| `QKX/` | Invoice processing, financial automation, payment systems | ~13 |
|
|
||||||
| `XT/` | Climate data, sustainability (Klimatkollen, Garbo) | ~2 |
|
|
||||||
|
|
||||||
See `~/dev/PROJECT_SUMMARY.md` for detailed descriptions of each project.
|
|
||||||
|
|
||||||
### Key active projects
|
|
||||||
|
|
||||||
- **super-koala** (`AGENTS/`) — multi-component agent stack with LangGraph, DSPy, MCP
|
|
||||||
- **azure-tiger** (`QKX/`) — invoice extraction → ISO 20022 payment instructions
|
|
||||||
- **gocrwl** (`AGENTS/`) — Go web crawler with containerized deployment
|
|
||||||
- **koala-ai-stack** (`AGENTS/`) — local AI server infrastructure management
|
|
||||||
- **klimatkollen** (`XT/`) — Swedish municipal climate data platform
|
|
||||||
|
|
||||||
## Knowledge base — actively use it
|
|
||||||
|
|
||||||
A persistent brain (BM25 search + LLM-synthesised Q&A) survives across sessions,
|
|
||||||
hosts, and harnesses. It holds 100+ hard-won entries: infra incident postmortems,
|
|
||||||
Go pitfalls, framework gotchas, design principles, ADRs. **It is not optional
|
|
||||||
reference material — query it actively, not just when explicitly told.**
|
|
||||||
|
|
||||||
### When to query (treat as a reflex)
|
|
||||||
|
|
||||||
- **Before** starting a non-trivial task — search for prior art with the symptom
|
|
||||||
AND the system component ("how did we solve X in Y?"). 5 seconds beats 5 hours.
|
|
||||||
- **When debugging** — search for the error string, the stack frame, the affected
|
|
||||||
service. Past you may have already paid this tax.
|
|
||||||
- **Before adopting** a pattern, library, framework, or model name — check if it
|
|
||||||
was tried and rejected, or what the integration footguns are.
|
|
||||||
- **When making architectural decisions** — search for the domain + "ADR" or
|
|
||||||
"decision" to find prior reasoning before re-deriving it.
|
|
||||||
- **When a recommendation feels novel** — challenge yourself: "has this been
|
|
||||||
documented?" The brain often has it.
|
|
||||||
|
|
||||||
### When to write
|
|
||||||
|
|
||||||
After you discover something that **future-you would forget** and that **isn't
|
|
||||||
recoverable from the code, git log, or PR description alone**:
|
|
||||||
|
|
||||||
- Bugs whose root cause is non-obvious and generalisable beyond this project.
|
|
||||||
- Framework / library / model-name quirks that bit you and would bite anyone.
|
|
||||||
- Design principles validated under fire (e.g. "every `_get` needs a `_list`").
|
|
||||||
- Postmortems for incidents: what broke, why, how diagnosed, what to do next time.
|
|
||||||
|
|
||||||
DON'T write project status, sprint progress, PR summaries, or "what I did this
|
|
||||||
session" — those rot fast and the originals are in git/gitea anyway. Brain
|
|
||||||
entries that age well are about *why*, *how to avoid*, and *what to do when*.
|
|
||||||
|
|
||||||
### How to access (per harness)
|
|
||||||
|
|
||||||
| Harness | Query | Write |
|
|
||||||
|---------|-------|-------|
|
|
||||||
| **Claude Code, Claude Desktop** | `brain_query` (BM25), `brain_answer` (LLM-synth + sources) MCP tools | `brain_write` MCP tool |
|
|
||||||
| **Crush, Pi, Antigravity, other MCP-capable** | same MCP server: `ingestion-brain` (via the `mcp__*_brain__*` namespace once authenticated) | same |
|
|
||||||
| **Anything HTTP-only (curl, scripts)** | `POST https://brain-mcp.d-ma.be/query` with `{"query":"..."}` (auth via `BRAIN_MCP_TOKEN`) | `POST .../write` with `{"content":"...","filename":"..."}` |
|
|
||||||
| **Browser / human inspection** | `https://gitea.d-ma.be/mathias/hyperguild` → `knowledge/` and `wiki/` markdown files |
|
|
||||||
|
|
||||||
- **Scoping**: defaults to `public` collection; client projects filter to `{client}` + `public`.
|
|
||||||
- **Routing**: brain_answer's LLM uses berget.ai as primary, iguana ollama as
|
|
||||||
fallback. Both are configurable in the `supervisor/ingestion-deployment.yaml`
|
|
||||||
on the koala k3s cluster; don't hardcode local-only model names into the
|
|
||||||
berget URL (see knowledge entry on namespace mismatches).
|
|
||||||
|
|
||||||
### Quick reflex checks
|
|
||||||
|
|
||||||
If you find yourself about to say any of these out loud, you owe yourself a brain query first:
|
|
||||||
|
|
||||||
- "I think the issue might be..."
|
|
||||||
- "Let me try X and see..."
|
|
||||||
- "I'll just write a script to..."
|
|
||||||
- "This is probably a new bug..."
|
|
||||||
- "Has anyone done this before?" — *yes, probably, go check.*
|
|
||||||
|
|
||||||
## Client work rules
|
|
||||||
|
|
||||||
When working on a project tagged with a client name:
|
|
||||||
1. Never send code, data, or context to cloud APIs — use local models only
|
|
||||||
2. Never reference other client projects or their data
|
|
||||||
3. Keep all artifacts within the client's git org / directory
|
|
||||||
4. Treat everything as confidential unless told otherwise
|
|
||||||
|
|
||||||
## Harness-agnostic principles
|
|
||||||
|
|
||||||
This context is designed to work with any AI coding tool:
|
|
||||||
- Claude Code, Cursor, Aider, Open WebUI, Charmbracelet Mods/Crush
|
|
||||||
- Pi Coding Agent, Mistral Vibe, Antigravity
|
|
||||||
- Any tool that accepts a system prompt or reads a markdown context file
|
|
||||||
|
|
||||||
The canonical source is always `.context/AGENT.md` (root) and `.context/PROJECT.md` (per-project).
|
|
||||||
Derived files are committed (see *How context propagates* below) so a `git pull` on any host yields full agent context with no setup.
|
|
||||||
|
|
||||||
## How context propagates
|
|
||||||
|
|
||||||
Canonical sources of truth:
|
|
||||||
- Universal: `~/dev/.context/AGENT.md` (this file)
|
|
||||||
- Project: `<repo>/.context/PROJECT.md` (per-repo)
|
|
||||||
|
|
||||||
Derived files (committed, regenerated by `task context:sync`):
|
|
||||||
- `CLAUDE.md`, `AGENTS.md`, `.cursorrules`, `.aider.conventions.md`,
|
|
||||||
`.context/system-prompt.txt`
|
|
||||||
|
|
||||||
Workflow:
|
|
||||||
1. Edit a canonical file. Run `task context:sync`. Commit canonical and
|
|
||||||
derived together. Push.
|
|
||||||
2. On any other host, `git pull` brings both. Claude Code (tree-walking)
|
|
||||||
uses `CLAUDE.md`; Crush / Pi / Antigravity (cwd-only) use `AGENTS.md`;
|
|
||||||
Cursor uses `.cursorrules`; Aider uses `.aider.conventions.md`.
|
|
||||||
3. `task check` runs `context:sync` then asserts `git status --porcelain`
|
|
||||||
is empty over the derived files (catches both modified-tracked drift
|
|
||||||
and missing-untracked adapters). A drift fails the check with a
|
|
||||||
message telling you to stage the regenerated files.
|
|
||||||
|
|
||||||
Behavior rules in this file and per-project rules in `PROJECT.md` apply
|
|
||||||
unconditionally on every host, every harness.
|
|
||||||
|
|
||||||
## Engineering Skills
|
|
||||||
|
|
||||||
Shared engineering skills are available in `~/dev/.skills/`. Load on demand via the index.
|
|
||||||
|
|
||||||
See `~/dev/.skills/SKILLS_INDEX.md` for the full list with descriptions and "use when" triggers.
|
|
||||||
|
|
||||||
Key skills:
|
|
||||||
- **TDD**: always write tests first — load `tdd` skill
|
|
||||||
- **Code Review**: load `code-review` skill before any review
|
|
||||||
- **SOLID/Clean Code**: load `solid` or `clean-code` skill for design work
|
|
||||||
- **Problem first**: load `problem-analysis` skill before coding non-trivial features
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
# Project context
|
|
||||||
|
|
||||||
<!-- Canonical project context. Edit this, run `task context:sync`.
|
|
||||||
Root agent context from ~/dev/.context/AGENT.md is automatically
|
|
||||||
prepended for harnesses that don't walk the directory tree. -->
|
|
||||||
|
|
||||||
## Identity
|
|
||||||
|
|
||||||
- **Name**: supervisor
|
|
||||||
- **Owner**: Mathias
|
|
||||||
- **Client**: personal
|
|
||||||
- **Repo**:
|
|
||||||
- **Status**: active
|
|
||||||
|
|
||||||
## Stack
|
|
||||||
|
|
||||||
- **Primary language**: Go
|
|
||||||
- **UI layer**: HTMX + Templ (when applicable)
|
|
||||||
- **Fallback languages**: Python, TypeScript (justify in PR if used)
|
|
||||||
- **Build**: Task (taskfile.dev), not Make
|
|
||||||
- **Containers**: Docker (compose for dev, k3s for deploy)
|
|
||||||
- **Target infra**: koala (GPU workloads), iguana (services), flamingo (edge)
|
|
||||||
|
|
||||||
## Conventions
|
|
||||||
|
|
||||||
### Code style
|
|
||||||
- Go: follow `golines`, `gofumpt`, `golangci-lint` with project config
|
|
||||||
- Tests: table-driven, in `_test.go` next to source, `testify` for assertions
|
|
||||||
- Errors: wrap with `fmt.Errorf("operation: %w", err)`, no naked returns
|
|
||||||
- Naming: stdlib conventions, no stuttering (`http.Client` not `http.HTTPClient`)
|
|
||||||
|
|
||||||
### Architecture preferences
|
|
||||||
- Prefer standard library over frameworks (net/http over gin/echo)
|
|
||||||
- Dependency injection via constructor functions, not containers
|
|
||||||
- Configuration via environment variables, parsed at startup into a typed struct
|
|
||||||
- Structured logging via `slog`
|
|
||||||
|
|
||||||
### Git
|
|
||||||
- Conventional commits: `feat:`, `fix:`, `chore:`, `docs:`, `refactor:`
|
|
||||||
- Branch naming: `feat/short-description`, `fix/short-description`
|
|
||||||
- PRs: one concern per PR, description explains *why* not *what*
|
|
||||||
|
|
||||||
### Security
|
|
||||||
- No secrets in code, ever — use env vars or SOPS-encrypted files
|
|
||||||
- Client data never leaves local network unless explicitly cleared
|
|
||||||
- Dependencies: audit with `govulncheck` before adding
|
|
||||||
|
|
||||||
## MCP endpoints
|
|
||||||
|
|
||||||
Two MCP servers are live, both reachable over Tailscale and via HTTPS domain:
|
|
||||||
|
|
||||||
- **`brain`** at `https://brain-mcp.d-ma.be/mcp` (NodePort `koala:30330`) —
|
|
||||||
`brain_query`, `brain_write`, `brain_ingest`, `brain_ingest_raw`,
|
|
||||||
`brain_answer`, `brain_classify`, `session_log`. Hosted by the ingestion
|
|
||||||
service. Auth: Dex JWT (claude.ai OAuth) or static `BRAIN_MCP_TOKEN`.
|
|
||||||
- **`routing`** at `http://koala:30310/mcp` — Mode 2 routing pod. Advertises
|
|
||||||
`review`, `debug`, `retrospective`, `trainer`; per-call routes to local model
|
|
||||||
or Claude based on brain `/pass-rate`. Bearer auth via `ROUTING_MCP_TOKEN`
|
|
||||||
(opt-in). Only `mode client-local` registers this endpoint.
|
|
||||||
|
|
||||||
The supervisor MCP (`koala:30320`) was retired in Plan 7 (2026-05-12). Its
|
|
||||||
skill workers (`tdd`, `spec`) are now SKILL.md files; routed skills moved to
|
|
||||||
the routing pod; brain tools moved to the brain MCP.
|
|
||||||
|
|
||||||
The brain HTTP REST API (`/query`, `/write`, `/ingest`, `/ingest-raw`,
|
|
||||||
`/ingest-path`, `/backfill-refs`, `/pass-rate`) remains available on port 3300
|
|
||||||
for shell scripts and non-MCP clients.
|
|
||||||
|
|
||||||
`brain_answer(query)` performs BM25 retrieval + LLM synthesis (berget.ai
|
|
||||||
gemma4:31b → iguana fallback). `brain_classify(text)` infers doc type, title,
|
|
||||||
and tags. Both require `BRAIN_LLM_PRIMARY_URL` to be set in the ingestion pod.
|
|
||||||
|
|
||||||
## Agent instructions
|
|
||||||
|
|
||||||
When acting as a coding agent on this project:
|
|
||||||
|
|
||||||
1. Read this file and all `SKILL.md` files in `.skills/` before starting work
|
|
||||||
2. Run `task check` before committing (lint + test + vet)
|
|
||||||
3. If unsure about a convention, check `DECISIONS.md` or ask
|
|
||||||
4. Never modify files outside the project root without explicit permission
|
|
||||||
5. When adding a dependency, explain why in the commit message
|
|
||||||
6. For client projects: never send code or context to cloud APIs — use local models via LiteLLM
|
|
||||||
@@ -1,6 +1,6 @@
|
|||||||
name: cd
|
name: cd
|
||||||
|
|
||||||
on:
|
"on":
|
||||||
workflow_run:
|
workflow_run:
|
||||||
workflows: ["CI"]
|
workflows: ["CI"]
|
||||||
types: [completed]
|
types: [completed]
|
||||||
@@ -13,9 +13,9 @@ jobs:
|
|||||||
if: ${{ github.event.workflow_run.conclusion == 'success' && github.event.workflow_run.event == 'push' }}
|
if: ${{ github.event.workflow_run.conclusion == 'success' && github.event.workflow_run.event == 'push' }}
|
||||||
environment: staging
|
environment: staging
|
||||||
env:
|
env:
|
||||||
INGESTION_IMAGE: gitea.d-ma.be/mathias/ingestion
|
INGESTION_IMAGE: git.d-ma.be/mathias/ingestion
|
||||||
ROUTING_IMAGE: gitea.d-ma.be/mathias/routing
|
ROUTING_IMAGE: git.d-ma.be/mathias/routing
|
||||||
INFRA_REPO: git@gitea.d-ma.be:mathias/infra.git
|
INFRA_REPO: git@git.d-ma.be:mathias/infra.git
|
||||||
BUILDKIT_HOST: unix:///run/buildkit/buildkitd.sock
|
BUILDKIT_HOST: unix:///run/buildkit/buildkitd.sock
|
||||||
steps:
|
steps:
|
||||||
- name: Checkout
|
- name: Checkout
|
||||||
@@ -71,17 +71,17 @@ jobs:
|
|||||||
mkdir -p ~/.ssh
|
mkdir -p ~/.ssh
|
||||||
echo "${{ secrets.INFRA_DEPLOY_KEY }}" > ~/.ssh/infra_deploy_key
|
echo "${{ secrets.INFRA_DEPLOY_KEY }}" > ~/.ssh/infra_deploy_key
|
||||||
chmod 600 ~/.ssh/infra_deploy_key
|
chmod 600 ~/.ssh/infra_deploy_key
|
||||||
printf 'Host gitea.d-ma.be\n HostName 127.0.0.1\n Port 30022\n StrictHostKeyChecking no\n' >> ~/.ssh/config
|
printf 'Host git.d-ma.be\n HostName 127.0.0.1\n Port 30022\n StrictHostKeyChecking no\n' >> ~/.ssh/config
|
||||||
|
|
||||||
GIT_SSH_COMMAND="ssh -i ~/.ssh/infra_deploy_key -o IdentitiesOnly=yes" \
|
GIT_SSH_COMMAND="ssh -i ~/.ssh/infra_deploy_key -o IdentitiesOnly=yes" \
|
||||||
git clone "${INFRA_REPO}" /tmp/infra-update
|
git clone "${INFRA_REPO}" /tmp/infra-update
|
||||||
|
|
||||||
cd /tmp/infra-update
|
cd /tmp/infra-update
|
||||||
|
|
||||||
sed -i "s|gitea.d-ma.be/mathias/ingestion:.*|gitea.d-ma.be/mathias/ingestion:${IMAGE_TAG}|" \
|
sed -i "s|git.d-ma.be/mathias/ingestion:.*|git.d-ma.be/mathias/ingestion:${IMAGE_TAG}|" \
|
||||||
"k3s/apps/supervisor/ingestion-deployment.yaml"
|
"k3s/apps/supervisor/ingestion-deployment.yaml"
|
||||||
|
|
||||||
sed -i "s|gitea.d-ma.be/mathias/routing:.*|gitea.d-ma.be/mathias/routing:${IMAGE_TAG}|" \
|
sed -i "s|git.d-ma.be/mathias/routing:.*|git.d-ma.be/mathias/routing:${IMAGE_TAG}|" \
|
||||||
"k3s/apps/routing/deployment.yaml"
|
"k3s/apps/routing/deployment.yaml"
|
||||||
|
|
||||||
git config user.email "cd-bot@d-ma.be"
|
git config user.email "cd-bot@d-ma.be"
|
||||||
@@ -103,7 +103,7 @@ jobs:
|
|||||||
|
|
||||||
- name: Wait for Flux to apply new ingestion image
|
- name: Wait for Flux to apply new ingestion image
|
||||||
run: |
|
run: |
|
||||||
EXPECTED="gitea.d-ma.be/mathias/ingestion:${{ github.sha }}"
|
EXPECTED="git.d-ma.be/mathias/ingestion:${{ github.sha }}"
|
||||||
for i in $(seq 1 60); do
|
for i in $(seq 1 60); do
|
||||||
CURRENT=$(kubectl get deploy ingestion -n supervisor \
|
CURRENT=$(kubectl get deploy ingestion -n supervisor \
|
||||||
-o jsonpath='{.spec.template.spec.containers[0].image}' 2>/dev/null || echo "")
|
-o jsonpath='{.spec.template.spec.containers[0].image}' 2>/dev/null || echo "")
|
||||||
@@ -135,7 +135,7 @@ jobs:
|
|||||||
|
|
||||||
- name: Wait for Flux to apply new routing image
|
- name: Wait for Flux to apply new routing image
|
||||||
run: |
|
run: |
|
||||||
EXPECTED="gitea.d-ma.be/mathias/routing:${{ github.sha }}"
|
EXPECTED="git.d-ma.be/mathias/routing:${{ github.sha }}"
|
||||||
for i in $(seq 1 60); do
|
for i in $(seq 1 60); do
|
||||||
CURRENT=$(kubectl get deploy routing -n routing \
|
CURRENT=$(kubectl get deploy routing -n routing \
|
||||||
-o jsonpath='{.spec.template.spec.containers[0].image}' 2>/dev/null || echo "")
|
-o jsonpath='{.spec.template.spec.containers[0].image}' 2>/dev/null || echo "")
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
name: CI
|
name: CI
|
||||||
|
|
||||||
on:
|
"on":
|
||||||
push:
|
push:
|
||||||
branches: [main]
|
branches: [main]
|
||||||
tags: ["v*"]
|
tags: ["v*"]
|
||||||
|
|||||||
@@ -27,6 +27,14 @@ and climate/sustainability tech.
|
|||||||
|
|
||||||
These rules apply to every task across every project, regardless of harness.
|
These rules apply to every task across every project, regardless of harness.
|
||||||
|
|
||||||
|
0. **Pre-task ritual — before ANY implementation (non-negotiable).** Run this before writing a single line:
|
||||||
|
- **Query the brain** (`brain_query`) for the domain + symptom. If the result changes your approach, surface it before acting. 5 seconds beats 5 hours.
|
||||||
|
- **Load the relevant skill** — see trigger table in *Engineering Skills* below.
|
||||||
|
- **Write the failing test first.** Name the test before the function. If the target is untestable (e.g. `main()` wiring), extract the logic into a testable function first. No implementation without a red test.
|
||||||
|
- **State the observable success criterion** — what specific behavior, output, or passing test proves this is done?
|
||||||
|
|
||||||
|
**TDD is non-negotiable.** "Tests pass" is not proof of correctness — only proof the tests ran. Write tests that would catch the bug before writing code that fixes it.
|
||||||
|
|
||||||
1. **No assumptions.** Don't hide confusion — surface it. Surface tradeoffs explicitly.
|
1. **No assumptions.** Don't hide confusion — surface it. Surface tradeoffs explicitly.
|
||||||
Think before coding; if the problem is unclear, ask or state assumptions before acting.
|
Think before coding; if the problem is unclear, ask or state assumptions before acting.
|
||||||
2. **Minimum viable code.** Solve with the smallest change that works. Nothing
|
2. **Minimum viable code.** Solve with the smallest change that works. Nothing
|
||||||
@@ -49,6 +57,22 @@ These rules apply to every task across every project, regardless of harness.
|
|||||||
PR flow only when a human reviewer outside the project is required. Document
|
PR flow only when a human reviewer outside the project is required. Document
|
||||||
the reason in PROJECT.md.
|
the reason in PROJECT.md.
|
||||||
|
|
||||||
|
6. **Close the loop — every substantive task ends with the same ritual.** Shipping
|
||||||
|
the code is not the end of the task; capturing it is. Run this unprompted:
|
||||||
|
- **Tag + bump SemVer** on the change (annotated tag; minor for a feature or
|
||||||
|
new/changed ADR, patch for a fix; docs in the same commit). Check the repo's
|
||||||
|
actual last tag — stated versions in docs drift stale.
|
||||||
|
- **Push** main and the tag (CI is the gate).
|
||||||
|
- **Persist generalizable learnings to the brain** (`brain_write`, wing/hall) —
|
||||||
|
the reusable patterns and the footguns that would bite anyone again, never
|
||||||
|
project status. See *Knowledge base — when to write* below.
|
||||||
|
- **File discovered-but-deferred work as tracker issues** on the project's own
|
||||||
|
repo — token-budget gaps, recorded ADR limitations, v2 follow-ups. Don't let
|
||||||
|
"out of scope, recorded" rot in a commit message; make it a ticket with a
|
||||||
|
source pointer.
|
||||||
|
- Surface the brain entries and issue numbers in the closing summary so the
|
||||||
|
trail is auditable.
|
||||||
|
|
||||||
## Default stack
|
## Default stack
|
||||||
|
|
||||||
| Layer | Default | Fallback | Last resort |
|
| Layer | Default | Fallback | Last resort |
|
||||||
@@ -78,6 +102,26 @@ Exploratory: Rust, Zig — I'll tell you when I want these.
|
|||||||
- **Security**: no secrets in code, govulncheck before adding deps, SOPS for encrypted config
|
- **Security**: no secrets in code, govulncheck before adding deps, SOPS for encrypted config
|
||||||
- **Dependencies**: prefer stdlib. testify, slog, templ, sqlc, google.golang.org/adk (agent projects only) are pre-approved; anything else needs justification in the commit message
|
- **Dependencies**: prefer stdlib. testify, slog, templ, sqlc, google.golang.org/adk (agent projects only) are pre-approved; anything else needs justification in the commit message
|
||||||
|
|
||||||
|
## Secret handling (every harness, every command)
|
||||||
|
|
||||||
|
Tool output is persisted: terminal → `~/.claude/projects` transcripts →
|
||||||
|
claudewatcher → brain/wiki → gitea history. A secret printed once is
|
||||||
|
searchable forever, and clearing it means rotating the key. So:
|
||||||
|
|
||||||
|
1. **Never print, echo, log, or transform a secret to inspect it.** No
|
||||||
|
`base64`/`xxd`/`cat` of a key, and never pipe a secret through a transform
|
||||||
|
to defeat `op run`'s output masking (it masks raw values; base64 hides them
|
||||||
|
from the mask — that exact trick leaked a key on 2026-06-11).
|
||||||
|
2. **Secrets stay in the subprocess.** Reference them only as env vars consumed
|
||||||
|
*inside* `op run --env-file ~/.op-env -- <cmd>`. Never place a literal secret
|
||||||
|
in a command's argv (it lands in the tool call and the transcript).
|
||||||
|
3. **Existence check without revealing the value:** `[ -n "$X" ] && echo set` —
|
||||||
|
never `${X:-...}` (returns the value when set) and never echo a substring of it.
|
||||||
|
4. **Cross-host secrets:** run the secret-consuming command on the host that has
|
||||||
|
the secret; do not forward a raw key over ssh argv/stdout.
|
||||||
|
5. If a secret does leak into output, say so immediately and flag it for rotation —
|
||||||
|
don't bury it.
|
||||||
|
|
||||||
## Infrastructure
|
## Infrastructure
|
||||||
|
|
||||||
Three machines on Tailscale:
|
Three machines on Tailscale:
|
||||||
@@ -157,7 +201,7 @@ entries that age well are about *why*, *how to avoid*, and *what to do when*.
|
|||||||
| **Claude Code, Claude Desktop** | `brain_query` (BM25), `brain_answer` (LLM-synth + sources) MCP tools | `brain_write` MCP tool |
|
| **Claude Code, Claude Desktop** | `brain_query` (BM25), `brain_answer` (LLM-synth + sources) MCP tools | `brain_write` MCP tool |
|
||||||
| **Crush, Pi, Antigravity, other MCP-capable** | same MCP server: `ingestion-brain` (via the `mcp__*_brain__*` namespace once authenticated) | same |
|
| **Crush, Pi, Antigravity, other MCP-capable** | same MCP server: `ingestion-brain` (via the `mcp__*_brain__*` namespace once authenticated) | same |
|
||||||
| **Anything HTTP-only (curl, scripts)** | `POST https://brain-mcp.d-ma.be/query` with `{"query":"..."}` (auth via `BRAIN_MCP_TOKEN`) | `POST .../write` with `{"content":"...","filename":"..."}` |
|
| **Anything HTTP-only (curl, scripts)** | `POST https://brain-mcp.d-ma.be/query` with `{"query":"..."}` (auth via `BRAIN_MCP_TOKEN`) | `POST .../write` with `{"content":"...","filename":"..."}` |
|
||||||
| **Browser / human inspection** | `https://gitea.d-ma.be/mathias/hyperguild` → `knowledge/` and `wiki/` markdown files |
|
| **Browser / human inspection** | `https://git.d-ma.be/mathias/hyperguild` → `knowledge/` and `wiki/` markdown files |
|
||||||
|
|
||||||
- **Scoping**: defaults to `public` collection; client projects filter to `{client}` + `public`.
|
- **Scoping**: defaults to `public` collection; client projects filter to `{client}` + `public`.
|
||||||
- **Routing**: brain_answer's LLM uses berget.ai as primary, iguana ollama as
|
- **Routing**: brain_answer's LLM uses berget.ai as primary, iguana ollama as
|
||||||
@@ -219,15 +263,17 @@ unconditionally on every host, every harness.
|
|||||||
|
|
||||||
## Engineering Skills
|
## Engineering Skills
|
||||||
|
|
||||||
Shared engineering skills are available in `~/dev/.skills/`. Load on demand via the index.
|
Shared engineering skills are available in `~/dev/.skills/`. Load at task start — not "on demand" but on schedule, before writing code. See `~/dev/.skills/SKILLS_INDEX.md` for the full list.
|
||||||
|
|
||||||
See `~/dev/.skills/SKILLS_INDEX.md` for the full list with descriptions and "use when" triggers.
|
**Skill trigger table — load before starting, not after getting stuck:**
|
||||||
|
|
||||||
Key skills:
|
| Task type | Load |
|
||||||
- **TDD**: always write tests first — load `tdd` skill
|
|-----------|------|
|
||||||
- **Code Review**: load `code-review` skill before any review
|
| Any feature or bug fix | `tdd` |
|
||||||
- **SOLID/Clean Code**: load `solid` or `clean-code` skill for design work
|
| Refactor or design | `clean-code` or `solid` |
|
||||||
- **Problem first**: load `problem-analysis` skill before coding non-trivial features
|
| Debug | `problem-analysis` |
|
||||||
|
| Review code or PRs | `code-review` |
|
||||||
|
| Frame a problem before coding | `problem-analysis` |
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
|
|||||||
@@ -4,6 +4,74 @@ Record *why* things are the way they are. Future-you will thank present-you.
|
|||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
|
## 2026-05-28 — three active harnesses: hyperguild, agentsquad, Crush (extends earlier boundary decision)
|
||||||
|
|
||||||
|
**Context:** After wiring Crush to LiteLLM in May 2026, there are now three active harnesses.
|
||||||
|
The earlier boundary decision only covered hyperguild vs agentsquad. Crush's role was undefined.
|
||||||
|
|
||||||
|
**Decision:** Three harnesses, three distinct roles, shared skills layer.
|
||||||
|
|
||||||
|
| Harness | Engine | Primary use | Brain MCP? | Routing pod? | Skills? |
|
||||||
|
|---------|--------|-------------|------------|--------------|---------|
|
||||||
|
| **hyperguild** | Claude Code + MCP | Disciplined solo coding sessions, TDD/review/debug workflows | Yes | Yes | Yes (SKILL.md) |
|
||||||
|
| **agentsquad** | OpenCode + LiteLLM | Multi-agent task execution, executor/reviewer pipelines | No | No (own routing) | Yes (SKILL.md) |
|
||||||
|
| **Crush** | Charmbracelet TUI + LiteLLM | Interactive local coding, quick iterations on flamingo | No (not yet) | No (direct LiteLLM) | Yes (SKILL.md) |
|
||||||
|
|
||||||
|
**Crush specifics (as of 2026-05-28):**
|
||||||
|
- Config: `~/.config/crush/crush.json` on flamingo (see brain: `homelab/facts/crush-litellm-wiring-2026-05`)
|
||||||
|
- Connects directly to LiteLLM at `http://koala:4000/v1/` using `sk-local-123`
|
||||||
|
- Auth type: `openai-compat` (not `openai`)
|
||||||
|
- Does NOT go through the routing pod — model selection is manual in the Crush UI
|
||||||
|
- Brain MCP not wired — Crush has no MCP client capability today; revisit if Crush adds MCP support
|
||||||
|
|
||||||
|
**Shared across all three:**
|
||||||
|
- `mathias/skills` — any SKILL.md file works in all three harnesses
|
||||||
|
- LiteLLM proxy on koala (`http://koala:4000/v1/`) — Crush and agentsquad both route through it; hyperguild does too for local model calls
|
||||||
|
|
||||||
|
**Consequences:** No consolidation needed. crush.json must be kept in sync when litellm_config.yaml model names change. The `crush.json` canonical location is `~/.config/crush/crush.json` on flamingo — not yet tracked in a dotfiles repo (track as tech debt).
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 2026-05-28 — "field benchmark" for local models = pass-rate at scale (supersedes GOTTH eval suite)
|
||||||
|
|
||||||
|
**Context:** The GOTTH eval suite (45 offline prompts across 5 categories) was replaced by
|
||||||
|
a "field benchmark" in May 2026, but the replacement was never defined concretely.
|
||||||
|
|
||||||
|
**Decision:** The field benchmark is per-skill pass rate over real routing pod usage,
|
||||||
|
collected automatically by `internal/routing/passrate.go` and exposed at:
|
||||||
|
|
||||||
|
```
|
||||||
|
GET /pass-rate?skill=<name>&window=<duration>
|
||||||
|
```
|
||||||
|
|
||||||
|
No separate eval suite. No synthetic prompts. The benchmark runs itself once the routing
|
||||||
|
pod receives real traffic. Target: 30-day rolling window per skill, reviewed monthly.
|
||||||
|
|
||||||
|
**Bootstrap note:** With no session history, `passrate.go` returns `nil` and the router
|
||||||
|
defaults to the thinking model for every call. The fast-model path activates only after
|
||||||
|
real pass-rate data accumulates. Seed with real usage — do not pre-populate.
|
||||||
|
|
||||||
|
**Consequences:** Zero maintenance overhead for the benchmark. The tradeoff is that results
|
||||||
|
are only meaningful after ~2 weeks of real usage, and skills that are rarely invoked will
|
||||||
|
have statistically thin pass-rate data. Revisit if a skill has fewer than 20 calls in 30 days.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 2026-05-28 — brain injection in skill handlers: review is done, others unverified
|
||||||
|
|
||||||
|
**Context:** The April 2026 scope reset listed "brain_query injection into skill handlers"
|
||||||
|
as the top priority. As of 2026-05-28, `internal/skills/review/handlers.go` calls
|
||||||
|
`brain.Query(ctx, ...)` before dispatching to the LLM — confirmed in code review.
|
||||||
|
Status of debug, retrospective, and trainer handlers is unverified.
|
||||||
|
|
||||||
|
**Decision:** Treat review as the reference implementation. Verify debug, retrospective,
|
||||||
|
trainer against the same pattern before shipping new skill work. Tracked in issue #32.
|
||||||
|
|
||||||
|
**Consequences:** The April concern may be stale for review. A one-pass audit of the other
|
||||||
|
three skill handlers closes this fully.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
## 2026-04-08 — AGENTS.md as cross-tool standard, not CLAUDE.md
|
## 2026-04-08 — AGENTS.md as cross-tool standard, not CLAUDE.md
|
||||||
|
|
||||||
**Context**: Multiple tools (Crush, Pi, Antigravity) read `AGENTS.md` natively. Claude Code reads `CLAUDE.md`. Building on `CLAUDE.md` as the primary format locks into one vendor.
|
**Context**: Multiple tools (Crush, Pi, Antigravity) read `AGENTS.md` natively. Claude Code reads `CLAUDE.md`. Building on `CLAUDE.md` as the primary format locks into one vendor.
|
||||||
|
|||||||
@@ -5,14 +5,31 @@ Instead of letting Claude Code do whatever it wants, hyperguild enforces structu
|
|||||||
workflows (TDD red/green/refactor), logs every session, and accumulates learnings
|
workflows (TDD red/green/refactor), logs every session, and accumulates learnings
|
||||||
into a searchable brain.
|
into a searchable brain.
|
||||||
|
|
||||||
|
## Hypothesis
|
||||||
|
|
||||||
|
> We believe routing skill tasks through local models, backed by brain context,
|
||||||
|
> produces measurably better outcomes than raw Claude Code alone —
|
||||||
|
> measurable by per-skill pass rate over rolling 30-day windows
|
||||||
|
> (available at `GET /pass-rate?skill=<name>&window=30d` on the brain pod).
|
||||||
|
|
||||||
|
This is the falsifiable claim the routing pod and pass-rate infrastructure exist to test.
|
||||||
|
If per-skill pass rates don't improve over baseline (all-cloud) after 30 days of real
|
||||||
|
usage, the fast-model routing path should be reconsidered.
|
||||||
|
|
||||||
|
## Harness
|
||||||
|
|
||||||
|
**hyperguild = Claude Code + MCP.** This is a supervisor for Claude Code sessions specifically.
|
||||||
|
For multi-agent orchestration (OpenCode + LiteLLM, executor/reviewer pipelines), see
|
||||||
|
[agentsquad](http://gitea.d-ma.be/mathias/agentsquad) — a separate harness for a different
|
||||||
|
orchestration model. Skills (mathias/skills) are shared between both.
|
||||||
|
|
||||||
## How it works
|
## How it works
|
||||||
|
|
||||||
```
|
```
|
||||||
Your Claude Code session (in any project)
|
Your Claude Code session (in any project)
|
||||||
│
|
│
|
||||||
│ MCP over HTTP (Tailscale)
|
│ MCP over HTTP (Tailscale)
|
||||||
├──▶ supervisor :3200 (NodePort 30320 on koala) — skill workers: tdd, debug, spec, …
|
├──▶ routing :3210 (NodePort 30310 on koala) — review, debug, retrospective, trainer
|
||||||
├──▶ routing :3210 (NodePort 30310 on koala) — Mode 2 only: review, debug, retrospective, trainer
|
|
||||||
└──▶ brain :3300 (NodePort 30330 on koala) — brain_query, brain_write, brain_ingest, session_log
|
└──▶ brain :3300 (NodePort 30330 on koala) — brain_query, brain_write, brain_ingest, session_log
|
||||||
│
|
│
|
||||||
└─ also serves the legacy REST endpoints (/query, /write, /ingest, …)
|
└─ also serves the legacy REST endpoints (/query, /write, /ingest, …)
|
||||||
@@ -20,34 +37,28 @@ Your Claude Code session (in any project)
|
|||||||
▼
|
▼
|
||||||
brain/
|
brain/
|
||||||
├── sessions/ — JSONL log, one file per session_id
|
├── sessions/ — JSONL log, one file per session_id
|
||||||
├── wiki/ — searchable knowledge (full-text)
|
├── wiki/ — searchable knowledge (wing/hall layout)
|
||||||
│ ├── concepts/
|
│ ├── homelab/
|
||||||
│ ├── entities/
|
│ ├── claude-sessions/
|
||||||
│ └── sources/
|
│ └── ...
|
||||||
├── raw/ — retrospective output, staged for review
|
└── knowledge/ — legacy flat notes (migration pending: hyperguild#22)
|
||||||
└── training-data/ — SFT/DPO/RL data (Phase 2)
|
|
||||||
```
|
```
|
||||||
|
|
||||||
## Phase 1 tools (available now)
|
## Phase 1 tools (available now)
|
||||||
|
|
||||||
| Tool | What it does |
|
| Tool | What it does |
|
||||||
|------|-------------|
|
|------|-------------|
|
||||||
| `tdd_red` | Writes a failing test for a spec, verifies it fails |
|
|
||||||
| `tdd_green` | Writes the minimal implementation to make tests pass |
|
|
||||||
| `tdd_refactor` | Cleans up implementation while keeping tests green |
|
|
||||||
| `session_log` | Appends a structured entry to the session JSONL log |
|
| `session_log` | Appends a structured entry to the session JSONL log |
|
||||||
| `retrospective` | Reads the session log, identifies novel learnings, writes to brain/raw/ |
|
| `retrospective` | Reads the session log, identifies novel learnings, writes to brain |
|
||||||
|
| `review` | Structured code review via local model, brain-context injected |
|
||||||
|
| `debug` | Hypothesis-driven debugging via local model |
|
||||||
| `brain_query` | Full-text search over brain/wiki/ |
|
| `brain_query` | Full-text search over brain/wiki/ |
|
||||||
| `brain_write` | Writes a note to brain/raw/ (with optional YAML frontmatter) |
|
| `brain_write` | Writes a note to brain (with wing/hall routing) |
|
||||||
|
| `brain_answer` | BM25 + LLM synthesis — Q&A over brain corpus |
|
||||||
| `tier` | Returns the current connectivity tier (1=cloud, 2=LAN, 3=offline) |
|
| `tier` | Returns the current connectivity tier (1=cloud, 2=LAN, 3=offline) |
|
||||||
|
|
||||||
## Start the servers
|
> **Note:** `tdd_red/green/refactor` and `spec` were retired in Plan 7 (2026-05-12).
|
||||||
|
> They are now SKILL.md files in [mathias/skills](http://gitea.d-ma.be/mathias/skills).
|
||||||
```bash
|
|
||||||
# Requires goreman: go install github.com/mattn/goreman@latest
|
|
||||||
task start # starts ingestion (:3300) + supervisor (:3200) via goreman
|
|
||||||
task stop # kills both by port
|
|
||||||
```
|
|
||||||
|
|
||||||
## Connect a project
|
## Connect a project
|
||||||
|
|
||||||
@@ -56,9 +67,9 @@ Create `.mcp.json` in your project root:
|
|||||||
```json
|
```json
|
||||||
{
|
{
|
||||||
"mcpServers": {
|
"mcpServers": {
|
||||||
"supervisor": {
|
"routing": {
|
||||||
"type": "http",
|
"type": "http",
|
||||||
"url": "http://koala:30320/mcp"
|
"url": "http://koala:30310/mcp"
|
||||||
},
|
},
|
||||||
"brain": {
|
"brain": {
|
||||||
"type": "http",
|
"type": "http",
|
||||||
@@ -68,33 +79,29 @@ Create `.mcp.json` in your project root:
|
|||||||
}
|
}
|
||||||
```
|
```
|
||||||
|
|
||||||
Two MCP servers are exposed today, both reachable over Tailscale:
|
Two MCP servers are exposed, both reachable over Tailscale:
|
||||||
|
|
||||||
- **`supervisor`** at `koala:30320` — skill workers (`tdd_red/green/refactor`,
|
- **`routing`** at `koala:30310` — skill workers (`review`, `debug`, `retrospective`, `trainer`).
|
||||||
`review`, `debug`, `spec`, `retrospective`, `trainer`, `tier`).
|
Routes each call to fast local model or thinking model based on per-skill pass rate.
|
||||||
- **`brain`** at `koala:30330` — knowledge access (`brain_query`, `brain_write`,
|
- **`brain`** at `koala:30330` — knowledge access (`brain_query`, `brain_write`,
|
||||||
`brain_ingest`, `brain_ingest_raw`) and `session_log`. Hosted by the ingestion
|
`brain_ingest`, `brain_ingest_raw`, `brain_answer`, `brain_classify`) and `session_log`.
|
||||||
service directly, no separate pod.
|
|
||||||
|
|
||||||
No local binary or stdio shim is required — Claude Code talks to both via HTTP.
|
No local binary or stdio shim is required — Claude Code talks to both via HTTP.
|
||||||
|
|
||||||
Open Claude Code in your project — run `/mcp` to confirm both servers are listed.
|
Open Claude Code in your project — run `/mcp` to confirm both servers are listed.
|
||||||
|
|
||||||
## A typical TDD session
|
## A typical session
|
||||||
|
|
||||||
```
|
```
|
||||||
1. Call tdd_red → spec in, failing test file out
|
1. Call review → brain context injected + local model review → findings
|
||||||
2. Call tdd_green → test path in, implementation out
|
2. Call session_log → log each phase result
|
||||||
3. Call tdd_refactor → impl + test in, cleaned code out
|
3. Call retrospective → extracts learnings → brain
|
||||||
4. Call session_log → log each phase result
|
4. Future sessions: call brain_query / brain_answer to retrieve relevant context
|
||||||
5. Call retrospective → extracts learnings → brain/raw/
|
|
||||||
6. Review brain/raw/, move worthy notes to brain/wiki/concepts/
|
|
||||||
7. Future sessions: call brain_query to retrieve relevant context
|
|
||||||
```
|
```
|
||||||
|
|
||||||
## Tier detection
|
## Tier detection
|
||||||
|
|
||||||
The supervisor probes connectivity at call time:
|
The routing pod probes connectivity at call time:
|
||||||
|
|
||||||
| Tier | Label | Condition |
|
| Tier | Label | Condition |
|
||||||
|------|-------|-----------|
|
|------|-------|-----------|
|
||||||
@@ -102,31 +109,47 @@ The supervisor probes connectivity at call time:
|
|||||||
| 2 | lan-only | Can reach LiteLLM but not Anthropic |
|
| 2 | lan-only | Can reach LiteLLM but not Anthropic |
|
||||||
| 3 | airplane | No external connectivity |
|
| 3 | airplane | No external connectivity |
|
||||||
|
|
||||||
|
## Model routing
|
||||||
|
|
||||||
|
The routing pod selects models per skill call based on historical pass rate:
|
||||||
|
|
||||||
|
| Pass rate | Decision |
|
||||||
|
|-----------|----------|
|
||||||
|
| ≥ 0.90 (FLOOR) | Fast model (`HYPERGUILD_FAST_MODEL`) |
|
||||||
|
| ≤ 0.70 (CEIL) | Thinking model (`HYPERGUILD_THINKING_MODEL`) |
|
||||||
|
| between CEIL and FLOOR | Sample band — probabilistic routing |
|
||||||
|
| nil (no history yet) | Defaults to thinking model |
|
||||||
|
|
||||||
|
> **Bootstrap note:** With no session history, all calls route to the thinking model.
|
||||||
|
> The fast-model path activates only after real pass-rate data accumulates at `/pass-rate`.
|
||||||
|
> Seed with real usage — don't try to pre-populate.
|
||||||
|
|
||||||
## Key env vars
|
## Key env vars
|
||||||
|
|
||||||
| Variable | Default | Purpose |
|
| Variable | Default | Purpose |
|
||||||
|----------|---------|---------|
|
|----------|---------|---------|
|
||||||
| `INGEST_BRAIN_DIR` | `../brain` | Brain directory for ingestion server |
|
| `INGEST_BRAIN_DIR` | `../brain` | Brain directory for ingestion server |
|
||||||
| `INGEST_PORT` | `3300` | Ingestion server port |
|
| `INGEST_PORT` | `3300` | Ingestion server port |
|
||||||
| `SUPERVISOR_CONFIG_DIR` | `./config/supervisor` | Skill discipline files |
|
| `INGEST_BASE_URL` | `http://localhost:3300` | Routing pod → brain |
|
||||||
| `SUPERVISOR_SESSIONS_DIR` | `./brain/sessions` | JSONL session logs |
|
|
||||||
| `INGEST_BASE_URL` | `http://localhost:3300` | Supervisor → ingestion |
|
|
||||||
| `LITELLM_BASE_URL` | — | LiteLLM proxy for Tier 2 model routing |
|
| `LITELLM_BASE_URL` | — | LiteLLM proxy for Tier 2 model routing |
|
||||||
| `SUPERVISOR_MCP_TOKEN` | — | Optional bearer token for the supervisor MCP HTTP endpoint; when empty, no auth is enforced |
|
|
||||||
| `ROUTING_PORT` | `3210` | Routing pod's listen port |
|
| `ROUTING_PORT` | `3210` | Routing pod's listen port |
|
||||||
| `ROUTING_MCP_TOKEN` | — | Optional bearer token for the routing MCP HTTP endpoint |
|
| `ROUTING_MCP_TOKEN` | — | Optional bearer token; when empty, no auth enforced |
|
||||||
| `BRAIN_URL` | `http://ingestion.supervisor:3300` | Routing pod → brain (in-cluster) |
|
| `BRAIN_URL` | `http://ingestion.supervisor:3300` | Routing pod → brain (in-cluster) |
|
||||||
| `HYPERGUILD_FAST_MODEL` | `koala/qwen35-9b-fast` | Fast model for high-pass-rate skill calls |
|
| `HYPERGUILD_FAST_MODEL` | `koala/qwen35-9b-fast` | Fast model for high-pass-rate skill calls |
|
||||||
| `HYPERGUILD_THINKING_MODEL` | `iguana/gemma4-26b` | Thinking model for low-pass-rate skill calls |
|
| `HYPERGUILD_THINKING_MODEL` | `iguana/gemma4-26b` | Thinking model for low-pass-rate skill calls |
|
||||||
| `HYPERGUILD_ROUTE_LOCAL_FLOOR` | `0.90` | At/above pass rate, route to fast model |
|
| `HYPERGUILD_ROUTE_LOCAL_FLOOR` | `0.90` | Fast model threshold |
|
||||||
| `HYPERGUILD_ROUTE_LOCAL_CEIL` | `0.70` | Below pass rate, route to thinking model. Between CEIL and FLOOR is the sample band. |
|
| `HYPERGUILD_ROUTE_LOCAL_CEIL` | `0.70` | Thinking model threshold |
|
||||||
| `HYPERGUILD_PASS_RATE_TTL_SECONDS` | `60` | Per-skill pass-rate cache TTL |
|
| `HYPERGUILD_PASS_RATE_TTL_SECONDS` | `60` | Per-skill pass-rate cache TTL |
|
||||||
|
|
||||||
> **Operator note:** LiteLLM at `LITELLM_BASE_URL` must register both `HYPERGUILD_FAST_MODEL` and `HYPERGUILD_THINKING_MODEL` for routing to do useful work. If a model is missing, LiteLLM returns 4xx, the routing pod's fast route fails, the fail-open retry on the thinking model likely also fails (since both are missing), and the only signal is `final_status: "fail"` on `_routing` entries in the brain.
|
> **Operator note:** LiteLLM at `LITELLM_BASE_URL` must register both `HYPERGUILD_FAST_MODEL`
|
||||||
|
> and `HYPERGUILD_THINKING_MODEL`. If a model is missing, the fail-open retry also fails and
|
||||||
|
> the only signal is `final_status: "fail"` on `_routing` entries in the brain.
|
||||||
|
|
||||||
## Phase 2 (planned)
|
## Open issues
|
||||||
|
|
||||||
- `review` skill — structured code review with iron law enforcement
|
See [issues](http://gitea.d-ma.be/mathias/hyperguild/issues) — key open items:
|
||||||
- `debug` skill — hypothesis-driven debugging sessions
|
|
||||||
- `spec` skill — generates specs from conversations
|
- **#25** — skills platform overhaul (audit first, then lazy loading + brain feedback loop)
|
||||||
- `trainer` — extracts SFT/DPO pairs from session logs for fine-tuning
|
- **#24** — reduce context burn from skill listing
|
||||||
|
- **#22** — migrate legacy brain notes to wing/hall layout (one-shot script, low risk)
|
||||||
|
- **#31** — connect routing-mcp to claude.ai as custom connector
|
||||||
|
|||||||
+2
-4
@@ -17,8 +17,6 @@ tasks:
|
|||||||
cmds: [bash scripts/context-sync.sh claude]
|
cmds: [bash scripts/context-sync.sh claude]
|
||||||
context:sync:agents:
|
context:sync:agents:
|
||||||
cmds: [bash scripts/context-sync.sh agents]
|
cmds: [bash scripts/context-sync.sh agents]
|
||||||
context:sync:cursor:
|
|
||||||
cmds: [bash scripts/context-sync.sh cursor]
|
|
||||||
|
|
||||||
# ── Development ────────────────────────────────────────────────────────────
|
# ── Development ────────────────────────────────────────────────────────────
|
||||||
|
|
||||||
@@ -90,12 +88,12 @@ tasks:
|
|||||||
cmds:
|
cmds:
|
||||||
- task: context:sync
|
- task: context:sync
|
||||||
- cmd: |
|
- cmd: |
|
||||||
drift=$(git status --porcelain -- AGENTS.md CLAUDE.md .cursorrules .aider.conventions.md .context/system-prompt.txt 2>/dev/null)
|
drift=$(git status --porcelain -- AGENTS.md CLAUDE.md .context/system-prompt.txt 2>/dev/null)
|
||||||
if [ -n "$drift" ]; then
|
if [ -n "$drift" ]; then
|
||||||
echo "ERROR: derived adapters drifted from canonical context." >&2
|
echo "ERROR: derived adapters drifted from canonical context." >&2
|
||||||
echo "$drift" >&2
|
echo "$drift" >&2
|
||||||
echo "" >&2
|
echo "" >&2
|
||||||
echo "Run: git add AGENTS.md CLAUDE.md .cursorrules .aider.conventions.md .context/system-prompt.txt" >&2
|
echo "Run: git add AGENTS.md CLAUDE.md .context/system-prompt.txt" >&2
|
||||||
echo " git commit -m 'chore: re-sync context adapters'" >&2
|
echo " git commit -m 'chore: re-sync context adapters'" >&2
|
||||||
exit 1
|
exit 1
|
||||||
fi
|
fi
|
||||||
|
|||||||
@@ -0,0 +1,48 @@
|
|||||||
|
{"_meta":true,"note":"Agent-consumer column of brain-MCP intent analysis. consumer_type fixed=autonomous_agent. CAVEAT: canonical schema file brain-intent-extraction.md is NOT present on this host (koala) — only this session's own task prompt references it. The closed intent vocabulary below was RECONSTRUCTED from the task prompt's framing + brain/schema.md. Re-map intent labels if the canonical vocab differs. schema_source=reconstructed on every row.","closed_intent_vocab":["semantic_retrieval","lexical_lookup","check_prior_art","synthesized_answer","store_new_knowledge","update_or_supersede","ingest_raw_source","verify_write_landed","discover_capability","intent_unclear"],"intent_tool_match_values":["match","mismatch","partial"],"corpus":"~/.claude/projects/*/*.jsonl (Claude Code agent transcripts on koala). brain/sessions/*.jsonl empty. agentsquad docs/eval/*.jsonl are code-review eval results, NOT brain calls. No separate Crush logs found. Zero brain calls appear under any mcp__ name with a human typing the call — all brain acts are agent-initiated (CLAUDE.md reflex), so all qualify as autonomous_agent."}
|
||||||
|
{"id":"a01","session":"tapir-c","ts":"2026-06-?T15:01:51","tool":"brain_query","intent":"check_prior_art","intent_tool_match":"match","workaround":null,"observed_friction":null,"evidence":"single pre-task query 'YouTube Data API captions download ownership limitation timedtext adapter Go' — named-entity lexical lookup, fit BM25 well, no reformulation."}
|
||||||
|
{"id":"a02","session":"tapir","ts":"2026-06-05T21:44:10","tool":"brain_ingest","intent":"ingest_raw_source","intent_tool_match":"match","workaround":null,"observed_friction":"preceded by ToolSearch select:brain_ingest 14s earlier — tool not ambient, had to be discovered/loaded first.","evidence":"source=tapir-scheduled-discovery-session-2026-06-05, a session learnings dump."}
|
||||||
|
{"id":"a03","session":"tapir","ts":"2026-06-?T14:00:59","tool":"brain_ingest","intent":"ingest_raw_source","intent_tool_match":"match","workaround":null,"observed_friction":null,"evidence":"source=tapir-rls-identity-bootstrapping, first write of RLS lesson."}
|
||||||
|
{"id":"a04","session":"tapir","ts":"2026-06-?T14:02:08","tool":"brain_ingest","intent":"update_or_supersede","intent_tool_match":"mismatch","workaround":"RE-INGESTED same source name 'tapir-rls-identity-bootstrapping' 69s later with edited/condensed body. No update/patch/supersede verb exists, so the agent overwrote-by-re-ingest. Whether this dedups or creates a v2 duplicate is opaque to the agent.","observed_friction":"agent revised content within 70s of first write — classic edit-after-write with no edit primitive.","evidence":"two brain_ingest, identical source string, divergent content."}
|
||||||
|
{"id":"a05","session":"tapir","ts":"2026-06-?T14:56:27","tool":"brain_write","intent":"store_new_knowledge","intent_tool_match":"match","workaround":null,"observed_friction":null,"evidence":"postgres-cascade-skips-tables-without-fk.md — distinct new lesson."}
|
||||||
|
{"id":"a06","session":"tapir","ts":"2026-06-?T21:08:57","tool":"brain_query","intent":"check_prior_art","intent_tool_match":"match","workaround":null,"observed_friction":"preceded by ToolSearch select:brain_query — discovery tax again.","evidence":"'Dex passwords.dex.coreos.com CRD ...' keyword-rich, single shot."}
|
||||||
|
{"id":"a07","session":"AI-infra","ts":"2026-05-?T15:38:03","tool":"brain_query(HTTP-curl)","intent":"discover_capability","intent_tool_match":"mismatch","workaround":"raw `curl -X POST` to brain-mcp endpoint instead of MCP tool. Preceded by two ToolSearch ('brain knowledge memory' then 'brain') that did not yield a usable loaded tool, so agent fell back to HTTP.","observed_friction":"3-step ladder: ToolSearch 'brain knowledge memory' -> ToolSearch 'brain' -> curl. Agent did not know which act maps to which tool name.","evidence":"curl -s -o /tmp/brain-init.txt -w code:%{http_code} -X POST ..."}
|
||||||
|
{"id":"a08","session":"AI-infra","ts":"2026-05-?T04:46:13","tool":"brain_query(HTTP-curl)","intent":"discover_capability","intent_tool_match":"mismatch","workaround":"hand-set TOKEN=... then curl brain-test endpoint — probing whether the HTTP brain path is reachable/authed at all. MCP path not used.","observed_friction":"agent testing connectivity by hand; MCP auth/availability not trusted.","evidence":"TOKEN=...; curl -s -o /tmp/brain-test ..."}
|
||||||
|
{"id":"a09","session":"AI-infra","ts":"2026-05-?T05:23:35","tool":"brain_query","intent":"check_prior_art","intent_tool_match":"match","workaround":null,"observed_friction":"preceded by ToolSearch select:brain_query (after an earlier 05:03 ToolSearch 'brain ingestion knowledge wiki' that explored layers).","evidence":"'koala machine state RTX 5070 llama-swap' — named-entity recall, fits lexical."}
|
||||||
|
{"id":"a10","session":"AI-infra","ts":"2026-05-?T05:27:36","tool":"brain_query","intent":"check_prior_art","intent_tool_match":"match","workaround":null,"observed_friction":null,"evidence":"batch of 3 in ~1s (k3s/flux gitops; llama-swap ai-stack GPU; MCP Dex OAuth claude.ai) — parallel prior-art sweep, all named-entity."}
|
||||||
|
{"id":"a11","session":"AI-infra","ts":"2026-05-?T07:13:50","tool":"brain_query","intent":"check_prior_art","intent_tool_match":"match","workaround":null,"observed_friction":null,"evidence":"batch of 4 in ~2s before a debugging session (flux healthCheck; exit 255 restart loop; NVML mismatch; mirror rebase). Named symptoms, lexical fit OK on first pass."}
|
||||||
|
{"id":"a12","session":"AI-infra","ts":"2026-05-?T07:14:26","tool":"brain_write","intent":"store_new_knowledge","intent_tool_match":"match","workaround":null,"observed_friction":null,"evidence":"4 writes in ~35s (flux-healthcheck-stale; exit-255-unknown-reason-not-oom; nvidia-nvml-mismatch; mcp-static-bearer) — answers to the 4 queries just run, captured as lessons. Healthy query->fix->write loop."}
|
||||||
|
{"id":"a13","session":"AI-infra","ts":"2026-05-?T07:15:25","tool":"brain_query","intent":"verify_write_landed","intent_tool_match":"mismatch","workaround":"25s after writing exit-255-unknown-reason-not-oom.md, re-queried 'exit 255 unknown SIGKILL containerd' — reformulated terms (SIGKILL/containerd not in original query 'exit 255 unknown reason restart loop diagnosis'). Either confirming the fresh write is retrievable or re-searching because first lexical query missed. No read-after-write / get-by-id act exists.","observed_friction":"reformulation chain: 'exit 255 unknown reason restart loop diagnosis' -> 'exit 255 unknown SIGKILL containerd'. Same need, different keywords.","evidence":"query at 07:13:51 vs 07:15:25 bracketing the 07:14:37 write."}
|
||||||
|
{"id":"a14","session":"AI-infra","ts":"2026-05-?T07:22:55","tool":"brain_query","intent":"check_prior_art","intent_tool_match":"match","workaround":null,"observed_friction":null,"evidence":"batch of 5 in ~20s before a homelab security audit (piguard/iguana tailscale; unifi UCG firewall; SOPS age; ingress TLS cert-manager; koala UFW iptables). Named-entity sweep."}
|
||||||
|
{"id":"a15","session":"AI-infra","ts":"2026-05-?T09:27:53","tool":"brain_write","intent":"store_new_knowledge","intent_tool_match":"match","workaround":null,"observed_friction":null,"evidence":"audit-shortcut-tls-blocks-zero; policy-audit-mode-blocks-nothing — distinct new audit lessons."}
|
||||||
|
{"id":"a16","session":"AI-infra","ts":"2026-05-?T09:28:22","tool":"brain_write","intent":"store_new_knowledge","intent_tool_match":"match","workaround":null,"observed_friction":null,"evidence":"homelab-security-chains-not-bugs.md FIRST write (worked example: koala 2026-05-13)."}
|
||||||
|
{"id":"a17","session":"AI-infra","ts":"2026-05-?T10:32:49","tool":"brain_write","intent":"update_or_supersede","intent_tool_match":"mismatch","workaround":"RE-WROTE homelab-security-chains-not-bugs.md ~64min later with a different/expanded worked example (host-user dotfile, over-broad ClusterRole). Same filename, additive revision, no patch/append/supersede verb — agent overwrites and hopes the index replaces rather than duplicates.","observed_friction":"the in-between hour of audit work produced a better example; only way to fold it in was a full re-write of the same slug.","evidence":"two brain_write same filename at 09:28:22 and 10:32:49, divergent worked examples."}
|
||||||
|
{"id":"a18","session":"AI-infra","ts":"2026-05-?T10:18:25","tool":"brain_write","intent":"store_new_knowledge","intent_tool_match":"match","workaround":null,"observed_friction":null,"evidence":"homelab-document-accepted-risk-to-break-audit-cycle.md — distinct."}
|
||||||
|
{"id":"a19","session":"AI-infra","ts":"2026-05-?T10:32:55","tool":"brain_query","intent":"verify_write_landed","intent_tool_match":"mismatch","workaround":"6s after the homelab-chains re-write, queried 'RBAC MCP cluster pods log chain' — checking the chain reasoning is retrievable / finding the related entry. Read-after-write done via lexical search.","observed_friction":null,"evidence":"query immediately follows the 10:32:49 write."}
|
||||||
|
{"id":"a20","session":"AI-infra","ts":"2026-05-?T18:57:34","tool":"brain_query","intent":"check_prior_art","intent_tool_match":"match","workaround":null,"observed_friction":"preceded by ToolSearch select:brain_write,brain_query — re-discovered tools this session.","evidence":"'extension build pinned version major version upgrade postgres pgvector'."}
|
||||||
|
{"id":"a21","session":"AI-infra","ts":"2026-05-?T18:57:58","tool":"brain_write","intent":"store_new_knowledge","intent_tool_match":"match","workaround":null,"observed_friction":null,"evidence":"extension-version-lags-platform-major-upgrade.md."}
|
||||||
|
{"id":"a22","session":"AI-infra","ts":"2026-05-?T18:58:06","tool":"brain_query","intent":"verify_write_landed","intent_tool_match":"mismatch","workaround":"8s after writing extension-version-lags, re-queried 'pgvector postgres extension version compile error bump' — reformulated from the 18:57:34 query ('extension build pinned version...'). Lexical re-search to confirm the just-written lesson is findable, with different keyword guess.","observed_friction":"reformulation: 'extension build pinned version major version upgrade postgres pgvector' -> 'pgvector postgres extension version compile error bump'.","evidence":"write at 18:57:58 bracketed by queries 18:57:34 and 18:58:06."}
|
||||||
|
{"id":"a23","session":"AI-infra","ts":"2026-05-?T18:31:23","tool":"brain_write","intent":"store_new_knowledge","intent_tool_match":"match","workaround":null,"observed_friction":null,"evidence":"webfetch-readme-when-image-or-flag-uncertain.md FIRST write."}
|
||||||
|
{"id":"a24","session":"AI-infra","ts":"2026-05-?T18:34:23","tool":"brain_write","intent":"update_or_supersede","intent_tool_match":"mismatch","workaround":"RE-WROTE webfetch-readme-when-image-or-flag-uncertain.md 3min later, near-identical body. Looks like a retry/overwrite (uncertain the first landed, or minor edit). No idempotent upsert with confirmation, so agent re-fires the write.","observed_friction":"followed 7s later by a brain_query on the same topic ('OSS tool image registry CLI flag webhook path schema drift README pre-flight') — write-write-query, i.e. overwrite then verify-by-search.","evidence":"two brain_write same filename 18:31:23 / 18:34:23, then query 18:34:31."}
|
||||||
|
{"id":"a25","session":"AI-infra","ts":"2026-05-?T18:34:31","tool":"brain_query","intent":"verify_write_landed","intent_tool_match":"mismatch","workaround":"keyword-stuffed lexical query 'OSS tool image registry CLI flag webhook path schema drift README pre-flight' fired right after the webfetch-readme write — agent dumps every concept token hoping BM25 surfaces its own fresh note. This is semantic intent (find that conceptual lesson) coerced into a bag-of-keywords.","observed_friction":"query is a concatenation of the note's section headings — a tell that the agent is groping lexically for content it knows by meaning.","evidence":"query text mirrors the just-written note's bullet topics."}
|
||||||
|
{"id":"a26","session":"dev","ts":"2026-06-?T21:18:26","tool":"brain_answer","intent":"synthesized_answer","intent_tool_match":"match","workaround":null,"observed_friction":"preceded by ToolSearch select:brain_answer.","evidence":"'tapir transcript persistence shared cross-user dedup table RLS isolation ADR-021 ...' -> 22min later a brain_write (acted on the answer). Answer consumed, not re-queried. Healthy."}
|
||||||
|
{"id":"a27","session":"dev","ts":"2026-06-?T21:40:20","tool":"brain_write","intent":"store_new_knowledge","intent_tool_match":"match","workaround":null,"observed_friction":null,"evidence":"tapir-migration-and-rls-test-infra-gotchas, with wing/hall absent here (flat) — see schema-confusion note a40."}
|
||||||
|
{"id":"a28","session":"dev","ts":"2026-06-?T05:35:04","tool":"brain_answer","intent":"synthesized_answer","intent_tool_match":"partial","workaround":"asked 'gitea MCP not working workaround file issue via API which token ... how to authenticate gitea API' — a how-do-I question. Next brain act (05:39 query) is a different topic (tapir transcript), so the answer was apparently sufficient OR abandoned; ambiguous.","observed_friction":null,"evidence":"brain_answer then unrelated brain_query 4min later."}
|
||||||
|
{"id":"a29","session":"dev","ts":"2026-06-?T05:39:54","tool":"brain_query","intent":"check_prior_art","intent_tool_match":"match","workaround":null,"observed_friction":null,"evidence":"'tapir transcript persistence ADR-021 shared non-RLS'."}
|
||||||
|
{"id":"a30","session":"dev","ts":"2026-06-?T05:57:37","tool":"brain_write","intent":"store_new_knowledge","intent_tool_match":"match","workaround":null,"observed_friction":null,"evidence":"gitea-mcp-per-repo-tools-404-and-rest-fallback FIRST write."}
|
||||||
|
{"id":"a31","session":"dev","ts":"2026-06-?T05:57:55","tool":"brain_write","intent":"update_or_supersede","intent_tool_match":"mismatch","workaround":"RE-WROTE gitea-mcp-per-repo-tools-404-and-rest-fallback 18s later — overwrite/retry of same slug, no upsert confirmation.","observed_friction":"sub-20s gap = almost certainly a content tweak the agent could not express as an edit.","evidence":"two brain_write same filename 05:57:37 / 05:57:55."}
|
||||||
|
{"id":"a32","session":"dev","ts":"2026-05-?T11:51:47","tool":"brain_write","intent":"store_new_knowledge","intent_tool_match":"match","workaround":null,"observed_friction":null,"evidence":"infra-litellm-absorption-2026-05-16.md."}
|
||||||
|
{"id":"a33","session":"dev","ts":"2026-05-?T12:07:04","tool":"brain_query","intent":"check_prior_art","intent_tool_match":"match","workaround":null,"observed_friction":null,"evidence":"batch of 3 (litellm rebuild time piguard; docker compose orphaned volumes; prometheus_client ModuleNotFoundError) — lexical, error-string driven."}
|
||||||
|
{"id":"a34","session":"dev","ts":"2026-05-?T15:08:15","tool":"brain_answer","intent":"synthesized_answer","intent_tool_match":"mismatch","workaround":"THREE brain_answer at 15:08 (moved compose volumes? / pi rebuild time? / litellm ModuleNotFound prometheus) — the SAME three topics queried lexically an hour earlier (12:07) — were IMMEDIATELY followed at 15:09 by THREE brain_query on the same three topics. The agent asked the synthesizer, was unsatisfied, and fell straight back to raw lexical search. Strongest answer->query fallback in the corpus.","observed_friction":"answer/query duplication across one intent: agent hedges by firing both interfaces, trusting neither.","evidence":"15:08 answers vs 15:09 queries, topic-for-topic aligned."}
|
||||||
|
{"id":"a35","session":"dev","ts":"2026-05-?T15:09:16","tool":"brain_query","intent":"semantic_retrieval","intent_tool_match":"mismatch","workaround":"after the 3 brain_answer calls failed to satisfy, re-issued as lexical brain_query ('moved compose stack to new directory volumes disappeared empty'; 'raspberry pi docker build time arm slow'; 'how to enable prometheus metrics on litellm proxy callback'). The want is meaning-based ('did my volumes move?') but the only retrieval that 'worked' was keyword search — and these are full natural-language sentences crammed into a BM25 box.","observed_friction":"natural-language questions ('how to enable...', 'moved ... disappeared') passed to a lexical query tool — semantic intent, lexical interface.","evidence":"3 queries at 15:09 mirror the 3 answers at 15:08."}
|
||||||
|
{"id":"a36","session":"dev","ts":"2026-05-?T20:31:50","tool":"brain_answer","intent":"synthesized_answer","intent_tool_match":"match","workaround":null,"observed_friction":null,"evidence":"'What happened with the litellm migration on 2026-05-16?' — episodic recall question, answer fit; no re-query followed."}
|
||||||
|
{"id":"a37","session":"dev","ts":"2026-05-?T21:07:17","tool":"brain_query","intent":"semantic_retrieval","intent_tool_match":"mismatch","workaround":"FOUR-step reformulation chain over one Go bug: 'bytes.Buffer Bytes Reset aliasing slice sharing' -> 'go buffer reuse map backing array bug' -> [write go-bytes-buffer-bytes-reset-aliasing-trap.md] -> 'go map values all show same content after loop' -> 'bytes.Buffer Bytes returns same data every iteration'. The agent knows the SYMPTOM (all map values identical) and the CAUSE (Bytes() aliasing) but cannot phrase a single lexical query that bridges them — it wants concept retrieval and is forced to brute-force keyword variants.","observed_friction":"4 distinct phrasings of the same bug, two before and two after writing the lesson — also doubles as verify_write_landed on the trailing queries.","evidence":"21:07:17, 21:07:17, (write 21:08:01), 21:08:10, 21:08:19."}
|
||||||
|
{"id":"a38","session":"dev","ts":"2026-05-?T21:08:01","tool":"brain_write","intent":"store_new_knowledge","intent_tool_match":"match","workaround":null,"observed_friction":null,"evidence":"go-bytes-buffer-bytes-reset-aliasing-trap.md."}
|
||||||
|
{"id":"a39","session":"dev","ts":"2026-05-?T07:54:26","tool":"brain_write","intent":"store_new_knowledge","intent_tool_match":"match","workaround":null,"observed_friction":null,"evidence":"mcp-tool-design-get-needs-list-partner.md — a design principle."}
|
||||||
|
{"id":"a40","session":"dev","ts":"2026-06-?T06:25:00","tool":"brain_query","intent":"check_prior_art","intent_tool_match":"match","workaround":null,"observed_friction":null,"evidence":"'Dex to Authentik migration auth.d-ma.be issuer cutover OIDC subject ...'."}
|
||||||
|
{"id":"a41","session":"dev","ts":"2026-06-?T06:25:08","tool":"brain_answer","intent":"synthesized_answer","intent_tool_match":"mismatch","workaround":"brain_query (a40) and brain_answer (a41) fired ~8s apart on the SAME intent (Dex->Authentik subject-keyed token orphan). Agent runs lexical search AND synthesized answer in parallel for one question rather than choosing — it cannot predict which interface will return usable knowledge, so it pays both.","observed_friction":"query+answer doublet on one need.","evidence":"06:25:00 query then 06:25:08 answer, same topic."}
|
||||||
|
{"id":"a42","session":"dev","ts":"2026-06-?T13:48:49","tool":"brain_query","intent":"verify_write_landed","intent_tool_match":"mismatch","workaround":"'authentik cutover validation probe' then 6min later 'authentik cutover post-flip validation' — reformulated pair, likely searching for the agent's own earlier cutover notes / confirming validation steps are recorded. Lexical re-search standing in for recall-my-recent-context.","observed_friction":"reformulation: 'validation probe' -> 'post-flip validation'.","evidence":"13:48:49 and 13:54:56."}
|
||||||
|
{"id":"a43","session":"dev","ts":"2026-06-?T13:59:17","tool":"brain_write","intent":"store_new_knowledge","intent_tool_match":"match","workaround":null,"observed_friction":"preceded by ToolSearch select:brain_write.","evidence":"oidc-issuer-host-change-vs-idp-swap-subject."}
|
||||||
|
{"id":"a44","session":"dev","ts":"2026-06-?T13:59:50","tool":"brain_write","intent":"store_new_knowledge","intent_tool_match":"match","workaround":null,"observed_friction":null,"evidence":"cannot-move-ingress-host-across-namespaces-flux-dryrun FIRST write."}
|
||||||
|
{"id":"a45","session":"dev","ts":"2026-06-?T14:00:13","tool":"brain_write","intent":"update_or_supersede","intent_tool_match":"mismatch","workaround":"RE-WROTE cannot-move-ingress-host-across-namespaces-flux-dryrun 23s later — overwrite of same slug, no edit/upsert primitive.","observed_friction":"sub-30s gap = content correction expressed as a full re-write.","evidence":"two brain_write same filename 13:59:50 / 14:00:13."}
|
||||||
|
{"id":"a46","session":"dev","ts":"2026-06-?T13:53:54","tool":"brain_write(HTTP-staged)","intent":"store_new_knowledge","intent_tool_match":"mismatch","workaround":"after `ToolSearch select:mcp__claude_ai_brain__authenticate` (MCP auth flow), the agent staged the entry as `cat > /tmp/brain_entry.json` ({filename:'postgres-force-rls-cross-u...', content}) for a curl write rather than calling brain_write directly — MCP write path was not usable (auth/loading), so it dropped to the HTTP bodge.","observed_friction":"reached for an 'authenticate' tool, then abandoned MCP for hand-built JSON + curl. Matches known pattern: brain/op MCP auth lapses often.","evidence":"ToolSearch authenticate 13:53:27 -> cat /tmp/brain_entry.json 13:53:54."}
|
||||||
|
{"id":"a47","session":"template-go-agent","ts":"2026-05-?T18:46:26","tool":"BASH(not-a-brain-act)","intent":"intent_unclear","intent_tool_match":"match","workaround":null,"observed_friction":null,"evidence":"'brain_' substring was inside a git commit message body ('agent boundaries, network policy, agent s...'), NOT a brain call. Excluded from knowledge-act analysis; logged for audit completeness."}
|
||||||
@@ -0,0 +1,148 @@
|
|||||||
|
# Agent-Consumer Brain Intent Analysis — koala column
|
||||||
|
|
||||||
|
**Consumer:** `autonomous_agent` (all rows). **Host:** koala. **Date:** 2026-06-15.
|
||||||
|
**Raw rows:** `agent-intent-column.jsonl` (46 real knowledge-acts + 1 excluded false-positive).
|
||||||
|
|
||||||
|
## Caveat — canonical schema not on this host
|
||||||
|
|
||||||
|
The shared closed-vocabulary file `brain-intent-extraction.md` **does not exist on
|
||||||
|
koala** — the only reference to it is inside *this task's own prompt*. The intent
|
||||||
|
vocabulary below was **reconstructed** from the prompt's framing + `brain/schema.md`.
|
||||||
|
Every row carries `schema_source: reconstructed`. If the canonical vocab differs,
|
||||||
|
re-map the `intent` field; the `intent_tool_match` / `workaround` / `observed_friction`
|
||||||
|
evidence stands regardless of label names.
|
||||||
|
|
||||||
|
**Reconstructed closed vocab:** `semantic_retrieval`, `lexical_lookup`,
|
||||||
|
`check_prior_art`, `synthesized_answer`, `store_new_knowledge`, `update_or_supersede`,
|
||||||
|
`ingest_raw_source`, `verify_write_landed`, `discover_capability`, `intent_unclear`.
|
||||||
|
|
||||||
|
## Corpus
|
||||||
|
|
||||||
|
- `~/.claude/projects/*/*.jsonl` — Claude Code agent transcripts (98 files). **The only
|
||||||
|
source with brain calls.**
|
||||||
|
- `brain/sessions/*.jsonl` — empty (only `.gitkeep`).
|
||||||
|
- `agentsquad docs/eval/*.jsonl` — code-review eval results, **not** brain calls.
|
||||||
|
- No separate Crush session logs on this host.
|
||||||
|
- **Zero** brain calls were human-typed. Every brain act is agent-initiated (the
|
||||||
|
CLAUDE.md "query as reflex / close-the-loop write" behaviour), so all qualify as
|
||||||
|
`autonomous_agent`. The human gave the top-level task; the agent chose every brain act.
|
||||||
|
|
||||||
|
## 1. Intent histogram, split by `intent_tool_match`
|
||||||
|
|
||||||
|
| intent | match | mismatch | partial | total |
|
||||||
|
|---|---|---|---|---|
|
||||||
|
| check_prior_art | 10 | 0 | 0 | 10 |
|
||||||
|
| store_new_knowledge | 14 | 1 | 0 | 15 |
|
||||||
|
| update_or_supersede | 0 | 5 | 0 | 5 |
|
||||||
|
| synthesized_answer | 2 | 2 | 1 | 5 |
|
||||||
|
| verify_write_landed | 0 | 4 | 0 | 4 |
|
||||||
|
| semantic_retrieval | 0 | 3 | 0 | 3 |
|
||||||
|
| ingest_raw_source | 2 | 0 | 0 | 2 |
|
||||||
|
| discover_capability | 0 | 2 | 0 | 2 |
|
||||||
|
| **total** | **28** | **17** | **1** | **46** |
|
||||||
|
|
||||||
|
> Batch note: several rows collapse a same-second fan-out of identical-intent calls
|
||||||
|
> (a10=3, a11=4, a14=5, a33=3, a34=3, a35=3). Call-level the corpus is ~62 brain calls;
|
||||||
|
> the table counts the 46 distinct knowledge-acts. Frequency is deliberately *not* the
|
||||||
|
> point — the mismatch column is.
|
||||||
|
|
||||||
|
**37% of agent knowledge-acts (17/46) are interface mismatches.** Every mismatch falls
|
||||||
|
into one of four intents: `update_or_supersede`, `verify_write_landed`,
|
||||||
|
`semantic_retrieval`, `discover_capability` — plus one `store` that had to use HTTP.
|
||||||
|
|
||||||
|
## 2. Mismatch list, grouped by intent (primary deliverable)
|
||||||
|
|
||||||
|
### update_or_supersede → re-write same slug (5/5 mismatch) — HIGHEST VALUE
|
||||||
|
There is **no update / patch / append / supersede verb**. When an agent improves a note
|
||||||
|
it already wrote, the only move is to call `brain_write`/`brain_ingest` **again with the
|
||||||
|
same filename/source** and hope the index replaces rather than duplicates. Observed:
|
||||||
|
|
||||||
|
| slug | 1st write | 2nd write | gap | what changed |
|
||||||
|
|---|---|---|---|---|
|
||||||
|
| `tapir-rls-identity-bootstrapping` (ingest) | 14:00:59 | 14:02:08 | 69s | condensed body |
|
||||||
|
| `homelab-security-chains-not-bugs.md` | 09:28:22 | 10:32:49 | 64m | new worked example |
|
||||||
|
| `webfetch-readme-when-image-or-flag-uncertain.md` | 18:31:23 | 18:34:23 | 3m | near-identical (retry) |
|
||||||
|
| `gitea-mcp-per-repo-tools-404-and-rest-fallback` | 05:57:37 | 05:57:55 | 18s | content tweak |
|
||||||
|
| `cannot-move-ingress-host-across-namespaces-flux-dryrun` | 13:59:50 | 14:00:13 | 23s | content tweak |
|
||||||
|
|
||||||
|
Sub-30s gaps (3 of 5) read as "I wanted to edit but can only overwrite." The agent has
|
||||||
|
no way to know whether the second write deduped or created a contradictory v2 — opacity
|
||||||
|
the brain's own design principle (`mcp-tool-design-get-needs-list-partner.md`, written
|
||||||
|
*by one of these very agents*) would flag: every `_write` needs a `_get`/`_update` partner.
|
||||||
|
|
||||||
|
### verify_write_landed → lexical re-query (4/4 mismatch)
|
||||||
|
No read-after-write / get-by-id confirmation. After every substantive write, agents
|
||||||
|
re-query lexically to check the note is retrievable — and *reformulate the keywords*
|
||||||
|
because they can't predict what BM25 indexed:
|
||||||
|
- `exit-255` lesson: query `exit 255 unknown reason restart loop diagnosis` → write →
|
||||||
|
query `exit 255 unknown SIGKILL containerd`.
|
||||||
|
- `extension-version-lags`: query `extension build pinned version...pgvector` → write →
|
||||||
|
query `pgvector postgres extension version compile error bump`.
|
||||||
|
- `webfetch-readme`: write → write → query stuffed with the note's own section headings.
|
||||||
|
|
||||||
|
### semantic_retrieval → BM25 keyword-stuffing (3/3 mismatch)
|
||||||
|
Agent knows the *meaning* but not the *indexed words*, so it brute-forces phrasings of
|
||||||
|
one need against a lexical tool:
|
||||||
|
- **4-step chain on one Go bug:** `bytes.Buffer Bytes Reset aliasing slice sharing` →
|
||||||
|
`go buffer reuse map backing array bug` → (write) → `go map values all show same
|
||||||
|
content after loop` → `bytes.Buffer Bytes returns same data every iteration`. Symptom
|
||||||
|
and cause both known; no single lexical query bridges them.
|
||||||
|
- Natural-language questions (`how to enable prometheus metrics on litellm proxy
|
||||||
|
callback`, `moved compose stack to new directory volumes disappeared empty`) shoved
|
||||||
|
into `brain_query`.
|
||||||
|
|
||||||
|
### synthesized_answer → fall back to / hedge with brain_query (2 mismatch + 1 partial)
|
||||||
|
`brain_answer` is frequently **not trusted as terminal**:
|
||||||
|
- **Strongest signal:** 3× `brain_answer` at 15:08 (compose volumes / pi rebuild time /
|
||||||
|
litellm ModuleNotFound) → 3× `brain_query` at 15:09 on the *same three topics*. The
|
||||||
|
agent asked the synthesizer, was unsatisfied, and immediately re-ran raw search.
|
||||||
|
- Dex→Authentik: `brain_query` and `brain_answer` fired **8s apart on one question** —
|
||||||
|
the agent pays both interfaces because it can't predict which returns usable knowledge.
|
||||||
|
- (Counter-examples exist: `brain_answer` for episodic recall — "what happened with the
|
||||||
|
litellm migration on 2026-05-16?" — was consumed and not re-queried. So `answer`
|
||||||
|
works for *episodic/temporal* recall, fails for *how-do-I / does-X-hold* reasoning.)
|
||||||
|
|
||||||
|
### discover_capability + store-via-HTTP (3 mismatch)
|
||||||
|
brain tools are **not ambient** — they are deferred and must be `ToolSearch`-loaded each
|
||||||
|
session. Agents fumble the discovery (`ToolSearch 'brain knowledge memory'` →
|
||||||
|
`'brain'` → `'brain ingestion knowledge wiki'`) and, when MCP load/auth fails, drop to
|
||||||
|
**raw `curl` against `brain-mcp` / hand-built `/tmp/brain_entry.json`**. One agent even
|
||||||
|
`ToolSearch`-ed an `authenticate` tool, then abandoned MCP for the HTTP bodge — matching
|
||||||
|
the known "brain/op MCP auth lapses too often" footgun.
|
||||||
|
|
||||||
|
### Write-interface / layer schema confusion (cross-cutting)
|
||||||
|
`brain_write` was called with **three different param shapes** in the same corpus:
|
||||||
|
`{filename, type:"lesson", content}`, `{filename, content}` (no type), and
|
||||||
|
`{wing:"tapir", hall:"failures", filename, content}` — plus `brain_ingest {source,
|
||||||
|
content}`. Agents are unsure which verb and which layer (flat slug vs `wing`/`hall`
|
||||||
|
knowledge routing vs raw ingest) a given knowledge-act maps to. This is the
|
||||||
|
`knowledge/ vs wiki/` confusion expressed at the parameter level.
|
||||||
|
|
||||||
|
## 3. `intent_unclear` rate
|
||||||
|
|
||||||
|
**0 / 46 genuine brain acts (0%).** Agent intent is unusually legible because these are
|
||||||
|
Claude Code transcripts: the surrounding task, the query/filename strings, and the
|
||||||
|
write content all disambiguate. One row (`a47`) was tagged `intent_unclear` and
|
||||||
|
**excluded** — its `brain_` substring was inside a git commit message, not a brain call.
|
||||||
|
Example of the only ambiguity that arose: a `brain_answer` on "gitea MCP not working...
|
||||||
|
how to authenticate" followed by an unrelated query — can't tell if the answer satisfied
|
||||||
|
or was abandoned (`partial`, row a28).
|
||||||
|
|
||||||
|
## 4. The single biggest intent↔interface gap
|
||||||
|
|
||||||
|
**The brain offers one write verb and one lexical read verb, but autonomous agents
|
||||||
|
perform four distinct knowledge-acts against them — and three of the four have no fitting
|
||||||
|
interface.** The deepest gap is the **missing update/supersede path**: agents close every
|
||||||
|
task by writing a lesson (the CLAUDE.md ritual), routinely improve it minutes-to-an-hour
|
||||||
|
later, and — having no edit primitive — re-write the same slug blind, unable to tell
|
||||||
|
whether they corrected the entry or forked a contradiction into the index. This compounds
|
||||||
|
with the lexical-only read side: because there is no `get-by-id` or semantic retrieval,
|
||||||
|
agents can't even reliably *find their own just-written note* to check it, so they
|
||||||
|
keyword-stuff reformulated queries and hedge `brain_answer` with parallel `brain_query`.
|
||||||
|
The interface is built for *append-and-keyword-search*; the agents are trying to
|
||||||
|
*curate a living, deduplicated knowledge base*, and the seam between those two shows up
|
||||||
|
as the 5 blind re-writes, 4 read-after-write re-queries, and 3 semantic-as-lexical chains
|
||||||
|
that dominate the mismatch column.
|
||||||
|
|
||||||
|
---
|
||||||
|
*Evidence-only per task scope — no redesign proposed.*
|
||||||
@@ -0,0 +1,140 @@
|
|||||||
|
# Brain-MCP Intent↔Interface Findings — Unified (two-column merge)
|
||||||
|
|
||||||
|
**Status — 2026-06-16**
|
||||||
|
- ✅ **Agent column** filled from `agent-intent-column.jsonl` (46 acts, koala).
|
||||||
|
- ⏳ **Human column** = `PENDING`. Drop the Claude.ai-history analysis into
|
||||||
|
`human-intent-column.jsonl` (same dir, schema below), then fill the `PENDING`
|
||||||
|
cells and the synthesis blocks marked `<<SYNTH>>`.
|
||||||
|
- ⚠️ Canonical `brain-intent-extraction.md` still absent on koala. Vocab below is
|
||||||
|
the **reconstructed** lock both columns must share. If the real file surfaces,
|
||||||
|
re-map `intent` labels in *both* columns identically before merging.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Shared schema (LOCKED — both columns conform)
|
||||||
|
|
||||||
|
Per-call row, JSONL:
|
||||||
|
|
||||||
|
| field | values / form | notes |
|
||||||
|
|---|---|---|
|
||||||
|
| `id` | `a01..` (agent) / `h01..` (human) | column prefix kept distinct |
|
||||||
|
| `session` | string | source session/conversation id |
|
||||||
|
| `ts` | ISO-8601 | best-effort |
|
||||||
|
| `tool` | brain tool name (+ `(HTTP-curl)` / `(HTTP-staged)` suffix for bodges) | |
|
||||||
|
| `intent` | closed vocab ↓ | the knowledge-act WANTED |
|
||||||
|
| `intent_tool_match` | `match` \| `mismatch` \| `partial` | does the called tool fit the want |
|
||||||
|
| `consumer_type` | `autonomous_agent` \| `human_interactive` | fixed per column |
|
||||||
|
| `workaround` | string \| null | the bodge when mismatch — **primary signal** |
|
||||||
|
| `observed_friction` | string \| null | reformulation chains, discovery tax, hedging |
|
||||||
|
| `evidence` | string | excerpt anchoring the classification |
|
||||||
|
| `schema_source` | `reconstructed` | flip to `canonical` if real vocab lands |
|
||||||
|
|
||||||
|
### Closed intent vocab (LOCKED)
|
||||||
|
`semantic_retrieval`, `lexical_lookup`, `check_prior_art`, `synthesized_answer`,
|
||||||
|
`store_new_knowledge`, `update_or_supersede`, `ingest_raw_source`,
|
||||||
|
`verify_write_landed`, `discover_capability`, `intent_unclear`.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Master comparison — by intent
|
||||||
|
|
||||||
|
| intent | agent acts | agent mismatch | human acts | human mismatch | shared gap |
|
||||||
|
|---|---|---|---|---|---|
|
||||||
|
| check_prior_art | 10 | 0% | `PENDING` | `PENDING` | — |
|
||||||
|
| store_new_knowledge | 15 | 7% (1/15) | `PENDING` | `PENDING` | `<<SYNTH>>` |
|
||||||
|
| update_or_supersede | 5 | **100%** (5/5) | `PENDING` | `PENDING` | `<<SYNTH>>` no edit verb |
|
||||||
|
| synthesized_answer | 5 | 40% (2/5)+1 partial | `PENDING` | `PENDING` | `<<SYNTH>>` |
|
||||||
|
| verify_write_landed | 4 | **100%** (4/4) | `PENDING` | `PENDING` | `<<SYNTH>>` no read-after-write |
|
||||||
|
| semantic_retrieval | 3 | **100%** (3/3) | `PENDING` | `PENDING` | `<<SYNTH>>` lexical-only read |
|
||||||
|
| ingest_raw_source | 2 | 0% | `PENDING` | `PENDING` | — |
|
||||||
|
| discover_capability | 2 | **100%** (2/2) | `PENDING` | `PENDING` | agent-specific (ToolSearch/auth)? |
|
||||||
|
| intent_unclear | 0 | — | `PENDING` | `PENDING` | divergence expected ↓ |
|
||||||
|
| **TOTAL** | **46** | **37% (17)** | `PENDING` | `PENDING` | |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Per-intent merged findings
|
||||||
|
|
||||||
|
### update_or_supersede — agent: 5/5 mismatch (highest value)
|
||||||
|
**Agent:** no edit/patch/append verb. Agents re-write same slug blind:
|
||||||
|
`homelab-security-chains-not-bugs.md` (+64m), `tapir-rls-identity-bootstrapping`,
|
||||||
|
`webfetch-readme...`, `gitea-mcp-per-repo-tools-404...`,
|
||||||
|
`cannot-move-ingress-host...` — 3 of 5 sub-30s ("wanted edit, got overwrite").
|
||||||
|
Cannot tell if write deduped or forked a contradiction.
|
||||||
|
**Human:** `PENDING` — *look for: user editing a prior note, asking "update what I
|
||||||
|
saved about X", or expressing frustration that an old fact is stale/duplicated.*
|
||||||
|
**<<SYNTH>>** shared verdict once both filled.
|
||||||
|
|
||||||
|
### verify_write_landed — agent: 4/4 mismatch
|
||||||
|
**Agent:** no `get-by-id`/read-after-write. Agents lexically re-query their own
|
||||||
|
fresh note with reformulated keywords (`exit 255 unknown reason` → `...SIGKILL
|
||||||
|
containerd`; `extension build pinned...` → `pgvector ...compile error bump`).
|
||||||
|
**Human:** `PENDING` — *humans may not exhibit this (they trust the write UI
|
||||||
|
confirmation). If absent in human column, it's an agent-specific gap → flag.*
|
||||||
|
**<<SYNTH>>**.
|
||||||
|
|
||||||
|
### semantic_retrieval — agent: 3/3 mismatch
|
||||||
|
**Agent:** meaning known, indexed words unknown → BM25 keyword-stuffing. 4-step
|
||||||
|
chain on one Go `bytes.Buffer` bug; NL questions shoved into `brain_query`.
|
||||||
|
**Human:** `PENDING` — *humans likely hit this HARDER (they phrase conversationally).
|
||||||
|
Compare reformulation-chain length agent vs human.*
|
||||||
|
**<<SYNTH>>** — likely the strongest cross-consumer overlap.
|
||||||
|
|
||||||
|
### synthesized_answer — agent: 2 mismatch + 1 partial
|
||||||
|
**Agent:** `brain_answer` not trusted terminal — 3 answers → 3 same-topic queries
|
||||||
|
1min later; query+answer fired 8s apart hedging one need. Works for *episodic*
|
||||||
|
recall, fails for *how-do-I / does-X-hold*.
|
||||||
|
**Human:** `PENDING` — *humans may prefer `brain_answer` as primary (chat-native).
|
||||||
|
If human match-rate >> agent, the tool fits humans not agents → key divergence.*
|
||||||
|
**<<SYNTH>>**.
|
||||||
|
|
||||||
|
### store_new_knowledge — agent: 14/15 match
|
||||||
|
**Agent:** healthy, except 1 HTTP-staged bodge when MCP auth lapsed. Also surfaced
|
||||||
|
write-schema confusion: 3 param shapes (`{filename,type}` / `{filename}` /
|
||||||
|
`{wing,hall,filename}`) + `ingest{source}`.
|
||||||
|
**Human:** `PENDING` — *humans rarely write directly; expect low volume.*
|
||||||
|
**<<SYNTH>>**.
|
||||||
|
|
||||||
|
### check_prior_art / ingest_raw_source — agent: 0% mismatch
|
||||||
|
Lexical fits named-entity recall and raw-source capture. **Human:** `PENDING`.
|
||||||
|
|
||||||
|
### discover_capability — agent: 2/2 mismatch (agent-specific)
|
||||||
|
Brain tools deferred → `ToolSearch`-load each session; auth lapse → `curl` bodge.
|
||||||
|
**Likely has NO human analog** (humans get ambient connectors). Candidate for
|
||||||
|
"agent-only gap" bucket. **Human:** `PENDING` to confirm absent.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Cross-consumer divergence — questions to resolve at merge
|
||||||
|
|
||||||
|
1. **intent_unclear rate.** Agent = 0% (transcripts self-document). Human expected
|
||||||
|
higher (conversational, implicit). Big delta = the columns measure legibility
|
||||||
|
differently, not just intent.
|
||||||
|
2. **Where does each consumer's mismatch concentrate?** Agent mismatch is
|
||||||
|
write-side-heavy (supersede + verify-landed = 9/17). Hypothesis: human mismatch
|
||||||
|
is read-side-heavy (semantic + answer). If true → **the interface fails the two
|
||||||
|
consumers at opposite ends.**
|
||||||
|
3. **Agent-only gaps** (`discover_capability`, `verify_write_landed`) vs
|
||||||
|
**shared gaps** (`semantic_retrieval`, `update_or_supersede`). Shared gaps =
|
||||||
|
highest-priority evidence; agent-only = harness/auth issues.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Combined headline — `<<SYNTH>>` (fill when human column lands)
|
||||||
|
|
||||||
|
> Agent-side draft (to be reconciled with human-side):
|
||||||
|
> Brain = append + keyword-search; agents want a curated, dedup'd, self-verifying KB.
|
||||||
|
> Missing update/supersede path + lexical-only reads are the seam. **Open question
|
||||||
|
> for the merge: do humans hit the same read-side wall, making semantic-retrieval the
|
||||||
|
> universal gap — or do agents uniquely suffer the write-side (supersede / verify)
|
||||||
|
> wall that humans sidestep via the chat UI?**
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Drop-in checklist (when human column arrives)
|
||||||
|
1. Place `human-intent-column.jsonl` in this dir; conform to LOCKED schema.
|
||||||
|
2. Fill every `PENDING` cell in master table + per-intent blocks.
|
||||||
|
3. Resolve the 3 divergence questions with evidence.
|
||||||
|
4. Replace each `<<SYNTH>>` with the reconciled verdict; write the combined headline.
|
||||||
|
5. If canonical vocab surfaced: re-map both columns' `intent`, flip `schema_source`.
|
||||||
|
6. Commit as `docs(brain): merge human+agent intent columns`.
|
||||||
@@ -256,6 +256,19 @@ func main() {
|
|||||||
logger.Error("CLAUDE_SESSIONS_DIR set but BRAIN_PG_DSN missing — claudewatcher needs the cursor table")
|
logger.Error("CLAUDE_SESSIONS_DIR set but BRAIN_PG_DSN missing — claudewatcher needs the cursor table")
|
||||||
os.Exit(1)
|
os.Exit(1)
|
||||||
}
|
}
|
||||||
|
// Client-name guard. The env value is a regex alternation
|
||||||
|
// (e.g. "SEB|Mastercard"); we wrap it with word boundaries
|
||||||
|
// and case-insensitive flag so substrings inside longer
|
||||||
|
// identifiers don't false-match. Sourced from a SOPS secret
|
||||||
|
// so client identities never live in source.
|
||||||
|
if clientBlock := os.Getenv("CLAUDE_INGEST_CLIENT_BLOCK"); clientBlock != "" {
|
||||||
|
pattern := `(?i)\b(` + clientBlock + `)\b`
|
||||||
|
if err := claudewatcher.RegisterRule("client-name", pattern); err != nil {
|
||||||
|
logger.Error("claudewatcher client-block rule invalid", "err", err)
|
||||||
|
os.Exit(1)
|
||||||
|
}
|
||||||
|
logger.Info("claudewatcher client-block guard registered")
|
||||||
|
}
|
||||||
cursorStore, cerr := claudewatcher.NewCursorStore(ctx, pgDSN)
|
cursorStore, cerr := claudewatcher.NewCursorStore(ctx, pgDSN)
|
||||||
if cerr != nil {
|
if cerr != nil {
|
||||||
logger.Error("claudewatcher cursor init", "err", cerr)
|
logger.Error("claudewatcher cursor init", "err", cerr)
|
||||||
|
|||||||
@@ -0,0 +1,97 @@
|
|||||||
|
package api
|
||||||
|
|
||||||
|
import "strings"
|
||||||
|
|
||||||
|
// frontmatter is an ordered, line-preserving view of a note's YAML
|
||||||
|
// frontmatter block. It deliberately avoids a full YAML round-trip: the
|
||||||
|
// brain writes flat `key: value` frontmatter by hand, and a yaml.v3
|
||||||
|
// re-marshal would reorder keys and strip comments. Preserving the
|
||||||
|
// original lines verbatim keeps brain_update a surgical edit — only the
|
||||||
|
// keys it manages (updated_at, supersedes, supersede_reason) change.
|
||||||
|
type frontmatter struct {
|
||||||
|
lines []fmLine
|
||||||
|
}
|
||||||
|
|
||||||
|
// fmLine is one frontmatter line. For `key: value` lines, key and value
|
||||||
|
// are populated; for blank lines, comments, or anything that isn't a
|
||||||
|
// simple scalar pair, key is empty and raw holds the line verbatim.
|
||||||
|
type fmLine struct {
|
||||||
|
key string
|
||||||
|
value string
|
||||||
|
raw string
|
||||||
|
}
|
||||||
|
|
||||||
|
// parseFrontmatter splits src into its frontmatter block and body. A
|
||||||
|
// frontmatter block is recognised only when the file opens with a `---`
|
||||||
|
// fence and a closing `---` fence follows. Otherwise the whole input is
|
||||||
|
// the body and the returned frontmatter is empty.
|
||||||
|
func parseFrontmatter(src string) (frontmatter, string) {
|
||||||
|
var fm frontmatter
|
||||||
|
if !strings.HasPrefix(src, "---\n") {
|
||||||
|
return fm, src
|
||||||
|
}
|
||||||
|
rest := src[len("---\n"):]
|
||||||
|
end := strings.Index(rest, "\n---\n")
|
||||||
|
if end < 0 {
|
||||||
|
// Opening fence with no closing fence — treat as bodyless content.
|
||||||
|
return fm, src
|
||||||
|
}
|
||||||
|
block := rest[:end]
|
||||||
|
body := rest[end+len("\n---\n"):]
|
||||||
|
|
||||||
|
for _, line := range strings.Split(block, "\n") {
|
||||||
|
key, val, ok := strings.Cut(line, ":")
|
||||||
|
key = strings.TrimSpace(key)
|
||||||
|
if !ok || key == "" || strings.HasPrefix(strings.TrimSpace(line), "#") {
|
||||||
|
fm.lines = append(fm.lines, fmLine{raw: line})
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
fm.lines = append(fm.lines, fmLine{key: key, value: strings.TrimSpace(val)})
|
||||||
|
}
|
||||||
|
return fm, body
|
||||||
|
}
|
||||||
|
|
||||||
|
// get returns the value for key, or "" if absent.
|
||||||
|
func (f *frontmatter) get(key string) string {
|
||||||
|
for _, l := range f.lines {
|
||||||
|
if l.key == key {
|
||||||
|
return l.value
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return ""
|
||||||
|
}
|
||||||
|
|
||||||
|
// set overrides the value for an existing key in place, or appends a new
|
||||||
|
// `key: value` line when the key is absent.
|
||||||
|
func (f *frontmatter) set(key, value string) {
|
||||||
|
for i := range f.lines {
|
||||||
|
if f.lines[i].key == key {
|
||||||
|
f.lines[i].value = value
|
||||||
|
return
|
||||||
|
}
|
||||||
|
}
|
||||||
|
f.lines = append(f.lines, fmLine{key: key, value: value})
|
||||||
|
}
|
||||||
|
|
||||||
|
// render serialises the frontmatter back into a `---`-fenced block. An
|
||||||
|
// empty frontmatter renders to the empty string so bodies without a
|
||||||
|
// header stay header-less.
|
||||||
|
func (f *frontmatter) render() string {
|
||||||
|
if len(f.lines) == 0 {
|
||||||
|
return ""
|
||||||
|
}
|
||||||
|
var b strings.Builder
|
||||||
|
b.WriteString("---\n")
|
||||||
|
for _, l := range f.lines {
|
||||||
|
if l.key == "" {
|
||||||
|
b.WriteString(l.raw)
|
||||||
|
} else {
|
||||||
|
b.WriteString(l.key)
|
||||||
|
b.WriteString(": ")
|
||||||
|
b.WriteString(l.value)
|
||||||
|
}
|
||||||
|
b.WriteByte('\n')
|
||||||
|
}
|
||||||
|
b.WriteString("---\n")
|
||||||
|
return b.String()
|
||||||
|
}
|
||||||
@@ -0,0 +1,61 @@
|
|||||||
|
package api
|
||||||
|
|
||||||
|
import (
|
||||||
|
"strings"
|
||||||
|
"testing"
|
||||||
|
|
||||||
|
"github.com/stretchr/testify/assert"
|
||||||
|
)
|
||||||
|
|
||||||
|
func TestParseFrontmatterSplitsHeaderAndBody(t *testing.T) {
|
||||||
|
src := "---\nwing: jepa-fx\nhall: facts\ncreated_at: 2026-01-01T00:00:00Z\n---\n# Title\n\nbody text\n"
|
||||||
|
fm, body := parseFrontmatter(src)
|
||||||
|
|
||||||
|
assert.Equal(t, "jepa-fx", fm.get("wing"))
|
||||||
|
assert.Equal(t, "facts", fm.get("hall"))
|
||||||
|
assert.Equal(t, "2026-01-01T00:00:00Z", fm.get("created_at"))
|
||||||
|
assert.Equal(t, "# Title\n\nbody text\n", body)
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestParseFrontmatterNoHeader(t *testing.T) {
|
||||||
|
src := "# Just a body\n\nno frontmatter here\n"
|
||||||
|
fm, body := parseFrontmatter(src)
|
||||||
|
|
||||||
|
assert.Empty(t, fm.lines)
|
||||||
|
assert.Equal(t, src, body)
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestFrontmatterSetOverridesExistingKey(t *testing.T) {
|
||||||
|
fm, _ := parseFrontmatter("---\nwing: a\nupdated_at: old\n---\nbody\n")
|
||||||
|
fm.set("updated_at", "new")
|
||||||
|
|
||||||
|
assert.Equal(t, "new", fm.get("updated_at"))
|
||||||
|
// No duplicate key.
|
||||||
|
assert.Equal(t, 1, strings.Count(fm.render(), "updated_at:"))
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestFrontmatterSetAppendsNewKey(t *testing.T) {
|
||||||
|
fm, _ := parseFrontmatter("---\nwing: a\n---\nbody\n")
|
||||||
|
fm.set("supersedes", "abc123")
|
||||||
|
|
||||||
|
out := fm.render()
|
||||||
|
assert.Contains(t, out, "wing: a")
|
||||||
|
assert.Contains(t, out, "supersedes: abc123")
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestFrontmatterRenderPreservesCustomFields(t *testing.T) {
|
||||||
|
src := "---\nwing: a\nhall: facts\ncustom_field: keep-me\ntags: [x, y]\n---\nbody\n"
|
||||||
|
fm, _ := parseFrontmatter(src)
|
||||||
|
fm.set("updated_at", "2026-06-22T00:00:00Z")
|
||||||
|
|
||||||
|
out := fm.render()
|
||||||
|
assert.Contains(t, out, "custom_field: keep-me")
|
||||||
|
assert.Contains(t, out, "tags: [x, y]")
|
||||||
|
assert.Contains(t, out, "updated_at: 2026-06-22T00:00:00Z")
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestFrontmatterRenderRoundTrips(t *testing.T) {
|
||||||
|
src := "---\nwing: a\nhall: facts\n---\n"
|
||||||
|
fm, _ := parseFrontmatter(src)
|
||||||
|
assert.Equal(t, src, fm.render())
|
||||||
|
}
|
||||||
@@ -0,0 +1,131 @@
|
|||||||
|
package api
|
||||||
|
|
||||||
|
import (
|
||||||
|
"crypto/sha256"
|
||||||
|
"encoding/hex"
|
||||||
|
"fmt"
|
||||||
|
"os"
|
||||||
|
"path/filepath"
|
||||||
|
"strings"
|
||||||
|
"time"
|
||||||
|
|
||||||
|
"github.com/mathiasbq/hyperguild/ingestion/internal/brain"
|
||||||
|
)
|
||||||
|
|
||||||
|
// ContentHash returns the lowercase hex sha256 of b. It is the note's
|
||||||
|
// content_hash handle: brain_write / brain_update return it, brain_get
|
||||||
|
// recomputes it from the file on disk, and brain_update stamps the prior
|
||||||
|
// note's hash into the new note's `supersedes` frontmatter.
|
||||||
|
func ContentHash(b []byte) string {
|
||||||
|
sum := sha256.Sum256(b)
|
||||||
|
return hex.EncodeToString(sum[:])
|
||||||
|
}
|
||||||
|
|
||||||
|
// resolveWithin maps a brainDir-relative path to an absolute path and
|
||||||
|
// guarantees it does not escape brainDir. Returns the cleaned relPath
|
||||||
|
// (forward-slashed) and the absolute path.
|
||||||
|
func resolveWithin(brainDir, relPath string) (rel, abs string, err error) {
|
||||||
|
clean := filepath.Clean("/" + filepath.ToSlash(relPath))
|
||||||
|
rel = strings.TrimPrefix(clean, "/")
|
||||||
|
abs = filepath.Join(brainDir, filepath.FromSlash(rel))
|
||||||
|
check, err := filepath.Rel(brainDir, abs)
|
||||||
|
if err != nil || check == ".." || strings.HasPrefix(check, ".."+string(filepath.Separator)) {
|
||||||
|
return "", "", fmt.Errorf("path %q escapes brain dir", relPath)
|
||||||
|
}
|
||||||
|
return rel, abs, nil
|
||||||
|
}
|
||||||
|
|
||||||
|
// UpdateNoteOptions identifies the note to supersede and supplies its new
|
||||||
|
// body. Path takes precedence; otherwise the target is resolved from
|
||||||
|
// Wing/Hall/Slug via brain.NotePath.
|
||||||
|
type UpdateNoteOptions struct {
|
||||||
|
Path string // brainDir-relative path; takes precedence over wing/hall/slug
|
||||||
|
Wing string
|
||||||
|
Hall string
|
||||||
|
Slug string
|
||||||
|
Content string // new full body (whole-note replace)
|
||||||
|
Reason string // optional; stamped as supersede_reason
|
||||||
|
}
|
||||||
|
|
||||||
|
// UpdateNote supersedes an existing note in place. It replaces the body
|
||||||
|
// with opts.Content, preserves the existing frontmatter (created_at,
|
||||||
|
// wing, hall, and any custom fields), and stamps updated_at, supersedes
|
||||||
|
// (the prior content hash), and supersede_reason (when given).
|
||||||
|
//
|
||||||
|
// It never creates: if the target does not exist, it returns an error so
|
||||||
|
// the caller can fall back to brain_write. Returns the note's relPath,
|
||||||
|
// the new content hash, and the prior content hash.
|
||||||
|
//
|
||||||
|
// Embeddings are NOT refreshed here. The rewritten file's mtime advances,
|
||||||
|
// which the mtime-driven vectorstore.Sync ticker uses to re-embed it on
|
||||||
|
// its next pass — the same out-of-band mechanism brain_write relies on.
|
||||||
|
func UpdateNote(brainDir string, opts UpdateNoteOptions) (relPath, contentHash, priorHash string, err error) {
|
||||||
|
if opts.Content == "" {
|
||||||
|
return "", "", "", fmt.Errorf("content is required")
|
||||||
|
}
|
||||||
|
|
||||||
|
var rel string
|
||||||
|
if opts.Path != "" {
|
||||||
|
rel = opts.Path
|
||||||
|
} else {
|
||||||
|
full, perr := brain.NotePath(brainDir, opts.Wing, opts.Hall, opts.Slug)
|
||||||
|
if perr != nil {
|
||||||
|
return "", "", "", perr
|
||||||
|
}
|
||||||
|
rel, _ = filepath.Rel(brainDir, full)
|
||||||
|
rel = filepath.ToSlash(rel)
|
||||||
|
}
|
||||||
|
|
||||||
|
rel, abs, err := resolveWithin(brainDir, rel)
|
||||||
|
if err != nil {
|
||||||
|
return "", "", "", err
|
||||||
|
}
|
||||||
|
|
||||||
|
prior, err := os.ReadFile(abs)
|
||||||
|
if err != nil {
|
||||||
|
if os.IsNotExist(err) {
|
||||||
|
return "", "", "", fmt.Errorf("note %q does not exist: use brain_write to create", rel)
|
||||||
|
}
|
||||||
|
return "", "", "", fmt.Errorf("read target: %w", err)
|
||||||
|
}
|
||||||
|
priorHash = ContentHash(prior)
|
||||||
|
|
||||||
|
fm, _ := parseFrontmatter(string(prior))
|
||||||
|
fm.set("updated_at", time.Now().UTC().Format(time.RFC3339))
|
||||||
|
fm.set("supersedes", priorHash)
|
||||||
|
if opts.Reason != "" {
|
||||||
|
fm.set("supersede_reason", opts.Reason)
|
||||||
|
}
|
||||||
|
|
||||||
|
out := []byte(fm.render() + opts.Content)
|
||||||
|
if err := os.WriteFile(abs, out, 0o644); err != nil {
|
||||||
|
return "", "", "", fmt.Errorf("write: %w", err)
|
||||||
|
}
|
||||||
|
return rel, ContentHash(out), priorHash, nil
|
||||||
|
}
|
||||||
|
|
||||||
|
// ReadNote reads the note at the brainDir-relative relPath and returns
|
||||||
|
// its parsed frontmatter, body, and content hash. It is the read-after-
|
||||||
|
// write primitive behind brain_get: the hash it returns equals the hash
|
||||||
|
// brain_write / brain_update returned for the same bytes.
|
||||||
|
func ReadNote(brainDir, relPath string) (fm map[string]string, body, contentHash string, err error) {
|
||||||
|
_, abs, err := resolveWithin(brainDir, relPath)
|
||||||
|
if err != nil {
|
||||||
|
return nil, "", "", err
|
||||||
|
}
|
||||||
|
raw, err := os.ReadFile(abs)
|
||||||
|
if err != nil {
|
||||||
|
if os.IsNotExist(err) {
|
||||||
|
return nil, "", "", fmt.Errorf("note %q does not exist", relPath)
|
||||||
|
}
|
||||||
|
return nil, "", "", fmt.Errorf("read note: %w", err)
|
||||||
|
}
|
||||||
|
parsed, body := parseFrontmatter(string(raw))
|
||||||
|
fm = make(map[string]string, len(parsed.lines))
|
||||||
|
for _, l := range parsed.lines {
|
||||||
|
if l.key != "" {
|
||||||
|
fm[l.key] = l.value
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return fm, body, ContentHash(raw), nil
|
||||||
|
}
|
||||||
@@ -0,0 +1,130 @@
|
|||||||
|
package api
|
||||||
|
|
||||||
|
import (
|
||||||
|
"os"
|
||||||
|
"path/filepath"
|
||||||
|
"strings"
|
||||||
|
"testing"
|
||||||
|
|
||||||
|
"github.com/stretchr/testify/assert"
|
||||||
|
"github.com/stretchr/testify/require"
|
||||||
|
)
|
||||||
|
|
||||||
|
// seedNote writes a note directly to disk and returns its relPath.
|
||||||
|
func seedNote(t *testing.T, brainDir, rel, content string) string {
|
||||||
|
t.Helper()
|
||||||
|
full := filepath.Join(brainDir, filepath.FromSlash(rel))
|
||||||
|
require.NoError(t, os.MkdirAll(filepath.Dir(full), 0o755))
|
||||||
|
require.NoError(t, os.WriteFile(full, []byte(content), 0o644))
|
||||||
|
return rel
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestUpdateNoteSupersedesAndStamps(t *testing.T) {
|
||||||
|
brainDir := t.TempDir()
|
||||||
|
rel := seedNote(t, brainDir, "wiki/jepa-fx/facts/val-vol.md",
|
||||||
|
"---\nwing: jepa-fx\nhall: facts\ncreated_at: 2026-01-01T00:00:00Z\ncustom: keep-me\n---\n# Old\n\nold body\n")
|
||||||
|
|
||||||
|
relPath, hash, priorHash, err := UpdateNote(brainDir, UpdateNoteOptions{
|
||||||
|
Path: rel,
|
||||||
|
Content: "# New\n\nnew body\n",
|
||||||
|
Reason: "facts changed",
|
||||||
|
})
|
||||||
|
require.NoError(t, err)
|
||||||
|
assert.Equal(t, rel, relPath)
|
||||||
|
assert.NotEmpty(t, hash)
|
||||||
|
assert.NotEmpty(t, priorHash)
|
||||||
|
assert.NotEqual(t, hash, priorHash)
|
||||||
|
|
||||||
|
got, err := os.ReadFile(filepath.Join(brainDir, filepath.FromSlash(rel)))
|
||||||
|
require.NoError(t, err)
|
||||||
|
s := string(got)
|
||||||
|
// Body replaced.
|
||||||
|
assert.Contains(t, s, "# New")
|
||||||
|
assert.NotContains(t, s, "old body")
|
||||||
|
// Prior fields preserved.
|
||||||
|
assert.Contains(t, s, "wing: jepa-fx")
|
||||||
|
assert.Contains(t, s, "hall: facts")
|
||||||
|
assert.Contains(t, s, "created_at: 2026-01-01T00:00:00Z")
|
||||||
|
assert.Contains(t, s, "custom: keep-me")
|
||||||
|
// Supersession stamped.
|
||||||
|
assert.Contains(t, s, "updated_at:")
|
||||||
|
assert.Contains(t, s, "supersedes: "+priorHash)
|
||||||
|
assert.Contains(t, s, "supersede_reason: facts changed")
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestUpdateNoteResolvesByWingHallSlug(t *testing.T) {
|
||||||
|
brainDir := t.TempDir()
|
||||||
|
seedNote(t, brainDir, "wiki/jepa-fx/facts/val-vol.md",
|
||||||
|
"---\nwing: jepa-fx\nhall: facts\n---\nold\n")
|
||||||
|
|
||||||
|
relPath, _, _, err := UpdateNote(brainDir, UpdateNoteOptions{
|
||||||
|
Wing: "jepa-fx", Hall: "facts", Slug: "val-vol",
|
||||||
|
Content: "new\n",
|
||||||
|
})
|
||||||
|
require.NoError(t, err)
|
||||||
|
assert.Equal(t, "wiki/jepa-fx/facts/val-vol.md", relPath)
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestUpdateNoteErrorsOnMissingAndDoesNotCreate(t *testing.T) {
|
||||||
|
brainDir := t.TempDir()
|
||||||
|
|
||||||
|
_, _, _, err := UpdateNote(brainDir, UpdateNoteOptions{
|
||||||
|
Wing: "jepa-fx", Hall: "facts", Slug: "ghost",
|
||||||
|
Content: "x\n",
|
||||||
|
})
|
||||||
|
require.Error(t, err)
|
||||||
|
assert.Contains(t, err.Error(), "does not exist")
|
||||||
|
|
||||||
|
// No file created.
|
||||||
|
_, statErr := os.Stat(filepath.Join(brainDir, "wiki/jepa-fx/facts/ghost.md"))
|
||||||
|
assert.True(t, os.IsNotExist(statErr), "missing-target update must not create a note")
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestUpdateNoteRejectsTraversal(t *testing.T) {
|
||||||
|
brainDir := t.TempDir()
|
||||||
|
_, _, _, err := UpdateNote(brainDir, UpdateNoteOptions{
|
||||||
|
Path: "../escape.md",
|
||||||
|
Content: "x\n",
|
||||||
|
})
|
||||||
|
require.Error(t, err)
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestReadNoteReturnsFrontmatterBodyHash(t *testing.T) {
|
||||||
|
brainDir := t.TempDir()
|
||||||
|
rel := seedNote(t, brainDir, "wiki/jepa-fx/facts/n.md",
|
||||||
|
"---\nwing: jepa-fx\nhall: facts\n---\n# Body\n\ntext\n")
|
||||||
|
|
||||||
|
fm, body, hash, err := ReadNote(brainDir, rel)
|
||||||
|
require.NoError(t, err)
|
||||||
|
assert.Equal(t, "jepa-fx", fm["wing"])
|
||||||
|
assert.Equal(t, "facts", fm["hall"])
|
||||||
|
assert.Equal(t, "# Body\n\ntext\n", body)
|
||||||
|
|
||||||
|
// Hash matches ContentHash of the raw bytes on disk (round-trip).
|
||||||
|
raw, _ := os.ReadFile(filepath.Join(brainDir, filepath.FromSlash(rel)))
|
||||||
|
assert.Equal(t, ContentHash(raw), hash)
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestReadNoteRejectsTraversal(t *testing.T) {
|
||||||
|
brainDir := t.TempDir()
|
||||||
|
_, _, _, err := ReadNote(brainDir, "../../etc/passwd")
|
||||||
|
require.Error(t, err)
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestUpdateThenReadRoundTripsHash(t *testing.T) {
|
||||||
|
brainDir := t.TempDir()
|
||||||
|
rel := seedNote(t, brainDir, "wiki/a/facts/n.md", "---\nwing: a\nhall: facts\n---\nold\n")
|
||||||
|
|
||||||
|
_, hash, _, err := UpdateNote(brainDir, UpdateNoteOptions{Path: rel, Content: "new\n"})
|
||||||
|
require.NoError(t, err)
|
||||||
|
|
||||||
|
_, _, readHash, err := ReadNote(brainDir, rel)
|
||||||
|
require.NoError(t, err)
|
||||||
|
assert.Equal(t, hash, readHash, "update content_hash must round-trip through ReadNote")
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestContentHashStable(t *testing.T) {
|
||||||
|
assert.Equal(t, ContentHash([]byte("abc")), ContentHash([]byte("abc")))
|
||||||
|
assert.NotEqual(t, ContentHash([]byte("abc")), ContentHash([]byte("abd")))
|
||||||
|
assert.True(t, strings.HasPrefix(ContentHash([]byte("")), "")) // hex, non-panicking
|
||||||
|
}
|
||||||
@@ -1,6 +1,10 @@
|
|||||||
package claudewatcher
|
package claudewatcher
|
||||||
|
|
||||||
import "regexp"
|
import (
|
||||||
|
"fmt"
|
||||||
|
"regexp"
|
||||||
|
"sync"
|
||||||
|
)
|
||||||
|
|
||||||
// Scrubber drops any turn whose content matches a known-bad pattern.
|
// Scrubber drops any turn whose content matches a known-bad pattern.
|
||||||
// Fail-closed by design: we'd rather lose signal than ingest credentials
|
// Fail-closed by design: we'd rather lose signal than ingest credentials
|
||||||
@@ -34,21 +38,74 @@ var DefaultRules = []Rule{
|
|||||||
// specific match name in logs.
|
// specific match name in logs.
|
||||||
{Name: "authorization-header", RE: regexp.MustCompile(`(?i)Authorization\s*:\s*[A-Za-z]+\s+\S{8,}`)},
|
{Name: "authorization-header", RE: regexp.MustCompile(`(?i)Authorization\s*:\s*[A-Za-z]+\s+\S{8,}`)},
|
||||||
{Name: "bearer-token", RE: regexp.MustCompile(`(?i)Bearer\s+[A-Za-z0-9._\-]{16,}`)},
|
{Name: "bearer-token", RE: regexp.MustCompile(`(?i)Bearer\s+[A-Za-z0-9._\-]{16,}`)},
|
||||||
|
// JWT (header.payload.sig), e.g. a Dex/OAuth token dumped to stdout
|
||||||
|
// without a "Bearer " prefix. Both header and payload base64url-encode
|
||||||
|
// JSON, so both segments begin with "eyJ".
|
||||||
|
{Name: "jwt", RE: regexp.MustCompile(`eyJ[A-Za-z0-9_\-]{8,}\.eyJ[A-Za-z0-9_\-]{8,}\.[A-Za-z0-9_\-]{8,}`)},
|
||||||
{Name: "postgres-uri-with-password", RE: regexp.MustCompile(`postgres(?:ql)?://[^:\s/]+:[^@\s/]+@`)},
|
{Name: "postgres-uri-with-password", RE: regexp.MustCompile(`postgres(?:ql)?://[^:\s/]+:[^@\s/]+@`)},
|
||||||
{Name: "private-key", RE: regexp.MustCompile(`-----BEGIN[^-]*PRIVATE KEY-----`)},
|
{Name: "private-key", RE: regexp.MustCompile(`-----BEGIN[^-]*PRIVATE KEY-----`)},
|
||||||
{Name: "ssh-key", RE: regexp.MustCompile(`ssh-(?:rsa|ed25519|ecdsa)\s+[A-Za-z0-9+/=]{40,}`)},
|
{Name: "ssh-key", RE: regexp.MustCompile(`ssh-(?:rsa|ed25519|ecdsa)\s+[A-Za-z0-9+/=]{40,}`)},
|
||||||
{Name: "github-pat", RE: regexp.MustCompile(`\b(?:ghp|gho|ghu|ghr|gha)_[A-Za-z0-9]{30,}\b`)},
|
{Name: "github-pat", RE: regexp.MustCompile(`\b(?:ghp|gho|ghu|ghr|gha)_[A-Za-z0-9]{30,}\b`)},
|
||||||
{Name: "openai-sk", RE: regexp.MustCompile(`\bsk-(?:proj-)?[A-Za-z0-9]{32,}\b`)},
|
// 1Password service-account token (ops_<base64url>). Long, high-value root
|
||||||
|
// credential; guard the bare value (the _TOKEN= form also hits homelab-env-token).
|
||||||
|
{Name: "op-service-account", RE: regexp.MustCompile(`\bops_[A-Za-z0-9_\-]{40,}`)},
|
||||||
|
// No leading \b: a shell mangle can glue the key to a preceding word
|
||||||
|
// ("yes"+"sk-...") which has no word boundary, and that exact case
|
||||||
|
// leaked a LiteLLM master key past this rule (2026-06-11). Match the
|
||||||
|
// sk- shape wherever it appears; the {32,} length floor keeps short
|
||||||
|
// "task-"/"disk-" words from tripping it.
|
||||||
|
{Name: "openai-sk", RE: regexp.MustCompile(`sk-(?:proj-)?[A-Za-z0-9]{32,}`)},
|
||||||
{Name: "anthropic-sk", RE: regexp.MustCompile(`\bsk-ant-[A-Za-z0-9_\-]{32,}\b`)},
|
{Name: "anthropic-sk", RE: regexp.MustCompile(`\bsk-ant-[A-Za-z0-9_\-]{32,}\b`)},
|
||||||
{Name: "aws-access-key", RE: regexp.MustCompile(`\bAKIA[0-9A-Z]{16}\b`)},
|
{Name: "aws-access-key", RE: regexp.MustCompile(`\bAKIA[0-9A-Z]{16}\b`)},
|
||||||
{Name: "homelab-env-token", RE: regexp.MustCompile(`(?i)(?:_TOKEN|_PASSWORD|_API_KEY|_SECRET)\s*[:=]\s*['"]?[A-Za-z0-9._/+\-]{12,}`)},
|
{Name: "homelab-env-token", RE: regexp.MustCompile(`(?i)(?:_TOKEN|_PASSWORD|_API_KEY|_SECRET)\s*[:=]\s*['"]?[A-Za-z0-9._/+\-]{12,}`)},
|
||||||
{Name: "sops-encrypted-marker", RE: regexp.MustCompile(`ENC\[AES256_GCM,data:[A-Za-z0-9+/=]{8,}`)},
|
{Name: "sops-encrypted-marker", RE: regexp.MustCompile(`ENC\[AES256_GCM,data:[A-Za-z0-9+/=]{8,}`)},
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// extraRules is appended to DefaultRules at process startup via
|
||||||
|
// RegisterRule. The mutex guards concurrent RegisterRule calls (rare)
|
||||||
|
// against concurrent Scrub reads (hot path). Scrub takes a read lock
|
||||||
|
// only when extraRules is non-empty, so steady-state cost is zero
|
||||||
|
// when no client-name guard is configured.
|
||||||
|
var (
|
||||||
|
extraRulesMu sync.RWMutex
|
||||||
|
extraRules []Rule
|
||||||
|
)
|
||||||
|
|
||||||
|
// RegisterRule appends a runtime-configured regex to the scrubber's
|
||||||
|
// rule set. Used by main to inject client-name guards from
|
||||||
|
// CLAUDE_INGEST_CLIENT_BLOCK env var (or equivalent SOPS-encrypted
|
||||||
|
// secret) without baking client identities into source code.
|
||||||
|
//
|
||||||
|
// pattern is compiled as-is — callers wrap with `\b...\b` and case
|
||||||
|
// flags as needed. Duplicate names are accepted (rules are positional);
|
||||||
|
// the second registration just fires after the first.
|
||||||
|
func RegisterRule(name, pattern string) error {
|
||||||
|
re, err := regexp.Compile(pattern)
|
||||||
|
if err != nil {
|
||||||
|
return fmt.Errorf("compile rule %q: %w", name, err)
|
||||||
|
}
|
||||||
|
extraRulesMu.Lock()
|
||||||
|
extraRules = append(extraRules, Rule{Name: name, RE: re})
|
||||||
|
extraRulesMu.Unlock()
|
||||||
|
return nil
|
||||||
|
}
|
||||||
|
|
||||||
|
// ResetExtraRules clears every RegisterRule-added rule. Test-only.
|
||||||
|
func ResetExtraRules() {
|
||||||
|
extraRulesMu.Lock()
|
||||||
|
extraRules = nil
|
||||||
|
extraRulesMu.Unlock()
|
||||||
|
}
|
||||||
|
|
||||||
// Scrub reports the first matching rule, or empty when content is clean.
|
// Scrub reports the first matching rule, or empty when content is clean.
|
||||||
// Empty string is treated as clean. Caller decides what to do on a hit;
|
// Empty string is treated as clean. Caller decides what to do on a hit;
|
||||||
// the convention in claudewatcher is to drop the turn entirely and emit
|
// the convention in claudewatcher is to drop the turn entirely and emit
|
||||||
// a slog.Warn naming the rule.
|
// a slog.Warn naming the rule.
|
||||||
|
//
|
||||||
|
// Rule order: DefaultRules first (credential shapes), then runtime
|
||||||
|
// RegisterRule additions (client-name guards). Credential leaks
|
||||||
|
// outrank client-name hits in the log because they're strictly more
|
||||||
|
// dangerous.
|
||||||
func Scrub(content string) string {
|
func Scrub(content string) string {
|
||||||
if content == "" {
|
if content == "" {
|
||||||
return ""
|
return ""
|
||||||
@@ -58,5 +115,12 @@ func Scrub(content string) string {
|
|||||||
return r.Name
|
return r.Name
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
extraRulesMu.RLock()
|
||||||
|
defer extraRulesMu.RUnlock()
|
||||||
|
for _, r := range extraRules {
|
||||||
|
if r.RE.MatchString(content) {
|
||||||
|
return r.Name
|
||||||
|
}
|
||||||
|
}
|
||||||
return ""
|
return ""
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -25,6 +25,18 @@ func TestScrub_PoisonedFixtures(t *testing.T) {
|
|||||||
{"aws-access-key", "AWS_ACCESS_KEY_ID=AKIAIOSFODNN7EXAMPLE", "aws-access-key"},
|
{"aws-access-key", "AWS_ACCESS_KEY_ID=AKIAIOSFODNN7EXAMPLE", "aws-access-key"},
|
||||||
{"homelab-env", "POSTGRES_PASSWORD=hunter2supersecretvalue", "homelab-env-token"},
|
{"homelab-env", "POSTGRES_PASSWORD=hunter2supersecretvalue", "homelab-env-token"},
|
||||||
{"sops-marker", "value: ENC[AES256_GCM,data:abc123def456,iv:zzz]", "sops-encrypted-marker"},
|
{"sops-marker", "value: ENC[AES256_GCM,data:abc123def456,iv:zzz]", "sops-encrypted-marker"},
|
||||||
|
// Regression: a shell mangle glued the key to a preceding word
|
||||||
|
// ("yes"+"sk-..."), defeating the leading \b in the sk- rule and
|
||||||
|
// leaking a LiteLLM master key past the scrubber (2026-06-11).
|
||||||
|
{"sk-glued-to-word", "master key resolved: yessk-7181ca984603239d8c4819361bf33b94b9c3c07018791868", "openai-sk"},
|
||||||
|
{"sk-standalone-hex", "sk-7181ca984603239d8c4819361bf33b94b9c3c07018791868", "openai-sk"},
|
||||||
|
// Bare JWT not preceded by "Bearer" (e.g. a Dex token dumped to stdout).
|
||||||
|
{"jwt-bare", "token=eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9.eyJzdWIiOiIxMjM0NTY3ODkwIn0.dQw4w9WgXcQabcdef", "jwt"},
|
||||||
|
// 1Password service-account token (ops_<base64url>), env-assigned and bare.
|
||||||
|
// Both hit the dedicated op-service-account rule (ordered before the
|
||||||
|
// generic homelab-env-token). Guards ~/.zshrc reads etc. (2026-06-14).
|
||||||
|
{"op-sa-env", "export OP_SERVICE_ACCOUNT_TOKEN=ops_eyJzaWduSW5BZGRyZXNzIjoibXkuMXBhc3N3b3JkLmNvbSJ9", "op-service-account"},
|
||||||
|
{"op-sa-bare", "ops_eyJzaWduSW5BZGRyZXNzIjoibXkuMXBhc3N3b3JkLmNvbSIsInVzZXJBdXRoIjp7fX0aGVsbG8", "op-service-account"},
|
||||||
}
|
}
|
||||||
for _, tc := range cases {
|
for _, tc := range cases {
|
||||||
t.Run(tc.name, func(t *testing.T) {
|
t.Run(tc.name, func(t *testing.T) {
|
||||||
@@ -43,6 +55,9 @@ func TestScrub_CleanContentPassesThrough(t *testing.T) {
|
|||||||
"file at ~/.ssh/id_ed25519",
|
"file at ~/.ssh/id_ed25519",
|
||||||
"the function Authorization() takes no args",
|
"the function Authorization() takes no args",
|
||||||
"comment: see API key in 1Password",
|
"comment: see API key in 1Password",
|
||||||
|
// loosened sk- rule must not trip on short "task-"/"disk-" words
|
||||||
|
"run task-build then task-test in the pipeline",
|
||||||
|
"mounted /dev/disk-by-id/wwn-0x5000",
|
||||||
}
|
}
|
||||||
for _, c := range cases {
|
for _, c := range cases {
|
||||||
assert.Empty(t, Scrub(c), "expected clean for %q", c)
|
assert.Empty(t, Scrub(c), "expected clean for %q", c)
|
||||||
@@ -55,3 +70,63 @@ func TestScrub_FirstMatchWins(t *testing.T) {
|
|||||||
content := "Authorization: Bearer ghp_aBcD1234EfGh5678IjKl9012MnOp3456QrSt"
|
content := "Authorization: Bearer ghp_aBcD1234EfGh5678IjKl9012MnOp3456QrSt"
|
||||||
assert.Equal(t, "authorization-header", Scrub(content))
|
assert.Equal(t, "authorization-header", Scrub(content))
|
||||||
}
|
}
|
||||||
|
|
||||||
|
func TestRegisterRule_ClientNameGuard(t *testing.T) {
|
||||||
|
t.Cleanup(ResetExtraRules)
|
||||||
|
require := func(err error) {
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("unexpected err: %v", err)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
require(RegisterRule("client-name", `(?i)\b(SEB|Mastercard)\b`))
|
||||||
|
|
||||||
|
// Hits — case variations + word-boundary respect.
|
||||||
|
for _, hit := range []string{
|
||||||
|
"mentioned SEB in this commit",
|
||||||
|
"the Mastercard project deadline",
|
||||||
|
"working on mastercard scope",
|
||||||
|
"SEB internal review",
|
||||||
|
} {
|
||||||
|
assert.Equal(t, "client-name", Scrub(hit), "should match %q", hit)
|
||||||
|
}
|
||||||
|
|
||||||
|
// Misses — substring within a longer word should NOT match
|
||||||
|
// thanks to \b. "Sebastian" contains "seb" but \b prevents hit.
|
||||||
|
for _, miss := range []string{
|
||||||
|
"Sebastian wrote the docs",
|
||||||
|
"unrelated text",
|
||||||
|
"researcher",
|
||||||
|
"https://example.com/search?seb=1", // 'seb' bounded by ?=, still matches \b
|
||||||
|
} {
|
||||||
|
got := Scrub(miss)
|
||||||
|
if miss == "https://example.com/search?seb=1" {
|
||||||
|
// `seb=` has word-boundary at '='; this DOES match \bseb\b.
|
||||||
|
// Accept either outcome; document the tradeoff.
|
||||||
|
assert.Contains(t, []string{"", "client-name"}, got)
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
assert.Empty(t, got, "should NOT match %q", miss)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestRegisterRule_CredentialsTakePrecedence(t *testing.T) {
|
||||||
|
t.Cleanup(ResetExtraRules)
|
||||||
|
require := func(err error) {
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("unexpected err: %v", err)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
require(RegisterRule("client-name", `\b(SEB)\b`))
|
||||||
|
|
||||||
|
// Content matches both a credential rule AND a client rule —
|
||||||
|
// credential rule wins by ordering, so log triage points at the
|
||||||
|
// strictly more dangerous leak.
|
||||||
|
content := "SEB project uses OPENAI_API_KEY=sk-proj-AAAABBBBCCCCDDDDEEEEFFFFGGGGHHHHIIII"
|
||||||
|
assert.Equal(t, "openai-sk", Scrub(content))
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestRegisterRule_RejectsInvalidPattern(t *testing.T) {
|
||||||
|
t.Cleanup(ResetExtraRules)
|
||||||
|
err := RegisterRule("bad", "[unclosed")
|
||||||
|
assert.Error(t, err)
|
||||||
|
}
|
||||||
|
|||||||
@@ -0,0 +1,204 @@
|
|||||||
|
package mcp_test
|
||||||
|
|
||||||
|
import (
|
||||||
|
"context"
|
||||||
|
"encoding/json"
|
||||||
|
"os"
|
||||||
|
"path/filepath"
|
||||||
|
"strings"
|
||||||
|
"testing"
|
||||||
|
"time"
|
||||||
|
|
||||||
|
"github.com/mathiasbq/hyperguild/ingestion/internal/mcp"
|
||||||
|
"github.com/mathiasbq/hyperguild/ingestion/internal/vectorstore"
|
||||||
|
"github.com/stretchr/testify/assert"
|
||||||
|
"github.com/stretchr/testify/require"
|
||||||
|
)
|
||||||
|
|
||||||
|
// callResult parses the JSON text payload of a successful tool call.
|
||||||
|
func callResult(t *testing.T, resp map[string]any) map[string]any {
|
||||||
|
t.Helper()
|
||||||
|
require.Nil(t, resp["error"], "tool returned error: %v", resp["error"])
|
||||||
|
text := resp["result"].(map[string]any)["content"].([]any)[0].(map[string]any)["text"].(string)
|
||||||
|
var out map[string]any
|
||||||
|
require.NoError(t, json.Unmarshal([]byte(text), &out))
|
||||||
|
return out
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestBrainUpdateSupersedesExisting(t *testing.T) {
|
||||||
|
brainDir := t.TempDir()
|
||||||
|
srv := mcp.NewServer(brainDir, nil, nil, nil)
|
||||||
|
|
||||||
|
// Seed via brain_write so the note carries real frontmatter.
|
||||||
|
callResult(t, toolCall(t, srv, "brain_write", map[string]any{
|
||||||
|
"content": "# Old\n\nold body\n", "filename": "val-vol",
|
||||||
|
"wing": "jepa-fx", "hall": "facts",
|
||||||
|
}))
|
||||||
|
|
||||||
|
out := callResult(t, toolCall(t, srv, "brain_update", map[string]any{
|
||||||
|
"wing": "jepa-fx", "hall": "facts", "slug": "val-vol",
|
||||||
|
"content": "# New\n\nnew body\n", "reason": "facts changed",
|
||||||
|
}))
|
||||||
|
assert.Equal(t, "wiki/jepa-fx/facts/val-vol.md", out["path"])
|
||||||
|
assert.Equal(t, out["path"], out["id"])
|
||||||
|
assert.NotEmpty(t, out["content_hash"])
|
||||||
|
assert.Equal(t, true, out["superseded"])
|
||||||
|
|
||||||
|
got, err := os.ReadFile(filepath.Join(brainDir, "wiki/jepa-fx/facts/val-vol.md"))
|
||||||
|
require.NoError(t, err)
|
||||||
|
s := string(got)
|
||||||
|
assert.Contains(t, s, "# New")
|
||||||
|
assert.NotContains(t, s, "old body")
|
||||||
|
assert.Contains(t, s, "wing: jepa-fx")
|
||||||
|
assert.Contains(t, s, "supersede_reason: facts changed")
|
||||||
|
assert.Contains(t, s, "supersedes:")
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestBrainUpdateMissingTargetErrorsNoCreate(t *testing.T) {
|
||||||
|
brainDir := t.TempDir()
|
||||||
|
srv := mcp.NewServer(brainDir, nil, nil, nil)
|
||||||
|
|
||||||
|
resp := toolCall(t, srv, "brain_update", map[string]any{
|
||||||
|
"wing": "jepa-fx", "hall": "facts", "slug": "ghost",
|
||||||
|
"content": "x\n",
|
||||||
|
})
|
||||||
|
require.NotNil(t, resp["error"])
|
||||||
|
assert.Contains(t, resp["error"].(map[string]any)["message"].(string), "does not exist")
|
||||||
|
_, statErr := os.Stat(filepath.Join(brainDir, "wiki/jepa-fx/facts/ghost.md"))
|
||||||
|
assert.True(t, os.IsNotExist(statErr))
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestBrainUpdateByFullPath(t *testing.T) {
|
||||||
|
brainDir := t.TempDir()
|
||||||
|
srv := mcp.NewServer(brainDir, nil, nil, nil)
|
||||||
|
callResult(t, toolCall(t, srv, "brain_write", map[string]any{
|
||||||
|
"content": "old\n", "filename": "n", "wing": "a", "hall": "facts",
|
||||||
|
}))
|
||||||
|
|
||||||
|
out := callResult(t, toolCall(t, srv, "brain_update", map[string]any{
|
||||||
|
"slug": "wiki/a/facts/n.md", "content": "fresh\n",
|
||||||
|
}))
|
||||||
|
assert.Equal(t, "wiki/a/facts/n.md", out["path"])
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestBrainGetByIDAndPath(t *testing.T) {
|
||||||
|
brainDir := t.TempDir()
|
||||||
|
srv := mcp.NewServer(brainDir, nil, nil, nil)
|
||||||
|
w := callResult(t, toolCall(t, srv, "brain_write", map[string]any{
|
||||||
|
"content": "# Body\n\ntext\n", "filename": "n", "wing": "a", "hall": "facts",
|
||||||
|
}))
|
||||||
|
id := w["id"].(string)
|
||||||
|
hash := w["content_hash"].(string)
|
||||||
|
require.NotEmpty(t, id)
|
||||||
|
require.NotEmpty(t, hash)
|
||||||
|
|
||||||
|
// by id
|
||||||
|
g1 := callResult(t, toolCall(t, srv, "brain_get", map[string]any{"id": id}))
|
||||||
|
assert.Equal(t, id, g1["path"])
|
||||||
|
assert.Equal(t, hash, g1["content_hash"], "content_hash must round-trip write→get")
|
||||||
|
assert.Contains(t, g1["body"].(string), "# Body")
|
||||||
|
fm := g1["frontmatter"].(map[string]any)
|
||||||
|
assert.Equal(t, "a", fm["wing"])
|
||||||
|
|
||||||
|
// by path
|
||||||
|
g2 := callResult(t, toolCall(t, srv, "brain_get", map[string]any{"path": id}))
|
||||||
|
assert.Equal(t, hash, g2["content_hash"])
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestBrainGetMissingArgsErrors(t *testing.T) {
|
||||||
|
srv := mcp.NewServer(t.TempDir(), nil, nil, nil)
|
||||||
|
resp := toolCall(t, srv, "brain_get", map[string]any{})
|
||||||
|
require.NotNil(t, resp["error"])
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestBrainWriteReturnsHandle(t *testing.T) {
|
||||||
|
brainDir := t.TempDir()
|
||||||
|
srv := mcp.NewServer(brainDir, nil, nil, nil)
|
||||||
|
out := callResult(t, toolCall(t, srv, "brain_write", map[string]any{
|
||||||
|
"content": "# X\n\nbody\n", "filename": "x", "wing": "a", "hall": "facts",
|
||||||
|
}))
|
||||||
|
assert.Equal(t, "wiki/a/facts/x.md", out["path"])
|
||||||
|
assert.Equal(t, out["path"], out["id"])
|
||||||
|
assert.NotEmpty(t, out["content_hash"])
|
||||||
|
}
|
||||||
|
|
||||||
|
// --- retrieval-reflects-new-content: exercises the real mtime-driven Sync ---
|
||||||
|
|
||||||
|
type fakeVecStore struct {
|
||||||
|
chunks map[string][]float32
|
||||||
|
deleted []string
|
||||||
|
}
|
||||||
|
|
||||||
|
func (f *fakeVecStore) KnownPathsWithTime(_ context.Context) (map[string]time.Time, error) {
|
||||||
|
m := make(map[string]time.Time, len(f.chunks))
|
||||||
|
for p := range f.chunks {
|
||||||
|
m[p] = time.Unix(0, 0) // always stale → mtime(now) is always newer
|
||||||
|
}
|
||||||
|
return m, nil
|
||||||
|
}
|
||||||
|
|
||||||
|
func (f *fakeVecStore) Upsert(_ context.Context, path string, vec []float32) error {
|
||||||
|
f.chunks[path] = vec
|
||||||
|
return nil
|
||||||
|
}
|
||||||
|
|
||||||
|
func (f *fakeVecStore) Delete(_ context.Context, path string) error {
|
||||||
|
delete(f.chunks, path)
|
||||||
|
f.deleted = append(f.deleted, path)
|
||||||
|
return nil
|
||||||
|
}
|
||||||
|
|
||||||
|
type fakeEmbedder struct{ seen []string }
|
||||||
|
|
||||||
|
func (e *fakeEmbedder) Embed(_ context.Context, text string) ([]float32, error) {
|
||||||
|
e.seen = append(e.seen, text)
|
||||||
|
return []float32{1, 0, 0}, nil
|
||||||
|
}
|
||||||
|
|
||||||
|
// TestBrainUpdateReembedsNewContent proves the supersede contract end to
|
||||||
|
// end against the actual embedding mechanism: brain_update rewrites the
|
||||||
|
// file, advancing its mtime, and the next vectorstore.Sync pass re-embeds
|
||||||
|
// the NEW body and drops the stale chunk. No stub of the re-index path.
|
||||||
|
func TestBrainUpdateReembedsNewContent(t *testing.T) {
|
||||||
|
brainDir := t.TempDir()
|
||||||
|
srv := mcp.NewServer(brainDir, nil, nil, nil)
|
||||||
|
ctx := context.Background()
|
||||||
|
|
||||||
|
callResult(t, toolCall(t, srv, "brain_write", map[string]any{
|
||||||
|
"content": "# Note\n\nthe OLD distinctive payload\n",
|
||||||
|
"filename": "n", "wing": "a", "hall": "facts",
|
||||||
|
}))
|
||||||
|
|
||||||
|
store := &fakeVecStore{chunks: map[string][]float32{}}
|
||||||
|
emb := &fakeEmbedder{}
|
||||||
|
|
||||||
|
// First sync embeds the original content.
|
||||||
|
_, err := vectorstore.Sync(ctx, brainDir, store, emb)
|
||||||
|
require.NoError(t, err)
|
||||||
|
require.NotEmpty(t, store.chunks)
|
||||||
|
require.True(t, anyContains(emb.seen, "OLD distinctive payload"))
|
||||||
|
|
||||||
|
callResult(t, toolCall(t, srv, "brain_update", map[string]any{
|
||||||
|
"wing": "a", "hall": "facts", "slug": "n",
|
||||||
|
"content": "# Note\n\nthe NEW distinctive payload\n",
|
||||||
|
}))
|
||||||
|
|
||||||
|
emb.seen = nil // only watch what the second pass embeds
|
||||||
|
_, err = vectorstore.Sync(ctx, brainDir, store, emb)
|
||||||
|
require.NoError(t, err)
|
||||||
|
|
||||||
|
assert.True(t, anyContains(emb.seen, "NEW distinctive payload"),
|
||||||
|
"Sync must re-embed the superseded body; saw %v", emb.seen)
|
||||||
|
assert.False(t, anyContains(emb.seen, "OLD distinctive payload"),
|
||||||
|
"the old body must not be re-embedded")
|
||||||
|
assert.NotEmpty(t, store.deleted, "stale chunks must be deleted before re-embed")
|
||||||
|
}
|
||||||
|
|
||||||
|
func anyContains(ss []string, sub string) bool {
|
||||||
|
for _, s := range ss {
|
||||||
|
if strings.Contains(s, sub) {
|
||||||
|
return true
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return false
|
||||||
|
}
|
||||||
@@ -61,6 +61,26 @@ func (s *Server) tools() []map[string]any {
|
|||||||
"hall": enum("optional memory type (requires wing)", halls...),
|
"hall": enum("optional memory type (requires wing)", halls...),
|
||||||
}),
|
}),
|
||||||
},
|
},
|
||||||
|
{
|
||||||
|
"name": "brain_update",
|
||||||
|
"description": "Supersede an existing brain note in place: whole-note body replace + frontmatter re-stamp (updated_at, supersedes=prior content hash, supersede_reason). Errors if the target does not exist — use brain_write to create. Returns {id, path, content_hash, superseded}. Prior version recoverable from git.",
|
||||||
|
"inputSchema": schema([]string{"content"}, map[string]any{
|
||||||
|
"content": str("new full body (whole-note replace)"),
|
||||||
|
"slug": str("target note slug within wing/hall, OR a full brain-relative path (e.g. wiki/jepa-fx/facts/x.md)"),
|
||||||
|
"wing": str("wing of the target (required unless slug/path is a full path)"),
|
||||||
|
"hall": enum("hall of the target (required unless slug/path is a full path)", halls...),
|
||||||
|
"path": str("full brain-relative path to the target; takes precedence over slug/wing/hall"),
|
||||||
|
"reason": str("optional short note on why superseded — stamped into frontmatter"),
|
||||||
|
}),
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "brain_get",
|
||||||
|
"description": "Fetch a single brain note by id or path (both are the brain-relative path — the note handle). Returns {id, path, content_hash, frontmatter, body}. Read-after-write confirmation without a lexical re-query.",
|
||||||
|
"inputSchema": schema([]string{}, map[string]any{
|
||||||
|
"id": str("note id (brain-relative path) as returned by brain_write/brain_update"),
|
||||||
|
"path": str("brain-relative path to the note; equivalent to id"),
|
||||||
|
}),
|
||||||
|
},
|
||||||
{
|
{
|
||||||
"name": "brain_tunnel",
|
"name": "brain_tunnel",
|
||||||
"description": "Create an explicit bidirectional [[wikilink]] between two notes in different wings. Idempotent.",
|
"description": "Create an explicit bidirectional [[wikilink]] between two notes in different wings. Idempotent.",
|
||||||
@@ -222,7 +242,111 @@ func (s *Server) brainWrite(ctx context.Context, args json.RawMessage) (json.Raw
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
s.indexInGraph(ctx, "brain_write", relPath)
|
s.indexInGraph(ctx, "brain_write", relPath)
|
||||||
return json.Marshal(map[string]string{"path": relPath})
|
// Read-after-write handle: id == relPath, content_hash == sha256 of
|
||||||
|
// the bytes just written. path is kept for backward compatibility.
|
||||||
|
_, _, hash, _ := api.ReadNote(s.brainDir, relPath)
|
||||||
|
return json.Marshal(map[string]string{"id": relPath, "path": relPath, "content_hash": hash})
|
||||||
|
}
|
||||||
|
|
||||||
|
type brainUpdateArgs struct {
|
||||||
|
Slug string `json:"slug,omitempty"`
|
||||||
|
Wing string `json:"wing,omitempty"`
|
||||||
|
Hall string `json:"hall,omitempty"`
|
||||||
|
Path string `json:"path,omitempty"`
|
||||||
|
Content string `json:"content"`
|
||||||
|
Reason string `json:"reason,omitempty"`
|
||||||
|
}
|
||||||
|
|
||||||
|
// brainUpdate supersedes an existing note in place: whole-note body
|
||||||
|
// replace, frontmatter re-stamp (updated_at/supersedes/supersede_reason),
|
||||||
|
// graph re-index, and wing _index rebuild. It never creates — a missing
|
||||||
|
// target is an error so the caller can fall back to brain_write.
|
||||||
|
//
|
||||||
|
// Embedding re-sync is delegated to the out-of-band vectorstore.Sync
|
||||||
|
// ticker: the rewritten file's mtime advances, so the next pass re-embeds
|
||||||
|
// it. This mirrors brain_write, which likewise does not embed in-handler.
|
||||||
|
func (s *Server) brainUpdate(ctx context.Context, args json.RawMessage) (json.RawMessage, error) {
|
||||||
|
var a brainUpdateArgs
|
||||||
|
if err := json.Unmarshal(args, &a); err != nil {
|
||||||
|
return nil, fmt.Errorf("parse args: %w", err)
|
||||||
|
}
|
||||||
|
if a.Content == "" {
|
||||||
|
return nil, fmt.Errorf("content is required")
|
||||||
|
}
|
||||||
|
|
||||||
|
opts := api.UpdateNoteOptions{Content: a.Content, Reason: a.Reason}
|
||||||
|
switch {
|
||||||
|
case a.Path != "":
|
||||||
|
opts.Path = a.Path
|
||||||
|
case strings.Contains(a.Slug, "/"):
|
||||||
|
// slug carries a full path (issue #45: "slug ... OR full path").
|
||||||
|
opts.Path = a.Slug
|
||||||
|
default:
|
||||||
|
opts.Wing, opts.Hall, opts.Slug = a.Wing, a.Hall, a.Slug
|
||||||
|
}
|
||||||
|
|
||||||
|
relPath, hash, _, err := api.UpdateNote(s.brainDir, opts)
|
||||||
|
if err != nil {
|
||||||
|
return nil, err
|
||||||
|
}
|
||||||
|
|
||||||
|
// Best-effort wiki upkeep, mirroring brain_write: rebuild the wing
|
||||||
|
// _index and re-tunnel cross-wing matches against the new body. Both
|
||||||
|
// are idempotent and never block — the note is already superseded.
|
||||||
|
if wing := wingFromRelPath(relPath); wing != "" {
|
||||||
|
if err := brain.BuildWingIndex(s.brainDir, wing); err != nil {
|
||||||
|
slog.Warn("brain_update: auto-index failed", "wing", wing, "err", err)
|
||||||
|
}
|
||||||
|
if err := brain.AutoTunnel(s.brainDir, relPath, a.Content); err != nil {
|
||||||
|
slog.Warn("brain_update: auto-tunnel failed", "src", relPath, "err", err)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
s.indexInGraph(ctx, "brain_update", relPath)
|
||||||
|
|
||||||
|
return json.Marshal(map[string]any{
|
||||||
|
"id": relPath, "path": relPath, "content_hash": hash, "superseded": true,
|
||||||
|
})
|
||||||
|
}
|
||||||
|
|
||||||
|
// wingFromRelPath extracts the wing segment from a structured wiki path
|
||||||
|
// (wiki/<wing>/<hall>/<slug>.md). Returns "" for legacy/non-wiki paths.
|
||||||
|
func wingFromRelPath(relPath string) string {
|
||||||
|
parts := strings.Split(relPath, "/")
|
||||||
|
if len(parts) >= 4 && parts[0] == "wiki" {
|
||||||
|
return parts[1]
|
||||||
|
}
|
||||||
|
return ""
|
||||||
|
}
|
||||||
|
|
||||||
|
type brainGetArgs struct {
|
||||||
|
ID string `json:"id,omitempty"`
|
||||||
|
Path string `json:"path,omitempty"`
|
||||||
|
}
|
||||||
|
|
||||||
|
// brainGet fetches a note by id or path (both are the brainDir-relative
|
||||||
|
// path — the de-facto handle). Read-only; the create-path read-after-
|
||||||
|
// write primitive that lets callers confirm a write landed without a
|
||||||
|
// lexical re-query.
|
||||||
|
func (s *Server) brainGet(_ context.Context, args json.RawMessage) (json.RawMessage, error) {
|
||||||
|
var a brainGetArgs
|
||||||
|
if err := json.Unmarshal(args, &a); err != nil {
|
||||||
|
return nil, fmt.Errorf("parse args: %w", err)
|
||||||
|
}
|
||||||
|
target := a.Path
|
||||||
|
if target == "" {
|
||||||
|
target = a.ID
|
||||||
|
}
|
||||||
|
if target == "" {
|
||||||
|
return nil, fmt.Errorf("id or path is required")
|
||||||
|
}
|
||||||
|
fm, body, hash, err := api.ReadNote(s.brainDir, target)
|
||||||
|
if err != nil {
|
||||||
|
return nil, err
|
||||||
|
}
|
||||||
|
return json.Marshal(map[string]any{
|
||||||
|
"id": target, "path": target, "content_hash": hash,
|
||||||
|
"frontmatter": fm, "body": body,
|
||||||
|
})
|
||||||
}
|
}
|
||||||
|
|
||||||
// indexInGraph is a best-effort wrapper around graphsync.IndexDoc that
|
// indexInGraph is a best-effort wrapper around graphsync.IndexDoc that
|
||||||
|
|||||||
@@ -1,7 +1,7 @@
|
|||||||
// Package mcp implements an MCP HTTP handler for the ingestion service.
|
// Package mcp implements an MCP HTTP handler for the ingestion service.
|
||||||
// Exposed tools: brain_query, brain_write, brain_index, brain_tunnel,
|
// Exposed tools: brain_query, brain_write, brain_update, brain_get,
|
||||||
// brain_ingest, brain_ingest_raw, brain_answer, brain_classify,
|
// brain_index, brain_tunnel, brain_ingest, brain_ingest_raw,
|
||||||
// brain_graph, brain_context, session_log.
|
// brain_answer, brain_classify, brain_graph, brain_context, session_log.
|
||||||
package mcp
|
package mcp
|
||||||
|
|
||||||
import (
|
import (
|
||||||
@@ -177,6 +177,10 @@ func (s *Server) handleCall(ctx context.Context, name string, args json.RawMessa
|
|||||||
return s.brainQuery(ctx, args)
|
return s.brainQuery(ctx, args)
|
||||||
case "brain_write":
|
case "brain_write":
|
||||||
return s.brainWrite(ctx, args)
|
return s.brainWrite(ctx, args)
|
||||||
|
case "brain_update":
|
||||||
|
return s.brainUpdate(ctx, args)
|
||||||
|
case "brain_get":
|
||||||
|
return s.brainGet(ctx, args)
|
||||||
case "brain_index":
|
case "brain_index":
|
||||||
return s.brainIndex(ctx, args)
|
return s.brainIndex(ctx, args)
|
||||||
case "brain_tunnel":
|
case "brain_tunnel":
|
||||||
|
|||||||
@@ -55,7 +55,8 @@ func TestServerToolsList(t *testing.T) {
|
|||||||
names = append(names, t.(map[string]any)["name"].(string))
|
names = append(names, t.(map[string]any)["name"].(string))
|
||||||
}
|
}
|
||||||
assert.ElementsMatch(t, []string{
|
assert.ElementsMatch(t, []string{
|
||||||
"brain_query", "brain_write", "brain_index", "brain_tunnel",
|
"brain_query", "brain_write", "brain_update", "brain_get",
|
||||||
|
"brain_index", "brain_tunnel",
|
||||||
"brain_ingest_raw", "brain_ingest",
|
"brain_ingest_raw", "brain_ingest",
|
||||||
"brain_answer", "brain_classify", "brain_graph", "brain_context",
|
"brain_answer", "brain_classify", "brain_graph", "brain_context",
|
||||||
"session_log",
|
"session_log",
|
||||||
|
|||||||
@@ -77,9 +77,22 @@ func (s *Server) brainAnswer(ctx context.Context, args json.RawMessage) (json.Ra
|
|||||||
return nil, fmt.Errorf("search: %w", err)
|
return nil, fmt.Errorf("search: %w", err)
|
||||||
}
|
}
|
||||||
if s.reranker != nil && len(results) > 0 {
|
if s.reranker != nil && len(results) > 0 {
|
||||||
results, err = rerankResults(ctx, s.reranker, a.Query, results, 5)
|
reranked, rerr := rerankResults(ctx, s.reranker, a.Query, results, 5)
|
||||||
if err != nil {
|
if rerr != nil {
|
||||||
return nil, fmt.Errorf("rerank: %w", err)
|
return nil, fmt.Errorf("rerank: %w", rerr)
|
||||||
|
}
|
||||||
|
// The reranker is a filter, not a gate. The Qwen3-Reranker is a
|
||||||
|
// web-search cross-encoder: against a conversational / personal-
|
||||||
|
// intent query ("what am I optimizing toward?") it scores even
|
||||||
|
// on-topic notes as "no", which would collapse the whole answer to
|
||||||
|
// "no relevant content" despite BM25 having retrieved relevant
|
||||||
|
// content. When the reranker keeps nothing, fall back to the
|
||||||
|
// BM25/vector ordering (capped to the no-reranker depth) rather
|
||||||
|
// than returning an empty answer.
|
||||||
|
if len(reranked) > 0 {
|
||||||
|
results = reranked
|
||||||
|
} else if len(results) > 10 {
|
||||||
|
results = results[:10]
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
if len(results) == 0 {
|
if len(results) == 0 {
|
||||||
|
|||||||
@@ -98,6 +98,42 @@ func TestBrainAnswer_RerankerFiltersBeforeLLM(t *testing.T) {
|
|||||||
assert.NotContains(t, sawSources, "noise.md")
|
assert.NotContains(t, sawSources, "noise.md")
|
||||||
}
|
}
|
||||||
|
|
||||||
|
func TestBrainAnswer_RerankerKeepsNone_FallsBackToBM25(t *testing.T) {
|
||||||
|
brainDir := brainDirWithContent(t) // test.md BM25-matches "pass-rate logging"
|
||||||
|
|
||||||
|
// Reranker rejects every candidate ("no" to all) — models a
|
||||||
|
// web-search cross-encoder facing a conversational / personal-intent
|
||||||
|
// query, which is exactly when it wrongly scores on-topic notes as
|
||||||
|
// irrelevant. The answer must still synthesize from the BM25 hits, not
|
||||||
|
// collapse to "no relevant content".
|
||||||
|
rrSrv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, _ *http.Request) {
|
||||||
|
_ = json.NewEncoder(w).Encode(map[string]any{"response": "no", "done": true})
|
||||||
|
}))
|
||||||
|
defer rrSrv.Close()
|
||||||
|
|
||||||
|
var sawSources string
|
||||||
|
llm := func(_ context.Context, _, user string) (string, error) {
|
||||||
|
sawSources = user
|
||||||
|
return "fallback answer", nil
|
||||||
|
}
|
||||||
|
|
||||||
|
srv := mcp.NewServer(brainDir, nil, nil, llm).
|
||||||
|
WithReranker(reranker.New(rrSrv.URL, "qwen3"))
|
||||||
|
ts := httptest.NewServer(srv)
|
||||||
|
defer ts.Close()
|
||||||
|
|
||||||
|
rpc := callTool(t, ts, "brain_answer", map[string]any{"query": "pass-rate logging"})
|
||||||
|
require.Nil(t, rpc["error"])
|
||||||
|
|
||||||
|
content := rpc["result"].(map[string]any)["content"].([]any)[0].(map[string]any)["text"].(string)
|
||||||
|
var result map[string]any
|
||||||
|
require.NoError(t, json.Unmarshal([]byte(content), &result))
|
||||||
|
|
||||||
|
assert.Equal(t, "fallback answer", result["answer"])
|
||||||
|
assert.NotEmpty(t, result["sources"], "reranker keeping nothing must fall back to BM25, not empty")
|
||||||
|
assert.Contains(t, sawSources, "test.md")
|
||||||
|
}
|
||||||
|
|
||||||
func TestBrainAnswer_NoLLM(t *testing.T) {
|
func TestBrainAnswer_NoLLM(t *testing.T) {
|
||||||
srv := mcp.NewServer(t.TempDir(), nil, nil, nil)
|
srv := mcp.NewServer(t.TempDir(), nil, nil, nil)
|
||||||
ts := httptest.NewServer(srv)
|
ts := httptest.NewServer(srv)
|
||||||
|
|||||||
@@ -3,6 +3,7 @@ package vectorstore
|
|||||||
import (
|
import (
|
||||||
"fmt"
|
"fmt"
|
||||||
"strings"
|
"strings"
|
||||||
|
"unicode/utf8"
|
||||||
)
|
)
|
||||||
|
|
||||||
// NumberedChunk pairs a chunk's body with the storage path it will use
|
// NumberedChunk pairs a chunk's body with the storage path it will use
|
||||||
@@ -66,6 +67,70 @@ func ChunkMarkdown(content string, maxBytes int) []string {
|
|||||||
}
|
}
|
||||||
out = append(out, splitAtParagraphs(s, maxBytes)...)
|
out = append(out, splitAtParagraphs(s, maxBytes)...)
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// Final guarantee: no chunk exceeds maxBytes. A single heading-less,
|
||||||
|
// paragraph-less block (JSON-lines, minified content) survives the two
|
||||||
|
// passes above whole — splitAtParagraphs emits an over-budget paragraph
|
||||||
|
// rather than truncating prose. Hard-split any such chunk at line/rune
|
||||||
|
// boundaries so the embedder never rejects an over-context chunk.
|
||||||
|
final := make([]string, 0, len(out))
|
||||||
|
for _, c := range out {
|
||||||
|
if len(c) <= maxBytes {
|
||||||
|
final = append(final, c)
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
final = append(final, hardSplit(c, maxBytes)...)
|
||||||
|
}
|
||||||
|
return final
|
||||||
|
}
|
||||||
|
|
||||||
|
// hardSplit slices s into pieces no larger than maxBytes, breaking at line
|
||||||
|
// boundaries where possible and otherwise mid-line at a UTF-8 rune boundary.
|
||||||
|
// Last resort for content that has neither headings nor blank-line paragraphs.
|
||||||
|
func hardSplit(s string, maxBytes int) []string {
|
||||||
|
var out []string
|
||||||
|
var cur strings.Builder
|
||||||
|
flush := func() {
|
||||||
|
if cur.Len() > 0 {
|
||||||
|
out = append(out, cur.String())
|
||||||
|
cur.Reset()
|
||||||
|
}
|
||||||
|
}
|
||||||
|
for _, line := range strings.SplitAfter(s, "\n") {
|
||||||
|
if line == "" {
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
if len(line) > maxBytes {
|
||||||
|
flush()
|
||||||
|
out = append(out, runeSplit(line, maxBytes)...)
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
if cur.Len() > 0 && cur.Len()+len(line) > maxBytes {
|
||||||
|
flush()
|
||||||
|
}
|
||||||
|
cur.WriteString(line)
|
||||||
|
}
|
||||||
|
flush()
|
||||||
|
return out
|
||||||
|
}
|
||||||
|
|
||||||
|
// runeSplit slices s into <=maxBytes pieces without splitting a UTF-8 rune.
|
||||||
|
func runeSplit(s string, maxBytes int) []string {
|
||||||
|
var out []string
|
||||||
|
for len(s) > maxBytes {
|
||||||
|
cut := maxBytes
|
||||||
|
for cut > 0 && !utf8.RuneStart(s[cut]) {
|
||||||
|
cut--
|
||||||
|
}
|
||||||
|
if cut == 0 { // single rune wider than the budget; emit it whole
|
||||||
|
cut = maxBytes
|
||||||
|
}
|
||||||
|
out = append(out, s[:cut])
|
||||||
|
s = s[cut:]
|
||||||
|
}
|
||||||
|
if len(s) > 0 {
|
||||||
|
out = append(out, s)
|
||||||
|
}
|
||||||
return out
|
return out
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|||||||
@@ -30,6 +30,23 @@ func TestChunkMarkdown_SplitsAtHeadings(t *testing.T) {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
func TestChunkMarkdown_HardSplitsHeadinglessOversizedBlock(t *testing.T) {
|
||||||
|
// A document with no headings and no blank-line paragraph breaks (e.g.
|
||||||
|
// JSON-lines like wiki/telos/decisions/human-intent-column.md). The old
|
||||||
|
// chunker emitted it as one over-budget chunk → nomic-embed returned
|
||||||
|
// "input length exceeds the context length" (400). Every chunk must now
|
||||||
|
// fit the budget, with no content lost.
|
||||||
|
maxBytes := 200
|
||||||
|
src := strings.Repeat("x", 1000) // one 1000-byte blob, no headings, no \n\n
|
||||||
|
out := vectorstore.ChunkMarkdown(src, maxBytes)
|
||||||
|
|
||||||
|
require.Greater(t, len(out), 1, "oversized blob must be split")
|
||||||
|
for i, c := range out {
|
||||||
|
assert.LessOrEqual(t, len(c), maxBytes, "chunk %d over budget: %d bytes", i, len(c))
|
||||||
|
}
|
||||||
|
assert.Equal(t, 1000, strings.Count(strings.Join(out, ""), "x"), "no content lost")
|
||||||
|
}
|
||||||
|
|
||||||
func TestChunkMarkdown_FurtherSplitsOversizedSection(t *testing.T) {
|
func TestChunkMarkdown_FurtherSplitsOversizedSection(t *testing.T) {
|
||||||
// One H2 section with 4 paragraphs of ~80 chars each, limit 100.
|
// One H2 section with 4 paragraphs of ~80 chars each, limit 100.
|
||||||
src := "## big\n\n" +
|
src := "## big\n\n" +
|
||||||
|
|||||||
+3
-32
@@ -40,10 +40,9 @@ if [ -n "$ROOT_CONTEXT" ] && [ -f "$ROOT_CONTEXT" ]; then
|
|||||||
echo " Root context: $ROOT_CONTEXT"
|
echo " Root context: $ROOT_CONTEXT"
|
||||||
else
|
else
|
||||||
# No reachable root AGENT.md — common in CI's clean checkout. The root+project
|
# No reachable root AGENT.md — common in CI's clean checkout. The root+project
|
||||||
# adapters (AGENTS.md, .cursorrules, .aider.conventions.md, system-prompt.txt)
|
# adapters (AGENTS.md, system-prompt.txt) require the root context to
|
||||||
# require the root context to regenerate correctly, so we skip them entirely
|
# regenerate correctly, so we skip them entirely and only regenerate CLAUDE.md
|
||||||
# and only regenerate CLAUDE.md (which is project-only and inherits root via
|
# (which is project-only and inherits root via tree walk in Claude Code itself).
|
||||||
# tree walk in Claude Code itself).
|
|
||||||
echo " No root AGENT.md found — regenerating CLAUDE.md only"
|
echo " No root AGENT.md found — regenerating CLAUDE.md only"
|
||||||
echo "Syncing project context from $PROJECT_FILE..."
|
echo "Syncing project context from $PROJECT_FILE..."
|
||||||
cat "$PROJECT_FILE" > CLAUDE.md
|
cat "$PROJECT_FILE" > CLAUDE.md
|
||||||
@@ -78,30 +77,6 @@ generate_agents() {
|
|||||||
echo " → AGENTS.md (root + project; Crush, Pi, Antigravity)"
|
echo " → AGENTS.md (root + project; Crush, Pi, Antigravity)"
|
||||||
}
|
}
|
||||||
|
|
||||||
# ── Cursor ───────────────────────────────────────────────────
|
|
||||||
generate_cursor() {
|
|
||||||
{
|
|
||||||
echo "# Cursor rules — auto-generated"
|
|
||||||
echo "# Do not edit. Run: task context:sync"
|
|
||||||
echo ""
|
|
||||||
root_block
|
|
||||||
cat "$PROJECT_FILE"
|
|
||||||
} > .cursorrules
|
|
||||||
echo " → .cursorrules (root + project)"
|
|
||||||
}
|
|
||||||
|
|
||||||
# ── Aider ────────────────────────────────────────────────────
|
|
||||||
generate_aider() {
|
|
||||||
{ root_block; cat "$PROJECT_FILE"; } > .aider.conventions.md
|
|
||||||
if [ ! -f .aider.conf.yml ]; then
|
|
||||||
cat > .aider.conf.yml << 'YAML'
|
|
||||||
read: .aider.conventions.md
|
|
||||||
auto-commits: false
|
|
||||||
YAML
|
|
||||||
fi
|
|
||||||
echo " → .aider.conventions.md (root + project)"
|
|
||||||
}
|
|
||||||
|
|
||||||
# ── Generic system prompt (Open WebUI, Mods, etc.) ──────────
|
# ── Generic system prompt (Open WebUI, Mods, etc.) ──────────
|
||||||
generate_system_prompt() {
|
generate_system_prompt() {
|
||||||
{
|
{
|
||||||
@@ -142,8 +117,6 @@ echo "Syncing project context from $PROJECT_FILE..."
|
|||||||
if [ $# -eq 0 ]; then
|
if [ $# -eq 0 ]; then
|
||||||
generate_claude
|
generate_claude
|
||||||
generate_agents
|
generate_agents
|
||||||
generate_cursor
|
|
||||||
generate_aider
|
|
||||||
generate_system_prompt
|
generate_system_prompt
|
||||||
generate_mcp
|
generate_mcp
|
||||||
else
|
else
|
||||||
@@ -151,8 +124,6 @@ else
|
|||||||
case "$adapter" in
|
case "$adapter" in
|
||||||
claude) generate_claude ;;
|
claude) generate_claude ;;
|
||||||
agents) generate_agents ;;
|
agents) generate_agents ;;
|
||||||
cursor) generate_cursor ;;
|
|
||||||
aider) generate_aider ;;
|
|
||||||
prompt|system|openwebui|owui|generic) generate_system_prompt ;;
|
prompt|system|openwebui|owui|generic) generate_system_prompt ;;
|
||||||
mcp) generate_mcp ;;
|
mcp) generate_mcp ;;
|
||||||
*) echo "Unknown adapter: $adapter" ;;
|
*) echo "Unknown adapter: $adapter" ;;
|
||||||
|
|||||||
Reference in New Issue
Block a user