diff --git a/.aider.conventions.md b/.aider.conventions.md index a89d468..7b849a7 100644 --- a/.aider.conventions.md +++ b/.aider.conventions.md @@ -1,3 +1,282 @@ +# Agent context — Mathias workspace + + + +## Who I am + +I'm Mathias, a digital product manager and technology consultant based in Sweden. +I build software, research emerging tech, and deliver consulting engagements +for clients under NDA. I work across AI/ML, financial automation, web applications, +and climate/sustainability tech. + +## How I work with agents + +- I think like a product manager — I care about *why* before *how* +- I want agents to be opinionated and push back, not just execute blindly +- I prefer concise responses; skip ceremony and get to the point +- When I say "build this", I mean production-quality with tests, not a demo +- Ask me before making irreversible changes or adding heavy dependencies +- I work with confidential client data — never send it to cloud APIs unless I explicitly say it's OK + +## Behavior rules + +These rules apply to every task across every project, regardless of harness. + +0. **Pre-task ritual — before ANY implementation (non-negotiable).** Run this before writing a single line: + - **Query the brain** (`brain_query`) for the domain + symptom. If the result changes your approach, surface it before acting. 5 seconds beats 5 hours. + - **Load the relevant skill** — see trigger table in *Engineering Skills* below. + - **Write the failing test first.** Name the test before the function. If the target is untestable (e.g. `main()` wiring), extract the logic into a testable function first. No implementation without a red test. + - **State the observable success criterion** — what specific behavior, output, or passing test proves this is done? + + **TDD is non-negotiable.** "Tests pass" is not proof of correctness — only proof the tests ran. Write tests that would catch the bug before writing code that fixes it. + +1. **No assumptions.** Don't hide confusion — surface it. Surface tradeoffs explicitly. + Think before coding; if the problem is unclear, ask or state assumptions before acting. +2. **Minimum viable code.** Solve with the smallest change that works. Nothing + speculative, no "while we're here" cleanups, no premature abstractions. Simplicity first. +3. **Surgical changes.** Touch only what the task requires. Leave unrelated code, + files, and formatting alone. Diffs should be small and reviewable. +4. **Goal-driven execution.** Define clear success criteria up front for every task. + Loop — implement, verify, refine — until those criteria are met. Don't claim + completion without evidence (tests pass, command output, observed behavior). +5. **Trunk-Based Development — commit directly to main.** Every commit is one + logical change (one tool, one fix, one test) with passing tests. Main is always + deployable. Never create long-lived feature branches. + + **Exception — parallel agents on same repo:** If another agent is known to be + actively working on the same repo simultaneously, create a short-lived branch + (`agent/`), finish the task, and merge to main within the same + session. Do not leave agent branches open between sessions. + + **Exception — external contributor or client four-eyes requirement:** Use + PR flow only when a human reviewer outside the project is required. Document + the reason in PROJECT.md. + +6. **Close the loop — every substantive task ends with the same ritual.** Shipping + the code is not the end of the task; capturing it is. Run this unprompted: + - **Tag + bump SemVer** on the change (annotated tag; minor for a feature or + new/changed ADR, patch for a fix; docs in the same commit). Check the repo's + actual last tag — stated versions in docs drift stale. + - **Push** main and the tag (CI is the gate). + - **Persist generalizable learnings to the brain** (`brain_write`, wing/hall) — + the reusable patterns and the footguns that would bite anyone again, never + project status. See *Knowledge base — when to write* below. + - **File discovered-but-deferred work as tracker issues** on the project's own + repo — token-budget gaps, recorded ADR limitations, v2 follow-ups. Don't let + "out of scope, recorded" rot in a commit message; make it a ticket with a + source pointer. + - Surface the brain entries and issue numbers in the closing summary so the + trail is auditable. + +## Default stack + +| Layer | Default | Fallback | Last resort | +|-------|---------|----------|-------------| +| Language | Go | Python | TypeScript, Java, C | +| UI | HTMX + Templ | Server-rendered HTML | React (only if SPA is justified) | +| Build | Task (taskfile.dev) | Make | — | +| Containers | Docker Compose (dev), k3s (prod) | — | — | +| DB | PostgreSQL + sqlc | SQLite | — | +| Search | pgvector (vector), BM25 | Qdrant (when >1M vectors or hybrid retrieval) | — | +| Logging | slog (structured) | — | — | +| Testing | Table-driven, testify | — | — | +| Agents (Go) | google.golang.org/adk + pkg/litellm adapter | — | — | + +Exploratory: Rust, Zig — I'll tell you when I want these. + +## Code conventions + +- **Go style**: golines, gofumpt, golangci-lint +- **Errors**: `fmt.Errorf("operation: %w", err)` — never naked, never log-and-return +- **Naming**: stdlib conventions, no stuttering +- **Architecture**: prefer stdlib over frameworks, constructor injection, env-var config parsed into typed structs +- **Git**: conventional commits (`feat:`, `fix:`, `chore:`), commit directly to main, + one logical change per commit, CI is the quality gate +- **Never**: long-lived feature branches, PRs for solo work, direct push without + passing `task check` locally first +- **Security**: no secrets in code, govulncheck before adding deps, SOPS for encrypted config +- **Dependencies**: prefer stdlib. testify, slog, templ, sqlc, google.golang.org/adk (agent projects only) are pre-approved; anything else needs justification in the commit message + +## Secret handling (every harness, every command) + +Tool output is persisted: terminal → `~/.claude/projects` transcripts → +claudewatcher → brain/wiki → gitea history. A secret printed once is +searchable forever, and clearing it means rotating the key. So: + +1. **Never print, echo, log, or transform a secret to inspect it.** No + `base64`/`xxd`/`cat` of a key, and never pipe a secret through a transform + to defeat `op run`'s output masking (it masks raw values; base64 hides them + from the mask — that exact trick leaked a key on 2026-06-11). +2. **Secrets stay in the subprocess.** Reference them only as env vars consumed + *inside* `op run --env-file ~/.op-env -- `. Never place a literal secret + in a command's argv (it lands in the tool call and the transcript). +3. **Existence check without revealing the value:** `[ -n "$X" ] && echo set` — + never `${X:-...}` (returns the value when set) and never echo a substring of it. +4. **Cross-host secrets:** run the secret-consuming command on the host that has + the secret; do not forward a raw key over ssh argv/stdout. +5. If a secret does leak into output, say so immediately and flag it for rotation — + don't bury it. + +## Infrastructure + +Three machines on Tailscale: + +| Machine | Role | Key specs | +|---------|------|-----------| +| koala | GPU inference, heavy compute | RTX 5070, runs k3s + llama-swap + shared postgres18/pgvector | +| iguana | Services, builds | M2 Ultra Mac | +| flamingo | Daily driver, edge | Mac mini, ~/dev is here | + +- **Model routing**: LiteLLM in front of llama-swap (local) + cloud APIs (when permitted) +- **Orchestration**: k3s cluster across all three machines +- **Networking**: Tailscale mesh + +## Project landscape + +All development repos live at `~/dev/` (softlink from `~/Documents/local-dev/`). + +Organized in thematic folders: + +| Folder | Focus | Count | +|--------|-------|-------| +| `GO/` | Go web frameworks, API integrations, learning projects | ~10 | +| `AI/` | ML research, AI frameworks (FinRL, DSPy, crawl4ai) | ~6 | +| `AGENTS/` | Autonomous agents, coding agents, MCP servers, infra | ~15 | +| `QKX/` | Invoice processing, financial automation, payment systems | ~13 | +| `XT/` | Climate data, sustainability (Klimatkollen, Garbo) | ~2 | + +See `~/dev/PROJECT_SUMMARY.md` for detailed descriptions of each project. + +### Key active projects + +- **super-koala** (`AGENTS/`) — multi-component agent stack with LangGraph, DSPy, MCP +- **azure-tiger** (`QKX/`) — invoice extraction → ISO 20022 payment instructions +- **gocrwl** (`AGENTS/`) — Go web crawler with containerized deployment +- **koala-ai-stack** (`AGENTS/`) — local AI server infrastructure management +- **klimatkollen** (`XT/`) — Swedish municipal climate data platform + +## Knowledge base — actively use it + +A persistent brain (BM25 search + LLM-synthesised Q&A) survives across sessions, +hosts, and harnesses. It holds 100+ hard-won entries: infra incident postmortems, +Go pitfalls, framework gotchas, design principles, ADRs. **It is not optional +reference material — query it actively, not just when explicitly told.** + +### When to query (treat as a reflex) + +- **Before** starting a non-trivial task — search for prior art with the symptom + AND the system component ("how did we solve X in Y?"). 5 seconds beats 5 hours. +- **When debugging** — search for the error string, the stack frame, the affected + service. Past you may have already paid this tax. +- **Before adopting** a pattern, library, framework, or model name — check if it + was tried and rejected, or what the integration footguns are. +- **When making architectural decisions** — search for the domain + "ADR" or + "decision" to find prior reasoning before re-deriving it. +- **When a recommendation feels novel** — challenge yourself: "has this been + documented?" The brain often has it. + +### When to write + +After you discover something that **future-you would forget** and that **isn't +recoverable from the code, git log, or PR description alone**: + +- Bugs whose root cause is non-obvious and generalisable beyond this project. +- Framework / library / model-name quirks that bit you and would bite anyone. +- Design principles validated under fire (e.g. "every `_get` needs a `_list`"). +- Postmortems for incidents: what broke, why, how diagnosed, what to do next time. + +DON'T write project status, sprint progress, PR summaries, or "what I did this +session" — those rot fast and the originals are in git/gitea anyway. Brain +entries that age well are about *why*, *how to avoid*, and *what to do when*. + +### How to access (per harness) + +| Harness | Query | Write | +|---------|-------|-------| +| **Claude Code, Claude Desktop** | `brain_query` (BM25), `brain_answer` (LLM-synth + sources) MCP tools | `brain_write` MCP tool | +| **Crush, Pi, Antigravity, other MCP-capable** | same MCP server: `ingestion-brain` (via the `mcp__*_brain__*` namespace once authenticated) | same | +| **Anything HTTP-only (curl, scripts)** | `POST https://brain-mcp.d-ma.be/query` with `{"query":"..."}` (auth via `BRAIN_MCP_TOKEN`) | `POST .../write` with `{"content":"...","filename":"..."}` | +| **Browser / human inspection** | `https://git.d-ma.be/mathias/hyperguild` → `knowledge/` and `wiki/` markdown files | + +- **Scoping**: defaults to `public` collection; client projects filter to `{client}` + `public`. +- **Routing**: brain_answer's LLM uses berget.ai as primary, iguana ollama as + fallback. Both are configurable in the `supervisor/ingestion-deployment.yaml` + on the koala k3s cluster; don't hardcode local-only model names into the + berget URL (see knowledge entry on namespace mismatches). + +### Quick reflex checks + +If you find yourself about to say any of these out loud, you owe yourself a brain query first: + +- "I think the issue might be..." +- "Let me try X and see..." +- "I'll just write a script to..." +- "This is probably a new bug..." +- "Has anyone done this before?" — *yes, probably, go check.* + +## Client work rules + +When working on a project tagged with a client name: +1. Never send code, data, or context to cloud APIs — use local models only +2. Never reference other client projects or their data +3. Keep all artifacts within the client's git org / directory +4. Treat everything as confidential unless told otherwise + +## Harness-agnostic principles + +This context is designed to work with any AI coding tool: +- Claude Code, Cursor, Aider, Open WebUI, Charmbracelet Mods/Crush +- Pi Coding Agent, Mistral Vibe, Antigravity +- Any tool that accepts a system prompt or reads a markdown context file + +The canonical source is always `.context/AGENT.md` (root) and `.context/PROJECT.md` (per-project). +Derived files are committed (see *How context propagates* below) so a `git pull` on any host yields full agent context with no setup. + +## How context propagates + +Canonical sources of truth: +- Universal: `~/dev/.context/AGENT.md` (this file) +- Project: `/.context/PROJECT.md` (per-repo) + +Derived files (committed, regenerated by `task context:sync`): +- `CLAUDE.md`, `AGENTS.md`, `.cursorrules`, `.aider.conventions.md`, + `.context/system-prompt.txt` + +Workflow: +1. Edit a canonical file. Run `task context:sync`. Commit canonical and + derived together. Push. +2. On any other host, `git pull` brings both. Claude Code (tree-walking) + uses `CLAUDE.md`; Crush / Pi / Antigravity (cwd-only) use `AGENTS.md`; + Cursor uses `.cursorrules`; Aider uses `.aider.conventions.md`. +3. `task check` runs `context:sync` then asserts `git status --porcelain` + is empty over the derived files (catches both modified-tracked drift + and missing-untracked adapters). A drift fails the check with a + message telling you to stage the regenerated files. + +Behavior rules in this file and per-project rules in `PROJECT.md` apply +unconditionally on every host, every harness. + +## Engineering Skills + +Shared engineering skills live in the **`mathias/skills`** repo (`git.d-ma.be/mathias/skills`). Clone it to `~/dev/skills/` and run `SKILLS_CHECKOUT_DIR="$PWD" bash install.sh` there to wire every skill into your harnesses (Claude Code, Crush, Antigravity, Mistral Vibe) as native, on-demand skills. (Use `install.sh`, not `task install` — the latter is currently broken, skills#7.) Load at task start — not "on demand" but on schedule, before writing code. Browse `~/dev/skills/SKILLS_INDEX.md` for the full list. + +**Skill trigger table — load before starting, not after getting stuck:** + +| Task type | Load | +|-----------|------| +| Any feature or bug fix | `tdd` | +| Refactor or design | `clean-code` or `solid` | +| Debug | `problem-analysis` | +| Review code or PRs | `code-review` | +| Frame a problem before coding | `problem-analysis` | + +--- + # cad-atlas ## Identity @@ -7,7 +286,85 @@ - **Client**: personal - **Repo**: git.d-ma.be/mathias/cad-atlas - **Status**: active +- **Stack**: Go + Templ + HTMX + CDN Tailwind (template-go-web). Cross-project conventions: `~/dev/.context/AGENT.md`. -## Stack +## What this is -Go + Templ + HTMX + CDN Tailwind. See `~/dev/.context/AGENT.md` for cross-project conventions. +A visual **atlas of the Continuous Agentic Development (CAD) workflow** — the full path +from a captured signal to a deployed k3s pod, one screen, reel-style ("From Signal to Pod"). +It exists to (a) make the homelab's agentic delivery pipeline legible to a human, and +(b) render the CAD **audit chain** — which doubles as the regulated-industry audit artifact. + +> **Core thesis:** the CAD audit chain *is* the visualization data. +> `TELOS → goal → spec → issue → execution → attestation → deploy` is both the trace and +> the audit package. Phase C renders it once and serves two masters (observability + compliance). + +## Phases + +- **Phase A — static hero viz** (current). Self-contained `internal/web/static/cad-atlas.html`, + data-driven from a hand-authored `STAGES` array (ground-truth snapshot from brain, 2026-07-19). + Served at `/` by `internal/web/handler.go` via `go:embed`. Reel-parity: SVG spine with + arrowheads, animated pulse, dashed **feedback bus** (stage 08 → TELOS), replay + slow-mo. +- **Phase B — generated-from-source**. Parse brain docs + `.gitea/workflows` + infra manifests + → render the graph so it can't drift from config. +- **Phase C — live trace viewer**. Replace the static `STAGES` array with live reads of the + `assessor-loop` attestation ledger + brain `session_log` + Gitea run API + Flux events. + This is the prize: a real signal→pod trace viewer that is also the audit package. + +## The workflow it visualizes (9 stages) + +`00 Signals` (Applied AI Radar → mathias/signals) → `01 TELOS` (intention substrate) → +`02 Strategic session` (claude.ai frontier + LLM Council + Autoresearch Council) → +`03 Spec → Gitea issue` (agent-ready contract; Ed25519 admission #36; **var-go Oath**) → +`04 Human dispatch gate` (the only checkpoint; session-dispatch bridge → cad-dispatch.yml) → +`05 Execute · agentsquad` (serve/taskqueue, exec+review loop, risk LOW/MED/HIGH, dma-cli routing, +assessor-loop ledger) → `06 PR → CI` (go test/vet/lint/govulncheck + **var-go/oath gate**) → +`07 CD → pod` (Flux GitOps → k3s on koala) → `08 Loop back` (outcome scored vs TELOS goal). + +### Three orthogonal governance gates + +| Gate | Guards | Where | +|---|---|---| +| Ed25519 admission controller (#36) | spec **integrity** (issue untampered) | stage 03 | +| dispatch-allow (`.dispatch-allow` + `mathias/dispatch` allowlist) | repo **eligibility** (may agents run here) | stage 04/05 | +| **var-go Oath** (`cmd/vargo-gate`, commit status `var-go/oath`) | output **correctness** (PR satisfies the Oath; floor over reviewer, anti-rubber-stamp #55) | stage 06 | + +## Dogfooding + +This repo is built *through* the workflow it depicts. It is `dispatch-allow`-enabled, and its +own build increments are governed by a **var-go Oath** embedded in their spec issues (see the +Stage-03 tracking issue). Bootstrapping honesty (per swedsl honest-stub discipline): the Oath is +**defined** but `cmd/vargo-gate` is **not yet wired** into this repo's CI — until it is, the Oath +is advisory here. Wiring it is a first tracked task; disclosed in code, this doc, and CI config. + +## Brain references (source of truth — `brain_get `) + +- `knowledge/workflow-idea-to-running-service.md` — Double Diamond idea→service workflow +- `wiki/homelab/decisions/continuous-agentic-development-cad-concept-2026-06-16.md` — CAD definition +- `wiki/agentsquad/decisions/cad-dispatch-bridge.md` — claude.ai → agentsquad trigger path +- `wiki/agentsquad/facts/llm-council-design-and-first-runs-2026-06-21.md` — LLM Council +- `wiki/agentsquad/decisions/autoresearch-council-sibling-pipe.md` — Autoresearch Council +- `wiki/agentsquad/decisions/serve-http-task-api.md` — agentsquad serve/taskqueue +- `knowledge/var-go-anchor-to-span-spike-verdict.md` — var-go runner + CAD gate seam +- `knowledge/swedsl-vargo-sprint1-enforcement-teeth-verdict.md` — vargo-gate enforcement teeth +- `wiki/assessor-loop/decisions/assessor-loop-genesis.md` — attestation ledger (Phase C source) +- `wiki/homelab/facts/homelab-network-topology-reference.md` — koala/iguana/flamingo/piguard + +## External references + +- Inspiration reel — "From Inbox to Shipped" pipeline viz: https://www.instagram.com/reel/DY92L7bu27j/ +- karpathy/llm-council — origin of the Council pattern +- Double Diamond design process (Discover/Define/Develop/Deliver) + +## Run + +```bash +task check # lint + vet + test (CI gate) +task run # build + serve at http://localhost:8080 → the atlas +``` + +## Deploy note + +CD (`.gitea/workflows/cd.yml`) deploys to k3s namespace `cad-atlas` via Flux. Per the known +template-go-agent CD gap: the `deploy` job stays RED until `mathias/infra` has +`k3s/apps/cad-atlas/deployment.yaml`. `check` + `build` are the real bootstrap gate. diff --git a/.context/PROJECT.md b/.context/PROJECT.md index a89d468..222add0 100644 --- a/.context/PROJECT.md +++ b/.context/PROJECT.md @@ -7,7 +7,85 @@ - **Client**: personal - **Repo**: git.d-ma.be/mathias/cad-atlas - **Status**: active +- **Stack**: Go + Templ + HTMX + CDN Tailwind (template-go-web). Cross-project conventions: `~/dev/.context/AGENT.md`. -## Stack +## What this is -Go + Templ + HTMX + CDN Tailwind. See `~/dev/.context/AGENT.md` for cross-project conventions. +A visual **atlas of the Continuous Agentic Development (CAD) workflow** — the full path +from a captured signal to a deployed k3s pod, one screen, reel-style ("From Signal to Pod"). +It exists to (a) make the homelab's agentic delivery pipeline legible to a human, and +(b) render the CAD **audit chain** — which doubles as the regulated-industry audit artifact. + +> **Core thesis:** the CAD audit chain *is* the visualization data. +> `TELOS → goal → spec → issue → execution → attestation → deploy` is both the trace and +> the audit package. Phase C renders it once and serves two masters (observability + compliance). + +## Phases + +- **Phase A — static hero viz** (current). Self-contained `internal/web/static/cad-atlas.html`, + data-driven from a hand-authored `STAGES` array (ground-truth snapshot from brain, 2026-07-19). + Served at `/` by `internal/web/handler.go` via `go:embed`. Reel-parity: SVG spine with + arrowheads, animated pulse, dashed **feedback bus** (stage 08 → TELOS), replay + slow-mo. +- **Phase B — generated-from-source**. Parse brain docs + `.gitea/workflows` + infra manifests + → render the graph so it can't drift from config. +- **Phase C — live trace viewer**. Replace the static `STAGES` array with live reads of the + `assessor-loop` attestation ledger + brain `session_log` + Gitea run API + Flux events. + This is the prize: a real signal→pod trace viewer that is also the audit package. + +## The workflow it visualizes (9 stages) + +`00 Signals` (Applied AI Radar → mathias/signals) → `01 TELOS` (intention substrate) → +`02 Strategic session` (claude.ai frontier + LLM Council + Autoresearch Council) → +`03 Spec → Gitea issue` (agent-ready contract; Ed25519 admission #36; **var-go Oath**) → +`04 Human dispatch gate` (the only checkpoint; session-dispatch bridge → cad-dispatch.yml) → +`05 Execute · agentsquad` (serve/taskqueue, exec+review loop, risk LOW/MED/HIGH, dma-cli routing, +assessor-loop ledger) → `06 PR → CI` (go test/vet/lint/govulncheck + **var-go/oath gate**) → +`07 CD → pod` (Flux GitOps → k3s on koala) → `08 Loop back` (outcome scored vs TELOS goal). + +### Three orthogonal governance gates + +| Gate | Guards | Where | +|---|---|---| +| Ed25519 admission controller (#36) | spec **integrity** (issue untampered) | stage 03 | +| dispatch-allow (`.dispatch-allow` + `mathias/dispatch` allowlist) | repo **eligibility** (may agents run here) | stage 04/05 | +| **var-go Oath** (`cmd/vargo-gate`, commit status `var-go/oath`) | output **correctness** (PR satisfies the Oath; floor over reviewer, anti-rubber-stamp #55) | stage 06 | + +## Dogfooding + +This repo is built *through* the workflow it depicts. It is `dispatch-allow`-enabled, and its +own build increments are governed by a **var-go Oath** embedded in their spec issues (see the +Stage-03 tracking issue). Bootstrapping honesty (per swedsl honest-stub discipline): the Oath is +**defined** but `cmd/vargo-gate` is **not yet wired** into this repo's CI — until it is, the Oath +is advisory here. Wiring it is a first tracked task; disclosed in code, this doc, and CI config. + +## Brain references (source of truth — `brain_get `) + +- `knowledge/workflow-idea-to-running-service.md` — Double Diamond idea→service workflow +- `wiki/homelab/decisions/continuous-agentic-development-cad-concept-2026-06-16.md` — CAD definition +- `wiki/agentsquad/decisions/cad-dispatch-bridge.md` — claude.ai → agentsquad trigger path +- `wiki/agentsquad/facts/llm-council-design-and-first-runs-2026-06-21.md` — LLM Council +- `wiki/agentsquad/decisions/autoresearch-council-sibling-pipe.md` — Autoresearch Council +- `wiki/agentsquad/decisions/serve-http-task-api.md` — agentsquad serve/taskqueue +- `knowledge/var-go-anchor-to-span-spike-verdict.md` — var-go runner + CAD gate seam +- `knowledge/swedsl-vargo-sprint1-enforcement-teeth-verdict.md` — vargo-gate enforcement teeth +- `wiki/assessor-loop/decisions/assessor-loop-genesis.md` — attestation ledger (Phase C source) +- `wiki/homelab/facts/homelab-network-topology-reference.md` — koala/iguana/flamingo/piguard + +## External references + +- Inspiration reel — "From Inbox to Shipped" pipeline viz: https://www.instagram.com/reel/DY92L7bu27j/ +- karpathy/llm-council — origin of the Council pattern +- Double Diamond design process (Discover/Define/Develop/Deliver) + +## Run + +```bash +task check # lint + vet + test (CI gate) +task run # build + serve at http://localhost:8080 → the atlas +``` + +## Deploy note + +CD (`.gitea/workflows/cd.yml`) deploys to k3s namespace `cad-atlas` via Flux. Per the known +template-go-agent CD gap: the `deploy` job stays RED until `mathias/infra` has +`k3s/apps/cad-atlas/deployment.yaml`. `check` + `build` are the real bootstrap gate. diff --git a/.context/mcp.json b/.context/mcp.json index cd0f1bc..c9514c5 100644 --- a/.context/mcp.json +++ b/.context/mcp.json @@ -3,6 +3,24 @@ "knowledge": { "url": "http://localhost:3100/mcp", "description": "Project knowledge base — vector + graph retrieval" + }, + "brain": { + "type": "http", + "url": "https://brain-mcp.d-ma.be/mcp", + "headers": { + "Authorization": "Bearer ${BRAIN_MCP_TOKEN}" + } + }, + "gitea": { + "type": "http", + "url": "https://git-mcp.d-ma.be/mcp", + "headers": { + "Authorization": "Bearer ${GITEA_MCP_TOKEN}" + } + }, + "infra": { + "type": "http", + "url": "https://infra-mcp.d-ma.be/mcp" } } } diff --git a/.context/system-prompt.txt b/.context/system-prompt.txt index bf43d74..033b1bf 100644 --- a/.context/system-prompt.txt +++ b/.context/system-prompt.txt @@ -3,6 +3,285 @@ Follow all conventions from both the root agent context and project context. --- +# Agent context — Mathias workspace + + + +## Who I am + +I'm Mathias, a digital product manager and technology consultant based in Sweden. +I build software, research emerging tech, and deliver consulting engagements +for clients under NDA. I work across AI/ML, financial automation, web applications, +and climate/sustainability tech. + +## How I work with agents + +- I think like a product manager — I care about *why* before *how* +- I want agents to be opinionated and push back, not just execute blindly +- I prefer concise responses; skip ceremony and get to the point +- When I say "build this", I mean production-quality with tests, not a demo +- Ask me before making irreversible changes or adding heavy dependencies +- I work with confidential client data — never send it to cloud APIs unless I explicitly say it's OK + +## Behavior rules + +These rules apply to every task across every project, regardless of harness. + +0. **Pre-task ritual — before ANY implementation (non-negotiable).** Run this before writing a single line: + - **Query the brain** (`brain_query`) for the domain + symptom. If the result changes your approach, surface it before acting. 5 seconds beats 5 hours. + - **Load the relevant skill** — see trigger table in *Engineering Skills* below. + - **Write the failing test first.** Name the test before the function. If the target is untestable (e.g. `main()` wiring), extract the logic into a testable function first. No implementation without a red test. + - **State the observable success criterion** — what specific behavior, output, or passing test proves this is done? + + **TDD is non-negotiable.** "Tests pass" is not proof of correctness — only proof the tests ran. Write tests that would catch the bug before writing code that fixes it. + +1. **No assumptions.** Don't hide confusion — surface it. Surface tradeoffs explicitly. + Think before coding; if the problem is unclear, ask or state assumptions before acting. +2. **Minimum viable code.** Solve with the smallest change that works. Nothing + speculative, no "while we're here" cleanups, no premature abstractions. Simplicity first. +3. **Surgical changes.** Touch only what the task requires. Leave unrelated code, + files, and formatting alone. Diffs should be small and reviewable. +4. **Goal-driven execution.** Define clear success criteria up front for every task. + Loop — implement, verify, refine — until those criteria are met. Don't claim + completion without evidence (tests pass, command output, observed behavior). +5. **Trunk-Based Development — commit directly to main.** Every commit is one + logical change (one tool, one fix, one test) with passing tests. Main is always + deployable. Never create long-lived feature branches. + + **Exception — parallel agents on same repo:** If another agent is known to be + actively working on the same repo simultaneously, create a short-lived branch + (`agent/`), finish the task, and merge to main within the same + session. Do not leave agent branches open between sessions. + + **Exception — external contributor or client four-eyes requirement:** Use + PR flow only when a human reviewer outside the project is required. Document + the reason in PROJECT.md. + +6. **Close the loop — every substantive task ends with the same ritual.** Shipping + the code is not the end of the task; capturing it is. Run this unprompted: + - **Tag + bump SemVer** on the change (annotated tag; minor for a feature or + new/changed ADR, patch for a fix; docs in the same commit). Check the repo's + actual last tag — stated versions in docs drift stale. + - **Push** main and the tag (CI is the gate). + - **Persist generalizable learnings to the brain** (`brain_write`, wing/hall) — + the reusable patterns and the footguns that would bite anyone again, never + project status. See *Knowledge base — when to write* below. + - **File discovered-but-deferred work as tracker issues** on the project's own + repo — token-budget gaps, recorded ADR limitations, v2 follow-ups. Don't let + "out of scope, recorded" rot in a commit message; make it a ticket with a + source pointer. + - Surface the brain entries and issue numbers in the closing summary so the + trail is auditable. + +## Default stack + +| Layer | Default | Fallback | Last resort | +|-------|---------|----------|-------------| +| Language | Go | Python | TypeScript, Java, C | +| UI | HTMX + Templ | Server-rendered HTML | React (only if SPA is justified) | +| Build | Task (taskfile.dev) | Make | — | +| Containers | Docker Compose (dev), k3s (prod) | — | — | +| DB | PostgreSQL + sqlc | SQLite | — | +| Search | pgvector (vector), BM25 | Qdrant (when >1M vectors or hybrid retrieval) | — | +| Logging | slog (structured) | — | — | +| Testing | Table-driven, testify | — | — | +| Agents (Go) | google.golang.org/adk + pkg/litellm adapter | — | — | + +Exploratory: Rust, Zig — I'll tell you when I want these. + +## Code conventions + +- **Go style**: golines, gofumpt, golangci-lint +- **Errors**: `fmt.Errorf("operation: %w", err)` — never naked, never log-and-return +- **Naming**: stdlib conventions, no stuttering +- **Architecture**: prefer stdlib over frameworks, constructor injection, env-var config parsed into typed structs +- **Git**: conventional commits (`feat:`, `fix:`, `chore:`), commit directly to main, + one logical change per commit, CI is the quality gate +- **Never**: long-lived feature branches, PRs for solo work, direct push without + passing `task check` locally first +- **Security**: no secrets in code, govulncheck before adding deps, SOPS for encrypted config +- **Dependencies**: prefer stdlib. testify, slog, templ, sqlc, google.golang.org/adk (agent projects only) are pre-approved; anything else needs justification in the commit message + +## Secret handling (every harness, every command) + +Tool output is persisted: terminal → `~/.claude/projects` transcripts → +claudewatcher → brain/wiki → gitea history. A secret printed once is +searchable forever, and clearing it means rotating the key. So: + +1. **Never print, echo, log, or transform a secret to inspect it.** No + `base64`/`xxd`/`cat` of a key, and never pipe a secret through a transform + to defeat `op run`'s output masking (it masks raw values; base64 hides them + from the mask — that exact trick leaked a key on 2026-06-11). +2. **Secrets stay in the subprocess.** Reference them only as env vars consumed + *inside* `op run --env-file ~/.op-env -- `. Never place a literal secret + in a command's argv (it lands in the tool call and the transcript). +3. **Existence check without revealing the value:** `[ -n "$X" ] && echo set` — + never `${X:-...}` (returns the value when set) and never echo a substring of it. +4. **Cross-host secrets:** run the secret-consuming command on the host that has + the secret; do not forward a raw key over ssh argv/stdout. +5. If a secret does leak into output, say so immediately and flag it for rotation — + don't bury it. + +## Infrastructure + +Three machines on Tailscale: + +| Machine | Role | Key specs | +|---------|------|-----------| +| koala | GPU inference, heavy compute | RTX 5070, runs k3s + llama-swap + shared postgres18/pgvector | +| iguana | Services, builds | M2 Ultra Mac | +| flamingo | Daily driver, edge | Mac mini, ~/dev is here | + +- **Model routing**: LiteLLM in front of llama-swap (local) + cloud APIs (when permitted) +- **Orchestration**: k3s cluster across all three machines +- **Networking**: Tailscale mesh + +## Project landscape + +All development repos live at `~/dev/` (softlink from `~/Documents/local-dev/`). + +Organized in thematic folders: + +| Folder | Focus | Count | +|--------|-------|-------| +| `GO/` | Go web frameworks, API integrations, learning projects | ~10 | +| `AI/` | ML research, AI frameworks (FinRL, DSPy, crawl4ai) | ~6 | +| `AGENTS/` | Autonomous agents, coding agents, MCP servers, infra | ~15 | +| `QKX/` | Invoice processing, financial automation, payment systems | ~13 | +| `XT/` | Climate data, sustainability (Klimatkollen, Garbo) | ~2 | + +See `~/dev/PROJECT_SUMMARY.md` for detailed descriptions of each project. + +### Key active projects + +- **super-koala** (`AGENTS/`) — multi-component agent stack with LangGraph, DSPy, MCP +- **azure-tiger** (`QKX/`) — invoice extraction → ISO 20022 payment instructions +- **gocrwl** (`AGENTS/`) — Go web crawler with containerized deployment +- **koala-ai-stack** (`AGENTS/`) — local AI server infrastructure management +- **klimatkollen** (`XT/`) — Swedish municipal climate data platform + +## Knowledge base — actively use it + +A persistent brain (BM25 search + LLM-synthesised Q&A) survives across sessions, +hosts, and harnesses. It holds 100+ hard-won entries: infra incident postmortems, +Go pitfalls, framework gotchas, design principles, ADRs. **It is not optional +reference material — query it actively, not just when explicitly told.** + +### When to query (treat as a reflex) + +- **Before** starting a non-trivial task — search for prior art with the symptom + AND the system component ("how did we solve X in Y?"). 5 seconds beats 5 hours. +- **When debugging** — search for the error string, the stack frame, the affected + service. Past you may have already paid this tax. +- **Before adopting** a pattern, library, framework, or model name — check if it + was tried and rejected, or what the integration footguns are. +- **When making architectural decisions** — search for the domain + "ADR" or + "decision" to find prior reasoning before re-deriving it. +- **When a recommendation feels novel** — challenge yourself: "has this been + documented?" The brain often has it. + +### When to write + +After you discover something that **future-you would forget** and that **isn't +recoverable from the code, git log, or PR description alone**: + +- Bugs whose root cause is non-obvious and generalisable beyond this project. +- Framework / library / model-name quirks that bit you and would bite anyone. +- Design principles validated under fire (e.g. "every `_get` needs a `_list`"). +- Postmortems for incidents: what broke, why, how diagnosed, what to do next time. + +DON'T write project status, sprint progress, PR summaries, or "what I did this +session" — those rot fast and the originals are in git/gitea anyway. Brain +entries that age well are about *why*, *how to avoid*, and *what to do when*. + +### How to access (per harness) + +| Harness | Query | Write | +|---------|-------|-------| +| **Claude Code, Claude Desktop** | `brain_query` (BM25), `brain_answer` (LLM-synth + sources) MCP tools | `brain_write` MCP tool | +| **Crush, Pi, Antigravity, other MCP-capable** | same MCP server: `ingestion-brain` (via the `mcp__*_brain__*` namespace once authenticated) | same | +| **Anything HTTP-only (curl, scripts)** | `POST https://brain-mcp.d-ma.be/query` with `{"query":"..."}` (auth via `BRAIN_MCP_TOKEN`) | `POST .../write` with `{"content":"...","filename":"..."}` | +| **Browser / human inspection** | `https://git.d-ma.be/mathias/hyperguild` → `knowledge/` and `wiki/` markdown files | + +- **Scoping**: defaults to `public` collection; client projects filter to `{client}` + `public`. +- **Routing**: brain_answer's LLM uses berget.ai as primary, iguana ollama as + fallback. Both are configurable in the `supervisor/ingestion-deployment.yaml` + on the koala k3s cluster; don't hardcode local-only model names into the + berget URL (see knowledge entry on namespace mismatches). + +### Quick reflex checks + +If you find yourself about to say any of these out loud, you owe yourself a brain query first: + +- "I think the issue might be..." +- "Let me try X and see..." +- "I'll just write a script to..." +- "This is probably a new bug..." +- "Has anyone done this before?" — *yes, probably, go check.* + +## Client work rules + +When working on a project tagged with a client name: +1. Never send code, data, or context to cloud APIs — use local models only +2. Never reference other client projects or their data +3. Keep all artifacts within the client's git org / directory +4. Treat everything as confidential unless told otherwise + +## Harness-agnostic principles + +This context is designed to work with any AI coding tool: +- Claude Code, Cursor, Aider, Open WebUI, Charmbracelet Mods/Crush +- Pi Coding Agent, Mistral Vibe, Antigravity +- Any tool that accepts a system prompt or reads a markdown context file + +The canonical source is always `.context/AGENT.md` (root) and `.context/PROJECT.md` (per-project). +Derived files are committed (see *How context propagates* below) so a `git pull` on any host yields full agent context with no setup. + +## How context propagates + +Canonical sources of truth: +- Universal: `~/dev/.context/AGENT.md` (this file) +- Project: `/.context/PROJECT.md` (per-repo) + +Derived files (committed, regenerated by `task context:sync`): +- `CLAUDE.md`, `AGENTS.md`, `.cursorrules`, `.aider.conventions.md`, + `.context/system-prompt.txt` + +Workflow: +1. Edit a canonical file. Run `task context:sync`. Commit canonical and + derived together. Push. +2. On any other host, `git pull` brings both. Claude Code (tree-walking) + uses `CLAUDE.md`; Crush / Pi / Antigravity (cwd-only) use `AGENTS.md`; + Cursor uses `.cursorrules`; Aider uses `.aider.conventions.md`. +3. `task check` runs `context:sync` then asserts `git status --porcelain` + is empty over the derived files (catches both modified-tracked drift + and missing-untracked adapters). A drift fails the check with a + message telling you to stage the regenerated files. + +Behavior rules in this file and per-project rules in `PROJECT.md` apply +unconditionally on every host, every harness. + +## Engineering Skills + +Shared engineering skills live in the **`mathias/skills`** repo (`git.d-ma.be/mathias/skills`). Clone it to `~/dev/skills/` and run `SKILLS_CHECKOUT_DIR="$PWD" bash install.sh` there to wire every skill into your harnesses (Claude Code, Crush, Antigravity, Mistral Vibe) as native, on-demand skills. (Use `install.sh`, not `task install` — the latter is currently broken, skills#7.) Load at task start — not "on demand" but on schedule, before writing code. Browse `~/dev/skills/SKILLS_INDEX.md` for the full list. + +**Skill trigger table — load before starting, not after getting stuck:** + +| Task type | Load | +|-----------|------| +| Any feature or bug fix | `tdd` | +| Refactor or design | `clean-code` or `solid` | +| Debug | `problem-analysis` | +| Review code or PRs | `code-review` | +| Frame a problem before coding | `problem-analysis` | + +--- + # cad-atlas ## Identity @@ -12,9 +291,87 @@ Follow all conventions from both the root agent context and project context. - **Client**: personal - **Repo**: git.d-ma.be/mathias/cad-atlas - **Status**: active +- **Stack**: Go + Templ + HTMX + CDN Tailwind (template-go-web). Cross-project conventions: `~/dev/.context/AGENT.md`. -## Stack +## What this is -Go + Templ + HTMX + CDN Tailwind. See `~/dev/.context/AGENT.md` for cross-project conventions. +A visual **atlas of the Continuous Agentic Development (CAD) workflow** — the full path +from a captured signal to a deployed k3s pod, one screen, reel-style ("From Signal to Pod"). +It exists to (a) make the homelab's agentic delivery pipeline legible to a human, and +(b) render the CAD **audit chain** — which doubles as the regulated-industry audit artifact. + +> **Core thesis:** the CAD audit chain *is* the visualization data. +> `TELOS → goal → spec → issue → execution → attestation → deploy` is both the trace and +> the audit package. Phase C renders it once and serves two masters (observability + compliance). + +## Phases + +- **Phase A — static hero viz** (current). Self-contained `internal/web/static/cad-atlas.html`, + data-driven from a hand-authored `STAGES` array (ground-truth snapshot from brain, 2026-07-19). + Served at `/` by `internal/web/handler.go` via `go:embed`. Reel-parity: SVG spine with + arrowheads, animated pulse, dashed **feedback bus** (stage 08 → TELOS), replay + slow-mo. +- **Phase B — generated-from-source**. Parse brain docs + `.gitea/workflows` + infra manifests + → render the graph so it can't drift from config. +- **Phase C — live trace viewer**. Replace the static `STAGES` array with live reads of the + `assessor-loop` attestation ledger + brain `session_log` + Gitea run API + Flux events. + This is the prize: a real signal→pod trace viewer that is also the audit package. + +## The workflow it visualizes (9 stages) + +`00 Signals` (Applied AI Radar → mathias/signals) → `01 TELOS` (intention substrate) → +`02 Strategic session` (claude.ai frontier + LLM Council + Autoresearch Council) → +`03 Spec → Gitea issue` (agent-ready contract; Ed25519 admission #36; **var-go Oath**) → +`04 Human dispatch gate` (the only checkpoint; session-dispatch bridge → cad-dispatch.yml) → +`05 Execute · agentsquad` (serve/taskqueue, exec+review loop, risk LOW/MED/HIGH, dma-cli routing, +assessor-loop ledger) → `06 PR → CI` (go test/vet/lint/govulncheck + **var-go/oath gate**) → +`07 CD → pod` (Flux GitOps → k3s on koala) → `08 Loop back` (outcome scored vs TELOS goal). + +### Three orthogonal governance gates + +| Gate | Guards | Where | +|---|---|---| +| Ed25519 admission controller (#36) | spec **integrity** (issue untampered) | stage 03 | +| dispatch-allow (`.dispatch-allow` + `mathias/dispatch` allowlist) | repo **eligibility** (may agents run here) | stage 04/05 | +| **var-go Oath** (`cmd/vargo-gate`, commit status `var-go/oath`) | output **correctness** (PR satisfies the Oath; floor over reviewer, anti-rubber-stamp #55) | stage 06 | + +## Dogfooding + +This repo is built *through* the workflow it depicts. It is `dispatch-allow`-enabled, and its +own build increments are governed by a **var-go Oath** embedded in their spec issues (see the +Stage-03 tracking issue). Bootstrapping honesty (per swedsl honest-stub discipline): the Oath is +**defined** but `cmd/vargo-gate` is **not yet wired** into this repo's CI — until it is, the Oath +is advisory here. Wiring it is a first tracked task; disclosed in code, this doc, and CI config. + +## Brain references (source of truth — `brain_get `) + +- `knowledge/workflow-idea-to-running-service.md` — Double Diamond idea→service workflow +- `wiki/homelab/decisions/continuous-agentic-development-cad-concept-2026-06-16.md` — CAD definition +- `wiki/agentsquad/decisions/cad-dispatch-bridge.md` — claude.ai → agentsquad trigger path +- `wiki/agentsquad/facts/llm-council-design-and-first-runs-2026-06-21.md` — LLM Council +- `wiki/agentsquad/decisions/autoresearch-council-sibling-pipe.md` — Autoresearch Council +- `wiki/agentsquad/decisions/serve-http-task-api.md` — agentsquad serve/taskqueue +- `knowledge/var-go-anchor-to-span-spike-verdict.md` — var-go runner + CAD gate seam +- `knowledge/swedsl-vargo-sprint1-enforcement-teeth-verdict.md` — vargo-gate enforcement teeth +- `wiki/assessor-loop/decisions/assessor-loop-genesis.md` — attestation ledger (Phase C source) +- `wiki/homelab/facts/homelab-network-topology-reference.md` — koala/iguana/flamingo/piguard + +## External references + +- Inspiration reel — "From Inbox to Shipped" pipeline viz: https://www.instagram.com/reel/DY92L7bu27j/ +- karpathy/llm-council — origin of the Council pattern +- Double Diamond design process (Discover/Define/Develop/Deliver) + +## Run + +```bash +task check # lint + vet + test (CI gate) +task run # build + serve at http://localhost:8080 → the atlas +``` + +## Deploy note + +CD (`.gitea/workflows/cd.yml`) deploys to k3s namespace `cad-atlas` via Flux. Per the known +template-go-agent CD gap: the `deploy` job stays RED until `mathias/infra` has +`k3s/apps/cad-atlas/deployment.yaml`. `check` + `build` are the real bootstrap gate. --- diff --git a/.cursorrules b/.cursorrules index 433e0bd..66ddeb3 100644 --- a/.cursorrules +++ b/.cursorrules @@ -1,6 +1,285 @@ # Cursor rules — auto-generated # Do not edit. Run: task context:sync +# Agent context — Mathias workspace + + + +## Who I am + +I'm Mathias, a digital product manager and technology consultant based in Sweden. +I build software, research emerging tech, and deliver consulting engagements +for clients under NDA. I work across AI/ML, financial automation, web applications, +and climate/sustainability tech. + +## How I work with agents + +- I think like a product manager — I care about *why* before *how* +- I want agents to be opinionated and push back, not just execute blindly +- I prefer concise responses; skip ceremony and get to the point +- When I say "build this", I mean production-quality with tests, not a demo +- Ask me before making irreversible changes or adding heavy dependencies +- I work with confidential client data — never send it to cloud APIs unless I explicitly say it's OK + +## Behavior rules + +These rules apply to every task across every project, regardless of harness. + +0. **Pre-task ritual — before ANY implementation (non-negotiable).** Run this before writing a single line: + - **Query the brain** (`brain_query`) for the domain + symptom. If the result changes your approach, surface it before acting. 5 seconds beats 5 hours. + - **Load the relevant skill** — see trigger table in *Engineering Skills* below. + - **Write the failing test first.** Name the test before the function. If the target is untestable (e.g. `main()` wiring), extract the logic into a testable function first. No implementation without a red test. + - **State the observable success criterion** — what specific behavior, output, or passing test proves this is done? + + **TDD is non-negotiable.** "Tests pass" is not proof of correctness — only proof the tests ran. Write tests that would catch the bug before writing code that fixes it. + +1. **No assumptions.** Don't hide confusion — surface it. Surface tradeoffs explicitly. + Think before coding; if the problem is unclear, ask or state assumptions before acting. +2. **Minimum viable code.** Solve with the smallest change that works. Nothing + speculative, no "while we're here" cleanups, no premature abstractions. Simplicity first. +3. **Surgical changes.** Touch only what the task requires. Leave unrelated code, + files, and formatting alone. Diffs should be small and reviewable. +4. **Goal-driven execution.** Define clear success criteria up front for every task. + Loop — implement, verify, refine — until those criteria are met. Don't claim + completion without evidence (tests pass, command output, observed behavior). +5. **Trunk-Based Development — commit directly to main.** Every commit is one + logical change (one tool, one fix, one test) with passing tests. Main is always + deployable. Never create long-lived feature branches. + + **Exception — parallel agents on same repo:** If another agent is known to be + actively working on the same repo simultaneously, create a short-lived branch + (`agent/`), finish the task, and merge to main within the same + session. Do not leave agent branches open between sessions. + + **Exception — external contributor or client four-eyes requirement:** Use + PR flow only when a human reviewer outside the project is required. Document + the reason in PROJECT.md. + +6. **Close the loop — every substantive task ends with the same ritual.** Shipping + the code is not the end of the task; capturing it is. Run this unprompted: + - **Tag + bump SemVer** on the change (annotated tag; minor for a feature or + new/changed ADR, patch for a fix; docs in the same commit). Check the repo's + actual last tag — stated versions in docs drift stale. + - **Push** main and the tag (CI is the gate). + - **Persist generalizable learnings to the brain** (`brain_write`, wing/hall) — + the reusable patterns and the footguns that would bite anyone again, never + project status. See *Knowledge base — when to write* below. + - **File discovered-but-deferred work as tracker issues** on the project's own + repo — token-budget gaps, recorded ADR limitations, v2 follow-ups. Don't let + "out of scope, recorded" rot in a commit message; make it a ticket with a + source pointer. + - Surface the brain entries and issue numbers in the closing summary so the + trail is auditable. + +## Default stack + +| Layer | Default | Fallback | Last resort | +|-------|---------|----------|-------------| +| Language | Go | Python | TypeScript, Java, C | +| UI | HTMX + Templ | Server-rendered HTML | React (only if SPA is justified) | +| Build | Task (taskfile.dev) | Make | — | +| Containers | Docker Compose (dev), k3s (prod) | — | — | +| DB | PostgreSQL + sqlc | SQLite | — | +| Search | pgvector (vector), BM25 | Qdrant (when >1M vectors or hybrid retrieval) | — | +| Logging | slog (structured) | — | — | +| Testing | Table-driven, testify | — | — | +| Agents (Go) | google.golang.org/adk + pkg/litellm adapter | — | — | + +Exploratory: Rust, Zig — I'll tell you when I want these. + +## Code conventions + +- **Go style**: golines, gofumpt, golangci-lint +- **Errors**: `fmt.Errorf("operation: %w", err)` — never naked, never log-and-return +- **Naming**: stdlib conventions, no stuttering +- **Architecture**: prefer stdlib over frameworks, constructor injection, env-var config parsed into typed structs +- **Git**: conventional commits (`feat:`, `fix:`, `chore:`), commit directly to main, + one logical change per commit, CI is the quality gate +- **Never**: long-lived feature branches, PRs for solo work, direct push without + passing `task check` locally first +- **Security**: no secrets in code, govulncheck before adding deps, SOPS for encrypted config +- **Dependencies**: prefer stdlib. testify, slog, templ, sqlc, google.golang.org/adk (agent projects only) are pre-approved; anything else needs justification in the commit message + +## Secret handling (every harness, every command) + +Tool output is persisted: terminal → `~/.claude/projects` transcripts → +claudewatcher → brain/wiki → gitea history. A secret printed once is +searchable forever, and clearing it means rotating the key. So: + +1. **Never print, echo, log, or transform a secret to inspect it.** No + `base64`/`xxd`/`cat` of a key, and never pipe a secret through a transform + to defeat `op run`'s output masking (it masks raw values; base64 hides them + from the mask — that exact trick leaked a key on 2026-06-11). +2. **Secrets stay in the subprocess.** Reference them only as env vars consumed + *inside* `op run --env-file ~/.op-env -- `. Never place a literal secret + in a command's argv (it lands in the tool call and the transcript). +3. **Existence check without revealing the value:** `[ -n "$X" ] && echo set` — + never `${X:-...}` (returns the value when set) and never echo a substring of it. +4. **Cross-host secrets:** run the secret-consuming command on the host that has + the secret; do not forward a raw key over ssh argv/stdout. +5. If a secret does leak into output, say so immediately and flag it for rotation — + don't bury it. + +## Infrastructure + +Three machines on Tailscale: + +| Machine | Role | Key specs | +|---------|------|-----------| +| koala | GPU inference, heavy compute | RTX 5070, runs k3s + llama-swap + shared postgres18/pgvector | +| iguana | Services, builds | M2 Ultra Mac | +| flamingo | Daily driver, edge | Mac mini, ~/dev is here | + +- **Model routing**: LiteLLM in front of llama-swap (local) + cloud APIs (when permitted) +- **Orchestration**: k3s cluster across all three machines +- **Networking**: Tailscale mesh + +## Project landscape + +All development repos live at `~/dev/` (softlink from `~/Documents/local-dev/`). + +Organized in thematic folders: + +| Folder | Focus | Count | +|--------|-------|-------| +| `GO/` | Go web frameworks, API integrations, learning projects | ~10 | +| `AI/` | ML research, AI frameworks (FinRL, DSPy, crawl4ai) | ~6 | +| `AGENTS/` | Autonomous agents, coding agents, MCP servers, infra | ~15 | +| `QKX/` | Invoice processing, financial automation, payment systems | ~13 | +| `XT/` | Climate data, sustainability (Klimatkollen, Garbo) | ~2 | + +See `~/dev/PROJECT_SUMMARY.md` for detailed descriptions of each project. + +### Key active projects + +- **super-koala** (`AGENTS/`) — multi-component agent stack with LangGraph, DSPy, MCP +- **azure-tiger** (`QKX/`) — invoice extraction → ISO 20022 payment instructions +- **gocrwl** (`AGENTS/`) — Go web crawler with containerized deployment +- **koala-ai-stack** (`AGENTS/`) — local AI server infrastructure management +- **klimatkollen** (`XT/`) — Swedish municipal climate data platform + +## Knowledge base — actively use it + +A persistent brain (BM25 search + LLM-synthesised Q&A) survives across sessions, +hosts, and harnesses. It holds 100+ hard-won entries: infra incident postmortems, +Go pitfalls, framework gotchas, design principles, ADRs. **It is not optional +reference material — query it actively, not just when explicitly told.** + +### When to query (treat as a reflex) + +- **Before** starting a non-trivial task — search for prior art with the symptom + AND the system component ("how did we solve X in Y?"). 5 seconds beats 5 hours. +- **When debugging** — search for the error string, the stack frame, the affected + service. Past you may have already paid this tax. +- **Before adopting** a pattern, library, framework, or model name — check if it + was tried and rejected, or what the integration footguns are. +- **When making architectural decisions** — search for the domain + "ADR" or + "decision" to find prior reasoning before re-deriving it. +- **When a recommendation feels novel** — challenge yourself: "has this been + documented?" The brain often has it. + +### When to write + +After you discover something that **future-you would forget** and that **isn't +recoverable from the code, git log, or PR description alone**: + +- Bugs whose root cause is non-obvious and generalisable beyond this project. +- Framework / library / model-name quirks that bit you and would bite anyone. +- Design principles validated under fire (e.g. "every `_get` needs a `_list`"). +- Postmortems for incidents: what broke, why, how diagnosed, what to do next time. + +DON'T write project status, sprint progress, PR summaries, or "what I did this +session" — those rot fast and the originals are in git/gitea anyway. Brain +entries that age well are about *why*, *how to avoid*, and *what to do when*. + +### How to access (per harness) + +| Harness | Query | Write | +|---------|-------|-------| +| **Claude Code, Claude Desktop** | `brain_query` (BM25), `brain_answer` (LLM-synth + sources) MCP tools | `brain_write` MCP tool | +| **Crush, Pi, Antigravity, other MCP-capable** | same MCP server: `ingestion-brain` (via the `mcp__*_brain__*` namespace once authenticated) | same | +| **Anything HTTP-only (curl, scripts)** | `POST https://brain-mcp.d-ma.be/query` with `{"query":"..."}` (auth via `BRAIN_MCP_TOKEN`) | `POST .../write` with `{"content":"...","filename":"..."}` | +| **Browser / human inspection** | `https://git.d-ma.be/mathias/hyperguild` → `knowledge/` and `wiki/` markdown files | + +- **Scoping**: defaults to `public` collection; client projects filter to `{client}` + `public`. +- **Routing**: brain_answer's LLM uses berget.ai as primary, iguana ollama as + fallback. Both are configurable in the `supervisor/ingestion-deployment.yaml` + on the koala k3s cluster; don't hardcode local-only model names into the + berget URL (see knowledge entry on namespace mismatches). + +### Quick reflex checks + +If you find yourself about to say any of these out loud, you owe yourself a brain query first: + +- "I think the issue might be..." +- "Let me try X and see..." +- "I'll just write a script to..." +- "This is probably a new bug..." +- "Has anyone done this before?" — *yes, probably, go check.* + +## Client work rules + +When working on a project tagged with a client name: +1. Never send code, data, or context to cloud APIs — use local models only +2. Never reference other client projects or their data +3. Keep all artifacts within the client's git org / directory +4. Treat everything as confidential unless told otherwise + +## Harness-agnostic principles + +This context is designed to work with any AI coding tool: +- Claude Code, Cursor, Aider, Open WebUI, Charmbracelet Mods/Crush +- Pi Coding Agent, Mistral Vibe, Antigravity +- Any tool that accepts a system prompt or reads a markdown context file + +The canonical source is always `.context/AGENT.md` (root) and `.context/PROJECT.md` (per-project). +Derived files are committed (see *How context propagates* below) so a `git pull` on any host yields full agent context with no setup. + +## How context propagates + +Canonical sources of truth: +- Universal: `~/dev/.context/AGENT.md` (this file) +- Project: `/.context/PROJECT.md` (per-repo) + +Derived files (committed, regenerated by `task context:sync`): +- `CLAUDE.md`, `AGENTS.md`, `.cursorrules`, `.aider.conventions.md`, + `.context/system-prompt.txt` + +Workflow: +1. Edit a canonical file. Run `task context:sync`. Commit canonical and + derived together. Push. +2. On any other host, `git pull` brings both. Claude Code (tree-walking) + uses `CLAUDE.md`; Crush / Pi / Antigravity (cwd-only) use `AGENTS.md`; + Cursor uses `.cursorrules`; Aider uses `.aider.conventions.md`. +3. `task check` runs `context:sync` then asserts `git status --porcelain` + is empty over the derived files (catches both modified-tracked drift + and missing-untracked adapters). A drift fails the check with a + message telling you to stage the regenerated files. + +Behavior rules in this file and per-project rules in `PROJECT.md` apply +unconditionally on every host, every harness. + +## Engineering Skills + +Shared engineering skills live in the **`mathias/skills`** repo (`git.d-ma.be/mathias/skills`). Clone it to `~/dev/skills/` and run `SKILLS_CHECKOUT_DIR="$PWD" bash install.sh` there to wire every skill into your harnesses (Claude Code, Crush, Antigravity, Mistral Vibe) as native, on-demand skills. (Use `install.sh`, not `task install` — the latter is currently broken, skills#7.) Load at task start — not "on demand" but on schedule, before writing code. Browse `~/dev/skills/SKILLS_INDEX.md` for the full list. + +**Skill trigger table — load before starting, not after getting stuck:** + +| Task type | Load | +|-----------|------| +| Any feature or bug fix | `tdd` | +| Refactor or design | `clean-code` or `solid` | +| Debug | `problem-analysis` | +| Review code or PRs | `code-review` | +| Frame a problem before coding | `problem-analysis` | + +--- + # cad-atlas ## Identity @@ -10,7 +289,85 @@ - **Client**: personal - **Repo**: git.d-ma.be/mathias/cad-atlas - **Status**: active +- **Stack**: Go + Templ + HTMX + CDN Tailwind (template-go-web). Cross-project conventions: `~/dev/.context/AGENT.md`. -## Stack +## What this is -Go + Templ + HTMX + CDN Tailwind. See `~/dev/.context/AGENT.md` for cross-project conventions. +A visual **atlas of the Continuous Agentic Development (CAD) workflow** — the full path +from a captured signal to a deployed k3s pod, one screen, reel-style ("From Signal to Pod"). +It exists to (a) make the homelab's agentic delivery pipeline legible to a human, and +(b) render the CAD **audit chain** — which doubles as the regulated-industry audit artifact. + +> **Core thesis:** the CAD audit chain *is* the visualization data. +> `TELOS → goal → spec → issue → execution → attestation → deploy` is both the trace and +> the audit package. Phase C renders it once and serves two masters (observability + compliance). + +## Phases + +- **Phase A — static hero viz** (current). Self-contained `internal/web/static/cad-atlas.html`, + data-driven from a hand-authored `STAGES` array (ground-truth snapshot from brain, 2026-07-19). + Served at `/` by `internal/web/handler.go` via `go:embed`. Reel-parity: SVG spine with + arrowheads, animated pulse, dashed **feedback bus** (stage 08 → TELOS), replay + slow-mo. +- **Phase B — generated-from-source**. Parse brain docs + `.gitea/workflows` + infra manifests + → render the graph so it can't drift from config. +- **Phase C — live trace viewer**. Replace the static `STAGES` array with live reads of the + `assessor-loop` attestation ledger + brain `session_log` + Gitea run API + Flux events. + This is the prize: a real signal→pod trace viewer that is also the audit package. + +## The workflow it visualizes (9 stages) + +`00 Signals` (Applied AI Radar → mathias/signals) → `01 TELOS` (intention substrate) → +`02 Strategic session` (claude.ai frontier + LLM Council + Autoresearch Council) → +`03 Spec → Gitea issue` (agent-ready contract; Ed25519 admission #36; **var-go Oath**) → +`04 Human dispatch gate` (the only checkpoint; session-dispatch bridge → cad-dispatch.yml) → +`05 Execute · agentsquad` (serve/taskqueue, exec+review loop, risk LOW/MED/HIGH, dma-cli routing, +assessor-loop ledger) → `06 PR → CI` (go test/vet/lint/govulncheck + **var-go/oath gate**) → +`07 CD → pod` (Flux GitOps → k3s on koala) → `08 Loop back` (outcome scored vs TELOS goal). + +### Three orthogonal governance gates + +| Gate | Guards | Where | +|---|---|---| +| Ed25519 admission controller (#36) | spec **integrity** (issue untampered) | stage 03 | +| dispatch-allow (`.dispatch-allow` + `mathias/dispatch` allowlist) | repo **eligibility** (may agents run here) | stage 04/05 | +| **var-go Oath** (`cmd/vargo-gate`, commit status `var-go/oath`) | output **correctness** (PR satisfies the Oath; floor over reviewer, anti-rubber-stamp #55) | stage 06 | + +## Dogfooding + +This repo is built *through* the workflow it depicts. It is `dispatch-allow`-enabled, and its +own build increments are governed by a **var-go Oath** embedded in their spec issues (see the +Stage-03 tracking issue). Bootstrapping honesty (per swedsl honest-stub discipline): the Oath is +**defined** but `cmd/vargo-gate` is **not yet wired** into this repo's CI — until it is, the Oath +is advisory here. Wiring it is a first tracked task; disclosed in code, this doc, and CI config. + +## Brain references (source of truth — `brain_get `) + +- `knowledge/workflow-idea-to-running-service.md` — Double Diamond idea→service workflow +- `wiki/homelab/decisions/continuous-agentic-development-cad-concept-2026-06-16.md` — CAD definition +- `wiki/agentsquad/decisions/cad-dispatch-bridge.md` — claude.ai → agentsquad trigger path +- `wiki/agentsquad/facts/llm-council-design-and-first-runs-2026-06-21.md` — LLM Council +- `wiki/agentsquad/decisions/autoresearch-council-sibling-pipe.md` — Autoresearch Council +- `wiki/agentsquad/decisions/serve-http-task-api.md` — agentsquad serve/taskqueue +- `knowledge/var-go-anchor-to-span-spike-verdict.md` — var-go runner + CAD gate seam +- `knowledge/swedsl-vargo-sprint1-enforcement-teeth-verdict.md` — vargo-gate enforcement teeth +- `wiki/assessor-loop/decisions/assessor-loop-genesis.md` — attestation ledger (Phase C source) +- `wiki/homelab/facts/homelab-network-topology-reference.md` — koala/iguana/flamingo/piguard + +## External references + +- Inspiration reel — "From Inbox to Shipped" pipeline viz: https://www.instagram.com/reel/DY92L7bu27j/ +- karpathy/llm-council — origin of the Council pattern +- Double Diamond design process (Discover/Define/Develop/Deliver) + +## Run + +```bash +task check # lint + vet + test (CI gate) +task run # build + serve at http://localhost:8080 → the atlas +``` + +## Deploy note + +CD (`.gitea/workflows/cd.yml`) deploys to k3s namespace `cad-atlas` via Flux. Per the known +template-go-agent CD gap: the `deploy` job stays RED until `mathias/infra` has +`k3s/apps/cad-atlas/deployment.yaml`. `check` + `build` are the real bootstrap gate. diff --git a/.gitignore b/.gitignore index 902b7e7..a14f6ed 100644 --- a/.gitignore +++ b/.gitignore @@ -27,4 +27,4 @@ go.work.sum # Project-specific bin/ -*.templ.go +*_templ.go diff --git a/AGENTS.md b/AGENTS.md index a89d468..7b849a7 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -1,3 +1,282 @@ +# Agent context — Mathias workspace + + + +## Who I am + +I'm Mathias, a digital product manager and technology consultant based in Sweden. +I build software, research emerging tech, and deliver consulting engagements +for clients under NDA. I work across AI/ML, financial automation, web applications, +and climate/sustainability tech. + +## How I work with agents + +- I think like a product manager — I care about *why* before *how* +- I want agents to be opinionated and push back, not just execute blindly +- I prefer concise responses; skip ceremony and get to the point +- When I say "build this", I mean production-quality with tests, not a demo +- Ask me before making irreversible changes or adding heavy dependencies +- I work with confidential client data — never send it to cloud APIs unless I explicitly say it's OK + +## Behavior rules + +These rules apply to every task across every project, regardless of harness. + +0. **Pre-task ritual — before ANY implementation (non-negotiable).** Run this before writing a single line: + - **Query the brain** (`brain_query`) for the domain + symptom. If the result changes your approach, surface it before acting. 5 seconds beats 5 hours. + - **Load the relevant skill** — see trigger table in *Engineering Skills* below. + - **Write the failing test first.** Name the test before the function. If the target is untestable (e.g. `main()` wiring), extract the logic into a testable function first. No implementation without a red test. + - **State the observable success criterion** — what specific behavior, output, or passing test proves this is done? + + **TDD is non-negotiable.** "Tests pass" is not proof of correctness — only proof the tests ran. Write tests that would catch the bug before writing code that fixes it. + +1. **No assumptions.** Don't hide confusion — surface it. Surface tradeoffs explicitly. + Think before coding; if the problem is unclear, ask or state assumptions before acting. +2. **Minimum viable code.** Solve with the smallest change that works. Nothing + speculative, no "while we're here" cleanups, no premature abstractions. Simplicity first. +3. **Surgical changes.** Touch only what the task requires. Leave unrelated code, + files, and formatting alone. Diffs should be small and reviewable. +4. **Goal-driven execution.** Define clear success criteria up front for every task. + Loop — implement, verify, refine — until those criteria are met. Don't claim + completion without evidence (tests pass, command output, observed behavior). +5. **Trunk-Based Development — commit directly to main.** Every commit is one + logical change (one tool, one fix, one test) with passing tests. Main is always + deployable. Never create long-lived feature branches. + + **Exception — parallel agents on same repo:** If another agent is known to be + actively working on the same repo simultaneously, create a short-lived branch + (`agent/`), finish the task, and merge to main within the same + session. Do not leave agent branches open between sessions. + + **Exception — external contributor or client four-eyes requirement:** Use + PR flow only when a human reviewer outside the project is required. Document + the reason in PROJECT.md. + +6. **Close the loop — every substantive task ends with the same ritual.** Shipping + the code is not the end of the task; capturing it is. Run this unprompted: + - **Tag + bump SemVer** on the change (annotated tag; minor for a feature or + new/changed ADR, patch for a fix; docs in the same commit). Check the repo's + actual last tag — stated versions in docs drift stale. + - **Push** main and the tag (CI is the gate). + - **Persist generalizable learnings to the brain** (`brain_write`, wing/hall) — + the reusable patterns and the footguns that would bite anyone again, never + project status. See *Knowledge base — when to write* below. + - **File discovered-but-deferred work as tracker issues** on the project's own + repo — token-budget gaps, recorded ADR limitations, v2 follow-ups. Don't let + "out of scope, recorded" rot in a commit message; make it a ticket with a + source pointer. + - Surface the brain entries and issue numbers in the closing summary so the + trail is auditable. + +## Default stack + +| Layer | Default | Fallback | Last resort | +|-------|---------|----------|-------------| +| Language | Go | Python | TypeScript, Java, C | +| UI | HTMX + Templ | Server-rendered HTML | React (only if SPA is justified) | +| Build | Task (taskfile.dev) | Make | — | +| Containers | Docker Compose (dev), k3s (prod) | — | — | +| DB | PostgreSQL + sqlc | SQLite | — | +| Search | pgvector (vector), BM25 | Qdrant (when >1M vectors or hybrid retrieval) | — | +| Logging | slog (structured) | — | — | +| Testing | Table-driven, testify | — | — | +| Agents (Go) | google.golang.org/adk + pkg/litellm adapter | — | — | + +Exploratory: Rust, Zig — I'll tell you when I want these. + +## Code conventions + +- **Go style**: golines, gofumpt, golangci-lint +- **Errors**: `fmt.Errorf("operation: %w", err)` — never naked, never log-and-return +- **Naming**: stdlib conventions, no stuttering +- **Architecture**: prefer stdlib over frameworks, constructor injection, env-var config parsed into typed structs +- **Git**: conventional commits (`feat:`, `fix:`, `chore:`), commit directly to main, + one logical change per commit, CI is the quality gate +- **Never**: long-lived feature branches, PRs for solo work, direct push without + passing `task check` locally first +- **Security**: no secrets in code, govulncheck before adding deps, SOPS for encrypted config +- **Dependencies**: prefer stdlib. testify, slog, templ, sqlc, google.golang.org/adk (agent projects only) are pre-approved; anything else needs justification in the commit message + +## Secret handling (every harness, every command) + +Tool output is persisted: terminal → `~/.claude/projects` transcripts → +claudewatcher → brain/wiki → gitea history. A secret printed once is +searchable forever, and clearing it means rotating the key. So: + +1. **Never print, echo, log, or transform a secret to inspect it.** No + `base64`/`xxd`/`cat` of a key, and never pipe a secret through a transform + to defeat `op run`'s output masking (it masks raw values; base64 hides them + from the mask — that exact trick leaked a key on 2026-06-11). +2. **Secrets stay in the subprocess.** Reference them only as env vars consumed + *inside* `op run --env-file ~/.op-env -- `. Never place a literal secret + in a command's argv (it lands in the tool call and the transcript). +3. **Existence check without revealing the value:** `[ -n "$X" ] && echo set` — + never `${X:-...}` (returns the value when set) and never echo a substring of it. +4. **Cross-host secrets:** run the secret-consuming command on the host that has + the secret; do not forward a raw key over ssh argv/stdout. +5. If a secret does leak into output, say so immediately and flag it for rotation — + don't bury it. + +## Infrastructure + +Three machines on Tailscale: + +| Machine | Role | Key specs | +|---------|------|-----------| +| koala | GPU inference, heavy compute | RTX 5070, runs k3s + llama-swap + shared postgres18/pgvector | +| iguana | Services, builds | M2 Ultra Mac | +| flamingo | Daily driver, edge | Mac mini, ~/dev is here | + +- **Model routing**: LiteLLM in front of llama-swap (local) + cloud APIs (when permitted) +- **Orchestration**: k3s cluster across all three machines +- **Networking**: Tailscale mesh + +## Project landscape + +All development repos live at `~/dev/` (softlink from `~/Documents/local-dev/`). + +Organized in thematic folders: + +| Folder | Focus | Count | +|--------|-------|-------| +| `GO/` | Go web frameworks, API integrations, learning projects | ~10 | +| `AI/` | ML research, AI frameworks (FinRL, DSPy, crawl4ai) | ~6 | +| `AGENTS/` | Autonomous agents, coding agents, MCP servers, infra | ~15 | +| `QKX/` | Invoice processing, financial automation, payment systems | ~13 | +| `XT/` | Climate data, sustainability (Klimatkollen, Garbo) | ~2 | + +See `~/dev/PROJECT_SUMMARY.md` for detailed descriptions of each project. + +### Key active projects + +- **super-koala** (`AGENTS/`) — multi-component agent stack with LangGraph, DSPy, MCP +- **azure-tiger** (`QKX/`) — invoice extraction → ISO 20022 payment instructions +- **gocrwl** (`AGENTS/`) — Go web crawler with containerized deployment +- **koala-ai-stack** (`AGENTS/`) — local AI server infrastructure management +- **klimatkollen** (`XT/`) — Swedish municipal climate data platform + +## Knowledge base — actively use it + +A persistent brain (BM25 search + LLM-synthesised Q&A) survives across sessions, +hosts, and harnesses. It holds 100+ hard-won entries: infra incident postmortems, +Go pitfalls, framework gotchas, design principles, ADRs. **It is not optional +reference material — query it actively, not just when explicitly told.** + +### When to query (treat as a reflex) + +- **Before** starting a non-trivial task — search for prior art with the symptom + AND the system component ("how did we solve X in Y?"). 5 seconds beats 5 hours. +- **When debugging** — search for the error string, the stack frame, the affected + service. Past you may have already paid this tax. +- **Before adopting** a pattern, library, framework, or model name — check if it + was tried and rejected, or what the integration footguns are. +- **When making architectural decisions** — search for the domain + "ADR" or + "decision" to find prior reasoning before re-deriving it. +- **When a recommendation feels novel** — challenge yourself: "has this been + documented?" The brain often has it. + +### When to write + +After you discover something that **future-you would forget** and that **isn't +recoverable from the code, git log, or PR description alone**: + +- Bugs whose root cause is non-obvious and generalisable beyond this project. +- Framework / library / model-name quirks that bit you and would bite anyone. +- Design principles validated under fire (e.g. "every `_get` needs a `_list`"). +- Postmortems for incidents: what broke, why, how diagnosed, what to do next time. + +DON'T write project status, sprint progress, PR summaries, or "what I did this +session" — those rot fast and the originals are in git/gitea anyway. Brain +entries that age well are about *why*, *how to avoid*, and *what to do when*. + +### How to access (per harness) + +| Harness | Query | Write | +|---------|-------|-------| +| **Claude Code, Claude Desktop** | `brain_query` (BM25), `brain_answer` (LLM-synth + sources) MCP tools | `brain_write` MCP tool | +| **Crush, Pi, Antigravity, other MCP-capable** | same MCP server: `ingestion-brain` (via the `mcp__*_brain__*` namespace once authenticated) | same | +| **Anything HTTP-only (curl, scripts)** | `POST https://brain-mcp.d-ma.be/query` with `{"query":"..."}` (auth via `BRAIN_MCP_TOKEN`) | `POST .../write` with `{"content":"...","filename":"..."}` | +| **Browser / human inspection** | `https://git.d-ma.be/mathias/hyperguild` → `knowledge/` and `wiki/` markdown files | + +- **Scoping**: defaults to `public` collection; client projects filter to `{client}` + `public`. +- **Routing**: brain_answer's LLM uses berget.ai as primary, iguana ollama as + fallback. Both are configurable in the `supervisor/ingestion-deployment.yaml` + on the koala k3s cluster; don't hardcode local-only model names into the + berget URL (see knowledge entry on namespace mismatches). + +### Quick reflex checks + +If you find yourself about to say any of these out loud, you owe yourself a brain query first: + +- "I think the issue might be..." +- "Let me try X and see..." +- "I'll just write a script to..." +- "This is probably a new bug..." +- "Has anyone done this before?" — *yes, probably, go check.* + +## Client work rules + +When working on a project tagged with a client name: +1. Never send code, data, or context to cloud APIs — use local models only +2. Never reference other client projects or their data +3. Keep all artifacts within the client's git org / directory +4. Treat everything as confidential unless told otherwise + +## Harness-agnostic principles + +This context is designed to work with any AI coding tool: +- Claude Code, Cursor, Aider, Open WebUI, Charmbracelet Mods/Crush +- Pi Coding Agent, Mistral Vibe, Antigravity +- Any tool that accepts a system prompt or reads a markdown context file + +The canonical source is always `.context/AGENT.md` (root) and `.context/PROJECT.md` (per-project). +Derived files are committed (see *How context propagates* below) so a `git pull` on any host yields full agent context with no setup. + +## How context propagates + +Canonical sources of truth: +- Universal: `~/dev/.context/AGENT.md` (this file) +- Project: `/.context/PROJECT.md` (per-repo) + +Derived files (committed, regenerated by `task context:sync`): +- `CLAUDE.md`, `AGENTS.md`, `.cursorrules`, `.aider.conventions.md`, + `.context/system-prompt.txt` + +Workflow: +1. Edit a canonical file. Run `task context:sync`. Commit canonical and + derived together. Push. +2. On any other host, `git pull` brings both. Claude Code (tree-walking) + uses `CLAUDE.md`; Crush / Pi / Antigravity (cwd-only) use `AGENTS.md`; + Cursor uses `.cursorrules`; Aider uses `.aider.conventions.md`. +3. `task check` runs `context:sync` then asserts `git status --porcelain` + is empty over the derived files (catches both modified-tracked drift + and missing-untracked adapters). A drift fails the check with a + message telling you to stage the regenerated files. + +Behavior rules in this file and per-project rules in `PROJECT.md` apply +unconditionally on every host, every harness. + +## Engineering Skills + +Shared engineering skills live in the **`mathias/skills`** repo (`git.d-ma.be/mathias/skills`). Clone it to `~/dev/skills/` and run `SKILLS_CHECKOUT_DIR="$PWD" bash install.sh` there to wire every skill into your harnesses (Claude Code, Crush, Antigravity, Mistral Vibe) as native, on-demand skills. (Use `install.sh`, not `task install` — the latter is currently broken, skills#7.) Load at task start — not "on demand" but on schedule, before writing code. Browse `~/dev/skills/SKILLS_INDEX.md` for the full list. + +**Skill trigger table — load before starting, not after getting stuck:** + +| Task type | Load | +|-----------|------| +| Any feature or bug fix | `tdd` | +| Refactor or design | `clean-code` or `solid` | +| Debug | `problem-analysis` | +| Review code or PRs | `code-review` | +| Frame a problem before coding | `problem-analysis` | + +--- + # cad-atlas ## Identity @@ -7,7 +286,85 @@ - **Client**: personal - **Repo**: git.d-ma.be/mathias/cad-atlas - **Status**: active +- **Stack**: Go + Templ + HTMX + CDN Tailwind (template-go-web). Cross-project conventions: `~/dev/.context/AGENT.md`. -## Stack +## What this is -Go + Templ + HTMX + CDN Tailwind. See `~/dev/.context/AGENT.md` for cross-project conventions. +A visual **atlas of the Continuous Agentic Development (CAD) workflow** — the full path +from a captured signal to a deployed k3s pod, one screen, reel-style ("From Signal to Pod"). +It exists to (a) make the homelab's agentic delivery pipeline legible to a human, and +(b) render the CAD **audit chain** — which doubles as the regulated-industry audit artifact. + +> **Core thesis:** the CAD audit chain *is* the visualization data. +> `TELOS → goal → spec → issue → execution → attestation → deploy` is both the trace and +> the audit package. Phase C renders it once and serves two masters (observability + compliance). + +## Phases + +- **Phase A — static hero viz** (current). Self-contained `internal/web/static/cad-atlas.html`, + data-driven from a hand-authored `STAGES` array (ground-truth snapshot from brain, 2026-07-19). + Served at `/` by `internal/web/handler.go` via `go:embed`. Reel-parity: SVG spine with + arrowheads, animated pulse, dashed **feedback bus** (stage 08 → TELOS), replay + slow-mo. +- **Phase B — generated-from-source**. Parse brain docs + `.gitea/workflows` + infra manifests + → render the graph so it can't drift from config. +- **Phase C — live trace viewer**. Replace the static `STAGES` array with live reads of the + `assessor-loop` attestation ledger + brain `session_log` + Gitea run API + Flux events. + This is the prize: a real signal→pod trace viewer that is also the audit package. + +## The workflow it visualizes (9 stages) + +`00 Signals` (Applied AI Radar → mathias/signals) → `01 TELOS` (intention substrate) → +`02 Strategic session` (claude.ai frontier + LLM Council + Autoresearch Council) → +`03 Spec → Gitea issue` (agent-ready contract; Ed25519 admission #36; **var-go Oath**) → +`04 Human dispatch gate` (the only checkpoint; session-dispatch bridge → cad-dispatch.yml) → +`05 Execute · agentsquad` (serve/taskqueue, exec+review loop, risk LOW/MED/HIGH, dma-cli routing, +assessor-loop ledger) → `06 PR → CI` (go test/vet/lint/govulncheck + **var-go/oath gate**) → +`07 CD → pod` (Flux GitOps → k3s on koala) → `08 Loop back` (outcome scored vs TELOS goal). + +### Three orthogonal governance gates + +| Gate | Guards | Where | +|---|---|---| +| Ed25519 admission controller (#36) | spec **integrity** (issue untampered) | stage 03 | +| dispatch-allow (`.dispatch-allow` + `mathias/dispatch` allowlist) | repo **eligibility** (may agents run here) | stage 04/05 | +| **var-go Oath** (`cmd/vargo-gate`, commit status `var-go/oath`) | output **correctness** (PR satisfies the Oath; floor over reviewer, anti-rubber-stamp #55) | stage 06 | + +## Dogfooding + +This repo is built *through* the workflow it depicts. It is `dispatch-allow`-enabled, and its +own build increments are governed by a **var-go Oath** embedded in their spec issues (see the +Stage-03 tracking issue). Bootstrapping honesty (per swedsl honest-stub discipline): the Oath is +**defined** but `cmd/vargo-gate` is **not yet wired** into this repo's CI — until it is, the Oath +is advisory here. Wiring it is a first tracked task; disclosed in code, this doc, and CI config. + +## Brain references (source of truth — `brain_get `) + +- `knowledge/workflow-idea-to-running-service.md` — Double Diamond idea→service workflow +- `wiki/homelab/decisions/continuous-agentic-development-cad-concept-2026-06-16.md` — CAD definition +- `wiki/agentsquad/decisions/cad-dispatch-bridge.md` — claude.ai → agentsquad trigger path +- `wiki/agentsquad/facts/llm-council-design-and-first-runs-2026-06-21.md` — LLM Council +- `wiki/agentsquad/decisions/autoresearch-council-sibling-pipe.md` — Autoresearch Council +- `wiki/agentsquad/decisions/serve-http-task-api.md` — agentsquad serve/taskqueue +- `knowledge/var-go-anchor-to-span-spike-verdict.md` — var-go runner + CAD gate seam +- `knowledge/swedsl-vargo-sprint1-enforcement-teeth-verdict.md` — vargo-gate enforcement teeth +- `wiki/assessor-loop/decisions/assessor-loop-genesis.md` — attestation ledger (Phase C source) +- `wiki/homelab/facts/homelab-network-topology-reference.md` — koala/iguana/flamingo/piguard + +## External references + +- Inspiration reel — "From Inbox to Shipped" pipeline viz: https://www.instagram.com/reel/DY92L7bu27j/ +- karpathy/llm-council — origin of the Council pattern +- Double Diamond design process (Discover/Define/Develop/Deliver) + +## Run + +```bash +task check # lint + vet + test (CI gate) +task run # build + serve at http://localhost:8080 → the atlas +``` + +## Deploy note + +CD (`.gitea/workflows/cd.yml`) deploys to k3s namespace `cad-atlas` via Flux. Per the known +template-go-agent CD gap: the `deploy` job stays RED until `mathias/infra` has +`k3s/apps/cad-atlas/deployment.yaml`. `check` + `build` are the real bootstrap gate. diff --git a/CLAUDE.md b/CLAUDE.md index a89d468..222add0 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -7,7 +7,85 @@ - **Client**: personal - **Repo**: git.d-ma.be/mathias/cad-atlas - **Status**: active +- **Stack**: Go + Templ + HTMX + CDN Tailwind (template-go-web). Cross-project conventions: `~/dev/.context/AGENT.md`. -## Stack +## What this is -Go + Templ + HTMX + CDN Tailwind. See `~/dev/.context/AGENT.md` for cross-project conventions. +A visual **atlas of the Continuous Agentic Development (CAD) workflow** — the full path +from a captured signal to a deployed k3s pod, one screen, reel-style ("From Signal to Pod"). +It exists to (a) make the homelab's agentic delivery pipeline legible to a human, and +(b) render the CAD **audit chain** — which doubles as the regulated-industry audit artifact. + +> **Core thesis:** the CAD audit chain *is* the visualization data. +> `TELOS → goal → spec → issue → execution → attestation → deploy` is both the trace and +> the audit package. Phase C renders it once and serves two masters (observability + compliance). + +## Phases + +- **Phase A — static hero viz** (current). Self-contained `internal/web/static/cad-atlas.html`, + data-driven from a hand-authored `STAGES` array (ground-truth snapshot from brain, 2026-07-19). + Served at `/` by `internal/web/handler.go` via `go:embed`. Reel-parity: SVG spine with + arrowheads, animated pulse, dashed **feedback bus** (stage 08 → TELOS), replay + slow-mo. +- **Phase B — generated-from-source**. Parse brain docs + `.gitea/workflows` + infra manifests + → render the graph so it can't drift from config. +- **Phase C — live trace viewer**. Replace the static `STAGES` array with live reads of the + `assessor-loop` attestation ledger + brain `session_log` + Gitea run API + Flux events. + This is the prize: a real signal→pod trace viewer that is also the audit package. + +## The workflow it visualizes (9 stages) + +`00 Signals` (Applied AI Radar → mathias/signals) → `01 TELOS` (intention substrate) → +`02 Strategic session` (claude.ai frontier + LLM Council + Autoresearch Council) → +`03 Spec → Gitea issue` (agent-ready contract; Ed25519 admission #36; **var-go Oath**) → +`04 Human dispatch gate` (the only checkpoint; session-dispatch bridge → cad-dispatch.yml) → +`05 Execute · agentsquad` (serve/taskqueue, exec+review loop, risk LOW/MED/HIGH, dma-cli routing, +assessor-loop ledger) → `06 PR → CI` (go test/vet/lint/govulncheck + **var-go/oath gate**) → +`07 CD → pod` (Flux GitOps → k3s on koala) → `08 Loop back` (outcome scored vs TELOS goal). + +### Three orthogonal governance gates + +| Gate | Guards | Where | +|---|---|---| +| Ed25519 admission controller (#36) | spec **integrity** (issue untampered) | stage 03 | +| dispatch-allow (`.dispatch-allow` + `mathias/dispatch` allowlist) | repo **eligibility** (may agents run here) | stage 04/05 | +| **var-go Oath** (`cmd/vargo-gate`, commit status `var-go/oath`) | output **correctness** (PR satisfies the Oath; floor over reviewer, anti-rubber-stamp #55) | stage 06 | + +## Dogfooding + +This repo is built *through* the workflow it depicts. It is `dispatch-allow`-enabled, and its +own build increments are governed by a **var-go Oath** embedded in their spec issues (see the +Stage-03 tracking issue). Bootstrapping honesty (per swedsl honest-stub discipline): the Oath is +**defined** but `cmd/vargo-gate` is **not yet wired** into this repo's CI — until it is, the Oath +is advisory here. Wiring it is a first tracked task; disclosed in code, this doc, and CI config. + +## Brain references (source of truth — `brain_get `) + +- `knowledge/workflow-idea-to-running-service.md` — Double Diamond idea→service workflow +- `wiki/homelab/decisions/continuous-agentic-development-cad-concept-2026-06-16.md` — CAD definition +- `wiki/agentsquad/decisions/cad-dispatch-bridge.md` — claude.ai → agentsquad trigger path +- `wiki/agentsquad/facts/llm-council-design-and-first-runs-2026-06-21.md` — LLM Council +- `wiki/agentsquad/decisions/autoresearch-council-sibling-pipe.md` — Autoresearch Council +- `wiki/agentsquad/decisions/serve-http-task-api.md` — agentsquad serve/taskqueue +- `knowledge/var-go-anchor-to-span-spike-verdict.md` — var-go runner + CAD gate seam +- `knowledge/swedsl-vargo-sprint1-enforcement-teeth-verdict.md` — vargo-gate enforcement teeth +- `wiki/assessor-loop/decisions/assessor-loop-genesis.md` — attestation ledger (Phase C source) +- `wiki/homelab/facts/homelab-network-topology-reference.md` — koala/iguana/flamingo/piguard + +## External references + +- Inspiration reel — "From Inbox to Shipped" pipeline viz: https://www.instagram.com/reel/DY92L7bu27j/ +- karpathy/llm-council — origin of the Council pattern +- Double Diamond design process (Discover/Define/Develop/Deliver) + +## Run + +```bash +task check # lint + vet + test (CI gate) +task run # build + serve at http://localhost:8080 → the atlas +``` + +## Deploy note + +CD (`.gitea/workflows/cd.yml`) deploys to k3s namespace `cad-atlas` via Flux. Per the known +template-go-agent CD gap: the `deploy` job stays RED until `mathias/infra` has +`k3s/apps/cad-atlas/deployment.yaml`. `check` + `build` are the real bootstrap gate. diff --git a/README.md b/README.md index 5509738..31c059b 100644 --- a/README.md +++ b/README.md @@ -1,13 +1,28 @@ # cad-atlas -> Generated from `mathias/template-go-web`. +**A visual atlas of the Continuous Agentic Development (CAD) workflow — from a captured signal to a deployed k3s pod.** -## Bootstrap +One screen, reel-style: `Signals → TELOS → Strategic session → Spec/Oath → Human gate → agentsquad → CI → CD → pod`, looping back to TELOS. It renders the CAD **audit chain**, which doubles as the regulated-industry audit artifact. -After creating from template, run: +> The CAD audit chain *is* the visualization data. Phase C renders it once and serves two masters: observability + compliance. + +## Phases + +- **A — static hero viz** (current): self-contained `internal/web/static/cad-atlas.html`, served at `/`. SVG spine, animated pulse, dashed feedback bus, replay/slow-mo. +- **B — generated-from-source**: render from brain docs + workflows + infra manifests (no drift). +- **C — live trace viewer**: hydrate from the `assessor-loop` ledger + `session_log` + Gitea run API + Flux events. + +## Quickstart ```bash -go mod tidy # regenerate go.sum with real module path -task generate # generate templ files -task build # build the binary +task check # lint + vet + test (CI gate) +task run # → http://localhost:8080 ``` + +## Context + +Full project context, the 9-stage workflow, the three governance gates, the dogfooding model, and all **brain deep-links + external references** live in [`.context/PROJECT.md`](.context/PROJECT.md). Any clean agent session should read that first (`brain_get ` for the linked notes). + +## Governance (dogfooded) + +This repo is built *through* the workflow it depicts: `dispatch-allow`-enabled, build increments governed by a **var-go Oath** in their spec issues. Bootstrapping honesty: the Oath is defined but `cmd/vargo-gate` is not yet wired into this repo's CI — advisory until it is. See PROJECT.md. diff --git a/go.mod b/go.mod index 6bfb577..fa7f9bc 100644 --- a/go.mod +++ b/go.mod @@ -2,6 +2,4 @@ module git.d-ma.be/mathias/cad-atlas go 1.26 -require ( - github.com/a-h/templ v0.2.778 -) +require github.com/a-h/templ v0.3.1020 diff --git a/go.sum b/go.sum new file mode 100644 index 0000000..0857b6f --- /dev/null +++ b/go.sum @@ -0,0 +1,4 @@ +github.com/a-h/templ v0.3.1020 h1:ypAT/L5ySWEnZ6Zft/5yfoWXYYkhFNvEFOeeqecg4tw= +github.com/a-h/templ v0.3.1020/go.mod h1:A2DlK61v+K+NRoGnhmYbNYVmtYHcFO5/AisMvBdDxTM= +github.com/google/go-cmp v0.6.0 h1:ofyhxvXcZhMsU5ulbFiLKl/XBFqE1GSq7atu8tAmTRI= +github.com/google/go-cmp v0.6.0/go.mod h1:17dUlkBOakJ0+DkrSSNjCkIjxS6bF9zb3elmeNGIjoY= diff --git a/internal/web/handler.go b/internal/web/handler.go index 92cf8a1..adb6c19 100644 --- a/internal/web/handler.go +++ b/internal/web/handler.go @@ -2,13 +2,28 @@ package web import ( "context" + _ "embed" "net/http" ) +// atlasHTML is the Phase-A static hero visualization. Phase C replaces this +// self-contained file with a Templ view hydrated from live CAD trace data +// (assessor-loop ledger, session_log, Gitea run API, Flux events). +// +//go:embed static/cad-atlas.html +var atlasHTML []byte + +// NewHandler serves the CAD Atlas. Root ("/") returns the static atlas; +// /api/hello is a leftover template probe kept until Phase C wires real endpoints. func NewHandler() http.Handler { mux := http.NewServeMux() mux.HandleFunc("/", func(w http.ResponseWriter, r *http.Request) { - _ = Index().Render(context.Background(), w) + if r.URL.Path != "/" { + http.NotFound(w, r) + return + } + w.Header().Set("Content-Type", "text/html; charset=utf-8") + _, _ = w.Write(atlasHTML) }) mux.HandleFunc("/api/hello", func(w http.ResponseWriter, r *http.Request) { _ = Hello("world").Render(context.Background(), w) diff --git a/internal/web/handler_test.go b/internal/web/handler_test.go new file mode 100644 index 0000000..e5bbe0e --- /dev/null +++ b/internal/web/handler_test.go @@ -0,0 +1,43 @@ +package web + +import ( + "net/http" + "net/http/httptest" + "strings" + "testing" +) + +func TestRootServesAtlas(t *testing.T) { + srv := httptest.NewServer(NewHandler()) + defer srv.Close() + + resp, err := http.Get(srv.URL + "/") + if err != nil { + t.Fatalf("GET /: %v", err) + } + defer func() { _ = resp.Body.Close() }() + + if resp.StatusCode != http.StatusOK { + t.Fatalf("GET / status = %d, want 200", resp.StatusCode) + } + buf := make([]byte, 4096) + n, _ := resp.Body.Read(buf) + if !strings.Contains(string(buf[:n]), "CAD") { + t.Fatalf("GET / body missing atlas marker %q", "CAD") + } +} + +func TestUnknownPath404(t *testing.T) { + srv := httptest.NewServer(NewHandler()) + defer srv.Close() + + resp, err := http.Get(srv.URL + "/nope") + if err != nil { + t.Fatalf("GET /nope: %v", err) + } + defer func() { _ = resp.Body.Close() }() + + if resp.StatusCode != http.StatusNotFound { + t.Fatalf("GET /nope status = %d, want 404", resp.StatusCode) + } +} diff --git a/internal/web/static/cad-atlas.html b/internal/web/static/cad-atlas.html new file mode 100644 index 0000000..3b70b89 --- /dev/null +++ b/internal/web/static/cad-atlas.html @@ -0,0 +1,272 @@ + + + + + +CAD Atlas · From Signal to Pod + + + +
+

CAD Atlas · From Signal to Pod

+ one human gate · everything up- and downstream is agents · v0.3 static snapshot (→ live in Phase C) +
+ + +
+
+ +
SUBSTRATE
+ +
+
+ + + + + + + + + + +
+
+
+ +
+ CAD → CI → CD · intent→specify→dispatch · build→test→validate · deploy→ship. + Dashed violet = feedback bus (stage 08 → TELOS: deploy outcome scored vs originating goal). + Data: static inventory from brain (2026-07-19). Phase C swaps these arrays for live reads of + assessor-loop ledger · session_log · Gitea run API · Flux events. +
+ + + +