# Cursor rules — auto-generated # Do not edit. Run: task context:sync # Agent context — Mathias workspace ## Who I am I'm Mathias, a digital product manager and technology consultant based in Sweden. I build software, research emerging tech, and deliver consulting engagements for clients under NDA. I work across AI/ML, financial automation, web applications, and climate/sustainability tech. ## How I work with agents - I think like a product manager — I care about *why* before *how* - I want agents to be opinionated and push back, not just execute blindly - I prefer concise responses; skip ceremony and get to the point - When I say "build this", I mean production-quality with tests, not a demo - Ask me before making irreversible changes or adding heavy dependencies - I work with confidential client data — never send it to cloud APIs unless I explicitly say it's OK ## Behavior rules These rules apply to every task across every project, regardless of harness. 0. **Pre-task ritual — before ANY implementation (non-negotiable).** Run this before writing a single line: - **Query the brain** (`brain_query`) for the domain + symptom. If the result changes your approach, surface it before acting. 5 seconds beats 5 hours. - **Load the relevant skill** — see trigger table in *Engineering Skills* below. - **Write the failing test first.** Name the test before the function. If the target is untestable (e.g. `main()` wiring), extract the logic into a testable function first. No implementation without a red test. - **State the observable success criterion** — what specific behavior, output, or passing test proves this is done? **TDD is non-negotiable.** "Tests pass" is not proof of correctness — only proof the tests ran. Write tests that would catch the bug before writing code that fixes it. 1. **No assumptions.** Don't hide confusion — surface it. Surface tradeoffs explicitly. Think before coding; if the problem is unclear, ask or state assumptions before acting. 2. **Minimum viable code.** Solve with the smallest change that works. Nothing speculative, no "while we're here" cleanups, no premature abstractions. Simplicity first. 3. **Surgical changes.** Touch only what the task requires. Leave unrelated code, files, and formatting alone. Diffs should be small and reviewable. 4. **Goal-driven execution.** Define clear success criteria up front for every task. Loop — implement, verify, refine — until those criteria are met. Don't claim completion without evidence (tests pass, command output, observed behavior). 5. **Trunk-Based Development — commit directly to main.** Every commit is one logical change (one tool, one fix, one test) with passing tests. Main is always deployable. Never create long-lived feature branches. **Exception — parallel agents on same repo:** If another agent is known to be actively working on the same repo simultaneously, create a short-lived branch (`agent/`), finish the task, and merge to main within the same session. Do not leave agent branches open between sessions. **Exception — external contributor or client four-eyes requirement:** Use PR flow only when a human reviewer outside the project is required. Document the reason in PROJECT.md. 6. **Close the loop — every substantive task ends with the same ritual.** Shipping the code is not the end of the task; capturing it is. Run this unprompted: - **Tag + bump SemVer** on the change (annotated tag; minor for a feature or new/changed ADR, patch for a fix; docs in the same commit). Check the repo's actual last tag — stated versions in docs drift stale. - **Push** main and the tag (CI is the gate). - **Persist generalizable learnings to the brain** (`brain_write`, wing/hall) — the reusable patterns and the footguns that would bite anyone again, never project status. See *Knowledge base — when to write* below. - **File discovered-but-deferred work as tracker issues** on the project's own repo — token-budget gaps, recorded ADR limitations, v2 follow-ups. Don't let "out of scope, recorded" rot in a commit message; make it a ticket with a source pointer. - Surface the brain entries and issue numbers in the closing summary so the trail is auditable. ## Default stack | Layer | Default | Fallback | Last resort | |-------|---------|----------|-------------| | Language | Go | Python | TypeScript, Java, C | | UI | HTMX + Templ | Server-rendered HTML | React (only if SPA is justified) | | Build | Task (taskfile.dev) | Make | — | | Containers | Docker Compose (dev), k3s (prod) | — | — | | DB | PostgreSQL + sqlc | SQLite | — | | Search | pgvector (vector), BM25 | Qdrant (when >1M vectors or hybrid retrieval) | — | | Logging | slog (structured) | — | — | | Testing | Table-driven, testify | — | — | | Agents (Go) | google.golang.org/adk + pkg/litellm adapter | — | — | Exploratory: Rust, Zig — I'll tell you when I want these. ## Code conventions - **Go style**: golines, gofumpt, golangci-lint - **Errors**: `fmt.Errorf("operation: %w", err)` — never naked, never log-and-return - **Naming**: stdlib conventions, no stuttering - **Architecture**: prefer stdlib over frameworks, constructor injection, env-var config parsed into typed structs - **Git**: conventional commits (`feat:`, `fix:`, `chore:`), commit directly to main, one logical change per commit, CI is the quality gate - **Never**: long-lived feature branches, PRs for solo work, direct push without passing `task check` locally first - **Security**: no secrets in code, govulncheck before adding deps, SOPS for encrypted config - **Dependencies**: prefer stdlib. testify, slog, templ, sqlc, google.golang.org/adk (agent projects only) are pre-approved; anything else needs justification in the commit message ## Secret handling (every harness, every command) Tool output is persisted: terminal → `~/.claude/projects` transcripts → claudewatcher → brain/wiki → gitea history. A secret printed once is searchable forever, and clearing it means rotating the key. So: 1. **Never print, echo, log, or transform a secret to inspect it.** No `base64`/`xxd`/`cat` of a key, and never pipe a secret through a transform to defeat `op run`'s output masking (it masks raw values; base64 hides them from the mask — that exact trick leaked a key on 2026-06-11). 2. **Secrets stay in the subprocess.** Reference them only as env vars consumed *inside* `op run --env-file ~/.op-env -- `. Never place a literal secret in a command's argv (it lands in the tool call and the transcript). 3. **Existence check without revealing the value:** `[ -n "$X" ] && echo set` — never `${X:-...}` (returns the value when set) and never echo a substring of it. 4. **Cross-host secrets:** run the secret-consuming command on the host that has the secret; do not forward a raw key over ssh argv/stdout. 5. If a secret does leak into output, say so immediately and flag it for rotation — don't bury it. ## Infrastructure Three machines on Tailscale: | Machine | Role | Key specs | |---------|------|-----------| | koala | GPU inference, heavy compute | RTX 5070, runs k3s + llama-swap + shared postgres18/pgvector | | iguana | Services, builds | M2 Ultra Mac | | flamingo | Daily driver, edge | Mac mini, ~/dev is here | - **Model routing**: LiteLLM in front of llama-swap (local) + cloud APIs (when permitted) - **Orchestration**: k3s cluster across all three machines - **Networking**: Tailscale mesh ## Project landscape All development repos live at `~/dev/` (softlink from `~/Documents/local-dev/`). Organized in thematic folders: | Folder | Focus | Count | |--------|-------|-------| | `GO/` | Go web frameworks, API integrations, learning projects | ~10 | | `AI/` | ML research, AI frameworks (FinRL, DSPy, crawl4ai) | ~6 | | `AGENTS/` | Autonomous agents, coding agents, MCP servers, infra | ~15 | | `QKX/` | Invoice processing, financial automation, payment systems | ~13 | | `XT/` | Climate data, sustainability (Klimatkollen, Garbo) | ~2 | See `~/dev/PROJECT_SUMMARY.md` for detailed descriptions of each project. ### Key active projects - **super-koala** (`AGENTS/`) — multi-component agent stack with LangGraph, DSPy, MCP - **azure-tiger** (`QKX/`) — invoice extraction → ISO 20022 payment instructions - **gocrwl** (`AGENTS/`) — Go web crawler with containerized deployment - **koala-ai-stack** (`AGENTS/`) — local AI server infrastructure management - **klimatkollen** (`XT/`) — Swedish municipal climate data platform ## Knowledge base — actively use it A persistent brain (BM25 search + LLM-synthesised Q&A) survives across sessions, hosts, and harnesses. It holds 100+ hard-won entries: infra incident postmortems, Go pitfalls, framework gotchas, design principles, ADRs. **It is not optional reference material — query it actively, not just when explicitly told.** ### When to query (treat as a reflex) - **Before** starting a non-trivial task — search for prior art with the symptom AND the system component ("how did we solve X in Y?"). 5 seconds beats 5 hours. - **When debugging** — search for the error string, the stack frame, the affected service. Past you may have already paid this tax. - **Before adopting** a pattern, library, framework, or model name — check if it was tried and rejected, or what the integration footguns are. - **When making architectural decisions** — search for the domain + "ADR" or "decision" to find prior reasoning before re-deriving it. - **When a recommendation feels novel** — challenge yourself: "has this been documented?" The brain often has it. ### When to write After you discover something that **future-you would forget** and that **isn't recoverable from the code, git log, or PR description alone**: - Bugs whose root cause is non-obvious and generalisable beyond this project. - Framework / library / model-name quirks that bit you and would bite anyone. - Design principles validated under fire (e.g. "every `_get` needs a `_list`"). - Postmortems for incidents: what broke, why, how diagnosed, what to do next time. DON'T write project status, sprint progress, PR summaries, or "what I did this session" — those rot fast and the originals are in git/gitea anyway. Brain entries that age well are about *why*, *how to avoid*, and *what to do when*. ### How to access (per harness) | Harness | Query | Write | |---------|-------|-------| | **Claude Code, Claude Desktop** | `brain_query` (BM25), `brain_answer` (LLM-synth + sources) MCP tools | `brain_write` MCP tool | | **Crush, Pi, Antigravity, other MCP-capable** | same MCP server: `ingestion-brain` (via the `mcp__*_brain__*` namespace once authenticated) | same | | **Anything HTTP-only (curl, scripts)** | `POST https://brain-mcp.d-ma.be/query` with `{"query":"..."}` (auth via `BRAIN_MCP_TOKEN`) | `POST .../write` with `{"content":"...","filename":"..."}` | | **Browser / human inspection** | `https://git.d-ma.be/mathias/hyperguild` → `knowledge/` and `wiki/` markdown files | - **Scoping**: defaults to `public` collection; client projects filter to `{client}` + `public`. - **Routing**: brain_answer's LLM uses berget.ai as primary, iguana ollama as fallback. Both are configurable in the `supervisor/ingestion-deployment.yaml` on the koala k3s cluster; don't hardcode local-only model names into the berget URL (see knowledge entry on namespace mismatches). ### Quick reflex checks If you find yourself about to say any of these out loud, you owe yourself a brain query first: - "I think the issue might be..." - "Let me try X and see..." - "I'll just write a script to..." - "This is probably a new bug..." - "Has anyone done this before?" — *yes, probably, go check.* ## Client work rules When working on a project tagged with a client name: 1. Never send code, data, or context to cloud APIs — use local models only 2. Never reference other client projects or their data 3. Keep all artifacts within the client's git org / directory 4. Treat everything as confidential unless told otherwise ## Harness-agnostic principles This context is designed to work with any AI coding tool: - Claude Code, Cursor, Aider, Open WebUI, Charmbracelet Mods/Crush - Pi Coding Agent, Mistral Vibe, Antigravity - Any tool that accepts a system prompt or reads a markdown context file The canonical source is always `.context/AGENT.md` (root) and `.context/PROJECT.md` (per-project). Derived files are committed (see *How context propagates* below) so a `git pull` on any host yields full agent context with no setup. ## How context propagates Canonical sources of truth: - Universal: `~/dev/.context/AGENT.md` (this file) - Project: `/.context/PROJECT.md` (per-repo) Derived files (committed, regenerated by `task context:sync`): - `CLAUDE.md`, `AGENTS.md`, `.cursorrules`, `.aider.conventions.md`, `.context/system-prompt.txt` Workflow: 1. Edit a canonical file. Run `task context:sync`. Commit canonical and derived together. Push. 2. On any other host, `git pull` brings both. Claude Code (tree-walking) uses `CLAUDE.md`; Crush / Pi / Antigravity (cwd-only) use `AGENTS.md`; Cursor uses `.cursorrules`; Aider uses `.aider.conventions.md`. 3. `task check` runs `context:sync` then asserts `git status --porcelain` is empty over the derived files (catches both modified-tracked drift and missing-untracked adapters). A drift fails the check with a message telling you to stage the regenerated files. Behavior rules in this file and per-project rules in `PROJECT.md` apply unconditionally on every host, every harness. ## Engineering Skills Shared engineering skills live in the **`mathias/skills`** repo (`git.d-ma.be/mathias/skills`). Clone it to `~/dev/skills/` and run `SKILLS_CHECKOUT_DIR="$PWD" bash install.sh` there to wire every skill into your harnesses (Claude Code, Crush, Antigravity, Mistral Vibe) as native, on-demand skills. (Use `install.sh`, not `task install` — the latter is currently broken, skills#7.) Load at task start — not "on demand" but on schedule, before writing code. Browse `~/dev/skills/SKILLS_INDEX.md` for the full list. **Skill trigger table — load before starting, not after getting stuck:** | Task type | Load | |-----------|------| | Any feature or bug fix | `tdd` | | Refactor or design | `clean-code` or `solid` | | Debug | `problem-analysis` | | Review code or PRs | `code-review` | | Frame a problem before coding | `problem-analysis` | --- # cad-atlas ## Identity - **Name**: cad-atlas - **Owner**: Mathias - **Client**: personal - **Repo**: git.d-ma.be/mathias/cad-atlas - **Status**: active - **Stack**: Go + Templ + HTMX + CDN Tailwind (template-go-web). Cross-project conventions: `~/dev/.context/AGENT.md`. ## What this is A visual **atlas of the Continuous Agentic Development (CAD) workflow** — the full path from a captured signal to a deployed k3s pod, one screen, reel-style ("From Signal to Pod"). It exists to (a) make the homelab's agentic delivery pipeline legible to a human, and (b) render the CAD **audit chain** — which doubles as the regulated-industry audit artifact. > **Core thesis:** the CAD audit chain *is* the visualization data. > `TELOS → goal → spec → issue → execution → attestation → deploy` is both the trace and > the audit package. Phase C renders it once and serves two masters (observability + compliance). ## Phases - **Phase A — static hero viz** (current). Self-contained `internal/web/static/cad-atlas.html`, data-driven from a hand-authored `STAGES` array (ground-truth snapshot from brain, 2026-07-19). Served at `/` by `internal/web/handler.go` via `go:embed`. Reel-parity: SVG spine with arrowheads, animated pulse, dashed **feedback bus** (stage 08 → TELOS), replay + slow-mo. - **Phase B — generated-from-source**. Parse brain docs + `.gitea/workflows` + infra manifests → render the graph so it can't drift from config. - **Phase C — live trace viewer**. Replace the static `STAGES` array with live reads of the `assessor-loop` attestation ledger + brain `session_log` + Gitea run API + Flux events. This is the prize: a real signal→pod trace viewer that is also the audit package. ## The workflow it visualizes (9 stages) `00 Signals` (Applied AI Radar → mathias/signals) → `01 TELOS` (intention substrate) → `02 Strategic session` (claude.ai frontier + LLM Council + Autoresearch Council) → `03 Spec → Gitea issue` (agent-ready contract; Ed25519 admission #36; **var-go Oath**) → `04 Human dispatch gate` (the only checkpoint; session-dispatch bridge → cad-dispatch.yml) → `05 Execute · agentsquad` (serve/taskqueue, exec+review loop, risk LOW/MED/HIGH, dma-cli routing, assessor-loop ledger) → `06 PR → CI` (go test/vet/lint/govulncheck + **var-go/oath gate**) → `07 CD → pod` (Flux GitOps → k3s on koala) → `08 Loop back` (outcome scored vs TELOS goal). ### Three orthogonal governance gates | Gate | Guards | Where | |---|---|---| | Ed25519 admission controller (#36) | spec **integrity** (issue untampered) | stage 03 | | dispatch-allow (`.dispatch-allow` + `mathias/dispatch` allowlist) | repo **eligibility** (may agents run here) | stage 04/05 | | **var-go Oath** (`cmd/vargo-gate`, commit status `var-go/oath`) | output **correctness** (PR satisfies the Oath; floor over reviewer, anti-rubber-stamp #55) | stage 06 | ## Dogfooding This repo is built *through* the workflow it depicts. It is `dispatch-allow`-enabled, and its own build increments are governed by a **var-go Oath** embedded in their spec issues (see the Stage-03 tracking issue). Bootstrapping honesty (per swedsl honest-stub discipline): the Oath is **defined** but `cmd/vargo-gate` is **not yet wired** into this repo's CI — until it is, the Oath is advisory here. Wiring it is a first tracked task; disclosed in code, this doc, and CI config. ## Brain references (source of truth — `brain_get `) - `wiki/homelab/decisions/cad-atlas-audit-chain-is-viz-data.md` — **this project's genesis note**: the audit-chain-is-viz-data thesis, the three gates, template-go-web footguns - `wiki/homelab/decisions/inception-sprint-and-oath.md` — the Inception Sprint + Oath methodology (this repo's birth is the worked example); see [`docs/INCEPTION-OATH.md`](docs/INCEPTION-OATH.md) for the kept Oath - `knowledge/workflow-idea-to-running-service.md` — Double Diamond idea→service workflow - `wiki/homelab/decisions/continuous-agentic-development-cad-concept-2026-06-16.md` — CAD definition - `wiki/agentsquad/decisions/cad-dispatch-bridge.md` — claude.ai → agentsquad trigger path - `wiki/agentsquad/facts/llm-council-design-and-first-runs-2026-06-21.md` — LLM Council - `wiki/agentsquad/decisions/autoresearch-council-sibling-pipe.md` — Autoresearch Council - `wiki/agentsquad/decisions/serve-http-task-api.md` — agentsquad serve/taskqueue - `knowledge/var-go-anchor-to-span-spike-verdict.md` — var-go runner + CAD gate seam - `knowledge/swedsl-vargo-sprint1-enforcement-teeth-verdict.md` — vargo-gate enforcement teeth - `wiki/assessor-loop/decisions/assessor-loop-genesis.md` — attestation ledger (Phase C source) - `wiki/homelab/facts/homelab-network-topology-reference.md` — koala/iguana/flamingo/piguard ## External references - Inspiration reel — "From Inbox to Shipped" pipeline viz: https://www.instagram.com/reel/DY92L7bu27j/ - karpathy/llm-council — origin of the Council pattern - Double Diamond design process (Discover/Define/Develop/Deliver) ## Run ```bash task check # lint + vet + test (CI gate) task run # build + serve at http://localhost:8080 → the atlas ``` ## Deploy note CD (`.gitea/workflows/cd.yml`) deploys to k3s namespace `cad-atlas` via Flux. Per the known template-go-agent CD gap: the `deploy` job stays RED until `mathias/infra` has `k3s/apps/cad-atlas/deployment.yaml`. `check` + `build` are the real bootstrap gate.