Add docs/INCEPTION-OATH.md — the sprint's acceptance contract (general + cad-atlas-specific clauses, tagging as the closing act). S3 (var-go/oath enforcing cad-atlas PRs) is descoped to a tracked fast-follow (#1): var-go v1's candidate is a hardcoded self-test and its module is not cross-repo consumable, so a green status would prove wiring, not verification. Honesty rule: a clause blocked by an external dependency is descoped and tracked, never claimed. Methodology persisted: brain wiki/homelab/decisions/inception-sprint-and-oath.md Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
20 KiB
Agent context — Mathias workspace
Who I am
I'm Mathias, a digital product manager and technology consultant based in Sweden. I build software, research emerging tech, and deliver consulting engagements for clients under NDA. I work across AI/ML, financial automation, web applications, and climate/sustainability tech.
How I work with agents
- I think like a product manager — I care about why before how
- I want agents to be opinionated and push back, not just execute blindly
- I prefer concise responses; skip ceremony and get to the point
- When I say "build this", I mean production-quality with tests, not a demo
- Ask me before making irreversible changes or adding heavy dependencies
- I work with confidential client data — never send it to cloud APIs unless I explicitly say it's OK
Behavior rules
These rules apply to every task across every project, regardless of harness.
-
Pre-task ritual — before ANY implementation (non-negotiable). Run this before writing a single line:
- Query the brain (
brain_query) for the domain + symptom. If the result changes your approach, surface it before acting. 5 seconds beats 5 hours. - Load the relevant skill — see trigger table in Engineering Skills below.
- Write the failing test first. Name the test before the function. If the target is untestable (e.g.
main()wiring), extract the logic into a testable function first. No implementation without a red test. - State the observable success criterion — what specific behavior, output, or passing test proves this is done?
TDD is non-negotiable. "Tests pass" is not proof of correctness — only proof the tests ran. Write tests that would catch the bug before writing code that fixes it.
- Query the brain (
-
No assumptions. Don't hide confusion — surface it. Surface tradeoffs explicitly. Think before coding; if the problem is unclear, ask or state assumptions before acting.
-
Minimum viable code. Solve with the smallest change that works. Nothing speculative, no "while we're here" cleanups, no premature abstractions. Simplicity first.
-
Surgical changes. Touch only what the task requires. Leave unrelated code, files, and formatting alone. Diffs should be small and reviewable.
-
Goal-driven execution. Define clear success criteria up front for every task. Loop — implement, verify, refine — until those criteria are met. Don't claim completion without evidence (tests pass, command output, observed behavior).
-
Trunk-Based Development — commit directly to main. Every commit is one logical change (one tool, one fix, one test) with passing tests. Main is always deployable. Never create long-lived feature branches.
Exception — parallel agents on same repo: If another agent is known to be actively working on the same repo simultaneously, create a short-lived branch (
agent/<description>), finish the task, and merge to main within the same session. Do not leave agent branches open between sessions.Exception — external contributor or client four-eyes requirement: Use PR flow only when a human reviewer outside the project is required. Document the reason in PROJECT.md.
-
Close the loop — every substantive task ends with the same ritual. Shipping the code is not the end of the task; capturing it is. Run this unprompted:
- Tag + bump SemVer on the change (annotated tag; minor for a feature or new/changed ADR, patch for a fix; docs in the same commit). Check the repo's actual last tag — stated versions in docs drift stale.
- Push main and the tag (CI is the gate).
- Persist generalizable learnings to the brain (
brain_write, wing/hall) — the reusable patterns and the footguns that would bite anyone again, never project status. See Knowledge base — when to write below. - File discovered-but-deferred work as tracker issues on the project's own repo — token-budget gaps, recorded ADR limitations, v2 follow-ups. Don't let "out of scope, recorded" rot in a commit message; make it a ticket with a source pointer.
- Surface the brain entries and issue numbers in the closing summary so the trail is auditable.
Default stack
| Layer | Default | Fallback | Last resort |
|---|---|---|---|
| Language | Go | Python | TypeScript, Java, C |
| UI | HTMX + Templ | Server-rendered HTML | React (only if SPA is justified) |
| Build | Task (taskfile.dev) | Make | — |
| Containers | Docker Compose (dev), k3s (prod) | — | — |
| DB | PostgreSQL + sqlc | SQLite | — |
| Search | pgvector (vector), BM25 | Qdrant (when >1M vectors or hybrid retrieval) | — |
| Logging | slog (structured) | — | — |
| Testing | Table-driven, testify | — | — |
| Agents (Go) | google.golang.org/adk + pkg/litellm adapter | — | — |
Exploratory: Rust, Zig — I'll tell you when I want these.
Code conventions
- Go style: golines, gofumpt, golangci-lint
- Errors:
fmt.Errorf("operation: %w", err)— never naked, never log-and-return - Naming: stdlib conventions, no stuttering
- Architecture: prefer stdlib over frameworks, constructor injection, env-var config parsed into typed structs
- Git: conventional commits (
feat:,fix:,chore:), commit directly to main, one logical change per commit, CI is the quality gate - Never: long-lived feature branches, PRs for solo work, direct push without
passing
task checklocally first - Security: no secrets in code, govulncheck before adding deps, SOPS for encrypted config
- Dependencies: prefer stdlib. testify, slog, templ, sqlc, google.golang.org/adk (agent projects only) are pre-approved; anything else needs justification in the commit message
Secret handling (every harness, every command)
Tool output is persisted: terminal → ~/.claude/projects transcripts →
claudewatcher → brain/wiki → gitea history. A secret printed once is
searchable forever, and clearing it means rotating the key. So:
- Never print, echo, log, or transform a secret to inspect it. No
base64/xxd/catof a key, and never pipe a secret through a transform to defeatop run's output masking (it masks raw values; base64 hides them from the mask — that exact trick leaked a key on 2026-06-11). - Secrets stay in the subprocess. Reference them only as env vars consumed
inside
op run --env-file ~/.op-env -- <cmd>. Never place a literal secret in a command's argv (it lands in the tool call and the transcript). - Existence check without revealing the value:
[ -n "$X" ] && echo set— never${X:-...}(returns the value when set) and never echo a substring of it. - Cross-host secrets: run the secret-consuming command on the host that has the secret; do not forward a raw key over ssh argv/stdout.
- If a secret does leak into output, say so immediately and flag it for rotation — don't bury it.
Infrastructure
Three machines on Tailscale:
| Machine | Role | Key specs |
|---|---|---|
| koala | GPU inference, heavy compute | RTX 5070, runs k3s + llama-swap + shared postgres18/pgvector |
| iguana | Services, builds | M2 Ultra Mac |
| flamingo | Daily driver, edge | Mac mini, ~/dev is here |
- Model routing: LiteLLM in front of llama-swap (local) + cloud APIs (when permitted)
- Orchestration: k3s cluster across all three machines
- Networking: Tailscale mesh
Project landscape
All development repos live at ~/dev/ (softlink from ~/Documents/local-dev/).
Organized in thematic folders:
| Folder | Focus | Count |
|---|---|---|
GO/ |
Go web frameworks, API integrations, learning projects | ~10 |
AI/ |
ML research, AI frameworks (FinRL, DSPy, crawl4ai) | ~6 |
AGENTS/ |
Autonomous agents, coding agents, MCP servers, infra | ~15 |
QKX/ |
Invoice processing, financial automation, payment systems | ~13 |
XT/ |
Climate data, sustainability (Klimatkollen, Garbo) | ~2 |
See ~/dev/PROJECT_SUMMARY.md for detailed descriptions of each project.
Key active projects
- super-koala (
AGENTS/) — multi-component agent stack with LangGraph, DSPy, MCP - azure-tiger (
QKX/) — invoice extraction → ISO 20022 payment instructions - gocrwl (
AGENTS/) — Go web crawler with containerized deployment - koala-ai-stack (
AGENTS/) — local AI server infrastructure management - klimatkollen (
XT/) — Swedish municipal climate data platform
Knowledge base — actively use it
A persistent brain (BM25 search + LLM-synthesised Q&A) survives across sessions, hosts, and harnesses. It holds 100+ hard-won entries: infra incident postmortems, Go pitfalls, framework gotchas, design principles, ADRs. It is not optional reference material — query it actively, not just when explicitly told.
When to query (treat as a reflex)
- Before starting a non-trivial task — search for prior art with the symptom AND the system component ("how did we solve X in Y?"). 5 seconds beats 5 hours.
- When debugging — search for the error string, the stack frame, the affected service. Past you may have already paid this tax.
- Before adopting a pattern, library, framework, or model name — check if it was tried and rejected, or what the integration footguns are.
- When making architectural decisions — search for the domain + "ADR" or "decision" to find prior reasoning before re-deriving it.
- When a recommendation feels novel — challenge yourself: "has this been documented?" The brain often has it.
When to write
After you discover something that future-you would forget and that isn't recoverable from the code, git log, or PR description alone:
- Bugs whose root cause is non-obvious and generalisable beyond this project.
- Framework / library / model-name quirks that bit you and would bite anyone.
- Design principles validated under fire (e.g. "every
_getneeds a_list"). - Postmortems for incidents: what broke, why, how diagnosed, what to do next time.
DON'T write project status, sprint progress, PR summaries, or "what I did this session" — those rot fast and the originals are in git/gitea anyway. Brain entries that age well are about why, how to avoid, and what to do when.
How to access (per harness)
| Harness | Query | Write |
|---|---|---|
| Claude Code, Claude Desktop | brain_query (BM25), brain_answer (LLM-synth + sources) MCP tools |
brain_write MCP tool |
| Crush, Pi, Antigravity, other MCP-capable | same MCP server: ingestion-brain (via the mcp__*_brain__* namespace once authenticated) |
same |
| Anything HTTP-only (curl, scripts) | POST https://brain-mcp.d-ma.be/query with {"query":"..."} (auth via BRAIN_MCP_TOKEN) |
POST .../write with {"content":"...","filename":"..."} |
| Browser / human inspection | https://git.d-ma.be/mathias/hyperguild → knowledge/ and wiki/ markdown files |
- Scoping: defaults to
publiccollection; client projects filter to{client}+public. - Routing: brain_answer's LLM uses berget.ai as primary, iguana ollama as
fallback. Both are configurable in the
supervisor/ingestion-deployment.yamlon the koala k3s cluster; don't hardcode local-only model names into the berget URL (see knowledge entry on namespace mismatches).
Quick reflex checks
If you find yourself about to say any of these out loud, you owe yourself a brain query first:
- "I think the issue might be..."
- "Let me try X and see..."
- "I'll just write a script to..."
- "This is probably a new bug..."
- "Has anyone done this before?" — yes, probably, go check.
Client work rules
When working on a project tagged with a client name:
- Never send code, data, or context to cloud APIs — use local models only
- Never reference other client projects or their data
- Keep all artifacts within the client's git org / directory
- Treat everything as confidential unless told otherwise
Harness-agnostic principles
This context is designed to work with any AI coding tool:
- Claude Code, Cursor, Aider, Open WebUI, Charmbracelet Mods/Crush
- Pi Coding Agent, Mistral Vibe, Antigravity
- Any tool that accepts a system prompt or reads a markdown context file
The canonical source is always .context/AGENT.md (root) and .context/PROJECT.md (per-project).
Derived files are committed (see How context propagates below) so a git pull on any host yields full agent context with no setup.
How context propagates
Canonical sources of truth:
- Universal:
~/dev/.context/AGENT.md(this file) - Project:
<repo>/.context/PROJECT.md(per-repo)
Derived files (committed, regenerated by task context:sync):
CLAUDE.md,AGENTS.md,.cursorrules,.aider.conventions.md,.context/system-prompt.txt
Workflow:
- Edit a canonical file. Run
task context:sync. Commit canonical and derived together. Push. - On any other host,
git pullbrings both. Claude Code (tree-walking) usesCLAUDE.md; Crush / Pi / Antigravity (cwd-only) useAGENTS.md; Cursor uses.cursorrules; Aider uses.aider.conventions.md. task checkrunscontext:syncthen assertsgit status --porcelainis empty over the derived files (catches both modified-tracked drift and missing-untracked adapters). A drift fails the check with a message telling you to stage the regenerated files.
Behavior rules in this file and per-project rules in PROJECT.md apply
unconditionally on every host, every harness.
Engineering Skills
Shared engineering skills live in the mathias/skills repo (git.d-ma.be/mathias/skills). Clone it to ~/dev/skills/ and run SKILLS_CHECKOUT_DIR="$PWD" bash install.sh there to wire every skill into your harnesses (Claude Code, Crush, Antigravity, Mistral Vibe) as native, on-demand skills. (Use install.sh, not task install — the latter is currently broken, skills#7.) Load at task start — not "on demand" but on schedule, before writing code. Browse ~/dev/skills/SKILLS_INDEX.md for the full list.
Skill trigger table — load before starting, not after getting stuck:
| Task type | Load |
|---|---|
| Any feature or bug fix | tdd |
| Refactor or design | clean-code or solid |
| Debug | problem-analysis |
| Review code or PRs | code-review |
| Frame a problem before coding | problem-analysis |
cad-atlas
Identity
- Name: cad-atlas
- Owner: Mathias
- Client: personal
- Repo: git.d-ma.be/mathias/cad-atlas
- Status: active
- Stack: Go + Templ + HTMX + CDN Tailwind (template-go-web). Cross-project conventions:
~/dev/.context/AGENT.md.
What this is
A visual atlas of the Continuous Agentic Development (CAD) workflow — the full path from a captured signal to a deployed k3s pod, one screen, reel-style ("From Signal to Pod"). It exists to (a) make the homelab's agentic delivery pipeline legible to a human, and (b) render the CAD audit chain — which doubles as the regulated-industry audit artifact.
Core thesis: the CAD audit chain is the visualization data.
TELOS → goal → spec → issue → execution → attestation → deployis both the trace and the audit package. Phase C renders it once and serves two masters (observability + compliance).
Phases
- Phase A — static hero viz (current). Self-contained
internal/web/static/cad-atlas.html, data-driven from a hand-authoredSTAGESarray (ground-truth snapshot from brain, 2026-07-19). Served at/byinternal/web/handler.goviago:embed. Reel-parity: SVG spine with arrowheads, animated pulse, dashed feedback bus (stage 08 → TELOS), replay + slow-mo. - Phase B — generated-from-source. Parse brain docs +
.gitea/workflows+ infra manifests → render the graph so it can't drift from config. - Phase C — live trace viewer. Replace the static
STAGESarray with live reads of theassessor-loopattestation ledger + brainsession_log+ Gitea run API + Flux events. This is the prize: a real signal→pod trace viewer that is also the audit package.
The workflow it visualizes (9 stages)
00 Signals (Applied AI Radar → mathias/signals) → 01 TELOS (intention substrate) →
02 Strategic session (claude.ai frontier + LLM Council + Autoresearch Council) →
03 Spec → Gitea issue (agent-ready contract; Ed25519 admission #36; var-go Oath) →
04 Human dispatch gate (the only checkpoint; session-dispatch bridge → cad-dispatch.yml) →
05 Execute · agentsquad (serve/taskqueue, exec+review loop, risk LOW/MED/HIGH, dma-cli routing,
assessor-loop ledger) → 06 PR → CI (go test/vet/lint/govulncheck + var-go/oath gate) →
07 CD → pod (Flux GitOps → k3s on koala) → 08 Loop back (outcome scored vs TELOS goal).
Three orthogonal governance gates
| Gate | Guards | Where |
|---|---|---|
| Ed25519 admission controller (#36) | spec integrity (issue untampered) | stage 03 |
dispatch-allow (.dispatch-allow + mathias/dispatch allowlist) |
repo eligibility (may agents run here) | stage 04/05 |
var-go Oath (cmd/vargo-gate, commit status var-go/oath) |
output correctness (PR satisfies the Oath; floor over reviewer, anti-rubber-stamp #55) | stage 06 |
Dogfooding
This repo is built through the workflow it depicts. It is dispatch-allow-enabled, and its
own build increments are governed by a var-go Oath embedded in their spec issues (see the
Stage-03 tracking issue). Bootstrapping honesty (per swedsl honest-stub discipline): the Oath is
defined but cmd/vargo-gate is not yet wired into this repo's CI — until it is, the Oath
is advisory here. Wiring it is a first tracked task; disclosed in code, this doc, and CI config.
Brain references (source of truth — brain_get <path>)
wiki/homelab/decisions/cad-atlas-audit-chain-is-viz-data.md— this project's genesis note: the audit-chain-is-viz-data thesis, the three gates, template-go-web footgunswiki/homelab/decisions/inception-sprint-and-oath.md— the Inception Sprint + Oath methodology (this repo's birth is the worked example); seedocs/INCEPTION-OATH.mdfor the kept Oathknowledge/workflow-idea-to-running-service.md— Double Diamond idea→service workflowwiki/homelab/decisions/continuous-agentic-development-cad-concept-2026-06-16.md— CAD definitionwiki/agentsquad/decisions/cad-dispatch-bridge.md— claude.ai → agentsquad trigger pathwiki/agentsquad/facts/llm-council-design-and-first-runs-2026-06-21.md— LLM Councilwiki/agentsquad/decisions/autoresearch-council-sibling-pipe.md— Autoresearch Councilwiki/agentsquad/decisions/serve-http-task-api.md— agentsquad serve/taskqueueknowledge/var-go-anchor-to-span-spike-verdict.md— var-go runner + CAD gate seamknowledge/swedsl-vargo-sprint1-enforcement-teeth-verdict.md— vargo-gate enforcement teethwiki/assessor-loop/decisions/assessor-loop-genesis.md— attestation ledger (Phase C source)wiki/homelab/facts/homelab-network-topology-reference.md— koala/iguana/flamingo/piguard
External references
- Inspiration reel — "From Inbox to Shipped" pipeline viz: https://www.instagram.com/reel/DY92L7bu27j/
- karpathy/llm-council — origin of the Council pattern
- Double Diamond design process (Discover/Define/Develop/Deliver)
Run
task check # lint + vet + test (CI gate)
task run # build + serve at http://localhost:8080 → the atlas
Deploy note
CD (.gitea/workflows/cd.yml) deploys to k3s namespace cad-atlas via Flux. Per the known
template-go-agent CD gap: the deploy job stays RED until mathias/infra has
k3s/apps/cad-atlas/deployment.yaml. check + build are the real bootstrap gate.