Files
mathiasandClaude Opus 4.8 5048450b79
CD / Detect unsubstituted template (push) Successful in 1s
CD / Lint / Test / Vet (push) Successful in 5s
CD / Build & Import (push) Successful in 18s
CD / Deploy via GitOps (push) Has been skipped
feat(atlas): UX sprint — progressive disclosure (Plain default + transitions)
Makes the atlas self-explanatory. Every stage/node gains a plain-language layer
(plain_title + plain "what happens" + jargon-free node text); the previous
technical copy demotes to a subtitle + on-demand detail. A Plain⇄Technical
toggle (default Plain, persisted) flips the whole atlas. Biggest win: every
transition arrow is now LABELLED with "what must be true to advance" (the gated-
flow story that was invisible), gate hops (human @04, CI @06) styled distinctly.
Spine repositioned into a uniform header band so labels never collide with copy.

Copy grounded in a fresh-eyes UX review (docs/UX-REVIEW.md, reviewer≠implementer).
Data model: plain_title/plain/trans_label/trans on Stage, plain on Node — guarded
by a test (every stage has plain_title + a transition). Live overlays unchanged.

Verified: build/vet/lint(0)/test green; Plain render screenshot-checked.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-20 08:22:21 +02:00

16 KiB
Raw Permalink Blame History

CAD Atlas — Fresh-Eyes UX Review

Reviewer role: fresh-eyes UX (not implementer). This document critiques the copy and information architecture of the "From Signal to Pod" atlas and specifies the progressive-disclosure layer for the coming sprint. It does not change code.

Sources reviewed:

  • internal/atlas/atlas.json (authored stages + nodes)
  • internal/web/handler.go (which stages get live data overlaid)
  • .context/PROJECT.md (ground-truth meaning of each stage/gate)

1. Diagnosis — why the current atlas is hard for a naive viewer

The atlas is written by the person who built the pipeline, for the person who built the pipeline. Almost every node names a mechanism (Ed25519 admission controller, var-go Oath, dma-cli, assessor-loop ledger, agentsquad, ISC, TELOS) rather than the thing that happens to a piece of work. A compliance officer or a new engineer cannot answer the two questions they actually have: "what is happening to the work at this step?" and "why does it move to the next step?"

The second question is completely unanswered. The atlas renders nine stages connected by arrows, but the arrows carry zero copy. There is no statement of what has to be true for work to advance — which is exactly where the interesting governance lives (a human sign-off at 04, a green CI gate at 06, an integrity check at 03). The pipeline's whole selling point is "auditable, gated flow," yet the gates between stages are invisible. A viewer sees a row of jargon boxes and an implied left-to-right drift, with no sense of what earns each hop.

Two smaller aggravators: (a) the reel is dense — 20+ nodes, colour-coded pills whose meaning is never keyed, and Swedish-homelab proper nouns (koala, iguana, flamingo) that a stakeholder can't decode; (b) Stage 06 is authored-empty (it's generated live from CI), so on a cold/offline load that column can read as "nothing happens here," which is the opposite of the truth — CI is a governance gate.

The fix is not to dumb it down. It is to make plain-language the default layer — a one-line "what happens + why" per stage, and a labelled "what must be true to advance" per arrow — and demote today's precise, correct technical copy to an on-demand layer.


2. Per-stage plain-language map (0008)

stage plain_title plain_what (one jargon-free sentence) keep_technical (on-demand)
00 Signals Notice what's happening New ideas and developments worth reacting to are collected — mostly an automated daily/weekly scan of AI news, plus things saved by hand. "Signals → mathias/signals"; Applied AI Radar (Tier-1 daily / Tier-2 weekly), verified-primary bar; brain capture; aspirational inbox surfaces (not built).
01 TELOS Why we're here The mission, goals, and problems we're actually trying to solve live here — every piece of work downstream has to trace back to one of these goals. "TELOS — intention substrate", wiki/telos/, brain_query wing=telos.
02 Strategic session Think it through A human and AI models work out what to do and why, debating hard calls and writing down the decision and what "done" will mean. "Strategic session" — claude.ai frontier + brain MCP; ADRs/specs; ISC acceptance criteria; LLM Council (fan-out → anonymous cross-review → chairman synth); Autoresearch Council.
03 Spec → Gitea issue Write the work order The decision is turned into a precise, self-contained work order that an AI agent can execute unsupervised — with a pass/fail definition of done, a risk rating, and a tamper-proof seal. "Spec → Gitea issue" — binary ISC, risk tier LOW/MED/HIGH, reg-risk assessment, no open human deps; Ed25519 admission controller (#36); var-go Oath (single fenced block, fail-closed).
04 Human dispatch gate Human says go A person reviews the work order and its risk and decides whether to release it — this is the one and only checkpoint where work does not move on its own. "Human dispatch gate — the only checkpoint"; ratify plan + risk tier; Session-Dispatch bridge (claude.ai MCP → workflow_run_triggercad-dispatch.yml → agentsquad); dispatch-allow eligibility.
05 Execute · agentsquad Agents do the work AI agents actually build the thing — one writes, a second independent one reviews it to avoid marking its own homework — and every step is logged for the audit trail. "Execute · agentsquad" on koala; Task API (POST /tasks); executor+reviewer loop (ADK Go + LiteLLM, reviewer on distinct tier); dma-cli routing + 3-layer scope guardrail; assessor-loop attestation ledger + brain session_log.
06 PR → CI Automatic quality checks The proposed change is run through automated tests and safety checks — including a check that it actually satisfies the work order's definition of done — and only a clean pass lets it continue. "PR → CI" — Gitea Actions cd.yml; go test/vet/lint/govulncheck; var-go/oath gate (correctness floor over the reviewer, anti-rubber-stamp #55). Nodes generated live from the latest CI run.
07 CD → pod Ship it Once everything is green, the change is deployed automatically to the live server — with the rule that merging code alone doesn't ship it; the release has to be pointed at the new version. "CD → pod" — Flux GitOps → k3s on koala; "push ≠ deploy: bump tag in mathias/infra"; ntfy on deploy. Deploy + Flux state overlaid live.
08 Loop back Did it work? The result is scored against the goal that started it and fed back into the mission board, so the next round of planning learns from what shipped. "Loop back → TELOS (feedback bus)"; session_log + attestation → brain; outcome scored vs originating goal; arc partly manual (improvement target).

3. Per-node plain restatements

The existing d text stays as the technical detail layer. Each plain line below is the jargon-free default. Kept accurate to PROJECT.md.

Stage 00 — Signals

  • Applied AI Radar"An automated scan reads AI news every day (and deeper every week) and keeps only claims backed by a real paper, benchmark, code, or named lab."
  • Manual capture"Anything interesting spotted by hand gets saved into the same inbox."
  • Aspirational surfaces"Planned-but-not-built: sending ideas in by Telegram, voice, or a URL." (mark clearly as a gap / not yet real.)

Stage 01 — TELOS

  • Intention substrate"The master list of mission, goals, problems, and current status — the yardstick everything downstream is measured against."

Stage 02 — Strategic session

  • Design · ADRs · specs"A human and a top-tier AI model figure out the approach and write down the decision plus what a finished result must prove."
  • LLM Council"For hard calls, several AI models answer independently, anonymously critique each other, and a 'chair' model synthesises one verdict — reduces any single model's bias."
  • Autoresearch Council"A parallel version of the same review that vets research findings before they're allowed through."

Stage 03 — Spec → Gitea issue

  • Contract enforced"The work order must have a clear pass/fail test, a risk rating, a regulatory-risk note, and no unfinished human dependencies before it counts as agent-ready."
  • Admission controller"The work order is cryptographically signed when created, so any later tampering is detectable and the eventual change can be checked against it."
  • var-go Oath"A machine-checkable 'definition of done' is embedded in the work order — exactly one, or the order is rejected — later used to prove the result actually meets the spec."

Stage 04 — Human dispatch gate

  • Human triggers execution"A person confirms the plan and its risk level, then releases the work — nothing runs until they do."
  • Session-Dispatch bridge"The approval flips a switch that hands the signed work order over to the agents to start execution."

Stage 05 — Execute · agentsquad

  • Task API"A request kicks off a job and hands back an id you can poll for progress."
  • Executor + reviewer loop"One agent does the work; a second, independent agent on a different model reviews it — so nothing marks its own homework."
  • dma-cli · routing + scope"A router sends each agent to the right AI backend and enforces what it is and isn't allowed to touch, with a confirmation gate as a guardrail."
  • assessor-loop ledger"Every step is recorded in a tamper-evident log so the whole run can be audited afterwards."

Stage 06 — PR → CI (nodes generated live from the latest CI run — no authored nodes)

  • Live jobs render here; the plain framing for the column is: "Automated tests and safety checks run on the proposed change, including a check that it truly satisfies the work order — only a clean pass moves on."

Stage 07 — CD → pod

  • Deploy on green"When all checks pass, the release system rolls the new version onto the live server automatically — but only once the release is pointed at that version (merging code alone doesn't ship it)." (live deploy + Flux status also shown.)

Stage 08 — Loop back

  • Close the loop"The outcome is scored against the goal that started it and written back to the mission board, so future planning learns from what actually shipped."

4. Transitions — the key deliverable

For each arrow: what moves the work forward, and what must be true for it to advance. These should be rendered on the arrows themselves (see §5). Today they are blank.

  • 00 → 01 — "Does it matter to us?" A raw signal only advances if it connects to something we actually care about. Most captured signals stop here; the few that touch the mission get pulled up against a goal.

  • 01 → 02 — "Worth a session?" A goal or problem on the board becomes the seed for a design session when it's decided it's worth working on now. The goal is the input the session must trace back to.

  • 02 → 03 — "Decision reached." Once the debate converges on a decision (and what "done" will mean), it advances only when that thinking is written down as a concrete, testable specification — not while the answer is still open.

  • 03 → 04 — "Order written, sealed, agent-ready." Work advances to the gate only when the spec is a complete contract: a pass/fail test, a risk tier, a regulatory note, no open human dependencies, one embedded Oath, and a valid cryptographic signature. A malformed or unsigned order fails closed and does not reach the gate.

  • 04 → 05 — "A human said go." This is the hard stop. Nothing crosses automatically. A person must review the plan and risk and explicitly release it, and the repo must be on the allow-list, before any agent starts. This is the single human checkpoint in the whole pipeline.

  • 05 → 06 — "Agents produced a change." Work advances when the agents finish and open a proposed change (a PR) with its audit log attached. Until there's a concrete change to test, nothing moves.

  • 06 → 07 — "All checks green." The change advances only if every automated check passes — tests, linters, security scan, and the Oath check proving it meets the original work order. Any red gate stops it here; a passing reviewer is not enough to override a failed Oath.

  • 07 → 08 — "It's live." Once the new version is actually running on the server, the deployed outcome becomes the input to scoring. Advancing means "shipped and observable," not just "merged."

  • 08 → TELOS (feedback bus, dashed) — "What did we learn?" The scored outcome flows back into the mission board so goals, problems, and priorities update. This is the loop that makes the pipeline a cycle rather than a line. Note per PROJECT.md this arc is partly manual today and is an explicit improvement target — the dashed styling should read as "aspirational / not fully automated," not just decorative.


5. Progressive-disclosure recommendations

Default (Plain) layer — what everyone sees on load:

  • Each stage column shows: the plain_title as the headline, the technical title as a smaller subtitle, and the one-line plain_what directly under it.
  • Each node shows its plain one-liner as the primary text. The current d string is hidden by default.
  • Each arrow shows a short transition label (the bolded phrase from §4, e.g. "A human said go", "All checks green") — this is the single biggest comprehension win and must ship in the default layer, not behind a toggle.

On hover / expand (per node):

  • Reveal the technical d text, the tags, and the pill's meaning.
  • On the arrow, hovering expands the short label into the full "what must be true to advance" sentence from §4.

Plain ⇄ Technical toggle (global):

  • A single top-level switch, defaulting to Plain. Persist the choice (localStorage).
  • Plain: plain_title headline, plain_what, plain node lines, short arrow labels. Proper nouns (koala/iguana/agentsquad/TELOS) suppressed or shown only as a footnote.
  • Technical: today's exact copy — titles, d strings, tags, substrate host specs — i.e. the atlas as it exists now. Nothing is lost; the current view becomes "Technical."
  • The toggle should crossfade in place, not reflow the whole layout, so a viewer can flip back and forth and map plain↔technical on the same node.

Visual cues for the currently-bare transitions:

  • Give every arrow a label chip sitting on the spine. Gate arrows (04→05 human, 06→07 CI) get a distinct treatment — a lock/shield glyph and a stronger colour — because those are the governance moments the whole atlas exists to show.
  • Make the 08 → TELOS feedback bus visibly different (dashed + "partly manual" tag) so its aspirational status is honest, matching PROJECT.md's dogfooding-honesty discipline.
  • Add a small, always-visible legend keying the pill colours and the three gate types (integrity / eligibility / correctness), since colour is currently unexplained.
  • For Stage 06 (authored-empty, generated live): when no live CI data is present, show the plain_what placeholder ("Automated tests and safety checks run…") rather than an empty column, so it never reads as "nothing happens here."

Three-gate overlay (stretch, high value for the compliance audience):

  • A "show governance gates" toggle that highlights the three orthogonal gates on top of the pipeline: integrity (03, Ed25519), eligibility (04/05, dispatch-allow), correctness (06, Oath). This directly serves the "audit chain is the viz data" thesis for a compliance/exec viewer.

6. Prioritized punch list (top 8 by comprehension impact)

  1. Label every arrow with a plain "what must be true to advance" phrase (§4). Biggest miss, biggest win — turns a row of boxes into a story of gated flow. Default layer.
  2. Add a plain_what one-liner per stage as the default column copy (§2), with the technical title demoted to subtitle.
  3. Ship the Plain ⇄ Technical global toggle, defaulting to Plain, persisting choice; current copy becomes the Technical view (nothing thrown away).
  4. Rewrite node primary text to the plain lines (§3); move existing d to hover/expand.
  5. Visually distinguish the two real gates (04 human, 06 CI) with lock/shield glyphs and stronger colour so the checkpoints read as checkpoints.
  6. Add a legend keying pill colours and the three gate types — colour currently carries meaning nobody can decode.
  7. Fix the Stage-06 empty-column problem: show a plain placeholder when live CI data is absent, so the CI gate never looks like a no-op.
  8. Make the 08→TELOS feedback bus honestly aspirational (dashed + "partly manual" tag), and suppress homelab proper nouns (koala/iguana/flamingo/agentsquad/TELOS) in Plain mode, surfacing them only in Technical or a footnote.