Author SHA1 Message Date
mathiasandClaude Opus 4.8 b7d8cfc3cc feat: add telos-load and regulatory-risk-assessment skills
release / tag (push) Successful in 1s
- telos-load: harness-agnostic TELOS session context loading (closes #4)
- regulatory-risk-assessment: CAD compliance gate + risk register (closes #3)

Bump-Type: minor

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-18 18:17:44 +02:00
mathias e553dae354 feat(skills): add dream skill (#2)
release / tag (push) Failing after 2s
2026-05-29 06:50:09 +00:00
5 changed files with 353 additions and 192 deletions
+4
View File
@@ -23,6 +23,8 @@ This index lists all available engineering skills. Load the full SKILL.md on dem
| `session-retrospective` | Surface non-obvious learnings from a session log | After a coding session ends, before context is lost | | `session-retrospective` | Surface non-obvious learnings from a session log | After a coding session ends, before context is lost |
| `trainer` | Two-phase brain curation (score candidates, write past quality gate) | Periodic brain pruning and curation | | `trainer` | Two-phase brain curation (score candidates, write past quality gate) | Periodic brain pruning and curation |
| `grill-me` | Structured plan interrogation — Quick Poke, Full Grill, Pre-mortem | Stress-testing a plan before committing; end of Diamond 1; before promoting to pre-prod | | `grill-me` | Structured plan interrogation — Quick Poke, Full Grill, Pre-mortem | Stress-testing a plan before committing; end of Diamond 1; before promoting to pre-prod |
| `telos-load` | Load TELOS intention substrate at session start | Starting any koala session; before architectural decisions; CAD pipeline entry |
| `regulatory-risk-assessment` | Structured risk register for regulated-industry features | Filing a CAD issue (needs Risk: level); features touching payments, auth, external APIs, user data |
## Wiring into tools ## Wiring into tools
@@ -89,3 +91,5 @@ grill-me ──→ planning (after plan is sharpened)
| "Grill me / stress-test this / is this ready?" | `grill-me` | | "Grill me / stress-test this / is this ready?" | `grill-me` |
| "End of Diamond 1 — should we build this?" | `grill-me` (Full Grill) | | "End of Diamond 1 — should we build this?" | `grill-me` (Full Grill) |
| "About to promote to pre-prod" | `grill-me` (Pre-mortem) | | "About to promote to pre-prod" | `grill-me` (Pre-mortem) |
| "What are the risks?" / "compliance gate" / "risk register" | `regulatory-risk-assessment` |
| "Start of session" / "load TELOS" / "what are we optimizing toward?" | `telos-load` |
+168
View File
@@ -0,0 +1,168 @@
---
name: dream
description: >
Run a "dream" — a reflective memory consolidation pass over an agent's memory
directory. Use this skill whenever the user says "dream", "run a dream", "consolidate
my memory files", "clean up my MEMORY.md", or asks Claude to do a memory maintenance
pass, prune stale notes, or reorganize topic files. Also trigger when the user wants
to rebuild a memory index, merge duplicate facts, or convert relative dates in notes
to absolute ones. This is an agentic, multi-phase workflow — always use this skill
rather than improvising the steps.
---
# Dream — Memory Consolidation Skill
You are performing a **dream**: a reflective, agentic pass over a memory directory.
Your goal is to synthesize recent signal into durable, well-organized memory so that
future sessions can orient quickly.
---
## Pre-flight
Before starting, confirm:
1. **Where is the memory directory?** Ask the user if not obvious from context.
Common locations: `~/memory/`, `~/.agent/memory/`, `./memory/`, a path in an env var like `$MEMORY_DIR`.
2. **Are there transcripts or daily logs to scan?** Ask if not obvious.
3. **Any topics to skip or treat as sensitive?**
Once confirmed, proceed through the four phases in order. Narrate each phase briefly as you go.
---
## Phase 1 — Orient
**Goal**: Get a map of what exists before touching anything.
```bash
ls -la <memory_dir>/
cat <memory_dir>/MEMORY.md
```
For each file listed (excluding MEMORY.md):
- Read or skim it (first 4060 lines is usually enough unless it's small).
- Note: topic, approximate recency, any obvious staleness or duplication.
Build a mental inventory:
- Files present, rough line counts
- Topics covered
- Any files that look abandoned, mislabeled, or overlapping
---
## Phase 2 — Gather Recent Signal
**Goal**: Find new facts, corrections, and drift since the last dream.
Check in this order:
1. **Daily logs** — read recent entries (last 714 days).
Look for: new decisions, changed preferences, completed projects, new relationships/tools.
2. **Drifted facts** — scan existing topic files for statements that may now be false.
Examples: "currently evaluating X" (did they pick one?), "planning to do Y" (done or dropped?), relative dates like "last week" or "recently".
3. **Transcripts** — only grep narrowly if there's a specific gap.
Avoid bulk-reading transcripts; it's slow and noisy. Use targeted patterns:
```bash
grep -r "decided\|switched to\|no longer\|now using\|moved to" <transcripts_dir>/ | tail -40
```
Collect a list of **updates to make**: new facts, corrections, removals.
---
## Phase 3 — Consolidate
**Goal**: Apply the updates. Leave memory files cleaner and more accurate than you found them.
For each topic file:
- **Merge duplicates**: if the same fact appears in two files, keep it in the more specific one and remove from the general one.
- **Convert relative dates**: replace "last week", "recently", "a few months ago" with an absolute date (use the current date as reference; estimate if necessary and note the uncertainty).
- **Delete contradicted facts**: if a new fact supersedes an old one, remove the old one outright — don't leave both.
- **Tighten language**: convert vague hedges ("probably uses", "might be") to definite statements where the evidence supports it, or remove if genuinely unknown.
- **Add new facts** from Phase 2 to the appropriate topic file. Create a new topic file if no good home exists.
After editing files, do a final pass:
```bash
grep -n "last week\|recently\|a few months\|soon\|currently planning" <memory_dir>/*.md
```
Clean up any remaining relative time references.
---
## Phase 4 — Prune and Index
**Goal**: Rebuild MEMORY.md as a clean, navigable index under 200 lines.
**MEMORY.md structure**:
```markdown
# Memory Index
_Last updated: YYYY-MM-DD_
## Overview
One short paragraph: who this agent is, primary context, most important standing facts.
## Topic Files
| File | Contents | Last updated |
|------|----------|-------------|
| person.md | Identity, preferences, background | YYYY-MM-DD |
| projects.md | Active and recent projects | YYYY-MM-DD |
| tools.md | Stack, infra, dev environment | YYYY-MM-DD |
| ... | ... | ... |
## Quick Facts
- Bullet list of the 1015 most frequently-needed facts (role, location, key tools, etc.)
## Recent Changes
- Bullet list of what changed in this dream (so the next session knows what's fresh)
```
Rules:
- Remove any pointers to files that no longer exist.
- Add pointers for any new files created in Phase 3.
- Keep Quick Facts ≤ 15 bullets — this is for speed, not completeness.
- Recent Changes section replaces itself each dream (don't accumulate).
- Total MEMORY.md length: **200 lines max**.
---
## Output
After completing all four phases, return a **dream summary** to the user:
```
## Dream complete — YYYY-MM-DD
### What changed
- [file]: [what was updated]
- MEMORY.md: rebuilt index, N topic files indexed
### Facts added
- ...
### Facts removed / corrected
- ...
### Files created
- ...
### Files deleted or merged
- ...
### Still uncertain / needs follow-up
- ...
```
Keep it concise — a few bullets per section, not exhaustive diffs. The goal is for the user to quickly confirm the dream went well and catch any mistakes.
---
## Tips and Edge Cases
- **No MEMORY.md exists yet**: create one from scratch using the structure above. Treat all existing files as "first time indexed".
- **Memory dir is empty**: create MEMORY.md and a starter `scratch.md` noting the date and that the memory system is new.
- **Conflicting facts with no clear resolution**: note both in the file with a date stamp and flag in "Still uncertain" section of the summary.
- **Large transcript dumps**: resist reading them in full. Grep is your friend. If you must read, read the last N lines only.
- **Files with sensitive content**: if the user flagged topics to skip, skip them entirely — don't even open them.
-192
View File
@@ -1,192 +0,0 @@
---
name: experiment-spec
description: Write a rigorous experiment spec for a research phase before any code or training runs. Use instead of feature-spec for scientific/ML research projects. Enforces falsifiable hypothesis, quantitative acceptance criteria, baseline comparison, and null-result protocol.
---
# Experiment Spec
## Overview
An experiment spec is the scientific contract for one research phase or experiment, written before any implementation or training begins. It is the research analogue of `feature-spec` — same discipline, different vocabulary.
**Core principle:** If you cannot write a falsifiable hypothesis with a quantitative acceptance criterion, you do not understand the experiment well enough to run it.
## When to Use
- Starting a new research phase (Phase 0, Phase 1, etc.)
- Running any experiment that will produce metrics used to make a go/no-go decision
- Any time the question is "does X work?" rather than "build X"
- Before touching training code, data, or hyperparameters for a new question
**When NOT to use:**
- Implementing a specific component whose behaviour is already defined by a phase spec (use `feature-spec` instead)
- Exploratory data analysis with no hypothesis (use a notebook; note it as EDA)
- Bug fixes or refactors
## Iron Laws
1. **The hypothesis must be falsifiable.** "JEPA is promising" is not a hypothesis. "JEPA embeddings will achieve silhouette > 0.35 on held-out data" is. If you cannot state conditions under which the hypothesis is false, it is not a hypothesis.
2. **Acceptance criteria must be quantitative and pre-registered.** Write the number before you run the experiment. Moving the goalposts after seeing results is p-hacking.
3. **A baseline is mandatory.** Every experiment must compare against at least one simpler baseline. "Better than nothing" is not a baseline.
4. **A null-result protocol is mandatory.** State what you will conclude and do if the hypothesis is rejected. "Try harder" is not a protocol.
5. **The training cutoff is sacred.** No post-cutoff data informs any decision in the spec or implementation.
## Spec Template
```markdown
# Experiment Spec: [Phase N — Short Name]
## Hypothesis
> "[Falsifiable claim]: We believe [X] will produce [Y], measurable by [Z]."
State conditions under which this hypothesis is FALSE.
## Background
Why this experiment? What does it build on? What prior result or decision motivates it?
(24 sentences. Reference DECISIONS.md or brain wing entries where relevant.)
## Design
### Data
- Source, date range, pairs/assets, features used
- Train / validation / test split (respect training cutoff)
### Model / method
- Architecture, configuration, key hyperparameters
- What is being varied vs. held fixed
### Baseline
- What simpler method is being compared against?
- Why is this the right baseline?
### Ablations (if any)
- What variants will be run to isolate the effect being studied?
## Acceptance Criteria
- [ ] [Primary criterion — quantitative threshold on primary metric]
- [ ] [Baseline comparison — e.g. "exceeds baseline by >X%"]
- [ ] [Reproducibility — reruns within ±Y% of reported metric]
- [ ] [Collapse/sanity check — e.g. "PC1/rolling-HV correlation < 0.85"]
## Out of Scope
What this experiment explicitly does NOT answer, even if related.
Anything plausibly in scope that is deferred goes here.
## Null Result Protocol
If the primary acceptance criterion is NOT met:
- What do we conclude?
- What is the next step? (Investigate X, pivot to Y, terminate programme)
- What gets written to the brain and results/summaries/?
## Risks
What could go wrong, and how would it be detected?
At least one risk must be listed.
```
## Worked Example
```markdown
# Experiment Spec: Phase 0 — SSL Feasibility Gate
## Hypothesis
> We believe that a masked autoencoder (MAE) trained on FX hourly data will produce
> latent embeddings that show structural separability by volatility regime without
> explicit regime labels, measurable by silhouette score > 0.20 on held-out 2023 data.
This hypothesis is FALSE if silhouette score ≤ 0.20 on the held-out evaluation.
## Background
Before investing in JEPA-specific machinery, we need to confirm that SSL-based
representation learning can find regime structure in FX time-series at all. MAE is
the simplest SSL baseline — if it cannot find structure, JEPA will not either.
Added post Full Grill (2026-05-27). See DECISIONS.md: "Phase 0: SSL feasibility gate".
## Design
### Data
- Source: DUKASCopy, EUR/USD hourly, 20082022 (train), 2023 (held-out test)
- Features: log-return, rolling 20-period HV, VIX (daily interpolated to hourly)
- Regime label (for evaluation only, not training): rolling 30-day HV percentile,
binary high/low threshold at 50th percentile
### Model
- Masked Autoencoder: 1D temporal masking (mask contiguous 24h window)
- Encoder: 3-layer 1D CNN + positional encoding
- Decoder: 2-layer MLP reconstructing masked segment
- Context window: 120 hours (5 days)
### Baseline
- PCA on raw feature vectors (same window) — tests whether any dimensionality
reduction shows regime structure, not just SSL
### Ablations
- Masking horizon: K ∈ {8h, 24h, 72h} — does horizon affect embedding quality?
## Acceptance Criteria
- [ ] Silhouette score > 0.20 on held-out 2023 data (k-means, k=3, vs. HV regime label)
- [ ] MAE silhouette exceeds PCA baseline silhouette
- [ ] Rerun within ±10% of reported silhouette
- [ ] PC1 / rolling-HV correlation < 0.95 (not purely encoding volatility level)
## Out of Scope
- JEPA implementation (Phase 1)
- Multi-pair training (Phase 1+)
- VaR or ES computation
- Any use of post-2023 data
## Null Result Protocol
If silhouette ≤ 0.20:
- Conclude: SSL cannot reliably find regime structure in EUR/USD hourly data with
these features at this resolution
- Next step: investigate whether (a) hourly resolution is too noisy (try daily),
(b) 3 features are insufficient, or (c) regime label definition is too coarse
- Record result in results/summaries/phase-0-null.md and brain wing jepa-fx/failures/
## Risks
- Encoder collapses to near-constant output: detect via reconstruction loss plateau
in first 10 epochs; mitigation: add batch norm, reduce learning rate
- Regime label too coarse (binary HV): silhouette may be low even with good structure;
mitigation: also evaluate with 4-class label (HV quartiles)
```
## Common Failure Modes
| Failure mode | What it looks like | Fix |
|---|---|---|
| Non-falsifiable hypothesis | "JEPA shows promise" | Rewrite with a number |
| Post-hoc criteria | Threshold chosen after seeing results | Write the number first, commit the spec |
| No baseline | Silhouette of 0.30 sounds good until PCA achieves 0.35 | Always include a dumber method |
| Missing null protocol | "We'll figure it out if it fails" | Write it now — it forces clarity about what you're actually betting on |
| Cutoff violation | Architecture choice informed by 2024 data patterns | Never open the test set during development |
## Brain MCP Integration
**At spec start:**
- `brain_query wing=jepa-fx hall=decisions` — load current architectural decisions
- `brain_query wing=jepa-fx hall=failures` — load known failure modes; address them in Risks section
**After spec is approved:**
- `brain_write` to `jepa-fx/hypotheses/` with the hypothesis and acceptance criteria
**After experiment concludes:**
- `brain_write` to `jepa-fx/failures/` with any new failure modes discovered
- `session_log` with outcome
## Cross-References
- Use `feature-spec` for implementing a specific component within an already-specced phase
- Use `grill-me` on the spec before running the experiment if the hypothesis feels shaky
- Use `tdd` once the spec is approved — each acceptance criterion maps to a test
- Use `session-retrospective` after the experiment concludes
+105
View File
@@ -0,0 +1,105 @@
# regulatory-risk-assessment
Produce a structured regulatory risk assessment for a feature, component,
or integration — generating a risk register entry that satisfies the
audit requirements of regulated-industry clients (banking, finance,
insurance, PSD2/PSR, DORA, AML/KYC contexts).
**Use when:**
- Filing a Gitea issue dispatched via CAD (needs Risk: LOW/MEDIUM/HIGH)
- Designing a feature touching payments, auth, data storage, external APIs
- Preparing a client deliverable in a regulated industry
- "what are the risks?" / "compliance gate" / "risk register" in session
**Do not use for:** routine refactoring, docs-only changes, internal
tooling with no external data or user impact.
## What this skill produces
A docs/risk-register.md section with this schema:
### R-[DOMAIN]-[NN] — [Short risk title]
| Field | Value |
|-------|-------|
| Risk | What could go wrong (concrete, specific) |
| Regulatory reference | Which obligation/regulation, if any |
| Likelihood | H / M / L |
| Impact | H / M / L |
| Overall | H / M / L (highest of likelihood x impact) |
| Mitigation | What we are doing about it |
| Validation | Test name or Gitea issue number |
| Status | open / mitigated / accepted |
Risk ID namespace:
R-AUTH-NN authentication and authorization
R-DATA-NN data storage, retention, privacy
R-API-NN external API integration
R-PAY-NN payment and financial transactions
R-INFRA-NN infrastructure and availability
R-AGENT-NN agentic / AI execution
R-COMP-NN compliance and regulatory obligation
## Mechanics
Step 1 — Scope the assessment
1. What external systems does this touch?
2. What user data does it read, write, or transmit?
3. What happens on silent failure? Loud failure?
4. What is the blast radius of a worst-case bug?
5. Is a regulation implicated?
Step 2 — Enumerate risks (common examples)
Auth/OAuth: token theft, refresh failure, insufficient scope
Email/Gmail: misclassification archives HUMAN thread, PII in logs
Payment/PAIN.001: wrong amount, duplicate submission, missing field
Agentic: irreversible action without approval, prompt injection, spirals
Infra: single point of failure, secret in logs
Step 3 — Declare overall risk level
Highest individual risk = feature overall level (LOW/MEDIUM/HIGH).
This is the **Risk:** declaration in the CAD agent-ready issue contract.
Step 4 — Write validations
Every mitigation needs a validation:
- Specific test name (TestXxx)
- Gitea issue number
- Manual verification step with acceptance criteria
Not acceptable: "will test later", "review manually"
Step 5 — Update docs/risk-register.md
Append entries. Create if absent:
# Risk Register
All entries follow R-[DOMAIN]-[NN] schema.
See skills/regulatory-risk-assessment/SKILL.md for conventions.
Last updated: [date]
## Integration with assessor-loop
For complex regulated-industry features (PSD2/PSR, DORA, AML), route to
mathias/assessor-loop for deep obligation decomposition first.
Use this skill standalone for internal tooling or general engineering risk.
## Integration with CAD
Every CAD-dispatched issue must include:
**Risk:** LOW | MEDIUM | HIGH
CAD pre-flight (agentsquad#29) rejects issues without it.
If genuinely no risks: declare LOW and note why.
## Example entry
### R-DATA-01 — Email misclassification archives a HUMAN thread
| Field | Value |
|-------|-------|
| Risk | LLM labels a real person's email as NOISE, causing auto-archive |
| Regulatory reference | None (internal) |
| Likelihood | M |
| Impact | H |
| Overall | H |
| Mitigation | Phase 1 read-only; HUMAN class never auto-archived in any phase |
| Validation | TestClassifier_HumanThreadNeverArchived; 1-week Phase 1 review |
| Status | open |
+76
View File
@@ -0,0 +1,76 @@
# telos-load
Load TELOS — the intention substrate — into a session so every decision
can be traced back to a goal, and every goal to a problem.
**Use when:** starting any session on koala (Claude Code, Crush,
Antigravity), or whenever a session needs to know what we are optimizing
toward. If you find yourself making architectural decisions without knowing
the active goals, load TELOS first.
**Do not skip:** a session without TELOS context is flying blind. It may
produce technically correct output that is strategically wrong.
## What TELOS is
TELOS is a wing in brain (wiki/telos/decisions/) containing 9 files:
PROBLEMS, MISSION, GOALS, CHALLENGES, PROJECTS, STRATEGIES, BELIEFS,
WRONG, STATUS. Together they answer: what are we working against, what
are we trying to build, and how do we approach the work?
Full aggregate: wiki/telos/decisions/principal-telos.md
## Loading TELOS by harness
### Claude Code (per-project CLAUDE.md)
Add to the project CLAUDE.md or ~/.claude/CLAUDE.md:
## Intention context (TELOS)
At session start, call brain_context with wing=telos and limit=8.
Fallback if brain MCP unavailable:
wiki/telos/decisions/principal-telos.md in the brain repo.
The global ~/.claude/CLAUDE.md was wired on koala on 2026-06-16.
### Crush
Location on koala: ~/.config/crush/CRUSH.md (auto-loaded via
global_context_paths). Add:
## Intention context (TELOS)
At session start: query brain MCP with wing=telos, limit=8.
If brain MCP unavailable: read wiki/telos/decisions/principal-telos.md
### Antigravity
Add @brain_context wing=telos directive at top of system instructions.
### Fallback — no brain MCP
cat ~/dev/AI/brain/wiki/telos/decisions/principal-telos.md
Or @-import in CLAUDE.md:
@~/dev/AI/brain/wiki/telos/decisions/principal-telos.md
## Verification
brain_query wing=telos limit=3
→ Expected: PROBLEMS, MISSION, GOALS returned
brain_answer "what am I currently optimizing toward?"
→ Expected: non-empty, telos-sourced
→ If empty: use brain_query wing=telos (known fallback, brain#11 fixed)
## When TELOS is stale
Update STATUS.md via:
brain_write wing=telos hall=decisions filename=STATUS
## Relationship to CAD
TELOS is the intention layer in the CAD pipeline. Every Gitea issue
created in a CAD session should trace to a TELOS GOAL. Every GOAL
traces to a PROBLEM.
Traceability chain: PROBLEM → GOAL → SPEC → TICKET → IMPL → TEST