From 5ff79891b538c8c9f343e650e9de6946ee469148 Mon Sep 17 00:00:00 2001 From: mathias Date: Wed, 27 May 2026 21:58:50 +0000 Subject: [PATCH] feat: add experiment-spec skill --- experiment-spec/SKILL.md | 192 +++++++++++++++++++++++++++++++++++++++ 1 file changed, 192 insertions(+) create mode 100644 experiment-spec/SKILL.md diff --git a/experiment-spec/SKILL.md b/experiment-spec/SKILL.md new file mode 100644 index 0000000..96da443 --- /dev/null +++ b/experiment-spec/SKILL.md @@ -0,0 +1,192 @@ +--- +name: experiment-spec +description: Write a rigorous experiment spec for a research phase before any code or training runs. Use instead of feature-spec for scientific/ML research projects. Enforces falsifiable hypothesis, quantitative acceptance criteria, baseline comparison, and null-result protocol. +--- + +# Experiment Spec + +## Overview + +An experiment spec is the scientific contract for one research phase or experiment, written before any implementation or training begins. It is the research analogue of `feature-spec` — same discipline, different vocabulary. + +**Core principle:** If you cannot write a falsifiable hypothesis with a quantitative acceptance criterion, you do not understand the experiment well enough to run it. + +## When to Use + +- Starting a new research phase (Phase 0, Phase 1, etc.) +- Running any experiment that will produce metrics used to make a go/no-go decision +- Any time the question is "does X work?" rather than "build X" +- Before touching training code, data, or hyperparameters for a new question + +**When NOT to use:** +- Implementing a specific component whose behaviour is already defined by a phase spec (use `feature-spec` instead) +- Exploratory data analysis with no hypothesis (use a notebook; note it as EDA) +- Bug fixes or refactors + +## Iron Laws + +1. **The hypothesis must be falsifiable.** "JEPA is promising" is not a hypothesis. "JEPA embeddings will achieve silhouette > 0.35 on held-out data" is. If you cannot state conditions under which the hypothesis is false, it is not a hypothesis. +2. **Acceptance criteria must be quantitative and pre-registered.** Write the number before you run the experiment. Moving the goalposts after seeing results is p-hacking. +3. **A baseline is mandatory.** Every experiment must compare against at least one simpler baseline. "Better than nothing" is not a baseline. +4. **A null-result protocol is mandatory.** State what you will conclude and do if the hypothesis is rejected. "Try harder" is not a protocol. +5. **The training cutoff is sacred.** No post-cutoff data informs any decision in the spec or implementation. + +## Spec Template + +```markdown +# Experiment Spec: [Phase N — Short Name] + +## Hypothesis + +> "[Falsifiable claim]: We believe [X] will produce [Y], measurable by [Z]." + +State conditions under which this hypothesis is FALSE. + +## Background + +Why this experiment? What does it build on? What prior result or decision motivates it? +(2–4 sentences. Reference DECISIONS.md or brain wing entries where relevant.) + +## Design + +### Data +- Source, date range, pairs/assets, features used +- Train / validation / test split (respect training cutoff) + +### Model / method +- Architecture, configuration, key hyperparameters +- What is being varied vs. held fixed + +### Baseline +- What simpler method is being compared against? +- Why is this the right baseline? + +### Ablations (if any) +- What variants will be run to isolate the effect being studied? + +## Acceptance Criteria + +- [ ] [Primary criterion — quantitative threshold on primary metric] +- [ ] [Baseline comparison — e.g. "exceeds baseline by >X%"] +- [ ] [Reproducibility — reruns within ±Y% of reported metric] +- [ ] [Collapse/sanity check — e.g. "PC1/rolling-HV correlation < 0.85"] + +## Out of Scope + +What this experiment explicitly does NOT answer, even if related. +Anything plausibly in scope that is deferred goes here. + +## Null Result Protocol + +If the primary acceptance criterion is NOT met: +- What do we conclude? +- What is the next step? (Investigate X, pivot to Y, terminate programme) +- What gets written to the brain and results/summaries/? + +## Risks + +What could go wrong, and how would it be detected? +At least one risk must be listed. +``` + +## Worked Example + +```markdown +# Experiment Spec: Phase 0 — SSL Feasibility Gate + +## Hypothesis + +> We believe that a masked autoencoder (MAE) trained on FX hourly data will produce +> latent embeddings that show structural separability by volatility regime without +> explicit regime labels, measurable by silhouette score > 0.20 on held-out 2023 data. + +This hypothesis is FALSE if silhouette score ≤ 0.20 on the held-out evaluation. + +## Background + +Before investing in JEPA-specific machinery, we need to confirm that SSL-based +representation learning can find regime structure in FX time-series at all. MAE is +the simplest SSL baseline — if it cannot find structure, JEPA will not either. +Added post Full Grill (2026-05-27). See DECISIONS.md: "Phase 0: SSL feasibility gate". + +## Design + +### Data +- Source: DUKASCopy, EUR/USD hourly, 2008–2022 (train), 2023 (held-out test) +- Features: log-return, rolling 20-period HV, VIX (daily interpolated to hourly) +- Regime label (for evaluation only, not training): rolling 30-day HV percentile, + binary high/low threshold at 50th percentile + +### Model +- Masked Autoencoder: 1D temporal masking (mask contiguous 24h window) +- Encoder: 3-layer 1D CNN + positional encoding +- Decoder: 2-layer MLP reconstructing masked segment +- Context window: 120 hours (5 days) + +### Baseline +- PCA on raw feature vectors (same window) — tests whether any dimensionality + reduction shows regime structure, not just SSL + +### Ablations +- Masking horizon: K ∈ {8h, 24h, 72h} — does horizon affect embedding quality? + +## Acceptance Criteria + +- [ ] Silhouette score > 0.20 on held-out 2023 data (k-means, k=3, vs. HV regime label) +- [ ] MAE silhouette exceeds PCA baseline silhouette +- [ ] Rerun within ±10% of reported silhouette +- [ ] PC1 / rolling-HV correlation < 0.95 (not purely encoding volatility level) + +## Out of Scope + +- JEPA implementation (Phase 1) +- Multi-pair training (Phase 1+) +- VaR or ES computation +- Any use of post-2023 data + +## Null Result Protocol + +If silhouette ≤ 0.20: +- Conclude: SSL cannot reliably find regime structure in EUR/USD hourly data with + these features at this resolution +- Next step: investigate whether (a) hourly resolution is too noisy (try daily), + (b) 3 features are insufficient, or (c) regime label definition is too coarse +- Record result in results/summaries/phase-0-null.md and brain wing jepa-fx/failures/ + +## Risks + +- Encoder collapses to near-constant output: detect via reconstruction loss plateau + in first 10 epochs; mitigation: add batch norm, reduce learning rate +- Regime label too coarse (binary HV): silhouette may be low even with good structure; + mitigation: also evaluate with 4-class label (HV quartiles) +``` + +## Common Failure Modes + +| Failure mode | What it looks like | Fix | +|---|---|---| +| Non-falsifiable hypothesis | "JEPA shows promise" | Rewrite with a number | +| Post-hoc criteria | Threshold chosen after seeing results | Write the number first, commit the spec | +| No baseline | Silhouette of 0.30 sounds good until PCA achieves 0.35 | Always include a dumber method | +| Missing null protocol | "We'll figure it out if it fails" | Write it now — it forces clarity about what you're actually betting on | +| Cutoff violation | Architecture choice informed by 2024 data patterns | Never open the test set during development | + +## Brain MCP Integration + +**At spec start:** +- `brain_query wing=jepa-fx hall=decisions` — load current architectural decisions +- `brain_query wing=jepa-fx hall=failures` — load known failure modes; address them in Risks section + +**After spec is approved:** +- `brain_write` to `jepa-fx/hypotheses/` with the hypothesis and acceptance criteria + +**After experiment concludes:** +- `brain_write` to `jepa-fx/failures/` with any new failure modes discovered +- `session_log` with outcome + +## Cross-References + +- Use `feature-spec` for implementing a specific component within an already-specced phase +- Use `grill-me` on the spec before running the experiment if the hypothesis feels shaky +- Use `tdd` once the spec is approved — each acceptance criterion maps to a test +- Use `session-retrospective` after the experiment concludes