feat: add experiment-spec skill for research projects #1

Merged
mathias merged 1 commits from feat/experiment-spec-skill into main 2026-06-22 17:33:26 +00:00
+192
View File
@@ -0,0 +1,192 @@
---
name: experiment-spec
description: Write a rigorous experiment spec for a research phase before any code or training runs. Use instead of feature-spec for scientific/ML research projects. Enforces falsifiable hypothesis, quantitative acceptance criteria, baseline comparison, and null-result protocol.
---
# Experiment Spec
## Overview
An experiment spec is the scientific contract for one research phase or experiment, written before any implementation or training begins. It is the research analogue of `feature-spec` — same discipline, different vocabulary.
**Core principle:** If you cannot write a falsifiable hypothesis with a quantitative acceptance criterion, you do not understand the experiment well enough to run it.
## When to Use
- Starting a new research phase (Phase 0, Phase 1, etc.)
- Running any experiment that will produce metrics used to make a go/no-go decision
- Any time the question is "does X work?" rather than "build X"
- Before touching training code, data, or hyperparameters for a new question
**When NOT to use:**
- Implementing a specific component whose behaviour is already defined by a phase spec (use `feature-spec` instead)
- Exploratory data analysis with no hypothesis (use a notebook; note it as EDA)
- Bug fixes or refactors
## Iron Laws
1. **The hypothesis must be falsifiable.** "JEPA is promising" is not a hypothesis. "JEPA embeddings will achieve silhouette > 0.35 on held-out data" is. If you cannot state conditions under which the hypothesis is false, it is not a hypothesis.
2. **Acceptance criteria must be quantitative and pre-registered.** Write the number before you run the experiment. Moving the goalposts after seeing results is p-hacking.
3. **A baseline is mandatory.** Every experiment must compare against at least one simpler baseline. "Better than nothing" is not a baseline.
4. **A null-result protocol is mandatory.** State what you will conclude and do if the hypothesis is rejected. "Try harder" is not a protocol.
5. **The training cutoff is sacred.** No post-cutoff data informs any decision in the spec or implementation.
## Spec Template
```markdown
# Experiment Spec: [Phase N — Short Name]
## Hypothesis
> "[Falsifiable claim]: We believe [X] will produce [Y], measurable by [Z]."
State conditions under which this hypothesis is FALSE.
## Background
Why this experiment? What does it build on? What prior result or decision motivates it?
(24 sentences. Reference DECISIONS.md or brain wing entries where relevant.)
## Design
### Data
- Source, date range, pairs/assets, features used
- Train / validation / test split (respect training cutoff)
### Model / method
- Architecture, configuration, key hyperparameters
- What is being varied vs. held fixed
### Baseline
- What simpler method is being compared against?
- Why is this the right baseline?
### Ablations (if any)
- What variants will be run to isolate the effect being studied?
## Acceptance Criteria
- [ ] [Primary criterion — quantitative threshold on primary metric]
- [ ] [Baseline comparison — e.g. "exceeds baseline by >X%"]
- [ ] [Reproducibility — reruns within ±Y% of reported metric]
- [ ] [Collapse/sanity check — e.g. "PC1/rolling-HV correlation < 0.85"]
## Out of Scope
What this experiment explicitly does NOT answer, even if related.
Anything plausibly in scope that is deferred goes here.
## Null Result Protocol
If the primary acceptance criterion is NOT met:
- What do we conclude?
- What is the next step? (Investigate X, pivot to Y, terminate programme)
- What gets written to the brain and results/summaries/?
## Risks
What could go wrong, and how would it be detected?
At least one risk must be listed.
```
## Worked Example
```markdown
# Experiment Spec: Phase 0 — SSL Feasibility Gate
## Hypothesis
> We believe that a masked autoencoder (MAE) trained on FX hourly data will produce
> latent embeddings that show structural separability by volatility regime without
> explicit regime labels, measurable by silhouette score > 0.20 on held-out 2023 data.
This hypothesis is FALSE if silhouette score ≤ 0.20 on the held-out evaluation.
## Background
Before investing in JEPA-specific machinery, we need to confirm that SSL-based
representation learning can find regime structure in FX time-series at all. MAE is
the simplest SSL baseline — if it cannot find structure, JEPA will not either.
Added post Full Grill (2026-05-27). See DECISIONS.md: "Phase 0: SSL feasibility gate".
## Design
### Data
- Source: DUKASCopy, EUR/USD hourly, 20082022 (train), 2023 (held-out test)
- Features: log-return, rolling 20-period HV, VIX (daily interpolated to hourly)
- Regime label (for evaluation only, not training): rolling 30-day HV percentile,
binary high/low threshold at 50th percentile
### Model
- Masked Autoencoder: 1D temporal masking (mask contiguous 24h window)
- Encoder: 3-layer 1D CNN + positional encoding
- Decoder: 2-layer MLP reconstructing masked segment
- Context window: 120 hours (5 days)
### Baseline
- PCA on raw feature vectors (same window) — tests whether any dimensionality
reduction shows regime structure, not just SSL
### Ablations
- Masking horizon: K ∈ {8h, 24h, 72h} — does horizon affect embedding quality?
## Acceptance Criteria
- [ ] Silhouette score > 0.20 on held-out 2023 data (k-means, k=3, vs. HV regime label)
- [ ] MAE silhouette exceeds PCA baseline silhouette
- [ ] Rerun within ±10% of reported silhouette
- [ ] PC1 / rolling-HV correlation < 0.95 (not purely encoding volatility level)
## Out of Scope
- JEPA implementation (Phase 1)
- Multi-pair training (Phase 1+)
- VaR or ES computation
- Any use of post-2023 data
## Null Result Protocol
If silhouette ≤ 0.20:
- Conclude: SSL cannot reliably find regime structure in EUR/USD hourly data with
these features at this resolution
- Next step: investigate whether (a) hourly resolution is too noisy (try daily),
(b) 3 features are insufficient, or (c) regime label definition is too coarse
- Record result in results/summaries/phase-0-null.md and brain wing jepa-fx/failures/
## Risks
- Encoder collapses to near-constant output: detect via reconstruction loss plateau
in first 10 epochs; mitigation: add batch norm, reduce learning rate
- Regime label too coarse (binary HV): silhouette may be low even with good structure;
mitigation: also evaluate with 4-class label (HV quartiles)
```
## Common Failure Modes
| Failure mode | What it looks like | Fix |
|---|---|---|
| Non-falsifiable hypothesis | "JEPA shows promise" | Rewrite with a number |
| Post-hoc criteria | Threshold chosen after seeing results | Write the number first, commit the spec |
| No baseline | Silhouette of 0.30 sounds good until PCA achieves 0.35 | Always include a dumber method |
| Missing null protocol | "We'll figure it out if it fails" | Write it now — it forces clarity about what you're actually betting on |
| Cutoff violation | Architecture choice informed by 2024 data patterns | Never open the test set during development |
## Brain MCP Integration
**At spec start:**
- `brain_query wing=jepa-fx hall=decisions` — load current architectural decisions
- `brain_query wing=jepa-fx hall=failures` — load known failure modes; address them in Risks section
**After spec is approved:**
- `brain_write` to `jepa-fx/hypotheses/` with the hypothesis and acceptance criteria
**After experiment concludes:**
- `brain_write` to `jepa-fx/failures/` with any new failure modes discovered
- `session_log` with outcome
## Cross-References
- Use `feature-spec` for implementing a specific component within an already-specced phase
- Use `grill-me` on the spec before running the experiment if the hypothesis feels shaky
- Use `tdd` once the spec is approved — each acceptance criterion maps to a test
- Use `session-retrospective` after the experiment concludes