Files
skills/experiment-spec/SKILL.md
T

193 lines
7.7 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
---
name: experiment-spec
description: Write a rigorous experiment spec for a research phase before any code or training runs. Use instead of feature-spec for scientific/ML research projects. Enforces falsifiable hypothesis, quantitative acceptance criteria, baseline comparison, and null-result protocol.
---
# Experiment Spec
## Overview
An experiment spec is the scientific contract for one research phase or experiment, written before any implementation or training begins. It is the research analogue of `feature-spec` — same discipline, different vocabulary.
**Core principle:** If you cannot write a falsifiable hypothesis with a quantitative acceptance criterion, you do not understand the experiment well enough to run it.
## When to Use
- Starting a new research phase (Phase 0, Phase 1, etc.)
- Running any experiment that will produce metrics used to make a go/no-go decision
- Any time the question is "does X work?" rather than "build X"
- Before touching training code, data, or hyperparameters for a new question
**When NOT to use:**
- Implementing a specific component whose behaviour is already defined by a phase spec (use `feature-spec` instead)
- Exploratory data analysis with no hypothesis (use a notebook; note it as EDA)
- Bug fixes or refactors
## Iron Laws
1. **The hypothesis must be falsifiable.** "JEPA is promising" is not a hypothesis. "JEPA embeddings will achieve silhouette > 0.35 on held-out data" is. If you cannot state conditions under which the hypothesis is false, it is not a hypothesis.
2. **Acceptance criteria must be quantitative and pre-registered.** Write the number before you run the experiment. Moving the goalposts after seeing results is p-hacking.
3. **A baseline is mandatory.** Every experiment must compare against at least one simpler baseline. "Better than nothing" is not a baseline.
4. **A null-result protocol is mandatory.** State what you will conclude and do if the hypothesis is rejected. "Try harder" is not a protocol.
5. **The training cutoff is sacred.** No post-cutoff data informs any decision in the spec or implementation.
## Spec Template
```markdown
# Experiment Spec: [Phase N — Short Name]
## Hypothesis
> "[Falsifiable claim]: We believe [X] will produce [Y], measurable by [Z]."
State conditions under which this hypothesis is FALSE.
## Background
Why this experiment? What does it build on? What prior result or decision motivates it?
(24 sentences. Reference DECISIONS.md or brain wing entries where relevant.)
## Design
### Data
- Source, date range, pairs/assets, features used
- Train / validation / test split (respect training cutoff)
### Model / method
- Architecture, configuration, key hyperparameters
- What is being varied vs. held fixed
### Baseline
- What simpler method is being compared against?
- Why is this the right baseline?
### Ablations (if any)
- What variants will be run to isolate the effect being studied?
## Acceptance Criteria
- [ ] [Primary criterion — quantitative threshold on primary metric]
- [ ] [Baseline comparison — e.g. "exceeds baseline by >X%"]
- [ ] [Reproducibility — reruns within ±Y% of reported metric]
- [ ] [Collapse/sanity check — e.g. "PC1/rolling-HV correlation < 0.85"]
## Out of Scope
What this experiment explicitly does NOT answer, even if related.
Anything plausibly in scope that is deferred goes here.
## Null Result Protocol
If the primary acceptance criterion is NOT met:
- What do we conclude?
- What is the next step? (Investigate X, pivot to Y, terminate programme)
- What gets written to the brain and results/summaries/?
## Risks
What could go wrong, and how would it be detected?
At least one risk must be listed.
```
## Worked Example
```markdown
# Experiment Spec: Phase 0 — SSL Feasibility Gate
## Hypothesis
> We believe that a masked autoencoder (MAE) trained on FX hourly data will produce
> latent embeddings that show structural separability by volatility regime without
> explicit regime labels, measurable by silhouette score > 0.20 on held-out 2023 data.
This hypothesis is FALSE if silhouette score ≤ 0.20 on the held-out evaluation.
## Background
Before investing in JEPA-specific machinery, we need to confirm that SSL-based
representation learning can find regime structure in FX time-series at all. MAE is
the simplest SSL baseline — if it cannot find structure, JEPA will not either.
Added post Full Grill (2026-05-27). See DECISIONS.md: "Phase 0: SSL feasibility gate".
## Design
### Data
- Source: DUKASCopy, EUR/USD hourly, 20082022 (train), 2023 (held-out test)
- Features: log-return, rolling 20-period HV, VIX (daily interpolated to hourly)
- Regime label (for evaluation only, not training): rolling 30-day HV percentile,
binary high/low threshold at 50th percentile
### Model
- Masked Autoencoder: 1D temporal masking (mask contiguous 24h window)
- Encoder: 3-layer 1D CNN + positional encoding
- Decoder: 2-layer MLP reconstructing masked segment
- Context window: 120 hours (5 days)
### Baseline
- PCA on raw feature vectors (same window) — tests whether any dimensionality
reduction shows regime structure, not just SSL
### Ablations
- Masking horizon: K ∈ {8h, 24h, 72h} — does horizon affect embedding quality?
## Acceptance Criteria
- [ ] Silhouette score > 0.20 on held-out 2023 data (k-means, k=3, vs. HV regime label)
- [ ] MAE silhouette exceeds PCA baseline silhouette
- [ ] Rerun within ±10% of reported silhouette
- [ ] PC1 / rolling-HV correlation < 0.95 (not purely encoding volatility level)
## Out of Scope
- JEPA implementation (Phase 1)
- Multi-pair training (Phase 1+)
- VaR or ES computation
- Any use of post-2023 data
## Null Result Protocol
If silhouette ≤ 0.20:
- Conclude: SSL cannot reliably find regime structure in EUR/USD hourly data with
these features at this resolution
- Next step: investigate whether (a) hourly resolution is too noisy (try daily),
(b) 3 features are insufficient, or (c) regime label definition is too coarse
- Record result in results/summaries/phase-0-null.md and brain wing jepa-fx/failures/
## Risks
- Encoder collapses to near-constant output: detect via reconstruction loss plateau
in first 10 epochs; mitigation: add batch norm, reduce learning rate
- Regime label too coarse (binary HV): silhouette may be low even with good structure;
mitigation: also evaluate with 4-class label (HV quartiles)
```
## Common Failure Modes
| Failure mode | What it looks like | Fix |
|---|---|---|
| Non-falsifiable hypothesis | "JEPA shows promise" | Rewrite with a number |
| Post-hoc criteria | Threshold chosen after seeing results | Write the number first, commit the spec |
| No baseline | Silhouette of 0.30 sounds good until PCA achieves 0.35 | Always include a dumber method |
| Missing null protocol | "We'll figure it out if it fails" | Write it now — it forces clarity about what you're actually betting on |
| Cutoff violation | Architecture choice informed by 2024 data patterns | Never open the test set during development |
## Brain MCP Integration
**At spec start:**
- `brain_query wing=jepa-fx hall=decisions` — load current architectural decisions
- `brain_query wing=jepa-fx hall=failures` — load known failure modes; address them in Risks section
**After spec is approved:**
- `brain_write` to `jepa-fx/hypotheses/` with the hypothesis and acceptance criteria
**After experiment concludes:**
- `brain_write` to `jepa-fx/failures/` with any new failure modes discovered
- `session_log` with outcome
## Cross-References
- Use `feature-spec` for implementing a specific component within an already-specced phase
- Use `grill-me` on the spec before running the experiment if the hypothesis feels shaky
- Use `tdd` once the spec is approved — each acceptance criterion maps to a test
- Use `session-retrospective` after the experiment concludes