Files
2026-06-22 17:33:26 +00:00

7.7 KiB
Raw Permalink Blame History

name, description
name description
experiment-spec Write a rigorous experiment spec for a research phase before any code or training runs. Use instead of feature-spec for scientific/ML research projects. Enforces falsifiable hypothesis, quantitative acceptance criteria, baseline comparison, and null-result protocol.

Experiment Spec

Overview

An experiment spec is the scientific contract for one research phase or experiment, written before any implementation or training begins. It is the research analogue of feature-spec — same discipline, different vocabulary.

Core principle: If you cannot write a falsifiable hypothesis with a quantitative acceptance criterion, you do not understand the experiment well enough to run it.

When to Use

  • Starting a new research phase (Phase 0, Phase 1, etc.)
  • Running any experiment that will produce metrics used to make a go/no-go decision
  • Any time the question is "does X work?" rather than "build X"
  • Before touching training code, data, or hyperparameters for a new question

When NOT to use:

  • Implementing a specific component whose behaviour is already defined by a phase spec (use feature-spec instead)
  • Exploratory data analysis with no hypothesis (use a notebook; note it as EDA)
  • Bug fixes or refactors

Iron Laws

  1. The hypothesis must be falsifiable. "JEPA is promising" is not a hypothesis. "JEPA embeddings will achieve silhouette > 0.35 on held-out data" is. If you cannot state conditions under which the hypothesis is false, it is not a hypothesis.
  2. Acceptance criteria must be quantitative and pre-registered. Write the number before you run the experiment. Moving the goalposts after seeing results is p-hacking.
  3. A baseline is mandatory. Every experiment must compare against at least one simpler baseline. "Better than nothing" is not a baseline.
  4. A null-result protocol is mandatory. State what you will conclude and do if the hypothesis is rejected. "Try harder" is not a protocol.
  5. The training cutoff is sacred. No post-cutoff data informs any decision in the spec or implementation.

Spec Template

# Experiment Spec: [Phase N — Short Name]

## Hypothesis

> "[Falsifiable claim]: We believe [X] will produce [Y], measurable by [Z]."

State conditions under which this hypothesis is FALSE.

## Background

Why this experiment? What does it build on? What prior result or decision motivates it?
(24 sentences. Reference DECISIONS.md or brain wing entries where relevant.)

## Design

### Data
- Source, date range, pairs/assets, features used
- Train / validation / test split (respect training cutoff)

### Model / method
- Architecture, configuration, key hyperparameters
- What is being varied vs. held fixed

### Baseline
- What simpler method is being compared against?
- Why is this the right baseline?

### Ablations (if any)
- What variants will be run to isolate the effect being studied?

## Acceptance Criteria

- [ ] [Primary criterion — quantitative threshold on primary metric]
- [ ] [Baseline comparison — e.g. "exceeds baseline by >X%"]
- [ ] [Reproducibility — reruns within ±Y% of reported metric]
- [ ] [Collapse/sanity check — e.g. "PC1/rolling-HV correlation < 0.85"]

## Out of Scope

What this experiment explicitly does NOT answer, even if related.
Anything plausibly in scope that is deferred goes here.

## Null Result Protocol

If the primary acceptance criterion is NOT met:
- What do we conclude?
- What is the next step? (Investigate X, pivot to Y, terminate programme)
- What gets written to the brain and results/summaries/?

## Risks

What could go wrong, and how would it be detected?
At least one risk must be listed.

Worked Example

# Experiment Spec: Phase 0 — SSL Feasibility Gate

## Hypothesis

> We believe that a masked autoencoder (MAE) trained on FX hourly data will produce
> latent embeddings that show structural separability by volatility regime without
> explicit regime labels, measurable by silhouette score > 0.20 on held-out 2023 data.

This hypothesis is FALSE if silhouette score ≤ 0.20 on the held-out evaluation.

## Background

Before investing in JEPA-specific machinery, we need to confirm that SSL-based
representation learning can find regime structure in FX time-series at all. MAE is
the simplest SSL baseline — if it cannot find structure, JEPA will not either.
Added post Full Grill (2026-05-27). See DECISIONS.md: "Phase 0: SSL feasibility gate".

## Design

### Data
- Source: DUKASCopy, EUR/USD hourly, 20082022 (train), 2023 (held-out test)
- Features: log-return, rolling 20-period HV, VIX (daily interpolated to hourly)
- Regime label (for evaluation only, not training): rolling 30-day HV percentile,
  binary high/low threshold at 50th percentile

### Model
- Masked Autoencoder: 1D temporal masking (mask contiguous 24h window)
- Encoder: 3-layer 1D CNN + positional encoding
- Decoder: 2-layer MLP reconstructing masked segment
- Context window: 120 hours (5 days)

### Baseline
- PCA on raw feature vectors (same window) — tests whether any dimensionality
  reduction shows regime structure, not just SSL

### Ablations
- Masking horizon: K ∈ {8h, 24h, 72h} — does horizon affect embedding quality?

## Acceptance Criteria

- [ ] Silhouette score > 0.20 on held-out 2023 data (k-means, k=3, vs. HV regime label)
- [ ] MAE silhouette exceeds PCA baseline silhouette
- [ ] Rerun within ±10% of reported silhouette
- [ ] PC1 / rolling-HV correlation < 0.95 (not purely encoding volatility level)

## Out of Scope

- JEPA implementation (Phase 1)
- Multi-pair training (Phase 1+)
- VaR or ES computation
- Any use of post-2023 data

## Null Result Protocol

If silhouette ≤ 0.20:
- Conclude: SSL cannot reliably find regime structure in EUR/USD hourly data with
  these features at this resolution
- Next step: investigate whether (a) hourly resolution is too noisy (try daily),
  (b) 3 features are insufficient, or (c) regime label definition is too coarse
- Record result in results/summaries/phase-0-null.md and brain wing jepa-fx/failures/

## Risks

- Encoder collapses to near-constant output: detect via reconstruction loss plateau
  in first 10 epochs; mitigation: add batch norm, reduce learning rate
- Regime label too coarse (binary HV): silhouette may be low even with good structure;
  mitigation: also evaluate with 4-class label (HV quartiles)

Common Failure Modes

Failure mode What it looks like Fix
Non-falsifiable hypothesis "JEPA shows promise" Rewrite with a number
Post-hoc criteria Threshold chosen after seeing results Write the number first, commit the spec
No baseline Silhouette of 0.30 sounds good until PCA achieves 0.35 Always include a dumber method
Missing null protocol "We'll figure it out if it fails" Write it now — it forces clarity about what you're actually betting on
Cutoff violation Architecture choice informed by 2024 data patterns Never open the test set during development

Brain MCP Integration

At spec start:

  • brain_query wing=jepa-fx hall=decisions — load current architectural decisions
  • brain_query wing=jepa-fx hall=failures — load known failure modes; address them in Risks section

After spec is approved:

  • brain_write to jepa-fx/hypotheses/ with the hypothesis and acceptance criteria

After experiment concludes:

  • brain_write to jepa-fx/failures/ with any new failure modes discovered
  • session_log with outcome

Cross-References

  • Use feature-spec for implementing a specific component within an already-specced phase
  • Use grill-me on the spec before running the experiment if the hypothesis feels shaky
  • Use tdd once the spec is approved — each acceptance criterion maps to a test
  • Use session-retrospective after the experiment concludes