From a27200cc5ee3aa2db4b2b424680325330605abe4 Mon Sep 17 00:00:00 2001 From: mathias Date: Wed, 27 May 2026 21:59:43 +0000 Subject: [PATCH] docs: add Phase 1 experiment spec (JEPA representation PoC) --- specs/phase-1-representation-poc.md | 82 +++++++++++++++++++++++++++++ 1 file changed, 82 insertions(+) create mode 100644 specs/phase-1-representation-poc.md diff --git a/specs/phase-1-representation-poc.md b/specs/phase-1-representation-poc.md new file mode 100644 index 0000000..64fd6c8 --- /dev/null +++ b/specs/phase-1-representation-poc.md @@ -0,0 +1,82 @@ +# Experiment Spec: Phase 1 — JEPA Representation PoC + +## Hypothesis + +> We believe that a TS-JEPA encoder trained on G10 FX hourly data (2008–2022) will +> produce latent market-state embeddings that are structurally separable by volatility +> regime without explicit regime labels, measurable by silhouette score > 0.35 on +> k-means clusters evaluated against realised-volatility regime labels on held-out +> 2023 data including at least one structural break. + +This hypothesis is FALSE if silhouette score ≤ 0.35 OR linear probe R² ≤ 0.40 on +held-out evaluation. + +**Prerequisites:** Phase 0 passed (MAE silhouette > 0.20, TS-JEPA reproduced). + +## Background + +Phase 0 confirmed that SSL-based representation learning can find regime structure in +FX time-series. Phase 1 tests whether JEPA's specific inductive bias (predict target +embeddings from context embeddings, never reconstruct raw data) produces richer +representations than a simple MAE — and whether those representations are useful for +risk management tasks (measurable via linear probe). + +## Design + +### Data +- **Source:** DUKASCopy, all G10 pairs (EUR/USD, GBP/USD, USD/JPY, USD/CHF, AUD/USD, USD/CAD, NZD/USD, EUR/GBP, EUR/JPY, EUR/CHF), hourly +- **Train:** 2008-01-01 – 2022-12-31 (all 10 pairs, jointly) +- **Held-out test:** 2023-01-01 – 2023-12-31 (sealed until final evaluation) +- **Features:** log-return, rolling 20-period HV, VIX (daily → hourly) +- **Regime label (evaluation only):** rolling 30-day HV percentile; binary + 4-class (quartiles) + +### Model +- **Architecture:** TS-JEPA (Ennadir et al., 2025) + - Context encoder: maps observed window → latent embedding + - Target encoder: EMA of context encoder (momentum β ≈ 0.996); stop-gradient + - Predictor: shallow MLP bridging context → target embedding +- **Masking horizons:** ablate K ∈ {1h, 8h, 24h}; context window = 120h +- **Training:** multi-pair joint training (one model, all G10 pairs) + +### Baseline +- Phase 0 MAE (best masking horizon from Phase 0) +- PCA on raw features (Phase 0 baseline) + +### Ablations +1. TS-JEPA vs MAE — is JEPA's no-reconstruction inductive bias adding value? +2. Single-pair (EUR/USD only) vs. multi-pair — does joint training improve representations? +3. Masking horizon K: {1h, 8h, 24h} + +## Acceptance Criteria + +- [ ] Silhouette score > 0.35 on held-out 2023 data (k-means k=3–5, binary HV label) +- [ ] Linear probe R² > 0.40 on frozen embeddings vs. realised-vol decile +- [ ] PC1 / rolling-HV correlation < 0.85 (encoder learning more than volatility level) +- [ ] TS-JEPA silhouette exceeds Phase 0 MAE silhouette by > 5% +- [ ] Rerun ×3 within ±10% of reported silhouette + +## Out of Scope + +- VaR, ES, distributional forecasting (Phase 3) +- Exotic pairs beyond G10 +- Options pricing, alpha generation, live trading +- Any data after 2022-12-31 for training; test set opened only for final evaluation +- Regime detection backtesting (Phase 2 — embedding drift as early warning) + +## Null Result Protocol + +If primary criteria not met: +1. If silhouette > 0.20 but ≤ 0.35: JEPA shows partial structure; not sufficient for Phase 2. Investigate whether multi-pair training, longer context window, or additional features close the gap. One retry permitted with documented rationale. +2. If silhouette ≤ 0.20: regression from Phase 0; investigate JEPA training stability (collapse risk). Do not proceed. +3. If linear probe R² ≤ 0.40 despite good silhouette: embeddings are structured but not encoding risk-relevant information. Record as a finding; reconsider feature set. +4. Record all results in `results/summaries/phase-1-[pass|null].md` +5. Write findings to brain: `brain_write wing=jepa-fx hall=failures` + +## Risks + +| Risk | Canary | Mitigation | +|---|---|---| +| EMA encoder collapses | All embeddings converge to near-zero; loss goes to ~0 early | Verify EMA momentum schedule; check stop-gradient implementation | +| Encoder learns only EUR/USD vol | PC1 dominated by EUR/USD HV even in multi-pair model | Evaluate per-pair silhouette; if EUR/USD dominates, weight loss by pair | +| Phase 0 silhouette was data-split artefact | Phase 1 MAE baseline doesn't reproduce Phase 0 numbers | Fix random seeds; document split methodology in Phase 0 | +| Insufficient regime diversity (2008–2022) | Embedding clusters don't separate 2023 structural break | Verify 2023 test includes high-vol episode; add 4-class label as backup |