4.3 KiB
Experiment Spec: Phase 1 — JEPA Representation PoC
Hypothesis
We believe that a TS-JEPA encoder trained on G10 FX hourly data (2008–2022) will produce latent market-state embeddings that are structurally separable by volatility regime without explicit regime labels, measurable by silhouette score > 0.35 on k-means clusters evaluated against realised-volatility regime labels on held-out 2023 data including at least one structural break.
This hypothesis is FALSE if silhouette score ≤ 0.35 OR linear probe R² ≤ 0.40 on held-out evaluation.
Prerequisites: Phase 0 passed (MAE silhouette > 0.20, TS-JEPA reproduced).
Background
Phase 0 confirmed that SSL-based representation learning can find regime structure in FX time-series. Phase 1 tests whether JEPA's specific inductive bias (predict target embeddings from context embeddings, never reconstruct raw data) produces richer representations than a simple MAE — and whether those representations are useful for risk management tasks (measurable via linear probe).
Design
Data
- Source: DUKASCopy, all G10 pairs (EUR/USD, GBP/USD, USD/JPY, USD/CHF, AUD/USD, USD/CAD, NZD/USD, EUR/GBP, EUR/JPY, EUR/CHF), hourly
- Train: 2008-01-01 – 2022-12-31 (all 10 pairs, jointly)
- Held-out test: 2023-01-01 – 2023-12-31 (sealed until final evaluation)
- Features: log-return, rolling 20-period HV, VIX (daily → hourly)
- Regime label (evaluation only): rolling 30-day HV percentile; binary + 4-class (quartiles)
Model
- Architecture: TS-JEPA (Ennadir et al., 2025)
- Context encoder: maps observed window → latent embedding
- Target encoder: EMA of context encoder (momentum β ≈ 0.996); stop-gradient
- Predictor: shallow MLP bridging context → target embedding
- Masking horizons: ablate K ∈ {1h, 8h, 24h}; context window = 120h
- Training: multi-pair joint training (one model, all G10 pairs)
Baseline
- Phase 0 MAE (best masking horizon from Phase 0)
- PCA on raw features (Phase 0 baseline)
Ablations
- TS-JEPA vs MAE — is JEPA's no-reconstruction inductive bias adding value?
- Single-pair (EUR/USD only) vs. multi-pair — does joint training improve representations?
- Masking horizon K: {1h, 8h, 24h}
Acceptance Criteria
- Silhouette score > 0.35 on held-out 2023 data (k-means k=3–5, binary HV label)
- Linear probe R² > 0.40 on frozen embeddings vs. realised-vol decile
- PC1 / rolling-HV correlation < 0.85 (encoder learning more than volatility level)
- TS-JEPA silhouette exceeds Phase 0 MAE silhouette by > 5%
- Rerun ×3 within ±10% of reported silhouette
Out of Scope
- VaR, ES, distributional forecasting (Phase 3)
- Exotic pairs beyond G10
- Options pricing, alpha generation, live trading
- Any data after 2022-12-31 for training; test set opened only for final evaluation
- Regime detection backtesting (Phase 2 — embedding drift as early warning)
Null Result Protocol
If primary criteria not met:
- If silhouette > 0.20 but ≤ 0.35: JEPA shows partial structure; not sufficient for Phase 2. Investigate whether multi-pair training, longer context window, or additional features close the gap. One retry permitted with documented rationale.
- If silhouette ≤ 0.20: regression from Phase 0; investigate JEPA training stability (collapse risk). Do not proceed.
- If linear probe R² ≤ 0.40 despite good silhouette: embeddings are structured but not encoding risk-relevant information. Record as a finding; reconsider feature set.
- Record all results in
results/summaries/phase-1-[pass|null].md - Write findings to brain:
brain_write wing=jepa-fx hall=failures
Risks
| Risk | Canary | Mitigation |
|---|---|---|
| EMA encoder collapses | All embeddings converge to near-zero; loss goes to ~0 early | Verify EMA momentum schedule; check stop-gradient implementation |
| Encoder learns only EUR/USD vol | PC1 dominated by EUR/USD HV even in multi-pair model | Evaluate per-pair silhouette; if EUR/USD dominates, weight loss by pair |
| Phase 0 silhouette was data-split artefact | Phase 1 MAE baseline doesn't reproduce Phase 0 numbers | Fix random seeds; document split methodology in Phase 0 |
| Insufficient regime diversity (2008–2022) | Embedding clusters don't separate 2023 structural break | Verify 2023 test includes high-vol episode; add 4-class label as backup |