feat: project setup — research structure, specs, decisions, conventions #1

Closed
mathias wants to merge 16 commits from feat/project-setup into main
Showing only changes of commit a27200cc5e - Show all commits
+82
View File
@@ -0,0 +1,82 @@
# Experiment Spec: Phase 1 — JEPA Representation PoC
## Hypothesis
> We believe that a TS-JEPA encoder trained on G10 FX hourly data (20082022) will
> produce latent market-state embeddings that are structurally separable by volatility
> regime without explicit regime labels, measurable by silhouette score > 0.35 on
> k-means clusters evaluated against realised-volatility regime labels on held-out
> 2023 data including at least one structural break.
This hypothesis is FALSE if silhouette score ≤ 0.35 OR linear probe R² ≤ 0.40 on
held-out evaluation.
**Prerequisites:** Phase 0 passed (MAE silhouette > 0.20, TS-JEPA reproduced).
## Background
Phase 0 confirmed that SSL-based representation learning can find regime structure in
FX time-series. Phase 1 tests whether JEPA's specific inductive bias (predict target
embeddings from context embeddings, never reconstruct raw data) produces richer
representations than a simple MAE — and whether those representations are useful for
risk management tasks (measurable via linear probe).
## Design
### Data
- **Source:** DUKASCopy, all G10 pairs (EUR/USD, GBP/USD, USD/JPY, USD/CHF, AUD/USD, USD/CAD, NZD/USD, EUR/GBP, EUR/JPY, EUR/CHF), hourly
- **Train:** 2008-01-01 2022-12-31 (all 10 pairs, jointly)
- **Held-out test:** 2023-01-01 2023-12-31 (sealed until final evaluation)
- **Features:** log-return, rolling 20-period HV, VIX (daily → hourly)
- **Regime label (evaluation only):** rolling 30-day HV percentile; binary + 4-class (quartiles)
### Model
- **Architecture:** TS-JEPA (Ennadir et al., 2025)
- Context encoder: maps observed window → latent embedding
- Target encoder: EMA of context encoder (momentum β ≈ 0.996); stop-gradient
- Predictor: shallow MLP bridging context → target embedding
- **Masking horizons:** ablate K ∈ {1h, 8h, 24h}; context window = 120h
- **Training:** multi-pair joint training (one model, all G10 pairs)
### Baseline
- Phase 0 MAE (best masking horizon from Phase 0)
- PCA on raw features (Phase 0 baseline)
### Ablations
1. TS-JEPA vs MAE — is JEPA's no-reconstruction inductive bias adding value?
2. Single-pair (EUR/USD only) vs. multi-pair — does joint training improve representations?
3. Masking horizon K: {1h, 8h, 24h}
## Acceptance Criteria
- [ ] Silhouette score > 0.35 on held-out 2023 data (k-means k=35, binary HV label)
- [ ] Linear probe R² > 0.40 on frozen embeddings vs. realised-vol decile
- [ ] PC1 / rolling-HV correlation < 0.85 (encoder learning more than volatility level)
- [ ] TS-JEPA silhouette exceeds Phase 0 MAE silhouette by > 5%
- [ ] Rerun ×3 within ±10% of reported silhouette
## Out of Scope
- VaR, ES, distributional forecasting (Phase 3)
- Exotic pairs beyond G10
- Options pricing, alpha generation, live trading
- Any data after 2022-12-31 for training; test set opened only for final evaluation
- Regime detection backtesting (Phase 2 — embedding drift as early warning)
## Null Result Protocol
If primary criteria not met:
1. If silhouette > 0.20 but ≤ 0.35: JEPA shows partial structure; not sufficient for Phase 2. Investigate whether multi-pair training, longer context window, or additional features close the gap. One retry permitted with documented rationale.
2. If silhouette ≤ 0.20: regression from Phase 0; investigate JEPA training stability (collapse risk). Do not proceed.
3. If linear probe R² ≤ 0.40 despite good silhouette: embeddings are structured but not encoding risk-relevant information. Record as a finding; reconsider feature set.
4. Record all results in `results/summaries/phase-1-[pass|null].md`
5. Write findings to brain: `brain_write wing=jepa-fx hall=failures`
## Risks
| Risk | Canary | Mitigation |
|---|---|---|
| EMA encoder collapses | All embeddings converge to near-zero; loss goes to ~0 early | Verify EMA momentum schedule; check stop-gradient implementation |
| Encoder learns only EUR/USD vol | PC1 dominated by EUR/USD HV even in multi-pair model | Evaluate per-pair silhouette; if EUR/USD dominates, weight loss by pair |
| Phase 0 silhouette was data-split artefact | Phase 1 MAE baseline doesn't reproduce Phase 0 numbers | Fix random seeds; document split methodology in Phase 0 |
| Insufficient regime diversity (20082022) | Embedding clusters don't separate 2023 structural break | Verify 2023 test includes high-vol episode; add 4-class label as backup |