From e616575979d02738f6c17a053d3cf4000e98daee Mon Sep 17 00:00:00 2001 From: mathias Date: Mon, 22 Jun 2026 18:10:32 +0000 Subject: [PATCH] docs: add DECISIONS.md and rewrite PROJECT.md (#8) --- .context/PROJECT.md | 103 +++++++++++++++++++-- DECISIONS.md | 215 ++++++++++++++++++++++++++++++++++++++++++++ 2 files changed, 312 insertions(+), 6 deletions(-) create mode 100644 DECISIONS.md diff --git a/.context/PROJECT.md b/.context/PROJECT.md index 7779a58..8d5f930 100644 --- a/.context/PROJECT.md +++ b/.context/PROJECT.md @@ -1,13 +1,104 @@ -# hostexecutor +# jepa-fx-risk ## Identity -- **Name**: hostexecutor +- **Name**: jepa-fx-risk - **Owner**: Mathias -- **Client**: personal -- **Repo**: gitea.d-ma.be/mathias/hostexecutor -- **Status**: active +- **Client**: personal research +- **Repo**: gitea.d-ma.be/mathias/jepa-fx-risk +- **Status**: active — Phase 0 + +## Purpose + +Research project: apply JEPA-based self-supervised representation learning to +FX risk management for a corporate bank with an internal global FX trading desk. + +Primary tasks: FX volatility forecasting and VaR/CVaR estimation. +Target output: internal PoC for the trading desk. + +## Architecture decision (see DECISIONS.md ADR-001) + +**TS-JEPA + SIGReg** — TS-JEPA temporal patchwise architecture (Ennadir et al., +arXiv:2509.25449) with EMA replaced by Sketched Isotropic Gaussian Regularization +(SIGReg, Balestriero & LeCun, arXiv:2511.08544). Single search axis: λ ∈ [0.01, 1.0]. + +Phase 2 hypothesis: MTS-JEPA multi-resolution objective (arXiv:2602.04643). ## Stack -Go + Templ + HTMX + CDN Tailwind. See `~/dev/.context/AGENT.md` for cross-project conventions. +**Go** (`src/`): data pipeline (DUKASCopy fetch + hourly processing), evaluation +harness (silhouette, linear probe R², collapse diagnostic, Kupiec/Christoffersen), +results dashboard (Templ + HTMX + CDN Tailwind). + +**Python** (`model/`): all training and embedding export only. PyTorch cu130 +(Blackwell sm_120 compatible). No evaluation logic in Python. + +**Infra**: koala (Arch Linux, Blackwell GPU 12 GB VRAM) for training. +iguana (Mac Studio M2 Ultra) + LiteLLM on piguard for autoresearch agent LLM. + +## Repository layout + +``` +jepa-fx-risk/ +├── src/ # Go — data pipeline + eval harness + dashboard +│ ├── data/ # DUKASCopy fetch, hourly processing, validation +│ └── eval/ # silhouette, linear probe, collapse, backtest +├── model/ # Python — training only +│ ├── train.py # TS-JEPA + SIGReg backbone (autoresearch edits this) +│ ├── prepare.py # LOCKED — data loading, tokenization, export +│ └── requirements.txt +├── specs/ # Research specs (one per phase/experiment type) +├── experiments/ # Per-run outputs: embeddings, metrics.json, git tag +├── results/summaries/ # Human-readable outcome per experiment +├── program.md # Autoresearch agenda — researcher edits this +├── DECISIONS.md # Architecture Decision Records +└── Taskfile.yml # task data:fetch, task experiment:run, task eval:* +``` + +## Phase structure + +- **Phase 0** (current): MAE baseline on EUR/USD hourly 2008–2022. + Gate: silhouette > 0.20, MAE > PCA, ±10% over 3 reruns. +- **Phase 1**: TS-JEPA + SIGReg autoresearch sweep, 50 experiments. + Gate: val_vol_r2 > GARCH baseline AND Kupiec p > 0.05. +- **Phase 2**: MTS-JEPA multi-resolution hypothesis. +- **Phase 3**: Internal bank tick/position data (future, out of current scope). + +## Data + +DUKASCopy hourly OHLCV, 10 G10 pairs, 2008–2022 train / 2023 val / 2024 test. +Features: log-return, log rolling-20-period HV, VIX (daily interpolated). +Weekend gaps handled explicitly. See ADR-002. + +## Key conventions + +- `prepare.py` is LOCKED — never modified by agents or autoresearch +- Evaluation metrics are always computed by the Go harness, never in Python +- Every experiment gets a git tag: `exp/YYYYMMDD-description` +- Null results are recorded explicitly in `results/summaries/` — do not iterate silently +- `program.md` is the only file the researcher edits to steer autoresearch +- CI runs `task check` (lint + vet + test) only — no GPU, no training + +## Evaluation metrics + +Primary (autoresearch optimizes): `val_vol_r2` — linear probe R² on 1-day +realized volatility from frozen embeddings, computed by Go harness. + +Secondary (logged, not optimized): Kupiec p-value (VaR 99% backtest on EUR/USD). +Must co-move with val_vol_r2 — checked from experiment 1. + +Diagnostics: silhouette score (regime clustering), PC1/HV correlation (collapse check). + +## Benchmarks to beat (trading desk comparison) + +- GARCH(1,1) — volatility forecasting baseline +- Historical Simulation VaR (250-day rolling) — Basel default +- EWMA RiskMetrics (λ=0.94) + +## Agent guidance + +Read `DECISIONS.md` before making architecture suggestions. +Do not modify `prepare.py` or `src/eval/`. +Do not suggest changing the Python/Go separation. +Training runs are always manual via `task experiment:run` — never triggered by CI. +When implementing, follow Go conventions in `.skills/go-patterns/SKILL.md`. diff --git a/DECISIONS.md b/DECISIONS.md new file mode 100644 index 0000000..6d69aa2 --- /dev/null +++ b/DECISIONS.md @@ -0,0 +1,215 @@ +# Architecture Decision Records + +This file records significant technical and research decisions for `jepa-fx-risk`. +Each record is immutable once merged — append new records rather than editing old ones. +Format: ID · Date · Status · Context · Decision · Rationale · Consequences. + +--- + +## ADR-001 · Architecture: TS-JEPA + SIGReg as Phase 1 backbone + +**Date:** 2026-05-28 +**Status:** Accepted +**Supersedes:** informal decision to use TS-JEPA standalone (pre-ADR) + +### Context + +Four JEPA variants were evaluated for FX volatility forecasting and VaR/CVaR estimation: + +| Variant | Origin | Key property | +|---|---|---| +| TS-JEPA | Ennadir et al., Sep 2025 | Time-series native; EMA collapse prevention | +| LeJEPA | Balestriero & LeCun, Nov 2025 | Proven optimal embeddings (isotropic Gaussian); SIGReg | +| MTS-JEPA | He et al., Feb 2026 | Multi-resolution + codebook; no public code | +| Var-JEPA | Multiple, Mar 2026 | ELBO-based UQ; no public code | + +Key constraints: 12 GB VRAM (Blackwell, koala), hourly DUKASCopy data, internal PoC target, +autoresearch loop requires a clean single-scalar search space, trading desk requires an +explainable theoretical story. + +### Decision + +Use **TS-JEPA architecture with SIGReg replacing EMA** as the Phase 1 backbone. + +Concretely: +- Start from the TS-JEPA open-source implementation (arXiv:2509.25449, GitHub) +- Remove the EMA target-network mechanism +- Replace it with Sketched Isotropic Gaussian Regularization (SIGReg) from LeJEPA + (arXiv:2511.08544), controlled by a single λ hyperparameter +- Keep TS-JEPA's temporal patchwise masking and Transformer encoder unchanged + +### Rationale + +**Why not pure TS-JEPA:** EMA is a heuristic; λ interacts with EMA momentum and +learning rate, creating a three-way search space that is hard to navigate with autoresearch. +EMA also has no theoretical non-stationarity guarantee. + +**Why not pure LeJEPA:** The reference implementation targets vision (multi-crop views). +Adapting it to temporal patchwise masking requires non-trivial surgery and moves away from +open code. TS-JEPA's masking is already the right inductive bias for time series. + +**Why the hybrid:** SIGReg is architecture-agnostic — it operates on the embedding +distribution, not the encoder structure. Swapping EMA for SIGReg is a ~20-line change to +TS-JEPA's training loop. The result is: +- Time-series native (TS-JEPA masking + patch structure) +- Provably collapse-free without heuristics (SIGReg) +- Single search axis for autoresearch (λ ∈ [0.01, 1.0]) +- Non-stationarity robustness proven formally (arXiv:2602.19373 extends LeJEPA + guarantees to non-stationary target distributions — directly relevant to FX) +- Explainable to a model validation team: "embeddings are provably optimal for + downstream prediction under distributional uncertainty" + +**Why not MTS-JEPA or Var-JEPA now:** Both lack public code (as of May 2026). +MTS-JEPA's multi-resolution objective is the right next hypothesis (see ADR-003). +Var-JEPA's ELBO-based UQ is a compelling future direction for CVaR estimation. + +### Consequences + +- Phase 0 (MAE baseline) is unaffected — it precedes the JEPA architecture choice +- Issue #3 (TS-JEPA reproduction) is still the right first step; SIGReg is added after + reproduction is confirmed +- The autoresearch `program.md` primary search axis is λ (SIGReg weight) +- Secondary axes: masking block size, patch stride, encoder depth +- `model/requirements.txt` must include the SIGReg implementation (≈20 lines, + can be vendored directly) + +--- + +## ADR-002 · Data: DUKASCopy hourly G10 FX as primary training data + +**Date:** 2026-05-28 +**Status:** Accepted + +### Context + +Data scale is the most dangerous assumption for any SSL/JEPA approach. Daily FX data +(~5,000 samples over 20 years) is insufficient for self-supervised pretraining. +Two alternatives were considered: daily public data (yfinance) vs. hourly tick data +(DUKASCopy, free, rate-limited HTTP API). + +### Decision + +Use **DUKASCopy hourly OHLCV** as the primary data source. + +- 10 G10 pairs: EURUSD, GBPUSD, USDJPY, USDCHF, AUDUSD, NZDUSD, USDCAD, + EURGBP, EURJPY, GBPJPY +- Training window: 2008-01-01 – 2022-12-31 (~175,000 samples per pair) +- Validation window: 2023-01-01 – 2023-12-31 (~2,600 samples) +- Test window: 2024-01-01 – 2024-12-31 (held out, never seen during development) +- Features per bar: log-return, log rolling-20-period HV, VIX (daily interpolated) +- Weekend gaps handled explicitly — no interpolation across market close + +### Rationale + +Hourly data gives ~35× more samples than daily. This is the minimum threshold for +JEPA-style SSL to show a training signal within 10-minute autoresearch experiments. +DUKASCopy is free, reliable, and provides consistent tick-level source data back to 2003. + +### Consequences + +- The Go data pipeline (Issue #2) is the critical path for everything else +- Phase 0 MAE baseline trains on the same 2008-2022 window +- Daily data (yfinance) may still be used for VIX and rate differentials as auxiliary features + +--- + +## ADR-003 · Research roadmap: Phase structure and JEPA variant progression + +**Date:** 2026-05-28 +**Status:** Accepted + +### Decision + +Three-phase research roadmap: + +**Phase 0 — SSL feasibility gate (MAE baseline)** +Implement a 1D temporal MAE (not JEPA) on EUR/USD hourly 2008-2022. +Gate criteria: silhouette > 0.20 on 2023 held-out, MAE > PCA baseline, ±10% over 3 reruns. +Purpose: validate that the data and eval harness work before committing to JEPA complexity. +If gate fails: follow null result protocol in `specs/phase-0-ssl-feasibility.md`. + +**Phase 1 — TS-JEPA + SIGReg autoresearch sweep** +Primary architecture per ADR-001. +Autoresearch loop: `program.md`-driven, 10-min experiments, 50-experiment budget. +Primary metric: `val_vol_r2` (linear probe R² on 1-day realized volatility). +Gate criteria: `val_vol_r2` > GARCH-implied baseline AND Kupiec p-value > 0.05 on +EUR/USD VaR 99%. +Kupiec is logged from experiment 1 to verify it co-moves with `val_vol_r2`. + +**Phase 2 — MTS-JEPA multi-resolution hypothesis** +Introduce parallel multi-scale predictive pathways (1h, 8h, 24h context windows) +adapted from MTS-JEPA (arXiv:2602.04643). +Hypothesis: multi-scale representations improve regime detection (silhouette) and +reduce VaR exceedance clustering (Christoffersen test). +Prerequisite: Phase 1 gate passed AND MTS-JEPA code available or reproducible from paper. +Time-box: if MTS-JEPA code not available within 4 weeks of Phase 2 start, implement +multi-resolution masking from scratch using Phase 1 backbone as base. + +**Phase 3 — Internal bank data (future)** +Replace DUKASCopy pipeline with internal tick feed adapter. +Fine-tune heads only; backbone frozen or lightly fine-tuned. +Out of scope for current PoC cycle. + +### Consequences + +- Issue #5 (Phase 0 MAE) is the unblocked next executable step +- Phase 1 autoresearch is blocked until Phase 0 passes its gate +- Var-JEPA (ELBO-based UQ) is a named future hypothesis for CVaR estimation in Phase 2+ + but not on the critical path + +--- + +## ADR-004 · Evaluation: Go harness + Python training separation + +**Date:** 2026-05-28 +**Status:** Accepted + +### Decision + +Hard separation between training (Python) and evaluation (Go): + +- **Python** (`model/`): all training, embedding export, model checkpointing +- **Go** (`src/eval/`): all evaluation metrics — silhouette, linear probe R², collapse + diagnostic, Kupiec/Christoffersen backtests +- Interface: Python exports embedding matrices + labels to `experiments/RUNID/` as + `.npy` files; Go eval harness reads them and writes `metrics.json` + +### Rationale + +Go evaluation gives deterministic, fast, auditable metric computation with proper +unit tests. It decouples the experimental loop from the training framework, making +it possible to re-evaluate any past experiment without re-running training. +The Go layer also serves as the foundation for the eventual trading desk dashboard. + +### Consequences + +- All acceptance criteria in Issues #4 and #5 are specified in terms of Go eval outputs +- `val_vol_r2` (the autoresearch optimization metric) is computed by the Go harness, + not inside the Python training loop +- Python training loop calls `task eval:probe` as a subprocess after each experiment + to get the scalar fed back to autoresearch + +--- + +## ADR-005 · Compute: Blackwell GPU on koala, PyTorch cu130 + +**Date:** 2026-05-28 +**Status:** Accepted + +### Decision + +All GPU training runs on koala (Arch Linux, Blackwell GPU, 12 GB VRAM). +PyTorch install: `pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu130` +(CUDA 13.0 wheel — required for sm_120 Blackwell support; stable as of May 2026). +Driver requirement: NVIDIA R570+, CUDA toolkit 12.8+. + +Ollama on iguana (Mac Studio M2 Ultra) serves the autoresearch agent LLM via the +existing LiteLLM proxy on piguard. Agent calls never hit koala directly. + +### Consequences + +- `model/requirements.txt` must NOT pin torch to a cu124 or earlier wheel +- CI (Issue #7) must NOT run GPU tests — CPU-only for unit tests, GPU only via + `task experiment:run` on koala +- 12 GB VRAM is sufficient for <5M parameter models at batch=64; monitor if + autoresearch explores larger architectures