# Architecture Decision Records This file records significant technical and research decisions for `jepa-fx-risk`. Each record is immutable once merged — append new records rather than editing old ones. Format: ID · Date · Status · Context · Decision · Rationale · Consequences. --- ## ADR-001 · Architecture: TS-JEPA + SIGReg as Phase 1 backbone **Date:** 2026-05-28 **Status:** Accepted **Supersedes:** informal decision to use TS-JEPA standalone (pre-ADR) ### Context Four JEPA variants were evaluated for FX volatility forecasting and VaR/CVaR estimation: | Variant | Origin | Key property | |---|---|---| | TS-JEPA | Ennadir et al., Sep 2025 | Time-series native; EMA collapse prevention | | LeJEPA | Balestriero & LeCun, Nov 2025 | Proven optimal embeddings (isotropic Gaussian); SIGReg | | MTS-JEPA | He et al., Feb 2026 | Multi-resolution + codebook; no public code | | Var-JEPA | Multiple, Mar 2026 | ELBO-based UQ; no public code | Key constraints: 12 GB VRAM (Blackwell, koala), hourly DUKASCopy data, internal PoC target, autoresearch loop requires a clean single-scalar search space, trading desk requires an explainable theoretical story. ### Decision Use **TS-JEPA architecture with SIGReg replacing EMA** as the Phase 1 backbone. Concretely: - Start from the TS-JEPA open-source implementation (arXiv:2509.25449, GitHub) - Remove the EMA target-network mechanism - Replace it with Sketched Isotropic Gaussian Regularization (SIGReg) from LeJEPA (arXiv:2511.08544), controlled by a single λ hyperparameter - Keep TS-JEPA's temporal patchwise masking and Transformer encoder unchanged ### Rationale **Why not pure TS-JEPA:** EMA is a heuristic; λ interacts with EMA momentum and learning rate, creating a three-way search space that is hard to navigate with autoresearch. EMA also has no theoretical non-stationarity guarantee. **Why not pure LeJEPA:** The reference implementation targets vision (multi-crop views). Adapting it to temporal patchwise masking requires non-trivial surgery and moves away from open code. TS-JEPA's masking is already the right inductive bias for time series. **Why the hybrid:** SIGReg is architecture-agnostic — it operates on the embedding distribution, not the encoder structure. Swapping EMA for SIGReg is a ~20-line change to TS-JEPA's training loop. The result is: - Time-series native (TS-JEPA masking + patch structure) - Provably collapse-free without heuristics (SIGReg) - Single search axis for autoresearch (λ ∈ [0.01, 1.0]) - Non-stationarity robustness proven formally (arXiv:2602.19373 extends LeJEPA guarantees to non-stationary target distributions — directly relevant to FX) - Explainable to a model validation team: "embeddings are provably optimal for downstream prediction under distributional uncertainty" **Why not MTS-JEPA or Var-JEPA now:** Both lack public code (as of May 2026). MTS-JEPA's multi-resolution objective is the right next hypothesis (see ADR-003). Var-JEPA's ELBO-based UQ is a compelling future direction for CVaR estimation. ### Consequences - Phase 0 (MAE baseline) is unaffected — it precedes the JEPA architecture choice - Issue #3 (TS-JEPA reproduction) is still the right first step; SIGReg is added after reproduction is confirmed - The autoresearch `program.md` primary search axis is λ (SIGReg weight) - Secondary axes: masking block size, patch stride, encoder depth - `model/requirements.txt` must include the SIGReg implementation (≈20 lines, can be vendored directly) --- ## ADR-002 · Data: DUKASCopy hourly G10 FX as primary training data **Date:** 2026-05-28 **Status:** Accepted ### Context Data scale is the most dangerous assumption for any SSL/JEPA approach. Daily FX data (~5,000 samples over 20 years) is insufficient for self-supervised pretraining. Two alternatives were considered: daily public data (yfinance) vs. hourly tick data (DUKASCopy, free, rate-limited HTTP API). ### Decision Use **DUKASCopy hourly OHLCV** as the primary data source. - 10 G10 pairs: EURUSD, GBPUSD, USDJPY, USDCHF, AUDUSD, NZDUSD, USDCAD, EURGBP, EURJPY, GBPJPY - Training window: 2008-01-01 – 2022-12-31 (~175,000 samples per pair) - Validation window: 2023-01-01 – 2023-12-31 (~2,600 samples) - Test window: 2024-01-01 – 2024-12-31 (held out, never seen during development) - Features per bar: log-return, log rolling-20-period HV, VIX (daily interpolated) - Weekend gaps handled explicitly — no interpolation across market close ### Rationale Hourly data gives ~35× more samples than daily. This is the minimum threshold for JEPA-style SSL to show a training signal within 10-minute autoresearch experiments. DUKASCopy is free, reliable, and provides consistent tick-level source data back to 2003. ### Consequences - The Go data pipeline (Issue #2) is the critical path for everything else - Phase 0 MAE baseline trains on the same 2008-2022 window - Daily data (yfinance) may still be used for VIX and rate differentials as auxiliary features --- ## ADR-003 · Research roadmap: Phase structure and JEPA variant progression **Date:** 2026-05-28 **Status:** Accepted ### Decision Three-phase research roadmap: **Phase 0 — SSL feasibility gate (MAE baseline)** Implement a 1D temporal MAE (not JEPA) on EUR/USD hourly 2008-2022. Gate criteria: silhouette > 0.20 on 2023 held-out, MAE > PCA baseline, ±10% over 3 reruns. Purpose: validate that the data and eval harness work before committing to JEPA complexity. If gate fails: follow null result protocol in `specs/phase-0-ssl-feasibility.md`. **Phase 1 — TS-JEPA + SIGReg autoresearch sweep** Primary architecture per ADR-001. Autoresearch loop: `program.md`-driven, 10-min experiments, 50-experiment budget. Primary metric: `val_vol_r2` (linear probe R² on 1-day realized volatility). Gate criteria: `val_vol_r2` > GARCH-implied baseline AND Kupiec p-value > 0.05 on EUR/USD VaR 99%. Kupiec is logged from experiment 1 to verify it co-moves with `val_vol_r2`. **Phase 2 — MTS-JEPA multi-resolution hypothesis** Introduce parallel multi-scale predictive pathways (1h, 8h, 24h context windows) adapted from MTS-JEPA (arXiv:2602.04643). Hypothesis: multi-scale representations improve regime detection (silhouette) and reduce VaR exceedance clustering (Christoffersen test). Prerequisite: Phase 1 gate passed AND MTS-JEPA code available or reproducible from paper. Time-box: if MTS-JEPA code not available within 4 weeks of Phase 2 start, implement multi-resolution masking from scratch using Phase 1 backbone as base. **Phase 3 — Internal bank data (future)** Replace DUKASCopy pipeline with internal tick feed adapter. Fine-tune heads only; backbone frozen or lightly fine-tuned. Out of scope for current PoC cycle. ### Consequences - Issue #5 (Phase 0 MAE) is the unblocked next executable step - Phase 1 autoresearch is blocked until Phase 0 passes its gate - Var-JEPA (ELBO-based UQ) is a named future hypothesis for CVaR estimation in Phase 2+ but not on the critical path --- ## ADR-004 · Evaluation: Go harness + Python training separation **Date:** 2026-05-28 **Status:** Accepted ### Decision Hard separation between training (Python) and evaluation (Go): - **Python** (`model/`): all training, embedding export, model checkpointing - **Go** (`src/eval/`): all evaluation metrics — silhouette, linear probe R², collapse diagnostic, Kupiec/Christoffersen backtests - Interface: Python exports embedding matrices + labels to `experiments/RUNID/` as `.npy` files; Go eval harness reads them and writes `metrics.json` ### Rationale Go evaluation gives deterministic, fast, auditable metric computation with proper unit tests. It decouples the experimental loop from the training framework, making it possible to re-evaluate any past experiment without re-running training. The Go layer also serves as the foundation for the eventual trading desk dashboard. ### Consequences - All acceptance criteria in Issues #4 and #5 are specified in terms of Go eval outputs - `val_vol_r2` (the autoresearch optimization metric) is computed by the Go harness, not inside the Python training loop - Python training loop calls `task eval:probe` as a subprocess after each experiment to get the scalar fed back to autoresearch --- ## ADR-005 · Compute: Blackwell GPU on koala, PyTorch cu130 **Date:** 2026-05-28 **Status:** Accepted ### Decision All GPU training runs on koala (Arch Linux, Blackwell GPU, 12 GB VRAM). PyTorch install: `pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu130` (CUDA 13.0 wheel — required for sm_120 Blackwell support; stable as of May 2026). Driver requirement: NVIDIA R570+, CUDA toolkit 12.8+. Ollama on iguana (Mac Studio M2 Ultra) serves the autoresearch agent LLM via the existing LiteLLM proxy on piguard. Agent calls never hit koala directly. ### Consequences - `model/requirements.txt` must NOT pin torch to a cu124 or earlier wheel - CI (Issue #7) must NOT run GPU tests — CPU-only for unit tests, GPU only via `task experiment:run` on koala - 12 GB VRAM is sufficient for <5M parameter models at batch=64; monitor if autoresearch explores larger architectures