generated from mathias/template-go-web
docs: add DECISIONS.md and rewrite PROJECT.md (#8)
This commit was merged in pull request #8.
This commit is contained in:
+97
-6
@@ -1,13 +1,104 @@
|
||||
# hostexecutor
|
||||
# jepa-fx-risk
|
||||
|
||||
## Identity
|
||||
|
||||
- **Name**: hostexecutor
|
||||
- **Name**: jepa-fx-risk
|
||||
- **Owner**: Mathias
|
||||
- **Client**: personal
|
||||
- **Repo**: gitea.d-ma.be/mathias/hostexecutor
|
||||
- **Status**: active
|
||||
- **Client**: personal research
|
||||
- **Repo**: gitea.d-ma.be/mathias/jepa-fx-risk
|
||||
- **Status**: active — Phase 0
|
||||
|
||||
## Purpose
|
||||
|
||||
Research project: apply JEPA-based self-supervised representation learning to
|
||||
FX risk management for a corporate bank with an internal global FX trading desk.
|
||||
|
||||
Primary tasks: FX volatility forecasting and VaR/CVaR estimation.
|
||||
Target output: internal PoC for the trading desk.
|
||||
|
||||
## Architecture decision (see DECISIONS.md ADR-001)
|
||||
|
||||
**TS-JEPA + SIGReg** — TS-JEPA temporal patchwise architecture (Ennadir et al.,
|
||||
arXiv:2509.25449) with EMA replaced by Sketched Isotropic Gaussian Regularization
|
||||
(SIGReg, Balestriero & LeCun, arXiv:2511.08544). Single search axis: λ ∈ [0.01, 1.0].
|
||||
|
||||
Phase 2 hypothesis: MTS-JEPA multi-resolution objective (arXiv:2602.04643).
|
||||
|
||||
## Stack
|
||||
|
||||
Go + Templ + HTMX + CDN Tailwind. See `~/dev/.context/AGENT.md` for cross-project conventions.
|
||||
**Go** (`src/`): data pipeline (DUKASCopy fetch + hourly processing), evaluation
|
||||
harness (silhouette, linear probe R², collapse diagnostic, Kupiec/Christoffersen),
|
||||
results dashboard (Templ + HTMX + CDN Tailwind).
|
||||
|
||||
**Python** (`model/`): all training and embedding export only. PyTorch cu130
|
||||
(Blackwell sm_120 compatible). No evaluation logic in Python.
|
||||
|
||||
**Infra**: koala (Arch Linux, Blackwell GPU 12 GB VRAM) for training.
|
||||
iguana (Mac Studio M2 Ultra) + LiteLLM on piguard for autoresearch agent LLM.
|
||||
|
||||
## Repository layout
|
||||
|
||||
```
|
||||
jepa-fx-risk/
|
||||
├── src/ # Go — data pipeline + eval harness + dashboard
|
||||
│ ├── data/ # DUKASCopy fetch, hourly processing, validation
|
||||
│ └── eval/ # silhouette, linear probe, collapse, backtest
|
||||
├── model/ # Python — training only
|
||||
│ ├── train.py # TS-JEPA + SIGReg backbone (autoresearch edits this)
|
||||
│ ├── prepare.py # LOCKED — data loading, tokenization, export
|
||||
│ └── requirements.txt
|
||||
├── specs/ # Research specs (one per phase/experiment type)
|
||||
├── experiments/ # Per-run outputs: embeddings, metrics.json, git tag
|
||||
├── results/summaries/ # Human-readable outcome per experiment
|
||||
├── program.md # Autoresearch agenda — researcher edits this
|
||||
├── DECISIONS.md # Architecture Decision Records
|
||||
└── Taskfile.yml # task data:fetch, task experiment:run, task eval:*
|
||||
```
|
||||
|
||||
## Phase structure
|
||||
|
||||
- **Phase 0** (current): MAE baseline on EUR/USD hourly 2008–2022.
|
||||
Gate: silhouette > 0.20, MAE > PCA, ±10% over 3 reruns.
|
||||
- **Phase 1**: TS-JEPA + SIGReg autoresearch sweep, 50 experiments.
|
||||
Gate: val_vol_r2 > GARCH baseline AND Kupiec p > 0.05.
|
||||
- **Phase 2**: MTS-JEPA multi-resolution hypothesis.
|
||||
- **Phase 3**: Internal bank tick/position data (future, out of current scope).
|
||||
|
||||
## Data
|
||||
|
||||
DUKASCopy hourly OHLCV, 10 G10 pairs, 2008–2022 train / 2023 val / 2024 test.
|
||||
Features: log-return, log rolling-20-period HV, VIX (daily interpolated).
|
||||
Weekend gaps handled explicitly. See ADR-002.
|
||||
|
||||
## Key conventions
|
||||
|
||||
- `prepare.py` is LOCKED — never modified by agents or autoresearch
|
||||
- Evaluation metrics are always computed by the Go harness, never in Python
|
||||
- Every experiment gets a git tag: `exp/YYYYMMDD-description`
|
||||
- Null results are recorded explicitly in `results/summaries/` — do not iterate silently
|
||||
- `program.md` is the only file the researcher edits to steer autoresearch
|
||||
- CI runs `task check` (lint + vet + test) only — no GPU, no training
|
||||
|
||||
## Evaluation metrics
|
||||
|
||||
Primary (autoresearch optimizes): `val_vol_r2` — linear probe R² on 1-day
|
||||
realized volatility from frozen embeddings, computed by Go harness.
|
||||
|
||||
Secondary (logged, not optimized): Kupiec p-value (VaR 99% backtest on EUR/USD).
|
||||
Must co-move with val_vol_r2 — checked from experiment 1.
|
||||
|
||||
Diagnostics: silhouette score (regime clustering), PC1/HV correlation (collapse check).
|
||||
|
||||
## Benchmarks to beat (trading desk comparison)
|
||||
|
||||
- GARCH(1,1) — volatility forecasting baseline
|
||||
- Historical Simulation VaR (250-day rolling) — Basel default
|
||||
- EWMA RiskMetrics (λ=0.94)
|
||||
|
||||
## Agent guidance
|
||||
|
||||
Read `DECISIONS.md` before making architecture suggestions.
|
||||
Do not modify `prepare.py` or `src/eval/`.
|
||||
Do not suggest changing the Python/Go separation.
|
||||
Training runs are always manual via `task experiment:run` — never triggered by CI.
|
||||
When implementing, follow Go conventions in `.skills/go-patterns/SKILL.md`.
|
||||
|
||||
+215
@@ -0,0 +1,215 @@
|
||||
# Architecture Decision Records
|
||||
|
||||
This file records significant technical and research decisions for `jepa-fx-risk`.
|
||||
Each record is immutable once merged — append new records rather than editing old ones.
|
||||
Format: ID · Date · Status · Context · Decision · Rationale · Consequences.
|
||||
|
||||
---
|
||||
|
||||
## ADR-001 · Architecture: TS-JEPA + SIGReg as Phase 1 backbone
|
||||
|
||||
**Date:** 2026-05-28
|
||||
**Status:** Accepted
|
||||
**Supersedes:** informal decision to use TS-JEPA standalone (pre-ADR)
|
||||
|
||||
### Context
|
||||
|
||||
Four JEPA variants were evaluated for FX volatility forecasting and VaR/CVaR estimation:
|
||||
|
||||
| Variant | Origin | Key property |
|
||||
|---|---|---|
|
||||
| TS-JEPA | Ennadir et al., Sep 2025 | Time-series native; EMA collapse prevention |
|
||||
| LeJEPA | Balestriero & LeCun, Nov 2025 | Proven optimal embeddings (isotropic Gaussian); SIGReg |
|
||||
| MTS-JEPA | He et al., Feb 2026 | Multi-resolution + codebook; no public code |
|
||||
| Var-JEPA | Multiple, Mar 2026 | ELBO-based UQ; no public code |
|
||||
|
||||
Key constraints: 12 GB VRAM (Blackwell, koala), hourly DUKASCopy data, internal PoC target,
|
||||
autoresearch loop requires a clean single-scalar search space, trading desk requires an
|
||||
explainable theoretical story.
|
||||
|
||||
### Decision
|
||||
|
||||
Use **TS-JEPA architecture with SIGReg replacing EMA** as the Phase 1 backbone.
|
||||
|
||||
Concretely:
|
||||
- Start from the TS-JEPA open-source implementation (arXiv:2509.25449, GitHub)
|
||||
- Remove the EMA target-network mechanism
|
||||
- Replace it with Sketched Isotropic Gaussian Regularization (SIGReg) from LeJEPA
|
||||
(arXiv:2511.08544), controlled by a single λ hyperparameter
|
||||
- Keep TS-JEPA's temporal patchwise masking and Transformer encoder unchanged
|
||||
|
||||
### Rationale
|
||||
|
||||
**Why not pure TS-JEPA:** EMA is a heuristic; λ interacts with EMA momentum and
|
||||
learning rate, creating a three-way search space that is hard to navigate with autoresearch.
|
||||
EMA also has no theoretical non-stationarity guarantee.
|
||||
|
||||
**Why not pure LeJEPA:** The reference implementation targets vision (multi-crop views).
|
||||
Adapting it to temporal patchwise masking requires non-trivial surgery and moves away from
|
||||
open code. TS-JEPA's masking is already the right inductive bias for time series.
|
||||
|
||||
**Why the hybrid:** SIGReg is architecture-agnostic — it operates on the embedding
|
||||
distribution, not the encoder structure. Swapping EMA for SIGReg is a ~20-line change to
|
||||
TS-JEPA's training loop. The result is:
|
||||
- Time-series native (TS-JEPA masking + patch structure)
|
||||
- Provably collapse-free without heuristics (SIGReg)
|
||||
- Single search axis for autoresearch (λ ∈ [0.01, 1.0])
|
||||
- Non-stationarity robustness proven formally (arXiv:2602.19373 extends LeJEPA
|
||||
guarantees to non-stationary target distributions — directly relevant to FX)
|
||||
- Explainable to a model validation team: "embeddings are provably optimal for
|
||||
downstream prediction under distributional uncertainty"
|
||||
|
||||
**Why not MTS-JEPA or Var-JEPA now:** Both lack public code (as of May 2026).
|
||||
MTS-JEPA's multi-resolution objective is the right next hypothesis (see ADR-003).
|
||||
Var-JEPA's ELBO-based UQ is a compelling future direction for CVaR estimation.
|
||||
|
||||
### Consequences
|
||||
|
||||
- Phase 0 (MAE baseline) is unaffected — it precedes the JEPA architecture choice
|
||||
- Issue #3 (TS-JEPA reproduction) is still the right first step; SIGReg is added after
|
||||
reproduction is confirmed
|
||||
- The autoresearch `program.md` primary search axis is λ (SIGReg weight)
|
||||
- Secondary axes: masking block size, patch stride, encoder depth
|
||||
- `model/requirements.txt` must include the SIGReg implementation (≈20 lines,
|
||||
can be vendored directly)
|
||||
|
||||
---
|
||||
|
||||
## ADR-002 · Data: DUKASCopy hourly G10 FX as primary training data
|
||||
|
||||
**Date:** 2026-05-28
|
||||
**Status:** Accepted
|
||||
|
||||
### Context
|
||||
|
||||
Data scale is the most dangerous assumption for any SSL/JEPA approach. Daily FX data
|
||||
(~5,000 samples over 20 years) is insufficient for self-supervised pretraining.
|
||||
Two alternatives were considered: daily public data (yfinance) vs. hourly tick data
|
||||
(DUKASCopy, free, rate-limited HTTP API).
|
||||
|
||||
### Decision
|
||||
|
||||
Use **DUKASCopy hourly OHLCV** as the primary data source.
|
||||
|
||||
- 10 G10 pairs: EURUSD, GBPUSD, USDJPY, USDCHF, AUDUSD, NZDUSD, USDCAD,
|
||||
EURGBP, EURJPY, GBPJPY
|
||||
- Training window: 2008-01-01 – 2022-12-31 (~175,000 samples per pair)
|
||||
- Validation window: 2023-01-01 – 2023-12-31 (~2,600 samples)
|
||||
- Test window: 2024-01-01 – 2024-12-31 (held out, never seen during development)
|
||||
- Features per bar: log-return, log rolling-20-period HV, VIX (daily interpolated)
|
||||
- Weekend gaps handled explicitly — no interpolation across market close
|
||||
|
||||
### Rationale
|
||||
|
||||
Hourly data gives ~35× more samples than daily. This is the minimum threshold for
|
||||
JEPA-style SSL to show a training signal within 10-minute autoresearch experiments.
|
||||
DUKASCopy is free, reliable, and provides consistent tick-level source data back to 2003.
|
||||
|
||||
### Consequences
|
||||
|
||||
- The Go data pipeline (Issue #2) is the critical path for everything else
|
||||
- Phase 0 MAE baseline trains on the same 2008-2022 window
|
||||
- Daily data (yfinance) may still be used for VIX and rate differentials as auxiliary features
|
||||
|
||||
---
|
||||
|
||||
## ADR-003 · Research roadmap: Phase structure and JEPA variant progression
|
||||
|
||||
**Date:** 2026-05-28
|
||||
**Status:** Accepted
|
||||
|
||||
### Decision
|
||||
|
||||
Three-phase research roadmap:
|
||||
|
||||
**Phase 0 — SSL feasibility gate (MAE baseline)**
|
||||
Implement a 1D temporal MAE (not JEPA) on EUR/USD hourly 2008-2022.
|
||||
Gate criteria: silhouette > 0.20 on 2023 held-out, MAE > PCA baseline, ±10% over 3 reruns.
|
||||
Purpose: validate that the data and eval harness work before committing to JEPA complexity.
|
||||
If gate fails: follow null result protocol in `specs/phase-0-ssl-feasibility.md`.
|
||||
|
||||
**Phase 1 — TS-JEPA + SIGReg autoresearch sweep**
|
||||
Primary architecture per ADR-001.
|
||||
Autoresearch loop: `program.md`-driven, 10-min experiments, 50-experiment budget.
|
||||
Primary metric: `val_vol_r2` (linear probe R² on 1-day realized volatility).
|
||||
Gate criteria: `val_vol_r2` > GARCH-implied baseline AND Kupiec p-value > 0.05 on
|
||||
EUR/USD VaR 99%.
|
||||
Kupiec is logged from experiment 1 to verify it co-moves with `val_vol_r2`.
|
||||
|
||||
**Phase 2 — MTS-JEPA multi-resolution hypothesis**
|
||||
Introduce parallel multi-scale predictive pathways (1h, 8h, 24h context windows)
|
||||
adapted from MTS-JEPA (arXiv:2602.04643).
|
||||
Hypothesis: multi-scale representations improve regime detection (silhouette) and
|
||||
reduce VaR exceedance clustering (Christoffersen test).
|
||||
Prerequisite: Phase 1 gate passed AND MTS-JEPA code available or reproducible from paper.
|
||||
Time-box: if MTS-JEPA code not available within 4 weeks of Phase 2 start, implement
|
||||
multi-resolution masking from scratch using Phase 1 backbone as base.
|
||||
|
||||
**Phase 3 — Internal bank data (future)**
|
||||
Replace DUKASCopy pipeline with internal tick feed adapter.
|
||||
Fine-tune heads only; backbone frozen or lightly fine-tuned.
|
||||
Out of scope for current PoC cycle.
|
||||
|
||||
### Consequences
|
||||
|
||||
- Issue #5 (Phase 0 MAE) is the unblocked next executable step
|
||||
- Phase 1 autoresearch is blocked until Phase 0 passes its gate
|
||||
- Var-JEPA (ELBO-based UQ) is a named future hypothesis for CVaR estimation in Phase 2+
|
||||
but not on the critical path
|
||||
|
||||
---
|
||||
|
||||
## ADR-004 · Evaluation: Go harness + Python training separation
|
||||
|
||||
**Date:** 2026-05-28
|
||||
**Status:** Accepted
|
||||
|
||||
### Decision
|
||||
|
||||
Hard separation between training (Python) and evaluation (Go):
|
||||
|
||||
- **Python** (`model/`): all training, embedding export, model checkpointing
|
||||
- **Go** (`src/eval/`): all evaluation metrics — silhouette, linear probe R², collapse
|
||||
diagnostic, Kupiec/Christoffersen backtests
|
||||
- Interface: Python exports embedding matrices + labels to `experiments/RUNID/` as
|
||||
`.npy` files; Go eval harness reads them and writes `metrics.json`
|
||||
|
||||
### Rationale
|
||||
|
||||
Go evaluation gives deterministic, fast, auditable metric computation with proper
|
||||
unit tests. It decouples the experimental loop from the training framework, making
|
||||
it possible to re-evaluate any past experiment without re-running training.
|
||||
The Go layer also serves as the foundation for the eventual trading desk dashboard.
|
||||
|
||||
### Consequences
|
||||
|
||||
- All acceptance criteria in Issues #4 and #5 are specified in terms of Go eval outputs
|
||||
- `val_vol_r2` (the autoresearch optimization metric) is computed by the Go harness,
|
||||
not inside the Python training loop
|
||||
- Python training loop calls `task eval:probe` as a subprocess after each experiment
|
||||
to get the scalar fed back to autoresearch
|
||||
|
||||
---
|
||||
|
||||
## ADR-005 · Compute: Blackwell GPU on koala, PyTorch cu130
|
||||
|
||||
**Date:** 2026-05-28
|
||||
**Status:** Accepted
|
||||
|
||||
### Decision
|
||||
|
||||
All GPU training runs on koala (Arch Linux, Blackwell GPU, 12 GB VRAM).
|
||||
PyTorch install: `pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu130`
|
||||
(CUDA 13.0 wheel — required for sm_120 Blackwell support; stable as of May 2026).
|
||||
Driver requirement: NVIDIA R570+, CUDA toolkit 12.8+.
|
||||
|
||||
Ollama on iguana (Mac Studio M2 Ultra) serves the autoresearch agent LLM via the
|
||||
existing LiteLLM proxy on piguard. Agent calls never hit koala directly.
|
||||
|
||||
### Consequences
|
||||
|
||||
- `model/requirements.txt` must NOT pin torch to a cu124 or earlier wheel
|
||||
- CI (Issue #7) must NOT run GPU tests — CPU-only for unit tests, GPU only via
|
||||
`task experiment:run` on koala
|
||||
- 12 GB VRAM is sufficient for <5M parameter models at batch=64; monitor if
|
||||
autoresearch explores larger architectures
|
||||
Reference in New Issue
Block a user