diff --git a/docs/homelab-integration.md b/docs/homelab-integration.md new file mode 100644 index 0000000..3dce04e --- /dev/null +++ b/docs/homelab-integration.md @@ -0,0 +1,82 @@ +# Tapir — Homelab Integration Facts + +The concrete endpoints, conventions, and identifiers Tapir depends on, so an independent +session doesn't have to rediscover them. **Verify anything marked "confirm" before relying on +it** — endpoints and aliases drift, and this file is a snapshot (2026-06-02), not a live source. + +## Local AI (the Primary in `llm.Router`) + +- **LiteLLM gateway:** `http://koala:4000/v1/` (OpenAI-compatible). In-cluster callers use this; + there is also a NodePort (`31234`) for off-cluster access — **confirm** which Tapir should use + based on where it runs. +- **Auth key:** `sk-local-123` (the local LiteLLM gateway key). Treat as a secret via the + SecretStore port; do not hardcode in committed code. +- **Model alias format:** `host/name`, e.g. `koala/qwen3-coder-30b`, `koala/phi4-mini`, + `iguana/devstral`, `iguana/deepseek-r1-14b`. **Not** the `ollama/` prefix form. +- **Which alias for summarization:** NOT yet decided. Tapir summarizes transcript text, so a + capable general/instruct model on koala or iguana is the candidate — pick during the build and + record the choice (an ADR if it's load-bearing). Do not assume a coder alias is right for prose + summarization. + +This maps directly onto the copied `llm` package: `Client` is the OpenAI-compatible caller, +`Router.Primary` points at this gateway with a chosen alias, `Router.Fallback` is the user's BYO. + +## BYO AI (the Fallback in `llm.Router`) + +- Per-user, opt-in: Anthropic / OpenAI / Gemini. All reachable as OpenAI-compatible or via a thin + adapter. The user's key is stored only as a secret reference (data-model `AI_CREDENTIAL`). +- Used **only** when Primary fails and the user has configured a provider (see + `docs/use-cases/ai_routing.feature`). With no BYO, work is queued for retry — content is never + sent externally. + +## Brain sink (optional delivery target) + +- Reached via **brain-mcp** over HTTP, calling the `brain_ingest` tool (ADR-005). Do **not** use + the filesystem `brain` package from `hyperguild/ingestion`. +- **brain-mcp base URL:** **confirm** — not pinned in this snapshot. The homelab convention is + `*-mcp.d-ma.be` ingress hostnames (e.g. `git-mcp.d-ma.be`, `brain-mcp.d-ma.be` is the likely + form), but verify before wiring. +- brain distinguishes `knowledge/` (session-derived) from `wiki/` (stable, manually promoted). + Tapir summaries are session-derived signal → `knowledge/`-style ingestion. `brain_ingest` runs + the LLM extraction pipeline; there is a raw-ingest path (`brain_ingest_raw`) and a reliable + `brain_write` fallback when the local extraction model is unavailable — relevant if ingestion + reliability becomes an issue. + +## Auth / identity (Stage 1+, dormant at Stage 0) + +- **Dex** (`auth.d-ma.be`) is the homelab OIDC provider. When Tapir needs real user identity + (Stage 1, Future B), it authenticates via Dex — not a new auth system (ADR-002). +- If Tapir ever exposes an MCP surface, use the **`mcp-chassis`** shared Go library + (`gitea.d-ma.be/mathias/mcp-chassis`) for the Dex-JWT validator + Bearer middleware + RFC 9728 + metadata, rather than hand-rolling it. + +## Secrets (OAuth tokens + BYO keys) + +- Convention: **ESO + 1Password** (the homelab's External Secrets Operator against a 1Password + vault). Tables store only an opaque `*_secret_ref`; the secret material lives in the vault and + is surfaced to the workload via ESO. +- **Exact ref format / vault item naming for Tapir:** **confirm / decide during build.** Follow + the pattern existing services use (e.g. how `gitea-mcp` references `GITEA_MCP_DEFAULT_TOKEN`) + rather than inventing a new scheme. + +## Hosts (for reference) + +- **koala** — Arch, RTX 5070, k3s single-node, Gitea, GPU inference. Where Tapir most likely runs + and where the local models live. +- **iguana** — Mac Studio M2 Ultra, ollama + mlx-whisper, also serves models via the gateway. +- **piguard** — Pi, NPM perimeter (public HTTPS edge). Note: the May-2026 architecture review has + LiteLLM relocating from piguard into k3s/ai-stack on koala — **confirm current location** if + latency or endpoint matters. + +## Deployment / GitOps (when Tapir reaches deploy) + +- The homelab is **Flux GitOps**: manifests in `mathias/infra` under `k3s/`, Flux watches `main`. +- Tapir's k3s manifests will live in `infra/k3s/apps/tapir/` (by convention) — not in this repo. + This repo is the application; `infra` is the deployment source of truth. +- Multi-tenancy primitives (NetworkPolicy per namespace, Kyverno, postgres role-per-tenant, + `tenant=` label) exist in the architecture review (SC7/P6) and apply at Stage 1 — not Stage 0. + +--- + +_Snapshot date 2026-06-02. Items marked **confirm** were not verified to a pinned source at +snapshot time — check brain or the live cluster before depending on them._