Files
tapir/docs/homelab-integration.md
T
mathias 32e06843f1 docs: add homelab integration facts reference
Pins the concrete endpoints/conventions Tapir depends on so independent sessions
don't rediscover them: LiteLLM gateway (http://koala:4000/v1/, sk-local-123, host/name
alias format), BYO fallback, brain-mcp HTTP sink, Dex identity, ESO+1Password secrets,
hosts, and Flux GitOps deploy location. Verified facts recorded as such; unconfirmed
items (brain-mcp URL, exact secret-ref naming, post-relocation LiteLLM location,
summarization model alias) explicitly flagged 'confirm' rather than invented.
2026-06-02 10:37:55 +00:00

4.7 KiB

Tapir — Homelab Integration Facts

The concrete endpoints, conventions, and identifiers Tapir depends on, so an independent session doesn't have to rediscover them. Verify anything marked "confirm" before relying on it — endpoints and aliases drift, and this file is a snapshot (2026-06-02), not a live source.

Local AI (the Primary in llm.Router)

  • LiteLLM gateway: http://koala:4000/v1/ (OpenAI-compatible). In-cluster callers use this; there is also a NodePort (31234) for off-cluster access — confirm which Tapir should use based on where it runs.
  • Auth key: sk-local-123 (the local LiteLLM gateway key). Treat as a secret via the SecretStore port; do not hardcode in committed code.
  • Model alias format: host/name, e.g. koala/qwen3-coder-30b, koala/phi4-mini, iguana/devstral, iguana/deepseek-r1-14b. Not the ollama/ prefix form.
  • Which alias for summarization: NOT yet decided. Tapir summarizes transcript text, so a capable general/instruct model on koala or iguana is the candidate — pick during the build and record the choice (an ADR if it's load-bearing). Do not assume a coder alias is right for prose summarization.

This maps directly onto the copied llm package: Client is the OpenAI-compatible caller, Router.Primary points at this gateway with a chosen alias, Router.Fallback is the user's BYO.

BYO AI (the Fallback in llm.Router)

  • Per-user, opt-in: Anthropic / OpenAI / Gemini. All reachable as OpenAI-compatible or via a thin adapter. The user's key is stored only as a secret reference (data-model AI_CREDENTIAL).
  • Used only when Primary fails and the user has configured a provider (see docs/use-cases/ai_routing.feature). With no BYO, work is queued for retry — content is never sent externally.

Brain sink (optional delivery target)

  • Reached via brain-mcp over HTTP, calling the brain_ingest tool (ADR-005). Do not use the filesystem brain package from hyperguild/ingestion.
  • brain-mcp base URL: confirm — not pinned in this snapshot. The homelab convention is *-mcp.d-ma.be ingress hostnames (e.g. git-mcp.d-ma.be, brain-mcp.d-ma.be is the likely form), but verify before wiring.
  • brain distinguishes knowledge/ (session-derived) from wiki/ (stable, manually promoted). Tapir summaries are session-derived signal → knowledge/-style ingestion. brain_ingest runs the LLM extraction pipeline; there is a raw-ingest path (brain_ingest_raw) and a reliable brain_write fallback when the local extraction model is unavailable — relevant if ingestion reliability becomes an issue.

Auth / identity (Stage 1+, dormant at Stage 0)

  • Dex (auth.d-ma.be) is the homelab OIDC provider. When Tapir needs real user identity (Stage 1, Future B), it authenticates via Dex — not a new auth system (ADR-002).
  • If Tapir ever exposes an MCP surface, use the mcp-chassis shared Go library (gitea.d-ma.be/mathias/mcp-chassis) for the Dex-JWT validator + Bearer middleware + RFC 9728 metadata, rather than hand-rolling it.

Secrets (OAuth tokens + BYO keys)

  • Convention: ESO + 1Password (the homelab's External Secrets Operator against a 1Password vault). Tables store only an opaque *_secret_ref; the secret material lives in the vault and is surfaced to the workload via ESO.
  • Exact ref format / vault item naming for Tapir: confirm / decide during build. Follow the pattern existing services use (e.g. how gitea-mcp references GITEA_MCP_DEFAULT_TOKEN) rather than inventing a new scheme.

Hosts (for reference)

  • koala — Arch, RTX 5070, k3s single-node, Gitea, GPU inference. Where Tapir most likely runs and where the local models live.
  • iguana — Mac Studio M2 Ultra, ollama + mlx-whisper, also serves models via the gateway.
  • piguard — Pi, NPM perimeter (public HTTPS edge). Note: the May-2026 architecture review has LiteLLM relocating from piguard into k3s/ai-stack on koala — confirm current location if latency or endpoint matters.

Deployment / GitOps (when Tapir reaches deploy)

  • The homelab is Flux GitOps: manifests in mathias/infra under k3s/, Flux watches main.
  • Tapir's k3s manifests will live in infra/k3s/apps/tapir/ (by convention) — not in this repo. This repo is the application; infra is the deployment source of truth.
  • Multi-tenancy primitives (NetworkPolicy per namespace, Kyverno, postgres role-per-tenant, tenant= label) exist in the architecture review (SC7/P6) and apply at Stage 1 — not Stage 0.

Snapshot date 2026-06-02. Items marked confirm were not verified to a pinned source at snapshot time — check brain or the live cluster before depending on them.