Pins the concrete endpoints/conventions Tapir depends on so independent sessions don't rediscover them: LiteLLM gateway (http://koala:4000/v1/, sk-local-123, host/name alias format), BYO fallback, brain-mcp HTTP sink, Dex identity, ESO+1Password secrets, hosts, and Flux GitOps deploy location. Verified facts recorded as such; unconfirmed items (brain-mcp URL, exact secret-ref naming, post-relocation LiteLLM location, summarization model alias) explicitly flagged 'confirm' rather than invented.
4.7 KiB
4.7 KiB
Tapir — Homelab Integration Facts
The concrete endpoints, conventions, and identifiers Tapir depends on, so an independent session doesn't have to rediscover them. Verify anything marked "confirm" before relying on it — endpoints and aliases drift, and this file is a snapshot (2026-06-02), not a live source.
Local AI (the Primary in llm.Router)
- LiteLLM gateway:
http://koala:4000/v1/(OpenAI-compatible). In-cluster callers use this; there is also a NodePort (31234) for off-cluster access — confirm which Tapir should use based on where it runs. - Auth key:
sk-local-123(the local LiteLLM gateway key). Treat as a secret via the SecretStore port; do not hardcode in committed code. - Model alias format:
host/name, e.g.koala/qwen3-coder-30b,koala/phi4-mini,iguana/devstral,iguana/deepseek-r1-14b. Not theollama/prefix form. - Which alias for summarization: NOT yet decided. Tapir summarizes transcript text, so a capable general/instruct model on koala or iguana is the candidate — pick during the build and record the choice (an ADR if it's load-bearing). Do not assume a coder alias is right for prose summarization.
This maps directly onto the copied llm package: Client is the OpenAI-compatible caller,
Router.Primary points at this gateway with a chosen alias, Router.Fallback is the user's BYO.
BYO AI (the Fallback in llm.Router)
- Per-user, opt-in: Anthropic / OpenAI / Gemini. All reachable as OpenAI-compatible or via a thin
adapter. The user's key is stored only as a secret reference (data-model
AI_CREDENTIAL). - Used only when Primary fails and the user has configured a provider (see
docs/use-cases/ai_routing.feature). With no BYO, work is queued for retry — content is never sent externally.
Brain sink (optional delivery target)
- Reached via brain-mcp over HTTP, calling the
brain_ingesttool (ADR-005). Do not use the filesystembrainpackage fromhyperguild/ingestion. - brain-mcp base URL: confirm — not pinned in this snapshot. The homelab convention is
*-mcp.d-ma.beingress hostnames (e.g.git-mcp.d-ma.be,brain-mcp.d-ma.beis the likely form), but verify before wiring. - brain distinguishes
knowledge/(session-derived) fromwiki/(stable, manually promoted). Tapir summaries are session-derived signal →knowledge/-style ingestion.brain_ingestruns the LLM extraction pipeline; there is a raw-ingest path (brain_ingest_raw) and a reliablebrain_writefallback when the local extraction model is unavailable — relevant if ingestion reliability becomes an issue.
Auth / identity (Stage 1+, dormant at Stage 0)
- Dex (
auth.d-ma.be) is the homelab OIDC provider. When Tapir needs real user identity (Stage 1, Future B), it authenticates via Dex — not a new auth system (ADR-002). - If Tapir ever exposes an MCP surface, use the
mcp-chassisshared Go library (gitea.d-ma.be/mathias/mcp-chassis) for the Dex-JWT validator + Bearer middleware + RFC 9728 metadata, rather than hand-rolling it.
Secrets (OAuth tokens + BYO keys)
- Convention: ESO + 1Password (the homelab's External Secrets Operator against a 1Password
vault). Tables store only an opaque
*_secret_ref; the secret material lives in the vault and is surfaced to the workload via ESO. - Exact ref format / vault item naming for Tapir: confirm / decide during build. Follow
the pattern existing services use (e.g. how
gitea-mcpreferencesGITEA_MCP_DEFAULT_TOKEN) rather than inventing a new scheme.
Hosts (for reference)
- koala — Arch, RTX 5070, k3s single-node, Gitea, GPU inference. Where Tapir most likely runs and where the local models live.
- iguana — Mac Studio M2 Ultra, ollama + mlx-whisper, also serves models via the gateway.
- piguard — Pi, NPM perimeter (public HTTPS edge). Note: the May-2026 architecture review has LiteLLM relocating from piguard into k3s/ai-stack on koala — confirm current location if latency or endpoint matters.
Deployment / GitOps (when Tapir reaches deploy)
- The homelab is Flux GitOps: manifests in
mathias/infraunderk3s/, Flux watchesmain. - Tapir's k3s manifests will live in
infra/k3s/apps/tapir/(by convention) — not in this repo. This repo is the application;infrais the deployment source of truth. - Multi-tenancy primitives (NetworkPolicy per namespace, Kyverno, postgres role-per-tenant,
tenant=label) exist in the architecture review (SC7/P6) and apply at Stage 1 — not Stage 0.
Snapshot date 2026-06-02. Items marked confirm were not verified to a pinned source at snapshot time — check brain or the live cluster before depending on them.