Wire the metrics package into the live paths and serve it:
- summarizer: per-endpoint latency by model/outcome(success|error|parse_error)/fallback + slog.
- youtube.FetchTranscript: latency by outcome (captions|none|rate_limited) + slog.
- chat: answer latency by model + slog.
- llm usage hook → token counts (prompt|completion) per model, wired in buildSummarizer/buildChat.
- oidc callback: login counter.
- cmdServe: wrap Router in metrics.HTTPMiddleware (request count + latency by bounded
route pattern) and serve /metrics on TAPIR_METRICS_ADDR (default :9090), a SEPARATE
port — never on the public app mux.
BDD: observability.feature scenarios un-pended + mapped. TDD: summarizer wiring tested
black-box via the /metrics scrape; metrics-not-on-public-mux asserted.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
A read-only deeper-dive over a video's already-stored transcript (ADR-021).
Safe by construction: the Service has no VideoSource and no caption-fetch
dependency — only a Completer factory over the existing LiteLLM gateway — so it
cannot reach YouTube or the rate gate. Reuses the summarizer's truncation
discipline (TAPIR_MAX_TRANSCRIPT_CHARS), reporting the cut so the UI can be
honest about a bounded transcript. Model defaults to the summary's model and is
switchable among an offered, local-first list; an un-offered alias is forced
back to the default so chat can never call the gateway with an arbitrary model.
Ephemeral: history is carried per-request, nothing persisted.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>