Compare commits

...
23 Commits
Author SHA1 Message Date
mathiasandClaude Opus 4.8 cd461b95f8 feat(observability): instrument AI + HTTP paths, serve /metrics on a side port (ADR-030, #15)
CI / Build & Import (push) Successful in 11s
CI / Lint / Test / Vet (push) Successful in 10s
Wire the metrics package into the live paths and serve it:
- summarizer: per-endpoint latency by model/outcome(success|error|parse_error)/fallback + slog.
- youtube.FetchTranscript: latency by outcome (captions|none|rate_limited) + slog.
- chat: answer latency by model + slog.
- llm usage hook → token counts (prompt|completion) per model, wired in buildSummarizer/buildChat.
- oidc callback: login counter.
- cmdServe: wrap Router in metrics.HTTPMiddleware (request count + latency by bounded
  route pattern) and serve /metrics on TAPIR_METRICS_ADDR (default :9090), a SEPARATE
  port — never on the public app mux.

BDD: observability.feature scenarios un-pended + mapped. TDD: summarizer wiring tested
black-box via the /metrics scrape; metrics-not-on-public-mux asserted.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-12 08:49:51 +02:00
mathiasandClaude Opus 4.8 9c2a04406b feat(llm): usage hook to surface token counts (ADR-030, #15)
WithUsageHook callback fires with model + prompt/completion tokens parsed from the
response usage block. Keeps the copied stdlib-only llm package decoupled from
metrics (ADR-004) — the caller wires it to internal/metrics. TDD covered.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-12 08:36:02 +02:00
mathiasandClaude Opus 4.8 a5a8cf6f6d feat(metrics): Prometheus collectors + typed API + HTTP middleware (ADR-030, #15)
New internal/metrics package: AI metrics (summarize/caption/chat latency, LLM
tokens), session metrics (requests+latency by bounded route pattern, logins), and
a /metrics Handler. Adapters call a typed API; never touch prometheus types.
Dep: github.com/prometheus/client_golang (standard Go client; cluster runs
prometheus-operator). TDD: collectors + middleware route-pattern cardinality
covered.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-12 08:32:41 +02:00
mathias 64d11af9ef docs(bdd): observability.feature scenarios (@pending until TDD, #15) 2026-06-12 08:28:46 +02:00
mathias b590d2708d docs(adr): ADR-030 observability — slog + Prometheus metrics (issue #15) 2026-06-12 08:27:56 +02:00
mathiasandClaude Opus 4.8 7314895ec4 fix(auth): stateless session cookie — stop logging users out on deploy (ADR-029)
CI / Lint / Test / Vet (push) Successful in 10s
CI / Build & Import (push) Successful in 11s
Pilot feedback: lots of re-logging-in on iPhone. Three causes: sessions lived in
an in-memory map (wiped on every pod restart/deploy), a 1h TTL (idle >1h forced
re-login on a check-back-tomorrow reader), and a session cookie with no Max-Age
(dropped on Safari close). Each re-login is the full IdP redirect dance.

Make sessions stateless: identity + absolute expiry live inside the existing
HMAC-signed cookie (no server table), TTL 1h → 30 days sliding, cookie now
persistent (Max-Age). Survives restarts (test: a cookie from one instance is
accepted by a fresh instance with the same secret), browser-close, and idle.
Trade: no server-side revocation — logout clears the cookie client-side; rotating
tapir-session-secret is the global logout lever. Accepted for the Stage-0 reader.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-11 23:07:13 +02:00
mathiasandClaude Opus 4.8 36dd182fb5 feat(report): Stage-0 usage gate counts from a baseline date (default 2026-06-11)
CI / Build & Import (push) Successful in 11s
CI / Lint / Test / Vet (push) Successful in 10s
The return-usage gate (ADR-016) counted distinct active weeks over ALL history,
so pre-launch noise — testing churn and the period the pilot sat blocked on zero
summaries — would inflate the signal. Add a baseline: ActiveWeeks(ctx, since)
filters login_events + summary_actions to seen_at/acted_at >= since. The report
command sets it to TAPIR_USAGE_GATE_START (YYYY-MM-DD, default 2026-06-11 — the
morning the pilot was unblocked) and prints the baseline. The gate now measures
whether users RETURN once it genuinely works.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-11 22:52:15 +02:00
mathias 40808f2d4b docs(onboard): BDD scenarios + architecture for the ADR-028 burst
CI / Lint / Test / Vet (push) Successful in 10s
CI / Build & Import (push) Successful in 10s
- connect_account.feature: refine the burst scenario to "best recent" (likely-good
  selection, skips too-short/too-long) and add a scenario for the stronger model.
- scenario_coverage: remap to TestOnboardBurstVideoIDs + add the model scenario
  (the old NewestUnsummarizedVideoIDs test was removed).
- architecture.md: new "Connect-time onboarding burst" subsection — junk-avoiding
  selection over persisted duration, the burst-only stronger-model chain, and why
  has-captions/cached-first are not selection signals.
2026-06-11 19:54:29 +02:00
mathias 9232f49555 feat(onboard): burst picks likely-good videos, summarizes with stronger model (ADR-028)
Wire the onboarding burst to its quality-aware selection and a stronger model:
- main.go onboard uses OnboardBurstVideoIDs (junk-avoiding) instead of pure
  newest-first, bounded by MinVideoSeconds / OnboardMaxVideoSeconds.
- buildBurstProcessor builds a burst-only summarizer chain led by the onboard
  model (burstChainModels: onboard -> standard ADR-022 chain, deduped, NDA lever
  intact), over the same store/cache/sink. Collapses onto the shared Processor
  when the onboard model is empty/equal-to-primary or config is incomplete.
- summarizerEndpoint + newYouTubeSource extracted so standard and burst wiring
  share one definition.
- Remove now-superseded NewestUnsummarizedVideoIDs: OnboardBurstVideoIDs(.,0,0)
  is identical pure-newest behaviour and its test covers RLS + ordering.
2026-06-11 19:54:29 +02:00
mathias 582c1a2065 feat(store): OnboardBurstVideoIDs — junk-avoiding burst selection (ADR-028)
Newest-first but quality-aware: excludes a video when its duration is KNOWN and
outside [minSeconds, maxSeconds], dropping Shorts and multi-hour livestream VODs
that waste a scarce caption fetch on a poor first impression. NULL/unknown
duration is kept (degrade-open) but ranked after known-good rows. 0/0 bounds
disable the filter (pure newest-first, the reversibility lever). RLS-scoped.
NewestUnsummarizedVideoIDs is left in place for callers that want pure-newest.
2026-06-11 19:54:29 +02:00
mathias 9f0d8cf198 feat(discovery): persist video duration_s instead of discarding it (ADR-028)
ADR-023's filterLowValue already fetches each candidate's duration via the
cheap videos.list quota call to drop Shorts/live, then threw it away — the
videos.duration_s column (migration 001) was never written. Carry it onto the
kept domain.Video and have UpsertVideo persist it, COALESCE-preserving a known
value so an unknown (0) re-upsert never clobbers it (the channel_title backfill
stance, migration 014). This is the enabling change for length-aware burst
selection. No new migration — the column already exists.
2026-06-11 19:54:29 +02:00
mathias d21077303d feat(config): add OnboardSummarizerModel + OnboardMaxVideoSeconds (ADR-028)
Two knobs for the onboarding-burst quality work:
- TAPIR_ONBOARD_SUMMARIZER_MODEL (default iguana/gemma4-26b): the stronger model
  the burst leads its chain with; empty collapses the burst onto the shared
  processor (lookupOr, so explicit-empty disables).
- TAPIR_ONBOARD_MAX_VIDEO_SECONDS (default 14400/4h): upper duration bound for
  burst picks; 0 disables, negative clamps to 0.

Table-driven tests cover defaults, explicit, disable, and invalid input.
2026-06-11 19:54:29 +02:00
mathias 9b2ee2e765 docs: spec onboarding-wow-burst + ADR-028 (better picks, stronger model)
Phase-1 live-DB investigation found the connect-time burst (ADR-018) fires
but delivers a weak first impression: pure newest-first selection picks junk
(a livestream + regional news for pilot user Jonte), and all burst summaries
run on the weak koala/phi4-mini instead of the validated iguana/gemma4-26b.
Cached-first was investigated and rejected — ~3% cross-user overlap, 0
cached-and-unsummarized, newest-20 all uncached (newest-first and cached-first
are structurally incompatible).

ADR-028 records: persist duration_s at discovery (ADR-023 already fetches it,
just stops discarding), junk-avoiding burst selection (drop known too-short/
too-long), and a stronger burst-only summarizer chain. Not a throughput change.
2026-06-11 19:54:29 +02:00
mathias b8fbc5a805 docs: spec first-session wow — verify & improve onboarding summary burst (investigate-first)
CI / Lint / Test / Vet (push) Successful in 11s
CI / Build & Import (push) Successful in 10s
The concern is new-user first-contact: see some GOOD summaries fast or they
don't return. Reframed as curation/latency for ~3 videos, NOT a 429/throughput
problem (3 fetches is nowhere near the wall). Phase 1 (report-and-stop) verifies
whether the existing cap-3 onboarding burst even fires today, what it delivers,
and — critically — how much transcript-cache overlap exists between users (drives
the blend). Phase 2 levers: cached-transcript-first (instant, zero-fetch),
likely-good selection (has-captions/good-length, not just newest), and optionally
the stronger model for the burst's few summaries. Blend deferred to the
maintainer post-Phase-1. Explicitly NOT bulk-fetch, NOT credentials (ADR-010/026
dead end), NOT a client extension.
2026-06-11 15:36:09 +00:00
mathiasandClaude Opus 4.8 19ca4282a8 refactor(web): dock chat inline below the summary, integrated view (ADR-027)
CI / Lint / Test / Vet (push) Successful in 10s
CI / Build & Import (push) Successful in 10s
The chat was a separate page — opening it left the summary behind, the very
context you're asking about. Now the chat docks open IN PLACE below the summary:
"Dig deeper" reveals the chat section (HTMX, outerHTML over the closed dock) so
the summary stays on screen above it; no navigation. The no-JS fallback renders
the full summary AND the open chat on one page (the same integrated view), so
progressive enhancement holds.

Extracts a shared summaryBody templ so the detail page and the chat page render
one identical summary, not two divergent ones. GET /v/{id}/chat returns just the
open chat section as a fragment for the inline reveal, or the full summary+chat
page for a no-JS navigation; POST swaps the panel inline or re-renders the whole
page. Tests assert summary+chat coexist on the page and the reveal is a fragment.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-11 14:02:39 +02:00
mathiasandClaude Opus 4.8 71df696448 feat(web): per-video chat over the stored transcript (ADR-027)
CI / Lint / Test / Vet (push) Successful in 10s
CI / Build & Import (push) Successful in 10s
A deeper-dive chat entered from the summary view: ask questions about a video
against its already-stored transcript (ADR-021), no caption fetch, ever.

Safety by construction — the load-bearing property. The chat handlers reach the
chat.Service only after reading the SHARED stored transcript via Store.GetTranscript
(a pure DB read); the service holds no VideoSource. So an enabled chat cannot
trigger a caption fetch, touch the rate gate, or reach YouTube. A video with no
stored transcript gets an honest "not available" — no fetch, no model call. The
key web test wires the summarize/fetch collaborators as tripwires that fail the
test if chat ever routes into them, and asserts the model answered from the
stored text.

Model defaults to the summary's own model and is switchable among the ADR-022
chain (phi4-mini → gemma4-26b → mistral-small); switching re-runs against the
same transcript — deliberate model-comparison instrumentation. The cloud model
is absent from the switcher when TAPIR_CLOUD_FALLBACK_MODEL="" (the local-first /
NDA lever), honoured the same way the summarizer honours it. Reuses the existing
LiteLLM gateway client (a chat is a different call, not a new integration) and
the TAPIR_MAX_TRANSCRIPT_CHARS truncation, surfacing an honest bounded-context
note when a long transcript is cut.

Ephemeral v1: the multi-turn conversation rides in hidden request fields; no
table, no migration, nothing persisted. Entry is RLS-scoped through
GetSummaryByVideo, so chat is reachable only from the user's own summary view.
Show-source verification and on-demand fetch are deferred (ADR-027).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-11 09:26:58 +02:00
mathiasandClaude Opus 4.8 cc3cda4ab8 feat(chat): stored-transcript QA service (ADR-027)
A read-only deeper-dive over a video's already-stored transcript (ADR-021).
Safe by construction: the Service has no VideoSource and no caption-fetch
dependency — only a Completer factory over the existing LiteLLM gateway — so it
cannot reach YouTube or the rate gate. Reuses the summarizer's truncation
discipline (TAPIR_MAX_TRANSCRIPT_CHARS), reporting the cut so the UI can be
honest about a bounded transcript. Model defaults to the summary's model and is
switchable among an offered, local-first list; an un-offered alias is forced
back to the default so chat can never call the gateway with an arbitrary model.
Ephemeral: history is carried per-request, nothing persisted.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-11 09:16:45 +02:00
mathiasandClaude Opus 4.8 69a49bc603 docs(decisions): add ADR-027 — chat with a video's stored transcript
Records the deeper-dive chat decision before its build: stored-transcript-only
(safe by construction — no caption fetch, no rate gate, no YouTube), entered
from the summary view, model defaulting to the summary's model and switchable
among the ADR-022 chain. Ephemeral v1; show-source verification deferred to v2.
Inserted before "Rejected alternatives" so the decision precedes the feature.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-11 09:13:31 +02:00
mathias 137804b0b1 docs: spec chat-with-transcript (ADR-027)
CI / Lint / Test / Vet (push) Successful in 11s
CI / Build & Import (push) Successful in 10s
Per-video chat against the ADR-021 stored transcript, entered from the summary
view, born from observed demand (maintainer read real summaries, some made him
want to dig deeper). HARD constraint: stored-transcript-only — never fetches
captions, never touches the rate gate or YouTube, safe by construction. Default
model = the summary's model, user-switchable among the ADR-022 chain models
(doubles as model-comparison instrumentation). Ephemeral v1 (no persisted
history); chat-only/trust-the-model with show-source verification recorded as the
natural v2. ADR-027 to be appended to DECISIONS.md as the first commit.
2026-06-11 06:38:16 +00:00
mathiasandClaude Opus 4.8 a9be5f285b feat(summarizer): move local fallback off koala to iguana/gemma4-26b
CI / Build & Import (push) Successful in 10s
CI / Lint / Test / Vet (push) Successful in 10s
koala now carries other GPU loads, so the first fallback should not run there.
Change the default chain to koala/phi4-mini → iguana/gemma4-26b → berget/mistral-small:
the local fallback now runs on iguana (M2 Ultra headroom, different host = different
egress IP for the rare fallback fetch). gemma4-26b is the brain-validated homelab
general-purpose model (agentsquad H2/H3 executor) and returned valid summary JSON on
the real prompt in a smoke test (~37s incl. cold-load — fine for a fallback path).

Pure config default (TAPIR_FALLBACK_MODEL); chain mechanism (ADR-022) unchanged.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-11 07:54:56 +02:00
mathiasandClaude Opus 4.8 fe56e2fe01 test: make embedded-postgres per-process so concurrent CI runs don't collide
CI / Lint / Test / Vet (push) Successful in 9s
CI / Build & Import (push) Successful in 10s
A push to main and its version tag fire two CI runs for the same commit. Both ran
`go test ./...`, which starts embedded-postgres on a FIXED port (54329/54330) and
a shared data dir. -p 1 serialises packages WITHIN a run, not across two
concurrent runs — so when the two runs overlapped they collided on the port/data
dir and BOTH failed the Lint/Test job (no image built). Prior commits passed only
because their two runs happened not to overlap.

Derive the port and runtime/data dirs from the PID; share only CachePath so the
PG archive downloads once. Proven: two concurrent `go test` of the store package
now both pass. Unblocks the v0.21.0 (Pillar A) build.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-11 07:26:33 +02:00
mathiasandClaude Opus 4.8 beeb5bc31b feat(gate): foreground caption fetches take priority over the background sweep (ADR-026)
CI / Lint / Test / Vet (push) Failing after 9s
CI / Build & Import (push) Has been skipped
A user waiting on a Summarize click shared the per-IP caption gate equally with
the background firehose, so on a busy IP the click was slow or 429'd. Add a
context-marked priority lane: the web path (engineProcessor.ProcessVideo) marks
its context foreground; the gate serves foreground immediately while background
fetches yield until no foreground is pending. Threaded via a context value (no
new signatures) + a process-wide foregroundPending counter. Clicks are rare, so
the background barely loses throughput; the waiting human gets the cleaner slot.

Drops the credentials probe: ADR-010 and captions.go already settle it — the
timedtext/InnerTube path rejects authenticated requests and the OAuth token does
not authenticate it anyway, so auth cannot help and can hurt. Documented in
ADR-026 rather than built.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-10 22:10:15 +02:00
mathiasandClaude Opus 4.8 09eb31d1fe fix(scheduler): derive rotation offset from wall-clock, not a reset-on-restart counter
CI / Lint / Test / Vet (push) Successful in 11s
CI / Build & Import (push) Successful in 14s
The lead-user rotation used an in-memory pass counter reset to 0 on every pod
restart, so the first-listed user always re-took the lead after a restart — a
deploy-heavy session re-starved the last user (the pilot stalled at 2 summaries
because each deploy reset his every-other-pass lead before the 2h tick fired).
Derive the offset from wall-clock (floor(now/interval)) so it advances with real
time and is identical across restarts: rotation stays fair however often the pod
bounces.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-10 21:59:12 +02:00
49 changed files with 3838 additions and 597 deletions
+266 -4
View File
@@ -827,10 +827,18 @@ The fix is resilience around it, not replacing it.
**Decision.** **Decision.**
1. **Ordered endpoint chain (`summarizer.NewChain`).** Endpoints are tried in order; the first to 1. **Ordered endpoint chain (`summarizer.NewChain`).** Endpoints are tried in order; the first to
return a *parseable* summary wins. Default chain: return a *parseable* summary wins. Default chain:
`koala/phi4-mini` (primary, local) → `koala/phi4-14b` (fallback, local) → `koala/phi4-mini` (primary, local) → `iguana/gemma4-26b` (fallback, local on a
`berget/mistral-small` (worst-case, external). All three are reached through the **one** LiteLLM *different host*) → `berget/mistral-small` (worst-case, external). All three are reached through
gateway by alias — the gateway already fronts both llama-swap and berget — so a fallback is a the **one** LiteLLM gateway by alias — the gateway already fronts both llama-swap and berget — so
different alias, not a second client config. a fallback is a different alias, not a second client config.
**Update 2026-06-11:** the local fallback moved from `koala/phi4-14b` to `iguana/gemma4-26b`.
koala now carries other GPU loads, so keeping the fallback on koala competed with them; iguana
(M2 Ultra) has the headroom, and a different host is also a different egress IP for the rare
fallback fetch. `gemma4-26b` is the brain-validated homelab general-purpose model (agentsquad
H2/H3 executor) and returned valid summary JSON on the real prompt in a smoke test
(~37s incl. cold-load — fine for a path hit only when the fast primary fails). Pure config:
`TAPIR_FALLBACK_MODEL`.
2. **A parse failure advances the chain, same as a transport error.** "Reliably summarized" means 2. **A parse failure advances the chain, same as a transport error.** "Reliably summarized" means
*parseable summary returned*, not *HTTP 200*. This is the behaviour the old Primary→Fallback *parseable summary returned*, not *HTTP 200*. This is the behaviour the old Primary→Fallback
shape missed. shape missed.
@@ -958,6 +966,259 @@ schema change (reuses `transcript_status` from migration 007).
--- ---
## ADR-026 — Foreground caption fetches take priority; the credentials probe is dead
**Status:** Accepted (2026-06-10). **Pillar A of the manual-mode UX work** (Pillar B was
ADR-025). Builds on ADR-014 (the shared per-IP gate).
**Context.** Every caption fetch — the background sweep and the web click-path — shared one
process-wide rate gate equally. So a user waiting on a "Summarize" click competed with the
firehose for both pacing and the scarce pre-429 window; on a busy IP the click was slow or
429'd while the background churned.
**Decision.** A context-marked priority lane. The web path
(`engineProcessor.ProcessVideo`) wraps its context with `ForegroundContext`; the gate gives
foreground fetches a token immediately, while **background fetches yield** — they wait until no
foreground fetch is pending before taking a token. Threaded via a context value (not new
signatures) and a process-wide `foregroundPending` counter. Clicks are rare and bursty, so the
background barely loses throughput; the waiting human gets the next (and cleanest) slot.
**Credentials probe — rejected, not built.** The idea was to fetch captions with the user's
auth in manual mode to dodge 429s. It is a dead end, already settled by ADR-010 and the code:
the caption path is *deliberately anonymous* because the InnerTube/timedtext endpoints **reject
or break on authenticated requests** (`captions.go`: "no OAuth token — it can break the
timedtext endpoint"). The user's OAuth (a Data API credential) does not authenticate InnerTube
at all, and the official `captions.download` is owner-only (403 on third-party). So auth cannot
help here and can actively hurt. No probe needed — building one would only re-confirm the ADR.
**Reversibility.** Context-marker + a yield loop in the gate; removing the marker collapses to
the prior equal-share behaviour. No schema or API change.
---
## ADR-027 — Chat with a video's stored transcript (deeper-dive, on an already-summarized video)
**Status:** Accepted (2026-06-11). **Consumes ADR-021** (the shared, video-keyed transcript
store) for the first time beyond summarization; **uses the ADR-022 chain models**; relates to
ADR-012 (isolation) and ADR-016 (the Stage-0 gate).
**Context — observed demand, not hypothetical.** The maintainer read 10+ real pilot summaries
and reported the reactions: *many good; some he wanted to dig deeper into; some less useful*
(the "less useful" split between weak-*model* output and uninteresting-*video* content). The
middle reaction is the signal: a good summary that makes the reader want *more* is the summary
succeeding at triage and then hitting a wall — there is nowhere to go deeper short of watching
the video. That want is the feature. It is also the cheapest possible feature to satisfy
honestly, because ADR-021 already persists the transcript: the deeper-dive runs entirely on
stored public-content text + local models, touching **no** caption fetch and **no** YouTube.
**Decision.** Add a per-video chat that lets the user ask questions against a video's stored
transcript.
1. **Entry from the summary view only.** A "dig deeper / ask" affordance on a summarized video —
the chat lives exactly where the "I want more" reaction happens. No standalone chat surface.
2. **Stored-transcript-only (load-bearing constraint).** Chat is available **only** for videos
that already have a stored transcript. It never triggers a caption fetch, so it cannot touch
the rate gate, the 429 surface, or YouTube at all — the entire account-safety constraint that
governs the rest of Tapir is satisfied *by construction* here, not by careful gating. (Entry
being "from a summary" guarantees the transcript exists.) On-demand fetch for un-stored videos
is explicitly deferred.
3. **Model = the summary's model by default; user-switchable among the ADR-022 chain models**
(`phi4-mini` / `gemma4-26b` / `mistral-small` to start). This is deliberate: it doubles as
live model-comparison instrumentation — ask the same question of the same transcript under two
models and the difference is directly felt. This is the mechanism by which the maintainer
learns *which* model is worth defaulting to, and it is the multi-model-analysis direction
ADR-021 anticipated, arriving as a user-facing capability.
- Chat is a **read-bounded retrieval/QA task** (the user supplies the focus), which is
*easier* than summarization (the model must decide what matters). So a model that summarizes
mediocrely may chat well — chat is plausibly a partial remedy for the weak-summary case, not
an inheritor of it.
4. **Ephemeral chat (v1).** No persisted history; chat is per-session. Persisting per-user,
RLS-scoped history is deferred until there is evidence anyone wants to revisit a conversation.
5. **Chat-only, trust-the-model (v1) — with a recorded limitation.** The chat does not expose the
raw transcript for verification in v1 (kept simple). **Known limitation:** because some
summaries were weak-model output, the user has reason not to fully trust a chat answer's
fidelity to the transcript, and v1 gives no in-UI way to check. The model-switcher partially
compensates (two models disagreeing on the same question is itself a signal). A
"show source / view transcript" verification path is the natural **v2** and is *not*
foreclosed — ADR-021's stored transcript already makes it cheap. Recorded so v2 is a known
next step, not a rediscovery.
**Why this is safe and in-scope.** It adds no caption-fetch surface (stored-only), no new
non-RLS table (transcripts already shared per ADR-021; ephemeral chat stores nothing), and no
auth change. It is additive to the read path. The one genuine product expansion — Tapir becomes
an interactive transcript-QA tool, not only a summarizer — is justified by *observed* demand from
real reading, which is exactly the kind of evidence the Stage-0 discipline asks for before
building.
**Relation to the Stage-0 gate.** This is **not** a return-nudge (those stay deferred, ADR-020) —
it adds nothing that prompts the user to return; it deepens the value *once they are already
reading*. It does not contaminate the unprompted-return signal. If anything it strengthens the
"useful to me" case the gate measures, by giving a good summary somewhere to lead.
**Reversibility.** Additive read-path feature over the unchanged engine + the ADR-021 store.
Removing the summary-view affordance removes the feature; nothing else depends on it. Ephemeral =
no migration, no stored state to unwind. Spec: `docs/specs/chat-with-transcript.md`.
---
## ADR-028 — Onboarding burst: pick likely-good videos, summarize them with a stronger model
**Status:** Accepted (2026-06-11). **Refines ADR-018** (the connect-time burst) and **ADR-020**
(recency-bounded auto-summarize). **Builds on ADR-022** (the endpoint chain), **ADR-023**
(the discovery-time `videos.list` enrichment), and **ADR-021** (the shared transcript cache).
Triggered by a Phase-1 investigation of the live pilot DB.
**Context.** A new user's first session decides whether they return (the Stage-0 gate, ADR-016).
The connect-time burst (ADR-018: summarize ≤`TAPIR_ONBOARD_SUMMARIZE_COUNT` newest videos so the
feed isn't empty) *fires* in production, but a live-DB investigation of the second pilot user
("Jonte") found it delivers a weak first impression for two reasons, and ruled out a third idea:
1. **Junk picks.** Selection was pure newest-first (`NewestUnsummarizedVideoIDs`,
`ORDER BY published_at DESC`) with **no quality signal**. Jonte's live burst-3 were a
stock-ticker **livestream** + two regional news clips — newest, not best. The cheap signals
that *could* gate this (duration, live status) are fetched by ADR-023's `videos.list`
enrichment at discovery and then **thrown away**: the `videos.duration_s` column (migration
001) was never written.
2. **Weakest model on the first impression.** All burst summaries ran on `koala/phi4-mini` — the
documented weak link (ADR-022 was born from its failures). The stronger, brain-validated
`iguana/gemma4-26b` was never used, even though the burst is only ~3 summaries.
3. **Cached-first instant summaries — REJECTED.** The idea: skip the fetch, summarize
already-cached transcripts (ADR-021) instantly. The pilot numbers kill it — only **11 videos**
overlap between the two users (~3% of each ~350400-video library), **0** cached-and-
unsummarized, and a new user's newest-20 unsummarized are **20/20 NOT cached**. Newest-first
and cached-first are structurally incompatible: fresh uploads are exactly what nobody has
fetched. An empty lever at pilot scale.
**Decision.**
1. **Persist `duration_s` at discovery.** `filterLowValue` (ADR-023) already has each candidate's
duration in hand; carry it onto the kept `domain.Video` and have `UpsertVideo` write it,
COALESCE-preserving a known value (the channel-title backfill stance, migration 014). No new
migration — the column exists. The connect-triggered discovery pass runs *before* the burst,
so a fresh user's candidates are enriched in time.
2. **Junk-avoiding selection.** A new `OnboardBurstVideoIDs(userID, limit, minSeconds, maxSeconds)`
keeps the newest-first order but drops a video when its duration is *known* and outside
`[minSeconds, maxSeconds]``minSeconds` = `TAPIR_MIN_VIDEO_SECONDS` (60, the Shorts floor),
`maxSeconds` = new `TAPIR_ONBOARD_MAX_VIDEO_SECONDS` (default 14400 = 4h, to drop multi-hour
livestream VODs that pass the live filter once ended). A NULL duration is **unknown** — kept
(degrade-open) but ranked after known-good rows. **has-captions stays un-gateable pre-fetch**
(only knowable after a gate fetch or a ~0-probability cache hit); selection only *avoids
known-junk*, it does not *promise* captions.
3. **Stronger model for the burst only.** `TAPIR_ONBOARD_SUMMARIZER_MODEL` (default
`iguana/gemma4-26b`) leads a burst-specific summarizer chain (onboard model first, then the
standard ADR-022 chain as resilience, deduped), wrapped in a burst-specific processor over the
*same* store/cache/sink — a pure wiring choice; the engine and ports are unchanged (ADR-003).
Empty or equal-to-primary collapses the burst back onto the shared processor.
**Not a throughput change.** The caption rate gate (ADR-014) and the foreground priority lane
(ADR-026) are untouched — same pacing, same cap. This changes *which* ≤3 videos the burst spends
its fetches on and *which model* summarizes them, never how fast or how many. The engine's
existing read-stored-first (ADR-021) is unchanged and still yields a free instant summary on the
rare cache hit — we simply do not *select* for cache hits.
**Consequences.** Better odds of a strong first session: the burst avoids the obvious junk and
runs the better model on the one impression that decides return. The selection improvement is
forward-looking — existing rows have NULL `duration_s` until their next discovery pass backfills
it (lazy, like channel_title); a brand-new user benefits immediately because connect-discovery
runs first. `duration_s` becoming live also unblocks future length-aware features (feed sorting,
"long read" badges) for free.
**Reversibility.** Pure config + wiring + one column write + one query, no migration.
`TAPIR_ONBOARD_MAX_VIDEO_SECONDS=0` (and `TAPIR_MIN_VIDEO_SECONDS=0`) restores pure newest-first;
`TAPIR_ONBOARD_SUMMARIZER_MODEL=""` collapses the burst back to the shared processor.
Spec: `docs/specs/onboarding-wow-burst.md`.
---
## ADR-029 — Stateless session cookie (survives restarts, browser-close, idle)
**Status:** Accepted (2026-06-11). Triggered by pilot feedback: "lots of clicking to log in again
on iPhone." Supersedes the in-memory session store in the ADR-011 login.
**Context.** Three compounding causes made users re-login constantly:
1. **In-memory session store** (`sessionStore` map) — wiped on every pod restart, so each deploy
logged everyone out. During the active build period that was ~15 logouts.
2. **1-hour session TTL** — for a "check back tomorrow" reader, idle > 1h forced a re-login on
nearly every visit.
3. **No cookie Max-Age** — a session cookie (deleted on browser/app close); iPhone Safari closing
the tab dropped it.
Each re-login is the full Dex/Authentik redirect dance — many taps on mobile.
**Decision.** Make the session **stateless**: the identity (subject + email) and an absolute
expiry live INSIDE the existing HMAC-signed (HS256) cookie — no server-side table. Plus:
- **30-day sliding TTL** (was 1h), re-signed on each request so an active user never lapses.
- **Persistent cookie** (`Max-Age` set) so it survives browser/app close.
The cookie is HttpOnly + Secure + SameSite=Lax; the HMAC (keyed by the stable ESO
`tapir-session-secret`, which does NOT rotate per deploy) makes it tamper-proof. The payload is
identity, not secrets — the OIDC access/ID tokens are still discarded after callback.
**Consequences.** A deploy/restart no longer logs anyone out (proven by a test: a cookie issued by
one instance is accepted by a fresh instance with the same secret); works across replicas for
free. **Trade:** no server-side revocation — `logout` clears the cookie client-side, but a copied
cookie stays valid until expiry. Accepted for the Stage-0 reader pilot; revisit (server-side
revocation list, or shorter TTL + refresh) if it ever holds sensitive actions. Rotating
`tapir-session-secret` invalidates all sessions — the global logout lever.
**Not addressed here:** the tap-count of the IdP login page itself is Authentik's UX; with
re-login now rare (30-day idle or explicit logout), it matters far less.
---
## ADR-030 — Observability: slog timing + Prometheus metrics (AI-focused)
**Status:** Proposed (2026-06-11). Issue #15. **Draft for review — no code yet.**
**Context / requirements.** Nothing measures the activities that drive Tapir's performance and
UX, and the Stage-0 eval gate (ADR-016) needs a *performance* dimension to sit beside the
return-usage one. We need timing for: caption fetches (the scarce op), summarization (which model
won, how long, fallbacks), Q&A latency, LLM token spend, and basic session/usage (request rate,
latency by route, logins). Requirements:
- R1: structured `slog` timing at each AI call site (human-readable, already the logging stack).
- R2: Prometheus metrics for the same, scrapeable by the cluster's prometheus-operator.
- R3: **AI metrics are the priority** — summarize latency by `model`/`outcome`/`fallback`,
caption-fetch latency by `outcome`, chat latency by `model`, and LLM `tokens` by model+kind.
- R4: HTTP/session metrics via middleware — request count + latency by route, logins.
- R5: bounded label cardinality (no per-user, no raw-path labels).
- R6: `/metrics` must NOT be publicly exposed.
**Decision / architecture.**
1. **New package `internal/metrics`** owns all Prometheus collectors + a typed API
(`ObserveSummarize`, `ObserveCaptionFetch`, `ObserveChat`, `RecordTokens`, `IncLogin`,
`HTTPMiddleware`, `Handler`). Adapters call this API; they never import prometheus types.
2. **New dependency `github.com/prometheus/client_golang`.** Justification: it is *the* standard
Go Prometheus client and the cluster already runs prometheus-operator; hand-rolling exposition
is not worth it. (Needs the dep-justification note in the commit per repo rules.)
3. **The copied `llm` package stays stdlib-only (ADR-004).** It must not import `internal/metrics`.
Token usage is surfaced via an **optional callback** `llm.WithUsageHook(func(model string, prompt, completion int))`
set at wiring time (`buildSummarizer`/`buildChat`) to `metrics.RecordTokens`; `llm.Client` only
gains parsing of the response `usage` block. Our own adapters (`summarizer`, `youtube`, `chat`)
may import `internal/metrics` directly.
4. **HTTP middleware** reads `r.Pattern` AFTER routing (Go 1.22 sets it during ServeMux match), so
the `route` label is the bounded registered pattern (`GET /v/{videoId}`), satisfying R5;
unmatched → `other`.
5. **Dedicated metrics port** (`TAPIR_METRICS_ADDR`, default `:9090`) served by a second
`http.Server` in `cmdServe`; `/metrics` is never on the public app mux (R6). A **PodMonitor**
in `mathias/infra` scrapes it; the deployment exposes the port.
6. **slog** elapsed fields are emitted alongside each metric at the call sites (R1).
**Hook points (where the instrumentation lands).**
- `summarizer.Summarize` — per-endpoint timing + outcome (`success`/`parse_error`/`error`) + fallback flag.
- `youtube.FetchTranscript` — fetch timing + outcome from `domain.Transcript.Source`.
- `chat.Service` answer — timing by model.
- `llm.Client.Complete` — parse `usage`, fire the usage hook.
- `oidc.handleCallback``IncLogin`.
- `cmdServe` — wrap `Router()` in `metrics.HTTPMiddleware`; start the metrics server.
**Out of scope / later.** Persisting per-summary latency into Postgres for `tapir report`
(derive UX latency — publish/discovery → summary — from existing timestamps first; only persist
op-latency if the scrape proves insufficient). SPA view (#16) and visual refresh (#17).
**Reversibility.** Additive: a new package + middleware + a metrics port. Removing the PodMonitor
stops scraping; the app is unaffected. No schema change.
**Next steps (gated):** on approval of this ADR → BDD scenarios (`docs/use-cases/observability.feature`
+ scenario-coverage map) → TDD → implement → SemVer + docs + PodMonitor.
---
## Rejected alternatives ## Rejected alternatives
Approaches considered during the 2026-06-02 planning + grill session and **deliberately not Approaches considered during the 2026-06-02 planning + grill session and **deliberately not
@@ -981,6 +1242,7 @@ maps to the ADR that settles it.
| Feedback-based Stage 0 gate (friends saying it's useful) | Politeness bias makes asked-for feedback the least reliable signal; return-usage is the real test | ADR-016 | | Feedback-based Stage 0 gate (friends saying it's useful) | Politeness bias makes asked-for feedback the least reliable signal; return-usage is the real test | ADR-016 |
| Reverse the Dex-write invite flow (Google OIDC only) | Some intended Future-B users won't use Google; OIDC-only leaves them with no onboarding path — invite flow is load-bearing | ADR-017 | | Reverse the Dex-write invite flow (Google OIDC only) | Some intended Future-B users won't use Google; OIDC-only leaves them with no onboarding path — invite flow is load-bearing | ADR-017 |
| k8s CronJob for scheduled discovery (vs in-process) | At Future-B scale the in-process scheduler is simpler to deploy; CronJob's failure-isolation benefit was weighed and traded away knowingly (revisit if >1 replica or load grows) | ADR-018 | | k8s CronJob for scheduled discovery (vs in-process) | At Future-B scale the in-process scheduler is simpler to deploy; CronJob's failure-isolation benefit was weighed and traded away knowingly (revisit if >1 replica or load grows) | ADR-018 |
| Cached-transcript-first onboarding burst (instant, zero-fetch picks) | Live pilot DB: ~3% cross-user video overlap, 0 cached-and-unsummarized, a new user's newest-20 are 20/20 uncached — newest-first and cached-first are structurally incompatible. Empty lever at pilot scale | ADR-028 |
If a future case genuinely reopens one of these, that's a new ADR superseding the relevant one — If a future case genuinely reopens one of these, that's a new ADR superseding the relevant one —
not a silent reversal. not a silent reversal.
+64
View File
@@ -0,0 +1,64 @@
package main
import (
"strings"
"testing"
"gitea.d-ma.be/mathias/tapir/internal/config"
)
// TestChatModelsReuseTheChainLocalFirst: the switcher offers the ADR-022 chain in
// order, deduped — primary, local fallback, cloud.
func TestChatModelsReuseTheChainLocalFirst(t *testing.T) {
got := chatModels(config.Config{
SummarizerModel: "koala/phi4-mini",
FallbackModel: "iguana/gemma4-26b",
CloudFallbackModel: "berget/mistral-small",
})
want := []string{"koala/phi4-mini", "iguana/gemma4-26b", "berget/mistral-small"}
if strings.Join(got, ",") != strings.Join(want, ",") {
t.Fatalf("chatModels = %v, want %v", got, want)
}
}
// TestChatCloudModelAbsentWhenDisabled: the local-first / NDA lever — with the
// cloud fallback empty (TAPIR_CLOUD_FALLBACK_MODEL=""), no external model is
// offered in the switcher, so chat content never leaves the local stack (ADR-027,
// honouring ADR-022's "content stays local" guarantee).
func TestChatCloudModelAbsentWhenDisabled(t *testing.T) {
got := chatModels(config.Config{
SummarizerModel: "koala/phi4-mini",
FallbackModel: "iguana/gemma4-26b",
CloudFallbackModel: "",
})
for _, m := range got {
if strings.HasPrefix(m, "berget/") || strings.Contains(m, "mistral") {
t.Fatalf("cloud model %q offered though the cloud fallback is disabled", m)
}
}
want := []string{"koala/phi4-mini", "iguana/gemma4-26b"}
if strings.Join(got, ",") != strings.Join(want, ",") {
t.Fatalf("chatModels = %v, want %v", got, want)
}
}
// TestChatModelsDedup: a config that reuses one alias across slots collapses to a
// single switcher entry (no duplicate options).
func TestChatModelsDedup(t *testing.T) {
got := chatModels(config.Config{
SummarizerModel: "koala/phi4-mini",
FallbackModel: "koala/phi4-mini",
CloudFallbackModel: "",
})
if len(got) != 1 || got[0] != "koala/phi4-mini" {
t.Fatalf("chatModels = %v, want a single deduped entry", got)
}
}
// TestBuildChatNilWithoutGateway: no gateway → no chat backend (routes unmounted,
// read path unaffected).
func TestBuildChatNilWithoutGateway(t *testing.T) {
if c := buildChat(config.Config{GatewayURL: ""}); c != nil {
t.Fatal("buildChat must return nil without a gateway URL")
}
}
+52 -11
View File
@@ -28,6 +28,7 @@ import (
"gitea.d-ma.be/mathias/tapir/internal/adapters/youtube" "gitea.d-ma.be/mathias/tapir/internal/adapters/youtube"
"gitea.d-ma.be/mathias/tapir/internal/auth" "gitea.d-ma.be/mathias/tapir/internal/auth"
"gitea.d-ma.be/mathias/tapir/internal/config" "gitea.d-ma.be/mathias/tapir/internal/config"
"gitea.d-ma.be/mathias/tapir/internal/metrics"
"gitea.d-ma.be/mathias/tapir/internal/runner" "gitea.d-ma.be/mathias/tapir/internal/runner"
"gitea.d-ma.be/mathias/tapir/internal/web" "gitea.d-ma.be/mathias/tapir/internal/web"
"gitea.d-ma.be/mathias/tapir/internal/web/oidc" "gitea.d-ma.be/mathias/tapir/internal/web/oidc"
@@ -197,6 +198,15 @@ func cmdServe(ctx context.Context, log *slog.Logger) error {
secretStore := secrets.NewFileStore(cfg.SecretsFile) secretStore := secrets.NewFileStore(cfg.SecretsFile)
app := &web.App{Store: st, Identity: st, Auth: authn, Secrets: secretStore, Log: log, RecencyWindow: cfg.AutoSummarizeWindow} app := &web.App{Store: st, Identity: st, Auth: authn, Secrets: secretStore, Log: log, RecencyWindow: cfg.AutoSummarizeWindow}
// Per-video deeper-dive chat over the STORED transcript (ADR-027). Enabled
// whenever a gateway is configured — it needs no YouTube credentials because it
// never fetches. Guarded so a typed-nil never lands in the interface field
// (which would mount the routes over a nil backend).
if c := buildChat(cfg); c != nil {
app.Chat = c
log.Info("web chat enabled (stored-transcript only)", "models", chatModels(cfg))
}
// User onboarding is handled by the IdP (Authentik invite flow), not Tapir — // User onboarding is handled by the IdP (Authentik invite flow), not Tapir —
// the Dex local-password provisioning path was removed (ADR-019). An // the Dex local-password provisioning path was removed (ADR-019). An
// authenticated subject with no Tapir user is routed to /register. // authenticated subject with no Tapir user is routed to /register.
@@ -253,23 +263,36 @@ func cmdServe(ctx context.Context, log *slog.Logger) error {
// One lock shared by the scheduler and connect-triggered passes (#6) so // One lock shared by the scheduler and connect-triggered passes (#6) so
// they never fetch concurrently — the single-fetcher invariant (ADR-018). // they never fetch concurrently — the single-fetcher invariant (ADR-018).
runUser := serialize(&sync.Mutex{}, rawRunUser) runUser := serialize(&sync.Mutex{}, rawRunUser)
// Onboarding burst (Feature 1): after the connect-triggered discovery pass, // Onboarding burst (Feature 1, refined by ADR-028): after the connect-triggered
// summarize up to OnboardSummarizeCount of the user's NEWEST unsummarized // discovery pass, summarize up to OnboardSummarizeCount of the user's newest
// videos so a fresh account gets real summaries in its first session. Hard // LIKELY-GOOD unsummarized videos so a fresh account gets a strong first
// cap; explicit, so it bypasses the recency window — but every fetch still // session. Selection avoids known-junk (Shorts/over-long/livestream VODs via
// goes through globalFetchGate via the Processor. No-op when disabled // the persisted duration); the burst leads its chain with the stronger onboard
// (count 0) or queue-only (no Processor). // model. Hard cap; explicit, so it bypasses the recency window — but every
// fetch still goes through globalFetchGate. No-op when disabled (count 0) or
// queue-only (no processor).
//
// burstProcessor leads with the stronger model (ADR-028); it collapses onto the
// shared Processor when the onboard model is empty/equal-to-primary or the
// engine config is incomplete.
burstProcessor := app.Processor
if burstEngine, berr := buildBurstProcessor(cfg, st); berr != nil {
return berr
} else if burstEngine != nil {
burstProcessor = &engineProcessor{engine: burstEngine, store: st}
log.Info("onboarding burst uses a stronger model", "onboard_model", cfg.OnboardSummarizerModel)
}
onboard := func(ctx context.Context, userID string) { onboard := func(ctx context.Context, userID string) {
if cfg.OnboardSummarizeCount <= 0 || app.Processor == nil { if cfg.OnboardSummarizeCount <= 0 || burstProcessor == nil {
return return
} }
ids, err := st.NewestUnsummarizedVideoIDs(ctx, userID, cfg.OnboardSummarizeCount) ids, err := st.OnboardBurstVideoIDs(ctx, userID, cfg.OnboardSummarizeCount, cfg.MinVideoSeconds, cfg.OnboardMaxVideoSeconds)
if err != nil { if err != nil {
log.Warn("onboarding: list newest unsummarized", "user", userID, "err", err) log.Warn("onboarding: list burst candidates", "user", userID, "err", err)
return return
} }
for _, id := range ids { for _, id := range ids {
if err := app.Processor.ProcessVideo(ctx, userID, id); err != nil { if err := burstProcessor.ProcessVideo(ctx, userID, id); err != nil {
log.Warn("onboarding: summarize", "user", userID, "video", id, "err", err) log.Warn("onboarding: summarize", "user", userID, "video", id, "err", err)
} }
} }
@@ -288,16 +311,34 @@ func cmdServe(ctx context.Context, log *slog.Logger) error {
srv := &http.Server{ srv := &http.Server{
Addr: cfg.HTTPAddr, Addr: cfg.HTTPAddr,
Handler: app.Router(), Handler: metrics.HTTPMiddleware(app.Router()),
ReadHeaderTimeout: 10 * time.Second, ReadHeaderTimeout: 10 * time.Second,
} }
// Prometheus /metrics on a SEPARATE port (ADR-030) — never on the public app
// mux, so a scrape is in-cluster only. Empty TAPIR_METRICS_ADDR disables it.
var metricsSrv *http.Server
if cfg.MetricsAddr != "" {
mmux := http.NewServeMux()
mmux.Handle("GET /metrics", metrics.Handler())
metricsSrv = &http.Server{Addr: cfg.MetricsAddr, Handler: mmux, ReadHeaderTimeout: 10 * time.Second}
go func() {
log.Info("serving metrics", "addr", cfg.MetricsAddr)
if err := metricsSrv.ListenAndServe(); err != nil && !errors.Is(err, http.ErrServerClosed) {
log.Error("metrics server", "err", err)
}
}()
}
// Graceful shutdown on signal: stop accepting, drain in-flight requests. // Graceful shutdown on signal: stop accepting, drain in-flight requests.
go func() { go func() {
<-ctx.Done() <-ctx.Done()
shutdownCtx, cancel := context.WithTimeout(context.Background(), 10*time.Second) shutdownCtx, cancel := context.WithTimeout(context.Background(), 10*time.Second)
defer cancel() defer cancel()
_ = srv.Shutdown(shutdownCtx) _ = srv.Shutdown(shutdownCtx)
if metricsSrv != nil {
_ = metricsSrv.Shutdown(shutdownCtx)
}
}() }()
log.Info("serving web ui", "addr", cfg.HTTPAddr, "user", cfg.UserID) log.Info("serving web ui", "addr", cfg.HTTPAddr, "user", cfg.UserID)
+133 -16
View File
@@ -5,6 +5,7 @@ import (
"fmt" "fmt"
"strings" "strings"
"gitea.d-ma.be/mathias/tapir/internal/adapters/chat"
"gitea.d-ma.be/mathias/tapir/internal/adapters/llm" "gitea.d-ma.be/mathias/tapir/internal/adapters/llm"
"gitea.d-ma.be/mathias/tapir/internal/adapters/secrets" "gitea.d-ma.be/mathias/tapir/internal/adapters/secrets"
"gitea.d-ma.be/mathias/tapir/internal/adapters/store" "gitea.d-ma.be/mathias/tapir/internal/adapters/store"
@@ -12,6 +13,7 @@ import (
"gitea.d-ma.be/mathias/tapir/internal/adapters/youtube" "gitea.d-ma.be/mathias/tapir/internal/adapters/youtube"
"gitea.d-ma.be/mathias/tapir/internal/config" "gitea.d-ma.be/mathias/tapir/internal/config"
"gitea.d-ma.be/mathias/tapir/internal/domain" "gitea.d-ma.be/mathias/tapir/internal/domain"
"gitea.d-ma.be/mathias/tapir/internal/metrics"
"gitea.d-ma.be/mathias/tapir/internal/ports" "gitea.d-ma.be/mathias/tapir/internal/ports"
"gitea.d-ma.be/mathias/tapir/internal/usecase" "gitea.d-ma.be/mathias/tapir/internal/usecase"
"gitea.d-ma.be/mathias/tapir/internal/web" "gitea.d-ma.be/mathias/tapir/internal/web"
@@ -51,13 +53,7 @@ func (f videoFetcher) FetchVideo(ctx context.Context, userID, videoID string) (d
// alias, not a second client config. Empty model entries are skipped, so a // alias, not a second client config. Empty model entries are skipped, so a
// client deployment can set the cloud fallback empty to keep content local. // client deployment can set the cloud fallback empty to keep content local.
func buildSummarizer(cfg config.Config) *summarizer.Summarizer { func buildSummarizer(cfg config.Config) *summarizer.Summarizer {
mk := func(model string) summarizer.Endpoint { mk := summarizerEndpoint(cfg)
return summarizer.Endpoint{
Client: llm.New(cfg.GatewayURL, cfg.GatewayKey, model, cfg.SummarizerTimeout, llm.WithMaxTokens(cfg.SummaryMaxTokens)),
Provider: providerOf(model),
Model: model,
}
}
eps := []summarizer.Endpoint{mk(cfg.SummarizerModel)} eps := []summarizer.Endpoint{mk(cfg.SummarizerModel)}
if cfg.FallbackModel != "" && cfg.FallbackModel != cfg.SummarizerModel { if cfg.FallbackModel != "" && cfg.FallbackModel != cfg.SummarizerModel {
eps = append(eps, mk(cfg.FallbackModel)) eps = append(eps, mk(cfg.FallbackModel))
@@ -68,6 +64,98 @@ func buildSummarizer(cfg config.Config) *summarizer.Summarizer {
return summarizer.NewChain(eps, cfg.MaxTranscriptChars) return summarizer.NewChain(eps, cfg.MaxTranscriptChars)
} }
// summarizerEndpoint returns a constructor for a chain endpoint over the one
// LiteLLM gateway, varying only the model alias (the gateway fronts both
// llama-swap and berget). Shared by the standard and burst chains.
func summarizerEndpoint(cfg config.Config) func(model string) summarizer.Endpoint {
return func(model string) summarizer.Endpoint {
return summarizer.Endpoint{
Client: llm.New(cfg.GatewayURL, cfg.GatewayKey, model, cfg.SummarizerTimeout, llm.WithMaxTokens(cfg.SummaryMaxTokens), llm.WithUsageHook(metrics.RecordTokens)),
Provider: providerOf(model),
Model: model,
}
}
}
// burstChainModels is the ordered, deduped model list for the onboarding burst
// (ADR-028): the stronger onboard model leads, then the standard ADR-022 chain
// (primary → local fallback → cloud) follows as resilience. Empty entries are
// dropped and duplicates collapsed, so the NDA lever (empty cloud fallback) keeps
// the burst chain fully local exactly as the standard chain does.
func burstChainModels(cfg config.Config) []string {
var models []string
add := func(m string) {
if m == "" {
return
}
for _, e := range models {
if e == m {
return
}
}
models = append(models, m)
}
add(cfg.OnboardSummarizerModel)
add(cfg.SummarizerModel)
add(cfg.FallbackModel)
add(cfg.CloudFallbackModel)
return models
}
// buildBurstSummarizer builds the onboarding-burst summarizer chain (ADR-028):
// the onboard model first, then the standard chain as fallback, deduped.
func buildBurstSummarizer(cfg config.Config) *summarizer.Summarizer {
mk := summarizerEndpoint(cfg)
var eps []summarizer.Endpoint
for _, m := range burstChainModels(cfg) {
eps = append(eps, mk(m))
}
return summarizer.NewChain(eps, cfg.MaxTranscriptChars)
}
// chatModels is the ordered, local-first set of models offered in the chat
// switcher (ADR-027), reusing the ADR-022 chain: primary → local fallback →
// cloud. Empty entries are dropped and duplicates collapsed, so a client/NDA
// deployment that sets the cloud fallback empty simply has no external model in
// the switcher — the same local-first lever the summarizer honours.
func chatModels(cfg config.Config) []string {
var models []string
add := func(m string) {
if m == "" {
return
}
for _, e := range models {
if e == m {
return
}
}
models = append(models, m)
}
add(cfg.SummarizerModel)
add(cfg.FallbackModel)
add(cfg.CloudFallbackModel)
return models
}
// buildChat wires the per-video chat service (ADR-027): a Completer factory over
// the SAME LiteLLM gateway the summarizer uses (a different alias per model, not a
// second client config) and the same transcript-truncation budget. It returns nil
// when no gateway is configured — chat is simply not mounted, the read path is
// unaffected. It deliberately takes NO YouTube source: chat is stored-only.
func buildChat(cfg config.Config) *chat.Service {
if cfg.GatewayURL == "" {
return nil
}
models := chatModels(cfg)
if len(models) == 0 {
return nil
}
newClient := func(model string) chat.Completer {
return llm.New(cfg.GatewayURL, cfg.GatewayKey, model, cfg.SummarizerTimeout, llm.WithMaxTokens(cfg.SummaryMaxTokens), llm.WithUsageHook(metrics.RecordTokens))
}
return chat.New(newClient, models, cfg.MaxTranscriptChars)
}
// providerOf maps a model alias to the domain AIProvider recorded on summaries. // providerOf maps a model alias to the domain AIProvider recorded on summaries.
// A "berget/" alias is an external provider; everything else is the local stack. // A "berget/" alias is an external provider; everything else is the local stack.
func providerOf(model string) string { func providerOf(model string) string {
@@ -82,15 +170,7 @@ func buildProcessor(cfg config.Config, st *store.Store) (*usecase.Engine, error)
return nil, nil return nil, nil
} }
secretStore := secrets.NewFileStore(cfg.SecretsFile) src := newYouTubeSource(cfg, secrets.NewFileStore(cfg.SecretsFile))
src := youtube.New(youtube.Config{
ClientID: cfg.YTClientID,
ClientSecret: cfg.YTClientSecret,
TokenSecretRef: cfg.YTTokenRef,
PreferredLanguages: []string{"en"},
MinVideoSeconds: cfg.MinVideoSeconds,
}, secretStore)
sum := buildSummarizer(cfg) sum := buildSummarizer(cfg)
// The store is both the summary sink and the shared transcript cache (ADR-021): // The store is both the summary sink and the shared transcript cache (ADR-021):
@@ -101,6 +181,38 @@ func buildProcessor(cfg config.Config, st *store.Store) (*usecase.Engine, error)
return eng, nil return eng, nil
} }
// newYouTubeSource builds the captions-first VideoSource shared by the standard
// and burst processors — same per-process YouTube credentials and ADR-023 Shorts
// filter; only the summarizer chain differs between them.
func newYouTubeSource(cfg config.Config, secretStore ports.SecretStore) ports.VideoSource {
return youtube.New(youtube.Config{
ClientID: cfg.YTClientID,
ClientSecret: cfg.YTClientSecret,
TokenSecretRef: cfg.YTTokenRef,
PreferredLanguages: []string{"en"},
MinVideoSeconds: cfg.MinVideoSeconds,
}, secretStore)
}
// buildBurstProcessor wires a processor whose summarizer leads with the stronger
// onboard model (ADR-028), used only by the connect-time burst over the SAME
// store / transcript cache / sink — a wiring choice; the engine and ports are
// unchanged. Returns (nil, nil) — the collapse lever — when the onboard model is
// empty or equal to the primary (the burst then reuses the shared processor), or
// when the engine config is incomplete (queue-only, same as buildProcessor).
func buildBurstProcessor(cfg config.Config, st *store.Store) (*usecase.Engine, error) {
if cfg.OnboardSummarizerModel == "" || cfg.OnboardSummarizerModel == cfg.SummarizerModel {
return nil, nil
}
if cfg.GatewayURL == "" || cfg.YTClientID == "" || cfg.YTClientSecret == "" || cfg.SecretsFile == "" {
return nil, nil
}
src := newYouTubeSource(cfg, secrets.NewFileStore(cfg.SecretsFile))
eng := usecase.NewEngine(src, buildBurstSummarizer(cfg), st)
eng.Transcripts = st
return eng, nil
}
// engineProcessor adapts the engine (which works in terms of a domain.Video) to // engineProcessor adapts the engine (which works in terms of a domain.Video) to
// the web.Processor port (which works in terms of a stored video id): it loads the // the web.Processor port (which works in terms of a stored video id): it loads the
// video row, runs the engine, and — on a produced summary — clears the manual // video row, runs the engine, and — on a produced summary — clears the manual
@@ -113,6 +225,11 @@ type engineProcessor struct {
} }
func (p *engineProcessor) ProcessVideo(ctx context.Context, userID, videoID string) error { func (p *engineProcessor) ProcessVideo(ctx context.Context, userID, videoID string) error {
// This is the user-initiated (foreground) path — a click on "Summarize",
// "Try now", or a pasted URL. Mark the context so the caption gate gives it
// priority over the background sweep (ADR-026, Pillar A).
ctx = youtube.ForegroundContext(ctx)
row, err := p.store.GetVideoRow(ctx, userID, videoID) row, err := p.store.GetVideoRow(ctx, userID, videoID)
if err != nil { if err != nil {
return fmt.Errorf("load video %q: %w", videoID, err) return fmt.Errorf("load video %q: %w", videoID, err)
+62
View File
@@ -45,3 +45,65 @@ func TestBuildProcessorNilOnIncompleteConfig(t *testing.T) {
}) })
} }
} }
// TestBurstChainModelsLeadsWithOnboardModel: the onboarding burst chain (ADR-028)
// leads with the stronger onboard model, then falls back through the standard
// ADR-022 chain (primary -> local fallback -> cloud), deduped.
func TestBurstChainModelsLeadsWithOnboardModel(t *testing.T) {
got := burstChainModels(config.Config{
OnboardSummarizerModel: "iguana/gemma4-26b",
SummarizerModel: "koala/phi4-mini",
FallbackModel: "iguana/gemma4-26b", // also the onboard model -> dedup
CloudFallbackModel: "berget/mistral-small",
})
want := []string{"iguana/gemma4-26b", "koala/phi4-mini", "berget/mistral-small"}
if len(got) != len(want) {
t.Fatalf("burstChainModels = %v, want %v", got, want)
}
for i := range want {
if got[i] != want[i] {
t.Fatalf("burstChainModels = %v, want %v", got, want)
}
}
}
// TestBurstChainModelsCloudAbsentWhenDisabled: the NDA lever holds for the burst
// too — empty cloud fallback keeps the burst chain fully local.
func TestBurstChainModelsCloudAbsentWhenDisabled(t *testing.T) {
got := burstChainModels(config.Config{
OnboardSummarizerModel: "iguana/gemma4-26b",
SummarizerModel: "koala/phi4-mini",
CloudFallbackModel: "",
})
for _, m := range got {
if m == "" || m == "berget/mistral-small" {
t.Fatalf("cloud model leaked into burst chain: %v", got)
}
}
}
// TestBuildBurstProcessorNilWhenCollapsed: an empty or primary-equal onboard model
// collapses the burst onto the shared processor (buildBurstProcessor returns nil).
func TestBuildBurstProcessorNilWhenCollapsed(t *testing.T) {
base := config.Config{
GatewayURL: "http://gw/v1",
YTClientID: "id",
YTClientSecret: "secret",
SecretsFile: "/tmp/secrets.json",
SummarizerModel: "koala/phi4-mini",
}
t.Run("empty onboard model", func(t *testing.T) {
base.OnboardSummarizerModel = ""
eng, err := buildBurstProcessor(base, nil)
if err != nil || eng != nil {
t.Fatalf("buildBurstProcessor = (%v, %v), want (nil, nil)", eng, err)
}
})
t.Run("onboard model equals primary", func(t *testing.T) {
base.OnboardSummarizerModel = "koala/phi4-mini"
eng, err := buildBurstProcessor(base, nil)
if err != nil || eng != nil {
t.Fatalf("buildBurstProcessor = (%v, %v), want (nil, nil)", eng, err)
}
})
}
+33 -3
View File
@@ -6,6 +6,7 @@ import (
"io" "io"
"os" "os"
"text/tabwriter" "text/tabwriter"
"time"
"gitea.d-ma.be/mathias/tapir/internal/adapters/store" "gitea.d-ma.be/mathias/tapir/internal/adapters/store"
) )
@@ -14,6 +15,27 @@ import (
// weeks. The gate passes when any user reaches it. // weeks. The gate passes when any user reaches it.
const gateThreshold = 2 const gateThreshold = 2
// defaultGateStart is the date Stage-0 return-usage tracking begins: the morning
// the pilot was actually unblocked and summaries started flowing (2026-06-11).
// Activity before this — testing, the period the pilot was stuck on zero — is
// noise and must not count toward the gate. Override with TAPIR_USAGE_GATE_START
// (YYYY-MM-DD). The gate measures whether users RETURN once it genuinely works.
const defaultGateStart = "2026-06-11"
// gateStart resolves the baseline date from TAPIR_USAGE_GATE_START or the default,
// parsed as a UTC calendar day.
func gateStart() (time.Time, error) {
v := os.Getenv("TAPIR_USAGE_GATE_START")
if v == "" {
v = defaultGateStart
}
t, err := time.Parse("2006-01-02", v)
if err != nil {
return time.Time{}, fmt.Errorf("TAPIR_USAGE_GATE_START=%q: want YYYY-MM-DD: %w", v, err)
}
return t, nil
}
// runReport prints the Stage-0 usage gate: per-user distinct active weeks (reads // runReport prints the Stage-0 usage gate: per-user distinct active weeks (reads
// UNION acts) and the pass/fail verdict. Read-only, cross-user — needs only // UNION acts) and the pass/fail verdict. Read-only, cross-user — needs only
// TAPIR_DB_DSN (not TAPIR_USER_ID; the report enumerates all users itself). // TAPIR_DB_DSN (not TAPIR_USER_ID; the report enumerates all users itself).
@@ -28,16 +50,24 @@ func runReport(ctx context.Context, _ []string) error {
} }
defer s.Close() defer s.Close()
rows, err := s.ActiveWeeks(ctx) since, err := gateStart()
if err != nil { if err != nil {
return err return err
} }
return formatReport(os.Stdout, rows)
rows, err := s.ActiveWeeks(ctx, since)
if err != nil {
return err
}
return formatReport(os.Stdout, rows, since)
} }
// formatReport renders the per-user week counts and the gate verdict. Pure: no DB, // formatReport renders the per-user week counts and the gate verdict. Pure: no DB,
// no env — so the layout and verdict logic are unit-testable without Postgres. // no env — so the layout and verdict logic are unit-testable without Postgres.
func formatReport(w io.Writer, rows []store.UserActiveWeeks) error { func formatReport(w io.Writer, rows []store.UserActiveWeeks, since time.Time) error {
if _, err := fmt.Fprintf(w, "Counting usage since %s (Stage-0 gate baseline)\n\n", since.Format("2006-01-02")); err != nil {
return err
}
if len(rows) == 0 { if len(rows) == 0 {
_, err := fmt.Fprintln(w, "no users yet") _, err := fmt.Fprintln(w, "no users yet")
return err return err
+7 -3
View File
@@ -3,12 +3,15 @@ package main
import ( import (
"strings" "strings"
"testing" "testing"
"time"
"github.com/stretchr/testify/require" "github.com/stretchr/testify/require"
"gitea.d-ma.be/mathias/tapir/internal/adapters/store" "gitea.d-ma.be/mathias/tapir/internal/adapters/store"
) )
var testSince = time.Date(2026, 6, 11, 0, 0, 0, 0, time.UTC)
func TestFormatReportColumnsAndGatePass(t *testing.T) { func TestFormatReportColumnsAndGatePass(t *testing.T) {
rows := []store.UserActiveWeeks{ rows := []store.UserActiveWeeks{
{UserID: "user-a", DisplayName: "Ada", ActiveWeeks: 3}, {UserID: "user-a", DisplayName: "Ada", ActiveWeeks: 3},
@@ -16,9 +19,10 @@ func TestFormatReportColumnsAndGatePass(t *testing.T) {
} }
var b strings.Builder var b strings.Builder
require.NoError(t, formatReport(&b, rows)) require.NoError(t, formatReport(&b, rows, testSince))
out := b.String() out := b.String()
require.Contains(t, out, "since 2026-06-11", "report states the gate baseline date")
require.Contains(t, out, "USER") require.Contains(t, out, "USER")
require.Contains(t, out, "ACTIVE_WEEKS") require.Contains(t, out, "ACTIVE_WEEKS")
require.Contains(t, out, "Ada") require.Contains(t, out, "Ada")
@@ -34,12 +38,12 @@ func TestFormatReportGateNotMet(t *testing.T) {
rows := []store.UserActiveWeeks{{UserID: "user-a", ActiveWeeks: 1}} rows := []store.UserActiveWeeks{{UserID: "user-a", ActiveWeeks: 1}}
var b strings.Builder var b strings.Builder
require.NoError(t, formatReport(&b, rows)) require.NoError(t, formatReport(&b, rows, testSince))
require.Contains(t, b.String(), "NOT YET MET", "no user at >= 2 weeks fails the gate") require.Contains(t, b.String(), "NOT YET MET", "no user at >= 2 weeks fails the gate")
} }
func TestFormatReportEmpty(t *testing.T) { func TestFormatReportEmpty(t *testing.T) {
var b strings.Builder var b strings.Builder
require.NoError(t, formatReport(&b, nil)) require.NoError(t, formatReport(&b, nil, testSince))
require.Contains(t, b.String(), "no users yet") require.Contains(t, b.String(), "no users yet")
} }
+12 -3
View File
@@ -143,8 +143,18 @@ func runScheduler(
return // disabled return // disabled
} }
pass := 0 // Derive the rotation offset from wall-clock, NOT an in-memory counter. A
// counter reset to 0 on every pod restart always hands the lead to the
// first-listed user — so frequent deploys re-starve whoever is last (exactly
// what happened to the first pilot user during a deploy-heavy session). A
// time-based offset advances with real time and is identical across restarts,
// so the lead rotates fairly regardless of how often the pod bounces.
runPass := func() {
pass := int(time.Now().Unix() / int64(interval/time.Second))
runDiscoveryPass(ctx, pass, lister, runUser, log) runDiscoveryPass(ctx, pass, lister, runUser, log)
}
runPass()
ticker := time.NewTicker(interval) ticker := time.NewTicker(interval)
defer ticker.Stop() defer ticker.Stop()
@@ -153,8 +163,7 @@ func runScheduler(
case <-ctx.Done(): case <-ctx.Done():
return return
case <-ticker.C: case <-ticker.C:
pass++ runPass()
runDiscoveryPass(ctx, pass, lister, runUser, log)
} }
} }
} }
+26
View File
@@ -250,6 +250,32 @@ After (newest-first): `[chanB-new, chanA-mid, chanA-old, chanB-null]`
The set of *processed* videos now also excludes auto-mode back-catalogue beyond the recency The set of *processed* videos now also excludes auto-mode back-catalogue beyond the recency
window (those stay listed, summarised on demand); within the processed set, only order changes. window (those stay listed, summarised on demand); within the processed set, only order changes.
### Connect-time onboarding burst (ADR-018 → ADR-028)
On a successful YouTube connect, `ConnectHandler` enqueues a connect-triggered discovery pass;
the `discoveryTrigger` runs that pass and then fires the **onboarding burst** — a third entry path
that summarises up to `TAPIR_ONBOARD_SUMMARIZE_COUNT` (default 3, hard-capped) of the new user's
videos so the first session is not empty. The burst still flows through `globalFetchGate` (it is
not a throughput change); ADR-028 sharpened *which* videos and *which model*:
- **Selection** is `OnboardBurstVideoIDs`, not pure newest-first. It keeps newest-first order but
excludes a video whose **known** duration is outside `[TAPIR_MIN_VIDEO_SECONDS,
TAPIR_ONBOARD_MAX_VIDEO_SECONDS]` (drops Shorts and multi-hour livestream VODs). An unknown
(NULL) duration is degrade-open — kept, but ranked after known-good rows. The connect-triggered
discovery pass runs *before* the burst, and ADR-023's `videos.list` enrichment now **persists**
`duration_s` (instead of discarding it after the Shorts filter), so a fresh user's candidates
carry a duration in time for selection.
- **Model**: the burst runs through a dedicated summarizer chain led by
`TAPIR_ONBOARD_SUMMARIZER_MODEL` (default `iguana/gemma4-26b`, the stronger local model), with
the standard ADR-022 chain following as fallback. This is a wiring choice — a second
`engineProcessor` over the same store / transcript cache / sink; the engine and ports are
unchanged. Empty / equal-to-primary collapses it back onto the shared processor.
`has-captions` is deliberately **not** a selection signal — it is only knowable after a gate fetch
(or a ~0-probability cache hit at pilot scale), so the burst can avoid known-junk but cannot
promise captions. Cached-transcript-first selection was investigated and rejected (ADR-028:
~3% cross-user overlap).
--- ---
## Sequence — core use case: new video summarized ## Sequence — core use case: new video summarized
+6 -1
View File
@@ -30,7 +30,9 @@ it** — endpoints and aliases drift, and this file is a snapshot (2026-06-06),
- **Summarizer fallback chain (ADR-022).** The primary alias is the *first* of an ordered chain; - **Summarizer fallback chain (ADR-022).** The primary alias is the *first* of an ordered chain;
on failure or unparseable output the summarizer advances to the next model. All reached through on failure or unparseable output the summarizer advances to the next model. All reached through
the same gateway by alias. the same gateway by alias.
- `TAPIR_FALLBACK_MODEL` — local fallback. **Default `koala/phi4-14b`.** Empty disables it. - `TAPIR_FALLBACK_MODEL` — local fallback. **Default `iguana/gemma4-26b`** — on iguana, NOT
koala, so the fallback does not compete with koala's other GPU loads (and runs from a different
egress IP). Empty disables it.
- `TAPIR_CLOUD_FALLBACK_MODEL` — worst-case EXTERNAL fallback. **Default `berget/mistral-small`.** - `TAPIR_CLOUD_FALLBACK_MODEL` — worst-case EXTERNAL fallback. **Default `berget/mistral-small`.**
**Set this empty (`""`) for any client/NDA deployment** so content never leaves the local **Set this empty (`""`) for any client/NDA deployment** so content never leaves the local
stack — the chain then contains only local endpoints. stack — the chain then contains only local endpoints.
@@ -222,6 +224,9 @@ knobs plus one load-bearing deployment constraint:
- `TAPIR_DISCOVERY_INTERVAL` — Go duration, e.g. `2h`. The cadence the serve process runs a - `TAPIR_DISCOVERY_INTERVAL` — Go duration, e.g. `2h`. The cadence the serve process runs a
discovery pass for every registered user (run-once-on-startup, then every interval). discovery pass for every registered user (run-once-on-startup, then every interval).
**Unset or `0` = disabled** (dev/tests never auto-fetch). **Unset or `0` = disabled** (dev/tests never auto-fetch).
- `TAPIR_USAGE_GATE_START``YYYY-MM-DD`, default **`2026-06-11`** (the morning the pilot was
unblocked and summaries started flowing). `tapir report` counts return-usage (distinct active
weeks, ADR-016) only from this date, so pre-launch testing and the blocked period are excluded.
- `TAPIR_FETCH_RATE` — Go duration, default `2s`. The **process-wide per-egress-IP caption-fetch - `TAPIR_FETCH_RATE` — Go duration, default `2s`. The **process-wide per-egress-IP caption-fetch
rate gate** (ADR-014 item 2). Every caption fetch — scheduler runners *and* the web "Summarize" rate gate** (ADR-014 item 2). Every caption fetch — scheduler runners *and* the web "Summarize"
click-path — serialises through this one limiter so the pod cannot collectively trip 429s. `0` click-path — serialises through this one limiter so the pod cannot collectively trip 429s. `0`
+89
View File
@@ -0,0 +1,89 @@
# Spec — Chat with a video's stored transcript (ADR-027)
**Repo:** tapir · **Size:** medium · **Solo session.** Implements ADR-027. Read CLAUDE.md,
DECISIONS.md (ADR-021 transcript store, ADR-022 model chain, ADR-012 isolation, ADR-027), and
`docs/ui-spec.md` first. TBD, conventional commits, `task check` green per commit, `templ
generate` after view changes.
## FIRST: append ADR-027 to DECISIONS.md
ADR-027 text is provided separately (planning thread). Insert immediately before the
`## Rejected alternatives` heading, as the first commit, so the decision precedes the build.
## What this is
A per-video chat letting the user ask questions against a video's **already-stored** transcript,
entered from the summary view. Born from observed demand: the maintainer read real summaries and
some made him want to dig deeper — this gives that "I want more" reaction somewhere to go, without
watching the video.
## HARD CONSTRAINT — stored-transcript-only (the safety property)
Chat is available **ONLY** for videos that already have a stored transcript (ADR-021). It must
**never** trigger a caption fetch, never touch the rate gate, never reach YouTube. Entry being
"from a summarized video" guarantees the transcript exists. If somehow invoked on a video with no
stored transcript → show "transcript not available for chat", NO fetch. This is what makes the
feature safe by construction; do not add an on-demand-fetch path (explicitly deferred).
## 1. Entry point
- A "Dig deeper" / "Ask about this" affordance on the **summary view** of a summarized video
(not the list cards — the detail/summary page). Quiet, consistent with the existing card-state
styling.
- Opens a chat panel/view scoped to that one video, with its stored transcript as context.
## 2. The chat
- Read the stored transcript for the video (via the ADR-021 `TranscriptStore`, keyed by
`(provider, provider_video_id)`). No fetch.
- Send transcript + the user's question + minimal system framing to the chosen model via the
**existing LiteLLM gateway** (the same client the summarizer uses — a chat is a different
call, not a new integration).
- Stream or return the answer; render in the chat panel. HTMX/no-JS ethos — match the existing
app (the summarize status uses HTMX polling; chat can use a simple POST-and-render or HTMX
streaming if clean).
- **Transcript truncation:** reuse/respect `TAPIR_MAX_TRANSCRIPT_CHARS` (ADR-022) so a long
transcript fits the model context. If truncated, the chat should be honest that it's working
from a bounded portion (a quiet note), since answers about the tail of a long video may be
incomplete.
## 3. Model selection (the instrumentation win)
- **Default model = the model that produced this video's summary.** (Store/lookup which chain
model summarized it — if not already recorded, this is a small addition; if recording it is
non-trivial, default to the chain primary and note the gap.)
- **User can switch** among the ADR-022 chain models (`phi4-mini`, `gemma4-26b`,
`mistral-small` to start) via a simple selector in the chat panel. Switching re-runs against
the same transcript — this is deliberate model-comparison instrumentation.
- Respect the local-first / NDA posture: if `TAPIR_CLOUD_FALLBACK_MODEL=""` (cloud disabled),
the external model is NOT offered in the switcher — only local models. Chat must honor the same
"content stays local" guarantee as ADR-022.
## 4. Ephemeral (v1)
- No persisted chat history. Conversation lives for the session/page. No new table, no migration.
- (Multi-turn within a session is fine — keep the running messages in the request/page state —
but nothing is written to the DB.)
## 5. Isolation
- The transcript is shared/non-RLS (ADR-021) — fine, it's public content. But the chat is invoked
by a user about a video **in their feed**; confirm the entry path is reachable only for the
requesting user's own videos (the summary view is already RLS-scoped). Chat adds no new
user-data surface (ephemeral), so there's nothing new to RLS — but the test should confirm a
user can only open chat from their own summary view, not arbitrary video ids.
## Tests
- Chat on a video with a stored transcript → answer returned; assert NO caption-fetch / no
YouTube call occurs (the safety property — this is the key assertion).
- Chat invoked on a video with no stored transcript → honest "not available", NO fetch.
- Model switch → re-runs against the same transcript with the selected model; cloud model absent
from the switcher when `TAPIR_CLOUD_FALLBACK_MODEL=""`.
- Truncation honored for a long transcript; the bounded-context note shows.
- Entry is reachable only from the user's own summary view (isolation).
## Out of scope / deferred (record, don't build)
- **Persisted chat history** (per-user, RLS-scoped) — deferred until evidence anyone revisits a
conversation.
- **Show-source / transcript-verification UI** — the natural v2 (ADR-027 records it); v1 is
chat-only/trust-the-model. ADR-021's stored transcript makes v2 cheap when wanted.
- **On-demand fetch** for un-stored videos — would reintroduce the caption-fetch surface the
stored-only constraint removes. Not now.
- Anything that nudges the user to return (ADR-020 — gate contamination).
## Boundaries
Stored-transcript-only (HARD). No rate-gate/fetch surface. No auth changes. No new persisted
state in v1. Reuse the existing gateway client + truncation config; don't build a new model
integration.
+128
View File
@@ -0,0 +1,128 @@
# Spec — Onboarding "wow" burst: better picks, stronger model
**Repo:** tapir · **Size:** medium · **Solo session** (not a swarm).
> **Status: built (v0.25.0, ADR-028).** This supersedes the original investigate-first brief
> (committed as the prior version of this file): Phase 1 was run against the live pilot DB and its
> findings are folded into "Why this exists" below; Phase 2 was built as described here. The one
> brief lever NOT built — the honest "the rest fill in over the coming days" framing copy — is
> listed under *Explicitly NOT in this slice*.
**Why this exists.** A new user's first session decides whether they return (the Stage-0 gate,
VISION.md). On connect, Tapir fires a capped burst (≤`TAPIR_ONBOARD_SUMMARIZE_COUNT`, default 3)
that summarizes the user's newest unsummarized videos so the feed isn't empty (the burst itself
works — wired in `cmd/tapir/discovery.go``cmd/tapir/main.go` `onboard`). A Phase-1
investigation of the live pilot DB found the burst *fires* but delivers a **weak first
impression** for two concrete reasons, and ruled out a third idea:
1. **Picks are junk.** Selection is pure newest-first (`videos.NewestUnsummarizedVideoIDs`,
`ORDER BY published_at DESC`) with **zero quality signal**. Pilot user "Jonte"'s live burst-3
were a stock-ticker **livestream** + two regional news clips — the newest, not the best.
2. **Weakest model on the first impression.** All of Jonte's summaries ran on
`koala/phi4-mini` (the documented weak link — ADR-022 was born from its failures). The
stronger, brain-validated `iguana/gemma4-26b` was never used for the burst.
3. **Cached-first is empty at pilot scale — REJECTED.** The idea (summarize already-cached
transcripts instantly, zero fetch) dies on the numbers: only **11 videos** overlap between the
two pilot users (~3% of each library), **0** cached-and-unsummarized, and a new user's
newest-20 unsummarized are **20/20 NOT cached** — newest-first and cached-first are
structurally incompatible (fresh uploads are exactly what nobody has fetched yet). Not built.
This is a **curation/latency problem for ~3 videos, NOT a throughput/429 problem** — fetching 3
captions is nowhere near the rate limit. Nothing here fetches harder or pressures the rate gate;
it picks the right few videos and runs a better model on them.
Read `CLAUDE.md`, `DECISIONS.md` (esp. ADR-014, ADR-018, ADR-020, ADR-021, ADR-022, ADR-023,
and the new **ADR-028**), and `VISION.md` (the Stage-0 gate) first. TBD — commit directly to
`main`, one logical change per commit, conventional commits, `task check` green before each
commit, `templ generate` if any view changes (none expected).
## Decisions already made (do not reopen)
- **Not a throughput change.** The caption rate gate (ADR-014) is untouched — same pacing, same
priority lane (ADR-026). This slice changes *which* ≤3 videos the burst spends its fetches on
and *which model* summarizes them, never how fast or how many.
- **Cached-first is dropped** (ADR-028, the 3% overlap). The engine's existing read-stored-first
(ADR-021, `resolveTranscript`) stays — it already gives a free instant summary on the rare
cache hit, transparently. We do not *select* for cache hits.
- **has-captions is not a pre-fetch signal.** It is only knowable after a gate fetch (or a cache
hit, ~0 for new videos). Selection can only *avoid known-junk* (Shorts/live/over-long) — it
cannot *guarantee* captions. The spec is honest about this: better odds, not a promise.
- **No credentialed caption fetch** (ADR-010/ADR-026 dead end). **No client extension.**
## 1. Persist `duration_s` at discovery (the enabling change)
The `videos.duration_s` column exists (migration 001) but is **never written** — ADR-023's
`filterLowValue` (`internal/adapters/youtube/youtube.go`) already fetches each candidate's
duration via the cheap quota `videos.list` call, uses it to drop Shorts/live, then **discards
it**. Stop discarding:
- Add `DurationSeconds int` to `domain.Video`.
- In `filterLowValue`, set `DurationSeconds` on each kept video from the `videos.list` `meta`.
- `UpsertVideo` writes `duration_s`, **COALESCE-preserving** a known value (never overwrite a
real duration with 0/unknown), mirroring the `channel_title` backfill stance (migration 014).
- No new migration — the column is already there.
Consequence: a fresh user's connect-triggered discovery pass runs **before** the onboard burst
(`Enqueue`: `run()` then `onboard()`), so duration is populated for the burst's candidates at
connect. Existing rows backfill on their next discovery pass; until then their `duration_s` is
NULL and treated as "unknown" (§2).
## 2. Junk-avoiding burst selection
New store method, RLS-scoped via `withUser`:
```
OnboardBurstVideoIDs(ctx, userID string, limit, minSeconds, maxSeconds int) ([]string, error)
```
- Same base as the old `NewestUnsummarizedVideoIDs`: the user's videos with no summary yet,
`ORDER BY published_at DESC NULLS LAST, seen_at DESC`, `LIMIT limit`.
- **Exclude known-junk**: a row is dropped only when `duration_s IS NOT NULL` **and**
(`duration_s < minSeconds` OR `duration_s > maxSeconds`). A NULL duration is **unknown** — kept
(degrade-open: never starve the burst because metadata is missing), but ordered *after* rows
with a known-good duration so a freshly-enriched good pick wins when both exist.
- `minSeconds` reuses `TAPIR_MIN_VIDEO_SECONDS` (default 60 — the Shorts floor, ADR-023).
`maxSeconds` is new: `TAPIR_ONBOARD_MAX_VIDEO_SECONDS` (default 14400 = 4h) — drops the
multi-hour livestream VODs that pass the live filter once ended.
- `minSeconds<=0` and `maxSeconds<=0` each disable that bound (so `0/0` == the old
newest-first behaviour, the reversibility lever).
- The burst switches to this method; `NewestUnsummarizedVideoIDs` is removed (fully superseded —
`OnboardBurstVideoIDs(., 0, 0)` is identical pure-newest behaviour).
## 3. Stronger model for the burst
The burst summarizes only ≤3 videos, so a slower, stronger model is affordable exactly here.
- New config `TAPIR_ONBOARD_SUMMARIZER_MODEL` (default `iguana/gemma4-26b` — the brain-validated
homelab general-purpose model, already the ADR-022 fallback).
- Build a **burst-specific summarizer chain** that puts the onboard model **first**, then the
standard chain (primary → local fallback → cloud) as resilience, deduped. Wrap it in a
burst-specific `engineProcessor` reusing the same store/transcript-cache/sink — a pure wiring
choice, engine and ports unchanged (Clean Architecture, ADR-003).
- The `onboard` closure uses the burst processor instead of `app.Processor`.
- **Collapse cleanly**: when `OnboardSummarizerModel` is empty or equals `SummarizerModel`, the
onboard path reuses `app.Processor` (no separate chain) — the reversibility lever.
- Local-first preserved: the onboard model is a local alias; the cloud endpoint stays last in the
chain, so a client/NDA deployment with `TAPIR_CLOUD_FALLBACK_MODEL=""` keeps burst content
local too.
## 4. Behaviour spec + docs
- Add scenarios to `docs/use-cases/connect_account.feature` (the connect → burst flow): burst
skips a too-long/live video in favour of a reasonable-length one; burst summarizes with the
stronger model first. Map them in `scenarioCoverage` so `TestScenarioCoverage` stays green.
- Update `docs/architecture/architecture.md` (the onboarding-burst section) to describe the
junk-avoiding selection + the burst model override.
- ADR-028 in `DECISIONS.md` records the rationale (incl. the rejected cached-first lever).
## Success criteria
- `task check` green (fmt, vet, lint, `go test -p 1 ./...`).
- A unit test proves `OnboardBurstVideoIDs` drops a known too-long / sub-min video and keeps a
good one, newest-first, RLS-scoped, unsummarized-only.
- A test proves discovery persists `duration_s` and does not clobber it on re-upsert.
- A test proves the burst chain leads with the onboard model (then the standard chain).
- Config defaults + bounds tested (`OnboardMaxVideoSeconds`, `OnboardSummarizerModel`).
- No change to the rate gate, fetch pacing, or burst cap. `0/0` + empty model == prior behaviour.
## Explicitly NOT in this slice
- Cached-first selection (rejected, ADR-028).
- Any caption-availability *guarantee* (impossible pre-fetch).
- **Honest "taster" framing copy** ("summaries of a few of your videos to get you started — the
rest fill in over the coming days"). A good lever from the original brief, but it's a UI/copy
change with no backend dependency; deferred to a UI pass, tracked as an issue.
- Backfilling `duration_s` for existing rows via a migration (it backfills lazily on discovery).
- Return-nudges / digests (ADR-020: poisons the unprompted-return signal).
- Raising fetch throughput, multi-IP, or Whisper (out of scope; the gate is deliberate).
+69
View File
@@ -0,0 +1,69 @@
Feature: Chat with a video's stored transcript
As a reader whose summary made me want to dig deeper
I want to ask questions about the video without watching it
So that I can go further on the ones worth it, without leaving the reader
# ADR-027. The load-bearing constraint is safety-by-construction: chat runs
# ONLY against an already-stored transcript (ADR-021) and never fetches captions,
# never touches the rate gate, never reaches YouTube. Entry is from the summary
# view of one's OWN video; the conversation is ephemeral (no persisted history).
Background:
Given I have a summarized video with a stored transcript
Scenario: A summary view offers a deeper-dive into the video
When I view the summary
Then I see a "dig deeper" affordance that opens a chat about this video
And it opens the chat in place, below the summary, without leaving the page
Scenario: The summary and the chat are on one page
When I open the chat
Then the summary stays visible alongside the chat
And I can read the summary while I ask questions
Scenario: Ask a question answered from the stored transcript
When I ask a question in the chat
Then the answer is produced from the stored transcript
And no caption fetch and no YouTube call occurs
Scenario: Chat never fetches captions or reaches YouTube
When I ask a question in the chat
Then Tapir reads only the stored transcript
And the caption-fetch and video-fetch paths are never invoked
Scenario: A video with no stored transcript offers no chat
Given a video that has no stored transcript
When I open the chat for it
Then I am told chat is not available
And no fetch is attempted and no model is called
Scenario: The default model is the summary's model and is switchable
When I open the chat
Then the model defaults to the model that produced the summary
And I can switch among the offered chain models
Scenario: Switching models re-runs against the same transcript
When I ask a question with a different chain model selected
Then the chosen model answers
And it answers against the same stored transcript
Scenario: The cloud model is hidden when cloud is disabled
Given the cloud fallback model is disabled
Then the chat switcher offers only local models
Scenario: A long transcript is bounded and the chat says so
Given the stored transcript is longer than the model budget
When I ask a question
Then the answer is produced from a bounded portion
And the chat notes that it worked from a bounded portion
Scenario: A multi-turn conversation is ephemeral
When I ask a follow-up question
Then the prior turn is carried into the answer
And nothing about the conversation is written to the database
Scenario: Chat is reachable only from my own summary view
Given another user has a summarized video with a stored transcript
When I try to open the chat for their video
Then I get a not-found response
And no model is called
+9 -2
View File
@@ -24,12 +24,19 @@ Feature: Connect and manage video accounts
Then a discovery pass for my account is triggered right away Then a discovery pass for my account is triggered right away
And I do not have to wait for the next scheduled pass to see my videos And I do not have to wait for the next scheduled pass to see my videos
Scenario: Connecting summarizes my newest videos right away Scenario: Connecting summarizes my best recent videos right away
Given I have no connected video accounts Given I have no connected video accounts
When I connect my YouTube account When I connect my YouTube account
Then up to the onboarding cap of my newest videos are summarized through the rate gate Then up to the onboarding cap of my newest likely-good videos are summarized through the rate gate
And videos whose known duration is too short or too long are skipped
And the rest are left to the scheduled recency-bounded pass And the rest are left to the scheduled recency-bounded pass
Scenario: The onboarding burst summarizes with a stronger model
Given I have no connected video accounts
When I connect my YouTube account
Then the burst summarizes with the stronger onboarding model first
And the standard summarizer chain still follows as a fallback
Scenario: Tokens are never stored in the clear Scenario: Tokens are never stored in the clear
When I connect any video account When I connect any video account
Then no OAuth token value is stored in the database Then no OAuth token value is stored in the database
+46
View File
@@ -0,0 +1,46 @@
Feature: Observability — timing and metrics for performance and UX (ADR-030, #15)
As the maintainer running Tapir for pilot users
I want timing and Prometheus metrics for the activities that drive performance and UX
So that I can see latency, model behaviour, and usage and feed the Stage-0 eval gate
# AI metrics are the priority (ADR-030 R3). Each scenario maps to a Go test in
# test/acceptance/scenario_coverage_test.go (the BDD name-coverage gate).
Scenario: Summarization latency is recorded per endpoint
Given the summarizer runs a transcript through its endpoint chain
When an endpoint returns a parseable summary
Then the summarize latency is recorded with the model, outcome "success", and whether it was a fallback
Scenario: A failing summarizer endpoint records its failure outcome
Given the summarizer runs a transcript through its endpoint chain
When an endpoint errors or returns unparseable output
Then the summarize latency is recorded with outcome "error" or "parse_error" before the chain advances
Scenario: Caption fetch latency is recorded by outcome
Given a caption fetch is attempted for a video
When it resolves to captions, no captions, or a rate limit
Then the caption-fetch latency is recorded labelled by that outcome
Scenario: LLM token usage is recorded from the completion
Given an LLM completion returns a usage block with prompt and completion tokens
When the client finishes the call
Then the prompt and completion tokens are recorded for that model
Scenario: Q&A answer latency is recorded
Given a user asks a question about a video
When the answer is produced from the stored transcript
Then the chat answer latency is recorded for the answering model
Scenario: HTTP requests are counted by route, method, and status
Given the metrics HTTP middleware wraps the app
When a request is served against a registered route
Then it is counted and timed under the bounded route pattern, not the raw path
Scenario: A successful login is counted
Given a user completes the OIDC callback and a session is established
Then the login counter is incremented
Scenario: The metrics endpoint is not on the public app port
Given the service is running
When the public app mux is inspected
Then it exposes no /metrics route metrics are served on the dedicated metrics port only
+11 -1
View File
@@ -9,23 +9,33 @@ require (
github.com/go-jose/go-jose/v4 v4.1.4 github.com/go-jose/go-jose/v4 v4.1.4
github.com/golang-migrate/migrate/v4 v4.19.1 github.com/golang-migrate/migrate/v4 v4.19.1
github.com/jackc/pgx/v5 v5.9.2 github.com/jackc/pgx/v5 v5.9.2
github.com/prometheus/client_golang v1.23.2
github.com/prometheus/client_model v0.6.2
github.com/stretchr/testify v1.11.1 github.com/stretchr/testify v1.11.1
golang.org/x/crypto v0.45.0
golang.org/x/oauth2 v0.36.0 golang.org/x/oauth2 v0.36.0
golang.org/x/time v0.15.0 golang.org/x/time v0.15.0
) )
require ( require (
github.com/beorn7/perks v1.0.1 // indirect
github.com/cespare/xxhash/v2 v2.3.0 // indirect
github.com/davecgh/go-spew v1.1.2-0.20180830191138-d8f796af33cc // indirect github.com/davecgh/go-spew v1.1.2-0.20180830191138-d8f796af33cc // indirect
github.com/jackc/pgerrcode v0.0.0-20220416144525-469b46aa5efa // indirect github.com/jackc/pgerrcode v0.0.0-20220416144525-469b46aa5efa // indirect
github.com/jackc/pgpassfile v1.0.0 // indirect github.com/jackc/pgpassfile v1.0.0 // indirect
github.com/jackc/pgservicefile v0.0.0-20240606120523-5a60cdf6a761 // indirect github.com/jackc/pgservicefile v0.0.0-20240606120523-5a60cdf6a761 // indirect
github.com/jackc/puddle/v2 v2.2.2 // indirect github.com/jackc/puddle/v2 v2.2.2 // indirect
github.com/kylelemons/godebug v1.1.0 // indirect
github.com/lib/pq v1.10.9 // indirect github.com/lib/pq v1.10.9 // indirect
github.com/munnerz/goautoneg v0.0.0-20191010083416-a7dc8b61c822 // indirect
github.com/pmezard/go-difflib v1.0.1-0.20181226105442-5d4384ee4fb2 // indirect github.com/pmezard/go-difflib v1.0.1-0.20181226105442-5d4384ee4fb2 // indirect
github.com/prometheus/common v0.66.1 // indirect
github.com/prometheus/procfs v0.16.1 // indirect
github.com/rogpeppe/go-internal v1.15.0 // indirect github.com/rogpeppe/go-internal v1.15.0 // indirect
github.com/xi2/xz v0.0.0-20171230120015-48954b6210f8 // indirect github.com/xi2/xz v0.0.0-20171230120015-48954b6210f8 // indirect
go.yaml.in/yaml/v2 v2.4.2 // indirect
golang.org/x/sync v0.18.0 // indirect golang.org/x/sync v0.18.0 // indirect
golang.org/x/sys v0.41.0 // indirect
golang.org/x/text v0.31.0 // indirect golang.org/x/text v0.31.0 // indirect
google.golang.org/protobuf v1.36.8 // indirect
gopkg.in/yaml.v3 v3.0.1 // indirect gopkg.in/yaml.v3 v3.0.1 // indirect
) )
+26 -6
View File
@@ -4,6 +4,10 @@ github.com/Microsoft/go-winio v0.6.2 h1:F2VQgta7ecxGYO8k3ZZz3RS8fVIXVxONVUPlNERo
github.com/Microsoft/go-winio v0.6.2/go.mod h1:yd8OoFMLzJbo9gZq8j5qaps8bJ9aShtEA8Ipt1oGCvU= github.com/Microsoft/go-winio v0.6.2/go.mod h1:yd8OoFMLzJbo9gZq8j5qaps8bJ9aShtEA8Ipt1oGCvU=
github.com/a-h/templ v0.3.1020 h1:ypAT/L5ySWEnZ6Zft/5yfoWXYYkhFNvEFOeeqecg4tw= github.com/a-h/templ v0.3.1020 h1:ypAT/L5ySWEnZ6Zft/5yfoWXYYkhFNvEFOeeqecg4tw=
github.com/a-h/templ v0.3.1020/go.mod h1:A2DlK61v+K+NRoGnhmYbNYVmtYHcFO5/AisMvBdDxTM= github.com/a-h/templ v0.3.1020/go.mod h1:A2DlK61v+K+NRoGnhmYbNYVmtYHcFO5/AisMvBdDxTM=
github.com/beorn7/perks v1.0.1 h1:VlbKKnNfV8bJzeqoa4cOKqO6bYr3WgKZxO8Z16+hsOM=
github.com/beorn7/perks v1.0.1/go.mod h1:G2ZrVWU2WbWT9wwq4/hrbKbnv/1ERSJQ0ibhJ6rlkpw=
github.com/cespare/xxhash/v2 v2.3.0 h1:UL815xU9SqsFlibzuggzjXhog7bL6oX9BbNZnL2UFvs=
github.com/cespare/xxhash/v2 v2.3.0/go.mod h1:VGX0DQ3Q6kWi7AoAeZDth3/j3BFtOZR5XLFGgcrjCOs=
github.com/containerd/errdefs v1.0.0 h1:tg5yIfIlQIrxYtu9ajqY42W3lpS19XqdxRQeEwYG8PI= github.com/containerd/errdefs v1.0.0 h1:tg5yIfIlQIrxYtu9ajqY42W3lpS19XqdxRQeEwYG8PI=
github.com/containerd/errdefs v1.0.0/go.mod h1:+YBYIdtsnF4Iw6nWZhJcqGSg/dwvV7tyJ/kCkyJ2k+M= github.com/containerd/errdefs v1.0.0/go.mod h1:+YBYIdtsnF4Iw6nWZhJcqGSg/dwvV7tyJ/kCkyJ2k+M=
github.com/containerd/errdefs/pkg v0.3.0 h1:9IKJ06FvyNlexW690DXuQNx2KA2cUJXx151Xdx3ZPPE= github.com/containerd/errdefs/pkg v0.3.0 h1:9IKJ06FvyNlexW690DXuQNx2KA2cUJXx151Xdx3ZPPE=
@@ -37,8 +41,8 @@ github.com/gogo/protobuf v1.3.2 h1:Ov1cvc58UF3b5XjBnZv7+opcTcQFZebYjWzi34vdm4Q=
github.com/gogo/protobuf v1.3.2/go.mod h1:P1XiOD3dCwIKUDQYPy72D8LYyHL2YPYrpS2s69NZV8Q= github.com/gogo/protobuf v1.3.2/go.mod h1:P1XiOD3dCwIKUDQYPy72D8LYyHL2YPYrpS2s69NZV8Q=
github.com/golang-migrate/migrate/v4 v4.19.1 h1:OCyb44lFuQfYXYLx1SCxPZQGU7mcaZ7gH9yH4jSFbBA= github.com/golang-migrate/migrate/v4 v4.19.1 h1:OCyb44lFuQfYXYLx1SCxPZQGU7mcaZ7gH9yH4jSFbBA=
github.com/golang-migrate/migrate/v4 v4.19.1/go.mod h1:CTcgfjxhaUtsLipnLoQRWCrjYXycRz/g5+RWDuYgPrE= github.com/golang-migrate/migrate/v4 v4.19.1/go.mod h1:CTcgfjxhaUtsLipnLoQRWCrjYXycRz/g5+RWDuYgPrE=
github.com/google/go-cmp v0.6.0 h1:ofyhxvXcZhMsU5ulbFiLKl/XBFqE1GSq7atu8tAmTRI= github.com/google/go-cmp v0.7.0 h1:wk8382ETsv4JYUZwIsn6YpYiWiBsYLSJiTsyBybVuN8=
github.com/google/go-cmp v0.6.0/go.mod h1:17dUlkBOakJ0+DkrSSNjCkIjxS6bF9zb3elmeNGIjoY= github.com/google/go-cmp v0.7.0/go.mod h1:pXiqmnSA92OHEEa9HXL2W4E7lf9JzCmGVUdgjX3N/iU=
github.com/jackc/pgerrcode v0.0.0-20220416144525-469b46aa5efa h1:s+4MhCQ6YrzisK6hFJUX53drDT4UsSW3DEhKn0ifuHw= github.com/jackc/pgerrcode v0.0.0-20220416144525-469b46aa5efa h1:s+4MhCQ6YrzisK6hFJUX53drDT4UsSW3DEhKn0ifuHw=
github.com/jackc/pgerrcode v0.0.0-20220416144525-469b46aa5efa/go.mod h1:a/s9Lp5W7n/DD0VrVoyJ00FbP2ytTPDVOivvn2bMlds= github.com/jackc/pgerrcode v0.0.0-20220416144525-469b46aa5efa/go.mod h1:a/s9Lp5W7n/DD0VrVoyJ00FbP2ytTPDVOivvn2bMlds=
github.com/jackc/pgpassfile v1.0.0 h1:/6Hmqy13Ss2zCq62VdNG8tM1wchn8zjSGOBJ6icpsIM= github.com/jackc/pgpassfile v1.0.0 h1:/6Hmqy13Ss2zCq62VdNG8tM1wchn8zjSGOBJ6icpsIM=
@@ -49,10 +53,14 @@ github.com/jackc/pgx/v5 v5.9.2 h1:3ZhOzMWnR4yJ+RW1XImIPsD1aNSz4T4fyP7zlQb56hw=
github.com/jackc/pgx/v5 v5.9.2/go.mod h1:mal1tBGAFfLHvZzaYh77YS/eC6IX9OWbRV1QIIM0Jn4= github.com/jackc/pgx/v5 v5.9.2/go.mod h1:mal1tBGAFfLHvZzaYh77YS/eC6IX9OWbRV1QIIM0Jn4=
github.com/jackc/puddle/v2 v2.2.2 h1:PR8nw+E/1w0GLuRFSmiioY6UooMp6KJv0/61nB7icHo= github.com/jackc/puddle/v2 v2.2.2 h1:PR8nw+E/1w0GLuRFSmiioY6UooMp6KJv0/61nB7icHo=
github.com/jackc/puddle/v2 v2.2.2/go.mod h1:vriiEXHvEE654aYKXXjOvZM39qJ0q+azkZFrfEOc3H4= github.com/jackc/puddle/v2 v2.2.2/go.mod h1:vriiEXHvEE654aYKXXjOvZM39qJ0q+azkZFrfEOc3H4=
github.com/kr/pretty v0.3.0 h1:WgNl7dwNpEZ6jJ9k1snq4pZsg7DOEN8hP9Xw0Tsjwk0= github.com/klauspost/compress v1.18.0 h1:c/Cqfb0r+Yi+JtIEq73FWXVkRonBlf0CRNYc8Zttxdo=
github.com/kr/pretty v0.3.0/go.mod h1:640gp4NfQd8pI5XOwp5fnNeVWj67G7CFk/SaSQn7NBk= github.com/klauspost/compress v1.18.0/go.mod h1:2Pp+KzxcywXVXMr50+X0Q/Lsb43OQHYWRCY2AiWywWQ=
github.com/kr/pretty v0.3.1 h1:flRD4NNwYAUpkphVc1HcthR4KEIFJ65n8Mw5qdRn3LE=
github.com/kr/pretty v0.3.1/go.mod h1:hoEshYVHaxMs3cyo3Yncou5ZscifuDolrwPKZanG3xk=
github.com/kr/text v0.2.0 h1:5Nx0Ya0ZqY2ygV366QzturHI13Jq95ApcVaJBhpS+AY= github.com/kr/text v0.2.0 h1:5Nx0Ya0ZqY2ygV366QzturHI13Jq95ApcVaJBhpS+AY=
github.com/kr/text v0.2.0/go.mod h1:eLer722TekiGuMkidMxC/pM04lWEeraHUUmBw8l2grE= github.com/kr/text v0.2.0/go.mod h1:eLer722TekiGuMkidMxC/pM04lWEeraHUUmBw8l2grE=
github.com/kylelemons/godebug v1.1.0 h1:RPNrshWIDI6G2gRW9EHilWtl7Z6Sb1BR0xunSBf0SNc=
github.com/kylelemons/godebug v1.1.0/go.mod h1:9/0rRGxNHcop5bhtWyNeEfOS8JIWk580+fNqagV/RAw=
github.com/lib/pq v1.10.9 h1:YXG7RB+JIjhP29X+OtkiDnYaXQwpS4JEWq7dtCCRUEw= github.com/lib/pq v1.10.9 h1:YXG7RB+JIjhP29X+OtkiDnYaXQwpS4JEWq7dtCCRUEw=
github.com/lib/pq v1.10.9/go.mod h1:AlVN5x4E4T544tWzH6hKfbfQvm3HdbOxrmggDNAPY9o= github.com/lib/pq v1.10.9/go.mod h1:AlVN5x4E4T544tWzH6hKfbfQvm3HdbOxrmggDNAPY9o=
github.com/moby/docker-image-spec v1.3.1 h1:jMKff3w6PgbfSa69GfNg+zN/XLhfXJGnEx3Nl2EsFP0= github.com/moby/docker-image-spec v1.3.1 h1:jMKff3w6PgbfSa69GfNg+zN/XLhfXJGnEx3Nl2EsFP0=
@@ -61,6 +69,8 @@ github.com/moby/term v0.5.0 h1:xt8Q1nalod/v7BqbG21f8mQPqH+xAaC9C3N3wfWbVP0=
github.com/moby/term v0.5.0/go.mod h1:8FzsFHVUBGZdbDsJw/ot+X+d5HLUbvklYLJ9uGfcI3Y= github.com/moby/term v0.5.0/go.mod h1:8FzsFHVUBGZdbDsJw/ot+X+d5HLUbvklYLJ9uGfcI3Y=
github.com/morikuni/aec v1.0.0 h1:nP9CBfwrvYnBRgY6qfDQkygYDmYwOilePFkwzv4dU8A= github.com/morikuni/aec v1.0.0 h1:nP9CBfwrvYnBRgY6qfDQkygYDmYwOilePFkwzv4dU8A=
github.com/morikuni/aec v1.0.0/go.mod h1:BbKIizmSmc5MMPqRYbxO4ZU0S0+P200+tUnFx7PXmsc= github.com/morikuni/aec v1.0.0/go.mod h1:BbKIizmSmc5MMPqRYbxO4ZU0S0+P200+tUnFx7PXmsc=
github.com/munnerz/goautoneg v0.0.0-20191010083416-a7dc8b61c822 h1:C3w9PqII01/Oq1c1nUAm88MOHcQC9l5mIlSMApZMrHA=
github.com/munnerz/goautoneg v0.0.0-20191010083416-a7dc8b61c822/go.mod h1:+n7T8mK8HuQTcFwEeznm/DIxMOiR9yIdICNftLE1DvQ=
github.com/opencontainers/go-digest v1.0.0 h1:apOUWs51W5PlhuyGyz9FCeeBIOUDA/6nW8Oi/yOhh5U= github.com/opencontainers/go-digest v1.0.0 h1:apOUWs51W5PlhuyGyz9FCeeBIOUDA/6nW8Oi/yOhh5U=
github.com/opencontainers/go-digest v1.0.0/go.mod h1:0JzlMkj0TRzQZfJkVvzbP0HBR3IKzErnv2BNG4W4MAM= github.com/opencontainers/go-digest v1.0.0/go.mod h1:0JzlMkj0TRzQZfJkVvzbP0HBR3IKzErnv2BNG4W4MAM=
github.com/opencontainers/image-spec v1.1.0 h1:8SG7/vwALn54lVB/0yZ/MMwhFrPYtpEHQb2IpWsCzug= github.com/opencontainers/image-spec v1.1.0 h1:8SG7/vwALn54lVB/0yZ/MMwhFrPYtpEHQb2IpWsCzug=
@@ -70,6 +80,14 @@ github.com/pkg/errors v0.9.1/go.mod h1:bwawxfHBFNV+L2hUp1rHADufV3IMtnDRdf1r5NINE
github.com/pmezard/go-difflib v1.0.0/go.mod h1:iKH77koFhYxTK1pcRnkKkqfTogsbg7gZNVY4sRDYZ/4= github.com/pmezard/go-difflib v1.0.0/go.mod h1:iKH77koFhYxTK1pcRnkKkqfTogsbg7gZNVY4sRDYZ/4=
github.com/pmezard/go-difflib v1.0.1-0.20181226105442-5d4384ee4fb2 h1:Jamvg5psRIccs7FGNTlIRMkT8wgtp5eCXdBlqhYGL6U= github.com/pmezard/go-difflib v1.0.1-0.20181226105442-5d4384ee4fb2 h1:Jamvg5psRIccs7FGNTlIRMkT8wgtp5eCXdBlqhYGL6U=
github.com/pmezard/go-difflib v1.0.1-0.20181226105442-5d4384ee4fb2/go.mod h1:iKH77koFhYxTK1pcRnkKkqfTogsbg7gZNVY4sRDYZ/4= github.com/pmezard/go-difflib v1.0.1-0.20181226105442-5d4384ee4fb2/go.mod h1:iKH77koFhYxTK1pcRnkKkqfTogsbg7gZNVY4sRDYZ/4=
github.com/prometheus/client_golang v1.23.2 h1:Je96obch5RDVy3FDMndoUsjAhG5Edi49h0RJWRi/o0o=
github.com/prometheus/client_golang v1.23.2/go.mod h1:Tb1a6LWHB3/SPIzCoaDXI4I8UHKeFTEQ1YCr+0Gyqmg=
github.com/prometheus/client_model v0.6.2 h1:oBsgwpGs7iVziMvrGhE53c/GrLUsZdHnqNwqPLxwZyk=
github.com/prometheus/client_model v0.6.2/go.mod h1:y3m2F6Gdpfy6Ut/GBsUqTWZqCUvMVzSfMLjcu6wAwpE=
github.com/prometheus/common v0.66.1 h1:h5E0h5/Y8niHc5DlaLlWLArTQI7tMrsfQjHV+d9ZoGs=
github.com/prometheus/common v0.66.1/go.mod h1:gcaUsgf3KfRSwHY4dIMXLPV0K/Wg1oZ8+SbZk/HH/dA=
github.com/prometheus/procfs v0.16.1 h1:hZ15bTNuirocR6u0JZ6BAHHmwS1p8B4P6MRqxtzMyRg=
github.com/prometheus/procfs v0.16.1/go.mod h1:teAbpZRB1iIAJYREa1LsoWUXykVXA1KlTmWl8x/U+Is=
github.com/rogpeppe/go-internal v1.15.0 h1:D0RCU5rMAp+SpgkiNdrjfJ+LX4J1M32V2NeCY7EJ6hc= github.com/rogpeppe/go-internal v1.15.0 h1:D0RCU5rMAp+SpgkiNdrjfJ+LX4J1M32V2NeCY7EJ6hc=
github.com/rogpeppe/go-internal v1.15.0/go.mod h1:DrUVZyrJU+txYW5/1kwtXQSMFio52ZOxX7yM1VHvnxs= github.com/rogpeppe/go-internal v1.15.0/go.mod h1:DrUVZyrJU+txYW5/1kwtXQSMFio52ZOxX7yM1VHvnxs=
github.com/stretchr/objx v0.1.0/go.mod h1:HFkY916IF+rwdDfMAkV7OtwuqBVzrE8GR6GFx+wExME= github.com/stretchr/objx v0.1.0/go.mod h1:HFkY916IF+rwdDfMAkV7OtwuqBVzrE8GR6GFx+wExME=
@@ -91,8 +109,8 @@ go.opentelemetry.io/otel/trace v1.37.0 h1:HLdcFNbRQBE2imdSEgm/kwqmQj1Or1l/7bW6mx
go.opentelemetry.io/otel/trace v1.37.0/go.mod h1:TlgrlQ+PtQO5XFerSPUYG0JSgGyryXewPGyayAWSBS0= go.opentelemetry.io/otel/trace v1.37.0/go.mod h1:TlgrlQ+PtQO5XFerSPUYG0JSgGyryXewPGyayAWSBS0=
go.uber.org/goleak v1.3.0 h1:2K3zAYmnTNqV73imy9J1T3WC+gmCePx2hEGkimedGto= go.uber.org/goleak v1.3.0 h1:2K3zAYmnTNqV73imy9J1T3WC+gmCePx2hEGkimedGto=
go.uber.org/goleak v1.3.0/go.mod h1:CoHD4mav9JJNrW/WLlf7HGZPjdw8EucARQHekz1X6bE= go.uber.org/goleak v1.3.0/go.mod h1:CoHD4mav9JJNrW/WLlf7HGZPjdw8EucARQHekz1X6bE=
golang.org/x/crypto v0.45.0 h1:jMBrvKuj23MTlT0bQEOBcAE0mjg8mK9RXFhRH6nyF3Q= go.yaml.in/yaml/v2 v2.4.2 h1:DzmwEr2rDGHl7lsFgAHxmNz/1NlQ7xLIrlN2h5d1eGI=
golang.org/x/crypto v0.45.0/go.mod h1:XTGrrkGJve7CYK7J8PEww4aY7gM3qMCElcJQ8n8JdX4= go.yaml.in/yaml/v2 v2.4.2/go.mod h1:081UH+NErpNdqlCXm3TtEran0rJZGxAYx9hb/ELlsPU=
golang.org/x/oauth2 v0.36.0 h1:peZ/1z27fi9hUOFCAZaHyrpWG5lwe0RJEEEeH0ThlIs= golang.org/x/oauth2 v0.36.0 h1:peZ/1z27fi9hUOFCAZaHyrpWG5lwe0RJEEEeH0ThlIs=
golang.org/x/oauth2 v0.36.0/go.mod h1:YDBUJMTkDnJS+A4BP4eZBjCqtokkg1hODuPjwiGPO7Q= golang.org/x/oauth2 v0.36.0/go.mod h1:YDBUJMTkDnJS+A4BP4eZBjCqtokkg1hODuPjwiGPO7Q=
golang.org/x/sync v0.18.0 h1:kr88TuHDroi+UVf+0hZnirlk8o8T+4MrK6mr60WkH/I= golang.org/x/sync v0.18.0 h1:kr88TuHDroi+UVf+0hZnirlk8o8T+4MrK6mr60WkH/I=
@@ -103,6 +121,8 @@ golang.org/x/text v0.31.0 h1:aC8ghyu4JhP8VojJ2lEHBnochRno1sgL6nEi9WGFGMM=
golang.org/x/text v0.31.0/go.mod h1:tKRAlv61yKIjGGHX/4tP1LTbc13YSec1pxVEWXzfoeM= golang.org/x/text v0.31.0/go.mod h1:tKRAlv61yKIjGGHX/4tP1LTbc13YSec1pxVEWXzfoeM=
golang.org/x/time v0.15.0 h1:bbrp8t3bGUeFOx08pvsMYRTCVSMk89u4tKbNOZbp88U= golang.org/x/time v0.15.0 h1:bbrp8t3bGUeFOx08pvsMYRTCVSMk89u4tKbNOZbp88U=
golang.org/x/time v0.15.0/go.mod h1:Y4YMaQmXwGQZoFaVFk4YpCt4FLQMYKZe9oeV/f4MSno= golang.org/x/time v0.15.0/go.mod h1:Y4YMaQmXwGQZoFaVFk4YpCt4FLQMYKZe9oeV/f4MSno=
google.golang.org/protobuf v1.36.8 h1:xHScyCOEuuwZEc6UtSOvPbAT4zRh0xcNRYekJwfqyMc=
google.golang.org/protobuf v1.36.8/go.mod h1:fuxRtAxBytpl4zzqUh6/eyUujkJdNiuEkXntxiD/uRU=
gopkg.in/check.v1 v0.0.0-20161208181325-20d25e280405/go.mod h1:Co6ibVJAznAaIkqp8huTwlJQCZ016jof/cbN4VW5Yz0= gopkg.in/check.v1 v0.0.0-20161208181325-20d25e280405/go.mod h1:Co6ibVJAznAaIkqp8huTwlJQCZ016jof/cbN4VW5Yz0=
gopkg.in/check.v1 v1.0.0-20201130134442-10cb98267c6c h1:Hei/4ADfdWqJk1ZMxUNpqntNwaWcugrBjAiHlqqRiVk= gopkg.in/check.v1 v1.0.0-20201130134442-10cb98267c6c h1:Hei/4ADfdWqJk1ZMxUNpqntNwaWcugrBjAiHlqqRiVk=
gopkg.in/check.v1 v1.0.0-20201130134442-10cb98267c6c/go.mod h1:JHkPIbrfpd72SG/EVd6muEfDQjcINNoR0C8j2r3qZ4Q= gopkg.in/check.v1 v1.0.0-20201130134442-10cb98267c6c/go.mod h1:JHkPIbrfpd72SG/EVd6muEfDQjcINNoR0C8j2r3qZ4Q=
+179
View File
@@ -0,0 +1,179 @@
// Package chat implements the per-video deeper-dive chat (ADR-027): a read-only
// QA over a video's ALREADY-STORED transcript (ADR-021). It is the enforcement
// point for the feature's load-bearing safety property — stored-transcript-only:
// the Service has NO VideoSource and NO caption-fetch dependency, only a
// Completer factory, so it CANNOT reach YouTube or the rate gate by construction.
// The caller supplies the stored transcript text; chat never fetches.
//
// It reuses the same LiteLLM gateway as the summarizer (a chat is a different
// call, not a new integration) and the same transcript-truncation discipline
// (TAPIR_MAX_TRANSCRIPT_CHARS) so a long transcript fits a small-context model.
package chat
import (
"context"
"fmt"
"log/slog"
"strings"
"time"
"unicode/utf8"
"gitea.d-ma.be/mathias/tapir/internal/metrics"
)
// Completer is the minimal LLM chat surface the Service needs. *llm.Client
// satisfies it; tests use a fake. It is the SAME surface the summarizer uses.
type Completer interface {
Complete(ctx context.Context, system, user string) (string, error)
}
// Turn is one completed exchange in an ephemeral, session-only conversation
// (ADR-027 v1: nothing is persisted).
type Turn struct {
Question string
Answer string
}
// Request is one chat turn: the chosen model, the stored transcript text, the
// prior turns (for multi-turn context within the session), and the new question.
type Request struct {
Model string
Transcript string
History []Turn
Question string
}
// Reply is the model's answer plus whether the transcript was bounded to fit the
// model context (so the UI can be honest that an answer about the tail of a long
// video may be incomplete).
type Reply struct {
Answer string
Truncated bool
}
// Service answers questions against a stored transcript via a switchable set of
// models. models is the ordered, local-first list offered to the user (the cloud
// model is simply absent when disabled — see cmd wiring); maxChars bounds the
// transcript sent to any model (0 = unbounded). newClient builds a Completer for
// a chosen model alias (the same gateway, a different alias).
type Service struct {
newClient func(model string) Completer
models []string
maxChars int
}
// New constructs a Service. models must be non-empty and already filtered to the
// offerable set (cloud excluded when disabled) and de-duplicated by the caller.
func New(newClient func(model string) Completer, models []string, maxChars int) *Service {
return &Service{newClient: newClient, models: models, maxChars: maxChars}
}
// Models returns a copy of the offerable model list (local-first order).
func (s *Service) Models() []string {
out := make([]string, len(s.models))
copy(out, s.models)
return out
}
// offers reports whether model is in the offerable set — the guard that keeps an
// arbitrary, un-offered alias (e.g. a forged form value) from reaching the gateway.
func (s *Service) offers(model string) bool {
for _, m := range s.models {
if m == model {
return true
}
}
return false
}
// DefaultModel resolves the model a fresh chat opens with: the summary's own
// model when it is still an offered option (the ADR-027 default — chat continues
// in the model that produced the summary), otherwise the first offered model.
// Returns "" only when no models are configured.
func (s *Service) DefaultModel(summaryModel string) string {
if summaryModel != "" && s.offers(summaryModel) {
return summaryModel
}
if len(s.models) > 0 {
return s.models[0]
}
return ""
}
// Answer runs one chat turn. The model is forced back to a default if the request
// names an un-offered alias, so chat can never call the gateway with an arbitrary
// model. The transcript is truncated up front (reporting whether it was cut) and
// passed as system context; the running conversation is the user message.
func (s *Service) Answer(ctx context.Context, req Request) (Reply, error) {
if len(s.models) == 0 {
return Reply{}, fmt.Errorf("chat: no models configured")
}
model := req.Model
if !s.offers(model) {
model = s.DefaultModel("")
}
transcript, truncated := truncate(req.Transcript, s.maxChars)
system := buildSystem(transcript, truncated)
user := buildUser(req.History, req.Question)
start := time.Now()
out, err := s.newClient(model).Complete(ctx, system, user)
if err != nil {
return Reply{}, fmt.Errorf("chat: %s: %w", model, err)
}
dur := time.Since(start)
metrics.ObserveChat(model, dur)
slog.Default().Info("chat answer", "model", model, "elapsed_ms", dur.Milliseconds())
answer := strings.TrimSpace(out)
if answer == "" {
return Reply{}, fmt.Errorf("chat: %s returned an empty answer", model)
}
return Reply{Answer: answer, Truncated: truncated}, nil
}
const systemPreamble = `You are Tapir, answering questions about ONE video using ONLY the transcript below.
Ground every answer in the transcript. If the transcript does not contain the answer, say so plainly rather than guessing.`
const truncatedNote = `
The transcript below is truncated to fit the model — if a question seems to concern something missing, note it may be beyond the available portion.`
// buildSystem frames the model as a transcript-grounded QA assistant and embeds
// the (possibly truncated) transcript as context.
func buildSystem(transcript string, truncated bool) string {
var b strings.Builder
b.WriteString(systemPreamble)
if truncated {
b.WriteString(truncatedNote)
}
b.WriteString("\n\nTranscript:\n")
b.WriteString(transcript)
return b.String()
}
// buildUser renders the running conversation as the user message: prior turns as
// Q/A pairs followed by the new question. Folding history into one message keeps
// the Completer surface (a single system+user call) unchanged — no new llm method.
func buildUser(history []Turn, question string) string {
var b strings.Builder
for _, t := range history {
fmt.Fprintf(&b, "Q: %s\nA: %s\n\n", t.Question, t.Answer)
}
fmt.Fprintf(&b, "Q: %s", question)
return b.String()
}
// truncate caps content to max bytes on a UTF-8 rune boundary, reporting whether
// it cut. It mirrors the summarizer's truncation discipline (ADR-022) but returns
// the cut flag so the chat UI can be honest about a bounded transcript. A
// non-positive max (or content already within budget) returns content unchanged.
func truncate(content string, max int) (string, bool) {
if max <= 0 || len(content) <= max {
return content, false
}
cut := max
for cut > 0 && !utf8.RuneStart(content[cut]) {
cut--
}
return content[:cut], true
}
+164
View File
@@ -0,0 +1,164 @@
package chat
import (
"context"
"errors"
"strings"
"testing"
)
// recordingCompleter captures the system+user it was asked with and returns a
// canned answer (or error). It also records which model alias built it.
type recordingCompleter struct {
model string
lastSystem string
lastUser string
answer string
err error
calls *int
}
func (c *recordingCompleter) Complete(_ context.Context, system, user string) (string, error) {
*c.calls++
c.lastSystem = system
c.lastUser = user
if c.err != nil {
return "", c.err
}
return c.answer, nil
}
// factory builds a recordingCompleter per model and records the last one built so
// the test can assert which model alias was actually used for the gateway call.
type factory struct {
answer string
err error
calls int
used *recordingCompleter
}
func (f *factory) make(model string) Completer {
c := &recordingCompleter{model: model, answer: f.answer, err: f.err, calls: &f.calls}
f.used = c
return c
}
func TestModelsAreOfferedLocalFirstAndCopied(t *testing.T) {
f := &factory{answer: "ok"}
s := New(f.make, []string{"koala/phi4-mini", "iguana/gemma4-26b"}, 0)
got := s.Models()
want := []string{"koala/phi4-mini", "iguana/gemma4-26b"}
if len(got) != len(want) || got[0] != want[0] || got[1] != want[1] {
t.Fatalf("Models() = %v, want %v", got, want)
}
// Mutating the returned slice must not corrupt the Service's list.
got[0] = "tampered"
if s.Models()[0] != "koala/phi4-mini" {
t.Fatal("Models() leaked its backing slice")
}
}
func TestDefaultModelIsTheSummarysModelWhenOffered(t *testing.T) {
f := &factory{answer: "ok"}
s := New(f.make, []string{"koala/phi4-mini", "iguana/gemma4-26b", "berget/mistral-small"}, 0)
if got := s.DefaultModel("iguana/gemma4-26b"); got != "iguana/gemma4-26b" {
t.Fatalf("DefaultModel(summary) = %q, want the summary's model", got)
}
// A summary model no longer offered (e.g. cloud disabled) falls back to first.
if got := s.DefaultModel("berget/old-model"); got != "koala/phi4-mini" {
t.Fatalf("DefaultModel(un-offered) = %q, want the first offered model", got)
}
// No summary model recorded → first offered.
if got := s.DefaultModel(""); got != "koala/phi4-mini" {
t.Fatalf("DefaultModel(\"\") = %q, want the first offered model", got)
}
}
func TestAnswerGroundsOnTranscriptAndCarriesHistory(t *testing.T) {
f := &factory{answer: " The video is about attention. "}
s := New(f.make, []string{"koala/phi4-mini"}, 0)
reply, err := s.Answer(context.Background(), Request{
Model: "koala/phi4-mini",
Transcript: "ATTENTION-TRANSCRIPT-MARKER",
History: []Turn{{Question: "who", Answer: "the host"}},
Question: "what is it about",
})
if err != nil {
t.Fatalf("Answer: %v", err)
}
if reply.Answer != "The video is about attention." {
t.Fatalf("answer not trimmed: %q", reply.Answer)
}
if reply.Truncated {
t.Fatal("short transcript must not report truncated")
}
// The transcript rides in the system prompt; the conversation in the user msg.
if !strings.Contains(f.used.lastSystem, "ATTENTION-TRANSCRIPT-MARKER") {
t.Fatal("transcript not grounded into the system prompt")
}
if !strings.Contains(f.used.lastUser, "Q: who") || !strings.Contains(f.used.lastUser, "A: the host") {
t.Fatalf("history not carried into the user message: %q", f.used.lastUser)
}
if !strings.Contains(f.used.lastUser, "what is it about") {
t.Fatal("new question missing from the user message")
}
}
func TestAnswerTruncatesLongTranscriptAndReportsIt(t *testing.T) {
f := &factory{answer: "answer"}
s := New(f.make, []string{"koala/phi4-mini"}, 10)
reply, err := s.Answer(context.Background(), Request{
Model: "koala/phi4-mini",
Transcript: strings.Repeat("x", 500),
Question: "summarize",
})
if err != nil {
t.Fatalf("Answer: %v", err)
}
if !reply.Truncated {
t.Fatal("a transcript past maxChars must report Truncated")
}
if strings.Count(f.used.lastSystem, "x") != 10 {
t.Fatalf("transcript not bounded to maxChars: got %d x's", strings.Count(f.used.lastSystem, "x"))
}
}
func TestAnswerForcesAnUnofferedModelBackToDefault(t *testing.T) {
f := &factory{answer: "answer"}
s := New(f.make, []string{"koala/phi4-mini", "iguana/gemma4-26b"}, 0)
// A forged/un-offered model must never reach the gateway as-is — it is forced
// to the default offered model (the cloud-absent guarantee depends on this).
_, err := s.Answer(context.Background(), Request{
Model: "berget/secret-cloud-model",
Transcript: "t",
Question: "q",
})
if err != nil {
t.Fatalf("Answer: %v", err)
}
if f.used.model != "koala/phi4-mini" {
t.Fatalf("un-offered model reached the gateway as %q, want the default", f.used.model)
}
}
func TestAnswerPropagatesCompleterError(t *testing.T) {
f := &factory{err: errors.New("gateway down")}
s := New(f.make, []string{"koala/phi4-mini"}, 0)
_, err := s.Answer(context.Background(), Request{Model: "koala/phi4-mini", Transcript: "t", Question: "q"})
if err == nil {
t.Fatal("expected the gateway error to propagate")
}
}
func TestAnswerRejectsEmptyModelSet(t *testing.T) {
s := New(func(string) Completer { return nil }, nil, 0)
if _, err := s.Answer(context.Background(), Request{Question: "q"}); err == nil {
t.Fatal("expected an error when no models are configured")
}
}
+16
View File
@@ -32,6 +32,7 @@ type Client struct {
model string model string
maxTokens int maxTokens int
httpClient *http.Client httpClient *http.Client
usageHook func(model string, prompt, completion int)
} }
// Option configures a Client at construction. Variadic so the existing 4-arg // Option configures a Client at construction. Variadic so the existing 4-arg
@@ -50,6 +51,14 @@ func WithMaxTokens(n int) Option {
} }
} }
// WithUsageHook registers a callback fired after a successful completion with the
// model and the prompt/completion token counts from the response usage block. It
// keeps this copied, stdlib-only package (ADR-004) decoupled from metrics: the
// caller wires it to internal/metrics, the client imports nothing. nil is ignored.
func WithUsageHook(fn func(model string, prompt, completion int)) Option {
return func(c *Client) { c.usageHook = fn }
}
// New constructs a Client. // New constructs a Client.
func New(baseURL, apiKey, model string, timeout time.Duration, opts ...Option) *Client { func New(baseURL, apiKey, model string, timeout time.Duration, opts ...Option) *Client {
c := &Client{ c := &Client{
@@ -81,6 +90,10 @@ type chatResponse struct {
Choices []struct { Choices []struct {
Message message `json:"message"` Message message `json:"message"`
} `json:"choices"` } `json:"choices"`
Usage struct {
PromptTokens int `json:"prompt_tokens"`
CompletionTokens int `json:"completion_tokens"`
} `json:"usage"`
} }
// Complete sends a system + user message and returns the assistant's reply. // Complete sends a system + user message and returns the assistant's reply.
@@ -152,5 +165,8 @@ func (c *Client) Complete(ctx context.Context, system, user string) (string, err
if len(cr.Choices) == 0 { if len(cr.Choices) == 0 {
return "", fmt.Errorf("LLM returned no choices") return "", fmt.Errorf("LLM returned no choices")
} }
if c.usageHook != nil {
c.usageHook(c.model, cr.Usage.PromptTokens, cr.Usage.CompletionTokens)
}
return cr.Choices[0].Message.Content, nil return cr.Choices[0].Message.Content, nil
} }
+24
View File
@@ -85,6 +85,30 @@ func TestClient_WithMaxTokens(t *testing.T) {
} }
} }
// TestClient_UsageHookRecordsTokens: the usage hook fires with the model and the
// prompt/completion token counts parsed from the response usage block.
func TestClient_UsageHookRecordsTokens(t *testing.T) {
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, _ *http.Request) {
_ = json.NewEncoder(w).Encode(map[string]any{
"choices": []map[string]any{{"message": map[string]any{"content": "ok"}}},
"usage": map[string]any{"prompt_tokens": 123, "completion_tokens": 45},
})
}))
defer srv.Close()
var gotModel string
var gotPrompt, gotCompletion int
c := New(srv.URL, "", "test-model", 10*time.Second, WithUsageHook(func(model string, p, comp int) {
gotModel, gotPrompt, gotCompletion = model, p, comp
}))
if _, err := c.Complete(context.Background(), "sys", "user"); err != nil {
t.Fatalf("Complete: %v", err)
}
if gotModel != "test-model" || gotPrompt != 123 || gotCompletion != 45 {
t.Errorf("usage hook got (%q, %d, %d), want (test-model, 123, 45)", gotModel, gotPrompt, gotCompletion)
}
}
func TestClient_ReturnsErrorOnNon200(t *testing.T) { func TestClient_ReturnsErrorOnNon200(t *testing.T) {
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
http.Error(w, "overloaded", http.StatusServiceUnavailable) http.Error(w, "overloaded", http.StatusServiceUnavailable)
+12 -6
View File
@@ -4,6 +4,7 @@ import (
"context" "context"
"fmt" "fmt"
"sort" "sort"
"time"
"github.com/jackc/pgx/v5" "github.com/jackc/pgx/v5"
) )
@@ -32,7 +33,12 @@ type UserActiveWeeks struct {
// Scope note: the enumeration covers users with a Dex identity (the web users the // Scope note: the enumeration covers users with a Dex identity (the web users the
// gate is about). A CLI-only user created by the store sink without an identity // gate is about). A CLI-only user created by the store sink without an identity
// row would not appear — out of scope for this gate. // row would not appear — out of scope for this gate.
func (s *Store) ActiveWeeks(ctx context.Context) ([]UserActiveWeeks, error) { // ActiveWeeks counts each user's distinct active weeks from `since` onward. A zero
// `since` means no lower bound (count all history). The Stage-0 gate baseline is
// set by the caller (the report command) to the date real usage tracking began,
// so pre-launch noise — testing, the period the pilot was blocked — does not count
// toward the return-usage signal (ADR-016).
func (s *Store) ActiveWeeks(ctx context.Context, since time.Time) ([]UserActiveWeeks, error) {
userIDs, err := s.identityUserIDs(ctx) userIDs, err := s.identityUserIDs(ctx)
if err != nil { if err != nil {
return nil, err return nil, err
@@ -40,7 +46,7 @@ func (s *Store) ActiveWeeks(ctx context.Context) ([]UserActiveWeeks, error) {
out := make([]UserActiveWeeks, 0, len(userIDs)) out := make([]UserActiveWeeks, 0, len(userIDs))
for _, uid := range userIDs { for _, uid := range userIDs {
row, err := s.activeWeeksFor(ctx, uid) row, err := s.activeWeeksFor(ctx, uid, since)
if err != nil { if err != nil {
return nil, err return nil, err
} }
@@ -85,18 +91,18 @@ func (s *Store) identityUserIDs(ctx context.Context) ([]string, error) {
// activeWeeksFor counts one user's distinct active weeks (reads UNION acts) and // activeWeeksFor counts one user's distinct active weeks (reads UNION acts) and
// reads their display name, RLS-scoped via withUser. The UNION dedups a week that // reads their display name, RLS-scoped via withUser. The UNION dedups a week that
// has both a login and an action so it counts once. // has both a login and an action so it counts once.
func (s *Store) activeWeeksFor(ctx context.Context, userID string) (UserActiveWeeks, error) { func (s *Store) activeWeeksFor(ctx context.Context, userID string, since time.Time) (UserActiveWeeks, error) {
res := UserActiveWeeks{UserID: userID} res := UserActiveWeeks{UserID: userID}
if err := s.withUser(ctx, userID, func(tx pgx.Tx) error { if err := s.withUser(ctx, userID, func(tx pgx.Tx) error {
if err := tx.QueryRow(ctx, if err := tx.QueryRow(ctx,
`WITH weeks AS ( `WITH weeks AS (
SELECT date_trunc('week', seen_at) AS wk SELECT date_trunc('week', seen_at) AS wk
FROM login_events WHERE user_id = $1 FROM login_events WHERE user_id = $1 AND seen_at >= $2
UNION UNION
SELECT date_trunc('week', acted_at) SELECT date_trunc('week', acted_at)
FROM summary_actions WHERE user_id = $1 FROM summary_actions WHERE user_id = $1 AND acted_at >= $2
) )
SELECT count(DISTINCT wk) FROM weeks`, userID).Scan(&res.ActiveWeeks); err != nil { SELECT count(DISTINCT wk) FROM weeks`, userID, since).Scan(&res.ActiveWeeks); err != nil {
return fmt.Errorf("store: count active weeks: %w", err) return fmt.Errorf("store: count active weeks: %w", err)
} }
if err := tx.QueryRow(ctx, if err := tx.QueryRow(ctx,
+29 -2
View File
@@ -3,6 +3,7 @@ package store_test
import ( import (
"context" "context"
"testing" "testing"
"time"
"github.com/jackc/pgx/v5/pgxpool" "github.com/jackc/pgx/v5/pgxpool"
"github.com/stretchr/testify/require" "github.com/stretchr/testify/require"
@@ -54,7 +55,7 @@ func TestActiveWeeksCountsDistinctWeeksAcrossReadsAndActs(t *testing.T) {
($1, 'vid-2', 'saved', '2026-01-19T18:00:00Z')`, userA) ($1, 'vid-2', 'saved', '2026-01-19T18:00:00Z')`, userA)
require.NoError(t, err) require.NoError(t, err)
got, err := s.ActiveWeeks(ctx) got, err := s.ActiveWeeks(ctx, time.Time{}) // zero since = no lower bound
require.NoError(t, err) require.NoError(t, err)
require.Len(t, got, 2, "both identity users must appear") require.Len(t, got, 2, "both identity users must appear")
@@ -72,7 +73,33 @@ func TestActiveWeeksEmptyWhenNoUsers(t *testing.T) {
s := newStore(t) s := newStore(t)
resetDB(t, rawPool(t)) resetDB(t, rawPool(t))
got, err := s.ActiveWeeks(ctx) got, err := s.ActiveWeeks(ctx, time.Time{})
require.NoError(t, err) require.NoError(t, err)
require.Empty(t, got) require.Empty(t, got)
} }
// TestActiveWeeksExcludesBeforeGateStart proves the baseline cutoff: activity
// before `since` does not count, so pre-launch noise (testing, the pilot's blocked
// period) is excluded from the Stage-0 return-usage gate (ADR-016).
func TestActiveWeeksExcludesBeforeGateStart(t *testing.T) {
ctx := context.Background()
s := newStore(t)
p := rawPool(t)
resetDB(t, p)
seedReportUser(t, p, userA, "subject-a", "Ada")
// One read well before the baseline, two reads in distinct weeks after it.
_, err := p.Exec(ctx,
`INSERT INTO login_events (user_id, seen_at) VALUES
($1, '2026-05-01T09:00:00Z'),
($1, '2026-06-12T09:00:00Z'),
($1, '2026-06-19T09:00:00Z')`, userA)
require.NoError(t, err)
since := time.Date(2026, 6, 11, 0, 0, 0, 0, time.UTC)
got, err := s.ActiveWeeks(ctx, since)
require.NoError(t, err)
require.Len(t, got, 1)
require.Equal(t, 2, got[0].ActiveWeeks, "only the two post-baseline weeks count; the May read is excluded")
}
+15 -2
View File
@@ -4,6 +4,7 @@ import (
"context" "context"
"fmt" "fmt"
"os" "os"
"path/filepath"
"testing" "testing"
embeddedpostgres "github.com/fergusstrange/embedded-postgres" embeddedpostgres "github.com/fergusstrange/embedded-postgres"
@@ -24,11 +25,22 @@ var _ ports.Sink = (*store.Store)(nil)
var dsn string var dsn string
func TestMain(m *testing.M) { func TestMain(m *testing.M) {
const port = 54329 // Port + runtime/data dirs are per-process (PID-derived) so two concurrent
// `go test` invocations — e.g. a push-run and a tag-run firing together in CI —
// don't collide on a fixed port or a shared data dir (which silently failed
// both runs). CachePath is shared so the PG archive is downloaded once, not
// per process. Base 54000 keeps this package's range distinct from web's.
port := uint32(54000 + os.Getpid()%1000)
dsn = fmt.Sprintf("postgres://postgres:postgres@localhost:%d/postgres?sslmode=disable", port) dsn = fmt.Sprintf("postgres://postgres:postgres@localhost:%d/postgres?sslmode=disable", port)
rt := filepath.Join(os.TempDir(), fmt.Sprintf("tapir-epg-store-%d", os.Getpid()))
pg := embeddedpostgres.NewDatabase( pg := embeddedpostgres.NewDatabase(
embeddedpostgres.DefaultConfig().Port(port), embeddedpostgres.DefaultConfig().
Port(port).
RuntimePath(rt).
DataPath(filepath.Join(rt, "data")).
BinariesPath(filepath.Join(rt, "bin")).
CachePath(filepath.Join(os.TempDir(), "tapir-epg-cache")),
) )
if err := pg.Start(); err != nil { if err := pg.Start(); err != nil {
fmt.Fprintf(os.Stderr, "embedded-postgres start: %v\n", err) fmt.Fprintf(os.Stderr, "embedded-postgres start: %v\n", err)
@@ -40,6 +52,7 @@ func TestMain(m *testing.M) {
if err := pg.Stop(); err != nil { if err := pg.Stop(); err != nil {
fmt.Fprintf(os.Stderr, "embedded-postgres stop: %v\n", err) fmt.Fprintf(os.Stderr, "embedded-postgres stop: %v\n", err)
} }
_ = os.RemoveAll(rt)
os.Exit(code) os.Exit(code)
} }
+35 -14
View File
@@ -46,15 +46,16 @@ func (s *Store) UpsertVideo(ctx context.Context, v domain.Video) (string, error)
} }
if err := tx.QueryRow(ctx, if err := tx.QueryRow(ctx,
`INSERT INTO videos (user_id, provider, provider_video_id, title, url, published_at, channel_title) `INSERT INTO videos (user_id, provider, provider_video_id, title, url, published_at, channel_title, duration_s)
VALUES ($1, $2, $3, $4, $5, $6, $7) VALUES ($1, $2, $3, $4, $5, $6, $7, $8)
ON CONFLICT (user_id, provider, provider_video_id) DO UPDATE SET ON CONFLICT (user_id, provider, provider_video_id) DO UPDATE SET
title = EXCLUDED.title, title = EXCLUDED.title,
url = EXCLUDED.url, url = EXCLUDED.url,
published_at = EXCLUDED.published_at, published_at = EXCLUDED.published_at,
channel_title = COALESCE(NULLIF(EXCLUDED.channel_title, ''), videos.channel_title) channel_title = COALESCE(NULLIF(EXCLUDED.channel_title, ''), videos.channel_title),
duration_s = COALESCE(EXCLUDED.duration_s, videos.duration_s)
RETURNING id`, RETURNING id`,
v.UserID, provider, v.ProviderVideoID, v.Title, v.URL, nullTime(v.PublishedAt), v.ChannelTitle, v.UserID, provider, v.ProviderVideoID, v.Title, v.URL, nullTime(v.PublishedAt), v.ChannelTitle, nullDuration(v.DurationSeconds),
).Scan(&id); err != nil { ).Scan(&id); err != nil {
return fmt.Errorf("store: upsert video: %w", err) return fmt.Errorf("store: upsert video: %w", err)
} }
@@ -74,12 +75,27 @@ func nullTime(t time.Time) *time.Time {
return &t return &t
} }
// NewestUnsummarizedVideoIDs returns up to limit of the user's videos that have // nullDuration maps an unknown duration (0) to SQL NULL so the upsert's
// no summary yet, newest first (published_at DESC, NULLS LAST). It caps the // COALESCE(EXCLUDED.duration_s, videos.duration_s) preserves a previously-known
// connect-time onboarding burst (Feature 1) at a fixed count: the caller marks // value instead of clobbering it with 0 (ADR-028; the channel_title backfill
// these for summarization through the shared rate gate. RLS-scoped via withUser, // stance, migration 014).
// so it only ever sees the requesting user's rows. limit <= 0 returns nil. func nullDuration(seconds int) *int {
func (s *Store) NewestUnsummarizedVideoIDs(ctx context.Context, userID string, limit int) ([]string, error) { if seconds <= 0 {
return nil
}
return &seconds
}
// OnboardBurstVideoIDs returns up to limit of the user's unsummarized videos for
// the connect-time onboarding burst (ADR-028), newest-first but quality-aware: a
// video is excluded when its duration is KNOWN and outside [minSeconds, maxSeconds]
// — dropping Shorts (below min) and multi-hour livestream VODs (above max) that
// would waste a scarce caption fetch on a poor first impression. A NULL/unknown
// duration is kept (degrade-open) but ranked AFTER known-good rows, so a freshly
// enriched good pick wins when both exist. minSeconds<=0 / maxSeconds<=0 each
// disable that bound (0/0 == pure newest-first, the reversibility lever).
// RLS-scoped via withUser; limit <= 0 returns nil.
func (s *Store) OnboardBurstVideoIDs(ctx context.Context, userID string, limit, minSeconds, maxSeconds int) ([]string, error) {
if limit <= 0 { if limit <= 0 {
return nil, nil return nil, nil
} }
@@ -92,16 +108,21 @@ func (s *Store) NewestUnsummarizedVideoIDs(ctx context.Context, userID string, l
AND NOT EXISTS ( AND NOT EXISTS (
SELECT 1 FROM summaries su SELECT 1 FROM summaries su
WHERE su.user_id = v.user_id AND su.video_id = v.id) WHERE su.user_id = v.user_id AND su.video_id = v.id)
ORDER BY v.published_at DESC NULLS LAST, v.seen_at DESC AND NOT (
LIMIT $2`, userID, limit) v.duration_s IS NOT NULL
AND ( ($3 > 0 AND v.duration_s < $3)
OR ($4 > 0 AND v.duration_s > $4) ))
ORDER BY (v.duration_s IS NOT NULL) DESC,
v.published_at DESC NULLS LAST, v.seen_at DESC
LIMIT $2`, userID, limit, minSeconds, maxSeconds)
if err != nil { if err != nil {
return fmt.Errorf("store: newest unsummarized: %w", err) return fmt.Errorf("store: onboard burst videos: %w", err)
} }
defer rows.Close() defer rows.Close()
for rows.Next() { for rows.Next() {
var id string var id string
if err := rows.Scan(&id); err != nil { if err := rows.Scan(&id); err != nil {
return fmt.Errorf("store: scan newest unsummarized: %w", err) return fmt.Errorf("store: scan onboard burst video: %w", err)
} }
ids = append(ids, id) ids = append(ids, id)
} }
+63 -14
View File
@@ -52,6 +52,36 @@ func TestUpsertVideo_ReturnsStableID(t *testing.T) {
require.Equal(t, 1, count, "must not duplicate the row") require.Equal(t, 1, count, "must not duplicate the row")
} }
func TestUpsertVideo_PersistsAndPreservesDuration(t *testing.T) {
ctx := context.Background()
s := newStore(t)
resetDB(t, rawPool(t))
// First upsert carries a known duration (ADR-028: discovery enriches it).
v := ytVideo(userA, "dur0000001x", "with duration")
v.DurationSeconds = 750
id, err := s.UpsertVideo(ctx, v)
require.NoError(t, err)
p := rawPool(t)
readDuration := func() *int {
var d *int
require.NoError(t, p.QueryRow(ctx, `SELECT duration_s FROM videos WHERE id = $1`, id).Scan(&d))
return d
}
require.NotNil(t, readDuration())
require.Equal(t, 750, *readDuration(), "duration must persist")
// A later upsert that does NOT know the duration (0) must not clobber it —
// the channel_title backfill stance (migration 014): COALESCE-preserve.
v2 := ytVideo(userA, "dur0000001x", "title updated, duration unknown")
v2.DurationSeconds = 0
_, err = s.UpsertVideo(ctx, v2)
require.NoError(t, err)
require.NotNil(t, readDuration(), "a 0/unknown re-upsert must not erase a known duration")
require.Equal(t, 750, *readDuration())
}
func TestUpsertVideo_IDMatchesSummaryDedup(t *testing.T) { func TestUpsertVideo_IDMatchesSummaryDedup(t *testing.T) {
ctx := context.Background() ctx := context.Background()
s := newStore(t) s := newStore(t)
@@ -82,36 +112,55 @@ func TestUpsertVideo_PerUserIsolation(t *testing.T) {
require.NotEqual(t, idA, idB, "same provider video for two users must be two distinct rows") require.NotEqual(t, idA, idB, "same provider video for two users must be two distinct rows")
} }
func TestNewestUnsummarizedVideoIDs(t *testing.T) { func TestOnboardBurstVideoIDs(t *testing.T) {
ctx := context.Background() ctx := context.Background()
s := newStore(t) s := newStore(t)
resetDB(t, rawPool(t)) resetDB(t, rawPool(t))
mk := func(user, pid string, day int) string { mk := func(user, pid string, day, dur int) string {
v := ytVideo(user, pid, pid) v := ytVideo(user, pid, pid)
v.PublishedAt = time.Date(2026, 6, day, 12, 0, 0, 0, time.UTC) v.PublishedAt = time.Date(2026, 6, day, 12, 0, 0, 0, time.UTC)
v.DurationSeconds = dur // 0 == unknown (NULL)
id, err := s.UpsertVideo(ctx, v) id, err := s.UpsertVideo(ctx, v)
require.NoError(t, err) require.NoError(t, err)
return id return id
} }
_ = mk(userA, "a1vid000001", 1) summarized := mk(userA, "summ0000001", 6, 600) // newest known-good, but already summarized
id2 := mk(userA, "a2vid000002", 2) good1 := mk(userA, "good0000001", 5, 600) // 10m, newest UNsummarized known-good
id3 := mk(userA, "a3vid000003", 3) tooLong := mk(userA, "toolong0001", 4, 20000) // > maxSeconds -> dropped
id4 := mk(userA, "a4vid000004", 4) tooShort := mk(userA, "tooshort001", 3, 30) // < minSeconds -> dropped
mk(userB, "b1vid000009", 9) // userB's newest — must never leak via RLS unknown := mk(userA, "unknown0001", 2, 0) // NULL duration -> kept, ranked last
good2 := mk(userA, "good0000002", 1, 800) // known-good but oldest
mk(userB, "bvid0000009", 9, 600) // userB -> must not leak via RLS
// The newest (v4) is summarized, so it's excluded from "unsummarized". // The newest video is summarized, so it is excluded from the burst.
require.NoError(t, s.Deliver(ctx, summary(userA, id4, "done"))) require.NoError(t, s.Deliver(ctx, summary(userA, summarized, "done")))
// Cap 2, newest-first unsummarized: v3 then v2 (v4 excluded; userB excluded). const minSec, maxSec = 60, 14400
got, err := s.NewestUnsummarizedVideoIDs(ctx, userA, 2)
// Known-good ranked before unknown, each newest-first within its group; the
// too-long and too-short videos are excluded by their known duration.
got, err := s.OnboardBurstVideoIDs(ctx, userA, 5, minSec, maxSec)
require.NoError(t, err) require.NoError(t, err)
require.Equal(t, []string{id3, id2}, got) require.Equal(t, []string{good1, good2, unknown}, got,
"known-good first (newest-first), then unknown-duration; junk excluded")
none, err := s.NewestUnsummarizedVideoIDs(ctx, userA, 0) // Cap is honoured.
capped, err := s.OnboardBurstVideoIDs(ctx, userA, 2, minSec, maxSec)
require.NoError(t, err) require.NoError(t, err)
require.Empty(t, none, "limit 0 returns nothing") require.Equal(t, []string{good1, good2}, capped)
// Bounds disabled (0/0) == pure newest-first, nothing excluded.
all, err := s.OnboardBurstVideoIDs(ctx, userA, 10, 0, 0)
require.NoError(t, err)
require.ElementsMatch(t, []string{good1, tooLong, tooShort, unknown, good2}, all,
"0/0 bounds disable the duration filter (prior newest-first behaviour)")
// limit <= 0 returns nothing.
none, err := s.OnboardBurstVideoIDs(ctx, userA, 0, minSec, maxSec)
require.NoError(t, err)
require.Empty(t, none)
} }
func TestUpsertVideoPersistsChannelAndDistinctChannels(t *testing.T) { func TestUpsertVideoPersistsChannelAndDistinctChannels(t *testing.T) {
+10 -1
View File
@@ -13,11 +13,13 @@ import (
"encoding/json" "encoding/json"
"errors" "errors"
"fmt" "fmt"
"log/slog"
"strings" "strings"
"time" "time"
"unicode/utf8" "unicode/utf8"
"gitea.d-ma.be/mathias/tapir/internal/domain" "gitea.d-ma.be/mathias/tapir/internal/domain"
"gitea.d-ma.be/mathias/tapir/internal/metrics"
) )
// Completer is the minimal LLM chat surface the Summarizer needs. // Completer is the minimal LLM chat surface the Summarizer needs.
@@ -95,16 +97,23 @@ func (s *Summarizer) Summarize(ctx context.Context, v domain.Video, t domain.Tra
var errs []error var errs []error
for i, ep := range s.endpoints { for i, ep := range s.endpoints {
fallback := i > 0
start := time.Now()
out, err := ep.Client.Complete(ctx, systemPrompt, user) out, err := ep.Client.Complete(ctx, systemPrompt, user)
dur := time.Since(start)
if err != nil { if err != nil {
metrics.ObserveSummarize(ep.Model, "error", fallback, dur)
errs = append(errs, fmt.Errorf("%s/%s call: %w", ep.Provider, ep.Model, err)) errs = append(errs, fmt.Errorf("%s/%s call: %w", ep.Provider, ep.Model, err))
continue continue
} }
sum, perr := s.build(v, ep, i > 0, out) sum, perr := s.build(v, ep, fallback, out)
if perr != nil { if perr != nil {
metrics.ObserveSummarize(ep.Model, "parse_error", fallback, dur)
errs = append(errs, fmt.Errorf("%s/%s output: %w", ep.Provider, ep.Model, perr)) errs = append(errs, fmt.Errorf("%s/%s output: %w", ep.Provider, ep.Model, perr))
continue continue
} }
metrics.ObserveSummarize(ep.Model, "success", fallback, dur)
slog.Default().Info("summarized", "model", ep.Model, "fallback", fallback, "elapsed_ms", dur.Milliseconds())
return sum, nil return sum, nil
} }
return domain.Summary{}, fmt.Errorf("summarize: all %d endpoint(s) failed: %w", len(s.endpoints), errors.Join(errs...)) return domain.Summary{}, fmt.Errorf("summarize: all %d endpoint(s) failed: %w", len(s.endpoints), errors.Join(errs...))
@@ -7,13 +7,36 @@ package summarizer
import ( import (
"context" "context"
"errors" "errors"
"net/http"
"net/http/httptest"
"strings" "strings"
"testing" "testing"
"gitea.d-ma.be/mathias/tapir/internal/domain" "gitea.d-ma.be/mathias/tapir/internal/domain"
"gitea.d-ma.be/mathias/tapir/internal/metrics"
"gitea.d-ma.be/mathias/tapir/internal/ports" "gitea.d-ma.be/mathias/tapir/internal/ports"
) )
// TestSummarizerRecordsMetric verifies the summarizer→metrics wiring (ADR-030)
// black-box: after a successful summarize, the public /metrics scrape shows a
// success observation for that endpoint's model.
func TestSummarizerRecordsMetric(t *testing.T) {
const model = "metrics-test-model"
s := New(Endpoint{Client: &fakeClient{reply: goodReply}, Provider: "local", Model: model}, nil)
if _, err := s.Summarize(context.Background(), testVideo(), testTranscript()); err != nil {
t.Fatalf("Summarize: %v", err)
}
rec := httptest.NewRecorder()
metrics.Handler().ServeHTTP(rec, httptest.NewRequest(http.MethodGet, "/metrics", nil))
body := rec.Body.String()
if !strings.Contains(body, `tapir_summarize_duration_seconds`) ||
!strings.Contains(body, `model="`+model+`"`) ||
!strings.Contains(body, `outcome="success"`) {
t.Errorf("metrics scrape missing summarize success for %s", model)
}
}
// compile-time check: Summarizer satisfies the port. // compile-time check: Summarizer satisfies the port.
var _ ports.Summarizer = (*Summarizer)(nil) var _ ports.Summarizer = (*Summarizer)(nil)
+31
View File
@@ -7,10 +7,13 @@ import (
"encoding/xml" "encoding/xml"
"fmt" "fmt"
"io" "io"
"log/slog"
"net/http" "net/http"
"strings" "strings"
"time"
"gitea.d-ma.be/mathias/tapir/internal/domain" "gitea.d-ma.be/mathias/tapir/internal/domain"
"gitea.d-ma.be/mathias/tapir/internal/metrics"
) )
// defaultPlayerBaseURL is the InnerTube / watch-page host. Overridable via // defaultPlayerBaseURL is the InnerTube / watch-page host. Overridable via
@@ -43,7 +46,35 @@ const maxCaptionBytes = 16 << 20 // 16 MiB
// fetch, or an unparseable body all yield SourceNone rather than an error. Only // fetch, or an unparseable body all yield SourceNone rather than an error. Only
// genuine transport (network) faults return an error. Audio download and // genuine transport (network) faults return an error. Audio download and
// speech-to-text remain absent (ADR-007). // speech-to-text remain absent (ADR-007).
// FetchTranscript times the caption fetch and records its latency by outcome
// (ADR-030) before returning. Transport errors are surfaced to the caller and not
// recorded as an outcome (logged upstream); the three resolved outcomes
// captions|none|rate_limited are the ones that consume the scarce fetch budget.
func (a *Adapter) FetchTranscript(ctx context.Context, v domain.Video) (domain.Transcript, error) { func (a *Adapter) FetchTranscript(ctx context.Context, v domain.Video) (domain.Transcript, error) {
start := time.Now()
tr, err := a.fetchTranscript(ctx, v)
if err == nil {
dur := time.Since(start)
outcome := captionOutcome(tr.Source)
metrics.ObserveCaptionFetch(outcome, dur)
slog.Default().Info("caption fetch", "video", v.ProviderVideoID, "outcome", outcome, "elapsed_ms", dur.Milliseconds())
}
return tr, err
}
// captionOutcome maps a transcript source to the metric outcome label.
func captionOutcome(s domain.TranscriptSource) string {
switch s {
case domain.SourceCaptions:
return "captions"
case domain.SourceRateLimited:
return "rate_limited"
default:
return "none"
}
}
func (a *Adapter) fetchTranscript(ctx context.Context, v domain.Video) (domain.Transcript, error) {
client := a.plainClient() client := a.plainClient()
tracks, err := a.captionTracks(ctx, client, v.ProviderVideoID) tracks, err := a.captionTracks(ctx, client, v.ProviderVideoID)
+47
View File
@@ -2,6 +2,7 @@ package youtube
import ( import (
"context" "context"
"sync/atomic"
"time" "time"
"golang.org/x/time/rate" "golang.org/x/time/rate"
@@ -29,9 +30,55 @@ func SetFetchRate(interval time.Duration) {
globalFetchGate = rate.NewLimiter(rate.Every(interval), 1) globalFetchGate = rate.NewLimiter(rate.Every(interval), 1)
} }
// foregroundPending counts in-flight foreground (user-initiated) caption fetches.
// The background sweep yields the gate while this is non-zero so a human waiting
// on a click gets the next slot — and, on a near-throttled IP, the pre-429 window
// — instead of competing equally with the firehose (ADR-026, Pillar A). Clicks are
// rare and bursty, so background barely notices; the win to the click is large.
var foregroundPending atomic.Int64
// fgCtxKey marks a context as foreground (user-initiated). Unexported; set via
// ForegroundContext and read via isForeground so only this package owns the key.
type fgCtxKey struct{}
// ForegroundContext marks ctx as a user-initiated (foreground) fetch so the gate
// gives it priority. The web "Summarize"/paste/retry path wraps its context with
// this; the background scheduler leaves it unset.
func ForegroundContext(ctx context.Context) context.Context {
return context.WithValue(ctx, fgCtxKey{}, true)
}
func isForeground(ctx context.Context) bool {
v, _ := ctx.Value(fgCtxKey{}).(bool)
return v
}
// fgYieldPoll is how often a background waiter re-checks whether a foreground
// fetch is still pending. Short enough to feel immediate, long enough not to spin.
const fgYieldPoll = 200 * time.Millisecond
// WaitFetchGate blocks until the process-wide gate allows one timedtext fetch, // WaitFetchGate blocks until the process-wide gate allows one timedtext fetch,
// respecting ctx cancellation. Called from httpDo before every live outbound // respecting ctx cancellation. Called from httpDo before every live outbound
// caption fetch so the scheduler and the click-path share the same egress budget. // caption fetch so the scheduler and the click-path share the same egress budget.
//
// Foreground (user-initiated) fetches take priority: they register as pending and
// acquire a token immediately. Background fetches first yield — they wait until no
// foreground fetch is pending — so a live click is never stuck behind the
// background sweep and gets the cleaner slot against the per-IP limit (ADR-026).
func WaitFetchGate(ctx context.Context) error { func WaitFetchGate(ctx context.Context) error {
if isForeground(ctx) {
foregroundPending.Add(1)
defer foregroundPending.Add(-1)
return globalFetchGate.Wait(ctx)
}
// Background: defer to any pending foreground fetch before taking a token.
for foregroundPending.Load() > 0 {
select {
case <-ctx.Done():
return ctx.Err()
case <-time.After(fgYieldPoll):
}
}
return globalFetchGate.Wait(ctx) return globalFetchGate.Wait(ctx)
} }
+34
View File
@@ -82,3 +82,37 @@ func TestSetFetchRateZeroIsUnlimited(t *testing.T) {
require.NoError(t, WaitFetchGate(context.Background())) require.NoError(t, WaitFetchGate(context.Background()))
} }
} }
func TestForegroundContextMarker(t *testing.T) {
require.False(t, isForeground(context.Background()), "plain context is background")
require.True(t, isForeground(ForegroundContext(context.Background())), "marked context is foreground")
}
// TestWaitFetchGateForegroundProceedsImmediately: a foreground fetch acquires a
// token without yielding, even when background callers exist.
func TestWaitFetchGateForegroundProceedsImmediately(t *testing.T) {
SetFetchRate(0) // unlimited limiter — isolate the yield logic from pacing
foregroundPending.Store(0)
t.Cleanup(func() { foregroundPending.Store(0) })
ctx, cancel := context.WithTimeout(context.Background(), time.Second)
defer cancel()
require.NoError(t, WaitFetchGate(ForegroundContext(ctx)), "foreground proceeds immediately")
}
// TestWaitFetchGateBackgroundYieldsToForeground: while a foreground fetch is
// pending, a background fetch yields (does not take a token) until the foreground
// clears — proven by a background wait timing out against its own deadline, then
// succeeding once the foreground is done.
func TestWaitFetchGateBackgroundYieldsToForeground(t *testing.T) {
SetFetchRate(0)
foregroundPending.Store(1) // simulate a foreground fetch in flight
t.Cleanup(func() { foregroundPending.Store(0) })
ctx, cancel := context.WithTimeout(context.Background(), 250*time.Millisecond)
defer cancel()
require.Error(t, WaitFetchGate(ctx), "background yields (blocks) while foreground is pending")
foregroundPending.Store(0) // foreground done
require.NoError(t, WaitFetchGate(context.Background()), "background proceeds once foreground clears")
}
+4
View File
@@ -302,6 +302,10 @@ func (a *Adapter) filterLowValue(ctx context.Context, client *http.Client, video
if m.seconds > 0 && m.seconds < a.cfg.MinVideoSeconds { if m.seconds > 0 && m.seconds < a.cfg.MinVideoSeconds {
continue // Short / sub-threshold clip continue // Short / sub-threshold clip
} }
// Carry the duration we already fetched onto the kept video so the store
// can persist it (ADR-028) — the burst's length-aware selection depends on
// it. Discarding it here was the gap the onboarding investigation found.
v.DurationSeconds = m.seconds
kept = append(kept, v) kept = append(kept, v)
} }
return kept return kept
@@ -224,6 +224,11 @@ func TestNewVideosFiltersShortsAndLive(t *testing.T) {
if len(vids) != 1 || vids[0].ProviderVideoID != "long1" { if len(vids) != 1 || vids[0].ProviderVideoID != "long1" {
t.Fatalf("expected only long1 to survive the filter, got %+v", vids) t.Fatalf("expected only long1 to survive the filter, got %+v", vids)
} }
// The duration fetched for the filter is carried onto the kept video so the
// store can persist it (ADR-028) instead of discarding it.
if vids[0].DurationSeconds != 750 {
t.Fatalf("kept video DurationSeconds = %d, want 750 (PT12M30S)", vids[0].DurationSeconds)
}
} }
// TestNewVideosNoFilterWhenDisabled: MinVideoSeconds=0 keeps the pre-ADR-023 // TestNewVideosNoFilterWhenDisabled: MinVideoSeconds=0 keeps the pre-ADR-023
+39 -2
View File
@@ -33,7 +33,10 @@ type Config struct {
SummarizerModel string SummarizerModel string
// FallbackModel is the LOCAL fallback alias tried when the primary fails or // FallbackModel is the LOCAL fallback alias tried when the primary fails or
// returns unparseable output (ADR-022). Kept local so content stays on the // returns unparseable output (ADR-022). Kept local so content stays on the
// homelab stack. Empty disables it. Default a bigger-context local model. // homelab stack. Default is an IGUANA model (not koala) so the fallback runs
// on a different host than the koala primary — koala carries other loads, and
// a different host also means a different egress IP for the (rare) fallback.
// Empty disables it.
FallbackModel string FallbackModel string
// CloudFallbackModel is the worst-case EXTERNAL fallback alias, tried only // CloudFallbackModel is the worst-case EXTERNAL fallback alias, tried only
// after every local endpoint has failed (ADR-022). For client deployments set // after every local endpoint has failed (ADR-022). For client deployments set
@@ -117,6 +120,21 @@ type Config struct {
// caption rate gate (ADR-014) — the cap bounds count, never the pacing. Default 3. // caption rate gate (ADR-014) — the cap bounds count, never the pacing. Default 3.
OnboardSummarizeCount int OnboardSummarizeCount int
// OnboardSummarizerModel is the summarizer alias the connect-time burst leads
// its chain with (ADR-028) — a stronger model is affordable on the ≤3 summaries
// that form a new user's first impression. It heads a burst-specific chain;
// the standard chain (ADR-022) follows as resilience. Empty (or equal to
// SummarizerModel) collapses the burst back onto the shared processor — the
// reversibility lever. Default iguana/gemma4-26b (the brain-validated model).
OnboardSummarizerModel string
// OnboardMaxVideoSeconds upper-bounds the duration of a video the onboarding
// burst will pick (ADR-028), so the burst does not spend a scarce caption fetch
// on a multi-hour livestream VOD that passed the live filter once it ended. Only
// a KNOWN duration outside [MinVideoSeconds, this] is dropped; a NULL/unknown
// duration is kept (degrade-open). 0 disables the upper bound. Default 14400 (4h).
OnboardMaxVideoSeconds int
// DiscoveryInterval, when > 0, makes `serve` run in-process scheduled discovery // DiscoveryInterval, when > 0, makes `serve` run in-process scheduled discovery
// for ALL users on that cadence (ADR-018). Zero/unset = disabled, so dev and // for ALL users on that cadence (ADR-018). Zero/unset = disabled, so dev and
// tests never auto-fetch. Single-replica assumption — see cmdServe. // tests never auto-fetch. Single-replica assumption — see cmdServe.
@@ -125,6 +143,11 @@ type Config struct {
// HTTPAddr is the listen address for `tapir serve` (the Stage-0 web UI). // HTTPAddr is the listen address for `tapir serve` (the Stage-0 web UI).
HTTPAddr string HTTPAddr string
// MetricsAddr is the listen address for the Prometheus /metrics endpoint
// (ADR-030). A SEPARATE port from HTTPAddr so /metrics is never exposed on the
// public app — only scraped in-cluster. Empty disables the metrics server.
MetricsAddr string
// PublicURL is the externally-reachable base URL of the deployed service, // PublicURL is the externally-reachable base URL of the deployed service,
// e.g. "https://tapir.d-ma.be". Used to build absolute links handed to humans // e.g. "https://tapir.d-ma.be". Used to build absolute links handed to humans
// (the `tapir invite` URL). No trailing slash is assumed — callers trim it. // (the `tapir invite` URL). No trailing slash is assumed — callers trim it.
@@ -148,7 +171,7 @@ func (c Config) DexConfigured() bool { return strings.TrimSpace(c.OIDCIssuer) !=
const ( const (
defaultGatewayURL = "http://koala:30401/v1" defaultGatewayURL = "http://koala:30401/v1"
defaultSummarizerModel = "koala/phi4-mini" defaultSummarizerModel = "koala/phi4-mini"
defaultFallbackModel = "koala/phi4-14b" defaultFallbackModel = "iguana/gemma4-26b"
defaultCloudFallbackModel = "berget/mistral-small" defaultCloudFallbackModel = "berget/mistral-small"
defaultSummaryMaxTokens = 1500 defaultSummaryMaxTokens = 1500
defaultMaxTranscriptChars = 18000 defaultMaxTranscriptChars = 18000
@@ -160,12 +183,15 @@ const (
defaultYTConnectRedirectURL = "https://tapir.d-ma.be/oauth/youtube/callback" defaultYTConnectRedirectURL = "https://tapir.d-ma.be/oauth/youtube/callback"
defaultOAuthRedirectAddr = "localhost:8080" defaultOAuthRedirectAddr = "localhost:8080"
defaultHTTPAddr = ":8080" defaultHTTPAddr = ":8080"
defaultMetricsAddr = ":9090"
defaultFetchBackoff = time.Hour defaultFetchBackoff = time.Hour
defaultFetchRate = 2 * time.Second defaultFetchRate = 2 * time.Second
defaultPublicURL = "https://tapir.d-ma.be" defaultPublicURL = "https://tapir.d-ma.be"
defaultAutoSummarizeWindow = 7 * 24 * time.Hour defaultAutoSummarizeWindow = 7 * 24 * time.Hour
defaultOnboardSummarizeCount = 3 defaultOnboardSummarizeCount = 3
maxOnboardSummarizeCount = 5 maxOnboardSummarizeCount = 5
defaultOnboardSummarizerModel = "iguana/gemma4-26b"
defaultOnboardMaxVideoSeconds = 14400 // 4h
) )
// Load reads the environment into a Config, applying defaults. It does not // Load reads the environment into a Config, applying defaults. It does not
@@ -178,6 +204,7 @@ func Load() (Config, error) {
GatewayURL: envOr("TAPIR_GATEWAY_URL", defaultGatewayURL), GatewayURL: envOr("TAPIR_GATEWAY_URL", defaultGatewayURL),
GatewayKey: os.Getenv("TAPIR_GATEWAY_KEY"), GatewayKey: os.Getenv("TAPIR_GATEWAY_KEY"),
SummarizerModel: envOr("TAPIR_SUMMARIZER_MODEL", defaultSummarizerModel), SummarizerModel: envOr("TAPIR_SUMMARIZER_MODEL", defaultSummarizerModel),
OnboardSummarizerModel: lookupOr("TAPIR_ONBOARD_SUMMARIZER_MODEL", defaultOnboardSummarizerModel),
FallbackModel: lookupOr("TAPIR_FALLBACK_MODEL", defaultFallbackModel), FallbackModel: lookupOr("TAPIR_FALLBACK_MODEL", defaultFallbackModel),
CloudFallbackModel: lookupOr("TAPIR_CLOUD_FALLBACK_MODEL", defaultCloudFallbackModel), CloudFallbackModel: lookupOr("TAPIR_CLOUD_FALLBACK_MODEL", defaultCloudFallbackModel),
DBDSN: os.Getenv("TAPIR_DB_DSN"), DBDSN: os.Getenv("TAPIR_DB_DSN"),
@@ -188,6 +215,7 @@ func Load() (Config, error) {
SecretsFile: envOr("TAPIR_SECRETS_FILE", defaultSecretsFile()), SecretsFile: envOr("TAPIR_SECRETS_FILE", defaultSecretsFile()),
OAuthRedirectAddr: envOr("TAPIR_OAUTH_REDIRECT_ADDR", defaultOAuthRedirectAddr), OAuthRedirectAddr: envOr("TAPIR_OAUTH_REDIRECT_ADDR", defaultOAuthRedirectAddr),
HTTPAddr: envOr("TAPIR_HTTP_ADDR", defaultHTTPAddr), HTTPAddr: envOr("TAPIR_HTTP_ADDR", defaultHTTPAddr),
MetricsAddr: lookupOr("TAPIR_METRICS_ADDR", defaultMetricsAddr),
PublicURL: envOr("TAPIR_PUBLIC_URL", defaultPublicURL), PublicURL: envOr("TAPIR_PUBLIC_URL", defaultPublicURL),
OIDCIssuer: os.Getenv("TAPIR_OIDC_ISSUER"), OIDCIssuer: os.Getenv("TAPIR_OIDC_ISSUER"),
DexClientID: os.Getenv("TAPIR_DEX_CLIENT_ID"), DexClientID: os.Getenv("TAPIR_DEX_CLIENT_ID"),
@@ -283,6 +311,15 @@ func Load() (Config, error) {
} }
c.OnboardSummarizeCount = onboard c.OnboardSummarizeCount = onboard
onboardMax, err := intOr("TAPIR_ONBOARD_MAX_VIDEO_SECONDS", defaultOnboardMaxVideoSeconds)
if err != nil {
return Config{}, err
}
if onboardMax < 0 {
onboardMax = 0
}
c.OnboardMaxVideoSeconds = onboardMax
return c, nil return c, nil
} }
+59
View File
@@ -214,3 +214,62 @@ func TestValidateForAuth_PassesWhenComplete(t *testing.T) {
t.Errorf("ValidateForAuth: unexpected error %v", err) t.Errorf("ValidateForAuth: unexpected error %v", err)
} }
} }
func TestLoad_OnboardSummarizerModel(t *testing.T) {
cases := []struct {
name, env string
set bool
want string
}{
{"default", "", false, defaultOnboardSummarizerModel},
{"explicit", "koala/some-model", true, "koala/some-model"},
{"empty disables (collapses to shared processor)", "", true, ""},
}
for _, c := range cases {
t.Run(c.name, func(t *testing.T) {
env := map[string]string{}
if c.set {
env["TAPIR_ONBOARD_SUMMARIZER_MODEL"] = c.env
}
setEnv(t, env)
cfg, err := Load()
if err != nil {
t.Fatalf("Load: %v", err)
}
if cfg.OnboardSummarizerModel != c.want {
t.Fatalf("OnboardSummarizerModel = %q, want %q", cfg.OnboardSummarizerModel, c.want)
}
})
}
}
func TestLoad_OnboardMaxVideoSeconds(t *testing.T) {
cases := []struct {
name, env string
want int
}{
{"default", "", defaultOnboardMaxVideoSeconds},
{"explicit", "7200", 7200},
{"zero disables", "0", 0},
{"negative clamps to zero", "-9", 0},
}
for _, c := range cases {
t.Run(c.name, func(t *testing.T) {
setEnv(t, map[string]string{"TAPIR_ONBOARD_MAX_VIDEO_SECONDS": c.env})
cfg, err := Load()
if err != nil {
t.Fatalf("Load: %v", err)
}
if cfg.OnboardMaxVideoSeconds != c.want {
t.Fatalf("OnboardMaxVideoSeconds = %d, want %d", cfg.OnboardMaxVideoSeconds, c.want)
}
})
}
}
func TestLoad_OnboardMaxVideoSecondsInvalid(t *testing.T) {
setEnv(t, map[string]string{"TAPIR_ONBOARD_MAX_VIDEO_SECONDS": "long"})
if _, err := Load(); err == nil {
t.Fatal("Load: want error for non-numeric TAPIR_ONBOARD_MAX_VIDEO_SECONDS")
}
}
+5
View File
@@ -77,6 +77,11 @@ type Video struct {
URL string URL string
PublishedAt time.Time PublishedAt time.Time
SeenAt time.Time SeenAt time.Time
// DurationSeconds is the video length in seconds, when known (fetched by the
// ADR-023 videos.list enrichment at discovery). 0 means unknown — the store
// preserves a previously-known value rather than overwriting it with 0, and
// the onboarding burst (ADR-028) treats unknown as degrade-open (kept).
DurationSeconds int
} }
// Transcript is the text of a video (or a record that none was available). // Transcript is the text of a video (or a record that none was available).
+139
View File
@@ -0,0 +1,139 @@
// Package metrics is Tapir's Prometheus instrumentation (ADR-030, issue #15). It
// owns the collectors and a small typed API the rest of the app calls — adapters
// never touch prometheus types directly. Two themes:
//
// - HTTP/session: request count + latency by route (the matched pattern, so
// cardinality stays bounded), and logins.
// - AI (the priority): summarization latency by model/outcome/fallback, caption
// fetch latency by outcome, chat latency by model, and LLM token usage.
//
// Handler() is served on a dedicated port (never the public app port) so a scrape
// is in-cluster only. slog timing lines are emitted at the call sites too.
package metrics
import (
"net/http"
"strconv"
"time"
"github.com/prometheus/client_golang/prometheus"
"github.com/prometheus/client_golang/prometheus/promauto"
"github.com/prometheus/client_golang/prometheus/promhttp"
)
// latencyBuckets spans sub-second UI calls up to multi-minute model calls (a cold
// local model load is tens of seconds; the cloud fallback can be longer).
var latencyBuckets = []float64{0.05, 0.1, 0.25, 0.5, 1, 2, 5, 10, 20, 30, 60, 120, 300}
var (
httpRequests = promauto.NewCounterVec(prometheus.CounterOpts{
Name: "tapir_http_requests_total",
Help: "HTTP requests by method, matched route pattern, and status code.",
}, []string{"method", "route", "code"})
httpDuration = promauto.NewHistogramVec(prometheus.HistogramOpts{
Name: "tapir_http_request_duration_seconds",
Help: "HTTP request latency by method and matched route pattern.",
Buckets: []float64{0.005, 0.01, 0.025, 0.05, 0.1, 0.25, 0.5, 1, 2, 5},
}, []string{"method", "route"})
logins = promauto.NewCounter(prometheus.CounterOpts{
Name: "tapir_logins_total",
Help: "Successful OIDC logins (session established).",
})
summarizeDuration = promauto.NewHistogramVec(prometheus.HistogramOpts{
Name: "tapir_summarize_duration_seconds",
Help: "Per-endpoint summarization latency by model, outcome (success|parse_error|error), and whether it was a fallback.",
Buckets: latencyBuckets,
}, []string{"model", "outcome", "fallback"})
captionFetchDuration = promauto.NewHistogramVec(prometheus.HistogramOpts{
Name: "tapir_caption_fetch_duration_seconds",
Help: "Caption fetch latency by outcome (captions|none|rate_limited).",
Buckets: latencyBuckets,
}, []string{"outcome"})
chatDuration = promauto.NewHistogramVec(prometheus.HistogramOpts{
Name: "tapir_chat_duration_seconds",
Help: "Per-video Q&A answer latency by model.",
Buckets: latencyBuckets,
}, []string{"model"})
llmTokens = promauto.NewCounterVec(prometheus.CounterOpts{
Name: "tapir_llm_tokens_total",
Help: "LLM tokens consumed by model and kind (prompt|completion).",
}, []string{"model", "kind"})
)
// Handler serves the Prometheus exposition format. Mount on the dedicated metrics
// port, never the public app mux.
func Handler() http.Handler { return promhttp.Handler() }
// IncLogin records a successful login.
func IncLogin() { logins.Inc() }
// ObserveSummarize records one summarization endpoint attempt.
func ObserveSummarize(model, outcome string, fallback bool, d time.Duration) {
summarizeDuration.WithLabelValues(model, outcome, strconv.FormatBool(fallback)).Observe(d.Seconds())
}
// ObserveCaptionFetch records one caption fetch by outcome.
func ObserveCaptionFetch(outcome string, d time.Duration) {
captionFetchDuration.WithLabelValues(outcome).Observe(d.Seconds())
}
// ObserveChat records one Q&A answer latency.
func ObserveChat(model string, d time.Duration) {
chatDuration.WithLabelValues(model).Observe(d.Seconds())
}
// RecordTokens records LLM token usage from a completion's usage block. Zero
// counts are skipped so a provider that omits usage adds nothing.
func RecordTokens(model string, prompt, completion int) {
if prompt > 0 {
llmTokens.WithLabelValues(model, "prompt").Add(float64(prompt))
}
if completion > 0 {
llmTokens.WithLabelValues(model, "completion").Add(float64(completion))
}
}
// HTTPMiddleware records request count + latency. It reads r.Pattern AFTER the
// inner handler routes (Go 1.22 sets it during ServeMux matching), so the label is
// the bounded registered pattern (e.g. "GET /v/{videoId}"), never the raw path
// with its high-cardinality ids. Unmatched requests bucket as "other".
func HTTPMiddleware(next http.Handler) http.Handler {
return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
start := time.Now()
sw := &statusWriter{ResponseWriter: w, code: http.StatusOK}
next.ServeHTTP(sw, r)
route := r.Pattern
if route == "" {
route = "other"
}
httpRequests.WithLabelValues(r.Method, route, strconv.Itoa(sw.code)).Inc()
httpDuration.WithLabelValues(r.Method, route).Observe(time.Since(start).Seconds())
})
}
// statusWriter captures the response status for the request-count label.
type statusWriter struct {
http.ResponseWriter
code int
wroteHeader bool
}
func (s *statusWriter) WriteHeader(code int) {
if !s.wroteHeader {
s.code = code
s.wroteHeader = true
}
s.ResponseWriter.WriteHeader(code)
}
func (s *statusWriter) Write(b []byte) (int, error) {
s.wroteHeader = true // an implicit 200
return s.ResponseWriter.Write(b)
}
+92
View File
@@ -0,0 +1,92 @@
package metrics
import (
"net/http"
"net/http/httptest"
"testing"
"time"
"github.com/prometheus/client_golang/prometheus"
"github.com/prometheus/client_golang/prometheus/testutil"
dto "github.com/prometheus/client_model/go"
"github.com/stretchr/testify/require"
)
// histCount reads a histogram child's observation count (testutil.ToFloat64 only
// works on counters/gauges; a histogram's WithLabelValues child is an Observer).
func histCount(t *testing.T, o prometheus.Observer) uint64 {
t.Helper()
m, ok := o.(prometheus.Metric)
require.True(t, ok, "histogram child must be a prometheus.Metric")
var d dto.Metric
require.NoError(t, m.Write(&d))
return d.GetHistogram().GetSampleCount()
}
// TestObserveSummarizeRecordsModelOutcomeFallback: a success observation lands on
// the right model/outcome/fallback series.
func TestObserveSummarizeRecordsModelOutcomeFallback(t *testing.T) {
before := histCount(t, summarizeDuration.WithLabelValues("koala/phi4-mini", "success", "false"))
ObserveSummarize("koala/phi4-mini", "success", false, 1200*time.Millisecond)
after := histCount(t, summarizeDuration.WithLabelValues("koala/phi4-mini", "success", "false"))
require.Equal(t, before+1, after, "one success observation recorded for the model")
}
// TestObserveSummarizeRecordsFailureOutcomes: error and parse_error are distinct
// series so a fallback chain's failures are visible.
func TestObserveSummarizeRecordsFailureOutcomes(t *testing.T) {
e0 := histCount(t, summarizeDuration.WithLabelValues("m", "error", "false"))
p0 := histCount(t, summarizeDuration.WithLabelValues("m", "parse_error", "false"))
ObserveSummarize("m", "error", false, time.Second)
ObserveSummarize("m", "parse_error", false, time.Second)
require.Equal(t, e0+1, histCount(t, summarizeDuration.WithLabelValues("m", "error", "false")))
require.Equal(t, p0+1, histCount(t, summarizeDuration.WithLabelValues("m", "parse_error", "false")))
}
func TestObserveCaptionFetchByOutcome(t *testing.T) {
b := histCount(t, captionFetchDuration.WithLabelValues("captions"))
ObserveCaptionFetch("captions", 3*time.Second)
require.Equal(t, b+1, histCount(t, captionFetchDuration.WithLabelValues("captions")))
}
func TestChatAnswerLatencyRecorded(t *testing.T) {
b := histCount(t, chatDuration.WithLabelValues("iguana/gemma4-26b"))
ObserveChat("iguana/gemma4-26b", 2*time.Second)
require.Equal(t, b+1, histCount(t, chatDuration.WithLabelValues("iguana/gemma4-26b")))
}
// TestRecordTokens: prompt + completion land on their kind series; zero is skipped.
func TestRecordTokens(t *testing.T) {
p0 := testutil.ToFloat64(llmTokens.WithLabelValues("m", "prompt"))
c0 := testutil.ToFloat64(llmTokens.WithLabelValues("m", "completion"))
RecordTokens("m", 100, 40)
RecordTokens("m", 0, 0) // skipped, no panic
require.Equal(t, p0+100, testutil.ToFloat64(llmTokens.WithLabelValues("m", "prompt")))
require.Equal(t, c0+40, testutil.ToFloat64(llmTokens.WithLabelValues("m", "completion")))
}
func TestLoginCounted(t *testing.T) {
b := testutil.ToFloat64(logins)
IncLogin()
require.Equal(t, b+1, testutil.ToFloat64(logins))
}
// TestHTTPMiddlewareRecordsByRoutePattern: the request is counted under the bounded
// registered pattern (r.Pattern after routing), not the raw path with its ids.
func TestHTTPMiddlewareRecordsByRoutePattern(t *testing.T) {
mux := http.NewServeMux()
mux.HandleFunc("GET /v/{videoId}", func(w http.ResponseWriter, _ *http.Request) {
w.WriteHeader(http.StatusTeapot)
})
h := HTTPMiddleware(mux)
before := testutil.ToFloat64(httpRequests.WithLabelValues("GET", "GET /v/{videoId}", "418"))
rec := httptest.NewRecorder()
h.ServeHTTP(rec, httptest.NewRequest(http.MethodGet, "/v/abc-123", nil))
require.Equal(t, http.StatusTeapot, rec.Code)
after := testutil.ToFloat64(httpRequests.WithLabelValues("GET", "GET /v/{videoId}", "418"))
require.Equal(t, before+1, after, "counted under the pattern, not /v/abc-123")
require.Equal(t, float64(0), testutil.ToFloat64(httpRequests.WithLabelValues("GET", "/v/abc-123", "418")),
"raw path must never be a label value")
}
+227
View File
@@ -0,0 +1,227 @@
package web
import (
"context"
"errors"
"net/http"
"strings"
"gitea.d-ma.be/mathias/tapir/internal/adapters/chat"
"gitea.d-ma.be/mathias/tapir/internal/adapters/store"
)
// Chatter is the per-video chat backend (ADR-027). *chat.Service satisfies it;
// tests substitute a fake. It carries NO caption-fetch dependency — the chat
// handlers reach it only after reading an already-stored transcript, so an
// enabled chat cannot trigger a fetch, touch the rate gate, or reach YouTube.
type Chatter interface {
// Models returns the offerable models, local-first (cloud absent when disabled).
Models() []string
// DefaultModel resolves the model a fresh chat opens with given the summary's
// model (the ADR-027 default), falling back to the first offered model.
DefaultModel(summaryModel string) string
// Answer runs one chat turn against the supplied stored transcript text.
Answer(ctx context.Context, req chat.Request) (chat.Reply, error)
}
// maxHistoryTurns bounds the ephemeral conversation carried per request, so a long
// back-and-forth cannot grow the prompt without limit (the transcript already
// dominates the budget). Older turns drop off the front.
const maxHistoryTurns = 8
// chatView is everything the chat templates render: the video identity for links
// and titles, the model switcher state, the running (ephemeral) conversation, and
// the honest flags — Available is false when no usable transcript is stored
// (ADR-027: honest "not available", never a fetch), Truncated when the transcript
// was bounded to fit the model, Error for a transient model failure.
type chatView struct {
VideoID string
Title string
Available bool
Models []string
Selected string
History []chat.Turn
Truncated bool
Error string
}
// handleChat renders the chat page for a summarized video (GET). Entry is scoped
// through GetSummaryByVideo, which is RLS/user-scoped: a video that is not the
// requesting user's own resolves to ErrNotFound → 404, so chat is reachable only
// from the user's own summary view (ADR-027 isolation). The transcript is read
// from the SHARED store (ADR-021) — a pure DB read, never a caption fetch; an
// absent/text-less transcript renders the honest "not available" state, no fetch.
func (a *App) handleChat(w http.ResponseWriter, r *http.Request) {
userID, ok := a.currentUserID(w, r)
if !ok {
return
}
row, ok := a.loadOwnedSummary(w, r, userID)
if !ok {
return
}
_, hasText, ok := a.readTranscript(w, r, *row)
if !ok {
return
}
view := chatView{
VideoID: row.VideoID,
Title: displayTitle(*row),
Available: hasText,
Models: a.Chat.Models(),
Selected: a.Chat.DefaultModel(row.AIModel),
}
// HTMX (the in-place reveal from the summary) gets just the open chat section,
// swapped over the closed dock so the summary above it stays put. A no-JS
// navigation gets the full page: the whole summary plus the open chat.
if isHTMX(r) {
a.render(w, r, chatSection(view))
return
}
a.render(w, r, ChatPage(*row, view))
}
// handleChatMessage answers one question against the stored transcript (POST).
// It reads the transcript from the store (no fetch), runs the chosen model over
// it plus the prior turns, appends the answer, and returns the refreshed chat
// panel (HTMX) or the whole page (no-JS). A model failure is surfaced inline,
// not as a 500 — the conversation and the question are preserved for a retry.
func (a *App) handleChatMessage(w http.ResponseWriter, r *http.Request) {
userID, ok := a.currentUserID(w, r)
if !ok {
return
}
row, ok := a.loadOwnedSummary(w, r, userID)
if !ok {
return
}
if err := r.ParseForm(); err != nil {
http.Error(w, "bad form", http.StatusBadRequest)
return
}
transcript, hasText, ok := a.readTranscript(w, r, *row)
if !ok {
return
}
model := a.resolveModel(r.FormValue("model"), row.AIModel)
history := parseHistory(r.Form["hq"], r.Form["ha"])
question := strings.TrimSpace(r.FormValue("question"))
view := chatView{
VideoID: row.VideoID,
Title: displayTitle(*row),
Available: hasText,
Models: a.Chat.Models(),
Selected: model,
History: history,
}
switch {
case !hasText:
// Honest "not available" — no fetch, no model call (ADR-027).
case question == "":
// A model switch with no text just re-renders — no wasted round-trip.
default:
reply, err := a.Chat.Answer(r.Context(), chat.Request{
Model: model,
Transcript: transcript,
History: history,
Question: question,
})
if err != nil {
a.logger().Error("chat answer", "video", row.VideoID, "model", model, "err", err)
view.Error = "That model couldn't answer just now. Try again, or switch models."
} else {
view.History = appendTurn(history, chat.Turn{Question: question, Answer: reply.Answer})
view.Truncated = reply.Truncated
}
}
a.renderChatTurn(w, r, *row, view)
}
// loadOwnedSummary fetches the summary for the path's video scoped to userID, or
// writes the right response (404 on not-found/not-owned, 500 on error) and reports
// false. It is the single isolation gate for both chat handlers.
func (a *App) loadOwnedSummary(w http.ResponseWriter, r *http.Request, userID string) (*store.SummaryRow, bool) {
videoID := r.PathValue("videoId")
row, err := a.Store.GetSummaryByVideo(r.Context(), userID, videoID)
if errors.Is(err, store.ErrNotFound) {
http.NotFound(w, r)
return nil, false
}
if err != nil {
a.serverError(w, r, "chat get summary", err)
return nil, false
}
return row, true
}
// readTranscript reads the shared stored transcript for a row (ADR-021) and
// reports whether it carries usable text. It is a pure DB read — NO caption fetch,
// the property the whole feature's safety rests on. ok is false only on a store
// error (after a 500 is written); a missing/text-less transcript is (",", false,
// true) — the honest "not available" case, handled by the caller, not an error.
func (a *App) readTranscript(w http.ResponseWriter, r *http.Request, row store.SummaryRow) (content string, hasText, ok bool) {
t, found, err := a.Store.GetTranscript(r.Context(), row.Channel, row.ProviderVideoID)
if err != nil {
a.serverError(w, r, "chat get transcript", err)
return "", false, false
}
if !found || !t.HasText() {
return "", false, true
}
return t.Content, true, true
}
// renderChatTurn returns the chat panel fragment for an HTMX answer (swapped in
// place within the open dock), or the full chat page otherwise — the no-JS POST
// re-renders the whole summary + open chat with the new turn.
func (a *App) renderChatTurn(w http.ResponseWriter, r *http.Request, row store.SummaryRow, v chatView) {
if isHTMX(r) {
a.render(w, r, chatPanel(v))
return
}
a.render(w, r, ChatPage(row, v))
}
// resolveModel keeps the posted model only when it is an offered option; anything
// else (a forged value, or a model dropped because cloud is disabled) falls back
// to the default. The chat.Service enforces the same guard before the gateway;
// this keeps the rendered switcher honest too.
func (a *App) resolveModel(posted, summaryModel string) string {
for _, m := range a.Chat.Models() {
if m == posted {
return posted
}
}
return a.Chat.DefaultModel(summaryModel)
}
// parseHistory zips the parallel hidden hq/ha fields back into ordered turns,
// keeping only the most recent maxHistoryTurns. net/url preserves the submission
// order of repeated fields, so the pairing is stable.
func parseHistory(qs, as []string) []chat.Turn {
n := len(qs)
if len(as) < n {
n = len(as)
}
turns := make([]chat.Turn, 0, n)
for i := 0; i < n; i++ {
turns = append(turns, chat.Turn{Question: qs[i], Answer: as[i]})
}
return capHistory(turns)
}
// appendTurn adds a completed exchange and re-bounds the conversation.
func appendTurn(history []chat.Turn, t chat.Turn) []chat.Turn {
return capHistory(append(history, t))
}
// capHistory keeps the last maxHistoryTurns turns (drops the oldest).
func capHistory(turns []chat.Turn) []chat.Turn {
if len(turns) <= maxHistoryTurns {
return turns
}
return turns[len(turns)-maxHistoryTurns:]
}
+375
View File
@@ -0,0 +1,375 @@
package web_test
import (
"context"
"net/http"
"net/http/httptest"
"net/url"
"strings"
"testing"
"time"
"github.com/jackc/pgx/v5/pgxpool"
"github.com/stretchr/testify/require"
"gitea.d-ma.be/mathias/tapir/internal/adapters/chat"
"gitea.d-ma.be/mathias/tapir/internal/adapters/store"
"gitea.d-ma.be/mathias/tapir/internal/domain"
"gitea.d-ma.be/mathias/tapir/internal/web"
)
// videoZ is a video id used by the isolation test for a DIFFERENT user's video.
const videoZ = "33333333-3333-3333-3333-333333333333"
// --- chat test doubles -----------------------------------------------------
// fakeChatter is a web.Chatter that records the request it received and returns a
// canned reply. It performs NO network and NO fetch — it stands in for the real
// chat.Service so the handler behaviour is what's under test.
type fakeChatter struct {
models []string
reply chat.Reply
err error
gotReq chat.Request
callCount int
}
func (f *fakeChatter) Models() []string { return f.models }
func (f *fakeChatter) DefaultModel(summaryModel string) string {
for _, m := range f.models {
if m == summaryModel {
return summaryModel
}
}
if len(f.models) > 0 {
return f.models[0]
}
return ""
}
func (f *fakeChatter) Answer(_ context.Context, req chat.Request) (chat.Reply, error) {
f.callCount++
f.gotReq = req
if f.err != nil {
return chat.Reply{}, f.err
}
return f.reply, nil
}
// tripwireProcessor and tripwireFetcher are the YouTube-reaching collaborators
// (summarize → caption fetch, and the paste metadata fetch). Wired into the App
// for the safety test, they fail it the instant chat routes into either — the
// behavioural proof that chat never triggers a fetch (ADR-027).
type tripwireProcessor struct{ t *testing.T }
func (p tripwireProcessor) ProcessVideo(context.Context, string, string) error {
p.t.Fatal("chat triggered summarization (→ caption fetch) — must never happen (ADR-027)")
return nil
}
type tripwireFetcher struct{ t *testing.T }
func (f tripwireFetcher) FetchVideo(context.Context, string, string) (domain.Video, error) {
f.t.Fatal("chat triggered a YouTube video fetch — must never happen (ADR-027)")
return domain.Video{}, nil
}
// --- helpers ---------------------------------------------------------------
// newChatApp builds the App under test as the registered stub user, with a Chat
// backend wired. Mirrors newApp but adds chat (and any extra wiring via mutate).
func newChatApp(t *testing.T, chatter web.Chatter, mutate func(*web.App)) *web.App {
t.Helper()
s := newStore(t)
app := &web.App{
Store: s,
Identity: s,
Auth: web.StubAuth{U: web.User{Subject: stubSubject}},
Chat: chatter,
}
if mutate != nil {
mutate(app)
}
return app
}
// seededProviderVideoID mirrors seedVideo's derivation so the chat path's
// (provider, providerVideoID) transcript key matches the seeded video row.
func seededProviderVideoID(videoID string) string { return "pv-" + videoID[:8] }
// seedTranscript stores a shared (provider, providerVideoID) transcript — the
// ADR-021 stored content the chat reads. Captions source = usable text.
func seedTranscript(t *testing.T, s *store.Store, videoID, content string) {
t.Helper()
err := s.SaveTranscript(context.Background(), "youtube", seededProviderVideoID(videoID),
domain.Transcript{Source: domain.SourceCaptions, Language: "en", Content: content})
require.NoError(t, err)
}
func getChat(t *testing.T, app *web.App, videoID string) *httptest.ResponseRecorder {
t.Helper()
return do(t, app, httptest.NewRequest(http.MethodGet, "/v/"+videoID+"/chat", nil))
}
func postChat(t *testing.T, app *web.App, videoID string, form url.Values, htmx bool) *httptest.ResponseRecorder {
t.Helper()
req := httptest.NewRequest(http.MethodPost, "/v/"+videoID+"/chat", strings.NewReader(form.Encode()))
req.Header.Set("Content-Type", "application/x-www-form-urlencoded")
if htmx {
req.Header.Set("HX-Request", "true")
}
return do(t, app, req)
}
// --- scenarios -------------------------------------------------------------
// The "dig deeper" affordance appears on a summary detail view only when chat is
// enabled, and links to that video's chat — the single entry point (ADR-027 §1).
func TestChatEntryAffordanceOnSummaryView(t *testing.T) {
ctx := context.Background()
withChat := newChatApp(t, &fakeChatter{models: []string{"phi4-mini"}}, nil)
p := rawPool(t)
resetDB(t, p)
require.NoError(t, deliver(ctx, withChat, videoX, "body x"))
seedVideo(t, p, videoX, "X Title", "https://x", time.Time{})
html := body(t, do(t, withChat, httptest.NewRequest(http.MethodGet, "/v/"+videoX, nil)))
require.Contains(t, html, "Dig deeper", "the deeper-dive affordance is shown when chat is enabled")
require.Contains(t, html, "/v/"+videoX+"/chat", "it links to this video's chat")
require.Contains(t, html, `id="chat-section"`, "the dock lives on the detail page")
require.Contains(t, html, `hx-get="/v/`+videoX+`/chat"`, "it opens the chat in place (HTMX), not a navigation")
// With no chat backend wired the affordance is absent (routes unmounted).
noChat := newApp(t)
require.NoError(t, deliver(ctx, noChat, videoX, "body x"))
html = body(t, do(t, noChat, httptest.NewRequest(http.MethodGet, "/v/"+videoX, nil)))
require.NotContains(t, html, "Dig deeper", "no affordance when chat is disabled")
}
// The summary and the chat live together (the integrated UX): the no-JS chat page
// renders the full summary alongside the chat, and the HTMX reveal returns just
// the open chat section as a fragment so it docks in below the summary already on
// screen — the summary is never navigated away from.
func TestChatIntegratedWithSummaryOnSamePage(t *testing.T) {
ctx := context.Background()
chatter := &fakeChatter{models: []string{"phi4-mini", "gemma4-26b"}}
app := newChatApp(t, chatter, nil)
p := rawPool(t)
resetDB(t, p)
require.NoError(t, deliver(ctx, app, videoX, "SUMMARY-BODY-MARKER"))
seedVideo(t, p, videoX, "X Title", "https://x", time.Time{})
seedTranscript(t, app.Store.(*store.Store), videoX, "the transcript")
// No-JS full page: the summary payload and the chat are on one page.
html := body(t, getChat(t, app, videoX))
require.Contains(t, html, "SUMMARY-BODY-MARKER", "the summary text is shown on the chat page")
require.Contains(t, html, "takeaway one", "takeaways shown alongside the chat")
require.Contains(t, html, "highlight one", "highlights shown alongside the chat")
require.Contains(t, html, "Ask about this video", "the chat sits on the same page as the summary")
require.Contains(t, html, `name="question"`, "the ask form is present")
// HTMX reveal: the open chat section ONLY (a fragment) — no full-page chrome and
// no duplicated summary, so it swaps in below the summary already rendered.
req := httptest.NewRequest(http.MethodGet, "/v/"+videoX+"/chat", nil)
req.Header.Set("HX-Request", "true")
frag := body(t, do(t, app, req))
require.NotContains(t, frag, "<html", "the reveal is a fragment, not a full page")
require.NotContains(t, frag, "SUMMARY-BODY-MARKER", "the reveal does not re-send the summary (it's already on screen)")
require.Contains(t, frag, `id="chat-section"`, "the fragment replaces the dock in place")
require.Contains(t, frag, `name="question"`, "the ask form is in the revealed section")
}
// THE KEY SAFETY ASSERTION (ADR-027): a chat answer is produced entirely from the
// stored transcript — the model receives the stored text, and neither the
// summarize→fetch path nor the YouTube fetch path is ever touched. The tripwire
// collaborators t.Fatal the test if chat reaches them.
func TestChatAnswersFromStoredTranscriptWithoutAnyFetch(t *testing.T) {
ctx := context.Background()
const transcript = "STORED-TRANSCRIPT-MARKER: the host explains attention budgets."
chatter := &fakeChatter{
models: []string{"phi4-mini"},
reply: chat.Reply{Answer: "It is about attention budgets."},
}
app := newChatApp(t, chatter, func(a *web.App) {
a.Processor = tripwireProcessor{t} // fails the test if chat summarizes/fetches
a.Fetcher = tripwireFetcher{t} // fails the test if chat fetches video metadata
})
p := rawPool(t)
resetDB(t, p)
require.NoError(t, deliver(ctx, app, videoX, "summary body"))
seedVideo(t, p, videoX, "X Title", "https://x", time.Time{})
seedTranscript(t, app.Store.(*store.Store), videoX, transcript)
form := url.Values{"question": {"what is it about?"}, "model": {"phi4-mini"}}
rec := postChat(t, app, videoX, form, true)
require.Equal(t, http.StatusOK, rec.Code)
require.Contains(t, body(t, rec), "It is about attention budgets.", "the answer is rendered")
require.Equal(t, 1, chatter.callCount, "the model was asked exactly once")
require.Equal(t, transcript, chatter.gotReq.Transcript,
"the model answered from the STORED transcript, not a fetched one")
// The tripwires never firing IS the no-fetch / no-YouTube proof.
}
// A video with no stored transcript yields an honest "not available" — and never
// a fetch, never a model call (ADR-027 §2: do not add an on-demand-fetch path).
func TestChatUnavailableWhenNoStoredTranscript(t *testing.T) {
ctx := context.Background()
chatter := &fakeChatter{models: []string{"phi4-mini"}}
app := newChatApp(t, chatter, func(a *web.App) {
a.Processor = tripwireProcessor{t}
a.Fetcher = tripwireFetcher{t}
})
p := rawPool(t)
resetDB(t, p)
require.NoError(t, deliver(ctx, app, videoX, "summary body"))
seedVideo(t, p, videoX, "X Title", "https://x", time.Time{})
// NB: no seedTranscript — the transcript is absent.
rec := getChat(t, app, videoX)
require.Equal(t, http.StatusOK, rec.Code)
require.Contains(t, body(t, rec), "isn't available", "honest not-available copy")
// Asking anyway still triggers nothing: no model call, no fetch.
rec = postChat(t, app, videoX, url.Values{"question": {"hi"}, "model": {"phi4-mini"}}, true)
require.Equal(t, http.StatusOK, rec.Code)
require.Equal(t, 0, chatter.callCount, "no model call without a stored transcript")
}
// Chat is reachable ONLY from the user's own summary view: opening chat for a
// video that belongs to another user 404s (the summary read is RLS-scoped), so a
// guessed/arbitrary video id is not a chat surface (ADR-027 §5 isolation).
func TestChatOnlyReachableForOwnVideo(t *testing.T) {
chatter := &fakeChatter{models: []string{"phi4-mini"}}
app := newChatApp(t, chatter, func(a *web.App) {
a.Processor = tripwireProcessor{t}
a.Fetcher = tripwireFetcher{t}
})
p := rawPool(t)
resetDB(t, p)
// A summarized video with a stored transcript owned by ANOTHER user.
const other = "22222222-2222-2222-2222-222222222222"
seedForeignSummaryWithTranscript(t, p, other, videoZ, "Foreign Title")
rec := getChat(t, app, videoZ)
require.Equal(t, http.StatusNotFound, rec.Code, "cannot open chat for another user's video")
require.Equal(t, 0, chatter.callCount, "no model call for a non-owned video")
rec = postChat(t, app, videoZ, url.Values{"question": {"hi"}, "model": {"phi4-mini"}}, true)
require.Equal(t, http.StatusNotFound, rec.Code, "cannot post chat to another user's video")
}
// The switcher offers exactly the backend's models, defaults to the summary's own
// model, and a switch re-runs against the SAME transcript with the chosen model
// (ADR-027 §3 — model-comparison instrumentation). A model dropped from the offer
// set (e.g. cloud disabled) is not rendered.
func TestChatModelSwitcherDefaultAndSwitch(t *testing.T) {
ctx := context.Background()
// The seeded summary's model is "phi4-mini" (see handlers_test summary()).
chatter := &fakeChatter{
models: []string{"phi4-mini", "gemma4-26b"}, // note: no cloud model offered
reply: chat.Reply{Answer: "answer"},
}
app := newChatApp(t, chatter, nil)
p := rawPool(t)
resetDB(t, p)
require.NoError(t, deliver(ctx, app, videoX, "summary body"))
seedVideo(t, p, videoX, "X Title", "https://x", time.Time{})
seedTranscript(t, app.Store.(*store.Store), videoX, "the transcript text")
html := body(t, getChat(t, app, videoX))
require.Contains(t, html, "gemma4-26b", "every offered model is in the switcher")
require.NotContains(t, html, "mistral", "a non-offered (cloud-disabled) model is absent")
require.Contains(t, html, `value="phi4-mini" selected`, "defaults to the summary's own model")
// Switching to gemma re-runs against the same stored transcript.
form := url.Values{"question": {"q"}, "model": {"gemma4-26b"}}
rec := postChat(t, app, videoX, form, true)
require.Equal(t, http.StatusOK, rec.Code)
require.Equal(t, "gemma4-26b", chatter.gotReq.Model, "the chosen model answers")
require.Equal(t, "the transcript text", chatter.gotReq.Transcript, "against the same transcript")
}
// A truncated transcript surfaces the honest bounded-context note (ADR-027 §2).
func TestChatTruncationNoteShown(t *testing.T) {
ctx := context.Background()
chatter := &fakeChatter{
models: []string{"phi4-mini"},
reply: chat.Reply{Answer: "answer", Truncated: true},
}
app := newChatApp(t, chatter, nil)
p := rawPool(t)
resetDB(t, p)
require.NoError(t, deliver(ctx, app, videoX, "summary body"))
seedVideo(t, p, videoX, "X Title", "https://x", time.Time{})
seedTranscript(t, app.Store.(*store.Store), videoX, "long transcript")
rec := postChat(t, app, videoX, url.Values{"question": {"q"}, "model": {"phi4-mini"}}, true)
require.Contains(t, body(t, rec), "bounded portion", "the truncation note is shown")
}
// A multi-turn conversation is carried in the request (hidden fields), not the DB:
// prior turns ride back into the next answer, and nothing is persisted (ADR-027 §4).
func TestChatMultiTurnHistoryIsEphemeral(t *testing.T) {
ctx := context.Background()
chatter := &fakeChatter{models: []string{"phi4-mini"}, reply: chat.Reply{Answer: "second answer"}}
app := newChatApp(t, chatter, nil)
p := rawPool(t)
resetDB(t, p)
require.NoError(t, deliver(ctx, app, videoX, "summary body"))
seedVideo(t, p, videoX, "X Title", "https://x", time.Time{})
seedTranscript(t, app.Store.(*store.Store), videoX, "the transcript")
// Second turn posts the prior exchange as hidden history fields.
form := url.Values{
"question": {"follow-up question"},
"model": {"phi4-mini"},
"hq": {"first question"},
"ha": {"first answer"},
}
rec := postChat(t, app, videoX, form, true)
require.Equal(t, http.StatusOK, rec.Code)
require.Len(t, chatter.gotReq.History, 1, "the prior turn was carried into the request")
require.Equal(t, "first question", chatter.gotReq.History[0].Question)
require.Equal(t, "first answer", chatter.gotReq.History[0].Answer)
html := body(t, rec)
require.Contains(t, html, "second answer", "the new answer renders")
require.Contains(t, html, "first question", "the conversation persists in page state")
// Nothing was written to the DB — there is no chat table; the summaries/actions
// are unchanged by a chat turn.
row, err := app.Store.GetSummaryByVideo(ctx, userID, videoX)
require.NoError(t, err)
require.Equal(t, "summary body", row.Summary, "chat never mutates stored data")
}
// --- foreign-user seeding (isolation) --------------------------------------
// seedForeignSummaryWithTranscript creates a fully-summarized, transcript-backed
// video owned by a DIFFERENT user via the raw pool (which bypasses RLS), so the
// isolation test can confirm the requesting user cannot open chat for it.
func seedForeignSummaryWithTranscript(t *testing.T, p *pgxpool.Pool, otherUserID, videoID, title string) {
t.Helper()
ctx := context.Background()
_, err := p.Exec(ctx, `INSERT INTO users (id) VALUES ($1) ON CONFLICT (id) DO NOTHING`, otherUserID)
require.NoError(t, err)
_, err = p.Exec(ctx,
`INSERT INTO videos (id, user_id, provider, provider_video_id, title, url)
VALUES ($1, $2, 'youtube', $3, $4, 'https://z')`,
videoID, otherUserID, seededProviderVideoID(videoID), title)
require.NoError(t, err)
_, err = p.Exec(ctx,
`INSERT INTO summaries (user_id, video_id, summary, ai_provider, ai_model)
VALUES ($1, $2, 'foreign summary', 'local', 'phi4-mini')`,
otherUserID, videoID)
require.NoError(t, err)
_, err = p.Exec(ctx,
`INSERT INTO transcripts (provider, provider_video_id, source, language, content)
VALUES ('youtube', $1, 'captions', 'en', 'foreign transcript')`,
seededProviderVideoID(videoID))
require.NoError(t, err)
}
+16 -1
View File
@@ -26,6 +26,10 @@ type Store interface {
DistinctChannels(ctx context.Context, userID string) ([]string, error) DistinctChannels(ctx context.Context, userID string) ([]string, error)
GetSummaryByVideo(ctx context.Context, userID, videoID string) (*store.SummaryRow, error) GetSummaryByVideo(ctx context.Context, userID, videoID string) (*store.SummaryRow, error)
GetVideoRow(ctx context.Context, userID, videoID string) (*store.SummaryRow, error) GetVideoRow(ctx context.Context, userID, videoID string) (*store.SummaryRow, error)
// GetTranscript reads the shared, stored transcript keyed by (provider,
// providerVideoID) — ADR-021. It is a pure DB read: it never fetches captions,
// so the chat path (ADR-027) reaches it without any caption-fetch surface.
GetTranscript(ctx context.Context, provider, providerVideoID string) (domain.Transcript, bool, error)
ActionsFor(ctx context.Context, userID string, videoIDs []string) (map[string][]string, error) ActionsFor(ctx context.Context, userID string, videoIDs []string) (map[string][]string, error)
SetAction(ctx context.Context, userID, videoID, action string) error SetAction(ctx context.Context, userID, videoID, action string) error
ClearAction(ctx context.Context, userID, videoID, action string) error ClearAction(ctx context.Context, userID, videoID, action string) error
@@ -91,6 +95,11 @@ type App struct {
// Fetcher, when non-nil, resolves an arbitrary YouTube video id to metadata for // Fetcher, when non-nil, resolves an arbitrary YouTube video id to metadata for
// the paste-a-URL flow (Feature 2). Nil = the /paste route is not mounted. // the paste-a-URL flow (Feature 2). Nil = the /paste route is not mounted.
Fetcher VideoFetcher Fetcher VideoFetcher
// Chat, when non-nil, answers per-video questions against a video's STORED
// transcript (ADR-027). Nil = the /v/{id}/chat routes are not mounted and the
// summary view shows no "dig deeper" affordance. It holds no caption-fetch
// dependency, so an enabled chat cannot reach YouTube or the rate gate.
Chat Chatter
// Processing tracks in-flight immediate summarizations so the status endpoint // Processing tracks in-flight immediate summarizations so the status endpoint
// shows the animation until the summary lands. The zero value is ready to use. // shows the animation until the summary lands. The zero value is ready to use.
Processing ProcessingSet Processing ProcessingSet
@@ -149,6 +158,12 @@ func (a *App) Router() http.Handler {
app.HandleFunc("POST /paste", a.handlePaste) app.HandleFunc("POST /paste", a.handlePaste)
} }
app.HandleFunc("GET /v/{videoId}/status", a.handleStatus) app.HandleFunc("GET /v/{videoId}/status", a.handleStatus)
// Per-video deeper-dive chat over the STORED transcript (ADR-027). Mounted only
// when a Chat backend is wired; it never fetches captions.
if a.Chat != nil {
app.HandleFunc("GET /v/{videoId}/chat", a.handleChat)
app.HandleFunc("POST /v/{videoId}/chat", a.handleChatMessage)
}
app.HandleFunc("GET /register", a.handleRegisterForm) app.HandleFunc("GET /register", a.handleRegisterForm)
app.HandleFunc("POST /register", a.handleRegister) app.HandleFunc("POST /register", a.handleRegister)
@@ -265,7 +280,7 @@ func (a *App) handleDetail(w http.ResponseWriter, r *http.Request) {
a.serverError(w, r, "get summary", err) a.serverError(w, r, "get summary", err)
return return
} }
a.render(w, r, DetailPage(*row)) a.render(w, r, DetailPage(*row, a.Chat != nil))
} }
// handleAction toggles one action: re-clicking an active verb clears it, else it // handleAction toggles one action: re-clicking an active verb clears it, else it
+25 -2
View File
@@ -7,6 +7,7 @@ import (
"net/http" "net/http"
"net/http/httptest" "net/http/httptest"
"os" "os"
"path/filepath"
"strings" "strings"
"testing" "testing"
"time" "time"
@@ -26,10 +27,22 @@ import (
var dsn string var dsn string
func TestMain(m *testing.M) { func TestMain(m *testing.M) {
const port = 54330 // distinct from the store package's embedded PG (54329) // Per-process port + dirs so concurrent `go test` runs (e.g. a push-run and a
// tag-run in CI) never collide on a fixed port or shared data dir. Base 55000
// keeps web's range distinct from the store package (54000). Shared CachePath
// downloads the PG archive once.
port := uint32(55000 + os.Getpid()%1000)
dsn = fmt.Sprintf("postgres://postgres:postgres@localhost:%d/postgres?sslmode=disable", port) dsn = fmt.Sprintf("postgres://postgres:postgres@localhost:%d/postgres?sslmode=disable", port)
pg := embeddedpostgres.NewDatabase(embeddedpostgres.DefaultConfig().Port(port)) rt := filepath.Join(os.TempDir(), fmt.Sprintf("tapir-epg-web-%d", os.Getpid()))
pg := embeddedpostgres.NewDatabase(
embeddedpostgres.DefaultConfig().
Port(port).
RuntimePath(rt).
DataPath(filepath.Join(rt, "data")).
BinariesPath(filepath.Join(rt, "bin")).
CachePath(filepath.Join(os.TempDir(), "tapir-epg-cache")),
)
if err := pg.Start(); err != nil { if err := pg.Start(); err != nil {
fmt.Fprintf(os.Stderr, "embedded-postgres start: %v\n", err) fmt.Fprintf(os.Stderr, "embedded-postgres start: %v\n", err)
os.Exit(1) os.Exit(1)
@@ -38,6 +51,7 @@ func TestMain(m *testing.M) {
if err := pg.Stop(); err != nil { if err := pg.Stop(); err != nil {
fmt.Fprintf(os.Stderr, "embedded-postgres stop: %v\n", err) fmt.Fprintf(os.Stderr, "embedded-postgres stop: %v\n", err)
} }
_ = os.RemoveAll(rt)
os.Exit(code) os.Exit(code)
} }
@@ -510,3 +524,12 @@ func TestListAutoModeBannerCopy(t *testing.T) {
require.Contains(t, html, "land gradually") require.Contains(t, html, "land gradually")
require.NotContains(t, html, "are not summarized automatically") require.NotContains(t, html, "are not summarized automatically")
} }
// TestMetricsNotOnPublicMux: the public app router exposes no /metrics route —
// Prometheus is served on the dedicated metrics port only (ADR-030, security R6).
func TestMetricsNotOnPublicMux(t *testing.T) {
app := newApp(t)
resetDB(t, rawPool(t))
rec := do(t, app, httptest.NewRequest(http.MethodGet, "/metrics", nil))
require.Equal(t, http.StatusNotFound, rec.Code, "/metrics must not be on the public mux")
}
+38 -44
View File
@@ -7,10 +7,11 @@
// Authentication is real (Dex OIDC) and is the only gate: any Dex-authenticated // Authentication is real (Dex OIDC) and is the only gate: any Dex-authenticated
// subject may sign in (ADR-012 dropped ADR-011's single-subject allowlist). // subject may sign in (ADR-012 dropped ADR-011's single-subject allowlist).
// Authorization/registration is layered on top in internal/web (an authenticated // Authorization/registration is layered on top in internal/web (an authenticated
// subject with no tapir user is routed to registration). Sessions are server-side // subject with no tapir user is routed to registration). Sessions are STATELESS
// (in-memory, fine for the single Stage-1 replica) addressed by an HMAC-signed // (ADR-029): the identity + expiry live inside an HMAC-signed (HS256) HttpOnly
// (HS256) HttpOnly Secure SameSite=Lax cookie with a short TTL and sliding // Secure SameSite=Lax persistent cookie with a long sliding TTL — no server-side
// refresh. Tokens are never logged. // table, so a deploy/restart never logs anyone out and the cookie also survives
// browser-close. Tokens are never logged; logout clears the cookie client-side.
// //
// This is mcp-chassis's cousin but NOT the same code: mcp-chassis validates // This is mcp-chassis's cousin but NOT the same code: mcp-chassis validates
// inbound Bearer JWTs for MCP APIs; this is a browser session login. // inbound Bearer JWTs for MCP APIs; this is a browser session login.
@@ -26,6 +27,7 @@ import (
"github.com/coreos/go-oidc/v3/oidc" "github.com/coreos/go-oidc/v3/oidc"
"golang.org/x/oauth2" "golang.org/x/oauth2"
"gitea.d-ma.be/mathias/tapir/internal/metrics"
"gitea.d-ma.be/mathias/tapir/internal/web" "gitea.d-ma.be/mathias/tapir/internal/web"
) )
@@ -47,7 +49,11 @@ type Config struct {
} }
const ( const (
defaultSessionTTL = time.Hour // defaultSessionTTL is generous and sliding: Tapir is a "check back tomorrow"
// reader, so a short TTL meant a re-login (full IdP redirect dance) on almost
// every visit. 30 days, slid forward on each request, keeps a regular user
// logged in indefinitely while an abandoned session still lapses.
defaultSessionTTL = 30 * 24 * time.Hour
pendingTTL = 10 * time.Minute pendingTTL = 10 * time.Minute
sessionCookie = "tapir_session" sessionCookie = "tapir_session"
loginPath = "/auth/login" loginPath = "/auth/login"
@@ -59,7 +65,6 @@ type DexAuth struct {
oauth *oauth2.Config oauth *oauth2.Config
verifier *oidc.IDTokenVerifier verifier *oidc.IDTokenVerifier
sessions *sessionStore
pending *pendingStore pending *pendingStore
secret []byte secret []byte
sessionTTL time.Duration sessionTTL time.Duration
@@ -127,7 +132,6 @@ func New(ctx context.Context, cfg Config, opts ...Option) (*DexAuth, error) {
RedirectURL: cfg.RedirectURL, RedirectURL: cfg.RedirectURL,
Scopes: []string{oidc.ScopeOpenID, "profile", "email"}, Scopes: []string{oidc.ScopeOpenID, "profile", "email"},
}, },
sessions: newSessionStore(),
pending: newPendingStore(), pending: newPendingStore(),
secret: []byte(cfg.SessionSecret), secret: []byte(cfg.SessionSecret),
sessionTTL: defaultSessionTTL, sessionTTL: defaultSessionTTL,
@@ -159,31 +163,31 @@ func (d *DexAuth) Middleware(h http.Handler) http.Handler {
h.ServeHTTP(w, r) h.ServeHTTP(w, r)
return return
} }
sid, ok := d.sessionID(r) c, err := r.Cookie(sessionCookie)
if err != nil {
d.redirectUnauthenticated(w, r)
return
}
user, _, ok := d.decodeSession(c.Value, d.now())
if !ok { if !ok {
d.redirectUnauthenticated(w, r) d.redirectUnauthenticated(w, r)
return return
} }
if _, ok := d.sessions.get(sid, d.now()); !ok { // Sliding refresh: re-issue the cookie with a fresh expiry so an active
d.redirectUnauthenticated(w, r) // user never lapses (the expiry lives in the cookie, so sliding = re-sign).
return d.setSessionCookie(w, d.encodeSession(user, d.now().Add(d.sessionTTL)))
}
d.sessions.refresh(sid, d.now().Add(d.sessionTTL)) // sliding refresh
h.ServeHTTP(w, r) h.ServeHTTP(w, r)
}) })
} }
// CurrentUser resolves the authenticated principal from the session cookie. // CurrentUser resolves the authenticated principal from the stateless cookie.
func (d *DexAuth) CurrentUser(r *http.Request) (web.User, bool) { func (d *DexAuth) CurrentUser(r *http.Request) (web.User, bool) {
sid, ok := d.sessionID(r) c, err := r.Cookie(sessionCookie)
if !ok { if err != nil {
return web.User{}, false return web.User{}, false
} }
data, ok := d.sessions.get(sid, d.now()) user, _, ok := d.decodeSession(c.Value, d.now())
if !ok { return user, ok
return web.User{}, false
}
return data.user, true
} }
func (d *DexAuth) handleLogin(w http.ResponseWriter, r *http.Request) { func (d *DexAuth) handleLogin(w http.ResponseWriter, r *http.Request) {
@@ -248,23 +252,16 @@ func (d *DexAuth) handleCallback(w http.ResponseWriter, r *http.Request) {
} }
_ = idToken.Claims(&claims) // email is best-effort; subject is the identity _ = idToken.Claims(&claims) // email is best-effort; subject is the identity
sid, err := randToken() user := web.User{Subject: idToken.Subject, Email: claims.Email}
if err != nil { d.setSessionCookie(w, d.encodeSession(user, d.now().Add(d.sessionTTL)))
http.Error(w, "internal error", http.StatusInternalServerError) metrics.IncLogin()
return
}
d.sessions.put(sid, sessionData{
user: web.User{Subject: idToken.Subject, Email: claims.Email},
expiry: d.now().Add(d.sessionTTL),
})
d.setSessionCookie(w, sid)
http.Redirect(w, r, "/", http.StatusFound) http.Redirect(w, r, "/", http.StatusFound)
} }
func (d *DexAuth) handleLogout(w http.ResponseWriter, r *http.Request) { func (d *DexAuth) handleLogout(w http.ResponseWriter, r *http.Request) {
if sid, ok := d.sessionID(r); ok { // Stateless sessions: clearing the cookie logs the browser out. There is no
d.sessions.delete(sid) // server-side record to delete (ADR-029); a copy of the cookie stays valid
} // until its expiry — an accepted trade for the Stage-0 reader app.
d.clearSessionCookie(w) d.clearSessionCookie(w)
// Land on the public landing page, not the login endpoint: a just-logged-out // Land on the public landing page, not the login endpoint: a just-logged-out
// visitor should see /welcome, not be bounced straight back into a Dex login. // visitor should see /welcome, not be bounced straight back into a Dex login.
@@ -287,22 +284,19 @@ func (d *DexAuth) redirectToLogin(w http.ResponseWriter, r *http.Request) {
http.Redirect(w, r, loginPath, http.StatusFound) http.Redirect(w, r, loginPath, http.StatusFound)
} }
func (d *DexAuth) sessionID(r *http.Request) (string, bool) { // setSessionCookie writes the signed session value as a PERSISTENT cookie
c, err := r.Cookie(sessionCookie) // (Max-Age set), so it survives the browser/app being closed — a session cookie
if err != nil { // (no Max-Age) was dropped on iPhone Safari close, forcing re-login. value is the
return "", false // already-signed payload from encodeSession.
} func (d *DexAuth) setSessionCookie(w http.ResponseWriter, value string) {
return d.unsign(c.Value)
}
func (d *DexAuth) setSessionCookie(w http.ResponseWriter, sid string) {
http.SetCookie(w, &http.Cookie{ http.SetCookie(w, &http.Cookie{
Name: sessionCookie, Name: sessionCookie,
Value: d.sign(sid), Value: value,
Path: "/", Path: "/",
HttpOnly: true, HttpOnly: true,
Secure: !d.insecure, Secure: !d.insecure,
SameSite: http.SameSiteLaxMode, SameSite: http.SameSiteLaxMode,
MaxAge: int(d.sessionTTL.Seconds()),
}) })
} }
+33 -4
View File
@@ -306,13 +306,42 @@ func TestLogoutClearsSession(t *testing.T) {
require.Equal(t, http.StatusFound, rec.Code) require.Equal(t, http.StatusFound, rec.Code)
require.Equal(t, "/welcome", rec.Header().Get("Location"), "logout lands on the public page") require.Equal(t, "/welcome", rec.Header().Get("Location"), "logout lands on the public page")
cleared := sessionCookie(t, rec.Result()) cleared := sessionCookie(t, rec.Result())
require.Less(t, cleared.MaxAge, 0, "logout expires the cookie") require.Less(t, cleared.MaxAge, 0, "logout expires the cookie so the browser drops it")
require.Empty(t, cleared.Value, "logout blanks the cookie value")
// The server-side session is gone: the original cookie no longer resolves. // Sessions are stateless (ADR-029): logout clears the cookie client-side, so a
// request carrying the cleared (empty) cookie is unauthenticated. The original
// signed cookie remains technically valid until its expiry — the accepted
// trade for no server-side store; the browser no longer holds it.
check := httptest.NewRequest(http.MethodGet, "/", nil) check := httptest.NewRequest(http.MethodGet, "/", nil)
check.AddCookie(cookie) check.AddCookie(cleared)
_, ok := auth.CurrentUser(check) _, ok := auth.CurrentUser(check)
require.False(t, ok) require.False(t, ok, "the cleared cookie does not authenticate")
}
// TestSessionSurvivesRestart is the core of ADR-029: a cookie issued by one
// process is accepted by a FRESH instance with the same session secret — so a
// deploy/pod-restart no longer logs users out (the old in-memory store did).
func TestSessionSurvivesRestart(t *testing.T) {
f := newFakeIssuer(t)
auth1 := newAuth(t, f)
cookie := authenticate(t, auth1, f)
auth2 := newAuth(t, f) // simulate a redeploy: new process, same SessionSecret
req := httptest.NewRequest(http.MethodGet, "/", nil)
req.AddCookie(cookie)
user, ok := auth2.CurrentUser(req)
require.True(t, ok, "a session must survive a restart (stateless signed cookie)")
require.Equal(t, testSubject, user.Subject)
}
// TestSessionCookieIsPersistent: the cookie carries a positive Max-Age so it
// survives the browser/app being closed (a session cookie was dropped on iOS).
func TestSessionCookieIsPersistent(t *testing.T) {
f := newFakeIssuer(t)
auth := newAuth(t, f)
cookie := authenticate(t, auth, f)
require.Greater(t, cookie.MaxAge, 0, "session cookie must be persistent (Max-Age set)")
} }
func TestExpiredSessionRejected(t *testing.T) { func TestExpiredSessionRejected(t *testing.T) {
+31 -42
View File
@@ -6,6 +6,7 @@ import (
"crypto/sha256" "crypto/sha256"
"encoding/base64" "encoding/base64"
"encoding/hex" "encoding/hex"
"encoding/json"
"fmt" "fmt"
"strings" "strings"
"sync" "sync"
@@ -53,56 +54,44 @@ func (d *DexAuth) unsign(signed string) (string, bool) {
return value, true return value, true
} }
// sessionData is the server-side session record. // sessionClaims is the self-contained session payload carried INSIDE the signed
type sessionData struct { // cookie — there is no server-side session table. This is deliberate (ADR-029):
user web.User // an in-memory store was wiped on every pod restart, logging every user out on
expiry time.Time // each deploy, and a stateless cookie also survives browser-close and works
// across replicas. It holds only the identity (subject + email, not secret) and
// an absolute expiry; the HMAC tag (sign/unsign) makes it tamper-proof.
type sessionClaims struct {
Sub string `json:"s"`
Email string `json:"e"`
Exp int64 `json:"x"` // unix seconds; absolute expiry
} }
// sessionStore is an in-memory session table. Single replica at Stage 0, so an // encodeSession produces the signed cookie value for a user with the given expiry.
// in-process map is sufficient; it is safe for concurrent use. func (d *DexAuth) encodeSession(u web.User, exp time.Time) string {
type sessionStore struct { b, _ := json.Marshal(sessionClaims{Sub: u.Subject, Email: u.Email, Exp: exp.Unix()})
mu sync.Mutex return d.sign(base64.RawURLEncoding.EncodeToString(b))
m map[string]sessionData
} }
func newSessionStore() *sessionStore { return &sessionStore{m: make(map[string]sessionData)} } // decodeSession verifies the cookie's HMAC, parses the claims, and checks expiry.
// It returns the user and the absolute expiry on success.
func (s *sessionStore) put(id string, d sessionData) { func (d *DexAuth) decodeSession(cookieValue string, now time.Time) (web.User, time.Time, bool) {
s.mu.Lock() payload, ok := d.unsign(cookieValue)
defer s.mu.Unlock()
s.m[id] = d
}
// get returns the session if present and unexpired; expired entries are evicted.
func (s *sessionStore) get(id string, now time.Time) (sessionData, bool) {
s.mu.Lock()
defer s.mu.Unlock()
d, ok := s.m[id]
if !ok { if !ok {
return sessionData{}, false return web.User{}, time.Time{}, false
} }
if !now.Before(d.expiry) { raw, err := base64.RawURLEncoding.DecodeString(payload)
delete(s.m, id) if err != nil {
return sessionData{}, false return web.User{}, time.Time{}, false
} }
return d, true var c sessionClaims
} if err := json.Unmarshal(raw, &c); err != nil {
return web.User{}, time.Time{}, false
// refresh slides an existing session's expiry forward; a no-op for unknown ids.
func (s *sessionStore) refresh(id string, expiry time.Time) {
s.mu.Lock()
defer s.mu.Unlock()
if d, ok := s.m[id]; ok {
d.expiry = expiry
s.m[id] = d
} }
} exp := time.Unix(c.Exp, 0)
if !now.Before(exp) {
func (s *sessionStore) delete(id string) { return web.User{}, time.Time{}, false // expired
s.mu.Lock() }
defer s.mu.Unlock() return web.User{Subject: c.Sub, Email: c.Email}, exp, true
delete(s.m, id)
} }
// pendingData holds the nonce bound to an in-flight authorization request. // pendingData holds the nonce bound to an in-flight authorization request.
+30
View File
@@ -184,6 +184,12 @@ func statusURL(videoID string) templ.SafeURL {
return templ.SafeURL("/v/" + videoID + "/status") return templ.SafeURL("/v/" + videoID + "/status")
} }
// chatURL builds the per-video chat path (GET renders the page, POST answers) —
// the deeper-dive over the stored transcript (ADR-027).
func chatURL(videoID string) templ.SafeURL {
return templ.SafeURL("/v/" + videoID + "/chat")
}
// Charmbracelet-inspired palette for the summarizing animation (TapirSpinner) — // Charmbracelet-inspired palette for the summarizing animation (TapirSpinner) —
// a charm purple box, pink tapir, mint snout/eyes/progress. Kept as named consts // a charm purple box, pink tapir, mint snout/eyes/progress. Kept as named consts
// so the inline span colours and the CSS track/fill share one source of truth. // so the inline span colours and the CSS track/fill share one source of truth.
@@ -725,6 +731,30 @@ a.btn, a.btn:visited { color: var(--accent-fg); }
.detail ul { margin: 0; padding-left: 1.2rem; line-height: 1.6; } .detail ul { margin: 0; padding-left: 1.2rem; line-height: 1.6; }
.detail li { margin-bottom: var(--s1); } .detail li { margin-bottom: var(--s1); }
/* deeper-dive chat (ADR-027) — docks in place below the summary */
.chat-dock { margin-top: var(--s5); border-top: 1px solid var(--line); padding-top: var(--s4); }
.chat-dock .chat-open { display: inline-block; }
.chat-heading { font-size: 1.1rem; margin: 0 0 var(--s2); }
.chat-scope { margin: 0 0 var(--s3); font-size: .9rem; }
.chat-panel { display: flex; flex-direction: column; gap: var(--s3); }
.chat-log { display: flex; flex-direction: column; gap: var(--s3); }
.chat-turn { border-radius: var(--radius); padding: var(--s2) var(--s3); }
.chat-turn p { margin: 0; }
.chat-q { background: var(--accent-weak); color: var(--fg); align-self: flex-end; max-width: 85%; }
.chat-a { background: var(--card); border: 1px solid var(--line); }
.chat-a .body { white-space: pre-wrap; line-height: 1.6; }
.chat-note { margin: 0; font-size: .82rem; font-style: italic; }
.chat-error { margin: 0; color: #8a1c10; font-size: .9rem; }
@media (prefers-color-scheme: dark) { .chat-error { color: #f3b5ae; } }
.chat-form { display: flex; flex-direction: column; gap: var(--s2); margin: var(--s2) 0 0; }
.chat-model { flex-direction: column; display: flex; gap: var(--s1); font-size: .78rem; text-transform: uppercase; letter-spacing: .04em; color: var(--muted); align-items: flex-start; }
.chat-model select { font: inherit; text-transform: none; letter-spacing: 0; padding: .4rem .55rem; border: 1px solid var(--line); border-radius: var(--radius); background: var(--card); color: var(--fg); }
.chat-model-hint { text-transform: none; letter-spacing: 0; font-size: .78rem; }
.chat-form textarea { font: inherit; padding: .55rem; border: 1px solid var(--line); border-radius: var(--radius); background: var(--card); color: var(--fg); resize: vertical; }
.chat-form textarea:focus-visible { outline: 2px solid var(--accent); outline-offset: 1px; border-color: var(--accent); }
.chat-form .btn { align-self: flex-start; }
.chat-thinking { font-style: italic; }
/* action toggles — watched|skipped form one segmented control (they are mutually /* action toggles — watched|skipped form one segmented control (they are mutually
exclusive), "saved" sits apart as an independent toggle */ exclusive), "saved" sits apart as an independent toggle */
.actions { display: flex; gap: var(--s3); margin: var(--s4) 0; flex-wrap: wrap; align-items: center; } .actions { display: flex; gap: var(--s3); margin: var(--s4) 0; flex-wrap: wrap; align-items: center; }
+117 -11
View File
@@ -410,13 +410,11 @@ templ noCaptionsCard(r store.SummaryRow) {
</li> </li>
} }
// DetailPage is the full summary view: text, highlights, takeaways, metadata, // summaryBody is the summary payload shared by the detail page and the no-JS
// and the action button group. // chat page (so the chat page shows the same summary, not a separate view):
templ DetailPage(r store.SummaryRow) { // metadata, embed, source, the action toggles, then the attention-saving order
@Layout("Tapir — " + displayTitle(r)) { // Takeaways → Highlights → Summary (UX review A8).
<article class="detail"> templ summaryBody(r store.SummaryRow) {
<p class="back"><a href="/"> Summaries</a></p>
<h1>{ displayTitle(r) }</h1>
<p class="meta"> <p class="meta">
if detailMeta(r) != "" { if detailMeta(r) != "" {
<span>{ detailMeta(r) }</span> <span>{ detailMeta(r) }</span>
@@ -441,10 +439,6 @@ templ DetailPage(r store.SummaryRow) {
<p class="source"><a href={ externalURL(r.URL) } rel="noopener noreferrer">watch on source </a></p> <p class="source"><a href={ externalURL(r.URL) } rel="noopener noreferrer">watch on source </a></p>
} }
@ActionButtons(r.VideoID, actionSet(r.Actions)) @ActionButtons(r.VideoID, actionSet(r.Actions))
// Lead with the attention-saving payload: Takeaways ("is this worth my
// time?") first, then Highlights, then the full Summary last (UX review
// A8). Takeaways/Highlights are conditional, so a video without them falls
// through to the Summary leading naturally.
if len(r.Takeaways) > 0 { if len(r.Takeaways) > 0 {
<section> <section>
<h2>Takeaways</h2> <h2>Takeaways</h2>
@@ -469,10 +463,122 @@ templ DetailPage(r store.SummaryRow) {
<h2>Summary</h2> <h2>Summary</h2>
<p class="body">{ r.Summary }</p> <p class="body">{ r.Summary }</p>
</section> </section>
}
// DetailPage is the full summary view: the summary payload, then (when chat is
// enabled) the deeper-dive dock (ADR-027) — a reveal that opens the chat IN PLACE
// below the summary, so the summary stays on screen as the context being asked
// about rather than being navigated away from.
templ DetailPage(r store.SummaryRow, chatEnabled bool) {
@Layout("Tapir — " + displayTitle(r)) {
<article class="detail">
<p class="back"><a href="/"> Summaries</a></p>
<h1>{ displayTitle(r) }</h1>
@summaryBody(r)
if chatEnabled {
@chatReveal(r.VideoID)
}
</article> </article>
} }
} }
// chatReveal is the CLOSED dock at the foot of the summary: a quiet affordance,
// not a loud CTA (it deepens value for a reader already here, never nudges). With
// JS it swaps itself for the open chat section in place (HTMX, summary stays
// above); without JS the same href navigates to the full chat page, which renders
// the summary alongside the chat. Either way the summary is never lost.
templ chatReveal(videoID string) {
<section id="chat-section" class="chat-dock">
<a
class="btn-secondary chat-open"
href={ chatURL(videoID) }
hx-get={ string(chatURL(videoID)) }
hx-target="#chat-section"
hx-swap="outerHTML"
>
Dig deeper ask about this video
</a>
</section>
}
// chatSection is the OPEN dock: heading + scope note + the chat panel, swapped in
// over the closed reveal (same #chat-section id, outerHTML). It is the HTMX reveal
// response AND the inline chat block on the no-JS chat page.
templ chatSection(v chatView) {
<section id="chat-section" class="chat-dock chat-dock-open">
<h2 class="chat-heading">Ask about this video</h2>
<p class="chat-scope muted">Answers come only from this video's stored transcript Tapir never fetches anything new here.</p>
@chatPanel(v)
</section>
}
// ChatPage is the no-JS full-page render of the chat: the whole summary followed
// by the open chat dock, so a visitor without JS sees the same integrated view
// (summary beside the conversation) that JS users get inline via the reveal.
templ ChatPage(r store.SummaryRow, v chatView) {
@Layout("Tapir — " + displayTitle(r)) {
<article class="detail">
<p class="back"><a href="/"> Summaries</a></p>
<h1>{ displayTitle(r) }</h1>
@summaryBody(r)
@chatSection(v)
</article>
}
}
// chatPanel is the conversation + ask form, swapped in place on each answer
// (HTMX targets #chat-panel, outerHTML). When no transcript is stored it shows the
// honest "not available" state and no form (ADR-027: never a fetch). The prior
// turns ride as hidden hq/ha fields so the ephemeral conversation survives the
// round-trip without any persisted state.
templ chatPanel(v chatView) {
<div id="chat-panel" class="chat-panel">
if !v.Available {
<p class="chat-unavailable muted">Chat isn't available for this video its transcript isn't stored, and chat never fetches new captions. Summarize the video first to store its transcript.</p>
} else {
if len(v.History) > 0 {
<div class="chat-log">
for _, t := range v.History {
<div class="chat-turn chat-q"><p>{ t.Question }</p></div>
<div class="chat-turn chat-a"><p class="body">{ t.Answer }</p></div>
}
</div>
}
if v.Truncated {
<p class="chat-note muted">Working from a bounded portion of a long transcript answers about the end of the video may be incomplete.</p>
}
if v.Error != "" {
<p class="chat-error" role="alert">{ v.Error }</p>
}
<form
class="chat-form"
method="post"
action={ chatURL(v.VideoID) }
hx-post={ string(chatURL(v.VideoID)) }
hx-target="#chat-panel"
hx-swap="outerHTML"
>
for _, t := range v.History {
<input type="hidden" name="hq" value={ t.Question }/>
<input type="hidden" name="ha" value={ t.Answer }/>
}
<label class="chat-model">
Model
<select name="model">
for _, m := range v.Models {
<option value={ m } selected?={ m == v.Selected }>{ m }</option>
}
</select>
<span class="chat-model-hint muted">Switch models to compare answers on the same transcript.</span>
</label>
<textarea name="question" rows="3" placeholder="Ask a question about this video…" required aria-label="Your question"></textarea>
<button type="submit" class="btn">Ask</button>
<span class="htmx-indicator chat-thinking">thinking…</span>
</form>
}
</div>
}
// RegisterPage is the explicit registration step (ADR-012): an authenticated Dex // RegisterPage is the explicit registration step (ADR-012): an authenticated Dex
// subject with no tapir user picks a display name to create their account. // subject with no tapir user picks a display name to create their account.
// errMsg, when set, reports a validation problem on the prior POST. // errMsg, when set, reports a validation problem on the prior POST.
File diff suppressed because it is too large Load Diff
+25 -1
View File
@@ -25,6 +25,16 @@ import (
// fails if a scenario is unmapped, a mapped test is missing, or an entry no // fails if a scenario is unmapped, a mapped test is missing, or an entry no
// longer matches a real non-pending scenario. // longer matches a real non-pending scenario.
var scenarioCoverage = map[string]string{ var scenarioCoverage = map[string]string{
// observability.feature (ADR-030, #15)
"Summarization latency is recorded per endpoint": "TestSummarizerRecordsMetric",
"A failing summarizer endpoint records its failure outcome": "TestObserveSummarizeRecordsFailureOutcomes",
"Caption fetch latency is recorded by outcome": "TestObserveCaptionFetchByOutcome",
"LLM token usage is recorded from the completion": "TestClient_UsageHookRecordsTokens",
"Q&A answer latency is recorded": "TestChatAnswerLatencyRecorded",
"HTTP requests are counted by route, method, and status": "TestHTTPMiddlewareRecordsByRoutePattern",
"A successful login is counted": "TestLoginCounted",
"The metrics endpoint is not on the public app port": "TestMetricsNotOnPublicMux",
// ai_routing.feature // ai_routing.feature
"Local AI produces the summary": "TestSummarize_LocalSucceeds", "Local AI produces the summary": "TestSummarize_LocalSucceeds",
"Local AI fails and the user has a BYO provider configured": "TestSummarize_FallsBackToBYO", "Local AI fails and the user has a BYO provider configured": "TestSummarize_FallsBackToBYO",
@@ -40,7 +50,8 @@ var scenarioCoverage = map[string]string{
// connect_account.feature // connect_account.feature
"Connect a YouTube account": "TestCallbackExchangesAndRecordsConnection", "Connect a YouTube account": "TestCallbackExchangesAndRecordsConnection",
"Connecting an account discovers videos immediately": "TestCallbackTriggersDiscovery", "Connecting an account discovers videos immediately": "TestCallbackTriggersDiscovery",
"Connecting summarizes my newest videos right away": "TestNewestUnsummarizedVideoIDs", "Connecting summarizes my best recent videos right away": "TestOnboardBurstVideoIDs",
"The onboarding burst summarizes with a stronger model": "TestBurstChainModelsLeadsWithOnboardModel",
"Tokens are never stored in the clear": "TestCallbackExchangesAndRecordsConnection", "Tokens are never stored in the clear": "TestCallbackExchangesAndRecordsConnection",
// paste_url.feature // paste_url.feature
@@ -63,6 +74,19 @@ var scenarioCoverage = map[string]string{
"A returning subject passes straight through": "TestRegisteredSubjectPassesThrough", "A returning subject passes straight through": "TestRegisteredSubjectPassesThrough",
"Deleting an account removes only my data and leaves other users untouched": "TestDeleteAccountWipesDataAndSecretsAndLogsOut", "Deleting an account removes only my data and leaves other users untouched": "TestDeleteAccountWipesDataAndSecretsAndLogsOut",
// chat_transcript.feature (ADR-027)
"A summary view offers a deeper-dive into the video": "TestChatEntryAffordanceOnSummaryView",
"The summary and the chat are on one page": "TestChatIntegratedWithSummaryOnSamePage",
"Ask a question answered from the stored transcript": "TestChatAnswersFromStoredTranscriptWithoutAnyFetch",
"Chat never fetches captions or reaches YouTube": "TestChatAnswersFromStoredTranscriptWithoutAnyFetch",
"A video with no stored transcript offers no chat": "TestChatUnavailableWhenNoStoredTranscript",
"The default model is the summary's model and is switchable": "TestChatModelSwitcherDefaultAndSwitch",
"Switching models re-runs against the same transcript": "TestChatModelSwitcherDefaultAndSwitch",
"The cloud model is hidden when cloud is disabled": "TestChatCloudModelAbsentWhenDisabled",
"A long transcript is bounded and the chat says so": "TestChatTruncationNoteShown",
"A multi-turn conversation is ephemeral": "TestChatMultiTurnHistoryIsEphemeral",
"Chat is reachable only from my own summary view": "TestChatOnlyReachableForOwnVideo",
// summarize_new_video.feature // summarize_new_video.feature
"A subscribed channel posts a video that has captions": "TestSubscribedVideoWithCaptionsIsSummarizedAndDelivered", "A subscribed channel posts a video that has captions": "TestSubscribedVideoWithCaptionsIsSummarizedAndDelivered",
"A subscribed channel posts a video with no usable transcript": "TestVideoWithNoTranscriptIsSkipped", "A subscribed channel posts a video with no usable transcript": "TestVideoWithNoTranscriptIsSkipped",