Compare commits

...
161 Commits
Author SHA1 Message Date
mathiasandClaude Opus 5 489b4bb2a1 chore(spike): recover the /bygge toolchain from tmpfs (#28)
CI / Lint / Test / Vet (push) Successful in 37s
CI / Build & Import (push) Failing after 5s
CI / Deploy via GitOps (push) Skipped
analyze_srt.py, build_page.py and transcode.yaml produced the /bygge
prototype. They were sitting in a session scratchpad on tmpfs, one reboot
from gone, while spike issues #29-#31 were written as if they had to be
built from scratch.

Committed as found, defects documented rather than fixed: max_tokens=6000
truncates anything past ~6 minutes of audio, the page title is hardcoded,
and the schema still carries fields no consumer renders.

The four-check validator is the part worth keeping — the coverage-gap and
Swedish-number-grounding checks catch errors that timestamp and citation
checks structurally cannot.

Real transcripts and analyses stay out: this repo is public and the
recordings are a named person discussing a client's project. Only the
synthetic fixture is committed, which exercises all four checks with no
model call.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BeAp5LTscnz2W7Ubt4eJ6m
2026-08-13 11:37:03 +02:00
mathias 21e7d6c74b fix(ci): smoke test hung 14min — buildah run doesn't die under plain timeout
CI / Lint / Test / Vet (push) Successful in 15s
CI / Build & Import (push) Successful in 20s
CI / Deploy via GitOps (push) Successful in 5s
Bare `/tapir` runs the long-running server, same as the old ctr-based smoke
test. ctr's --rm reliably force-killed it; a plain `timeout N buildah run`
does not — it only signals the wrapper, and the container process can
survive that and keep the log pipe open, hanging the whole job (observed
live: run 155, 14min before failure). Backgrounds the run and tears it down
with `buildah rm -f`, which forcibly kills regardless of wrapper state, and
captures output via a file instead of a blocking pipe.

Refs infra#132.
2026-07-26 06:17:18 +00:00
mathias 8814ba6673 ci(smoke): use buildah run instead of sudo k3s ctr for smoke test
CI / Lint / Test / Vet (push) Successful in 12s
CI / Build & Import (push) Failing after 13m42s
CI / Deploy via GitOps (push) Has been skipped
Same fix as cobalt-dingo — removes the sudo/host-containerd dependency so
this still works once the act_runner is containerized (infra#132).

Refs infra#132.
2026-07-26 05:57:37 +00:00
mathias 4b557a4325 ci(deploy): add Flux GitOps deploy job, mirroring cobalt-dingo's pattern
CI / Lint / Test / Vet (push) Successful in 27s
CI / Build & Import (push) Successful in 18s
CI / Deploy via GitOps (push) Successful in 5s
Fixes silent stale-deploy gap: CI built+pushed images but nothing bumped
k3s/apps/tapir/deployment.yaml, so merged features sat CI-green with zero
production effect (infra#111, infra#168). Flux native image-automation
can't scan localhost:5000 from inside k3s pods, so this patches the infra
repo directly via the existing INFRA_DEPLOY_KEY org secret (same key
cobalt-dingo and brain-gardener already use) on every push to main.

Refs infra#111.
2026-07-25 21:21:02 +00:00
mathiasandClaude Opus 4.8 72cb25111f test(store): address migrations by version, not step count (#8)
CI / Lint / Test / Vet (push) Successful in 12s
CI / Build & Import (push) Successful in 12s
The up/down migration tests stepped a hard-coded number of Steps(-N)/Steps(+N)
down from HEAD and back. The counts assumed a specific latest migration, so
adding one shifted every count by one and unrelated tests (010/011/014) went
red with confusing off-by-one symptoms — a papercut on every new migration.

Drive the schema to an exact version with m.Migrate(version) via two helpers
(headVersion, migrateTo). Each test now steps to just below its target by
version, asserts the down effect, steps up to the target, asserts the up
effect, then restores to the captured HEAD. A migration added on top changes
HEAD but shifts no count, so no test needs editing.

Verified by adding a throwaway migration 017 on top: all four tests stayed
green with zero edits.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QbdxXWxLefS5AwLN5eyze
2026-07-02 14:49:02 +02:00
mathiasandClaude Sonnet 4.6 38f222c931 chore: rename Go module path gitea.d-ma.be → git.d-ma.be
CI / Lint / Test / Vet (push) Successful in 11s
CI / Build & Import (push) Successful in 12s
Infra ADR-0004 renamed the Gitea host. Bulk replace across go.mod and
all .go import paths. Build and tests pass unchanged.

Closes #20

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Dt6aHEDWRjkK14Voi6HnGh
2026-07-02 14:37:33 +02:00
mathias e042b6d26d chore: add .mcp.json (brain + gitea) — wire tapir for the single-harness workflow
CI / Lint / Test / Vet (push) Successful in 16s
CI / Build & Import (push) Successful in 12s
tapir had no MCP config at all — a hyperguild session here couldn't reach brain
or Gitea, breaking the report-back path before it could start. Matches the
shape now standard post-consolidation (hyperguild#75/#76, infra#178 template
refresh). Prerequisite for tapir#20 being run as the first real workflow test.
2026-07-02 07:00:53 +00:00
mathias c85a32e770 Merge pull request 'LANGUAGE.md: fix LiteLLM gateway never-say (stale endpoints, not piguard)' (#21) from language-md-litellm-fix into main
CI / Lint / Test / Vet (push) Successful in 12s
CI / Build & Import (push) Successful in 10s
2026-06-16 15:45:45 +00:00
mathias b246c0e688 fix(language): LiteLLM gateway never-say — stale endpoints, not piguard
CI / Lint / Test / Vet (pull_request) Successful in 12s
CI / Build & Import (pull_request) Has been skipped
piguard is legitimately in the request path as the reverse proxy; only the
stale endpoint forms (piguard:4000, koala:4000) are wrong. Per brain#6 review
amendment #3 / tapir#18 follow-on.
2026-06-16 15:35:31 +00:00
mathias c7896cb3cc Merge pull request 'LANGUAGE.md pilot: hand-written vocabulary + caveman rubric (Phase 1)' (#19) from language-md-pilot into main
CI / Lint / Test / Vet (push) Successful in 12s
CI / Build & Import (push) Successful in 10s
2026-06-16 15:22:14 +00:00
mathiasandClaude Opus 4.8 b79fb892c8 docs(claude): point to LANGUAGE.md + caveman rubric
CI / Lint / Test / Vet (pull_request) Successful in 13s
CI / Build & Import (pull_request) Has been skipped
Wires the Phase 1 vocabulary pilot into the canonical agent instructions
(tapir#18). CLAUDE.md is canonical here (no .context/ source, no generation
header).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-12 22:54:57 +02:00
mathias c315e3e003 feat(language): add hand-written vocabulary + caveman rubric pilot
Phase 1 of the ubiquitous-language system (tapir#18). Tapir-specific terms
verified against domain/ports code + README/VISION; homelab-core terms (brain
sink, LiteLLM gateway) derived from brain#5 glossary. Tooling-free pilot.
2026-06-12 20:52:23 +00:00
mathiasandClaude Opus 4.8 64e3368f5f feat(web): charm-reader visual refresh with light/dark theme toggle (ADR-032, #17)
CI / Lint / Test / Vet (push) Successful in 11s
CI / Build & Import (push) Successful in 10s
The UI read flat and boring. Reskin to one charm/TUI-inspired layout in two
palettes (CSS custom properties): a warm "reader" light theme (sketch B) and a
"cozy terminal" dark theme (sketch C).

- Palette chosen in cascade order: :root light default; an OS-preference dark
  block scoped to :root:not([data-theme]) so it applies only absent an explicit
  choice; and :root[data-theme="dark"|"light"] set by a header toggle that
  outranks the media query by specificity and persists in localStorage (guarded,
  degrades to OS default). A <head> init script applies the stored choice before
  paint, so no flash of the wrong palette.
- Charm touches via existing classes (no templ structure churn): monospace meta
  lines, accent uppercase section dividers with a trailing rule, pill buttons, a
  lifted/accent-edged expanded card.
- Error/danger shades become --err-* tokens so they follow the theme, replacing
  three per-block prefers-color-scheme dark overrides.
- Theme toggle wired into Layout and PublicLayout headers.

BDD: docs/use-cases/visual_theme.feature un-pended, mapped in scenarioCoverage.
TDD: internal/web/visual_theme_test.go (palettes, OS default, persisted toggle,
expanded-card embed). Verified light+dark on list/reader/welcome via web-shot.
Sketches kept as the design record. ui-spec.md as-built row added.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-12 12:54:06 +02:00
mathias b01f8d5783 docs(bdd): visual_theme.feature scenarios (@pending until TDD, #17) 2026-06-12 12:29:31 +02:00
mathias 8ca767c144 docs(adr): ADR-032 visual refresh — light(B)+dark(C) charm-reader themes + sketches (issue #17) 2026-06-12 11:54:54 +02:00
mathiasandClaude Opus 4.8 f98b640531 feat(web): inline-expand summary + Q&A in the list (ADR-031, #16)
CI / Lint / Test / Vet (push) Successful in 11s
CI / Build & Import (push) Successful in 11s
Click a summarized card → its full summary (summaryBody) + chat dock (chatReveal)
expand in place via HTMX (GET /v/{id}/expand → expandedCard), collapse back via
GET /v/{id}/card → compact VideoCard. Same <li id>, outerHTML swap — the existing
list-fragment pattern. The card title carries href=/v/{id} as the no-JS fallback
(detail page stays for no-JS + deep links); only summarized cards expand. Reuses
summaryBody + chatReveal so the expanded card never drifts from the detail page.

BDD: inline_expand.feature un-pended + mapped. TDD: 6 handler/fragment tests
(expand/collapse fragments, chat dock, summarized-only, no-JS href, shared body).
Minimal CSS only — the TUI/charm restyle is #17.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-12 10:47:47 +02:00
mathias 607a8cbe8d docs(bdd): inline_expand.feature scenarios (@pending until TDD, #16) 2026-06-12 10:38:40 +02:00
mathias 54d60e53b9 docs(adr): ADR-031 inline-expand summary + Q&A in the list (issue #16) 2026-06-12 10:37:55 +02:00
mathias 9298e0c686 docs(homelab): TAPIR_METRICS_ADDR + key metric series (ADR-030, #15)
CI / Lint / Test / Vet (push) Successful in 10s
CI / Build & Import (push) Successful in 10s
2026-06-12 08:50:36 +02:00
mathiasandClaude Opus 4.8 cd461b95f8 feat(observability): instrument AI + HTTP paths, serve /metrics on a side port (ADR-030, #15)
CI / Build & Import (push) Successful in 11s
CI / Lint / Test / Vet (push) Successful in 10s
Wire the metrics package into the live paths and serve it:
- summarizer: per-endpoint latency by model/outcome(success|error|parse_error)/fallback + slog.
- youtube.FetchTranscript: latency by outcome (captions|none|rate_limited) + slog.
- chat: answer latency by model + slog.
- llm usage hook → token counts (prompt|completion) per model, wired in buildSummarizer/buildChat.
- oidc callback: login counter.
- cmdServe: wrap Router in metrics.HTTPMiddleware (request count + latency by bounded
  route pattern) and serve /metrics on TAPIR_METRICS_ADDR (default :9090), a SEPARATE
  port — never on the public app mux.

BDD: observability.feature scenarios un-pended + mapped. TDD: summarizer wiring tested
black-box via the /metrics scrape; metrics-not-on-public-mux asserted.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-12 08:49:51 +02:00
mathiasandClaude Opus 4.8 9c2a04406b feat(llm): usage hook to surface token counts (ADR-030, #15)
WithUsageHook callback fires with model + prompt/completion tokens parsed from the
response usage block. Keeps the copied stdlib-only llm package decoupled from
metrics (ADR-004) — the caller wires it to internal/metrics. TDD covered.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-12 08:36:02 +02:00
mathiasandClaude Opus 4.8 a5a8cf6f6d feat(metrics): Prometheus collectors + typed API + HTTP middleware (ADR-030, #15)
New internal/metrics package: AI metrics (summarize/caption/chat latency, LLM
tokens), session metrics (requests+latency by bounded route pattern, logins), and
a /metrics Handler. Adapters call a typed API; never touch prometheus types.
Dep: github.com/prometheus/client_golang (standard Go client; cluster runs
prometheus-operator). TDD: collectors + middleware route-pattern cardinality
covered.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-12 08:32:41 +02:00
mathias 64d11af9ef docs(bdd): observability.feature scenarios (@pending until TDD, #15) 2026-06-12 08:28:46 +02:00
mathias b590d2708d docs(adr): ADR-030 observability — slog + Prometheus metrics (issue #15) 2026-06-12 08:27:56 +02:00
mathiasandClaude Opus 4.8 7314895ec4 fix(auth): stateless session cookie — stop logging users out on deploy (ADR-029)
CI / Lint / Test / Vet (push) Successful in 10s
CI / Build & Import (push) Successful in 11s
Pilot feedback: lots of re-logging-in on iPhone. Three causes: sessions lived in
an in-memory map (wiped on every pod restart/deploy), a 1h TTL (idle >1h forced
re-login on a check-back-tomorrow reader), and a session cookie with no Max-Age
(dropped on Safari close). Each re-login is the full IdP redirect dance.

Make sessions stateless: identity + absolute expiry live inside the existing
HMAC-signed cookie (no server table), TTL 1h → 30 days sliding, cookie now
persistent (Max-Age). Survives restarts (test: a cookie from one instance is
accepted by a fresh instance with the same secret), browser-close, and idle.
Trade: no server-side revocation — logout clears the cookie client-side; rotating
tapir-session-secret is the global logout lever. Accepted for the Stage-0 reader.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-11 23:07:13 +02:00
mathiasandClaude Opus 4.8 36dd182fb5 feat(report): Stage-0 usage gate counts from a baseline date (default 2026-06-11)
CI / Build & Import (push) Successful in 11s
CI / Lint / Test / Vet (push) Successful in 10s
The return-usage gate (ADR-016) counted distinct active weeks over ALL history,
so pre-launch noise — testing churn and the period the pilot sat blocked on zero
summaries — would inflate the signal. Add a baseline: ActiveWeeks(ctx, since)
filters login_events + summary_actions to seen_at/acted_at >= since. The report
command sets it to TAPIR_USAGE_GATE_START (YYYY-MM-DD, default 2026-06-11 — the
morning the pilot was unblocked) and prints the baseline. The gate now measures
whether users RETURN once it genuinely works.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-11 22:52:15 +02:00
mathias 40808f2d4b docs(onboard): BDD scenarios + architecture for the ADR-028 burst
CI / Lint / Test / Vet (push) Successful in 10s
CI / Build & Import (push) Successful in 10s
- connect_account.feature: refine the burst scenario to "best recent" (likely-good
  selection, skips too-short/too-long) and add a scenario for the stronger model.
- scenario_coverage: remap to TestOnboardBurstVideoIDs + add the model scenario
  (the old NewestUnsummarizedVideoIDs test was removed).
- architecture.md: new "Connect-time onboarding burst" subsection — junk-avoiding
  selection over persisted duration, the burst-only stronger-model chain, and why
  has-captions/cached-first are not selection signals.
2026-06-11 19:54:29 +02:00
mathias 9232f49555 feat(onboard): burst picks likely-good videos, summarizes with stronger model (ADR-028)
Wire the onboarding burst to its quality-aware selection and a stronger model:
- main.go onboard uses OnboardBurstVideoIDs (junk-avoiding) instead of pure
  newest-first, bounded by MinVideoSeconds / OnboardMaxVideoSeconds.
- buildBurstProcessor builds a burst-only summarizer chain led by the onboard
  model (burstChainModels: onboard -> standard ADR-022 chain, deduped, NDA lever
  intact), over the same store/cache/sink. Collapses onto the shared Processor
  when the onboard model is empty/equal-to-primary or config is incomplete.
- summarizerEndpoint + newYouTubeSource extracted so standard and burst wiring
  share one definition.
- Remove now-superseded NewestUnsummarizedVideoIDs: OnboardBurstVideoIDs(.,0,0)
  is identical pure-newest behaviour and its test covers RLS + ordering.
2026-06-11 19:54:29 +02:00
mathias 582c1a2065 feat(store): OnboardBurstVideoIDs — junk-avoiding burst selection (ADR-028)
Newest-first but quality-aware: excludes a video when its duration is KNOWN and
outside [minSeconds, maxSeconds], dropping Shorts and multi-hour livestream VODs
that waste a scarce caption fetch on a poor first impression. NULL/unknown
duration is kept (degrade-open) but ranked after known-good rows. 0/0 bounds
disable the filter (pure newest-first, the reversibility lever). RLS-scoped.
NewestUnsummarizedVideoIDs is left in place for callers that want pure-newest.
2026-06-11 19:54:29 +02:00
mathias 9f0d8cf198 feat(discovery): persist video duration_s instead of discarding it (ADR-028)
ADR-023's filterLowValue already fetches each candidate's duration via the
cheap videos.list quota call to drop Shorts/live, then threw it away — the
videos.duration_s column (migration 001) was never written. Carry it onto the
kept domain.Video and have UpsertVideo persist it, COALESCE-preserving a known
value so an unknown (0) re-upsert never clobbers it (the channel_title backfill
stance, migration 014). This is the enabling change for length-aware burst
selection. No new migration — the column already exists.
2026-06-11 19:54:29 +02:00
mathias d21077303d feat(config): add OnboardSummarizerModel + OnboardMaxVideoSeconds (ADR-028)
Two knobs for the onboarding-burst quality work:
- TAPIR_ONBOARD_SUMMARIZER_MODEL (default iguana/gemma4-26b): the stronger model
  the burst leads its chain with; empty collapses the burst onto the shared
  processor (lookupOr, so explicit-empty disables).
- TAPIR_ONBOARD_MAX_VIDEO_SECONDS (default 14400/4h): upper duration bound for
  burst picks; 0 disables, negative clamps to 0.

Table-driven tests cover defaults, explicit, disable, and invalid input.
2026-06-11 19:54:29 +02:00
mathias 9b2ee2e765 docs: spec onboarding-wow-burst + ADR-028 (better picks, stronger model)
Phase-1 live-DB investigation found the connect-time burst (ADR-018) fires
but delivers a weak first impression: pure newest-first selection picks junk
(a livestream + regional news for pilot user Jonte), and all burst summaries
run on the weak koala/phi4-mini instead of the validated iguana/gemma4-26b.
Cached-first was investigated and rejected — ~3% cross-user overlap, 0
cached-and-unsummarized, newest-20 all uncached (newest-first and cached-first
are structurally incompatible).

ADR-028 records: persist duration_s at discovery (ADR-023 already fetches it,
just stops discarding), junk-avoiding burst selection (drop known too-short/
too-long), and a stronger burst-only summarizer chain. Not a throughput change.
2026-06-11 19:54:29 +02:00
mathias b8fbc5a805 docs: spec first-session wow — verify & improve onboarding summary burst (investigate-first)
CI / Lint / Test / Vet (push) Successful in 11s
CI / Build & Import (push) Successful in 10s
The concern is new-user first-contact: see some GOOD summaries fast or they
don't return. Reframed as curation/latency for ~3 videos, NOT a 429/throughput
problem (3 fetches is nowhere near the wall). Phase 1 (report-and-stop) verifies
whether the existing cap-3 onboarding burst even fires today, what it delivers,
and — critically — how much transcript-cache overlap exists between users (drives
the blend). Phase 2 levers: cached-transcript-first (instant, zero-fetch),
likely-good selection (has-captions/good-length, not just newest), and optionally
the stronger model for the burst's few summaries. Blend deferred to the
maintainer post-Phase-1. Explicitly NOT bulk-fetch, NOT credentials (ADR-010/026
dead end), NOT a client extension.
2026-06-11 15:36:09 +00:00
mathiasandClaude Opus 4.8 19ca4282a8 refactor(web): dock chat inline below the summary, integrated view (ADR-027)
CI / Lint / Test / Vet (push) Successful in 10s
CI / Build & Import (push) Successful in 10s
The chat was a separate page — opening it left the summary behind, the very
context you're asking about. Now the chat docks open IN PLACE below the summary:
"Dig deeper" reveals the chat section (HTMX, outerHTML over the closed dock) so
the summary stays on screen above it; no navigation. The no-JS fallback renders
the full summary AND the open chat on one page (the same integrated view), so
progressive enhancement holds.

Extracts a shared summaryBody templ so the detail page and the chat page render
one identical summary, not two divergent ones. GET /v/{id}/chat returns just the
open chat section as a fragment for the inline reveal, or the full summary+chat
page for a no-JS navigation; POST swaps the panel inline or re-renders the whole
page. Tests assert summary+chat coexist on the page and the reveal is a fragment.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-11 14:02:39 +02:00
mathiasandClaude Opus 4.8 71df696448 feat(web): per-video chat over the stored transcript (ADR-027)
CI / Lint / Test / Vet (push) Successful in 10s
CI / Build & Import (push) Successful in 10s
A deeper-dive chat entered from the summary view: ask questions about a video
against its already-stored transcript (ADR-021), no caption fetch, ever.

Safety by construction — the load-bearing property. The chat handlers reach the
chat.Service only after reading the SHARED stored transcript via Store.GetTranscript
(a pure DB read); the service holds no VideoSource. So an enabled chat cannot
trigger a caption fetch, touch the rate gate, or reach YouTube. A video with no
stored transcript gets an honest "not available" — no fetch, no model call. The
key web test wires the summarize/fetch collaborators as tripwires that fail the
test if chat ever routes into them, and asserts the model answered from the
stored text.

Model defaults to the summary's own model and is switchable among the ADR-022
chain (phi4-mini → gemma4-26b → mistral-small); switching re-runs against the
same transcript — deliberate model-comparison instrumentation. The cloud model
is absent from the switcher when TAPIR_CLOUD_FALLBACK_MODEL="" (the local-first /
NDA lever), honoured the same way the summarizer honours it. Reuses the existing
LiteLLM gateway client (a chat is a different call, not a new integration) and
the TAPIR_MAX_TRANSCRIPT_CHARS truncation, surfacing an honest bounded-context
note when a long transcript is cut.

Ephemeral v1: the multi-turn conversation rides in hidden request fields; no
table, no migration, nothing persisted. Entry is RLS-scoped through
GetSummaryByVideo, so chat is reachable only from the user's own summary view.
Show-source verification and on-demand fetch are deferred (ADR-027).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-11 09:26:58 +02:00
mathiasandClaude Opus 4.8 cc3cda4ab8 feat(chat): stored-transcript QA service (ADR-027)
A read-only deeper-dive over a video's already-stored transcript (ADR-021).
Safe by construction: the Service has no VideoSource and no caption-fetch
dependency — only a Completer factory over the existing LiteLLM gateway — so it
cannot reach YouTube or the rate gate. Reuses the summarizer's truncation
discipline (TAPIR_MAX_TRANSCRIPT_CHARS), reporting the cut so the UI can be
honest about a bounded transcript. Model defaults to the summary's model and is
switchable among an offered, local-first list; an un-offered alias is forced
back to the default so chat can never call the gateway with an arbitrary model.
Ephemeral: history is carried per-request, nothing persisted.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-11 09:16:45 +02:00
mathiasandClaude Opus 4.8 69a49bc603 docs(decisions): add ADR-027 — chat with a video's stored transcript
Records the deeper-dive chat decision before its build: stored-transcript-only
(safe by construction — no caption fetch, no rate gate, no YouTube), entered
from the summary view, model defaulting to the summary's model and switchable
among the ADR-022 chain. Ephemeral v1; show-source verification deferred to v2.
Inserted before "Rejected alternatives" so the decision precedes the feature.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-11 09:13:31 +02:00
mathias 137804b0b1 docs: spec chat-with-transcript (ADR-027)
CI / Lint / Test / Vet (push) Successful in 11s
CI / Build & Import (push) Successful in 10s
Per-video chat against the ADR-021 stored transcript, entered from the summary
view, born from observed demand (maintainer read real summaries, some made him
want to dig deeper). HARD constraint: stored-transcript-only — never fetches
captions, never touches the rate gate or YouTube, safe by construction. Default
model = the summary's model, user-switchable among the ADR-022 chain models
(doubles as model-comparison instrumentation). Ephemeral v1 (no persisted
history); chat-only/trust-the-model with show-source verification recorded as the
natural v2. ADR-027 to be appended to DECISIONS.md as the first commit.
2026-06-11 06:38:16 +00:00
mathiasandClaude Opus 4.8 a9be5f285b feat(summarizer): move local fallback off koala to iguana/gemma4-26b
CI / Build & Import (push) Successful in 10s
CI / Lint / Test / Vet (push) Successful in 10s
koala now carries other GPU loads, so the first fallback should not run there.
Change the default chain to koala/phi4-mini → iguana/gemma4-26b → berget/mistral-small:
the local fallback now runs on iguana (M2 Ultra headroom, different host = different
egress IP for the rare fallback fetch). gemma4-26b is the brain-validated homelab
general-purpose model (agentsquad H2/H3 executor) and returned valid summary JSON on
the real prompt in a smoke test (~37s incl. cold-load — fine for a fallback path).

Pure config default (TAPIR_FALLBACK_MODEL); chain mechanism (ADR-022) unchanged.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-11 07:54:56 +02:00
mathiasandClaude Opus 4.8 fe56e2fe01 test: make embedded-postgres per-process so concurrent CI runs don't collide
CI / Lint / Test / Vet (push) Successful in 9s
CI / Build & Import (push) Successful in 10s
A push to main and its version tag fire two CI runs for the same commit. Both ran
`go test ./...`, which starts embedded-postgres on a FIXED port (54329/54330) and
a shared data dir. -p 1 serialises packages WITHIN a run, not across two
concurrent runs — so when the two runs overlapped they collided on the port/data
dir and BOTH failed the Lint/Test job (no image built). Prior commits passed only
because their two runs happened not to overlap.

Derive the port and runtime/data dirs from the PID; share only CachePath so the
PG archive downloads once. Proven: two concurrent `go test` of the store package
now both pass. Unblocks the v0.21.0 (Pillar A) build.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-11 07:26:33 +02:00
mathiasandClaude Opus 4.8 beeb5bc31b feat(gate): foreground caption fetches take priority over the background sweep (ADR-026)
CI / Lint / Test / Vet (push) Failing after 9s
CI / Build & Import (push) Has been skipped
A user waiting on a Summarize click shared the per-IP caption gate equally with
the background firehose, so on a busy IP the click was slow or 429'd. Add a
context-marked priority lane: the web path (engineProcessor.ProcessVideo) marks
its context foreground; the gate serves foreground immediately while background
fetches yield until no foreground is pending. Threaded via a context value (no
new signatures) + a process-wide foregroundPending counter. Clicks are rare, so
the background barely loses throughput; the waiting human gets the cleaner slot.

Drops the credentials probe: ADR-010 and captions.go already settle it — the
timedtext/InnerTube path rejects authenticated requests and the OAuth token does
not authenticate it anyway, so auth cannot help and can hurt. Documented in
ADR-026 rather than built.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-10 22:10:15 +02:00
mathiasandClaude Opus 4.8 09eb31d1fe fix(scheduler): derive rotation offset from wall-clock, not a reset-on-restart counter
CI / Lint / Test / Vet (push) Successful in 11s
CI / Build & Import (push) Successful in 14s
The lead-user rotation used an in-memory pass counter reset to 0 on every pod
restart, so the first-listed user always re-took the lead after a restart — a
deploy-heavy session re-starved the last user (the pilot stalled at 2 summaries
because each deploy reset his every-other-pass lead before the 2h tick fired).
Derive the offset from wall-clock (floor(now/interval)) so it advances with real
time and is identical across restarts: rotation stays fair however often the pod
bounces.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-10 21:59:12 +02:00
mathiasandClaude Opus 4.8 1665a1e7c4 feat(web): honest, state-aware summarize status with charm spinner (ADR-025)
CI / Build & Import (push) Successful in 10s
CI / Lint / Test / Vet (push) Successful in 12s
Clicking Summarize polled /status, which knew only "spinning" or "done". The web
ProcessVideo recorded an outcome only on success, so a 429'd or caption-less
click left transcript_status unset and the poll silently reverted to the
Summarize button — the rate limit was invisible and the click felt broken.

- ProcessVideo now records rate_limited / none / fetched (mirrors the runner); a
  rate-limited video keeps its requested flag so the background sweep retries it.
- /status is state-aware: summary card (done), working spinner (in-flight),
  a calm "waiting on rate limit, will retry" card that keeps polling so the
  summary lands on its own (no re-click), and a terminal "no captions" card.
- Charm status text (Claude-Code / Crush inspired): the spinner cycles playful
  tapir-themed gerunds via CSS only (no JS), decorative + an sr-only stable
  status line for a11y.

Pillar A (foreground fetch priority lane) is a separate follow-up.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-10 21:52:33 +02:00
mathiasandClaude Opus 4.8 5219561a91 feat(discovery): per-channel caption-availability memory (ADR-024)
CI / Lint / Test / Vet (push) Successful in 30s
CI / Build & Import (push) Successful in 11s
After transcript caching (ADR-021) and the Shorts filter (ADR-023), the remaining
caption waste is the first fetch on every new video of a channel that never has
English captions — each costs one rate-limited fetch to resolve to "none", and on
a throttled IP churns the backoff machinery first.

Remember, per (user, channel), a streak of consecutive no-caption outcomes
(channel_caption_state, migration 016, RLS-scoped). Once it reaches
TAPIR_CHANNEL_CAPTIONLESS_THRESHOLD (default 5) the channel is suppressed — videos
discovered/listed but not caption-fetched — for TAPIR_CHANNEL_CAPTIONLESS_WINDOW
(default 14d), then one is re-probed (auto-recovery). A successful fetch resets
the streak; a 429 does not count; an explicit manual request bypasses suppression.
threshold=0 disables.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-10 20:27:25 +02:00
mathiasandClaude Opus 4.8 9db06d8a63 feat(discovery): drop Shorts and livestreams before the caption fetch (ADR-023)
CI / Lint / Test / Vet (push) Successful in 12s
CI / Build & Import (push) Successful in 10s
The scarce resource is the per-IP timedtext caption fetch (ADR-014); the pilot's
candidate set was mostly Shorts/clips/livestreams, each burning a fetch (a "none"
result is a completed fetch — it costs budget even when it yields nothing).

NewVideos now enriches candidates with one cheap Data API videos.list call
(contentDetails.duration + snippet.liveBroadcastContent — the quota API, a
DIFFERENT limit from the timedtext 429) and drops, before returning: videos
shorter than TAPIR_MIN_VIDEO_SECONDS (default 60) and any live/upcoming
broadcast. Dropped videos are never persisted, so the list declutters too.

Degrade-open: MinVideoSeconds=0 disables it (no quota call); a videos.list error
returns candidates unfiltered so discovery never breaks on a metadata hiccup. The
paste-a-URL path (VideoByID) is not filtered — an explicit request is honoured.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-10 19:34:45 +02:00
mathiasandClaude Opus 4.8 1aa8a97f95 fix(scheduler): rotate over connected users only for true fair share
CI / Lint / Test / Vet (push) Successful in 12s
CI / Build & Import (push) Successful in 10s
The lead-user rotation rotated the full ListAllUsers set, so a connectionless
orphan identity ate a rotation slot — collapsing onto the next real user and
skewing the lead share (two real users got 2/3 vs 1/3 instead of 50/50). Filter
to connected users BEFORE rotating so the rotation is over exactly the users that
consume the caption budget. A dead identity can no longer skew fairness.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-10 18:28:54 +02:00
mathiasandClaude Opus 4.8 cc69a912f4 fix(scheduler): rotate lead user each pass so caption budget is shared
CI / Lint / Test / Vet (push) Successful in 12s
CI / Build & Import (push) Successful in 10s
Caption fetches share one per-egress-IP rate budget; whoever runs first each pass
spends the pre-throttle window before YouTube starts 429ing. ListAllUsers order
is unspecified and was stable, so the last-listed user was permanently starved —
a friendly-pilot user got 0 fetches in 12h (all rate_limited) while the
first-listed user got every successful fetch. rotateUsers left-rotates the user
order by pass index so each user leads 1/N passes and the lead slot is shared.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-10 18:07:01 +02:00
mathiasandClaude Opus 4.8 f4a0544903 fix(scheduler): cache transcripts on the scheduled path (ADR-021 regression)
buildUserRunner built the engine without engine.Transcripts = st, so the
scheduler — unlike the web "Summarize now" path — never read or wrote the shared
transcript cache. Every discovery pass re-fetched transcripts it had already
fetched, burning the scarce per-egress-IP caption budget (ADR-014) on redundant
work and starving other users' first-time fetches. The transcripts table was
empty despite summaries existing. Wire the cache on this path too.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-10 18:07:01 +02:00
mathiasandClaude Opus 4.8 e2a52789b9 feat(web): mode-aware backlog banner — stop telling manual users summaries auto-arrive
CI / Lint / Test / Vet (push) Successful in 12s
CI / Build & Import (push) Successful in 10s
The first pilot user sat in Manual mode reading "new summaries land gradually,
check back tomorrow" — copy that only makes sense in Automatic mode. Manual mode
never auto-summarizes, so the banner promised delivery that would never come.

The list page now reads the user's summarize mode and shows mode-correct copy:
- Auto: unchanged "land gradually" backlog note.
- Manual: "new videos appear here but are not summarized automatically — use the
  Summarize button" plus a "Switch to Automatic" link to /account.
The connected-but-empty first-run state is likewise mode-aware.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-10 09:14:52 +02:00
mathias e696b6405b docs(build-state): v0.15.0, summarizer is now a resilient chain (ADR-022)
CI / Lint / Test / Vet (push) Successful in 12s
CI / Build & Import (push) Successful in 10s
2026-06-10 08:59:26 +02:00
mathiasandClaude Opus 4.8 e9b5a3f3e7 feat(summarizer): resilient endpoint chain with local→cloud fallback (ADR-022)
CI / Lint / Test / Vet (push) Successful in 11s
CI / Build & Import (push) Successful in 10s
The first friendly-pilot live run produced zero summaries: koala/phi4-mini hit
three silent failure modes — 8k context overflow on long transcripts (HTTP 400),
intermittent malformed JSON (highlights as a bare string), and no fallback wired
at all (summarizer.New(primary, nil)).

Keep phi4-mini as the fast primary and add resilience around it:

- Ordered endpoint chain (summarizer.NewChain): phi4-mini → koala/phi4-14b
  (local) → berget/mistral-small (worst-case external). All reached through the
  one LiteLLM gateway by alias.
- A parse failure now advances the chain like a transport error — the old
  Primary→Fallback shape returned the parse error without trying anyone else.
- Tolerant parse: highlights/takeaways coerce string→[]string, absorbing the
  common small-model quirk without spending a fallback round-trip.
- Transcript truncation (TAPIR_MAX_TRANSCRIPT_CHARS=18000) prevents the overflow
  rather than recovering from it; validated to fit phi4-mini's 8k window.
- Bounded completion budget (TAPIR_SUMMARY_MAX_TOKENS=1500) — the old 8192 budget
  itself contributed to the overflow.

Local-first guarantee preserved by ordering: external endpoint is tried only
after every local one fails. TAPIR_CLOUD_FALLBACK_MODEL="" disables it entirely
for client/NDA deployments.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-10 08:36:16 +02:00
mathiasandClaude Opus 4.8 0ba78e8868 docs: refresh build-state for transcript persistence (ADR-021)
CI / Lint / Test / Vet (push) Successful in 11s
CI / Build & Import (push) Successful in 10s
Update the CLAUDE.md orientation block: last tag v0.14.0, migrations
001–015, and a transcript-persistence bullet (shared non-RLS store,
engine reads stored-first). The stale "v0.9.0" reference is corrected.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-09 23:39:27 +02:00
mathiasandClaude Opus 4.8 821d5f99cd docs(bdd): scenario for transcript reuse — re-analysis never re-fetches
Capture the ADR-021 promise as a mapped BDD scenario: re-analyzing a
stored video reads the stored transcript and does not fetch captions.
Since paste-a-URL and the onboarding burst summarize through the same
engine chokepoint (resolveTranscript, store-first), this one scenario
covers their reuse path too — there is exactly one gated caption entry
point (youtube.FetchTranscript → WaitFetchGate) and one engine caller in
front of it, so the dedup is structural, not per-feature.

Mapped to TestProcessNewVideo_SecondSummarizeDoesNotRefetch.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-09 23:37:54 +02:00
mathiasandClaude Opus 4.8 5c70408e75 feat(usecase): read stored transcript before fetching (ADR-021)
The engine now resolves transcripts store-first: a stored transcript —
including a stored SourceNone — is summarized without touching YouTube,
so re-analysis never re-fetches. On a miss it fetches through the source
(caption call still gated, ADR-014) and persists the terminal outcome for
the next analysis by any user. A transient SourceRateLimited is surfaced
to the runner for per-user backoff but never cached, so persistence can
never mask a 429 as a permanent "no transcript".

The TranscriptStore is optional (nil → fetch every time), keeping the
pure-core and scaffold wiring valid. cmd/tapir wires the store as both
summary sink and transcript cache, so `tapir run` and the web summarize
path (incl. paste + onboarding) all share the dedup.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-09 23:35:04 +02:00
mathiasandClaude Opus 4.8 cb6917ca59 feat(store): shared, video-keyed transcript persistence (ADR-021)
Reshape the dead per-user transcripts table (PK videos.id, user_id,
RLS-FORCEd — never read or written by app code) into the shared public
caption store ADR-021 specifies: keyed by (provider, provider_video_id),
no user_id, NOT RLS-scoped. Migration 015 (reversible). Add
ports.TranscriptStore + Store.GetTranscript/SaveTranscript via the raw
pool (no withUser): public content, shared across users by construction.
SaveTranscript persists only terminal outcomes (captions/none) and
refuses SourceRateLimited so a transient 429 can never be stored as a
false permanent absence (ADR-014).

Flip the isolation proof: transcripts leaves the RLS-scoped set;
TestTranscriptsTableIsSharedNotRLS asserts it is the SINGLE non-RLS
surface (writable/readable with no user scope, no user_id column, RLS off
on it alone, still on every user-owned table) — the proof the
public-content classification was applied exactly here and leaked nowhere.
appPool made idempotent so two tests can build it. Adjust the 010/011/014
up-down migration tests for the new HEAD. account.go: user deletion no
longer strips shared transcripts. Reconcile data-model.md + CLAUDE.md.

Wiring the engine to read-stored-first is the next commit.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-09 23:32:12 +02:00
mathiasandClaude Opus 4.8 099b2d4c68 docs(decisions): ADR-021 — shared, video-keyed transcript persistence
Persist transcripts in a single shared table keyed by
(provider, provider_video_id) — public caption content, NOT RLS-scoped —
so re-analysis (re-summarize, paste of an already-seen video, a second
user with overlapping subs) never re-fetches from YouTube. The avoided
cost is the rate-gated, reputation-risky caption fetch (ADR-010/014), not
LLM re-summarization, which is why this reopens the transcripts half of
the "no global cross-tenant table" rejection while videos stay per-user.
Summaries remain RLS-scoped (ADR-012 unchanged). The gate is neither
bypassed nor weakened — persistence reduces fetch frequency, not pacing.

Annotate the rejected-alternatives row to record the partial reopen.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-09 23:24:11 +02:00
mathiasandClaude Opus 4.8 f66c1bcdcc feat(web): real channel filter — multi-select of the user's channels
CI / Build & Import (push) Successful in 10s
CI / Lint / Test / Vet (push) Successful in 11s
The free-text 'channel' filter was dead: it exact-matched SummaryRow.Channel,
which is just the provider ('youtube'), because videos never stored their source
channel. Now they do.

- migration 014: videos.channel_title (nullable; existing rows backfill on the
  next discovery pass, pasted videos immediately).
- discovery (NewVideos) + paste (VideoByID) populate channel_title; UpsertVideo
  persists it, preserving an existing title when an update arrives empty.
- store.DistinctChannels lists a user's channels (RLS-scoped); SummaryRow carries
  ChannelTitle via the shared projection.
- Filter: single Channel -> Channels []string, matching on ChannelTitle; the feed
  renders a multi-select of DistinctChannels (hidden until channels exist).
- migrate tests: 014 reversibility + fixed the relative-step counts in the 010/011
  up/down tests (014 shifted the topology).

TDD throughout: channel persist + distinct, adapter channel wiring, multi-channel
filter match, handler channel filter, migration up/down.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-09 23:02:54 +02:00
mathiasandClaude Opus 4.8 1e65c3b413 fix(web): show paste box to any connected user, not only on an empty feed
CI / Lint / Test / Vet (push) Successful in 11s
CI / Build & Import (push) Successful in 10s
hasConnected was computed only inside the buckets.empty() branch (it was added
for the empty-state copy), so a connected user WITH videos got hasConnected=false
and never saw the paste box (#2-regression of the v0.12.0 paste UI). Compute it
on every list render. Test: connected user with a non-empty feed sees the box.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-09 22:17:00 +02:00
mathiasandClaude Opus 4.8 87c978774f docs(bdd): scenarios for paste-a-URL + onboarding burst, mapped to tests
CI / Lint / Test / Vet (push) Successful in 10s
CI / Build & Import (push) Successful in 9s
Adds paste_url.feature (valid/invalid/not-found/dedup, +@pending no-captions)
and an onboarding-burst scenario on connect; all non-pending scenarios mapped in
the coverage gate to their existing Go tests.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-09 22:06:40 +02:00
mathiasandClaude Opus 4.8 70a9f1d4cd feat(web): in-feed paste box + onboarding-aware connect confirmation
5b: connected users get a 'Summarize any video' URL input on the feed; submit
posts to /paste (HTMX) and swaps the resulting card / inline error into the feed.
7: the connect flash now sets expectations for the async onboarding burst —
'finding your subscriptions, your newest videos will appear below as summarized'.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-09 22:06:40 +02:00
mathiasandClaude Opus 4.8 1d5b2c6365 feat(serve): wire paste fetcher + connect-time onboarding burst
CI / Lint / Test / Vet (push) Successful in 11s
CI / Build & Import (push) Successful in 10s
5c: app.Fetcher = a per-user YouTube videoFetcher, so POST /paste mounts and
resolves arbitrary-video metadata (Feature 2 goes live).

6: the connect trigger now runs an onboarding burst after discovery — summarize
up to TAPIR_ONBOARD_SUMMARIZE_COUNT of the user's newest unsummarized videos via
the gated Processor (Feature 1). Hard cap; explicit so it bypasses recency; every
fetch still through globalFetchGate. No-op when count=0 or queue-only.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-09 21:58:39 +02:00
mathiasandClaude Opus 4.8 62acfee2ed feat(web): paste-a-URL handler — add an arbitrary video + summarize (Feature 2)
POST /paste: parse the video id, fetch metadata via the VideoFetcher port (Data
API, ungated), upsert a subscription-less row scoped to the user (idempotent =
dedup), and — unless already summarized — RequestSummarize + start immediate
processing through the SAME globalFetchGate as the Summarize button. Explicit
paste overrides the recency window; a captionless video degrades to the honest
'no transcript' terminal state via the engine (ADR-010). Invalid URL -> 400,
not-found -> 404, both add nothing. Route mounts only when a Fetcher is wired.

Moves the video-not-found sentinel to domain (shared by adapter + web, no
cross-adapter coupling). Tests: valid add+queue, invalid, not-found, dedup.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-09 21:54:28 +02:00
mathiasandClaude Opus 4.8 2907801aca feat(youtube): VideoByID for arbitrary-video metadata (paste-a-URL)
videos.list (part=snippet) for a single id, including channels the user does
not follow. Data API call (1 quota unit), NOT the rate-limited caption path —
ungated metadata; only the later transcript fetch hits globalFetchGate. Returns
a subscription-less domain.Video scoped to the user, or ErrVideoNotFound for a
deleted/private/typo'd id. Foundation for Feature 2.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-09 21:43:28 +02:00
mathiasandClaude Opus 4.8 59050c4db6 feat(store): NewestUnsummarizedVideoIDs for the onboarding cap
Returns up to limit of a user's newest videos (published_at DESC, NULLS LAST)
that have no summary yet. RLS-scoped via withUser — the test proves a second
user's newer video never leaks. Drives the connect-time onboarding burst
(Feature 1); the caller routes each through the shared rate gate.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-09 21:39:40 +02:00
mathiasandClaude Opus 4.8 c320ed88aa feat(config): add TAPIR_ONBOARD_SUMMARIZE_COUNT (default 3, hard cap 5)
Bounds the connect-time onboarding summary burst (Feature 1). Hard-capped at 5
and clamped (negative->0, >cap->cap) so onboarding can never bulk-fetch; 0
disables. The cap bounds COUNT only — every fetch still flows through the shared
caption rate gate (ADR-014).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-09 21:35:05 +02:00
mathiasandClaude Opus 4.8 bddd75d92e feat(web): parse YouTube video id from pasted URL forms
Pure parser for watch?v=, youtu.be/, shorts/, embed/, and bare ids; rejects
non-YouTube hosts and malformed input. Foundation for paste-a-URL summarize
(Feature 2). No fetch, no gate interaction — parsing only.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-09 21:35:05 +02:00
mathiasandClaude Opus 4.8 e4c701c6f1 feat(discovery): trigger a discovery pass on YouTube connect
CI / Lint / Test / Vet (push) Successful in 11s
CI / Build & Import (push) Successful in 10s
A newly connected account showed no videos until the next 2h scheduled pass —
the gap that made onboarding look broken (a second user connected, saw nothing,
read as failure). The connect callback now fires an out-of-band discovery pass
for the connecting user, so videos appear promptly.

Concurrency: scheduled and connect-triggered passes share one lock (serialize),
preserving the single-fetcher invariant (ADR-018). A trigger interleaves between
the scheduler's per-user passes rather than fetching concurrently or waiting for
a whole pass. The trigger runs on the server ctx (survives the redirect) and is
non-blocking for the request goroutine.

Scope: connect-trigger only. The optional login-refresh / "Discover now" button
from #6 are intentionally not built — an unconditional login hook risks 429
storms (per the ticket's own recommendation); defer until wanted.

TDD: TestCallbackTriggersDiscovery, TestSerializeRunsOneAtATime,
TestDiscoveryTriggerEnqueueRunsUser; new BDD scenario mapped.

Refs #6

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-09 21:15:46 +02:00
mathias b527db9739 docs: bump last-tag reference to v0.9.0
CI / Lint / Test / Vet (push) Successful in 15s
CI / Build & Import (push) Successful in 11s
2026-06-09 21:03:44 +02:00
mathiasandClaude Opus 4.8 c5f556d1d6 fix(scheduler): skip discovery for users with no video connection
CI / Lint / Test / Vet (push) Successful in 11s
CI / Build & Import (push) Successful in 10s
The scheduler enumerates every user_identities row (ListAllUsers) and ran a
discovery pass for each — including users who never connected a video source.
Their per-user runner then tried to resolve a YouTube refresh token that was
never minted, logging a spurious "secrets: ref not found:
youtube/<uid>/refresh_token" every tick (e.g. stale Dex-era orphan identities
left by the Authentik migration).

Skip users whose ConnectionsForUser is empty before running their pass. Removes
the recurring noise — which actively misled a debug session into thinking a
healthy onboarded user was broken — with no change to connected users.

Refs #7

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-09 20:49:11 +02:00
mathiasandClaude Opus 4.8 a884e7e9c5 test(bdd): add scenario name-coverage gate (no godog)
CI / Lint / Test / Vet (push) Successful in 12s
CI / Build & Import (push) Successful in 10s
Close the gap where docs/use-cases/*.feature claimed to be the behavior spec
but nothing executed them — so scenarios drifted (the stale "auto summarizes
every new video" and "manual is the default" were proof).

Decision (per issue #5 BDD-runner fork): no godog — keep .feature as design
records, add a cheap name-coverage gate instead. TestScenarioCoverage parses
every scenario and asserts each non-@pending one maps to an existing Go test in
the scenarioCoverage manifest; it flags unmapped scenarios, missing/renamed
tests, and stale entries. It checks the link, not that the test exercises the
scenario (the deliberate trade for skipping godog).

Also:
- Fix the stale ADR-018 drift: "Manual is the default" -> auto is the default
  for new users; added an explicit default scenario + a plain manual scenario.
- Tag 4 documented-but-unbuilt/untested scenarios @pending with reasons (Vimeo
  connect, BYO config flow, logout->welcome, re-register-after-delete) so they
  are tracked without a false coverage claim.
- CLAUDE.md BDD section now describes the real setup (design records + the gate
  + @pending convention) instead of claiming an executable spec.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-08 22:59:25 +02:00
mathiasandClaude Opus 4.8 27fd33c99c docs: reconcile requirements + architecture with ADR-020
CI / Lint / Test / Vet (push) Successful in 11s
CI / Build & Import (push) Successful in 10s
Bring the living docs current with the recency-bounded auto-summarize + sparse
honesty + feed IA bundle (ADR-020):

- requirements (BDD): summarize_mode.feature — auto now summarizes RECENT new
  videos; added a scenario for older videos (listed, on-demand), recency note.
- architecture.md: summarization-mode + new list-surface paragraph; scheduler
  diagram + two-path table + three-phase pass now show the recency pre-filter;
  dropped stale "Summarize now".
- data-model.md: auto_summarize is recent-only, older on-demand.
- README.md: one-line recency note on the serve scheduler.
- ui-spec.md: appended the as-built ADR-020 row (supersedes earlier copy/sort).
- specs/{video-card-states,newest-first-ordering,scheduled-discovery}.md:
  superseded/extended banners pointing at ADR-020 (kept as design records).

Docs-only; task check green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-08 20:21:14 +02:00
mathiasandClaude Opus 4.8 2c96926ff7 docs: correct stale last-tag reference (v0.4.0 -> v0.8.0)
CI / Lint / Test / Vet (push) Successful in 12s
CI / Build & Import (push) Successful in 10s
CLAUDE.md "Current build state" still cited v0.4.0; the repo is at v0.8.0.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-08 15:50:06 +02:00
mathiasandClaude Opus 4.8 2cda62b3ad docs(adr): record ADR-020 recency-bounded auto-summarize + sparse honesty
CI / Lint / Test / Vet (push) Successful in 12s
CI / Build & Import (push) Successful in 10s
Document the architecture decision behind this bundle: bound auto-summarize to
a recency window (refines ADR-018; bounds load against the ADR-014 gate without
fetching harder), surface scarcity honestly, and collapse the un-summarized
back-catalogue in a single feed. Records the return-nudge as a deliberate
non-goal (it would contaminate the Stage-0 unprompted-return signal, ADR-016).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-08 14:04:37 +02:00
mathiasandClaude Opus 4.8 f29927f50d test(web): guard against the removed over-promise card copy
Extend the card copy guard so "Try now", "Summarize now", "Fetching soon", and
"the next run" can't silently return to any card state (UX review honesty pass).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-08 14:03:08 +02:00
mathiasandClaude Opus 4.8 51aa5d940c feat(web): segment watched/skipped action toggles
Watched and skipped are mutually exclusive (the store clears one when the other
is set), but rendered as three independent-looking buttons the exclusivity was
invisible. Group watched|skipped into a single segmented control and keep Saved
apart as an independent toggle (UX review C5). HTMX posting and the active/✓/
aria-pressed semantics are unchanged; extracted a shared actionButton component.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-08 14:02:31 +02:00
mathiasandClaude Opus 4.8 f775441a62 feat(web): drop the empty terms checkbox from registration
The register step asked the user to accept "the terms of use" with no terms
linked anywhere — ceremony accepting nothing on a friends-only tool (UX review
C4). Remove the checkbox and the server-side acceptance requirement; only a
display name is required now.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-08 14:01:28 +02:00
mathiasandClaude Opus 4.8 12fb031b6c feat(web): add a back link to the detail page
The summary detail page only returned to the list via the brand logo. Add an
explicit "← Summaries" link at the top (UX review C3).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-08 14:00:17 +02:00
mathiasandClaude Opus 4.8 980638d80a feat(web): slim the list filters and hide them when empty
At current scale the date-range pickers are dead weight (UX review C1/C2):

- Drop the From/To date inputs from the filter bar; keep the Channel field and
  the "Summarized only" toggle. (Filter still parses from/to for hand-built
  URLs and apply() compatibility — only the UI is removed.)
- Hide the filter bar entirely on a genuinely empty account (no rows AND no
  active filter) so the connect CTA stands alone; a filter that matches nothing
  still shows the bar so it can be cleared.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-08 13:59:44 +02:00
mathiasandClaude Opus 4.8 40b703e02a feat(web): collapse older + caption-less videos in the list
Stop the un-summarized back-catalogue from burying the readable summaries
(UX review B3/B4). One feed, with a noise-collapse — not sections:

- Summarized + recent un-summarized videos lead inline as cards.
- Un-summarized videos older than the recency window collapse into a single
  "Show N older videos — summarize on demand" disclosure (they will not
  auto-fill; they are manual-only). Window comes from App.RecencyWindow
  (= cfg.AutoSummarizeWindow); 0 disables the collapse (all inline).
- Caption-less videos collapse into one honest line ("N videos have no
  captions and can't be summarized") instead of N dead terminal cards.

bucketRows is a pure classifier (cutoff-driven; undated rows never age out);
App gains RecencyWindow + an injectable clock for the cutoff.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-08 13:55:07 +02:00
mathiasandClaude Opus 4.8 3df0459fed feat(store): order video list by published_at, NULLS LAST
ListVideos now sorts summarized-first, then published_at DESC with undated
videos last, then seen_at DESC as a tiebreak (was seen_at only). Aligns the
list with the recency framing — newest content surfaces first — so the
recency-bounded feed reads coherently (UX review B2).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-08 13:50:30 +02:00
mathiasandClaude Opus 4.8 2384c47b81 feat(runner): bound auto-summarize to a recency window
In automatic mode the scheduler now only summarizes videos published within
TAPIR_AUTO_SUMMARIZE_WINDOW (default ~7d). Older videos are still discovered
and listed — they keep the manual "Summarize" affordance — but are not
auto-processed, so a large back-catalogue (the maintainer's ~256-deep queue)
stops self-inflicting 429s against the per-IP caption gate each cycle (UX
review B1, recency design).

- runner.WithAutoWindow + Stats.SkippedTooOld; tooOld() treats a zero window
  as disabled and an undated video as never-aged-out (processed, not stranded).
- An explicit manual request bypasses the bound even in auto mode (requested
  videos are loaded in auto mode when a window is active).
- Wired through cmdRun, the scheduler's per-user runner, sumStats, and pass
  logging. config: TAPIR_AUTO_SUMMARIZE_WINDOW (default 168h), .env.example.
- Account copy (A7) updated to match: "Automatic summarizes new videos from
  about the last week; older videos stay browsable — summarize on demand."

The rate gate is untouched; the manual path still serialises through it. This
bounds auto LOAD, it does not fetch harder.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-08 13:48:04 +02:00
mathiasandClaude Opus 4.8 4a0a56e152 feat(web): lead summary detail with takeaways
The product promise is "decide what's worth your time", but the detail page
buried Takeaways — the verdict that answers that — below the full Summary.
Reorder to Takeaways → Highlights → Summary so the attention-saving payload
leads (UX review A8). Data already existed; this is a section reorder only.
Conditional sections mean a video without takeaways/highlights still leads
with the Summary naturally.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-08 13:44:20 +02:00
mathiasandClaude Opus 4.8 9bf1c31605 fix(web): correct stale welcome-page invite copy
The landing page promised "if you have an invite link, it will set up your
account automatically" — but invites moved to Authentik (ADR-019); Tapir no
longer handles invite links and "Get Started" goes straight to OIDC. Replace
with honest "invite-only — if you've been invited, sign in" and set the
gradual-fill expectation before the login wall (UX review A5). Test pins the
stale phrase out.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-08 13:43:32 +02:00
mathiasandClaude Opus 4.8 a1a5217d77 feat(web): honest sparse-state and queue copy
Make the sparse reality legible instead of implying abundance or imminence
(UX review A1-A4, A6):

- Empty-connected state drops the impossible "Run `tapir run`" instruction
  (no shell for web users; discovery is in-process since ADR-018) for a
  passive "summaries appear gradually, check back later".
- Pipeline bar reframes counts by what the user can do: "N ready · M in queue
  · K no captions" (was "summarized / fetching soon / pending").
- A one-line note explains captions are fetched slowly on purpose to respect
  YouTube's limits — turning confusing emptiness into intentional design.
- Card state for throttled videos reads "In queue", not "Fetching soon…"
  (256 items behind a per-IP gate are not all imminent — ADR-014).
- Quiet nudge button drops the over-promising "now": "Summarize", not
  "Summarize now". On click the card still honestly becomes "Queued".
- Queued card says "summarizing shortly", not "waiting for the next run"
  (no scheduler jargon).

Pure copy/label — no logic, DB, or fetch-rate change. The rate gate is
untouched; scarcity is surfaced, never engineered around.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-08 13:42:40 +02:00
mathiasandClaude Opus 4.8 9c7e3be984 docs(ux): add Stage-0 recency-bounded heuristic review
Prioritized UX findings for the product as it actually is — sparse feed,
respected caption rate limit, recency-bounded auto-summarize (incoming),
single-user. 15 findings, NOW/LATER tagged. P0s target the first-contact
return-cliff that the Stage-0 gate depends on.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-08 13:11:41 +02:00
mathiasandClaude Opus 4.8 8e45f21d23 docs: ADR-019 (Authentik owns invites), supersede ADR-017
CI / Lint / Test / Vet (push) Successful in 13s
CI / Build & Import (push) Successful in 10s
Record the invite-provisioning removal; mark ADR-017 superseded; fix
ui-spec invite-onboarding + auth-delegation sections to reflect Authentik.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-07 23:01:14 +02:00
mathiasandClaude Opus 4.8 e7c2e575d3 refactor: remove Dex local-password invite provisioning (ADR-019)
Authentik owns invites now (infra ADR-0001). Delete adapters/dex, the
/invite set-password UI, the tapir invite CLI, the InvitationStore/
DexPasswordCreator ports + App wiring, the invite Templ pages, and the
invite Taskfile target. New users are invited via Authentik, log in via
OIDC, and hit the existing /register gate. invitations table (mig 009)
left in place (append-only; harmless). task check green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-07 23:01:14 +02:00
mathias 9cd3f7e934 fix(dex): passwordName must match Dex's internal passwordID() — maps non-[a-z0-9-] to '-'
CI / Lint / Test / Vet (push) Failing after 12s
CI / Build & Import (push) Has been skipped
Tapir used human-readable substitutions ('@' -> '-at-', '.' -> '-dot-') when
deriving the Password CR name from an email. Dex's internal passwordID() maps
every non-[a-z0-9-] character to plain '-'. This caused a name mismatch:
Tapir wrote the CR as 'mathias-at-d-ma-dot-be', Dex looked it up as
'mathias-d-ma-be', got not-found, and returned 'Invalid credentials' on every
invite login — while static configmap passwords (a different code path) worked
fine. Diagnosed by adding the email to staticPasswords and confirming login
succeeded, proving the kubernetes CR lookup was the failure point.
2026-06-07 11:37:24 +02:00
mathias c812c71ecc fix(dex): store raw bcrypt hash in Password CR, not base64-encoded
CI / Lint / Test / Vet (push) Successful in 26s
CI / Build & Import (push) Successful in 12s
The original NOTE claimed Dex's kubernetes storage types Hash as []byte,
requiring the bcrypt string to be base64-encoded before storage. This was
wrong: Dex v2.41 stores and compares the hash field as a plain string. The
base64-encoding caused every invite login to fail with 'Invalid credentials'
because Dex passed the base64 bytes (starting with 'J' not '$') directly to
bcrypt. Static passwords in the configmap always used raw bcrypt strings and
worked fine — confirming the dynamic CR encoding was the bug.
2026-06-07 09:28:11 +02:00
mathias 084b73907d feat(task): add invite task — task invite EMAIL=user@example.com
CI / Lint / Test / Vet (push) Successful in 12s
CI / Build & Import (push) Successful in 10s
2026-06-07 07:35:47 +02:00
mathias 0b04e487ad docs: update card-state model — 'Summarize now' verb, no-captions state, five-state table
CI / Lint / Test / Vet (push) Successful in 12s
CI / Build & Import (push) Successful in 10s
2026-06-06 22:42:37 +02:00
mathias 2fedc45443 feat(web): unify card states — one 'Summarize now' verb, honest no-captions state
CI / Lint / Test / Vet (push) Successful in 13s
CI / Build & Import (push) Successful in 10s
Five explicit footer states, status-primary:
1. Summarized — chip + actions, no button (unchanged)
2. No captions (TranscriptStatus=="none") — NEW: 'No transcript available' muted
   text, no button, no POST URL. Removes the dead-end 'Summarize' button that
   tried and failed when there were no captions to fetch.
3. Queued (SummarizeRequested) — chip + muted text, no button (unchanged)
4. Rate-limited — 'Fetching soon…' + quiet 'Summarize now' → /retry-now
5. Pending — 'Not summarized' + quiet 'Summarize now' → /summarize

One verb ('Summarize now'), one quiet style (.btn-quiet, renamed from .btn-retry
which was state-specific). User doesn't see the internal pipeline distinction;
both buttons post to their existing handlers unchanged. Form class renamed
card-nudge-form. Dropped engineer-facing tooltip; user-facing hint added.
'Try now' wording removed entirely.
2026-06-06 22:31:41 +02:00
mathias 58cd68c1eb docs: spec video-card state unification — one "Summarize now" verb + honest no-captions state
CI / Lint / Test / Vet (push) Successful in 12s
CI / Build & Import (push) Successful in 10s
The card shows two verbs ("Try now" for rate-limited, "Summarize" for pending)
for one user intent, and — the real bug — a no-captions video falls into the
pending branch and wrongly shows a Summarize button that can only fail. Spec
collapses to one quiet "Summarize now" verb wherever a nudge is possible (both
handlers unchanged underneath), adds an honest no-button "No transcript
available" state, and keeps the card status-first (buttons are exceptions in auto
mode). View-layer only.
2026-06-06 19:58:45 +00:00
mathias e472015c76 docs: replace evasion framing with honest onboarding-prioritisation rationale
CI / Lint / Test / Vet (push) Successful in 12s
CI / Build & Import (push) Successful in 10s
'Try now' and the newest-first batch implement onboarding prioritisation:
foreground (user-clicked 'Try now') summarises a chosen video on demand;
background batch summarises newest-first; both honour the shared rate gate.

Remove any prior framing that described 'Try now' as making traffic 'look
organic to YouTube' or as rate-limit evasion — that was not the rationale
and contradicts ADR-014's explicit account-safety constraint.

Correct statement: rate limiting is respected, not evaded. TAPIR_FETCH_RATE
and TAPIR_FETCH_BACKOFF are honest rate controls; they govern how fast Tapir
fetches captions, not how the requests appear to YouTube.

Architecture: add two-path model table (foreground/background, both through
globalFetchGate) and newest-first batch ordering doc (three-phase RunOnce,
before/after example).

ui-spec: add 'Try now' row with correct rationale; add pipeline stats bar row;
update Summarized-only filter row to mention sort-to-top.
2026-06-06 21:29:28 +02:00
mathias 0c0225f9c6 feat(runner): process candidates newest-first within each pass (ADR-018)
Restructures RunOnce from per-channel inline processing to collect-sort-process:

Phase 1 — discover, persist (UpsertVideo), apply pre-filters (seen/manual/backoff)
           and collect surviving candidates with their discovery position.
Phase 2 — sort candidates by published_at DESC, NULLS LAST, pos ASC tiebreak
           so videos with no publish date never jump ahead of dated content.
Phase 3 — process in sorted order through the unchanged globalFetchGate.

Before (per-channel): chanA=[v-old, v-mid], chanB=[v-new, v-null]
                    → [v-old, v-mid, v-new, v-null]
After  (newest-first): [v-new, v-mid, v-old, v-null]

Same set of videos processed; only the order changes within a pass. All existing
behaviour is preserved: failure isolation, backoff skip, manual mode,
channel-unavailable, stats. In-memory sort; no new table or persisted queue.

The ordering is onboarding prioritisation — new users get summaries of their most
recent, relevant videos first; the back-catalogue fills in behind across subsequent
passes. Both this background batch and the foreground 'Try now' button honour the
shared globalFetchGate: rate limiting is respected, not evaded.
2026-06-06 21:29:28 +02:00
mathias 8403e8e524 docs: spec newest-first batch ordering + honest Try-now/prioritisation docs
CI / Lint / Test / Vet (push) Successful in 12s
CI / Build & Import (push) Successful in 11s
Product intent: new users get summaries of their newest videos fast while the
back-catalogue fills behind, within the shared rate gate. The background batch
currently processes in subscription/channel order, not newest-first — this spec
closes that gap (collect candidates, sort published_at DESC NULLS LAST, process
in order, gate unchanged). Also corrects the docs to describe Try-now as
onboarding prioritisation, explicitly removing the prior "looks organic to
YouTube" traffic-disguising framing — rate limiting is respected, not evaded.
2026-06-06 19:20:27 +00:00
mathias 57e29ca06c feat(web): pipeline stats bar, summarized-first sort, Try now button for rate-limited videos
CI / Lint / Test / Vet (push) Successful in 15s
CI / Build & Import (push) Successful in 10s
Three UX improvements for the pending-transcript state:
1. Summarized videos sort to top (ORDER BY (s.id IS NOT NULL) DESC, seen_at DESC)
   so completed summaries are always immediately visible without filtering.
   ListVideos default limit raised from 50 to 500 to show the full backlog.
2. Pipeline stats bar above the video list: '2 summarized · 256 fetching soon · 12
   no captions' — computed from the unfiltered row set, hidden when everything is
   summarized.
3. 'Try now' button on rate-limited cards replaces the passive 'Retrying later'
   chip. POST /v/{id}/retry-now clears rate_limited_at then calls ProcessVideo
   through the shared globalFetchGate — same rate limiting as the scheduler, safe
   under concurrent use.
2026-06-06 19:20:20 +02:00
mathias 24f2a69eaa fix(web): gofmt view.go
CI / Lint / Test / Vet (push) Successful in 12s
CI / Build & Import (push) Successful in 9s
2026-06-06 11:31:28 +02:00
mathias 241eebfd9f docs(homelab): add TAPIR_FETCH_BACKOFF config, update snapshot date 2026-06-06 11:31:11 +02:00
mathias 317b0d4834 docs(ui-spec): add Dex auth details, invite onboarding, summarized-only filter, unavailable channels 2026-06-06 11:29:14 +02:00
mathias 252a4ebd9e docs(architecture): add in-process scheduler sequence, rate gate description, auto_summarize default fix 2026-06-06 11:27:18 +02:00
mathias 1ad1966672 docs(data-model): add CHANNEL_ERRORS + LOGIN_EVENTS entities, transcript_status columns, auto_summarize default update 2026-06-06 11:25:52 +02:00
mathias ccadcecfef docs(readme): add tapir serve, fix TAPIR_DISCOVERY_INTERVAL reference 2026-06-06 11:24:23 +02:00
mathias 4c5d3cca81 feat(web): add 'Summarized only' filter checkbox to video list
CI / Lint / Test / Vet (push) Failing after 3s
CI / Build & Import (push) Has been skipped
2026-06-06 11:19:18 +02:00
mathias 10233ee881 fix(scheduler): include ChannelUnavailable in sumStats + pass-complete log
CI / Lint / Test / Vet (push) Successful in 12s
CI / Build & Import (push) Successful in 9s
2026-06-06 10:22:37 +02:00
mathias f1e9739900 feat(store,runner,web): channel unavailability notice (migration 013)
CI / Lint / Test / Vet (push) Successful in 26s
CI / Build & Import (push) Successful in 11s
YouTube channels that 404 on playlist discovery (deleted/private) are now:
1. Wrapped in domain.ErrChannelUnavailable by the YouTube adapter (instead of
   a generic error), so the runner can identify them without string-matching.
2. Stored per-user in channel_errors (migration 013, RLS-guarded) via runner's
   new UpsertChannelError path — removed from the generic Errors counter,
   counted separately as ChannelUnavailable.
3. Shown on the account page under "Unavailable channels" with name, chip-warn
   badge, and first-seen date, so users know why some subscribed channels
   produce no videos.
2026-06-06 10:09:52 +02:00
mathias 940f80899a fix(store): migration 012 — back-fill auto_summarize via RLS bypass
CI / Lint / Test / Vet (push) Successful in 12s
CI / Build & Import (push) Successful in 11s
Migration 011's UPDATE ran without tapir.current_user_id set, so FORCE RLS
blocked all rows and 0 users were updated (skipped_manual=607 in scheduler).
Migration 012 temporarily drops FORCE so the table owner can run the UPDATE,
then restores it.
2026-06-06 10:01:13 +02:00
mathiasandClaude Opus 4.8 f35c2a85a5 docs: scheduled-discovery env + single-replica constraint; VISION gate-clock reset
CI / Lint / Test / Vet (push) Successful in 13s
CI / Build & Import (push) Successful in 11s
homelab-integration.md gains a "Scheduled discovery" section documenting
TAPIR_DISCOVERY_INTERVAL and TAPIR_FETCH_RATE and the load-bearing
single-replica constraint (in-process scheduler → replicas: 1 is required;
>1 double-runs discovery). VISION Stage 0 carries a pointer to ADR-018's
gate-clock reset so nothing in docs implies the window started before
unprompted use was possible.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-05 23:43:17 +02:00
mathiasandClaude Opus 4.8 5d029a2823 feat(store): default auto_summarize ON for new users (ADR-018)
Migration 011 flips the auto_summarize column default to TRUE and brings
existing rows (maintainer + current registrations) along. Onboarded friends
now get zero-friction discovery: scheduled discovery (ADR-018) both discovers
AND summarizes new videos, so a user's list fills and summarizes itself
instead of presenting an empty list of manual Summarize buttons.

Safe only because the process-wide caption-fetch rate gate (ADR-014 item 2,
prior commit) now exists — auto + scheduled + multi-user would otherwise
self-inflict 429s every cycle. The down migration reverts the default but
intentionally leaves existing rows as-is (no surprise manual regression on
rollback). RegisterUser already lets the column default drive the value, so
no app change is needed; the account-page manual toggle still works.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-05 23:42:23 +02:00
mathiasandClaude Opus 4.8 f6623afd41 feat(serve): in-process scheduled discovery for all users (ADR-018)
Stage 0's "returns and reads in >=2 weeks" gate can't be met while discovery
is host-side manual (`tapir run`): a newly onboarded user sees an empty list
and never comes back. Make Tapir watch on its own.

cmdServe launches a background goroutine (when TAPIR_DISCOVERY_INTERVAL > 0)
that runs a discovery pass for ALL users on that cadence: enumerate via the
un-RLS'd ListAllUsers, then run each user's pass through the EXISTING
runner.Runner — the only new code is the per-user loop, not a new scheduler.
Run-once-on-startup then ticked; ctx-cancelled on SIGTERM; per-user failures
(including buildUserRunner errors) are logged and skipped so one bad user
never aborts the rest. interval <= 0 disables it entirely (dev/tests).

buildUserRunner binds each runner to that user's own YouTube refresh token
(web.YouTubeTokenRef) — the Stage-1 per-tenant ref — reusing buildProcessor's
engine wiring. SetFetchRate is also wired in cmdServe so the click-path shares
the gate.

SINGLE-REPLICA is now load-bearing: the loop lives in the web process, so >1
replica double-runs discovery (429s + duplicate work). Documented in cmdServe
and warned at startup.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-05 23:39:39 +02:00
mathiasandClaude Opus 4.8 149ec2adae feat(store): ListAllUsers for scheduler user enumeration
The in-process scheduler (ADR-018) needs to enumerate every user to run a
discovery pass each. user_identities is the un-RLS'd map; add ListAllUsers as
a plain pool query (no withUser) — the same enumerate-then-act pattern
UserBySubject and the login_events gate query established. Scoping it to a
single user would defeat the point; user_identities carries no RLS by design.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-05 23:36:44 +02:00
mathiasandClaude Opus 4.8 f5021a8436 feat(youtube): process-wide caption-fetch rate gate (ADR-014 item 2)
ADR-014 item 2 — a single per-egress-IP rate gate shared by every caption
fetch — was specced but only per-video backoff (rate_limited_at) shipped.
Build the real gate now: it is load-bearing once ADR-018 puts auto-summarize
on an in-process schedule across multiple users (all fetches leave one pod's
egress IP, concurrently with live "Summarize" clicks — without a shared gate
that self-inflicts 429s every cycle).

globalFetchGate (golang.org/x/time/rate, default 2s/req burst 1) is consulted
in httpDo before every live outbound fetch — player, watch-page, timedtext —
so the scheduler runners and the web click-path serialise through one limiter
regardless of how many users/goroutines are upstream. The test seam
(a.transport != nil) skips the gate so fakes are not throttled.

TAPIR_FETCH_RATE (Go duration, default 2s, 0 = unlimited) wires SetFetchRate in
cmdRun; the existing per-video backoff stays as the complementary 429 handler.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-05 23:36:17 +02:00
mathias 88294d38bc docs: add ADR-018 — in-process scheduled discovery, auto-summarize, gate-clock reset
CI / Lint / Test / Vet (push) Successful in 12s
CI / Build & Import (push) Successful in 10s
Records the Stage-0 usability decision: tapir serve runs discovery for all users
on an interval (reusing runner.Loop), auto-summarize defaults ON so the list fills
itself, and the gate clock resets to when this ships (unprompted use was
impossible before, so the prior window measured nothing — framed as starting the
clock when the experiment can run, not dodging a failing gate). Records the
in-process-vs-CronJob tradeoff, the load-bearing single-replica constraint, and
that it makes ADR-014 item 2 (process-wide rate gate) a hard requirement folded
into the build. Cross-referenced ADR-014/016; added CronJob to rejected-alts.
2026-06-05 13:03:18 +00:00
mathias ca9bf2657f docs: spec in-process scheduled discovery + auto-summarize + rate-gate finish
CI / Lint / Test / Vet (push) Successful in 12s
CI / Build & Import (push) Successful in 13s
The Stage-0 usability fix: tapir serve runs discovery for all users on an
interval (reusing the existing Runner.Loop, enumerate-users-then-withUser),
auto-summarize defaults ON for Future-B users so the list fills itself, and the
ADR-014 process-wide per-egress-IP rate gate is confirmed/finished in the same
slice because in-process + auto + multi-user makes it load-bearing. Records the
single-replica constraint as load-bearing, and a fallback (auto-summarize OFF
until the gate exists) so the dangerous combination never ships half-built.
2026-06-05 12:59:50 +00:00
mathias 4706c508e9 docs: ratify ADR-017 KEEP — invite flow is load-bearing (some users can't use Google OIDC)
CI / Lint / Test / Vet (push) Successful in 12s
CI / Build & Import (push) Successful in 10s
Records the deliberate keep-or-reverse decision on the Dex-write invite flow.
Decision: KEEP. Deciding fact: not all intended Future-B users will use Google
accounts, so Google OIDC alone can't onboard them — the invite flow is
load-bearing, not redundant. Trust-surface cost accepted deliberately, explicitly
NOT as a precedent for widening further, and explicitly NOT by adding delete RBAC
to fix the orphan gap. Open items reframed as tracked follow-ups (verify RBAC
against the real manifest; accept orphan for Future B, revisit before Future C).
Added the reversal to rejected-alternatives.
2026-06-05 12:47:39 +00:00
mathias b4da97e0e2 docs: resolve ADR-014 migration-008 false alarm (cols shipped in 007)
CI / Lint / Test / Vet (push) Successful in 13s
CI / Build & Import (push) Successful in 10s
Reconciliation finding: the v0.6.0 report's "migration 008 videos.rate_limited_at"
was a mislabel. Both transcript_status and rate_limited_at shipped in migration
007; the sequence legitimately skips 008, nothing was lost, and the runner's
column reads are sound. Updated the ADR-014 implementation note to state this
(was flagged as an unresolved discrepancy). The item-2 (shared per-egress-IP rate
gate vs per-video backoff) flag stays open — still unconfirmed.
2026-06-04 18:54:56 +00:00
mathiasandClaude Opus 4.8 561ba79360 feat(cli): tapir report — Stage-0 usage gate query
CI / Lint / Test / Vet (push) Successful in 12s
CI / Build & Import (push) Successful in 10s
ActiveWeeks computes per-user distinct active weeks (reads login_events UNION
acts summary_actions) for the gate (VISION/ADR-016: usage in >=2 distinct
weeks). Because the user-owned tables are FORCE RLS under a non-superuser owner,
a single cross-user query is deny-all; instead it enumerates users from the
un-RLS'd identity map and counts each inside withUser — no privilege escalation,
no policy change.

`tapir report` prints the per-user table and the pass/fail verdict (needs only
TAPIR_DB_DSN). Pure formatter + store query are unit-tested, including the
cross-table shared-week dedup.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-03 23:46:03 +02:00
mathiasandClaude Opus 4.8 2fac735837 feat(web): stamp login_events on every gated request
The registration gate, once it resolves the authenticated subject to a tapir
user_id, calls StampLogin (store-throttled to one row per user per day). Best-
effort: a stamp failure is logged and swallowed so it never breaks the request.
This is what makes the read-side Stage-0 usage signal actually accrue.

Tests cover the happy-path stamp, the same-day throttle, and that an
unregistered subject (redirected to /register) is never stamped.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-03 23:46:03 +02:00
mathiasandClaude Opus 4.8 de54cd33b2 feat(store): throttled per-day StampLogin + login_events delete cascade
StampLogin appends one login_events row per user per day via an atomic
INSERT ... SELECT ... WHERE NOT EXISTS, run through withUser so the throttle
probe is itself RLS-scoped to the caller. DeleteUser now deletes login_events
explicitly (no FK = no cascade — the summary_actions footgun, repeated).

Extends the two-user RLS isolation proof and the delete-account proof to cover
login_events, and adds throttle / new-day / user-scoping tests.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-03 23:46:03 +02:00
mathiasandClaude Opus 4.8 b070347597 feat(store): add append-only login_events table with forced RLS
Stage-0 usage measurement (VISION/ADR-016): summary_actions captures acts
(watch/skip/save) but not reads. A reader who logs in weekly and clicks
nothing is invisible — for a reading product that return is the signal the
gate ("usage in >=2 distinct weeks") is defined on. login_events records
THAT a user was active, append-only, one row per user per active day.

Per-user isolation via the same GUC-keyed FORCE RLS policy as migration 003.
No FK to users (mirrors summary_actions) — the cascade footgun is handled by
DeleteUser in a later commit. Adds an up/down reversibility test and registers
the table in both truncate helpers.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-03 23:46:03 +02:00
mathias 99743af182 docs: add ADR-017 (Dex-write invite flow) retroactively; annotate ADR-002/013/014
CI / Lint / Test / Vet (push) Successful in 12s
CI / Build & Import (push) Successful in 10s
Reconciliation pass after parallel agent sessions shipped v0.6.0/v0.7.0.

ADR-017 documents the v0.7.0 invite flow, which gave Tapir scoped create+get on
passwords.dex.coreos.com in the auth namespace — Tapir now WRITES to the shared
identity provider. This shipped with no ADR; recorded retroactively with the
principle-reversal named (partially supersedes ADR-002/013), the security analysis
(bounded RBAC, but a real larger trust surface), and the open gaps (orphaned Dex
accounts on delete; plaintext invite tokens). Cross-referenced in ADR-002 and
ADR-013 status lines so a future reader isn't misled.

ADR-014 annotated: 429 handling shipped in v0.6.0 but decision item 2 (shared
per-egress-IP rate gate) appears realised as per-VIDEO backoff, not a process-wide
IP gate — flagged not-confirmed-done. Also flags the missing migration 008 /
rate_limited_at discrepancy (v0.6.0 report cited 008; tree jumps 007->009).

No code changed in this commit — audit trail only.
2026-06-03 21:38:49 +00:00
mathias 070491261d fix(web): hide nav auth links on public pages (/welcome, /invite)
CI / Lint / Test / Vet (push) Successful in 11s
CI / Build & Import (push) Successful in 10s
Logged-out visitors on /welcome and /invite should not see Account or
Log out. Split Layout into Layout (authenticated, full nav) and
PublicLayout (public, brand-only header). WelcomePage + InvitePage
variants now use PublicLayout.
2026-06-03 23:25:24 +02:00
mathiasandClaude Opus 4.8 72bf8a5553 docs(env): document TAPIR_PUBLIC_URL for tapir invite
CI / Lint / Test / Vet (push) Successful in 12s
CI / Build & Import (push) Successful in 10s
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-03 23:21:06 +02:00
mathiasandClaude Opus 4.8 dece5dec44 feat(web): public /invite/{token} set-password + account-creation flow
The Stage-1 onboarding path: an invited user opens their emailed link,
sets a password, and Tapir creates their Dex local-password account so
they can log in. Mounted on root OUTSIDE Auth.Middleware — the visitor
has no Dex session yet; the token in the path is the capability.

handleInviteForm previews the token (no consume) and shows the form, or
a clear "expired / already used" page. handleInviteSubmit validates the
password BEFORE consuming the token (a typo is retryable), then claims
the invite exactly once, bcrypt-hashes (cost 12), and creates the Dex
account — mapping ErrPasswordExists -> "log in instead" and ErrForbidden
-> "contact the administrator". Off-cluster (App.Dex nil) it degrades to
a "deployed-only" message without burning the token. On success it sets
an account_created flash and redirects to /auth/login.

Welcome sub-text now states access is invite-only. Handlers depend on
narrow ports (InvitationStore, DexPasswordCreator) so tests use fakes;
cmdServe wires the store + an in-cluster dex.PasswordClient.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-03 23:19:22 +02:00
mathiasandClaude Opus 4.8 893886a60a feat(cli): tapir invite <email> + TAPIR_PUBLIC_URL config
Mints a single-use invitation and prints the absolute claim URL for the
operator to send. The URL base is TAPIR_PUBLIC_URL (default
https://tapir.d-ma.be). runInvite is factored from config/store wiring so
it's unit-tested against a fake inviter — no Postgres.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-03 23:15:24 +02:00
mathiasandClaude Opus 4.8 e44485df16 feat(dex): in-cluster Password CR client for local-password accounts
Writes passwords.dex.coreos.com CRs against the in-cluster Kubernetes
API using the pod's service-account token + cluster CA (no kubectl /
client-go dependency). NewPasswordClient returns ErrNotInCluster off
cluster so the web layer degrades gracefully in dev.

Load-bearing: Dex's kubernetes storage types Password.Hash as []byte,
which k8s JSON-marshals as base64 — so the `hash` field carries the
base64 of the bcrypt string, not the raw string. Storing the raw string
makes Dex's base64-decode-on-login produce garbage and every login fail.

409 -> ErrPasswordExists, 401/403 -> ErrForbidden (RBAC missing) so the
handler can give precise messages. Tested against an httptest TLS server.

bcrypt cost-12 hashing lives in the web handler; golang.org/x/crypto was
already a transitive dep (now promoted in go.sum).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-03 23:14:08 +02:00
mathiasandClaude Opus 4.8 8b7ef07ba3 feat(store): invitations table + create/peek/claim methods
Stage-1 email onboarding: Mathias mints an invite, the recipient claims
it to set a Dex password. Invitations exist before their user, so the
table carries no user_id FK and is deliberately outside RLS — the
32-byte crypto-random token is the capability (single-use, time-boxed).

ClaimInvitation consumes atomically (UPDATE ... WHERE used_at IS NULL
... RETURNING) so concurrent claims of one token can't both succeed.
PeekInvitation validates the link for the form without consuming it.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-03 23:12:40 +02:00
mathias 1cf58768ed docs: spec the Stage 0 usage-measurement build (login events)
CI / Lint / Test / Vet (push) Successful in 10s
CI / Build & Import (push) Successful in 10s
Small tapir slice to make the gate measurable as written: append-only
login_events (RLS, per-user-per-day throttle) + a union query over reads
(login_events) and acts (summary_actions) for distinct-active-weeks. Carries the
honesty caveats (unprompted not measurable; data accrues from deploy; week-bucket
noise at low N) and the delete-cascade footgun (no FK, needs explicit delete +
test) from the prior delete work. Out of scope: analytics, prompt-tracking,
dashboards.
2026-06-03 20:59:49 +00:00
mathias f45ba35e25 docs: VISION Stage 0 — keep "unprompted" as ideal, note measurement gap
CI / Lint / Test / Vet (push) Has been cancelled
CI / Build & Import (push) Has been cancelled
Reframes "unprompted" from an enforced criterion to a named measurement
limitation: organic-vs-prompted returns aren't distinguishable from any data
Tapir holds, so in practice all returns are counted and the result read with that
caveat (a nudged return is a weaker signal). Adds a "how it's measured" note
pointing at summary_actions (acts) + a new append-only login-events table
(read-returns), which accrue from deploy onward. Honest about the gap rather than
silently dropping the word.
2026-06-03 20:59:14 +00:00
mathiasandClaude Opus 4.8 943554a96c feat(web): "Retrying later" badge for rate-limited videos
CI / Lint / Test / Vet (push) Successful in 10s
CI / Build & Import (push) Successful in 10s
A discovered-but-unsummarized video whose caption fetch was rate-limited now
shows a passive  "Retrying later" chip (dim CharmDim styling, not the accent)
instead of the Summarize button — the user cannot fix a 429, the runner retries
automatically once the backoff window expires. Regenerated views_templ.go.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-03 22:56:48 +02:00
mathiasandClaude Opus 4.8 40a614c8d4 feat(runner): 429 backoff — skip still-throttled videos, persist status
After ProcessNewVideo the runner records transcript_status per outcome:
rate_limited (stamps the backoff clock), none, or fetched. Before fetching, a
video inside the TAPIR_FETCH_BACKOFF window is skipped (SkippedRateLimited) so a
just-429'd caption endpoint is not re-hit; once the window expires it retries.

Backoff/clock injected via variadic Options (WithBackoff, WithClock) so existing
New call sites and the fake-driven loop tests stay valid. Backoff 0 = always retry.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-03 22:56:07 +02:00
mathiasandClaude Opus 4.8 ce2fc62ef8 feat(config): TAPIR_FETCH_BACKOFF for rate-limit retry window
Adds FetchBackoff (Go duration, default 1h) controlling how long the run loop
waits before re-fetching a transcript that returned HTTP 429. Zero means always
retry. Not required by ValidateForRun — a zero/unset value is a valid policy.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-03 22:54:24 +02:00
mathiasandClaude Opus 4.8 0ceacc8230 feat(usecase): surface TranscriptSource on ProcessResult
The engine already distinguishes SourceNone from SourceRateLimited internally
but collapsed both into Skipped. Expose the source string so the runner can
persist the right transcript_status and apply rate-limit backoff, without the
engine taking on any store/retry concern (dependencies still point inward).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-03 22:53:11 +02:00
mathiasandClaude Opus 4.8 1e81965519 feat(store): TranscriptStatus read field + status setter/loader
Surfaces videos.transcript_status (migration 007) on SummaryRow and adds
SetTranscriptStatus / GetTranscriptStatus / RateLimitedVideoIDs.

SetTranscriptStatus is the single choke point for the rate-limit lifecycle:
"rate_limited" stamps rate_limited_at = NOW(), every other status clears it,
so the runner's backoff window and the UI badge read one consistent source.
RateLimitedVideoIDs is the per-pass loader (mirrors SeenVideoIDs) the runner
uses to skip still-throttled videos without re-hitting the caption endpoint.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-03 22:52:48 +02:00
mathias f50c072d65 fix(web): add Log out to the persistent nav header
CI / Lint / Test / Vet (push) Successful in 12s
CI / Build & Import (push) Successful in 10s
Logout was only reachable from /welcome. Users who are logged in had no way
to sign out from any app page (list, detail, account). Added to the shared
nav alongside Account.
2026-06-03 22:49:14 +02:00
mathias e6f508824b docs: add ADR-016 — Stage 0 gate revised to "me or a friend", behavioural
CI / Lint / Test / Vet (push) Successful in 19s
CI / Build & Import (push) Successful in 11s
Records the gate change: Stage 0 now passes when either the maintainer or an
onboarded friend returns unprompted in >=2 separate weeks. Behavioural (return
usage), not feedback-based, to resist politeness bias. Includes an honest
self-scrutiny note that this is a guardrail edit made while the original gate was
unmet — examined on that basis and proceeding because it broadens who supplies the
signal without softening the kind of signal required. Adds feedback-based-gate to
rejected alternatives.
2026-06-03 20:46:40 +00:00
mathiasandClaude Opus 4.8 689500c85e feat(store): migration 007 — per-video transcript status
CI / Lint / Test / Vet (push) Successful in 20s
CI / Build & Import (push) Successful in 10s
Adds videos.transcript_status (NULL|none|fetched|rate_limited) and
videos.rate_limited_at, so the runner can record a 429 and skip re-fetching a
still-throttled video until a backoff window elapses. Columns inherit the
existing videos RLS policy (migration 003); no policy change needed.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-03 22:45:29 +02:00
mathiasandClaude Opus 4.8 4678d473b8 feat(youtube): map caption 429 to SourceRateLimited
The baseUrl fetch mapped every non-200 to SourceNone, recording a 429 as a
permanent "no captions". 429 is the IP being rate-limited, not an absent
transcript. Return SourceRateLimited (still a graceful degrade, no error) so
the runner can retry after a backoff window. Other non-200s (403/404/5xx)
stay SourceNone.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-03 22:45:29 +02:00
mathiasandClaude Opus 4.8 c63b2de66d fix(web): actionable empty state for the summary list
Videos rows are created by `tapir run`, not when a YouTube account is
connected, so a freshly-connected account correctly shows an empty list —
but the old empty state ("No videos yet") gave no clue why or what to do.
Split it on whether the user has any connection:

- connected, no videos: a distinct accent callout telling them to run
  `tapir run` to discover subscriptions.
- not connected: a prompt with a Connect YouTube button.

handleList fetches connections only when the list is empty. Includes
web-shot captures of all three states under docs/ux-review/fixes/.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-03 22:45:29 +02:00
mathiasandClaude Opus 4.8 27aa319f1d feat(domain): add SourceRateLimited transcript source
429 from the caption endpoint means the IP is rate-limited (retry later),
not that the video has no captions. Distinguishing it from SourceNone is the
prerequisite for the runner's backoff/retry logic.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-03 22:45:29 +02:00
mathiasandClaude Opus 4.8 61d4d5bc4a fix(web): clarify landing CTA copy for new users
The "Get Started" button drops users straight into the shared Dex flow,
which has no separate "register" option — registration completes
automatically after first login. Users new to Tapir had no signal that
signing in is also how they sign up. Reword the sub-text to say so
explicitly, keeping the single Dex CTA (sign-in and sign-up are one flow).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-03 22:45:29 +02:00
mathiasandClaude Opus 4.8 483730cd03 fix(web): make anchor-styled buttons readable
The .btn class sets color:var(--accent-fg), but the generic a{} and
a:visited{} rules outrank it on <a> elements, so anchor buttons (the
landing "Get Started" CTA, "Connect YouTube") rendered their label
accent-on-accent — invisible. Add a.btn / a.btn:visited to restore the
button foreground.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-03 22:45:29 +02:00
mathias 477701fea2 docs: revise Stage 0 gate to "useful to me or a friend" (behavioural)
CI / Lint / Test / Vet (push) Successful in 11s
CI / Build & Import (push) Successful in 10s
Replaces the original "useful to me, specifically" gate with "me OR a friend
returns unprompted in >=2 separate weeks" — friendly-user signal counts, but the
test stays behavioural (return usage) not feedback-based, to resist politeness
bias. Folds the old Stage 1 ("a trusted user returns") into the new Stage 0 (they
were near-identical), renumbers hardening to Stage 1, and updates the drift
signals (the gate can be softened by mistaking polite feedback for evidence;
multi-user shipping ahead of the gate was a recorded exception per ADR-012, not a
precedent). Rationale recorded in ADR-016.
2026-06-03 20:44:19 +00:00
mathias 17fad140a6 docs: add ADR-015 — per-user credentials envelope-encrypted in PG18
CI / Lint / Test / Vet (push) Successful in 10s
CI / Build & Import (push) Successful in 10s
Records the infra#88 spike decision: per-user OAuth tokens are runtime app-state,
not config, so they live envelope-encrypted in PG18 under RLS (key from 1P via the
existing read-only SA) rather than in the vault. Infra creds stay ESO/1Password —
two mechanisms because they're two different things. Includes the falsification
conditions (frequent rotation; estate audit policy; key-rotation cost) so the
choice is earned not assumed. Adds the two rejected candidates (vault-write SA;
Supabase) to the rejected-alternatives table. Full reasoning in the #88 decision
doc; build + reboot-validation in #89.
2026-06-03 20:29:50 +00:00
mathias eb24a24b9c chore(ci): remove mirror job (SSH key rotation pending)
CI / Lint / Test / Vet (push) Successful in 10s
CI / Build & Import (push) Successful in 10s
Mirror to github.com was failing with 'unsupported in libcrypto' on every run
(OpenSSL 3.6 / OpenSSH 10 dropped support for the existing key format). Removed
rather than leave it polluting the CI signal. Re-add when the deploy key is
rotated to ed25519.
2026-06-03 22:21:22 +02:00
mathiasandClaude Opus 4.8 21e6ddd61e docs(ui-spec): record as-built deviations and additions
The Stage-0 ui-spec (ADR-011) predated multi-user and several UX
features. Appended a "Deviations and additions (as-built)" table —
without rewriting the spec — recording each feature shipped beyond it
(multi-user+RLS, registration gate, per-user YouTube connect, account
management, immediate web summarization, charmbracelet spinner,
auto/manual mode, public landing page) with the why and the
commit/ADR that covers each. Preserves the intent-vs-reality split.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-03 22:20:04 +02:00
mathiasandClaude Opus 4.8 152aab7a4a docs(use-cases): add registration, summarize-mode, landing scenarios
The .feature spec lagged shipped behaviour. Added three files, no
duplication of existing scenarios:
- registration.feature: new Dex subject -> registration gate (users +
  user_identities), returning subject straight through, account delete
  is tapir-side only and leaves other users intact, clean
  re-registration (ADR-012, ADR-013).
- summarize_mode.feature: auto summarizes every new video; manual
  (default) leaves them unsummarized until queued; queued video is
  processed and the flag cleared (migration 006).
- landing_page.feature: unauthenticated / -> /welcome, Get Started for
  guests, summary link + logout for authed users, logout -> /welcome.

Scoped to built features only — no Vimeo/Whisper/billing scenarios.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-03 22:19:15 +02:00
mathiasandClaude Opus 4.8 1018dc0df9 docs(architecture): document the tapir serve web surface (ADR-011/012)
The C4 view predated the web transport. Added a "Web surface" section
with a component diagram and prose covering: tapir serve (HTMX+Templ
over the unchanged store), oidc authenticate-only session, registration
gate (users + user_identities), per-user YouTube web connect callback,
account disconnect/delete (ADR-013 tapir-side only), immediate
summarization via background goroutine + HTMX status poll, and
auto/manual summarize mode (migration 006). Relabeled the L2 http node
to tapir serve. Corrected the out-of-scope isolation bullet: RLS is live
(ADR-012, migration 003), not deferred. ADR-003 stance preserved — the
engine/ports/sinks core is untouched; the web is a new transport.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-03 22:18:23 +02:00
mathiasandClaude Opus 4.8 74f4fd7f2a docs(data-model): reconcile schema with migrations 002-006
The ER diagram and notes predated migrations 002-006. Brought them to
code truth:
- add summary_actions (002), user_identities (004), video_connections
  (005) with real columns/constraints; rename token_secret_ref ->
  token_ref to match migration 005.
- add users.auto_summarize and videos.summarize_requested (006).
- note FORCE RLS coverage and sink_deliveries' EXISTS-derived policy
  (003); user_identities NOT RLS'd.
- mark AI_CREDENTIAL and SUBSCRIPTION as planned (no table exists; only
  videos.subscription_id, no FK).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-03 22:17:15 +02:00
mathiasandClaude Opus 4.8 0cc441d6ce docs(data-model): isolation enforcement is live (RLS), not dormant
The "Isolation invariant" section said enforcement was dormant at
Stage 0. ADR-012 turned it on: Postgres RLS ENABLE+FORCE on every
user-owned table (migration 003_rls.up.sql), keyed off the
tapir.current_user_id GUC set by the store's withUser helper, deny-by-
default on an unset GUC. The Stage-2 two-user isolation test
(internal/adapters/store/rls_test.go) is pulled forward and passing.
Kept the ADR-011 single-user history honest.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-03 22:15:13 +02:00
mathiasandClaude Opus 4.8 8415a97d15 docs(claude): update build state section to v0.4.0 reality
The "Current build state" section described a scaffolded, intentionally
RED repo (ErrNotImplemented, go 1.23, unverified confirm items). All
stale: build is v0.4.0 green, go 1.26.1 (go.mod), Stage 1 multi-user
with RLS shipped (ADR-012, migrations 001-006). Rewrote to describe the
actual adapters, cmd subcommands (incl. serve), and how to build/run.
Replaced the "confirm" list with the resolved homelab facts (LiteLLM
koala:30401, model-as-config) now pinned in homelab-integration.md.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-03 22:14:42 +02:00
mathiasandClaude Opus 4.8 fb425cbf9a chore(decisions): reorder ADR-010 to numeric position
ADR-010 sat between ADR-008 and ADR-009. Moved ADR-009 (TBD) ahead of
ADR-010 (timedtext caption acquisition) so the file reads 001..014 in
numeric order. Pure structural move, no content change.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-03 22:13:50 +02:00
mathiasandClaude Opus 4.8 e77edf58ef docs(web): update auth.go comments for multi-user reality (ADR-012)
Package comment said "Stage-0 ... (ADR-011)" and User.Subject said
"single-user allowlist (ADR-011)". Both stale: ADR-012 opened Stage 1
(multi-user, RLS-enforced isolation). Subject is now the user_identities
lookup key (migration 004) resolving to a per-user UUID; an unknown
subject hits the registration gate.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-03 22:13:13 +02:00
mathiasandClaude Opus 4.8 f15f57f9ed test(web): cover the /welcome landing page and its routing
CI / Lint / Test / Vet (push) Successful in 10s
CI / Build & Import (push) Successful in 11s
CI / Mirror to GitHub (push) Failing after 3s
Add a configurable fakeAuth (StubAuth can't express the logged-out case)
and assert: /welcome renders the logged-out CTA with no session and the
logged-in controls with one, unauthenticated / redirects to /welcome,
unauthenticated deep links redirect to /auth/login, and an authenticated
root still renders the list.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-03 21:51:21 +02:00
mathiasandClaude Opus 4.8 d208110002 feat(web): mount /welcome handler outside the auth guard
Wire GET /welcome on the root mux alongside /healthz, outside
Auth.Middleware. handleWelcome peeks the session via CurrentUser (no
redirect) and renders the logged-out or logged-in WelcomePage variant.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-03 21:50:36 +02:00
mathiasandClaude Opus 4.8 3a27bf1126 feat(web): WelcomePage landing component
Add the public landing page: a static Charm-box tapir mascot (reusing the
shared tapirLine helpers and charm palette), a tagline, and a single
'Get Started' CTA into the shared Dex flow — sign-in and sign-up are the
same URL (ADR-012). Logged-in visitors get a greeting plus links back into
the app and to log out. Regenerated views_templ.go committed alongside.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-03 21:50:21 +02:00
mathiasandClaude Opus 4.8 8ca374e657 fix(web): logout redirects to /welcome, not /auth/login
Logout was bouncing the just-logged-out visitor straight back into a Dex
login. Land them on the public /welcome page instead — an intentional UX
fix. Cookie clearing and server-side session deletion are unchanged.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-03 21:49:08 +02:00
mathiasandClaude Opus 4.8 0fdf2f7218 feat(web): route unauthenticated root to /welcome, deep links to login
An unauthenticated visit to / now lands on the public /welcome page
instead of bouncing straight to Dex. Deeper guarded paths still redirect
to /auth/login so the post-login round-trip returns the visitor to the
page they asked for. isPublicPath runs first, so no redirect loop.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-03 21:48:51 +02:00
mathiasandClaude Opus 4.8 d83943c86a feat(web): make /welcome a public path
The landing page must render without a session. Add /welcome to
isPublicPath so the auth middleware lets it through (alongside /healthz
and /auth/*), and assert the bypass in the public-paths test.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-03 21:48:25 +02:00
mathias 6b817f11b9 docs: add landing-page + doc-reconciliation build spec
CI / Lint / Test / Vet (push) Successful in 11s
CI / Mirror to GitHub (push) Failing after 3s
CI / Build & Import (push) Successful in 10s
Workstream A: public /welcome landing page (bubbletea aesthetic, one Dex login
flow, logged-in shortcuts) — with the oidc.go facts verified against main, incl.
the two corrections that only surface from reading the code (logout must redirect
to /welcome not /auth/login; bare-/ vs deep-link redirect split).

Workstream B: reconcile the guardrail docs against deployed reality (v0.4.0) —
auth.go comments, data-model isolation status + migrations 002-006 schema,
architecture web surface, use-case scenarios for the Stage-1 features, ADR
ordering, and a requirements-vs-shipped deviation check. Structured as two
parallel workstreams so the doc audit isn't done cursorily alongside the build.
2026-06-03 19:18:40 +00:00
mathias 672a0c8580 docs: add ADR-013 (delete semantics) and ADR-014 (429 handling + UX)
CI / Lint / Test / Vet (push) Successful in 10s
CI / Mirror to GitHub (push) Failing after 3s
CI / Build & Import (push) Successful in 10s
ADR-013 records the deliberate choice that account deletion is Tapir-side only
(cascade + secret purge), leaving the shared Dex identity intact — clean
re-registration, but a noted GDPR-shaped gap if Future C ever arrives.

ADR-014 specifies timedtext 429 handling: Retry-After-aware backoff, a single
per-egress-IP rate gate shared by the batch and click paths, and honest in-flight
UX (summarizing / queued-waiting / no-transcript) so a rate-limited fetch never
presents as a stuck spinner or error. Whisper stays deferred pending measurement
of the sustainable rate, which this work finally makes measurable.
2026-06-03 19:03:52 +00:00
155 changed files with 15428 additions and 1284 deletions
+28
View File
@@ -46,3 +46,31 @@ TAPIR_SECRETS_FILE=
# --- run loop -------------------------------------------------------------
# Empty/0 = single pass. Set (e.g. 15m) to poll on that cadence.
TAPIR_POLL_INTERVAL=
# How long to wait before re-fetching a transcript that returned HTTP 429
# (rate_limited). Inside the window the video is skipped without hitting the
# caption endpoint; after it expires the video is retried. 0 = always retry.
# Go duration; default 1h.
TAPIR_FETCH_BACKOFF=
# Recency bound for AUTO summarization: in automatic mode only videos published
# within this window of now are summarized; older ones are discovered + listed
# but wait for a manual "Summarize" (so a back-catalogue doesn't self-inflict
# 429s). An explicit request bypasses it. Go duration; default 168h (~7d).
# 0 = no bound (summarize every unseen video).
TAPIR_AUTO_SUMMARIZE_WINDOW=
# Minimum interval between outbound caption fetches across the WHOLE process —
# the shared per-egress-IP rate gate (ADR-014). Scheduler runners and the web
# "Summarize" click-path serialise through it so they cannot collectively trip
# 429s. Go duration; default 2s. 0 = unlimited (dev/tests).
TAPIR_FETCH_RATE=
# --- scheduled discovery (tapir serve, ADR-018) ---------------------------
# When > 0, `serve` runs in-process discovery for ALL users on this cadence
# (e.g. 2h): one runner pass per user per tick, run-once-on-startup then ticked.
# Empty/0 = disabled (dev/tests never auto-fetch). SINGLE-REPLICA assumption —
# >1 replica double-runs discovery. Go duration.
TAPIR_DISCOVERY_INTERVAL=
# --- invitations (tapir invite) -------------------------------------------
# Public base URL used to build the invite link `tapir invite <email>` prints.
# Default https://tapir.d-ma.be; no trailing slash needed.
TAPIR_PUBLIC_URL=
+131 -19
View File
@@ -78,36 +78,148 @@ jobs:
${REGISTRY}/${{ env.IMAGE }}:${{ steps.meta.outputs.version-tag }} || true
echo "Image pushed to ${REF}"
# Run the just-built local image briefly via buildah, not k3s ctr —
# avoids sudo/host-containerd access so this still works once the
# runner itself is containerized (infra#132). Tests the local
# buildah-store image directly, no registry round-trip needed.
#
# Bare `/tapir` (no subcommand) runs the long-running server, same as
# the old ctr-based smoke test -- ctr's --rm reliably force-killed it,
# but a plain `timeout N buildah run` does NOT: it only signals the
# `buildah run` wrapper, and the actual container process can survive
# that and keep the output pipe open, hanging the whole job (hit this
# live: 14min hang, infra#132). Backgrounding the run + `buildah rm -f`
# decouples "is the script blocked" from "did the process exit" --
# rm -f forcibly tears down the container regardless of wrapper state.
- name: Smoke test
run: |
REGISTRY="localhost:5000"
REF="${REGISTRY}/${{ env.IMAGE }}:${{ steps.meta.outputs.sha-tag }}"
CNAME="smoke-${{ steps.meta.outputs.sha-tag }}"
sudo k3s ctr images pull --plain-http ${REF}
OUTPUT=$(timeout 5 sudo k3s ctr run --rm ${REF} ${CNAME} /tapir 2>&1 || true)
sudo k3s ctr containers delete ${CNAME} 2>/dev/null || true
CONTAINER=$(buildah from ${REF})
LOG=$(mktemp)
buildah run "$CONTAINER" -- /tapir > "$LOG" 2>&1 &
RUNPID=$!
sleep 5
kill -9 "$RUNPID" 2>/dev/null || true
wait "$RUNPID" 2>/dev/null || true
buildah rm -f "$CONTAINER" >/dev/null 2>&1 || true
OUTPUT=$(cat "$LOG"); rm -f "$LOG"
echo "$OUTPUT" | grep -q "tapir" \
&& echo "Smoke test passed" \
|| echo "Smoke test inconclusive: $OUTPUT"
# ── 3. Mirror to GitHub (deploy intentionally omitted until manifests exist)
mirror:
name: Mirror to GitHub
# ── 3. Deploy via infra repo + Flux ────────────────────────────────────────
# Flux native image-automation can't scan localhost:5000 from inside k3s pods
# (mathias/infra k3s/flux/flux-system/image-automation.yaml) — this job
# mirrors cobalt-dingo's proven pattern instead: patch the infra repo's
# manifest directly on every push to main, then annotate Flux for a fast
# reconcile. Fixes infra#111 (image built+pushed but manifest bump was
# manual, so merged features silently didn't deploy).
deploy:
name: Deploy via GitOps
needs: build
runs-on: self-hosted
if: github.ref == 'refs/heads/main' && github.event_name == 'push'
environment: staging
steps:
- uses: actions/checkout@v4
with:
fetch-depth: 0
- name: Push to GitHub
- name: Update image tag in infra repo
env:
IMAGE_TAG: ${{ needs.build.outputs.image-tag }}
DEPLOY_KEY: ${{ secrets.INFRA_DEPLOY_KEY }}
run: |
set -euo pipefail
# INFRA_DEPLOY_KEY is a Gitea org secret (mathias org), already
# configured per docs/cd-pipeline.md in the infra repo — same key
# cobalt-dingo and brain-gardener use, no new secret needed.
mkdir -p ~/.ssh
echo '${{ secrets.GH_DEPLOY_KEY }}' > ~/.ssh/id_rsa_gh_mirror
chmod 600 ~/.ssh/id_rsa_gh_mirror
ssh-keyscan github.com >> ~/.ssh/known_hosts 2>/dev/null
GIT_SSH_COMMAND="ssh -i ~/.ssh/id_rsa_gh_mirror -o IdentitiesOnly=yes" \
git push git@github.com:mathiasb/tapir.git HEAD:main
rm ~/.ssh/id_rsa_gh_mirror
echo "Mirrored to GitHub"
echo "$DEPLOY_KEY" > ~/.ssh/id_infra
chmod 600 ~/.ssh/id_infra
ssh-keyscan -p 30022 10.0.1.20 >> ~/.ssh/known_hosts 2>/dev/null
export GIT_SSH_COMMAND="ssh -i ~/.ssh/id_infra -o IdentitiesOnly=yes"
rm -rf /tmp/infra
git clone -b main ssh://git@10.0.1.20:30022/mathias/infra.git /tmp/infra
cd /tmp/infra
DEPLOYMENT="k3s/apps/tapir/deployment.yaml"
# In-place update of the image tag. sed (not yq) so we don't
# depend on additional tooling on the runner — same as cobalt-dingo.
sed -i "s|image: localhost:5000/tapir:.*|image: localhost:5000/tapir:${IMAGE_TAG}|" "$DEPLOYMENT"
# Verify the patch took effect.
grep -q "localhost:5000/tapir:${IMAGE_TAG}" "$DEPLOYMENT" \
|| { echo "✗ image tag patch failed"; exit 1; }
if git diff --quiet "$DEPLOYMENT"; then
echo " image tag unchanged — skipping push"
else
git -c user.name="tapir CI" \
-c user.email="ci@tapir.local" \
commit -m "chore(deploy): tapir → ${IMAGE_TAG}" "$DEPLOYMENT"
git push origin main
echo "✓ pushed to infra repo"
fi
shred -u ~/.ssh/id_infra
- name: Trigger Flux reconcile (immediate)
run: |
# Without these annotations, Flux would still pick up the change
# within 30s (the apps Kustomization interval). The annotations
# cut latency to ~1s.
kubectl -n flux-system annotate gitrepository flux-system \
reconcile.fluxcd.io/requestedAt="$(date +%s)" --overwrite
kubectl -n flux-system annotate kustomization apps \
reconcile.fluxcd.io/requestedAt="$(date +%s)" --overwrite
- name: Wait for Flux to apply new image
env:
IMAGE_TAG: ${{ needs.build.outputs.image-tag }}
run: |
# Poll the Deployment spec until it reflects the new tag.
# Bound to 60s so a stuck Flux doesn't hang CI.
EXPECTED="localhost:5000/tapir:${IMAGE_TAG}"
for i in $(seq 1 60); do
CURRENT=$(kubectl get deploy tapir -n tapir \
-o jsonpath='{.spec.template.spec.containers[0].image}' 2>/dev/null || echo "")
if [ "$CURRENT" = "$EXPECTED" ]; then
echo "✓ Flux applied new image after ${i}s"
break
fi
sleep 1
done
# Final assertion (in case the loop exited without matching).
kubectl get deploy tapir -n tapir \
-o jsonpath='{.spec.template.spec.containers[0].image}' \
| grep -qx "$EXPECTED" \
|| { echo "✗ Flux did not apply new image within 60s"; exit 1; }
- name: Verify rollout
run: |
kubectl rollout status deployment/tapir \
--namespace tapir \
--timeout=120s \
|| {
echo "── pod status ──"
kubectl get pods -n tapir -o wide
echo "── events ──"
kubectl get events -n tapir --sort-by='.lastTimestamp' | tail -20
echo "── describe ──"
kubectl describe pods -n tapir -l app=tapir | tail -40
exit 1
}
- name: Confirm pod running new image
env:
IMAGE_TAG: ${{ needs.build.outputs.image-tag }}
run: |
kubectl get pods -n tapir \
-l app=tapir \
--field-selector=status.phase=Running \
-o jsonpath='{.items[*].spec.containers[0].image}' \
| grep -q "localhost:5000/tapir:${IMAGE_TAG}" \
&& echo "✓ pod running new image" \
|| { echo "✗ pod image mismatch"; exit 1; }
# ── 4. Mirror to GitHub — skipped for now (SSH key rotation pending) ─
+11
View File
@@ -11,3 +11,14 @@
.env.*
!.env.example
*.local
# Spike media (#28): real recordings, their transcripts and derived analyses are
# private third-party content and this repo is public. Only synthetic fixtures
# are committed — see scripts/spike-media/README.md.
/scripts/spike-media/*.mov
/scripts/spike-media/*.mp4
/scripts/spike-media/*.wav
/scripts/spike-media/*.srt
/scripts/spike-media/*.json
/scripts/spike-media/*.html
!/scripts/spike-media/fixtures/
+18
View File
@@ -0,0 +1,18 @@
{
"mcpServers": {
"brain": {
"type": "http",
"url": "https://brain-mcp.d-ma.be/mcp",
"headers": {
"Authorization": "Bearer ${BRAIN_MCP_TOKEN}"
}
},
"gitea": {
"type": "http",
"url": "https://git-mcp.d-ma.be/mcp",
"headers": {
"Authorization": "Bearer ${GITEA_MCP_TOKEN}"
}
}
}
}
+43 -19
View File
@@ -12,6 +12,7 @@ docs it indexes.
3. `DECISIONS.md` — the ADRs. Decisions are settled here; do not re-litigate without a new ADR.
4. `docs/architecture/architecture.md`, `docs/data-model.md`, `docs/use-cases/*.feature`.
5. `docs/homelab-integration.md` — the concrete endpoints/conventions you'll need.
6. `LANGUAGE.md` — the project vocabulary. Apply the caveman rubric before destructive operations.
## How to work in this repo
@@ -46,8 +47,9 @@ These caused real mistakes that were caught and corrected; the corrections are l
(See `DECISIONS.md` for full rationale. Listed here so you don't propose them.)
- **No Supabase** — reuse Dex / ESO+1Password / Postgres (ADR-002).
- **No global cross-tenant video/transcript table** — per-user isolation (data-model). Dedup
across users is a Future C concern, not a Stage 0/1 default.
- **No global cross-tenant *video* table** — videos stay per-user (data-model). Transcripts ARE
shared since ADR-021 (public caption content, keyed by `(provider, provider_video_id)`, non-RLS)
so re-analysis never re-fetches; the *videos* half of cross-tenant dedup stays a Future C concern.
- **No audio-download + speech-to-text in the core path** — captions-first (ADR-007). STT is a
deferred, bounded optional component.
- **No public SaaS / sign-up / billing / Google OAuth verification at scale** — Future C,
@@ -60,8 +62,14 @@ These caused real mistakes that were caught and corrected; the corrections are l
stores, and sinks are adapters. Adding a video provider or a sink = a new adapter implementing
the interface, nothing in the engine changes. This is what keeps "standalone vs homelab" a
wiring choice (ADR-003).
- **BDD.** The `docs/use-cases/*.feature` files are the behavior spec. New behavior gets a
scenario; the use-case core is tested through fake adapters, not live YouTube/brain.
- **BDD.** The `docs/use-cases/*.feature` files are the behavior spec (design records — there is
no godog runner). New behavior gets a scenario; the use-case core is tested through fake
adapters, not live YouTube/brain. A name-coverage gate (`test/acceptance/scenario_coverage_test.go`,
`TestScenarioCoverage`) keeps the two from drifting: every non-`@pending` scenario must be
mapped to an existing Go test in `scenarioCoverage`. When you add a scenario, either map it to
its covering test or tag it `@pending` in the `.feature` with a one-line reason. It checks the
*link*, not that the test exercises the scenario — that's the deliberate trade for not running
godog (see issue #5 / the BDD-runner decision).
## Skills (engineering discipline)
@@ -79,23 +87,39 @@ Skills live in the canonical library `mathias/skills` and are wired into this re
## Current build state (start here for the first task)
The repo is **scaffolded and intentionally RED**:
The repo is **green and shipping** — last tag `v0.15.0`. `task check` passes (fmt, vet, lint,
`go test -p 1 ./...`). Go is `1.26.1` (see `go.mod`).
- Clean Architecture skeleton exists: `internal/domain` (entities), `internal/ports`
(interfaces), `internal/usecase` (engine), `cmd/tapir` (entrypoint stub),
`internal/adapters` (empty — concrete adapters go here).
- `usecase.Engine.ProcessNewVideo` returns `ErrNotImplemented`.
- `test/acceptance/summarize_new_video_test.go` translates the first two Gherkin scenarios and
**fails** against the stub. `task check` is therefore red on `test`.
- **First build task:** implement `ProcessNewVideo` (resolve transcript -> summarize -> deliver to
sinks | skip on no-transcript) to make the acceptance tests green, following the `.feature`
files. Then add the AI-router `Summarizer` (copy `llm` per ADR-004), the YouTube `VideoSource`
adapter (captions-first), and the store + brain sinks.
- Clean Architecture core is implemented: `internal/domain` (entities), `internal/ports`
(interfaces), `internal/usecase.Engine.ProcessNewVideo` (resolve transcript → summarize →
deliver to sinks | skip on no-transcript). The acceptance tests in `test/acceptance/` are
green against it.
- Adapters present under `internal/adapters/`: `youtube` (captions-first `VideoSource`,
timedtext/InnerTube acquisition per ADR-010), `summarizer` + `llm` (the copied AI router,
now a resilient endpoint chain — local primary → local fallback → external worst-case,
parse-failure-aware, ADR-004 + ADR-022), `store` (Postgres, golang-migrate migrations 001015),
`secrets` (file-backed `SecretStore`). The brain HTTP sink (ADR-005) is the remaining
optional sink.
- Stage 1 is open (ADR-012): multi-user with **DB-enforced** isolation — Postgres RLS `FORCE`d
on all user-owned tables (migration 003), two-user isolation test in
`internal/adapters/store/rls_test.go`. Registration gate, per-user YouTube web connect, and
account management (disconnect / delete, ADR-013) all shipped.
- Transcript persistence (ADR-021, migration 015): transcripts are a **shared, non-RLS** store
keyed by `(provider, provider_video_id)` — the single exception to the isolation boundary
(`TestTranscriptsTableIsSharedNotRLS`). The engine reads stored transcripts before any caption
fetch (`usecase.resolveTranscript`), so re-analysis — re-summarize, paste-a-URL, onboarding
burst — never re-touches YouTube. Per-user summaries/videos stay RLS-scoped.
- `cmd/tapir` subcommands: `list`, `show`, `auth` (interactive host-side OAuth), `run` (batch
watch→summarize), `serve` (the HTMX+Templ web reader/writer under `internal/web`, a new
transport over the unchanged engine/ports — ADR-003). `tapir env` prints config.
- **Build/run:** `task check` is the gate; `task build` produces the binary. Local dev uses
`StubAuth` (allow-all) and a `TAPIR_DB_DSN` Postgres; the deployed service uses Dex OIDC.
**Unverified setup items** (see `docs/homelab-integration.md`, marked `confirm`): the Go version
in `go.mod` (1.23 — match the koala runner; estate elsewhere uses 1.26.1), the brain-mcp URL, the
exact ESO secret-ref naming, and the summarization model alias. Resolve against the live cluster
before depending on them, and pin answers back into `docs/homelab-integration.md`.
**Setup facts** (resolved — see `docs/homelab-integration.md` for the live values): LiteLLM is
off-cluster at `koala:30401/v1/` with `LITELLM_MASTER_KEY` from 1Password; the summarization
model is config (`TAPIR_SUMMARIZER_MODEL`, default `koala/phi4-mini`), never hardcoded. The
brain-mcp base URL and ESO ref scheme are pinned in that doc; check it before wiring rather than
re-deriving.
## Provenance (where this design came from)
+1048 -18
View File
File diff suppressed because it is too large Load Diff
+26
View File
@@ -0,0 +1,26 @@
# Vocabulary (hand-written pilot 2026-06 — will be generated by langgen in Phase 2)
| Term | Means | Never say |
|---|---|---|
| transcript | Captions-first text of a video; the source for summarizing. Shared store keyed by `(provider, video_id)`. | "subtitles file", "the audio" |
| video | Per-user record of a seen video; the unit of work. Not globally deduped. | "the shared video" |
| subscription | A watched channel on a user's connection that Tapir polls for new videos. | "feed", "follow" |
| summary | A video's produced output: summary text + highlights + takeaways + AI provenance. | "transcript" |
| highlight | A notable point pulled from a video (`Summary.Highlights`). | "takeaway" |
| takeaway | An actionable conclusion from a video (`Summary.Takeaways`). | "highlight" |
| sink | A delivery destination for a summary (`store`, `brain`). New one = new adapter. | "the database" |
| AI router | Local-first chain: local Primary → local fallback → external worst-case. | "the API" |
| BYO-AI fallback | User's own external AI key, used only when local fails; opt-in. Without it, content is never sent externally. | "the default AI" |
| brain | Persistent homelab knowledge store; in Tapir, one optional HTTP sink (`brain_ingest`), not the filesystem package. | "the database", "the store" |
| LiteLLM gateway | Local AI gateway (`koala:30401/v1`); the Primary in the AI router. | "piguard:4000", "koala:4000", "the cloud" |
| connection | A connected video account a user authorizes via OAuth; subscriptions hang off it. | "login", "session" |
## Caveman rubric
Before any HIGH/CRITICAL operation, output one line:
`caveman: me <verb> <object>, not <excluded thing>`
Valid iff: (1) only vocabulary terms + plain verbs, (2) a stranger could identify
the exact operation, (3) names one thing explicitly NOT being done.
Examples:
- `caveman: me delete user connection, not the shared transcript`
- `caveman: me send transcript to BYO-AI fallback, not the LiteLLM gateway`
+14 -2
View File
@@ -37,7 +37,7 @@ model, behavior specs) remain the source of intent.
## Running the Stage-0 demo
Tapir runs on **your own** YouTube account: authorize once, then run the
Tapir runs on your YouTube account(s): authorize once, then run the
watch→summarize→deliver loop. All configuration is via `TAPIR_*` environment
variables — copy [`.env.example`](.env.example) to `.env` and fill it in (no
secrets are committed; at demo time source them from op, e.g. `op run -- ...`).
@@ -56,7 +56,7 @@ go build -o bin/tapir ./cmd/tapir
./bin/tapir auth
# 3. run: detect new videos across your subscriptions, summarize, deliver to the
# store. Unset TAPIR_POLL_INTERVAL = single pass; set it (e.g. 15m) to loop.
# store. Single pass; set TAPIR_DISCOVERY_INTERVAL (e.g. 2h) for the serve loop.
./bin/tapir run
```
@@ -68,6 +68,18 @@ URI matches `TAPIR_OAUTH_REDIRECT_ADDR`. The summarizer model
(`TAPIR_SUMMARIZER_MODEL`, default `koala/phi4-mini`) is overridable; pick the
final alias when the gateway is reachable (see `docs/homelab-integration.md`).
### Web surface (`tapir serve`)
`tapir serve` starts the HTMX+Templ web UI on `:8080`. Users log in via Dex OIDC (local
password or Google); a new Dex subject is routed to `/register` to create a Tapir account.
Stage 1 is multi-user: each user connects their own YouTube account from the browser and
manages their own summaries under DB-enforced RLS isolation. When `TAPIR_DISCOVERY_INTERVAL`
is set (e.g. `2h`), the serve process runs a scheduled discovery pass for every registered
user automatically — no CronJob required. In auto mode only videos published within
`TAPIR_AUTO_SUMMARIZE_WINDOW` (default ~7d, ADR-020) are summarised automatically; older videos
are listed and summarised on demand, so a large back-catalogue doesn't keep re-driving the
caption rate gate. See `docs/homelab-integration.md` for the full config reference.
### Headless on koala
koala has no browser and no interactive `op` session, so the two interactive
+55 -27
View File
@@ -44,50 +44,74 @@ fallback — their key, their choice.
## Who it is for
- **Now (the first customer):** the maintainer — one person, their own subscriptions,
summaries delivered to their own store and brain.
- **Soon (Future B):** a small number of known, trusted users (friends / beta) — each with
their own account, isolated data, optional BYO-AI.
- **Now (the first customers):** the maintainer and a small number of known, trusted
friends — each with their own account, isolated data, optional BYO-AI. The maintainer is
the first customer; friendly users provide the earliest real-world signal.
- **Maybe (Future C, explicitly not built yet):** a public multi-tenant service. Deferred
until there is evidence of sustained personal use **and** real demand. Building for C
before that evidence is a known anti-goal.
until there is evidence of sustained use **and** real demand. Building for C before that
evidence is a known anti-goal.
## Definition of Success
Success is staged. Each stage has a single, falsifiable headline test. We do not advance
to the next stage's ambition until the current stage's test passes.
### Stage 0 — Useful to me (the gate)
### Stage 0 — Useful to me or a friend (the gate)
> **Headline test:** For four consecutive weeks, the maintainer reads Tapir-produced
> summaries for their own subscriptions at least weekly, and at least once acts on a
> summary (watches / skips / saves a video *because of* the summary).
> **Headline test:** Over a 34 week window, *either* the maintainer *or* at least one
> onboarded friend returns to Tapir and reads/acts on summaries in **≥2 separate weeks**.
> The test is *return usage* (behavioural), not stated approval. The ideal signal is an
> **unprompted** return (organic, not because the maintainer nudged them) — but see the
> measurement note below: we currently cannot distinguish prompted from organic returns, so
> in practice we count all returns and read the result with that caveat.
- Captions-first summarization works end-to-end for the maintainer's real subscriptions.
- Summaries land in the maintainer's store and (optionally) brain.
- Captions-first summarization works end-to-end for real subscriptions (the maintainer's
and onboarded friends').
- Summaries land in each user's own store and (optionally) brain.
- Local-first AI produces summaries of acceptable quality without manual intervention
most of the time.
- **This is the gate.** Multi-user, BYO-AI-for-others, and any SaaS ambition stay deferred
until Stage 0 holds. (Ties to the 2026-07-01 self-use check-in.)
- **Why behavioural, not feedback.** Friend *feedback* is gathered and genuinely valuable —
but it is **not** the gate. Asked-for feedback from friendly users is the least reliable
signal in product development (politeness bias); whether they *come back* is the thing we
actually care about. So the gate measures returns, not nice words.
- **Measurement note — "unprompted" is an ideal we can't yet measure.** Whether a return was
organic or prompted by a nudge is not captured by any data Tapir holds (it's context only
the maintainer has). Rather than waive the standard, we name the gap: *unprompted* return
is the signal we genuinely want; *returns* (prompted or not) is what the data can show. A
return that needed a nudge is a weaker signal than one that didn't, and the result is read
with that in mind. If distinguishing them ever matters enough, the maintainer tracks nudges
manually or a future build records prompt events — neither is in scope now.
- **Why "me OR a friend".** This replaces the original "useful to *me*, specifically" gate
(2026-06-03 decision, recorded in DECISIONS.md ADR-016). Getting signal from friendly
users is valuable enough to count — but the bar stays behavioural so it can't be cleared
by a polite reaction. (Ties to the 2026-07-01 check-in.)
- **How it's measured.** Return usage is read from two sources: `summary_actions` (timestamped
watch/skip/save per user) answers "acted in ≥2 distinct weeks"; an append-only login-events
table (see infra/Tapir build) answers "returned/read in ≥2 distinct weeks" even without an
action click — the honest signal for a *reading* product. Login events accrue only from their
deploy date onward, so the gate window's data begins then.
- **Gate-clock reset (ADR-018).** The 34 week window starts when in-process scheduled discovery
+ auto-summarize ship — before that, unprompted use was impossible, so the prior window
measured nothing (this is starting the clock when the experiment can actually run, not a reset
to dodge a failing gate). The 2026-07-01 check-in referenced above moves accordingly to ~34
weeks after this deploys. See DECISIONS.md ADR-018.
- **This is the gate.** Hardening (Stage 1) and any SaaS ambition stay deferred until this
behavioural signal exists. Note: multi-user machinery was deliberately built *ahead* of
this gate (ADR-012) with isolation enforced — that was an explicit, recorded call, not a
sign the gate had passed. The gate is about *evidence of use*, which is still open.
### Stage 1 — Useful to a few (Future B)
> **Headline test:** At least one trusted user other than the maintainer connects their
> own account and, within their first month, keeps using it (returns to read summaries in
> ≥2 separate weeks) without the maintainer hand-holding each summary.
- Multiple users, each with isolated accounts, credentials, and summaries.
- A new user can self-connect a YouTube/Vimeo account and get summaries with no code change.
- Optional BYO-AI works per-user.
- No cross-user data leakage — demonstrable, not assumed.
### Stage 2 — Trustworthy at rest (hardening, still Future B)
### Stage 1 — Trustworthy at rest (hardening, Future B)
> **Headline test:** Credentials (OAuth tokens, BYO-AI keys) are encrypted at rest via the
> homelab's existing secrets convention; a documented, rehearsed recovery path exists; and
> a deliberate isolation test (user A cannot read user B's data) passes in CI or a
> documented manual drill.
- Per-user data isolation is enforced and tested (delivered early via ADR-012 RLS).
- Per-user credentials are encrypted at rest (ADR-015 envelope encryption; build in infra#89).
- A new user can self-connect a YouTube/Vimeo account and get summaries with no code change.
- Optional BYO-AI works per-user.
### Non-goals (current)
- Public sign-up / billing / a marketing surface.
@@ -98,7 +122,11 @@ to the next stage's ambition until the current stage's test passes.
## How we will know we are drifting
- We are building Stage 1+ machinery before the Stage 0 gate has passed.
- We declare the Stage 0 gate "passed" on the strength of polite feedback rather than
behavioural return-usage (the politeness-bias trap the gate is designed to resist).
- We build Stage 1 hardening or Future C machinery while the Stage 0 use-evidence is still
absent. (Multi-user machinery already shipped ahead of the gate via ADR-012 — a recorded,
deliberate exception, not a precedent for more.)
- A user's content reaches a third-party model without that user's explicit, per-user opt-in.
- "Brain ingestion" starts dictating the architecture instead of being one sink behind an
interface.
+64
View File
@@ -0,0 +1,64 @@
package main
import (
"strings"
"testing"
"git.d-ma.be/mathias/tapir/internal/config"
)
// TestChatModelsReuseTheChainLocalFirst: the switcher offers the ADR-022 chain in
// order, deduped — primary, local fallback, cloud.
func TestChatModelsReuseTheChainLocalFirst(t *testing.T) {
got := chatModels(config.Config{
SummarizerModel: "koala/phi4-mini",
FallbackModel: "iguana/gemma4-26b",
CloudFallbackModel: "berget/mistral-small",
})
want := []string{"koala/phi4-mini", "iguana/gemma4-26b", "berget/mistral-small"}
if strings.Join(got, ",") != strings.Join(want, ",") {
t.Fatalf("chatModels = %v, want %v", got, want)
}
}
// TestChatCloudModelAbsentWhenDisabled: the local-first / NDA lever — with the
// cloud fallback empty (TAPIR_CLOUD_FALLBACK_MODEL=""), no external model is
// offered in the switcher, so chat content never leaves the local stack (ADR-027,
// honouring ADR-022's "content stays local" guarantee).
func TestChatCloudModelAbsentWhenDisabled(t *testing.T) {
got := chatModels(config.Config{
SummarizerModel: "koala/phi4-mini",
FallbackModel: "iguana/gemma4-26b",
CloudFallbackModel: "",
})
for _, m := range got {
if strings.HasPrefix(m, "berget/") || strings.Contains(m, "mistral") {
t.Fatalf("cloud model %q offered though the cloud fallback is disabled", m)
}
}
want := []string{"koala/phi4-mini", "iguana/gemma4-26b"}
if strings.Join(got, ",") != strings.Join(want, ",") {
t.Fatalf("chatModels = %v, want %v", got, want)
}
}
// TestChatModelsDedup: a config that reuses one alias across slots collapses to a
// single switcher entry (no duplicate options).
func TestChatModelsDedup(t *testing.T) {
got := chatModels(config.Config{
SummarizerModel: "koala/phi4-mini",
FallbackModel: "koala/phi4-mini",
CloudFallbackModel: "",
})
if len(got) != 1 || got[0] != "koala/phi4-mini" {
t.Fatalf("chatModels = %v, want a single deduped entry", got)
}
}
// TestBuildChatNilWithoutGateway: no gateway → no chat backend (routes unmounted,
// read path unaffected).
func TestBuildChatNilWithoutGateway(t *testing.T) {
if c := buildChat(config.Config{GatewayURL: ""}); c != nil {
t.Fatal("buildChat must return nil without a gateway URL")
}
}
+50
View File
@@ -0,0 +1,50 @@
package main
import (
"context"
"log/slog"
"sync"
"git.d-ma.be/mathias/tapir/internal/runner"
)
// discoveryRunner runs one user's discovery pass.
type discoveryRunner func(ctx context.Context, userID string) (runner.Stats, error)
// serialize wraps run so calls never overlap: every discovery pass — scheduled
// or connect-triggered (#6) — acquires the same lock, preserving the
// one-fetcher-at-a-time invariant the scheduler relies on (ADR-018, the
// single-replica assumption). Locking is per-user, so a connect-triggered pass
// interleaves between the scheduler's users instead of waiting for a whole pass.
func serialize(mu *sync.Mutex, run discoveryRunner) discoveryRunner {
return func(ctx context.Context, userID string) (runner.Stats, error) {
mu.Lock()
defer mu.Unlock()
return run(ctx, userID)
}
}
// discoveryTrigger fires an out-of-band discovery pass for one user without
// blocking the caller (the connect HTTP handler). The pass runs on the server's
// long-lived ctx — not the request ctx — so it survives the post-connect
// redirect. run is the serialized runner, so a trigger never overlaps the
// scheduler. Satisfies web.DiscoveryTrigger.
type discoveryTrigger struct {
ctx context.Context
run discoveryRunner
// onboard, when set, runs after the discovery pass to summarize a capped number
// of the user's newest videos (Feature 1). Optional.
onboard func(ctx context.Context, userID string)
log *slog.Logger
}
func (t *discoveryTrigger) Enqueue(userID string) {
go func() {
if _, err := t.run(t.ctx, userID); err != nil {
t.log.Warn("discovery: connect-triggered pass had errors", "user", userID, "err", err)
}
if t.onboard != nil {
t.onboard(t.ctx, userID)
}
}()
}
+62
View File
@@ -0,0 +1,62 @@
package main
import (
"context"
"fmt"
"sync"
"sync/atomic"
"testing"
"time"
"github.com/stretchr/testify/require"
"git.d-ma.be/mathias/tapir/internal/runner"
)
// serialize must guarantee at most one discovery pass runs at a time, so a
// connect-triggered pass never fetches concurrently with the scheduler.
func TestSerializeRunsOneAtATime(t *testing.T) {
var active, maxActive int32
run := func(_ context.Context, _ string) (runner.Stats, error) {
n := atomic.AddInt32(&active, 1)
for { // record the high-water mark of concurrent runs
m := atomic.LoadInt32(&maxActive)
if n <= m || atomic.CompareAndSwapInt32(&maxActive, m, n) {
break
}
}
time.Sleep(2 * time.Millisecond)
atomic.AddInt32(&active, -1)
return runner.Stats{}, nil
}
s := serialize(&sync.Mutex{}, run)
var wg sync.WaitGroup
for i := 0; i < 20; i++ {
wg.Add(1)
go func(i int) { defer wg.Done(); _, _ = s(context.Background(), fmt.Sprintf("u%d", i)) }(i)
}
wg.Wait()
require.Equal(t, int32(1), atomic.LoadInt32(&maxActive),
"serialize must run at most one pass at a time")
}
// Enqueue runs the user's pass out-of-band (non-blocking) on the trigger's ctx.
func TestDiscoveryTriggerEnqueueRunsUser(t *testing.T) {
done := make(chan string, 1)
run := func(_ context.Context, userID string) (runner.Stats, error) {
done <- userID
return runner.Stats{}, nil
}
tr := &discoveryTrigger{ctx: context.Background(), run: run, log: quietLog()}
tr.Enqueue("u1")
select {
case got := <-done:
require.Equal(t, "u1", got)
case <-time.After(2 * time.Second):
t.Fatal("Enqueue did not run the user's pass")
}
}
+1 -1
View File
@@ -5,7 +5,7 @@ import (
"fmt"
"os"
"gitea.d-ma.be/mathias/tapir/internal/adapters/store"
"git.d-ma.be/mathias/tapir/internal/adapters/store"
)
// Env var names for the read-only CLI. DSN and user id are never hardcoded — the
+1 -1
View File
@@ -7,7 +7,7 @@ import (
"github.com/stretchr/testify/require"
"gitea.d-ma.be/mathias/tapir/internal/adapters/store"
"git.d-ma.be/mathias/tapir/internal/adapters/store"
)
func TestFormatListColumnsAndOrdering(t *testing.T) {
+1 -1
View File
@@ -9,7 +9,7 @@ import (
"strings"
"text/tabwriter"
"gitea.d-ma.be/mathias/tapir/internal/adapters/store"
"git.d-ma.be/mathias/tapir/internal/adapters/store"
)
// runList prints the user's stored summaries as a table, most recent first.
+132 -12
View File
@@ -20,15 +20,18 @@ import (
"net/http"
"os"
"os/signal"
"sync"
"time"
"gitea.d-ma.be/mathias/tapir/internal/adapters/secrets"
"gitea.d-ma.be/mathias/tapir/internal/adapters/store"
"gitea.d-ma.be/mathias/tapir/internal/auth"
"gitea.d-ma.be/mathias/tapir/internal/config"
"gitea.d-ma.be/mathias/tapir/internal/runner"
"gitea.d-ma.be/mathias/tapir/internal/web"
"gitea.d-ma.be/mathias/tapir/internal/web/oidc"
"git.d-ma.be/mathias/tapir/internal/adapters/secrets"
"git.d-ma.be/mathias/tapir/internal/adapters/store"
"git.d-ma.be/mathias/tapir/internal/adapters/youtube"
"git.d-ma.be/mathias/tapir/internal/auth"
"git.d-ma.be/mathias/tapir/internal/config"
"git.d-ma.be/mathias/tapir/internal/metrics"
"git.d-ma.be/mathias/tapir/internal/runner"
"git.d-ma.be/mathias/tapir/internal/web"
"git.d-ma.be/mathias/tapir/internal/web/oidc"
)
func main() {
@@ -53,6 +56,8 @@ func main() {
err = cmdRun(ctx, log)
case "serve":
err = cmdServe(ctx, log)
case "report":
err = runReport(ctx, os.Args[2:])
default:
usage()
os.Exit(2)
@@ -73,6 +78,7 @@ usage:
tapir serve run the web UI (read summaries, record watch/skip/save)
tapir list [-limit N] list stored summaries, recent first
tapir show <video-id> show one summary in full
tapir report Stage-0 usage gate: per-user distinct active weeks
configuration is via TAPIR_* environment variables (see .env.example).
`)
@@ -125,10 +131,18 @@ func cmdRun(ctx context.Context, log *slog.Logger) error {
if engine == nil {
return fmt.Errorf("run: incomplete summarization config (gateway, youtube credentials, secrets file)")
}
r := runner.New(engine.Source, st, engine, cfg.UserID, log)
// Process-wide caption-fetch rate gate (ADR-014 item 2): the batch path shares
// the same per-egress-IP limiter as the web click-path.
youtube.SetFetchRate(cfg.FetchRate)
r := runner.New(engine.Source, st, engine, cfg.UserID, log,
runner.WithBackoff(cfg.FetchBackoff),
runner.WithAutoWindow(cfg.AutoSummarizeWindow),
runner.WithCaptionMemory(cfg.ChannelCaptionlessThreshold, cfg.ChannelCaptionlessWindow))
log.Info("starting run", "user", cfg.UserID, "model", cfg.SummarizerModel,
"gateway", cfg.GatewayURL, "poll_interval", cfg.PollInterval)
"gateway", cfg.GatewayURL, "poll_interval", cfg.PollInterval, "fetch_backoff", cfg.FetchBackoff,
"fetch_rate", cfg.FetchRate, "auto_window", cfg.AutoSummarizeWindow)
return r.Loop(ctx, cfg.PollInterval)
}
@@ -152,6 +166,11 @@ func cmdServe(ctx context.Context, log *slog.Logger) error {
}
defer st.Close()
// Process-wide caption-fetch rate gate (ADR-014 item 2): the web click-path
// and the scheduled-discovery runners share one per-egress-IP limiter so they
// cannot collectively trip 429s. Must be set before either path fetches.
youtube.SetFetchRate(cfg.FetchRate)
// Auth seam (handlers depend on web.Auth only). With Dex configured
// (TAPIR_OIDC_ISSUER set) serve uses real OIDC login — any Dex subject may
// authenticate, then registers a tapir user (ADR-012); otherwise it falls
@@ -177,7 +196,20 @@ func cmdServe(ctx context.Context, log *slog.Logger) error {
// The file-backed SecretStore is shared by the connect flow (writes tokens)
// and account management (deletes them on disconnect / delete-account).
secretStore := secrets.NewFileStore(cfg.SecretsFile)
app := &web.App{Store: st, Identity: st, Auth: authn, Secrets: secretStore, Log: log}
app := &web.App{Store: st, Identity: st, Auth: authn, Secrets: secretStore, Log: log, RecencyWindow: cfg.AutoSummarizeWindow}
// Per-video deeper-dive chat over the STORED transcript (ADR-027). Enabled
// whenever a gateway is configured — it needs no YouTube credentials because it
// never fetches. Guarded so a typed-nil never lands in the interface field
// (which would mount the routes over a nil backend).
if c := buildChat(cfg); c != nil {
app.Chat = c
log.Info("web chat enabled (stored-transcript only)", "models", chatModels(cfg))
}
// User onboarding is handled by the IdP (Authentik invite flow), not Tapir —
// the Dex local-password provisioning path was removed (ADR-019). An
// authenticated subject with no Tapir user is routed to /register.
// Web-initiated YouTube connect (ADR-006). Mounted only when the OAuth client
// credentials are present; the refresh token persists through the SecretStore
@@ -189,7 +221,10 @@ func cmdServe(ctx context.Context, log *slog.Logger) error {
ClientSecret: cfg.YTClientSecret,
RedirectURL: cfg.YTConnectRedirectURL,
}, secretStore, st, log)
log.Info("web youtube connect enabled", "redirect", cfg.YTConnectRedirectURL)
// Paste-a-URL (Feature 2): same YouTube credentials, per-user adapter built
// per request. Mounting the /paste route keys off app.Fetcher being set.
app.Fetcher = videoFetcher{cfg: cfg, secrets: secretStore}
log.Info("web youtube connect + paste enabled", "redirect", cfg.YTConnectRedirectURL)
}
// Immediate summarization for the web "Summarize" button. When the engine can
@@ -207,18 +242,103 @@ func cmdServe(ctx context.Context, log *slog.Logger) error {
log.Info("web summarization is queue-only (incomplete engine config)")
}
// In-process scheduled discovery (ADR-018): when enabled, a background
// goroutine runs a discovery pass for ALL users on TAPIR_DISCOVERY_INTERVAL,
// reusing the per-user runner.Runner. Cancelled by the same ctx as the server.
//
// SINGLE-REPLICA ASSUMPTION (load-bearing): this loop lives in the web process.
// Running serve at >1 replica would make every replica fetch every user in
// parallel — duplicate work and self-inflicted 429s. replicas: 1 is required in
// the deployment manifest; scaling up needs a CronJob or leader election first.
if cfg.DiscoveryInterval > 0 {
log.Info("scheduled discovery enabled", "interval", cfg.DiscoveryInterval, "fetch_rate", cfg.FetchRate)
log.Warn("scheduled discovery assumes a SINGLE replica — running serve at >1 replica double-runs discovery (ADR-018)")
rawRunUser := func(ctx context.Context, userID string) (runner.Stats, error) {
r, err := buildUserRunner(cfg, st, secretStore, userID, log)
if err != nil {
return runner.Stats{}, err
}
return r.RunOnce(ctx)
}
// One lock shared by the scheduler and connect-triggered passes (#6) so
// they never fetch concurrently — the single-fetcher invariant (ADR-018).
runUser := serialize(&sync.Mutex{}, rawRunUser)
// Onboarding burst (Feature 1, refined by ADR-028): after the connect-triggered
// discovery pass, summarize up to OnboardSummarizeCount of the user's newest
// LIKELY-GOOD unsummarized videos so a fresh account gets a strong first
// session. Selection avoids known-junk (Shorts/over-long/livestream VODs via
// the persisted duration); the burst leads its chain with the stronger onboard
// model. Hard cap; explicit, so it bypasses the recency window — but every
// fetch still goes through globalFetchGate. No-op when disabled (count 0) or
// queue-only (no processor).
//
// burstProcessor leads with the stronger model (ADR-028); it collapses onto the
// shared Processor when the onboard model is empty/equal-to-primary or the
// engine config is incomplete.
burstProcessor := app.Processor
if burstEngine, berr := buildBurstProcessor(cfg, st); berr != nil {
return berr
} else if burstEngine != nil {
burstProcessor = &engineProcessor{engine: burstEngine, store: st}
log.Info("onboarding burst uses a stronger model", "onboard_model", cfg.OnboardSummarizerModel)
}
onboard := func(ctx context.Context, userID string) {
if cfg.OnboardSummarizeCount <= 0 || burstProcessor == nil {
return
}
ids, err := st.OnboardBurstVideoIDs(ctx, userID, cfg.OnboardSummarizeCount, cfg.MinVideoSeconds, cfg.OnboardMaxVideoSeconds)
if err != nil {
log.Warn("onboarding: list burst candidates", "user", userID, "err", err)
return
}
for _, id := range ids {
if err := burstProcessor.ProcessVideo(ctx, userID, id); err != nil {
log.Warn("onboarding: summarize", "user", userID, "video", id, "err", err)
}
}
if len(ids) > 0 {
log.Info("onboarding burst complete", "user", userID, "summarized", len(ids), "cap", cfg.OnboardSummarizeCount)
}
}
if app.Connect != nil {
app.Connect.Discovery = &discoveryTrigger{ctx: ctx, run: runUser, onboard: onboard, log: log}
log.Info("connect-triggered discovery enabled", "onboard_cap", cfg.OnboardSummarizeCount)
}
go runScheduler(ctx, cfg.DiscoveryInterval, st, runUser, log)
} else {
log.Info("scheduled discovery disabled (TAPIR_DISCOVERY_INTERVAL unset or 0)")
}
srv := &http.Server{
Addr: cfg.HTTPAddr,
Handler: app.Router(),
Handler: metrics.HTTPMiddleware(app.Router()),
ReadHeaderTimeout: 10 * time.Second,
}
// Prometheus /metrics on a SEPARATE port (ADR-030) — never on the public app
// mux, so a scrape is in-cluster only. Empty TAPIR_METRICS_ADDR disables it.
var metricsSrv *http.Server
if cfg.MetricsAddr != "" {
mmux := http.NewServeMux()
mmux.Handle("GET /metrics", metrics.Handler())
metricsSrv = &http.Server{Addr: cfg.MetricsAddr, Handler: mmux, ReadHeaderTimeout: 10 * time.Second}
go func() {
log.Info("serving metrics", "addr", cfg.MetricsAddr)
if err := metricsSrv.ListenAndServe(); err != nil && !errors.Is(err, http.ErrServerClosed) {
log.Error("metrics server", "err", err)
}
}()
}
// Graceful shutdown on signal: stop accepting, drain in-flight requests.
go func() {
<-ctx.Done()
shutdownCtx, cancel := context.WithTimeout(context.Background(), 10*time.Second)
defer cancel()
_ = srv.Shutdown(shutdownCtx)
if metricsSrv != nil {
_ = metricsSrv.Shutdown(shutdownCtx)
}
}()
log.Info("serving web ui", "addr", cfg.HTTPAddr, "user", cfg.UserID)
+212 -19
View File
@@ -3,17 +3,40 @@ package main
import (
"context"
"fmt"
"strings"
"gitea.d-ma.be/mathias/tapir/internal/adapters/llm"
"gitea.d-ma.be/mathias/tapir/internal/adapters/secrets"
"gitea.d-ma.be/mathias/tapir/internal/adapters/store"
"gitea.d-ma.be/mathias/tapir/internal/adapters/summarizer"
"gitea.d-ma.be/mathias/tapir/internal/adapters/youtube"
"gitea.d-ma.be/mathias/tapir/internal/config"
"gitea.d-ma.be/mathias/tapir/internal/domain"
"gitea.d-ma.be/mathias/tapir/internal/usecase"
"git.d-ma.be/mathias/tapir/internal/adapters/chat"
"git.d-ma.be/mathias/tapir/internal/adapters/llm"
"git.d-ma.be/mathias/tapir/internal/adapters/secrets"
"git.d-ma.be/mathias/tapir/internal/adapters/store"
"git.d-ma.be/mathias/tapir/internal/adapters/summarizer"
"git.d-ma.be/mathias/tapir/internal/adapters/youtube"
"git.d-ma.be/mathias/tapir/internal/config"
"git.d-ma.be/mathias/tapir/internal/domain"
"git.d-ma.be/mathias/tapir/internal/metrics"
"git.d-ma.be/mathias/tapir/internal/ports"
"git.d-ma.be/mathias/tapir/internal/usecase"
"git.d-ma.be/mathias/tapir/internal/web"
)
// videoFetcher adapts the YouTube adapter to web.VideoFetcher for the paste flow
// (Feature 2). It builds a per-user adapter bound to that user's token ref and
// resolves a single video's metadata via the Data API — ungated; only the later
// transcript fetch goes through globalFetchGate.
type videoFetcher struct {
cfg config.Config
secrets ports.SecretStore
}
func (f videoFetcher) FetchVideo(ctx context.Context, userID, videoID string) (domain.Video, error) {
a := youtube.New(youtube.Config{
ClientID: f.cfg.YTClientID,
ClientSecret: f.cfg.YTClientSecret,
TokenSecretRef: web.YouTubeTokenRef(userID),
}, f.secrets)
return a.VideoByID(ctx, userID, videoID)
}
// buildProcessor wires the summarization engine — YouTube source (captions-first),
// AI-router summarizer, store sink — shared by `tapir run` and the web
// "Summarize now" path so the wiring lives in one place. It returns (nil, nil) —
@@ -22,28 +45,172 @@ import (
// queue-only fallback: the web UI keeps working (the button just queues) and
// `tapir run` reports the gap via its own ValidateForRun. Missing engine config
// is never an error here.
// buildSummarizer wires the summarization endpoint chain (ADR-022) shared by the
// web "Summarize now" path and the scheduler's per-user runners. The chain is:
// primary (local, fast) → local fallback → cloud fallback (worst case). Each
// endpoint reaches the same LiteLLM gateway with a different model alias — the
// gateway fronts both llama-swap and berget — so a fallback is just a different
// alias, not a second client config. Empty model entries are skipped, so a
// client deployment can set the cloud fallback empty to keep content local.
func buildSummarizer(cfg config.Config) *summarizer.Summarizer {
mk := summarizerEndpoint(cfg)
eps := []summarizer.Endpoint{mk(cfg.SummarizerModel)}
if cfg.FallbackModel != "" && cfg.FallbackModel != cfg.SummarizerModel {
eps = append(eps, mk(cfg.FallbackModel))
}
if cfg.CloudFallbackModel != "" && cfg.CloudFallbackModel != cfg.SummarizerModel {
eps = append(eps, mk(cfg.CloudFallbackModel))
}
return summarizer.NewChain(eps, cfg.MaxTranscriptChars)
}
// summarizerEndpoint returns a constructor for a chain endpoint over the one
// LiteLLM gateway, varying only the model alias (the gateway fronts both
// llama-swap and berget). Shared by the standard and burst chains.
func summarizerEndpoint(cfg config.Config) func(model string) summarizer.Endpoint {
return func(model string) summarizer.Endpoint {
return summarizer.Endpoint{
Client: llm.New(cfg.GatewayURL, cfg.GatewayKey, model, cfg.SummarizerTimeout, llm.WithMaxTokens(cfg.SummaryMaxTokens), llm.WithUsageHook(metrics.RecordTokens)),
Provider: providerOf(model),
Model: model,
}
}
}
// burstChainModels is the ordered, deduped model list for the onboarding burst
// (ADR-028): the stronger onboard model leads, then the standard ADR-022 chain
// (primary → local fallback → cloud) follows as resilience. Empty entries are
// dropped and duplicates collapsed, so the NDA lever (empty cloud fallback) keeps
// the burst chain fully local exactly as the standard chain does.
func burstChainModels(cfg config.Config) []string {
var models []string
add := func(m string) {
if m == "" {
return
}
for _, e := range models {
if e == m {
return
}
}
models = append(models, m)
}
add(cfg.OnboardSummarizerModel)
add(cfg.SummarizerModel)
add(cfg.FallbackModel)
add(cfg.CloudFallbackModel)
return models
}
// buildBurstSummarizer builds the onboarding-burst summarizer chain (ADR-028):
// the onboard model first, then the standard chain as fallback, deduped.
func buildBurstSummarizer(cfg config.Config) *summarizer.Summarizer {
mk := summarizerEndpoint(cfg)
var eps []summarizer.Endpoint
for _, m := range burstChainModels(cfg) {
eps = append(eps, mk(m))
}
return summarizer.NewChain(eps, cfg.MaxTranscriptChars)
}
// chatModels is the ordered, local-first set of models offered in the chat
// switcher (ADR-027), reusing the ADR-022 chain: primary → local fallback →
// cloud. Empty entries are dropped and duplicates collapsed, so a client/NDA
// deployment that sets the cloud fallback empty simply has no external model in
// the switcher — the same local-first lever the summarizer honours.
func chatModels(cfg config.Config) []string {
var models []string
add := func(m string) {
if m == "" {
return
}
for _, e := range models {
if e == m {
return
}
}
models = append(models, m)
}
add(cfg.SummarizerModel)
add(cfg.FallbackModel)
add(cfg.CloudFallbackModel)
return models
}
// buildChat wires the per-video chat service (ADR-027): a Completer factory over
// the SAME LiteLLM gateway the summarizer uses (a different alias per model, not a
// second client config) and the same transcript-truncation budget. It returns nil
// when no gateway is configured — chat is simply not mounted, the read path is
// unaffected. It deliberately takes NO YouTube source: chat is stored-only.
func buildChat(cfg config.Config) *chat.Service {
if cfg.GatewayURL == "" {
return nil
}
models := chatModels(cfg)
if len(models) == 0 {
return nil
}
newClient := func(model string) chat.Completer {
return llm.New(cfg.GatewayURL, cfg.GatewayKey, model, cfg.SummarizerTimeout, llm.WithMaxTokens(cfg.SummaryMaxTokens), llm.WithUsageHook(metrics.RecordTokens))
}
return chat.New(newClient, models, cfg.MaxTranscriptChars)
}
// providerOf maps a model alias to the domain AIProvider recorded on summaries.
// A "berget/" alias is an external provider; everything else is the local stack.
func providerOf(model string) string {
if strings.HasPrefix(model, "berget/") {
return "berget"
}
return "local"
}
func buildProcessor(cfg config.Config, st *store.Store) (*usecase.Engine, error) {
if cfg.GatewayURL == "" || cfg.YTClientID == "" || cfg.YTClientSecret == "" || cfg.SecretsFile == "" {
return nil, nil
}
secretStore := secrets.NewFileStore(cfg.SecretsFile)
src := youtube.New(youtube.Config{
src := newYouTubeSource(cfg, secrets.NewFileStore(cfg.SecretsFile))
sum := buildSummarizer(cfg)
// The store is both the summary sink and the shared transcript cache (ADR-021):
// the engine reads stored transcripts before any caption fetch and writes
// resolved ones back, so re-analysis never re-touches YouTube.
eng := usecase.NewEngine(src, sum, st)
eng.Transcripts = st
return eng, nil
}
// newYouTubeSource builds the captions-first VideoSource shared by the standard
// and burst processors — same per-process YouTube credentials and ADR-023 Shorts
// filter; only the summarizer chain differs between them.
func newYouTubeSource(cfg config.Config, secretStore ports.SecretStore) ports.VideoSource {
return youtube.New(youtube.Config{
ClientID: cfg.YTClientID,
ClientSecret: cfg.YTClientSecret,
TokenSecretRef: cfg.YTTokenRef,
PreferredLanguages: []string{"en"},
MinVideoSeconds: cfg.MinVideoSeconds,
}, secretStore)
}
// Local Primary only; no BYO fallback for the demo (fallback nil).
primary := summarizer.Endpoint{
Client: llm.New(cfg.GatewayURL, cfg.GatewayKey, cfg.SummarizerModel, cfg.SummarizerTimeout),
Provider: "local",
Model: cfg.SummarizerModel,
// buildBurstProcessor wires a processor whose summarizer leads with the stronger
// onboard model (ADR-028), used only by the connect-time burst over the SAME
// store / transcript cache / sink — a wiring choice; the engine and ports are
// unchanged. Returns (nil, nil) — the collapse lever — when the onboard model is
// empty or equal to the primary (the burst then reuses the shared processor), or
// when the engine config is incomplete (queue-only, same as buildProcessor).
func buildBurstProcessor(cfg config.Config, st *store.Store) (*usecase.Engine, error) {
if cfg.OnboardSummarizerModel == "" || cfg.OnboardSummarizerModel == cfg.SummarizerModel {
return nil, nil
}
sum := summarizer.New(primary, nil)
return usecase.NewEngine(src, sum, st), nil
if cfg.GatewayURL == "" || cfg.YTClientID == "" || cfg.YTClientSecret == "" || cfg.SecretsFile == "" {
return nil, nil
}
src := newYouTubeSource(cfg, secrets.NewFileStore(cfg.SecretsFile))
eng := usecase.NewEngine(src, buildBurstSummarizer(cfg), st)
eng.Transcripts = st
return eng, nil
}
// engineProcessor adapts the engine (which works in terms of a domain.Video) to
@@ -58,6 +225,11 @@ type engineProcessor struct {
}
func (p *engineProcessor) ProcessVideo(ctx context.Context, userID, videoID string) error {
// This is the user-initiated (foreground) path — a click on "Summarize",
// "Try now", or a pasted URL. Mark the context so the caption gate gives it
// priority over the background sweep (ADR-026, Pillar A).
ctx = youtube.ForegroundContext(ctx)
row, err := p.store.GetVideoRow(ctx, userID, videoID)
if err != nil {
return fmt.Errorf("load video %q: %w", videoID, err)
@@ -77,7 +249,28 @@ func (p *engineProcessor) ProcessVideo(ctx context.Context, userID, videoID stri
if err != nil {
return fmt.Errorf("process video %q: %w", videoID, err)
}
if res.Summary != nil {
// Record the outcome so the status endpoint can show honest state (ADR-025):
// a 429'd or caption-less click used to leave transcript_status unset, so the
// poll silently reverted to the "Summarize" button. Mirror the runner: stamp
// rate_limited / none / fetched. A rate-limited video keeps its requested flag
// so the background sweep retries it; none and fetched are terminal here.
switch {
case res.Skipped && res.TranscriptSource == string(domain.SourceRateLimited):
if err := p.store.SetTranscriptStatus(ctx, userID, videoID, "rate_limited"); err != nil {
return fmt.Errorf("set rate_limited status %q: %w", videoID, err)
}
case res.Skipped:
if err := p.store.SetTranscriptStatus(ctx, userID, videoID, "none"); err != nil {
return fmt.Errorf("set none status %q: %w", videoID, err)
}
if err := p.store.ClearSummarizeRequested(ctx, userID, videoID); err != nil {
return fmt.Errorf("clear summarize flag %q: %w", videoID, err)
}
case res.Summary != nil:
if err := p.store.SetTranscriptStatus(ctx, userID, videoID, "fetched"); err != nil {
return fmt.Errorf("set fetched status %q: %w", videoID, err)
}
if err := p.store.ClearSummarizeRequested(ctx, userID, videoID); err != nil {
return fmt.Errorf("clear summarize flag %q: %w", videoID, err)
}
+63 -1
View File
@@ -3,7 +3,7 @@ package main
import (
"testing"
"gitea.d-ma.be/mathias/tapir/internal/config"
"git.d-ma.be/mathias/tapir/internal/config"
)
// TestBuildProcessorNilOnIncompleteConfig asserts the queue-only fallback: when a
@@ -45,3 +45,65 @@ func TestBuildProcessorNilOnIncompleteConfig(t *testing.T) {
})
}
}
// TestBurstChainModelsLeadsWithOnboardModel: the onboarding burst chain (ADR-028)
// leads with the stronger onboard model, then falls back through the standard
// ADR-022 chain (primary -> local fallback -> cloud), deduped.
func TestBurstChainModelsLeadsWithOnboardModel(t *testing.T) {
got := burstChainModels(config.Config{
OnboardSummarizerModel: "iguana/gemma4-26b",
SummarizerModel: "koala/phi4-mini",
FallbackModel: "iguana/gemma4-26b", // also the onboard model -> dedup
CloudFallbackModel: "berget/mistral-small",
})
want := []string{"iguana/gemma4-26b", "koala/phi4-mini", "berget/mistral-small"}
if len(got) != len(want) {
t.Fatalf("burstChainModels = %v, want %v", got, want)
}
for i := range want {
if got[i] != want[i] {
t.Fatalf("burstChainModels = %v, want %v", got, want)
}
}
}
// TestBurstChainModelsCloudAbsentWhenDisabled: the NDA lever holds for the burst
// too — empty cloud fallback keeps the burst chain fully local.
func TestBurstChainModelsCloudAbsentWhenDisabled(t *testing.T) {
got := burstChainModels(config.Config{
OnboardSummarizerModel: "iguana/gemma4-26b",
SummarizerModel: "koala/phi4-mini",
CloudFallbackModel: "",
})
for _, m := range got {
if m == "" || m == "berget/mistral-small" {
t.Fatalf("cloud model leaked into burst chain: %v", got)
}
}
}
// TestBuildBurstProcessorNilWhenCollapsed: an empty or primary-equal onboard model
// collapses the burst onto the shared processor (buildBurstProcessor returns nil).
func TestBuildBurstProcessorNilWhenCollapsed(t *testing.T) {
base := config.Config{
GatewayURL: "http://gw/v1",
YTClientID: "id",
YTClientSecret: "secret",
SecretsFile: "/tmp/secrets.json",
SummarizerModel: "koala/phi4-mini",
}
t.Run("empty onboard model", func(t *testing.T) {
base.OnboardSummarizerModel = ""
eng, err := buildBurstProcessor(base, nil)
if err != nil || eng != nil {
t.Fatalf("buildBurstProcessor = (%v, %v), want (nil, nil)", eng, err)
}
})
t.Run("onboard model equals primary", func(t *testing.T) {
base.OnboardSummarizerModel = "koala/phi4-mini"
eng, err := buildBurstProcessor(base, nil)
if err != nil || eng != nil {
t.Fatalf("buildBurstProcessor = (%v, %v), want (nil, nil)", eng, err)
}
})
}
+97
View File
@@ -0,0 +1,97 @@
package main
import (
"context"
"fmt"
"io"
"os"
"text/tabwriter"
"time"
"git.d-ma.be/mathias/tapir/internal/adapters/store"
)
// gateThreshold is the Stage-0 gate (VISION/ADR-016): usage in >= 2 distinct
// weeks. The gate passes when any user reaches it.
const gateThreshold = 2
// defaultGateStart is the date Stage-0 return-usage tracking begins: the morning
// the pilot was actually unblocked and summaries started flowing (2026-06-11).
// Activity before this — testing, the period the pilot was stuck on zero — is
// noise and must not count toward the gate. Override with TAPIR_USAGE_GATE_START
// (YYYY-MM-DD). The gate measures whether users RETURN once it genuinely works.
const defaultGateStart = "2026-06-11"
// gateStart resolves the baseline date from TAPIR_USAGE_GATE_START or the default,
// parsed as a UTC calendar day.
func gateStart() (time.Time, error) {
v := os.Getenv("TAPIR_USAGE_GATE_START")
if v == "" {
v = defaultGateStart
}
t, err := time.Parse("2006-01-02", v)
if err != nil {
return time.Time{}, fmt.Errorf("TAPIR_USAGE_GATE_START=%q: want YYYY-MM-DD: %w", v, err)
}
return t, nil
}
// runReport prints the Stage-0 usage gate: per-user distinct active weeks (reads
// UNION acts) and the pass/fail verdict. Read-only, cross-user — needs only
// TAPIR_DB_DSN (not TAPIR_USER_ID; the report enumerates all users itself).
func runReport(ctx context.Context, _ []string) error {
dsn := os.Getenv(envDSN)
if dsn == "" {
return fmt.Errorf("%s is required", envDSN)
}
s, err := store.New(ctx, dsn)
if err != nil {
return err
}
defer s.Close()
since, err := gateStart()
if err != nil {
return err
}
rows, err := s.ActiveWeeks(ctx, since)
if err != nil {
return err
}
return formatReport(os.Stdout, rows, since)
}
// formatReport renders the per-user week counts and the gate verdict. Pure: no DB,
// no env — so the layout and verdict logic are unit-testable without Postgres.
func formatReport(w io.Writer, rows []store.UserActiveWeeks, since time.Time) error {
if _, err := fmt.Fprintf(w, "Counting usage since %s (Stage-0 gate baseline)\n\n", since.Format("2006-01-02")); err != nil {
return err
}
if len(rows) == 0 {
_, err := fmt.Fprintln(w, "no users yet")
return err
}
tw := tabwriter.NewWriter(w, 0, 4, 2, ' ', 0)
_, _ = fmt.Fprintln(tw, "USER\tNAME\tACTIVE_WEEKS\tGATE")
passed := false
for _, r := range rows {
gate := "-"
if r.ActiveWeeks >= gateThreshold {
gate = "PASS"
passed = true
}
_, _ = fmt.Fprintf(tw, "%s\t%s\t%d\t%s\n", r.UserID, orDash(r.DisplayName), r.ActiveWeeks, gate)
}
if err := tw.Flush(); err != nil {
return err
}
verdict := fmt.Sprintf("\nGate (usage in >= %d distinct weeks): NOT YET MET\n", gateThreshold)
if passed {
verdict = fmt.Sprintf("\nGate (usage in >= %d distinct weeks): PASSED\n", gateThreshold)
}
_, err := fmt.Fprint(w, verdict)
return err
}
+49
View File
@@ -0,0 +1,49 @@
package main
import (
"strings"
"testing"
"time"
"github.com/stretchr/testify/require"
"git.d-ma.be/mathias/tapir/internal/adapters/store"
)
var testSince = time.Date(2026, 6, 11, 0, 0, 0, 0, time.UTC)
func TestFormatReportColumnsAndGatePass(t *testing.T) {
rows := []store.UserActiveWeeks{
{UserID: "user-a", DisplayName: "Ada", ActiveWeeks: 3},
{UserID: "user-b", DisplayName: "", ActiveWeeks: 1},
}
var b strings.Builder
require.NoError(t, formatReport(&b, rows, testSince))
out := b.String()
require.Contains(t, out, "since 2026-06-11", "report states the gate baseline date")
require.Contains(t, out, "USER")
require.Contains(t, out, "ACTIVE_WEEKS")
require.Contains(t, out, "Ada")
require.Contains(t, out, "PASSED", "a user at >= 2 weeks passes the gate")
// The >=2 user is marked PASS; the 1-week user is not.
require.Contains(t, lineContaining(t, out, "user-a"), "PASS")
require.NotContains(t, lineContaining(t, out, "user-b"), "PASS")
require.Contains(t, lineContaining(t, out, "user-b"), "-", "no display name falls back to dash")
}
func TestFormatReportGateNotMet(t *testing.T) {
rows := []store.UserActiveWeeks{{UserID: "user-a", ActiveWeeks: 1}}
var b strings.Builder
require.NoError(t, formatReport(&b, rows, testSince))
require.Contains(t, b.String(), "NOT YET MET", "no user at >= 2 weeks fails the gate")
}
func TestFormatReportEmpty(t *testing.T) {
var b strings.Builder
require.NoError(t, formatReport(&b, nil, testSince))
require.Contains(t, b.String(), "no users yet")
}
+206
View File
@@ -0,0 +1,206 @@
package main
import (
"context"
"fmt"
"log/slog"
"time"
"git.d-ma.be/mathias/tapir/internal/adapters/store"
"git.d-ma.be/mathias/tapir/internal/adapters/youtube"
"git.d-ma.be/mathias/tapir/internal/config"
"git.d-ma.be/mathias/tapir/internal/ports"
"git.d-ma.be/mathias/tapir/internal/runner"
"git.d-ma.be/mathias/tapir/internal/usecase"
"git.d-ma.be/mathias/tapir/internal/web"
)
// buildUserRunner constructs a runner.Runner for one user, reusing the same
// engine wiring as buildProcessor but bound to that user's own YouTube refresh
// token (web.YouTubeTokenRef(userID)) — the Stage-1 per-tenant ref, not the
// Stage-0 single ref. It returns an error (not nil) when the global config can't
// support live summarization (gateway, YouTube client creds, secrets file), so
// the scheduler can skip that user gracefully. A user who simply hasn't connected
// YouTube yet builds fine here; their token ref fails to resolve at RunOnce time,
// surfacing as a per-user error the scheduler logs and skips.
func buildUserRunner(cfg config.Config, st *store.Store, secretStore ports.SecretStore, userID string, log *slog.Logger) (*runner.Runner, error) {
if cfg.GatewayURL == "" || cfg.YTClientID == "" || cfg.YTClientSecret == "" || cfg.SecretsFile == "" {
return nil, fmt.Errorf("buildUserRunner: incomplete summarization config (gateway, youtube credentials, secrets file)")
}
src := youtube.New(youtube.Config{
ClientID: cfg.YTClientID,
ClientSecret: cfg.YTClientSecret,
TokenSecretRef: web.YouTubeTokenRef(userID),
PreferredLanguages: []string{"en"},
MinVideoSeconds: cfg.MinVideoSeconds,
}, secretStore)
engine := usecase.NewEngine(src, buildSummarizer(cfg), st)
// Share the transcript cache (ADR-021) on the scheduler path too — without
// this every scheduled pass re-fetches transcripts it already had, burning the
// scarce per-IP caption budget (ADR-014) and starving other users. The web
// "Summarize now" path already sets this; the scheduler omitting it was a bug.
engine.Transcripts = st
return runner.New(src, st, engine, userID, log,
runner.WithBackoff(cfg.FetchBackoff),
runner.WithAutoWindow(cfg.AutoSummarizeWindow),
runner.WithCaptionMemory(cfg.ChannelCaptionlessThreshold, cfg.ChannelCaptionlessWindow)), nil
}
// userLister enumerates every registered user and reports a user's video
// connections. *store.Store satisfies it via ListAllUsers + ConnectionsForUser.
// A small local interface keeps the scheduler testable with a fake.
type userLister interface {
ListAllUsers(ctx context.Context) ([]store.UserIdentity, error)
ConnectionsForUser(ctx context.Context, userID string) ([]store.Connection, error)
}
// runDiscoveryPass runs one discovery pass for every user. runUser performs a
// single user's pass (production: build a runner and RunOnce). Per-user failures
// — including a buildUserRunner error or a RunOnce error — are logged and skipped
// so one bad user, channel, or video never aborts the others (ADR-018 failure
// isolation). Returns the stats summed across users.
func runDiscoveryPass(
ctx context.Context,
pass int,
lister userLister,
runUser func(context.Context, string) (runner.Stats, error),
log *slog.Logger,
) runner.Stats {
users, err := lister.ListAllUsers(ctx)
if err != nil {
log.Error("scheduler: list users failed", "err", err)
return runner.Stats{}
}
// Keep only users with a video connection. A pass for a connectionless user
// (e.g. a stale Dex-era orphan identity) only tries to resolve a token that
// was never minted, logging a spurious "ref not found" every tick. Filtering
// here — BEFORE rotation — also keeps fairness honest: rotation is over the
// users that actually consume the caption budget, so a dead identity can't eat
// a rotation slot and skew the lead share.
var connected []store.UserIdentity
for _, u := range users {
if ctx.Err() != nil {
return runner.Stats{} // shutting down
}
conns, err := lister.ConnectionsForUser(ctx, u.UserID)
if err != nil {
log.Warn("scheduler: list connections failed", "user", u.UserID, "err", err)
continue
}
if len(conns) == 0 {
log.Debug("scheduler: skipping user with no video connections", "user", u.UserID)
continue
}
connected = append(connected, u)
}
// Rotate who goes first each pass. Caption fetches share one per-egress-IP
// rate budget (ADR-014); whoever runs first each pass spends the pre-throttle
// window, so a FIXED order permanently starves whoever is last (a new pilot
// user got 0 fetches for 12h while the first-listed user got all of them).
// Rotation over the connected set gives each real user the lead in turn.
connected = rotateUsers(connected, pass)
log.Info("scheduler: starting discovery pass", "users", len(connected))
var total runner.Stats
for _, u := range connected {
if ctx.Err() != nil {
break // shutting down: stop enumerating
}
stats, err := runUser(ctx, u.UserID)
total = sumStats(total, stats)
if err != nil {
log.Warn("scheduler: user discovery pass had errors", "user", u.UserID, "err", err)
}
}
log.Info("scheduler: pass complete",
"candidates", total.Candidates, "summarized", total.Summarized,
"skipped_seen", total.SkippedSeen, "skipped_no_text", total.SkippedNoText,
"skipped_manual", total.SkippedManual, "skipped_too_old", total.SkippedTooOld,
"skipped_rate_limited", total.SkippedRateLimited,
"skipped_no_caption_channel", total.SkippedNoCaptionChannel,
"channel_unavailable", total.ChannelUnavailable, "errors", total.Errors)
return total
}
// runScheduler runs a discovery pass on startup, then once every interval until
// ctx is cancelled (pod SIGTERM exits the loop cleanly). A non-positive interval
// disables scheduling entirely (no startup pass) so dev/tests never auto-fetch.
// It reuses the existing runner.Runner via runUser — the only new behaviour over
// runner.Loop is iterating all users per tick (ADR-018).
func runScheduler(
ctx context.Context,
interval time.Duration,
lister userLister,
runUser func(context.Context, string) (runner.Stats, error),
log *slog.Logger,
) {
if interval <= 0 {
return // disabled
}
// Derive the rotation offset from wall-clock, NOT an in-memory counter. A
// counter reset to 0 on every pod restart always hands the lead to the
// first-listed user — so frequent deploys re-starve whoever is last (exactly
// what happened to the first pilot user during a deploy-heavy session). A
// time-based offset advances with real time and is identical across restarts,
// so the lead rotates fairly regardless of how often the pod bounces.
runPass := func() {
pass := int(time.Now().Unix() / int64(interval/time.Second))
runDiscoveryPass(ctx, pass, lister, runUser, log)
}
runPass()
ticker := time.NewTicker(interval)
defer ticker.Stop()
for {
select {
case <-ctx.Done():
return
case <-ticker.C:
runPass()
}
}
}
// rotateUsers left-rotates users by pass positions so a different user leads each
// pass. With n users, user i leads on every pass where pass ≡ i (mod n). A pass
// offset that is negative or exceeds n is normalised. Order within the rotation
// is otherwise preserved, so the set of users run is unchanged — only who is
// first (and thus wins the scarce caption-fetch budget) rotates.
func rotateUsers(users []store.UserIdentity, pass int) []store.UserIdentity {
n := len(users)
if n <= 1 {
return users
}
off := ((pass % n) + n) % n
if off == 0 {
return users
}
out := make([]store.UserIdentity, 0, n)
out = append(out, users[off:]...)
out = append(out, users[:off]...)
return out
}
// sumStats adds two passes' stats field-wise, so runDiscoveryPass can report a
// per-tick aggregate across all users.
func sumStats(a, b runner.Stats) runner.Stats {
return runner.Stats{
Candidates: a.Candidates + b.Candidates,
Summarized: a.Summarized + b.Summarized,
SkippedSeen: a.SkippedSeen + b.SkippedSeen,
SkippedNoText: a.SkippedNoText + b.SkippedNoText,
SkippedManual: a.SkippedManual + b.SkippedManual,
SkippedTooOld: a.SkippedTooOld + b.SkippedTooOld,
SkippedRateLimited: a.SkippedRateLimited + b.SkippedRateLimited,
SkippedNoCaptionChannel: a.SkippedNoCaptionChannel + b.SkippedNoCaptionChannel,
ChannelUnavailable: a.ChannelUnavailable + b.ChannelUnavailable,
Errors: a.Errors + b.Errors,
}
}
+216
View File
@@ -0,0 +1,216 @@
package main
import (
"context"
"errors"
"io"
"log/slog"
"sync"
"testing"
"time"
"github.com/stretchr/testify/require"
"git.d-ma.be/mathias/tapir/internal/adapters/store"
"git.d-ma.be/mathias/tapir/internal/runner"
)
func quietLog() *slog.Logger {
return slog.New(slog.NewTextHandler(io.Discard, nil))
}
// fakeLister returns a fixed user set (or an error) for the scheduler under test.
type fakeLister struct {
users []store.UserIdentity
err error
noConn map[string]bool // users that have NOT connected a video source
}
func (f fakeLister) ListAllUsers(context.Context) ([]store.UserIdentity, error) {
return f.users, f.err
}
// ConnectionsForUser reports a single youtube connection for every user except
// those in noConn, which return zero — the connection-less case the scheduler
// must skip instead of running (and failing to resolve a token for).
func (f fakeLister) ConnectionsForUser(_ context.Context, userID string) ([]store.Connection, error) {
if f.noConn[userID] {
return nil, nil
}
return []store.Connection{{Provider: "youtube"}}, nil
}
// countingRunUser records how many passes each user got, optionally failing for
// specific users, under a mutex so it is safe across the scheduler goroutine.
type countingRunUser struct {
mu sync.Mutex
calls map[string]int
order []string // userIDs in the order they were run, across all passes
failFor map[string]bool
}
func newCountingRunUser(failFor ...string) *countingRunUser {
c := &countingRunUser{calls: map[string]int{}, failFor: map[string]bool{}}
for _, u := range failFor {
c.failFor[u] = true
}
return c
}
func (c *countingRunUser) run(_ context.Context, userID string) (runner.Stats, error) {
c.mu.Lock()
defer c.mu.Unlock()
c.calls[userID]++
c.order = append(c.order, userID)
if c.failFor[userID] {
return runner.Stats{Errors: 1}, errors.New("boom")
}
return runner.Stats{Summarized: 1}, nil
}
func (c *countingRunUser) runOrder() []string {
c.mu.Lock()
defer c.mu.Unlock()
return append([]string(nil), c.order...)
}
func (c *countingRunUser) count(userID string) int {
c.mu.Lock()
defer c.mu.Unlock()
return c.calls[userID]
}
func (c *countingRunUser) total() int {
c.mu.Lock()
defer c.mu.Unlock()
n := 0
for _, v := range c.calls {
n += v
}
return n
}
func usersN(ids ...string) []store.UserIdentity {
out := make([]store.UserIdentity, len(ids))
for i, id := range ids {
out[i] = store.UserIdentity{UserID: id, DexSubject: "dex|" + id}
}
return out
}
func TestDiscoveryPassRunsEveryUserOnce(t *testing.T) {
lister := fakeLister{users: usersN("a", "b", "c")}
rc := newCountingRunUser()
stats := runDiscoveryPass(context.Background(), 0, lister, rc.run, quietLog())
require.Equal(t, 1, rc.count("a"))
require.Equal(t, 1, rc.count("b"))
require.Equal(t, 1, rc.count("c"))
require.Equal(t, 3, stats.Summarized, "stats are summed across users")
}
// Caption fetches share one per-IP budget; a fixed user order starves whoever is
// last. Each pass must rotate which user leads so the lead slot is shared.
func TestDiscoveryPassRotatesLeadUser(t *testing.T) {
lister := fakeLister{users: usersN("a", "b", "c")}
rc := newCountingRunUser()
runDiscoveryPass(context.Background(), 0, lister, rc.run, quietLog())
runDiscoveryPass(context.Background(), 1, lister, rc.run, quietLog())
runDiscoveryPass(context.Background(), 2, lister, rc.run, quietLog())
require.Equal(t, []string{"a", "b", "c", "b", "c", "a", "c", "a", "b"}, rc.runOrder(),
"each pass left-rotates the user order so every user leads in turn")
// Fairness: over a full rotation cycle every user ran the same number of times.
require.Equal(t, 3, rc.count("a"))
require.Equal(t, 3, rc.count("b"))
require.Equal(t, 3, rc.count("c"))
}
// A connectionless orphan must not consume a rotation slot: rotation is over the
// connected users only, so two real users alternate the lead 50/50 even with a
// dead identity listed between them.
func TestDiscoveryPassRotationIgnoresConnectionlessUsers(t *testing.T) {
lister := fakeLister{users: usersN("a", "orphan", "c"), noConn: map[string]bool{"orphan": true}}
rc := newCountingRunUser()
runDiscoveryPass(context.Background(), 0, lister, rc.run, quietLog())
runDiscoveryPass(context.Background(), 1, lister, rc.run, quietLog())
require.Equal(t, []string{"a", "c", "c", "a"}, rc.runOrder(),
"only connected users rotate; the orphan never runs and never holds a slot")
require.Equal(t, 0, rc.count("orphan"))
}
func TestDiscoveryPassSkipsUsersWithoutConnections(t *testing.T) {
// b never connected a video source (e.g. a stale Dex-era orphan identity).
// It must be skipped silently — not run and logged as a token error every pass.
lister := fakeLister{users: usersN("a", "b", "c"), noConn: map[string]bool{"b": true}}
rc := newCountingRunUser()
stats := runDiscoveryPass(context.Background(), 0, lister, rc.run, quietLog())
require.Equal(t, 1, rc.count("a"))
require.Equal(t, 0, rc.count("b"), "a user with no connection must be skipped, not run")
require.Equal(t, 1, rc.count("c"))
require.Equal(t, 2, stats.Summarized, "only connected users contribute")
require.Equal(t, 0, stats.Errors, "skipping is silent — no spurious error stat")
}
func TestDiscoveryPassOneUserFailureDoesNotStopOthers(t *testing.T) {
lister := fakeLister{users: usersN("a", "b", "c")}
rc := newCountingRunUser("b") // user b's pass errors
stats := runDiscoveryPass(context.Background(), 0, lister, rc.run, quietLog())
require.Equal(t, 1, rc.count("a"))
require.Equal(t, 1, rc.count("b"))
require.Equal(t, 1, rc.count("c"), "a failing user must not abort the rest")
require.Equal(t, 2, stats.Summarized) // a + c
require.Equal(t, 1, stats.Errors) // b
}
func TestDiscoveryPassListerErrorIsContained(t *testing.T) {
lister := fakeLister{err: errors.New("db down")}
rc := newCountingRunUser()
stats := runDiscoveryPass(context.Background(), 0, lister, rc.run, quietLog())
require.Equal(t, 0, rc.total(), "no users enumerated → no passes")
require.Equal(t, runner.Stats{}, stats)
}
func TestSchedulerIntervalZeroDisablesEntirely(t *testing.T) {
lister := fakeLister{users: usersN("a", "b")}
rc := newCountingRunUser()
runScheduler(context.Background(), 0, lister, rc.run, quietLog())
require.Equal(t, 0, rc.total(), "interval 0 must not run even a startup pass")
}
func TestSchedulerRunsStartupPassThenStopsOnCancel(t *testing.T) {
lister := fakeLister{users: usersN("a", "b", "c")}
rc := newCountingRunUser()
ctx, cancel := context.WithCancel(context.Background())
done := make(chan struct{})
go func() {
// A long interval so only the startup pass runs before we cancel.
runScheduler(ctx, time.Hour, lister, rc.run, quietLog())
close(done)
}()
// The startup pass is synchronous at the top of runScheduler; once all three
// users have a pass it has completed and the loop is parked on the ticker.
require.Eventually(t, func() bool { return rc.total() == 3 }, time.Second, 5*time.Millisecond)
cancel()
select {
case <-done:
case <-time.After(time.Second):
t.Fatal("scheduler did not exit after ctx cancel")
}
require.Equal(t, 3, rc.total(), "no extra passes fired between startup and cancel")
}
+1 -1
View File
@@ -8,7 +8,7 @@ import (
"os"
"strings"
"gitea.d-ma.be/mathias/tapir/internal/adapters/store"
"git.d-ma.be/mathias/tapir/internal/adapters/store"
)
// runShow prints the full summary for one video: text, highlights, takeaways,
+188 -3
View File
@@ -45,7 +45,7 @@ adapter behind an interface (Clean Architecture ports & adapters).
```mermaid
graph TB
subgraph tapir["Tapir (Go)"]
http["HTTP server<br/>OAuth callbacks +<br/>user-facing API"]
http["tapir serve<br/>(HTMX+Templ web surface:<br/>read summaries, connect,<br/>account, summarize)"]
watcher["Watcher<br/>detects new videos<br/>(WebSub + poll)"]
engine["Summarization engine<br/>(use-case core)"]
resolver["Transcript resolver<br/>(captions-first)"]
@@ -95,6 +95,189 @@ two codebases (ADR-003).
---
## Web surface — `tapir serve` (Stage 1, ADR-011 → ADR-012)
A later transport added over the **unchanged** engine/ports/sinks core (ADR-003): `tapir serve`
is an HTMX+Templ reader/writer (`internal/web`) over the existing `store`. It added no business
logic to the engine — it reads the store and, for one action, kicks the existing engine. ADR-011
shipped it single-user; ADR-012 opened multi-user with DB-enforced (RLS) isolation.
```mermaid
graph TB
browser["Browser<br/>(Dex-authenticated user)"]
subgraph web["internal/web (tapir serve)"]
oidc["oidc<br/>Dex OIDC session<br/>(authenticate-only)"]
gate["registration gate<br/>new subject -> /register"]
pages["summary list + detail<br/>(read) + actions"]
connect["/oauth/youtube/callback<br/>per-user token connect"]
account["account<br/>(disconnect, delete)"]
summarize["Summarize button<br/>-> background goroutine"]
end
store[("store<br/>(Postgres, RLS per user)")]
engine["Summarization engine<br/>(unchanged core)"]
secrets["SecretStore<br/>(per-user token refs)"]
browser --> oidc
oidc --> gate
gate --> pages
pages --> store
connect --> secrets
connect --> store
account --> store
account --> secrets
summarize -->|background| engine
summarize -->|HTMX status poll| store
engine --> store
```
- **Dex OIDC session layer** (`internal/web/oidc`) — **authenticate-only** (ADR-012). It proves
*who*; authorization/isolation is the DB's job (RLS), not the session's.
- **Registration gate** — a Dex subject with no `users` row is routed to `/register`, which
creates the `users` row + the `user_identities` mapping (migration 004). Returning subjects
pass straight through.
- **Web-initiated YouTube connect** — `/oauth/youtube/connect``/oauth/youtube/callback`
persists a **per-user** refresh-token ref (`youtube/<userID>/refresh_token`) via `SecretStore`
and a `video_connections` row (ADR-006, migration 005). Distinct from the CLI `tapir auth`.
- **Account management** — `/account` offers disconnect and **delete account**. Delete removes
only Tapir-side state (cascade across the user's tables + secret refs); the shared Dex identity
is left intact (ADR-013).
- **Immediate summarization** — the web "Summarize" button (`POST /v/{id}/summarize`) fires the
engine in a **background goroutine** inside `serve`; the page HTMX-polls `/v/{id}/status`,
showing a Charmbracelet spinner while in-flight (and an honest "queued/waiting" state under
rate-limiting — ADR-014).
- **Summarization mode** — `users.auto_summarize` (migration 006). Default is **true** for new
users (migration 011, ADR-018); all existing rows were back-filled via migration 012. Auto:
new videos **published within the recency window** (`TAPIR_AUTO_SUMMARIZE_WINDOW`, default ~7d,
ADR-020) are summarized automatically; older videos are discovered and listed but wait for an
explicit "Summarize". Manual: new videos appear unsummarized; the button sets
`videos.summarize_requested`, which the next `tapir run` processes and clears. A manual request
bypasses the recency bound. Both the click path and the batch `tapir run` drive the same
unchanged engine.
- **List surface (ADR-020)** — the list reads `ListVideos` ordered summarized-first, then
`published_at DESC NULLS LAST`. The web layer collapses the noise so summaries are not buried:
un-summarized videos older than the recency window fold into one "Show N older videos"
disclosure, and caption-less videos collapse to a single count line. Copy surfaces scarcity
honestly (queue counts, gradual-fill note) — it never implies the feed is fuller than it is.
The engine, ports, and sink adapters are **untouched** by all of the above — the web surface only
reads the store and triggers the existing engine. Adding it changed wiring, not the core (ADR-003).
---
## In-process scheduler (ADR-018)
`cmdServe` launches a background goroutine when `TAPIR_DISCOVERY_INTERVAL > 0`. On each tick
it calls `store.ListAllUsers` (un-RLS'd admin query), builds a per-user `runner.Runner`, and
calls `RunOnce` for each registered user in sequence.
```mermaid
sequenceDiagram
participant S as Scheduler goroutine
participant DB as Postgres (RLS)
participant YT as YouTube timedtext
participant LLM as LiteLLM gateway
loop every TAPIR_DISCOVERY_INTERVAL
S->>DB: ListAllUsers() [un-RLS'd]
loop per user
S->>DB: GetAutoSummarize(userID)
S->>YT: ListSubscriptions + NewVideos
Note over S,DB: auto: skip videos published before<br/>TAPIR_AUTO_SUMMARIZE_WINDOW (ADR-020);<br/>older ones listed, await manual request
Note over S,YT: WaitFetchGate(ctx) throttles<br/>all fetches to TAPIR_FETCH_RATE
alt transcript available
S->>LLM: Summarize
S->>DB: Deliver(summary)
else 429
S->>DB: SetTranscriptStatus(rate_limited)
end
end
end
```
**Single-replica constraint (load-bearing).** The scheduler lives in the web process;
`replicas: 1` in the k3s deployment manifest is not cosmetic — running `tapir serve` at
>1 replica makes every replica run the full discovery loop, causing every registered user
to be fetched in parallel from the same egress IP (429s + duplicate work). Do not scale
`serve` past 1 replica without first moving discovery to a k8s CronJob or adding leader
election. The process logs a `Warn` at startup when scheduled discovery is enabled as a
reminder.
---
## Process-wide timedtext rate gate
**`internal/adapters/youtube/gate.go`** (ADR-014 item 2): a single `rate.Limiter`
(`golang.org/x/time/rate`) shared across **all** Adapter instances. Every `httpDo` call for
a caption fetch passes through `WaitFetchGate(ctx)` before hitting YouTube. This serialises
the scheduler loop AND the web click-path through the same per-egress-IP budget. Configured
via `TAPIR_FETCH_RATE` (Go duration, default `2s`). Setting it to `0` disables the gate
(dev/tests only).
This is the precondition that makes scheduled auto-summarize safe: without the gate, a
multi-user scheduler pass could fire many concurrent timedtext requests from the same IP
within seconds, triggering 429s for all users.
### Two-path summarisation model
Both paths share `globalFetchGate` — rate limiting is **respected in both**, not routed around.
| Path | Trigger | Order | Rationale |
|------|---------|-------|-----------|
| **Foreground** | User clicks "Summarize" on any non-summarized card (`POST /v/{id}/retry-now` for rate-limited; `POST /v/{id}/summarize` for pending) | Single chosen video | On-demand value: user picks a specific video to read now — bypasses the recency bound |
| **Background batch** | Scheduled discovery pass every `TAPIR_DISCOVERY_INTERVAL` | **Newest-first across all channels** (see below), **bounded to the recency window** (ADR-020) | Onboarding prioritisation within bounded load: recent videos auto-fill; the older back-catalogue stays on-demand |
The rationale for both paths is **onboarding prioritisation under an honest, bounded load** — a
new user gets summaries of their most recent videos automatically, while the older back-catalogue
is listed but summarised only on demand, so it never re-drives the shared rate gate every cycle.
### Newest-first batch ordering (ADR-018)
Within each scheduled pass, `RunOnce` uses a three-phase structure:
1. **Discover + persist**: walk all channels, `UpsertVideo` every candidate (so it appears in
the list), apply pre-filters (seen/manual/backoff/**recency**), collect surviving candidates.
The recency pre-filter (ADR-020) drops auto-mode videos published before
`now - TAPIR_AUTO_SUMMARIZE_WINDOW` unless they are explicitly requested; an undated video is
never aged out. They remain persisted/listed — only auto-summarisation is skipped.
2. **Sort**: order candidates `published_at DESC, NULLS LAST, discovery_pos ASC`. Videos with
no publish date (schema 001: nullable) sort after all dated content. The sort is in-memory
(`slices.SortStableFunc`) — at current scale this is fine.
3. **Process**: feed candidates to the engine in sorted order through `globalFetchGate`.
Before (per-channel inline): `[chanA-old, chanA-mid, chanB-new, chanB-null]`
After (newest-first): `[chanB-new, chanA-mid, chanA-old, chanB-null]`
The set of *processed* videos now also excludes auto-mode back-catalogue beyond the recency
window (those stay listed, summarised on demand); within the processed set, only order changes.
### Connect-time onboarding burst (ADR-018 → ADR-028)
On a successful YouTube connect, `ConnectHandler` enqueues a connect-triggered discovery pass;
the `discoveryTrigger` runs that pass and then fires the **onboarding burst** — a third entry path
that summarises up to `TAPIR_ONBOARD_SUMMARIZE_COUNT` (default 3, hard-capped) of the new user's
videos so the first session is not empty. The burst still flows through `globalFetchGate` (it is
not a throughput change); ADR-028 sharpened *which* videos and *which model*:
- **Selection** is `OnboardBurstVideoIDs`, not pure newest-first. It keeps newest-first order but
excludes a video whose **known** duration is outside `[TAPIR_MIN_VIDEO_SECONDS,
TAPIR_ONBOARD_MAX_VIDEO_SECONDS]` (drops Shorts and multi-hour livestream VODs). An unknown
(NULL) duration is degrade-open — kept, but ranked after known-good rows. The connect-triggered
discovery pass runs *before* the burst, and ADR-023's `videos.list` enrichment now **persists**
`duration_s` (instead of discarding it after the Shorts filter), so a fresh user's candidates
carry a duration in time for selection.
- **Model**: the burst runs through a dedicated summarizer chain led by
`TAPIR_ONBOARD_SUMMARIZER_MODEL` (default `iguana/gemma4-26b`, the stronger local model), with
the standard ADR-022 chain following as fallback. This is a wiring choice — a second
`engineProcessor` over the same store / transcript cache / sink; the engine and ports are
unchanged. Empty / equal-to-primary collapses it back onto the shared processor.
`has-captions` is deliberately **not** a selection signal — it is only knowable after a gate fetch
(or a ~0-probability cache hit at pilot scale), so the burst can avoid known-junk but cannot
promise captions. Cached-transcript-first selection was investigated and rejected (ADR-028:
~3% cross-user overlap).
---
## Sequence — core use case: new video summarized
```mermaid
@@ -191,6 +374,8 @@ Gherkin features in `docs/use-cases/`).
- Audio-download + speech-to-text resolver (ADR-007) — would be an additional `VideoSource`
fallback path, drawn when built.
- Multi-tenant isolation primitives (per-tenant Postgres role, NetworkPolicy, tenant label)
— activate at Stage 1 (ADR-002); single-user Stage 0 doesn't exercise them.
- Per-user isolation is **live, not deferred**: Postgres RLS `FORCE`d on every user-owned table
(ADR-012, migration 003), realising ADR-002's per-tenant intent at the DB layer. The coarser
multi-tenant primitives (per-namespace NetworkPolicy, Kyverno, tenant label) remain a
Stage-2 hardening item, not exercised yet.
- Public SaaS surface (sign-up, billing) — Future C, not built (ADR-008).
+131 -36
View File
@@ -10,12 +10,15 @@ only opaque references to them; the secret material lives in ESO/1Password (ADR-
## Design decisions baked into this model
- **Per-user isolation, not a shared global video table.** The earlier draft proposed a
global `videos`/`transcripts` table deduped across tenants. Rejected for Future B: it
reintroduces exactly the cross-domain coupling the homelab architecture review is
removing, and at 15 users the cost of occasionally re-summarizing the same video is
trivial compared to the isolation it would cost. Each user's data is self-contained.
(Revisit only if Future C makes GPU/transcription cost dominate — a new ADR, not a default.)
- **Per-user isolation for everything except transcripts.** The earlier draft proposed a
global `videos`/`transcripts` table deduped across tenants. **Videos** stay per-user and
RLS-scoped — a shared video table reintroduces exactly the cross-domain coupling the homelab
architecture review is removing. **Transcripts**, however, are now shared (ADR-021): keyed by
`(provider, provider_video_id)`, no `user_id`, **not** RLS-scoped. The cost avoided there is
not LLM re-summarization but a rate-gated, reputation-risky caption fetch (ADR-010/014), which
is paid per re-fetch regardless of user count — so persisting public caption content once and
sharing it strictly beats the coupling it removes. Everything else each user owns is
self-contained; `rls_test.go` proves transcripts is the single exception.
- **Secrets by reference only.** Tables hold a `secret_ref` (opaque string/UUID resolved via
the `SecretStore` port), never tokens or keys.
- **The brain sink is just a delivery target.** No brain-specific tables. Whether a summary
@@ -23,27 +26,39 @@ only opaque references to them; the secret material lives in ESO/1Password (ADR-
## Entities
Solid entities below are **persisted today** (migrations 001013). `AI_CREDENTIAL` and
`SUBSCRIPTION` are **planned, not yet a table** — kept in the model for intent; see the notes.
```mermaid
erDiagram
USER ||--|| USER_IDENTITY : "logs in via (Dex subject)"
USER ||--o{ VIDEO_CONNECTION : has
USER ||--o{ AI_CREDENTIAL : has
VIDEO_CONNECTION ||--o{ SUBSCRIPTION : exposes
USER ||--o{ SUMMARY_ACTION : records
USER ||--o{ AI_CREDENTIAL : "has (planned)"
VIDEO_CONNECTION ||--o{ SUBSCRIPTION : "exposes (planned)"
SUBSCRIPTION ||--o{ VIDEO : "produces (per user)"
VIDEO ||--o| TRANSCRIPT : "has at most one"
VIDEO }o--o| TRANSCRIPT : "shares one by (provider, provider_video_id) — not FK (ADR-021)"
VIDEO ||--o| SUMMARY : "has at most one"
SUMMARY ||--o{ SINK_DELIVERY : "delivered via"
USER ||--o{ CHANNEL_ERROR : "reports unavailable channels"
USER {
uuid id PK
text display_name
bool auto_summarize "default true for new users (migration 011, ADR-018)"
timestamptz created_at
}
USER_IDENTITY {
text dex_subject PK
uuid user_id FK "UNIQUE -> USER, ON DELETE CASCADE; NOT RLS-enabled"
timestamptz created_at
}
VIDEO_CONNECTION {
uuid id PK
uuid user_id FK
uuid user_id FK "-> USER, ON DELETE CASCADE"
text provider "youtube | vimeo"
text provider_account
text token_secret_ref "-> SecretStore, never the token"
text provider_account "nullable"
text token_ref "-> SecretStore, never the token"
text status "active | revoked | error"
timestamptz connected_at
}
@@ -66,28 +81,31 @@ erDiagram
}
VIDEO {
uuid id PK
uuid user_id FK
uuid subscription_id FK
uuid user_id FK "-> USER, ON DELETE CASCADE"
uuid subscription_id "nullable; no FK at Stage 0"
text provider
text provider_video_id
text title
int duration_s
timestamptz published_at
text url
bool summarize_requested "default false -> manual-mode queue flag (migration 006)"
timestamptz seen_at
text transcript_status "none|rate_limited|fetched (migration 007)"
timestamptz rate_limited_at "backoff clock for 429 retries (migration 007)"
}
TRANSCRIPT {
uuid video_id PK_FK
uuid user_id FK
text provider PK "part of shared key (ADR-021)"
text provider_video_id PK "part of shared key — the cross-user dedup key"
text source "captions | none"
text language
text content "null when source = none"
timestamptz resolved_at
timestamptz fetched_at
}
SUMMARY {
uuid id PK
uuid user_id FK
uuid video_id FK
uuid video_id "no FK to videos; (user_id, video_id) UNIQUE is the dedup key"
text summary
jsonb highlights
jsonb takeaways
@@ -98,41 +116,117 @@ erDiagram
}
SINK_DELIVERY {
uuid id PK
uuid summary_id FK
uuid summary_id FK "-> SUMMARY, ON DELETE CASCADE; ownership derived via this FK"
text sink "store | brain"
text status "pending | delivered | error"
text detail "nullable; error message etc"
timestamptz updated_at
}
SUMMARY_ACTION {
uuid id PK
uuid user_id FK "-> USER"
text video_id "TEXT, not FK (mirrors summaries' standalone key)"
text action "watched | skipped | saved"
timestamptz acted_at
}
CHANNEL_ERROR {
uuid user_id FK
text channel_id
text channel_name
timestamptz first_seen
timestamptz last_seen
}
```
`SUMMARY_ACTION` has `UNIQUE (user_id, video_id, action)`; `VIDEO_CONNECTION` has
`UNIQUE (user_id, provider)` (one connection per provider — reconnect upserts in place).
RLS (`ENABLE` + `FORCE`) is on **every solid user-owned table above**`users`, `videos`,
`transcripts`, `summaries`, `summary_actions`, `video_connections`, `channel_errors`.
`sink_deliveries` is RLS'd via an `EXISTS` on its parent summary; `user_identities` is
intentionally **not** RLS'd (auth plumbing). See the *Isolation invariant* section for the
mechanism.
## Notes per entity
- **USER** — at Stage 0 there is exactly one row. At Stage 1, identity comes via Dex; this
table holds the Tapir-side profile keyed to the Dex subject.
- **VIDEO_CONNECTION** — a connected YouTube/Vimeo account. `token_secret_ref` resolves to
the OAuth refresh token via `SecretStore`. Revocation flips `status`, doesn't delete history.
- **AI_CREDENTIAL** — optional, per provider, per user (ADR-004's Fallback). Absent for users
who only use the local stack. One row per provider max.
- **SUBSCRIPTION** — a watched channel. `websub_expires` tracks the YouTube push lease so the
watcher knows when to re-subscribe; null for poll-based (Vimeo).
- **USER** — one row per registered user (Stage 1, ADR-012; no longer single-row). The Tapir-side
profile; the Dex identity is held separately in `USER_IDENTITY`, not on this row. `auto_summarize`
(migration 006) is the per-user mode flag: `TRUE` = auto-summarize new videos **published within
the recency window** (`TAPIR_AUTO_SUMMARIZE_WINDOW`, default ~7d, ADR-020); older videos are
listed but summarised on demand. Default is **true** for new users (migration 011, ADR-018);
existing rows were back-filled via migration 012 with RLS bypass.
- **USER_IDENTITY** (migration 004) — the `dex_subject → user_id` map. `dex_subject` is the PK,
`user_id` a `UNIQUE` FK to `users` with `ON DELETE CASCADE`. This is the bridge resolved at login
*before* a `user_id` is known, so it is **deliberately not RLS-enabled** (it holds no user data;
RLS here would deadlock the lookup that yields the id used for scoping). Account deletion cascades
the mapping away (ADR-013).
- **VIDEO_CONNECTION** (migration 005) — a connected YouTube/Vimeo account. `token_ref` resolves to
the OAuth refresh token via `SecretStore` (per-user scheme `youtube/<userID>/refresh_token`).
`UNIQUE (user_id, provider)`: one connection per provider, reconnect upserts. Revocation/disconnect
flips `status`, doesn't delete history. FORCE RLS'd.
- **AI_CREDENTIAL** — *planned, no table yet.* Optional, per provider, per user (ADR-004's Fallback).
BYO keys are currently resolved via `SecretStore` refs without a dedicated table; this entity is
modelled for when per-credential metadata is needed.
- **SUBSCRIPTION** — *planned, no table yet.* A watched channel; `websub_expires` would track the
YouTube push lease. At Stage 0/1 `videos.subscription_id` is a nullable column with **no FK** (the
subscriptions table is not part of the shipped store-sink slice — migration 001).
- **VIDEO** — one row per (user, video) — note `user_id`, reflecting the per-user-isolation
decision. The same video seen by two users is two rows. `seen_at` is when Tapir detected it.
- **TRANSCRIPT** — at most one per video. `source = none` records "checked, no usable
transcript" so the watcher doesn't reprocess (ADR-007). `content` null in that case.
`summarize_requested` (migration 006) is the manual-mode queue flag: the web "Summarize" button
sets it `TRUE`; the next `tapir run` picks it up, summarizes, and clears it back to `FALSE`.
`transcript_status` and `rate_limited_at` (migration 007) track caption-fetch outcomes for
rate-limit backoff: `NULL` = not attempted; `rate_limited` = 429 seen, skip until
`NOW() - rate_limited_at > TAPIR_FETCH_BACKOFF`; `fetched` = resolved; `none` = no transcript.
- **TRANSCRIPT** — shared public caption content, one row per `(provider, provider_video_id)`,
**not** RLS-scoped and carrying no `user_id` (ADR-021). Two users who watch the same video
share the one row; the summarize path reads it before any caption fetch, so re-analysis never
re-touches YouTube (ADR-010/014). `source = none` records "checked, no usable transcript" so
no one reprocesses (ADR-007); `content` null in that case. A transient 429 is never stored
here — it stays a per-user retry via `VIDEO.transcript_status`.
- **SUMMARY** — at most one per video. `fallback_used` + `ai_provider`/`ai_model` make the
"is local good enough?" question queryable (the Stage 0 quality signal). `highlights`/
`takeaways` as jsonb to stay schema-flexible while the output format settles.
- **SINK_DELIVERY** — one row per (summary, sink) attempt. This is where "also sent to brain"
lives — no brain tables, just a delivery row with `sink = brain`. Sinks fail independently;
a failed brain delivery doesn't fail the store delivery.
a failed brain delivery doesn't fail the store delivery. No own `user_id`; RLS ownership is
derived from the parent summary via `EXISTS` (migration 003).
- **SUMMARY_ACTION** (migration 002) — records the maintainer's act on a summary (watch / skip /
save) — the column that makes the Stage-0 headline metric ("acts on ≥1 summary") queryable
(ui-spec.md §5, ADR-011). `video_id` is `TEXT` and **not** FK-constrained, mirroring summaries'
standalone `(user_id, video_id)` key. `UNIQUE (user_id, video_id, action)`. FORCE RLS'd.
- **CHANNEL_ERRORS** (migration 013) — channels that returned HTTP 404 (deleted or private) on
the most recent discovery pass. Upserted per scheduler pass (`last_seen` refreshed each run);
surfaced on the account page as a warning. Cascades on user deletion. Primary key is
`(user_id, channel_id)`. FORCE RLS'd.
- **LOGIN_EVENTS** (migration 010) — throttled one-row-per-(user, date) login stamp. Used by the
Stage-0 gate query (VISION §Stage 0: "returned and used in ≥2 distinct weeks").
## Isolation invariant (Stage 1+)
## Isolation invariant (Stage 1+) — LIVE
Every user-owned table carries `user_id`. At Stage 1, this is enforced at the DB layer via a
per-tenant Postgres role + row grants (architecture review SC7), not only in application code.
At Stage 0 (single user) the column exists but the enforcement is dormant. The isolation test
in VISION Stage 2 asserts user A cannot read user B's rows.
Every user-owned table carries `user_id`, and isolation is **enforced at the DB layer**, not
only in application code. ADR-011 shipped this surface single-user (one allowlisted subject,
enforcement dormant); **ADR-012 opened Stage 1 and turned enforcement on in the same slice.**
Enforcement is **Postgres Row-Level Security** (migration `003_rls.up.sql`):
- RLS is `ENABLE`d **and** `FORCE`d on every user-owned table — `users`, `videos`,
`transcripts`, `summaries`, `summary_actions`, `video_connections`, `channel_errors`.
`FORCE` is load-bearing: the app connects as the table **owner** (`tapir` role), and owners
bypass RLS unless forced.
- Each policy keys off the per-request GUC `tapir.current_user_id`, set transaction-locally by
the store's `withUser` helper via `set_config('tapir.current_user_id', $1, true)` — it
auto-resets on commit/rollback, so it never leaks across a pooled connection.
- `current_setting('tapir.current_user_id', true)` uses `missing_ok = true`: an **unset** GUC
yields `NULL`, the predicate matches no rows, and access **denies by default**.
- `sink_deliveries` has no `user_id`; its policy derives ownership from the parent summary via
`EXISTS (SELECT 1 FROM summaries …)`.
- `user_identities` (the Dex-subject → user_id map) is **deliberately not RLS-enabled** — it is
auth plumbing read *before* a user_id is known; putting RLS there would deadlock. It holds no
user data.
The Stage-2 isolation bar is **pulled forward, not deferred**: `internal/adapters/store/rls_test.go`
runs two users against a non-superuser, non-`BYPASSRLS` role and asserts user A reads/writes zero
of user B's rows across every table. It ships green with the multi-user features (ADR-012); no
multi-user feature merges ahead of it passing.
## Job / processing state
@@ -144,7 +238,8 @@ queue, doesn't replace it). Deferred until there's a reason.
## Explicitly out of scope (Future C)
- Global cross-tenant video/transcript dedup (rejected above).
- Global cross-tenant *video* dedup (rejected above). Note: cross-tenant *transcript* sharing
is now in scope and shipped (ADR-021); only the videos half stays per-user.
- Sharding / per-tenant physical databases.
- Soft-delete + full audit trail on connections/credentials (a Stage 2 hardening item; add
via ADR when Stage 2 work starts).
+60 -2
View File
@@ -2,7 +2,7 @@
The concrete endpoints, conventions, and identifiers Tapir depends on, so an independent
session doesn't have to rediscover them. **Verify anything marked "confirm" before relying on
it** — endpoints and aliases drift, and this file is a snapshot (2026-06-02), not a live source.
it** — endpoints and aliases drift, and this file is a snapshot (2026-06-06), not a live source.
## Local AI (the Primary in `llm.Router`)
@@ -27,6 +27,30 @@ it** — endpoints and aliases drift, and this file is a snapshot (2026-06-02),
`iguana/deepseek-r1-14b`) is preferred for summary quality if its latency/output is acceptable.
The `max_tokens` fix below means thinking models no longer return empty content, so they are now
viable choices, not blocked ones. Do not assume a coder alias is right for prose.
- **Summarizer fallback chain (ADR-022).** The primary alias is the *first* of an ordered chain;
on failure or unparseable output the summarizer advances to the next model. All reached through
the same gateway by alias.
- `TAPIR_FALLBACK_MODEL` — local fallback. **Default `iguana/gemma4-26b`** — on iguana, NOT
koala, so the fallback does not compete with koala's other GPU loads (and runs from a different
egress IP). Empty disables it.
- `TAPIR_CLOUD_FALLBACK_MODEL` — worst-case EXTERNAL fallback. **Default `berget/mistral-small`.**
**Set this empty (`""`) for any client/NDA deployment** so content never leaves the local
stack — the chain then contains only local endpoints.
- `TAPIR_SUMMARY_MAX_TOKENS` — per-summary completion budget. **Default `1500`.** Small on
purpose: with the old 8192 budget, prompt + completion overflowed `phi4-mini`'s 8k window.
- `TAPIR_MAX_TRANSCRIPT_CHARS` — transcript truncation budget sent to the model. **Default
`18000`** (~fits an 8k-context model). `0` disables truncation. Prevents the context-overflow
HTTP 400 a long transcript caused on `phi4-mini`.
- **Discovery low-value filter (ADR-023).** `TAPIR_MIN_VIDEO_SECONDS`**default `60`**. At
discovery, `NewVideos` enriches candidates with one cheap `videos.list` call (quota API, NOT
the timedtext 429 path) and drops videos shorter than this plus any live/upcoming broadcast,
so the scarce caption-fetch budget isn't spent on Shorts. `0` disables the filter. The
paste-a-URL path is never filtered.
- **Per-channel caption memory (ADR-024).** `TAPIR_CHANNEL_CAPTIONLESS_THRESHOLD` — **default
`5`** consecutive no-caption results before a channel is suppressed (its videos listed but not
caption-fetched). `TAPIR_CHANNEL_CAPTIONLESS_WINDOW`**default `336h`** (14d) suppression
before one video is re-probed. `THRESHOLD=0` disables. A successful fetch resets the channel;
a 429 does not count; an explicit manual request bypasses suppression.
- **Thinking models need an explicit `max_tokens`.** qwen3 / deepseek-r1 spend the budget on
reasoning and return **empty content** if `max_tokens` is too low (or unset). The summarizer's
parser treats an empty summary as an error for exactly this reason. **Done (2026-06-02, Worker F):**
@@ -160,7 +184,7 @@ allow per-provider when a user connects one.
---
_Snapshot date 2026-06-02. Items marked **confirm** were not verified to a pinned source at
_Snapshot date 2026-06-06. Items marked **confirm** were not verified to a pinned source at
snapshot time — check brain or the live cluster before depending on them._
## Stage 1 — multi-user facts (verified 2026-06-03)
@@ -191,3 +215,37 @@ snapshot time — check brain or the live cluster before depending on them._
- `user_identities(dex_subject → user_id)` table is **intentionally NOT RLS-enabled**
(it's auth plumbing, holds no user data; data isolation is on the user-owned tables).
All data access after subject resolution goes through `withUser`.
## Scheduled discovery (ADR-018, verified 2026-06-05)
`tapir serve` runs discovery for **all users** in-process on a timer (no CronJob). Three env
knobs plus one load-bearing deployment constraint:
- `TAPIR_DISCOVERY_INTERVAL` — Go duration, e.g. `2h`. The cadence the serve process runs a
discovery pass for every registered user (run-once-on-startup, then every interval).
**Unset or `0` = disabled** (dev/tests never auto-fetch).
- `TAPIR_USAGE_GATE_START``YYYY-MM-DD`, default **`2026-06-11`** (the morning the pilot was
unblocked and summaries started flowing). `tapir report` counts return-usage (distinct active
weeks, ADR-016) only from this date, so pre-launch testing and the blocked period are excluded.
- `TAPIR_METRICS_ADDR` — listen address for the Prometheus `/metrics` endpoint (ADR-030).
**Default `:9090`** — a SEPARATE port from `TAPIR_HTTP_ADDR` so metrics are never on the public
app; scraped in-cluster only (PodMonitor). Empty disables the metrics server. Key series:
`tapir_summarize_duration_seconds{model,outcome,fallback}`, `tapir_caption_fetch_duration_seconds{outcome}`,
`tapir_chat_duration_seconds{model}`, `tapir_llm_tokens_total{model,kind}`,
`tapir_http_request_duration_seconds{method,route}`, `tapir_logins_total`.
- `TAPIR_FETCH_RATE` — Go duration, default `2s`. The **process-wide per-egress-IP caption-fetch
rate gate** (ADR-014 item 2). Every caption fetch — scheduler runners *and* the web "Summarize"
click-path — serialises through this one limiter so the pod cannot collectively trip 429s. `0`
= unlimited (dev/tests). This is the precondition that makes auto-summarize-on-a-schedule safe;
do not raise it aggressively without watching for 429s.
- `TAPIR_FETCH_BACKOFF=4h` — per-video rate-limit retry window; default `1h`. A video that
returns HTTP 429 on a caption fetch is skipped for this duration before being retried. The
scheduler checks `NOW() - rate_limited_at > TAPIR_FETCH_BACKOFF` before attempting to fetch
a video marked `transcript_status = rate_limited`. Longer values reduce 429 pressure at the
cost of slower recovery after a throttling episode.
- **SINGLE-REPLICA WARNING (load-bearing).** The scheduler lives in the web process, so
`replicas: 1` in the deployment manifest is load-bearing: running `tapir serve` at >1 replica
makes **every** replica run the discovery loop → every user fetched in parallel from the same
egress IP (429s + duplicate work). Do **not** scale `serve` past 1 replica without first moving
discovery to a k8s CronJob or adding leader election. The process logs a `Warn` at startup when
scheduled discovery is enabled, as a reminder.
+63
View File
@@ -0,0 +1,63 @@
<!DOCTYPE html><html lang="en"><head><meta charset="utf-8">
<meta name="viewport" content="width=device-width, initial-scale=1">
<title>tapir — A · TUI panels</title>
<style>
:root{--bg:#0d0d12;--panel:#14141c;--line:#2a2a3a;--mint:#0EF9B6;--pink:#F740A0;--purple:#7653fc;--cream:#F5E9D6;--dim:#6b6b86;--text:#d9d9e6}
*{box-sizing:border-box}
body{margin:0;background:var(--bg);color:var(--text);font:14px/1.5 ui-monospace,SFMono-Regular,Menlo,"Cascadia Code",monospace}
.wrap{max-width:760px;margin:0 auto;padding:18px 14px 60px}
header{display:flex;align-items:center;gap:10px;border-bottom:1px solid var(--line);padding-bottom:12px;margin-bottom:18px}
.brand{color:var(--mint);font-weight:700;letter-spacing:.5px}
.brand b{color:var(--cream)}
.tip{margin-left:auto;color:var(--dim);font-size:12px}
.note{color:var(--dim);font-size:12.5px;border-left:2px solid var(--purple);padding:6px 10px;margin:0 0 16px;background:#11111a}
/* lipgloss-style bordered panel */
.card{border:1px solid var(--line);border-radius:8px;background:var(--panel);padding:12px 14px;margin:0 0 12px;position:relative}
.card::before{content:"";position:absolute;left:0;top:10px;bottom:10px;width:3px;border-radius:3px;background:var(--purple)}
.card.ready::before{background:var(--mint)}
.title{color:var(--cream);font-size:15px;font-weight:600;margin:0 0 4px}
.meta{color:var(--dim);font-size:12px}
.chip{display:inline-block;border:1px solid var(--mint);color:var(--mint);border-radius:4px;padding:0 6px;font-size:11px;margin-left:6px}
.chip.q{border-color:var(--pink);color:var(--pink)}
.preview{color:var(--dim);margin:6px 0 0;font-size:13px}
/* expanded */
.exp{border-color:var(--purple)}
.exp .head{display:flex;justify-content:space-between;align-items:baseline}
.collapse{color:var(--mint);font-size:12px;text-decoration:none;border:1px solid var(--line);border-radius:4px;padding:1px 7px}
.sec h2{color:var(--mint);font-size:12px;text-transform:uppercase;letter-spacing:1px;margin:16px 0 6px;border-bottom:1px dashed var(--line);padding-bottom:3px}
.sec ul{margin:0;padding-left:18px}.sec li{margin:3px 0}
.body{color:var(--text)}
.dock{margin-top:16px;border-top:1px solid var(--line);padding-top:12px}
.ask{display:inline-block;background:linear-gradient(90deg,var(--purple),var(--pink));color:#fff;border:0;border-radius:6px;padding:7px 12px;font:inherit;font-size:13px;cursor:pointer}
pre.tapir{margin:0;color:var(--mint);font-size:10px;line-height:1.05}
</style></head><body><div class="wrap">
<header>
<pre class="tapir"> ▄█▓▓█▄ ∩
█▓( ◕ ◕ )▓█──┘</pre>
<span class="brand"><b>tapir</b> · watch less, know more</span>
<span class="tip">261 in queue</span>
</header>
<p class="note">Tapir fetches captions slowly on purpose, to respect YouTube's limits — new summaries land gradually.</p>
<div class="card ready"><div class="title">How the Attention Economy Rewires Your Brain</div>
<div class="meta">youtube · 2026-06-11 · 18 min <span class="chip">ready</span></div>
<div class="preview">A tour of the incentive loops behind infinite feeds and three concrete ways to claw back focus…</div></div>
<article class="card exp ready">
<div class="head"><div class="title">Postgres 18 — What's Actually New</div><a class="collapse" href="#">collapse ↑</a></div>
<div class="meta">youtube · 2026-06-10 · 42 min · phi4-mini</div>
<div class="sec"><h2>Takeaways</h2><ul>
<li>Async I/O cuts cold-cache read latency materially on NVMe.</li>
<li>Skip-scan makes more multicolumn indexes usable without rewrites.</li>
<li>Upgrade path is smooth; test the new planner stats first.</li></ul></div>
<div class="sec"><h2>Highlights</h2><ul>
<li>Async I/O subsystem (effective_io_concurrency now matters more).</li>
<li>B-tree skip scan for leading-column gaps.</li>
<li>Better partition-wise joins.</li></ul></div>
<div class="sec"><h2>Summary</h2><p class="body">Postgres 18 is an incremental but meaningful release: the headline is the new asynchronous I/O path, with skip-scan and planner improvements close behind. For most homelabs the upgrade is low-risk and worth it for the read-latency wins.</p></div>
<div class="dock"><button class="ask">Dig deeper — ask about this video →</button></div>
</article>
<div class="card"><div class="title">RAG is dead, right?? — Conference Talk</div>
<div class="meta">youtube · 2026-06-09 · 31 min <span class="chip q">queued</span></div></div>
</div></body></html>
+62
View File
@@ -0,0 +1,62 @@
<!DOCTYPE html><html lang="en"><head><meta charset="utf-8">
<meta name="viewport" content="width=device-width, initial-scale=1">
<title>tapir — B · reader + charm accents</title>
<style>
:root{--bg:#faf7f2;--card:#fff;--ink:#1c1b22;--soft:#6a6878;--line:#e7e2d8;--mint:#0bbf8c;--purple:#6a4cf0;--pink:#e0379a}
*{box-sizing:border-box}
body{margin:0;background:var(--bg);color:var(--ink);font:16px/1.6 -apple-system,BlinkMacSystemFont,"Segoe UI",Inter,sans-serif}
.mono{font-family:ui-monospace,SFMono-Regular,Menlo,monospace}
.wrap{max-width:680px;margin:0 auto;padding:20px 16px 60px}
header{display:flex;align-items:center;gap:10px;margin-bottom:18px}
.brand{font-weight:800;font-size:18px;letter-spacing:-.3px}
.brand .dot{color:var(--mint)}
.tag{color:var(--soft);font-size:13px}
.tip{margin-left:auto;color:var(--soft);font-size:12px}
.note{color:var(--soft);font-size:13px;background:#fff;border:1px solid var(--line);border-left:3px solid var(--mint);border-radius:8px;padding:8px 12px;margin:0 0 18px}
.card{background:var(--card);border:1px solid var(--line);border-radius:12px;padding:16px 18px;margin:0 0 14px;box-shadow:0 1px 2px rgba(20,18,40,.04)}
.title{font-size:18px;font-weight:700;letter-spacing:-.2px;margin:0 0 4px;line-height:1.3}
.meta{color:var(--soft);font-size:13px}
.meta .mono{font-size:12.5px}
.chip{display:inline-block;background:rgba(11,191,140,.12);color:var(--mint);border-radius:999px;padding:1px 9px;font-size:12px;font-weight:600;margin-left:6px}
.chip.q{background:rgba(224,55,154,.12);color:var(--pink)}
.preview{color:var(--soft);margin:8px 0 0}
.exp{border-color:#d9d0ee;box-shadow:0 6px 24px rgba(106,76,240,.10)}
.head{display:flex;justify-content:space-between;align-items:baseline;gap:10px}
.collapse{color:var(--purple);font-size:13px;font-weight:600;text-decoration:none;white-space:nowrap}
.sec h2{font-size:12px;text-transform:uppercase;letter-spacing:1.2px;color:var(--purple);margin:18px 0 6px}
.sec ul{margin:0;padding-left:20px}.sec li{margin:5px 0}
.body{line-height:1.7}
.dock{margin-top:18px;border-top:1px solid var(--line);padding-top:14px}
.ask{display:inline-flex;align-items:center;gap:6px;background:var(--ink);color:#fff;border:0;border-radius:999px;padding:9px 16px;font:inherit;font-size:14px;font-weight:600;cursor:pointer}
pre.tapir{margin:0;color:var(--mint);font-size:10px;line-height:1.05}
</style></head><body><div class="wrap">
<header>
<pre class="tapir"> ▄█▓▓█▄ ∩
█▓( ◕ ◕ )▓█──┘</pre>
<span class="brand">tapir<span class="dot">.</span></span><span class="tag">watch less, know more</span>
<span class="tip mono">261 in queue</span>
</header>
<p class="note">Tapir fetches captions slowly on purpose, to respect YouTube's limits — new summaries land gradually.</p>
<div class="card"><div class="title">How the Attention Economy Rewires Your Brain</div>
<div class="meta mono">youtube · 2026-06-11 · 18 min <span class="chip">ready</span></div>
<div class="preview">A tour of the incentive loops behind infinite feeds and three concrete ways to claw back focus…</div></div>
<article class="card exp">
<div class="head"><div class="title">Postgres 18 — What's Actually New</div><a class="collapse" href="#">collapse ↑</a></div>
<div class="meta mono">youtube · 2026-06-10 · 42 min · phi4-mini</div>
<div class="sec"><h2>Takeaways</h2><ul>
<li>Async I/O cuts cold-cache read latency materially on NVMe.</li>
<li>Skip-scan makes more multicolumn indexes usable without rewrites.</li>
<li>Upgrade path is smooth; test the new planner stats first.</li></ul></div>
<div class="sec"><h2>Highlights</h2><ul>
<li>Async I/O subsystem (effective_io_concurrency now matters more).</li>
<li>B-tree skip scan for leading-column gaps.</li>
<li>Better partition-wise joins.</li></ul></div>
<div class="sec"><h2>Summary</h2><p class="body">Postgres 18 is an incremental but meaningful release: the headline is the new asynchronous I/O path, with skip-scan and planner improvements close behind. For most homelabs the upgrade is low-risk and worth it for the read-latency wins.</p></div>
<div class="dock"><button class="ask">✦ Ask about this video</button></div>
</article>
<div class="card"><div class="title">RAG is dead, right?? — Conference Talk</div>
<div class="meta mono">youtube · 2026-06-09 · 31 min <span class="chip q">queued</span></div></div>
</div></body></html>
+65
View File
@@ -0,0 +1,65 @@
<!DOCTYPE html><html lang="en"><head><meta charset="utf-8">
<meta name="viewport" content="width=device-width, initial-scale=1">
<title>tapir — C · cozy terminal</title>
<style>
:root{--bg:#16131d;--card:#1e1a28;--ink:#ece7f5;--soft:#9a92b4;--line:#322b44;--mint:#2ee6b6;--purple:#9d7bff;--pink:#ff6bdb;--cream:#f3ead8}
*{box-sizing:border-box}
body{margin:0;background:radial-gradient(1200px 600px at 70% -10%,#231b33,transparent),var(--bg);color:var(--ink);font:15px/1.6 -apple-system,BlinkMacSystemFont,"Segoe UI",Inter,sans-serif}
.mono{font-family:ui-monospace,SFMono-Regular,Menlo,monospace}
.wrap{max-width:700px;margin:0 auto;padding:20px 16px 60px}
header{display:flex;align-items:center;gap:12px;margin-bottom:18px}
.brand{font-weight:800;font-size:18px}.brand .c{color:var(--mint)}
.tag{color:var(--soft);font-size:13px}
.tip{margin-left:auto;color:var(--soft);font-size:12px}
.note{color:var(--soft);font-size:13px;background:#1b1726;border:1px solid var(--line);border-radius:10px;padding:9px 12px;margin:0 0 18px}
.note b{color:var(--mint);font-weight:600}
.card{position:relative;background:var(--card);border:1px solid var(--line);border-radius:12px;padding:14px 16px 14px 18px;margin:0 0 13px;overflow:hidden}
.card::before{content:"";position:absolute;left:0;top:0;bottom:0;width:4px;background:var(--purple)}
.card.ready::before{background:linear-gradient(var(--mint),var(--purple))}
.title{font-size:16px;font-weight:700;margin:0 0 4px;color:var(--cream)}
.meta{color:var(--soft);font-size:12.5px}
.chip{display:inline-block;border-radius:999px;padding:1px 9px;font-size:11px;font-weight:600;margin-left:6px;background:rgba(46,230,182,.14);color:var(--mint)}
.chip.q{background:rgba(157,123,255,.18);color:var(--purple)}
.preview{color:var(--soft);margin:7px 0 0;font-size:14px}
.exp{border-color:#473a63;box-shadow:0 10px 40px rgba(0,0,0,.35)}
.head{display:flex;justify-content:space-between;align-items:baseline;gap:10px}
.collapse{color:var(--mint);font-size:12px;text-decoration:none;border:1px solid var(--line);border-radius:6px;padding:2px 8px;white-space:nowrap}
.sec h2{font-size:11px;text-transform:uppercase;letter-spacing:1.3px;color:var(--mint);margin:18px 0 7px;display:flex;align-items:center;gap:8px}
.sec h2::after{content:"";flex:1;height:1px;background:var(--line)}
.sec ul{margin:0;padding-left:18px}.sec li{margin:5px 0}
.body{line-height:1.7;color:var(--ink)}
.dock{margin-top:18px;border-top:1px dashed var(--line);padding-top:14px;display:flex;align-items:center;gap:10px}
.ask{display:inline-flex;align-items:center;gap:7px;background:linear-gradient(90deg,var(--purple),var(--mint));color:#10101a;border:0;border-radius:999px;padding:9px 16px;font:inherit;font-weight:700;font-size:14px;cursor:pointer}
.dockhint{color:var(--soft);font-size:12px}
pre.tapir{margin:0;color:var(--mint);font-size:11px;line-height:1.05}
</style></head><body><div class="wrap">
<header>
<pre class="tapir"> ▄█▓▓█▄ ∩
█▓( ◕ ◕ )▓█──┘</pre>
<span class="brand">tapir<span class="c">_</span></span><span class="tag">watch less, know more</span>
<span class="tip mono">261 in queue</span>
</header>
<p class="note">Tapir fetches captions slowly on purpose, to respect YouTube's limits — <b>new summaries land gradually</b>.</p>
<div class="card ready"><div class="title">How the Attention Economy Rewires Your Brain</div>
<div class="meta mono">youtube · 2026-06-11 · 18 min <span class="chip">ready</span></div>
<div class="preview">A tour of the incentive loops behind infinite feeds and three concrete ways to claw back focus…</div></div>
<article class="card exp ready">
<div class="head"><div class="title">Postgres 18 — What's Actually New</div><a class="collapse" href="#">collapse ↑</a></div>
<div class="meta mono">youtube · 2026-06-10 · 42 min · phi4-mini</div>
<div class="sec"><h2>Takeaways</h2><ul>
<li>Async I/O cuts cold-cache read latency materially on NVMe.</li>
<li>Skip-scan makes more multicolumn indexes usable without rewrites.</li>
<li>Upgrade path is smooth; test the new planner stats first.</li></ul></div>
<div class="sec"><h2>Highlights</h2><ul>
<li>Async I/O subsystem (effective_io_concurrency now matters more).</li>
<li>B-tree skip scan for leading-column gaps.</li>
<li>Better partition-wise joins.</li></ul></div>
<div class="sec"><h2>Summary</h2><p class="body">Postgres 18 is an incremental but meaningful release: the headline is the new asynchronous I/O path, with skip-scan and planner improvements close behind. For most homelabs the upgrade is low-risk and worth it for the read-latency wins.</p></div>
<div class="dock"><button class="ask">◆ Ask about this video</button><span class="dockhint mono">answers only from this video's transcript</span></div>
</article>
<div class="card"><div class="title">RAG is dead, right?? — Conference Talk</div>
<div class="meta mono">youtube · 2026-06-09 · 31 min <span class="chip q">queued</span></div></div>
</div></body></html>
+89
View File
@@ -0,0 +1,89 @@
# Spec — Chat with a video's stored transcript (ADR-027)
**Repo:** tapir · **Size:** medium · **Solo session.** Implements ADR-027. Read CLAUDE.md,
DECISIONS.md (ADR-021 transcript store, ADR-022 model chain, ADR-012 isolation, ADR-027), and
`docs/ui-spec.md` first. TBD, conventional commits, `task check` green per commit, `templ
generate` after view changes.
## FIRST: append ADR-027 to DECISIONS.md
ADR-027 text is provided separately (planning thread). Insert immediately before the
`## Rejected alternatives` heading, as the first commit, so the decision precedes the build.
## What this is
A per-video chat letting the user ask questions against a video's **already-stored** transcript,
entered from the summary view. Born from observed demand: the maintainer read real summaries and
some made him want to dig deeper — this gives that "I want more" reaction somewhere to go, without
watching the video.
## HARD CONSTRAINT — stored-transcript-only (the safety property)
Chat is available **ONLY** for videos that already have a stored transcript (ADR-021). It must
**never** trigger a caption fetch, never touch the rate gate, never reach YouTube. Entry being
"from a summarized video" guarantees the transcript exists. If somehow invoked on a video with no
stored transcript → show "transcript not available for chat", NO fetch. This is what makes the
feature safe by construction; do not add an on-demand-fetch path (explicitly deferred).
## 1. Entry point
- A "Dig deeper" / "Ask about this" affordance on the **summary view** of a summarized video
(not the list cards — the detail/summary page). Quiet, consistent with the existing card-state
styling.
- Opens a chat panel/view scoped to that one video, with its stored transcript as context.
## 2. The chat
- Read the stored transcript for the video (via the ADR-021 `TranscriptStore`, keyed by
`(provider, provider_video_id)`). No fetch.
- Send transcript + the user's question + minimal system framing to the chosen model via the
**existing LiteLLM gateway** (the same client the summarizer uses — a chat is a different
call, not a new integration).
- Stream or return the answer; render in the chat panel. HTMX/no-JS ethos — match the existing
app (the summarize status uses HTMX polling; chat can use a simple POST-and-render or HTMX
streaming if clean).
- **Transcript truncation:** reuse/respect `TAPIR_MAX_TRANSCRIPT_CHARS` (ADR-022) so a long
transcript fits the model context. If truncated, the chat should be honest that it's working
from a bounded portion (a quiet note), since answers about the tail of a long video may be
incomplete.
## 3. Model selection (the instrumentation win)
- **Default model = the model that produced this video's summary.** (Store/lookup which chain
model summarized it — if not already recorded, this is a small addition; if recording it is
non-trivial, default to the chain primary and note the gap.)
- **User can switch** among the ADR-022 chain models (`phi4-mini`, `gemma4-26b`,
`mistral-small` to start) via a simple selector in the chat panel. Switching re-runs against
the same transcript — this is deliberate model-comparison instrumentation.
- Respect the local-first / NDA posture: if `TAPIR_CLOUD_FALLBACK_MODEL=""` (cloud disabled),
the external model is NOT offered in the switcher — only local models. Chat must honor the same
"content stays local" guarantee as ADR-022.
## 4. Ephemeral (v1)
- No persisted chat history. Conversation lives for the session/page. No new table, no migration.
- (Multi-turn within a session is fine — keep the running messages in the request/page state —
but nothing is written to the DB.)
## 5. Isolation
- The transcript is shared/non-RLS (ADR-021) — fine, it's public content. But the chat is invoked
by a user about a video **in their feed**; confirm the entry path is reachable only for the
requesting user's own videos (the summary view is already RLS-scoped). Chat adds no new
user-data surface (ephemeral), so there's nothing new to RLS — but the test should confirm a
user can only open chat from their own summary view, not arbitrary video ids.
## Tests
- Chat on a video with a stored transcript → answer returned; assert NO caption-fetch / no
YouTube call occurs (the safety property — this is the key assertion).
- Chat invoked on a video with no stored transcript → honest "not available", NO fetch.
- Model switch → re-runs against the same transcript with the selected model; cloud model absent
from the switcher when `TAPIR_CLOUD_FALLBACK_MODEL=""`.
- Truncation honored for a long transcript; the bounded-context note shows.
- Entry is reachable only from the user's own summary view (isolation).
## Out of scope / deferred (record, don't build)
- **Persisted chat history** (per-user, RLS-scoped) — deferred until evidence anyone revisits a
conversation.
- **Show-source / transcript-verification UI** — the natural v2 (ADR-027 records it); v1 is
chat-only/trust-the-model. ADR-021's stored transcript makes v2 cheap when wanted.
- **On-demand fetch** for un-stored videos — would reintroduce the caption-fetch surface the
stored-only constraint removes. Not now.
- Anything that nudges the user to return (ADR-020 — gate contamination).
## Boundaries
Stored-transcript-only (HARD). No rate-gate/fetch surface. No auth changes. No new persisted
state in v1. Reuse the existing gateway client + truncation config; don't build a new model
integration.
@@ -0,0 +1,153 @@
# Spec — Landing page + documentation reconciliation
**Date:** 2026-06-03
**Status:** Ready to build
**Scope:** Two parallel workstreams — (A) a public landing page; (B) reconciling the
requirements / use-case / architecture / data-model docs against the deployed reality
(v0.4.0). These are separate concerns; do not let one worker do both, or the audit gets
done cursorily.
All work: read `CLAUDE.md` + `DECISIONS.md` first. TBD — commit directly to `main`, one
logical change per commit, conventional commits, `task check` green before every commit.
After editing any `.templ`, run `templ generate` (the repo commits both `views.templ` and the
generated `views_templ.go`).
---
## Workstream A — Public landing page
### Goal
A public landing page at `/welcome`, in the established bubbletea aesthetic, that lets a
visitor sign in (one Dex flow) and, if already logged in, jump to their Tapir page or log out.
New public transport surface only — no engine/core change (ADR-003).
### Verified facts (read from `internal/web/oidc/oidc.go` @ main — do not re-guess)
- Auth endpoints are exactly `/auth/login`, `/auth/callback`, `/auth/logout`.
- `isPublicPath(p)` = `p == "/healthz" || strings.HasPrefix(p, "/auth/")` — the single
public-route chokepoint inside `DexAuth.Middleware`.
- `DexAuth.CurrentUser(r) (web.User, bool)` reads the session cookie and does NOT redirect —
this is the "peek" the landing page uses to branch logged-in vs logged-out.
- `handleCallback` redirects to `/` on success (correct — leave as-is).
- `handleLogout` currently redirects to `loginPath` (`/auth/login`) — this is wrong for this
feature (see A3).
- There is NO separate "sign up" against Dex/OIDC: one authorization flow. Registration is
Tapir's own `/register` step (ADR-012), reached after first login for an unknown subject.
### Tasks
**A1 — make `/welcome` public.** In `oidc.go`, extend `isPublicPath`:
```go
func isPublicPath(p string) bool {
return p == "/healthz" || p == "/welcome" || strings.HasPrefix(p, "/auth/")
}
```
**A2 — unauthenticated bare-`/` → `/welcome`; deep links unchanged.** In `DexAuth.Middleware`,
the unauthenticated branch currently always calls `redirectToLogin`. Change it so that when
`r.URL.Path == "/"` an unauthenticated visitor is redirected to `/welcome`; for any other
guarded path keep `redirectToLogin` (so a shared `/v/{id}` deep link still bounces through Dex
and returns to the destination). Keep the `isPublicPath` check first (redirect-loop guard).
**A3 — logout lands on `/welcome`, not login.** In `handleLogout`, change the final redirect
from `loginPath` to `/welcome`. As written it sends the user to `/auth/login`, which
immediately starts a fresh Dex login — visibly failing to log out. This intentionally breaks
the existing logout test (oidc_test.go) which asserts redirect to `/auth/login`; update that
test to expect `/welcome`. That break is expected, not a regression.
**A4 — mount the landing handler** in `internal/web/handlers.go` `Router()`, on `root`,
OUTSIDE `Auth.Middleware`, alongside `/healthz`:
```go
root.HandleFunc("GET /welcome", a.handleWelcome)
```
`handleWelcome` peeks `a.Auth.CurrentUser(r)` and renders `WelcomePage(user, ok)`. Not behind
`Auth.Middleware` or `registrationGate`.
**A5 — `WelcomePage` templ component** in `views.templ`. Reuse the existing shared
layout/header partial and the established aesthetic (#7653FC purple rounded ╭─╮╰─╯ box, pink
tapir mascot, #0EF9B6 mint accents) — match the existing pages, do not reinvent styling.
- Logged out (`ok == false`): tapir mascot + tagline; one primary CTA **"Get Started"** →
`/auth/login`; honest sub-text: "New here? You'll set up your account right after signing in
— returning users go straight through." One button only (see verified facts: no separate
Dex sign-up; two buttons to the same URL would mislead).
- Logged in (`ok == true`): "Go to my Tapir" → `/`; "Log Out" → `/auth/logout`. May greet via
`user.Email`.
**A6 — tests** (extend `handlers_test.go` patterns). Note `StubAuth.CurrentUser` always returns
true; for the logged-out case use a fake Auth returning `(web.User{}, false)`.
- `GET /welcome`, no session → "Get Started" → `/auth/login`.
- `GET /welcome`, with session → "Go to my Tapir" + "Log Out".
- Unauthenticated `GET /` → 302 `/welcome`.
- Unauthenticated `GET /v/{id}` → still 302 `/auth/login` (deep link preserved).
- Authenticated `GET /` → still serves the list, unchanged.
- oidc: `handleLogout` → 302 `/welcome` (update the existing test).
**A out of scope:** no Dex config change, no new auth/session logic, no sign-up backend.
---
## Workstream B — Documentation reconciliation
### Why
The guardrail docs were written before Stage 1 and the web surface. Several now describe the
opposite of the deployed reality (v0.4.0). Stale guardrail docs are worse than none — a future
cold session (human or agent) trusts them. This workstream brings requirements, use cases,
architecture, and data-model back in sync with `main`. Each fix is one commit; cite the ADR or
migration that is the source of truth.
### Known drift to fix (verified this session — not exhaustive; the worker confirms against code)
**B1 — `internal/web/auth.go` comments.** The `User.Subject` doc and package doc still say
"single-user allowlist (ADR-011)" / "Stage-0". Code is multi-user (ADR-012). Update the
comments to describe the current multi-user reality; reference ADR-012.
**B2 — `docs/data-model.md` isolation status.** It says isolation enforcement is "dormant at
Stage 0". It is now LIVE: Postgres RLS, `FORCE`d on all user-owned tables, with a passing
two-user isolation test (ADR-012, migration 003). Rewrite that section to describe enforced
RLS as the current state; keep the history honest (was dormant at Stage 0, enforced from
Stage 1).
**B3 — `docs/data-model.md` schema completeness.** The doc predates migrations 002006. Add
the entities/columns that now exist: `summary_actions` (002), RLS (003), `user_identities`
(004, dex_subject→user_id), `video_connections` (005), `users.auto_summarize` +
`videos.summarize_requested` (006). The ER section should match the live schema. Cross-check
against `internal/adapters/store/migrations/*.up.sql` — those are ground truth.
**B4 — `docs/architecture/architecture.md`.** Predates the entire web surface. Update the C4
container diagram and text to include: `tapir serve` (HTMX+Templ web reader/writer), the Dex
OIDC session layer (`internal/web/oidc`), registration gate, web-initiated YouTube connect,
account management, and the immediate-processing path (web "Summarize" button → background
goroutine → status poll). The engine/ports/sinks core is unchanged (ADR-003) — show the web
surface as a new transport over the same core, not a core change.
**B5 — `docs/use-cases/*.feature`.** Add scenarios for the behaviours now live and unspecced:
register (new subject → registration → user row; returning user straight through), connect
YouTube (web OAuth), disconnect, delete-account (cascade + secret purge, Dex untouched —
ADR-013), manual-vs-auto summarize mode + the Summarize button, and the landing page
(logged-out CTA; logged-in shortcuts). Keep them as executable-style Gherkin consistent with
the existing files.
**B6 — `DECISIONS.md` ADR ordering (cosmetic).** ADR-010 sits before ADR-009/011 (append
order). Reorder to numeric while you're in the file. Pure tidy, no content change.
**B7 — requirements check.** If a requirements doc exists (e.g. `docs/ui-spec.md`, referenced
by ADR-011), reconcile it with what shipped: note where the build deviated (e.g. the spinner /
immediate processing / summarize mode were beyond the original spec) so the spec reflects
reality or explicitly records the deviation. Do not silently rewrite history — record
deviations as deviations.
### B working method
- Source of truth order: migrations + code > ADRs > prose docs. When a prose doc disagrees
with code, the code wins and the doc is corrected (unless the code is the bug — then flag it,
don't quietly doc around it).
- One logical doc per commit. Cite the ADR/migration that justifies each change in the commit
body.
- This is an audit, not a rewrite: preserve the docs' structure and the "rejected alternatives
/ history" honesty. The goal is *current and trustworthy*, not *pretty*.
---
## Coordination
A and B touch mostly different files (A: oidc.go, handlers.go, views.templ, tests; B: docs/* +
auth.go comments). The one overlap is `auth.go` (B1 edits comments) vs A (reads it) — no
conflict. Run A and B in parallel; commit independently to `main`.
If anything in B reveals that code, not docs, is wrong (e.g. an isolation gap, a migration that
doesn't match the data-model intent), STOP and surface it — that's a finding, not a doc edit.
+77
View File
@@ -0,0 +1,77 @@
# Spec — Newest-first batch ordering + honest "Try now" / prioritisation docs
> **Extended by ADR-020 (2026-06-08).** This spec covers the *batch processing* order within a
> pass. ADR-020 adds (a) a recency pre-filter — auto mode skips videos published before
> `TAPIR_AUTO_SUMMARIZE_WINDOW`, listed but summarised on demand — and (b) the same
> `published_at DESC NULLS LAST` ordering on the **list read** (`ListVideos`), which previously
> sorted by `seen_at`. See `DECISIONS.md` ADR-020.
**Repo:** tapir · **Size:** small · **Solo session.**
**Why.** Product intent (maintainer, 2026-06-06): a new user should get summaries of their
**newest** videos quickly, while the older back-catalogue fills in behind — all within the one
shared rate gate. Today the foreground path ("Try now" button) lets a user hand-pick a video,
but the **background batch processes in subscription/channel order, not newest-first** — so a new
user with a large candidate set sees the batch summarise whatever channel is first in their
subscription list, not their newest videos. This slice makes the batch agree with the intent, and
fixes the docs to describe the real rationale (onboarding prioritisation), not the
traffic-disguising framing a prior session wrote.
Read `CLAUDE.md` + `DECISIONS.md` (ADR-014, ADR-018) first. TBD, conventional commits,
`task check` green per commit, `templ generate` if views change.
## 1. Newest-first batch ordering (the build)
In `internal/runner/runner.go` `RunOnce`: today the loop processes each video inline while
walking subscriptions channel-by-channel (`for sub → NewVideos → for v → process`). Change so
that, within a pass, **candidates are processed newest-first across ALL channels**:
- Collect the candidate videos across channels first (after dedup/seen/manual/rate-limit
filtering as today), then **sort by `published_at` descending before processing**, then process
in that order through the engine + shared `globalFetchGate`.
- **`published_at` is nullable** (schema 001). Sort **NULLS LAST** — videos with no publish date
must not jump ahead of dated newest videos. Decide a stable tiebreak (e.g. `seen_at DESC`) for
equal/again-null dates.
- Keep all existing behaviour: per-item failure isolation, the rate-limit backoff skip, manual
mode, channel-unavailable handling, stats. Ordering is the only change — not what gets
processed, just the order.
- At 868 candidates a collect-then-sort in memory is fine; do **not** build a streaming/external
sort. Keep it simple.
- The shared rate gate (`globalFetchGate`) is unchanged and still governs fetch pacing — ordering
does not bypass or weaken it.
**Optional (only if cheap and clearly correct):** a soft cap so the *first* pass for a brand-new
user summarises the newest N (e.g. 20) quickly and defers the long tail to subsequent passes — so
onboarding value lands fast without waiting for the whole sorted set. If this adds real
complexity, SKIP it and just do the newest-first ordering; the ordering alone delivers the intent.
## 2. Tests
- Given candidates across multiple channels with mixed `published_at` (incl. some NULL), assert
the processing order is newest-first, NULLS LAST, with the chosen tiebreak. Use the existing
fake VideoStore/Processor pattern in `runner_test.go`.
- Assert ordering does not change *which* videos are processed vs. today (same set, new order).
- Rate-gate / backoff / manual-mode behaviour unchanged (existing tests stay green).
## 3. Docs — describe the REAL rationale (replace prior framing)
The "Try now" button and the discovery batch together implement **onboarding prioritisation**:
foreground (user-clicked "Try now") summarises a specific video on demand; background batch
summarises newest-first; both honour the shared rate gate. **Update the docs to state this intent
— and explicitly REMOVE/replace any framing that describes "Try now" as making traffic "look
organic to YouTube" or evading rate limits.** That is not the rationale. The rationale is: *get
the user a few summaries of their newest, most relevant videos fast; process the back-catalogue in
the background; always within the honest shared rate limit.* Rate limiting is **respected**, not
evaded.
- `docs/ui-spec.md`: "Try now" = on-demand foreground summarisation of a chosen (typically newer)
video; rationale = fast onboarding value, not traffic shaping.
- `docs/architecture/architecture.md`: document the two-path model — foreground on-demand vs.
background newest-first batch, both through `globalFetchGate` — and the newest-first ordering.
- Any requirements/use-case doc mentioning discovery order: state newest-first.
- If a brain note or `wiki` entry captured the "looks organic" rationale, correct it there too.
## Boundaries
- Do NOT increase fetch rate or weaken the rate gate. Account-safety constraint stands: the
caption endpoint is unofficial (ADR-010) and must be treated with honest backoff, never evasion.
- Do NOT touch RLS, credentials, or the Dex surface.
- Ordering change is within a pass only — no persisted priority queue, no new table.
## Out of scope
Per-user configurable ordering; priority weighting beyond newest-first; the soft-cap if it proves
non-trivial.
+128
View File
@@ -0,0 +1,128 @@
# Spec — Onboarding "wow" burst: better picks, stronger model
**Repo:** tapir · **Size:** medium · **Solo session** (not a swarm).
> **Status: built (v0.25.0, ADR-028).** This supersedes the original investigate-first brief
> (committed as the prior version of this file): Phase 1 was run against the live pilot DB and its
> findings are folded into "Why this exists" below; Phase 2 was built as described here. The one
> brief lever NOT built — the honest "the rest fill in over the coming days" framing copy — is
> listed under *Explicitly NOT in this slice*.
**Why this exists.** A new user's first session decides whether they return (the Stage-0 gate,
VISION.md). On connect, Tapir fires a capped burst (≤`TAPIR_ONBOARD_SUMMARIZE_COUNT`, default 3)
that summarizes the user's newest unsummarized videos so the feed isn't empty (the burst itself
works — wired in `cmd/tapir/discovery.go``cmd/tapir/main.go` `onboard`). A Phase-1
investigation of the live pilot DB found the burst *fires* but delivers a **weak first
impression** for two concrete reasons, and ruled out a third idea:
1. **Picks are junk.** Selection is pure newest-first (`videos.NewestUnsummarizedVideoIDs`,
`ORDER BY published_at DESC`) with **zero quality signal**. Pilot user "Jonte"'s live burst-3
were a stock-ticker **livestream** + two regional news clips — the newest, not the best.
2. **Weakest model on the first impression.** All of Jonte's summaries ran on
`koala/phi4-mini` (the documented weak link — ADR-022 was born from its failures). The
stronger, brain-validated `iguana/gemma4-26b` was never used for the burst.
3. **Cached-first is empty at pilot scale — REJECTED.** The idea (summarize already-cached
transcripts instantly, zero fetch) dies on the numbers: only **11 videos** overlap between the
two pilot users (~3% of each library), **0** cached-and-unsummarized, and a new user's
newest-20 unsummarized are **20/20 NOT cached** — newest-first and cached-first are
structurally incompatible (fresh uploads are exactly what nobody has fetched yet). Not built.
This is a **curation/latency problem for ~3 videos, NOT a throughput/429 problem** — fetching 3
captions is nowhere near the rate limit. Nothing here fetches harder or pressures the rate gate;
it picks the right few videos and runs a better model on them.
Read `CLAUDE.md`, `DECISIONS.md` (esp. ADR-014, ADR-018, ADR-020, ADR-021, ADR-022, ADR-023,
and the new **ADR-028**), and `VISION.md` (the Stage-0 gate) first. TBD — commit directly to
`main`, one logical change per commit, conventional commits, `task check` green before each
commit, `templ generate` if any view changes (none expected).
## Decisions already made (do not reopen)
- **Not a throughput change.** The caption rate gate (ADR-014) is untouched — same pacing, same
priority lane (ADR-026). This slice changes *which* ≤3 videos the burst spends its fetches on
and *which model* summarizes them, never how fast or how many.
- **Cached-first is dropped** (ADR-028, the 3% overlap). The engine's existing read-stored-first
(ADR-021, `resolveTranscript`) stays — it already gives a free instant summary on the rare
cache hit, transparently. We do not *select* for cache hits.
- **has-captions is not a pre-fetch signal.** It is only knowable after a gate fetch (or a cache
hit, ~0 for new videos). Selection can only *avoid known-junk* (Shorts/live/over-long) — it
cannot *guarantee* captions. The spec is honest about this: better odds, not a promise.
- **No credentialed caption fetch** (ADR-010/ADR-026 dead end). **No client extension.**
## 1. Persist `duration_s` at discovery (the enabling change)
The `videos.duration_s` column exists (migration 001) but is **never written** — ADR-023's
`filterLowValue` (`internal/adapters/youtube/youtube.go`) already fetches each candidate's
duration via the cheap quota `videos.list` call, uses it to drop Shorts/live, then **discards
it**. Stop discarding:
- Add `DurationSeconds int` to `domain.Video`.
- In `filterLowValue`, set `DurationSeconds` on each kept video from the `videos.list` `meta`.
- `UpsertVideo` writes `duration_s`, **COALESCE-preserving** a known value (never overwrite a
real duration with 0/unknown), mirroring the `channel_title` backfill stance (migration 014).
- No new migration — the column is already there.
Consequence: a fresh user's connect-triggered discovery pass runs **before** the onboard burst
(`Enqueue`: `run()` then `onboard()`), so duration is populated for the burst's candidates at
connect. Existing rows backfill on their next discovery pass; until then their `duration_s` is
NULL and treated as "unknown" (§2).
## 2. Junk-avoiding burst selection
New store method, RLS-scoped via `withUser`:
```
OnboardBurstVideoIDs(ctx, userID string, limit, minSeconds, maxSeconds int) ([]string, error)
```
- Same base as the old `NewestUnsummarizedVideoIDs`: the user's videos with no summary yet,
`ORDER BY published_at DESC NULLS LAST, seen_at DESC`, `LIMIT limit`.
- **Exclude known-junk**: a row is dropped only when `duration_s IS NOT NULL` **and**
(`duration_s < minSeconds` OR `duration_s > maxSeconds`). A NULL duration is **unknown** — kept
(degrade-open: never starve the burst because metadata is missing), but ordered *after* rows
with a known-good duration so a freshly-enriched good pick wins when both exist.
- `minSeconds` reuses `TAPIR_MIN_VIDEO_SECONDS` (default 60 — the Shorts floor, ADR-023).
`maxSeconds` is new: `TAPIR_ONBOARD_MAX_VIDEO_SECONDS` (default 14400 = 4h) — drops the
multi-hour livestream VODs that pass the live filter once ended.
- `minSeconds<=0` and `maxSeconds<=0` each disable that bound (so `0/0` == the old
newest-first behaviour, the reversibility lever).
- The burst switches to this method; `NewestUnsummarizedVideoIDs` is removed (fully superseded —
`OnboardBurstVideoIDs(., 0, 0)` is identical pure-newest behaviour).
## 3. Stronger model for the burst
The burst summarizes only ≤3 videos, so a slower, stronger model is affordable exactly here.
- New config `TAPIR_ONBOARD_SUMMARIZER_MODEL` (default `iguana/gemma4-26b` — the brain-validated
homelab general-purpose model, already the ADR-022 fallback).
- Build a **burst-specific summarizer chain** that puts the onboard model **first**, then the
standard chain (primary → local fallback → cloud) as resilience, deduped. Wrap it in a
burst-specific `engineProcessor` reusing the same store/transcript-cache/sink — a pure wiring
choice, engine and ports unchanged (Clean Architecture, ADR-003).
- The `onboard` closure uses the burst processor instead of `app.Processor`.
- **Collapse cleanly**: when `OnboardSummarizerModel` is empty or equals `SummarizerModel`, the
onboard path reuses `app.Processor` (no separate chain) — the reversibility lever.
- Local-first preserved: the onboard model is a local alias; the cloud endpoint stays last in the
chain, so a client/NDA deployment with `TAPIR_CLOUD_FALLBACK_MODEL=""` keeps burst content
local too.
## 4. Behaviour spec + docs
- Add scenarios to `docs/use-cases/connect_account.feature` (the connect → burst flow): burst
skips a too-long/live video in favour of a reasonable-length one; burst summarizes with the
stronger model first. Map them in `scenarioCoverage` so `TestScenarioCoverage` stays green.
- Update `docs/architecture/architecture.md` (the onboarding-burst section) to describe the
junk-avoiding selection + the burst model override.
- ADR-028 in `DECISIONS.md` records the rationale (incl. the rejected cached-first lever).
## Success criteria
- `task check` green (fmt, vet, lint, `go test -p 1 ./...`).
- A unit test proves `OnboardBurstVideoIDs` drops a known too-long / sub-min video and keeps a
good one, newest-first, RLS-scoped, unsummarized-only.
- A test proves discovery persists `duration_s` and does not clobber it on re-upsert.
- A test proves the burst chain leads with the onboard model (then the standard chain).
- Config defaults + bounds tested (`OnboardMaxVideoSeconds`, `OnboardSummarizerModel`).
- No change to the rate gate, fetch pacing, or burst cap. `0/0` + empty model == prior behaviour.
## Explicitly NOT in this slice
- Cached-first selection (rejected, ADR-028).
- Any caption-availability *guarantee* (impossible pre-fetch).
- **Honest "taster" framing copy** ("summaries of a few of your videos to get you started — the
rest fill in over the coming days"). A good lever from the original brief, but it's a UI/copy
change with no backend dependency; deferred to a UI pass, tracked as an issue.
- Backfilling `duration_s` for existing rows via a migration (it backfills lazily on discovery).
- Return-nudges / digests (ADR-020: poisons the unprompted-return signal).
- Raising fetch throughput, multi-IP, or Whisper (out of scope; the gate is deliberate).
+92
View File
@@ -0,0 +1,92 @@
# Spec — In-process scheduled discovery + auto-summarize + rate-gate finish
> **Extended by ADR-020 (2026-06-08).** Auto-summarize is no longer "every unseen video": the
> scheduler now skips videos published before `TAPIR_AUTO_SUMMARIZE_WINDOW` (default ~7d) unless
> explicitly requested, so a back-catalogue does not re-drive the rate gate every cycle. See
> `DECISIONS.md` ADR-020.
**Repo:** tapir · **Size:** medium · **Solo session** (not a swarm).
**Why this exists.** The Stage-0 gate ("me or a friend returns and reads/acts in ≥2 separate
weeks") cannot be met because the system is not usable *unprompted*: discovery (`tapir run`) is
host-side manual, so a newly onboarded user sees an empty list and never comes back. This slice
makes Tapir watch on its own — the thing that makes the gate experiment actually runnable.
Read `CLAUDE.md` + `DECISIONS.md` (esp. ADR-012, ADR-014, and the new ADR-018) first. TBD —
commit directly to `main`, one logical change per commit, conventional commits, `task check`
green before each commit. `templ generate` if any view changes.
## Decisions already made (do not reopen)
- **In-process scheduler**, NOT a k8s CronJob (maintainer's call: simpler deploy, acceptable
coupling at 3 users). The known cost — discovery shares the web process's lifetime and egress
— is accepted and recorded in ADR-018.
- **Auto-summarize ON** for the maintainer + onboarded friends (zero-friction: the list fills
and summarizes itself).
- **Gate clock resets** to when this ships (ADR-018) — until unprompted use is possible, the
prior window measured nothing.
## 1. In-process scheduled discovery (core)
- In `tapir serve` startup, launch a background goroutine that runs discovery for ALL users on
an interval: env `TAPIR_DISCOVERY_INTERVAL` (Go duration, e.g. `2h`). **Unset or 0 = disabled**
(so dev/tests never auto-fetch).
- **Reuse the existing `runner.Runner` + `Loop`/`RunOnce`. Do NOT write a new scheduler.** The
per-user Runner already exists; the new work is **iterating users** and running one pass each
per tick. Enumerate users from the un-RLS'd `user_identities` (the same enumerate-then-act
pattern the login_events gate query established), then run each user's pass **inside that
user's RLS scope** (`withUser`).
- **Stateless timing:** run-once-on-startup, then every interval — exactly the existing `Loop`
shape. Do NOT persist schedule state; a pod restart just restarts the cycle. Acceptable at this
scale. Do not build cron-in-Go.
- **Graceful shutdown:** the goroutine respects `ctx` cancellation so a pod term doesn't wedge.
- **Failure isolation in the loop:** one user's pass failing (or one channel/video) must not
abort the other users or crash `serve` — log and continue. (RunOnce already collects per-item
errors; preserve that at the per-user level too.)
## 2. Auto-summarize default ON for Future-B users
- New registrations default `auto_summarize = true` (so onboarded friends get zero-friction);
keep the account-page toggle so a user can switch to manual. One-off update existing user rows
to `true` as well (maintainer + any current users).
- Consequence (intended): scheduled discovery both discovers AND summarizes new videos — which
is the point, and is why §3 is mandatory in the same slice.
## 3. Finish/confirm the ADR-014 shared per-egress-IP rate gate (NOW load-bearing)
- In-process scheduling + auto-summarize + multiple users = all caption fetches leave the **one
web pod's egress IP**, concurrently with any live "Summarize" button clicks. The timedtext
endpoint rate-limits per IP (ADR-010/014). Without a shared gate this self-inflicts 429s every
cycle.
- **Confirm in code whether ADR-014 item 2 (a single PROCESS-WIDE rate gate) exists.**
Reconciliation flagged it as possibly built only as per-*video* backoff. If it is not a
process-wide gate, **build it now**: ONE shared limiter (token-bucket / min-interval) that
every timedtext/caption fetch passes through — scheduler loop AND click-path alike. Per-process,
not per-user, not per-video.
- Keep the existing per-video 429 backoff (`rate_limited_at` + retry window) — complementary: the
gate prevents tripping 429; the backoff handles it if one still happens.
- Honest UX (ADR-014 item 3) still applies: a fetch waiting on the gate shows "queued/waiting",
never a stuck spinner.
## 4. Gate-clock reset — already recorded in ADR-018; verify VISION reflects it
- ADR-018 (committed) resets the Stage-0 34 week window to start when this ships, and revises the
check-in date. VISION Stage 0 carries a pointer to it. The build doesn't re-decide this; just
ensure nothing in docs still implies the clock started earlier.
## Tests
- **Scheduler:** fake clock + fake Runner → N users each get one pass per tick; one user's failure
doesn't stop the others; `ctx` cancel stops the loop; interval=0 disables it entirely.
- **Rate gate:** concurrent fetches (scheduler + simulated click) are serialized/limited through
the ONE gate — assert max-in-flight / min-interval honored regardless of caller.
- **Auto-summarize default:** new registration → `auto_summarize = true`; account-page toggle
still flips it.
## Out of scope / known constraints
- No CronJob / k8s objects (in-process chosen).
- **SINGLE-REPLICA ASSUMPTION (load-bearing).** In-process scheduling means if `tapir serve` ever
runs >1 replica, every replica runs the discovery loop → every user fetched in parallel (429s +
duplicate work). At 3 users this is single-replica, fine — but the build MUST note this
constraint in ADR-018 / deploy docs so a future scale-up doesn't silently double-run.
- No persisted schedules, no multi-pod coordination.
## Fallback if the session runs short
Ship the scheduler with **auto-summarize OFF** (discovery only; manual Summarize button) until the
process-wide rate gate (§3) is confirmed/built. NEVER ship auto-summarize-on-a-schedule without the
gate — that combination self-inflicts 429s for every user every cycle. Auto-summarize ON is gated
on §3 being done.
+77
View File
@@ -0,0 +1,77 @@
# Spec — Stage 0 usage measurement (login events)
**Date:** 2026-06-03
**Status:** Ready to build · **Repo:** tapir · **Size:** small (one migration + middleware + query)
**Why:** The Stage 0 gate (VISION, ADR-016) is *return usage in ≥2 separate weeks*. `summary_actions`
captures *acts* (watch/skip/save) but not *reads* — a friend who logs in weekly and reads summaries
without clicking anything is invisible. For a **reading** product that is the most important signal.
This adds the missing data so the gate is measurable as written. Solo session, not a swarm.
Read `CLAUDE.md` + ADR-016 first. TBD, conventional commits, `task check` green before each commit.
## Scope (resist sprawl — this is NOT analytics)
A lightweight, append-only record of *when each user was active*, enough to answer
"returned/read in ≥N distinct weeks". Not page-level events, not click tracking, not a funnel.
### 1. Migration — `login_events` (append-only)
```
login_events (
id UUID PK default gen_random_uuid(),
user_id UUID NOT NULL, -- per-user; RLS like every user-owned table
seen_at TIMESTAMPTZ NOT NULL default NOW()
)
INDEX (user_id, seen_at)
```
- **RLS:** `FORCE ROW LEVEL SECURITY`, same policy/pattern as the other user-owned tables (the
`tapir.current_user_id` GUC via the `withUser` seam — match migration 003). A reporting query that
needs cross-user counts runs as the owner/maintainer outside the per-user scope, or via a dedicated
read — decide consistently with how existing admin-ish reads are done.
- Append-only: no updates, no deletes except the user-delete cascade. **Add to the delete-account
cascade** (ADR-013) — `login_events` has no FK (mirrors `summary_actions`), so `DeleteUser` needs an
explicit delete for it, and the delete test must assert it's covered. *Do not forget this* — it's the
exact footgun the last delete work caught.
### 2. Middleware — throttled stamp
- In the authenticated request path (after `CurrentUserID` resolves, inside the registration-gated
app — NOT on `/welcome`/`/healthz`/`/auth`), record one `login_events` row **per user per day**
(throttle: skip if a row exists for this user with `seen_at` ≥ start-of-today). One insert per active
day, not per request — keeps the table small and the signal clean.
- Throttle check must itself be RLS-scoped (`withUser`). Keep it cheap (indexed lookup).
### 3. Query — the gate report
Provide a query (and optionally a tiny `tapir report` CLI subcommand or an admin page — your call,
CLI is fine) answering, per user:
```sql
-- distinct active weeks from reads (login_events) AND acts (summary_actions), unioned
WITH weeks AS (
SELECT user_id, date_trunc('week', seen_at) AS wk FROM login_events
UNION
SELECT user_id, date_trunc('week', acted_at) FROM summary_actions
)
SELECT user_id, COUNT(DISTINCT wk) AS active_weeks
FROM weeks GROUP BY user_id
ORDER BY active_weeks DESC;
```
Gate passes when any user_id (maintainer or friend) reaches `active_weeks >= 2` within the window.
## Honesty caveats to carry (from VISION/ADR-016)
- **"Unprompted" is not measurable here.** login_events records *that* a user returned, not *why*. A
nudged return looks identical to an organic one. This build does not close that gap and must not
claim to — the VISION measurement note stands: count returns, read a nudged return as weaker signal.
(If prompt-tracking is ever wanted, that's a separate decision, not this build.)
- **Data accrues from deploy onward.** The gate window's read-data starts when this ships — so ship
soon (maintainer's call) rather than batching with the infra tooling session.
- **`date_trunc('week')` is ISO/timezone-sensitive** and noisy at low volume (N=3). Two visits days
apart can fall in the same or different weeks. Acceptable, but don't over-read a single-week-margin
pass/fail.
## Out of scope
Page/event analytics; prompt-vs-organic tracking; dashboards beyond the one gate query; anything
touching the engine or sinks (this is web/store only — ADR-003 holds).
## Tests
- Migration up/down; RLS on `login_events` (extend the two-user isolation test to cover it).
- Throttle: N requests same day → 1 row; next day → 2nd row.
- `DeleteUser` removes the user's `login_events` and leaves others' intact (extend the delete test).
- The gate query returns correct distinct-week counts across a seeded reads+acts fixture.
+77
View File
@@ -0,0 +1,77 @@
# Spec — Unify video-card states: one "Summarize now" verb, honest no-captions state
> **Superseded in part by ADR-020 (2026-06-08).** The five card states still hold, but the copy
> changed: the nudge verb is now **"Summarize"** (not "Summarize now"), the rate-limited state
> reads **"In queue"** (not "Fetching soon…"), and the queued state reads **"summarizing
> shortly"** (not "waiting for the next run"). The list also now collapses older un-summarized
> and caption-less videos. See `DECISIONS.md` ADR-020 and `views.templ` (`VideoCard`) for the
> current copy; this doc is kept as the original design record.
**Repo:** tapir · **Size:** small, **view-layer only** (`views.templ` + a little CSS in
`view.go`; regenerate `views_templ.go`). No handler, store, or DB change. The two existing
handlers (`/summarize`, `/retry-now`) stay exactly as they are — only what the card *shows*
changes.
**Why.** The video card today presents two different verbs — "Try now" (`.btn-retry`, on
rate-limited videos) and "Summarize" (`.btn-secondary`, on pending videos) — for what the user
experiences as one intent: *"summarize this video now."* The user doesn't know or care about the
internal pipeline state (rate-limited vs. manual-queue); two differently-labelled, differently-
styled buttons leak that state machine into the UI as a choice. Also: a **no-captions video
(`TranscriptStatus == "none"`) currently falls into the `else` branch and wrongly shows a
"Summarize" button** that, if clicked, tries and fails — there are no captions to fetch. That's
the confusing dead-end to remove.
In normal use the maintainer is in **auto mode**, so these buttons are *exceptions*, not the main
path — videos summarize themselves. So the card should be **status-first**: the state is what the
user reads constantly; the manual nudge is a small, quiet affordance for impatience, not a
prominent call-to-action.
## The honest per-state card model (footer of `VideoCard`)
Restructure the footer branch in `VideoCard` (in `internal/web/views.templ`) to these states.
The branch ORDER matters (summarized first, then terminal/no-action states, then actionable):
1. **Summarized** — preview + provider chip + fallback badge + actions. **No button.** (unchanged)
2. **No captions** (`r.TranscriptStatus == "none"`) — **NEW branch.** Quiet status text, e.g.
`No transcript available` (use a muted `.card-state`/`.chip-retry`-style treatment, NOT a
button). This is a terminal honest dead-end — the user can do nothing, so offer nothing.
3. **Queued / requested** (`r.SummarizeRequested`) — "Queued · waiting for the next run". **No
button.** (unchanged)
4. **Rate-limited** (`r.TranscriptStatus == "rate_limited"`) — quiet status (keep the
"fetching soon" sense) + a **quiet "Summarize now"** button POSTing to `retryNowURL` (clears
backoff then processes). Same quiet style as state 5.
5. **Pending** (else — discovered, not yet attempted) — quiet "Not summarized" + a **quiet
"Summarize now"** button POSTing to `summarizeURL` (flips the queue flag then processes).
## Unify the verb and the style
- **One label everywhere a manual nudge is offered: "Summarize now"** (states 4 and 5). Drop the
"Try now" wording entirely.
- **One quiet style** for both: use the understated `.btn-retry` pattern (small, pill, outline,
transparent bg) — NOT `.btn-secondary`/`.btn` (heavier). Rename the CSS class to something
state-neutral (e.g. `.btn-quiet` or `.btn-summarize-now`) so it no longer reads as
"retry"-specific; keep the same visual. The point: the nudge is subtle, status is primary.
- Keep both `<form>`s posting to their respective existing handler URLs
(`retryNowURL` for rate-limited, `summarizeURL` for pending) with the existing HTMX
attributes (`hx-post`, `hx-target=#video-{id}`, `hx-swap=outerHTML`) — only the button
label/class change. The backend side-effect difference (clear-backoff vs. set-flag) stays
invisible to the user, which is correct.
- Drop the engineer-facing `title="Fetch transcript now through the shared rate gate"` tooltip;
if a hint is wanted, make it user-facing ("Summarize this one now").
## Quietness check (the design intent)
The summary content and the per-state *status* are the card's primary information. The "Summarize
now" button is a minor affordance. Do not make it a prominent solid-accent CTA — it must read as
"you can nudge this if you're impatient", not "action required". Status text uses muted styling;
the button uses the quiet outline style.
## Tests
- `videocard_internal_test.go` (exists): assert each of the 5 states renders the expected
footer — summarized (no button), no-captions (status, NO button, no `summarize`/`retry-now`
URL present), queued (no button), rate-limited ("Summarize now" → retry-now URL), pending
("Summarize now" → summarize URL). The key new assertion: **a `none`-status video renders no
action button and no POST URL.**
- Assert the label string "Try now" no longer appears anywhere in rendered output.
## Out of scope
Handler/DB changes; the detail-page action buttons (watched/skipped/saved — unrelated); the
pipeline stats bar wording; auto/manual mode behaviour. Verb/label/style/no-captions-state only.
+36 -2
View File
@@ -72,7 +72,14 @@ summary_actions
and join into the existing `SummaryRow` reads so list/detail show current state.
- This column is what makes the Stage-0 metric ("did I act on a summary?") queryable.
## 6. Auth (Dex OIDC, single-user authz)
## 6. Auth (Dex OIDC)
Authentication is delegated to the homelab OIDC provider at `TAPIR_OIDC_ISSUER`
**Authentik** since the Dex→Authentik migration (infra ADR-0001; ADR-019). It offers a
Google upstream and Authentik-managed accounts (incl. its invite flow); Tapir no longer
provisions accounts itself. Any authenticated subject can register a Tapir account
(ADR-012: allowlist removed). The `oidc`/`DexAuth` package keeps its name for now (rename
deferred, ADR-019).
- **Flow:** standard Authorization Code. Use `coreos/go-oidc` + `golang.org/x/oauth2`
(justify the deps in the commit; both are the homelab-standard OIDC libs and small).
@@ -92,7 +99,7 @@ summary_actions
`TAPIR_OIDC_ISSUER` (`https://auth.d-ma.be`), `TAPIR_DEX_CLIENT_ID`, `TAPIR_DEX_CLIENT_SECRET`,
`TAPIR_OIDC_REDIRECT_URL` (`https://tapir.d-ma.be/auth/callback`), `TAPIR_SESSION_SECRET`.
Reuses existing `TAPIR_DB_DSN`, `TAPIR_USER_ID` (the StubAuth dev subject only). No secrets
committed. (`TAPIR_ALLOWED_SUBJECT` was removed by ADR-012.)
committed. (`TAPIR_ALLOWED_SUBJECT` was removed by ADR-012; use the keys above.)
## 8. Deployment — k3s + Flux GitOps
@@ -147,3 +154,30 @@ Gate (lane A) commits first; B/C/D follow.
`task check` green per lane; B/C/D rebase on A. Deploy (D) lands last, after the binary serves
locally.
---
## Deviations and additions (as-built)
This spec describes the **Stage-0 single-user reader** (ADR-011). What actually shipped through
v0.4.0 went further — Stage 1 (ADR-012) opened multi-user, and several UX features were added on
top. Recorded here (append-only; the spec above is left intact) so intent and reality stay
distinguishable.
| As-built feature | What it is | Why | Covered by |
|------------------|-----------|-----|------------|
| **Multi-user + RLS isolation** | Several Dex users per deployment; isolation enforced by Postgres RLS, not the single-subject allowlist of §6. | Maintainer opened Stage 1 ahead of the formal Stage-0 gate, with DB-enforced isolation as the guardrail that keeps it safe. | ADR-012; migration 003 (`6775e5f`, `f28fdc0`, `2ae66da`) |
| **Registration gate** | A Dex subject with no `users` row is routed to `/register`, which creates the `users` row + a `user_identities` mapping. (§2 listed "sign-up / user CRUD" as a non-goal.) | Explicit registration is how a multi-user surface stays honest — no just-in-time row creation. | ADR-012; `f396e01` |
| **Per-user YouTube web connect** | `/oauth/youtube/connect``/oauth/youtube/callback` stores a per-user refresh-token ref + a `video_connections` row. (The spec assumed a host-side `tapir auth` only.) | Multi-user means each user connects their own account from the browser. | ADR-006, ADR-012; migration 005 (`0c9531a`, `2aad79b`) |
| **Account management** | `/account` page with **disconnect** and **delete account**; delete removes only Tapir-side state and leaves the Dex identity intact. (§2 listed isolation/CRUD as non-goals.) | A real account needs a way out; deletion semantics are deliberately Tapir-side only. | ADR-013; `22eafcf`, `c7624d9`, `17d5e8c` |
| **Immediate web summarization** | A quiet "Summarize now" button on non-summarized video cards. Pending cards POST to `/v/{id}/summarize` (queues + triggers engine); rate-limited cards POST to `/v/{id}/retry-now` (clears backoff + triggers engine). Both use the same `.btn-quiet` style and label — the internal pipeline distinction is invisible to the user. The page HTMX-polls `GET /v/{id}/status` while processing. Videos with `TranscriptStatus == "none"` show "No transcript available" with no button — this is a terminal honest state. (§2 said "triggering runs from the browser … do NOT build".) | Reading a list you can't act on is half a product; on-demand summarize closes the loop without waiting for a batch `tapir run`. | ADR-012, ADR-014; `25215cb`, `8c6c7ca` |
| **Charmbracelet tapir spinner** | An animated in-flight indicator (charm palette) shown while a summarize is processing; an honest "queued/waiting" state under rate-limiting rather than a stuck spinner. | The spinner must tell the truth when the timedtext endpoint rate-limits (429), not imply imminence. | ADR-014; `25215cb`, `a4aeb5e` |
| **Auto/manual summarization mode** | Per-user `auto_summarize`; manual lists new videos unsummarized and queues via `summarize_requested`; a mode toggle at `/account/summarize-mode`. Default is **true** for new users (migration 011, ADR-018); existing rows back-filled via migration 012. | Control over compute/noise — only summarize what the user cares about. | migration 006 (`748d5eb`, `bdbdce7`, `3014ee0`, `a269d4a`); migration 011/012 |
| **Public landing page** | `/welcome` mounted **outside** the auth guard; unauthenticated `/` redirects there; logout returns there (not `/auth/login`). (The spec guarded everything except `/healthz` and `/auth/*`.) | A first-time visitor needs a public "what is this / get started" page before the login wall. | `d83943c`, `0fdf2f7`, `3a27bf1`, `d208110`, `8ca374e`, `f15f57f` |
| **Invite onboarding** | **Removed from Tapir (ADR-019).** Invites are owned by the IdP (Authentik) now, not Tapir — the Dex local-password provisioning path (`tapir invite` CLI, `/invite/{token}` web flow, `internal/adapters/dex`) was deleted when the homelab migrated Dex→Authentik (infra ADR-0001). A new user is invited via Authentik's invite flow, logs into Tapir via OIDC, and is captured by the existing `/register` (display-name) gate. | Onboarding belongs to the identity provider; keeps Tapir out of the shared identity provider's write path. | ADR-019; infra ADR-0001 |
| **Summarized-only filter** | `?summarized=1` query param on the list view. When set, only videos with a completed summary (`SummaryRow.Summarized = true`) are shown. Rendered as a "Summarized only" checkbox in the filter form. Summarized videos also sort to the top of the unfiltered list (`ORDER BY (s.id IS NOT NULL) DESC, seen_at DESC`). | Lets users focus on videos that are ready to read; newly landing summaries are visible at the top without filtering. | `internal/web/view.go` (`Filter.OnlySummarized`, `ListVideos` ORDER BY) |
| **"Summarize now" foreground path** | Unified quiet nudge button on actionable non-summarized cards. Five explicit card states — (1) summarized: chip + no button; (2) no captions (`transcript_status = 'none'`): "No transcript available", no button; (3) queued: "Queued" chip, no button; (4) rate-limited: "Fetching soon…" + "Summarize now" → `POST /v/{id}/retry-now` (clears `rate_limited_at`, triggers engine); (5) pending: "Not summarized" + "Summarize now" → `POST /v/{id}/summarize` (queues + triggers engine). One verb, one style (`.btn-quiet`); backend difference invisible to user. Both handlers call `ProcessVideo` through `globalFetchGate`. Rate gate respected, not bypassed — this is onboarding prioritisation. | Fast onboarding value; honest dead-end for no-captions videos (no button that fails). | `internal/web/handlers.go` (`handleRetryNow`, `handleRequestSummarize`); `internal/web/views.templ` (`VideoCard`) |
| **Pipeline stats bar** | A one-line status bar above the video list: `N summarized · M fetching soon · K no captions`. Computed from the unfiltered row set; hidden when all videos are summarized. Gives the user a clear read on pipeline state without any interaction. | Replaces the "why is nothing happening?" confusion when most videos are pending or rate-limited. | `internal/web/view.go` (`PipelineStats`, `pipelineStats`) |
| **Unavailable channels (account page)** | The `/account` page shows a "Unavailable channels" section when any channels returned HTTP 404 on the last discovery pass. Lists channel name, an "unavailable" badge, and the first-seen date. Data sourced from the `channel_errors` table (migration 013). | Surfaces silent failures so users know why some subscribed channels produce no new videos. | migration 013; `internal/web/account.go`; `internal/adapters/youtube/youtube.go` (`domain.ErrChannelUnavailable`) |
| **Visual refresh — charm-reader theme + light/dark toggle (ADR-032)** | One layout in two palettes expressed as CSS custom properties: a warm "reader" light theme (sketch B) and a "cozy terminal" dark theme (sketch C). Palette is chosen in cascade order — `:root` light default, an OS-preference dark block scoped to `:root:not([data-theme])` so it only applies absent an explicit choice, and `:root[data-theme="dark"\|"light"]` set by a header toggle that outranks the media query and persists in `localStorage` (guarded; degrades to OS default). An init script in `<head>` applies the stored choice before paint (no flash). Charm touches: monospace meta lines, accent uppercase section dividers with a trailing rule, pill buttons, a lifted/accent-edged expanded card. Error/danger shades became `--err-*` tokens so they follow the theme without per-block dark overrides. Sketches kept as the design record under `docs/sketches/`. | The UI read "flat and boring"; the charm/TUI aesthetic makes it distinctive and gives a real light/dark choice rather than OS-only. | ADR-032; `docs/use-cases/visual_theme.feature`; `internal/web/visual_theme_test.go`; `internal/web/view.go` (`stylesheet`, `themeScript`), `internal/web/views.templ` (Layout/PublicLayout) |
| **Recency window + sparse-state honesty (ADR-020)** | Supersedes the copy/sort in the rows above. Auto-summarize is bounded to videos published within `TAPIR_AUTO_SUMMARIZE_WINDOW` (~7d); older un-summarized videos collapse behind a single "Show N older videos — summarize on demand" disclosure, and caption-less videos collapse to a one-line count (not N cards). List order is now `summarized-first, published_at DESC NULLS LAST`. Copy reframed for honest scarcity: pipeline bar reads "N ready · M in queue · K no captions" (no "fetching soon"); a gradual-fill note explains the rate limit; the nudge verb is "Summarize" (not "Summarize now"); the queued card says "summarizing shortly"; the empty-connected state drops the impossible `tapir run` instruction. Detail leads with Takeaways. Filters slimmed (no date pickers; hidden when empty); watched/skipped segmented; back link on detail; empty terms checkbox removed. | Make the sparse reality legible and honest instead of implying abundance/imminence; bound auto load so the back-catalogue doesn't re-drive the caption gate. Never fetch harder — scarcity is surfaced, not engineered around. | ADR-020; `2384c47`, `3df0459`, `40b703e`, `a1a5217`, `4a0a56e`, `9bf1c31`, `980638d`, `12fb031`, `f775441`, `51aa5d9` |
+13 -1
View File
@@ -31,5 +31,17 @@ Feature: Local-first AI with optional BYO fallback
When any transcript is summarized
Then my content is only ever sent to the local AI stack
# "Reliably" is operationalized as: Primary returned without error within timeout.
Scenario: A model returns unparseable output and the next endpoint succeeds
Given the local AI stack is available
But the primary model returns output that cannot be parsed into a summary
And a fallback model is configured
When a transcript is summarized
Then Tapir falls back to the next model in the chain
And the summary records fallback_used as true
# "Reliably" is operationalized as: an endpoint returned a PARSEABLE summary
# within timeout. A 200 with malformed JSON (or highlights emitted as a bare
# string) counts as a failure and advances the chain (ADR-022). Endpoints are
# tried in order, locals first, so the external worst-case model only ever sees
# content after every local endpoint has failed.
# Quality scoring may be added later without changing these scenarios.
+69
View File
@@ -0,0 +1,69 @@
Feature: Chat with a video's stored transcript
As a reader whose summary made me want to dig deeper
I want to ask questions about the video without watching it
So that I can go further on the ones worth it, without leaving the reader
# ADR-027. The load-bearing constraint is safety-by-construction: chat runs
# ONLY against an already-stored transcript (ADR-021) and never fetches captions,
# never touches the rate gate, never reaches YouTube. Entry is from the summary
# view of one's OWN video; the conversation is ephemeral (no persisted history).
Background:
Given I have a summarized video with a stored transcript
Scenario: A summary view offers a deeper-dive into the video
When I view the summary
Then I see a "dig deeper" affordance that opens a chat about this video
And it opens the chat in place, below the summary, without leaving the page
Scenario: The summary and the chat are on one page
When I open the chat
Then the summary stays visible alongside the chat
And I can read the summary while I ask questions
Scenario: Ask a question answered from the stored transcript
When I ask a question in the chat
Then the answer is produced from the stored transcript
And no caption fetch and no YouTube call occurs
Scenario: Chat never fetches captions or reaches YouTube
When I ask a question in the chat
Then Tapir reads only the stored transcript
And the caption-fetch and video-fetch paths are never invoked
Scenario: A video with no stored transcript offers no chat
Given a video that has no stored transcript
When I open the chat for it
Then I am told chat is not available
And no fetch is attempted and no model is called
Scenario: The default model is the summary's model and is switchable
When I open the chat
Then the model defaults to the model that produced the summary
And I can switch among the offered chain models
Scenario: Switching models re-runs against the same transcript
When I ask a question with a different chain model selected
Then the chosen model answers
And it answers against the same stored transcript
Scenario: The cloud model is hidden when cloud is disabled
Given the cloud fallback model is disabled
Then the chat switcher offers only local models
Scenario: A long transcript is bounded and the chat says so
Given the stored transcript is longer than the model budget
When I ask a question
Then the answer is produced from a bounded portion
And the chat notes that it worked from a bounded portion
Scenario: A multi-turn conversation is ephemeral
When I ask a follow-up question
Then the prior turn is carried into the answer
And nothing about the conversation is written to the database
Scenario: Chat is reachable only from my own summary view
Given another user has a summarized video with a stored transcript
When I try to open the chat for their video
Then I get a not-found response
And no model is called
+24
View File
@@ -10,12 +10,33 @@ Feature: Connect and manage video accounts
And my refresh token is stored only as a secret reference
And my subscriptions are synced
@pending
# Vimeo connect is not built yet (provider label exists; no connect flow or test).
Scenario: Connect a Vimeo account
Given I have no connected video accounts
When I connect my Vimeo account
Then the connection is stored with status "active"
And my subscriptions are synced
Scenario: Connecting an account discovers videos immediately
Given I have no connected video accounts
When I connect my YouTube account
Then a discovery pass for my account is triggered right away
And I do not have to wait for the next scheduled pass to see my videos
Scenario: Connecting summarizes my best recent videos right away
Given I have no connected video accounts
When I connect my YouTube account
Then up to the onboarding cap of my newest likely-good videos are summarized through the rate gate
And videos whose known duration is too short or too long are skipped
And the rest are left to the scheduled recency-bounded pass
Scenario: The onboarding burst summarizes with a stronger model
Given I have no connected video accounts
When I connect my YouTube account
Then the burst summarizes with the stronger onboarding model first
And the standard summarizer chain still follows as a fallback
Scenario: Tokens are never stored in the clear
When I connect any video account
Then no OAuth token value is stored in the database
@@ -28,6 +49,9 @@ Feature: Connect and manage video accounts
And no new videos are watched for that connection
And my existing summaries remain readable
@pending
# Per-provider BYO credential config is not built as a web flow yet (the summarizer
# supports a fallback endpoint, but there is no user-facing BYO setup + its test).
Scenario Outline: BYO AI credential is optional and per-provider
When I configure a BYO provider "<provider>"
Then the credential is stored only as a secret reference
+37
View File
@@ -0,0 +1,37 @@
Feature: Inline-expand summary + Q&A in the list (ADR-031, #16)
As a reader skimming my summaries
I want to open a summary and its Q&A in place in the list
So that I get the full read and follow-up without leaving the list (SPA-like, no page hop)
# HTMX inline-expand, no SPA framework (ADR-031). Each scenario maps to a Go test
# in scenario_coverage_test.go (the BDD name-coverage gate).
Scenario: A summarized card expands to the full summary in place
Given a summarized video in my list
When I expand its card
Then the full summary, highlights, and takeaways are returned as an in-place card fragment, not a full page
Scenario: An expanded card collapses back to the compact card
Given an expanded card
When I collapse it
Then the compact card fragment is returned in its place
Scenario: The expanded card offers the Q&A dock
Given chat is enabled
When a summarized card is expanded
Then the expanded card includes the deeper-dive chat affordance for that video
Scenario: Only a summarized card offers expand
Given a discovered but not-yet-summarized card
When the card is rendered
Then it shows its summarize/queue footer and no expand affordance
Scenario: With JS off the card still reaches the full summary
Given a summarized card
When it is rendered
Then its expand affordance carries an href to the detail page as a no-JS fallback
Scenario: The detail page and the expanded card show the same summary
Given a summarized video
When I view it on the detail page and as an expanded card
Then both render the same summary body (one shared fragment, no drift)
+31
View File
@@ -0,0 +1,31 @@
Feature: Public landing page
As a first-time visitor
I want a public welcome page before I log in
So that I understand what Tapir is and how to get started without hitting a login wall
Scenario: An unauthenticated visit to the root is sent to the welcome page
Given I am not logged in
When I open the root path "/"
Then I am redirected to "/welcome"
Scenario: The welcome page invites an unauthenticated visitor to start
Given I am not logged in
When I open "/welcome"
Then I see a "Get Started" call to action
Scenario: An authenticated user on the welcome page sees their way in and out
Given I am logged in
When I open "/welcome"
Then I see a link to my summaries
And I see a way to log out
@pending
# Behaviour ships (logout redirects to /welcome) but is not unit-tested: logout lives in
# the OIDC Auth impl and StubAuth has no routes to exercise it cheaply.
Scenario: Logging out returns to the welcome page
Given I am logged in
When I log out
Then I am returned to "/welcome"
# /welcome is mounted outside the auth guard so it is reachable without a session;
# the root and all data routes stay behind it (commits around the WelcomePage work).
+46
View File
@@ -0,0 +1,46 @@
Feature: Observability — timing and metrics for performance and UX (ADR-030, #15)
As the maintainer running Tapir for pilot users
I want timing and Prometheus metrics for the activities that drive performance and UX
So that I can see latency, model behaviour, and usage and feed the Stage-0 eval gate
# AI metrics are the priority (ADR-030 R3). Each scenario maps to a Go test in
# test/acceptance/scenario_coverage_test.go (the BDD name-coverage gate).
Scenario: Summarization latency is recorded per endpoint
Given the summarizer runs a transcript through its endpoint chain
When an endpoint returns a parseable summary
Then the summarize latency is recorded with the model, outcome "success", and whether it was a fallback
Scenario: A failing summarizer endpoint records its failure outcome
Given the summarizer runs a transcript through its endpoint chain
When an endpoint errors or returns unparseable output
Then the summarize latency is recorded with outcome "error" or "parse_error" before the chain advances
Scenario: Caption fetch latency is recorded by outcome
Given a caption fetch is attempted for a video
When it resolves to captions, no captions, or a rate limit
Then the caption-fetch latency is recorded labelled by that outcome
Scenario: LLM token usage is recorded from the completion
Given an LLM completion returns a usage block with prompt and completion tokens
When the client finishes the call
Then the prompt and completion tokens are recorded for that model
Scenario: Q&A answer latency is recorded
Given a user asks a question about a video
When the answer is produced from the stored transcript
Then the chat answer latency is recorded for the answering model
Scenario: HTTP requests are counted by route, method, and status
Given the metrics HTTP middleware wraps the app
When a request is served against a registered route
Then it is counted and timed under the bounded route pattern, not the raw path
Scenario: A successful login is counted
Given a user completes the OIDC callback and a session is established
Then the login counter is incremented
Scenario: The metrics endpoint is not on the public app port
Given the service is running
When the public app mux is inspected
Then it exposes no /metrics route metrics are served on the dedicated metrics port only
+30
View File
@@ -0,0 +1,30 @@
Feature: Paste a YouTube URL to summarize any video
As a user
I want to paste a YouTube link and get a summary
So that I can pull the specific video I want now, even from channels I don't follow
Scenario: Paste a valid YouTube URL
Given I am connected
When I paste a valid YouTube video URL
Then the video is added to my feed scoped to me
And it is queued for summarization through the shared rate gate
Scenario: Pasting an invalid link is rejected
When I paste something that is not a YouTube video URL
Then I get a clear error and nothing is added
Scenario: Pasting a video that cannot be found is honest
When I paste a URL whose video cannot be found
Then I am told it couldn't be found and nothing is added
Scenario: Pasting the same video twice does not duplicate it
Given I have pasted a video
When I paste the same video again
Then my feed still has exactly one entry for it
@pending
# Covered by the engine's ADR-010 no-transcript terminal state (degrade-never-error);
# there is no paste-specific test for it.
Scenario: A pasted video with no captions resolves honestly
When I paste a video that has no captions
Then it resolves to the "no transcript available" terminal state
+48
View File
@@ -0,0 +1,48 @@
Feature: Register and manage a multi-user account
As one of a handful of trusted users
I want my own account, isolated from everyone else's
So that Tapir can serve several people from one deployment without leaking data
# Stage 1 (ADR-012): Dex authenticates, Tapir authorizes per user. A Dex subject
# with no users row is a new user and must register before reaching any data.
Scenario: A new Dex subject is routed to registration
Given I am authenticated by Dex with a subject that has no Tapir account
When I open any page that requires an account
Then I am routed to the registration page
And no summaries are shown until I register
Scenario: Registering creates the account and its identity mapping
Given I am authenticated by Dex with a subject that has no Tapir account
When I complete registration
Then a user row is created for me
And a user_identities row maps my Dex subject to that user
And I am taken into the app as a registered user
Scenario: A returning subject passes straight through
Given I am authenticated by Dex with a subject that already has a Tapir account
When I open the app
Then I am not asked to register again
And I see my own summaries
Scenario: Deleting an account removes only my data and leaves other users untouched
Given I am a registered user with summaries, a connected account, and recorded actions
And another user exists with their own summaries
When I delete my account
Then all of my rows are removed across every user-owned table
And my stored secret references are removed
And the other user's data remains intact
And my Dex identity is left intact
@pending
# Re-registration after delete is supported by design (delete leaves the Dex identity,
# ADR-013) but has no dedicated end-to-end test yet.
Scenario: A deleted user can register again as a fresh account
Given I deleted my Tapir account but my Dex identity still exists
When I sign in again
Then I am routed to the registration page as a new user
And registering creates a fresh user row with none of my old data
# Isolation is DB-enforced (Postgres RLS, ADR-012, migration 003): a user can never
# read or write another user's rows even if an application WHERE clause is wrong.
# Deletion is Tapir-side only — the shared Dex directory is never modified (ADR-013).
+47
View File
@@ -0,0 +1,47 @@
Feature: Choose how new videos get summarized
As a user who wants control over compute and noise
I want to pick whether new videos are summarized automatically or on demand
So that I only spend summarization on the videos I actually care about
Background:
Given I am a registered user with a connected video account
Scenario: Auto mode summarizes recent new videos automatically
Given my summarization mode is "auto"
When a subscribed channel posts a new video with captions within the recency window
Then Tapir summarizes it without my asking
And the summary appears in my list
Scenario: Auto mode lists older videos without summarizing them
Given my summarization mode is "auto"
When discovery finds a video published before the recency window
Then the video appears in my list with no summary
And it is not summarized automatically
And I can still summarize it on demand with "Summarize"
Scenario: Automatic is the default for a new user
Given I have just registered
Then my summarization mode is "auto"
Scenario: Manual mode leaves new videos unsummarized
Given my summarization mode is "manual"
When a subscribed channel posts a new video with captions
Then the video appears in my list with no summary
And nothing is summarized until I request it
Scenario: Requesting a summary in manual mode queues it for the next run
Given my summarization mode is "manual"
And a new video is in my list with no summary
When I click "Summarize" on that video
Then the video is marked as requested
And the next run summarizes it
And the request flag is cleared after it is processed
# auto_summarize is a per-user setting and summarize_requested is a per-video queue
# flag (migration 006). The web button sets the flag; `tapir run` processes both the
# auto videos and the manually queued ones, then clears the flag.
#
# Recency bound (ADR-020): in auto mode the scheduler only summarizes videos published
# within TAPIR_AUTO_SUMMARIZE_WINDOW (default ~7d); older videos are discovered and
# listed but wait for an explicit "Summarize" — so a back-catalogue does not re-drive
# the per-IP caption gate (ADR-014) every cycle. A manual request bypasses the bound.
@@ -31,5 +31,16 @@ Feature: Summarize new videos from subscribed channels
When the watcher sees "Designing for Attention" again
Then Tapir does not produce a second summary for it
Scenario: Re-analyzing a stored video does not re-fetch its transcript
Given a transcript for "Designing for Attention" is already stored
When the video is summarized again
Then Tapir reads the stored transcript
And Tapir does not fetch captions from YouTube
# Captions-first is the core path (ADR-007). Audio-download + speech-to-text is
# deferred and intentionally has no scenario here yet.
#
# Transcript persistence (ADR-021): the stored transcript is shared, keyed by
# (provider, provider_video_id) and read before any caption fetch, so the
# re-analysis scenario above also covers paste-a-URL and the onboarding burst —
# both summarize through the same engine chokepoint.
+29
View File
@@ -0,0 +1,29 @@
Feature: Visual refresh — one charm-reader layout, light + dark themes (ADR-032, #17)
As a reader who likes a TUI/charm aesthetic
I want a fresh look with a light and a dark theme
So that the app feels distinctive and reads well in either mode
# Direction B (light) + C (dark) are one layout in two palettes (CSS variables),
# plus a persisted toggle and an OS-preference default. Colours are reviewed via
# the mockups in docs/sketches/, not unit-tested; these scenarios cover the
# testable structure. Each maps to a Go test in scenario_coverage_test.go.
Scenario: Light and dark themes share one layout via CSS variables
Given the app stylesheet
When it is rendered
Then it defines a light palette on the root and a dark palette under data-theme="dark", with no duplicate markup
Scenario: Without a stored choice the theme follows the OS preference
Given a visitor with no saved theme
When the page loads
Then a prefers-color-scheme dark block applies the dark palette automatically
Scenario: A persisted toggle switches light and dark
Given any page
When it is rendered
Then it includes a theme-toggle control and a small script that flips data-theme and persists the choice
Scenario: The expanded card embeds the video player
Given a summarized video with a valid provider id
When its card is expanded
Then the expanded card includes the embedded video player
+207
View File
@@ -0,0 +1,207 @@
# Tapir — Heuristic Review (Stage-0, sparse + recency-bounded)
_Findings document, not a build spec. The maintainer filters; a spec follows separately.
Reviewed against `VISION.md` (Stage-0 gate), `docs/ui-spec.md`, `internal/web/views.templ` +
`view.go`, the prior `UX-REVIEW.md` pass, and the current screenshots. Written for the product
**as it actually is**: ~283 discovered, ~15 summarized, ~256 in queue behind a respected per-IP
caption rate limit, ~12 no-captions; single-user (maintainer) with friends pending; recency-
bounded auto-summarize about to ship (auto = recent ~7d, older browsable + manual on demand)._
## Reviewer stance
The Stage-0 gate is **return usage**. So every finding is judged by one question: does this make
the maintainer (or a friend) come back to an honestly-sparse feed? The visual layer is already
decent — dark mode, cards, the constrained reader were all fixed in the prior pass. The open
problems are **expectation-setting, honesty-of-scale, and the recency feed** — not pixels.
**The single biggest risk:** a new user connects, sees `15 summarized · 256 fetching soon`,
nothing visibly moves, and never returns. The whole gate dies at that moment. Most P0s below
attack that one moment.
## Tagging
- **`[NOW]`** — improves the product as it is today (sparse, recency-bounded, single-user). Ships
in the upcoming bundle.
- **`[LATER]`** — improves the product we hope it becomes (abundance, engagement, multiple users).
Valuable but premature until real usage validates the core loop.
Severity: 🔴 breaks the core loop · 🟠 hurts it · 🟡 noticeable · 🔵 polish.
---
## P0 — fix before/with the recency ship
### 1. 🔴 Empty-connected state tells the user to run a CLI command they can't run — `[NOW]`
**Problem.** After connecting YouTube, the empty list says: *"Run `tapir run` to discover your
subscriptions."* A friend on the web has no shell. And post-ADR-018 discovery is an in-process
scheduled loop — so the instruction is wrong *even for the maintainer*. It is the first thing a
newly-onboarded user sees, and it is an impossible, stale instruction.
**Principle.** Match between system and the real world; help users recognize, not be blocked
(Nielsen #1, #2, #9).
**Proposal.** Replace with a passive, honest "we're working" state: *"Your account is connected.
Tapir is finding your subscriptions and fetching captions — summaries appear here gradually. Check
back later."* No command. No imperative the user can't satisfy.
**Evidence.** `views.templ` `summaryList``.empty-connected`.
### 2. 🔴 No expectation set for *gradual* fill — the return loop breaks at the cliff — `[NOW]`
**Problem.** Captions are rate-limited by design; the backlog trickles over days. Nothing tells
the user this. A first visit shows few/no summaries and no "come back" framing. The Stage-0 gate
is literally about returns, and the product never asks for one or explains why patience is
warranted.
**Principle.** Visibility of system status (#1). And: the gate can't be cleared if the UX doesn't
survive first contact.
**Proposal.** One honest sentence near the pipeline bar / empty state: *"Tapir fetches captions
slowly on purpose, to respect YouTube's limits. New summaries land gradually — usually best to
check back tomorrow."* Turns confusing emptiness into intentional design. Highest-leverage change
in the review.
### 3. 🟠 "256 fetching soon" overstates imminence — a lie of scale — `[NOW]`
**Problem.** The pipeline bar and rate-limited cards both say "fetching soon." For 256 items
behind a per-IP throttle, "soon" is false — they trickle over days/weeks. Exactly the abundance-
implying language the mandate forbids, inverted: it makes the *queue* look imminent.
**Principle.** Honesty of scarcity (project mandate); #1.
**Proposal.** Relabel by scale. Bar: `15 ready · 256 in queue · 12 no captions`. Card state:
"In queue" / "Waiting its turn", not "Fetching soon…". Reserve "soon" for items actually next.
**Evidence.** `views.templ` `pipelineBar`; card State 4.
### 4. 🟠 Recency boundary is invisible in the feed — `[NOW, ships with recency]`
**Problem.** Once auto-summarize is bounded to ~7d, a 6-month-old pending video and a 2-day-old
pending video render identically ("Not summarized" + button). But only one is in the auto path;
the other will *never* process unless clicked. The user can't tell "be patient, this is coming"
from "this is yours to trigger or ignore."
**Principle.** Visibility of system status; predictability (#1).
**Proposal.** Bucket the list into two sections: **Recent** (auto, will fill itself) and
**Older — browse / summarize on demand**. Card status language should encode *which side of the
line it's on*, not the pipeline internals. Central IA decision of the recency change.
### 5. 🟠 The un-summarized mass buries the ~15 readable summaries — `[NOW]`
**Problem.** ~268 of 283 cards are not readable (pending / queued / no-captions). Summarized-first
ordering helps, but the page is still 95% noise below the fold. "Attention is the scarce resource"
is the product's own principle — and the default view violates it.
**Principle.** Aesthetic/minimalist design; signal-to-noise (#8).
**Proposal.** Default view = readable summaries + the Recent bucket. Collapse the older
un-summarized mass behind *"Show 256 older un-summarized videos."* Collapse the 12 no-caption
videos into a single line: *"12 videos have no captions"* (terminal, never readable — they don't
deserve 12 full cards).
---
## P1 — high, near-term
### 6. 🟠 "Summarize now" overpromises against the rate gate — `[NOW]`
**Problem.** Button says "Summarize now" (title: "Summarize this one now"). Backend queues it
behind the shared per-IP gate. Click 10 older videos and they all sit at "Fetching soon…". The
verb sells immediacy the system can't honor.
**Principle.** Honesty; match system/reality (#1).
**Proposal.** Drop "now" → "Summarize". On click the card should honestly become "Queued" (it
already can). Optionally show queue position once the queue is real. Don't engineer the gate
harder — just stop the verb from lying.
**Evidence.** `views.templ` card States 4 & 5; `handlers.go` `handleRetryNow`,
`handleRequestSummarize`.
### 7. 🟠 Welcome-page copy is stale and misleading post-ADR-019 — `[NOW]`
**Problem.** Landing sub-copy: *"If you have an invite link, it will set up your account
automatically."* Invites moved to Authentik (ADR-019); Tapir no longer handles invite links. The
"Get Started" button goes straight to OIDC. The copy promises a flow that no longer exists.
**Principle.** Match between system and reality (#2); honesty.
**Proposal.** Rewrite: *"Tapir is invite-only right now. If you've been invited, sign in below."*
Single CTA. Also set the gradual-fill expectation here (ties to #2) so it lands before the wall,
not after.
**Evidence.** `views.templ` `WelcomePage``.welcome-sub`.
### 8. 🟡 "Queued · waiting for the next run" leaks system jargon — `[NOW]`
**Problem.** "the next run" exposes the discovery-loop concept; a user doesn't know what a "run"
is.
**Principle.** Speak the user's language (#2).
**Proposal.** "Queued — summarizing shortly." Hide the scheduler.
**Evidence.** `views.templ` card State 3.
### 9. 🟡 Filters render before there's anything to filter — `[NOW]`
**Problem.** The fresh/empty state shows the full Channel/From/To/Filter bar *above* "Connect
YouTube." Power tooling stacked on top of the one action that matters.
**Principle.** Progressive disclosure; minimalist design (#8).
**Proposal.** Hide the filter bar when there are 0 rows (and arguably below ~20). Show the connect
CTA alone.
**Evidence.** `fixes/03-empty-fresh.png`; `ListPage` renders `filterForm` unconditionally.
### 10. 🟡 Date-range filters are dead weight at this scale — `[NOW]` demote / `[LATER]` rebuild
**Problem.** From/To date pickers + a free-text exact-match Channel field are corpus-scale tools.
With 15 summaries they're noise; channel-as-freetext is unguessable. (Confirm the prior review's
#7 is resolved — that `Channel` shows a real channel name, not the `provider` string; seeded
screenshots suggest it is, live data may differ.)
**Principle.** Match tool to task; minimalist design.
**Proposal.** `[NOW]`: reduce to the "Summarized only" toggle (+ maybe channel chips derived from
present rows). Drop date pickers until the corpus justifies them. `[LATER]`: real channel facets +
search when there's volume.
---
## P2 — reading experience & polish
### 11. 🟡 Detail leads with Summary; the attention-saving payload (Takeaways) is last — `[NOW]`
**Problem.** Product promise is "decide what's worth your time." The element that answers that —
Takeaways / verdict — sits at the bottom. The user reads a full summary to reach the point.
**Principle.** Lead with the user's actual job-to-be-done.
**Proposal.** Reorder or add a one-line TL;DR/verdict at top. Takeaways → Highlights → Summary, or
a "Worth watching?" lede. Data already exists; a reorder, not new machinery.
**Evidence.** `views.templ` `DetailPage`.
### 12. 🔵 No "back to Summaries" on detail — `[NOW]`
**Problem.** Only the brand returns home, losing any filter context.
**Proposal.** Explicit "← Summaries" link on the detail page.
### 13. 🔵 Auto/manual toggle copy will be wrong after recency — `[NOW, with recency]`
**Problem.** Account copy: *"Automatic summarizes every new video as it is discovered."* Becomes
false once auto is bounded to ~7d.
**Proposal.** *"Automatic summarizes new videos from the last ~7 days. Older videos stay
browsable — summarize them on demand."*
**Evidence.** `views.templ` `AccountPage` summarization section.
### 14. 🔵 Register step asks acceptance of nonexistent terms — `[NOW]`
**Problem.** "I accept the terms of use" — no terms linked. For a friends-only tool, ceremony
accepting nothing.
**Proposal.** Either link real terms or drop the checkbox at this stage.
**Evidence.** `views.templ` `RegisterPage`.
### 15. 🔵 watched/skipped mutual exclusivity unsignaled — `[NOW]` low
**Problem.** Three independent-looking buttons; watched↔skipped are exclusive (carryover from
prior review #14).
**Proposal.** Segmented control for watched/skipped; keep Saved separate.
---
## `[LATER]` — premature until the core loop is validated
_(abundance / engagement / multiple users)_
- **Return-nudges (digest email / push).** 🔴 **Caution, not just defer.** A notification that
drives returns *contaminates the exact signal the gate measures* — VISION wants *unprompted*
returns and admits it can't distinguish prompted from organic. Building a nudge now poisons the
experiment. Defer until after the gate reads. `[LATER]`
- **Full-text search across summaries** — needs volume to matter. `[LATER]`
- **Channel facets / saved filters / sorting** — corpus-scale tooling (#10). `[LATER]`
- **Read/unread + "new since last visit"** — genuinely helps returns, but only meaningful once
there is throughput to be "new." `[LATER]`
- **Saved/queue view, collections** — engagement surface; no payoff at 15 items. `[LATER]`
- **Richer card previews (top-takeaway as preview, 23 lines)** — triage aid that only pays off
with many cards to triage. `[LATER]`
- **Backlog progress heartbeat ("N summarized this week", queue burn-down)** — rewards returning,
but needs real throughput to show motion; a static "last updated X ago" is the only `[NOW]`-worthy
slice. `[LATER]`
- **Onboarding tour / multi-step welcome** — over-built for one user + a few friends. `[LATER]`
---
## What's already right (don't regress)
Pipeline-bar concept, the honest "No transcript available" terminal state, the queued/waiting
spinner instead of a fake-imminent one, the Unavailable-channels surface, the constrained-width
reader, dark mode, the account danger-zone behind a disclosure. The honesty instincts are present
— the P0s are about making that honesty *legible and correctly-scaled*, not adding it.
## Two judgment calls to settle first
1. **#4 (Recent/Older split) and #5 (collapse the older mass) are one decision viewed twice.**
Settle the recency-feed IA once and both fall out.
2. **The return-nudge caution (`[LATER]` list, item 1) is the one to put in writing now** — before
someone "helpfully" ships an email digest to juice the gate and destroys the signal.
Binary file not shown.

After

Width:  |  Height:  |  Size: 116 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 73 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 58 KiB

+13 -1
View File
@@ -1,4 +1,4 @@
module gitea.d-ma.be/mathias/tapir
module git.d-ma.be/mathias/tapir
go 1.26.1
@@ -9,21 +9,33 @@ require (
github.com/go-jose/go-jose/v4 v4.1.4
github.com/golang-migrate/migrate/v4 v4.19.1
github.com/jackc/pgx/v5 v5.9.2
github.com/prometheus/client_golang v1.23.2
github.com/prometheus/client_model v0.6.2
github.com/stretchr/testify v1.11.1
golang.org/x/oauth2 v0.36.0
golang.org/x/time v0.15.0
)
require (
github.com/beorn7/perks v1.0.1 // indirect
github.com/cespare/xxhash/v2 v2.3.0 // indirect
github.com/davecgh/go-spew v1.1.2-0.20180830191138-d8f796af33cc // indirect
github.com/jackc/pgerrcode v0.0.0-20220416144525-469b46aa5efa // indirect
github.com/jackc/pgpassfile v1.0.0 // indirect
github.com/jackc/pgservicefile v0.0.0-20240606120523-5a60cdf6a761 // indirect
github.com/jackc/puddle/v2 v2.2.2 // indirect
github.com/kylelemons/godebug v1.1.0 // indirect
github.com/lib/pq v1.10.9 // indirect
github.com/munnerz/goautoneg v0.0.0-20191010083416-a7dc8b61c822 // indirect
github.com/pmezard/go-difflib v1.0.1-0.20181226105442-5d4384ee4fb2 // indirect
github.com/prometheus/common v0.66.1 // indirect
github.com/prometheus/procfs v0.16.1 // indirect
github.com/rogpeppe/go-internal v1.15.0 // indirect
github.com/xi2/xz v0.0.0-20171230120015-48954b6210f8 // indirect
go.yaml.in/yaml/v2 v2.4.2 // indirect
golang.org/x/sync v0.18.0 // indirect
golang.org/x/sys v0.41.0 // indirect
golang.org/x/text v0.31.0 // indirect
google.golang.org/protobuf v1.36.8 // indirect
gopkg.in/yaml.v3 v3.0.1 // indirect
)
+28 -4
View File
@@ -4,6 +4,10 @@ github.com/Microsoft/go-winio v0.6.2 h1:F2VQgta7ecxGYO8k3ZZz3RS8fVIXVxONVUPlNERo
github.com/Microsoft/go-winio v0.6.2/go.mod h1:yd8OoFMLzJbo9gZq8j5qaps8bJ9aShtEA8Ipt1oGCvU=
github.com/a-h/templ v0.3.1020 h1:ypAT/L5ySWEnZ6Zft/5yfoWXYYkhFNvEFOeeqecg4tw=
github.com/a-h/templ v0.3.1020/go.mod h1:A2DlK61v+K+NRoGnhmYbNYVmtYHcFO5/AisMvBdDxTM=
github.com/beorn7/perks v1.0.1 h1:VlbKKnNfV8bJzeqoa4cOKqO6bYr3WgKZxO8Z16+hsOM=
github.com/beorn7/perks v1.0.1/go.mod h1:G2ZrVWU2WbWT9wwq4/hrbKbnv/1ERSJQ0ibhJ6rlkpw=
github.com/cespare/xxhash/v2 v2.3.0 h1:UL815xU9SqsFlibzuggzjXhog7bL6oX9BbNZnL2UFvs=
github.com/cespare/xxhash/v2 v2.3.0/go.mod h1:VGX0DQ3Q6kWi7AoAeZDth3/j3BFtOZR5XLFGgcrjCOs=
github.com/containerd/errdefs v1.0.0 h1:tg5yIfIlQIrxYtu9ajqY42W3lpS19XqdxRQeEwYG8PI=
github.com/containerd/errdefs v1.0.0/go.mod h1:+YBYIdtsnF4Iw6nWZhJcqGSg/dwvV7tyJ/kCkyJ2k+M=
github.com/containerd/errdefs/pkg v0.3.0 h1:9IKJ06FvyNlexW690DXuQNx2KA2cUJXx151Xdx3ZPPE=
@@ -37,8 +41,8 @@ github.com/gogo/protobuf v1.3.2 h1:Ov1cvc58UF3b5XjBnZv7+opcTcQFZebYjWzi34vdm4Q=
github.com/gogo/protobuf v1.3.2/go.mod h1:P1XiOD3dCwIKUDQYPy72D8LYyHL2YPYrpS2s69NZV8Q=
github.com/golang-migrate/migrate/v4 v4.19.1 h1:OCyb44lFuQfYXYLx1SCxPZQGU7mcaZ7gH9yH4jSFbBA=
github.com/golang-migrate/migrate/v4 v4.19.1/go.mod h1:CTcgfjxhaUtsLipnLoQRWCrjYXycRz/g5+RWDuYgPrE=
github.com/google/go-cmp v0.6.0 h1:ofyhxvXcZhMsU5ulbFiLKl/XBFqE1GSq7atu8tAmTRI=
github.com/google/go-cmp v0.6.0/go.mod h1:17dUlkBOakJ0+DkrSSNjCkIjxS6bF9zb3elmeNGIjoY=
github.com/google/go-cmp v0.7.0 h1:wk8382ETsv4JYUZwIsn6YpYiWiBsYLSJiTsyBybVuN8=
github.com/google/go-cmp v0.7.0/go.mod h1:pXiqmnSA92OHEEa9HXL2W4E7lf9JzCmGVUdgjX3N/iU=
github.com/jackc/pgerrcode v0.0.0-20220416144525-469b46aa5efa h1:s+4MhCQ6YrzisK6hFJUX53drDT4UsSW3DEhKn0ifuHw=
github.com/jackc/pgerrcode v0.0.0-20220416144525-469b46aa5efa/go.mod h1:a/s9Lp5W7n/DD0VrVoyJ00FbP2ytTPDVOivvn2bMlds=
github.com/jackc/pgpassfile v1.0.0 h1:/6Hmqy13Ss2zCq62VdNG8tM1wchn8zjSGOBJ6icpsIM=
@@ -49,10 +53,14 @@ github.com/jackc/pgx/v5 v5.9.2 h1:3ZhOzMWnR4yJ+RW1XImIPsD1aNSz4T4fyP7zlQb56hw=
github.com/jackc/pgx/v5 v5.9.2/go.mod h1:mal1tBGAFfLHvZzaYh77YS/eC6IX9OWbRV1QIIM0Jn4=
github.com/jackc/puddle/v2 v2.2.2 h1:PR8nw+E/1w0GLuRFSmiioY6UooMp6KJv0/61nB7icHo=
github.com/jackc/puddle/v2 v2.2.2/go.mod h1:vriiEXHvEE654aYKXXjOvZM39qJ0q+azkZFrfEOc3H4=
github.com/kr/pretty v0.3.0 h1:WgNl7dwNpEZ6jJ9k1snq4pZsg7DOEN8hP9Xw0Tsjwk0=
github.com/kr/pretty v0.3.0/go.mod h1:640gp4NfQd8pI5XOwp5fnNeVWj67G7CFk/SaSQn7NBk=
github.com/klauspost/compress v1.18.0 h1:c/Cqfb0r+Yi+JtIEq73FWXVkRonBlf0CRNYc8Zttxdo=
github.com/klauspost/compress v1.18.0/go.mod h1:2Pp+KzxcywXVXMr50+X0Q/Lsb43OQHYWRCY2AiWywWQ=
github.com/kr/pretty v0.3.1 h1:flRD4NNwYAUpkphVc1HcthR4KEIFJ65n8Mw5qdRn3LE=
github.com/kr/pretty v0.3.1/go.mod h1:hoEshYVHaxMs3cyo3Yncou5ZscifuDolrwPKZanG3xk=
github.com/kr/text v0.2.0 h1:5Nx0Ya0ZqY2ygV366QzturHI13Jq95ApcVaJBhpS+AY=
github.com/kr/text v0.2.0/go.mod h1:eLer722TekiGuMkidMxC/pM04lWEeraHUUmBw8l2grE=
github.com/kylelemons/godebug v1.1.0 h1:RPNrshWIDI6G2gRW9EHilWtl7Z6Sb1BR0xunSBf0SNc=
github.com/kylelemons/godebug v1.1.0/go.mod h1:9/0rRGxNHcop5bhtWyNeEfOS8JIWk580+fNqagV/RAw=
github.com/lib/pq v1.10.9 h1:YXG7RB+JIjhP29X+OtkiDnYaXQwpS4JEWq7dtCCRUEw=
github.com/lib/pq v1.10.9/go.mod h1:AlVN5x4E4T544tWzH6hKfbfQvm3HdbOxrmggDNAPY9o=
github.com/moby/docker-image-spec v1.3.1 h1:jMKff3w6PgbfSa69GfNg+zN/XLhfXJGnEx3Nl2EsFP0=
@@ -61,6 +69,8 @@ github.com/moby/term v0.5.0 h1:xt8Q1nalod/v7BqbG21f8mQPqH+xAaC9C3N3wfWbVP0=
github.com/moby/term v0.5.0/go.mod h1:8FzsFHVUBGZdbDsJw/ot+X+d5HLUbvklYLJ9uGfcI3Y=
github.com/morikuni/aec v1.0.0 h1:nP9CBfwrvYnBRgY6qfDQkygYDmYwOilePFkwzv4dU8A=
github.com/morikuni/aec v1.0.0/go.mod h1:BbKIizmSmc5MMPqRYbxO4ZU0S0+P200+tUnFx7PXmsc=
github.com/munnerz/goautoneg v0.0.0-20191010083416-a7dc8b61c822 h1:C3w9PqII01/Oq1c1nUAm88MOHcQC9l5mIlSMApZMrHA=
github.com/munnerz/goautoneg v0.0.0-20191010083416-a7dc8b61c822/go.mod h1:+n7T8mK8HuQTcFwEeznm/DIxMOiR9yIdICNftLE1DvQ=
github.com/opencontainers/go-digest v1.0.0 h1:apOUWs51W5PlhuyGyz9FCeeBIOUDA/6nW8Oi/yOhh5U=
github.com/opencontainers/go-digest v1.0.0/go.mod h1:0JzlMkj0TRzQZfJkVvzbP0HBR3IKzErnv2BNG4W4MAM=
github.com/opencontainers/image-spec v1.1.0 h1:8SG7/vwALn54lVB/0yZ/MMwhFrPYtpEHQb2IpWsCzug=
@@ -70,6 +80,14 @@ github.com/pkg/errors v0.9.1/go.mod h1:bwawxfHBFNV+L2hUp1rHADufV3IMtnDRdf1r5NINE
github.com/pmezard/go-difflib v1.0.0/go.mod h1:iKH77koFhYxTK1pcRnkKkqfTogsbg7gZNVY4sRDYZ/4=
github.com/pmezard/go-difflib v1.0.1-0.20181226105442-5d4384ee4fb2 h1:Jamvg5psRIccs7FGNTlIRMkT8wgtp5eCXdBlqhYGL6U=
github.com/pmezard/go-difflib v1.0.1-0.20181226105442-5d4384ee4fb2/go.mod h1:iKH77koFhYxTK1pcRnkKkqfTogsbg7gZNVY4sRDYZ/4=
github.com/prometheus/client_golang v1.23.2 h1:Je96obch5RDVy3FDMndoUsjAhG5Edi49h0RJWRi/o0o=
github.com/prometheus/client_golang v1.23.2/go.mod h1:Tb1a6LWHB3/SPIzCoaDXI4I8UHKeFTEQ1YCr+0Gyqmg=
github.com/prometheus/client_model v0.6.2 h1:oBsgwpGs7iVziMvrGhE53c/GrLUsZdHnqNwqPLxwZyk=
github.com/prometheus/client_model v0.6.2/go.mod h1:y3m2F6Gdpfy6Ut/GBsUqTWZqCUvMVzSfMLjcu6wAwpE=
github.com/prometheus/common v0.66.1 h1:h5E0h5/Y8niHc5DlaLlWLArTQI7tMrsfQjHV+d9ZoGs=
github.com/prometheus/common v0.66.1/go.mod h1:gcaUsgf3KfRSwHY4dIMXLPV0K/Wg1oZ8+SbZk/HH/dA=
github.com/prometheus/procfs v0.16.1 h1:hZ15bTNuirocR6u0JZ6BAHHmwS1p8B4P6MRqxtzMyRg=
github.com/prometheus/procfs v0.16.1/go.mod h1:teAbpZRB1iIAJYREa1LsoWUXykVXA1KlTmWl8x/U+Is=
github.com/rogpeppe/go-internal v1.15.0 h1:D0RCU5rMAp+SpgkiNdrjfJ+LX4J1M32V2NeCY7EJ6hc=
github.com/rogpeppe/go-internal v1.15.0/go.mod h1:DrUVZyrJU+txYW5/1kwtXQSMFio52ZOxX7yM1VHvnxs=
github.com/stretchr/objx v0.1.0/go.mod h1:HFkY916IF+rwdDfMAkV7OtwuqBVzrE8GR6GFx+wExME=
@@ -91,6 +109,8 @@ go.opentelemetry.io/otel/trace v1.37.0 h1:HLdcFNbRQBE2imdSEgm/kwqmQj1Or1l/7bW6mx
go.opentelemetry.io/otel/trace v1.37.0/go.mod h1:TlgrlQ+PtQO5XFerSPUYG0JSgGyryXewPGyayAWSBS0=
go.uber.org/goleak v1.3.0 h1:2K3zAYmnTNqV73imy9J1T3WC+gmCePx2hEGkimedGto=
go.uber.org/goleak v1.3.0/go.mod h1:CoHD4mav9JJNrW/WLlf7HGZPjdw8EucARQHekz1X6bE=
go.yaml.in/yaml/v2 v2.4.2 h1:DzmwEr2rDGHl7lsFgAHxmNz/1NlQ7xLIrlN2h5d1eGI=
go.yaml.in/yaml/v2 v2.4.2/go.mod h1:081UH+NErpNdqlCXm3TtEran0rJZGxAYx9hb/ELlsPU=
golang.org/x/oauth2 v0.36.0 h1:peZ/1z27fi9hUOFCAZaHyrpWG5lwe0RJEEEeH0ThlIs=
golang.org/x/oauth2 v0.36.0/go.mod h1:YDBUJMTkDnJS+A4BP4eZBjCqtokkg1hODuPjwiGPO7Q=
golang.org/x/sync v0.18.0 h1:kr88TuHDroi+UVf+0hZnirlk8o8T+4MrK6mr60WkH/I=
@@ -99,6 +119,10 @@ golang.org/x/sys v0.41.0 h1:Ivj+2Cp/ylzLiEU89QhWblYnOE9zerudt9Ftecq2C6k=
golang.org/x/sys v0.41.0/go.mod h1:OgkHotnGiDImocRcuBABYBEXf8A9a87e/uXjp9XT3ks=
golang.org/x/text v0.31.0 h1:aC8ghyu4JhP8VojJ2lEHBnochRno1sgL6nEi9WGFGMM=
golang.org/x/text v0.31.0/go.mod h1:tKRAlv61yKIjGGHX/4tP1LTbc13YSec1pxVEWXzfoeM=
golang.org/x/time v0.15.0 h1:bbrp8t3bGUeFOx08pvsMYRTCVSMk89u4tKbNOZbp88U=
golang.org/x/time v0.15.0/go.mod h1:Y4YMaQmXwGQZoFaVFk4YpCt4FLQMYKZe9oeV/f4MSno=
google.golang.org/protobuf v1.36.8 h1:xHScyCOEuuwZEc6UtSOvPbAT4zRh0xcNRYekJwfqyMc=
google.golang.org/protobuf v1.36.8/go.mod h1:fuxRtAxBytpl4zzqUh6/eyUujkJdNiuEkXntxiD/uRU=
gopkg.in/check.v1 v0.0.0-20161208181325-20d25e280405/go.mod h1:Co6ibVJAznAaIkqp8huTwlJQCZ016jof/cbN4VW5Yz0=
gopkg.in/check.v1 v1.0.0-20201130134442-10cb98267c6c h1:Hei/4ADfdWqJk1ZMxUNpqntNwaWcugrBjAiHlqqRiVk=
gopkg.in/check.v1 v1.0.0-20201130134442-10cb98267c6c/go.mod h1:JHkPIbrfpd72SG/EVd6muEfDQjcINNoR0C8j2r3qZ4Q=
+179
View File
@@ -0,0 +1,179 @@
// Package chat implements the per-video deeper-dive chat (ADR-027): a read-only
// QA over a video's ALREADY-STORED transcript (ADR-021). It is the enforcement
// point for the feature's load-bearing safety property — stored-transcript-only:
// the Service has NO VideoSource and NO caption-fetch dependency, only a
// Completer factory, so it CANNOT reach YouTube or the rate gate by construction.
// The caller supplies the stored transcript text; chat never fetches.
//
// It reuses the same LiteLLM gateway as the summarizer (a chat is a different
// call, not a new integration) and the same transcript-truncation discipline
// (TAPIR_MAX_TRANSCRIPT_CHARS) so a long transcript fits a small-context model.
package chat
import (
"context"
"fmt"
"log/slog"
"strings"
"time"
"unicode/utf8"
"git.d-ma.be/mathias/tapir/internal/metrics"
)
// Completer is the minimal LLM chat surface the Service needs. *llm.Client
// satisfies it; tests use a fake. It is the SAME surface the summarizer uses.
type Completer interface {
Complete(ctx context.Context, system, user string) (string, error)
}
// Turn is one completed exchange in an ephemeral, session-only conversation
// (ADR-027 v1: nothing is persisted).
type Turn struct {
Question string
Answer string
}
// Request is one chat turn: the chosen model, the stored transcript text, the
// prior turns (for multi-turn context within the session), and the new question.
type Request struct {
Model string
Transcript string
History []Turn
Question string
}
// Reply is the model's answer plus whether the transcript was bounded to fit the
// model context (so the UI can be honest that an answer about the tail of a long
// video may be incomplete).
type Reply struct {
Answer string
Truncated bool
}
// Service answers questions against a stored transcript via a switchable set of
// models. models is the ordered, local-first list offered to the user (the cloud
// model is simply absent when disabled — see cmd wiring); maxChars bounds the
// transcript sent to any model (0 = unbounded). newClient builds a Completer for
// a chosen model alias (the same gateway, a different alias).
type Service struct {
newClient func(model string) Completer
models []string
maxChars int
}
// New constructs a Service. models must be non-empty and already filtered to the
// offerable set (cloud excluded when disabled) and de-duplicated by the caller.
func New(newClient func(model string) Completer, models []string, maxChars int) *Service {
return &Service{newClient: newClient, models: models, maxChars: maxChars}
}
// Models returns a copy of the offerable model list (local-first order).
func (s *Service) Models() []string {
out := make([]string, len(s.models))
copy(out, s.models)
return out
}
// offers reports whether model is in the offerable set — the guard that keeps an
// arbitrary, un-offered alias (e.g. a forged form value) from reaching the gateway.
func (s *Service) offers(model string) bool {
for _, m := range s.models {
if m == model {
return true
}
}
return false
}
// DefaultModel resolves the model a fresh chat opens with: the summary's own
// model when it is still an offered option (the ADR-027 default — chat continues
// in the model that produced the summary), otherwise the first offered model.
// Returns "" only when no models are configured.
func (s *Service) DefaultModel(summaryModel string) string {
if summaryModel != "" && s.offers(summaryModel) {
return summaryModel
}
if len(s.models) > 0 {
return s.models[0]
}
return ""
}
// Answer runs one chat turn. The model is forced back to a default if the request
// names an un-offered alias, so chat can never call the gateway with an arbitrary
// model. The transcript is truncated up front (reporting whether it was cut) and
// passed as system context; the running conversation is the user message.
func (s *Service) Answer(ctx context.Context, req Request) (Reply, error) {
if len(s.models) == 0 {
return Reply{}, fmt.Errorf("chat: no models configured")
}
model := req.Model
if !s.offers(model) {
model = s.DefaultModel("")
}
transcript, truncated := truncate(req.Transcript, s.maxChars)
system := buildSystem(transcript, truncated)
user := buildUser(req.History, req.Question)
start := time.Now()
out, err := s.newClient(model).Complete(ctx, system, user)
if err != nil {
return Reply{}, fmt.Errorf("chat: %s: %w", model, err)
}
dur := time.Since(start)
metrics.ObserveChat(model, dur)
slog.Default().Info("chat answer", "model", model, "elapsed_ms", dur.Milliseconds())
answer := strings.TrimSpace(out)
if answer == "" {
return Reply{}, fmt.Errorf("chat: %s returned an empty answer", model)
}
return Reply{Answer: answer, Truncated: truncated}, nil
}
const systemPreamble = `You are Tapir, answering questions about ONE video using ONLY the transcript below.
Ground every answer in the transcript. If the transcript does not contain the answer, say so plainly rather than guessing.`
const truncatedNote = `
The transcript below is truncated to fit the model — if a question seems to concern something missing, note it may be beyond the available portion.`
// buildSystem frames the model as a transcript-grounded QA assistant and embeds
// the (possibly truncated) transcript as context.
func buildSystem(transcript string, truncated bool) string {
var b strings.Builder
b.WriteString(systemPreamble)
if truncated {
b.WriteString(truncatedNote)
}
b.WriteString("\n\nTranscript:\n")
b.WriteString(transcript)
return b.String()
}
// buildUser renders the running conversation as the user message: prior turns as
// Q/A pairs followed by the new question. Folding history into one message keeps
// the Completer surface (a single system+user call) unchanged — no new llm method.
func buildUser(history []Turn, question string) string {
var b strings.Builder
for _, t := range history {
fmt.Fprintf(&b, "Q: %s\nA: %s\n\n", t.Question, t.Answer)
}
fmt.Fprintf(&b, "Q: %s", question)
return b.String()
}
// truncate caps content to max bytes on a UTF-8 rune boundary, reporting whether
// it cut. It mirrors the summarizer's truncation discipline (ADR-022) but returns
// the cut flag so the chat UI can be honest about a bounded transcript. A
// non-positive max (or content already within budget) returns content unchanged.
func truncate(content string, max int) (string, bool) {
if max <= 0 || len(content) <= max {
return content, false
}
cut := max
for cut > 0 && !utf8.RuneStart(content[cut]) {
cut--
}
return content[:cut], true
}
+164
View File
@@ -0,0 +1,164 @@
package chat
import (
"context"
"errors"
"strings"
"testing"
)
// recordingCompleter captures the system+user it was asked with and returns a
// canned answer (or error). It also records which model alias built it.
type recordingCompleter struct {
model string
lastSystem string
lastUser string
answer string
err error
calls *int
}
func (c *recordingCompleter) Complete(_ context.Context, system, user string) (string, error) {
*c.calls++
c.lastSystem = system
c.lastUser = user
if c.err != nil {
return "", c.err
}
return c.answer, nil
}
// factory builds a recordingCompleter per model and records the last one built so
// the test can assert which model alias was actually used for the gateway call.
type factory struct {
answer string
err error
calls int
used *recordingCompleter
}
func (f *factory) make(model string) Completer {
c := &recordingCompleter{model: model, answer: f.answer, err: f.err, calls: &f.calls}
f.used = c
return c
}
func TestModelsAreOfferedLocalFirstAndCopied(t *testing.T) {
f := &factory{answer: "ok"}
s := New(f.make, []string{"koala/phi4-mini", "iguana/gemma4-26b"}, 0)
got := s.Models()
want := []string{"koala/phi4-mini", "iguana/gemma4-26b"}
if len(got) != len(want) || got[0] != want[0] || got[1] != want[1] {
t.Fatalf("Models() = %v, want %v", got, want)
}
// Mutating the returned slice must not corrupt the Service's list.
got[0] = "tampered"
if s.Models()[0] != "koala/phi4-mini" {
t.Fatal("Models() leaked its backing slice")
}
}
func TestDefaultModelIsTheSummarysModelWhenOffered(t *testing.T) {
f := &factory{answer: "ok"}
s := New(f.make, []string{"koala/phi4-mini", "iguana/gemma4-26b", "berget/mistral-small"}, 0)
if got := s.DefaultModel("iguana/gemma4-26b"); got != "iguana/gemma4-26b" {
t.Fatalf("DefaultModel(summary) = %q, want the summary's model", got)
}
// A summary model no longer offered (e.g. cloud disabled) falls back to first.
if got := s.DefaultModel("berget/old-model"); got != "koala/phi4-mini" {
t.Fatalf("DefaultModel(un-offered) = %q, want the first offered model", got)
}
// No summary model recorded → first offered.
if got := s.DefaultModel(""); got != "koala/phi4-mini" {
t.Fatalf("DefaultModel(\"\") = %q, want the first offered model", got)
}
}
func TestAnswerGroundsOnTranscriptAndCarriesHistory(t *testing.T) {
f := &factory{answer: " The video is about attention. "}
s := New(f.make, []string{"koala/phi4-mini"}, 0)
reply, err := s.Answer(context.Background(), Request{
Model: "koala/phi4-mini",
Transcript: "ATTENTION-TRANSCRIPT-MARKER",
History: []Turn{{Question: "who", Answer: "the host"}},
Question: "what is it about",
})
if err != nil {
t.Fatalf("Answer: %v", err)
}
if reply.Answer != "The video is about attention." {
t.Fatalf("answer not trimmed: %q", reply.Answer)
}
if reply.Truncated {
t.Fatal("short transcript must not report truncated")
}
// The transcript rides in the system prompt; the conversation in the user msg.
if !strings.Contains(f.used.lastSystem, "ATTENTION-TRANSCRIPT-MARKER") {
t.Fatal("transcript not grounded into the system prompt")
}
if !strings.Contains(f.used.lastUser, "Q: who") || !strings.Contains(f.used.lastUser, "A: the host") {
t.Fatalf("history not carried into the user message: %q", f.used.lastUser)
}
if !strings.Contains(f.used.lastUser, "what is it about") {
t.Fatal("new question missing from the user message")
}
}
func TestAnswerTruncatesLongTranscriptAndReportsIt(t *testing.T) {
f := &factory{answer: "answer"}
s := New(f.make, []string{"koala/phi4-mini"}, 10)
reply, err := s.Answer(context.Background(), Request{
Model: "koala/phi4-mini",
Transcript: strings.Repeat("x", 500),
Question: "summarize",
})
if err != nil {
t.Fatalf("Answer: %v", err)
}
if !reply.Truncated {
t.Fatal("a transcript past maxChars must report Truncated")
}
if strings.Count(f.used.lastSystem, "x") != 10 {
t.Fatalf("transcript not bounded to maxChars: got %d x's", strings.Count(f.used.lastSystem, "x"))
}
}
func TestAnswerForcesAnUnofferedModelBackToDefault(t *testing.T) {
f := &factory{answer: "answer"}
s := New(f.make, []string{"koala/phi4-mini", "iguana/gemma4-26b"}, 0)
// A forged/un-offered model must never reach the gateway as-is — it is forced
// to the default offered model (the cloud-absent guarantee depends on this).
_, err := s.Answer(context.Background(), Request{
Model: "berget/secret-cloud-model",
Transcript: "t",
Question: "q",
})
if err != nil {
t.Fatalf("Answer: %v", err)
}
if f.used.model != "koala/phi4-mini" {
t.Fatalf("un-offered model reached the gateway as %q, want the default", f.used.model)
}
}
func TestAnswerPropagatesCompleterError(t *testing.T) {
f := &factory{err: errors.New("gateway down")}
s := New(f.make, []string{"koala/phi4-mini"}, 0)
_, err := s.Answer(context.Background(), Request{Model: "koala/phi4-mini", Transcript: "t", Question: "q"})
if err == nil {
t.Fatal("expected the gateway error to propagate")
}
}
func TestAnswerRejectsEmptyModelSet(t *testing.T) {
s := New(func(string) Completer { return nil }, nil, 0)
if _, err := s.Answer(context.Background(), Request{Question: "q"}); err == nil {
t.Fatal("expected an error when no models are configured")
}
}
+38 -2
View File
@@ -32,17 +32,46 @@ type Client struct {
model string
maxTokens int
httpClient *http.Client
usageHook func(model string, prompt, completion int)
}
// Option configures a Client at construction. Variadic so the existing 4-arg
// call sites stay valid as new knobs are added.
type Option func(*Client)
// WithMaxTokens overrides the per-request completion budget. The summarizer uses
// this to cap completion for small-context models (e.g. koala/phi4-mini, 8k):
// with the default 8192 budget, prompt + max_tokens overflows an 8k context and
// the gateway returns HTTP 400. A non-positive n is ignored (keeps the default).
func WithMaxTokens(n int) Option {
return func(c *Client) {
if n > 0 {
c.maxTokens = n
}
}
}
// WithUsageHook registers a callback fired after a successful completion with the
// model and the prompt/completion token counts from the response usage block. It
// keeps this copied, stdlib-only package (ADR-004) decoupled from metrics: the
// caller wires it to internal/metrics, the client imports nothing. nil is ignored.
func WithUsageHook(fn func(model string, prompt, completion int)) Option {
return func(c *Client) { c.usageHook = fn }
}
// New constructs a Client.
func New(baseURL, apiKey, model string, timeout time.Duration) *Client {
return &Client{
func New(baseURL, apiKey, model string, timeout time.Duration, opts ...Option) *Client {
c := &Client{
baseURL: strings.TrimRight(baseURL, "/"),
apiKey: apiKey,
model: model,
maxTokens: defaultMaxTokens,
httpClient: &http.Client{Timeout: timeout},
}
for _, opt := range opts {
opt(c)
}
return c
}
type chatRequest struct {
@@ -61,6 +90,10 @@ type chatResponse struct {
Choices []struct {
Message message `json:"message"`
} `json:"choices"`
Usage struct {
PromptTokens int `json:"prompt_tokens"`
CompletionTokens int `json:"completion_tokens"`
} `json:"usage"`
}
// Complete sends a system + user message and returns the assistant's reply.
@@ -132,5 +165,8 @@ func (c *Client) Complete(ctx context.Context, system, user string) (string, err
if len(cr.Choices) == 0 {
return "", fmt.Errorf("LLM returned no choices")
}
if c.usageHook != nil {
c.usageHook(c.model, cr.Usage.PromptTokens, cr.Usage.CompletionTokens)
}
return cr.Choices[0].Message.Content, nil
}
+45
View File
@@ -64,6 +64,51 @@ func TestClient_SendsMaxTokens(t *testing.T) {
}
}
// TestClient_WithMaxTokens overrides the completion budget — the summarizer caps
// it small so prompt + max_tokens fits a small-context model's window (8k).
func TestClient_WithMaxTokens(t *testing.T) {
var body chatRequest
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
_ = json.NewDecoder(r.Body).Decode(&body)
_ = json.NewEncoder(w).Encode(map[string]any{
"choices": []map[string]any{{"message": map[string]any{"content": "ok"}}},
})
}))
defer srv.Close()
c := New(srv.URL, "", "test-model", 10*time.Second, WithMaxTokens(1500))
if _, err := c.Complete(context.Background(), "sys", "user"); err != nil {
t.Fatalf("Complete: %v", err)
}
if body.MaxTokens != 1500 {
t.Errorf("max_tokens = %d, want 1500", body.MaxTokens)
}
}
// TestClient_UsageHookRecordsTokens: the usage hook fires with the model and the
// prompt/completion token counts parsed from the response usage block.
func TestClient_UsageHookRecordsTokens(t *testing.T) {
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, _ *http.Request) {
_ = json.NewEncoder(w).Encode(map[string]any{
"choices": []map[string]any{{"message": map[string]any{"content": "ok"}}},
"usage": map[string]any{"prompt_tokens": 123, "completion_tokens": 45},
})
}))
defer srv.Close()
var gotModel string
var gotPrompt, gotCompletion int
c := New(srv.URL, "", "test-model", 10*time.Second, WithUsageHook(func(model string, p, comp int) {
gotModel, gotPrompt, gotCompletion = model, p, comp
}))
if _, err := c.Complete(context.Background(), "sys", "user"); err != nil {
t.Fatalf("Complete: %v", err)
}
if gotModel != "test-model" || gotPrompt != 123 || gotCompletion != 45 {
t.Errorf("usage hook got (%q, %d, %d), want (test-model, 123, 45)", gotModel, gotPrompt, gotCompletion)
}
}
func TestClient_ReturnsErrorOnNon200(t *testing.T) {
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
http.Error(w, "overloaded", http.StatusServiceUnavailable)
+1 -1
View File
@@ -18,7 +18,7 @@ import (
"path/filepath"
"sync"
"gitea.d-ma.be/mathias/tapir/internal/ports"
"git.d-ma.be/mathias/tapir/internal/ports"
)
// FileStore is a SecretStore backed by a single 0600 JSON file mapping opaque
+1 -1
View File
@@ -7,7 +7,7 @@ import (
"path/filepath"
"testing"
"gitea.d-ma.be/mathias/tapir/internal/adapters/secrets"
"git.d-ma.be/mathias/tapir/internal/adapters/secrets"
)
func TestPutThenGet(t *testing.T) {
+15 -7
View File
@@ -10,13 +10,17 @@ import (
// DeleteUser permanently removes a user and all of their data. It runs through
// withUser so RLS confines every statement to the calling user's own rows.
//
// Deleting the users row cascades (ON DELETE CASCADE) to videos, transcripts,
// summaries (→ sink_deliveries), video_connections, and the user_identities map
// referential-integrity cascades bypass RLS, so a user's child rows are removed
// even though the deleting connection is scoped. summary_actions is the exception:
// it carries a user_id but has NO foreign key to users (migration 002), so the
// cascade does not reach it; it is deleted explicitly in the same scoped
// transaction. Deleting an absent user is a no-op (idempotent).
// Deleting the users row cascades (ON DELETE CASCADE) to videos, summaries
// (→ sink_deliveries), video_connections, and the user_identities map
// referential-integrity cascades bypass RLS, so a user's child rows are removed
// even though the deleting connection is scoped. Transcripts are NOT removed:
// since ADR-021 they are shared public content keyed by (provider,
// provider_video_id) with no user_id, so another user may still reference the
// same row — a user deletion must not strip shared caption content. summary_actions and login_events
// are the exceptions: each carries a user_id but has NO foreign key to users
// (migrations 002 and 010), so the cascade does not reach them; they are deleted
// explicitly in the same scoped transaction. Deleting an absent user is a no-op
// (idempotent).
//
// This is tapir-side only (decision 2026-06-03): it removes all tapir data; the
// Dex login identity is left untouched — a later login simply re-enters
@@ -28,6 +32,10 @@ func (s *Store) DeleteUser(ctx context.Context, userID string) error {
`DELETE FROM summary_actions WHERE user_id = $1`, userID); err != nil {
return fmt.Errorf("store: delete summary_actions: %w", err)
}
if _, err := tx.Exec(ctx,
`DELETE FROM login_events WHERE user_id = $1`, userID); err != nil {
return fmt.Errorf("store: delete login_events: %w", err)
}
if _, err := tx.Exec(ctx,
`DELETE FROM users WHERE id = $1`, userID); err != nil {
return fmt.Errorf("store: delete user: %w", err)
@@ -0,0 +1,83 @@
package store
import (
"context"
"fmt"
"time"
"github.com/jackc/pgx/v5"
)
// CaptionlessChannels returns the set of channel ids currently suppressed for the
// user — channels whose recent videos all yielded no captions, within their
// suppression window (ADR-024). The runner skips caption fetches for these
// channels' videos. A channel whose window has expired is not returned, so its
// next video is re-probed (auto-recovery).
func (s *Store) CaptionlessChannels(ctx context.Context, userID string) (map[string]bool, error) {
out := map[string]bool{}
err := s.withUser(ctx, userID, func(tx pgx.Tx) error {
rows, err := tx.Query(ctx, `
SELECT channel_id FROM channel_caption_state
WHERE user_id = $1 AND captionless_until IS NOT NULL AND captionless_until > now()`,
userID)
if err != nil {
return fmt.Errorf("store: caption-less channels: %w", err)
}
defer rows.Close()
for rows.Next() {
var ch string
if err := rows.Scan(&ch); err != nil {
return fmt.Errorf("store: scan caption-less channel: %w", err)
}
out[ch] = true
}
return rows.Err()
})
return out, err
}
// RecordChannelCaptionOutcome updates a channel's caption-availability memory
// after a fetch attempt (ADR-024). hadCaptions resets the channel (consecutive
// count to 0, suppression cleared). Otherwise the consecutive no-caption count is
// incremented; once it reaches threshold the channel is suppressed for window.
// threshold <= 0 is a no-op (feature disabled). An empty channelID is ignored
// (some sources may not carry one).
func (s *Store) RecordChannelCaptionOutcome(ctx context.Context, userID, channelID string, hadCaptions bool, threshold int, window time.Duration) error {
if channelID == "" || threshold <= 0 {
return nil
}
return s.withUser(ctx, userID, func(tx pgx.Tx) error {
if hadCaptions {
_, err := tx.Exec(ctx, `
INSERT INTO channel_caption_state (user_id, channel_id, consecutive_none, captionless_until, updated_at)
VALUES ($1, $2, 0, NULL, now())
ON CONFLICT (user_id, channel_id)
DO UPDATE SET consecutive_none = 0, captionless_until = NULL, updated_at = now()`,
userID, channelID)
if err != nil {
return fmt.Errorf("store: reset channel caption state: %w", err)
}
return nil
}
// No captions: increment the streak; suppress once it reaches threshold.
// captionless_until is set from the NEW count inside the same statement so
// the decision is atomic with the increment.
until := time.Now().Add(window)
_, err := tx.Exec(ctx, `
INSERT INTO channel_caption_state (user_id, channel_id, consecutive_none, captionless_until, updated_at)
VALUES ($1, $2, 1, CASE WHEN 1 >= $3 THEN $4::timestamptz ELSE NULL END, now())
ON CONFLICT (user_id, channel_id)
DO UPDATE SET
consecutive_none = channel_caption_state.consecutive_none + 1,
captionless_until = CASE
WHEN channel_caption_state.consecutive_none + 1 >= $3 THEN $4::timestamptz
ELSE channel_caption_state.captionless_until
END,
updated_at = now()`,
userID, channelID, threshold, until)
if err != nil {
return fmt.Errorf("store: record channel no-caption: %w", err)
}
return nil
})
}
@@ -0,0 +1,67 @@
package store_test
import (
"context"
"testing"
"time"
"github.com/stretchr/testify/require"
)
func TestChannelCaptionMemory_SuppressesAfterThreshold(t *testing.T) {
ctx := context.Background()
s := newStore(t)
super := rawPool(t)
resetDB(t, super)
seedUser(t, super, userA)
const threshold = 3
window := time.Hour
// Below threshold: not yet suppressed.
require.NoError(t, s.RecordChannelCaptionOutcome(ctx, userA, "chanX", false, threshold, window))
require.NoError(t, s.RecordChannelCaptionOutcome(ctx, userA, "chanX", false, threshold, window))
got, err := s.CaptionlessChannels(ctx, userA)
require.NoError(t, err)
require.NotContains(t, got, "chanX", "2 < threshold 3: not suppressed yet")
// Crossing the threshold suppresses the channel.
require.NoError(t, s.RecordChannelCaptionOutcome(ctx, userA, "chanX", false, threshold, window))
got, err = s.CaptionlessChannels(ctx, userA)
require.NoError(t, err)
require.Contains(t, got, "chanX", "3 consecutive no-caption results suppress the channel")
// A successful caption fetch resets it.
require.NoError(t, s.RecordChannelCaptionOutcome(ctx, userA, "chanX", true, threshold, window))
got, err = s.CaptionlessChannels(ctx, userA)
require.NoError(t, err)
require.NotContains(t, got, "chanX", "a captioned video clears suppression")
}
func TestChannelCaptionMemory_WindowExpiryReProbes(t *testing.T) {
ctx := context.Background()
s := newStore(t)
super := rawPool(t)
resetDB(t, super)
seedUser(t, super, userA)
// A negative window means captionless_until lands in the past — modelling an
// elapsed suppression window, which must make the channel eligible again.
require.NoError(t, s.RecordChannelCaptionOutcome(ctx, userA, "chanY", false, 1, -time.Hour))
got, err := s.CaptionlessChannels(ctx, userA)
require.NoError(t, err)
require.NotContains(t, got, "chanY", "an expired window re-enables the channel for a re-probe")
}
func TestChannelCaptionMemory_DisabledThresholdIsNoOp(t *testing.T) {
ctx := context.Background()
s := newStore(t)
super := rawPool(t)
resetDB(t, super)
seedUser(t, super, userA)
require.NoError(t, s.RecordChannelCaptionOutcome(ctx, userA, "chanZ", false, 0, time.Hour))
got, err := s.CaptionlessChannels(ctx, userA)
require.NoError(t, err)
require.Empty(t, got, "threshold 0 disables the memory — nothing recorded")
}
+59
View File
@@ -0,0 +1,59 @@
package store
import (
"context"
"fmt"
"time"
"github.com/jackc/pgx/v5"
)
// ChannelError is a channel that returned HTTP 404 on a discovery pass.
type ChannelError struct {
ChannelID string
ChannelName string
FirstSeen time.Time
LastSeen time.Time
}
// UpsertChannelError records (or refreshes) a 404 channel for the current user.
// Called by the runner inside a withUser scope; RLS guards user isolation.
func (s *Store) UpsertChannelError(ctx context.Context, userID, channelID, channelName string) error {
return s.withUser(ctx, userID, func(tx pgx.Tx) error {
_, err := tx.Exec(ctx, `
INSERT INTO channel_errors (user_id, channel_id, channel_name)
VALUES ($1, $2, $3)
ON CONFLICT (user_id, channel_id)
DO UPDATE SET channel_name = EXCLUDED.channel_name, last_seen = now()`,
userID, channelID, channelName)
if err != nil {
return fmt.Errorf("store: upsert channel error: %w", err)
}
return nil
})
}
// ListChannelErrors returns all 404-flagged channels for the user, newest first.
func (s *Store) ListChannelErrors(ctx context.Context, userID string) ([]ChannelError, error) {
var out []ChannelError
err := s.withUser(ctx, userID, func(tx pgx.Tx) error {
rows, err := tx.Query(ctx, `
SELECT channel_id, channel_name, first_seen, last_seen
FROM channel_errors
WHERE user_id = $1
ORDER BY last_seen DESC`, userID)
if err != nil {
return fmt.Errorf("store: list channel errors: %w", err)
}
defer rows.Close()
for rows.Next() {
var ce ChannelError
if err := rows.Scan(&ce.ChannelID, &ce.ChannelName, &ce.FirstSeen, &ce.LastSeen); err != nil {
return fmt.Errorf("store: scan channel error: %w", err)
}
out = append(out, ce)
}
return rows.Err()
})
return out, err
}
+1 -1
View File
@@ -6,7 +6,7 @@ import (
"github.com/stretchr/testify/require"
"gitea.d-ma.be/mathias/tapir/internal/adapters/store"
"git.d-ma.be/mathias/tapir/internal/adapters/store"
)
// seedUserRow inserts a bare users row (FK target for a connection) as the
+35
View File
@@ -31,6 +31,41 @@ func (s *Store) UserBySubject(ctx context.Context, subject string) (userID strin
return userID, true, nil
}
// UserIdentity is one (userID, dexSubject) pair from the un-RLS'd
// user_identities map — the unit the scheduler enumerates to run a discovery
// pass per user (ADR-018).
type UserIdentity struct {
UserID string
DexSubject string
}
// ListAllUsers returns every (userID, dexSubject) pair from user_identities. It
// runs as a plain pool query WITHOUT withUser — intentional and legitimate:
// user_identities is un-RLS'd auth plumbing (like UserBySubject), and the
// scheduler enumerating all users to run their discovery passes is an admin
// operation that cannot be scoped to any single user. Order is unspecified.
func (s *Store) ListAllUsers(ctx context.Context) ([]UserIdentity, error) {
rows, err := s.pool.Query(ctx,
`SELECT user_id, dex_subject FROM user_identities`)
if err != nil {
return nil, fmt.Errorf("store: list all users: %w", err)
}
defer rows.Close()
var users []UserIdentity
for rows.Next() {
var u UserIdentity
if err := rows.Scan(&u.UserID, &u.DexSubject); err != nil {
return nil, fmt.Errorf("store: scan user identity: %w", err)
}
users = append(users, u)
}
if err := rows.Err(); err != nil {
return nil, fmt.Errorf("store: iterate user identities: %w", err)
}
return users, nil
}
// RegisterUser creates the tapir user for a Dex subject and the identity mapping
// that points to it, returning the new user_id. It errors with
// ErrSubjectRegistered if the subject already maps.
+34 -1
View File
@@ -7,7 +7,7 @@ import (
"github.com/stretchr/testify/require"
"gitea.d-ma.be/mathias/tapir/internal/adapters/store"
"git.d-ma.be/mathias/tapir/internal/adapters/store"
)
const (
@@ -85,6 +85,39 @@ func TestRegisterUserRejectsDuplicateSubject(t *testing.T) {
require.Equal(t, first, got)
}
func TestListAllUsersReturnsEveryIdentity(t *testing.T) {
ctx := context.Background()
s := newStore(t)
resetDB(t, rawPool(t))
empty, err := s.ListAllUsers(ctx)
require.NoError(t, err)
require.Empty(t, empty, "no registrations yet → empty slice")
const subjectC = "dex|carol-789"
idA, err := s.RegisterUser(ctx, subjectA, "Alice")
require.NoError(t, err)
idB, err := s.RegisterUser(ctx, subjectB, "Bob")
require.NoError(t, err)
idC, err := s.RegisterUser(ctx, subjectC, "Carol")
require.NoError(t, err)
users, err := s.ListAllUsers(ctx)
require.NoError(t, err)
require.Len(t, users, 3)
// Order is unspecified; compare as a set of (userID, subject) pairs.
got := make(map[string]string, len(users))
for _, u := range users {
got[u.DexSubject] = u.UserID
}
require.Equal(t, map[string]string{
subjectA: idA,
subjectB: idB,
subjectC: idC,
}, got)
}
func TestRegisterUserDistinctSubjectsGetDistinctUsers(t *testing.T) {
ctx := context.Background()
s := newStore(t)
+41
View File
@@ -0,0 +1,41 @@
package store
import (
"context"
"fmt"
"github.com/jackc/pgx/v5"
)
// StampLogin records that the user was active today, throttled to one row per
// user per day. It is the read-side counterpart to SetAction: the middleware
// calls it on every authenticated request, but the append happens at most once a
// day so login_events stays small and the signal clean (one row = one active
// day, not one request).
//
// The check-and-insert is a single atomic statement: the INSERT ... SELECT ...
// WHERE NOT EXISTS only writes when no row for this user has seen_at in today
// (date_trunc('day', NOW()), server timezone). It runs through withUser, so the
// NOT EXISTS probe is itself RLS-scoped to the calling user via the
// tapir.current_user_id GUC — one user's stamp can never be suppressed or
// triggered by another user's rows. The explicit user_id predicate also keeps the
// probe on the (user_id, seen_at) index.
//
// A unique constraint is deliberately not used: under concurrent same-day
// requests the worst case is two rows for one day, which the gate query collapses
// to a single week bucket anyway — not worth a write-blocking constraint.
func (s *Store) StampLogin(ctx context.Context, userID string) error {
if err := s.withUser(ctx, userID, func(tx pgx.Tx) error {
_, err := tx.Exec(ctx,
`INSERT INTO login_events (user_id)
SELECT $1
WHERE NOT EXISTS (
SELECT 1 FROM login_events
WHERE user_id = $1 AND seen_at >= date_trunc('day', NOW())
)`, userID)
return err
}); err != nil {
return fmt.Errorf("store: stamp login: %w", err)
}
return nil
}
+83
View File
@@ -0,0 +1,83 @@
package store_test
import (
"context"
"testing"
"github.com/jackc/pgx/v5/pgxpool"
"github.com/stretchr/testify/require"
)
// countLoginEvents counts a user's login_events via the superuser pool, which
// bypasses RLS — so the assertion sees the true row count regardless of scope.
func countLoginEvents(t *testing.T, p *pgxpool.Pool, userID string) int {
t.Helper()
var n int
require.NoError(t, p.QueryRow(context.Background(),
`SELECT count(*) FROM login_events WHERE user_id = $1`, userID).Scan(&n))
return n
}
// TestStampLoginThrottlesToOnePerDay: repeated stamps within the same day insert
// exactly one row — the throttle that keeps login_events one-row-per-active-day.
func TestStampLoginThrottlesToOnePerDay(t *testing.T) {
ctx := context.Background()
s := newStore(t)
super := rawPool(t)
resetDB(t, super)
_, err := super.Exec(ctx, `INSERT INTO users (id) VALUES ($1)`, userA)
require.NoError(t, err)
for i := 0; i < 3; i++ {
require.NoError(t, s.StampLogin(ctx, userA))
}
require.Equal(t, 1, countLoginEvents(t, super, userA),
"three same-day stamps must collapse to one row")
}
// TestStampLoginRecordsOncePerNewDay: with yesterday's row already present, a
// stamp today is NOT throttled — it appends the day's row, so distinct active days
// accumulate (the substrate the gate's distinct-week count reads).
func TestStampLoginRecordsOncePerNewDay(t *testing.T) {
ctx := context.Background()
s := newStore(t)
super := rawPool(t)
resetDB(t, super)
_, err := super.Exec(ctx, `INSERT INTO users (id) VALUES ($1)`, userA)
require.NoError(t, err)
// Seed an event dated yesterday (before today's start), so the throttle's
// "row exists with seen_at >= start-of-today" probe finds nothing for today.
_, err = super.Exec(ctx,
`INSERT INTO login_events (user_id, seen_at) VALUES ($1, NOW() - INTERVAL '1 day')`, userA)
require.NoError(t, err)
require.NoError(t, s.StampLogin(ctx, userA))
require.Equal(t, 2, countLoginEvents(t, super, userA),
"a stamp on a new day must append a second row")
// A second stamp the same day is throttled again.
require.NoError(t, s.StampLogin(ctx, userA))
require.Equal(t, 2, countLoginEvents(t, super, userA),
"the same-day repeat must not add a third row")
}
// TestStampLoginIsUserScoped: one user's stamp lands only on that user's rows —
// the throttle probe is RLS-scoped, so user B's existing same-day row neither
// suppresses nor is touched by user A's stamp.
func TestStampLoginIsUserScoped(t *testing.T) {
ctx := context.Background()
s := newStore(t)
super := rawPool(t)
resetDB(t, super)
_, err := super.Exec(ctx, `INSERT INTO users (id) VALUES ($1), ($2)`, userA, userB)
require.NoError(t, err)
// B already has a same-day row; it must not throttle A's first stamp.
_, err = super.Exec(ctx, `INSERT INTO login_events (user_id) VALUES ($1)`, userB)
require.NoError(t, err)
require.NoError(t, s.StampLogin(ctx, userA))
require.Equal(t, 1, countLoginEvents(t, super, userA), "A's stamp must record despite B's same-day row")
require.Equal(t, 1, countLoginEvents(t, super, userB), "A's stamp must not touch B's rows")
}
+165
View File
@@ -0,0 +1,165 @@
package store_test
import (
"context"
"database/sql"
"errors"
"os"
"testing"
"github.com/golang-migrate/migrate/v4"
migratepgx "github.com/golang-migrate/migrate/v4/database/pgx/v5"
"github.com/golang-migrate/migrate/v4/source/iofs"
"github.com/stretchr/testify/require"
_ "github.com/jackc/pgx/v5/stdlib" // register the "pgx" database/sql driver
)
// fileMigrator builds a golang-migrate instance from the on-disk migration files
// (not the embedded FS the production Migrate uses), so a test can step the schema
// up and down. os.DirFS(".") is rooted at the package dir; the SQL lives under
// "migrations". Mirrors store.Migrate's construction otherwise.
func fileMigrator(t *testing.T) *migrate.Migrate {
t.Helper()
db, err := sql.Open("pgx", dsn)
require.NoError(t, err)
t.Cleanup(func() { _ = db.Close() })
drv, err := migratepgx.WithInstance(db, &migratepgx.Config{})
require.NoError(t, err)
src, err := iofs.New(os.DirFS("."), "migrations")
require.NoError(t, err)
m, err := migrate.NewWithInstance("iofs", src, "pgx", drv)
require.NoError(t, err)
t.Cleanup(func() { _, _ = m.Close() })
return m
}
// headVersion reports the current (HEAD) schema version so a test can restore
// to it after stepping down, without hard-coding what HEAD is. Adding a
// migration on top changes HEAD but no test that uses this needs editing.
func headVersion(t *testing.T, m *migrate.Migrate) uint {
t.Helper()
v, dirty, err := m.Version()
require.NoError(t, err)
require.False(t, dirty, "schema must not be dirty")
return v
}
// migrateTo drives the schema to an exact version *by version number*, not by
// step count. This is the whole point of the migrate-test design: a migration
// added above the target does not shift any count here, so unrelated tests stay
// green (see issue #8). ErrNoChange (already at that version) is not a failure.
func migrateTo(t *testing.T, m *migrate.Migrate, version uint) {
t.Helper()
err := m.Migrate(version)
if errors.Is(err, migrate.ErrNoChange) {
return
}
require.NoError(t, err)
}
// loginEventsExists reports whether the login_events relation is present.
func loginEventsExists(t *testing.T) bool {
t.Helper()
var reg *string
require.NoError(t, rawPool(t).QueryRow(context.Background(),
`SELECT to_regclass('public.login_events')::text`).Scan(&reg))
return reg != nil
}
// TestMigration010LoginEventsUpDown proves migration 010 is reversible: the down
// migration drops login_events cleanly and the up migration recreates it. A rotten
// down migration (forgotten DROP, dangling policy) would fail here rather than in
// production during a rollback. The test restores the schema to latest before
// returning so the shared embedded-postgres stays at HEAD for sibling tests.
func TestMigration010LoginEventsUpDown(t *testing.T) {
newStore(t) // ensure the schema is migrated to latest (011 applied)
require.True(t, loginEventsExists(t), "login_events must exist at latest migration")
m := fileMigrator(t)
head := headVersion(t, m)
migrateTo(t, m, 9) // just below 010 — everything above steps down
require.False(t, loginEventsExists(t), "login_events must be gone after the down migration")
migrateTo(t, m, 10) // up 010
require.True(t, loginEventsExists(t), "login_events must be restored after the up migration")
migrateTo(t, m, head) // restore to HEAD for sibling tests
}
// autoSummarizeDefault reads the users.auto_summarize column default as text
// ("true"/"false"), so the migration's default flip is verifiable directly.
func autoSummarizeDefault(t *testing.T) string {
t.Helper()
var def string
require.NoError(t, rawPool(t).QueryRow(context.Background(),
`SELECT column_default FROM information_schema.columns
WHERE table_name = 'users' AND column_name = 'auto_summarize'`).Scan(&def))
return def
}
// TestMigration011AutoSummarizeDefaultUpDown proves migration 011 is reversible:
// up sets the auto_summarize column default to TRUE (ADR-018), down restores
// FALSE. The down intentionally does not revert existing rows — only the default.
func TestMigration011AutoSummarizeDefaultUpDown(t *testing.T) {
newStore(t) // latest (013 applied)
require.Equal(t, "true", autoSummarizeDefault(t), "011 sets the default to TRUE")
m := fileMigrator(t)
head := headVersion(t, m)
migrateTo(t, m, 10) // just below 011 — reverts the column default
require.Equal(t, "false", autoSummarizeDefault(t), "default is FALSE after the down migration")
migrateTo(t, m, 11) // up 011 re-applies the TRUE default
require.Equal(t, "true", autoSummarizeDefault(t))
migrateTo(t, m, head) // restore to HEAD for sibling tests
}
// channelTitleExists reports whether videos.channel_title is present.
func channelTitleExists(t *testing.T) bool {
t.Helper()
var exists bool
require.NoError(t, rawPool(t).QueryRow(context.Background(),
`SELECT EXISTS (SELECT 1 FROM information_schema.columns
WHERE table_name = 'videos' AND column_name = 'channel_title')`).Scan(&exists))
return exists
}
// TestMigration014VideoChannelTitleUpDown proves 014 is reversible: down drops
// videos.channel_title, up recreates it.
func TestMigration014VideoChannelTitleUpDown(t *testing.T) {
newStore(t) // latest (014 applied)
require.True(t, channelTitleExists(t), "channel_title exists at latest migration")
m := fileMigrator(t)
head := headVersion(t, m)
migrateTo(t, m, 13) // just below 014 — drops channel_title
require.False(t, channelTitleExists(t), "channel_title must be gone after the down migration")
migrateTo(t, m, 14) // up 014 recreates channel_title
require.True(t, channelTitleExists(t), "channel_title must be restored after the up migration")
migrateTo(t, m, head) // restore to HEAD for sibling tests
}
// TestMigration012FixAutoSummarizeRLS proves 012 runs cleanly and flips any
// remaining auto_summarize=FALSE rows to TRUE (the back-fill blocked by RLS in 011).
func TestMigration012FixAutoSummarizeRLS(t *testing.T) {
newStore(t) // apply all migrations including 012
require.Equal(t, "true", autoSummarizeDefault(t), "column default is TRUE after 012")
// Round-trip: down 012, then up 012 — must be idempotent.
m := fileMigrator(t)
head := headVersion(t, m)
migrateTo(t, m, 11) // down 012 must not error
migrateTo(t, m, 12) // up 012 must re-apply cleanly
require.Equal(t, "true", autoSummarizeDefault(t), "default still TRUE after 012 re-applied")
migrateTo(t, m, head) // restore to HEAD for sibling tests
}
@@ -0,0 +1,2 @@
ALTER TABLE videos DROP COLUMN IF EXISTS rate_limited_at;
ALTER TABLE videos DROP COLUMN IF EXISTS transcript_status;
@@ -0,0 +1,17 @@
-- Migration 007: per-video transcript fetch status, for rate-limit backoff.
--
-- transcript_status records the outcome of the last transcript attempt:
-- NULL = not yet attempted
-- 'none' = checked, no usable transcript (permanent — SourceNone)
-- 'fetched' = transcript resolved and summarized (summary_id not null)
-- 'rate_limited'= the caption endpoint returned 429; retry after a backoff window
--
-- rate_limited_at stamps WHEN the 429 was seen, so the runner can skip re-fetching
-- a still-throttled video until NOW() - rate_limited_at exceeds TAPIR_FETCH_BACKOFF.
-- It is cleared (set NULL) whenever the status moves off 'rate_limited'.
--
-- No RLS policy changes needed: videos already has ENABLE + FORCE ROW LEVEL
-- SECURITY (migration 003) with the videos_isolation policy. New columns inherit
-- that protection automatically.
ALTER TABLE videos ADD COLUMN transcript_status TEXT;
ALTER TABLE videos ADD COLUMN rate_limited_at TIMESTAMPTZ;
@@ -0,0 +1 @@
DROP TABLE IF EXISTS invitations;
@@ -0,0 +1,22 @@
-- Migration 009: invitations — an email-based invite to join Tapir (Stage-1
-- onboarding gate). Mathias mints one with `tapir invite <email>`; the recipient
-- visits /invite/{token}, sets a password, and Tapir creates their Dex account.
--
-- Deliberately NOT user-owned and NOT under RLS: an invitation exists BEFORE the
-- user does, so there is no user_id to scope by and no authenticated user context
-- when the invite is created (host CLI) or consumed (public /invite handler, no
-- Dex session). The token itself is the capability — a 32-byte crypto-random,
-- single-use, time-boxed secret. Hence no `user_id` FK and no ENABLE/FORCE ROW
-- LEVEL SECURITY here (unlike every user-owned table in migrations 003/005).
CREATE TABLE invitations (
id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
email TEXT NOT NULL,
token TEXT NOT NULL UNIQUE,
created_at TIMESTAMPTZ NOT NULL DEFAULT NOW(),
expires_at TIMESTAMPTZ NOT NULL,
used_at TIMESTAMPTZ
);
-- Lookups are by token (both the claim and the form preview); the UNIQUE
-- constraint already creates an index, this names one explicitly for clarity.
CREATE INDEX idx_invitations_token ON invitations(token);
@@ -0,0 +1,4 @@
DROP POLICY IF EXISTS login_events_isolation ON login_events;
ALTER TABLE login_events NO FORCE ROW LEVEL SECURITY;
ALTER TABLE login_events DISABLE ROW LEVEL SECURITY;
DROP TABLE IF EXISTS login_events;
@@ -0,0 +1,33 @@
-- Migration 010: login_events records THAT a user was active (returned and read)
-- on a given day — the Stage-0 signal summary_actions misses. summary_actions
-- captures *acts* (watch/skip/save); a reader who logs in weekly and clicks
-- nothing is otherwise invisible, yet for a reading product that return IS the
-- signal the gate ("usage in >=2 distinct weeks", VISION/ADR-016) is defined on.
--
-- Append-only: one row per user per active day (the request-path throttle in the
-- web layer enforces that cadence), never updated. Per-user isolation like every
-- user-owned table.
--
-- NO foreign key to users (mirrors summary_actions, migration 002): user_id is
-- carried for RLS/scoping but the table is decoupled so a stamp never blocks on a
-- users row. The cost of that decoupling: the users-row cascade does NOT reach
-- login_events, so DeleteUser must delete it explicitly (see account.go) — the
-- exact footgun the summary_actions delete work caught.
CREATE TABLE login_events (
id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
user_id UUID NOT NULL,
seen_at TIMESTAMPTZ NOT NULL DEFAULT NOW()
);
-- (user_id, seen_at) serves both the per-user-per-day throttle lookup
-- (seen_at >= start-of-today) and the gate query's per-user week bucketing.
CREATE INDEX idx_login_events_user_seen ON login_events(user_id, seen_at);
-- RLS: identical GUC-keyed policy/pattern to migration 003. FORCE so the table
-- owner (tapir, non-superuser in prod) is subject to it; an unset GUC yields NULL
-- → no rows match → deny-all.
ALTER TABLE login_events ENABLE ROW LEVEL SECURITY;
ALTER TABLE login_events FORCE ROW LEVEL SECURITY;
CREATE POLICY login_events_isolation ON login_events
FOR ALL
USING (user_id = current_setting('tapir.current_user_id', true)::uuid);
@@ -0,0 +1,5 @@
-- Revert the column default to FALSE. Existing rows are intentionally NOT
-- reverted: flipping live users back to manual on a rollback would be a
-- surprising regression (they may have come to rely on auto). The default change
-- is the reversible part; data stays as the user left it.
ALTER TABLE users ALTER COLUMN auto_summarize SET DEFAULT FALSE;
@@ -0,0 +1,17 @@
-- Migration 011: flip auto_summarize default to TRUE (ADR-018, Future-B).
--
-- Scheduled discovery (ADR-018) makes Tapir watch unprompted. For onboarded
-- friends that only delivers zero-friction value if the list also SUMMARIZES
-- itself — a manual default would mean the scheduler discovers videos a user
-- still has to click through one by one, which is the empty-list problem again.
-- So new users default to AUTO. The account-page toggle still lets a user switch
-- to manual (SetAutoSummarize), so this only changes the out-of-the-box state.
--
-- Safe only because the process-wide caption-fetch rate gate (ADR-014 item 2)
-- now exists: auto + scheduled + multi-user would otherwise self-inflict 429s
-- every cycle. The gate is the precondition for shipping this default.
ALTER TABLE users ALTER COLUMN auto_summarize SET DEFAULT TRUE;
-- Bring existing rows (maintainer + any current registrations) onto the new
-- default so they benefit immediately, not just users created after this point.
UPDATE users SET auto_summarize = TRUE WHERE auto_summarize = FALSE;
@@ -0,0 +1 @@
-- No data revert: do not flip users back to manual on rollback.
@@ -0,0 +1,6 @@
-- Migration 011's UPDATE ran without tapir.current_user_id set, so FORCE RLS
-- blocked all rows and zero users were updated. Temporarily drop FORCE so the
-- table owner (tapir role) can bypass RLS for this back-fill, then restore it.
ALTER TABLE users NO FORCE ROW LEVEL SECURITY;
UPDATE users SET auto_summarize = TRUE WHERE auto_summarize = FALSE;
ALTER TABLE users FORCE ROW LEVEL SECURITY;
@@ -0,0 +1 @@
DROP TABLE IF EXISTS channel_errors;
@@ -0,0 +1,22 @@
-- channel_errors: channels that returned HTTP 404 (deleted/private) on the most
-- recent scheduler pass. Surfaced on the account page so users know why some
-- subscribed channels produce no videos. Upserted per-pass; cleared when the
-- channel starts returning results again (runner calls UpsertChannelError only
-- on 404, so a recovered channel simply stops appearing after its row ages out
-- or the user takes action). ON DELETE CASCADE keeps rows tidy on account deletion.
CREATE TABLE channel_errors (
user_id UUID NOT NULL REFERENCES users(id) ON DELETE CASCADE,
channel_id TEXT NOT NULL,
channel_name TEXT NOT NULL,
first_seen TIMESTAMPTZ NOT NULL DEFAULT now(),
last_seen TIMESTAMPTZ NOT NULL DEFAULT now(),
PRIMARY KEY (user_id, channel_id)
);
CREATE INDEX idx_channel_errors_user_id ON channel_errors(user_id);
ALTER TABLE channel_errors ENABLE ROW LEVEL SECURITY;
ALTER TABLE channel_errors FORCE ROW LEVEL SECURITY;
CREATE POLICY channel_errors_isolation ON channel_errors
FOR ALL
USING (user_id = current_setting('tapir.current_user_id', true)::uuid);
@@ -0,0 +1 @@
ALTER TABLE videos DROP COLUMN channel_title;
@@ -0,0 +1,8 @@
-- Store the source channel's title per video so the list can offer a real
-- channel filter (multi-select of the user's channels) instead of the dead
-- free-text field that only ever matched the provider string. Nullable: existing
-- rows backfill on the next discovery pass (UpsertVideo writes it); pasted videos
-- get it immediately from videos.list. No FK to a channels table at Stage 0 — the
-- title is a denormalised display/filter value, consistent with the existing
-- subscription_id-stays-NULL stance (data-model.md).
ALTER TABLE videos ADD COLUMN channel_title TEXT;
@@ -0,0 +1,19 @@
-- Down 015: restore the per-user RLS-scoped transcripts shape (001 + 003).
DROP TABLE transcripts;
CREATE TABLE transcripts (
video_id UUID PRIMARY KEY REFERENCES videos(id) ON DELETE CASCADE,
user_id UUID NOT NULL REFERENCES users(id) ON DELETE CASCADE,
source TEXT NOT NULL,
language TEXT,
content TEXT,
resolved_at TIMESTAMPTZ NOT NULL DEFAULT NOW()
);
CREATE INDEX idx_transcripts_user_id ON transcripts(user_id);
ALTER TABLE transcripts ENABLE ROW LEVEL SECURITY;
ALTER TABLE transcripts FORCE ROW LEVEL SECURITY;
CREATE POLICY transcripts_isolation ON transcripts
FOR ALL
USING (user_id = current_setting('tapir.current_user_id', true)::uuid);
@@ -0,0 +1,30 @@
-- Migration 015: transcripts become SHARED public-content storage (ADR-021).
--
-- The per-user transcripts table from 001 (PK videos.id, user_id NOT NULL, RLS
-- FORCEd in 003) was dead: no application code ever read or wrote it — only the
-- transcript_status columns on `videos` (007) carried fetch outcomes. ADR-021
-- repurposes it as the single shared store of public caption content, keyed by
-- the cross-user dedup key (provider, provider_video_id) — the video's public
-- identity, not Tapir's per-user videos.id — so re-analysis never re-fetches
-- from YouTube (ADR-010/014).
--
-- It holds ONLY public caption content + the video's public id (nothing
-- user-identifying), so it is deliberately NOT RLS-scoped: no user_id, no
-- policy, no FORCE. This is the single, intentional exception to the ADR-012
-- isolation boundary; rls_test.go asserts the boundary is exactly here and
-- nowhere else. Dropping the old table drops its RLS policy with it; it held no
-- real data, so drop+recreate loses nothing.
DROP TABLE transcripts;
CREATE TABLE transcripts (
provider TEXT NOT NULL,
provider_video_id TEXT NOT NULL,
source TEXT NOT NULL, -- 'captions' (content set) | 'none' (no captions; content NULL)
language TEXT,
content TEXT,
fetched_at TIMESTAMPTZ NOT NULL DEFAULT NOW(),
PRIMARY KEY (provider, provider_video_id)
);
COMMENT ON TABLE transcripts IS
'Shared public caption content keyed by (provider, provider_video_id). NOT RLS-scoped — public content only, de-facto cross-user dedup (ADR-021).';
@@ -0,0 +1 @@
DROP TABLE channel_caption_state;
@@ -0,0 +1,30 @@
-- Migration 016: per-(user, channel) caption-availability memory (ADR-024).
--
-- Some channels never publish English captions (foreign-language news, music,
-- etc.). Each of their new videos still costs ONE rate-limited caption fetch
-- (ADR-014) before resolving to "none" — and on a throttled egress IP that fetch
-- may 429 and churn through the backoff machinery first. This table remembers
-- channels that repeatedly yield no captions so discovery can stop attempting
-- their videos, freeing the scarce fetch budget for channels that do have them.
--
-- consecutive_none counts no-caption outcomes in a row; a successful fetch resets
-- it to 0. Once it crosses the threshold the channel is suppressed until
-- captionless_until, after which one video is re-probed (auto-recovery for a
-- channel that starts adding captions). Per-user + RLS-scoped, consistent with
-- the rest of the user-owned schema (subscriptions are per-user; ADR-012).
CREATE TABLE channel_caption_state (
user_id UUID NOT NULL REFERENCES users(id) ON DELETE CASCADE,
channel_id TEXT NOT NULL,
consecutive_none INT NOT NULL DEFAULT 0,
captionless_until TIMESTAMPTZ,
updated_at TIMESTAMPTZ NOT NULL DEFAULT now(),
PRIMARY KEY (user_id, channel_id)
);
CREATE INDEX idx_channel_caption_state_user_id ON channel_caption_state(user_id);
ALTER TABLE channel_caption_state ENABLE ROW LEVEL SECURITY;
ALTER TABLE channel_caption_state FORCE ROW LEVEL SECURITY;
CREATE POLICY channel_caption_state_isolation ON channel_caption_state
FOR ALL
USING (user_id = current_setting('tapir.current_user_id', true)::uuid);
+17 -5
View File
@@ -29,6 +29,7 @@ type SummaryRow struct {
ProviderVideoID string // videos.provider_video_id; empty when no videos row
Title string // videos.title; empty when no videos row
Channel string // videos.provider for now; empty when no videos row
ChannelTitle string // videos.channel_title; the source channel, for display + filtering
URL string // videos.url; empty when no videos row
PublishedAt time.Time // videos.published_at; zero when absent
Summary string
@@ -49,6 +50,10 @@ type SummaryRow struct {
// set by the web "Summarize" button and cleared by the next `tapir run`. Only
// populated by ListVideos/GetVideoRow (summary-only reads leave it false).
SummarizeRequested bool
// TranscriptStatus mirrors videos.transcript_status (migration 007): "" (unset),
// "none", "rate_limited", or "fetched". Drives the "Retrying later" list badge.
// Only populated by ListVideos/GetVideoRow ("" on summary-only reads).
TranscriptStatus string
}
// selectSummary is the shared projection for both reads. videos is LEFT JOINed
@@ -132,24 +137,29 @@ const selectVideo = `
COALESCE(s.fallback_used, FALSE),
COALESCE(s.created_at, v.seen_at),
(s.id IS NOT NULL) AS summarized,
v.summarize_requested
v.summarize_requested,
COALESCE(v.transcript_status, ''),
COALESCE(v.channel_title, '')
FROM videos v
LEFT JOIN summaries s ON s.video_id = v.id AND s.user_id = v.user_id`
// ListVideos returns ALL of the user's videos — summarized and not — most recent
// first by seen_at, capped at limit (non-positive defaults to 50). Unsummarized
// ListVideos returns ALL of the user's videos — summarized first, then by
// published_at DESC with undated videos last, then seen_at DESC as a tiebreak —
// capped at limit (non-positive defaults to 500). The published_at ordering
// aligns the list with the recency framing (newest content first); seen_at
// breaks ties and orders same/!undated rows deterministically. Unsummarized
// videos come back with Summarized=false and empty summary fields, so the list
// view can render them with a "Summarize" affordance. Scoped by user_id.
func (s *Store) ListVideos(ctx context.Context, userID string, limit int) ([]SummaryRow, error) {
if limit <= 0 {
limit = 50
limit = 500
}
var out []SummaryRow
if err := s.withUser(ctx, userID, func(tx pgx.Tx) error {
rows, err := tx.Query(ctx,
selectVideo+`
WHERE v.user_id = $1
ORDER BY v.seen_at DESC
ORDER BY (s.id IS NOT NULL) DESC, v.published_at DESC NULLS LAST, v.seen_at DESC
LIMIT $2`,
userID, limit)
if err != nil {
@@ -245,6 +255,8 @@ func scanVideoRow(rows pgx.Row) (SummaryRow, error) {
&row.CreatedAt,
&row.Summarized,
&row.SummarizeRequested,
&row.TranscriptStatus,
&row.ChannelTitle,
); err != nil {
return SummaryRow{}, fmt.Errorf("store: scan video: %w", err)
}
+1 -1
View File
@@ -8,7 +8,7 @@ import (
"github.com/jackc/pgx/v5/pgxpool"
"github.com/stretchr/testify/require"
"gitea.d-ma.be/mathias/tapir/internal/adapters/store"
"git.d-ma.be/mathias/tapir/internal/adapters/store"
)
// seedVideo inserts a videos row whose id matches a summary's video_id, so the
+117
View File
@@ -0,0 +1,117 @@
package store
import (
"context"
"fmt"
"sort"
"time"
"github.com/jackc/pgx/v5"
)
// UserActiveWeeks is one row of the Stage-0 gate report: how many DISTINCT
// calendar weeks a user was active in, counting reads (login_events) AND acts
// (summary_actions) together. The gate (VISION/ADR-016) passes when any user
// reaches ActiveWeeks >= 2.
type UserActiveWeeks struct {
UserID string
DisplayName string
ActiveWeeks int
}
// ActiveWeeks computes per-user distinct-active-weeks for the gate report, most
// active first.
//
// Why per-user iteration rather than one cross-user GROUP BY: the user-owned
// tables are FORCE RLS (migration 003/010) and the production role is a non-
// superuser owner, so a single un-scoped query sees nothing (deny-all). Instead we
// enumerate users from the deliberately un-RLS'd identity map (user_identities,
// migration 004) and count each user's weeks inside withUser, where the GUC scopes
// login_events + summary_actions to that user. No privilege escalation, no policy
// change — the same isolation seam every other read flows through.
//
// Scope note: the enumeration covers users with a Dex identity (the web users the
// gate is about). A CLI-only user created by the store sink without an identity
// row would not appear — out of scope for this gate.
// ActiveWeeks counts each user's distinct active weeks from `since` onward. A zero
// `since` means no lower bound (count all history). The Stage-0 gate baseline is
// set by the caller (the report command) to the date real usage tracking began,
// so pre-launch noise — testing, the period the pilot was blocked — does not count
// toward the return-usage signal (ADR-016).
func (s *Store) ActiveWeeks(ctx context.Context, since time.Time) ([]UserActiveWeeks, error) {
userIDs, err := s.identityUserIDs(ctx)
if err != nil {
return nil, err
}
out := make([]UserActiveWeeks, 0, len(userIDs))
for _, uid := range userIDs {
row, err := s.activeWeeksFor(ctx, uid, since)
if err != nil {
return nil, err
}
out = append(out, row)
}
// Most active first; user_id as a stable tie-break for deterministic output.
sort.SliceStable(out, func(i, j int) bool {
if out[i].ActiveWeeks != out[j].ActiveWeeks {
return out[i].ActiveWeeks > out[j].ActiveWeeks
}
return out[i].UserID < out[j].UserID
})
return out, nil
}
// identityUserIDs lists every tapir user_id from the un-RLS'd identity map. It
// runs directly on the pool (no withUser): user_identities carries no user data
// and is intentionally not RLS-enabled, so it is the one table that can be read
// pre-scope to discover who exists.
func (s *Store) identityUserIDs(ctx context.Context) ([]string, error) {
rows, err := s.pool.Query(ctx, `SELECT user_id FROM user_identities ORDER BY user_id`)
if err != nil {
return nil, fmt.Errorf("store: list identity users: %w", err)
}
defer rows.Close()
var ids []string
for rows.Next() {
var id string
if err := rows.Scan(&id); err != nil {
return nil, fmt.Errorf("store: scan identity user: %w", err)
}
ids = append(ids, id)
}
if err := rows.Err(); err != nil {
return nil, fmt.Errorf("store: iterate identity users: %w", err)
}
return ids, nil
}
// activeWeeksFor counts one user's distinct active weeks (reads UNION acts) and
// reads their display name, RLS-scoped via withUser. The UNION dedups a week that
// has both a login and an action so it counts once.
func (s *Store) activeWeeksFor(ctx context.Context, userID string, since time.Time) (UserActiveWeeks, error) {
res := UserActiveWeeks{UserID: userID}
if err := s.withUser(ctx, userID, func(tx pgx.Tx) error {
if err := tx.QueryRow(ctx,
`WITH weeks AS (
SELECT date_trunc('week', seen_at) AS wk
FROM login_events WHERE user_id = $1 AND seen_at >= $2
UNION
SELECT date_trunc('week', acted_at)
FROM summary_actions WHERE user_id = $1 AND acted_at >= $2
)
SELECT count(DISTINCT wk) FROM weeks`, userID, since).Scan(&res.ActiveWeeks); err != nil {
return fmt.Errorf("store: count active weeks: %w", err)
}
if err := tx.QueryRow(ctx,
`SELECT COALESCE(display_name, '') FROM users WHERE id = $1`, userID).Scan(&res.DisplayName); err != nil {
return fmt.Errorf("store: read display name: %w", err)
}
return nil
}); err != nil {
return UserActiveWeeks{}, err
}
return res, nil
}
+105
View File
@@ -0,0 +1,105 @@
package store_test
import (
"context"
"testing"
"time"
"github.com/jackc/pgx/v5/pgxpool"
"github.com/stretchr/testify/require"
)
// seedReportUser inserts a user + its identity mapping (the enumeration source
// ActiveWeeks reads). display_name is optional.
func seedReportUser(t *testing.T, p *pgxpool.Pool, userID, subject, name string) {
t.Helper()
ctx := context.Background()
_, err := p.Exec(ctx,
`INSERT INTO users (id, display_name) VALUES ($1, NULLIF($2, ''))`, userID, name)
require.NoError(t, err)
_, err = p.Exec(ctx,
`INSERT INTO user_identities (dex_subject, user_id) VALUES ($1, $2)`, subject, userID)
require.NoError(t, err)
}
// TestActiveWeeksCountsDistinctWeeksAcrossReadsAndActs is the gate-query proof.
// It seeds, with fixed timestamps in known ISO weeks:
// - user A: reads in week of Jan 5 and Jan 12, acts in week of Jan 12 (dup) and
// Jan 19 → the UNION across both tables collapses the shared week → 3 distinct.
// - user B: a single read in the week of Jan 5 → 1 distinct (below the gate).
//
// It verifies the count is correct, dedups the cross-table shared week, and orders
// most-active first.
func TestActiveWeeksCountsDistinctWeeksAcrossReadsAndActs(t *testing.T) {
ctx := context.Background()
s := newStore(t)
p := rawPool(t)
resetDB(t, p) // TRUNCATE ... users CASCADE also clears user_identities
seedReportUser(t, p, userA, "subject-a", "Ada")
seedReportUser(t, p, userB, "subject-b", "")
// Reads (login_events) — fixed dates in distinct ISO weeks.
_, err := p.Exec(ctx,
`INSERT INTO login_events (user_id, seen_at) VALUES
($1, '2026-01-05T09:00:00Z'),
($1, '2026-01-12T09:00:00Z'),
($2, '2026-01-05T09:00:00Z')`, userA, userB)
require.NoError(t, err)
// Acts (summary_actions) — one in A's week-of-Jan-12 (shared with a read, must
// dedup) and one in a new week (Jan 19).
_, err = p.Exec(ctx,
`INSERT INTO summary_actions (user_id, video_id, action, acted_at) VALUES
($1, 'vid-1', 'watched', '2026-01-12T18:00:00Z'),
($1, 'vid-2', 'saved', '2026-01-19T18:00:00Z')`, userA)
require.NoError(t, err)
got, err := s.ActiveWeeks(ctx, time.Time{}) // zero since = no lower bound
require.NoError(t, err)
require.Len(t, got, 2, "both identity users must appear")
require.Equal(t, userA, got[0].UserID, "most-active user first")
require.Equal(t, "Ada", got[0].DisplayName)
require.Equal(t, 3, got[0].ActiveWeeks, "3 distinct weeks across reads+acts, shared week deduped")
require.Equal(t, userB, got[1].UserID)
require.Equal(t, 1, got[1].ActiveWeeks, "single read = 1 distinct week (below gate)")
}
// TestActiveWeeksEmptyWhenNoUsers: no identities → no rows (not an error).
func TestActiveWeeksEmptyWhenNoUsers(t *testing.T) {
ctx := context.Background()
s := newStore(t)
resetDB(t, rawPool(t))
got, err := s.ActiveWeeks(ctx, time.Time{})
require.NoError(t, err)
require.Empty(t, got)
}
// TestActiveWeeksExcludesBeforeGateStart proves the baseline cutoff: activity
// before `since` does not count, so pre-launch noise (testing, the pilot's blocked
// period) is excluded from the Stage-0 return-usage gate (ADR-016).
func TestActiveWeeksExcludesBeforeGateStart(t *testing.T) {
ctx := context.Background()
s := newStore(t)
p := rawPool(t)
resetDB(t, p)
seedReportUser(t, p, userA, "subject-a", "Ada")
// One read well before the baseline, two reads in distinct weeks after it.
_, err := p.Exec(ctx,
`INSERT INTO login_events (user_id, seen_at) VALUES
($1, '2026-05-01T09:00:00Z'),
($1, '2026-06-12T09:00:00Z'),
($1, '2026-06-19T09:00:00Z')`, userA)
require.NoError(t, err)
since := time.Date(2026, 6, 11, 0, 0, 0, 0, time.UTC)
got, err := s.ActiveWeeks(ctx, since)
require.NoError(t, err)
require.Len(t, got, 1)
require.Equal(t, 2, got[0].ActiveWeeks, "only the two post-baseline weeks count; the May read is excluded")
}
+80 -15
View File
@@ -22,9 +22,11 @@ import (
// (no GUC set → zero rows) proves the enforcement path is live, not bypassed.
// userIsolatedTables are the tables that carry a user_id and whose policy keys
// directly off the tapir.current_user_id GUC.
// directly off the tapir.current_user_id GUC. transcripts is deliberately ABSENT
// — ADR-021 made it shared public content (non-RLS); TestTranscriptsTableIsSharedNotRLS
// proves that is the only place the isolation boundary moved.
var userIsolatedTables = []string{
"users", "videos", "transcripts", "summaries", "summary_actions", "video_connections",
"users", "videos", "summaries", "summary_actions", "login_events", "video_connections",
}
// allIsolatedTables adds sink_deliveries, whose ownership is derived from its
@@ -38,9 +40,10 @@ type seeded struct {
summaryID string
}
// seedUser inserts one full chain (user → video → transcript → summary →
// action → delivery) as the superuser pool, which bypasses RLS so both users'
// data lands regardless of the GUC.
// seedUser inserts one full chain (user → video → summary → action → delivery)
// as the superuser pool, which bypasses RLS so both users' data lands regardless
// of the GUC. Transcripts are NOT seeded here: they are shared, non-RLS public
// content (ADR-021), so they have no place in a per-user isolation chain.
func seedUser(t *testing.T, p *pgxpool.Pool, userID string) seeded {
t.Helper()
ctx := context.Background()
@@ -54,11 +57,6 @@ func seedUser(t *testing.T, p *pgxpool.Pool, userID string) seeded {
VALUES ($1, 'youtube', $2, 'title') RETURNING id`,
userID, "vid-"+userID).Scan(&videoID))
_, err = p.Exec(ctx,
`INSERT INTO transcripts (video_id, user_id, source, content)
VALUES ($1, $2, 'captions', 'words')`, videoID, userID)
require.NoError(t, err)
var summaryID string
require.NoError(t, p.QueryRow(ctx,
`INSERT INTO summaries (user_id, video_id, summary) VALUES ($1, $2, 'sum')
@@ -69,6 +67,10 @@ func seedUser(t *testing.T, p *pgxpool.Pool, userID string) seeded {
VALUES ($1, $2, 'watched')`, userID, videoID)
require.NoError(t, err)
_, err = p.Exec(ctx,
`INSERT INTO login_events (user_id) VALUES ($1)`, userID)
require.NoError(t, err)
_, err = p.Exec(ctx,
`INSERT INTO sink_deliveries (summary_id, sink, status)
VALUES ($1, 'store', 'delivered')`, summaryID)
@@ -88,9 +90,16 @@ func appPool(t *testing.T, super *pgxpool.Pool) *pgxpool.Pool {
t.Helper()
ctx := context.Background()
// Idempotent across test runs (schema/role persist for the TestMain PG).
_, _ = super.Exec(ctx, `DROP ROLE IF EXISTS app`)
_, err := super.Exec(ctx, `CREATE ROLE app LOGIN PASSWORD 'app'`)
// Idempotent across tests AND runs: the role persists for the TestMain PG and
// owns granted privileges, so a plain DROP ROLE fails once any GRANT exists
// (and more than one test now builds an app pool). Create only if absent; the
// GRANTs below are themselves idempotent.
_, err := super.Exec(ctx,
`DO $$ BEGIN
IF NOT EXISTS (SELECT FROM pg_roles WHERE rolname = 'app') THEN
CREATE ROLE app LOGIN PASSWORD 'app';
END IF;
END $$`)
require.NoError(t, err)
_, err = super.Exec(ctx, `GRANT USAGE ON SCHEMA public TO app`)
require.NoError(t, err)
@@ -182,13 +191,14 @@ func TestRLSEnforcesPerUserIsolation(t *testing.T) {
{"update users", `UPDATE users SET display_name = 'hacked' WHERE id = $1`, b.userID},
{"update videos", `UPDATE videos SET title = 'hacked' WHERE user_id = $1`, b.userID},
{"queue videos summarize", `UPDATE videos SET summarize_requested = TRUE WHERE id = $1`, b.videoID},
{"update transcripts", `UPDATE transcripts SET content = 'hacked' WHERE user_id = $1`, b.userID},
{"update summaries", `UPDATE summaries SET summary = 'hacked' WHERE user_id = $1`, b.userID},
{"update summary_actions", `UPDATE summary_actions SET action = 'skipped' WHERE user_id = $1`, b.userID},
{"update login_events", `UPDATE login_events SET seen_at = NOW() WHERE user_id = $1`, b.userID},
{"update sink_deliveries", `UPDATE sink_deliveries SET status = 'hacked' WHERE summary_id = $1`, b.summaryID},
{"update video_connections", `UPDATE video_connections SET token_ref = 'hacked' WHERE user_id = $1`, b.userID},
{"delete summaries", `DELETE FROM summaries WHERE user_id = $1`, b.userID},
{"delete summary_actions", `DELETE FROM summary_actions WHERE user_id = $1`, b.userID},
{"delete login_events", `DELETE FROM login_events WHERE user_id = $1`, b.userID},
{"delete sink_deliveries", `DELETE FROM sink_deliveries WHERE summary_id = $1`, b.summaryID},
{"delete video_connections", `DELETE FROM video_connections WHERE user_id = $1`, b.userID},
}
@@ -205,17 +215,20 @@ func TestRLSEnforcesPerUserIsolation(t *testing.T) {
`SELECT summary FROM summaries WHERE user_id = $1`, b.userID).Scan(&bSummary))
require.Equal(t, "sum", bSummary, "B's summary must be untouched by A's writes")
var bSummaries, bActions, bDeliveries, bConnections int
var bSummaries, bActions, bLogins, bDeliveries, bConnections int
require.NoError(t, super.QueryRow(ctx,
`SELECT count(*) FROM summaries WHERE user_id = $1`, b.userID).Scan(&bSummaries))
require.NoError(t, super.QueryRow(ctx,
`SELECT count(*) FROM summary_actions WHERE user_id = $1`, b.userID).Scan(&bActions))
require.NoError(t, super.QueryRow(ctx,
`SELECT count(*) FROM login_events WHERE user_id = $1`, b.userID).Scan(&bLogins))
require.NoError(t, super.QueryRow(ctx,
fmt.Sprintf(`SELECT count(*) FROM sink_deliveries WHERE summary_id = '%s'`, b.summaryID)).Scan(&bDeliveries))
require.NoError(t, super.QueryRow(ctx,
`SELECT count(*) FROM video_connections WHERE user_id = $1 AND token_ref <> 'hacked'`, b.userID).Scan(&bConnections))
require.Equal(t, 1, bSummaries, "A's DELETE must not have removed B's summary")
require.Equal(t, 1, bActions, "A's DELETE must not have removed B's action")
require.Equal(t, 1, bLogins, "A's DELETE must not have removed B's login event")
require.Equal(t, 1, bDeliveries, "A's DELETE must not have removed B's delivery")
require.Equal(t, 1, bConnections, "A's writes must not have touched B's connection")
@@ -226,3 +239,55 @@ func TestRLSEnforcesPerUserIsolation(t *testing.T) {
_ = a // a's ids are seeded for the symmetric read assertions above
}
// TestTranscriptsTableIsSharedNotRLS is the ADR-021 isolation proof: transcripts
// is the ONE shared, non-RLS surface, and the public-content classification
// leaked to nothing else. It is the inverse of TestRLSEnforcesPerUserIsolation —
// where that asserts deny-all on every user-owned table, this asserts transcripts
// is readable and writable with no user scope at all, holds no user_id, and is
// the single table with row-level security switched off.
func TestTranscriptsTableIsSharedNotRLS(t *testing.T) {
newStore(t)
super := rawPool(t)
resetDB(t, super)
app := appPool(t, super)
ctx := context.Background()
// 1. Shared + non-RLS: with NO GUC set, the app role both writes and reads a
// transcript. On an RLS table this would be deny-all (zero rows), exactly as
// the main isolation test asserts for every user-owned table.
_, err := app.Exec(ctx,
`INSERT INTO transcripts (provider, provider_video_id, source, content)
VALUES ('youtube', 'shared-vid', 'captions', 'public words')`)
require.NoError(t, err, "app role must write shared transcript content with no user scope")
require.Equal(t, 1, scopedCount(t, app, "", "transcripts"),
"transcripts must be readable with NO user scope — it is shared, non-RLS (ADR-021)")
// 2. No user_id column: the table holds only public caption content + the
// video's public id, nothing user-identifying.
var hasUserID bool
require.NoError(t, super.QueryRow(ctx,
`SELECT EXISTS (SELECT 1 FROM information_schema.columns
WHERE table_name = 'transcripts' AND column_name = 'user_id')`).Scan(&hasUserID))
require.False(t, hasUserID, "transcripts must carry no user_id (ADR-021 public content)")
// 3. The boundary is EXACTLY here: every user-owned table still has row-level
// security enabled; transcripts alone has it off. This is the proof the
// non-RLS classification was applied to transcripts and leaked nowhere else.
for _, table := range allIsolatedTables {
require.True(t, rlsEnabled(t, super, table),
"%s must still enforce row-level security — isolation must not have regressed", table)
}
require.False(t, rlsEnabled(t, super, "transcripts"),
"transcripts must be the single table with row-level security OFF (the one shared surface)")
}
// rlsEnabled reports whether a public table has ROW LEVEL SECURITY enabled.
func rlsEnabled(t *testing.T, p *pgxpool.Pool, table string) bool {
t.Helper()
var enabled bool
require.NoError(t, p.QueryRow(context.Background(),
`SELECT relrowsecurity FROM pg_class
WHERE relname = $1 AND relnamespace = 'public'::regnamespace`, table).Scan(&enabled))
return enabled
}
+1 -1
View File
@@ -24,7 +24,7 @@ import (
_ "github.com/jackc/pgx/v5/stdlib" // register the "pgx" database/sql driver for migrate
"gitea.d-ma.be/mathias/tapir/internal/domain"
"git.d-ma.be/mathias/tapir/internal/domain"
)
//go:embed migrations/*.sql
+19 -6
View File
@@ -4,15 +4,16 @@ import (
"context"
"fmt"
"os"
"path/filepath"
"testing"
embeddedpostgres "github.com/fergusstrange/embedded-postgres"
"github.com/jackc/pgx/v5/pgxpool"
"github.com/stretchr/testify/require"
"gitea.d-ma.be/mathias/tapir/internal/adapters/store"
"gitea.d-ma.be/mathias/tapir/internal/domain"
"gitea.d-ma.be/mathias/tapir/internal/ports"
"git.d-ma.be/mathias/tapir/internal/adapters/store"
"git.d-ma.be/mathias/tapir/internal/domain"
"git.d-ma.be/mathias/tapir/internal/ports"
)
// Static check: Store satisfies the Sink port.
@@ -24,11 +25,22 @@ var _ ports.Sink = (*store.Store)(nil)
var dsn string
func TestMain(m *testing.M) {
const port = 54329
// Port + runtime/data dirs are per-process (PID-derived) so two concurrent
// `go test` invocations — e.g. a push-run and a tag-run firing together in CI —
// don't collide on a fixed port or a shared data dir (which silently failed
// both runs). CachePath is shared so the PG archive is downloaded once, not
// per process. Base 54000 keeps this package's range distinct from web's.
port := uint32(54000 + os.Getpid()%1000)
dsn = fmt.Sprintf("postgres://postgres:postgres@localhost:%d/postgres?sslmode=disable", port)
rt := filepath.Join(os.TempDir(), fmt.Sprintf("tapir-epg-store-%d", os.Getpid()))
pg := embeddedpostgres.NewDatabase(
embeddedpostgres.DefaultConfig().Port(port),
embeddedpostgres.DefaultConfig().
Port(port).
RuntimePath(rt).
DataPath(filepath.Join(rt, "data")).
BinariesPath(filepath.Join(rt, "bin")).
CachePath(filepath.Join(os.TempDir(), "tapir-epg-cache")),
)
if err := pg.Start(); err != nil {
fmt.Fprintf(os.Stderr, "embedded-postgres start: %v\n", err)
@@ -40,6 +52,7 @@ func TestMain(m *testing.M) {
if err := pg.Stop(); err != nil {
fmt.Fprintf(os.Stderr, "embedded-postgres stop: %v\n", err)
}
_ = os.RemoveAll(rt)
os.Exit(code)
}
@@ -74,7 +87,7 @@ func rawPool(t *testing.T) *pgxpool.Pool {
func resetDB(t *testing.T, p *pgxpool.Pool) {
t.Helper()
_, err := p.Exec(context.Background(),
`TRUNCATE summary_actions, sink_deliveries, summaries, transcripts, videos, users CASCADE`)
`TRUNCATE login_events, summary_actions, sink_deliveries, summaries, transcripts, videos, users CASCADE`)
require.NoError(t, err)
}
+56 -1
View File
@@ -3,11 +3,12 @@ package store_test
import (
"context"
"testing"
"time"
"github.com/jackc/pgx/v5/pgxpool"
"github.com/stretchr/testify/require"
"gitea.d-ma.be/mathias/tapir/internal/adapters/store"
"git.d-ma.be/mathias/tapir/internal/adapters/store"
)
// seedBareVideo inserts a videos row with no summary, so the all-videos read and
@@ -45,6 +46,27 @@ func TestAutoSummarizeRoundTripDefaultsFalse(t *testing.T) {
require.False(t, got, "set back to manual round-trips")
}
func TestRegisteredUserDefaultsAutoSummarizeOn(t *testing.T) {
ctx := context.Background()
s := newStore(t)
resetDB(t, rawPool(t))
// A real registration creates the users row, so the column default (migration
// 011: TRUE) drives the mode — onboarded friends get auto out of the box.
id, err := s.RegisterUser(ctx, subjectA, "Alice")
require.NoError(t, err)
got, err := s.GetAutoSummarize(ctx, id)
require.NoError(t, err)
require.True(t, got, "new registrations default to auto-summarize (ADR-018)")
// The account-page toggle still works: a user can switch to manual.
require.NoError(t, s.SetAutoSummarize(ctx, id, false))
got, err = s.GetAutoSummarize(ctx, id)
require.NoError(t, err)
require.False(t, got, "the manual toggle still flips it off")
}
func TestRequestSummarizeSetsFlag(t *testing.T) {
ctx := context.Background()
s := newStore(t)
@@ -104,6 +126,39 @@ func TestListVideosReturnsSummarizedAndUnsummarized(t *testing.T) {
require.True(t, byID[videoY].SummarizeRequested, "queued video carries the flag")
}
// TestListVideosOrderedByPublishedDescNullsLast: summarized videos sort first
// (regardless of their date), then unsummarized by published_at DESC with
// undated (NULL) videos last — the recency-aligned list order (UX review B2).
func TestListVideosOrderedByPublishedDescNullsLast(t *testing.T) {
ctx := context.Background()
s := newStore(t)
p := rawPool(t)
resetDB(t, p)
const (
vSummOld = "cccccccc-cccc-cccc-cccc-cccccccccccc" // summarized, oldest date
vNewer = "dddddddd-dddd-dddd-dddd-dddddddddddd" // unsummarized, newest
vOlder = "eeeeeeee-eeee-eeee-eeee-eeeeeeeeeeee" // unsummarized, older
vUndated = "ffffffff-ffff-ffff-ffff-ffffffffffff" // unsummarized, no date
)
// Deliver first so the userA row exists (the videos FK needs it); the
// summarized video carries the OLDEST date yet must still sort first because
// it is summarized, proving summarized-first dominates the date sort.
require.NoError(t, s.Deliver(ctx, summary(userA, vSummOld, "body")))
seedVideo(t, p, userA, vSummOld, "Summarized Old", "youtube", "https://s", time.Date(2025, 1, 1, 0, 0, 0, 0, time.UTC))
seedVideo(t, p, userA, vNewer, "Newer", "youtube", "https://n", time.Date(2026, 3, 1, 0, 0, 0, 0, time.UTC))
seedVideo(t, p, userA, vOlder, "Older", "youtube", "https://o", time.Date(2026, 1, 1, 0, 0, 0, 0, time.UTC))
seedVideo(t, p, userA, vUndated, "Undated", "youtube", "https://u", time.Time{})
rows, err := s.ListVideos(ctx, userA, 50)
require.NoError(t, err)
require.Len(t, rows, 4)
got := []string{rows[0].VideoID, rows[1].VideoID, rows[2].VideoID, rows[3].VideoID}
require.Equal(t, []string{vSummOld, vNewer, vOlder, vUndated}, got,
"summarized first, then published_at DESC, NULL dates last")
}
func TestListVideosIsUserScoped(t *testing.T) {
ctx := context.Background()
s := newStore(t)
+66
View File
@@ -0,0 +1,66 @@
package store
import (
"context"
"errors"
"fmt"
"github.com/jackc/pgx/v5"
"git.d-ma.be/mathias/tapir/internal/domain"
)
// GetTranscript returns the shared, stored transcript for a video keyed by the
// cross-user dedup key (provider, providerVideoID), and whether one exists
// (ADR-021). It reads via the raw pool, NOT withUser: the table holds public
// content with no user_id and no RLS policy, so it is shared across users by
// construction. A stored SourceNone is a real hit (ok == true, HasText() ==
// false) — a known caption-less video, so the caller skips without re-fetching.
func (s *Store) GetTranscript(ctx context.Context, provider, providerVideoID string) (domain.Transcript, bool, error) {
var source, lang, content string
err := s.pool.QueryRow(ctx,
`SELECT source, COALESCE(language, ''), COALESCE(content, '')
FROM transcripts WHERE provider = $1 AND provider_video_id = $2`,
provider, providerVideoID).Scan(&source, &lang, &content)
if errors.Is(err, pgx.ErrNoRows) {
return domain.Transcript{}, false, nil
}
if err != nil {
return domain.Transcript{}, false, fmt.Errorf("store: get transcript: %w", err)
}
return domain.Transcript{
Source: domain.TranscriptSource(source),
Language: lang,
Content: content,
}, true, nil
}
// SaveTranscript upserts the shared transcript for (provider, providerVideoID).
// Only terminal outcomes belong here: SourceCaptions (with text) or SourceNone
// (no captions). A transient SourceRateLimited is rejected so persistence never
// masks a 429 as a permanent absence — that stays a per-user retry (ADR-014).
// Last write wins on conflict (a later re-fetch may correct an entry). It writes
// via the raw pool, NOT withUser — public content, shared, non-RLS (ADR-021).
func (s *Store) SaveTranscript(ctx context.Context, provider, providerVideoID string, t domain.Transcript) error {
switch t.Source {
case domain.SourceCaptions, domain.SourceNone:
// terminal — persist
case domain.SourceRateLimited:
return fmt.Errorf("store: refusing to persist transient rate-limited transcript for %s/%s", provider, providerVideoID)
default:
return fmt.Errorf("store: invalid transcript source %q", t.Source)
}
_, err := s.pool.Exec(ctx,
`INSERT INTO transcripts (provider, provider_video_id, source, language, content)
VALUES ($1, $2, $3, NULLIF($4, ''), NULLIF($5, ''))
ON CONFLICT (provider, provider_video_id)
DO UPDATE SET source = EXCLUDED.source,
language = EXCLUDED.language,
content = EXCLUDED.content,
fetched_at = NOW()`,
provider, providerVideoID, string(t.Source), t.Language, t.Content)
if err != nil {
return fmt.Errorf("store: save transcript: %w", err)
}
return nil
}
@@ -0,0 +1,102 @@
package store
import (
"context"
"errors"
"fmt"
"time"
"github.com/jackc/pgx/v5"
)
// validTranscriptStatuses bounds SetTranscriptStatus input. "" clears the status
// (column NULL); the three named states mirror migration 007's documented values.
var validTranscriptStatuses = map[string]bool{
"": true,
"none": true,
"rate_limited": true,
"fetched": true,
}
// SetTranscriptStatus records the outcome of the last transcript attempt for a
// video (migration 007). When status is "rate_limited" it also stamps
// rate_limited_at = NOW() so the runner can back off; every other status clears
// that timestamp. "" unsets the status (column NULL). An unknown status is
// rejected. Scoped via withUser, so RLS confines the UPDATE to the caller's own
// video; ErrNotFound when the user has no such video.
func (s *Store) SetTranscriptStatus(ctx context.Context, userID, videoID, status string) error {
if !validTranscriptStatuses[status] {
return fmt.Errorf("store: invalid transcript status %q", status)
}
return s.withUser(ctx, userID, func(tx pgx.Tx) error {
ct, err := tx.Exec(ctx,
`UPDATE videos
SET transcript_status = NULLIF($1, ''),
rate_limited_at = CASE WHEN $1 = 'rate_limited' THEN NOW() ELSE NULL END
WHERE id = $2`,
status, videoID)
if err != nil {
return fmt.Errorf("store: set transcript status: %w", err)
}
if ct.RowsAffected() == 0 {
return ErrNotFound
}
return nil
})
}
// GetTranscriptStatus returns a video's transcript_status ("" when unset/NULL).
// Returns ErrNotFound when the user has no such video. Scoped via withUser.
func (s *Store) GetTranscriptStatus(ctx context.Context, userID, videoID string) (string, error) {
var status string
if err := s.withUser(ctx, userID, func(tx pgx.Tx) error {
err := tx.QueryRow(ctx,
`SELECT COALESCE(transcript_status, '') FROM videos WHERE id = $1`, videoID).Scan(&status)
if errors.Is(err, pgx.ErrNoRows) {
return ErrNotFound
}
return err
}); err != nil {
if errors.Is(err, ErrNotFound) {
return "", ErrNotFound
}
return "", fmt.Errorf("store: get transcript status: %w", err)
}
return status, nil
}
// RateLimitedVideoIDs returns the user's videos currently in the "rate_limited"
// state, mapped to when the 429 was stamped (rate_limited_at). The run loop loads
// it once per pass (mirroring SeenVideoIDs) to skip re-fetching a video still
// inside the backoff window, saving caption requests. Scoped by user_id.
func (s *Store) RateLimitedVideoIDs(ctx context.Context, userID string) (map[string]time.Time, error) {
out := make(map[string]time.Time)
if err := s.withUser(ctx, userID, func(tx pgx.Tx) error {
rows, err := tx.Query(ctx,
`SELECT id, rate_limited_at FROM videos
WHERE user_id = $1 AND transcript_status = 'rate_limited' AND rate_limited_at IS NOT NULL`,
userID)
if err != nil {
return fmt.Errorf("store: rate limited video ids: %w", err)
}
defer rows.Close()
for rows.Next() {
var (
id string
at time.Time
)
if err := rows.Scan(&id, &at); err != nil {
return fmt.Errorf("store: scan rate limited id: %w", err)
}
out[id] = at
}
if err := rows.Err(); err != nil {
return fmt.Errorf("store: iterate rate limited ids: %w", err)
}
return nil
}); err != nil {
return nil, err
}
return out, nil
}
@@ -0,0 +1,80 @@
package store_test
import (
"context"
"testing"
"github.com/stretchr/testify/require"
"git.d-ma.be/mathias/tapir/internal/adapters/store"
)
func TestSetTranscriptStatus_RoundTrip(t *testing.T) {
ctx := context.Background()
s := newStore(t)
resetDB(t, rawPool(t))
id, err := s.UpsertVideo(ctx, ytVideo(userA, "rt12345", "round trip"))
require.NoError(t, err)
// Unset by default.
got, err := s.GetTranscriptStatus(ctx, userA, id)
require.NoError(t, err)
require.Equal(t, "", got)
for _, status := range []string{"none", "fetched", "rate_limited", ""} {
require.NoError(t, s.SetTranscriptStatus(ctx, userA, id, status))
got, err := s.GetTranscriptStatus(ctx, userA, id)
require.NoError(t, err)
require.Equal(t, status, got)
}
}
func TestSetTranscriptStatus_RejectsInvalid(t *testing.T) {
ctx := context.Background()
s := newStore(t)
resetDB(t, rawPool(t))
id, err := s.UpsertVideo(ctx, ytVideo(userA, "bad12345", "bad status"))
require.NoError(t, err)
require.Error(t, s.SetTranscriptStatus(ctx, userA, id, "bogus"))
// The rejected write left the status untouched.
got, err := s.GetTranscriptStatus(ctx, userA, id)
require.NoError(t, err)
require.Equal(t, "", got)
}
func TestSetTranscriptStatus_NotFound(t *testing.T) {
ctx := context.Background()
s := newStore(t)
resetDB(t, rawPool(t))
require.ErrorIs(t, s.SetTranscriptStatus(ctx, userA, videoX, "fetched"), store.ErrNotFound)
_, err := s.GetTranscriptStatus(ctx, userA, videoX)
require.ErrorIs(t, err, store.ErrNotFound)
}
func TestRateLimitedVideoIDs_StampsAndClears(t *testing.T) {
ctx := context.Background()
s := newStore(t)
resetDB(t, rawPool(t))
id, err := s.UpsertVideo(ctx, ytVideo(userA, "rl12345", "rate limited"))
require.NoError(t, err)
// Marking rate_limited stamps rate_limited_at, so the video appears.
require.NoError(t, s.SetTranscriptStatus(ctx, userA, id, "rate_limited"))
rl, err := s.RateLimitedVideoIDs(ctx, userA)
require.NoError(t, err)
require.Contains(t, rl, id)
require.False(t, rl[id].IsZero(), "rate_limited_at must be stamped")
// Moving off rate_limited clears the timestamp, so it drops out.
require.NoError(t, s.SetTranscriptStatus(ctx, userA, id, "fetched"))
rl, err = s.RateLimitedVideoIDs(ctx, userA)
require.NoError(t, err)
require.NotContains(t, rl, id)
}
@@ -0,0 +1,87 @@
package store_test
import (
"context"
"testing"
"github.com/stretchr/testify/require"
"git.d-ma.be/mathias/tapir/internal/adapters/store"
"git.d-ma.be/mathias/tapir/internal/domain"
"git.d-ma.be/mathias/tapir/internal/ports"
)
// Static check: Store satisfies the shared TranscriptStore port (ADR-021).
var _ ports.TranscriptStore = (*store.Store)(nil)
func TestSaveAndGetTranscript_RoundTrip(t *testing.T) {
s := newStore(t)
resetDB(t, rawPool(t))
ctx := context.Background()
want := domain.Transcript{Source: domain.SourceCaptions, Language: "en", Content: "the words"}
require.NoError(t, s.SaveTranscript(ctx, "youtube", "vid-1", want))
got, ok, err := s.GetTranscript(ctx, "youtube", "vid-1")
require.NoError(t, err)
require.True(t, ok, "a saved transcript must be found")
require.Equal(t, domain.SourceCaptions, got.Source)
require.Equal(t, "en", got.Language)
require.Equal(t, "the words", got.Content)
require.True(t, got.HasText())
}
func TestGetTranscript_Miss(t *testing.T) {
s := newStore(t)
resetDB(t, rawPool(t))
_, ok, err := s.GetTranscript(context.Background(), "youtube", "absent")
require.NoError(t, err, "a miss is not an error")
require.False(t, ok)
}
// A stored "no captions" outcome is a real hit: callers must skip without
// re-fetching, so ok is true even though there is no text (ADR-021 / ADR-007).
func TestSaveAndGetTranscript_NoneIsAStoredHit(t *testing.T) {
s := newStore(t)
resetDB(t, rawPool(t))
ctx := context.Background()
require.NoError(t, s.SaveTranscript(ctx, "youtube", "vid-none", domain.Transcript{Source: domain.SourceNone}))
got, ok, err := s.GetTranscript(ctx, "youtube", "vid-none")
require.NoError(t, err)
require.True(t, ok, "a stored SourceNone is a hit, not a miss")
require.Equal(t, domain.SourceNone, got.Source)
require.False(t, got.HasText())
}
// A transient 429 must never be persisted as a terminal transcript, or a later
// read would mask the rate-limit as a permanent "no transcript" (ADR-014).
func TestSaveTranscript_RejectsRateLimited(t *testing.T) {
s := newStore(t)
resetDB(t, rawPool(t))
err := s.SaveTranscript(context.Background(), "youtube", "vid-429",
domain.Transcript{Source: domain.SourceRateLimited})
require.Error(t, err)
_, ok, _ := s.GetTranscript(context.Background(), "youtube", "vid-429")
require.False(t, ok, "a rejected rate-limited save must leave nothing stored")
}
func TestSaveTranscript_UpsertLastWriteWins(t *testing.T) {
s := newStore(t)
resetDB(t, rawPool(t))
ctx := context.Background()
require.NoError(t, s.SaveTranscript(ctx, "youtube", "vid-up", domain.Transcript{Source: domain.SourceNone}))
require.NoError(t, s.SaveTranscript(ctx, "youtube", "vid-up",
domain.Transcript{Source: domain.SourceCaptions, Language: "en", Content: "now resolved"}))
got, ok, err := s.GetTranscript(ctx, "youtube", "vid-up")
require.NoError(t, err)
require.True(t, ok)
require.Equal(t, domain.SourceCaptions, got.Source)
require.Equal(t, "now resolved", got.Content)
}

Some files were not shown because too many files have changed in this diff Show More