Compare commits

..
93 Commits
Author SHA1 Message Date
mathiasandClaude Opus 4.8 0ba78e8868 docs: refresh build-state for transcript persistence (ADR-021)
CI / Lint / Test / Vet (push) Successful in 11s
CI / Build & Import (push) Successful in 10s
Update the CLAUDE.md orientation block: last tag v0.14.0, migrations
001–015, and a transcript-persistence bullet (shared non-RLS store,
engine reads stored-first). The stale "v0.9.0" reference is corrected.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-09 23:39:27 +02:00
mathiasandClaude Opus 4.8 821d5f99cd docs(bdd): scenario for transcript reuse — re-analysis never re-fetches
Capture the ADR-021 promise as a mapped BDD scenario: re-analyzing a
stored video reads the stored transcript and does not fetch captions.
Since paste-a-URL and the onboarding burst summarize through the same
engine chokepoint (resolveTranscript, store-first), this one scenario
covers their reuse path too — there is exactly one gated caption entry
point (youtube.FetchTranscript → WaitFetchGate) and one engine caller in
front of it, so the dedup is structural, not per-feature.

Mapped to TestProcessNewVideo_SecondSummarizeDoesNotRefetch.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-09 23:37:54 +02:00
mathiasandClaude Opus 4.8 5c70408e75 feat(usecase): read stored transcript before fetching (ADR-021)
The engine now resolves transcripts store-first: a stored transcript —
including a stored SourceNone — is summarized without touching YouTube,
so re-analysis never re-fetches. On a miss it fetches through the source
(caption call still gated, ADR-014) and persists the terminal outcome for
the next analysis by any user. A transient SourceRateLimited is surfaced
to the runner for per-user backoff but never cached, so persistence can
never mask a 429 as a permanent "no transcript".

The TranscriptStore is optional (nil → fetch every time), keeping the
pure-core and scaffold wiring valid. cmd/tapir wires the store as both
summary sink and transcript cache, so `tapir run` and the web summarize
path (incl. paste + onboarding) all share the dedup.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-09 23:35:04 +02:00
mathiasandClaude Opus 4.8 cb6917ca59 feat(store): shared, video-keyed transcript persistence (ADR-021)
Reshape the dead per-user transcripts table (PK videos.id, user_id,
RLS-FORCEd — never read or written by app code) into the shared public
caption store ADR-021 specifies: keyed by (provider, provider_video_id),
no user_id, NOT RLS-scoped. Migration 015 (reversible). Add
ports.TranscriptStore + Store.GetTranscript/SaveTranscript via the raw
pool (no withUser): public content, shared across users by construction.
SaveTranscript persists only terminal outcomes (captions/none) and
refuses SourceRateLimited so a transient 429 can never be stored as a
false permanent absence (ADR-014).

Flip the isolation proof: transcripts leaves the RLS-scoped set;
TestTranscriptsTableIsSharedNotRLS asserts it is the SINGLE non-RLS
surface (writable/readable with no user scope, no user_id column, RLS off
on it alone, still on every user-owned table) — the proof the
public-content classification was applied exactly here and leaked nowhere.
appPool made idempotent so two tests can build it. Adjust the 010/011/014
up-down migration tests for the new HEAD. account.go: user deletion no
longer strips shared transcripts. Reconcile data-model.md + CLAUDE.md.

Wiring the engine to read-stored-first is the next commit.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-09 23:32:12 +02:00
mathiasandClaude Opus 4.8 099b2d4c68 docs(decisions): ADR-021 — shared, video-keyed transcript persistence
Persist transcripts in a single shared table keyed by
(provider, provider_video_id) — public caption content, NOT RLS-scoped —
so re-analysis (re-summarize, paste of an already-seen video, a second
user with overlapping subs) never re-fetches from YouTube. The avoided
cost is the rate-gated, reputation-risky caption fetch (ADR-010/014), not
LLM re-summarization, which is why this reopens the transcripts half of
the "no global cross-tenant table" rejection while videos stay per-user.
Summaries remain RLS-scoped (ADR-012 unchanged). The gate is neither
bypassed nor weakened — persistence reduces fetch frequency, not pacing.

Annotate the rejected-alternatives row to record the partial reopen.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-09 23:24:11 +02:00
mathiasandClaude Opus 4.8 f66c1bcdcc feat(web): real channel filter — multi-select of the user's channels
CI / Build & Import (push) Successful in 10s
CI / Lint / Test / Vet (push) Successful in 11s
The free-text 'channel' filter was dead: it exact-matched SummaryRow.Channel,
which is just the provider ('youtube'), because videos never stored their source
channel. Now they do.

- migration 014: videos.channel_title (nullable; existing rows backfill on the
  next discovery pass, pasted videos immediately).
- discovery (NewVideos) + paste (VideoByID) populate channel_title; UpsertVideo
  persists it, preserving an existing title when an update arrives empty.
- store.DistinctChannels lists a user's channels (RLS-scoped); SummaryRow carries
  ChannelTitle via the shared projection.
- Filter: single Channel -> Channels []string, matching on ChannelTitle; the feed
  renders a multi-select of DistinctChannels (hidden until channels exist).
- migrate tests: 014 reversibility + fixed the relative-step counts in the 010/011
  up/down tests (014 shifted the topology).

TDD throughout: channel persist + distinct, adapter channel wiring, multi-channel
filter match, handler channel filter, migration up/down.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-09 23:02:54 +02:00
mathiasandClaude Opus 4.8 1e65c3b413 fix(web): show paste box to any connected user, not only on an empty feed
CI / Lint / Test / Vet (push) Successful in 11s
CI / Build & Import (push) Successful in 10s
hasConnected was computed only inside the buckets.empty() branch (it was added
for the empty-state copy), so a connected user WITH videos got hasConnected=false
and never saw the paste box (#2-regression of the v0.12.0 paste UI). Compute it
on every list render. Test: connected user with a non-empty feed sees the box.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-09 22:17:00 +02:00
mathiasandClaude Opus 4.8 87c978774f docs(bdd): scenarios for paste-a-URL + onboarding burst, mapped to tests
CI / Lint / Test / Vet (push) Successful in 10s
CI / Build & Import (push) Successful in 9s
Adds paste_url.feature (valid/invalid/not-found/dedup, +@pending no-captions)
and an onboarding-burst scenario on connect; all non-pending scenarios mapped in
the coverage gate to their existing Go tests.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-09 22:06:40 +02:00
mathiasandClaude Opus 4.8 70a9f1d4cd feat(web): in-feed paste box + onboarding-aware connect confirmation
5b: connected users get a 'Summarize any video' URL input on the feed; submit
posts to /paste (HTMX) and swaps the resulting card / inline error into the feed.
7: the connect flash now sets expectations for the async onboarding burst —
'finding your subscriptions, your newest videos will appear below as summarized'.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-09 22:06:40 +02:00
mathiasandClaude Opus 4.8 1d5b2c6365 feat(serve): wire paste fetcher + connect-time onboarding burst
CI / Lint / Test / Vet (push) Successful in 11s
CI / Build & Import (push) Successful in 10s
5c: app.Fetcher = a per-user YouTube videoFetcher, so POST /paste mounts and
resolves arbitrary-video metadata (Feature 2 goes live).

6: the connect trigger now runs an onboarding burst after discovery — summarize
up to TAPIR_ONBOARD_SUMMARIZE_COUNT of the user's newest unsummarized videos via
the gated Processor (Feature 1). Hard cap; explicit so it bypasses recency; every
fetch still through globalFetchGate. No-op when count=0 or queue-only.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-09 21:58:39 +02:00
mathiasandClaude Opus 4.8 62acfee2ed feat(web): paste-a-URL handler — add an arbitrary video + summarize (Feature 2)
POST /paste: parse the video id, fetch metadata via the VideoFetcher port (Data
API, ungated), upsert a subscription-less row scoped to the user (idempotent =
dedup), and — unless already summarized — RequestSummarize + start immediate
processing through the SAME globalFetchGate as the Summarize button. Explicit
paste overrides the recency window; a captionless video degrades to the honest
'no transcript' terminal state via the engine (ADR-010). Invalid URL -> 400,
not-found -> 404, both add nothing. Route mounts only when a Fetcher is wired.

Moves the video-not-found sentinel to domain (shared by adapter + web, no
cross-adapter coupling). Tests: valid add+queue, invalid, not-found, dedup.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-09 21:54:28 +02:00
mathiasandClaude Opus 4.8 2907801aca feat(youtube): VideoByID for arbitrary-video metadata (paste-a-URL)
videos.list (part=snippet) for a single id, including channels the user does
not follow. Data API call (1 quota unit), NOT the rate-limited caption path —
ungated metadata; only the later transcript fetch hits globalFetchGate. Returns
a subscription-less domain.Video scoped to the user, or ErrVideoNotFound for a
deleted/private/typo'd id. Foundation for Feature 2.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-09 21:43:28 +02:00
mathiasandClaude Opus 4.8 59050c4db6 feat(store): NewestUnsummarizedVideoIDs for the onboarding cap
Returns up to limit of a user's newest videos (published_at DESC, NULLS LAST)
that have no summary yet. RLS-scoped via withUser — the test proves a second
user's newer video never leaks. Drives the connect-time onboarding burst
(Feature 1); the caller routes each through the shared rate gate.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-09 21:39:40 +02:00
mathiasandClaude Opus 4.8 c320ed88aa feat(config): add TAPIR_ONBOARD_SUMMARIZE_COUNT (default 3, hard cap 5)
Bounds the connect-time onboarding summary burst (Feature 1). Hard-capped at 5
and clamped (negative->0, >cap->cap) so onboarding can never bulk-fetch; 0
disables. The cap bounds COUNT only — every fetch still flows through the shared
caption rate gate (ADR-014).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-09 21:35:05 +02:00
mathiasandClaude Opus 4.8 bddd75d92e feat(web): parse YouTube video id from pasted URL forms
Pure parser for watch?v=, youtu.be/, shorts/, embed/, and bare ids; rejects
non-YouTube hosts and malformed input. Foundation for paste-a-URL summarize
(Feature 2). No fetch, no gate interaction — parsing only.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-09 21:35:05 +02:00
mathiasandClaude Opus 4.8 e4c701c6f1 feat(discovery): trigger a discovery pass on YouTube connect
CI / Lint / Test / Vet (push) Successful in 11s
CI / Build & Import (push) Successful in 10s
A newly connected account showed no videos until the next 2h scheduled pass —
the gap that made onboarding look broken (a second user connected, saw nothing,
read as failure). The connect callback now fires an out-of-band discovery pass
for the connecting user, so videos appear promptly.

Concurrency: scheduled and connect-triggered passes share one lock (serialize),
preserving the single-fetcher invariant (ADR-018). A trigger interleaves between
the scheduler's per-user passes rather than fetching concurrently or waiting for
a whole pass. The trigger runs on the server ctx (survives the redirect) and is
non-blocking for the request goroutine.

Scope: connect-trigger only. The optional login-refresh / "Discover now" button
from #6 are intentionally not built — an unconditional login hook risks 429
storms (per the ticket's own recommendation); defer until wanted.

TDD: TestCallbackTriggersDiscovery, TestSerializeRunsOneAtATime,
TestDiscoveryTriggerEnqueueRunsUser; new BDD scenario mapped.

Refs #6

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-09 21:15:46 +02:00
mathias b527db9739 docs: bump last-tag reference to v0.9.0
CI / Lint / Test / Vet (push) Successful in 15s
CI / Build & Import (push) Successful in 11s
2026-06-09 21:03:44 +02:00
mathiasandClaude Opus 4.8 c5f556d1d6 fix(scheduler): skip discovery for users with no video connection
CI / Lint / Test / Vet (push) Successful in 11s
CI / Build & Import (push) Successful in 10s
The scheduler enumerates every user_identities row (ListAllUsers) and ran a
discovery pass for each — including users who never connected a video source.
Their per-user runner then tried to resolve a YouTube refresh token that was
never minted, logging a spurious "secrets: ref not found:
youtube/<uid>/refresh_token" every tick (e.g. stale Dex-era orphan identities
left by the Authentik migration).

Skip users whose ConnectionsForUser is empty before running their pass. Removes
the recurring noise — which actively misled a debug session into thinking a
healthy onboarded user was broken — with no change to connected users.

Refs #7

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-09 20:49:11 +02:00
mathiasandClaude Opus 4.8 a884e7e9c5 test(bdd): add scenario name-coverage gate (no godog)
CI / Lint / Test / Vet (push) Successful in 12s
CI / Build & Import (push) Successful in 10s
Close the gap where docs/use-cases/*.feature claimed to be the behavior spec
but nothing executed them — so scenarios drifted (the stale "auto summarizes
every new video" and "manual is the default" were proof).

Decision (per issue #5 BDD-runner fork): no godog — keep .feature as design
records, add a cheap name-coverage gate instead. TestScenarioCoverage parses
every scenario and asserts each non-@pending one maps to an existing Go test in
the scenarioCoverage manifest; it flags unmapped scenarios, missing/renamed
tests, and stale entries. It checks the link, not that the test exercises the
scenario (the deliberate trade for skipping godog).

Also:
- Fix the stale ADR-018 drift: "Manual is the default" -> auto is the default
  for new users; added an explicit default scenario + a plain manual scenario.
- Tag 4 documented-but-unbuilt/untested scenarios @pending with reasons (Vimeo
  connect, BYO config flow, logout->welcome, re-register-after-delete) so they
  are tracked without a false coverage claim.
- CLAUDE.md BDD section now describes the real setup (design records + the gate
  + @pending convention) instead of claiming an executable spec.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-08 22:59:25 +02:00
mathiasandClaude Opus 4.8 27fd33c99c docs: reconcile requirements + architecture with ADR-020
CI / Lint / Test / Vet (push) Successful in 11s
CI / Build & Import (push) Successful in 10s
Bring the living docs current with the recency-bounded auto-summarize + sparse
honesty + feed IA bundle (ADR-020):

- requirements (BDD): summarize_mode.feature — auto now summarizes RECENT new
  videos; added a scenario for older videos (listed, on-demand), recency note.
- architecture.md: summarization-mode + new list-surface paragraph; scheduler
  diagram + two-path table + three-phase pass now show the recency pre-filter;
  dropped stale "Summarize now".
- data-model.md: auto_summarize is recent-only, older on-demand.
- README.md: one-line recency note on the serve scheduler.
- ui-spec.md: appended the as-built ADR-020 row (supersedes earlier copy/sort).
- specs/{video-card-states,newest-first-ordering,scheduled-discovery}.md:
  superseded/extended banners pointing at ADR-020 (kept as design records).

Docs-only; task check green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-08 20:21:14 +02:00
mathiasandClaude Opus 4.8 2c96926ff7 docs: correct stale last-tag reference (v0.4.0 -> v0.8.0)
CI / Lint / Test / Vet (push) Successful in 12s
CI / Build & Import (push) Successful in 10s
CLAUDE.md "Current build state" still cited v0.4.0; the repo is at v0.8.0.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-08 15:50:06 +02:00
mathiasandClaude Opus 4.8 2cda62b3ad docs(adr): record ADR-020 recency-bounded auto-summarize + sparse honesty
CI / Lint / Test / Vet (push) Successful in 12s
CI / Build & Import (push) Successful in 10s
Document the architecture decision behind this bundle: bound auto-summarize to
a recency window (refines ADR-018; bounds load against the ADR-014 gate without
fetching harder), surface scarcity honestly, and collapse the un-summarized
back-catalogue in a single feed. Records the return-nudge as a deliberate
non-goal (it would contaminate the Stage-0 unprompted-return signal, ADR-016).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-08 14:04:37 +02:00
mathiasandClaude Opus 4.8 f29927f50d test(web): guard against the removed over-promise card copy
Extend the card copy guard so "Try now", "Summarize now", "Fetching soon", and
"the next run" can't silently return to any card state (UX review honesty pass).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-08 14:03:08 +02:00
mathiasandClaude Opus 4.8 51aa5d940c feat(web): segment watched/skipped action toggles
Watched and skipped are mutually exclusive (the store clears one when the other
is set), but rendered as three independent-looking buttons the exclusivity was
invisible. Group watched|skipped into a single segmented control and keep Saved
apart as an independent toggle (UX review C5). HTMX posting and the active/✓/
aria-pressed semantics are unchanged; extracted a shared actionButton component.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-08 14:02:31 +02:00
mathiasandClaude Opus 4.8 f775441a62 feat(web): drop the empty terms checkbox from registration
The register step asked the user to accept "the terms of use" with no terms
linked anywhere — ceremony accepting nothing on a friends-only tool (UX review
C4). Remove the checkbox and the server-side acceptance requirement; only a
display name is required now.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-08 14:01:28 +02:00
mathiasandClaude Opus 4.8 12fb031b6c feat(web): add a back link to the detail page
The summary detail page only returned to the list via the brand logo. Add an
explicit "← Summaries" link at the top (UX review C3).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-08 14:00:17 +02:00
mathiasandClaude Opus 4.8 980638d80a feat(web): slim the list filters and hide them when empty
At current scale the date-range pickers are dead weight (UX review C1/C2):

- Drop the From/To date inputs from the filter bar; keep the Channel field and
  the "Summarized only" toggle. (Filter still parses from/to for hand-built
  URLs and apply() compatibility — only the UI is removed.)
- Hide the filter bar entirely on a genuinely empty account (no rows AND no
  active filter) so the connect CTA stands alone; a filter that matches nothing
  still shows the bar so it can be cleared.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-08 13:59:44 +02:00
mathiasandClaude Opus 4.8 40b703e02a feat(web): collapse older + caption-less videos in the list
Stop the un-summarized back-catalogue from burying the readable summaries
(UX review B3/B4). One feed, with a noise-collapse — not sections:

- Summarized + recent un-summarized videos lead inline as cards.
- Un-summarized videos older than the recency window collapse into a single
  "Show N older videos — summarize on demand" disclosure (they will not
  auto-fill; they are manual-only). Window comes from App.RecencyWindow
  (= cfg.AutoSummarizeWindow); 0 disables the collapse (all inline).
- Caption-less videos collapse into one honest line ("N videos have no
  captions and can't be summarized") instead of N dead terminal cards.

bucketRows is a pure classifier (cutoff-driven; undated rows never age out);
App gains RecencyWindow + an injectable clock for the cutoff.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-08 13:55:07 +02:00
mathiasandClaude Opus 4.8 3df0459fed feat(store): order video list by published_at, NULLS LAST
ListVideos now sorts summarized-first, then published_at DESC with undated
videos last, then seen_at DESC as a tiebreak (was seen_at only). Aligns the
list with the recency framing — newest content surfaces first — so the
recency-bounded feed reads coherently (UX review B2).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-08 13:50:30 +02:00
mathiasandClaude Opus 4.8 2384c47b81 feat(runner): bound auto-summarize to a recency window
In automatic mode the scheduler now only summarizes videos published within
TAPIR_AUTO_SUMMARIZE_WINDOW (default ~7d). Older videos are still discovered
and listed — they keep the manual "Summarize" affordance — but are not
auto-processed, so a large back-catalogue (the maintainer's ~256-deep queue)
stops self-inflicting 429s against the per-IP caption gate each cycle (UX
review B1, recency design).

- runner.WithAutoWindow + Stats.SkippedTooOld; tooOld() treats a zero window
  as disabled and an undated video as never-aged-out (processed, not stranded).
- An explicit manual request bypasses the bound even in auto mode (requested
  videos are loaded in auto mode when a window is active).
- Wired through cmdRun, the scheduler's per-user runner, sumStats, and pass
  logging. config: TAPIR_AUTO_SUMMARIZE_WINDOW (default 168h), .env.example.
- Account copy (A7) updated to match: "Automatic summarizes new videos from
  about the last week; older videos stay browsable — summarize on demand."

The rate gate is untouched; the manual path still serialises through it. This
bounds auto LOAD, it does not fetch harder.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-08 13:48:04 +02:00
mathiasandClaude Opus 4.8 4a0a56e152 feat(web): lead summary detail with takeaways
The product promise is "decide what's worth your time", but the detail page
buried Takeaways — the verdict that answers that — below the full Summary.
Reorder to Takeaways → Highlights → Summary so the attention-saving payload
leads (UX review A8). Data already existed; this is a section reorder only.
Conditional sections mean a video without takeaways/highlights still leads
with the Summary naturally.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-08 13:44:20 +02:00
mathiasandClaude Opus 4.8 9bf1c31605 fix(web): correct stale welcome-page invite copy
The landing page promised "if you have an invite link, it will set up your
account automatically" — but invites moved to Authentik (ADR-019); Tapir no
longer handles invite links and "Get Started" goes straight to OIDC. Replace
with honest "invite-only — if you've been invited, sign in" and set the
gradual-fill expectation before the login wall (UX review A5). Test pins the
stale phrase out.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-08 13:43:32 +02:00
mathiasandClaude Opus 4.8 a1a5217d77 feat(web): honest sparse-state and queue copy
Make the sparse reality legible instead of implying abundance or imminence
(UX review A1-A4, A6):

- Empty-connected state drops the impossible "Run `tapir run`" instruction
  (no shell for web users; discovery is in-process since ADR-018) for a
  passive "summaries appear gradually, check back later".
- Pipeline bar reframes counts by what the user can do: "N ready · M in queue
  · K no captions" (was "summarized / fetching soon / pending").
- A one-line note explains captions are fetched slowly on purpose to respect
  YouTube's limits — turning confusing emptiness into intentional design.
- Card state for throttled videos reads "In queue", not "Fetching soon…"
  (256 items behind a per-IP gate are not all imminent — ADR-014).
- Quiet nudge button drops the over-promising "now": "Summarize", not
  "Summarize now". On click the card still honestly becomes "Queued".
- Queued card says "summarizing shortly", not "waiting for the next run"
  (no scheduler jargon).

Pure copy/label — no logic, DB, or fetch-rate change. The rate gate is
untouched; scarcity is surfaced, never engineered around.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-08 13:42:40 +02:00
mathiasandClaude Opus 4.8 9c7e3be984 docs(ux): add Stage-0 recency-bounded heuristic review
Prioritized UX findings for the product as it actually is — sparse feed,
respected caption rate limit, recency-bounded auto-summarize (incoming),
single-user. 15 findings, NOW/LATER tagged. P0s target the first-contact
return-cliff that the Stage-0 gate depends on.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-08 13:11:41 +02:00
mathiasandClaude Opus 4.8 8e45f21d23 docs: ADR-019 (Authentik owns invites), supersede ADR-017
CI / Lint / Test / Vet (push) Successful in 13s
CI / Build & Import (push) Successful in 10s
Record the invite-provisioning removal; mark ADR-017 superseded; fix
ui-spec invite-onboarding + auth-delegation sections to reflect Authentik.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-07 23:01:14 +02:00
mathiasandClaude Opus 4.8 e7c2e575d3 refactor: remove Dex local-password invite provisioning (ADR-019)
Authentik owns invites now (infra ADR-0001). Delete adapters/dex, the
/invite set-password UI, the tapir invite CLI, the InvitationStore/
DexPasswordCreator ports + App wiring, the invite Templ pages, and the
invite Taskfile target. New users are invited via Authentik, log in via
OIDC, and hit the existing /register gate. invitations table (mig 009)
left in place (append-only; harmless). task check green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-07 23:01:14 +02:00
mathias 9cd3f7e934 fix(dex): passwordName must match Dex's internal passwordID() — maps non-[a-z0-9-] to '-'
CI / Lint / Test / Vet (push) Failing after 12s
CI / Build & Import (push) Has been skipped
Tapir used human-readable substitutions ('@' -> '-at-', '.' -> '-dot-') when
deriving the Password CR name from an email. Dex's internal passwordID() maps
every non-[a-z0-9-] character to plain '-'. This caused a name mismatch:
Tapir wrote the CR as 'mathias-at-d-ma-dot-be', Dex looked it up as
'mathias-d-ma-be', got not-found, and returned 'Invalid credentials' on every
invite login — while static configmap passwords (a different code path) worked
fine. Diagnosed by adding the email to staticPasswords and confirming login
succeeded, proving the kubernetes CR lookup was the failure point.
2026-06-07 11:37:24 +02:00
mathias c812c71ecc fix(dex): store raw bcrypt hash in Password CR, not base64-encoded
CI / Lint / Test / Vet (push) Successful in 26s
CI / Build & Import (push) Successful in 12s
The original NOTE claimed Dex's kubernetes storage types Hash as []byte,
requiring the bcrypt string to be base64-encoded before storage. This was
wrong: Dex v2.41 stores and compares the hash field as a plain string. The
base64-encoding caused every invite login to fail with 'Invalid credentials'
because Dex passed the base64 bytes (starting with 'J' not '$') directly to
bcrypt. Static passwords in the configmap always used raw bcrypt strings and
worked fine — confirming the dynamic CR encoding was the bug.
2026-06-07 09:28:11 +02:00
mathias 084b73907d feat(task): add invite task — task invite EMAIL=user@example.com
CI / Lint / Test / Vet (push) Successful in 12s
CI / Build & Import (push) Successful in 10s
2026-06-07 07:35:47 +02:00
mathias 0b04e487ad docs: update card-state model — 'Summarize now' verb, no-captions state, five-state table
CI / Lint / Test / Vet (push) Successful in 12s
CI / Build & Import (push) Successful in 10s
2026-06-06 22:42:37 +02:00
mathias 2fedc45443 feat(web): unify card states — one 'Summarize now' verb, honest no-captions state
CI / Lint / Test / Vet (push) Successful in 13s
CI / Build & Import (push) Successful in 10s
Five explicit footer states, status-primary:
1. Summarized — chip + actions, no button (unchanged)
2. No captions (TranscriptStatus=="none") — NEW: 'No transcript available' muted
   text, no button, no POST URL. Removes the dead-end 'Summarize' button that
   tried and failed when there were no captions to fetch.
3. Queued (SummarizeRequested) — chip + muted text, no button (unchanged)
4. Rate-limited — 'Fetching soon…' + quiet 'Summarize now' → /retry-now
5. Pending — 'Not summarized' + quiet 'Summarize now' → /summarize

One verb ('Summarize now'), one quiet style (.btn-quiet, renamed from .btn-retry
which was state-specific). User doesn't see the internal pipeline distinction;
both buttons post to their existing handlers unchanged. Form class renamed
card-nudge-form. Dropped engineer-facing tooltip; user-facing hint added.
'Try now' wording removed entirely.
2026-06-06 22:31:41 +02:00
mathias 58cd68c1eb docs: spec video-card state unification — one "Summarize now" verb + honest no-captions state
CI / Lint / Test / Vet (push) Successful in 12s
CI / Build & Import (push) Successful in 10s
The card shows two verbs ("Try now" for rate-limited, "Summarize" for pending)
for one user intent, and — the real bug — a no-captions video falls into the
pending branch and wrongly shows a Summarize button that can only fail. Spec
collapses to one quiet "Summarize now" verb wherever a nudge is possible (both
handlers unchanged underneath), adds an honest no-button "No transcript
available" state, and keeps the card status-first (buttons are exceptions in auto
mode). View-layer only.
2026-06-06 19:58:45 +00:00
mathias e472015c76 docs: replace evasion framing with honest onboarding-prioritisation rationale
CI / Lint / Test / Vet (push) Successful in 12s
CI / Build & Import (push) Successful in 10s
'Try now' and the newest-first batch implement onboarding prioritisation:
foreground (user-clicked 'Try now') summarises a chosen video on demand;
background batch summarises newest-first; both honour the shared rate gate.

Remove any prior framing that described 'Try now' as making traffic 'look
organic to YouTube' or as rate-limit evasion — that was not the rationale
and contradicts ADR-014's explicit account-safety constraint.

Correct statement: rate limiting is respected, not evaded. TAPIR_FETCH_RATE
and TAPIR_FETCH_BACKOFF are honest rate controls; they govern how fast Tapir
fetches captions, not how the requests appear to YouTube.

Architecture: add two-path model table (foreground/background, both through
globalFetchGate) and newest-first batch ordering doc (three-phase RunOnce,
before/after example).

ui-spec: add 'Try now' row with correct rationale; add pipeline stats bar row;
update Summarized-only filter row to mention sort-to-top.
2026-06-06 21:29:28 +02:00
mathias 0c0225f9c6 feat(runner): process candidates newest-first within each pass (ADR-018)
Restructures RunOnce from per-channel inline processing to collect-sort-process:

Phase 1 — discover, persist (UpsertVideo), apply pre-filters (seen/manual/backoff)
           and collect surviving candidates with their discovery position.
Phase 2 — sort candidates by published_at DESC, NULLS LAST, pos ASC tiebreak
           so videos with no publish date never jump ahead of dated content.
Phase 3 — process in sorted order through the unchanged globalFetchGate.

Before (per-channel): chanA=[v-old, v-mid], chanB=[v-new, v-null]
                    → [v-old, v-mid, v-new, v-null]
After  (newest-first): [v-new, v-mid, v-old, v-null]

Same set of videos processed; only the order changes within a pass. All existing
behaviour is preserved: failure isolation, backoff skip, manual mode,
channel-unavailable, stats. In-memory sort; no new table or persisted queue.

The ordering is onboarding prioritisation — new users get summaries of their most
recent, relevant videos first; the back-catalogue fills in behind across subsequent
passes. Both this background batch and the foreground 'Try now' button honour the
shared globalFetchGate: rate limiting is respected, not evaded.
2026-06-06 21:29:28 +02:00
mathias 8403e8e524 docs: spec newest-first batch ordering + honest Try-now/prioritisation docs
CI / Lint / Test / Vet (push) Successful in 12s
CI / Build & Import (push) Successful in 11s
Product intent: new users get summaries of their newest videos fast while the
back-catalogue fills behind, within the shared rate gate. The background batch
currently processes in subscription/channel order, not newest-first — this spec
closes that gap (collect candidates, sort published_at DESC NULLS LAST, process
in order, gate unchanged). Also corrects the docs to describe Try-now as
onboarding prioritisation, explicitly removing the prior "looks organic to
YouTube" traffic-disguising framing — rate limiting is respected, not evaded.
2026-06-06 19:20:27 +00:00
mathias 57e29ca06c feat(web): pipeline stats bar, summarized-first sort, Try now button for rate-limited videos
CI / Lint / Test / Vet (push) Successful in 15s
CI / Build & Import (push) Successful in 10s
Three UX improvements for the pending-transcript state:
1. Summarized videos sort to top (ORDER BY (s.id IS NOT NULL) DESC, seen_at DESC)
   so completed summaries are always immediately visible without filtering.
   ListVideos default limit raised from 50 to 500 to show the full backlog.
2. Pipeline stats bar above the video list: '2 summarized · 256 fetching soon · 12
   no captions' — computed from the unfiltered row set, hidden when everything is
   summarized.
3. 'Try now' button on rate-limited cards replaces the passive 'Retrying later'
   chip. POST /v/{id}/retry-now clears rate_limited_at then calls ProcessVideo
   through the shared globalFetchGate — same rate limiting as the scheduler, safe
   under concurrent use.
2026-06-06 19:20:20 +02:00
mathias 24f2a69eaa fix(web): gofmt view.go
CI / Lint / Test / Vet (push) Successful in 12s
CI / Build & Import (push) Successful in 9s
2026-06-06 11:31:28 +02:00
mathias 241eebfd9f docs(homelab): add TAPIR_FETCH_BACKOFF config, update snapshot date 2026-06-06 11:31:11 +02:00
mathias 317b0d4834 docs(ui-spec): add Dex auth details, invite onboarding, summarized-only filter, unavailable channels 2026-06-06 11:29:14 +02:00
mathias 252a4ebd9e docs(architecture): add in-process scheduler sequence, rate gate description, auto_summarize default fix 2026-06-06 11:27:18 +02:00
mathias 1ad1966672 docs(data-model): add CHANNEL_ERRORS + LOGIN_EVENTS entities, transcript_status columns, auto_summarize default update 2026-06-06 11:25:52 +02:00
mathias ccadcecfef docs(readme): add tapir serve, fix TAPIR_DISCOVERY_INTERVAL reference 2026-06-06 11:24:23 +02:00
mathias 4c5d3cca81 feat(web): add 'Summarized only' filter checkbox to video list
CI / Lint / Test / Vet (push) Failing after 3s
CI / Build & Import (push) Has been skipped
2026-06-06 11:19:18 +02:00
mathias 10233ee881 fix(scheduler): include ChannelUnavailable in sumStats + pass-complete log
CI / Lint / Test / Vet (push) Successful in 12s
CI / Build & Import (push) Successful in 9s
2026-06-06 10:22:37 +02:00
mathias f1e9739900 feat(store,runner,web): channel unavailability notice (migration 013)
CI / Lint / Test / Vet (push) Successful in 26s
CI / Build & Import (push) Successful in 11s
YouTube channels that 404 on playlist discovery (deleted/private) are now:
1. Wrapped in domain.ErrChannelUnavailable by the YouTube adapter (instead of
   a generic error), so the runner can identify them without string-matching.
2. Stored per-user in channel_errors (migration 013, RLS-guarded) via runner's
   new UpsertChannelError path — removed from the generic Errors counter,
   counted separately as ChannelUnavailable.
3. Shown on the account page under "Unavailable channels" with name, chip-warn
   badge, and first-seen date, so users know why some subscribed channels
   produce no videos.
2026-06-06 10:09:52 +02:00
mathias 940f80899a fix(store): migration 012 — back-fill auto_summarize via RLS bypass
CI / Lint / Test / Vet (push) Successful in 12s
CI / Build & Import (push) Successful in 11s
Migration 011's UPDATE ran without tapir.current_user_id set, so FORCE RLS
blocked all rows and 0 users were updated (skipped_manual=607 in scheduler).
Migration 012 temporarily drops FORCE so the table owner can run the UPDATE,
then restores it.
2026-06-06 10:01:13 +02:00
mathiasandClaude Opus 4.8 f35c2a85a5 docs: scheduled-discovery env + single-replica constraint; VISION gate-clock reset
CI / Lint / Test / Vet (push) Successful in 13s
CI / Build & Import (push) Successful in 11s
homelab-integration.md gains a "Scheduled discovery" section documenting
TAPIR_DISCOVERY_INTERVAL and TAPIR_FETCH_RATE and the load-bearing
single-replica constraint (in-process scheduler → replicas: 1 is required;
>1 double-runs discovery). VISION Stage 0 carries a pointer to ADR-018's
gate-clock reset so nothing in docs implies the window started before
unprompted use was possible.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-05 23:43:17 +02:00
mathiasandClaude Opus 4.8 5d029a2823 feat(store): default auto_summarize ON for new users (ADR-018)
Migration 011 flips the auto_summarize column default to TRUE and brings
existing rows (maintainer + current registrations) along. Onboarded friends
now get zero-friction discovery: scheduled discovery (ADR-018) both discovers
AND summarizes new videos, so a user's list fills and summarizes itself
instead of presenting an empty list of manual Summarize buttons.

Safe only because the process-wide caption-fetch rate gate (ADR-014 item 2,
prior commit) now exists — auto + scheduled + multi-user would otherwise
self-inflict 429s every cycle. The down migration reverts the default but
intentionally leaves existing rows as-is (no surprise manual regression on
rollback). RegisterUser already lets the column default drive the value, so
no app change is needed; the account-page manual toggle still works.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-05 23:42:23 +02:00
mathiasandClaude Opus 4.8 f6623afd41 feat(serve): in-process scheduled discovery for all users (ADR-018)
Stage 0's "returns and reads in >=2 weeks" gate can't be met while discovery
is host-side manual (`tapir run`): a newly onboarded user sees an empty list
and never comes back. Make Tapir watch on its own.

cmdServe launches a background goroutine (when TAPIR_DISCOVERY_INTERVAL > 0)
that runs a discovery pass for ALL users on that cadence: enumerate via the
un-RLS'd ListAllUsers, then run each user's pass through the EXISTING
runner.Runner — the only new code is the per-user loop, not a new scheduler.
Run-once-on-startup then ticked; ctx-cancelled on SIGTERM; per-user failures
(including buildUserRunner errors) are logged and skipped so one bad user
never aborts the rest. interval <= 0 disables it entirely (dev/tests).

buildUserRunner binds each runner to that user's own YouTube refresh token
(web.YouTubeTokenRef) — the Stage-1 per-tenant ref — reusing buildProcessor's
engine wiring. SetFetchRate is also wired in cmdServe so the click-path shares
the gate.

SINGLE-REPLICA is now load-bearing: the loop lives in the web process, so >1
replica double-runs discovery (429s + duplicate work). Documented in cmdServe
and warned at startup.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-05 23:39:39 +02:00
mathiasandClaude Opus 4.8 149ec2adae feat(store): ListAllUsers for scheduler user enumeration
The in-process scheduler (ADR-018) needs to enumerate every user to run a
discovery pass each. user_identities is the un-RLS'd map; add ListAllUsers as
a plain pool query (no withUser) — the same enumerate-then-act pattern
UserBySubject and the login_events gate query established. Scoping it to a
single user would defeat the point; user_identities carries no RLS by design.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-05 23:36:44 +02:00
mathiasandClaude Opus 4.8 f5021a8436 feat(youtube): process-wide caption-fetch rate gate (ADR-014 item 2)
ADR-014 item 2 — a single per-egress-IP rate gate shared by every caption
fetch — was specced but only per-video backoff (rate_limited_at) shipped.
Build the real gate now: it is load-bearing once ADR-018 puts auto-summarize
on an in-process schedule across multiple users (all fetches leave one pod's
egress IP, concurrently with live "Summarize" clicks — without a shared gate
that self-inflicts 429s every cycle).

globalFetchGate (golang.org/x/time/rate, default 2s/req burst 1) is consulted
in httpDo before every live outbound fetch — player, watch-page, timedtext —
so the scheduler runners and the web click-path serialise through one limiter
regardless of how many users/goroutines are upstream. The test seam
(a.transport != nil) skips the gate so fakes are not throttled.

TAPIR_FETCH_RATE (Go duration, default 2s, 0 = unlimited) wires SetFetchRate in
cmdRun; the existing per-video backoff stays as the complementary 429 handler.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-05 23:36:17 +02:00
mathias 88294d38bc docs: add ADR-018 — in-process scheduled discovery, auto-summarize, gate-clock reset
CI / Lint / Test / Vet (push) Successful in 12s
CI / Build & Import (push) Successful in 10s
Records the Stage-0 usability decision: tapir serve runs discovery for all users
on an interval (reusing runner.Loop), auto-summarize defaults ON so the list fills
itself, and the gate clock resets to when this ships (unprompted use was
impossible before, so the prior window measured nothing — framed as starting the
clock when the experiment can run, not dodging a failing gate). Records the
in-process-vs-CronJob tradeoff, the load-bearing single-replica constraint, and
that it makes ADR-014 item 2 (process-wide rate gate) a hard requirement folded
into the build. Cross-referenced ADR-014/016; added CronJob to rejected-alts.
2026-06-05 13:03:18 +00:00
mathias ca9bf2657f docs: spec in-process scheduled discovery + auto-summarize + rate-gate finish
CI / Lint / Test / Vet (push) Successful in 12s
CI / Build & Import (push) Successful in 13s
The Stage-0 usability fix: tapir serve runs discovery for all users on an
interval (reusing the existing Runner.Loop, enumerate-users-then-withUser),
auto-summarize defaults ON for Future-B users so the list fills itself, and the
ADR-014 process-wide per-egress-IP rate gate is confirmed/finished in the same
slice because in-process + auto + multi-user makes it load-bearing. Records the
single-replica constraint as load-bearing, and a fallback (auto-summarize OFF
until the gate exists) so the dangerous combination never ships half-built.
2026-06-05 12:59:50 +00:00
mathias 4706c508e9 docs: ratify ADR-017 KEEP — invite flow is load-bearing (some users can't use Google OIDC)
CI / Lint / Test / Vet (push) Successful in 12s
CI / Build & Import (push) Successful in 10s
Records the deliberate keep-or-reverse decision on the Dex-write invite flow.
Decision: KEEP. Deciding fact: not all intended Future-B users will use Google
accounts, so Google OIDC alone can't onboard them — the invite flow is
load-bearing, not redundant. Trust-surface cost accepted deliberately, explicitly
NOT as a precedent for widening further, and explicitly NOT by adding delete RBAC
to fix the orphan gap. Open items reframed as tracked follow-ups (verify RBAC
against the real manifest; accept orphan for Future B, revisit before Future C).
Added the reversal to rejected-alternatives.
2026-06-05 12:47:39 +00:00
mathias b4da97e0e2 docs: resolve ADR-014 migration-008 false alarm (cols shipped in 007)
CI / Lint / Test / Vet (push) Successful in 13s
CI / Build & Import (push) Successful in 10s
Reconciliation finding: the v0.6.0 report's "migration 008 videos.rate_limited_at"
was a mislabel. Both transcript_status and rate_limited_at shipped in migration
007; the sequence legitimately skips 008, nothing was lost, and the runner's
column reads are sound. Updated the ADR-014 implementation note to state this
(was flagged as an unresolved discrepancy). The item-2 (shared per-egress-IP rate
gate vs per-video backoff) flag stays open — still unconfirmed.
2026-06-04 18:54:56 +00:00
mathiasandClaude Opus 4.8 561ba79360 feat(cli): tapir report — Stage-0 usage gate query
CI / Lint / Test / Vet (push) Successful in 12s
CI / Build & Import (push) Successful in 10s
ActiveWeeks computes per-user distinct active weeks (reads login_events UNION
acts summary_actions) for the gate (VISION/ADR-016: usage in >=2 distinct
weeks). Because the user-owned tables are FORCE RLS under a non-superuser owner,
a single cross-user query is deny-all; instead it enumerates users from the
un-RLS'd identity map and counts each inside withUser — no privilege escalation,
no policy change.

`tapir report` prints the per-user table and the pass/fail verdict (needs only
TAPIR_DB_DSN). Pure formatter + store query are unit-tested, including the
cross-table shared-week dedup.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-03 23:46:03 +02:00
mathiasandClaude Opus 4.8 2fac735837 feat(web): stamp login_events on every gated request
The registration gate, once it resolves the authenticated subject to a tapir
user_id, calls StampLogin (store-throttled to one row per user per day). Best-
effort: a stamp failure is logged and swallowed so it never breaks the request.
This is what makes the read-side Stage-0 usage signal actually accrue.

Tests cover the happy-path stamp, the same-day throttle, and that an
unregistered subject (redirected to /register) is never stamped.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-03 23:46:03 +02:00
mathiasandClaude Opus 4.8 de54cd33b2 feat(store): throttled per-day StampLogin + login_events delete cascade
StampLogin appends one login_events row per user per day via an atomic
INSERT ... SELECT ... WHERE NOT EXISTS, run through withUser so the throttle
probe is itself RLS-scoped to the caller. DeleteUser now deletes login_events
explicitly (no FK = no cascade — the summary_actions footgun, repeated).

Extends the two-user RLS isolation proof and the delete-account proof to cover
login_events, and adds throttle / new-day / user-scoping tests.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-03 23:46:03 +02:00
mathiasandClaude Opus 4.8 b070347597 feat(store): add append-only login_events table with forced RLS
Stage-0 usage measurement (VISION/ADR-016): summary_actions captures acts
(watch/skip/save) but not reads. A reader who logs in weekly and clicks
nothing is invisible — for a reading product that return is the signal the
gate ("usage in >=2 distinct weeks") is defined on. login_events records
THAT a user was active, append-only, one row per user per active day.

Per-user isolation via the same GUC-keyed FORCE RLS policy as migration 003.
No FK to users (mirrors summary_actions) — the cascade footgun is handled by
DeleteUser in a later commit. Adds an up/down reversibility test and registers
the table in both truncate helpers.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-03 23:46:03 +02:00
mathias 99743af182 docs: add ADR-017 (Dex-write invite flow) retroactively; annotate ADR-002/013/014
CI / Lint / Test / Vet (push) Successful in 12s
CI / Build & Import (push) Successful in 10s
Reconciliation pass after parallel agent sessions shipped v0.6.0/v0.7.0.

ADR-017 documents the v0.7.0 invite flow, which gave Tapir scoped create+get on
passwords.dex.coreos.com in the auth namespace — Tapir now WRITES to the shared
identity provider. This shipped with no ADR; recorded retroactively with the
principle-reversal named (partially supersedes ADR-002/013), the security analysis
(bounded RBAC, but a real larger trust surface), and the open gaps (orphaned Dex
accounts on delete; plaintext invite tokens). Cross-referenced in ADR-002 and
ADR-013 status lines so a future reader isn't misled.

ADR-014 annotated: 429 handling shipped in v0.6.0 but decision item 2 (shared
per-egress-IP rate gate) appears realised as per-VIDEO backoff, not a process-wide
IP gate — flagged not-confirmed-done. Also flags the missing migration 008 /
rate_limited_at discrepancy (v0.6.0 report cited 008; tree jumps 007->009).

No code changed in this commit — audit trail only.
2026-06-03 21:38:49 +00:00
mathias 070491261d fix(web): hide nav auth links on public pages (/welcome, /invite)
CI / Lint / Test / Vet (push) Successful in 11s
CI / Build & Import (push) Successful in 10s
Logged-out visitors on /welcome and /invite should not see Account or
Log out. Split Layout into Layout (authenticated, full nav) and
PublicLayout (public, brand-only header). WelcomePage + InvitePage
variants now use PublicLayout.
2026-06-03 23:25:24 +02:00
mathiasandClaude Opus 4.8 72bf8a5553 docs(env): document TAPIR_PUBLIC_URL for tapir invite
CI / Lint / Test / Vet (push) Successful in 12s
CI / Build & Import (push) Successful in 10s
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-03 23:21:06 +02:00
mathiasandClaude Opus 4.8 dece5dec44 feat(web): public /invite/{token} set-password + account-creation flow
The Stage-1 onboarding path: an invited user opens their emailed link,
sets a password, and Tapir creates their Dex local-password account so
they can log in. Mounted on root OUTSIDE Auth.Middleware — the visitor
has no Dex session yet; the token in the path is the capability.

handleInviteForm previews the token (no consume) and shows the form, or
a clear "expired / already used" page. handleInviteSubmit validates the
password BEFORE consuming the token (a typo is retryable), then claims
the invite exactly once, bcrypt-hashes (cost 12), and creates the Dex
account — mapping ErrPasswordExists -> "log in instead" and ErrForbidden
-> "contact the administrator". Off-cluster (App.Dex nil) it degrades to
a "deployed-only" message without burning the token. On success it sets
an account_created flash and redirects to /auth/login.

Welcome sub-text now states access is invite-only. Handlers depend on
narrow ports (InvitationStore, DexPasswordCreator) so tests use fakes;
cmdServe wires the store + an in-cluster dex.PasswordClient.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-03 23:19:22 +02:00
mathiasandClaude Opus 4.8 893886a60a feat(cli): tapir invite <email> + TAPIR_PUBLIC_URL config
Mints a single-use invitation and prints the absolute claim URL for the
operator to send. The URL base is TAPIR_PUBLIC_URL (default
https://tapir.d-ma.be). runInvite is factored from config/store wiring so
it's unit-tested against a fake inviter — no Postgres.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-03 23:15:24 +02:00
mathiasandClaude Opus 4.8 e44485df16 feat(dex): in-cluster Password CR client for local-password accounts
Writes passwords.dex.coreos.com CRs against the in-cluster Kubernetes
API using the pod's service-account token + cluster CA (no kubectl /
client-go dependency). NewPasswordClient returns ErrNotInCluster off
cluster so the web layer degrades gracefully in dev.

Load-bearing: Dex's kubernetes storage types Password.Hash as []byte,
which k8s JSON-marshals as base64 — so the `hash` field carries the
base64 of the bcrypt string, not the raw string. Storing the raw string
makes Dex's base64-decode-on-login produce garbage and every login fail.

409 -> ErrPasswordExists, 401/403 -> ErrForbidden (RBAC missing) so the
handler can give precise messages. Tested against an httptest TLS server.

bcrypt cost-12 hashing lives in the web handler; golang.org/x/crypto was
already a transitive dep (now promoted in go.sum).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-03 23:14:08 +02:00
mathiasandClaude Opus 4.8 8b7ef07ba3 feat(store): invitations table + create/peek/claim methods
Stage-1 email onboarding: Mathias mints an invite, the recipient claims
it to set a Dex password. Invitations exist before their user, so the
table carries no user_id FK and is deliberately outside RLS — the
32-byte crypto-random token is the capability (single-use, time-boxed).

ClaimInvitation consumes atomically (UPDATE ... WHERE used_at IS NULL
... RETURNING) so concurrent claims of one token can't both succeed.
PeekInvitation validates the link for the form without consuming it.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-03 23:12:40 +02:00
mathias 1cf58768ed docs: spec the Stage 0 usage-measurement build (login events)
CI / Lint / Test / Vet (push) Successful in 10s
CI / Build & Import (push) Successful in 10s
Small tapir slice to make the gate measurable as written: append-only
login_events (RLS, per-user-per-day throttle) + a union query over reads
(login_events) and acts (summary_actions) for distinct-active-weeks. Carries the
honesty caveats (unprompted not measurable; data accrues from deploy; week-bucket
noise at low N) and the delete-cascade footgun (no FK, needs explicit delete +
test) from the prior delete work. Out of scope: analytics, prompt-tracking,
dashboards.
2026-06-03 20:59:49 +00:00
mathias f45ba35e25 docs: VISION Stage 0 — keep "unprompted" as ideal, note measurement gap
CI / Lint / Test / Vet (push) Has been cancelled
CI / Build & Import (push) Has been cancelled
Reframes "unprompted" from an enforced criterion to a named measurement
limitation: organic-vs-prompted returns aren't distinguishable from any data
Tapir holds, so in practice all returns are counted and the result read with that
caveat (a nudged return is a weaker signal). Adds a "how it's measured" note
pointing at summary_actions (acts) + a new append-only login-events table
(read-returns), which accrue from deploy onward. Honest about the gap rather than
silently dropping the word.
2026-06-03 20:59:14 +00:00
mathiasandClaude Opus 4.8 943554a96c feat(web): "Retrying later" badge for rate-limited videos
CI / Lint / Test / Vet (push) Successful in 10s
CI / Build & Import (push) Successful in 10s
A discovered-but-unsummarized video whose caption fetch was rate-limited now
shows a passive  "Retrying later" chip (dim CharmDim styling, not the accent)
instead of the Summarize button — the user cannot fix a 429, the runner retries
automatically once the backoff window expires. Regenerated views_templ.go.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-03 22:56:48 +02:00
mathiasandClaude Opus 4.8 40a614c8d4 feat(runner): 429 backoff — skip still-throttled videos, persist status
After ProcessNewVideo the runner records transcript_status per outcome:
rate_limited (stamps the backoff clock), none, or fetched. Before fetching, a
video inside the TAPIR_FETCH_BACKOFF window is skipped (SkippedRateLimited) so a
just-429'd caption endpoint is not re-hit; once the window expires it retries.

Backoff/clock injected via variadic Options (WithBackoff, WithClock) so existing
New call sites and the fake-driven loop tests stay valid. Backoff 0 = always retry.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-03 22:56:07 +02:00
mathiasandClaude Opus 4.8 ce2fc62ef8 feat(config): TAPIR_FETCH_BACKOFF for rate-limit retry window
Adds FetchBackoff (Go duration, default 1h) controlling how long the run loop
waits before re-fetching a transcript that returned HTTP 429. Zero means always
retry. Not required by ValidateForRun — a zero/unset value is a valid policy.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-03 22:54:24 +02:00
mathiasandClaude Opus 4.8 0ceacc8230 feat(usecase): surface TranscriptSource on ProcessResult
The engine already distinguishes SourceNone from SourceRateLimited internally
but collapsed both into Skipped. Expose the source string so the runner can
persist the right transcript_status and apply rate-limit backoff, without the
engine taking on any store/retry concern (dependencies still point inward).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-03 22:53:11 +02:00
mathiasandClaude Opus 4.8 1e81965519 feat(store): TranscriptStatus read field + status setter/loader
Surfaces videos.transcript_status (migration 007) on SummaryRow and adds
SetTranscriptStatus / GetTranscriptStatus / RateLimitedVideoIDs.

SetTranscriptStatus is the single choke point for the rate-limit lifecycle:
"rate_limited" stamps rate_limited_at = NOW(), every other status clears it,
so the runner's backoff window and the UI badge read one consistent source.
RateLimitedVideoIDs is the per-pass loader (mirrors SeenVideoIDs) the runner
uses to skip still-throttled videos without re-hitting the caption endpoint.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-03 22:52:48 +02:00
mathias f50c072d65 fix(web): add Log out to the persistent nav header
CI / Lint / Test / Vet (push) Successful in 12s
CI / Build & Import (push) Successful in 10s
Logout was only reachable from /welcome. Users who are logged in had no way
to sign out from any app page (list, detail, account). Added to the shared
nav alongside Account.
2026-06-03 22:49:14 +02:00
mathias e6f508824b docs: add ADR-016 — Stage 0 gate revised to "me or a friend", behavioural
CI / Lint / Test / Vet (push) Successful in 19s
CI / Build & Import (push) Successful in 11s
Records the gate change: Stage 0 now passes when either the maintainer or an
onboarded friend returns unprompted in >=2 separate weeks. Behavioural (return
usage), not feedback-based, to resist politeness bias. Includes an honest
self-scrutiny note that this is a guardrail edit made while the original gate was
unmet — examined on that basis and proceeding because it broadens who supplies the
signal without softening the kind of signal required. Adds feedback-based-gate to
rejected alternatives.
2026-06-03 20:46:40 +00:00
mathiasandClaude Opus 4.8 689500c85e feat(store): migration 007 — per-video transcript status
CI / Lint / Test / Vet (push) Successful in 20s
CI / Build & Import (push) Successful in 10s
Adds videos.transcript_status (NULL|none|fetched|rate_limited) and
videos.rate_limited_at, so the runner can record a 429 and skip re-fetching a
still-throttled video until a backoff window elapses. Columns inherit the
existing videos RLS policy (migration 003); no policy change needed.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-03 22:45:29 +02:00
mathiasandClaude Opus 4.8 4678d473b8 feat(youtube): map caption 429 to SourceRateLimited
The baseUrl fetch mapped every non-200 to SourceNone, recording a 429 as a
permanent "no captions". 429 is the IP being rate-limited, not an absent
transcript. Return SourceRateLimited (still a graceful degrade, no error) so
the runner can retry after a backoff window. Other non-200s (403/404/5xx)
stay SourceNone.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-03 22:45:29 +02:00
mathiasandClaude Opus 4.8 c63b2de66d fix(web): actionable empty state for the summary list
Videos rows are created by `tapir run`, not when a YouTube account is
connected, so a freshly-connected account correctly shows an empty list —
but the old empty state ("No videos yet") gave no clue why or what to do.
Split it on whether the user has any connection:

- connected, no videos: a distinct accent callout telling them to run
  `tapir run` to discover subscriptions.
- not connected: a prompt with a Connect YouTube button.

handleList fetches connections only when the list is empty. Includes
web-shot captures of all three states under docs/ux-review/fixes/.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-03 22:45:29 +02:00
mathiasandClaude Opus 4.8 27aa319f1d feat(domain): add SourceRateLimited transcript source
429 from the caption endpoint means the IP is rate-limited (retry later),
not that the video has no captions. Distinguishing it from SourceNone is the
prerequisite for the runner's backoff/retry logic.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-03 22:45:29 +02:00
mathiasandClaude Opus 4.8 61d4d5bc4a fix(web): clarify landing CTA copy for new users
The "Get Started" button drops users straight into the shared Dex flow,
which has no separate "register" option — registration completes
automatically after first login. Users new to Tapir had no signal that
signing in is also how they sign up. Reword the sub-text to say so
explicitly, keeping the single Dex CTA (sign-in and sign-up are one flow).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-03 22:45:29 +02:00
mathiasandClaude Opus 4.8 483730cd03 fix(web): make anchor-styled buttons readable
The .btn class sets color:var(--accent-fg), but the generic a{} and
a:visited{} rules outrank it on <a> elements, so anchor buttons (the
landing "Get Started" CTA, "Connect YouTube") rendered their label
accent-on-accent — invisible. Add a.btn / a.btn:visited to restore the
button foreground.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-03 22:45:29 +02:00
mathias 477701fea2 docs: revise Stage 0 gate to "useful to me or a friend" (behavioural)
CI / Lint / Test / Vet (push) Successful in 11s
CI / Build & Import (push) Successful in 10s
Replaces the original "useful to me, specifically" gate with "me OR a friend
returns unprompted in >=2 separate weeks" — friendly-user signal counts, but the
test stays behavioural (return usage) not feedback-based, to resist politeness
bias. Folds the old Stage 1 ("a trusted user returns") into the new Stage 0 (they
were near-identical), renumbers hardening to Stage 1, and updates the drift
signals (the gate can be softened by mistaking polite feedback for evidence;
multi-user shipping ahead of the gate was a recorded exception per ADR-012, not a
precedent). Rationale recorded in ADR-016.
2026-06-03 20:44:19 +00:00
mathias 17fad140a6 docs: add ADR-015 — per-user credentials envelope-encrypted in PG18
CI / Lint / Test / Vet (push) Successful in 10s
CI / Build & Import (push) Successful in 10s
Records the infra#88 spike decision: per-user OAuth tokens are runtime app-state,
not config, so they live envelope-encrypted in PG18 under RLS (key from 1P via the
existing read-only SA) rather than in the vault. Infra creds stay ESO/1Password —
two mechanisms because they're two different things. Includes the falsification
conditions (frequent rotation; estate audit policy; key-rotation cost) so the
choice is earned not assumed. Adds the two rejected candidates (vault-write SA;
Supabase) to the rejected-alternatives table. Full reasoning in the #88 decision
doc; build + reboot-validation in #89.
2026-06-03 20:29:50 +00:00
99 changed files with 7151 additions and 854 deletions
+28
View File
@@ -46,3 +46,31 @@ TAPIR_SECRETS_FILE=
# --- run loop -------------------------------------------------------------
# Empty/0 = single pass. Set (e.g. 15m) to poll on that cadence.
TAPIR_POLL_INTERVAL=
# How long to wait before re-fetching a transcript that returned HTTP 429
# (rate_limited). Inside the window the video is skipped without hitting the
# caption endpoint; after it expires the video is retried. 0 = always retry.
# Go duration; default 1h.
TAPIR_FETCH_BACKOFF=
# Recency bound for AUTO summarization: in automatic mode only videos published
# within this window of now are summarized; older ones are discovered + listed
# but wait for a manual "Summarize" (so a back-catalogue doesn't self-inflict
# 429s). An explicit request bypasses it. Go duration; default 168h (~7d).
# 0 = no bound (summarize every unseen video).
TAPIR_AUTO_SUMMARIZE_WINDOW=
# Minimum interval between outbound caption fetches across the WHOLE process —
# the shared per-egress-IP rate gate (ADR-014). Scheduler runners and the web
# "Summarize" click-path serialise through it so they cannot collectively trip
# 429s. Go duration; default 2s. 0 = unlimited (dev/tests).
TAPIR_FETCH_RATE=
# --- scheduled discovery (tapir serve, ADR-018) ---------------------------
# When > 0, `serve` runs in-process discovery for ALL users on this cadence
# (e.g. 2h): one runner pass per user per tick, run-once-on-startup then ticked.
# Empty/0 = disabled (dev/tests never auto-fetch). SINGLE-REPLICA assumption —
# >1 replica double-runs discovery. Go duration.
TAPIR_DISCOVERY_INTERVAL=
# --- invitations (tapir invite) -------------------------------------------
# Public base URL used to build the invite link `tapir invite <email>` prints.
# Default https://tapir.d-ma.be; no trailing slash needed.
TAPIR_PUBLIC_URL=
+18 -6
View File
@@ -46,8 +46,9 @@ These caused real mistakes that were caught and corrected; the corrections are l
(See `DECISIONS.md` for full rationale. Listed here so you don't propose them.)
- **No Supabase** — reuse Dex / ESO+1Password / Postgres (ADR-002).
- **No global cross-tenant video/transcript table** — per-user isolation (data-model). Dedup
across users is a Future C concern, not a Stage 0/1 default.
- **No global cross-tenant *video* table** — videos stay per-user (data-model). Transcripts ARE
shared since ADR-021 (public caption content, keyed by `(provider, provider_video_id)`, non-RLS)
so re-analysis never re-fetches; the *videos* half of cross-tenant dedup stays a Future C concern.
- **No audio-download + speech-to-text in the core path** — captions-first (ADR-007). STT is a
deferred, bounded optional component.
- **No public SaaS / sign-up / billing / Google OAuth verification at scale** — Future C,
@@ -60,8 +61,14 @@ These caused real mistakes that were caught and corrected; the corrections are l
stores, and sinks are adapters. Adding a video provider or a sink = a new adapter implementing
the interface, nothing in the engine changes. This is what keeps "standalone vs homelab" a
wiring choice (ADR-003).
- **BDD.** The `docs/use-cases/*.feature` files are the behavior spec. New behavior gets a
scenario; the use-case core is tested through fake adapters, not live YouTube/brain.
- **BDD.** The `docs/use-cases/*.feature` files are the behavior spec (design records — there is
no godog runner). New behavior gets a scenario; the use-case core is tested through fake
adapters, not live YouTube/brain. A name-coverage gate (`test/acceptance/scenario_coverage_test.go`,
`TestScenarioCoverage`) keeps the two from drifting: every non-`@pending` scenario must be
mapped to an existing Go test in `scenarioCoverage`. When you add a scenario, either map it to
its covering test or tag it `@pending` in the `.feature` with a one-line reason. It checks the
*link*, not that the test exercises the scenario — that's the deliberate trade for not running
godog (see issue #5 / the BDD-runner decision).
## Skills (engineering discipline)
@@ -79,7 +86,7 @@ Skills live in the canonical library `mathias/skills` and are wired into this re
## Current build state (start here for the first task)
The repo is **green and shipping** — last tag `v0.4.0`. `task check` passes (fmt, vet, lint,
The repo is **green and shipping** — last tag `v0.14.0`. `task check` passes (fmt, vet, lint,
`go test -p 1 ./...`). Go is `1.26.1` (see `go.mod`).
- Clean Architecture core is implemented: `internal/domain` (entities), `internal/ports`
@@ -88,13 +95,18 @@ The repo is **green and shipping** — last tag `v0.4.0`. `task check` passes (f
green against it.
- Adapters present under `internal/adapters/`: `youtube` (captions-first `VideoSource`,
timedtext/InnerTube acquisition per ADR-010), `summarizer` + `llm` (the copied AI router,
Primary→Fallback per ADR-004), `store` (Postgres, golang-migrate migrations 001006),
Primary→Fallback per ADR-004), `store` (Postgres, golang-migrate migrations 001015),
`secrets` (file-backed `SecretStore`). The brain HTTP sink (ADR-005) is the remaining
optional sink.
- Stage 1 is open (ADR-012): multi-user with **DB-enforced** isolation — Postgres RLS `FORCE`d
on all user-owned tables (migration 003), two-user isolation test in
`internal/adapters/store/rls_test.go`. Registration gate, per-user YouTube web connect, and
account management (disconnect / delete, ADR-013) all shipped.
- Transcript persistence (ADR-021, migration 015): transcripts are a **shared, non-RLS** store
keyed by `(provider, provider_video_id)` — the single exception to the isolation boundary
(`TestTranscriptsTableIsSharedNotRLS`). The engine reads stored transcripts before any caption
fetch (`usecase.resolveTranscript`), so re-analysis — re-summarize, paste-a-URL, onboarding
burst — never re-touches YouTube. Per-user summaries/videos stay RLS-scoped.
- `cmd/tapir` subcommands: `list`, `show`, `auth` (interactive host-side OAuth), `run` (batch
watch→summarize), `serve` (the HTMX+Templ web reader/writer under `internal/web`, a new
transport over the unchanged engine/ports — ADR-003). `tapir env` prints config.
+428 -8
View File
@@ -29,7 +29,9 @@ Go's concurrency model fits the watcher/worker shape).
## ADR-002 — No Supabase; reuse existing Dex / ESO / Postgres conventions
**Status:** Accepted (2026-06-02)
**Status:** Accepted (2026-06-02). The "Tapir does not write to the shared identity provider"
implication is **partially superseded by ADR-017** (invite flow writes Dex Password CRs); the
no-Supabase / reuse-existing-primitives decision stands.
**Context.** The draft proposed self-hosted Supabase for auth + RLS + secrets (Vault). The
homelab already runs **Dex** (OIDC), **ESO + 1Password** (secrets), and a **postgres18**
@@ -292,7 +294,11 @@ explicit call, with isolation as the guardrail that keeps it safe.
## ADR-013 — Account deletion is Tapir-side only; the Dex identity is left intact
**Status:** Accepted (2026-06-03)
**Status:** Accepted (2026-06-03). The "Tapir never holds write access to the shared identity
provider" rationale is **partially superseded by ADR-017** (the invite flow now *creates* Dex
Password CRs). The deletion-behaviour decision itself — delete Tapir-side state, leave the Dex
identity intact — still stands; ADR-017 only changes the create side, not delete. See ADR-017
for the now-asymmetric posture (Tapir can create Dex accounts but still does not delete them).
**Context.** Stage 1 (ADR-012) added account deletion. A registered user is two things: a
`users` row (plus all their data, cascade-linked) in Tapir's Postgres, and a subject identity
@@ -308,7 +314,9 @@ deprovision the Dex identity. The maintainer chose (a).
- All of that user's secrets in the SecretStore (the per-user YouTube refresh-token refs).
The **Dex identity is deliberately left intact.** Tapir does not deprovision, disable, or
modify the shared Dex directory.
modify the shared Dex directory **on deletion**. (Note per ADR-017: Tapir *does* now create Dex
Password CRs on invite — so the create and delete sides are deliberately asymmetric, and a
deleted user's Dex Password CR persists. See ADR-017 consequences.)
**Consequences.**
- **Clean re-registration:** a deleted user who logs in again arrives as a Dex-authenticated
@@ -319,9 +327,13 @@ modify the shared Dex directory.
the identity carries no Tapir content. **But if Tapir ever moves toward Future C (real
external/public users), this is a GDPR-shaped gap** — a true "delete my account" there must
also deprovision or anonymise the Dex identity, which is a new ADR and likely a Dex-admin
integration Tapir does not currently have.
- **Blast radius stays small:** Tapir never holds write access to the shared identity provider,
consistent with the estate's blast-radius-minimisation posture (ADR-002, architecture review).
integration Tapir does not currently have. **ADR-017 widens this gap:** Tapir now *creates*
the Dex Password CR but does not delete it, so a deleted Tapir user leaves an orphaned Dex
local-password account. Tracked as a known gap in ADR-017.
- **Blast radius:** ADR-013 originally claimed "Tapir never holds write access to the shared
identity provider." **ADR-017 changes this** — Tapir now holds scoped `create`+`get` on
`passwords.dex.coreos.com` in the `auth` namespace. The blast radius is no longer zero; it is
bounded by that RBAC. See ADR-017 for the security analysis.
**Reversibility.** Adding Dex deprovisioning later is a superseding ADR; nothing about the
current choice blocks it. Recorded now because "deletion is partial by design" is a deliberate
@@ -331,7 +343,11 @@ semantic that future-Tapir (and any compliance review) must know was chosen, not
## ADR-014 — Timedtext 429 handling: per-host backoff + honest in-flight UX, before any Whisper reconsideration
**Status:** Accepted (2026-06-03)
**Status:** Accepted (2026-06-03). **Partially implemented as of v0.6.0** — see the
implementation note at the end; the shared per-egress-IP rate gate (decision item 2) may not be
fully realised. Verify against `internal/runner` + the youtube adapter before treating as done.
**ADR-018 (in-process scheduled discovery + auto-summarize) makes decision item 2 load-bearing
and folds its confirm/finish into that build.**
**Context.** ADR-010 acquires captions from the unauthenticated `timedtext` baseUrl. Live runs
show that endpoint **rate-limits per source IP (HTTP 429) under volume** — many videos fetched
@@ -386,6 +402,405 @@ clean data then tells you whether Whisper is warranted.
dedicated egress IP / outbound proxy is worth it; CronJob-driven `tapir run` interaction with
the rate gate (the batch path moves into k3s per the deferred CronJob item).
**Implementation note (2026-06-03, v0.6.0 — added during reconciliation).** A v0.6.0 release
shipped 429 handling: `domain.SourceRateLimited` (429 no longer collapsed to `SourceNone`),
youtube adapter maps 429 → `SourceRateLimited`, **migration 007 added `videos.transcript_status`
+ `videos.rate_limited_at`** (durable per-video retry state), `TAPIR_FETCH_BACKOFF` config,
runner records rate-limited + skips within the backoff window, and a "⏳ Retrying later" badge.
This realises decision items 1, 3, and 4 well. **Item 2 (a single shared per-egress-IP rate
gate) appears to be realised as per-*video* backoff state, NOT a process-wide IP gate** —
multiple videos can still each hit the endpoint and collectively trip the per-IP 429. Treat item
2 as **not yet confirmed done**; verify in `internal/runner`/youtube adapter and, if absent, it
remains open. (Resolved during reconciliation: the v0.6.0 report's reference to "migration 008
videos.rate_limited_at" was a **mislabel** — both columns shipped in migration **007**, not a
missing 008. No migration was lost; the sequence legitimately skips 008. `rate_limited_at`
exists and the runner's column reads are sound.) **ADR-018 makes item 2 a hard requirement:
in-process scheduled discovery + auto-summarize drives all users' fetches through one pod egress,
so the process-wide gate is confirmed/built as part of that slice.**
---
## ADR-015 — Per-user credentials: envelope-encrypted in PG18, not vault-stored
**Status:** Accepted (2026-06-03)
**Context.** The Stage-0/1 SecretStore (`internal/adapters/secrets/file.go`) holds per-user
YouTube OAuth refresh tokens as a flat key-value JSON map on a PVC — explicitly a stand-in for
"op/ESO later" (ADR-002, ADR-006). infra#86 proposed migrating it to an ESO-backed store. The
decision spike (infra#88) found that framing subtly wrong: **ESO syncs vault→cluster at
deploy/refresh time; it is not a runtime write API.** Per-user tokens are written *at runtime,
per end-user* (every YouTube connect; on token rotation) — they are application state, not
configuration. The homelab 1Password SA is also read-only, so a vault-write path would require
a new write-capable SA, widening Tapir's blast radius to shared estate infra to store what is
fundamentally Tapir's own row-data. Reading the actual SecretStore confirmed the shape: a
3-method port (`Get`/`Put`/`Delete`) over opaque refs, written interactively per user.
**Decision.** Per-user credentials are stored **envelope-encrypted in PG18**, not in the vault:
1. Tokens are encrypted with a **single app-level envelope key** and stored as ciphertext in
PG18, under the Row-Level Security already enforced and tested (ADR-012). Reads/writes go
through the existing `withUser` RLS-scoped seam.
2. The **envelope key** is the only secret in 1Password — fetched via the **existing read-only
SA** (confirmed working). No new write-capable SA; no per-user vault items.
3. The `ports.SecretStore` port is unchanged (`Get`/`Put`/`Delete`). The implementation swaps
`FileStore` (PVC JSON) for a `PGStore` (encrypted rows). Every consumer — connect,
disconnect, delete-account — is untouched (the port abstraction holds, ADR-003 spirit).
4. **Infra/operator credentials** (Dex client secret, MCP-auth tokens, service tokens) stay an
**ESO/1Password** concern. This ADR governs *per-user runtime* credentials only. The two
classes use two mechanisms deliberately — because they are two different things (runtime
app-state vs deploy-time config), not as a compromise. The "one mechanism" question
(maintainer's initial preference) was answered in #88 by correctly *classifying* the
secrets rather than unifying their storage.
**Consequences.**
- Runtime credential writes are normal RLS'd DB writes — no ESO sync latency, no indirection,
no write-SA blast radius. The interactive connect→store→use flow works without a vault
round-trip.
- Keeps PG18 and keeps ADR-002 intact (Supabase was considered and rejected again in #88
adding a datastore to hold a few encrypted strings PG18 already holds).
- Adds an encrypt/decrypt seam and an **envelope-key rotation** responsibility (re-encrypt the
per-user rows under a new key). infra#89 (build) must implement and test rotation, not assume
it — this is the real engineering cost of the choice.
- The vault's involvement shrinks to one static key via the SA already trusted for reads.
- **Supersedes** the "PVC stand-in for op/ESO" intent recorded in `secrets/file.go` and
`docs/homelab-integration.md` for the *per-user* secret path (the ESO/1Password reference in
ADR-006 stands for the *infra-cred* path).
**Reversibility / falsification (from infra#88).** Revisit if: per-user tokens need
high-frequency rotation writes (weak — PG18 handles it); an estate compliance policy requires
all credentials in 1P for a single audit surface (maintainer-knowable, not currently believed
to hold — would favour the vault-write path on policy grounds); or envelope-key rotation proves
operationally worse than per-secret vault rotation (the real cost #89 must prove). If none hold,
this stands. Full reasoning + rejected candidates (write-capable SA; Supabase): the infra#88
decision doc (`infra/docs/superpowers/handoffs/`). Build + reboot-validation: infra#89.
---
## ADR-016 — Stage 0 gate revised: "useful to me OR a friend", behavioural not feedback
**Status:** Accepted (2026-06-03). Revises the Stage 0 definition in VISION.md (supersedes the
original "useful to me, specifically" gate and folds in the old Stage 1 "a trusted user returns"
test). **The clock-start is further revised by ADR-018** (starts when scheduled discovery ships,
since unprompted use was not possible before then).
**Context.** The original Stage 0 gate was "the maintainer reads summaries weekly for four weeks
and acts on one." The maintainer chose to change it to include friendly users, reasoning that
early signal from friendly users is valuable. Two sub-decisions shaped the final form:
- *Me OR a friend* (not AND): either the maintainer or an onboarded friend showing use clears it.
- *Behavioural, not feedback*: the test is **return usage**, not stated approval.
**Decision.** Stage 0 passes when, over a 34 week window, **either the maintainer or at least
one onboarded friend returns to Tapir unprompted and reads/acts on summaries in ≥2 separate
weeks.** Friend feedback is gathered and valued but is **not** the gate.
**Why behavioural, not feedback (the load-bearing part).** Asked-for feedback from friendly
users is the least reliable signal in product development — politeness bias means a friend you
onboarded will tend to say encouraging things regardless of real value. The thing actually worth
knowing is whether they *come back on their own*. So the gate measures returns, not nice words.
This deliberately resists the most common way a principled gate dies: being declared "passed" on
the strength of a polite reaction.
**Honest note on what this change does.** This is a *guardrail edit made while the original gate
was unmet* (Stage 0 had barely started; build had run well ahead of use-evidence). That is
precisely the pattern that warrants scrutiny — redrawing a gate around work already done. It was
examined on that basis and proceeds because: (a) the new gate is **not softer in kind** — it
stays behavioural and sustained, merely broadening *who* can supply the signal; (b) friendly-user
signal is genuinely valuable; (c) the politeness-bias guard keeps it from collapsing into
"someone said it's nice." It is *not* a licence to treat the already-shipped Stage-1 machinery as
evidence the gate passed — use-evidence remains open.
**Consequences.**
- VISION.md Stage 0 rewritten; old Stage 1 ("a trusted user returns") folded in (it was
near-identical to the new test); hardening renumbered to Stage 1.
- New drift signal added: declaring the gate passed on polite feedback rather than return-usage.
- The 2026-07-01 check-in now asks "is anyone (me or a friend) coming back unprompted?", not
"am I using it weekly?". (Check-in date itself revised by ADR-018.)
**Reversibility.** A superseding ADR could tighten it back to maintainer-only or raise it to
require multiple returning users. Recorded with the full rationale (including the self-scrutiny
about editing a gate while it's unmet) so the reasoning survives, not just the new wording.
---
## ADR-017 — Invite flow: Tapir creates Dex local-password accounts (write access to the shared identity provider)
**Status:** ~~Accepted~~ **SUPERSEDED by [ADR-019](#adr-019--authentik-owns-invites-tapir-stops-provisioning-accounts) (2026-06-07).** The Dex local-password
invite provisioning was removed when the homelab IdP migrated Dex→Authentik
(infra ADR-0001); Authentik now owns invites. Original record below.
Accepted (2026-06-03), **recorded retroactively during reconciliation, then
deliberately ratified KEEP (2026-06-03).** This capability **shipped in v0.7.0 without an ADR**
code, RBAC, and a deployed ServiceAccount landed before any decision record existed. This ADR
documents what shipped and honestly records that the decision-before-code discipline was not
followed here (see "Process note"). **Partially supersedes ADR-002 and ADR-013** (the "Tapir
holds no write access to the shared identity provider" posture).
**Keep-or-reverse, decided (2026-06-03).** After the retroactive recording, the maintainer
weighed keep vs. reverse (drop the RBAC + invite flow, rely on Google OIDC only). **Decision:
KEEP.** Deciding fact: not all intended Future-B users have / will use Google accounts, so Google
OIDC alone *cannot* onboard them — the invite flow is therefore **load-bearing, not a redundant
convenience**, and reversing it would leave some intended users with no onboarding path. The
trust-surface cost (scoped `create` on `passwords` in `auth`) is accepted deliberately in
exchange. The opposing argument (this contradicts ADR-015's "no write credential to shared infra"
logic, applied there to secrets) was considered and outweighed *only* because the capability is
genuinely necessary for real users here — it is **not** a precedent for widening the surface
further. **Explicitly NOT chosen:** adding `delete` to the RBAC to close the orphan gap (that
widens the surface in the wrong direction; accept the orphan at Future B instead — see open
items).
**Context.** Stage 1 onboarding needs a way for a friend to get a login. Two paths shipped:
Google OIDC via Dex (no Tapir code — Dex handles it), and an **invite flow** where Tapir itself
provisions a Dex **local-password** account. The invite flow is what this ADR is about.
**What shipped (reconstructed from `internal/adapters/dex/dex.go`, `internal/web/invite.go`,
migration `009_invitations`).**
1. `tapir invite <email>` (host CLI) writes an `invitations` row: a 32-byte crypto-random,
single-use, 7-day-expiry `token` (the token is the capability), `email`, `used_at`.
The `invitations` table is deliberately **not** RLS/`user_id`-scoped — the invitee has no
user yet and the token itself is the secret. (Sound; documented in the migration.)
2. Recipient visits `/invite/{token}` (public, no session), sets a password (validated *before*
the token is consumed, so a typo doesn't burn it), the token is claimed exactly once.
3. Tapir bcrypt-hashes (cost 12) and **POSTs a `passwords.dex.coreos.com` Custom Resource into
the `auth` namespace** via the in-cluster Kubernetes API, authenticating as the `tapir`
ServiceAccount. TLS validated against the mounted cluster CA. Off-cluster (dev) it returns
`ErrNotInCluster` and degrades without consuming the invite.
4. RBAC applied this deploy: the `tapir` ServiceAccount has **`create`+`get` on
`passwords.dex.coreos.com` in the `auth` namespace** (and only that).
5. Dex (kubernetes storage) then serves local-password login for that email; Tapir's
registration gate creates the profile on first login.
**Decision (as ratified now).** Accept the invite flow as built: Tapir *may* hold scoped write
access to Dex's `passwords` resource in the `auth` namespace, for the purpose of invite-based
local-account creation. This is a deliberate, bounded reversal of the prior "no write access to
the shared identity provider" posture (ADR-002/013).
**Why this is acceptable (the case for keeping it).**
- The RBAC is **minimally scoped**: `create`+`get` on one resource type in one namespace, not
broad Dex/cluster write. Blast radius is bounded and inspectable.
- It enables friend onboarding **without Google OAuth app verification** (ADR-008 deferred that),
which is genuinely useful for Future B — and necessary for users who won't use Google (the
deciding fact in the keep decision above).
- The code is careful: validate-before-consume, single-use tokens, graceful off-cluster degrade,
sentinel errors mapped to clear messages, the base64-bcrypt storage gotcha handled.
**Consequences / known gaps (the case to watch).**
- **Reverses a load-bearing principle.** ADR-002 and ADR-013 leaned on "Tapir never writes to the
shared identity provider" for blast-radius minimisation. That is no longer true. Anyone reading
those ADRs without this one would be misled — hence the cross-references added to both.
- **Create/delete asymmetry → orphaned Dex accounts.** Tapir now *creates* Dex Password CRs but
(per ADR-013) does not *delete* them on account deletion. A deleted Tapir user leaves an
orphaned Dex local-password account that can still authenticate (though it would hit the
registration gate with no profile). This widens the ADR-013 right-to-erasure gap. **Open —
accepted for Future B, revisit before Future C (do NOT add `delete` RBAC just to fix this).**
- **`create` on a shared-namespace identity resource** is a meaningfully larger trust surface than
the rest of Tapir. If the `tapir` pod is compromised, the attacker can mint Dex local accounts.
Bounded by the namespace/resource scope, but real — worth a deliberate look before Future C.
- Token custody for invites is in PG (`invitations.token`) in plaintext; single-use + 7-day TTL
bound the exposure, but a DB read yields live invite tokens until claimed/expired.
**Process note (why this ADR is retroactive).** This capability was built and deployed by
parallel agent sessions while the main planning thread was elsewhere, and shipped with no ADR —
the first time in this project a guardrail-reversing change skipped the decision-before-code
discipline. Recorded here not to rubber-stamp it but to restore the audit trail: the decision is
now visible, its principle-reversal is named, and its open gaps (orphaned accounts, the larger
trust surface) are tracked rather than buried in a release note. The maintainer ratified KEEP
after the fact (see status block) on the deciding fact that some intended users cannot use Google
OIDC.
**Open items this ADR creates (tracked in infra):**
- Confirm the RBAC really is `create`+`get` only (not broader) against the deployed manifest in
`infra` — a security claim currently resting on a sibling report, not a verified manifest read.
- Accept the orphaned-Dex-account gap for Future B; revisit (orphan cleanup + the whole Dex-write
surface) before any Future C move. Do **not** add `delete` RBAC solely to fix the orphan.
- Review the larger trust surface before any Future C move.
---
## ADR-018 — Make Tapir usable unprompted: in-process scheduled discovery, auto-summarize default, gate-clock reset
**Status:** Accepted (2026-06-03)
**Context.** The Stage-0 gate (ADR-016) measures whether the maintainer or a friend *returns and
reads/acts* over weeks. But the system could not actually be used that way: discovery (`tapir
run`) was **host-side manual**, so a newly onboarded user saw an empty list and had no reason to
return. The gate was structurally unmeetable — not because the product failed, but because the
*workflow* depended on the maintainer SSHing in to trigger each pass. The runner already has a
`Loop(ctx, interval)` (run-on-start, then every tick) and per-user mode/dedup/backoff; what was
missing was *invocation* — nothing called it for the deployed users on a schedule.
**Decision.**
1. **In-process scheduled discovery.** `tapir serve` launches a background goroutine that runs
discovery for **all users** on an interval (`TAPIR_DISCOVERY_INTERVAL`; unset/0 = off). It
enumerates users (un-RLS'd `user_identities`) and runs each user's pass inside `withUser`,
reusing the existing `runner` — not a new scheduler. Stateless timing; ctx-cancellable;
per-user failures isolated.
2. **Auto-summarize default ON** for Future-B users (new registrations default `true`; existing
rows updated), so discovery both populates *and* summarizes — the list fills itself. The
per-user manual toggle remains.
3. **Gate-clock reset.** The Stage-0 34 week window (ADR-016) **starts when this ships**, because
unprompted use was impossible before it. This is starting the clock when the experiment can
actually run, **not** a reset to dodge a failing gate (the prior window measured nothing —
there was no way to use the system unprompted). The 2026-07-01 check-in moves accordingly to
~34 weeks after this deploys.
**Why in-process and not a k8s CronJob.** Maintainer's call: simpler deploy (no second deployable),
acceptable at Future-B scale (13 users). The CronJob's advantage — failure isolation between
discovery and the web/read path — was weighed and traded away knowingly.
**Consequences / constraints.**
- **Discovery shares the web process's lifetime and egress.** A wedged discovery pass can degrade
the reading UI (the coupling a CronJob would have avoided). Accepted at this scale.
- **SINGLE-REPLICA ASSUMPTION (load-bearing).** If `tapir serve` ever runs >1 replica, *every*
replica runs the discovery loop → every user fetched in parallel (429s + duplicate work). Tapir
must stay single-replica while in-process scheduling is enabled, or this is revisited (move to
CronJob, or add leader-election). Recorded so a future scale-up doesn't silently double-run.
- **Makes ADR-014 item 2 (shared per-egress-IP rate gate) load-bearing.** In-process + auto +
multi-user drives all caption fetches through one pod egress, concurrent with click-path
Summarize. The build confirms/finishes the process-wide rate gate in the same slice; without it
the system self-inflicts 429s every cycle. **Auto-summarize ON is gated on the rate gate
existing** (fallback: ship discovery with auto OFF until it does).
**Reversibility.** Disable via `TAPIR_DISCOVERY_INTERVAL=0` (reverts to manual `tapir run`).
Moving to a CronJob later is a superseding ADR; the per-user `runner` is unchanged either way.
Spec: `docs/specs/scheduled-discovery.md`.
---
## ADR-019 — Authentik owns invites; Tapir stops provisioning accounts
**Status:** Accepted (2026-06-07). **Supersedes ADR-017** (Dex local-password invite
provisioning).
**Context:** infra ADR-0001 migrated the homelab IdP Dex→Authentik. Authentik provides
first-class invite flows; the Dex local-password connector never consulted the Password
CRs Tapir wrote (the defect that triggered the migration). Tapir-web's OIDC issuer now
points at Authentik (infra ADR-0001 step 3).
**Decision:** Tapir no longer provisions accounts. The Dex-password invite path is removed:
`internal/adapters/dex`, the public `/invite/{token}` set-password UI (`internal/web/invite.go`),
the `tapir invite` CLI (`cmd/tapir/invite.go`), the `InvitationStore`/`DexPasswordCreator`
ports + `App.Invitations`/`App.Dex` wiring, the invite Templ pages, and the
`tapir invite` Taskfile target. New users are invited via Authentik's flow, log into Tapir
via OIDC, and are captured by Tapir's existing provider-agnostic `/register` (display name).
Login + Google moved by config only (Authentik per-app issuer); the OIDC adapter is unchanged.
**Consequences:** smaller Tapir blast surface — no writes to the shared identity provider, no
configmap/CR access, the dedicated `passwords.dex.coreos.com` RBAC + ServiceAccount are
removed (infra side, coupled change). The `invitations` table (migration 009) is left in
place — migrations are append-only and the unused table is harmless; a future migration may
drop it. The `*_DEX_*` config/identity names (`TAPIR_DEX_CLIENT_*`, `dex_subject`, the
`DexAuth`/`oidc` package) are now misnomers; renaming is deferred (cosmetic, not behavioural).
**Rejected:** keeping Tapir's `/invite` UI but calling Authentik's API on claim — couples
Tapir to Authentik's admin API + a token for no real gain; Authentik's own invite flow is
the supported path.
---
## ADR-020 — Recency-bounded auto-summarize + honest sparse-state surface
**Status:** Accepted (2026-06-08). **Refines ADR-018** (auto-summarize) and **ADR-014**
(per-IP caption rate gate).
**Context.** Tapir is operational but sparse: at real subscription volume the maintainer's
account holds ~283 discovered videos, ~15 summarized, ~256 behind the respected per-IP caption
rate gate, ~12 no-captions. Two problems follow. (1) **Load:** ADR-018 auto-summarizes *every*
unseen video, so a large back-catalogue re-drives the whole queue through the gate every cycle —
self-inflicted 429s with no user value (nobody is waiting on a 6-month-old video). (2) **First
contact:** a new user sees a mostly-empty feed with no moving parts and a UI that implied
abundance/imminence ("fetching soon" ×256, "Run `tapir run`", "Summarize now"); the Stage-0 gate
is *return usage*, and the experience died at the first visit. Source: a UX heuristic review
(`docs/ux-review/UX-REVIEW-stage0-recency.md`).
**Decision.**
1. **Recency bound on auto-summarize.** In automatic mode the scheduler only summarizes videos
published within `TAPIR_AUTO_SUMMARIZE_WINDOW` (default ~7d). Older videos are still discovered
and listed but not auto-processed — they keep the manual "Summarize" affordance. An explicit
manual request bypasses the bound. `0` disables it (pre-recency behaviour). This bounds auto
*load*; it does not fetch harder — the gate (ADR-014) is untouched and the manual path still
serialises through it.
2. **Honest sparse-state surface.** Copy is reframed to surface scarcity truthfully, never to
look fuller: "N ready · M in queue · K no captions" (not "fetching soon"); a one-line "captions
are fetched slowly on purpose" note; "Summarize" (not "Summarize now"); the empty-connected
state stops printing an impossible CLI command.
3. **Feed IA = one list, noise-collapsed.** Summarized + recent un-summarized cards lead inline;
the older un-summarized back-catalogue collapses behind a single "Show N older videos"
disclosure; caption-less videos collapse to a one-line count instead of N dead cards. List is
ordered summarized-first, then `published_at DESC NULLS LAST`.
**Consequences.** The auto path's per-cycle fetch volume is bounded by recent uploads, not the
whole back-catalogue, so steady-state 429 pressure drops sharply. Older videos become explicitly
on-demand — a deliberate honesty trade (the user chooses to spend a scarce fetch on old content).
The single-replica assumption (ADR-018) is unchanged.
**Deliberately NOT done (premature until the Stage-0 loop is validated).** Return-nudges
(digest email / push) — a nudge contaminates the *unprompted*-return signal the gate measures
(ADR-016); building it now poisons the experiment. Also deferred: full-text search, channel
facets, read/unread, saved views — all need summary abundance to matter.
**Reversibility.** `TAPIR_AUTO_SUMMARIZE_WINDOW=0` restores summarize-every-unseen; the feed
collapse keys off the same window (`App.RecencyWindow=0` → everything inline).
---
## ADR-021 — Persist transcripts as shared, video-keyed public content (re-analysis never re-fetches)
**Status:** Accepted (2026-06-09). **Reopens the transcripts half of** the "Global cross-tenant
`videos`/`transcripts` table" rejection (data-model.md). **Builds on ADR-007** (captions-first),
**ADR-010/ADR-014** (the per-IP caption rate gate), and **ADR-012** (per-user RLS isolation).
**Context.** Every summarization fetches the transcript fresh through the caption path, even when
the exact same transcript was fetched moments ago — for the same user re-summarizing, or for a
second user who happens to watch the same video. The caption fetch is the one genuinely scarce,
genuinely risky operation in the system: YouTube's timedtext endpoint is unofficial and per-IP
rate-limited (ADR-010), and tripping it risks the maintainer's Google standing (ADR-014). So the
operation we most want to *avoid repeating* is the one we currently repeat unconditionally. A
transcript is **public content** — the same words YouTube serves to anyone — and carries nothing
user-identifying. The per-user isolation that protects summaries, feeds, and tokens (ADR-012) is
the wrong shape for it: it forces a re-fetch per user for data that is identical across users.
The original rejection ("Global cross-tenant `videos`/`transcripts` table") bundled videos and
transcripts together and rejected both on the grounds that "at 15 users, re-summarizing is
cheaper than the coupling." That reasoning holds for **videos** (per-user feed rows, genuinely
user-scoped) but not for **transcripts**: the cost being avoided is not LLM re-summarization, it
is a *rate-gated, reputation-risky network fetch*, and that cost is paid per re-fetch regardless
of user count. One re-fetch avoided is strictly worth more than the coupling it removes.
**Decision.**
1. **A single shared `transcripts` table, keyed by the cross-user dedup key
`(provider, provider_video_id)`** — the stable public identity of the video, not Tapir's
internal per-user `videos.id`. Columns: the key, `source` (`captions`/`none`), `language`,
`content`, `fetched_at`. It holds **only public caption content + the video's public id**
nothing user-identifying — and is therefore **NOT RLS-scoped**: no `user_id`, no policy, no
`FORCE ROW LEVEL SECURITY`. This is the deliberate, single exception to the ADR-012 isolation
boundary, and the only one.
2. **Summarize path becomes read-stored-first.** Have a stored transcript for this video? →
summarize from the stored text, **no caption fetch**. No stored transcript? → fetch *through
the unchanged gate* (ADR-014) → store it → summarize. The gate is neither bypassed nor
weakened; persistence reduces how *often* we reach it, never how *fast*.
3. **De-facto cross-user dedup is the intended behaviour, not a feature with a switch.** Two
users who share a video share the one transcript row. A permanent `source = 'none'` (no
captions) is stored too, so a known-caption-less video is not re-fetched by anyone. A
transient 429 (`SourceRateLimited`) is **never** stored as terminal — it stays a per-user
retry via the existing `transcript_status` backoff (ADR-014), so persistence cannot mask a
rate-limit into a false "no transcript."
4. **Per-user `summaries` stay RLS-scoped (ADR-012 unchanged)** and reference the transcript by
video id. Videos stay per-user. Only transcripts go shared.
**Consequences.** Re-analysis (re-summarize, different model, paste of an already-seen video,
onboarding of a second user with overlapping subscriptions) never re-touches YouTube — the
primary win, and it *reduces* aggregate caption-gate pressure, reinforcing ADR-010/ADR-014 rather
than straining them. The isolation surface gains exactly one non-RLS table; an isolation test
asserts the boundary is *exactly* there and has not leaked to any user-owned table (this is the
proof the public-content classification was implemented as designed). It also unblocks
multi-model / customizable analysis (re-run analysis on stored text for free) — enabling that is
this ADR's point; building it is separate.
**Reversibility.** The read-stored-first check is the only behavioural coupling; removing it
restores fetch-every-time. The down-migration recreates the per-user RLS-scoped transcripts shape
(001/003). No user-facing surface depends on cross-user sharing — sharing is the *storage shape*,
never exposed in the UI.
---
## Rejected alternatives
@@ -402,10 +817,15 @@ maps to the ADR that settles it.
| Lifting shared packages into a `brain-common` module | Couples Tapir's release cycle to the monolith for negligible code savings | ADR-004 |
| Importing/replicating the filesystem `brain` package | Assumes co-location with the brain git checkout; wrong for a standalone networked service | ADR-005 |
| Reusing `ingestion`'s `oauth` package for YouTube/Vimeo | Same name, opposite direction — it's inbound MCP-server auth, not outbound provider OAuth | ADR-006 |
| Global cross-tenant `videos`/`transcripts` table (dedup) | Reintroduces the cross-domain DB coupling the homelab review is removing; at 15 users, re-summarizing is cheaper than the coupling | data-model.md |
| Global cross-tenant `videos`/`transcripts` table (dedup) | Reintroduces the cross-domain DB coupling the homelab review is removing; at 15 users, re-summarizing is cheaper than the coupling. **Transcripts half reopened by ADR-021** — the avoided cost there is a rate-gated, reputation-risky *caption fetch*, not LLM re-summarization, so it outweighs the coupling; **videos stay per-user.** | data-model.md, **ADR-021** (transcripts only) |
| Audio-download + Whisper STT in the core path | ToS-grey, breakage-prone (yt-dlp), contends for koala GPU with the JEPA PoC; captions alone test the core hypothesis | ADR-007 |
| Building multi-tenant SaaS / Google OAuth verification now | "Real users soon" was lowered to Future B; SaaS machinery before the Stage 0 self-use gate is the primary documented anti-goal | ADR-008, VISION |
| Delegating the S5 reuse spike to an agent swarm | A 1-hour sequential read-and-judge with a single coupled conclusion; orchestration overhead exceeds the work, and it's Diamond-1 judgment the maintainer wanted to own | (process note) |
| Vault-write SA for per-user OAuth tokens (ESO as runtime write path) | ESO syncs vault→cluster at deploy time, not a runtime write API; a write-SA widens blast radius to shared infra to store app row-data | ADR-015, infra#88 |
| Supabase for per-user credential storage | Adds a second datastore for a few encrypted strings PG18 already holds; reopens ADR-002 | ADR-015, infra#88 |
| Feedback-based Stage 0 gate (friends saying it's useful) | Politeness bias makes asked-for feedback the least reliable signal; return-usage is the real test | ADR-016 |
| Reverse the Dex-write invite flow (Google OIDC only) | Some intended Future-B users won't use Google; OIDC-only leaves them with no onboarding path — invite flow is load-bearing | ADR-017 |
| k8s CronJob for scheduled discovery (vs in-process) | At Future-B scale the in-process scheduler is simpler to deploy; CronJob's failure-isolation benefit was weighed and traded away knowingly (revisit if >1 replica or load grows) | ADR-018 |
If a future case genuinely reopens one of these, that's a new ADR superseding the relevant one —
not a silent reversal.
+14 -2
View File
@@ -37,7 +37,7 @@ model, behavior specs) remain the source of intent.
## Running the Stage-0 demo
Tapir runs on **your own** YouTube account: authorize once, then run the
Tapir runs on your YouTube account(s): authorize once, then run the
watch→summarize→deliver loop. All configuration is via `TAPIR_*` environment
variables — copy [`.env.example`](.env.example) to `.env` and fill it in (no
secrets are committed; at demo time source them from op, e.g. `op run -- ...`).
@@ -56,7 +56,7 @@ go build -o bin/tapir ./cmd/tapir
./bin/tapir auth
# 3. run: detect new videos across your subscriptions, summarize, deliver to the
# store. Unset TAPIR_POLL_INTERVAL = single pass; set it (e.g. 15m) to loop.
# store. Single pass; set TAPIR_DISCOVERY_INTERVAL (e.g. 2h) for the serve loop.
./bin/tapir run
```
@@ -68,6 +68,18 @@ URI matches `TAPIR_OAUTH_REDIRECT_ADDR`. The summarizer model
(`TAPIR_SUMMARIZER_MODEL`, default `koala/phi4-mini`) is overridable; pick the
final alias when the gateway is reachable (see `docs/homelab-integration.md`).
### Web surface (`tapir serve`)
`tapir serve` starts the HTMX+Templ web UI on `:8080`. Users log in via Dex OIDC (local
password or Google); a new Dex subject is routed to `/register` to create a Tapir account.
Stage 1 is multi-user: each user connects their own YouTube account from the browser and
manages their own summaries under DB-enforced RLS isolation. When `TAPIR_DISCOVERY_INTERVAL`
is set (e.g. `2h`), the serve process runs a scheduled discovery pass for every registered
user automatically — no CronJob required. In auto mode only videos published within
`TAPIR_AUTO_SUMMARIZE_WINDOW` (default ~7d, ADR-020) are summarised automatically; older videos
are listed and summarised on demand, so a large back-catalogue doesn't keep re-driving the
caption rate gate. See `docs/homelab-integration.md` for the full config reference.
### Headless on koala
koala has no browser and no interactive `op` session, so the two interactive
+55 -27
View File
@@ -44,50 +44,74 @@ fallback — their key, their choice.
## Who it is for
- **Now (the first customer):** the maintainer — one person, their own subscriptions,
summaries delivered to their own store and brain.
- **Soon (Future B):** a small number of known, trusted users (friends / beta) — each with
their own account, isolated data, optional BYO-AI.
- **Now (the first customers):** the maintainer and a small number of known, trusted
friends — each with their own account, isolated data, optional BYO-AI. The maintainer is
the first customer; friendly users provide the earliest real-world signal.
- **Maybe (Future C, explicitly not built yet):** a public multi-tenant service. Deferred
until there is evidence of sustained personal use **and** real demand. Building for C
before that evidence is a known anti-goal.
until there is evidence of sustained use **and** real demand. Building for C before that
evidence is a known anti-goal.
## Definition of Success
Success is staged. Each stage has a single, falsifiable headline test. We do not advance
to the next stage's ambition until the current stage's test passes.
### Stage 0 — Useful to me (the gate)
### Stage 0 — Useful to me or a friend (the gate)
> **Headline test:** For four consecutive weeks, the maintainer reads Tapir-produced
> summaries for their own subscriptions at least weekly, and at least once acts on a
> summary (watches / skips / saves a video *because of* the summary).
> **Headline test:** Over a 34 week window, *either* the maintainer *or* at least one
> onboarded friend returns to Tapir and reads/acts on summaries in **≥2 separate weeks**.
> The test is *return usage* (behavioural), not stated approval. The ideal signal is an
> **unprompted** return (organic, not because the maintainer nudged them) — but see the
> measurement note below: we currently cannot distinguish prompted from organic returns, so
> in practice we count all returns and read the result with that caveat.
- Captions-first summarization works end-to-end for the maintainer's real subscriptions.
- Summaries land in the maintainer's store and (optionally) brain.
- Captions-first summarization works end-to-end for real subscriptions (the maintainer's
and onboarded friends').
- Summaries land in each user's own store and (optionally) brain.
- Local-first AI produces summaries of acceptable quality without manual intervention
most of the time.
- **This is the gate.** Multi-user, BYO-AI-for-others, and any SaaS ambition stay deferred
until Stage 0 holds. (Ties to the 2026-07-01 self-use check-in.)
- **Why behavioural, not feedback.** Friend *feedback* is gathered and genuinely valuable —
but it is **not** the gate. Asked-for feedback from friendly users is the least reliable
signal in product development (politeness bias); whether they *come back* is the thing we
actually care about. So the gate measures returns, not nice words.
- **Measurement note — "unprompted" is an ideal we can't yet measure.** Whether a return was
organic or prompted by a nudge is not captured by any data Tapir holds (it's context only
the maintainer has). Rather than waive the standard, we name the gap: *unprompted* return
is the signal we genuinely want; *returns* (prompted or not) is what the data can show. A
return that needed a nudge is a weaker signal than one that didn't, and the result is read
with that in mind. If distinguishing them ever matters enough, the maintainer tracks nudges
manually or a future build records prompt events — neither is in scope now.
- **Why "me OR a friend".** This replaces the original "useful to *me*, specifically" gate
(2026-06-03 decision, recorded in DECISIONS.md ADR-016). Getting signal from friendly
users is valuable enough to count — but the bar stays behavioural so it can't be cleared
by a polite reaction. (Ties to the 2026-07-01 check-in.)
- **How it's measured.** Return usage is read from two sources: `summary_actions` (timestamped
watch/skip/save per user) answers "acted in ≥2 distinct weeks"; an append-only login-events
table (see infra/Tapir build) answers "returned/read in ≥2 distinct weeks" even without an
action click — the honest signal for a *reading* product. Login events accrue only from their
deploy date onward, so the gate window's data begins then.
- **Gate-clock reset (ADR-018).** The 34 week window starts when in-process scheduled discovery
+ auto-summarize ship — before that, unprompted use was impossible, so the prior window
measured nothing (this is starting the clock when the experiment can actually run, not a reset
to dodge a failing gate). The 2026-07-01 check-in referenced above moves accordingly to ~34
weeks after this deploys. See DECISIONS.md ADR-018.
- **This is the gate.** Hardening (Stage 1) and any SaaS ambition stay deferred until this
behavioural signal exists. Note: multi-user machinery was deliberately built *ahead* of
this gate (ADR-012) with isolation enforced — that was an explicit, recorded call, not a
sign the gate had passed. The gate is about *evidence of use*, which is still open.
### Stage 1 — Useful to a few (Future B)
> **Headline test:** At least one trusted user other than the maintainer connects their
> own account and, within their first month, keeps using it (returns to read summaries in
> ≥2 separate weeks) without the maintainer hand-holding each summary.
- Multiple users, each with isolated accounts, credentials, and summaries.
- A new user can self-connect a YouTube/Vimeo account and get summaries with no code change.
- Optional BYO-AI works per-user.
- No cross-user data leakage — demonstrable, not assumed.
### Stage 2 — Trustworthy at rest (hardening, still Future B)
### Stage 1 — Trustworthy at rest (hardening, Future B)
> **Headline test:** Credentials (OAuth tokens, BYO-AI keys) are encrypted at rest via the
> homelab's existing secrets convention; a documented, rehearsed recovery path exists; and
> a deliberate isolation test (user A cannot read user B's data) passes in CI or a
> documented manual drill.
- Per-user data isolation is enforced and tested (delivered early via ADR-012 RLS).
- Per-user credentials are encrypted at rest (ADR-015 envelope encryption; build in infra#89).
- A new user can self-connect a YouTube/Vimeo account and get summaries with no code change.
- Optional BYO-AI works per-user.
### Non-goals (current)
- Public sign-up / billing / a marketing surface.
@@ -98,7 +122,11 @@ to the next stage's ambition until the current stage's test passes.
## How we will know we are drifting
- We are building Stage 1+ machinery before the Stage 0 gate has passed.
- We declare the Stage 0 gate "passed" on the strength of polite feedback rather than
behavioural return-usage (the politeness-bias trap the gate is designed to resist).
- We build Stage 1 hardening or Future C machinery while the Stage 0 use-evidence is still
absent. (Multi-user machinery already shipped ahead of the gate via ADR-012 — a recorded,
deliberate exception, not a precedent for more.)
- A user's content reaches a third-party model without that user's explicit, per-user opt-in.
- "Brain ingestion" starts dictating the architecture instead of being one sink behind an
interface.
+50
View File
@@ -0,0 +1,50 @@
package main
import (
"context"
"log/slog"
"sync"
"gitea.d-ma.be/mathias/tapir/internal/runner"
)
// discoveryRunner runs one user's discovery pass.
type discoveryRunner func(ctx context.Context, userID string) (runner.Stats, error)
// serialize wraps run so calls never overlap: every discovery pass — scheduled
// or connect-triggered (#6) — acquires the same lock, preserving the
// one-fetcher-at-a-time invariant the scheduler relies on (ADR-018, the
// single-replica assumption). Locking is per-user, so a connect-triggered pass
// interleaves between the scheduler's users instead of waiting for a whole pass.
func serialize(mu *sync.Mutex, run discoveryRunner) discoveryRunner {
return func(ctx context.Context, userID string) (runner.Stats, error) {
mu.Lock()
defer mu.Unlock()
return run(ctx, userID)
}
}
// discoveryTrigger fires an out-of-band discovery pass for one user without
// blocking the caller (the connect HTTP handler). The pass runs on the server's
// long-lived ctx — not the request ctx — so it survives the post-connect
// redirect. run is the serialized runner, so a trigger never overlaps the
// scheduler. Satisfies web.DiscoveryTrigger.
type discoveryTrigger struct {
ctx context.Context
run discoveryRunner
// onboard, when set, runs after the discovery pass to summarize a capped number
// of the user's newest videos (Feature 1). Optional.
onboard func(ctx context.Context, userID string)
log *slog.Logger
}
func (t *discoveryTrigger) Enqueue(userID string) {
go func() {
if _, err := t.run(t.ctx, userID); err != nil {
t.log.Warn("discovery: connect-triggered pass had errors", "user", userID, "err", err)
}
if t.onboard != nil {
t.onboard(t.ctx, userID)
}
}()
}
+62
View File
@@ -0,0 +1,62 @@
package main
import (
"context"
"fmt"
"sync"
"sync/atomic"
"testing"
"time"
"github.com/stretchr/testify/require"
"gitea.d-ma.be/mathias/tapir/internal/runner"
)
// serialize must guarantee at most one discovery pass runs at a time, so a
// connect-triggered pass never fetches concurrently with the scheduler.
func TestSerializeRunsOneAtATime(t *testing.T) {
var active, maxActive int32
run := func(_ context.Context, _ string) (runner.Stats, error) {
n := atomic.AddInt32(&active, 1)
for { // record the high-water mark of concurrent runs
m := atomic.LoadInt32(&maxActive)
if n <= m || atomic.CompareAndSwapInt32(&maxActive, m, n) {
break
}
}
time.Sleep(2 * time.Millisecond)
atomic.AddInt32(&active, -1)
return runner.Stats{}, nil
}
s := serialize(&sync.Mutex{}, run)
var wg sync.WaitGroup
for i := 0; i < 20; i++ {
wg.Add(1)
go func(i int) { defer wg.Done(); _, _ = s(context.Background(), fmt.Sprintf("u%d", i)) }(i)
}
wg.Wait()
require.Equal(t, int32(1), atomic.LoadInt32(&maxActive),
"serialize must run at most one pass at a time")
}
// Enqueue runs the user's pass out-of-band (non-blocking) on the trigger's ctx.
func TestDiscoveryTriggerEnqueueRunsUser(t *testing.T) {
done := make(chan string, 1)
run := func(_ context.Context, userID string) (runner.Stats, error) {
done <- userID
return runner.Stats{}, nil
}
tr := &discoveryTrigger{ctx: context.Background(), run: run, log: quietLog()}
tr.Enqueue("u1")
select {
case got := <-done:
require.Equal(t, "u1", got)
case <-time.After(2 * time.Second):
t.Fatal("Enqueue did not run the user's pass")
}
}
+82 -4
View File
@@ -20,10 +20,12 @@ import (
"net/http"
"os"
"os/signal"
"sync"
"time"
"gitea.d-ma.be/mathias/tapir/internal/adapters/secrets"
"gitea.d-ma.be/mathias/tapir/internal/adapters/store"
"gitea.d-ma.be/mathias/tapir/internal/adapters/youtube"
"gitea.d-ma.be/mathias/tapir/internal/auth"
"gitea.d-ma.be/mathias/tapir/internal/config"
"gitea.d-ma.be/mathias/tapir/internal/runner"
@@ -53,6 +55,8 @@ func main() {
err = cmdRun(ctx, log)
case "serve":
err = cmdServe(ctx, log)
case "report":
err = runReport(ctx, os.Args[2:])
default:
usage()
os.Exit(2)
@@ -73,6 +77,7 @@ usage:
tapir serve run the web UI (read summaries, record watch/skip/save)
tapir list [-limit N] list stored summaries, recent first
tapir show <video-id> show one summary in full
tapir report Stage-0 usage gate: per-user distinct active weeks
configuration is via TAPIR_* environment variables (see .env.example).
`)
@@ -125,10 +130,17 @@ func cmdRun(ctx context.Context, log *slog.Logger) error {
if engine == nil {
return fmt.Errorf("run: incomplete summarization config (gateway, youtube credentials, secrets file)")
}
r := runner.New(engine.Source, st, engine, cfg.UserID, log)
// Process-wide caption-fetch rate gate (ADR-014 item 2): the batch path shares
// the same per-egress-IP limiter as the web click-path.
youtube.SetFetchRate(cfg.FetchRate)
r := runner.New(engine.Source, st, engine, cfg.UserID, log,
runner.WithBackoff(cfg.FetchBackoff),
runner.WithAutoWindow(cfg.AutoSummarizeWindow))
log.Info("starting run", "user", cfg.UserID, "model", cfg.SummarizerModel,
"gateway", cfg.GatewayURL, "poll_interval", cfg.PollInterval)
"gateway", cfg.GatewayURL, "poll_interval", cfg.PollInterval, "fetch_backoff", cfg.FetchBackoff,
"fetch_rate", cfg.FetchRate, "auto_window", cfg.AutoSummarizeWindow)
return r.Loop(ctx, cfg.PollInterval)
}
@@ -152,6 +164,11 @@ func cmdServe(ctx context.Context, log *slog.Logger) error {
}
defer st.Close()
// Process-wide caption-fetch rate gate (ADR-014 item 2): the web click-path
// and the scheduled-discovery runners share one per-egress-IP limiter so they
// cannot collectively trip 429s. Must be set before either path fetches.
youtube.SetFetchRate(cfg.FetchRate)
// Auth seam (handlers depend on web.Auth only). With Dex configured
// (TAPIR_OIDC_ISSUER set) serve uses real OIDC login — any Dex subject may
// authenticate, then registers a tapir user (ADR-012); otherwise it falls
@@ -177,7 +194,11 @@ func cmdServe(ctx context.Context, log *slog.Logger) error {
// The file-backed SecretStore is shared by the connect flow (writes tokens)
// and account management (deletes them on disconnect / delete-account).
secretStore := secrets.NewFileStore(cfg.SecretsFile)
app := &web.App{Store: st, Identity: st, Auth: authn, Secrets: secretStore, Log: log}
app := &web.App{Store: st, Identity: st, Auth: authn, Secrets: secretStore, Log: log, RecencyWindow: cfg.AutoSummarizeWindow}
// User onboarding is handled by the IdP (Authentik invite flow), not Tapir —
// the Dex local-password provisioning path was removed (ADR-019). An
// authenticated subject with no Tapir user is routed to /register.
// Web-initiated YouTube connect (ADR-006). Mounted only when the OAuth client
// credentials are present; the refresh token persists through the SecretStore
@@ -189,7 +210,10 @@ func cmdServe(ctx context.Context, log *slog.Logger) error {
ClientSecret: cfg.YTClientSecret,
RedirectURL: cfg.YTConnectRedirectURL,
}, secretStore, st, log)
log.Info("web youtube connect enabled", "redirect", cfg.YTConnectRedirectURL)
// Paste-a-URL (Feature 2): same YouTube credentials, per-user adapter built
// per request. Mounting the /paste route keys off app.Fetcher being set.
app.Fetcher = videoFetcher{cfg: cfg, secrets: secretStore}
log.Info("web youtube connect + paste enabled", "redirect", cfg.YTConnectRedirectURL)
}
// Immediate summarization for the web "Summarize" button. When the engine can
@@ -207,6 +231,60 @@ func cmdServe(ctx context.Context, log *slog.Logger) error {
log.Info("web summarization is queue-only (incomplete engine config)")
}
// In-process scheduled discovery (ADR-018): when enabled, a background
// goroutine runs a discovery pass for ALL users on TAPIR_DISCOVERY_INTERVAL,
// reusing the per-user runner.Runner. Cancelled by the same ctx as the server.
//
// SINGLE-REPLICA ASSUMPTION (load-bearing): this loop lives in the web process.
// Running serve at >1 replica would make every replica fetch every user in
// parallel — duplicate work and self-inflicted 429s. replicas: 1 is required in
// the deployment manifest; scaling up needs a CronJob or leader election first.
if cfg.DiscoveryInterval > 0 {
log.Info("scheduled discovery enabled", "interval", cfg.DiscoveryInterval, "fetch_rate", cfg.FetchRate)
log.Warn("scheduled discovery assumes a SINGLE replica — running serve at >1 replica double-runs discovery (ADR-018)")
rawRunUser := func(ctx context.Context, userID string) (runner.Stats, error) {
r, err := buildUserRunner(cfg, st, secretStore, userID, log)
if err != nil {
return runner.Stats{}, err
}
return r.RunOnce(ctx)
}
// One lock shared by the scheduler and connect-triggered passes (#6) so
// they never fetch concurrently — the single-fetcher invariant (ADR-018).
runUser := serialize(&sync.Mutex{}, rawRunUser)
// Onboarding burst (Feature 1): after the connect-triggered discovery pass,
// summarize up to OnboardSummarizeCount of the user's NEWEST unsummarized
// videos so a fresh account gets real summaries in its first session. Hard
// cap; explicit, so it bypasses the recency window — but every fetch still
// goes through globalFetchGate via the Processor. No-op when disabled
// (count 0) or queue-only (no Processor).
onboard := func(ctx context.Context, userID string) {
if cfg.OnboardSummarizeCount <= 0 || app.Processor == nil {
return
}
ids, err := st.NewestUnsummarizedVideoIDs(ctx, userID, cfg.OnboardSummarizeCount)
if err != nil {
log.Warn("onboarding: list newest unsummarized", "user", userID, "err", err)
return
}
for _, id := range ids {
if err := app.Processor.ProcessVideo(ctx, userID, id); err != nil {
log.Warn("onboarding: summarize", "user", userID, "video", id, "err", err)
}
}
if len(ids) > 0 {
log.Info("onboarding burst complete", "user", userID, "summarized", len(ids), "cap", cfg.OnboardSummarizeCount)
}
}
if app.Connect != nil {
app.Connect.Discovery = &discoveryTrigger{ctx: ctx, run: runUser, onboard: onboard, log: log}
log.Info("connect-triggered discovery enabled", "onboard_cap", cfg.OnboardSummarizeCount)
}
go runScheduler(ctx, cfg.DiscoveryInterval, st, runUser, log)
} else {
log.Info("scheduled discovery disabled (TAPIR_DISCOVERY_INTERVAL unset or 0)")
}
srv := &http.Server{
Addr: cfg.HTTPAddr,
Handler: app.Router(),
+26 -1
View File
@@ -11,9 +11,29 @@ import (
"gitea.d-ma.be/mathias/tapir/internal/adapters/youtube"
"gitea.d-ma.be/mathias/tapir/internal/config"
"gitea.d-ma.be/mathias/tapir/internal/domain"
"gitea.d-ma.be/mathias/tapir/internal/ports"
"gitea.d-ma.be/mathias/tapir/internal/usecase"
"gitea.d-ma.be/mathias/tapir/internal/web"
)
// videoFetcher adapts the YouTube adapter to web.VideoFetcher for the paste flow
// (Feature 2). It builds a per-user adapter bound to that user's token ref and
// resolves a single video's metadata via the Data API — ungated; only the later
// transcript fetch goes through globalFetchGate.
type videoFetcher struct {
cfg config.Config
secrets ports.SecretStore
}
func (f videoFetcher) FetchVideo(ctx context.Context, userID, videoID string) (domain.Video, error) {
a := youtube.New(youtube.Config{
ClientID: f.cfg.YTClientID,
ClientSecret: f.cfg.YTClientSecret,
TokenSecretRef: web.YouTubeTokenRef(userID),
}, f.secrets)
return a.VideoByID(ctx, userID, videoID)
}
// buildProcessor wires the summarization engine — YouTube source (captions-first),
// AI-router summarizer, store sink — shared by `tapir run` and the web
// "Summarize now" path so the wiring lives in one place. It returns (nil, nil) —
@@ -43,7 +63,12 @@ func buildProcessor(cfg config.Config, st *store.Store) (*usecase.Engine, error)
}
sum := summarizer.New(primary, nil)
return usecase.NewEngine(src, sum, st), nil
// The store is both the summary sink and the shared transcript cache (ADR-021):
// the engine reads stored transcripts before any caption fetch and writes
// resolved ones back, so re-analysis never re-touches YouTube.
eng := usecase.NewEngine(src, sum, st)
eng.Transcripts = st
return eng, nil
}
// engineProcessor adapts the engine (which works in terms of a domain.Video) to
+67
View File
@@ -0,0 +1,67 @@
package main
import (
"context"
"fmt"
"io"
"os"
"text/tabwriter"
"gitea.d-ma.be/mathias/tapir/internal/adapters/store"
)
// gateThreshold is the Stage-0 gate (VISION/ADR-016): usage in >= 2 distinct
// weeks. The gate passes when any user reaches it.
const gateThreshold = 2
// runReport prints the Stage-0 usage gate: per-user distinct active weeks (reads
// UNION acts) and the pass/fail verdict. Read-only, cross-user — needs only
// TAPIR_DB_DSN (not TAPIR_USER_ID; the report enumerates all users itself).
func runReport(ctx context.Context, _ []string) error {
dsn := os.Getenv(envDSN)
if dsn == "" {
return fmt.Errorf("%s is required", envDSN)
}
s, err := store.New(ctx, dsn)
if err != nil {
return err
}
defer s.Close()
rows, err := s.ActiveWeeks(ctx)
if err != nil {
return err
}
return formatReport(os.Stdout, rows)
}
// formatReport renders the per-user week counts and the gate verdict. Pure: no DB,
// no env — so the layout and verdict logic are unit-testable without Postgres.
func formatReport(w io.Writer, rows []store.UserActiveWeeks) error {
if len(rows) == 0 {
_, err := fmt.Fprintln(w, "no users yet")
return err
}
tw := tabwriter.NewWriter(w, 0, 4, 2, ' ', 0)
_, _ = fmt.Fprintln(tw, "USER\tNAME\tACTIVE_WEEKS\tGATE")
passed := false
for _, r := range rows {
gate := "-"
if r.ActiveWeeks >= gateThreshold {
gate = "PASS"
passed = true
}
_, _ = fmt.Fprintf(tw, "%s\t%s\t%d\t%s\n", r.UserID, orDash(r.DisplayName), r.ActiveWeeks, gate)
}
if err := tw.Flush(); err != nil {
return err
}
verdict := fmt.Sprintf("\nGate (usage in >= %d distinct weeks): NOT YET MET\n", gateThreshold)
if passed {
verdict = fmt.Sprintf("\nGate (usage in >= %d distinct weeks): PASSED\n", gateThreshold)
}
_, err := fmt.Fprint(w, verdict)
return err
}
+45
View File
@@ -0,0 +1,45 @@
package main
import (
"strings"
"testing"
"github.com/stretchr/testify/require"
"gitea.d-ma.be/mathias/tapir/internal/adapters/store"
)
func TestFormatReportColumnsAndGatePass(t *testing.T) {
rows := []store.UserActiveWeeks{
{UserID: "user-a", DisplayName: "Ada", ActiveWeeks: 3},
{UserID: "user-b", DisplayName: "", ActiveWeeks: 1},
}
var b strings.Builder
require.NoError(t, formatReport(&b, rows))
out := b.String()
require.Contains(t, out, "USER")
require.Contains(t, out, "ACTIVE_WEEKS")
require.Contains(t, out, "Ada")
require.Contains(t, out, "PASSED", "a user at >= 2 weeks passes the gate")
// The >=2 user is marked PASS; the 1-week user is not.
require.Contains(t, lineContaining(t, out, "user-a"), "PASS")
require.NotContains(t, lineContaining(t, out, "user-b"), "PASS")
require.Contains(t, lineContaining(t, out, "user-b"), "-", "no display name falls back to dash")
}
func TestFormatReportGateNotMet(t *testing.T) {
rows := []store.UserActiveWeeks{{UserID: "user-a", ActiveWeeks: 1}}
var b strings.Builder
require.NoError(t, formatReport(&b, rows))
require.Contains(t, b.String(), "NOT YET MET", "no user at >= 2 weeks fails the gate")
}
func TestFormatReportEmpty(t *testing.T) {
var b strings.Builder
require.NoError(t, formatReport(&b, nil))
require.Contains(t, b.String(), "no users yet")
}
+154
View File
@@ -0,0 +1,154 @@
package main
import (
"context"
"fmt"
"log/slog"
"time"
"gitea.d-ma.be/mathias/tapir/internal/adapters/llm"
"gitea.d-ma.be/mathias/tapir/internal/adapters/store"
"gitea.d-ma.be/mathias/tapir/internal/adapters/summarizer"
"gitea.d-ma.be/mathias/tapir/internal/adapters/youtube"
"gitea.d-ma.be/mathias/tapir/internal/config"
"gitea.d-ma.be/mathias/tapir/internal/ports"
"gitea.d-ma.be/mathias/tapir/internal/runner"
"gitea.d-ma.be/mathias/tapir/internal/usecase"
"gitea.d-ma.be/mathias/tapir/internal/web"
)
// buildUserRunner constructs a runner.Runner for one user, reusing the same
// engine wiring as buildProcessor but bound to that user's own YouTube refresh
// token (web.YouTubeTokenRef(userID)) — the Stage-1 per-tenant ref, not the
// Stage-0 single ref. It returns an error (not nil) when the global config can't
// support live summarization (gateway, YouTube client creds, secrets file), so
// the scheduler can skip that user gracefully. A user who simply hasn't connected
// YouTube yet builds fine here; their token ref fails to resolve at RunOnce time,
// surfacing as a per-user error the scheduler logs and skips.
func buildUserRunner(cfg config.Config, st *store.Store, secretStore ports.SecretStore, userID string, log *slog.Logger) (*runner.Runner, error) {
if cfg.GatewayURL == "" || cfg.YTClientID == "" || cfg.YTClientSecret == "" || cfg.SecretsFile == "" {
return nil, fmt.Errorf("buildUserRunner: incomplete summarization config (gateway, youtube credentials, secrets file)")
}
src := youtube.New(youtube.Config{
ClientID: cfg.YTClientID,
ClientSecret: cfg.YTClientSecret,
TokenSecretRef: web.YouTubeTokenRef(userID),
PreferredLanguages: []string{"en"},
}, secretStore)
primary := summarizer.Endpoint{
Client: llm.New(cfg.GatewayURL, cfg.GatewayKey, cfg.SummarizerModel, cfg.SummarizerTimeout),
Provider: "local",
Model: cfg.SummarizerModel,
}
engine := usecase.NewEngine(src, summarizer.New(primary, nil), st)
return runner.New(src, st, engine, userID, log,
runner.WithBackoff(cfg.FetchBackoff),
runner.WithAutoWindow(cfg.AutoSummarizeWindow)), nil
}
// userLister enumerates every registered user and reports a user's video
// connections. *store.Store satisfies it via ListAllUsers + ConnectionsForUser.
// A small local interface keeps the scheduler testable with a fake.
type userLister interface {
ListAllUsers(ctx context.Context) ([]store.UserIdentity, error)
ConnectionsForUser(ctx context.Context, userID string) ([]store.Connection, error)
}
// runDiscoveryPass runs one discovery pass for every user. runUser performs a
// single user's pass (production: build a runner and RunOnce). Per-user failures
// — including a buildUserRunner error or a RunOnce error — are logged and skipped
// so one bad user, channel, or video never aborts the others (ADR-018 failure
// isolation). Returns the stats summed across users.
func runDiscoveryPass(
ctx context.Context,
lister userLister,
runUser func(context.Context, string) (runner.Stats, error),
log *slog.Logger,
) runner.Stats {
users, err := lister.ListAllUsers(ctx)
if err != nil {
log.Error("scheduler: list users failed", "err", err)
return runner.Stats{}
}
log.Info("scheduler: starting discovery pass", "users", len(users))
var total runner.Stats
for _, u := range users {
if ctx.Err() != nil {
break // shutting down: stop enumerating
}
// Skip users with no video connection. A discovery pass for them only
// attempts to resolve a token that was never minted, logging a spurious
// "ref not found" every tick (e.g. stale Dex-era orphan identities).
conns, err := lister.ConnectionsForUser(ctx, u.UserID)
if err != nil {
log.Warn("scheduler: list connections failed", "user", u.UserID, "err", err)
continue
}
if len(conns) == 0 {
log.Debug("scheduler: skipping user with no video connections", "user", u.UserID)
continue
}
stats, err := runUser(ctx, u.UserID)
total = sumStats(total, stats)
if err != nil {
log.Warn("scheduler: user discovery pass had errors", "user", u.UserID, "err", err)
}
}
log.Info("scheduler: pass complete",
"candidates", total.Candidates, "summarized", total.Summarized,
"skipped_seen", total.SkippedSeen, "skipped_no_text", total.SkippedNoText,
"skipped_manual", total.SkippedManual, "skipped_too_old", total.SkippedTooOld,
"skipped_rate_limited", total.SkippedRateLimited,
"channel_unavailable", total.ChannelUnavailable, "errors", total.Errors)
return total
}
// runScheduler runs a discovery pass on startup, then once every interval until
// ctx is cancelled (pod SIGTERM exits the loop cleanly). A non-positive interval
// disables scheduling entirely (no startup pass) so dev/tests never auto-fetch.
// It reuses the existing runner.Runner via runUser — the only new behaviour over
// runner.Loop is iterating all users per tick (ADR-018).
func runScheduler(
ctx context.Context,
interval time.Duration,
lister userLister,
runUser func(context.Context, string) (runner.Stats, error),
log *slog.Logger,
) {
if interval <= 0 {
return // disabled
}
runDiscoveryPass(ctx, lister, runUser, log)
ticker := time.NewTicker(interval)
defer ticker.Stop()
for {
select {
case <-ctx.Done():
return
case <-ticker.C:
runDiscoveryPass(ctx, lister, runUser, log)
}
}
}
// sumStats adds two passes' stats field-wise, so runDiscoveryPass can report a
// per-tick aggregate across all users.
func sumStats(a, b runner.Stats) runner.Stats {
return runner.Stats{
Candidates: a.Candidates + b.Candidates,
Summarized: a.Summarized + b.Summarized,
SkippedSeen: a.SkippedSeen + b.SkippedSeen,
SkippedNoText: a.SkippedNoText + b.SkippedNoText,
SkippedManual: a.SkippedManual + b.SkippedManual,
SkippedTooOld: a.SkippedTooOld + b.SkippedTooOld,
SkippedRateLimited: a.SkippedRateLimited + b.SkippedRateLimited,
ChannelUnavailable: a.ChannelUnavailable + b.ChannelUnavailable,
Errors: a.Errors + b.Errors,
}
}
+175
View File
@@ -0,0 +1,175 @@
package main
import (
"context"
"errors"
"io"
"log/slog"
"sync"
"testing"
"time"
"github.com/stretchr/testify/require"
"gitea.d-ma.be/mathias/tapir/internal/adapters/store"
"gitea.d-ma.be/mathias/tapir/internal/runner"
)
func quietLog() *slog.Logger {
return slog.New(slog.NewTextHandler(io.Discard, nil))
}
// fakeLister returns a fixed user set (or an error) for the scheduler under test.
type fakeLister struct {
users []store.UserIdentity
err error
noConn map[string]bool // users that have NOT connected a video source
}
func (f fakeLister) ListAllUsers(context.Context) ([]store.UserIdentity, error) {
return f.users, f.err
}
// ConnectionsForUser reports a single youtube connection for every user except
// those in noConn, which return zero — the connection-less case the scheduler
// must skip instead of running (and failing to resolve a token for).
func (f fakeLister) ConnectionsForUser(_ context.Context, userID string) ([]store.Connection, error) {
if f.noConn[userID] {
return nil, nil
}
return []store.Connection{{Provider: "youtube"}}, nil
}
// countingRunUser records how many passes each user got, optionally failing for
// specific users, under a mutex so it is safe across the scheduler goroutine.
type countingRunUser struct {
mu sync.Mutex
calls map[string]int
failFor map[string]bool
}
func newCountingRunUser(failFor ...string) *countingRunUser {
c := &countingRunUser{calls: map[string]int{}, failFor: map[string]bool{}}
for _, u := range failFor {
c.failFor[u] = true
}
return c
}
func (c *countingRunUser) run(_ context.Context, userID string) (runner.Stats, error) {
c.mu.Lock()
defer c.mu.Unlock()
c.calls[userID]++
if c.failFor[userID] {
return runner.Stats{Errors: 1}, errors.New("boom")
}
return runner.Stats{Summarized: 1}, nil
}
func (c *countingRunUser) count(userID string) int {
c.mu.Lock()
defer c.mu.Unlock()
return c.calls[userID]
}
func (c *countingRunUser) total() int {
c.mu.Lock()
defer c.mu.Unlock()
n := 0
for _, v := range c.calls {
n += v
}
return n
}
func usersN(ids ...string) []store.UserIdentity {
out := make([]store.UserIdentity, len(ids))
for i, id := range ids {
out[i] = store.UserIdentity{UserID: id, DexSubject: "dex|" + id}
}
return out
}
func TestDiscoveryPassRunsEveryUserOnce(t *testing.T) {
lister := fakeLister{users: usersN("a", "b", "c")}
rc := newCountingRunUser()
stats := runDiscoveryPass(context.Background(), lister, rc.run, quietLog())
require.Equal(t, 1, rc.count("a"))
require.Equal(t, 1, rc.count("b"))
require.Equal(t, 1, rc.count("c"))
require.Equal(t, 3, stats.Summarized, "stats are summed across users")
}
func TestDiscoveryPassSkipsUsersWithoutConnections(t *testing.T) {
// b never connected a video source (e.g. a stale Dex-era orphan identity).
// It must be skipped silently — not run and logged as a token error every pass.
lister := fakeLister{users: usersN("a", "b", "c"), noConn: map[string]bool{"b": true}}
rc := newCountingRunUser()
stats := runDiscoveryPass(context.Background(), lister, rc.run, quietLog())
require.Equal(t, 1, rc.count("a"))
require.Equal(t, 0, rc.count("b"), "a user with no connection must be skipped, not run")
require.Equal(t, 1, rc.count("c"))
require.Equal(t, 2, stats.Summarized, "only connected users contribute")
require.Equal(t, 0, stats.Errors, "skipping is silent — no spurious error stat")
}
func TestDiscoveryPassOneUserFailureDoesNotStopOthers(t *testing.T) {
lister := fakeLister{users: usersN("a", "b", "c")}
rc := newCountingRunUser("b") // user b's pass errors
stats := runDiscoveryPass(context.Background(), lister, rc.run, quietLog())
require.Equal(t, 1, rc.count("a"))
require.Equal(t, 1, rc.count("b"))
require.Equal(t, 1, rc.count("c"), "a failing user must not abort the rest")
require.Equal(t, 2, stats.Summarized) // a + c
require.Equal(t, 1, stats.Errors) // b
}
func TestDiscoveryPassListerErrorIsContained(t *testing.T) {
lister := fakeLister{err: errors.New("db down")}
rc := newCountingRunUser()
stats := runDiscoveryPass(context.Background(), lister, rc.run, quietLog())
require.Equal(t, 0, rc.total(), "no users enumerated → no passes")
require.Equal(t, runner.Stats{}, stats)
}
func TestSchedulerIntervalZeroDisablesEntirely(t *testing.T) {
lister := fakeLister{users: usersN("a", "b")}
rc := newCountingRunUser()
runScheduler(context.Background(), 0, lister, rc.run, quietLog())
require.Equal(t, 0, rc.total(), "interval 0 must not run even a startup pass")
}
func TestSchedulerRunsStartupPassThenStopsOnCancel(t *testing.T) {
lister := fakeLister{users: usersN("a", "b", "c")}
rc := newCountingRunUser()
ctx, cancel := context.WithCancel(context.Background())
done := make(chan struct{})
go func() {
// A long interval so only the startup pass runs before we cancel.
runScheduler(ctx, time.Hour, lister, rc.run, quietLog())
close(done)
}()
// The startup pass is synchronous at the top of runScheduler; once all three
// users have a pass it has completed and the loop is parked on the ticker.
require.Eventually(t, func() bool { return rc.total() == 3 }, time.Second, 5*time.Millisecond)
cancel()
select {
case <-done:
case <-time.After(time.Second):
t.Fatal("scheduler did not exit after ctx cancel")
}
require.Equal(t, 3, rc.total(), "no extra passes fired between startup and cancel")
}
+101 -4
View File
@@ -145,16 +145,113 @@ graph TB
engine in a **background goroutine** inside `serve`; the page HTMX-polls `/v/{id}/status`,
showing a Charmbracelet spinner while in-flight (and an honest "queued/waiting" state under
rate-limiting — ADR-014).
- **Summarization mode** — `users.auto_summarize` (migration 006). Auto: every new video is
summarized. Manual (default): new videos appear unsummarized; the button sets
`videos.summarize_requested`, which the next `tapir run` processes and clears. Both the click
path and the batch `tapir run` drive the same unchanged engine.
- **Summarization mode** — `users.auto_summarize` (migration 006). Default is **true** for new
users (migration 011, ADR-018); all existing rows were back-filled via migration 012. Auto:
new videos **published within the recency window** (`TAPIR_AUTO_SUMMARIZE_WINDOW`, default ~7d,
ADR-020) are summarized automatically; older videos are discovered and listed but wait for an
explicit "Summarize". Manual: new videos appear unsummarized; the button sets
`videos.summarize_requested`, which the next `tapir run` processes and clears. A manual request
bypasses the recency bound. Both the click path and the batch `tapir run` drive the same
unchanged engine.
- **List surface (ADR-020)** — the list reads `ListVideos` ordered summarized-first, then
`published_at DESC NULLS LAST`. The web layer collapses the noise so summaries are not buried:
un-summarized videos older than the recency window fold into one "Show N older videos"
disclosure, and caption-less videos collapse to a single count line. Copy surfaces scarcity
honestly (queue counts, gradual-fill note) — it never implies the feed is fuller than it is.
The engine, ports, and sink adapters are **untouched** by all of the above — the web surface only
reads the store and triggers the existing engine. Adding it changed wiring, not the core (ADR-003).
---
## In-process scheduler (ADR-018)
`cmdServe` launches a background goroutine when `TAPIR_DISCOVERY_INTERVAL > 0`. On each tick
it calls `store.ListAllUsers` (un-RLS'd admin query), builds a per-user `runner.Runner`, and
calls `RunOnce` for each registered user in sequence.
```mermaid
sequenceDiagram
participant S as Scheduler goroutine
participant DB as Postgres (RLS)
participant YT as YouTube timedtext
participant LLM as LiteLLM gateway
loop every TAPIR_DISCOVERY_INTERVAL
S->>DB: ListAllUsers() [un-RLS'd]
loop per user
S->>DB: GetAutoSummarize(userID)
S->>YT: ListSubscriptions + NewVideos
Note over S,DB: auto: skip videos published before<br/>TAPIR_AUTO_SUMMARIZE_WINDOW (ADR-020);<br/>older ones listed, await manual request
Note over S,YT: WaitFetchGate(ctx) throttles<br/>all fetches to TAPIR_FETCH_RATE
alt transcript available
S->>LLM: Summarize
S->>DB: Deliver(summary)
else 429
S->>DB: SetTranscriptStatus(rate_limited)
end
end
end
```
**Single-replica constraint (load-bearing).** The scheduler lives in the web process;
`replicas: 1` in the k3s deployment manifest is not cosmetic — running `tapir serve` at
>1 replica makes every replica run the full discovery loop, causing every registered user
to be fetched in parallel from the same egress IP (429s + duplicate work). Do not scale
`serve` past 1 replica without first moving discovery to a k8s CronJob or adding leader
election. The process logs a `Warn` at startup when scheduled discovery is enabled as a
reminder.
---
## Process-wide timedtext rate gate
**`internal/adapters/youtube/gate.go`** (ADR-014 item 2): a single `rate.Limiter`
(`golang.org/x/time/rate`) shared across **all** Adapter instances. Every `httpDo` call for
a caption fetch passes through `WaitFetchGate(ctx)` before hitting YouTube. This serialises
the scheduler loop AND the web click-path through the same per-egress-IP budget. Configured
via `TAPIR_FETCH_RATE` (Go duration, default `2s`). Setting it to `0` disables the gate
(dev/tests only).
This is the precondition that makes scheduled auto-summarize safe: without the gate, a
multi-user scheduler pass could fire many concurrent timedtext requests from the same IP
within seconds, triggering 429s for all users.
### Two-path summarisation model
Both paths share `globalFetchGate` — rate limiting is **respected in both**, not routed around.
| Path | Trigger | Order | Rationale |
|------|---------|-------|-----------|
| **Foreground** | User clicks "Summarize" on any non-summarized card (`POST /v/{id}/retry-now` for rate-limited; `POST /v/{id}/summarize` for pending) | Single chosen video | On-demand value: user picks a specific video to read now — bypasses the recency bound |
| **Background batch** | Scheduled discovery pass every `TAPIR_DISCOVERY_INTERVAL` | **Newest-first across all channels** (see below), **bounded to the recency window** (ADR-020) | Onboarding prioritisation within bounded load: recent videos auto-fill; the older back-catalogue stays on-demand |
The rationale for both paths is **onboarding prioritisation under an honest, bounded load** — a
new user gets summaries of their most recent videos automatically, while the older back-catalogue
is listed but summarised only on demand, so it never re-drives the shared rate gate every cycle.
### Newest-first batch ordering (ADR-018)
Within each scheduled pass, `RunOnce` uses a three-phase structure:
1. **Discover + persist**: walk all channels, `UpsertVideo` every candidate (so it appears in
the list), apply pre-filters (seen/manual/backoff/**recency**), collect surviving candidates.
The recency pre-filter (ADR-020) drops auto-mode videos published before
`now - TAPIR_AUTO_SUMMARIZE_WINDOW` unless they are explicitly requested; an undated video is
never aged out. They remain persisted/listed — only auto-summarisation is skipped.
2. **Sort**: order candidates `published_at DESC, NULLS LAST, discovery_pos ASC`. Videos with
no publish date (schema 001: nullable) sort after all dated content. The sort is in-memory
(`slices.SortStableFunc`) — at current scale this is fine.
3. **Process**: feed candidates to the engine in sorted order through `globalFetchGate`.
Before (per-channel inline): `[chanA-old, chanA-mid, chanB-new, chanB-null]`
After (newest-first): `[chanB-new, chanA-mid, chanA-old, chanB-null]`
The set of *processed* videos now also excludes auto-mode back-catalogue beyond the recency
window (those stay listed, summarised on demand); within the processed set, only order changes.
---
## Sequence — core use case: new video summarized
```mermaid
+53 -22
View File
@@ -10,12 +10,15 @@ only opaque references to them; the secret material lives in ESO/1Password (ADR-
## Design decisions baked into this model
- **Per-user isolation, not a shared global video table.** The earlier draft proposed a
global `videos`/`transcripts` table deduped across tenants. Rejected for Future B: it
reintroduces exactly the cross-domain coupling the homelab architecture review is
removing, and at 15 users the cost of occasionally re-summarizing the same video is
trivial compared to the isolation it would cost. Each user's data is self-contained.
(Revisit only if Future C makes GPU/transcription cost dominate — a new ADR, not a default.)
- **Per-user isolation for everything except transcripts.** The earlier draft proposed a
global `videos`/`transcripts` table deduped across tenants. **Videos** stay per-user and
RLS-scoped — a shared video table reintroduces exactly the cross-domain coupling the homelab
architecture review is removing. **Transcripts**, however, are now shared (ADR-021): keyed by
`(provider, provider_video_id)`, no `user_id`, **not** RLS-scoped. The cost avoided there is
not LLM re-summarization but a rate-gated, reputation-risky caption fetch (ADR-010/014), which
is paid per re-fetch regardless of user count — so persisting public caption content once and
sharing it strictly beats the coupling it removes. Everything else each user owns is
self-contained; `rls_test.go` proves transcripts is the single exception.
- **Secrets by reference only.** Tables hold a `secret_ref` (opaque string/UUID resolved via
the `SecretStore` port), never tokens or keys.
- **The brain sink is just a delivery target.** No brain-specific tables. Whether a summary
@@ -23,7 +26,7 @@ only opaque references to them; the secret material lives in ESO/1Password (ADR-
## Entities
Solid entities below are **persisted today** (migrations 001006). `AI_CREDENTIAL` and
Solid entities below are **persisted today** (migrations 001013). `AI_CREDENTIAL` and
`SUBSCRIPTION` are **planned, not yet a table** — kept in the model for intent; see the notes.
```mermaid
@@ -34,14 +37,15 @@ erDiagram
USER ||--o{ AI_CREDENTIAL : "has (planned)"
VIDEO_CONNECTION ||--o{ SUBSCRIPTION : "exposes (planned)"
SUBSCRIPTION ||--o{ VIDEO : "produces (per user)"
VIDEO ||--o| TRANSCRIPT : "has at most one"
VIDEO }o--o| TRANSCRIPT : "shares one by (provider, provider_video_id) — not FK (ADR-021)"
VIDEO ||--o| SUMMARY : "has at most one"
SUMMARY ||--o{ SINK_DELIVERY : "delivered via"
USER ||--o{ CHANNEL_ERROR : "reports unavailable channels"
USER {
uuid id PK
text display_name
bool auto_summarize "default false -> manual mode out of the box (migration 006)"
bool auto_summarize "default true for new users (migration 011, ADR-018)"
timestamptz created_at
}
USER_IDENTITY {
@@ -87,14 +91,16 @@ erDiagram
text url
bool summarize_requested "default false -> manual-mode queue flag (migration 006)"
timestamptz seen_at
text transcript_status "none|rate_limited|fetched (migration 007)"
timestamptz rate_limited_at "backoff clock for 429 retries (migration 007)"
}
TRANSCRIPT {
uuid video_id PK_FK
uuid user_id FK
text provider PK "part of shared key (ADR-021)"
text provider_video_id PK "part of shared key — the cross-user dedup key"
text source "captions | none"
text language
text content "null when source = none"
timestamptz resolved_at
timestamptz fetched_at
}
SUMMARY {
uuid id PK
@@ -123,21 +129,31 @@ erDiagram
text action "watched | skipped | saved"
timestamptz acted_at
}
CHANNEL_ERROR {
uuid user_id FK
text channel_id
text channel_name
timestamptz first_seen
timestamptz last_seen
}
```
`SUMMARY_ACTION` has `UNIQUE (user_id, video_id, action)`; `VIDEO_CONNECTION` has
`UNIQUE (user_id, provider)` (one connection per provider — reconnect upserts in place).
RLS (`ENABLE` + `FORCE`) is on **every solid user-owned table above**`users`, `videos`,
`transcripts`, `summaries`, `summary_actions`, `video_connections`. `sink_deliveries` is
RLS'd via an `EXISTS` on its parent summary; `user_identities` is intentionally **not** RLS'd
(auth plumbing). See the *Isolation invariant* section for the mechanism.
`transcripts`, `summaries`, `summary_actions`, `video_connections`, `channel_errors`.
`sink_deliveries` is RLS'd via an `EXISTS` on its parent summary; `user_identities` is
intentionally **not** RLS'd (auth plumbing). See the *Isolation invariant* section for the
mechanism.
## Notes per entity
- **USER** — one row per registered user (Stage 1, ADR-012; no longer single-row). The Tapir-side
profile; the Dex identity is held separately in `USER_IDENTITY`, not on this row. `auto_summarize`
(migration 006) is the per-user mode flag: `FALSE` (default) = manual, `TRUE` = auto-summarize
every new video.
(migration 006) is the per-user mode flag: `TRUE` = auto-summarize new videos **published within
the recency window** (`TAPIR_AUTO_SUMMARIZE_WINDOW`, default ~7d, ADR-020); older videos are
listed but summarised on demand. Default is **true** for new users (migration 011, ADR-018);
existing rows were back-filled via migration 012 with RLS bypass.
- **USER_IDENTITY** (migration 004) — the `dex_subject → user_id` map. `dex_subject` is the PK,
`user_id` a `UNIQUE` FK to `users` with `ON DELETE CASCADE`. This is the bridge resolved at login
*before* a `user_id` is known, so it is **deliberately not RLS-enabled** (it holds no user data;
@@ -157,8 +173,15 @@ RLS'd via an `EXISTS` on its parent summary; `user_identities` is intentionally
decision. The same video seen by two users is two rows. `seen_at` is when Tapir detected it.
`summarize_requested` (migration 006) is the manual-mode queue flag: the web "Summarize" button
sets it `TRUE`; the next `tapir run` picks it up, summarizes, and clears it back to `FALSE`.
- **TRANSCRIPT** — at most one per video. `source = none` records "checked, no usable
transcript" so the watcher doesn't reprocess (ADR-007). `content` null in that case.
`transcript_status` and `rate_limited_at` (migration 007) track caption-fetch outcomes for
rate-limit backoff: `NULL` = not attempted; `rate_limited` = 429 seen, skip until
`NOW() - rate_limited_at > TAPIR_FETCH_BACKOFF`; `fetched` = resolved; `none` = no transcript.
- **TRANSCRIPT** — shared public caption content, one row per `(provider, provider_video_id)`,
**not** RLS-scoped and carrying no `user_id` (ADR-021). Two users who watch the same video
share the one row; the summarize path reads it before any caption fetch, so re-analysis never
re-touches YouTube (ADR-010/014). `source = none` records "checked, no usable transcript" so
no one reprocesses (ADR-007); `content` null in that case. A transient 429 is never stored
here — it stays a per-user retry via `VIDEO.transcript_status`.
- **SUMMARY** — at most one per video. `fallback_used` + `ai_provider`/`ai_model` make the
"is local good enough?" question queryable (the Stage 0 quality signal). `highlights`/
`takeaways` as jsonb to stay schema-flexible while the output format settles.
@@ -170,6 +193,12 @@ RLS'd via an `EXISTS` on its parent summary; `user_identities` is intentionally
save) — the column that makes the Stage-0 headline metric ("acts on ≥1 summary") queryable
(ui-spec.md §5, ADR-011). `video_id` is `TEXT` and **not** FK-constrained, mirroring summaries'
standalone `(user_id, video_id)` key. `UNIQUE (user_id, video_id, action)`. FORCE RLS'd.
- **CHANNEL_ERRORS** (migration 013) — channels that returned HTTP 404 (deleted or private) on
the most recent discovery pass. Upserted per scheduler pass (`last_seen` refreshed each run);
surfaced on the account page as a warning. Cascades on user deletion. Primary key is
`(user_id, channel_id)`. FORCE RLS'd.
- **LOGIN_EVENTS** (migration 010) — throttled one-row-per-(user, date) login stamp. Used by the
Stage-0 gate query (VISION §Stage 0: "returned and used in ≥2 distinct weeks").
## Isolation invariant (Stage 1+) — LIVE
@@ -180,8 +209,9 @@ enforcement dormant); **ADR-012 opened Stage 1 and turned enforcement on in the
Enforcement is **Postgres Row-Level Security** (migration `003_rls.up.sql`):
- RLS is `ENABLE`d **and** `FORCE`d on every user-owned table — `users`, `videos`,
`transcripts`, `summaries`, `summary_actions`, `video_connections`. `FORCE` is load-bearing:
the app connects as the table **owner** (`tapir` role), and owners bypass RLS unless forced.
`transcripts`, `summaries`, `summary_actions`, `video_connections`, `channel_errors`.
`FORCE` is load-bearing: the app connects as the table **owner** (`tapir` role), and owners
bypass RLS unless forced.
- Each policy keys off the per-request GUC `tapir.current_user_id`, set transaction-locally by
the store's `withUser` helper via `set_config('tapir.current_user_id', $1, true)` — it
auto-resets on commit/rollback, so it never leaks across a pooled connection.
@@ -208,7 +238,8 @@ queue, doesn't replace it). Deferred until there's a reason.
## Explicitly out of scope (Future C)
- Global cross-tenant video/transcript dedup (rejected above).
- Global cross-tenant *video* dedup (rejected above). Note: cross-tenant *transcript* sharing
is now in scope and shipped (ADR-021); only the videos half stays per-user.
- Sharding / per-tenant physical databases.
- Soft-delete + full audit trail on connections/credentials (a Stage 2 hardening item; add
via ADR when Stage 2 work starts).
+27 -2
View File
@@ -2,7 +2,7 @@
The concrete endpoints, conventions, and identifiers Tapir depends on, so an independent
session doesn't have to rediscover them. **Verify anything marked "confirm" before relying on
it** — endpoints and aliases drift, and this file is a snapshot (2026-06-02), not a live source.
it** — endpoints and aliases drift, and this file is a snapshot (2026-06-06), not a live source.
## Local AI (the Primary in `llm.Router`)
@@ -160,7 +160,7 @@ allow per-provider when a user connects one.
---
_Snapshot date 2026-06-02. Items marked **confirm** were not verified to a pinned source at
_Snapshot date 2026-06-06. Items marked **confirm** were not verified to a pinned source at
snapshot time — check brain or the live cluster before depending on them._
## Stage 1 — multi-user facts (verified 2026-06-03)
@@ -191,3 +191,28 @@ snapshot time — check brain or the live cluster before depending on them._
- `user_identities(dex_subject → user_id)` table is **intentionally NOT RLS-enabled**
(it's auth plumbing, holds no user data; data isolation is on the user-owned tables).
All data access after subject resolution goes through `withUser`.
## Scheduled discovery (ADR-018, verified 2026-06-05)
`tapir serve` runs discovery for **all users** in-process on a timer (no CronJob). Three env
knobs plus one load-bearing deployment constraint:
- `TAPIR_DISCOVERY_INTERVAL` — Go duration, e.g. `2h`. The cadence the serve process runs a
discovery pass for every registered user (run-once-on-startup, then every interval).
**Unset or `0` = disabled** (dev/tests never auto-fetch).
- `TAPIR_FETCH_RATE` — Go duration, default `2s`. The **process-wide per-egress-IP caption-fetch
rate gate** (ADR-014 item 2). Every caption fetch — scheduler runners *and* the web "Summarize"
click-path — serialises through this one limiter so the pod cannot collectively trip 429s. `0`
= unlimited (dev/tests). This is the precondition that makes auto-summarize-on-a-schedule safe;
do not raise it aggressively without watching for 429s.
- `TAPIR_FETCH_BACKOFF=4h` — per-video rate-limit retry window; default `1h`. A video that
returns HTTP 429 on a caption fetch is skipped for this duration before being retried. The
scheduler checks `NOW() - rate_limited_at > TAPIR_FETCH_BACKOFF` before attempting to fetch
a video marked `transcript_status = rate_limited`. Longer values reduce 429 pressure at the
cost of slower recovery after a throttling episode.
- **SINGLE-REPLICA WARNING (load-bearing).** The scheduler lives in the web process, so
`replicas: 1` in the deployment manifest is load-bearing: running `tapir serve` at >1 replica
makes **every** replica run the discovery loop → every user fetched in parallel from the same
egress IP (429s + duplicate work). Do **not** scale `serve` past 1 replica without first moving
discovery to a k8s CronJob or adding leader election. The process logs a `Warn` at startup when
scheduled discovery is enabled, as a reminder.
+77
View File
@@ -0,0 +1,77 @@
# Spec — Newest-first batch ordering + honest "Try now" / prioritisation docs
> **Extended by ADR-020 (2026-06-08).** This spec covers the *batch processing* order within a
> pass. ADR-020 adds (a) a recency pre-filter — auto mode skips videos published before
> `TAPIR_AUTO_SUMMARIZE_WINDOW`, listed but summarised on demand — and (b) the same
> `published_at DESC NULLS LAST` ordering on the **list read** (`ListVideos`), which previously
> sorted by `seen_at`. See `DECISIONS.md` ADR-020.
**Repo:** tapir · **Size:** small · **Solo session.**
**Why.** Product intent (maintainer, 2026-06-06): a new user should get summaries of their
**newest** videos quickly, while the older back-catalogue fills in behind — all within the one
shared rate gate. Today the foreground path ("Try now" button) lets a user hand-pick a video,
but the **background batch processes in subscription/channel order, not newest-first** — so a new
user with a large candidate set sees the batch summarise whatever channel is first in their
subscription list, not their newest videos. This slice makes the batch agree with the intent, and
fixes the docs to describe the real rationale (onboarding prioritisation), not the
traffic-disguising framing a prior session wrote.
Read `CLAUDE.md` + `DECISIONS.md` (ADR-014, ADR-018) first. TBD, conventional commits,
`task check` green per commit, `templ generate` if views change.
## 1. Newest-first batch ordering (the build)
In `internal/runner/runner.go` `RunOnce`: today the loop processes each video inline while
walking subscriptions channel-by-channel (`for sub → NewVideos → for v → process`). Change so
that, within a pass, **candidates are processed newest-first across ALL channels**:
- Collect the candidate videos across channels first (after dedup/seen/manual/rate-limit
filtering as today), then **sort by `published_at` descending before processing**, then process
in that order through the engine + shared `globalFetchGate`.
- **`published_at` is nullable** (schema 001). Sort **NULLS LAST** — videos with no publish date
must not jump ahead of dated newest videos. Decide a stable tiebreak (e.g. `seen_at DESC`) for
equal/again-null dates.
- Keep all existing behaviour: per-item failure isolation, the rate-limit backoff skip, manual
mode, channel-unavailable handling, stats. Ordering is the only change — not what gets
processed, just the order.
- At 868 candidates a collect-then-sort in memory is fine; do **not** build a streaming/external
sort. Keep it simple.
- The shared rate gate (`globalFetchGate`) is unchanged and still governs fetch pacing — ordering
does not bypass or weaken it.
**Optional (only if cheap and clearly correct):** a soft cap so the *first* pass for a brand-new
user summarises the newest N (e.g. 20) quickly and defers the long tail to subsequent passes — so
onboarding value lands fast without waiting for the whole sorted set. If this adds real
complexity, SKIP it and just do the newest-first ordering; the ordering alone delivers the intent.
## 2. Tests
- Given candidates across multiple channels with mixed `published_at` (incl. some NULL), assert
the processing order is newest-first, NULLS LAST, with the chosen tiebreak. Use the existing
fake VideoStore/Processor pattern in `runner_test.go`.
- Assert ordering does not change *which* videos are processed vs. today (same set, new order).
- Rate-gate / backoff / manual-mode behaviour unchanged (existing tests stay green).
## 3. Docs — describe the REAL rationale (replace prior framing)
The "Try now" button and the discovery batch together implement **onboarding prioritisation**:
foreground (user-clicked "Try now") summarises a specific video on demand; background batch
summarises newest-first; both honour the shared rate gate. **Update the docs to state this intent
— and explicitly REMOVE/replace any framing that describes "Try now" as making traffic "look
organic to YouTube" or evading rate limits.** That is not the rationale. The rationale is: *get
the user a few summaries of their newest, most relevant videos fast; process the back-catalogue in
the background; always within the honest shared rate limit.* Rate limiting is **respected**, not
evaded.
- `docs/ui-spec.md`: "Try now" = on-demand foreground summarisation of a chosen (typically newer)
video; rationale = fast onboarding value, not traffic shaping.
- `docs/architecture/architecture.md`: document the two-path model — foreground on-demand vs.
background newest-first batch, both through `globalFetchGate` — and the newest-first ordering.
- Any requirements/use-case doc mentioning discovery order: state newest-first.
- If a brain note or `wiki` entry captured the "looks organic" rationale, correct it there too.
## Boundaries
- Do NOT increase fetch rate or weaken the rate gate. Account-safety constraint stands: the
caption endpoint is unofficial (ADR-010) and must be treated with honest backoff, never evasion.
- Do NOT touch RLS, credentials, or the Dex surface.
- Ordering change is within a pass only — no persisted priority queue, no new table.
## Out of scope
Per-user configurable ordering; priority weighting beyond newest-first; the soft-cap if it proves
non-trivial.
+92
View File
@@ -0,0 +1,92 @@
# Spec — In-process scheduled discovery + auto-summarize + rate-gate finish
> **Extended by ADR-020 (2026-06-08).** Auto-summarize is no longer "every unseen video": the
> scheduler now skips videos published before `TAPIR_AUTO_SUMMARIZE_WINDOW` (default ~7d) unless
> explicitly requested, so a back-catalogue does not re-drive the rate gate every cycle. See
> `DECISIONS.md` ADR-020.
**Repo:** tapir · **Size:** medium · **Solo session** (not a swarm).
**Why this exists.** The Stage-0 gate ("me or a friend returns and reads/acts in ≥2 separate
weeks") cannot be met because the system is not usable *unprompted*: discovery (`tapir run`) is
host-side manual, so a newly onboarded user sees an empty list and never comes back. This slice
makes Tapir watch on its own — the thing that makes the gate experiment actually runnable.
Read `CLAUDE.md` + `DECISIONS.md` (esp. ADR-012, ADR-014, and the new ADR-018) first. TBD —
commit directly to `main`, one logical change per commit, conventional commits, `task check`
green before each commit. `templ generate` if any view changes.
## Decisions already made (do not reopen)
- **In-process scheduler**, NOT a k8s CronJob (maintainer's call: simpler deploy, acceptable
coupling at 3 users). The known cost — discovery shares the web process's lifetime and egress
— is accepted and recorded in ADR-018.
- **Auto-summarize ON** for the maintainer + onboarded friends (zero-friction: the list fills
and summarizes itself).
- **Gate clock resets** to when this ships (ADR-018) — until unprompted use is possible, the
prior window measured nothing.
## 1. In-process scheduled discovery (core)
- In `tapir serve` startup, launch a background goroutine that runs discovery for ALL users on
an interval: env `TAPIR_DISCOVERY_INTERVAL` (Go duration, e.g. `2h`). **Unset or 0 = disabled**
(so dev/tests never auto-fetch).
- **Reuse the existing `runner.Runner` + `Loop`/`RunOnce`. Do NOT write a new scheduler.** The
per-user Runner already exists; the new work is **iterating users** and running one pass each
per tick. Enumerate users from the un-RLS'd `user_identities` (the same enumerate-then-act
pattern the login_events gate query established), then run each user's pass **inside that
user's RLS scope** (`withUser`).
- **Stateless timing:** run-once-on-startup, then every interval — exactly the existing `Loop`
shape. Do NOT persist schedule state; a pod restart just restarts the cycle. Acceptable at this
scale. Do not build cron-in-Go.
- **Graceful shutdown:** the goroutine respects `ctx` cancellation so a pod term doesn't wedge.
- **Failure isolation in the loop:** one user's pass failing (or one channel/video) must not
abort the other users or crash `serve` — log and continue. (RunOnce already collects per-item
errors; preserve that at the per-user level too.)
## 2. Auto-summarize default ON for Future-B users
- New registrations default `auto_summarize = true` (so onboarded friends get zero-friction);
keep the account-page toggle so a user can switch to manual. One-off update existing user rows
to `true` as well (maintainer + any current users).
- Consequence (intended): scheduled discovery both discovers AND summarizes new videos — which
is the point, and is why §3 is mandatory in the same slice.
## 3. Finish/confirm the ADR-014 shared per-egress-IP rate gate (NOW load-bearing)
- In-process scheduling + auto-summarize + multiple users = all caption fetches leave the **one
web pod's egress IP**, concurrently with any live "Summarize" button clicks. The timedtext
endpoint rate-limits per IP (ADR-010/014). Without a shared gate this self-inflicts 429s every
cycle.
- **Confirm in code whether ADR-014 item 2 (a single PROCESS-WIDE rate gate) exists.**
Reconciliation flagged it as possibly built only as per-*video* backoff. If it is not a
process-wide gate, **build it now**: ONE shared limiter (token-bucket / min-interval) that
every timedtext/caption fetch passes through — scheduler loop AND click-path alike. Per-process,
not per-user, not per-video.
- Keep the existing per-video 429 backoff (`rate_limited_at` + retry window) — complementary: the
gate prevents tripping 429; the backoff handles it if one still happens.
- Honest UX (ADR-014 item 3) still applies: a fetch waiting on the gate shows "queued/waiting",
never a stuck spinner.
## 4. Gate-clock reset — already recorded in ADR-018; verify VISION reflects it
- ADR-018 (committed) resets the Stage-0 34 week window to start when this ships, and revises the
check-in date. VISION Stage 0 carries a pointer to it. The build doesn't re-decide this; just
ensure nothing in docs still implies the clock started earlier.
## Tests
- **Scheduler:** fake clock + fake Runner → N users each get one pass per tick; one user's failure
doesn't stop the others; `ctx` cancel stops the loop; interval=0 disables it entirely.
- **Rate gate:** concurrent fetches (scheduler + simulated click) are serialized/limited through
the ONE gate — assert max-in-flight / min-interval honored regardless of caller.
- **Auto-summarize default:** new registration → `auto_summarize = true`; account-page toggle
still flips it.
## Out of scope / known constraints
- No CronJob / k8s objects (in-process chosen).
- **SINGLE-REPLICA ASSUMPTION (load-bearing).** In-process scheduling means if `tapir serve` ever
runs >1 replica, every replica runs the discovery loop → every user fetched in parallel (429s +
duplicate work). At 3 users this is single-replica, fine — but the build MUST note this
constraint in ADR-018 / deploy docs so a future scale-up doesn't silently double-run.
- No persisted schedules, no multi-pod coordination.
## Fallback if the session runs short
Ship the scheduler with **auto-summarize OFF** (discovery only; manual Summarize button) until the
process-wide rate gate (§3) is confirmed/built. NEVER ship auto-summarize-on-a-schedule without the
gate — that combination self-inflicts 429s for every user every cycle. Auto-summarize ON is gated
on §3 being done.
+77
View File
@@ -0,0 +1,77 @@
# Spec — Stage 0 usage measurement (login events)
**Date:** 2026-06-03
**Status:** Ready to build · **Repo:** tapir · **Size:** small (one migration + middleware + query)
**Why:** The Stage 0 gate (VISION, ADR-016) is *return usage in ≥2 separate weeks*. `summary_actions`
captures *acts* (watch/skip/save) but not *reads* — a friend who logs in weekly and reads summaries
without clicking anything is invisible. For a **reading** product that is the most important signal.
This adds the missing data so the gate is measurable as written. Solo session, not a swarm.
Read `CLAUDE.md` + ADR-016 first. TBD, conventional commits, `task check` green before each commit.
## Scope (resist sprawl — this is NOT analytics)
A lightweight, append-only record of *when each user was active*, enough to answer
"returned/read in ≥N distinct weeks". Not page-level events, not click tracking, not a funnel.
### 1. Migration — `login_events` (append-only)
```
login_events (
id UUID PK default gen_random_uuid(),
user_id UUID NOT NULL, -- per-user; RLS like every user-owned table
seen_at TIMESTAMPTZ NOT NULL default NOW()
)
INDEX (user_id, seen_at)
```
- **RLS:** `FORCE ROW LEVEL SECURITY`, same policy/pattern as the other user-owned tables (the
`tapir.current_user_id` GUC via the `withUser` seam — match migration 003). A reporting query that
needs cross-user counts runs as the owner/maintainer outside the per-user scope, or via a dedicated
read — decide consistently with how existing admin-ish reads are done.
- Append-only: no updates, no deletes except the user-delete cascade. **Add to the delete-account
cascade** (ADR-013) — `login_events` has no FK (mirrors `summary_actions`), so `DeleteUser` needs an
explicit delete for it, and the delete test must assert it's covered. *Do not forget this* — it's the
exact footgun the last delete work caught.
### 2. Middleware — throttled stamp
- In the authenticated request path (after `CurrentUserID` resolves, inside the registration-gated
app — NOT on `/welcome`/`/healthz`/`/auth`), record one `login_events` row **per user per day**
(throttle: skip if a row exists for this user with `seen_at` ≥ start-of-today). One insert per active
day, not per request — keeps the table small and the signal clean.
- Throttle check must itself be RLS-scoped (`withUser`). Keep it cheap (indexed lookup).
### 3. Query — the gate report
Provide a query (and optionally a tiny `tapir report` CLI subcommand or an admin page — your call,
CLI is fine) answering, per user:
```sql
-- distinct active weeks from reads (login_events) AND acts (summary_actions), unioned
WITH weeks AS (
SELECT user_id, date_trunc('week', seen_at) AS wk FROM login_events
UNION
SELECT user_id, date_trunc('week', acted_at) FROM summary_actions
)
SELECT user_id, COUNT(DISTINCT wk) AS active_weeks
FROM weeks GROUP BY user_id
ORDER BY active_weeks DESC;
```
Gate passes when any user_id (maintainer or friend) reaches `active_weeks >= 2` within the window.
## Honesty caveats to carry (from VISION/ADR-016)
- **"Unprompted" is not measurable here.** login_events records *that* a user returned, not *why*. A
nudged return looks identical to an organic one. This build does not close that gap and must not
claim to — the VISION measurement note stands: count returns, read a nudged return as weaker signal.
(If prompt-tracking is ever wanted, that's a separate decision, not this build.)
- **Data accrues from deploy onward.** The gate window's read-data starts when this ships — so ship
soon (maintainer's call) rather than batching with the infra tooling session.
- **`date_trunc('week')` is ISO/timezone-sensitive** and noisy at low volume (N=3). Two visits days
apart can fall in the same or different weeks. Acceptable, but don't over-read a single-week-margin
pass/fail.
## Out of scope
Page/event analytics; prompt-vs-organic tracking; dashboards beyond the one gate query; anything
touching the engine or sinks (this is web/store only — ADR-003 holds).
## Tests
- Migration up/down; RLS on `login_events` (extend the two-user isolation test to cover it).
- Throttle: N requests same day → 1 row; next day → 2nd row.
- `DeleteUser` removes the user's `login_events` and leaves others' intact (extend the delete test).
- The gate query returns correct distinct-week counts across a seeded reads+acts fixture.
+77
View File
@@ -0,0 +1,77 @@
# Spec — Unify video-card states: one "Summarize now" verb, honest no-captions state
> **Superseded in part by ADR-020 (2026-06-08).** The five card states still hold, but the copy
> changed: the nudge verb is now **"Summarize"** (not "Summarize now"), the rate-limited state
> reads **"In queue"** (not "Fetching soon…"), and the queued state reads **"summarizing
> shortly"** (not "waiting for the next run"). The list also now collapses older un-summarized
> and caption-less videos. See `DECISIONS.md` ADR-020 and `views.templ` (`VideoCard`) for the
> current copy; this doc is kept as the original design record.
**Repo:** tapir · **Size:** small, **view-layer only** (`views.templ` + a little CSS in
`view.go`; regenerate `views_templ.go`). No handler, store, or DB change. The two existing
handlers (`/summarize`, `/retry-now`) stay exactly as they are — only what the card *shows*
changes.
**Why.** The video card today presents two different verbs — "Try now" (`.btn-retry`, on
rate-limited videos) and "Summarize" (`.btn-secondary`, on pending videos) — for what the user
experiences as one intent: *"summarize this video now."* The user doesn't know or care about the
internal pipeline state (rate-limited vs. manual-queue); two differently-labelled, differently-
styled buttons leak that state machine into the UI as a choice. Also: a **no-captions video
(`TranscriptStatus == "none"`) currently falls into the `else` branch and wrongly shows a
"Summarize" button** that, if clicked, tries and fails — there are no captions to fetch. That's
the confusing dead-end to remove.
In normal use the maintainer is in **auto mode**, so these buttons are *exceptions*, not the main
path — videos summarize themselves. So the card should be **status-first**: the state is what the
user reads constantly; the manual nudge is a small, quiet affordance for impatience, not a
prominent call-to-action.
## The honest per-state card model (footer of `VideoCard`)
Restructure the footer branch in `VideoCard` (in `internal/web/views.templ`) to these states.
The branch ORDER matters (summarized first, then terminal/no-action states, then actionable):
1. **Summarized** — preview + provider chip + fallback badge + actions. **No button.** (unchanged)
2. **No captions** (`r.TranscriptStatus == "none"`) — **NEW branch.** Quiet status text, e.g.
`No transcript available` (use a muted `.card-state`/`.chip-retry`-style treatment, NOT a
button). This is a terminal honest dead-end — the user can do nothing, so offer nothing.
3. **Queued / requested** (`r.SummarizeRequested`) — "Queued · waiting for the next run". **No
button.** (unchanged)
4. **Rate-limited** (`r.TranscriptStatus == "rate_limited"`) — quiet status (keep the
"fetching soon" sense) + a **quiet "Summarize now"** button POSTing to `retryNowURL` (clears
backoff then processes). Same quiet style as state 5.
5. **Pending** (else — discovered, not yet attempted) — quiet "Not summarized" + a **quiet
"Summarize now"** button POSTing to `summarizeURL` (flips the queue flag then processes).
## Unify the verb and the style
- **One label everywhere a manual nudge is offered: "Summarize now"** (states 4 and 5). Drop the
"Try now" wording entirely.
- **One quiet style** for both: use the understated `.btn-retry` pattern (small, pill, outline,
transparent bg) — NOT `.btn-secondary`/`.btn` (heavier). Rename the CSS class to something
state-neutral (e.g. `.btn-quiet` or `.btn-summarize-now`) so it no longer reads as
"retry"-specific; keep the same visual. The point: the nudge is subtle, status is primary.
- Keep both `<form>`s posting to their respective existing handler URLs
(`retryNowURL` for rate-limited, `summarizeURL` for pending) with the existing HTMX
attributes (`hx-post`, `hx-target=#video-{id}`, `hx-swap=outerHTML`) — only the button
label/class change. The backend side-effect difference (clear-backoff vs. set-flag) stays
invisible to the user, which is correct.
- Drop the engineer-facing `title="Fetch transcript now through the shared rate gate"` tooltip;
if a hint is wanted, make it user-facing ("Summarize this one now").
## Quietness check (the design intent)
The summary content and the per-state *status* are the card's primary information. The "Summarize
now" button is a minor affordance. Do not make it a prominent solid-accent CTA — it must read as
"you can nudge this if you're impatient", not "action required". Status text uses muted styling;
the button uses the quiet outline style.
## Tests
- `videocard_internal_test.go` (exists): assert each of the 5 states renders the expected
footer — summarized (no button), no-captions (status, NO button, no `summarize`/`retry-now`
URL present), queued (no button), rate-limited ("Summarize now" → retry-now URL), pending
("Summarize now" → summarize URL). The key new assertion: **a `none`-status video renders no
action button and no POST URL.**
- Assert the label string "Try now" no longer appears anywhere in rendered output.
## Out of scope
Handler/DB changes; the detail-page action buttons (watched/skipped/saved — unrelated); the
pipeline stats bar wording; auto/manual mode behaviour. Verb/label/style/no-captions-state only.
+17 -9
View File
@@ -72,7 +72,14 @@ summary_actions
and join into the existing `SummaryRow` reads so list/detail show current state.
- This column is what makes the Stage-0 metric ("did I act on a summary?") queryable.
## 6. Auth (Dex OIDC, single-user authz)
## 6. Auth (Dex OIDC)
Authentication is delegated to the homelab OIDC provider at `TAPIR_OIDC_ISSUER`
**Authentik** since the Dex→Authentik migration (infra ADR-0001; ADR-019). It offers a
Google upstream and Authentik-managed accounts (incl. its invite flow); Tapir no longer
provisions accounts itself. Any authenticated subject can register a Tapir account
(ADR-012: allowlist removed). The `oidc`/`DexAuth` package keeps its name for now (rename
deferred, ADR-019).
- **Flow:** standard Authorization Code. Use `coreos/go-oidc` + `golang.org/x/oauth2`
(justify the deps in the commit; both are the homelab-standard OIDC libs and small).
@@ -92,7 +99,7 @@ summary_actions
`TAPIR_OIDC_ISSUER` (`https://auth.d-ma.be`), `TAPIR_DEX_CLIENT_ID`, `TAPIR_DEX_CLIENT_SECRET`,
`TAPIR_OIDC_REDIRECT_URL` (`https://tapir.d-ma.be/auth/callback`), `TAPIR_SESSION_SECRET`.
Reuses existing `TAPIR_DB_DSN`, `TAPIR_USER_ID` (the StubAuth dev subject only). No secrets
committed. (`TAPIR_ALLOWED_SUBJECT` was removed by ADR-012.)
committed. (`TAPIR_ALLOWED_SUBJECT` was removed by ADR-012; use the keys above.)
## 8. Deployment — k3s + Flux GitOps
@@ -163,12 +170,13 @@ distinguishable.
| **Registration gate** | A Dex subject with no `users` row is routed to `/register`, which creates the `users` row + a `user_identities` mapping. (§2 listed "sign-up / user CRUD" as a non-goal.) | Explicit registration is how a multi-user surface stays honest — no just-in-time row creation. | ADR-012; `f396e01` |
| **Per-user YouTube web connect** | `/oauth/youtube/connect``/oauth/youtube/callback` stores a per-user refresh-token ref + a `video_connections` row. (The spec assumed a host-side `tapir auth` only.) | Multi-user means each user connects their own account from the browser. | ADR-006, ADR-012; migration 005 (`0c9531a`, `2aad79b`) |
| **Account management** | `/account` page with **disconnect** and **delete account**; delete removes only Tapir-side state and leaves the Dex identity intact. (§2 listed isolation/CRUD as non-goals.) | A real account needs a way out; deletion semantics are deliberately Tapir-side only. | ADR-013; `22eafcf`, `c7624d9`, `17d5e8c` |
| **Immediate web summarization** | A "Summarize" button (`POST /v/{id}/summarize`) runs the engine in a background goroutine inside `serve`; the page HTMX-polls `GET /v/{id}/status`. (§2 said "triggering runs from the browser … do NOT build".) | Reading a list you can't act on is half a product; on-demand summarize closes the loop without waiting for a batch `tapir run`. | ADR-012, ADR-014; `25215cb`, `8c6c7ca` |
| **Immediate web summarization** | A quiet "Summarize now" button on non-summarized video cards. Pending cards POST to `/v/{id}/summarize` (queues + triggers engine); rate-limited cards POST to `/v/{id}/retry-now` (clears backoff + triggers engine). Both use the same `.btn-quiet` style and label — the internal pipeline distinction is invisible to the user. The page HTMX-polls `GET /v/{id}/status` while processing. Videos with `TranscriptStatus == "none"` show "No transcript available" with no button — this is a terminal honest state. (§2 said "triggering runs from the browser … do NOT build".) | Reading a list you can't act on is half a product; on-demand summarize closes the loop without waiting for a batch `tapir run`. | ADR-012, ADR-014; `25215cb`, `8c6c7ca` |
| **Charmbracelet tapir spinner** | An animated in-flight indicator (charm palette) shown while a summarize is processing; an honest "queued/waiting" state under rate-limiting rather than a stuck spinner. | The spinner must tell the truth when the timedtext endpoint rate-limits (429), not imply imminence. | ADR-014; `25215cb`, `a4aeb5e` |
| **Auto/manual summarization mode** | Per-user `auto_summarize`; manual (default) lists new videos unsummarized and queues via `summarize_requested`; a mode toggle at `/account/summarize-mode`. | Control over compute/noise — only summarize what the user cares about. | migration 006 (`748d5eb`, `bdbdce7`, `3014ee0`, `a269d4a`) |
| **Auto/manual summarization mode** | Per-user `auto_summarize`; manual lists new videos unsummarized and queues via `summarize_requested`; a mode toggle at `/account/summarize-mode`. Default is **true** for new users (migration 011, ADR-018); existing rows back-filled via migration 012. | Control over compute/noise — only summarize what the user cares about. | migration 006 (`748d5eb`, `bdbdce7`, `3014ee0`, `a269d4a`); migration 011/012 |
| **Public landing page** | `/welcome` mounted **outside** the auth guard; unauthenticated `/` redirects there; logout returns there (not `/auth/login`). (The spec guarded everything except `/healthz` and `/auth/*`.) | A first-time visitor needs a public "what is this / get started" page before the login wall. | `d83943c`, `0fdf2f7`, `3a27bf1`, `d208110`, `8ca374e`, `f15f57f` |
The original Stage-0 goals (read summaries, record watch/skip/save actions, Dex login, GitOps
deploy) still hold — these are additions over that base, not replacements. The architecture
stance is unchanged: every item above is web-surface or store work; the engine/ports/sinks core
was not modified (ADR-003).
| **Invite onboarding** | **Removed from Tapir (ADR-019).** Invites are owned by the IdP (Authentik) now, not Tapir — the Dex local-password provisioning path (`tapir invite` CLI, `/invite/{token}` web flow, `internal/adapters/dex`) was deleted when the homelab migrated Dex→Authentik (infra ADR-0001). A new user is invited via Authentik's invite flow, logs into Tapir via OIDC, and is captured by the existing `/register` (display-name) gate. | Onboarding belongs to the identity provider; keeps Tapir out of the shared identity provider's write path. | ADR-019; infra ADR-0001 |
| **Summarized-only filter** | `?summarized=1` query param on the list view. When set, only videos with a completed summary (`SummaryRow.Summarized = true`) are shown. Rendered as a "Summarized only" checkbox in the filter form. Summarized videos also sort to the top of the unfiltered list (`ORDER BY (s.id IS NOT NULL) DESC, seen_at DESC`). | Lets users focus on videos that are ready to read; newly landing summaries are visible at the top without filtering. | `internal/web/view.go` (`Filter.OnlySummarized`, `ListVideos` ORDER BY) |
| **"Summarize now" foreground path** | Unified quiet nudge button on actionable non-summarized cards. Five explicit card states — (1) summarized: chip + no button; (2) no captions (`transcript_status = 'none'`): "No transcript available", no button; (3) queued: "Queued" chip, no button; (4) rate-limited: "Fetching soon…" + "Summarize now" → `POST /v/{id}/retry-now` (clears `rate_limited_at`, triggers engine); (5) pending: "Not summarized" + "Summarize now" → `POST /v/{id}/summarize` (queues + triggers engine). One verb, one style (`.btn-quiet`); backend difference invisible to user. Both handlers call `ProcessVideo` through `globalFetchGate`. Rate gate respected, not bypassed — this is onboarding prioritisation. | Fast onboarding value; honest dead-end for no-captions videos (no button that fails). | `internal/web/handlers.go` (`handleRetryNow`, `handleRequestSummarize`); `internal/web/views.templ` (`VideoCard`) |
| **Pipeline stats bar** | A one-line status bar above the video list: `N summarized · M fetching soon · K no captions`. Computed from the unfiltered row set; hidden when all videos are summarized. Gives the user a clear read on pipeline state without any interaction. | Replaces the "why is nothing happening?" confusion when most videos are pending or rate-limited. | `internal/web/view.go` (`PipelineStats`, `pipelineStats`) |
| **Unavailable channels (account page)** | The `/account` page shows a "Unavailable channels" section when any channels returned HTTP 404 on the last discovery pass. Lists channel name, an "unavailable" badge, and the first-seen date. Data sourced from the `channel_errors` table (migration 013). | Surfaces silent failures so users know why some subscribed channels produce no new videos. | migration 013; `internal/web/account.go`; `internal/adapters/youtube/youtube.go` (`domain.ErrChannelUnavailable`) |
| **Recency window + sparse-state honesty (ADR-020)** | Supersedes the copy/sort in the rows above. Auto-summarize is bounded to videos published within `TAPIR_AUTO_SUMMARIZE_WINDOW` (~7d); older un-summarized videos collapse behind a single "Show N older videos — summarize on demand" disclosure, and caption-less videos collapse to a one-line count (not N cards). List order is now `summarized-first, published_at DESC NULLS LAST`. Copy reframed for honest scarcity: pipeline bar reads "N ready · M in queue · K no captions" (no "fetching soon"); a gradual-fill note explains the rate limit; the nudge verb is "Summarize" (not "Summarize now"); the queued card says "summarizing shortly"; the empty-connected state drops the impossible `tapir run` instruction. Detail leads with Takeaways. Filters slimmed (no date pickers; hidden when empty); watched/skipped segmented; back link on detail; empty terms checkbox removed. | Make the sparse reality legible and honest instead of implying abundance/imminence; bound auto load so the back-catalogue doesn't re-drive the caption gate. Never fetch harder — scarcity is surfaced, not engineered around. | ADR-020; `2384c47`, `3df0459`, `40b703e`, `a1a5217`, `4a0a56e`, `9bf1c31`, `980638d`, `12fb031`, `f775441`, `51aa5d9` |
+17
View File
@@ -10,12 +10,26 @@ Feature: Connect and manage video accounts
And my refresh token is stored only as a secret reference
And my subscriptions are synced
@pending
# Vimeo connect is not built yet (provider label exists; no connect flow or test).
Scenario: Connect a Vimeo account
Given I have no connected video accounts
When I connect my Vimeo account
Then the connection is stored with status "active"
And my subscriptions are synced
Scenario: Connecting an account discovers videos immediately
Given I have no connected video accounts
When I connect my YouTube account
Then a discovery pass for my account is triggered right away
And I do not have to wait for the next scheduled pass to see my videos
Scenario: Connecting summarizes my newest videos right away
Given I have no connected video accounts
When I connect my YouTube account
Then up to the onboarding cap of my newest videos are summarized through the rate gate
And the rest are left to the scheduled recency-bounded pass
Scenario: Tokens are never stored in the clear
When I connect any video account
Then no OAuth token value is stored in the database
@@ -28,6 +42,9 @@ Feature: Connect and manage video accounts
And no new videos are watched for that connection
And my existing summaries remain readable
@pending
# Per-provider BYO credential config is not built as a web flow yet (the summarizer
# supports a fallback endpoint, but there is no user-facing BYO setup + its test).
Scenario Outline: BYO AI credential is optional and per-provider
When I configure a BYO provider "<provider>"
Then the credential is stored only as a secret reference
+3
View File
@@ -19,6 +19,9 @@ Feature: Public landing page
Then I see a link to my summaries
And I see a way to log out
@pending
# Behaviour ships (logout redirects to /welcome) but is not unit-tested: logout lives in
# the OIDC Auth impl and StubAuth has no routes to exercise it cheaply.
Scenario: Logging out returns to the welcome page
Given I am logged in
When I log out
+30
View File
@@ -0,0 +1,30 @@
Feature: Paste a YouTube URL to summarize any video
As a user
I want to paste a YouTube link and get a summary
So that I can pull the specific video I want now, even from channels I don't follow
Scenario: Paste a valid YouTube URL
Given I am connected
When I paste a valid YouTube video URL
Then the video is added to my feed scoped to me
And it is queued for summarization through the shared rate gate
Scenario: Pasting an invalid link is rejected
When I paste something that is not a YouTube video URL
Then I get a clear error and nothing is added
Scenario: Pasting a video that cannot be found is honest
When I paste a URL whose video cannot be found
Then I am told it couldn't be found and nothing is added
Scenario: Pasting the same video twice does not duplicate it
Given I have pasted a video
When I paste the same video again
Then my feed still has exactly one entry for it
@pending
# Covered by the engine's ADR-010 no-transcript terminal state (degrade-never-error);
# there is no paste-specific test for it.
Scenario: A pasted video with no captions resolves honestly
When I paste a video that has no captions
Then it resolves to the "no transcript available" terminal state
+3
View File
@@ -34,6 +34,9 @@ Feature: Register and manage a multi-user account
And the other user's data remains intact
And my Dex identity is left intact
@pending
# Re-registration after delete is supported by design (delete leaves the Dex identity,
# ADR-013) but has no dedicated end-to-end test yet.
Scenario: A deleted user can register again as a fresh account
Given I deleted my Tapir account but my Dex identity still exists
When I sign in again
+20 -5
View File
@@ -6,15 +6,25 @@ Feature: Choose how new videos get summarized
Background:
Given I am a registered user with a connected video account
Scenario: Auto mode summarizes every new video
Scenario: Auto mode summarizes recent new videos automatically
Given my summarization mode is "auto"
When a subscribed channel posts a new video with captions
When a subscribed channel posts a new video with captions within the recency window
Then Tapir summarizes it without my asking
And the summary appears in my list
Scenario: Manual mode is the default and leaves new videos unsummarized
Given I have not changed my summarization mode
Then my mode is "manual"
Scenario: Auto mode lists older videos without summarizing them
Given my summarization mode is "auto"
When discovery finds a video published before the recency window
Then the video appears in my list with no summary
And it is not summarized automatically
And I can still summarize it on demand with "Summarize"
Scenario: Automatic is the default for a new user
Given I have just registered
Then my summarization mode is "auto"
Scenario: Manual mode leaves new videos unsummarized
Given my summarization mode is "manual"
When a subscribed channel posts a new video with captions
Then the video appears in my list with no summary
And nothing is summarized until I request it
@@ -30,3 +40,8 @@ Feature: Choose how new videos get summarized
# auto_summarize is a per-user setting and summarize_requested is a per-video queue
# flag (migration 006). The web button sets the flag; `tapir run` processes both the
# auto videos and the manually queued ones, then clears the flag.
#
# Recency bound (ADR-020): in auto mode the scheduler only summarizes videos published
# within TAPIR_AUTO_SUMMARIZE_WINDOW (default ~7d); older videos are discovered and
# listed but wait for an explicit "Summarize" — so a back-catalogue does not re-drive
# the per-IP caption gate (ADR-014) every cycle. A manual request bypasses the bound.
@@ -31,5 +31,16 @@ Feature: Summarize new videos from subscribed channels
When the watcher sees "Designing for Attention" again
Then Tapir does not produce a second summary for it
Scenario: Re-analyzing a stored video does not re-fetch its transcript
Given a transcript for "Designing for Attention" is already stored
When the video is summarized again
Then Tapir reads the stored transcript
And Tapir does not fetch captions from YouTube
# Captions-first is the core path (ADR-007). Audio-download + speech-to-text is
# deferred and intentionally has no scenario here yet.
#
# Transcript persistence (ADR-021): the stored transcript is shared, keyed by
# (provider, provider_video_id) and read before any caption fetch, so the
# re-analysis scenario above also covers paste-a-URL and the onboarding burst —
# both summarize through the same engine chokepoint.
+207
View File
@@ -0,0 +1,207 @@
# Tapir — Heuristic Review (Stage-0, sparse + recency-bounded)
_Findings document, not a build spec. The maintainer filters; a spec follows separately.
Reviewed against `VISION.md` (Stage-0 gate), `docs/ui-spec.md`, `internal/web/views.templ` +
`view.go`, the prior `UX-REVIEW.md` pass, and the current screenshots. Written for the product
**as it actually is**: ~283 discovered, ~15 summarized, ~256 in queue behind a respected per-IP
caption rate limit, ~12 no-captions; single-user (maintainer) with friends pending; recency-
bounded auto-summarize about to ship (auto = recent ~7d, older browsable + manual on demand)._
## Reviewer stance
The Stage-0 gate is **return usage**. So every finding is judged by one question: does this make
the maintainer (or a friend) come back to an honestly-sparse feed? The visual layer is already
decent — dark mode, cards, the constrained reader were all fixed in the prior pass. The open
problems are **expectation-setting, honesty-of-scale, and the recency feed** — not pixels.
**The single biggest risk:** a new user connects, sees `15 summarized · 256 fetching soon`,
nothing visibly moves, and never returns. The whole gate dies at that moment. Most P0s below
attack that one moment.
## Tagging
- **`[NOW]`** — improves the product as it is today (sparse, recency-bounded, single-user). Ships
in the upcoming bundle.
- **`[LATER]`** — improves the product we hope it becomes (abundance, engagement, multiple users).
Valuable but premature until real usage validates the core loop.
Severity: 🔴 breaks the core loop · 🟠 hurts it · 🟡 noticeable · 🔵 polish.
---
## P0 — fix before/with the recency ship
### 1. 🔴 Empty-connected state tells the user to run a CLI command they can't run — `[NOW]`
**Problem.** After connecting YouTube, the empty list says: *"Run `tapir run` to discover your
subscriptions."* A friend on the web has no shell. And post-ADR-018 discovery is an in-process
scheduled loop — so the instruction is wrong *even for the maintainer*. It is the first thing a
newly-onboarded user sees, and it is an impossible, stale instruction.
**Principle.** Match between system and the real world; help users recognize, not be blocked
(Nielsen #1, #2, #9).
**Proposal.** Replace with a passive, honest "we're working" state: *"Your account is connected.
Tapir is finding your subscriptions and fetching captions — summaries appear here gradually. Check
back later."* No command. No imperative the user can't satisfy.
**Evidence.** `views.templ` `summaryList``.empty-connected`.
### 2. 🔴 No expectation set for *gradual* fill — the return loop breaks at the cliff — `[NOW]`
**Problem.** Captions are rate-limited by design; the backlog trickles over days. Nothing tells
the user this. A first visit shows few/no summaries and no "come back" framing. The Stage-0 gate
is literally about returns, and the product never asks for one or explains why patience is
warranted.
**Principle.** Visibility of system status (#1). And: the gate can't be cleared if the UX doesn't
survive first contact.
**Proposal.** One honest sentence near the pipeline bar / empty state: *"Tapir fetches captions
slowly on purpose, to respect YouTube's limits. New summaries land gradually — usually best to
check back tomorrow."* Turns confusing emptiness into intentional design. Highest-leverage change
in the review.
### 3. 🟠 "256 fetching soon" overstates imminence — a lie of scale — `[NOW]`
**Problem.** The pipeline bar and rate-limited cards both say "fetching soon." For 256 items
behind a per-IP throttle, "soon" is false — they trickle over days/weeks. Exactly the abundance-
implying language the mandate forbids, inverted: it makes the *queue* look imminent.
**Principle.** Honesty of scarcity (project mandate); #1.
**Proposal.** Relabel by scale. Bar: `15 ready · 256 in queue · 12 no captions`. Card state:
"In queue" / "Waiting its turn", not "Fetching soon…". Reserve "soon" for items actually next.
**Evidence.** `views.templ` `pipelineBar`; card State 4.
### 4. 🟠 Recency boundary is invisible in the feed — `[NOW, ships with recency]`
**Problem.** Once auto-summarize is bounded to ~7d, a 6-month-old pending video and a 2-day-old
pending video render identically ("Not summarized" + button). But only one is in the auto path;
the other will *never* process unless clicked. The user can't tell "be patient, this is coming"
from "this is yours to trigger or ignore."
**Principle.** Visibility of system status; predictability (#1).
**Proposal.** Bucket the list into two sections: **Recent** (auto, will fill itself) and
**Older — browse / summarize on demand**. Card status language should encode *which side of the
line it's on*, not the pipeline internals. Central IA decision of the recency change.
### 5. 🟠 The un-summarized mass buries the ~15 readable summaries — `[NOW]`
**Problem.** ~268 of 283 cards are not readable (pending / queued / no-captions). Summarized-first
ordering helps, but the page is still 95% noise below the fold. "Attention is the scarce resource"
is the product's own principle — and the default view violates it.
**Principle.** Aesthetic/minimalist design; signal-to-noise (#8).
**Proposal.** Default view = readable summaries + the Recent bucket. Collapse the older
un-summarized mass behind *"Show 256 older un-summarized videos."* Collapse the 12 no-caption
videos into a single line: *"12 videos have no captions"* (terminal, never readable — they don't
deserve 12 full cards).
---
## P1 — high, near-term
### 6. 🟠 "Summarize now" overpromises against the rate gate — `[NOW]`
**Problem.** Button says "Summarize now" (title: "Summarize this one now"). Backend queues it
behind the shared per-IP gate. Click 10 older videos and they all sit at "Fetching soon…". The
verb sells immediacy the system can't honor.
**Principle.** Honesty; match system/reality (#1).
**Proposal.** Drop "now" → "Summarize". On click the card should honestly become "Queued" (it
already can). Optionally show queue position once the queue is real. Don't engineer the gate
harder — just stop the verb from lying.
**Evidence.** `views.templ` card States 4 & 5; `handlers.go` `handleRetryNow`,
`handleRequestSummarize`.
### 7. 🟠 Welcome-page copy is stale and misleading post-ADR-019 — `[NOW]`
**Problem.** Landing sub-copy: *"If you have an invite link, it will set up your account
automatically."* Invites moved to Authentik (ADR-019); Tapir no longer handles invite links. The
"Get Started" button goes straight to OIDC. The copy promises a flow that no longer exists.
**Principle.** Match between system and reality (#2); honesty.
**Proposal.** Rewrite: *"Tapir is invite-only right now. If you've been invited, sign in below."*
Single CTA. Also set the gradual-fill expectation here (ties to #2) so it lands before the wall,
not after.
**Evidence.** `views.templ` `WelcomePage``.welcome-sub`.
### 8. 🟡 "Queued · waiting for the next run" leaks system jargon — `[NOW]`
**Problem.** "the next run" exposes the discovery-loop concept; a user doesn't know what a "run"
is.
**Principle.** Speak the user's language (#2).
**Proposal.** "Queued — summarizing shortly." Hide the scheduler.
**Evidence.** `views.templ` card State 3.
### 9. 🟡 Filters render before there's anything to filter — `[NOW]`
**Problem.** The fresh/empty state shows the full Channel/From/To/Filter bar *above* "Connect
YouTube." Power tooling stacked on top of the one action that matters.
**Principle.** Progressive disclosure; minimalist design (#8).
**Proposal.** Hide the filter bar when there are 0 rows (and arguably below ~20). Show the connect
CTA alone.
**Evidence.** `fixes/03-empty-fresh.png`; `ListPage` renders `filterForm` unconditionally.
### 10. 🟡 Date-range filters are dead weight at this scale — `[NOW]` demote / `[LATER]` rebuild
**Problem.** From/To date pickers + a free-text exact-match Channel field are corpus-scale tools.
With 15 summaries they're noise; channel-as-freetext is unguessable. (Confirm the prior review's
#7 is resolved — that `Channel` shows a real channel name, not the `provider` string; seeded
screenshots suggest it is, live data may differ.)
**Principle.** Match tool to task; minimalist design.
**Proposal.** `[NOW]`: reduce to the "Summarized only" toggle (+ maybe channel chips derived from
present rows). Drop date pickers until the corpus justifies them. `[LATER]`: real channel facets +
search when there's volume.
---
## P2 — reading experience & polish
### 11. 🟡 Detail leads with Summary; the attention-saving payload (Takeaways) is last — `[NOW]`
**Problem.** Product promise is "decide what's worth your time." The element that answers that —
Takeaways / verdict — sits at the bottom. The user reads a full summary to reach the point.
**Principle.** Lead with the user's actual job-to-be-done.
**Proposal.** Reorder or add a one-line TL;DR/verdict at top. Takeaways → Highlights → Summary, or
a "Worth watching?" lede. Data already exists; a reorder, not new machinery.
**Evidence.** `views.templ` `DetailPage`.
### 12. 🔵 No "back to Summaries" on detail — `[NOW]`
**Problem.** Only the brand returns home, losing any filter context.
**Proposal.** Explicit "← Summaries" link on the detail page.
### 13. 🔵 Auto/manual toggle copy will be wrong after recency — `[NOW, with recency]`
**Problem.** Account copy: *"Automatic summarizes every new video as it is discovered."* Becomes
false once auto is bounded to ~7d.
**Proposal.** *"Automatic summarizes new videos from the last ~7 days. Older videos stay
browsable — summarize them on demand."*
**Evidence.** `views.templ` `AccountPage` summarization section.
### 14. 🔵 Register step asks acceptance of nonexistent terms — `[NOW]`
**Problem.** "I accept the terms of use" — no terms linked. For a friends-only tool, ceremony
accepting nothing.
**Proposal.** Either link real terms or drop the checkbox at this stage.
**Evidence.** `views.templ` `RegisterPage`.
### 15. 🔵 watched/skipped mutual exclusivity unsignaled — `[NOW]` low
**Problem.** Three independent-looking buttons; watched↔skipped are exclusive (carryover from
prior review #14).
**Proposal.** Segmented control for watched/skipped; keep Saved separate.
---
## `[LATER]` — premature until the core loop is validated
_(abundance / engagement / multiple users)_
- **Return-nudges (digest email / push).** 🔴 **Caution, not just defer.** A notification that
drives returns *contaminates the exact signal the gate measures* — VISION wants *unprompted*
returns and admits it can't distinguish prompted from organic. Building a nudge now poisons the
experiment. Defer until after the gate reads. `[LATER]`
- **Full-text search across summaries** — needs volume to matter. `[LATER]`
- **Channel facets / saved filters / sorting** — corpus-scale tooling (#10). `[LATER]`
- **Read/unread + "new since last visit"** — genuinely helps returns, but only meaningful once
there is throughput to be "new." `[LATER]`
- **Saved/queue view, collections** — engagement surface; no payoff at 15 items. `[LATER]`
- **Richer card previews (top-takeaway as preview, 23 lines)** — triage aid that only pays off
with many cards to triage. `[LATER]`
- **Backlog progress heartbeat ("N summarized this week", queue burn-down)** — rewards returning,
but needs real throughput to show motion; a static "last updated X ago" is the only `[NOW]`-worthy
slice. `[LATER]`
- **Onboarding tour / multi-step welcome** — over-built for one user + a few friends. `[LATER]`
---
## What's already right (don't regress)
Pipeline-bar concept, the honest "No transcript available" terminal state, the queued/waiting
spinner instead of a fake-imminent one, the Unavailable-channels surface, the constrained-width
reader, dark mode, the account danger-zone behind a disclosure. The honesty instincts are present
— the P0s are about making that honesty *legible and correctly-scaled*, not adding it.
## Two judgment calls to settle first
1. **#4 (Recent/Older split) and #5 (collapse the older mass) are one decision viewed twice.**
Settle the recency-feed IA once and both fall out.
2. **The return-nudge caution (`[LATER]` list, item 1) is the one to put in writing now** — before
someone "helpfully" ships an email digest to juice the gate and destroys the signal.
Binary file not shown.

After

Width:  |  Height:  |  Size: 116 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 73 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 58 KiB

+2
View File
@@ -10,7 +10,9 @@ require (
github.com/golang-migrate/migrate/v4 v4.19.1
github.com/jackc/pgx/v5 v5.9.2
github.com/stretchr/testify v1.11.1
golang.org/x/crypto v0.45.0
golang.org/x/oauth2 v0.36.0
golang.org/x/time v0.15.0
)
require (
+4
View File
@@ -91,6 +91,8 @@ go.opentelemetry.io/otel/trace v1.37.0 h1:HLdcFNbRQBE2imdSEgm/kwqmQj1Or1l/7bW6mx
go.opentelemetry.io/otel/trace v1.37.0/go.mod h1:TlgrlQ+PtQO5XFerSPUYG0JSgGyryXewPGyayAWSBS0=
go.uber.org/goleak v1.3.0 h1:2K3zAYmnTNqV73imy9J1T3WC+gmCePx2hEGkimedGto=
go.uber.org/goleak v1.3.0/go.mod h1:CoHD4mav9JJNrW/WLlf7HGZPjdw8EucARQHekz1X6bE=
golang.org/x/crypto v0.45.0 h1:jMBrvKuj23MTlT0bQEOBcAE0mjg8mK9RXFhRH6nyF3Q=
golang.org/x/crypto v0.45.0/go.mod h1:XTGrrkGJve7CYK7J8PEww4aY7gM3qMCElcJQ8n8JdX4=
golang.org/x/oauth2 v0.36.0 h1:peZ/1z27fi9hUOFCAZaHyrpWG5lwe0RJEEEeH0ThlIs=
golang.org/x/oauth2 v0.36.0/go.mod h1:YDBUJMTkDnJS+A4BP4eZBjCqtokkg1hODuPjwiGPO7Q=
golang.org/x/sync v0.18.0 h1:kr88TuHDroi+UVf+0hZnirlk8o8T+4MrK6mr60WkH/I=
@@ -99,6 +101,8 @@ golang.org/x/sys v0.41.0 h1:Ivj+2Cp/ylzLiEU89QhWblYnOE9zerudt9Ftecq2C6k=
golang.org/x/sys v0.41.0/go.mod h1:OgkHotnGiDImocRcuBABYBEXf8A9a87e/uXjp9XT3ks=
golang.org/x/text v0.31.0 h1:aC8ghyu4JhP8VojJ2lEHBnochRno1sgL6nEi9WGFGMM=
golang.org/x/text v0.31.0/go.mod h1:tKRAlv61yKIjGGHX/4tP1LTbc13YSec1pxVEWXzfoeM=
golang.org/x/time v0.15.0 h1:bbrp8t3bGUeFOx08pvsMYRTCVSMk89u4tKbNOZbp88U=
golang.org/x/time v0.15.0/go.mod h1:Y4YMaQmXwGQZoFaVFk4YpCt4FLQMYKZe9oeV/f4MSno=
gopkg.in/check.v1 v0.0.0-20161208181325-20d25e280405/go.mod h1:Co6ibVJAznAaIkqp8huTwlJQCZ016jof/cbN4VW5Yz0=
gopkg.in/check.v1 v1.0.0-20201130134442-10cb98267c6c h1:Hei/4ADfdWqJk1ZMxUNpqntNwaWcugrBjAiHlqqRiVk=
gopkg.in/check.v1 v1.0.0-20201130134442-10cb98267c6c/go.mod h1:JHkPIbrfpd72SG/EVd6muEfDQjcINNoR0C8j2r3qZ4Q=
+15 -7
View File
@@ -10,13 +10,17 @@ import (
// DeleteUser permanently removes a user and all of their data. It runs through
// withUser so RLS confines every statement to the calling user's own rows.
//
// Deleting the users row cascades (ON DELETE CASCADE) to videos, transcripts,
// summaries (→ sink_deliveries), video_connections, and the user_identities map
// referential-integrity cascades bypass RLS, so a user's child rows are removed
// even though the deleting connection is scoped. summary_actions is the exception:
// it carries a user_id but has NO foreign key to users (migration 002), so the
// cascade does not reach it; it is deleted explicitly in the same scoped
// transaction. Deleting an absent user is a no-op (idempotent).
// Deleting the users row cascades (ON DELETE CASCADE) to videos, summaries
// (→ sink_deliveries), video_connections, and the user_identities map
// referential-integrity cascades bypass RLS, so a user's child rows are removed
// even though the deleting connection is scoped. Transcripts are NOT removed:
// since ADR-021 they are shared public content keyed by (provider,
// provider_video_id) with no user_id, so another user may still reference the
// same row — a user deletion must not strip shared caption content. summary_actions and login_events
// are the exceptions: each carries a user_id but has NO foreign key to users
// (migrations 002 and 010), so the cascade does not reach them; they are deleted
// explicitly in the same scoped transaction. Deleting an absent user is a no-op
// (idempotent).
//
// This is tapir-side only (decision 2026-06-03): it removes all tapir data; the
// Dex login identity is left untouched — a later login simply re-enters
@@ -28,6 +32,10 @@ func (s *Store) DeleteUser(ctx context.Context, userID string) error {
`DELETE FROM summary_actions WHERE user_id = $1`, userID); err != nil {
return fmt.Errorf("store: delete summary_actions: %w", err)
}
if _, err := tx.Exec(ctx,
`DELETE FROM login_events WHERE user_id = $1`, userID); err != nil {
return fmt.Errorf("store: delete login_events: %w", err)
}
if _, err := tx.Exec(ctx,
`DELETE FROM users WHERE id = $1`, userID); err != nil {
return fmt.Errorf("store: delete user: %w", err)
+59
View File
@@ -0,0 +1,59 @@
package store
import (
"context"
"fmt"
"time"
"github.com/jackc/pgx/v5"
)
// ChannelError is a channel that returned HTTP 404 on a discovery pass.
type ChannelError struct {
ChannelID string
ChannelName string
FirstSeen time.Time
LastSeen time.Time
}
// UpsertChannelError records (or refreshes) a 404 channel for the current user.
// Called by the runner inside a withUser scope; RLS guards user isolation.
func (s *Store) UpsertChannelError(ctx context.Context, userID, channelID, channelName string) error {
return s.withUser(ctx, userID, func(tx pgx.Tx) error {
_, err := tx.Exec(ctx, `
INSERT INTO channel_errors (user_id, channel_id, channel_name)
VALUES ($1, $2, $3)
ON CONFLICT (user_id, channel_id)
DO UPDATE SET channel_name = EXCLUDED.channel_name, last_seen = now()`,
userID, channelID, channelName)
if err != nil {
return fmt.Errorf("store: upsert channel error: %w", err)
}
return nil
})
}
// ListChannelErrors returns all 404-flagged channels for the user, newest first.
func (s *Store) ListChannelErrors(ctx context.Context, userID string) ([]ChannelError, error) {
var out []ChannelError
err := s.withUser(ctx, userID, func(tx pgx.Tx) error {
rows, err := tx.Query(ctx, `
SELECT channel_id, channel_name, first_seen, last_seen
FROM channel_errors
WHERE user_id = $1
ORDER BY last_seen DESC`, userID)
if err != nil {
return fmt.Errorf("store: list channel errors: %w", err)
}
defer rows.Close()
for rows.Next() {
var ce ChannelError
if err := rows.Scan(&ce.ChannelID, &ce.ChannelName, &ce.FirstSeen, &ce.LastSeen); err != nil {
return fmt.Errorf("store: scan channel error: %w", err)
}
out = append(out, ce)
}
return rows.Err()
})
return out, err
}
+35
View File
@@ -31,6 +31,41 @@ func (s *Store) UserBySubject(ctx context.Context, subject string) (userID strin
return userID, true, nil
}
// UserIdentity is one (userID, dexSubject) pair from the un-RLS'd
// user_identities map — the unit the scheduler enumerates to run a discovery
// pass per user (ADR-018).
type UserIdentity struct {
UserID string
DexSubject string
}
// ListAllUsers returns every (userID, dexSubject) pair from user_identities. It
// runs as a plain pool query WITHOUT withUser — intentional and legitimate:
// user_identities is un-RLS'd auth plumbing (like UserBySubject), and the
// scheduler enumerating all users to run their discovery passes is an admin
// operation that cannot be scoped to any single user. Order is unspecified.
func (s *Store) ListAllUsers(ctx context.Context) ([]UserIdentity, error) {
rows, err := s.pool.Query(ctx,
`SELECT user_id, dex_subject FROM user_identities`)
if err != nil {
return nil, fmt.Errorf("store: list all users: %w", err)
}
defer rows.Close()
var users []UserIdentity
for rows.Next() {
var u UserIdentity
if err := rows.Scan(&u.UserID, &u.DexSubject); err != nil {
return nil, fmt.Errorf("store: scan user identity: %w", err)
}
users = append(users, u)
}
if err := rows.Err(); err != nil {
return nil, fmt.Errorf("store: iterate user identities: %w", err)
}
return users, nil
}
// RegisterUser creates the tapir user for a Dex subject and the identity mapping
// that points to it, returning the new user_id. It errors with
// ErrSubjectRegistered if the subject already maps.
+33
View File
@@ -85,6 +85,39 @@ func TestRegisterUserRejectsDuplicateSubject(t *testing.T) {
require.Equal(t, first, got)
}
func TestListAllUsersReturnsEveryIdentity(t *testing.T) {
ctx := context.Background()
s := newStore(t)
resetDB(t, rawPool(t))
empty, err := s.ListAllUsers(ctx)
require.NoError(t, err)
require.Empty(t, empty, "no registrations yet → empty slice")
const subjectC = "dex|carol-789"
idA, err := s.RegisterUser(ctx, subjectA, "Alice")
require.NoError(t, err)
idB, err := s.RegisterUser(ctx, subjectB, "Bob")
require.NoError(t, err)
idC, err := s.RegisterUser(ctx, subjectC, "Carol")
require.NoError(t, err)
users, err := s.ListAllUsers(ctx)
require.NoError(t, err)
require.Len(t, users, 3)
// Order is unspecified; compare as a set of (userID, subject) pairs.
got := make(map[string]string, len(users))
for _, u := range users {
got[u.DexSubject] = u.UserID
}
require.Equal(t, map[string]string{
subjectA: idA,
subjectB: idB,
subjectC: idC,
}, got)
}
func TestRegisterUserDistinctSubjectsGetDistinctUsers(t *testing.T) {
ctx := context.Background()
s := newStore(t)
+41
View File
@@ -0,0 +1,41 @@
package store
import (
"context"
"fmt"
"github.com/jackc/pgx/v5"
)
// StampLogin records that the user was active today, throttled to one row per
// user per day. It is the read-side counterpart to SetAction: the middleware
// calls it on every authenticated request, but the append happens at most once a
// day so login_events stays small and the signal clean (one row = one active
// day, not one request).
//
// The check-and-insert is a single atomic statement: the INSERT ... SELECT ...
// WHERE NOT EXISTS only writes when no row for this user has seen_at in today
// (date_trunc('day', NOW()), server timezone). It runs through withUser, so the
// NOT EXISTS probe is itself RLS-scoped to the calling user via the
// tapir.current_user_id GUC — one user's stamp can never be suppressed or
// triggered by another user's rows. The explicit user_id predicate also keeps the
// probe on the (user_id, seen_at) index.
//
// A unique constraint is deliberately not used: under concurrent same-day
// requests the worst case is two rows for one day, which the gate query collapses
// to a single week bucket anyway — not worth a write-blocking constraint.
func (s *Store) StampLogin(ctx context.Context, userID string) error {
if err := s.withUser(ctx, userID, func(tx pgx.Tx) error {
_, err := tx.Exec(ctx,
`INSERT INTO login_events (user_id)
SELECT $1
WHERE NOT EXISTS (
SELECT 1 FROM login_events
WHERE user_id = $1 AND seen_at >= date_trunc('day', NOW())
)`, userID)
return err
}); err != nil {
return fmt.Errorf("store: stamp login: %w", err)
}
return nil
}
+83
View File
@@ -0,0 +1,83 @@
package store_test
import (
"context"
"testing"
"github.com/jackc/pgx/v5/pgxpool"
"github.com/stretchr/testify/require"
)
// countLoginEvents counts a user's login_events via the superuser pool, which
// bypasses RLS — so the assertion sees the true row count regardless of scope.
func countLoginEvents(t *testing.T, p *pgxpool.Pool, userID string) int {
t.Helper()
var n int
require.NoError(t, p.QueryRow(context.Background(),
`SELECT count(*) FROM login_events WHERE user_id = $1`, userID).Scan(&n))
return n
}
// TestStampLoginThrottlesToOnePerDay: repeated stamps within the same day insert
// exactly one row — the throttle that keeps login_events one-row-per-active-day.
func TestStampLoginThrottlesToOnePerDay(t *testing.T) {
ctx := context.Background()
s := newStore(t)
super := rawPool(t)
resetDB(t, super)
_, err := super.Exec(ctx, `INSERT INTO users (id) VALUES ($1)`, userA)
require.NoError(t, err)
for i := 0; i < 3; i++ {
require.NoError(t, s.StampLogin(ctx, userA))
}
require.Equal(t, 1, countLoginEvents(t, super, userA),
"three same-day stamps must collapse to one row")
}
// TestStampLoginRecordsOncePerNewDay: with yesterday's row already present, a
// stamp today is NOT throttled — it appends the day's row, so distinct active days
// accumulate (the substrate the gate's distinct-week count reads).
func TestStampLoginRecordsOncePerNewDay(t *testing.T) {
ctx := context.Background()
s := newStore(t)
super := rawPool(t)
resetDB(t, super)
_, err := super.Exec(ctx, `INSERT INTO users (id) VALUES ($1)`, userA)
require.NoError(t, err)
// Seed an event dated yesterday (before today's start), so the throttle's
// "row exists with seen_at >= start-of-today" probe finds nothing for today.
_, err = super.Exec(ctx,
`INSERT INTO login_events (user_id, seen_at) VALUES ($1, NOW() - INTERVAL '1 day')`, userA)
require.NoError(t, err)
require.NoError(t, s.StampLogin(ctx, userA))
require.Equal(t, 2, countLoginEvents(t, super, userA),
"a stamp on a new day must append a second row")
// A second stamp the same day is throttled again.
require.NoError(t, s.StampLogin(ctx, userA))
require.Equal(t, 2, countLoginEvents(t, super, userA),
"the same-day repeat must not add a third row")
}
// TestStampLoginIsUserScoped: one user's stamp lands only on that user's rows —
// the throttle probe is RLS-scoped, so user B's existing same-day row neither
// suppresses nor is touched by user A's stamp.
func TestStampLoginIsUserScoped(t *testing.T) {
ctx := context.Background()
s := newStore(t)
super := rawPool(t)
resetDB(t, super)
_, err := super.Exec(ctx, `INSERT INTO users (id) VALUES ($1), ($2)`, userA, userB)
require.NoError(t, err)
// B already has a same-day row; it must not throttle A's first stamp.
_, err = super.Exec(ctx, `INSERT INTO login_events (user_id) VALUES ($1)`, userB)
require.NoError(t, err)
require.NoError(t, s.StampLogin(ctx, userA))
require.Equal(t, 1, countLoginEvents(t, super, userA), "A's stamp must record despite B's same-day row")
require.Equal(t, 1, countLoginEvents(t, super, userB), "A's stamp must not touch B's rows")
}
+147
View File
@@ -0,0 +1,147 @@
package store_test
import (
"context"
"database/sql"
"os"
"testing"
"github.com/golang-migrate/migrate/v4"
migratepgx "github.com/golang-migrate/migrate/v4/database/pgx/v5"
"github.com/golang-migrate/migrate/v4/source/iofs"
"github.com/stretchr/testify/require"
_ "github.com/jackc/pgx/v5/stdlib" // register the "pgx" database/sql driver
)
// fileMigrator builds a golang-migrate instance from the on-disk migration files
// (not the embedded FS the production Migrate uses), so a test can step the schema
// up and down. os.DirFS(".") is rooted at the package dir; the SQL lives under
// "migrations". Mirrors store.Migrate's construction otherwise.
func fileMigrator(t *testing.T) *migrate.Migrate {
t.Helper()
db, err := sql.Open("pgx", dsn)
require.NoError(t, err)
t.Cleanup(func() { _ = db.Close() })
drv, err := migratepgx.WithInstance(db, &migratepgx.Config{})
require.NoError(t, err)
src, err := iofs.New(os.DirFS("."), "migrations")
require.NoError(t, err)
m, err := migrate.NewWithInstance("iofs", src, "pgx", drv)
require.NoError(t, err)
t.Cleanup(func() { _, _ = m.Close() })
return m
}
// loginEventsExists reports whether the login_events relation is present.
func loginEventsExists(t *testing.T) bool {
t.Helper()
var reg *string
require.NoError(t, rawPool(t).QueryRow(context.Background(),
`SELECT to_regclass('public.login_events')::text`).Scan(&reg))
return reg != nil
}
// TestMigration010LoginEventsUpDown proves migration 010 is reversible: the down
// migration drops login_events cleanly and the up migration recreates it. A rotten
// down migration (forgotten DROP, dangling policy) would fail here rather than in
// production during a rollback. The test restores the schema to latest before
// returning so the shared embedded-postgres stays at HEAD for sibling tests.
func TestMigration010LoginEventsUpDown(t *testing.T) {
newStore(t) // ensure the schema is migrated to latest (011 applied)
require.True(t, loginEventsExists(t), "login_events must exist at latest migration")
m := fileMigrator(t)
// 011..015 sit above 010; step them down first so 010 is exercised in isolation.
require.NoError(t, m.Steps(-1), "down 015 reshapes transcripts, login_events intact")
require.True(t, loginEventsExists(t), "015 down leaves login_events intact")
require.NoError(t, m.Steps(-1), "down 014 drops channel_title, login_events intact")
require.True(t, loginEventsExists(t), "014 down leaves login_events intact")
require.NoError(t, m.Steps(-1), "down 013 drops channel_errors, login_events intact")
require.True(t, loginEventsExists(t), "013 down leaves login_events intact")
require.NoError(t, m.Steps(-1), "down 012 is a no-op, login_events intact")
require.True(t, loginEventsExists(t), "012 down leaves login_events intact")
require.NoError(t, m.Steps(-1), "down 011 must not touch login_events")
require.True(t, loginEventsExists(t), "011 down leaves login_events intact")
require.NoError(t, m.Steps(-1), "down 010 must drop login_events")
require.False(t, loginEventsExists(t), "login_events must be gone after the down migration")
require.NoError(t, m.Steps(6), "up must recreate 010 then re-apply 011..015")
require.True(t, loginEventsExists(t), "login_events must be restored after the up migration")
}
// autoSummarizeDefault reads the users.auto_summarize column default as text
// ("true"/"false"), so the migration's default flip is verifiable directly.
func autoSummarizeDefault(t *testing.T) string {
t.Helper()
var def string
require.NoError(t, rawPool(t).QueryRow(context.Background(),
`SELECT column_default FROM information_schema.columns
WHERE table_name = 'users' AND column_name = 'auto_summarize'`).Scan(&def))
return def
}
// TestMigration011AutoSummarizeDefaultUpDown proves migration 011 is reversible:
// up sets the auto_summarize column default to TRUE (ADR-018), down restores
// FALSE. The down intentionally does not revert existing rows — only the default.
func TestMigration011AutoSummarizeDefaultUpDown(t *testing.T) {
newStore(t) // latest (013 applied)
require.Equal(t, "true", autoSummarizeDefault(t), "011 sets the default to TRUE")
m := fileMigrator(t)
require.NoError(t, m.Steps(-1), "down 015 reshapes transcripts")
require.NoError(t, m.Steps(-1), "down 014 drops channel_title")
require.NoError(t, m.Steps(-1), "down 013 drops channel_errors")
require.NoError(t, m.Steps(-1), "down 012 is a no-op")
require.NoError(t, m.Steps(-1), "down 011 reverts the column default")
require.Equal(t, "false", autoSummarizeDefault(t), "default is FALSE after the down migration")
require.NoError(t, m.Steps(1), "up 011 re-applies the TRUE default")
require.Equal(t, "true", autoSummarizeDefault(t))
require.NoError(t, m.Steps(1), "up 012 runs clean (no FORCE RLS on fresh schema)")
require.NoError(t, m.Steps(1), "up 013 creates channel_errors")
require.NoError(t, m.Steps(1), "up 014 recreates channel_title")
require.NoError(t, m.Steps(1), "up 015 reshapes transcripts to shared")
}
// channelTitleExists reports whether videos.channel_title is present.
func channelTitleExists(t *testing.T) bool {
t.Helper()
var exists bool
require.NoError(t, rawPool(t).QueryRow(context.Background(),
`SELECT EXISTS (SELECT 1 FROM information_schema.columns
WHERE table_name = 'videos' AND column_name = 'channel_title')`).Scan(&exists))
return exists
}
// TestMigration014VideoChannelTitleUpDown proves 014 is reversible: down drops
// videos.channel_title, up recreates it.
func TestMigration014VideoChannelTitleUpDown(t *testing.T) {
newStore(t) // latest (014 applied)
require.True(t, channelTitleExists(t), "channel_title exists at latest migration")
m := fileMigrator(t)
require.NoError(t, m.Steps(-1), "down 015 reshapes transcripts, channel_title intact")
require.True(t, channelTitleExists(t), "015 down leaves channel_title intact")
require.NoError(t, m.Steps(-1), "down 014 must drop channel_title")
require.False(t, channelTitleExists(t), "channel_title must be gone after the down migration")
require.NoError(t, m.Steps(1), "up 014 must recreate channel_title")
require.True(t, channelTitleExists(t), "channel_title must be restored after the up migration")
require.NoError(t, m.Steps(1), "up 015 restores the shared transcripts shape (HEAD)")
}
// TestMigration012FixAutoSummarizeRLS proves 012 runs cleanly and flips any
// remaining auto_summarize=FALSE rows to TRUE (the back-fill blocked by RLS in 011).
func TestMigration012FixAutoSummarizeRLS(t *testing.T) {
newStore(t) // apply all migrations including 012
require.Equal(t, "true", autoSummarizeDefault(t), "column default is TRUE after 012")
// Round-trip: down 012, then up 012 — must be idempotent.
m := fileMigrator(t)
require.NoError(t, m.Steps(-1), "down 012 must not error")
require.NoError(t, m.Steps(1), "up 012 must re-apply cleanly")
require.Equal(t, "true", autoSummarizeDefault(t), "default still TRUE after 012 re-applied")
}
@@ -0,0 +1,2 @@
ALTER TABLE videos DROP COLUMN IF EXISTS rate_limited_at;
ALTER TABLE videos DROP COLUMN IF EXISTS transcript_status;
@@ -0,0 +1,17 @@
-- Migration 007: per-video transcript fetch status, for rate-limit backoff.
--
-- transcript_status records the outcome of the last transcript attempt:
-- NULL = not yet attempted
-- 'none' = checked, no usable transcript (permanent — SourceNone)
-- 'fetched' = transcript resolved and summarized (summary_id not null)
-- 'rate_limited'= the caption endpoint returned 429; retry after a backoff window
--
-- rate_limited_at stamps WHEN the 429 was seen, so the runner can skip re-fetching
-- a still-throttled video until NOW() - rate_limited_at exceeds TAPIR_FETCH_BACKOFF.
-- It is cleared (set NULL) whenever the status moves off 'rate_limited'.
--
-- No RLS policy changes needed: videos already has ENABLE + FORCE ROW LEVEL
-- SECURITY (migration 003) with the videos_isolation policy. New columns inherit
-- that protection automatically.
ALTER TABLE videos ADD COLUMN transcript_status TEXT;
ALTER TABLE videos ADD COLUMN rate_limited_at TIMESTAMPTZ;
@@ -0,0 +1 @@
DROP TABLE IF EXISTS invitations;
@@ -0,0 +1,22 @@
-- Migration 009: invitations — an email-based invite to join Tapir (Stage-1
-- onboarding gate). Mathias mints one with `tapir invite <email>`; the recipient
-- visits /invite/{token}, sets a password, and Tapir creates their Dex account.
--
-- Deliberately NOT user-owned and NOT under RLS: an invitation exists BEFORE the
-- user does, so there is no user_id to scope by and no authenticated user context
-- when the invite is created (host CLI) or consumed (public /invite handler, no
-- Dex session). The token itself is the capability — a 32-byte crypto-random,
-- single-use, time-boxed secret. Hence no `user_id` FK and no ENABLE/FORCE ROW
-- LEVEL SECURITY here (unlike every user-owned table in migrations 003/005).
CREATE TABLE invitations (
id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
email TEXT NOT NULL,
token TEXT NOT NULL UNIQUE,
created_at TIMESTAMPTZ NOT NULL DEFAULT NOW(),
expires_at TIMESTAMPTZ NOT NULL,
used_at TIMESTAMPTZ
);
-- Lookups are by token (both the claim and the form preview); the UNIQUE
-- constraint already creates an index, this names one explicitly for clarity.
CREATE INDEX idx_invitations_token ON invitations(token);
@@ -0,0 +1,4 @@
DROP POLICY IF EXISTS login_events_isolation ON login_events;
ALTER TABLE login_events NO FORCE ROW LEVEL SECURITY;
ALTER TABLE login_events DISABLE ROW LEVEL SECURITY;
DROP TABLE IF EXISTS login_events;
@@ -0,0 +1,33 @@
-- Migration 010: login_events records THAT a user was active (returned and read)
-- on a given day — the Stage-0 signal summary_actions misses. summary_actions
-- captures *acts* (watch/skip/save); a reader who logs in weekly and clicks
-- nothing is otherwise invisible, yet for a reading product that return IS the
-- signal the gate ("usage in >=2 distinct weeks", VISION/ADR-016) is defined on.
--
-- Append-only: one row per user per active day (the request-path throttle in the
-- web layer enforces that cadence), never updated. Per-user isolation like every
-- user-owned table.
--
-- NO foreign key to users (mirrors summary_actions, migration 002): user_id is
-- carried for RLS/scoping but the table is decoupled so a stamp never blocks on a
-- users row. The cost of that decoupling: the users-row cascade does NOT reach
-- login_events, so DeleteUser must delete it explicitly (see account.go) — the
-- exact footgun the summary_actions delete work caught.
CREATE TABLE login_events (
id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
user_id UUID NOT NULL,
seen_at TIMESTAMPTZ NOT NULL DEFAULT NOW()
);
-- (user_id, seen_at) serves both the per-user-per-day throttle lookup
-- (seen_at >= start-of-today) and the gate query's per-user week bucketing.
CREATE INDEX idx_login_events_user_seen ON login_events(user_id, seen_at);
-- RLS: identical GUC-keyed policy/pattern to migration 003. FORCE so the table
-- owner (tapir, non-superuser in prod) is subject to it; an unset GUC yields NULL
-- → no rows match → deny-all.
ALTER TABLE login_events ENABLE ROW LEVEL SECURITY;
ALTER TABLE login_events FORCE ROW LEVEL SECURITY;
CREATE POLICY login_events_isolation ON login_events
FOR ALL
USING (user_id = current_setting('tapir.current_user_id', true)::uuid);
@@ -0,0 +1,5 @@
-- Revert the column default to FALSE. Existing rows are intentionally NOT
-- reverted: flipping live users back to manual on a rollback would be a
-- surprising regression (they may have come to rely on auto). The default change
-- is the reversible part; data stays as the user left it.
ALTER TABLE users ALTER COLUMN auto_summarize SET DEFAULT FALSE;
@@ -0,0 +1,17 @@
-- Migration 011: flip auto_summarize default to TRUE (ADR-018, Future-B).
--
-- Scheduled discovery (ADR-018) makes Tapir watch unprompted. For onboarded
-- friends that only delivers zero-friction value if the list also SUMMARIZES
-- itself — a manual default would mean the scheduler discovers videos a user
-- still has to click through one by one, which is the empty-list problem again.
-- So new users default to AUTO. The account-page toggle still lets a user switch
-- to manual (SetAutoSummarize), so this only changes the out-of-the-box state.
--
-- Safe only because the process-wide caption-fetch rate gate (ADR-014 item 2)
-- now exists: auto + scheduled + multi-user would otherwise self-inflict 429s
-- every cycle. The gate is the precondition for shipping this default.
ALTER TABLE users ALTER COLUMN auto_summarize SET DEFAULT TRUE;
-- Bring existing rows (maintainer + any current registrations) onto the new
-- default so they benefit immediately, not just users created after this point.
UPDATE users SET auto_summarize = TRUE WHERE auto_summarize = FALSE;
@@ -0,0 +1 @@
-- No data revert: do not flip users back to manual on rollback.
@@ -0,0 +1,6 @@
-- Migration 011's UPDATE ran without tapir.current_user_id set, so FORCE RLS
-- blocked all rows and zero users were updated. Temporarily drop FORCE so the
-- table owner (tapir role) can bypass RLS for this back-fill, then restore it.
ALTER TABLE users NO FORCE ROW LEVEL SECURITY;
UPDATE users SET auto_summarize = TRUE WHERE auto_summarize = FALSE;
ALTER TABLE users FORCE ROW LEVEL SECURITY;
@@ -0,0 +1 @@
DROP TABLE IF EXISTS channel_errors;
@@ -0,0 +1,22 @@
-- channel_errors: channels that returned HTTP 404 (deleted/private) on the most
-- recent scheduler pass. Surfaced on the account page so users know why some
-- subscribed channels produce no videos. Upserted per-pass; cleared when the
-- channel starts returning results again (runner calls UpsertChannelError only
-- on 404, so a recovered channel simply stops appearing after its row ages out
-- or the user takes action). ON DELETE CASCADE keeps rows tidy on account deletion.
CREATE TABLE channel_errors (
user_id UUID NOT NULL REFERENCES users(id) ON DELETE CASCADE,
channel_id TEXT NOT NULL,
channel_name TEXT NOT NULL,
first_seen TIMESTAMPTZ NOT NULL DEFAULT now(),
last_seen TIMESTAMPTZ NOT NULL DEFAULT now(),
PRIMARY KEY (user_id, channel_id)
);
CREATE INDEX idx_channel_errors_user_id ON channel_errors(user_id);
ALTER TABLE channel_errors ENABLE ROW LEVEL SECURITY;
ALTER TABLE channel_errors FORCE ROW LEVEL SECURITY;
CREATE POLICY channel_errors_isolation ON channel_errors
FOR ALL
USING (user_id = current_setting('tapir.current_user_id', true)::uuid);
@@ -0,0 +1 @@
ALTER TABLE videos DROP COLUMN channel_title;
@@ -0,0 +1,8 @@
-- Store the source channel's title per video so the list can offer a real
-- channel filter (multi-select of the user's channels) instead of the dead
-- free-text field that only ever matched the provider string. Nullable: existing
-- rows backfill on the next discovery pass (UpsertVideo writes it); pasted videos
-- get it immediately from videos.list. No FK to a channels table at Stage 0 — the
-- title is a denormalised display/filter value, consistent with the existing
-- subscription_id-stays-NULL stance (data-model.md).
ALTER TABLE videos ADD COLUMN channel_title TEXT;
@@ -0,0 +1,19 @@
-- Down 015: restore the per-user RLS-scoped transcripts shape (001 + 003).
DROP TABLE transcripts;
CREATE TABLE transcripts (
video_id UUID PRIMARY KEY REFERENCES videos(id) ON DELETE CASCADE,
user_id UUID NOT NULL REFERENCES users(id) ON DELETE CASCADE,
source TEXT NOT NULL,
language TEXT,
content TEXT,
resolved_at TIMESTAMPTZ NOT NULL DEFAULT NOW()
);
CREATE INDEX idx_transcripts_user_id ON transcripts(user_id);
ALTER TABLE transcripts ENABLE ROW LEVEL SECURITY;
ALTER TABLE transcripts FORCE ROW LEVEL SECURITY;
CREATE POLICY transcripts_isolation ON transcripts
FOR ALL
USING (user_id = current_setting('tapir.current_user_id', true)::uuid);
@@ -0,0 +1,30 @@
-- Migration 015: transcripts become SHARED public-content storage (ADR-021).
--
-- The per-user transcripts table from 001 (PK videos.id, user_id NOT NULL, RLS
-- FORCEd in 003) was dead: no application code ever read or wrote it — only the
-- transcript_status columns on `videos` (007) carried fetch outcomes. ADR-021
-- repurposes it as the single shared store of public caption content, keyed by
-- the cross-user dedup key (provider, provider_video_id) — the video's public
-- identity, not Tapir's per-user videos.id — so re-analysis never re-fetches
-- from YouTube (ADR-010/014).
--
-- It holds ONLY public caption content + the video's public id (nothing
-- user-identifying), so it is deliberately NOT RLS-scoped: no user_id, no
-- policy, no FORCE. This is the single, intentional exception to the ADR-012
-- isolation boundary; rls_test.go asserts the boundary is exactly here and
-- nowhere else. Dropping the old table drops its RLS policy with it; it held no
-- real data, so drop+recreate loses nothing.
DROP TABLE transcripts;
CREATE TABLE transcripts (
provider TEXT NOT NULL,
provider_video_id TEXT NOT NULL,
source TEXT NOT NULL, -- 'captions' (content set) | 'none' (no captions; content NULL)
language TEXT,
content TEXT,
fetched_at TIMESTAMPTZ NOT NULL DEFAULT NOW(),
PRIMARY KEY (provider, provider_video_id)
);
COMMENT ON TABLE transcripts IS
'Shared public caption content keyed by (provider, provider_video_id). NOT RLS-scoped — public content only, de-facto cross-user dedup (ADR-021).';
+17 -5
View File
@@ -29,6 +29,7 @@ type SummaryRow struct {
ProviderVideoID string // videos.provider_video_id; empty when no videos row
Title string // videos.title; empty when no videos row
Channel string // videos.provider for now; empty when no videos row
ChannelTitle string // videos.channel_title; the source channel, for display + filtering
URL string // videos.url; empty when no videos row
PublishedAt time.Time // videos.published_at; zero when absent
Summary string
@@ -49,6 +50,10 @@ type SummaryRow struct {
// set by the web "Summarize" button and cleared by the next `tapir run`. Only
// populated by ListVideos/GetVideoRow (summary-only reads leave it false).
SummarizeRequested bool
// TranscriptStatus mirrors videos.transcript_status (migration 007): "" (unset),
// "none", "rate_limited", or "fetched". Drives the "Retrying later" list badge.
// Only populated by ListVideos/GetVideoRow ("" on summary-only reads).
TranscriptStatus string
}
// selectSummary is the shared projection for both reads. videos is LEFT JOINed
@@ -132,24 +137,29 @@ const selectVideo = `
COALESCE(s.fallback_used, FALSE),
COALESCE(s.created_at, v.seen_at),
(s.id IS NOT NULL) AS summarized,
v.summarize_requested
v.summarize_requested,
COALESCE(v.transcript_status, ''),
COALESCE(v.channel_title, '')
FROM videos v
LEFT JOIN summaries s ON s.video_id = v.id AND s.user_id = v.user_id`
// ListVideos returns ALL of the user's videos — summarized and not — most recent
// first by seen_at, capped at limit (non-positive defaults to 50). Unsummarized
// ListVideos returns ALL of the user's videos — summarized first, then by
// published_at DESC with undated videos last, then seen_at DESC as a tiebreak —
// capped at limit (non-positive defaults to 500). The published_at ordering
// aligns the list with the recency framing (newest content first); seen_at
// breaks ties and orders same/!undated rows deterministically. Unsummarized
// videos come back with Summarized=false and empty summary fields, so the list
// view can render them with a "Summarize" affordance. Scoped by user_id.
func (s *Store) ListVideos(ctx context.Context, userID string, limit int) ([]SummaryRow, error) {
if limit <= 0 {
limit = 50
limit = 500
}
var out []SummaryRow
if err := s.withUser(ctx, userID, func(tx pgx.Tx) error {
rows, err := tx.Query(ctx,
selectVideo+`
WHERE v.user_id = $1
ORDER BY v.seen_at DESC
ORDER BY (s.id IS NOT NULL) DESC, v.published_at DESC NULLS LAST, v.seen_at DESC
LIMIT $2`,
userID, limit)
if err != nil {
@@ -245,6 +255,8 @@ func scanVideoRow(rows pgx.Row) (SummaryRow, error) {
&row.CreatedAt,
&row.Summarized,
&row.SummarizeRequested,
&row.TranscriptStatus,
&row.ChannelTitle,
); err != nil {
return SummaryRow{}, fmt.Errorf("store: scan video: %w", err)
}
+111
View File
@@ -0,0 +1,111 @@
package store
import (
"context"
"fmt"
"sort"
"github.com/jackc/pgx/v5"
)
// UserActiveWeeks is one row of the Stage-0 gate report: how many DISTINCT
// calendar weeks a user was active in, counting reads (login_events) AND acts
// (summary_actions) together. The gate (VISION/ADR-016) passes when any user
// reaches ActiveWeeks >= 2.
type UserActiveWeeks struct {
UserID string
DisplayName string
ActiveWeeks int
}
// ActiveWeeks computes per-user distinct-active-weeks for the gate report, most
// active first.
//
// Why per-user iteration rather than one cross-user GROUP BY: the user-owned
// tables are FORCE RLS (migration 003/010) and the production role is a non-
// superuser owner, so a single un-scoped query sees nothing (deny-all). Instead we
// enumerate users from the deliberately un-RLS'd identity map (user_identities,
// migration 004) and count each user's weeks inside withUser, where the GUC scopes
// login_events + summary_actions to that user. No privilege escalation, no policy
// change — the same isolation seam every other read flows through.
//
// Scope note: the enumeration covers users with a Dex identity (the web users the
// gate is about). A CLI-only user created by the store sink without an identity
// row would not appear — out of scope for this gate.
func (s *Store) ActiveWeeks(ctx context.Context) ([]UserActiveWeeks, error) {
userIDs, err := s.identityUserIDs(ctx)
if err != nil {
return nil, err
}
out := make([]UserActiveWeeks, 0, len(userIDs))
for _, uid := range userIDs {
row, err := s.activeWeeksFor(ctx, uid)
if err != nil {
return nil, err
}
out = append(out, row)
}
// Most active first; user_id as a stable tie-break for deterministic output.
sort.SliceStable(out, func(i, j int) bool {
if out[i].ActiveWeeks != out[j].ActiveWeeks {
return out[i].ActiveWeeks > out[j].ActiveWeeks
}
return out[i].UserID < out[j].UserID
})
return out, nil
}
// identityUserIDs lists every tapir user_id from the un-RLS'd identity map. It
// runs directly on the pool (no withUser): user_identities carries no user data
// and is intentionally not RLS-enabled, so it is the one table that can be read
// pre-scope to discover who exists.
func (s *Store) identityUserIDs(ctx context.Context) ([]string, error) {
rows, err := s.pool.Query(ctx, `SELECT user_id FROM user_identities ORDER BY user_id`)
if err != nil {
return nil, fmt.Errorf("store: list identity users: %w", err)
}
defer rows.Close()
var ids []string
for rows.Next() {
var id string
if err := rows.Scan(&id); err != nil {
return nil, fmt.Errorf("store: scan identity user: %w", err)
}
ids = append(ids, id)
}
if err := rows.Err(); err != nil {
return nil, fmt.Errorf("store: iterate identity users: %w", err)
}
return ids, nil
}
// activeWeeksFor counts one user's distinct active weeks (reads UNION acts) and
// reads their display name, RLS-scoped via withUser. The UNION dedups a week that
// has both a login and an action so it counts once.
func (s *Store) activeWeeksFor(ctx context.Context, userID string) (UserActiveWeeks, error) {
res := UserActiveWeeks{UserID: userID}
if err := s.withUser(ctx, userID, func(tx pgx.Tx) error {
if err := tx.QueryRow(ctx,
`WITH weeks AS (
SELECT date_trunc('week', seen_at) AS wk
FROM login_events WHERE user_id = $1
UNION
SELECT date_trunc('week', acted_at)
FROM summary_actions WHERE user_id = $1
)
SELECT count(DISTINCT wk) FROM weeks`, userID).Scan(&res.ActiveWeeks); err != nil {
return fmt.Errorf("store: count active weeks: %w", err)
}
if err := tx.QueryRow(ctx,
`SELECT COALESCE(display_name, '') FROM users WHERE id = $1`, userID).Scan(&res.DisplayName); err != nil {
return fmt.Errorf("store: read display name: %w", err)
}
return nil
}); err != nil {
return UserActiveWeeks{}, err
}
return res, nil
}
+78
View File
@@ -0,0 +1,78 @@
package store_test
import (
"context"
"testing"
"github.com/jackc/pgx/v5/pgxpool"
"github.com/stretchr/testify/require"
)
// seedReportUser inserts a user + its identity mapping (the enumeration source
// ActiveWeeks reads). display_name is optional.
func seedReportUser(t *testing.T, p *pgxpool.Pool, userID, subject, name string) {
t.Helper()
ctx := context.Background()
_, err := p.Exec(ctx,
`INSERT INTO users (id, display_name) VALUES ($1, NULLIF($2, ''))`, userID, name)
require.NoError(t, err)
_, err = p.Exec(ctx,
`INSERT INTO user_identities (dex_subject, user_id) VALUES ($1, $2)`, subject, userID)
require.NoError(t, err)
}
// TestActiveWeeksCountsDistinctWeeksAcrossReadsAndActs is the gate-query proof.
// It seeds, with fixed timestamps in known ISO weeks:
// - user A: reads in week of Jan 5 and Jan 12, acts in week of Jan 12 (dup) and
// Jan 19 → the UNION across both tables collapses the shared week → 3 distinct.
// - user B: a single read in the week of Jan 5 → 1 distinct (below the gate).
//
// It verifies the count is correct, dedups the cross-table shared week, and orders
// most-active first.
func TestActiveWeeksCountsDistinctWeeksAcrossReadsAndActs(t *testing.T) {
ctx := context.Background()
s := newStore(t)
p := rawPool(t)
resetDB(t, p) // TRUNCATE ... users CASCADE also clears user_identities
seedReportUser(t, p, userA, "subject-a", "Ada")
seedReportUser(t, p, userB, "subject-b", "")
// Reads (login_events) — fixed dates in distinct ISO weeks.
_, err := p.Exec(ctx,
`INSERT INTO login_events (user_id, seen_at) VALUES
($1, '2026-01-05T09:00:00Z'),
($1, '2026-01-12T09:00:00Z'),
($2, '2026-01-05T09:00:00Z')`, userA, userB)
require.NoError(t, err)
// Acts (summary_actions) — one in A's week-of-Jan-12 (shared with a read, must
// dedup) and one in a new week (Jan 19).
_, err = p.Exec(ctx,
`INSERT INTO summary_actions (user_id, video_id, action, acted_at) VALUES
($1, 'vid-1', 'watched', '2026-01-12T18:00:00Z'),
($1, 'vid-2', 'saved', '2026-01-19T18:00:00Z')`, userA)
require.NoError(t, err)
got, err := s.ActiveWeeks(ctx)
require.NoError(t, err)
require.Len(t, got, 2, "both identity users must appear")
require.Equal(t, userA, got[0].UserID, "most-active user first")
require.Equal(t, "Ada", got[0].DisplayName)
require.Equal(t, 3, got[0].ActiveWeeks, "3 distinct weeks across reads+acts, shared week deduped")
require.Equal(t, userB, got[1].UserID)
require.Equal(t, 1, got[1].ActiveWeeks, "single read = 1 distinct week (below gate)")
}
// TestActiveWeeksEmptyWhenNoUsers: no identities → no rows (not an error).
func TestActiveWeeksEmptyWhenNoUsers(t *testing.T) {
ctx := context.Background()
s := newStore(t)
resetDB(t, rawPool(t))
got, err := s.ActiveWeeks(ctx)
require.NoError(t, err)
require.Empty(t, got)
}
+80 -15
View File
@@ -22,9 +22,11 @@ import (
// (no GUC set → zero rows) proves the enforcement path is live, not bypassed.
// userIsolatedTables are the tables that carry a user_id and whose policy keys
// directly off the tapir.current_user_id GUC.
// directly off the tapir.current_user_id GUC. transcripts is deliberately ABSENT
// — ADR-021 made it shared public content (non-RLS); TestTranscriptsTableIsSharedNotRLS
// proves that is the only place the isolation boundary moved.
var userIsolatedTables = []string{
"users", "videos", "transcripts", "summaries", "summary_actions", "video_connections",
"users", "videos", "summaries", "summary_actions", "login_events", "video_connections",
}
// allIsolatedTables adds sink_deliveries, whose ownership is derived from its
@@ -38,9 +40,10 @@ type seeded struct {
summaryID string
}
// seedUser inserts one full chain (user → video → transcript → summary →
// action → delivery) as the superuser pool, which bypasses RLS so both users'
// data lands regardless of the GUC.
// seedUser inserts one full chain (user → video → summary → action → delivery)
// as the superuser pool, which bypasses RLS so both users' data lands regardless
// of the GUC. Transcripts are NOT seeded here: they are shared, non-RLS public
// content (ADR-021), so they have no place in a per-user isolation chain.
func seedUser(t *testing.T, p *pgxpool.Pool, userID string) seeded {
t.Helper()
ctx := context.Background()
@@ -54,11 +57,6 @@ func seedUser(t *testing.T, p *pgxpool.Pool, userID string) seeded {
VALUES ($1, 'youtube', $2, 'title') RETURNING id`,
userID, "vid-"+userID).Scan(&videoID))
_, err = p.Exec(ctx,
`INSERT INTO transcripts (video_id, user_id, source, content)
VALUES ($1, $2, 'captions', 'words')`, videoID, userID)
require.NoError(t, err)
var summaryID string
require.NoError(t, p.QueryRow(ctx,
`INSERT INTO summaries (user_id, video_id, summary) VALUES ($1, $2, 'sum')
@@ -69,6 +67,10 @@ func seedUser(t *testing.T, p *pgxpool.Pool, userID string) seeded {
VALUES ($1, $2, 'watched')`, userID, videoID)
require.NoError(t, err)
_, err = p.Exec(ctx,
`INSERT INTO login_events (user_id) VALUES ($1)`, userID)
require.NoError(t, err)
_, err = p.Exec(ctx,
`INSERT INTO sink_deliveries (summary_id, sink, status)
VALUES ($1, 'store', 'delivered')`, summaryID)
@@ -88,9 +90,16 @@ func appPool(t *testing.T, super *pgxpool.Pool) *pgxpool.Pool {
t.Helper()
ctx := context.Background()
// Idempotent across test runs (schema/role persist for the TestMain PG).
_, _ = super.Exec(ctx, `DROP ROLE IF EXISTS app`)
_, err := super.Exec(ctx, `CREATE ROLE app LOGIN PASSWORD 'app'`)
// Idempotent across tests AND runs: the role persists for the TestMain PG and
// owns granted privileges, so a plain DROP ROLE fails once any GRANT exists
// (and more than one test now builds an app pool). Create only if absent; the
// GRANTs below are themselves idempotent.
_, err := super.Exec(ctx,
`DO $$ BEGIN
IF NOT EXISTS (SELECT FROM pg_roles WHERE rolname = 'app') THEN
CREATE ROLE app LOGIN PASSWORD 'app';
END IF;
END $$`)
require.NoError(t, err)
_, err = super.Exec(ctx, `GRANT USAGE ON SCHEMA public TO app`)
require.NoError(t, err)
@@ -182,13 +191,14 @@ func TestRLSEnforcesPerUserIsolation(t *testing.T) {
{"update users", `UPDATE users SET display_name = 'hacked' WHERE id = $1`, b.userID},
{"update videos", `UPDATE videos SET title = 'hacked' WHERE user_id = $1`, b.userID},
{"queue videos summarize", `UPDATE videos SET summarize_requested = TRUE WHERE id = $1`, b.videoID},
{"update transcripts", `UPDATE transcripts SET content = 'hacked' WHERE user_id = $1`, b.userID},
{"update summaries", `UPDATE summaries SET summary = 'hacked' WHERE user_id = $1`, b.userID},
{"update summary_actions", `UPDATE summary_actions SET action = 'skipped' WHERE user_id = $1`, b.userID},
{"update login_events", `UPDATE login_events SET seen_at = NOW() WHERE user_id = $1`, b.userID},
{"update sink_deliveries", `UPDATE sink_deliveries SET status = 'hacked' WHERE summary_id = $1`, b.summaryID},
{"update video_connections", `UPDATE video_connections SET token_ref = 'hacked' WHERE user_id = $1`, b.userID},
{"delete summaries", `DELETE FROM summaries WHERE user_id = $1`, b.userID},
{"delete summary_actions", `DELETE FROM summary_actions WHERE user_id = $1`, b.userID},
{"delete login_events", `DELETE FROM login_events WHERE user_id = $1`, b.userID},
{"delete sink_deliveries", `DELETE FROM sink_deliveries WHERE summary_id = $1`, b.summaryID},
{"delete video_connections", `DELETE FROM video_connections WHERE user_id = $1`, b.userID},
}
@@ -205,17 +215,20 @@ func TestRLSEnforcesPerUserIsolation(t *testing.T) {
`SELECT summary FROM summaries WHERE user_id = $1`, b.userID).Scan(&bSummary))
require.Equal(t, "sum", bSummary, "B's summary must be untouched by A's writes")
var bSummaries, bActions, bDeliveries, bConnections int
var bSummaries, bActions, bLogins, bDeliveries, bConnections int
require.NoError(t, super.QueryRow(ctx,
`SELECT count(*) FROM summaries WHERE user_id = $1`, b.userID).Scan(&bSummaries))
require.NoError(t, super.QueryRow(ctx,
`SELECT count(*) FROM summary_actions WHERE user_id = $1`, b.userID).Scan(&bActions))
require.NoError(t, super.QueryRow(ctx,
`SELECT count(*) FROM login_events WHERE user_id = $1`, b.userID).Scan(&bLogins))
require.NoError(t, super.QueryRow(ctx,
fmt.Sprintf(`SELECT count(*) FROM sink_deliveries WHERE summary_id = '%s'`, b.summaryID)).Scan(&bDeliveries))
require.NoError(t, super.QueryRow(ctx,
`SELECT count(*) FROM video_connections WHERE user_id = $1 AND token_ref <> 'hacked'`, b.userID).Scan(&bConnections))
require.Equal(t, 1, bSummaries, "A's DELETE must not have removed B's summary")
require.Equal(t, 1, bActions, "A's DELETE must not have removed B's action")
require.Equal(t, 1, bLogins, "A's DELETE must not have removed B's login event")
require.Equal(t, 1, bDeliveries, "A's DELETE must not have removed B's delivery")
require.Equal(t, 1, bConnections, "A's writes must not have touched B's connection")
@@ -226,3 +239,55 @@ func TestRLSEnforcesPerUserIsolation(t *testing.T) {
_ = a // a's ids are seeded for the symmetric read assertions above
}
// TestTranscriptsTableIsSharedNotRLS is the ADR-021 isolation proof: transcripts
// is the ONE shared, non-RLS surface, and the public-content classification
// leaked to nothing else. It is the inverse of TestRLSEnforcesPerUserIsolation —
// where that asserts deny-all on every user-owned table, this asserts transcripts
// is readable and writable with no user scope at all, holds no user_id, and is
// the single table with row-level security switched off.
func TestTranscriptsTableIsSharedNotRLS(t *testing.T) {
newStore(t)
super := rawPool(t)
resetDB(t, super)
app := appPool(t, super)
ctx := context.Background()
// 1. Shared + non-RLS: with NO GUC set, the app role both writes and reads a
// transcript. On an RLS table this would be deny-all (zero rows), exactly as
// the main isolation test asserts for every user-owned table.
_, err := app.Exec(ctx,
`INSERT INTO transcripts (provider, provider_video_id, source, content)
VALUES ('youtube', 'shared-vid', 'captions', 'public words')`)
require.NoError(t, err, "app role must write shared transcript content with no user scope")
require.Equal(t, 1, scopedCount(t, app, "", "transcripts"),
"transcripts must be readable with NO user scope — it is shared, non-RLS (ADR-021)")
// 2. No user_id column: the table holds only public caption content + the
// video's public id, nothing user-identifying.
var hasUserID bool
require.NoError(t, super.QueryRow(ctx,
`SELECT EXISTS (SELECT 1 FROM information_schema.columns
WHERE table_name = 'transcripts' AND column_name = 'user_id')`).Scan(&hasUserID))
require.False(t, hasUserID, "transcripts must carry no user_id (ADR-021 public content)")
// 3. The boundary is EXACTLY here: every user-owned table still has row-level
// security enabled; transcripts alone has it off. This is the proof the
// non-RLS classification was applied to transcripts and leaked nowhere else.
for _, table := range allIsolatedTables {
require.True(t, rlsEnabled(t, super, table),
"%s must still enforce row-level security — isolation must not have regressed", table)
}
require.False(t, rlsEnabled(t, super, "transcripts"),
"transcripts must be the single table with row-level security OFF (the one shared surface)")
}
// rlsEnabled reports whether a public table has ROW LEVEL SECURITY enabled.
func rlsEnabled(t *testing.T, p *pgxpool.Pool, table string) bool {
t.Helper()
var enabled bool
require.NoError(t, p.QueryRow(context.Background(),
`SELECT relrowsecurity FROM pg_class
WHERE relname = $1 AND relnamespace = 'public'::regnamespace`, table).Scan(&enabled))
return enabled
}
+1 -1
View File
@@ -74,7 +74,7 @@ func rawPool(t *testing.T) *pgxpool.Pool {
func resetDB(t *testing.T, p *pgxpool.Pool) {
t.Helper()
_, err := p.Exec(context.Background(),
`TRUNCATE summary_actions, sink_deliveries, summaries, transcripts, videos, users CASCADE`)
`TRUNCATE login_events, summary_actions, sink_deliveries, summaries, transcripts, videos, users CASCADE`)
require.NoError(t, err)
}
@@ -3,6 +3,7 @@ package store_test
import (
"context"
"testing"
"time"
"github.com/jackc/pgx/v5/pgxpool"
"github.com/stretchr/testify/require"
@@ -45,6 +46,27 @@ func TestAutoSummarizeRoundTripDefaultsFalse(t *testing.T) {
require.False(t, got, "set back to manual round-trips")
}
func TestRegisteredUserDefaultsAutoSummarizeOn(t *testing.T) {
ctx := context.Background()
s := newStore(t)
resetDB(t, rawPool(t))
// A real registration creates the users row, so the column default (migration
// 011: TRUE) drives the mode — onboarded friends get auto out of the box.
id, err := s.RegisterUser(ctx, subjectA, "Alice")
require.NoError(t, err)
got, err := s.GetAutoSummarize(ctx, id)
require.NoError(t, err)
require.True(t, got, "new registrations default to auto-summarize (ADR-018)")
// The account-page toggle still works: a user can switch to manual.
require.NoError(t, s.SetAutoSummarize(ctx, id, false))
got, err = s.GetAutoSummarize(ctx, id)
require.NoError(t, err)
require.False(t, got, "the manual toggle still flips it off")
}
func TestRequestSummarizeSetsFlag(t *testing.T) {
ctx := context.Background()
s := newStore(t)
@@ -104,6 +126,39 @@ func TestListVideosReturnsSummarizedAndUnsummarized(t *testing.T) {
require.True(t, byID[videoY].SummarizeRequested, "queued video carries the flag")
}
// TestListVideosOrderedByPublishedDescNullsLast: summarized videos sort first
// (regardless of their date), then unsummarized by published_at DESC with
// undated (NULL) videos last — the recency-aligned list order (UX review B2).
func TestListVideosOrderedByPublishedDescNullsLast(t *testing.T) {
ctx := context.Background()
s := newStore(t)
p := rawPool(t)
resetDB(t, p)
const (
vSummOld = "cccccccc-cccc-cccc-cccc-cccccccccccc" // summarized, oldest date
vNewer = "dddddddd-dddd-dddd-dddd-dddddddddddd" // unsummarized, newest
vOlder = "eeeeeeee-eeee-eeee-eeee-eeeeeeeeeeee" // unsummarized, older
vUndated = "ffffffff-ffff-ffff-ffff-ffffffffffff" // unsummarized, no date
)
// Deliver first so the userA row exists (the videos FK needs it); the
// summarized video carries the OLDEST date yet must still sort first because
// it is summarized, proving summarized-first dominates the date sort.
require.NoError(t, s.Deliver(ctx, summary(userA, vSummOld, "body")))
seedVideo(t, p, userA, vSummOld, "Summarized Old", "youtube", "https://s", time.Date(2025, 1, 1, 0, 0, 0, 0, time.UTC))
seedVideo(t, p, userA, vNewer, "Newer", "youtube", "https://n", time.Date(2026, 3, 1, 0, 0, 0, 0, time.UTC))
seedVideo(t, p, userA, vOlder, "Older", "youtube", "https://o", time.Date(2026, 1, 1, 0, 0, 0, 0, time.UTC))
seedVideo(t, p, userA, vUndated, "Undated", "youtube", "https://u", time.Time{})
rows, err := s.ListVideos(ctx, userA, 50)
require.NoError(t, err)
require.Len(t, rows, 4)
got := []string{rows[0].VideoID, rows[1].VideoID, rows[2].VideoID, rows[3].VideoID}
require.Equal(t, []string{vSummOld, vNewer, vOlder, vUndated}, got,
"summarized first, then published_at DESC, NULL dates last")
}
func TestListVideosIsUserScoped(t *testing.T) {
ctx := context.Background()
s := newStore(t)
+66
View File
@@ -0,0 +1,66 @@
package store
import (
"context"
"errors"
"fmt"
"github.com/jackc/pgx/v5"
"gitea.d-ma.be/mathias/tapir/internal/domain"
)
// GetTranscript returns the shared, stored transcript for a video keyed by the
// cross-user dedup key (provider, providerVideoID), and whether one exists
// (ADR-021). It reads via the raw pool, NOT withUser: the table holds public
// content with no user_id and no RLS policy, so it is shared across users by
// construction. A stored SourceNone is a real hit (ok == true, HasText() ==
// false) — a known caption-less video, so the caller skips without re-fetching.
func (s *Store) GetTranscript(ctx context.Context, provider, providerVideoID string) (domain.Transcript, bool, error) {
var source, lang, content string
err := s.pool.QueryRow(ctx,
`SELECT source, COALESCE(language, ''), COALESCE(content, '')
FROM transcripts WHERE provider = $1 AND provider_video_id = $2`,
provider, providerVideoID).Scan(&source, &lang, &content)
if errors.Is(err, pgx.ErrNoRows) {
return domain.Transcript{}, false, nil
}
if err != nil {
return domain.Transcript{}, false, fmt.Errorf("store: get transcript: %w", err)
}
return domain.Transcript{
Source: domain.TranscriptSource(source),
Language: lang,
Content: content,
}, true, nil
}
// SaveTranscript upserts the shared transcript for (provider, providerVideoID).
// Only terminal outcomes belong here: SourceCaptions (with text) or SourceNone
// (no captions). A transient SourceRateLimited is rejected so persistence never
// masks a 429 as a permanent absence — that stays a per-user retry (ADR-014).
// Last write wins on conflict (a later re-fetch may correct an entry). It writes
// via the raw pool, NOT withUser — public content, shared, non-RLS (ADR-021).
func (s *Store) SaveTranscript(ctx context.Context, provider, providerVideoID string, t domain.Transcript) error {
switch t.Source {
case domain.SourceCaptions, domain.SourceNone:
// terminal — persist
case domain.SourceRateLimited:
return fmt.Errorf("store: refusing to persist transient rate-limited transcript for %s/%s", provider, providerVideoID)
default:
return fmt.Errorf("store: invalid transcript source %q", t.Source)
}
_, err := s.pool.Exec(ctx,
`INSERT INTO transcripts (provider, provider_video_id, source, language, content)
VALUES ($1, $2, $3, NULLIF($4, ''), NULLIF($5, ''))
ON CONFLICT (provider, provider_video_id)
DO UPDATE SET source = EXCLUDED.source,
language = EXCLUDED.language,
content = EXCLUDED.content,
fetched_at = NOW()`,
provider, providerVideoID, string(t.Source), t.Language, t.Content)
if err != nil {
return fmt.Errorf("store: save transcript: %w", err)
}
return nil
}
@@ -0,0 +1,102 @@
package store
import (
"context"
"errors"
"fmt"
"time"
"github.com/jackc/pgx/v5"
)
// validTranscriptStatuses bounds SetTranscriptStatus input. "" clears the status
// (column NULL); the three named states mirror migration 007's documented values.
var validTranscriptStatuses = map[string]bool{
"": true,
"none": true,
"rate_limited": true,
"fetched": true,
}
// SetTranscriptStatus records the outcome of the last transcript attempt for a
// video (migration 007). When status is "rate_limited" it also stamps
// rate_limited_at = NOW() so the runner can back off; every other status clears
// that timestamp. "" unsets the status (column NULL). An unknown status is
// rejected. Scoped via withUser, so RLS confines the UPDATE to the caller's own
// video; ErrNotFound when the user has no such video.
func (s *Store) SetTranscriptStatus(ctx context.Context, userID, videoID, status string) error {
if !validTranscriptStatuses[status] {
return fmt.Errorf("store: invalid transcript status %q", status)
}
return s.withUser(ctx, userID, func(tx pgx.Tx) error {
ct, err := tx.Exec(ctx,
`UPDATE videos
SET transcript_status = NULLIF($1, ''),
rate_limited_at = CASE WHEN $1 = 'rate_limited' THEN NOW() ELSE NULL END
WHERE id = $2`,
status, videoID)
if err != nil {
return fmt.Errorf("store: set transcript status: %w", err)
}
if ct.RowsAffected() == 0 {
return ErrNotFound
}
return nil
})
}
// GetTranscriptStatus returns a video's transcript_status ("" when unset/NULL).
// Returns ErrNotFound when the user has no such video. Scoped via withUser.
func (s *Store) GetTranscriptStatus(ctx context.Context, userID, videoID string) (string, error) {
var status string
if err := s.withUser(ctx, userID, func(tx pgx.Tx) error {
err := tx.QueryRow(ctx,
`SELECT COALESCE(transcript_status, '') FROM videos WHERE id = $1`, videoID).Scan(&status)
if errors.Is(err, pgx.ErrNoRows) {
return ErrNotFound
}
return err
}); err != nil {
if errors.Is(err, ErrNotFound) {
return "", ErrNotFound
}
return "", fmt.Errorf("store: get transcript status: %w", err)
}
return status, nil
}
// RateLimitedVideoIDs returns the user's videos currently in the "rate_limited"
// state, mapped to when the 429 was stamped (rate_limited_at). The run loop loads
// it once per pass (mirroring SeenVideoIDs) to skip re-fetching a video still
// inside the backoff window, saving caption requests. Scoped by user_id.
func (s *Store) RateLimitedVideoIDs(ctx context.Context, userID string) (map[string]time.Time, error) {
out := make(map[string]time.Time)
if err := s.withUser(ctx, userID, func(tx pgx.Tx) error {
rows, err := tx.Query(ctx,
`SELECT id, rate_limited_at FROM videos
WHERE user_id = $1 AND transcript_status = 'rate_limited' AND rate_limited_at IS NOT NULL`,
userID)
if err != nil {
return fmt.Errorf("store: rate limited video ids: %w", err)
}
defer rows.Close()
for rows.Next() {
var (
id string
at time.Time
)
if err := rows.Scan(&id, &at); err != nil {
return fmt.Errorf("store: scan rate limited id: %w", err)
}
out[id] = at
}
if err := rows.Err(); err != nil {
return fmt.Errorf("store: iterate rate limited ids: %w", err)
}
return nil
}); err != nil {
return nil, err
}
return out, nil
}
@@ -0,0 +1,80 @@
package store_test
import (
"context"
"testing"
"github.com/stretchr/testify/require"
"gitea.d-ma.be/mathias/tapir/internal/adapters/store"
)
func TestSetTranscriptStatus_RoundTrip(t *testing.T) {
ctx := context.Background()
s := newStore(t)
resetDB(t, rawPool(t))
id, err := s.UpsertVideo(ctx, ytVideo(userA, "rt12345", "round trip"))
require.NoError(t, err)
// Unset by default.
got, err := s.GetTranscriptStatus(ctx, userA, id)
require.NoError(t, err)
require.Equal(t, "", got)
for _, status := range []string{"none", "fetched", "rate_limited", ""} {
require.NoError(t, s.SetTranscriptStatus(ctx, userA, id, status))
got, err := s.GetTranscriptStatus(ctx, userA, id)
require.NoError(t, err)
require.Equal(t, status, got)
}
}
func TestSetTranscriptStatus_RejectsInvalid(t *testing.T) {
ctx := context.Background()
s := newStore(t)
resetDB(t, rawPool(t))
id, err := s.UpsertVideo(ctx, ytVideo(userA, "bad12345", "bad status"))
require.NoError(t, err)
require.Error(t, s.SetTranscriptStatus(ctx, userA, id, "bogus"))
// The rejected write left the status untouched.
got, err := s.GetTranscriptStatus(ctx, userA, id)
require.NoError(t, err)
require.Equal(t, "", got)
}
func TestSetTranscriptStatus_NotFound(t *testing.T) {
ctx := context.Background()
s := newStore(t)
resetDB(t, rawPool(t))
require.ErrorIs(t, s.SetTranscriptStatus(ctx, userA, videoX, "fetched"), store.ErrNotFound)
_, err := s.GetTranscriptStatus(ctx, userA, videoX)
require.ErrorIs(t, err, store.ErrNotFound)
}
func TestRateLimitedVideoIDs_StampsAndClears(t *testing.T) {
ctx := context.Background()
s := newStore(t)
resetDB(t, rawPool(t))
id, err := s.UpsertVideo(ctx, ytVideo(userA, "rl12345", "rate limited"))
require.NoError(t, err)
// Marking rate_limited stamps rate_limited_at, so the video appears.
require.NoError(t, s.SetTranscriptStatus(ctx, userA, id, "rate_limited"))
rl, err := s.RateLimitedVideoIDs(ctx, userA)
require.NoError(t, err)
require.Contains(t, rl, id)
require.False(t, rl[id].IsZero(), "rate_limited_at must be stamped")
// Moving off rate_limited clears the timestamp, so it drops out.
require.NoError(t, s.SetTranscriptStatus(ctx, userA, id, "fetched"))
rl, err = s.RateLimitedVideoIDs(ctx, userA)
require.NoError(t, err)
require.NotContains(t, rl, id)
}
@@ -0,0 +1,87 @@
package store_test
import (
"context"
"testing"
"github.com/stretchr/testify/require"
"gitea.d-ma.be/mathias/tapir/internal/adapters/store"
"gitea.d-ma.be/mathias/tapir/internal/domain"
"gitea.d-ma.be/mathias/tapir/internal/ports"
)
// Static check: Store satisfies the shared TranscriptStore port (ADR-021).
var _ ports.TranscriptStore = (*store.Store)(nil)
func TestSaveAndGetTranscript_RoundTrip(t *testing.T) {
s := newStore(t)
resetDB(t, rawPool(t))
ctx := context.Background()
want := domain.Transcript{Source: domain.SourceCaptions, Language: "en", Content: "the words"}
require.NoError(t, s.SaveTranscript(ctx, "youtube", "vid-1", want))
got, ok, err := s.GetTranscript(ctx, "youtube", "vid-1")
require.NoError(t, err)
require.True(t, ok, "a saved transcript must be found")
require.Equal(t, domain.SourceCaptions, got.Source)
require.Equal(t, "en", got.Language)
require.Equal(t, "the words", got.Content)
require.True(t, got.HasText())
}
func TestGetTranscript_Miss(t *testing.T) {
s := newStore(t)
resetDB(t, rawPool(t))
_, ok, err := s.GetTranscript(context.Background(), "youtube", "absent")
require.NoError(t, err, "a miss is not an error")
require.False(t, ok)
}
// A stored "no captions" outcome is a real hit: callers must skip without
// re-fetching, so ok is true even though there is no text (ADR-021 / ADR-007).
func TestSaveAndGetTranscript_NoneIsAStoredHit(t *testing.T) {
s := newStore(t)
resetDB(t, rawPool(t))
ctx := context.Background()
require.NoError(t, s.SaveTranscript(ctx, "youtube", "vid-none", domain.Transcript{Source: domain.SourceNone}))
got, ok, err := s.GetTranscript(ctx, "youtube", "vid-none")
require.NoError(t, err)
require.True(t, ok, "a stored SourceNone is a hit, not a miss")
require.Equal(t, domain.SourceNone, got.Source)
require.False(t, got.HasText())
}
// A transient 429 must never be persisted as a terminal transcript, or a later
// read would mask the rate-limit as a permanent "no transcript" (ADR-014).
func TestSaveTranscript_RejectsRateLimited(t *testing.T) {
s := newStore(t)
resetDB(t, rawPool(t))
err := s.SaveTranscript(context.Background(), "youtube", "vid-429",
domain.Transcript{Source: domain.SourceRateLimited})
require.Error(t, err)
_, ok, _ := s.GetTranscript(context.Background(), "youtube", "vid-429")
require.False(t, ok, "a rejected rate-limited save must leave nothing stored")
}
func TestSaveTranscript_UpsertLastWriteWins(t *testing.T) {
s := newStore(t)
resetDB(t, rawPool(t))
ctx := context.Background()
require.NoError(t, s.SaveTranscript(ctx, "youtube", "vid-up", domain.Transcript{Source: domain.SourceNone}))
require.NoError(t, s.SaveTranscript(ctx, "youtube", "vid-up",
domain.Transcript{Source: domain.SourceCaptions, Language: "en", Content: "now resolved"}))
got, ok, err := s.GetTranscript(ctx, "youtube", "vid-up")
require.NoError(t, err)
require.True(t, ok)
require.Equal(t, domain.SourceCaptions, got.Source)
require.Equal(t, "now resolved", got.Content)
}
+73 -6
View File
@@ -46,14 +46,15 @@ func (s *Store) UpsertVideo(ctx context.Context, v domain.Video) (string, error)
}
if err := tx.QueryRow(ctx,
`INSERT INTO videos (user_id, provider, provider_video_id, title, url, published_at)
VALUES ($1, $2, $3, $4, $5, $6)
`INSERT INTO videos (user_id, provider, provider_video_id, title, url, published_at, channel_title)
VALUES ($1, $2, $3, $4, $5, $6, $7)
ON CONFLICT (user_id, provider, provider_video_id) DO UPDATE SET
title = EXCLUDED.title,
url = EXCLUDED.url,
published_at = EXCLUDED.published_at
title = EXCLUDED.title,
url = EXCLUDED.url,
published_at = EXCLUDED.published_at,
channel_title = COALESCE(NULLIF(EXCLUDED.channel_title, ''), videos.channel_title)
RETURNING id`,
v.UserID, provider, v.ProviderVideoID, v.Title, v.URL, nullTime(v.PublishedAt),
v.UserID, provider, v.ProviderVideoID, v.Title, v.URL, nullTime(v.PublishedAt), v.ChannelTitle,
).Scan(&id); err != nil {
return fmt.Errorf("store: upsert video: %w", err)
}
@@ -72,3 +73,69 @@ func nullTime(t time.Time) *time.Time {
}
return &t
}
// NewestUnsummarizedVideoIDs returns up to limit of the user's videos that have
// no summary yet, newest first (published_at DESC, NULLS LAST). It caps the
// connect-time onboarding burst (Feature 1) at a fixed count: the caller marks
// these for summarization through the shared rate gate. RLS-scoped via withUser,
// so it only ever sees the requesting user's rows. limit <= 0 returns nil.
func (s *Store) NewestUnsummarizedVideoIDs(ctx context.Context, userID string, limit int) ([]string, error) {
if limit <= 0 {
return nil, nil
}
var ids []string
if err := s.withUser(ctx, userID, func(tx pgx.Tx) error {
rows, err := tx.Query(ctx,
`SELECT v.id
FROM videos v
WHERE v.user_id = $1
AND NOT EXISTS (
SELECT 1 FROM summaries su
WHERE su.user_id = v.user_id AND su.video_id = v.id)
ORDER BY v.published_at DESC NULLS LAST, v.seen_at DESC
LIMIT $2`, userID, limit)
if err != nil {
return fmt.Errorf("store: newest unsummarized: %w", err)
}
defer rows.Close()
for rows.Next() {
var id string
if err := rows.Scan(&id); err != nil {
return fmt.Errorf("store: scan newest unsummarized: %w", err)
}
ids = append(ids, id)
}
return rows.Err()
}); err != nil {
return nil, err
}
return ids, nil
}
// DistinctChannels returns the user's distinct, non-empty source channel titles
// (the channels they have videos from), alphabetically — the option list for the
// feed's channel filter. RLS-scoped via withUser.
func (s *Store) DistinctChannels(ctx context.Context, userID string) ([]string, error) {
var out []string
if err := s.withUser(ctx, userID, func(tx pgx.Tx) error {
rows, err := tx.Query(ctx,
`SELECT DISTINCT channel_title FROM videos
WHERE user_id = $1 AND channel_title IS NOT NULL AND channel_title <> ''
ORDER BY channel_title`, userID)
if err != nil {
return fmt.Errorf("store: distinct channels: %w", err)
}
defer rows.Close()
for rows.Next() {
var c string
if err := rows.Scan(&c); err != nil {
return fmt.Errorf("store: scan channel: %w", err)
}
out = append(out, c)
}
return rows.Err()
}); err != nil {
return nil, err
}
return out, nil
}
+58
View File
@@ -81,3 +81,61 @@ func TestUpsertVideo_PerUserIsolation(t *testing.T) {
require.NotEqual(t, idA, idB, "same provider video for two users must be two distinct rows")
}
func TestNewestUnsummarizedVideoIDs(t *testing.T) {
ctx := context.Background()
s := newStore(t)
resetDB(t, rawPool(t))
mk := func(user, pid string, day int) string {
v := ytVideo(user, pid, pid)
v.PublishedAt = time.Date(2026, 6, day, 12, 0, 0, 0, time.UTC)
id, err := s.UpsertVideo(ctx, v)
require.NoError(t, err)
return id
}
_ = mk(userA, "a1vid000001", 1)
id2 := mk(userA, "a2vid000002", 2)
id3 := mk(userA, "a3vid000003", 3)
id4 := mk(userA, "a4vid000004", 4)
mk(userB, "b1vid000009", 9) // userB's newest — must never leak via RLS
// The newest (v4) is summarized, so it's excluded from "unsummarized".
require.NoError(t, s.Deliver(ctx, summary(userA, id4, "done")))
// Cap 2, newest-first unsummarized: v3 then v2 (v4 excluded; userB excluded).
got, err := s.NewestUnsummarizedVideoIDs(ctx, userA, 2)
require.NoError(t, err)
require.Equal(t, []string{id3, id2}, got)
none, err := s.NewestUnsummarizedVideoIDs(ctx, userA, 0)
require.NoError(t, err)
require.Empty(t, none, "limit 0 returns nothing")
}
func TestUpsertVideoPersistsChannelAndDistinctChannels(t *testing.T) {
ctx := context.Background()
s := newStore(t)
resetDB(t, rawPool(t))
mk := func(pid, channel string) {
v := ytVideo(userA, pid, pid)
v.ChannelTitle = channel
_, err := s.UpsertVideo(ctx, v)
require.NoError(t, err)
}
mk("aa11111aaaa", "Acme Talks")
mk("bb22222bbbb", "Acme Talks") // same channel
mk("cc33333cccc", "Zeta Channel")
// userB's channel must not leak.
vb := ytVideo(userB, "dd44444dddd", "x")
vb.ChannelTitle = "Bravo Only"
_, err := s.UpsertVideo(ctx, vb)
require.NoError(t, err)
got, err := s.DistinctChannels(ctx, userA)
require.NoError(t, err)
require.Equal(t, []string{"Acme Talks", "Zeta Channel"}, got,
"distinct, alphabetical, user-scoped (no Bravo Only)")
}
+15
View File
@@ -60,6 +60,12 @@ func (a *Adapter) FetchTranscript(ctx context.Context, v domain.Video) (domain.T
if err != nil {
return domain.Transcript{}, fmt.Errorf("download caption track for %q: %w", v.ProviderVideoID, err)
}
if status == http.StatusTooManyRequests {
// 429 means the IP is rate-limited; record for retry, not a permanent
// absence. Degrade gracefully (no error, no text) like SourceNone, but
// flag it distinctly so the runner backs off and retries (ADR-007/010).
return domain.Transcript{VideoID: v.ID, UserID: v.UserID, Source: domain.SourceRateLimited}, nil
}
if status != http.StatusOK {
// Owner-only 403, region/age gate, or transient unavailability: not an error.
return noTranscript(v), nil
@@ -229,6 +235,15 @@ func (a *Adapter) httpDo(ctx context.Context, client *http.Client, method, url s
for k, v := range headers {
req.Header.Set(k, v)
}
// Process-wide rate gate (ADR-014 item 2): every live caption fetch — player,
// watch-page, and timedtext baseUrl — passes the shared per-egress-IP limiter
// so the scheduler and the click-path cannot collectively trip 429s. Skipped
// when a.transport is set (the test seam) so fakes are not throttled.
if a.transport == nil {
if err := WaitFetchGate(ctx); err != nil {
return nil, 0, fmt.Errorf("fetch gate %s %s: %w", method, url, err)
}
}
resp, err := client.Do(req)
if err != nil {
return nil, 0, fmt.Errorf("%s %s: %w", method, url, err)
+37
View File
@@ -0,0 +1,37 @@
package youtube
import (
"context"
"time"
"golang.org/x/time/rate"
)
// globalFetchGate is the process-wide rate limiter for outbound timedtext/caption
// fetches. A single instance is shared by ALL Adapter instances (the scheduler's
// per-user runners + the web click-path) so they cannot collectively exceed the
// per-egress-IP cap. ADR-014 item 2: the gate serialises/limits concurrent caption
// fetches regardless of how many users or goroutines are upstream. The 429 is per
// IP, not per user — so the gate is process-wide, not per-withUser, not per-video.
//
// Default 2s/req (burst 1): the first fetch passes immediately, subsequent fetches
// are spaced at least 2s apart. Production overrides via SetFetchRate from config.
var globalFetchGate = rate.NewLimiter(rate.Every(2*time.Second), 1)
// SetFetchRate replaces the process-wide gate's rate with one token per interval.
// Call once at startup from config (TAPIR_FETCH_RATE). A non-positive interval
// installs an unlimited gate (rate.Inf) — used in dev/tests so nothing throttles.
func SetFetchRate(interval time.Duration) {
if interval <= 0 {
globalFetchGate = rate.NewLimiter(rate.Inf, 1)
return
}
globalFetchGate = rate.NewLimiter(rate.Every(interval), 1)
}
// WaitFetchGate blocks until the process-wide gate allows one timedtext fetch,
// respecting ctx cancellation. Called from httpDo before every live outbound
// caption fetch so the scheduler and the click-path share the same egress budget.
func WaitFetchGate(ctx context.Context) error {
return globalFetchGate.Wait(ctx)
}
+84
View File
@@ -0,0 +1,84 @@
package youtube
import (
"context"
"sync"
"sync/atomic"
"testing"
"time"
"github.com/stretchr/testify/require"
"golang.org/x/time/rate"
)
// waitOn drives a test-scoped limiter the same way WaitFetchGate drives the
// global one, so these tests exercise the gate's behaviour without mutating the
// process-wide gate (which would pollute sibling tests / the click-path).
func waitOn(t *testing.T, lim *rate.Limiter) func(context.Context) error {
t.Helper()
return func(ctx context.Context) error { return lim.Wait(ctx) }
}
func TestFetchGateSerialisesConcurrentCallers(t *testing.T) {
const (
n = 5
interval = 10 * time.Millisecond
)
lim := rate.NewLimiter(rate.Every(interval), 1)
wait := waitOn(t, lim)
var (
inFlight, maxInFlight atomic.Int32
wg sync.WaitGroup
)
start := time.Now()
for i := 0; i < n; i++ {
wg.Add(1)
go func() {
defer wg.Done()
require.NoError(t, wait(context.Background()))
cur := inFlight.Add(1)
for {
old := maxInFlight.Load()
if cur <= old || maxInFlight.CompareAndSwap(old, cur) {
break
}
}
// Hold the "critical section" briefly so overlap would be observable.
time.Sleep(interval / 4)
inFlight.Add(-1)
}()
}
wg.Wait()
elapsed := time.Since(start)
require.Equal(t, int32(1), maxInFlight.Load(),
"the gate must admit at most one caller per interval — no overlap")
require.GreaterOrEqual(t, elapsed, time.Duration(n-1)*interval,
"N gated callers take at least (N-1)*interval wall time")
}
func TestFetchGateRespectsContextCancellation(t *testing.T) {
// A slow gate (1 token/hour, burst already spent) blocks; a cancelled ctx must
// unblock Wait with an error rather than hang.
lim := rate.NewLimiter(rate.Every(time.Hour), 1)
require.True(t, lim.Allow(), "spend the single burst token")
ctx, cancel := context.WithCancel(context.Background())
cancel()
require.Error(t, waitOn(t, lim)(ctx), "cancelled ctx must fail Wait, not block")
}
func TestSetFetchRateZeroIsUnlimited(t *testing.T) {
// Snapshot and restore the global so this test does not pollute the process.
prev := globalFetchGate
t.Cleanup(func() { globalFetchGate = prev })
SetFetchRate(0)
require.Equal(t, rate.Inf, globalFetchGate.Limit(), "0 interval = unlimited gate")
// An unlimited gate never blocks, even back-to-back.
for i := 0; i < 100; i++ {
require.NoError(t, WaitFetchGate(context.Background()))
}
}
@@ -0,0 +1,62 @@
package youtube
import (
"context"
"errors"
"net/http"
"testing"
"gitea.d-ma.be/mathias/tapir/internal/domain"
)
func TestVideoByID(t *testing.T) {
const id = "dQw4w9WgXcQ"
a, secrets := newTestAdapter(t, func(w http.ResponseWriter, r *http.Request) {
if r.URL.Path != "/videos" {
t.Errorf("unexpected path %q (must use videos.list)", r.URL.Path)
}
if got := r.URL.Query().Get("id"); got != id {
t.Errorf("expected id=%s, got %q", id, got)
}
if got := r.URL.Query().Get("part"); got != "snippet" {
t.Errorf("expected part=snippet, got %q", got)
}
_, _ = w.Write([]byte(`{"items":[{"snippet":{"title":"Never Gonna Give You Up","channelTitle":"Rick Astley","publishedAt":"2026-05-20T09:00:00Z"}}]}`))
})
v, err := a.VideoByID(context.Background(), "u1", id)
if err != nil {
t.Fatalf("VideoByID: %v", err)
}
if v.UserID != "u1" {
t.Errorf("UserID = %q, want u1", v.UserID)
}
if v.ProviderVideoID != id || v.Title != "Never Gonna Give You Up" {
t.Errorf("unexpected video: %+v", v)
}
if v.ChannelTitle != "Rick Astley" {
t.Errorf("ChannelTitle = %q, want Rick Astley", v.ChannelTitle)
}
if v.Provider != domain.ProviderYouTube || v.URL != "https://www.youtube.com/watch?v="+id {
t.Errorf("video not wired correctly: %+v", v)
}
if v.PublishedAt.IsZero() {
t.Errorf("expected publishedAt parsed, got zero")
}
if v.SubscriptionID != "" {
t.Errorf("a pasted video must have no subscription, got %q", v.SubscriptionID)
}
if secrets.byRef == nil {
t.Errorf("token must be resolved by reference through the SecretStore")
}
}
func TestVideoByIDNotFound(t *testing.T) {
a, _ := newTestAdapter(t, func(w http.ResponseWriter, _ *http.Request) {
_, _ = w.Write([]byte(`{"items":[]}`))
})
_, err := a.VideoByID(context.Background(), "u1", "missingvid0")
if !errors.Is(err, domain.ErrVideoNotFound) {
t.Fatalf("VideoByID for missing id = %v, want domain.ErrVideoNotFound", err)
}
}
+52
View File
@@ -216,6 +216,9 @@ func (a *Adapter) NewVideos(ctx context.Context, sub domain.Subscription) ([]dom
var resp playlistItemListResponse
if err := a.getJSON(ctx, client, "/playlistItems", q, &resp); err != nil {
if isHTTP404(err) {
return nil, &domain.ErrChannelUnavailable{ChannelID: sub.ChannelID, ChannelTitle: sub.ChannelTitle}
}
return nil, fmt.Errorf("new videos for channel %q: %w", sub.ChannelID, err)
}
@@ -231,6 +234,7 @@ func (a *Adapter) NewVideos(ctx context.Context, sub domain.Subscription) ([]dom
Provider: domain.ProviderYouTube,
ProviderVideoID: vid,
Title: item.Snippet.Title,
ChannelTitle: sub.ChannelTitle,
URL: "https://www.youtube.com/watch?v=" + vid,
PublishedAt: item.Snippet.PublishedAt,
})
@@ -241,6 +245,38 @@ func (a *Adapter) NewVideos(ctx context.Context, sub domain.Subscription) ([]dom
return videos, nil
}
// VideoByID fetches a single video's metadata (videos.list, snippet) for an
// arbitrary video id — including channels the user does not follow (paste-a-URL,
// Feature 2). This is a Data API call (1 quota unit), NOT the rate-limited
// caption path, so it is not gated: only the later transcript fetch goes through
// globalFetchGate. UserID is set on the result and SubscriptionID is left empty
// (a pasted video has no subscription parent). Returns ErrVideoNotFound when the
// id resolves to no video.
func (a *Adapter) VideoByID(ctx context.Context, userID, videoID string) (domain.Video, error) {
client, err := a.httpClient(ctx, a.cfg.TokenSecretRef)
if err != nil {
return domain.Video{}, err
}
q := url.Values{"part": {"snippet"}, "id": {videoID}}
var resp videoListResponse
if err := a.getJSON(ctx, client, "/videos", q, &resp); err != nil {
return domain.Video{}, fmt.Errorf("video by id %q: %w", videoID, err)
}
if len(resp.Items) == 0 {
return domain.Video{}, fmt.Errorf("video %q: %w", videoID, domain.ErrVideoNotFound)
}
it := resp.Items[0]
return domain.Video{
UserID: userID,
Provider: domain.ProviderYouTube,
ProviderVideoID: videoID,
Title: it.Snippet.Title,
ChannelTitle: it.Snippet.ChannelTitle,
URL: "https://www.youtube.com/watch?v=" + videoID,
PublishedAt: it.Snippet.PublishedAt,
}, nil
}
// uploadsPlaylistID derives a channel's uploads playlist id at zero API cost:
// a standard channel id "UCxxxx" maps to uploads playlist "UUxxxx". Returns
// ok=false for ids that don't follow this convention (caller falls back to
@@ -269,6 +305,12 @@ func (a *Adapter) resolveUploadsPlaylist(ctx context.Context, client *http.Clien
return resp.Items[0].ContentDetails.RelatedPlaylists.Uploads, nil
}
// isHTTP404 reports whether err came from a YouTube API call that returned HTTP 404.
// getRaw encodes the status as "youtube api <path>: status 404: ...".
func isHTTP404(err error) bool {
return err != nil && strings.Contains(err.Error(), "status 404")
}
// getJSON issues a GET and decodes a JSON body into out. A non-200 status is an
// error carrying a bounded slice of the response body for diagnosis.
func (a *Adapter) getJSON(ctx context.Context, client *http.Client, path string, q url.Values, out any) error {
@@ -333,6 +375,16 @@ type playlistItemListResponse struct {
} `json:"items"`
}
type videoListResponse struct {
Items []struct {
Snippet struct {
Title string `json:"title"`
ChannelTitle string `json:"channelTitle"`
PublishedAt time.Time `json:"publishedAt"`
} `json:"snippet"`
} `json:"items"`
}
type channelListResponse struct {
Items []struct {
ContentDetails struct {
+35 -1
View File
@@ -131,7 +131,7 @@ func TestNewVideos(t *testing.T) {
}`))
})
sub := domain.Subscription{ID: "s1", UserID: "u1", ChannelID: "UC_acme"}
sub := domain.Subscription{ID: "s1", UserID: "u1", ChannelID: "UC_acme", ChannelTitle: "Acme Channel"}
vids, err := a.NewVideos(context.Background(), sub)
if err != nil {
t.Fatalf("NewVideos: %v", err)
@@ -143,6 +143,9 @@ func TestNewVideos(t *testing.T) {
if v.ProviderVideoID != "vid1" || v.Title != "Designing for Attention" {
t.Errorf("unexpected video: %+v", v)
}
if v.ChannelTitle != "Acme Channel" {
t.Errorf("ChannelTitle = %q, want Acme Channel", v.ChannelTitle)
}
if v.Provider != domain.ProviderYouTube || v.URL != "https://www.youtube.com/watch?v=vid1" {
t.Errorf("video not wired correctly: %+v", v)
}
@@ -388,6 +391,37 @@ func TestFetchTranscriptBaseURLForbiddenDegrades(t *testing.T) {
}
}
// A 429 on the baseUrl fetch is the IP being rate-limited, NOT a permanent
// absence of captions: it returns SourceRateLimited (no error, no text) so the
// runner can record it and retry after a backoff window rather than recording a
// false "no transcript".
func TestFetchTranscriptRateLimitedReturnsSourceRateLimited(t *testing.T) {
a, _ := newTestAdapter(t, func(w http.ResponseWriter, r *http.Request) {
switch r.URL.Path {
case "/youtubei/v1/player":
base := "http://" + r.Host
_, _ = w.Write([]byte(`{"captions":{"playerCaptionsTracklistRenderer":{"captionTracks":[` +
`{"baseUrl":"` + base + `/api/timedtext?lang=en","languageCode":"en"}]}}}`))
case "/api/timedtext":
w.WriteHeader(http.StatusTooManyRequests)
}
})
tr, err := a.FetchTranscript(context.Background(), domain.Video{ID: "v1", UserID: "u1", ProviderVideoID: "vid1"})
if err != nil {
t.Fatalf("429 on baseUrl must degrade, not error: %v", err)
}
if tr.Source != domain.SourceRateLimited {
t.Fatalf("expected SourceRateLimited on 429, got %q", tr.Source)
}
if tr.HasText() {
t.Error("expected HasText() false for SourceRateLimited")
}
if tr.Content != "" {
t.Errorf("expected empty content on 429, got %q", tr.Content)
}
}
// An empty baseUrl on the selected track degrades to SourceNone, never an error.
func TestFetchTranscriptEmptyBaseURLDegrades(t *testing.T) {
a, _ := newTestAdapter(t, func(w http.ResponseWriter, r *http.Request) {
+100 -7
View File
@@ -13,6 +13,7 @@ import (
"os"
"path/filepath"
"sort"
"strconv"
"strings"
"time"
)
@@ -59,9 +60,46 @@ type Config struct {
// PollInterval, when > 0, makes `run` loop on that cadence; 0 means run once.
PollInterval time.Duration
// FetchBackoff is how long the run loop waits before re-fetching a transcript
// that previously returned HTTP 429 (rate_limited). Inside the window the video
// is skipped without hitting the caption endpoint, saving requests; after it
// expires the video is retried. Zero means "always retry" (no backoff).
FetchBackoff time.Duration
// FetchRate is the minimum interval between outbound caption fetches across the
// whole process — the shared per-egress-IP rate gate (ADR-014 item 2). It is
// the gate that makes auto-summarize-on-a-schedule safe: scheduler runners and
// the web click-path serialise through it. Zero = unlimited (dev/tests).
FetchRate time.Duration
// AutoSummarizeWindow bounds auto-summarization to recent videos: in automatic
// mode the scheduler only summarizes videos published within this window of now.
// Older videos are still discovered and listed, but wait for an explicit manual
// "Summarize" — so a large back-catalogue does not self-inflict 429s against the
// caption rate gate. Zero disables the bound (summarize every unseen video, the
// pre-recency behaviour). Default ~7 days.
AutoSummarizeWindow time.Duration
// OnboardSummarizeCount caps how many of a freshly-connected user's newest
// videos are summarized immediately on connect (the onboarding "it works"
// burst). HARD-capped at maxOnboardSummarizeCount so onboarding can never
// bulk-fetch; 0 disables the burst. Every fetch still flows through the shared
// caption rate gate (ADR-014) — the cap bounds count, never the pacing. Default 3.
OnboardSummarizeCount int
// DiscoveryInterval, when > 0, makes `serve` run in-process scheduled discovery
// for ALL users on that cadence (ADR-018). Zero/unset = disabled, so dev and
// tests never auto-fetch. Single-replica assumption — see cmdServe.
DiscoveryInterval time.Duration
// HTTPAddr is the listen address for `tapir serve` (the Stage-0 web UI).
HTTPAddr string
// PublicURL is the externally-reachable base URL of the deployed service,
// e.g. "https://tapir.d-ma.be". Used to build absolute links handed to humans
// (the `tapir invite` URL). No trailing slash is assumed — callers trim it.
PublicURL string
// Dex OIDC (web login, ADR-011/012). When OIDCIssuer is empty, `serve` falls
// back to the allow-all StubAuth (local dev). When set, serve uses Dex: any
// Dex-authenticated subject may sign in, then registers a tapir user (ADR-012).
@@ -78,13 +116,19 @@ func (c Config) DexConfigured() bool { return strings.TrimSpace(c.OIDCIssuer) !=
// Defaults (see docs/homelab-integration.md). All overridable via env.
const (
defaultGatewayURL = "http://koala:30401/v1"
defaultSummarizerModel = "koala/phi4-mini"
defaultSummarizerTimeout = 5 * time.Minute
defaultYTTokenRef = "youtube/refresh_token"
defaultYTConnectRedirectURL = "https://tapir.d-ma.be/oauth/youtube/callback"
defaultOAuthRedirectAddr = "localhost:8080"
defaultHTTPAddr = ":8080"
defaultGatewayURL = "http://koala:30401/v1"
defaultSummarizerModel = "koala/phi4-mini"
defaultSummarizerTimeout = 5 * time.Minute
defaultYTTokenRef = "youtube/refresh_token"
defaultYTConnectRedirectURL = "https://tapir.d-ma.be/oauth/youtube/callback"
defaultOAuthRedirectAddr = "localhost:8080"
defaultHTTPAddr = ":8080"
defaultFetchBackoff = time.Hour
defaultFetchRate = 2 * time.Second
defaultPublicURL = "https://tapir.d-ma.be"
defaultAutoSummarizeWindow = 7 * 24 * time.Hour
defaultOnboardSummarizeCount = 3
maxOnboardSummarizeCount = 5
)
// Load reads the environment into a Config, applying defaults. It does not
@@ -105,6 +149,7 @@ func Load() (Config, error) {
SecretsFile: envOr("TAPIR_SECRETS_FILE", defaultSecretsFile()),
OAuthRedirectAddr: envOr("TAPIR_OAUTH_REDIRECT_ADDR", defaultOAuthRedirectAddr),
HTTPAddr: envOr("TAPIR_HTTP_ADDR", defaultHTTPAddr),
PublicURL: envOr("TAPIR_PUBLIC_URL", defaultPublicURL),
OIDCIssuer: os.Getenv("TAPIR_OIDC_ISSUER"),
DexClientID: os.Getenv("TAPIR_DEX_CLIENT_ID"),
DexClientSecret: os.Getenv("TAPIR_DEX_CLIENT_SECRET"),
@@ -124,6 +169,42 @@ func Load() (Config, error) {
}
c.PollInterval = interval
backoff, err := durationOr("TAPIR_FETCH_BACKOFF", defaultFetchBackoff)
if err != nil {
return Config{}, err
}
c.FetchBackoff = backoff
fetchRate, err := durationOr("TAPIR_FETCH_RATE", defaultFetchRate)
if err != nil {
return Config{}, err
}
c.FetchRate = fetchRate
discovery, err := durationOr("TAPIR_DISCOVERY_INTERVAL", 0)
if err != nil {
return Config{}, err
}
c.DiscoveryInterval = discovery
autoWindow, err := durationOr("TAPIR_AUTO_SUMMARIZE_WINDOW", defaultAutoSummarizeWindow)
if err != nil {
return Config{}, err
}
c.AutoSummarizeWindow = autoWindow
onboard, err := intOr("TAPIR_ONBOARD_SUMMARIZE_COUNT", defaultOnboardSummarizeCount)
if err != nil {
return Config{}, err
}
if onboard < 0 {
onboard = 0
}
if onboard > maxOnboardSummarizeCount {
onboard = maxOnboardSummarizeCount
}
c.OnboardSummarizeCount = onboard
return c, nil
}
@@ -188,6 +269,18 @@ func envOr(key, fallback string) string {
return fallback
}
func intOr(key string, fallback int) (int, error) {
v := os.Getenv(key)
if v == "" {
return fallback, nil
}
n, err := strconv.Atoi(v)
if err != nil {
return 0, fmt.Errorf("config: %s=%q: %w", key, v, err)
}
return n, nil
}
func durationOr(key string, fallback time.Duration) (time.Duration, error) {
v := os.Getenv(key)
if v == "" {
+53 -7
View File
@@ -45,17 +45,57 @@ func TestLoad_AppliesDefaults(t *testing.T) {
if c.PollInterval != 0 {
t.Errorf("PollInterval = %v, want 0 (run once)", c.PollInterval)
}
if c.FetchBackoff != defaultFetchBackoff {
t.Errorf("FetchBackoff = %v, want default %v", c.FetchBackoff, defaultFetchBackoff)
}
if c.AutoSummarizeWindow != defaultAutoSummarizeWindow {
t.Errorf("AutoSummarizeWindow = %v, want default %v", c.AutoSummarizeWindow, defaultAutoSummarizeWindow)
}
}
func TestLoad_OnboardSummarizeCount(t *testing.T) {
cases := []struct {
name, env string
want int
}{
{"default", "", defaultOnboardSummarizeCount},
{"explicit", "4", 4},
{"zero disables", "0", 0},
{"clamped to hard cap", "50", maxOnboardSummarizeCount},
{"negative clamps to zero", "-3", 0},
}
for _, c := range cases {
t.Run(c.name, func(t *testing.T) {
setEnv(t, map[string]string{"TAPIR_ONBOARD_SUMMARIZE_COUNT": c.env})
cfg, err := Load()
if err != nil {
t.Fatalf("Load: %v", err)
}
if cfg.OnboardSummarizeCount != c.want {
t.Fatalf("OnboardSummarizeCount = %d, want %d", cfg.OnboardSummarizeCount, c.want)
}
})
}
}
func TestLoad_OnboardSummarizeCountInvalid(t *testing.T) {
setEnv(t, map[string]string{"TAPIR_ONBOARD_SUMMARIZE_COUNT": "three"})
if _, err := Load(); err == nil {
t.Fatal("Load: want error for non-numeric TAPIR_ONBOARD_SUMMARIZE_COUNT")
}
}
func TestLoad_ParsesValues(t *testing.T) {
setEnv(t, map[string]string{
"TAPIR_USER_ID": "11111111-1111-1111-1111-111111111111",
"TAPIR_GATEWAY_URL": "http://example/v1",
"TAPIR_GATEWAY_KEY": "sk-test",
"TAPIR_SUMMARIZER_MODEL": "iguana/deepseek-r1-14b",
"TAPIR_SUMMARIZER_TIMEOUT": "90s",
"TAPIR_DB_DSN": "postgres://x",
"TAPIR_POLL_INTERVAL": "10m",
"TAPIR_USER_ID": "11111111-1111-1111-1111-111111111111",
"TAPIR_GATEWAY_URL": "http://example/v1",
"TAPIR_GATEWAY_KEY": "sk-test",
"TAPIR_SUMMARIZER_MODEL": "iguana/deepseek-r1-14b",
"TAPIR_SUMMARIZER_TIMEOUT": "90s",
"TAPIR_DB_DSN": "postgres://x",
"TAPIR_POLL_INTERVAL": "10m",
"TAPIR_FETCH_BACKOFF": "30m",
"TAPIR_AUTO_SUMMARIZE_WINDOW": "48h",
})
c, err := Load()
@@ -77,6 +117,12 @@ func TestLoad_ParsesValues(t *testing.T) {
if c.PollInterval != 10*time.Minute {
t.Errorf("PollInterval = %v, want 10m", c.PollInterval)
}
if c.FetchBackoff != 30*time.Minute {
t.Errorf("FetchBackoff = %v, want 30m", c.FetchBackoff)
}
if c.AutoSummarizeWindow != 48*time.Hour {
t.Errorf("AutoSummarizeWindow = %v, want 48h", c.AutoSummarizeWindow)
}
}
func TestLoad_RejectsBadDuration(t *testing.T) {
+29 -1
View File
@@ -2,7 +2,28 @@
// the standard library — no providers, no storage, no AI. See docs/data-model.md.
package domain
import "time"
import (
"errors"
"fmt"
"time"
)
// ErrVideoNotFound is returned when a video id resolves to no video (deleted,
// private, or a typo'd paste). Defined in domain so adapters and the web layer
// share one sentinel without coupling to each other.
var ErrVideoNotFound = errors.New("video not found")
// ErrChannelUnavailable is returned by a VideoSource when a channel's upload
// playlist returns HTTP 404 — the channel was deleted or made private. The runner
// stores these so the account page can surface them to the user.
type ErrChannelUnavailable struct {
ChannelID string
ChannelTitle string
}
func (e *ErrChannelUnavailable) Error() string {
return fmt.Sprintf("channel %q (%s) unavailable: playlist not found", e.ChannelTitle, e.ChannelID)
}
// Provider identifies a video platform.
type Provider string
@@ -18,6 +39,12 @@ type TranscriptSource string
const (
SourceCaptions TranscriptSource = "captions"
SourceNone TranscriptSource = "none"
// SourceRateLimited records that the caption endpoint returned HTTP 429.
// Unlike SourceNone (a permanent absence), this is a transient "retry later":
// the IP is rate-limited, not the video caption-less. It carries no text
// (HasText is false), so the engine degrades the same as SourceNone, but the
// runner persists it distinctly to retry after a backoff window.
SourceRateLimited TranscriptSource = "rate_limited"
)
// User is the Tapir-side profile. At Stage 0 there is exactly one.
@@ -46,6 +73,7 @@ type Video struct {
Provider Provider
ProviderVideoID string
Title string
ChannelTitle string
URL string
PublishedAt time.Time
SeenAt time.Time
+20
View File
@@ -27,6 +27,26 @@ type Summarizer interface {
Summarize(ctx context.Context, v domain.Video, t domain.Transcript) (domain.Summary, error)
}
// TranscriptStore persists transcripts as shared, video-keyed public content
// (ADR-021). It is keyed by the cross-user dedup key (provider, providerVideoID)
// — the video's public identity, NOT Tapir's per-user videos.id — and holds only
// public caption content, so it is deliberately NOT user-scoped: two users who
// share a video share the one row. The engine reads it before any caption fetch
// so re-analysis never re-touches YouTube (ADR-010/014).
type TranscriptStore interface {
// GetTranscript returns the stored transcript for a video and whether one
// exists. A stored Source == SourceNone (captions permanently absent) is a
// real hit: ok is true and HasText() is false, so callers skip without
// re-fetching. A transient rate-limit is never stored, so it never appears
// here as a false absence.
GetTranscript(ctx context.Context, provider, providerVideoID string) (t domain.Transcript, ok bool, err error)
// SaveTranscript upserts the transcript for (provider, providerVideoID). Only
// terminal outcomes are persisted: SourceCaptions (with text) or SourceNone.
// SourceRateLimited must NOT be passed — it is a per-user retry (ADR-014), not
// a shared terminal state.
SaveTranscript(ctx context.Context, provider, providerVideoID string, t domain.Transcript) error
}
// Sink delivers a summary to a destination (user store, brain, ...).
// Implementations fail independently of one another.
type Sink interface {
+224 -54
View File
@@ -10,11 +10,13 @@
package runner
import (
"cmp"
"context"
"errors"
"fmt"
"log/slog"
"os"
"slices"
"time"
"gitea.d-ma.be/mathias/tapir/internal/domain"
@@ -22,6 +24,36 @@ import (
"gitea.d-ma.be/mathias/tapir/internal/usecase"
)
// passCandidate is a video that passed all pre-filters (seen/manual/backoff)
// and is queued for transcript fetch + summarization in this pass.
type passCandidate struct {
v domain.Video
pos int // discovery position — used as a stable tiebreak when published_at ties
}
// compareNewestFirst orders candidates by published_at descending, NULLS LAST,
// with pos ascending as a stable tiebreak. Videos with a zero published_at
// (schema 001: nullable) sort after all dated videos regardless of pos.
func compareNewestFirst(a, b passCandidate) int {
aNull := a.v.PublishedAt.IsZero()
bNull := b.v.PublishedAt.IsZero()
switch {
case aNull && bNull:
return cmp.Compare(a.pos, b.pos)
case aNull:
return 1 // a is null → after b
case bNull:
return -1 // b is null → after a
}
if !a.v.PublishedAt.Equal(b.v.PublishedAt) {
if a.v.PublishedAt.After(b.v.PublishedAt) {
return -1 // newer first
}
return 1
}
return cmp.Compare(a.pos, b.pos) // same timestamp: preserve discovery order
}
// VideoStore is the durable persistence the run loop needs: assign a stable id +
// metadata, read the already-summarized set, and (for manual summarization mode)
// read the user's mode + queued videos and clear a video's queue flag once it has
@@ -32,6 +64,16 @@ type VideoStore interface {
GetAutoSummarize(ctx context.Context, userID string) (bool, error)
RequestedVideoIDs(ctx context.Context, userID string) (map[string]bool, error)
ClearSummarizeRequested(ctx context.Context, userID, videoID string) error
// RateLimitedVideoIDs maps the user's still-throttled videos to when they were
// rate-limited, so the loop can back off without re-hitting the caption endpoint.
RateLimitedVideoIDs(ctx context.Context, userID string) (map[string]time.Time, error)
// SetTranscriptStatus records the outcome of a transcript attempt: "none",
// "rate_limited" (stamps the backoff clock), or "fetched".
SetTranscriptStatus(ctx context.Context, userID, videoID, status string) error
// UpsertChannelError records a channel that returned HTTP 404 (deleted/private).
// Called when NewVideos returns domain.ErrChannelUnavailable; best-effort, errors
// are logged and never abort the pass.
UpsertChannelError(ctx context.Context, userID, channelID, channelTitle string) error
}
// Processor runs the core use case for a single video. *usecase.Engine
@@ -43,44 +85,99 @@ type Processor interface {
// Runner walks a user's subscriptions, persists each candidate video, skips the
// ones already summarized (durably), and processes the rest through the engine.
type Runner struct {
src ports.VideoSource
store VideoStore
engine Processor
userID string
log *slog.Logger
src ports.VideoSource
store VideoStore
engine Processor
userID string
log *slog.Logger
backoff time.Duration // rate-limit retry window; 0 = always retry
autoWindow time.Duration // recency bound for auto-summarize; 0 = no bound
now func() time.Time // injectable clock (tests); defaults to time.Now
}
// Option configures a Runner at construction. Variadic so existing call sites
// stay valid as new knobs (backoff, clock) are added.
type Option func(*Runner)
// WithBackoff sets the rate-limit retry window. A video that returned HTTP 429 is
// skipped (no caption fetch) until this much time has passed; 0 = always retry.
func WithBackoff(d time.Duration) Option { return func(r *Runner) { r.backoff = d } }
// WithClock overrides the clock used for backoff comparisons. Tests inject a
// fixed time; production leaves the time.Now default.
func WithClock(now func() time.Time) Option { return func(r *Runner) { r.now = now } }
// WithAutoWindow bounds auto-summarization to videos published within d of now.
// In auto mode a video older than d is discovered and listed but not summarized
// automatically — it waits for an explicit manual request — so a large
// back-catalogue does not self-inflict 429s. An explicitly requested video
// bypasses the bound. 0 (the default) disables it (summarize every unseen video).
func WithAutoWindow(d time.Duration) Option { return func(r *Runner) { r.autoWindow = d } }
// New builds a Runner. A nil logger falls back to slog.Default.
func New(src ports.VideoSource, store VideoStore, engine Processor, userID string, log *slog.Logger) *Runner {
func New(src ports.VideoSource, store VideoStore, engine Processor, userID string, log *slog.Logger, opts ...Option) *Runner {
if log == nil {
log = slog.Default()
}
return &Runner{src: src, store: store, engine: engine, userID: userID, log: log}
r := &Runner{src: src, store: store, engine: engine, userID: userID, log: log, now: time.Now}
for _, opt := range opts {
opt(r)
}
if r.now == nil {
r.now = time.Now
}
return r
}
// Stats summarizes one RunOnce pass.
type Stats struct {
Candidates int
Summarized int
SkippedSeen int
SkippedNoText int
SkippedManual int // discovered but not queued, in manual mode
Errors int
Candidates int
Summarized int
SkippedSeen int
SkippedNoText int
SkippedManual int // discovered but not queued, in manual mode
SkippedTooOld int // auto mode: published outside the recency window (not requested)
SkippedRateLimited int // 429'd previously and still inside the backoff window
Errors int
ChannelUnavailable int // channels that returned HTTP 404 (deleted/private)
}
// RunOnce performs a single pass over the user's subscriptions. Per-item errors
// are logged and collected (one bad video or channel does not abort the pass)
// and returned joined alongside the Stats gathered.
// tooOld reports whether a video published at publishedAt falls outside the
// auto-summarize recency window. A zero window disables the bound, and a zero
// publishedAt (undated video) is never aged out — it cannot be dated, so it is
// processed rather than silently stranded.
func (r *Runner) tooOld(publishedAt time.Time) bool {
if r.autoWindow <= 0 || publishedAt.IsZero() {
return false
}
return r.now().Sub(publishedAt) > r.autoWindow
}
// RunOnce performs a single pass over the user's subscriptions in three phases:
//
// 1. Discovery: walk all channels, persist each candidate video (UpsertVideo),
// apply pre-filters (seen/manual/backoff) — same as before.
// 2. Sort: order the surviving candidates newest-first (published_at DESC, NULLS
// LAST) so new users get summaries of their most recent, relevant videos first;
// the back-catalogue fills in behind across subsequent passes.
// 3. Process: feed candidates to the engine in sorted order through the shared
// globalFetchGate — the gate is unchanged and still governs honest rate pacing.
//
// All existing behaviour is preserved: per-item failure isolation, the rate-limit
// backoff skip, manual mode, channel-unavailable handling, and stats accounting.
// Only the processing order changes within a pass.
func (r *Runner) RunOnce(ctx context.Context) (Stats, error) {
var (
stats Stats
errs []error
stats Stats
errs []error
candidates []passCandidate
pos int
)
// Throttle transcript fetches: unauthenticated caption scraping gets
// soft-throttled by YouTube under heavy back-to-back volume (captionTracks
// silently stripped from the player response). A small per-video delay keeps
// a full pass under the radar. TAPIR_FETCH_DELAY (Go duration), 0 = off.
// the fetch rate polite. TAPIR_FETCH_DELAY (Go duration), 0 = off.
fetchDelay, _ := time.ParseDuration(os.Getenv("TAPIR_FETCH_DELAY"))
seen, err := r.store.SeenVideoIDs(ctx, r.userID)
@@ -88,37 +185,62 @@ func (r *Runner) RunOnce(ctx context.Context) (Stats, error) {
return stats, fmt.Errorf("runner: load seen videos: %w", err)
}
// Summarization mode (per-user, ADR-012). Auto = summarize every unseen video
// (the original behavior). Manual = still discover/persist videos so the user
// sees them, but only summarize the ones explicitly queued via the web UI
// (summarize_requested). The queued set is loaded once per pass, like seen.
// Summarization mode (per-user, ADR-012). Auto = summarize every unseen video.
// Manual = discover/persist videos (visible in list) but only summarize ones
// explicitly queued via the web UI (summarize_requested).
auto, err := r.store.GetAutoSummarize(ctx, r.userID)
if err != nil {
return stats, fmt.Errorf("runner: load summarize mode: %w", err)
}
// requested is needed in manual mode (the queue) and in auto mode when a
// recency window is active (an explicit request bypasses the bound).
var requested map[string]bool
if !auto {
if !auto || r.autoWindow > 0 {
requested, err = r.store.RequestedVideoIDs(ctx, r.userID)
if err != nil {
return stats, fmt.Errorf("runner: load requested videos: %w", err)
}
}
// Rate-limit backoff: videos that 429'd on a prior pass, mapped to when.
// Within the backoff window they are skipped before any caption fetch.
// Disabled when backoff <= 0 ("always retry").
var rateLimited map[string]time.Time
if r.backoff > 0 {
rateLimited, err = r.store.RateLimitedVideoIDs(ctx, r.userID)
if err != nil {
return stats, fmt.Errorf("runner: load rate-limited videos: %w", err)
}
}
subs, err := r.src.ListSubscriptions(ctx, r.userID)
if err != nil {
return stats, fmt.Errorf("runner: list subscriptions: %w", err)
}
// ── Phase 1: discover, persist, filter ───────────────────────────────────
// Walk all channels. Persist every video (UpsertVideo) so it appears in the
// list regardless of whether it will be summarized this pass. Apply pre-filters
// and collect surviving candidates with their discovery position.
for _, sub := range subs {
vids, err := r.src.NewVideos(ctx, sub)
if err != nil {
errs = append(errs, fmt.Errorf("new videos for %q: %w", sub.ChannelTitle, err))
stats.Errors++
var unavail *domain.ErrChannelUnavailable
if errors.As(err, &unavail) {
stats.ChannelUnavailable++
r.log.Warn("channel unavailable (playlist 404)", "channel", sub.ChannelTitle, "channel_id", sub.ChannelID)
if storeErr := r.store.UpsertChannelError(ctx, r.userID, unavail.ChannelID, unavail.ChannelTitle); storeErr != nil {
r.log.Warn("failed to store channel error", "err", storeErr)
}
} else {
errs = append(errs, fmt.Errorf("new videos for %q: %w", sub.ChannelTitle, err))
stats.Errors++
}
continue
}
for _, v := range vids {
stats.Candidates++
v.UserID = r.userID // keep the dedup/FK key consistent with config
v.UserID = r.userID
id, err := r.store.UpsertVideo(ctx, v)
if err != nil {
@@ -132,43 +254,89 @@ func (r *Runner) RunOnce(ctx context.Context) (Stats, error) {
stats.SkippedSeen++
continue
}
seen[id] = true // also guard against the same video within this pass
seen[id] = true // guard against duplicates within this pass
// Manual mode: skip summarization for videos the user has not queued.
// Discovery already happened (UpsertVideo above), so the new video is
// Discovery already happened (UpsertVideo above), so the video is
// visible in the list; it just isn't summarized until requested.
if !auto && !requested[id] {
stats.SkippedManual++
continue
}
if fetchDelay > 0 {
time.Sleep(fetchDelay)
}
res, err := r.engine.ProcessNewVideo(ctx, v)
if err != nil {
errs = append(errs, fmt.Errorf("process %q: %w", v.ProviderVideoID, err))
stats.Errors++
// Recency bound (auto mode): summarize only recent videos automatically;
// older ones are discovered + listed (UpsertVideo above) but wait for an
// explicit manual request, so a large back-catalogue does not self-inflict
// 429s against the caption rate gate. A requested video bypasses the bound.
if auto && !requested[id] && r.tooOld(v.PublishedAt) {
stats.SkippedTooOld++
continue
}
switch {
case res.Skipped:
stats.SkippedNoText++
r.log.Info("skipped video (no transcript)", "video", v.ProviderVideoID, "title", v.Title)
case res.Summary != nil:
stats.Summarized++
// In manual mode the video was processed because it was queued;
// clear the flag so it is not re-summarized and the UI drops the
// "Queued" chip. (Auto mode never sets the flag.)
if !auto {
if err := r.store.ClearSummarizeRequested(ctx, r.userID, id); err != nil {
errs = append(errs, fmt.Errorf("clear summarize flag %q: %w", v.ProviderVideoID, err))
stats.Errors++
}
}
r.log.Info("summarized video", "video", v.ProviderVideoID, "title", v.Title,
"provider", res.Summary.AIProvider, "model", res.Summary.AIModel)
// Still inside the rate-limit backoff window: skip without fetching.
if at, ok := rateLimited[id]; ok && r.now().Sub(at) < r.backoff {
stats.SkippedRateLimited++
r.log.Info("skipped video (rate-limited, backing off)", "video", v.ProviderVideoID, "title", v.Title)
continue
}
candidates = append(candidates, passCandidate{v: v, pos: pos})
pos++
}
}
// ── Phase 2: sort newest-first, NULLS LAST ────────────────────────────────
// Within this pass, process the newest videos first so a new user gets
// summaries of their most recent content quickly; the back-catalogue fills in
// behind across subsequent passes. Both this background batch and the foreground
// "Try now" button honour the shared globalFetchGate — ordering is onboarding
// prioritisation, not rate-limit evasion.
slices.SortStableFunc(candidates, compareNewestFirst)
// ── Phase 3: process in sorted order ─────────────────────────────────────
for _, c := range candidates {
if fetchDelay > 0 {
time.Sleep(fetchDelay)
}
res, err := r.engine.ProcessNewVideo(ctx, c.v)
if err != nil {
errs = append(errs, fmt.Errorf("process %q: %w", c.v.ProviderVideoID, err))
stats.Errors++
continue
}
id := c.v.ID
switch {
case res.Skipped && res.TranscriptSource == string(domain.SourceRateLimited):
// Fresh 429: stamp the backoff clock so the next pass skips it.
stats.SkippedRateLimited++
if err := r.store.SetTranscriptStatus(ctx, r.userID, id, "rate_limited"); err != nil {
errs = append(errs, fmt.Errorf("set rate_limited status %q: %w", c.v.ProviderVideoID, err))
stats.Errors++
}
r.log.Info("skipped video (rate-limited)", "video", c.v.ProviderVideoID, "title", c.v.Title)
case res.Skipped:
stats.SkippedNoText++
if err := r.store.SetTranscriptStatus(ctx, r.userID, id, "none"); err != nil {
errs = append(errs, fmt.Errorf("set none status %q: %w", c.v.ProviderVideoID, err))
stats.Errors++
}
r.log.Info("skipped video (no transcript)", "video", c.v.ProviderVideoID, "title", c.v.Title)
case res.Summary != nil:
stats.Summarized++
if err := r.store.SetTranscriptStatus(ctx, r.userID, id, "fetched"); err != nil {
errs = append(errs, fmt.Errorf("set fetched status %q: %w", c.v.ProviderVideoID, err))
stats.Errors++
}
// In manual mode the video was explicitly queued; clear the flag so
// it is not re-summarized and the UI drops the "Queued" chip.
if !auto {
if err := r.store.ClearSummarizeRequested(ctx, r.userID, id); err != nil {
errs = append(errs, fmt.Errorf("clear summarize flag %q: %w", c.v.ProviderVideoID, err))
stats.Errors++
}
}
r.log.Info("summarized video", "video", c.v.ProviderVideoID, "title", c.v.Title,
"provider", res.Summary.AIProvider, "model", res.Summary.AIModel)
}
}
@@ -184,7 +352,9 @@ func (r *Runner) Loop(ctx context.Context, interval time.Duration) error {
r.log.Info("run pass complete",
"candidates", stats.Candidates, "summarized", stats.Summarized,
"skipped_seen", stats.SkippedSeen, "skipped_no_text", stats.SkippedNoText,
"skipped_manual", stats.SkippedManual, "errors", stats.Errors)
"skipped_manual", stats.SkippedManual, "skipped_too_old", stats.SkippedTooOld,
"skipped_rate_limited", stats.SkippedRateLimited,
"channel_unavailable", stats.ChannelUnavailable, "errors", stats.Errors)
if err != nil {
r.log.Warn("run pass had errors", "err", err)
}
+264 -5
View File
@@ -5,6 +5,7 @@ import (
"io"
"log/slog"
"testing"
"time"
"github.com/stretchr/testify/require"
@@ -43,11 +44,13 @@ func (f *fakeSource) FetchTranscript(_ context.Context, v domain.Video) (domain.
// auto controls the summarization mode; requested is the manual-mode queue keyed
// by store id; cleared records the ids whose queue flag the runner reset.
type fakeStore struct {
seen map[string]bool
upserted []domain.Video
auto bool
requested map[string]bool
cleared []string
seen map[string]bool
upserted []domain.Video
auto bool
requested map[string]bool
cleared []string
rateLimited map[string]time.Time // id -> when 429'd (seeds the backoff window)
statuses map[string]string // id -> last SetTranscriptStatus value
}
func (f *fakeStore) UpsertVideo(_ context.Context, v domain.Video) (string, error) {
@@ -80,6 +83,24 @@ func (f *fakeStore) ClearSummarizeRequested(_ context.Context, _, videoID string
return nil
}
func (f *fakeStore) RateLimitedVideoIDs(_ context.Context, _ string) (map[string]time.Time, error) {
cp := make(map[string]time.Time, len(f.rateLimited))
for k, v := range f.rateLimited {
cp[k] = v
}
return cp, nil
}
func (f *fakeStore) UpsertChannelError(_ context.Context, _, _, _ string) error { return nil }
func (f *fakeStore) SetTranscriptStatus(_ context.Context, _, videoID, status string) error {
if f.statuses == nil {
f.statuses = map[string]string{}
}
f.statuses[videoID] = status
return nil
}
type fakeSummarizer struct{}
func (fakeSummarizer) Summarize(_ context.Context, v domain.Video, _ domain.Transcript) (domain.Summary, error) {
@@ -102,6 +123,12 @@ func vid(provID, title string) domain.Video {
return domain.Video{UserID: testUser, Provider: domain.ProviderYouTube, ProviderVideoID: provID, Title: title}
}
func vidAt(provID, title string, publishedAt time.Time) domain.Video {
v := vid(provID, title)
v.PublishedAt = publishedAt
return v
}
func quietLogger() *slog.Logger {
return slog.New(slog.NewTextHandler(io.Discard, nil))
}
@@ -207,6 +234,152 @@ func TestRunOnce_ManualMode_ProcessesRequested(t *testing.T) {
require.Equal(t, []string{"id-v1"}, st.cleared, "the queue flag is cleared after summarizing")
}
// --- recency window (B1) ---------------------------------------------------
// TestRunOnce_AutoMode_SkipsOldVideos: with a recency window set, auto mode
// summarizes only videos published within the window; older ones are discovered
// (upserted) but not auto-summarized — they wait for a manual request.
func TestRunOnce_AutoMode_SkipsOldVideos(t *testing.T) {
base := time.Date(2026, 6, 8, 12, 0, 0, 0, time.UTC)
src := &fakeSource{
subs: []domain.Subscription{sub("chan1", "Channel One")},
videos: map[string][]domain.Video{"chan1": {
vidAt("recent", "Recent", base.Add(-24*time.Hour)), // 1d old → in window
vidAt("old", "Old", base.Add(-30*24*time.Hour)), // 30d old → out of window
}},
}
st := &fakeStore{seen: map[string]bool{}, auto: true}
sink := &recordingSink{}
eng := usecase.NewEngine(src, fakeSummarizer{}, sink)
r := runner.New(src, st, eng, testUser, quietLogger(),
runner.WithAutoWindow(7*24*time.Hour), runner.WithClock(func() time.Time { return base }))
stats, err := r.RunOnce(context.Background())
require.NoError(t, err)
require.Equal(t, 1, stats.Summarized, "only the recent video is auto-summarized")
require.Equal(t, 1, stats.SkippedTooOld, "the old video is skipped by the recency bound")
require.Len(t, sink.delivered, 1)
require.Equal(t, "id-recent", sink.delivered[0].VideoID)
require.Len(t, st.upserted, 2, "both videos are still discovered and listed")
}
// TestRunOnce_AutoMode_OldVideoRequestedBypassesWindow: an explicit manual
// request (summarize_requested) overrides the recency bound even in auto mode.
func TestRunOnce_AutoMode_OldVideoRequestedBypassesWindow(t *testing.T) {
base := time.Date(2026, 6, 8, 12, 0, 0, 0, time.UTC)
src := &fakeSource{
subs: []domain.Subscription{sub("chan1", "Channel One")},
videos: map[string][]domain.Video{"chan1": {vidAt("old", "Old", base.Add(-30*24*time.Hour))}},
}
st := &fakeStore{seen: map[string]bool{}, auto: true, requested: map[string]bool{"id-old": true}}
sink := &recordingSink{}
eng := usecase.NewEngine(src, fakeSummarizer{}, sink)
r := runner.New(src, st, eng, testUser, quietLogger(),
runner.WithAutoWindow(7*24*time.Hour), runner.WithClock(func() time.Time { return base }))
stats, err := r.RunOnce(context.Background())
require.NoError(t, err)
require.Equal(t, 1, stats.Summarized, "a requested old video is summarized despite the window")
require.Equal(t, 0, stats.SkippedTooOld)
require.Len(t, sink.delivered, 1)
}
// TestRunOnce_AutoWindowZero_SummarizesOld: a zero window disables the bound —
// the pre-recency behaviour (summarize every unseen video) is preserved.
func TestRunOnce_AutoWindowZero_SummarizesOld(t *testing.T) {
base := time.Date(2026, 6, 8, 12, 0, 0, 0, time.UTC)
src := &fakeSource{
subs: []domain.Subscription{sub("chan1", "Channel One")},
videos: map[string][]domain.Video{"chan1": {vidAt("old", "Old", base.Add(-365*24*time.Hour))}},
}
st := &fakeStore{seen: map[string]bool{}, auto: true}
sink := &recordingSink{}
eng := usecase.NewEngine(src, fakeSummarizer{}, sink)
r := runner.New(src, st, eng, testUser, quietLogger(),
runner.WithClock(func() time.Time { return base })) // no WithAutoWindow → 0
stats, err := r.RunOnce(context.Background())
require.NoError(t, err)
require.Equal(t, 1, stats.Summarized, "window disabled → old video summarized")
require.Equal(t, 0, stats.SkippedTooOld)
}
// TestRunOnce_AutoMode_UndatedVideoSummarized: a video with no published_at
// cannot be aged out — it is processed, not silently stranded.
func TestRunOnce_AutoMode_UndatedVideoSummarized(t *testing.T) {
base := time.Date(2026, 6, 8, 12, 0, 0, 0, time.UTC)
src := &fakeSource{
subs: []domain.Subscription{sub("chan1", "Channel One")},
videos: map[string][]domain.Video{"chan1": {vid("undated", "Undated")}}, // zero PublishedAt
}
st := &fakeStore{seen: map[string]bool{}, auto: true}
sink := &recordingSink{}
eng := usecase.NewEngine(src, fakeSummarizer{}, sink)
r := runner.New(src, st, eng, testUser, quietLogger(),
runner.WithAutoWindow(7*24*time.Hour), runner.WithClock(func() time.Time { return base }))
stats, err := r.RunOnce(context.Background())
require.NoError(t, err)
require.Equal(t, 1, stats.Summarized, "an undated video is processed, not aged out")
require.Equal(t, 0, stats.SkippedTooOld)
}
// noFetchSource fails the test if a transcript fetch happens — used to prove the
// runner skips a rate-limited video before touching the caption endpoint.
type noFetchSource struct{ *fakeSource }
func (noFetchSource) FetchTranscript(context.Context, domain.Video) (domain.Transcript, error) {
panic("FetchTranscript must not be called for a rate-limited video within the backoff window")
}
func TestRunOnce_SkipsRateLimitedWithinBackoff(t *testing.T) {
base := time.Date(2026, 6, 3, 12, 0, 0, 0, time.UTC)
src := &fakeSource{
subs: []domain.Subscription{sub("chan1", "Channel One")},
videos: map[string][]domain.Video{"chan1": {vid("v1", "Video 1")}},
}
// v1 was rate-limited 5m ago; backoff is 1h, so it is still inside the window.
st := &fakeStore{
seen: map[string]bool{},
auto: true,
rateLimited: map[string]time.Time{"id-v1": base.Add(-5 * time.Minute)},
}
eng := usecase.NewEngine(noFetchSource{src}, fakeSummarizer{}, &recordingSink{})
r := runner.New(noFetchSource{src}, st, eng, testUser, quietLogger(),
runner.WithBackoff(time.Hour), runner.WithClock(func() time.Time { return base }))
stats, err := r.RunOnce(context.Background())
require.NoError(t, err)
require.Equal(t, 1, stats.SkippedRateLimited, "still throttled -> skipped")
require.Equal(t, 0, stats.Summarized)
require.Empty(t, st.statuses, "no status write: the engine was never invoked")
}
func TestRunOnce_RetriesRateLimitedAfterBackoff(t *testing.T) {
base := time.Date(2026, 6, 3, 12, 0, 0, 0, time.UTC)
src := &fakeSource{
subs: []domain.Subscription{sub("chan1", "Channel One")},
videos: map[string][]domain.Video{"chan1": {vid("v1", "Video 1")}},
}
// v1 was rate-limited 2h ago; backoff is 1h, so the window has expired.
st := &fakeStore{
seen: map[string]bool{},
auto: true,
rateLimited: map[string]time.Time{"id-v1": base.Add(-2 * time.Hour)},
}
sink := &recordingSink{}
eng := usecase.NewEngine(src, fakeSummarizer{}, sink)
r := runner.New(src, st, eng, testUser, quietLogger(),
runner.WithBackoff(time.Hour), runner.WithClock(func() time.Time { return base }))
stats, err := r.RunOnce(context.Background())
require.NoError(t, err)
require.Equal(t, 0, stats.SkippedRateLimited, "window expired -> not skipped")
require.Equal(t, 1, stats.Summarized, "the video is retried and summarized")
require.Len(t, sink.delivered, 1)
require.Equal(t, "fetched", st.statuses["id-v1"], "status advances to fetched on success")
}
func TestRunOnce_UpsertsEveryCandidate(t *testing.T) {
src := &fakeSource{
subs: []domain.Subscription{sub("chan1", "Channel One")},
@@ -221,3 +394,89 @@ func TestRunOnce_UpsertsEveryCandidate(t *testing.T) {
require.NoError(t, err)
require.Len(t, st.upserted, 2, "every candidate is upserted, including seen ones")
}
// TestRunOnce_NewestFirstOrdering asserts that within a pass, candidates are
// processed newest-first (published_at DESC, NULLS LAST) across all channels,
// and that the set of processed videos is identical to what per-channel inline
// processing would produce (only the order differs).
//
// Fixture: two channels, four videos with mixed published_at (one NULL).
//
// Per-channel (before): chanA=[v-old, v-mid], chanB=[v-new, v-null]
// → [v-old, v-mid, v-new, v-null]
// Newest-first (after): [v-new, v-mid, v-old, v-null]
func TestRunOnce_NewestFirstOrdering(t *testing.T) {
old := time.Date(2024, 1, 1, 0, 0, 0, 0, time.UTC)
mid := time.Date(2024, 6, 1, 0, 0, 0, 0, time.UTC)
newt := time.Date(2024, 12, 1, 0, 0, 0, 0, time.UTC)
// zero time = NULL published_at (schema 001: nullable)
src := &fakeSource{
subs: []domain.Subscription{
sub("chanA", "Channel A"),
sub("chanB", "Channel B"),
},
videos: map[string][]domain.Video{
"chanA": {
vidAt("v-old", "Old Video", old),
vidAt("v-mid", "Mid Video", mid),
},
"chanB": {
vidAt("v-new", "New Video", newt),
vidAt("v-null", "No Date Video", time.Time{}), // NULL
},
},
}
st := &fakeStore{seen: map[string]bool{}, auto: true}
sink := &recordingSink{}
eng := usecase.NewEngine(src, fakeSummarizer{}, sink)
r := runner.New(src, st, eng, testUser, quietLogger())
stats, err := r.RunOnce(context.Background())
require.NoError(t, err)
// Same set: all 4 candidates processed regardless of order.
require.Equal(t, 4, stats.Candidates)
require.Equal(t, 4, stats.Summarized, "same set of videos processed as per-channel order")
require.Len(t, sink.delivered, 4)
// Build video-id → delivery-position map.
order := make(map[string]int, len(sink.delivered))
for i, s := range sink.delivered {
order[s.VideoID] = i
t.Logf("position %d: %s", i, s.VideoID)
}
require.Less(t, order["id-v-new"], order["id-v-mid"], "newest (Dec) before mid (Jun)")
require.Less(t, order["id-v-mid"], order["id-v-old"], "mid (Jun) before old (Jan)")
require.Less(t, order["id-v-old"], order["id-v-null"], "dated before NULL (NULLS LAST)")
}
// TestRunOnce_NewestFirstNullsOnly asserts that when all candidates have NULL
// published_at, discovery order (stable) is preserved as the tiebreak.
func TestRunOnce_NewestFirstNullsOnly(t *testing.T) {
src := &fakeSource{
subs: []domain.Subscription{
sub("chanA", "Channel A"),
sub("chanB", "Channel B"),
},
videos: map[string][]domain.Video{
"chanA": {vidAt("v1", "V1", time.Time{}), vidAt("v2", "V2", time.Time{})},
"chanB": {vidAt("v3", "V3", time.Time{})},
},
}
st := &fakeStore{seen: map[string]bool{}, auto: true}
sink := &recordingSink{}
eng := usecase.NewEngine(src, fakeSummarizer{}, sink)
r := runner.New(src, st, eng, testUser, quietLogger())
stats, err := r.RunOnce(context.Background())
require.NoError(t, err)
require.Equal(t, 3, stats.Summarized, "all null-date videos processed")
// Discovery order: chanA[v1, v2], chanB[v3] → [v1, v2, v3].
// All have NULL published_at so the sort is stable; discovery order must hold.
require.Equal(t, "id-v1", sink.delivered[0].VideoID)
require.Equal(t, "id-v2", sink.delivered[1].VideoID)
require.Equal(t, "id-v3", sink.delivered[2].VideoID)
}
+53 -6
View File
@@ -27,6 +27,13 @@ type Engine struct {
AI ports.Summarizer
Sinks []ports.Sink
// Transcripts, when set, is the shared transcript cache (ADR-021): the engine
// reads it before any caption fetch and writes resolved transcripts back, so
// re-analysis — the same user re-summarizing, or a second user with the same
// video — never re-touches YouTube (ADR-010/014). Optional: nil disables
// persistence (fetch every time), keeping the pure-core/scaffold wiring valid.
Transcripts ports.TranscriptStore
// processed dedups videos within this engine's lifetime so a video is not
// summarized twice when the watcher sees it again. Durable cross-restart
// dedup is the store's concern (a resolved TRANSCRIPT / existing SUMMARY,
@@ -44,22 +51,29 @@ func NewEngine(src ports.VideoSource, ai ports.Summarizer, sinks ...ports.Sink)
type ProcessResult struct {
Video domain.Video
Skipped bool
Reason string // set when Skipped (e.g. "no transcript")
Summary *domain.Summary // nil when Skipped
Reason string // set when Skipped (e.g. "no transcript")
// TranscriptSource is how the transcript resolved (or that there was none):
// the domain.TranscriptSource value as a string. The runner reads it to tell a
// permanent absence (SourceNone) from a transient 429 (SourceRateLimited) and
// persist the right transcript_status. Empty when a fetch error short-circuits.
TranscriptSource string
Summary *domain.Summary // nil when Skipped
}
// ProcessNewVideo runs the core use case for a single video:
// resolve transcript -> (summarize -> deliver) | skip.
// See docs/use-cases/summarize_new_video.feature.
func (e *Engine) ProcessNewVideo(ctx context.Context, v domain.Video) (ProcessResult, error) {
t, err := e.Source.FetchTranscript(ctx, v)
t, err := e.resolveTranscript(ctx, v)
if err != nil {
return ProcessResult{Video: v}, fmt.Errorf("fetch transcript: %w", err)
return ProcessResult{Video: v}, err
}
if !t.HasText() {
// No usable transcript: record the skip, produce no summary, deliver nothing
// (captions-first, ADR-007; the watcher uses this to avoid reprocessing).
return ProcessResult{Video: v, Skipped: true, Reason: "no transcript"}, nil
// Surface the source so the runner separates SourceNone (permanent) from
// SourceRateLimited (retry after a backoff window).
return ProcessResult{Video: v, Skipped: true, Reason: "no transcript", TranscriptSource: string(t.Source)}, nil
}
sum, err := e.AI.Summarize(ctx, v, t)
@@ -76,7 +90,40 @@ func (e *Engine) ProcessNewVideo(ctx context.Context, v domain.Video) (ProcessRe
}
}
return ProcessResult{Video: v, Summary: &sum}, errors.Join(errs...)
return ProcessResult{Video: v, Summary: &sum, TranscriptSource: string(t.Source)}, errors.Join(errs...)
}
// resolveTranscript returns v's transcript, reading the shared store first
// (ADR-021): a stored transcript — including a stored SourceNone (captions
// permanently absent) — is returned without touching YouTube, so re-analysis
// never re-fetches. On a store miss it fetches through the source (which gates
// the caption call, ADR-014) and persists the terminal outcome so the next
// analysis, for any user, reads from the store. A transient SourceRateLimited is
// returned to the caller (the runner stamps a per-user backoff) but never stored,
// so persistence can never mask a 429 as a permanent "no transcript". When no
// TranscriptStore is wired the engine simply fetches every time.
func (e *Engine) resolveTranscript(ctx context.Context, v domain.Video) (domain.Transcript, error) {
if e.Transcripts != nil {
stored, ok, err := e.Transcripts.GetTranscript(ctx, string(v.Provider), v.ProviderVideoID)
if err != nil {
return domain.Transcript{}, fmt.Errorf("get stored transcript: %w", err)
}
if ok {
return stored, nil
}
}
t, err := e.Source.FetchTranscript(ctx, v)
if err != nil {
return domain.Transcript{}, fmt.Errorf("fetch transcript: %w", err)
}
if e.Transcripts != nil && t.Source != domain.SourceRateLimited {
if err := e.Transcripts.SaveTranscript(ctx, string(v.Provider), v.ProviderVideoID, t); err != nil {
return domain.Transcript{}, fmt.Errorf("save transcript: %w", err)
}
}
return t, nil
}
// ProcessNewVideos walks a user's subscriptions and processes each newly seen
+187
View File
@@ -0,0 +1,187 @@
package usecase
import (
"context"
"testing"
"gitea.d-ma.be/mathias/tapir/internal/domain"
)
// These tests pin the ADR-021 read-stored-first behaviour at the engine core:
// a stored transcript is summarized without re-touching the source, a miss
// fetches once and persists, and a transient rate-limit is never cached.
type recordingSource struct {
transcript domain.Transcript
fetchCalls int
}
func (s *recordingSource) ListSubscriptions(context.Context, string) ([]domain.Subscription, error) {
return nil, nil
}
func (s *recordingSource) NewVideos(context.Context, domain.Subscription) ([]domain.Video, error) {
return nil, nil
}
func (s *recordingSource) FetchTranscript(context.Context, domain.Video) (domain.Transcript, error) {
s.fetchCalls++
return s.transcript, nil
}
type fakeTranscriptStore struct {
stored map[string]domain.Transcript
saves int
}
func newFakeTranscriptStore() *fakeTranscriptStore {
return &fakeTranscriptStore{stored: make(map[string]domain.Transcript)}
}
func (f *fakeTranscriptStore) key(provider, id string) string { return provider + "|" + id }
func (f *fakeTranscriptStore) GetTranscript(_ context.Context, provider, id string) (domain.Transcript, bool, error) {
t, ok := f.stored[f.key(provider, id)]
return t, ok, nil
}
func (f *fakeTranscriptStore) SaveTranscript(_ context.Context, provider, id string, t domain.Transcript) error {
f.saves++
f.stored[f.key(provider, id)] = t
return nil
}
type countingSummarizer struct{ calls int }
func (c *countingSummarizer) Summarize(_ context.Context, v domain.Video, _ domain.Transcript) (domain.Summary, error) {
c.calls++
return domain.Summary{VideoID: v.ID, UserID: v.UserID, Summary: "s", AIProvider: "local"}, nil
}
type nopSink struct{}
func (nopSink) Name() string { return "nop" }
func (nopSink) Deliver(context.Context, domain.Summary) error { return nil }
func testVideo() domain.Video {
return domain.Video{ID: "v1", UserID: "u1", Provider: domain.ProviderYouTube, ProviderVideoID: "yt1"}
}
func TestProcessNewVideo_StoredTranscriptSkipsFetch(t *testing.T) {
src := &recordingSource{}
ts := newFakeTranscriptStore()
ts.stored[ts.key("youtube", "yt1")] = domain.Transcript{Source: domain.SourceCaptions, Content: "stored words"}
sum := &countingSummarizer{}
eng := NewEngine(src, sum, nopSink{})
eng.Transcripts = ts
res, err := eng.ProcessNewVideo(context.Background(), testVideo())
if err != nil {
t.Fatalf("ProcessNewVideo: %v", err)
}
if src.fetchCalls != 0 {
t.Fatalf("stored transcript must not re-fetch from source; got %d fetches", src.fetchCalls)
}
if ts.saves != 0 {
t.Fatalf("a store hit must not re-save; got %d saves", ts.saves)
}
if sum.calls != 1 || res.Summary == nil {
t.Fatalf("expected a summary from the stored transcript; calls=%d summary=%v", sum.calls, res.Summary)
}
}
func TestProcessNewVideo_StoreMissFetchesAndPersists(t *testing.T) {
src := &recordingSource{transcript: domain.Transcript{Source: domain.SourceCaptions, Language: "en", Content: "fetched words"}}
ts := newFakeTranscriptStore()
sum := &countingSummarizer{}
eng := NewEngine(src, sum, nopSink{})
eng.Transcripts = ts
if _, err := eng.ProcessNewVideo(context.Background(), testVideo()); err != nil {
t.Fatalf("ProcessNewVideo: %v", err)
}
if src.fetchCalls != 1 {
t.Fatalf("a store miss must fetch exactly once; got %d", src.fetchCalls)
}
if ts.saves != 1 {
t.Fatalf("a fetched transcript must be persisted; got %d saves", ts.saves)
}
got, ok, _ := ts.GetTranscript(context.Background(), "youtube", "yt1")
if !ok || got.Content != "fetched words" {
t.Fatalf("persisted transcript not readable back: ok=%v content=%q", ok, got.Content)
}
}
// The second summarize of the same video reads the persisted transcript and does
// NOT re-fetch — the primary ADR-021 win, proven end to end at the engine.
func TestProcessNewVideo_SecondSummarizeDoesNotRefetch(t *testing.T) {
src := &recordingSource{transcript: domain.Transcript{Source: domain.SourceCaptions, Content: "words"}}
ts := newFakeTranscriptStore()
eng := NewEngine(src, &countingSummarizer{}, nopSink{})
eng.Transcripts = ts
for i := 0; i < 2; i++ {
if _, err := eng.ProcessNewVideo(context.Background(), testVideo()); err != nil {
t.Fatalf("pass %d: %v", i, err)
}
}
if src.fetchCalls != 1 {
t.Fatalf("the second summarize must reuse the stored transcript; got %d fetches", src.fetchCalls)
}
}
// A stored "no captions" outcome short-circuits before both fetch and summarize.
func TestProcessNewVideo_StoredNoneSkipsFetchAndSummarize(t *testing.T) {
src := &recordingSource{}
ts := newFakeTranscriptStore()
ts.stored[ts.key("youtube", "yt1")] = domain.Transcript{Source: domain.SourceNone}
sum := &countingSummarizer{}
eng := NewEngine(src, sum, nopSink{})
eng.Transcripts = ts
res, err := eng.ProcessNewVideo(context.Background(), testVideo())
if err != nil {
t.Fatalf("ProcessNewVideo: %v", err)
}
if !res.Skipped {
t.Fatal("a stored SourceNone must skip")
}
if src.fetchCalls != 0 || sum.calls != 0 {
t.Fatalf("stored none must neither fetch nor summarize; fetches=%d calls=%d", src.fetchCalls, sum.calls)
}
}
// A transient 429 is surfaced (so the runner backs off per-user) but never cached
// as a shared terminal state — otherwise it would mask a rate-limit as permanent.
func TestProcessNewVideo_RateLimitedIsNotPersisted(t *testing.T) {
src := &recordingSource{transcript: domain.Transcript{Source: domain.SourceRateLimited}}
ts := newFakeTranscriptStore()
eng := NewEngine(src, &countingSummarizer{}, nopSink{})
eng.Transcripts = ts
res, err := eng.ProcessNewVideo(context.Background(), testVideo())
if err != nil {
t.Fatalf("ProcessNewVideo: %v", err)
}
if !res.Skipped || res.TranscriptSource != string(domain.SourceRateLimited) {
t.Fatalf("expected a rate-limited skip; skipped=%v source=%q", res.Skipped, res.TranscriptSource)
}
if ts.saves != 0 {
t.Fatalf("a transient rate-limit must not be persisted; got %d saves", ts.saves)
}
}
// With no TranscriptStore wired the engine fetches every time (back-compat).
func TestProcessNewVideo_NilStoreFetchesEveryTime(t *testing.T) {
src := &recordingSource{transcript: domain.Transcript{Source: domain.SourceCaptions, Content: "words"}}
eng := NewEngine(src, &countingSummarizer{}, nopSink{})
for i := 0; i < 2; i++ {
if _, err := eng.ProcessNewVideo(context.Background(), testVideo()); err != nil {
t.Fatalf("pass %d: %v", i, err)
}
}
if src.fetchCalls != 2 {
t.Fatalf("nil store must fetch every time; got %d", src.fetchCalls)
}
}
+6 -1
View File
@@ -33,7 +33,12 @@ func (a *App) handleAccount(w http.ResponseWriter, r *http.Request) {
a.serverError(w, r, "summarize mode", err)
return
}
a.render(w, r, AccountPage(name, email, conns, auto, takeFlash(w, r)))
channelErrs, err := a.Store.ListChannelErrors(r.Context(), userID)
if err != nil {
a.serverError(w, r, "channel errors", err)
return
}
a.render(w, r, AccountPage(name, email, conns, auto, channelErrs, takeFlash(w, r)))
}
// handleDisconnect removes a provider connection: it deletes the OAuth token from
+19
View File
@@ -23,6 +23,15 @@ type Connections interface {
UpsertConnection(ctx context.Context, userID string, c store.Connection) error
}
// DiscoveryTrigger requests an out-of-band discovery pass for a user. The connect
// flow fires it the moment a YouTube account is linked so videos appear promptly
// instead of waiting for the next scheduled pass (#6). Enqueue must be
// non-blocking and safe to call from the request goroutine; the implementation
// owns serialization with the scheduler (one pass at a time). nil = no trigger.
type DiscoveryTrigger interface {
Enqueue(userID string)
}
// connectStateTTL bounds how long a generated CSRF state is valid between the
// connect redirect and the provider callback.
const connectStateTTL = 10 * time.Minute
@@ -43,6 +52,10 @@ type ConnectHandler struct {
Conns Connections
Log *slog.Logger
// Discovery, when set, is fired after a successful connect so the new
// connection's videos are discovered immediately (#6). Optional.
Discovery DiscoveryTrigger
states *connectStateStore
now func() time.Time
}
@@ -134,6 +147,12 @@ func (h *ConnectHandler) handleCallback(w http.ResponseWriter, r *http.Request)
return
}
// Discover this user's videos now rather than waiting for the next scheduled
// pass (#6). Non-blocking; the trigger serializes with the scheduler.
if h.Discovery != nil {
h.Discovery.Enqueue(userID)
}
setFlash(w, flashConnected)
http.Redirect(w, r, "/", http.StatusSeeOther)
}
+48
View File
@@ -45,6 +45,24 @@ func (c *fakeConns) UpsertConnection(_ context.Context, userID string, conn stor
return nil
}
// fakeTrigger records Enqueue calls so a test can assert connect fired discovery.
type fakeTrigger struct {
mu sync.Mutex
users []string
}
func (f *fakeTrigger) Enqueue(userID string) {
f.mu.Lock()
defer f.mu.Unlock()
f.users = append(f.users, userID)
}
func (f *fakeTrigger) seen() []string {
f.mu.Lock()
defer f.mu.Unlock()
return append([]string(nil), f.users...)
}
// tokenServer fakes Google's token endpoint, returning body for any POST.
func tokenServer(t *testing.T, body string) *httptest.Server {
t.Helper()
@@ -126,6 +144,36 @@ func TestCallbackExchangesAndRecordsConnection(t *testing.T) {
require.Equal(t, wantRef, conns.conn.TokenRef)
}
func TestCallbackTriggersDiscovery(t *testing.T) {
srv := tokenServer(t,
`{"access_token":"at","refresh_token":"rt-secret","token_type":"Bearer","expires_in":3600}`)
app := newConnectApp(t, srv.URL, &fakeWriter{}, &fakeConns{})
trig := &fakeTrigger{}
app.Connect.Discovery = trig
state := connectState(t, app)
rec := do(t, app, httptest.NewRequest(http.MethodGet,
"/oauth/youtube/callback?state="+state+"&code=the-code", nil))
require.Equal(t, http.StatusSeeOther, rec.Code)
require.Equal(t, []string{userID}, trig.seen(),
"a successful connect must trigger discovery for the connecting user")
}
func TestCallbackNoDiscoveryOnFailedConnect(t *testing.T) {
srv := tokenServer(t,
`{"access_token":"at","refresh_token":"rt","token_type":"Bearer","expires_in":3600}`)
app := newConnectApp(t, srv.URL, &fakeWriter{}, &fakeConns{})
trig := &fakeTrigger{}
app.Connect.Discovery = trig
// No state → CSRF reject → nothing connected, so no discovery.
rec := do(t, app, httptest.NewRequest(http.MethodGet,
"/oauth/youtube/callback?code=the-code", nil))
require.Equal(t, http.StatusBadRequest, rec.Code)
require.Empty(t, trig.seen(), "a failed connect must not trigger discovery")
}
func TestCallbackRejectsMissingState(t *testing.T) {
srv := tokenServer(t,
`{"access_token":"at","refresh_token":"rt","token_type":"Bearer","expires_in":3600}`)
+189 -8
View File
@@ -3,12 +3,16 @@ package web
import (
"context"
"errors"
"html/template"
"io"
"log/slog"
"net/http"
"time"
"github.com/a-h/templ"
"gitea.d-ma.be/mathias/tapir/internal/adapters/store"
"gitea.d-ma.be/mathias/tapir/internal/domain"
)
// Store is the read/write surface the web handlers depend on — a narrow port over
@@ -17,6 +21,9 @@ import (
// fake without a database.
type Store interface {
ListVideos(ctx context.Context, userID string, limit int) ([]store.SummaryRow, error)
// DistinctChannels lists the user's source channels — the options for the
// feed's channel multi-select filter.
DistinctChannels(ctx context.Context, userID string) ([]string, error)
GetSummaryByVideo(ctx context.Context, userID, videoID string) (*store.SummaryRow, error)
GetVideoRow(ctx context.Context, userID, videoID string) (*store.SummaryRow, error)
ActionsFor(ctx context.Context, userID string, videoIDs []string) (map[string][]string, error)
@@ -29,11 +36,27 @@ type Store interface {
SetAutoSummarize(ctx context.Context, userID string, enabled bool) error
RequestSummarize(ctx context.Context, userID, videoID string) error
// UpsertVideo persists a pasted video (idempotent on user+provider+video id,
// so it also dedups) and returns its durable store id.
UpsertVideo(ctx context.Context, v domain.Video) (string, error)
// Account management (the /account page, disconnect, delete-account).
ConnectionsForUser(ctx context.Context, userID string) ([]store.Connection, error)
DeleteConnection(ctx context.Context, userID, provider string) error
DeleteUser(ctx context.Context, userID string) error
DisplayName(ctx context.Context, userID string) (string, error)
// ListChannelErrors returns channels that returned HTTP 404 (deleted/private) on
// the most recent discovery pass, shown on the account page as a warning.
ListChannelErrors(ctx context.Context, userID string) ([]store.ChannelError, error)
// SetTranscriptStatus clears or updates a video's transcript backoff state.
// Used by handleRetryNow to clear rate_limited_at before immediate processing.
SetTranscriptStatus(ctx context.Context, userID, videoID, status string) error
// StampLogin records (throttled, one row per user per day) that the resolved
// user was active on this request — the read-side Stage-0 usage signal the
// registration gate stamps for every authenticated request.
StampLogin(ctx context.Context, userID string) error
}
// SecretRemover deletes secret material by its opaque ref. *secrets.FileStore
@@ -65,9 +88,38 @@ type App struct {
// background goroutine (the "Summarize" button kicks it off). Nil = queue-only:
// the button flips the DB flag and the next `tapir run` does the work.
Processor Processor
// Fetcher, when non-nil, resolves an arbitrary YouTube video id to metadata for
// the paste-a-URL flow (Feature 2). Nil = the /paste route is not mounted.
Fetcher VideoFetcher
// Processing tracks in-flight immediate summarizations so the status endpoint
// shows the animation until the summary lands. The zero value is ready to use.
Processing ProcessingSet
// RecencyWindow mirrors the auto-summarize recency bound: un-summarized videos
// published before now-RecencyWindow collapse into the "older videos"
// disclosure on the list, so the readable summaries are not buried (B3). Zero
// disables the collapse (everything stays inline).
RecencyWindow time.Duration
// Now is an injectable clock for the recency cutoff (tests fix it). Nil =
// time.Now.
Now func() time.Time
}
// now returns the App's clock (time.Now unless overridden for tests).
func (a *App) now() time.Time {
if a.Now != nil {
return a.Now()
}
return time.Now()
}
// recencyCutoff is the timestamp before which an un-summarized video counts as
// "older" and collapses into the disclosure. A zero RecencyWindow yields the zero
// time, which bucketRows treats as "collapse disabled".
func (a *App) recencyCutoff() time.Time {
if a.RecencyWindow <= 0 {
return time.Time{}
}
return a.now().Add(-a.RecencyWindow)
}
func (a *App) logger() *slog.Logger {
@@ -92,6 +144,10 @@ func (a *App) Router() http.Handler {
app.HandleFunc("GET /v/{videoId}", a.handleDetail)
app.HandleFunc("POST /v/{videoId}/action", a.handleAction)
app.HandleFunc("POST /v/{videoId}/summarize", a.handleRequestSummarize)
app.HandleFunc("POST /v/{videoId}/retry-now", a.handleRetryNow)
if a.Fetcher != nil {
app.HandleFunc("POST /paste", a.handlePaste)
}
app.HandleFunc("GET /v/{videoId}/status", a.handleStatus)
app.HandleFunc("GET /register", a.handleRegisterForm)
app.HandleFunc("POST /register", a.handleRegister)
@@ -143,23 +199,45 @@ func (a *App) handleList(w http.ResponseWriter, r *http.Request) {
}
q := r.URL.Query()
f := Filter{
Channel: q.Get("channel"),
From: q.Get("from"),
To: q.Get("to"),
Channels: nonEmptyStrings(q["channel"]),
From: q.Get("from"),
To: q.Get("to"),
OnlySummarized: q.Get("summarized") == "1",
}
rows, err := a.Store.ListVideos(r.Context(), userID, 0)
allRows, err := a.Store.ListVideos(r.Context(), userID, 0)
if err != nil {
a.serverError(w, r, "list videos", err)
return
}
rows = f.apply(rows)
stats := pipelineStats(allRows)
rows := f.apply(allRows)
buckets := bucketRows(rows, a.recencyCutoff())
if isHTMX(r) {
a.render(w, r, summaryList(rows))
// Channel options for the multi-select filter (the user's source channels).
channels, err := a.Store.DistinctChannels(r.Context(), userID)
if err != nil {
a.serverError(w, r, "distinct channels", err)
return
}
a.render(w, r, ListPage(rows, f, takeFlash(w, r)))
// hasConnected drives both the paste box (shown to ANY connected user, #2) and
// the empty-state copy (a fresh account with a connection but no discovery pass
// yet reads "connected, summaries land gradually" rather than "nothing here").
// Computed every render — not only when empty — so a user with videos still
// gets the paste box.
conns, err := a.Store.ConnectionsForUser(r.Context(), userID)
if err != nil {
a.serverError(w, r, "connections for user", err)
return
}
hasConnected := len(conns) > 0
if isHTMX(r) {
a.render(w, r, summaryList(buckets, hasConnected))
return
}
a.render(w, r, ListPage(buckets, f, stats, takeFlash(w, r), hasConnected, channels))
}
// handleDetail renders one summary in full (highlights, takeaways, action group).
@@ -267,6 +345,109 @@ func (a *App) handleRequestSummarize(w http.ResponseWriter, r *http.Request) {
a.render(w, r, VideoCard(*row))
}
// handlePaste handles "paste a YouTube URL" (Feature 2). It parses the video id,
// fetches metadata (Data API — ungated), upserts a subscription-less video row
// scoped to the user (idempotent, so it also dedups), and — if the video isn't
// already summarized — requests a summary and kicks off immediate processing
// through the SAME rate gate as the Summarize button. An explicit paste is a
// manual request, so it summarizes regardless of the recency window. A video that
// turns out to have no captions resolves to the honest "no transcript" terminal
// state via the engine (ADR-010), not an error here.
func (a *App) handlePaste(w http.ResponseWriter, r *http.Request) {
userID, ok := a.currentUserID(w, r)
if !ok {
return
}
videoID, err := parseYouTubeVideoID(r.FormValue("url"))
if err != nil {
a.pasteFailure(w, http.StatusBadRequest, "That doesn't look like a YouTube video link.")
return
}
v, err := a.Fetcher.FetchVideo(r.Context(), userID, videoID)
if errors.Is(err, domain.ErrVideoNotFound) {
a.pasteFailure(w, http.StatusNotFound, "That video couldn't be found — it may be private or removed.")
return
}
if err != nil {
a.serverError(w, r, "paste fetch", err)
return
}
id, err := a.Store.UpsertVideo(r.Context(), v)
if err != nil {
a.serverError(w, r, "paste upsert", err)
return
}
row, err := a.Store.GetVideoRow(r.Context(), userID, id)
if err != nil {
a.serverError(w, r, "paste get video", err)
return
}
// Dedup: already in the feed with a summary — surface the existing entry,
// don't re-summarize.
if row.Summarized {
a.render(w, r, VideoCard(*row))
return
}
// New or unsummarized: queue + (if a Processor is wired) summarize now, through
// the shared gate. RequestSummarize makes it durable even if the process dies.
if err := a.Store.RequestSummarize(r.Context(), userID, id); err != nil {
a.serverError(w, r, "paste request summarize", err)
return
}
if a.Processor != nil {
a.startProcessing(userID, id)
a.render(w, r, processingCard(*row))
return
}
a.render(w, r, VideoCard(*row))
}
// pasteFailure renders a minimal inline error fragment for the paste form (HTMX
// swaps it in). No templ dependency so it renders even on a bad-input fast path.
func (a *App) pasteFailure(w http.ResponseWriter, status int, msg string) {
w.Header().Set("Content-Type", "text/html; charset=utf-8")
w.WriteHeader(status)
_, _ = io.WriteString(w, `<p class="paste-error" role="alert">`+template.HTMLEscapeString(msg)+`</p>`)
}
// handleRetryNow handles the "Try now" button on rate-limited video cards. It
// clears the rate_limited_at backoff so the scheduler won't skip the video, then
// triggers an immediate ProcessVideo — same background path as handleRequestSummarize.
// The rate gate (globalFetchGate) still applies, so this is safe under concurrent use.
func (a *App) handleRetryNow(w http.ResponseWriter, r *http.Request) {
userID, ok := a.currentUserID(w, r)
if !ok {
return
}
videoID := r.PathValue("videoId")
// Clear the backoff so the scheduler won't skip this video on the next pass.
if err := a.Store.SetTranscriptStatus(r.Context(), userID, videoID, "none"); err != nil {
a.serverError(w, r, "clear rate limit", err)
return
}
if !isHTMX(r) {
http.Redirect(w, r, "/", http.StatusSeeOther)
return
}
row, err := a.Store.GetVideoRow(r.Context(), userID, videoID)
if err != nil {
a.serverError(w, r, "get video", err)
return
}
if a.Processor != nil {
a.startProcessing(userID, videoID)
a.render(w, r, processingCard(*row))
return
}
a.render(w, r, VideoCard(*row))
}
// startProcessing marks a video in-flight and summarizes it in the background.
// The goroutine uses a detached context — not the request's, which is cancelled
// when the handler returns — and clears the in-flight mark on completion. On
+85 -12
View File
@@ -72,7 +72,7 @@ func rawPool(t *testing.T) *pgxpool.Pool {
func truncateAll(t *testing.T, p *pgxpool.Pool) {
t.Helper()
_, err := p.Exec(context.Background(),
`TRUNCATE summary_actions, sink_deliveries, summaries, transcripts, videos, users CASCADE`)
`TRUNCATE login_events, summary_actions, sink_deliveries, summaries, transcripts, videos, users CASCADE`)
require.NoError(t, err)
}
@@ -207,13 +207,15 @@ func TestListChannelFilter(t *testing.T) {
resetDB(t, p)
require.NoError(t, deliver(ctx, app, videoX, "body x"))
seedVideo(t, p, videoX, "X Title", "https://x", time.Time{})
_, err := p.Exec(ctx, `UPDATE videos SET channel_title = 'Acme Channel' WHERE id = $1`, videoX)
require.NoError(t, err)
// Channel is "youtube" for seeded rows; a non-matching filter hides them.
rec := do(t, app, httptest.NewRequest(http.MethodGet, "/?channel=vimeo", nil))
// Selecting a different channel hides the row; selecting its channel shows it.
rec := do(t, app, httptest.NewRequest(http.MethodGet, "/?channel=Other+Channel", nil))
require.Equal(t, http.StatusOK, rec.Code)
require.NotContains(t, body(t, rec), "X Title")
rec = do(t, app, httptest.NewRequest(http.MethodGet, "/?channel=youtube", nil))
rec = do(t, app, httptest.NewRequest(http.MethodGet, "/?channel=Acme+Channel", nil))
require.Contains(t, body(t, rec), "X Title")
}
@@ -234,6 +236,18 @@ func TestDetailRendersHighlightsAndTakeaways(t *testing.T) {
require.Contains(t, html, "takeaway one")
require.Contains(t, html, `id="action-buttons"`, "action button group present")
require.Contains(t, html, "Watched")
require.Contains(t, html, `class="segmented"`, "watched/skipped render as a segmented control (UX review C5)")
require.Less(t, strings.Index(html, "Skipped"), strings.Index(html, "Saved"),
"Saved sits after the watched/skipped segment")
require.Contains(t, html, "← Summaries", "back link to the list (UX review C3)")
// The detail page leads with the attention-saving payload (UX review A8):
// Takeaways ("is it worth my time?") above Highlights above the full Summary.
takeaways := strings.Index(html, "takeaway one")
highlights := strings.Index(html, "highlight one")
summary := strings.Index(html, "the full summary body")
require.Less(t, takeaways, highlights, "Takeaways render before Highlights")
require.Less(t, highlights, summary, "Highlights render before the full Summary")
}
func TestDetailNotFound(t *testing.T) {
@@ -310,6 +324,64 @@ func TestListShowsSummarizeButtonForUnsummarized(t *testing.T) {
require.NotContains(t, html, "Queued", "not queued yet")
}
// TestListCollapsesOlderAndNoCaption verifies the feed IA (UX review B3/B4):
// summarized + recent un-summarized cards lead inline; older un-summarized
// videos collapse into a single disclosure; caption-less videos collapse into a
// one-line count instead of dead cards.
func TestListCollapsesOlderAndNoCaption(t *testing.T) {
ctx := context.Background()
app := newApp(t)
now := time.Date(2026, 6, 8, 12, 0, 0, 0, time.UTC)
app.RecencyWindow = 7 * 24 * time.Hour
app.Now = func() time.Time { return now }
p := rawPool(t)
resetDB(t, p)
const (
vRecent = "cccccccc-cccc-cccc-cccc-cccccccccccc" // 2d old, unsummarized → inline
vOld = "dddddddd-dddd-dddd-dddd-dddddddddddd" // 30d old, unsummarized → disclosure
vNoCap = "eeeeeeee-eeee-eeee-eeee-eeeeeeeeeeee" // caption-less → collapsed line
)
require.NoError(t, deliver(ctx, app, videoX, "body x")) // summarized, recent
seedVideo(t, p, videoX, "Summarized X", "https://x", now.Add(-24*time.Hour))
seedVideo(t, p, vRecent, "Recent Pending", "https://r", now.Add(-2*24*time.Hour))
seedVideo(t, p, vOld, "Old Pending", "https://o", now.Add(-30*24*time.Hour))
seedVideo(t, p, vNoCap, "No Caption Vid", "https://n", now.Add(-40*24*time.Hour))
_, err := p.Exec(ctx, `UPDATE videos SET transcript_status = 'none' WHERE id = $1`, vNoCap)
require.NoError(t, err)
rec := do(t, app, httptest.NewRequest(http.MethodGet, "/", nil))
require.Equal(t, http.StatusOK, rec.Code)
html := body(t, rec)
disclosure := strings.Index(html, "Show 1 older videos")
require.GreaterOrEqual(t, disclosure, 0, "older-videos disclosure present")
// Summarized + recent un-summarized lead inline, above the disclosure.
require.Less(t, strings.Index(html, "Summarized X"), disclosure, "summarized card is inline")
require.Less(t, strings.Index(html, "Recent Pending"), disclosure, "recent pending is inline")
// The older video is hidden inside the disclosure, after its summary.
require.Greater(t, strings.Index(html, "Old Pending"), disclosure, "older video lives in the disclosure")
// Caption-less video is a one-line count, never a card.
require.Contains(t, html, "have no captions")
require.NotContains(t, html, "No Caption Vid", "caption-less video is collapsed, not a card")
}
// TestListHidesFilterBarWhenEmpty: a genuinely empty list shows no filter bar
// (the connect CTA stands alone), but a filter that matches nothing still shows
// the bar so it can be cleared (UX review C1).
func TestListHidesFilterBarWhenEmpty(t *testing.T) {
app := newApp(t)
resetDB(t, rawPool(t))
rec := do(t, app, httptest.NewRequest(http.MethodGet, "/", nil))
require.Equal(t, http.StatusOK, rec.Code)
require.NotContains(t, body(t, rec), `class="filters"`, "no filter bar on an empty account")
rec = do(t, app, httptest.NewRequest(http.MethodGet, "/?channel=nope", nil))
require.Equal(t, http.StatusOK, rec.Code)
require.Contains(t, body(t, rec), `class="filters"`, "filtered-to-empty keeps the bar so it can be cleared")
}
func TestRequestSummarizeQueuesAndRendersCard(t *testing.T) {
ctx := context.Background()
app := newApp(t)
@@ -357,27 +429,28 @@ func TestSummarizeModeToggle(t *testing.T) {
app := newApp(t)
resetDB(t, rawPool(t))
// Account page defaults to manual.
// Account page defaults to automatic (ADR-018: onboarded users get
// zero-friction discovery — the list fills and summarizes itself).
rec := do(t, app, httptest.NewRequest(http.MethodGet, "/account", nil))
require.Equal(t, http.StatusOK, rec.Code)
html := body(t, rec)
require.Contains(t, html, "Manual", "default mode shown")
require.Contains(t, html, "Switch to automatic")
require.Contains(t, html, "Automatic", "default mode shown")
require.Contains(t, html, "Switch to manual")
// Toggle to automatic via HTMX returns the refreshed control.
// Toggle to manual via HTMX returns the refreshed control.
req := httptest.NewRequest(http.MethodPost, "/account/summarize-mode",
strings.NewReader("enabled=true"))
strings.NewReader("enabled=false"))
req.Header.Set("Content-Type", "application/x-www-form-urlencoded")
req.Header.Set("HX-Request", "true")
rec = do(t, app, req)
require.Equal(t, http.StatusOK, rec.Code)
html = body(t, rec)
require.Contains(t, html, "Automatic")
require.Contains(t, html, "Switch to manual")
require.Contains(t, html, "Manual")
require.Contains(t, html, "Switch to automatic")
got, err := app.Store.GetAutoSummarize(ctx, userID)
require.NoError(t, err)
require.True(t, got, "mode persisted")
require.False(t, got, "mode persisted")
}
func postSummarize(t *testing.T, app *web.App, videoID string, htmx bool) *httptest.ResponseRecorder {
+62
View File
@@ -0,0 +1,62 @@
package web
import (
"fmt"
"net/url"
"regexp"
"strings"
)
// youtubeVideoID matches a canonical YouTube video id: exactly 11 URL-safe chars.
var youtubeVideoID = regexp.MustCompile(`^[A-Za-z0-9_-]{11}$`)
// parseYouTubeVideoID extracts the 11-character video id from a pasted YouTube
// URL (watch?v=, youtu.be/, shorts/, embed/) or a bare id. It rejects non-YouTube
// hosts and anything that doesn't yield a valid id, so the paste flow never tries
// to fetch a video that can't exist (Feature 2).
func parseYouTubeVideoID(raw string) (string, error) {
s := strings.TrimSpace(raw)
if s == "" {
return "", fmt.Errorf("empty input")
}
// Bare id (no URL) — accept directly.
if youtubeVideoID.MatchString(s) {
return s, nil
}
// Accept scheme-less URLs (youtube.com/watch?v=...) by giving url.Parse a host.
if !strings.Contains(s, "://") {
s = "https://" + s
}
u, err := url.Parse(s)
if err != nil {
return "", fmt.Errorf("not a URL: %w", err)
}
host := strings.ToLower(u.Hostname())
isYouTube := host == "youtu.be" || host == "youtube.com" || strings.HasSuffix(host, ".youtube.com")
if !isYouTube {
return "", fmt.Errorf("not a YouTube URL: %q", host)
}
var id string
switch {
case host == "youtu.be":
// youtu.be/<id>
id = strings.Trim(u.Path, "/")
case u.Path == "/watch":
id = u.Query().Get("v")
default:
// /shorts/<id>, /embed/<id>
parts := strings.Split(strings.Trim(u.Path, "/"), "/")
if len(parts) == 2 && (parts[0] == "shorts" || parts[0] == "embed") {
id = parts[1]
}
}
if !youtubeVideoID.MatchString(id) {
return "", fmt.Errorf("no YouTube video id in %q", raw)
}
return id, nil
}
+134
View File
@@ -0,0 +1,134 @@
package web_test
import (
"context"
"net/http"
"net/http/httptest"
"net/url"
"strings"
"testing"
"time"
"github.com/stretchr/testify/require"
"gitea.d-ma.be/mathias/tapir/internal/domain"
)
// fakeFetcher is a web.VideoFetcher returning a fixed video (or an error),
// scoped to whatever (userID, videoID) the handler asks for.
type fakeFetcher struct {
title string
err error
calls int
}
func (f *fakeFetcher) FetchVideo(_ context.Context, userID, videoID string) (domain.Video, error) {
f.calls++
if f.err != nil {
return domain.Video{}, f.err
}
return domain.Video{
UserID: userID,
Provider: domain.ProviderYouTube,
ProviderVideoID: videoID,
Title: f.title,
URL: "https://www.youtube.com/watch?v=" + videoID,
}, nil
}
func pasteReq(rawURL string) *http.Request {
req := httptest.NewRequest(http.MethodPost, "/paste",
strings.NewReader("url="+url.QueryEscape(rawURL)))
req.Header.Set("Content-Type", "application/x-www-form-urlencoded")
return req
}
func TestPasteValidURLAddsAndRequests(t *testing.T) {
ctx := context.Background()
app := newApp(t)
resetDB(t, rawPool(t))
app.Fetcher = &fakeFetcher{title: "Pasted Talk"}
p := rawPool(t)
rec := do(t, app, pasteReq("https://youtu.be/dQw4w9WgXcQ"))
require.Equal(t, http.StatusOK, rec.Code)
var (
count, requested int
title string
)
require.NoError(t, p.QueryRow(ctx,
`SELECT count(*), coalesce(max(title),'') FROM videos
WHERE user_id=$1 AND provider_video_id='dQw4w9WgXcQ'`, userID).Scan(&count, &title))
require.Equal(t, 1, count, "pasted video added once, scoped to the user")
require.Equal(t, "Pasted Talk", title)
require.NoError(t, p.QueryRow(ctx,
`SELECT count(*) FROM videos
WHERE user_id=$1 AND provider_video_id='dQw4w9WgXcQ' AND summarize_requested`,
userID).Scan(&requested))
require.Equal(t, 1, requested, "pasted video is queued for summarization (through the gate)")
}
func TestPasteInvalidURLRejected(t *testing.T) {
app := newApp(t)
resetDB(t, rawPool(t))
app.Fetcher = &fakeFetcher{title: "x"}
rec := do(t, app, pasteReq("definitely not a url"))
require.Equal(t, http.StatusBadRequest, rec.Code)
var count int
require.NoError(t, rawPool(t).QueryRow(context.Background(),
`SELECT count(*) FROM videos WHERE user_id=$1`, userID).Scan(&count))
require.Equal(t, 0, count, "invalid input adds nothing")
}
func TestPasteVideoNotFound(t *testing.T) {
app := newApp(t)
resetDB(t, rawPool(t))
app.Fetcher = &fakeFetcher{err: domain.ErrVideoNotFound}
rec := do(t, app, pasteReq("https://youtu.be/dQw4w9WgXcQ"))
require.Equal(t, http.StatusNotFound, rec.Code)
var count int
require.NoError(t, rawPool(t).QueryRow(context.Background(),
`SELECT count(*) FROM videos WHERE user_id=$1`, userID).Scan(&count))
require.Equal(t, 0, count, "a not-found video adds nothing")
}
func TestPasteDedupNoDuplicate(t *testing.T) {
app := newApp(t)
resetDB(t, rawPool(t))
app.Fetcher = &fakeFetcher{title: "Pasted Talk"}
require.Equal(t, http.StatusOK, do(t, app, pasteReq("https://youtu.be/dQw4w9WgXcQ")).Code)
require.Equal(t, http.StatusOK, do(t, app, pasteReq("https://www.youtube.com/watch?v=dQw4w9WgXcQ")).Code)
var count int
require.NoError(t, rawPool(t).QueryRow(context.Background(),
`SELECT count(*) FROM videos WHERE user_id=$1 AND provider_video_id='dQw4w9WgXcQ'`,
userID).Scan(&count))
require.Equal(t, 1, count, "pasting the same video twice must not duplicate the row")
}
func TestListShowsPasteFormForConnectedUserWithVideos(t *testing.T) {
app := newApp(t)
resetDB(t, rawPool(t))
app.Fetcher = &fakeFetcher{title: "x"}
p := rawPool(t)
// Connected user with a non-empty feed (the case the bug missed: hasConnected
// was only computed for an empty feed).
_, err := p.Exec(context.Background(),
`INSERT INTO video_connections (user_id, provider, token_ref, status)
VALUES ($1, 'youtube', 'youtube/x/refresh_token', 'active')`, userID)
require.NoError(t, err)
seedVideo(t, p, "11111111-1111-1111-1111-111111111111", "A talk", "https://youtu.be/aaaaaaaaaaa", time.Now())
rec := do(t, app, httptest.NewRequest(http.MethodGet, "/", nil))
require.Equal(t, http.StatusOK, rec.Code)
require.Contains(t, body(t, rec), `action="/paste"`,
"a connected user must see the paste box even when the feed has videos")
}
+97
View File
@@ -0,0 +1,97 @@
package web
import (
"bytes"
"context"
"gitea.d-ma.be/mathias/tapir/internal/adapters/store"
"strings"
"testing"
)
func TestParseYouTubeVideoID(t *testing.T) {
const id = "dQw4w9WgXcQ"
ok := []struct {
name, in string
}{
{"watch", "https://www.youtube.com/watch?v=" + id},
{"watch no www", "https://youtube.com/watch?v=" + id},
{"watch m", "https://m.youtube.com/watch?v=" + id},
{"watch extra params", "https://www.youtube.com/watch?v=" + id + "&t=42s&list=PLxyz"},
{"watch param after", "https://www.youtube.com/watch?list=PLxyz&v=" + id},
{"short link", "https://youtu.be/" + id},
{"short link param", "https://youtu.be/" + id + "?si=abcd&t=1"},
{"shorts", "https://www.youtube.com/shorts/" + id},
{"embed", "https://www.youtube.com/embed/" + id},
{"bare id", id},
{"http scheme", "http://youtube.com/watch?v=" + id},
{"no scheme", "youtube.com/watch?v=" + id},
{"trailing space", " https://youtu.be/" + id + " "},
}
for _, c := range ok {
t.Run(c.name, func(t *testing.T) {
got, err := parseYouTubeVideoID(c.in)
if err != nil {
t.Fatalf("parseYouTubeVideoID(%q) error: %v", c.in, err)
}
if got != id {
t.Fatalf("parseYouTubeVideoID(%q) = %q, want %q", c.in, got, id)
}
})
}
bad := []struct {
name, in string
}{
{"empty", ""},
{"blank", " "},
{"vimeo", "https://vimeo.com/123456789"},
{"other host", "https://example.com/watch?v=" + id},
{"watch no id", "https://www.youtube.com/watch?v="},
{"short id", "https://youtu.be/abc"},
{"long id", "https://youtu.be/" + id + "extra"},
{"bad chars", "https://youtu.be/dQw4w9Wg!cQ"},
{"not a url", "just some text"},
{"channel url", "https://www.youtube.com/@somechannel"},
}
for _, c := range bad {
t.Run("reject "+c.name, func(t *testing.T) {
if got, err := parseYouTubeVideoID(c.in); err == nil {
t.Fatalf("parseYouTubeVideoID(%q) = %q, want error", c.in, got)
}
})
}
}
func TestListPageShowsPasteFormOnlyWhenConnected(t *testing.T) {
render := func(connected bool) string {
var buf bytes.Buffer
if err := ListPage(listBuckets{}, Filter{}, PipelineStats{}, "", connected, nil).Render(context.Background(), &buf); err != nil {
t.Fatalf("render: %v", err)
}
return buf.String()
}
html := render(true)
if !strings.Contains(html, `name="url"`) || !strings.Contains(html, `action="/paste"`) {
t.Errorf("connected feed must show the paste form")
}
if strings.Contains(render(false), `name="url"`) {
t.Errorf("disconnected feed must not show the paste form")
}
}
func TestFilterMatchesMultipleChannels(t *testing.T) {
f := Filter{Channels: []string{"Acme", "Zeta"}}
row := func(ch string) store.SummaryRow { return store.SummaryRow{ChannelTitle: ch, Summarized: true} }
rows := []store.SummaryRow{row("Acme"), row("Beta"), row("Zeta")}
got := f.apply(rows)
if len(got) != 2 || got[0].ChannelTitle != "Acme" || got[1].ChannelTitle != "Zeta" {
t.Fatalf("multi-channel filter = %+v, want Acme+Zeta only", got)
}
// Empty selection = no channel constraint (all pass).
if n := len(Filter{}.apply(rows)); n != 3 {
t.Fatalf("no channel filter should pass all rows, got %d", n)
}
}
+10
View File
@@ -3,6 +3,8 @@ package web
import (
"context"
"sync"
"gitea.d-ma.be/mathias/tapir/internal/domain"
)
// Processor runs the core summarization use case for a single already-discovered
@@ -14,6 +16,14 @@ type Processor interface {
ProcessVideo(ctx context.Context, userID, videoID string) error
}
// VideoFetcher resolves an arbitrary YouTube video id to its metadata for the
// paste-a-URL flow (Feature 2). It is a Data API call, NOT the rate-limited
// caption path. Returns domain.ErrVideoNotFound for a deleted/private/typo'd id.
// cmd/tapir wires a per-user YouTube adapter; nil disables the paste route.
type VideoFetcher interface {
FetchVideo(ctx context.Context, userID, videoID string) (domain.Video, error)
}
// ProcessingSet tracks the (user, video) ids currently being summarized in-process
// so the status endpoint can show the animation until the summary lands. It is
// ephemeral (single-instance Stage-1): a restart drops it, and the DB holds the
+45 -2
View File
@@ -7,6 +7,7 @@ import (
"strings"
"testing"
"github.com/jackc/pgx/v5/pgxpool"
"github.com/stretchr/testify/require"
)
@@ -45,7 +46,7 @@ func TestRegisterCreatesExactlyOneUserAndIdentity(t *testing.T) {
truncateAll(t, p)
req := httptest.NewRequest(http.MethodPost, "/register",
strings.NewReader("display_name=Newbie&accept_terms=yes"))
strings.NewReader("display_name=Newbie"))
req.Header.Set("Content-Type", "application/x-www-form-urlencoded")
rec := do(t, app, req)
require.Equal(t, http.StatusSeeOther, rec.Code)
@@ -74,12 +75,54 @@ func TestRegisterCreatesExactlyOneUserAndIdentity(t *testing.T) {
require.Equal(t, 1, totalIdents)
}
// loginEventCount counts login_events for the stub user via the raw pool.
func loginEventCount(t *testing.T, p *pgxpool.Pool) int {
t.Helper()
var n int
require.NoError(t, p.QueryRow(context.Background(),
`SELECT count(*) FROM login_events WHERE user_id = $1`, userID).Scan(&n))
return n
}
// TestGateStampsLoginEventThrottled: a gated request for a registered user stamps
// exactly one login event, and a same-day repeat is throttled to no new row — the
// read-side Stage-0 signal flowing from the registration gate.
func TestGateStampsLoginEventThrottled(t *testing.T) {
app := newApp(t) // stubSubject → userID
p := rawPool(t)
resetDB(t, p)
require.Equal(t, 0, loginEventCount(t, p))
rec := do(t, app, httptest.NewRequest(http.MethodGet, "/", nil))
require.Equal(t, http.StatusOK, rec.Code)
require.Equal(t, 1, loginEventCount(t, p), "a gated request must stamp one login event")
do(t, app, httptest.NewRequest(http.MethodGet, "/", nil))
require.Equal(t, 1, loginEventCount(t, p), "a same-day repeat must not stamp again")
}
// TestUnregisteredSubjectIsNotStamped: a subject with no tapir user is redirected
// to /register and never reaches the stamp (no user_id to attribute it to).
func TestUnregisteredSubjectIsNotStamped(t *testing.T) {
app := newAppAs(t, "unregistered-sub")
p := rawPool(t)
truncateAll(t, p)
rec := do(t, app, httptest.NewRequest(http.MethodGet, "/", nil))
require.Equal(t, http.StatusFound, rec.Code)
var n int
require.NoError(t, p.QueryRow(context.Background(),
`SELECT count(*) FROM login_events`).Scan(&n))
require.Equal(t, 0, n, "an unregistered subject must not stamp a login event")
}
func TestRegisterRejectsMissingFields(t *testing.T) {
app := newAppAs(t, "incomplete-subject")
truncateAll(t, rawPool(t))
req := httptest.NewRequest(http.MethodPost, "/register",
strings.NewReader("display_name=&accept_terms=")) // both missing
strings.NewReader("display_name=")) // missing display name
req.Header.Set("Content-Type", "application/x-www-form-urlencoded")
rec := do(t, app, req)
require.Equal(t, http.StatusBadRequest, rec.Code)
+9 -3
View File
@@ -64,6 +64,13 @@ func (a *App) registrationGate(h http.Handler) http.Handler {
http.Redirect(w, r, registerPath, http.StatusFound)
return
}
// Stamp the read-side usage signal (Stage-0 gate, ADR-016): the store
// throttles this to one row per user per day, so a stamp on every gated
// request is cheap. Best-effort — a stamp failure must never break the
// request the user actually came for, so it is logged and swallowed.
if err := a.Store.StampLogin(r.Context(), userID); err != nil {
a.logger().Warn("stamp login event", "user", userID, "err", err)
}
h.ServeHTTP(w, r.WithContext(withUserID(r.Context(), userID)))
})
}
@@ -108,10 +115,9 @@ func (a *App) handleRegister(w http.ResponseWriter, r *http.Request) {
return
}
displayName := strings.TrimSpace(r.FormValue("display_name"))
accepted := r.FormValue("accept_terms") != ""
if displayName == "" || !accepted {
if displayName == "" {
a.renderStatus(w, r, http.StatusBadRequest,
RegisterPage(user.Email, "Enter a display name and accept the terms to continue."))
RegisterPage(user.Email, "Enter a display name to continue."))
return
}
+158
View File
@@ -0,0 +1,158 @@
package web
import (
"context"
"strings"
"testing"
"gitea.d-ma.be/mathias/tapir/internal/adapters/store"
)
func renderVideoCard(t *testing.T, r store.SummaryRow) string {
t.Helper()
var sb strings.Builder
if err := VideoCard(r).Render(context.Background(), &sb); err != nil {
t.Fatalf("render VideoCard: %v", err)
}
return sb.String()
}
// noActionButton asserts the rendered HTML contains no form POST (no action URL
// of the form /v/.../summarize or /v/.../retry-now) and no btn-quiet button.
func noActionButton(t *testing.T, html, state string) {
t.Helper()
if strings.Contains(html, "/summarize") {
t.Errorf("state %q: expected no summarize URL, got:\n%s", state, html)
}
if strings.Contains(html, "/retry-now") {
t.Errorf("state %q: expected no retry-now URL, got:\n%s", state, html)
}
if strings.Contains(html, "btn-quiet") {
t.Errorf("state %q: expected no action button, got:\n%s", state, html)
}
}
// TestVideoCard_State1_Summarized — chip + no nudge button.
func TestVideoCard_State1_Summarized(t *testing.T) {
html := renderVideoCard(t, store.SummaryRow{
VideoID: "aaaaaaaa-aaaa-aaaa-aaaa-aaaaaaaaaaaa",
Title: "Done Video",
Summarized: true,
Summary: "A great talk about Go.",
AIProvider: "local",
})
if !strings.Contains(html, "local") {
t.Errorf("state 1: expected AI provider chip, got:\n%s", html)
}
noActionButton(t, html, "summarized")
if strings.Contains(html, "Summarize") {
t.Errorf("state 1: no nudge button on a summarized card, got:\n%s", html)
}
}
// TestVideoCard_State2_NoTranscript — terminal; muted status, NO button, NO POST URL.
func TestVideoCard_State2_NoTranscript(t *testing.T) {
html := renderVideoCard(t, store.SummaryRow{
VideoID: "aaaaaaaa-aaaa-aaaa-aaaa-aaaaaaaaaaaa",
Title: "Silent Video",
Summarized: false,
TranscriptStatus: "none",
})
if !strings.Contains(html, "No transcript available") {
t.Errorf("state 2: expected 'No transcript available' text, got:\n%s", html)
}
noActionButton(t, html, "none-transcript")
if strings.Contains(html, "Summarize") {
t.Errorf("state 2: no nudge button when there are no captions, got:\n%s", html)
}
}
// TestVideoCard_State3_Queued — "Queued" chip, no button.
func TestVideoCard_State3_Queued(t *testing.T) {
html := renderVideoCard(t, store.SummaryRow{
VideoID: "aaaaaaaa-aaaa-aaaa-aaaa-aaaaaaaaaaaa",
Title: "Queued Video",
Summarized: false,
SummarizeRequested: true,
})
if !strings.Contains(html, "Queued") {
t.Errorf("state 3: expected 'Queued' chip, got:\n%s", html)
}
noActionButton(t, html, "queued")
}
// TestVideoCard_State4_RateLimited — quiet status + "Summarize" → retry-now URL.
func TestVideoCard_State4_RateLimited(t *testing.T) {
html := renderVideoCard(t, store.SummaryRow{
VideoID: "aaaaaaaa-aaaa-aaaa-aaaa-aaaaaaaaaaaa",
Title: "Throttled Video",
Summarized: false,
TranscriptStatus: "rate_limited",
})
if !strings.Contains(html, "In queue") {
t.Errorf("state 4: expected 'In queue' status text, got:\n%s", html)
}
if !strings.Contains(html, "Summarize") {
t.Errorf("state 4: expected 'Summarize' button, got:\n%s", html)
}
if !strings.Contains(html, "retry-now") {
t.Errorf("state 4: expected retry-now URL in form action, got:\n%s", html)
}
if strings.Contains(html, "/summarize\"") {
t.Errorf("state 4: rate-limited card must not POST to /summarize, got:\n%s", html)
}
if strings.Contains(html, "Try now") {
t.Errorf("state 4: 'Try now' verb must not appear, got:\n%s", html)
}
}
// TestVideoCard_State5_Pending — quiet status + "Summarize" → summarize URL.
func TestVideoCard_State5_Pending(t *testing.T) {
html := renderVideoCard(t, store.SummaryRow{
VideoID: "aaaaaaaa-aaaa-aaaa-aaaa-aaaaaaaaaaaa",
Title: "Fresh Video",
Summarized: false,
})
if !strings.Contains(html, "Not summarized") {
t.Errorf("state 5: expected 'Not summarized' status text, got:\n%s", html)
}
if !strings.Contains(html, "Summarize") {
t.Errorf("state 5: expected 'Summarize' button, got:\n%s", html)
}
if !strings.Contains(html, "/summarize") {
t.Errorf("state 5: expected summarize URL in form action, got:\n%s", html)
}
if strings.Contains(html, "retry-now") {
t.Errorf("state 5: pending card must not POST to /retry-now, got:\n%s", html)
}
if strings.Contains(html, "Try now") {
t.Errorf("state 5: 'Try now' verb must not appear, got:\n%s", html)
}
}
// TestVideoCard_ForbiddenCopyAbsent — the over-promising / jargon phrases the
// honesty pass removed must not reappear in any rendered card state (UX review
// A3/A4/A6): "Try now", "Summarize now", "Fetching soon", "the next run".
func TestVideoCard_ForbiddenCopyAbsent(t *testing.T) {
forbidden := []string{"Try now", "Summarize now", "Fetching soon", "the next run"}
cases := []store.SummaryRow{
{VideoID: "a", Summarized: true, Summary: "s", AIProvider: "local"},
{VideoID: "b", TranscriptStatus: "none"},
{VideoID: "c", SummarizeRequested: true},
{VideoID: "d", TranscriptStatus: "rate_limited"},
{VideoID: "e"},
}
for _, r := range cases {
html := renderVideoCard(t, r)
for _, phrase := range forbidden {
if strings.Contains(html, phrase) {
t.Errorf("forbidden copy %q reappeared in card %q:\n%s", phrase, r.VideoID, html)
}
}
}
}
+167 -8
View File
@@ -2,6 +2,7 @@ package web
import (
"regexp"
"slices"
"strings"
"time"
"unicode/utf8"
@@ -333,7 +334,7 @@ type flashView struct {
// flashMessages maps each flash code to its banner. An unknown code renders no
// banner (flashFor returns ok=false), so a forged cookie value is inert.
var flashMessages = map[string]flashView{
flashConnected: {"success", "YouTube account connected."},
flashConnected: {"success", "YouTube account connected — finding your subscriptions. Your newest videos will appear below as they're summarized."},
flashConnectFailed: {"error", "Could not connect your YouTube account. Please try again."},
flashDisconnected: {"success", "Account disconnected."},
flashDeleted: {"success", "Your account and all its data were deleted."},
@@ -382,13 +383,126 @@ func disconnectURL(provider string) templ.SafeURL {
return templ.SafeURL("/account/disconnect/" + provider)
}
// PipelineStats summarises the user's video backlog so the list page can show
// a one-line status bar ("2 summaries · 256 fetching soon · 12 no captions").
type PipelineStats struct {
Summarized int
RateLimited int // in the backoff window, will be retried
NoText int // no caption track available
Pending int // discovered but not yet attempted
}
// pipelineStats computes a PipelineStats from all (unfiltered) rows.
func pipelineStats(rows []store.SummaryRow) PipelineStats {
var s PipelineStats
for _, r := range rows {
switch {
case r.Summarized:
s.Summarized++
case r.TranscriptStatus == "rate_limited":
s.RateLimited++
case r.TranscriptStatus == "none":
s.NoText++
default:
s.Pending++
}
}
return s
}
// retryNowURL builds the POST path for manual retry of a rate-limited video.
func retryNowURL(videoID string) templ.SafeURL {
return templ.SafeURL("/v/" + videoID + "/retry-now")
}
// listBuckets splits the (already filtered) video list into what the list view
// shows where, so the readable summaries are not buried under the un-summarized
// back-catalogue (UX review B3/B4). It is one feed with a noise-collapse, not
// separate sections:
// - Main: summarized videos + recent un-summarized ones — shown inline as cards.
// - Older: un-summarized videos published before the recency cutoff — collapsed
// behind a single "Show N older videos" disclosure (they will not auto-fill;
// they are summarize-on-demand).
// - NoCaption: count of un-summarized videos with no caption track — collapsed
// to one honest line instead of N dead terminal cards.
type listBuckets struct {
Main []store.SummaryRow
Older []store.SummaryRow
NoCaption int
}
// bucketRows classifies rows into the list buckets given a recency cutoff. A zero
// cutoff (recency collapse disabled) leaves Older empty — every un-summarized,
// captioned video stays inline. Order within each bucket is preserved.
func bucketRows(rows []store.SummaryRow, cutoff time.Time) listBuckets {
var b listBuckets
for _, r := range rows {
switch {
case r.Summarized:
b.Main = append(b.Main, r)
case r.TranscriptStatus == "none":
b.NoCaption++
case isOlder(r, cutoff):
b.Older = append(b.Older, r)
default:
b.Main = append(b.Main, r)
}
}
return b
}
// isOlder reports whether an un-summarized row falls before the recency cutoff.
// A zero cutoff (window disabled) or an undated row is never "older" — it cannot
// be aged out, so it stays inline rather than being hidden in the disclosure.
func isOlder(r store.SummaryRow, cutoff time.Time) bool {
if cutoff.IsZero() || r.PublishedAt.IsZero() {
return false
}
return r.PublishedAt.Before(cutoff)
}
// empty reports whether there is nothing to show at all (drives the empty state).
func (b listBuckets) empty() bool {
return len(b.Main) == 0 && len(b.Older) == 0 && b.NoCaption == 0
}
// Filter holds the list-view query parameters. Empty fields mean "no constraint".
// Dates are kept as the raw YYYY-MM-DD strings so the form re-renders the user's
// input verbatim; parsing happens in matchFilter.
type Filter struct {
Channel string
From string
To string
Channels []string // selected channel titles; empty = all channels
From string
To string
OnlySummarized bool // show only videos that have a summary
}
// active reports whether any filter constraint is set. Drives whether the filter
// bar is shown at all: on a genuinely empty account (no rows AND no active
// filter) the bar is hidden so the connect CTA stands alone (UX review C1); a
// filter that happens to match nothing still shows the bar so it can be cleared.
func (f Filter) active() bool {
return len(f.Channels) > 0 || f.From != "" || f.To != "" || f.OnlySummarized
}
// HasChannel reports whether a channel is currently selected (drives the
// multi-select's selected state in the view).
func (f Filter) HasChannel(c string) bool {
return slices.Contains(f.Channels, c)
}
// nonEmptyStrings drops blank entries. A channel multi-select submits real
// channel titles; this guards against a stray empty value reaching the filter.
func nonEmptyStrings(ss []string) []string {
out := ss[:0:0]
for _, s := range ss {
if strings.TrimSpace(s) != "" {
out = append(out, s)
}
}
if len(out) == 0 {
return nil
}
return out
}
// matches reports whether a row satisfies the filter. Channel is an exact match;
@@ -396,7 +510,10 @@ type Filter struct {
// (no constraint) — Stage-0 filtering is in-memory over the listed rows, not a
// store query.
func (f Filter) matches(r store.SummaryRow) bool {
if f.Channel != "" && r.Channel != f.Channel {
if f.OnlySummarized && !r.Summarized {
return false
}
if len(f.Channels) > 0 && !slices.Contains(f.Channels, r.ChannelTitle) {
return false
}
if from, ok := parseDate(f.From); ok {
@@ -426,7 +543,7 @@ func parseDate(s string) (time.Time, bool) {
// apply returns the subset of rows matching the filter, preserving order.
func (f Filter) apply(rows []store.SummaryRow) []store.SummaryRow {
if f == (Filter{}) {
if len(f.Channels) == 0 && f.From == "" && f.To == "" && !f.OnlySummarized {
return rows
}
out := rows[:0:0]
@@ -478,7 +595,13 @@ main { max-width: 60rem; margin: 0 auto; padding: var(--s4) var(--s3); }
.filters label { display: flex; flex-direction: column; font-size: .78rem; text-transform: uppercase; letter-spacing: .04em; color: var(--muted); gap: var(--s1); }
.filters input { font: inherit; padding: .4rem .55rem; border: 1px solid var(--line); border-radius: var(--radius); background: var(--card); color: var(--fg); min-width: 9rem; }
.filters input:focus-visible { outline: 2px solid var(--accent); outline-offset: 1px; border-color: var(--accent); }
.filter-check { flex-direction: row !important; align-items: center; gap: var(--s2) !important; padding-bottom: .45rem; }
.filter-check input[type=checkbox] { width: 1rem; height: 1rem; min-width: 0; padding: 0; accent-color: var(--accent); cursor: pointer; }
.btn { font: inherit; font-weight: 600; padding: .45rem 1rem; border: 1px solid var(--accent); border-radius: var(--radius); background: var(--accent); color: var(--accent-fg); cursor: pointer; }
/* anchors styled as buttons: the generic a{} / a:visited{} colour rules outrank
.btn on <a>, painting the label accent-on-accent (invisible). Restore the
button foreground for anchor buttons, visited included. */
a.btn, a.btn:visited { color: var(--accent-fg); }
.btn:hover { filter: brightness(1.05); }
.btn:active { transform: translateY(1px); }
@@ -490,6 +613,26 @@ main { max-width: 60rem; margin: 0 auto; padding: var(--s4) var(--s3); }
.card-preview { color: var(--muted); font-size: .9rem; line-height: 1.5; display: -webkit-box; -webkit-line-clamp: 1; line-clamp: 1; -webkit-box-orient: vertical; overflow: hidden; }
.card-foot { display: flex; gap: var(--s2); align-items: center; flex-wrap: wrap; margin-top: var(--s1); }
.chip { display: inline-block; padding: .15rem .55rem; border-radius: 999px; background: var(--accent-weak); color: var(--accent); font-size: .72rem; font-weight: 600; }
/* passive "retrying later" chip: dim/grey (CharmDim), not the accent — it is a
status, not an action the user can take. */
.chip-retry { background: rgba(108, 108, 108, .16); color: #6c6c6c; }
.pipeline-bar { display: flex; gap: var(--s3); align-items: center; flex-wrap: wrap; margin-bottom: var(--s3); font-size: .8rem; color: var(--muted); }
.pipeline-bar span { display: flex; align-items: center; gap: var(--s1); }
.pipeline-bar span + span::before { content: "·"; margin-right: var(--s1); }
.pipeline-note { margin: calc(-1 * var(--s2)) 0 var(--s3); font-size: .8rem; line-height: 1.5; max-width: 40rem; }
/* one-line count of caption-less videos (collapsed instead of N dead cards) */
.list-note { margin: var(--s3) 0 0; font-size: .85rem; }
/* older un-summarized back-catalogue, collapsed behind a disclosure so it does
not bury the readable summaries above it */
.older-videos { margin-top: var(--s4); }
.older-videos > summary { cursor: pointer; font-size: .85rem; font-weight: 600; color: var(--accent); padding: var(--s2) 0; list-style: revert; }
.older-videos > summary:hover { text-decoration: underline; }
.older-videos[open] > summary { margin-bottom: var(--s3); }
.older-videos .cards { margin-top: 0; }
.card-nudge-form { display: inline; }
.btn-quiet { font: inherit; font-size: .72rem; font-weight: 600; padding: .15rem .55rem; border-radius: 999px; border: 1px solid var(--accent); background: transparent; color: var(--accent); cursor: pointer; }
.btn-quiet:hover { background: var(--accent-weak); }
.chip-warn { background: rgba(255, 110, 156, .15); color: #FF6E9C; }
.card-state { color: var(--muted); font-size: .8rem; }
.badge { display: inline-block; padding: .15rem .55rem; border-radius: 999px; background: var(--badge-bg); color: var(--badge-fg); font-size: .72rem; font-weight: 600; }
@@ -531,6 +674,11 @@ main { max-width: 60rem; margin: 0 auto; padding: var(--s4) var(--s3); }
.empty { text-align: center; color: var(--muted); padding: var(--s5) var(--s4); border: 1px dashed var(--line); border-radius: var(--radius); background: var(--card); }
.empty strong { display: block; color: var(--fg); font-size: 1.05rem; margin-bottom: var(--s2); }
.empty code { background: var(--accent-weak); color: var(--accent); padding: .1rem .35rem; border-radius: .3rem; }
.empty p { margin: var(--s3) 0 0; }
/* connected-but-empty: a distinct accent callout, not a muted blank state, so a
fresh account knows the next step is to run tapir, not "something is broken". */
.empty-connected { border-style: solid; border-color: var(--accent); background: var(--accent-weak); color: var(--fg); }
.empty-connected strong { color: var(--accent); }
/* flash / notification banner */
.flash { padding: var(--s2) var(--s3); border-radius: var(--radius); margin-bottom: var(--s4); font-size: .92rem; border: 1px solid var(--line); }
@@ -546,6 +694,7 @@ main { max-width: 60rem; margin: 0 auto; padding: var(--s4) var(--s3); }
/* detail reader */
.detail { max-width: 38rem; }
.detail .back { margin: 0 0 var(--s3); font-size: .85rem; }
.detail h1 { font-size: 1.7rem; line-height: 1.25; margin: 0 0 var(--s2); }
.detail .meta { color: var(--muted); font-size: .9rem; margin: 0 0 var(--s2); display: flex; gap: var(--s2); align-items: center; flex-wrap: wrap; }
.detail .source { margin: 0 0 var(--s4); font-size: .9rem; }
@@ -557,8 +706,13 @@ main { max-width: 60rem; margin: 0 auto; padding: var(--s4) var(--s3); }
.detail ul { margin: 0; padding-left: 1.2rem; line-height: 1.6; }
.detail li { margin-bottom: var(--s1); }
/* action toggles */
.actions { display: flex; gap: var(--s2); margin: var(--s4) 0; flex-wrap: wrap; }
/* action toggles — watched|skipped form one segmented control (they are mutually
exclusive), "saved" sits apart as an independent toggle */
.actions { display: flex; gap: var(--s3); margin: var(--s4) 0; flex-wrap: wrap; align-items: center; }
.segmented { display: inline-flex; }
.segmented .action { border-radius: 0; border-right-width: 0; }
.segmented .action:first-child { border-top-left-radius: var(--radius); border-bottom-left-radius: var(--radius); }
.segmented .action:last-child { border-top-right-radius: var(--radius); border-bottom-right-radius: var(--radius); border-right-width: 1px; }
.actions .action { font: inherit; padding: .4rem .9rem; border: 1px solid var(--line); border-radius: var(--radius); background: var(--card); color: var(--fg); cursor: pointer; transition: border-color .15s, background .15s; }
.actions .action:hover { border-color: var(--accent); }
.actions .action:focus-visible { outline: 2px solid var(--accent); outline-offset: 1px; }
@@ -573,6 +727,11 @@ main { max-width: 60rem; margin: 0 auto; padding: var(--s4) var(--s3); }
.account-meta { display: grid; grid-template-columns: max-content 1fr; gap: var(--s1) var(--s3); margin: 0; }
.account-meta dt { color: var(--muted); font-size: .85rem; }
.account-meta dd { margin: 0; }
.channel-errors { }
.channel-error-list { list-style: none; margin: 0 0 var(--s3); padding: 0; display: grid; gap: var(--s1); }
.channel-error-list li { display: flex; align-items: center; gap: var(--s2); }
.channel-error-name { font-weight: 500; }
.channel-error-since { font-size: .8rem; }
.conn-list { list-style: none; margin: 0 0 var(--s3); padding: 0; display: grid; gap: var(--s2); }
.conn { background: var(--card); border: 1px solid var(--line); border-radius: var(--radius); padding: var(--s3); display: flex; flex-direction: column; gap: var(--s1); }
.conn-main { display: flex; gap: var(--s2); align-items: center; flex-wrap: wrap; }
+237 -66
View File
@@ -1,6 +1,7 @@
package web
import (
"fmt"
"strings"
"gitea.d-ma.be/mathias/tapir/internal/adapters/store"
@@ -22,7 +23,31 @@ templ Layout(title string) {
<body>
<header>
<a href="/" class="brand">Tapir</a>
<nav class="nav"><a href="/account">Account</a></nav>
<nav class="nav"><a href="/account">Account</a><a href="/auth/logout">Log out</a></nav>
</header>
<main>
{ children... }
</main>
</body>
</html>
}
// PublicLayout is the shell for unauthenticated pages (/welcome, /invite).
// Same structure as Layout but without the nav auth links — a visitor who is not
// logged in should not see "Account" or "Log out".
templ PublicLayout(title string) {
<!DOCTYPE html>
<html lang="en">
<head>
<meta charset="utf-8"/>
<meta name="viewport" content="width=device-width, initial-scale=1"/>
<title>{ title }</title>
<script src="/static/htmx.min.js" defer></script>
@templ.Raw(styleTag)
</head>
<body>
<header>
<a href="/" class="brand">Tapir</a>
</header>
<main>
{ children... }
@@ -36,7 +61,7 @@ templ Layout(title string) {
// single "Get Started" CTA into the shared Dex flow (sign-in and sign-up are the
// same URL). Logged in: a greeting plus links back into the app and to log out.
templ WelcomePage(user User, loggedIn bool) {
@Layout("Tapir — Watch less, know more") {
@PublicLayout("Tapir — Watch less, know more") {
<section class="welcome">
<div class="welcome-hero">
<pre aria-hidden="true">@templ.Raw(welcomeHero)</pre>
@@ -59,7 +84,7 @@ templ WelcomePage(user User, loggedIn bool) {
<div class="welcome-cta">
<a class="btn btn-lg" href="/auth/login">Get Started</a>
</div>
<p class="welcome-sub">Already have an account? You'll go straight through.</p>
<p class="welcome-sub">Tapir is invite-only right now. If you've been invited, sign in above. New summaries land gradually Tapir fetches captions slowly to respect YouTube's limits.</p>
}
</section>
}
@@ -79,17 +104,71 @@ templ flashBanner(code string) {
// #summary-list region; a non-HTMX request renders the whole page. flash carries
// a one-shot notification (e.g. "connected", "registered") surfaced on arrival
// after a POST→redirect.
templ ListPage(rows []store.SummaryRow, f Filter, flash string) {
templ ListPage(b listBuckets, f Filter, stats PipelineStats, flash string, hasConnected bool, channels []string) {
@Layout("Tapir — Summaries") {
@flashBanner(flash)
@filterForm(f)
if hasConnected {
@pasteForm()
}
if !b.empty() || f.active() {
@filterForm(f, channels)
}
if stats.RateLimited > 0 || stats.Pending > 0 || stats.NoText > 0 {
@pipelineBar(stats)
}
if stats.RateLimited+stats.Pending > 0 {
<p class="pipeline-note muted">
Tapir fetches captions slowly on purpose, to respect YouTube's limits
new summaries land gradually. Check back tomorrow.
</p>
}
<div id="summary-list">
@summaryList(rows)
@summaryList(b, hasConnected)
</div>
}
}
templ filterForm(f Filter) {
// pipelineBar is the one-line backlog status. Counts are framed by what the user
// can read NOW ("ready"), what is waiting behind the honest caption rate limit
// ("in queue" = pending + rate-limited, never "fetching soon" — see A3/ADR-014),
// and what is permanently unreadable ("no captions").
templ pipelineBar(s PipelineStats) {
<div class="pipeline-bar">
if s.Summarized > 0 {
<span>{ fmt.Sprintf("%d ready", s.Summarized) }</span>
}
if s.RateLimited+s.Pending > 0 {
<span>{ fmt.Sprintf("%d in queue", s.RateLimited+s.Pending) }</span>
}
if s.NoText > 0 {
<span class="muted">{ fmt.Sprintf("%d no captions", s.NoText) }</span>
}
</div>
}
// pasteForm lets a connected user summarize any YouTube video by pasting its URL
// (Feature 2). The result (a video card, or an inline error) swaps into
// #paste-result; the next list refresh shows it inline. Summarization runs
// through the shared caption rate gate like every other fetch.
templ pasteForm() {
<form
class="paste"
method="post"
action="/paste"
hx-post="/paste"
hx-target="#paste-result"
hx-swap="innerHTML"
>
<label>
Summarize any video
<input type="url" name="url" placeholder="Paste a YouTube link…" required/>
</label>
<button type="submit">Add</button>
</form>
<div id="paste-result"></div>
}
templ filterForm(f Filter, channels []string) {
<form
class="filters"
method="get"
@@ -99,37 +178,76 @@ templ filterForm(f Filter) {
hx-swap="innerHTML"
hx-indicator="#filter-indicator"
>
<label>Channel <input type="text" name="channel" value={ f.Channel } placeholder="any"/></label>
<label>From <input type="date" name="from" value={ f.From }/></label>
<label>To <input type="date" name="to" value={ f.To }/></label>
if len(channels) > 0 {
<label>
Channels
<select name="channel" multiple size="4">
for _, c := range channels {
<option value={ c } selected?={ f.HasChannel(c) }>{ c }</option>
}
</select>
</label>
}
<label class="filter-check">
<input type="checkbox" name="summarized" value="1" if f.OnlySummarized { checked }/>
Summarized only
</label>
<button type="submit" class="btn">Filter</button>
<span id="filter-indicator" class="htmx-indicator">filtering…</span>
</form>
}
// summaryList is the swappable list fragment: one card per video (summarized or
// not). Cards reflow to a single column on mobile; an empty list shows a friendly
// first-run state instead of a blank table.
templ summaryList(rows []store.SummaryRow) {
if len(rows) == 0 {
<div class="empty">
<strong>No videos yet</strong>
<span>Videos appear here as your subscriptions are processed run <code>tapir run</code> to fetch them. In manual mode, use the Summarize button to queue one.</span>
</div>
// summaryList is the swappable list fragment. It leads with readable summaries +
// recent un-summarized cards (b.Main), then collapses the noise so it does not
// bury the payload (UX review B3/B4): a one-line count of caption-less videos,
// and a single disclosure holding the older un-summarized back-catalogue. Cards
// reflow to a single column on mobile; an empty list shows a friendly first-run
// state instead of a blank table.
templ summaryList(b listBuckets, hasConnected bool) {
if b.empty() {
if hasConnected {
<div class="empty empty-connected">
<strong>Your account is connected</strong>
<span>Tapir is finding your subscriptions and fetching captions summaries appear here gradually. Check back later.</span>
</div>
} else {
<div class="empty">
<strong>No videos yet</strong>
<span>Connect your YouTube account to get started.</span>
<p><a class="btn" href="/oauth/youtube/connect">Connect YouTube</a></p>
</div>
}
} else {
<ul class="cards">
for _, r := range rows {
for _, r := range b.Main {
@VideoCard(r)
}
</ul>
if b.NoCaption > 0 {
<p class="list-note muted">{ fmt.Sprintf("%d video(s) have no captions and can't be summarized.", b.NoCaption) }</p>
}
if len(b.Older) > 0 {
<details class="older-videos">
<summary>{ fmt.Sprintf("Show %d older videos — summarize on demand", len(b.Older)) }</summary>
<ul class="cards">
for _, r := range b.Older {
@VideoCard(r)
}
</ul>
</details>
}
}
}
// VideoCard is one list card, also returned standalone by POST /v/{id}/summarize
// (HTMX swaps it in place via outerHTML). A summarized video links to its detail
// page and shows its provider chip / fallback badge / action state. An
// unsummarized video gets a muted "pending" treatment and either a "Summarize"
// button (to queue it) or a "Queued" chip when already requested.
// VideoCard is one list card, returned standalone by POST /v/{id}/summarize and
// /v/{id}/retry-now (HTMX swaps outerHTML). Five footer states, status-primary:
// 1. Summarized — preview + chip + actions; no button.
// 2. No captions (TranscriptStatus=="none") — terminal; "No transcript available"; no button.
// 3. Queued (SummarizeRequested) — "Queued · summarizing shortly"; no button.
// 4. Rate-limited — "In queue" + quiet "Summarize" → /retry-now.
// 5. Pending (else) — "Not summarized" + quiet "Summarize" → /summarize.
// States 4 and 5 use one verb ("Summarize") and one style (.btn-quiet); the
// backend side-effect difference (clear-backoff vs. set-flag) is invisible to users.
templ VideoCard(r store.SummaryRow) {
<li class={ "card", templ.KV("card-pending", !r.Summarized) } id={ "video-" + r.VideoID }>
if r.Summarized {
@@ -147,6 +265,7 @@ templ VideoCard(r store.SummaryRow) {
}
<div class="card-foot">
if r.Summarized {
// State 1: summarized — provider chip, fallback badge, action state.
if r.AIProvider != "" {
<span class="chip">{ r.AIProvider }</span>
}
@@ -156,18 +275,39 @@ templ VideoCard(r store.SummaryRow) {
if len(r.Actions) > 0 {
<span class="card-state">{ strings.Join(r.Actions, ", ") }</span>
}
} else if r.TranscriptStatus == "none" {
// State 2: no captions — terminal dead-end; nothing the user can do.
<span class="card-state muted">No transcript available</span>
} else if r.SummarizeRequested {
// State 3: queued — being summarized on the next pass; no scheduler jargon.
<span class="chip">Queued</span>
<span class="card-state muted">waiting for the next run</span>
<span class="card-state muted">summarizing shortly</span>
} else if r.TranscriptStatus == "rate_limited" {
// State 4: rate-limited — honest "in queue" status (NOT "fetching soon",
// which oversells imminence) + a quiet nudge → retry-now handler.
<span class="card-state muted">In queue</span>
<form
method="post"
action={ retryNowURL(r.VideoID) }
hx-post={ string(retryNowURL(r.VideoID)) }
hx-target={ "#video-" + r.VideoID }
hx-swap="outerHTML"
class="card-nudge-form"
>
<button type="submit" class="btn-quiet" title="Summarize this video">Summarize</button>
</form>
} else {
// State 5: pending — discovered, not yet attempted; nudge button → summarize handler.
<span class="card-state muted">Not summarized</span>
<form
method="post"
action={ summarizeURL(r.VideoID) }
hx-post={ string(summarizeURL(r.VideoID)) }
hx-target={ "#video-" + r.VideoID }
hx-swap="outerHTML"
class="card-nudge-form"
>
<button type="submit" class="btn-secondary">Summarize</button>
<button type="submit" class="btn-quiet" title="Summarize this video">Summarize</button>
</form>
}
</div>
@@ -215,6 +355,7 @@ templ processingCard(r store.SummaryRow) {
templ DetailPage(r store.SummaryRow) {
@Layout("Tapir — " + displayTitle(r)) {
<article class="detail">
<p class="back"><a href="/"> Summaries</a></p>
<h1>{ displayTitle(r) }</h1>
<p class="meta">
if detailMeta(r) != "" {
@@ -240,20 +381,10 @@ templ DetailPage(r store.SummaryRow) {
<p class="source"><a href={ externalURL(r.URL) } rel="noopener noreferrer">watch on source </a></p>
}
@ActionButtons(r.VideoID, actionSet(r.Actions))
<section>
<h2>Summary</h2>
<p class="body">{ r.Summary }</p>
</section>
if len(r.Highlights) > 0 {
<section>
<h2>Highlights</h2>
<ul>
for _, h := range r.Highlights {
<li>{ h }</li>
}
</ul>
</section>
}
// Lead with the attention-saving payload: Takeaways ("is this worth my
// time?") first, then Highlights, then the full Summary last (UX review
// A8). Takeaways/Highlights are conditional, so a video without them falls
// through to the Summary leading naturally.
if len(r.Takeaways) > 0 {
<section>
<h2>Takeaways</h2>
@@ -264,13 +395,27 @@ templ DetailPage(r store.SummaryRow) {
</ul>
</section>
}
if len(r.Highlights) > 0 {
<section>
<h2>Highlights</h2>
<ul>
for _, h := range r.Highlights {
<li>{ h }</li>
}
</ul>
</section>
}
<section>
<h2>Summary</h2>
<p class="body">{ r.Summary }</p>
</section>
</article>
}
}
// RegisterPage is the explicit registration step (ADR-012): an authenticated Dex
// subject with no tapir user picks a display name and accepts the terms to create
// their account. errMsg, when set, reports a validation problem on the prior POST.
// subject with no tapir user picks a display name to create their account.
// errMsg, when set, reports a validation problem on the prior POST.
templ RegisterPage(email, errMsg string) {
@Layout("Tapir — Register") {
<article class="register">
@@ -287,10 +432,6 @@ templ RegisterPage(email, errMsg string) {
Display name
<input type="text" name="display_name" required autofocus/>
</label>
<label class="checkbox">
<input type="checkbox" name="accept_terms" value="yes" required/>
I accept the terms of use
</label>
<button type="submit" class="btn">Register</button>
</form>
</article>
@@ -301,7 +442,7 @@ templ RegisterPage(email, errMsg string) {
// signed-in email, the user's connected video accounts (each with a Disconnect
// control), a Connect-YouTube link when none is connected, and the delete-account
// danger zone. flash surfaces a one-shot notification (disconnect/connect).
templ AccountPage(displayName, email string, conns []store.Connection, autoSummarize bool, flash string) {
templ AccountPage(displayName, email string, conns []store.Connection, autoSummarize bool, channelErrors []store.ChannelError, flash string) {
@Layout("Tapir — Account") {
@flashBanner(flash)
<article class="account">
@@ -317,12 +458,31 @@ templ AccountPage(displayName, email string, conns []store.Connection, autoSumma
<section>
<h2>Summarization</h2>
<p class="muted">
Automatic summarizes every new video as it is discovered. Manual lets you
pick which videos to summarize new videos appear in your list with a
Summarize button.
Automatic summarizes new videos from about the last week as they are
discovered. Older videos stay browsable summarize them on demand.
Manual lets you pick which videos to summarize every new video appears
in your list with a Summarize button.
</p>
@summarizeModeControl(autoSummarize)
</section>
if len(channelErrors) > 0 {
<section class="channel-errors">
<h2>Unavailable channels</h2>
<p class="muted">
{ fmt.Sprintf("%d channel(s) returned errors on the last discovery pass.", len(channelErrors)) }
These may have been deleted or made private on YouTube.
</p>
<ul class="channel-error-list">
for _, ce := range channelErrors {
<li>
<span class="channel-error-name">{ ce.ChannelName }</span>
<span class="chip chip-warn">unavailable</span>
<span class="muted channel-error-since">since { ce.FirstSeen.Format("2006-01-02") }</span>
</li>
}
</ul>
</section>
}
<section>
<h2>Connected accounts</h2>
if len(conns) == 0 {
@@ -404,20 +564,31 @@ templ ActionButtons(videoID string, active map[string]bool) {
hx-target="#action-buttons"
hx-swap="outerHTML"
>
for _, v := range actionVerbs {
<button
type="submit"
name="action"
value={ v }
class={ "action", templ.KV("active", active[v]) }
aria-pressed={ ariaPressed(active[v]) }
>
if active[v] {
{ "✓ " + actionLabel(v) }
} else {
{ actionLabel(v) }
}
</button>
}
// watched ↔ skipped are mutually exclusive (the store clears one when the
// other is set), so they read as a single segmented choice. "saved" is an
// independent toggle and sits apart (UX review C5).
<span class="segmented" role="group" aria-label="Watched or skipped">
@actionButton("watched", active["watched"])
@actionButton("skipped", active["skipped"])
</span>
@actionButton("saved", active["saved"])
</form>
}
// actionButton is one toggle button in the action group: a submit carrying its
// verb, marked active (accent fill + ✓ prefix + aria-pressed) when currently set.
templ actionButton(verb string, isActive bool) {
<button
type="submit"
name="action"
value={ verb }
class={ "action", templ.KV("active", isActive) }
aria-pressed={ ariaPressed(isActive) }
>
if isActive {
{ "✓ " + actionLabel(verb) }
} else {
{ actionLabel(verb) }
}
</button>
}
+1013 -549
View File
File diff suppressed because it is too large Load Diff
+5
View File
@@ -53,6 +53,11 @@ func TestWelcomeLoggedOut(t *testing.T) {
require.Contains(t, html, "Get Started", "logged-out CTA present")
require.Contains(t, html, `href="/auth/login"`, "CTA links into the Dex flow")
require.NotContains(t, html, "Go to my Tapir", "no logged-in controls")
// Invites are owned by Authentik now (ADR-019); Tapir no longer handles
// invite links. The old "invite link will set up your account automatically"
// promise is stale and misleading — it must be gone.
require.NotContains(t, html, "invite link", "stale Authentik-superseded invite copy must be removed")
require.Contains(t, html, "invite-only", "honest invite-only framing present")
}
func TestWelcomeLoggedIn(t *testing.T) {
+220
View File
@@ -0,0 +1,220 @@
package acceptance
// This is the name-coverage gate for the BDD spec (see docs/use-cases/*.feature).
// There is no godog runner — the .feature files are design records, and the real
// behaviour is covered by the hand-written Go tests across the module. This test
// keeps the two from drifting in the cheapest honest way: every non-@pending
// Scenario must have an entry in scenarioCoverage pointing at a Go test that
// actually exists. It does NOT prove the test exercises the scenario (only godog
// could); it catches the common drift — "added a scenario, forgot the test", a
// renamed/deleted covering test, or a scenario removed without cleaning the map.
//
// When you add a Scenario: either map it here to its covering test, or tag it
// @pending in the .feature with a one-line reason for why it has no test yet.
import (
"os"
"path/filepath"
"regexp"
"strings"
"testing"
)
// scenarioCoverage maps each non-@pending Scenario name to the Go test that
// covers it. Keep it in sync with docs/use-cases/*.feature — the test below
// fails if a scenario is unmapped, a mapped test is missing, or an entry no
// longer matches a real non-pending scenario.
var scenarioCoverage = map[string]string{
// ai_routing.feature
"Local AI produces the summary": "TestSummarize_LocalSucceeds",
"Local AI fails and the user has a BYO provider configured": "TestSummarize_FallsBackToBYO",
"Local AI fails and the user has no BYO provider": "TestSummarize_LocalFailsNoBYO_NoExternalSend",
"A user without BYO never has content sent externally": "TestSummarize_NoBYO_ContentOnlyLocal",
// landing_page.feature
"An unauthenticated visit to the root is sent to the welcome page": "TestUnauthenticatedRootRedirectsToWelcome",
"The welcome page invites an unauthenticated visitor to start": "TestWelcomeLoggedOut",
"An authenticated user on the welcome page sees their way in and out": "TestWelcomeLoggedIn",
// connect_account.feature
"Connect a YouTube account": "TestCallbackExchangesAndRecordsConnection",
"Connecting an account discovers videos immediately": "TestCallbackTriggersDiscovery",
"Connecting summarizes my newest videos right away": "TestNewestUnsummarizedVideoIDs",
"Tokens are never stored in the clear": "TestCallbackExchangesAndRecordsConnection",
// paste_url.feature
"Paste a valid YouTube URL": "TestPasteValidURLAddsAndRequests",
"Pasting an invalid link is rejected": "TestPasteInvalidURLRejected",
"Pasting a video that cannot be found is honest": "TestPasteVideoNotFound",
"Pasting the same video twice does not duplicate it": "TestPasteDedupNoDuplicate",
"Revoking a connection stops watching but keeps history": "TestDisconnectRemovesTokenAndConnectionKeepsAccount",
// summarize_mode.feature
"Auto mode summarizes recent new videos automatically": "TestRunOnce_AutoMode_SkipsOldVideos",
"Auto mode lists older videos without summarizing them": "TestRunOnce_AutoMode_SkipsOldVideos",
"Automatic is the default for a new user": "TestRegisteredUserDefaultsAutoSummarizeOn",
"Manual mode leaves new videos unsummarized": "TestRunOnce_ManualMode_SkipsUnrequested",
"Requesting a summary in manual mode queues it for the next run": "TestRunOnce_ManualMode_ProcessesRequested",
// registration.feature
"A new Dex subject is routed to registration": "TestUnregisteredSubjectRedirectedToRegister",
"Registering creates the account and its identity mapping": "TestRegisterCreatesExactlyOneUserAndIdentity",
"A returning subject passes straight through": "TestRegisteredSubjectPassesThrough",
"Deleting an account removes only my data and leaves other users untouched": "TestDeleteAccountWipesDataAndSecretsAndLogsOut",
// summarize_new_video.feature
"A subscribed channel posts a video that has captions": "TestSubscribedVideoWithCaptionsIsSummarizedAndDelivered",
"A subscribed channel posts a video with no usable transcript": "TestVideoWithNoTranscriptIsSkipped",
"A channel I am not subscribed to posts a video": "TestUnsubscribedChannelVideoIsNotProcessed",
"The same video is not summarized twice": "TestAlreadySummarizedVideoIsNotReprocessed",
"Re-analyzing a stored video does not re-fetch its transcript": "TestProcessNewVideo_SecondSummarizeDoesNotRefetch",
}
var (
scenarioRe = regexp.MustCompile(`^\s*Scenario(?: Outline)?:\s*(.+?)\s*$`)
testFuncRe = regexp.MustCompile(`^func (Test\w+)\(`)
)
// scenario is one parsed Gherkin scenario and whether it is @pending.
type scenario struct {
name string
pending bool
}
func TestScenarioCoverage(t *testing.T) {
root := moduleRoot(t)
scenarios := parseScenarios(t, filepath.Join(root, "docs", "use-cases"))
if len(scenarios) == 0 {
t.Fatal("no scenarios parsed from docs/use-cases — wrong path?")
}
tests := allTestFuncNames(t, root)
// Index scenario names for the reverse (stale-entry) check.
active := map[string]bool{} // non-pending scenario names
var pending []string
for _, s := range scenarios {
if s.pending {
pending = append(pending, s.name)
continue
}
active[s.name] = true
// 1. Every non-pending scenario must be mapped.
fn, ok := scenarioCoverage[s.name]
if !ok {
t.Errorf("scenario %q has no coverage entry — map it in scenarioCoverage to a covering test, or tag it @pending in the .feature", s.name)
continue
}
// 2. The mapped test must actually exist.
if !tests[fn] {
t.Errorf("scenario %q maps to %q, which is not a Test function anywhere in the module", s.name, fn)
}
}
// 3. No stale entries: every map key must be a real, non-pending scenario.
for name := range scenarioCoverage {
if !active[name] {
t.Errorf("scenarioCoverage has entry %q, which is not a current non-pending scenario (renamed, removed, or now @pending?)", name)
}
}
if len(pending) > 0 {
t.Logf("%d @pending scenario(s) without a test (tracked, not required): %s",
len(pending), strings.Join(pending, "; "))
}
}
// parseScenarios reads every *.feature under dir and returns its scenarios with
// their @pending status. A scenario is @pending when a `@pending` tag line
// precedes it (tags survive intervening comment lines, the layout these files
// use); the flag is consumed at the Scenario line and reset afterward.
func parseScenarios(t *testing.T, dir string) []scenario {
t.Helper()
entries, err := os.ReadDir(dir)
if err != nil {
t.Fatalf("read use-cases dir: %v", err)
}
var out []scenario
for _, e := range entries {
if e.IsDir() || !strings.HasSuffix(e.Name(), ".feature") {
continue
}
b, err := os.ReadFile(filepath.Join(dir, e.Name()))
if err != nil {
t.Fatalf("read %s: %v", e.Name(), err)
}
pending := false
for _, line := range strings.Split(string(b), "\n") {
trimmed := strings.TrimSpace(line)
if strings.HasPrefix(trimmed, "@") {
if strings.Contains(trimmed, "@pending") {
pending = true
}
continue
}
if m := scenarioRe.FindStringSubmatch(line); m != nil {
out = append(out, scenario{name: m[1], pending: pending})
pending = false
}
// comment (#) and step lines leave a set @pending intact until the
// scenario consumes it; a blank line between scenarios is harmless.
}
}
return out
}
// allTestFuncNames walks the module for `func TestXxx(` declarations, excluding
// this file (whose regex literal would otherwise look like a definition).
func allTestFuncNames(t *testing.T, root string) map[string]bool {
t.Helper()
self := "scenario_coverage_test.go"
names := map[string]bool{}
err := filepath.WalkDir(root, func(path string, d os.DirEntry, err error) error {
if err != nil {
return err
}
if d.IsDir() {
if d.Name() == ".git" {
return filepath.SkipDir
}
return nil
}
if !strings.HasSuffix(path, "_test.go") || filepath.Base(path) == self {
return nil
}
b, err := os.ReadFile(path)
if err != nil {
return err
}
for _, line := range strings.Split(string(b), "\n") {
if m := testFuncRe.FindStringSubmatch(line); m != nil {
names[m[1]] = true
}
}
return nil
})
if err != nil {
t.Fatalf("walk module: %v", err)
}
return names
}
// moduleRoot walks up from the working directory to the dir containing go.mod.
func moduleRoot(t *testing.T) string {
t.Helper()
dir, err := os.Getwd()
if err != nil {
t.Fatalf("getwd: %v", err)
}
for {
if _, err := os.Stat(filepath.Join(dir, "go.mod")); err == nil {
return dir
}
parent := filepath.Dir(dir)
if parent == dir {
t.Fatal("go.mod not found walking up from cwd")
}
dir = parent
}
}