Compare commits

..
15 Commits
Author SHA1 Message Date
mathiasandClaude Opus 4.8 0ba78e8868 docs: refresh build-state for transcript persistence (ADR-021)
CI / Lint / Test / Vet (push) Successful in 11s
CI / Build & Import (push) Successful in 10s
Update the CLAUDE.md orientation block: last tag v0.14.0, migrations
001–015, and a transcript-persistence bullet (shared non-RLS store,
engine reads stored-first). The stale "v0.9.0" reference is corrected.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-09 23:39:27 +02:00
mathiasandClaude Opus 4.8 821d5f99cd docs(bdd): scenario for transcript reuse — re-analysis never re-fetches
Capture the ADR-021 promise as a mapped BDD scenario: re-analyzing a
stored video reads the stored transcript and does not fetch captions.
Since paste-a-URL and the onboarding burst summarize through the same
engine chokepoint (resolveTranscript, store-first), this one scenario
covers their reuse path too — there is exactly one gated caption entry
point (youtube.FetchTranscript → WaitFetchGate) and one engine caller in
front of it, so the dedup is structural, not per-feature.

Mapped to TestProcessNewVideo_SecondSummarizeDoesNotRefetch.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-09 23:37:54 +02:00
mathiasandClaude Opus 4.8 5c70408e75 feat(usecase): read stored transcript before fetching (ADR-021)
The engine now resolves transcripts store-first: a stored transcript —
including a stored SourceNone — is summarized without touching YouTube,
so re-analysis never re-fetches. On a miss it fetches through the source
(caption call still gated, ADR-014) and persists the terminal outcome for
the next analysis by any user. A transient SourceRateLimited is surfaced
to the runner for per-user backoff but never cached, so persistence can
never mask a 429 as a permanent "no transcript".

The TranscriptStore is optional (nil → fetch every time), keeping the
pure-core and scaffold wiring valid. cmd/tapir wires the store as both
summary sink and transcript cache, so `tapir run` and the web summarize
path (incl. paste + onboarding) all share the dedup.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-09 23:35:04 +02:00
mathiasandClaude Opus 4.8 cb6917ca59 feat(store): shared, video-keyed transcript persistence (ADR-021)
Reshape the dead per-user transcripts table (PK videos.id, user_id,
RLS-FORCEd — never read or written by app code) into the shared public
caption store ADR-021 specifies: keyed by (provider, provider_video_id),
no user_id, NOT RLS-scoped. Migration 015 (reversible). Add
ports.TranscriptStore + Store.GetTranscript/SaveTranscript via the raw
pool (no withUser): public content, shared across users by construction.
SaveTranscript persists only terminal outcomes (captions/none) and
refuses SourceRateLimited so a transient 429 can never be stored as a
false permanent absence (ADR-014).

Flip the isolation proof: transcripts leaves the RLS-scoped set;
TestTranscriptsTableIsSharedNotRLS asserts it is the SINGLE non-RLS
surface (writable/readable with no user scope, no user_id column, RLS off
on it alone, still on every user-owned table) — the proof the
public-content classification was applied exactly here and leaked nowhere.
appPool made idempotent so two tests can build it. Adjust the 010/011/014
up-down migration tests for the new HEAD. account.go: user deletion no
longer strips shared transcripts. Reconcile data-model.md + CLAUDE.md.

Wiring the engine to read-stored-first is the next commit.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-09 23:32:12 +02:00
mathiasandClaude Opus 4.8 099b2d4c68 docs(decisions): ADR-021 — shared, video-keyed transcript persistence
Persist transcripts in a single shared table keyed by
(provider, provider_video_id) — public caption content, NOT RLS-scoped —
so re-analysis (re-summarize, paste of an already-seen video, a second
user with overlapping subs) never re-fetches from YouTube. The avoided
cost is the rate-gated, reputation-risky caption fetch (ADR-010/014), not
LLM re-summarization, which is why this reopens the transcripts half of
the "no global cross-tenant table" rejection while videos stay per-user.
Summaries remain RLS-scoped (ADR-012 unchanged). The gate is neither
bypassed nor weakened — persistence reduces fetch frequency, not pacing.

Annotate the rejected-alternatives row to record the partial reopen.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-09 23:24:11 +02:00
mathiasandClaude Opus 4.8 f66c1bcdcc feat(web): real channel filter — multi-select of the user's channels
CI / Build & Import (push) Successful in 10s
CI / Lint / Test / Vet (push) Successful in 11s
The free-text 'channel' filter was dead: it exact-matched SummaryRow.Channel,
which is just the provider ('youtube'), because videos never stored their source
channel. Now they do.

- migration 014: videos.channel_title (nullable; existing rows backfill on the
  next discovery pass, pasted videos immediately).
- discovery (NewVideos) + paste (VideoByID) populate channel_title; UpsertVideo
  persists it, preserving an existing title when an update arrives empty.
- store.DistinctChannels lists a user's channels (RLS-scoped); SummaryRow carries
  ChannelTitle via the shared projection.
- Filter: single Channel -> Channels []string, matching on ChannelTitle; the feed
  renders a multi-select of DistinctChannels (hidden until channels exist).
- migrate tests: 014 reversibility + fixed the relative-step counts in the 010/011
  up/down tests (014 shifted the topology).

TDD throughout: channel persist + distinct, adapter channel wiring, multi-channel
filter match, handler channel filter, migration up/down.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-09 23:02:54 +02:00
mathiasandClaude Opus 4.8 1e65c3b413 fix(web): show paste box to any connected user, not only on an empty feed
CI / Lint / Test / Vet (push) Successful in 11s
CI / Build & Import (push) Successful in 10s
hasConnected was computed only inside the buckets.empty() branch (it was added
for the empty-state copy), so a connected user WITH videos got hasConnected=false
and never saw the paste box (#2-regression of the v0.12.0 paste UI). Compute it
on every list render. Test: connected user with a non-empty feed sees the box.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-09 22:17:00 +02:00
mathiasandClaude Opus 4.8 87c978774f docs(bdd): scenarios for paste-a-URL + onboarding burst, mapped to tests
CI / Lint / Test / Vet (push) Successful in 10s
CI / Build & Import (push) Successful in 9s
Adds paste_url.feature (valid/invalid/not-found/dedup, +@pending no-captions)
and an onboarding-burst scenario on connect; all non-pending scenarios mapped in
the coverage gate to their existing Go tests.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-09 22:06:40 +02:00
mathiasandClaude Opus 4.8 70a9f1d4cd feat(web): in-feed paste box + onboarding-aware connect confirmation
5b: connected users get a 'Summarize any video' URL input on the feed; submit
posts to /paste (HTMX) and swaps the resulting card / inline error into the feed.
7: the connect flash now sets expectations for the async onboarding burst —
'finding your subscriptions, your newest videos will appear below as summarized'.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-09 22:06:40 +02:00
mathiasandClaude Opus 4.8 1d5b2c6365 feat(serve): wire paste fetcher + connect-time onboarding burst
CI / Lint / Test / Vet (push) Successful in 11s
CI / Build & Import (push) Successful in 10s
5c: app.Fetcher = a per-user YouTube videoFetcher, so POST /paste mounts and
resolves arbitrary-video metadata (Feature 2 goes live).

6: the connect trigger now runs an onboarding burst after discovery — summarize
up to TAPIR_ONBOARD_SUMMARIZE_COUNT of the user's newest unsummarized videos via
the gated Processor (Feature 1). Hard cap; explicit so it bypasses recency; every
fetch still through globalFetchGate. No-op when count=0 or queue-only.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-09 21:58:39 +02:00
mathiasandClaude Opus 4.8 62acfee2ed feat(web): paste-a-URL handler — add an arbitrary video + summarize (Feature 2)
POST /paste: parse the video id, fetch metadata via the VideoFetcher port (Data
API, ungated), upsert a subscription-less row scoped to the user (idempotent =
dedup), and — unless already summarized — RequestSummarize + start immediate
processing through the SAME globalFetchGate as the Summarize button. Explicit
paste overrides the recency window; a captionless video degrades to the honest
'no transcript' terminal state via the engine (ADR-010). Invalid URL -> 400,
not-found -> 404, both add nothing. Route mounts only when a Fetcher is wired.

Moves the video-not-found sentinel to domain (shared by adapter + web, no
cross-adapter coupling). Tests: valid add+queue, invalid, not-found, dedup.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-09 21:54:28 +02:00
mathiasandClaude Opus 4.8 2907801aca feat(youtube): VideoByID for arbitrary-video metadata (paste-a-URL)
videos.list (part=snippet) for a single id, including channels the user does
not follow. Data API call (1 quota unit), NOT the rate-limited caption path —
ungated metadata; only the later transcript fetch hits globalFetchGate. Returns
a subscription-less domain.Video scoped to the user, or ErrVideoNotFound for a
deleted/private/typo'd id. Foundation for Feature 2.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-09 21:43:28 +02:00
mathiasandClaude Opus 4.8 59050c4db6 feat(store): NewestUnsummarizedVideoIDs for the onboarding cap
Returns up to limit of a user's newest videos (published_at DESC, NULLS LAST)
that have no summary yet. RLS-scoped via withUser — the test proves a second
user's newer video never leaks. Drives the connect-time onboarding burst
(Feature 1); the caller routes each through the shared rate gate.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-09 21:39:40 +02:00
mathiasandClaude Opus 4.8 c320ed88aa feat(config): add TAPIR_ONBOARD_SUMMARIZE_COUNT (default 3, hard cap 5)
Bounds the connect-time onboarding summary burst (Feature 1). Hard-capped at 5
and clamped (negative->0, >cap->cap) so onboarding can never bulk-fetch; 0
disables. The cap bounds COUNT only — every fetch still flows through the shared
caption rate gate (ADR-014).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-09 21:35:05 +02:00
mathiasandClaude Opus 4.8 bddd75d92e feat(web): parse YouTube video id from pasted URL forms
Pure parser for watch?v=, youtu.be/, shorts/, embed/, and bare ids; rejects
non-YouTube hosts and malformed input. Foundation for paste-a-URL summarize
(Feature 2). No fetch, no gate interaction — parsing only.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-09 21:35:05 +02:00
40 changed files with 2259 additions and 674 deletions
+10 -4
View File
@@ -46,8 +46,9 @@ These caused real mistakes that were caught and corrected; the corrections are l
(See `DECISIONS.md` for full rationale. Listed here so you don't propose them.) (See `DECISIONS.md` for full rationale. Listed here so you don't propose them.)
- **No Supabase** — reuse Dex / ESO+1Password / Postgres (ADR-002). - **No Supabase** — reuse Dex / ESO+1Password / Postgres (ADR-002).
- **No global cross-tenant video/transcript table** — per-user isolation (data-model). Dedup - **No global cross-tenant *video* table** — videos stay per-user (data-model). Transcripts ARE
across users is a Future C concern, not a Stage 0/1 default. shared since ADR-021 (public caption content, keyed by `(provider, provider_video_id)`, non-RLS)
so re-analysis never re-fetches; the *videos* half of cross-tenant dedup stays a Future C concern.
- **No audio-download + speech-to-text in the core path** — captions-first (ADR-007). STT is a - **No audio-download + speech-to-text in the core path** — captions-first (ADR-007). STT is a
deferred, bounded optional component. deferred, bounded optional component.
- **No public SaaS / sign-up / billing / Google OAuth verification at scale** — Future C, - **No public SaaS / sign-up / billing / Google OAuth verification at scale** — Future C,
@@ -85,7 +86,7 @@ Skills live in the canonical library `mathias/skills` and are wired into this re
## Current build state (start here for the first task) ## Current build state (start here for the first task)
The repo is **green and shipping** — last tag `v0.9.0`. `task check` passes (fmt, vet, lint, The repo is **green and shipping** — last tag `v0.14.0`. `task check` passes (fmt, vet, lint,
`go test -p 1 ./...`). Go is `1.26.1` (see `go.mod`). `go test -p 1 ./...`). Go is `1.26.1` (see `go.mod`).
- Clean Architecture core is implemented: `internal/domain` (entities), `internal/ports` - Clean Architecture core is implemented: `internal/domain` (entities), `internal/ports`
@@ -94,13 +95,18 @@ The repo is **green and shipping** — last tag `v0.9.0`. `task check` passes (f
green against it. green against it.
- Adapters present under `internal/adapters/`: `youtube` (captions-first `VideoSource`, - Adapters present under `internal/adapters/`: `youtube` (captions-first `VideoSource`,
timedtext/InnerTube acquisition per ADR-010), `summarizer` + `llm` (the copied AI router, timedtext/InnerTube acquisition per ADR-010), `summarizer` + `llm` (the copied AI router,
Primary→Fallback per ADR-004), `store` (Postgres, golang-migrate migrations 001006), Primary→Fallback per ADR-004), `store` (Postgres, golang-migrate migrations 001015),
`secrets` (file-backed `SecretStore`). The brain HTTP sink (ADR-005) is the remaining `secrets` (file-backed `SecretStore`). The brain HTTP sink (ADR-005) is the remaining
optional sink. optional sink.
- Stage 1 is open (ADR-012): multi-user with **DB-enforced** isolation — Postgres RLS `FORCE`d - Stage 1 is open (ADR-012): multi-user with **DB-enforced** isolation — Postgres RLS `FORCE`d
on all user-owned tables (migration 003), two-user isolation test in on all user-owned tables (migration 003), two-user isolation test in
`internal/adapters/store/rls_test.go`. Registration gate, per-user YouTube web connect, and `internal/adapters/store/rls_test.go`. Registration gate, per-user YouTube web connect, and
account management (disconnect / delete, ADR-013) all shipped. account management (disconnect / delete, ADR-013) all shipped.
- Transcript persistence (ADR-021, migration 015): transcripts are a **shared, non-RLS** store
keyed by `(provider, provider_video_id)` — the single exception to the isolation boundary
(`TestTranscriptsTableIsSharedNotRLS`). The engine reads stored transcripts before any caption
fetch (`usecase.resolveTranscript`), so re-analysis — re-summarize, paste-a-URL, onboarding
burst — never re-touches YouTube. Per-user summaries/videos stay RLS-scoped.
- `cmd/tapir` subcommands: `list`, `show`, `auth` (interactive host-side OAuth), `run` (batch - `cmd/tapir` subcommands: `list`, `show`, `auth` (interactive host-side OAuth), `run` (batch
watch→summarize), `serve` (the HTMX+Templ web reader/writer under `internal/web`, a new watch→summarize), `serve` (the HTMX+Templ web reader/writer under `internal/web`, a new
transport over the unchanged engine/ports — ADR-003). `tapir env` prints config. transport over the unchanged engine/ports — ADR-003). `tapir env` prints config.
+61 -1
View File
@@ -743,6 +743,66 @@ collapse keys off the same window (`App.RecencyWindow=0` → everything inline).
--- ---
## ADR-021 — Persist transcripts as shared, video-keyed public content (re-analysis never re-fetches)
**Status:** Accepted (2026-06-09). **Reopens the transcripts half of** the "Global cross-tenant
`videos`/`transcripts` table" rejection (data-model.md). **Builds on ADR-007** (captions-first),
**ADR-010/ADR-014** (the per-IP caption rate gate), and **ADR-012** (per-user RLS isolation).
**Context.** Every summarization fetches the transcript fresh through the caption path, even when
the exact same transcript was fetched moments ago — for the same user re-summarizing, or for a
second user who happens to watch the same video. The caption fetch is the one genuinely scarce,
genuinely risky operation in the system: YouTube's timedtext endpoint is unofficial and per-IP
rate-limited (ADR-010), and tripping it risks the maintainer's Google standing (ADR-014). So the
operation we most want to *avoid repeating* is the one we currently repeat unconditionally. A
transcript is **public content** — the same words YouTube serves to anyone — and carries nothing
user-identifying. The per-user isolation that protects summaries, feeds, and tokens (ADR-012) is
the wrong shape for it: it forces a re-fetch per user for data that is identical across users.
The original rejection ("Global cross-tenant `videos`/`transcripts` table") bundled videos and
transcripts together and rejected both on the grounds that "at 15 users, re-summarizing is
cheaper than the coupling." That reasoning holds for **videos** (per-user feed rows, genuinely
user-scoped) but not for **transcripts**: the cost being avoided is not LLM re-summarization, it
is a *rate-gated, reputation-risky network fetch*, and that cost is paid per re-fetch regardless
of user count. One re-fetch avoided is strictly worth more than the coupling it removes.
**Decision.**
1. **A single shared `transcripts` table, keyed by the cross-user dedup key
`(provider, provider_video_id)`** — the stable public identity of the video, not Tapir's
internal per-user `videos.id`. Columns: the key, `source` (`captions`/`none`), `language`,
`content`, `fetched_at`. It holds **only public caption content + the video's public id**
nothing user-identifying — and is therefore **NOT RLS-scoped**: no `user_id`, no policy, no
`FORCE ROW LEVEL SECURITY`. This is the deliberate, single exception to the ADR-012 isolation
boundary, and the only one.
2. **Summarize path becomes read-stored-first.** Have a stored transcript for this video? →
summarize from the stored text, **no caption fetch**. No stored transcript? → fetch *through
the unchanged gate* (ADR-014) → store it → summarize. The gate is neither bypassed nor
weakened; persistence reduces how *often* we reach it, never how *fast*.
3. **De-facto cross-user dedup is the intended behaviour, not a feature with a switch.** Two
users who share a video share the one transcript row. A permanent `source = 'none'` (no
captions) is stored too, so a known-caption-less video is not re-fetched by anyone. A
transient 429 (`SourceRateLimited`) is **never** stored as terminal — it stays a per-user
retry via the existing `transcript_status` backoff (ADR-014), so persistence cannot mask a
rate-limit into a false "no transcript."
4. **Per-user `summaries` stay RLS-scoped (ADR-012 unchanged)** and reference the transcript by
video id. Videos stay per-user. Only transcripts go shared.
**Consequences.** Re-analysis (re-summarize, different model, paste of an already-seen video,
onboarding of a second user with overlapping subscriptions) never re-touches YouTube — the
primary win, and it *reduces* aggregate caption-gate pressure, reinforcing ADR-010/ADR-014 rather
than straining them. The isolation surface gains exactly one non-RLS table; an isolation test
asserts the boundary is *exactly* there and has not leaked to any user-owned table (this is the
proof the public-content classification was implemented as designed). It also unblocks
multi-model / customizable analysis (re-run analysis on stored text for free) — enabling that is
this ADR's point; building it is separate.
**Reversibility.** The read-stored-first check is the only behavioural coupling; removing it
restores fetch-every-time. The down-migration recreates the per-user RLS-scoped transcripts shape
(001/003). No user-facing surface depends on cross-user sharing — sharing is the *storage shape*,
never exposed in the UI.
---
## Rejected alternatives ## Rejected alternatives
Approaches considered during the 2026-06-02 planning + grill session and **deliberately not Approaches considered during the 2026-06-02 planning + grill session and **deliberately not
@@ -757,7 +817,7 @@ maps to the ADR that settles it.
| Lifting shared packages into a `brain-common` module | Couples Tapir's release cycle to the monolith for negligible code savings | ADR-004 | | Lifting shared packages into a `brain-common` module | Couples Tapir's release cycle to the monolith for negligible code savings | ADR-004 |
| Importing/replicating the filesystem `brain` package | Assumes co-location with the brain git checkout; wrong for a standalone networked service | ADR-005 | | Importing/replicating the filesystem `brain` package | Assumes co-location with the brain git checkout; wrong for a standalone networked service | ADR-005 |
| Reusing `ingestion`'s `oauth` package for YouTube/Vimeo | Same name, opposite direction — it's inbound MCP-server auth, not outbound provider OAuth | ADR-006 | | Reusing `ingestion`'s `oauth` package for YouTube/Vimeo | Same name, opposite direction — it's inbound MCP-server auth, not outbound provider OAuth | ADR-006 |
| Global cross-tenant `videos`/`transcripts` table (dedup) | Reintroduces the cross-domain DB coupling the homelab review is removing; at 15 users, re-summarizing is cheaper than the coupling | data-model.md | | Global cross-tenant `videos`/`transcripts` table (dedup) | Reintroduces the cross-domain DB coupling the homelab review is removing; at 15 users, re-summarizing is cheaper than the coupling. **Transcripts half reopened by ADR-021** — the avoided cost there is a rate-gated, reputation-risky *caption fetch*, not LLM re-summarization, so it outweighs the coupling; **videos stay per-user.** | data-model.md, **ADR-021** (transcripts only) |
| Audio-download + Whisper STT in the core path | ToS-grey, breakage-prone (yt-dlp), contends for koala GPU with the JEPA PoC; captions alone test the core hypothesis | ADR-007 | | Audio-download + Whisper STT in the core path | ToS-grey, breakage-prone (yt-dlp), contends for koala GPU with the JEPA PoC; captions alone test the core hypothesis | ADR-007 |
| Building multi-tenant SaaS / Google OAuth verification now | "Real users soon" was lowered to Future B; SaaS machinery before the Stage 0 self-use gate is the primary documented anti-goal | ADR-008, VISION | | Building multi-tenant SaaS / Google OAuth verification now | "Real users soon" was lowered to Future B; SaaS machinery before the Stage 0 self-use gate is the primary documented anti-goal | ADR-008, VISION |
| Delegating the S5 reuse spike to an agent swarm | A 1-hour sequential read-and-judge with a single coupled conclusion; orchestration overhead exceeds the work, and it's Diamond-1 judgment the maintainer wanted to own | (process note) | | Delegating the S5 reuse spike to an agent swarm | A 1-hour sequential read-and-judge with a single coupled conclusion; orchestration overhead exceeds the work, and it's Diamond-1 judgment the maintainer wanted to own | (process note) |
+6
View File
@@ -32,6 +32,9 @@ func serialize(mu *sync.Mutex, run discoveryRunner) discoveryRunner {
type discoveryTrigger struct { type discoveryTrigger struct {
ctx context.Context ctx context.Context
run discoveryRunner run discoveryRunner
// onboard, when set, runs after the discovery pass to summarize a capped number
// of the user's newest videos (Feature 1). Optional.
onboard func(ctx context.Context, userID string)
log *slog.Logger log *slog.Logger
} }
@@ -40,5 +43,8 @@ func (t *discoveryTrigger) Enqueue(userID string) {
if _, err := t.run(t.ctx, userID); err != nil { if _, err := t.run(t.ctx, userID); err != nil {
t.log.Warn("discovery: connect-triggered pass had errors", "user", userID, "err", err) t.log.Warn("discovery: connect-triggered pass had errors", "user", userID, "err", err)
} }
if t.onboard != nil {
t.onboard(t.ctx, userID)
}
}() }()
} }
+30 -3
View File
@@ -210,7 +210,10 @@ func cmdServe(ctx context.Context, log *slog.Logger) error {
ClientSecret: cfg.YTClientSecret, ClientSecret: cfg.YTClientSecret,
RedirectURL: cfg.YTConnectRedirectURL, RedirectURL: cfg.YTConnectRedirectURL,
}, secretStore, st, log) }, secretStore, st, log)
log.Info("web youtube connect enabled", "redirect", cfg.YTConnectRedirectURL) // Paste-a-URL (Feature 2): same YouTube credentials, per-user adapter built
// per request. Mounting the /paste route keys off app.Fetcher being set.
app.Fetcher = videoFetcher{cfg: cfg, secrets: secretStore}
log.Info("web youtube connect + paste enabled", "redirect", cfg.YTConnectRedirectURL)
} }
// Immediate summarization for the web "Summarize" button. When the engine can // Immediate summarization for the web "Summarize" button. When the engine can
@@ -249,9 +252,33 @@ func cmdServe(ctx context.Context, log *slog.Logger) error {
// One lock shared by the scheduler and connect-triggered passes (#6) so // One lock shared by the scheduler and connect-triggered passes (#6) so
// they never fetch concurrently — the single-fetcher invariant (ADR-018). // they never fetch concurrently — the single-fetcher invariant (ADR-018).
runUser := serialize(&sync.Mutex{}, rawRunUser) runUser := serialize(&sync.Mutex{}, rawRunUser)
// Onboarding burst (Feature 1): after the connect-triggered discovery pass,
// summarize up to OnboardSummarizeCount of the user's NEWEST unsummarized
// videos so a fresh account gets real summaries in its first session. Hard
// cap; explicit, so it bypasses the recency window — but every fetch still
// goes through globalFetchGate via the Processor. No-op when disabled
// (count 0) or queue-only (no Processor).
onboard := func(ctx context.Context, userID string) {
if cfg.OnboardSummarizeCount <= 0 || app.Processor == nil {
return
}
ids, err := st.NewestUnsummarizedVideoIDs(ctx, userID, cfg.OnboardSummarizeCount)
if err != nil {
log.Warn("onboarding: list newest unsummarized", "user", userID, "err", err)
return
}
for _, id := range ids {
if err := app.Processor.ProcessVideo(ctx, userID, id); err != nil {
log.Warn("onboarding: summarize", "user", userID, "video", id, "err", err)
}
}
if len(ids) > 0 {
log.Info("onboarding burst complete", "user", userID, "summarized", len(ids), "cap", cfg.OnboardSummarizeCount)
}
}
if app.Connect != nil { if app.Connect != nil {
app.Connect.Discovery = &discoveryTrigger{ctx: ctx, run: runUser, log: log} app.Connect.Discovery = &discoveryTrigger{ctx: ctx, run: runUser, onboard: onboard, log: log}
log.Info("connect-triggered discovery enabled") log.Info("connect-triggered discovery enabled", "onboard_cap", cfg.OnboardSummarizeCount)
} }
go runScheduler(ctx, cfg.DiscoveryInterval, st, runUser, log) go runScheduler(ctx, cfg.DiscoveryInterval, st, runUser, log)
} else { } else {
+26 -1
View File
@@ -11,9 +11,29 @@ import (
"gitea.d-ma.be/mathias/tapir/internal/adapters/youtube" "gitea.d-ma.be/mathias/tapir/internal/adapters/youtube"
"gitea.d-ma.be/mathias/tapir/internal/config" "gitea.d-ma.be/mathias/tapir/internal/config"
"gitea.d-ma.be/mathias/tapir/internal/domain" "gitea.d-ma.be/mathias/tapir/internal/domain"
"gitea.d-ma.be/mathias/tapir/internal/ports"
"gitea.d-ma.be/mathias/tapir/internal/usecase" "gitea.d-ma.be/mathias/tapir/internal/usecase"
"gitea.d-ma.be/mathias/tapir/internal/web"
) )
// videoFetcher adapts the YouTube adapter to web.VideoFetcher for the paste flow
// (Feature 2). It builds a per-user adapter bound to that user's token ref and
// resolves a single video's metadata via the Data API — ungated; only the later
// transcript fetch goes through globalFetchGate.
type videoFetcher struct {
cfg config.Config
secrets ports.SecretStore
}
func (f videoFetcher) FetchVideo(ctx context.Context, userID, videoID string) (domain.Video, error) {
a := youtube.New(youtube.Config{
ClientID: f.cfg.YTClientID,
ClientSecret: f.cfg.YTClientSecret,
TokenSecretRef: web.YouTubeTokenRef(userID),
}, f.secrets)
return a.VideoByID(ctx, userID, videoID)
}
// buildProcessor wires the summarization engine — YouTube source (captions-first), // buildProcessor wires the summarization engine — YouTube source (captions-first),
// AI-router summarizer, store sink — shared by `tapir run` and the web // AI-router summarizer, store sink — shared by `tapir run` and the web
// "Summarize now" path so the wiring lives in one place. It returns (nil, nil) — // "Summarize now" path so the wiring lives in one place. It returns (nil, nil) —
@@ -43,7 +63,12 @@ func buildProcessor(cfg config.Config, st *store.Store) (*usecase.Engine, error)
} }
sum := summarizer.New(primary, nil) sum := summarizer.New(primary, nil)
return usecase.NewEngine(src, sum, st), nil // The store is both the summary sink and the shared transcript cache (ADR-021):
// the engine reads stored transcripts before any caption fetch and writes
// resolved ones back, so re-analysis never re-touches YouTube.
eng := usecase.NewEngine(src, sum, st)
eng.Transcripts = st
return eng, nil
} }
// engineProcessor adapts the engine (which works in terms of a domain.Video) to // engineProcessor adapts the engine (which works in terms of a domain.Video) to
+21 -13
View File
@@ -10,12 +10,15 @@ only opaque references to them; the secret material lives in ESO/1Password (ADR-
## Design decisions baked into this model ## Design decisions baked into this model
- **Per-user isolation, not a shared global video table.** The earlier draft proposed a - **Per-user isolation for everything except transcripts.** The earlier draft proposed a
global `videos`/`transcripts` table deduped across tenants. Rejected for Future B: it global `videos`/`transcripts` table deduped across tenants. **Videos** stay per-user and
reintroduces exactly the cross-domain coupling the homelab architecture review is RLS-scoped — a shared video table reintroduces exactly the cross-domain coupling the homelab
removing, and at 15 users the cost of occasionally re-summarizing the same video is architecture review is removing. **Transcripts**, however, are now shared (ADR-021): keyed by
trivial compared to the isolation it would cost. Each user's data is self-contained. `(provider, provider_video_id)`, no `user_id`, **not** RLS-scoped. The cost avoided there is
(Revisit only if Future C makes GPU/transcription cost dominate — a new ADR, not a default.) not LLM re-summarization but a rate-gated, reputation-risky caption fetch (ADR-010/014), which
is paid per re-fetch regardless of user count — so persisting public caption content once and
sharing it strictly beats the coupling it removes. Everything else each user owns is
self-contained; `rls_test.go` proves transcripts is the single exception.
- **Secrets by reference only.** Tables hold a `secret_ref` (opaque string/UUID resolved via - **Secrets by reference only.** Tables hold a `secret_ref` (opaque string/UUID resolved via
the `SecretStore` port), never tokens or keys. the `SecretStore` port), never tokens or keys.
- **The brain sink is just a delivery target.** No brain-specific tables. Whether a summary - **The brain sink is just a delivery target.** No brain-specific tables. Whether a summary
@@ -34,7 +37,7 @@ erDiagram
USER ||--o{ AI_CREDENTIAL : "has (planned)" USER ||--o{ AI_CREDENTIAL : "has (planned)"
VIDEO_CONNECTION ||--o{ SUBSCRIPTION : "exposes (planned)" VIDEO_CONNECTION ||--o{ SUBSCRIPTION : "exposes (planned)"
SUBSCRIPTION ||--o{ VIDEO : "produces (per user)" SUBSCRIPTION ||--o{ VIDEO : "produces (per user)"
VIDEO ||--o| TRANSCRIPT : "has at most one" VIDEO }o--o| TRANSCRIPT : "shares one by (provider, provider_video_id) — not FK (ADR-021)"
VIDEO ||--o| SUMMARY : "has at most one" VIDEO ||--o| SUMMARY : "has at most one"
SUMMARY ||--o{ SINK_DELIVERY : "delivered via" SUMMARY ||--o{ SINK_DELIVERY : "delivered via"
USER ||--o{ CHANNEL_ERROR : "reports unavailable channels" USER ||--o{ CHANNEL_ERROR : "reports unavailable channels"
@@ -92,12 +95,12 @@ erDiagram
timestamptz rate_limited_at "backoff clock for 429 retries (migration 007)" timestamptz rate_limited_at "backoff clock for 429 retries (migration 007)"
} }
TRANSCRIPT { TRANSCRIPT {
uuid video_id PK_FK text provider PK "part of shared key (ADR-021)"
uuid user_id FK text provider_video_id PK "part of shared key — the cross-user dedup key"
text source "captions | none" text source "captions | none"
text language text language
text content "null when source = none" text content "null when source = none"
timestamptz resolved_at timestamptz fetched_at
} }
SUMMARY { SUMMARY {
uuid id PK uuid id PK
@@ -173,8 +176,12 @@ mechanism.
`transcript_status` and `rate_limited_at` (migration 007) track caption-fetch outcomes for `transcript_status` and `rate_limited_at` (migration 007) track caption-fetch outcomes for
rate-limit backoff: `NULL` = not attempted; `rate_limited` = 429 seen, skip until rate-limit backoff: `NULL` = not attempted; `rate_limited` = 429 seen, skip until
`NOW() - rate_limited_at > TAPIR_FETCH_BACKOFF`; `fetched` = resolved; `none` = no transcript. `NOW() - rate_limited_at > TAPIR_FETCH_BACKOFF`; `fetched` = resolved; `none` = no transcript.
- **TRANSCRIPT** — at most one per video. `source = none` records "checked, no usable - **TRANSCRIPT** — shared public caption content, one row per `(provider, provider_video_id)`,
transcript" so the watcher doesn't reprocess (ADR-007). `content` null in that case. **not** RLS-scoped and carrying no `user_id` (ADR-021). Two users who watch the same video
share the one row; the summarize path reads it before any caption fetch, so re-analysis never
re-touches YouTube (ADR-010/014). `source = none` records "checked, no usable transcript" so
no one reprocesses (ADR-007); `content` null in that case. A transient 429 is never stored
here — it stays a per-user retry via `VIDEO.transcript_status`.
- **SUMMARY** — at most one per video. `fallback_used` + `ai_provider`/`ai_model` make the - **SUMMARY** — at most one per video. `fallback_used` + `ai_provider`/`ai_model` make the
"is local good enough?" question queryable (the Stage 0 quality signal). `highlights`/ "is local good enough?" question queryable (the Stage 0 quality signal). `highlights`/
`takeaways` as jsonb to stay schema-flexible while the output format settles. `takeaways` as jsonb to stay schema-flexible while the output format settles.
@@ -231,7 +238,8 @@ queue, doesn't replace it). Deferred until there's a reason.
## Explicitly out of scope (Future C) ## Explicitly out of scope (Future C)
- Global cross-tenant video/transcript dedup (rejected above). - Global cross-tenant *video* dedup (rejected above). Note: cross-tenant *transcript* sharing
is now in scope and shipped (ADR-021); only the videos half stays per-user.
- Sharding / per-tenant physical databases. - Sharding / per-tenant physical databases.
- Soft-delete + full audit trail on connections/credentials (a Stage 2 hardening item; add - Soft-delete + full audit trail on connections/credentials (a Stage 2 hardening item; add
via ADR when Stage 2 work starts). via ADR when Stage 2 work starts).
+6
View File
@@ -24,6 +24,12 @@ Feature: Connect and manage video accounts
Then a discovery pass for my account is triggered right away Then a discovery pass for my account is triggered right away
And I do not have to wait for the next scheduled pass to see my videos And I do not have to wait for the next scheduled pass to see my videos
Scenario: Connecting summarizes my newest videos right away
Given I have no connected video accounts
When I connect my YouTube account
Then up to the onboarding cap of my newest videos are summarized through the rate gate
And the rest are left to the scheduled recency-bounded pass
Scenario: Tokens are never stored in the clear Scenario: Tokens are never stored in the clear
When I connect any video account When I connect any video account
Then no OAuth token value is stored in the database Then no OAuth token value is stored in the database
+30
View File
@@ -0,0 +1,30 @@
Feature: Paste a YouTube URL to summarize any video
As a user
I want to paste a YouTube link and get a summary
So that I can pull the specific video I want now, even from channels I don't follow
Scenario: Paste a valid YouTube URL
Given I am connected
When I paste a valid YouTube video URL
Then the video is added to my feed scoped to me
And it is queued for summarization through the shared rate gate
Scenario: Pasting an invalid link is rejected
When I paste something that is not a YouTube video URL
Then I get a clear error and nothing is added
Scenario: Pasting a video that cannot be found is honest
When I paste a URL whose video cannot be found
Then I am told it couldn't be found and nothing is added
Scenario: Pasting the same video twice does not duplicate it
Given I have pasted a video
When I paste the same video again
Then my feed still has exactly one entry for it
@pending
# Covered by the engine's ADR-010 no-transcript terminal state (degrade-never-error);
# there is no paste-specific test for it.
Scenario: A pasted video with no captions resolves honestly
When I paste a video that has no captions
Then it resolves to the "no transcript available" terminal state
@@ -31,5 +31,16 @@ Feature: Summarize new videos from subscribed channels
When the watcher sees "Designing for Attention" again When the watcher sees "Designing for Attention" again
Then Tapir does not produce a second summary for it Then Tapir does not produce a second summary for it
Scenario: Re-analyzing a stored video does not re-fetch its transcript
Given a transcript for "Designing for Attention" is already stored
When the video is summarized again
Then Tapir reads the stored transcript
And Tapir does not fetch captions from YouTube
# Captions-first is the core path (ADR-007). Audio-download + speech-to-text is # Captions-first is the core path (ADR-007). Audio-download + speech-to-text is
# deferred and intentionally has no scenario here yet. # deferred and intentionally has no scenario here yet.
#
# Transcript persistence (ADR-021): the stored transcript is shared, keyed by
# (provider, provider_video_id) and read before any caption fetch, so the
# re-analysis scenario above also covers paste-a-URL and the onboarding burst —
# both summarize through the same engine chokepoint.
+7 -4
View File
@@ -10,10 +10,13 @@ import (
// DeleteUser permanently removes a user and all of their data. It runs through // DeleteUser permanently removes a user and all of their data. It runs through
// withUser so RLS confines every statement to the calling user's own rows. // withUser so RLS confines every statement to the calling user's own rows.
// //
// Deleting the users row cascades (ON DELETE CASCADE) to videos, transcripts, // Deleting the users row cascades (ON DELETE CASCADE) to videos, summaries
// summaries (→ sink_deliveries), video_connections, and the user_identities map // (→ sink_deliveries), video_connections, and the user_identities map
// referential-integrity cascades bypass RLS, so a user's child rows are removed // referential-integrity cascades bypass RLS, so a user's child rows are removed
// even though the deleting connection is scoped. summary_actions and login_events // even though the deleting connection is scoped. Transcripts are NOT removed:
// since ADR-021 they are shared public content keyed by (provider,
// provider_video_id) with no user_id, so another user may still reference the
// same row — a user deletion must not strip shared caption content. summary_actions and login_events
// are the exceptions: each carries a user_id but has NO foreign key to users // are the exceptions: each carries a user_id but has NO foreign key to users
// (migrations 002 and 010), so the cascade does not reach them; they are deleted // (migrations 002 and 010), so the cascade does not reach them; they are deleted
// explicitly in the same scoped transaction. Deleting an absent user is a no-op // explicitly in the same scoped transaction. Deleting an absent user is a no-op
+37 -2
View File
@@ -53,7 +53,11 @@ func TestMigration010LoginEventsUpDown(t *testing.T) {
require.True(t, loginEventsExists(t), "login_events must exist at latest migration") require.True(t, loginEventsExists(t), "login_events must exist at latest migration")
m := fileMigrator(t) m := fileMigrator(t)
// 011, 012, 013 sit above 010; step them down first so 010 is exercised in isolation. // 011..015 sit above 010; step them down first so 010 is exercised in isolation.
require.NoError(t, m.Steps(-1), "down 015 reshapes transcripts, login_events intact")
require.True(t, loginEventsExists(t), "015 down leaves login_events intact")
require.NoError(t, m.Steps(-1), "down 014 drops channel_title, login_events intact")
require.True(t, loginEventsExists(t), "014 down leaves login_events intact")
require.NoError(t, m.Steps(-1), "down 013 drops channel_errors, login_events intact") require.NoError(t, m.Steps(-1), "down 013 drops channel_errors, login_events intact")
require.True(t, loginEventsExists(t), "013 down leaves login_events intact") require.True(t, loginEventsExists(t), "013 down leaves login_events intact")
require.NoError(t, m.Steps(-1), "down 012 is a no-op, login_events intact") require.NoError(t, m.Steps(-1), "down 012 is a no-op, login_events intact")
@@ -64,7 +68,7 @@ func TestMigration010LoginEventsUpDown(t *testing.T) {
require.NoError(t, m.Steps(-1), "down 010 must drop login_events") require.NoError(t, m.Steps(-1), "down 010 must drop login_events")
require.False(t, loginEventsExists(t), "login_events must be gone after the down migration") require.False(t, loginEventsExists(t), "login_events must be gone after the down migration")
require.NoError(t, m.Steps(4), "up must recreate 010 then re-apply 011, 012, 013") require.NoError(t, m.Steps(6), "up must recreate 010 then re-apply 011..015")
require.True(t, loginEventsExists(t), "login_events must be restored after the up migration") require.True(t, loginEventsExists(t), "login_events must be restored after the up migration")
} }
@@ -87,6 +91,8 @@ func TestMigration011AutoSummarizeDefaultUpDown(t *testing.T) {
require.Equal(t, "true", autoSummarizeDefault(t), "011 sets the default to TRUE") require.Equal(t, "true", autoSummarizeDefault(t), "011 sets the default to TRUE")
m := fileMigrator(t) m := fileMigrator(t)
require.NoError(t, m.Steps(-1), "down 015 reshapes transcripts")
require.NoError(t, m.Steps(-1), "down 014 drops channel_title")
require.NoError(t, m.Steps(-1), "down 013 drops channel_errors") require.NoError(t, m.Steps(-1), "down 013 drops channel_errors")
require.NoError(t, m.Steps(-1), "down 012 is a no-op") require.NoError(t, m.Steps(-1), "down 012 is a no-op")
require.NoError(t, m.Steps(-1), "down 011 reverts the column default") require.NoError(t, m.Steps(-1), "down 011 reverts the column default")
@@ -96,6 +102,35 @@ func TestMigration011AutoSummarizeDefaultUpDown(t *testing.T) {
require.Equal(t, "true", autoSummarizeDefault(t)) require.Equal(t, "true", autoSummarizeDefault(t))
require.NoError(t, m.Steps(1), "up 012 runs clean (no FORCE RLS on fresh schema)") require.NoError(t, m.Steps(1), "up 012 runs clean (no FORCE RLS on fresh schema)")
require.NoError(t, m.Steps(1), "up 013 creates channel_errors") require.NoError(t, m.Steps(1), "up 013 creates channel_errors")
require.NoError(t, m.Steps(1), "up 014 recreates channel_title")
require.NoError(t, m.Steps(1), "up 015 reshapes transcripts to shared")
}
// channelTitleExists reports whether videos.channel_title is present.
func channelTitleExists(t *testing.T) bool {
t.Helper()
var exists bool
require.NoError(t, rawPool(t).QueryRow(context.Background(),
`SELECT EXISTS (SELECT 1 FROM information_schema.columns
WHERE table_name = 'videos' AND column_name = 'channel_title')`).Scan(&exists))
return exists
}
// TestMigration014VideoChannelTitleUpDown proves 014 is reversible: down drops
// videos.channel_title, up recreates it.
func TestMigration014VideoChannelTitleUpDown(t *testing.T) {
newStore(t) // latest (014 applied)
require.True(t, channelTitleExists(t), "channel_title exists at latest migration")
m := fileMigrator(t)
require.NoError(t, m.Steps(-1), "down 015 reshapes transcripts, channel_title intact")
require.True(t, channelTitleExists(t), "015 down leaves channel_title intact")
require.NoError(t, m.Steps(-1), "down 014 must drop channel_title")
require.False(t, channelTitleExists(t), "channel_title must be gone after the down migration")
require.NoError(t, m.Steps(1), "up 014 must recreate channel_title")
require.True(t, channelTitleExists(t), "channel_title must be restored after the up migration")
require.NoError(t, m.Steps(1), "up 015 restores the shared transcripts shape (HEAD)")
} }
// TestMigration012FixAutoSummarizeRLS proves 012 runs cleanly and flips any // TestMigration012FixAutoSummarizeRLS proves 012 runs cleanly and flips any
@@ -0,0 +1 @@
ALTER TABLE videos DROP COLUMN channel_title;
@@ -0,0 +1,8 @@
-- Store the source channel's title per video so the list can offer a real
-- channel filter (multi-select of the user's channels) instead of the dead
-- free-text field that only ever matched the provider string. Nullable: existing
-- rows backfill on the next discovery pass (UpsertVideo writes it); pasted videos
-- get it immediately from videos.list. No FK to a channels table at Stage 0 — the
-- title is a denormalised display/filter value, consistent with the existing
-- subscription_id-stays-NULL stance (data-model.md).
ALTER TABLE videos ADD COLUMN channel_title TEXT;
@@ -0,0 +1,19 @@
-- Down 015: restore the per-user RLS-scoped transcripts shape (001 + 003).
DROP TABLE transcripts;
CREATE TABLE transcripts (
video_id UUID PRIMARY KEY REFERENCES videos(id) ON DELETE CASCADE,
user_id UUID NOT NULL REFERENCES users(id) ON DELETE CASCADE,
source TEXT NOT NULL,
language TEXT,
content TEXT,
resolved_at TIMESTAMPTZ NOT NULL DEFAULT NOW()
);
CREATE INDEX idx_transcripts_user_id ON transcripts(user_id);
ALTER TABLE transcripts ENABLE ROW LEVEL SECURITY;
ALTER TABLE transcripts FORCE ROW LEVEL SECURITY;
CREATE POLICY transcripts_isolation ON transcripts
FOR ALL
USING (user_id = current_setting('tapir.current_user_id', true)::uuid);
@@ -0,0 +1,30 @@
-- Migration 015: transcripts become SHARED public-content storage (ADR-021).
--
-- The per-user transcripts table from 001 (PK videos.id, user_id NOT NULL, RLS
-- FORCEd in 003) was dead: no application code ever read or wrote it — only the
-- transcript_status columns on `videos` (007) carried fetch outcomes. ADR-021
-- repurposes it as the single shared store of public caption content, keyed by
-- the cross-user dedup key (provider, provider_video_id) — the video's public
-- identity, not Tapir's per-user videos.id — so re-analysis never re-fetches
-- from YouTube (ADR-010/014).
--
-- It holds ONLY public caption content + the video's public id (nothing
-- user-identifying), so it is deliberately NOT RLS-scoped: no user_id, no
-- policy, no FORCE. This is the single, intentional exception to the ADR-012
-- isolation boundary; rls_test.go asserts the boundary is exactly here and
-- nowhere else. Dropping the old table drops its RLS policy with it; it held no
-- real data, so drop+recreate loses nothing.
DROP TABLE transcripts;
CREATE TABLE transcripts (
provider TEXT NOT NULL,
provider_video_id TEXT NOT NULL,
source TEXT NOT NULL, -- 'captions' (content set) | 'none' (no captions; content NULL)
language TEXT,
content TEXT,
fetched_at TIMESTAMPTZ NOT NULL DEFAULT NOW(),
PRIMARY KEY (provider, provider_video_id)
);
COMMENT ON TABLE transcripts IS
'Shared public caption content keyed by (provider, provider_video_id). NOT RLS-scoped — public content only, de-facto cross-user dedup (ADR-021).';
+4 -1
View File
@@ -29,6 +29,7 @@ type SummaryRow struct {
ProviderVideoID string // videos.provider_video_id; empty when no videos row ProviderVideoID string // videos.provider_video_id; empty when no videos row
Title string // videos.title; empty when no videos row Title string // videos.title; empty when no videos row
Channel string // videos.provider for now; empty when no videos row Channel string // videos.provider for now; empty when no videos row
ChannelTitle string // videos.channel_title; the source channel, for display + filtering
URL string // videos.url; empty when no videos row URL string // videos.url; empty when no videos row
PublishedAt time.Time // videos.published_at; zero when absent PublishedAt time.Time // videos.published_at; zero when absent
Summary string Summary string
@@ -137,7 +138,8 @@ const selectVideo = `
COALESCE(s.created_at, v.seen_at), COALESCE(s.created_at, v.seen_at),
(s.id IS NOT NULL) AS summarized, (s.id IS NOT NULL) AS summarized,
v.summarize_requested, v.summarize_requested,
COALESCE(v.transcript_status, '') COALESCE(v.transcript_status, ''),
COALESCE(v.channel_title, '')
FROM videos v FROM videos v
LEFT JOIN summaries s ON s.video_id = v.id AND s.user_id = v.user_id` LEFT JOIN summaries s ON s.video_id = v.id AND s.user_id = v.user_id`
@@ -254,6 +256,7 @@ func scanVideoRow(rows pgx.Row) (SummaryRow, error) {
&row.Summarized, &row.Summarized,
&row.SummarizeRequested, &row.SummarizeRequested,
&row.TranscriptStatus, &row.TranscriptStatus,
&row.ChannelTitle,
); err != nil { ); err != nil {
return SummaryRow{}, fmt.Errorf("store: scan video: %w", err) return SummaryRow{}, fmt.Errorf("store: scan video: %w", err)
} }
+70 -14
View File
@@ -22,9 +22,11 @@ import (
// (no GUC set → zero rows) proves the enforcement path is live, not bypassed. // (no GUC set → zero rows) proves the enforcement path is live, not bypassed.
// userIsolatedTables are the tables that carry a user_id and whose policy keys // userIsolatedTables are the tables that carry a user_id and whose policy keys
// directly off the tapir.current_user_id GUC. // directly off the tapir.current_user_id GUC. transcripts is deliberately ABSENT
// — ADR-021 made it shared public content (non-RLS); TestTranscriptsTableIsSharedNotRLS
// proves that is the only place the isolation boundary moved.
var userIsolatedTables = []string{ var userIsolatedTables = []string{
"users", "videos", "transcripts", "summaries", "summary_actions", "login_events", "video_connections", "users", "videos", "summaries", "summary_actions", "login_events", "video_connections",
} }
// allIsolatedTables adds sink_deliveries, whose ownership is derived from its // allIsolatedTables adds sink_deliveries, whose ownership is derived from its
@@ -38,9 +40,10 @@ type seeded struct {
summaryID string summaryID string
} }
// seedUser inserts one full chain (user → video → transcript → summary → // seedUser inserts one full chain (user → video → summary → action → delivery)
// action → delivery) as the superuser pool, which bypasses RLS so both users' // as the superuser pool, which bypasses RLS so both users' data lands regardless
// data lands regardless of the GUC. // of the GUC. Transcripts are NOT seeded here: they are shared, non-RLS public
// content (ADR-021), so they have no place in a per-user isolation chain.
func seedUser(t *testing.T, p *pgxpool.Pool, userID string) seeded { func seedUser(t *testing.T, p *pgxpool.Pool, userID string) seeded {
t.Helper() t.Helper()
ctx := context.Background() ctx := context.Background()
@@ -54,11 +57,6 @@ func seedUser(t *testing.T, p *pgxpool.Pool, userID string) seeded {
VALUES ($1, 'youtube', $2, 'title') RETURNING id`, VALUES ($1, 'youtube', $2, 'title') RETURNING id`,
userID, "vid-"+userID).Scan(&videoID)) userID, "vid-"+userID).Scan(&videoID))
_, err = p.Exec(ctx,
`INSERT INTO transcripts (video_id, user_id, source, content)
VALUES ($1, $2, 'captions', 'words')`, videoID, userID)
require.NoError(t, err)
var summaryID string var summaryID string
require.NoError(t, p.QueryRow(ctx, require.NoError(t, p.QueryRow(ctx,
`INSERT INTO summaries (user_id, video_id, summary) VALUES ($1, $2, 'sum') `INSERT INTO summaries (user_id, video_id, summary) VALUES ($1, $2, 'sum')
@@ -92,9 +90,16 @@ func appPool(t *testing.T, super *pgxpool.Pool) *pgxpool.Pool {
t.Helper() t.Helper()
ctx := context.Background() ctx := context.Background()
// Idempotent across test runs (schema/role persist for the TestMain PG). // Idempotent across tests AND runs: the role persists for the TestMain PG and
_, _ = super.Exec(ctx, `DROP ROLE IF EXISTS app`) // owns granted privileges, so a plain DROP ROLE fails once any GRANT exists
_, err := super.Exec(ctx, `CREATE ROLE app LOGIN PASSWORD 'app'`) // (and more than one test now builds an app pool). Create only if absent; the
// GRANTs below are themselves idempotent.
_, err := super.Exec(ctx,
`DO $$ BEGIN
IF NOT EXISTS (SELECT FROM pg_roles WHERE rolname = 'app') THEN
CREATE ROLE app LOGIN PASSWORD 'app';
END IF;
END $$`)
require.NoError(t, err) require.NoError(t, err)
_, err = super.Exec(ctx, `GRANT USAGE ON SCHEMA public TO app`) _, err = super.Exec(ctx, `GRANT USAGE ON SCHEMA public TO app`)
require.NoError(t, err) require.NoError(t, err)
@@ -186,7 +191,6 @@ func TestRLSEnforcesPerUserIsolation(t *testing.T) {
{"update users", `UPDATE users SET display_name = 'hacked' WHERE id = $1`, b.userID}, {"update users", `UPDATE users SET display_name = 'hacked' WHERE id = $1`, b.userID},
{"update videos", `UPDATE videos SET title = 'hacked' WHERE user_id = $1`, b.userID}, {"update videos", `UPDATE videos SET title = 'hacked' WHERE user_id = $1`, b.userID},
{"queue videos summarize", `UPDATE videos SET summarize_requested = TRUE WHERE id = $1`, b.videoID}, {"queue videos summarize", `UPDATE videos SET summarize_requested = TRUE WHERE id = $1`, b.videoID},
{"update transcripts", `UPDATE transcripts SET content = 'hacked' WHERE user_id = $1`, b.userID},
{"update summaries", `UPDATE summaries SET summary = 'hacked' WHERE user_id = $1`, b.userID}, {"update summaries", `UPDATE summaries SET summary = 'hacked' WHERE user_id = $1`, b.userID},
{"update summary_actions", `UPDATE summary_actions SET action = 'skipped' WHERE user_id = $1`, b.userID}, {"update summary_actions", `UPDATE summary_actions SET action = 'skipped' WHERE user_id = $1`, b.userID},
{"update login_events", `UPDATE login_events SET seen_at = NOW() WHERE user_id = $1`, b.userID}, {"update login_events", `UPDATE login_events SET seen_at = NOW() WHERE user_id = $1`, b.userID},
@@ -235,3 +239,55 @@ func TestRLSEnforcesPerUserIsolation(t *testing.T) {
_ = a // a's ids are seeded for the symmetric read assertions above _ = a // a's ids are seeded for the symmetric read assertions above
} }
// TestTranscriptsTableIsSharedNotRLS is the ADR-021 isolation proof: transcripts
// is the ONE shared, non-RLS surface, and the public-content classification
// leaked to nothing else. It is the inverse of TestRLSEnforcesPerUserIsolation —
// where that asserts deny-all on every user-owned table, this asserts transcripts
// is readable and writable with no user scope at all, holds no user_id, and is
// the single table with row-level security switched off.
func TestTranscriptsTableIsSharedNotRLS(t *testing.T) {
newStore(t)
super := rawPool(t)
resetDB(t, super)
app := appPool(t, super)
ctx := context.Background()
// 1. Shared + non-RLS: with NO GUC set, the app role both writes and reads a
// transcript. On an RLS table this would be deny-all (zero rows), exactly as
// the main isolation test asserts for every user-owned table.
_, err := app.Exec(ctx,
`INSERT INTO transcripts (provider, provider_video_id, source, content)
VALUES ('youtube', 'shared-vid', 'captions', 'public words')`)
require.NoError(t, err, "app role must write shared transcript content with no user scope")
require.Equal(t, 1, scopedCount(t, app, "", "transcripts"),
"transcripts must be readable with NO user scope — it is shared, non-RLS (ADR-021)")
// 2. No user_id column: the table holds only public caption content + the
// video's public id, nothing user-identifying.
var hasUserID bool
require.NoError(t, super.QueryRow(ctx,
`SELECT EXISTS (SELECT 1 FROM information_schema.columns
WHERE table_name = 'transcripts' AND column_name = 'user_id')`).Scan(&hasUserID))
require.False(t, hasUserID, "transcripts must carry no user_id (ADR-021 public content)")
// 3. The boundary is EXACTLY here: every user-owned table still has row-level
// security enabled; transcripts alone has it off. This is the proof the
// non-RLS classification was applied to transcripts and leaked nowhere else.
for _, table := range allIsolatedTables {
require.True(t, rlsEnabled(t, super, table),
"%s must still enforce row-level security — isolation must not have regressed", table)
}
require.False(t, rlsEnabled(t, super, "transcripts"),
"transcripts must be the single table with row-level security OFF (the one shared surface)")
}
// rlsEnabled reports whether a public table has ROW LEVEL SECURITY enabled.
func rlsEnabled(t *testing.T, p *pgxpool.Pool, table string) bool {
t.Helper()
var enabled bool
require.NoError(t, p.QueryRow(context.Background(),
`SELECT relrowsecurity FROM pg_class
WHERE relname = $1 AND relnamespace = 'public'::regnamespace`, table).Scan(&enabled))
return enabled
}
+66
View File
@@ -0,0 +1,66 @@
package store
import (
"context"
"errors"
"fmt"
"github.com/jackc/pgx/v5"
"gitea.d-ma.be/mathias/tapir/internal/domain"
)
// GetTranscript returns the shared, stored transcript for a video keyed by the
// cross-user dedup key (provider, providerVideoID), and whether one exists
// (ADR-021). It reads via the raw pool, NOT withUser: the table holds public
// content with no user_id and no RLS policy, so it is shared across users by
// construction. A stored SourceNone is a real hit (ok == true, HasText() ==
// false) — a known caption-less video, so the caller skips without re-fetching.
func (s *Store) GetTranscript(ctx context.Context, provider, providerVideoID string) (domain.Transcript, bool, error) {
var source, lang, content string
err := s.pool.QueryRow(ctx,
`SELECT source, COALESCE(language, ''), COALESCE(content, '')
FROM transcripts WHERE provider = $1 AND provider_video_id = $2`,
provider, providerVideoID).Scan(&source, &lang, &content)
if errors.Is(err, pgx.ErrNoRows) {
return domain.Transcript{}, false, nil
}
if err != nil {
return domain.Transcript{}, false, fmt.Errorf("store: get transcript: %w", err)
}
return domain.Transcript{
Source: domain.TranscriptSource(source),
Language: lang,
Content: content,
}, true, nil
}
// SaveTranscript upserts the shared transcript for (provider, providerVideoID).
// Only terminal outcomes belong here: SourceCaptions (with text) or SourceNone
// (no captions). A transient SourceRateLimited is rejected so persistence never
// masks a 429 as a permanent absence — that stays a per-user retry (ADR-014).
// Last write wins on conflict (a later re-fetch may correct an entry). It writes
// via the raw pool, NOT withUser — public content, shared, non-RLS (ADR-021).
func (s *Store) SaveTranscript(ctx context.Context, provider, providerVideoID string, t domain.Transcript) error {
switch t.Source {
case domain.SourceCaptions, domain.SourceNone:
// terminal — persist
case domain.SourceRateLimited:
return fmt.Errorf("store: refusing to persist transient rate-limited transcript for %s/%s", provider, providerVideoID)
default:
return fmt.Errorf("store: invalid transcript source %q", t.Source)
}
_, err := s.pool.Exec(ctx,
`INSERT INTO transcripts (provider, provider_video_id, source, language, content)
VALUES ($1, $2, $3, NULLIF($4, ''), NULLIF($5, ''))
ON CONFLICT (provider, provider_video_id)
DO UPDATE SET source = EXCLUDED.source,
language = EXCLUDED.language,
content = EXCLUDED.content,
fetched_at = NOW()`,
provider, providerVideoID, string(t.Source), t.Language, t.Content)
if err != nil {
return fmt.Errorf("store: save transcript: %w", err)
}
return nil
}
@@ -0,0 +1,87 @@
package store_test
import (
"context"
"testing"
"github.com/stretchr/testify/require"
"gitea.d-ma.be/mathias/tapir/internal/adapters/store"
"gitea.d-ma.be/mathias/tapir/internal/domain"
"gitea.d-ma.be/mathias/tapir/internal/ports"
)
// Static check: Store satisfies the shared TranscriptStore port (ADR-021).
var _ ports.TranscriptStore = (*store.Store)(nil)
func TestSaveAndGetTranscript_RoundTrip(t *testing.T) {
s := newStore(t)
resetDB(t, rawPool(t))
ctx := context.Background()
want := domain.Transcript{Source: domain.SourceCaptions, Language: "en", Content: "the words"}
require.NoError(t, s.SaveTranscript(ctx, "youtube", "vid-1", want))
got, ok, err := s.GetTranscript(ctx, "youtube", "vid-1")
require.NoError(t, err)
require.True(t, ok, "a saved transcript must be found")
require.Equal(t, domain.SourceCaptions, got.Source)
require.Equal(t, "en", got.Language)
require.Equal(t, "the words", got.Content)
require.True(t, got.HasText())
}
func TestGetTranscript_Miss(t *testing.T) {
s := newStore(t)
resetDB(t, rawPool(t))
_, ok, err := s.GetTranscript(context.Background(), "youtube", "absent")
require.NoError(t, err, "a miss is not an error")
require.False(t, ok)
}
// A stored "no captions" outcome is a real hit: callers must skip without
// re-fetching, so ok is true even though there is no text (ADR-021 / ADR-007).
func TestSaveAndGetTranscript_NoneIsAStoredHit(t *testing.T) {
s := newStore(t)
resetDB(t, rawPool(t))
ctx := context.Background()
require.NoError(t, s.SaveTranscript(ctx, "youtube", "vid-none", domain.Transcript{Source: domain.SourceNone}))
got, ok, err := s.GetTranscript(ctx, "youtube", "vid-none")
require.NoError(t, err)
require.True(t, ok, "a stored SourceNone is a hit, not a miss")
require.Equal(t, domain.SourceNone, got.Source)
require.False(t, got.HasText())
}
// A transient 429 must never be persisted as a terminal transcript, or a later
// read would mask the rate-limit as a permanent "no transcript" (ADR-014).
func TestSaveTranscript_RejectsRateLimited(t *testing.T) {
s := newStore(t)
resetDB(t, rawPool(t))
err := s.SaveTranscript(context.Background(), "youtube", "vid-429",
domain.Transcript{Source: domain.SourceRateLimited})
require.Error(t, err)
_, ok, _ := s.GetTranscript(context.Background(), "youtube", "vid-429")
require.False(t, ok, "a rejected rate-limited save must leave nothing stored")
}
func TestSaveTranscript_UpsertLastWriteWins(t *testing.T) {
s := newStore(t)
resetDB(t, rawPool(t))
ctx := context.Background()
require.NoError(t, s.SaveTranscript(ctx, "youtube", "vid-up", domain.Transcript{Source: domain.SourceNone}))
require.NoError(t, s.SaveTranscript(ctx, "youtube", "vid-up",
domain.Transcript{Source: domain.SourceCaptions, Language: "en", Content: "now resolved"}))
got, ok, err := s.GetTranscript(ctx, "youtube", "vid-up")
require.NoError(t, err)
require.True(t, ok)
require.Equal(t, domain.SourceCaptions, got.Source)
require.Equal(t, "now resolved", got.Content)
}
+71 -4
View File
@@ -46,14 +46,15 @@ func (s *Store) UpsertVideo(ctx context.Context, v domain.Video) (string, error)
} }
if err := tx.QueryRow(ctx, if err := tx.QueryRow(ctx,
`INSERT INTO videos (user_id, provider, provider_video_id, title, url, published_at) `INSERT INTO videos (user_id, provider, provider_video_id, title, url, published_at, channel_title)
VALUES ($1, $2, $3, $4, $5, $6) VALUES ($1, $2, $3, $4, $5, $6, $7)
ON CONFLICT (user_id, provider, provider_video_id) DO UPDATE SET ON CONFLICT (user_id, provider, provider_video_id) DO UPDATE SET
title = EXCLUDED.title, title = EXCLUDED.title,
url = EXCLUDED.url, url = EXCLUDED.url,
published_at = EXCLUDED.published_at published_at = EXCLUDED.published_at,
channel_title = COALESCE(NULLIF(EXCLUDED.channel_title, ''), videos.channel_title)
RETURNING id`, RETURNING id`,
v.UserID, provider, v.ProviderVideoID, v.Title, v.URL, nullTime(v.PublishedAt), v.UserID, provider, v.ProviderVideoID, v.Title, v.URL, nullTime(v.PublishedAt), v.ChannelTitle,
).Scan(&id); err != nil { ).Scan(&id); err != nil {
return fmt.Errorf("store: upsert video: %w", err) return fmt.Errorf("store: upsert video: %w", err)
} }
@@ -72,3 +73,69 @@ func nullTime(t time.Time) *time.Time {
} }
return &t return &t
} }
// NewestUnsummarizedVideoIDs returns up to limit of the user's videos that have
// no summary yet, newest first (published_at DESC, NULLS LAST). It caps the
// connect-time onboarding burst (Feature 1) at a fixed count: the caller marks
// these for summarization through the shared rate gate. RLS-scoped via withUser,
// so it only ever sees the requesting user's rows. limit <= 0 returns nil.
func (s *Store) NewestUnsummarizedVideoIDs(ctx context.Context, userID string, limit int) ([]string, error) {
if limit <= 0 {
return nil, nil
}
var ids []string
if err := s.withUser(ctx, userID, func(tx pgx.Tx) error {
rows, err := tx.Query(ctx,
`SELECT v.id
FROM videos v
WHERE v.user_id = $1
AND NOT EXISTS (
SELECT 1 FROM summaries su
WHERE su.user_id = v.user_id AND su.video_id = v.id)
ORDER BY v.published_at DESC NULLS LAST, v.seen_at DESC
LIMIT $2`, userID, limit)
if err != nil {
return fmt.Errorf("store: newest unsummarized: %w", err)
}
defer rows.Close()
for rows.Next() {
var id string
if err := rows.Scan(&id); err != nil {
return fmt.Errorf("store: scan newest unsummarized: %w", err)
}
ids = append(ids, id)
}
return rows.Err()
}); err != nil {
return nil, err
}
return ids, nil
}
// DistinctChannels returns the user's distinct, non-empty source channel titles
// (the channels they have videos from), alphabetically — the option list for the
// feed's channel filter. RLS-scoped via withUser.
func (s *Store) DistinctChannels(ctx context.Context, userID string) ([]string, error) {
var out []string
if err := s.withUser(ctx, userID, func(tx pgx.Tx) error {
rows, err := tx.Query(ctx,
`SELECT DISTINCT channel_title FROM videos
WHERE user_id = $1 AND channel_title IS NOT NULL AND channel_title <> ''
ORDER BY channel_title`, userID)
if err != nil {
return fmt.Errorf("store: distinct channels: %w", err)
}
defer rows.Close()
for rows.Next() {
var c string
if err := rows.Scan(&c); err != nil {
return fmt.Errorf("store: scan channel: %w", err)
}
out = append(out, c)
}
return rows.Err()
}); err != nil {
return nil, err
}
return out, nil
}
+58
View File
@@ -81,3 +81,61 @@ func TestUpsertVideo_PerUserIsolation(t *testing.T) {
require.NotEqual(t, idA, idB, "same provider video for two users must be two distinct rows") require.NotEqual(t, idA, idB, "same provider video for two users must be two distinct rows")
} }
func TestNewestUnsummarizedVideoIDs(t *testing.T) {
ctx := context.Background()
s := newStore(t)
resetDB(t, rawPool(t))
mk := func(user, pid string, day int) string {
v := ytVideo(user, pid, pid)
v.PublishedAt = time.Date(2026, 6, day, 12, 0, 0, 0, time.UTC)
id, err := s.UpsertVideo(ctx, v)
require.NoError(t, err)
return id
}
_ = mk(userA, "a1vid000001", 1)
id2 := mk(userA, "a2vid000002", 2)
id3 := mk(userA, "a3vid000003", 3)
id4 := mk(userA, "a4vid000004", 4)
mk(userB, "b1vid000009", 9) // userB's newest — must never leak via RLS
// The newest (v4) is summarized, so it's excluded from "unsummarized".
require.NoError(t, s.Deliver(ctx, summary(userA, id4, "done")))
// Cap 2, newest-first unsummarized: v3 then v2 (v4 excluded; userB excluded).
got, err := s.NewestUnsummarizedVideoIDs(ctx, userA, 2)
require.NoError(t, err)
require.Equal(t, []string{id3, id2}, got)
none, err := s.NewestUnsummarizedVideoIDs(ctx, userA, 0)
require.NoError(t, err)
require.Empty(t, none, "limit 0 returns nothing")
}
func TestUpsertVideoPersistsChannelAndDistinctChannels(t *testing.T) {
ctx := context.Background()
s := newStore(t)
resetDB(t, rawPool(t))
mk := func(pid, channel string) {
v := ytVideo(userA, pid, pid)
v.ChannelTitle = channel
_, err := s.UpsertVideo(ctx, v)
require.NoError(t, err)
}
mk("aa11111aaaa", "Acme Talks")
mk("bb22222bbbb", "Acme Talks") // same channel
mk("cc33333cccc", "Zeta Channel")
// userB's channel must not leak.
vb := ytVideo(userB, "dd44444dddd", "x")
vb.ChannelTitle = "Bravo Only"
_, err := s.UpsertVideo(ctx, vb)
require.NoError(t, err)
got, err := s.DistinctChannels(ctx, userA)
require.NoError(t, err)
require.Equal(t, []string{"Acme Talks", "Zeta Channel"}, got,
"distinct, alphabetical, user-scoped (no Bravo Only)")
}
@@ -0,0 +1,62 @@
package youtube
import (
"context"
"errors"
"net/http"
"testing"
"gitea.d-ma.be/mathias/tapir/internal/domain"
)
func TestVideoByID(t *testing.T) {
const id = "dQw4w9WgXcQ"
a, secrets := newTestAdapter(t, func(w http.ResponseWriter, r *http.Request) {
if r.URL.Path != "/videos" {
t.Errorf("unexpected path %q (must use videos.list)", r.URL.Path)
}
if got := r.URL.Query().Get("id"); got != id {
t.Errorf("expected id=%s, got %q", id, got)
}
if got := r.URL.Query().Get("part"); got != "snippet" {
t.Errorf("expected part=snippet, got %q", got)
}
_, _ = w.Write([]byte(`{"items":[{"snippet":{"title":"Never Gonna Give You Up","channelTitle":"Rick Astley","publishedAt":"2026-05-20T09:00:00Z"}}]}`))
})
v, err := a.VideoByID(context.Background(), "u1", id)
if err != nil {
t.Fatalf("VideoByID: %v", err)
}
if v.UserID != "u1" {
t.Errorf("UserID = %q, want u1", v.UserID)
}
if v.ProviderVideoID != id || v.Title != "Never Gonna Give You Up" {
t.Errorf("unexpected video: %+v", v)
}
if v.ChannelTitle != "Rick Astley" {
t.Errorf("ChannelTitle = %q, want Rick Astley", v.ChannelTitle)
}
if v.Provider != domain.ProviderYouTube || v.URL != "https://www.youtube.com/watch?v="+id {
t.Errorf("video not wired correctly: %+v", v)
}
if v.PublishedAt.IsZero() {
t.Errorf("expected publishedAt parsed, got zero")
}
if v.SubscriptionID != "" {
t.Errorf("a pasted video must have no subscription, got %q", v.SubscriptionID)
}
if secrets.byRef == nil {
t.Errorf("token must be resolved by reference through the SecretStore")
}
}
func TestVideoByIDNotFound(t *testing.T) {
a, _ := newTestAdapter(t, func(w http.ResponseWriter, _ *http.Request) {
_, _ = w.Write([]byte(`{"items":[]}`))
})
_, err := a.VideoByID(context.Background(), "u1", "missingvid0")
if !errors.Is(err, domain.ErrVideoNotFound) {
t.Fatalf("VideoByID for missing id = %v, want domain.ErrVideoNotFound", err)
}
}
+43
View File
@@ -234,6 +234,7 @@ func (a *Adapter) NewVideos(ctx context.Context, sub domain.Subscription) ([]dom
Provider: domain.ProviderYouTube, Provider: domain.ProviderYouTube,
ProviderVideoID: vid, ProviderVideoID: vid,
Title: item.Snippet.Title, Title: item.Snippet.Title,
ChannelTitle: sub.ChannelTitle,
URL: "https://www.youtube.com/watch?v=" + vid, URL: "https://www.youtube.com/watch?v=" + vid,
PublishedAt: item.Snippet.PublishedAt, PublishedAt: item.Snippet.PublishedAt,
}) })
@@ -244,6 +245,38 @@ func (a *Adapter) NewVideos(ctx context.Context, sub domain.Subscription) ([]dom
return videos, nil return videos, nil
} }
// VideoByID fetches a single video's metadata (videos.list, snippet) for an
// arbitrary video id — including channels the user does not follow (paste-a-URL,
// Feature 2). This is a Data API call (1 quota unit), NOT the rate-limited
// caption path, so it is not gated: only the later transcript fetch goes through
// globalFetchGate. UserID is set on the result and SubscriptionID is left empty
// (a pasted video has no subscription parent). Returns ErrVideoNotFound when the
// id resolves to no video.
func (a *Adapter) VideoByID(ctx context.Context, userID, videoID string) (domain.Video, error) {
client, err := a.httpClient(ctx, a.cfg.TokenSecretRef)
if err != nil {
return domain.Video{}, err
}
q := url.Values{"part": {"snippet"}, "id": {videoID}}
var resp videoListResponse
if err := a.getJSON(ctx, client, "/videos", q, &resp); err != nil {
return domain.Video{}, fmt.Errorf("video by id %q: %w", videoID, err)
}
if len(resp.Items) == 0 {
return domain.Video{}, fmt.Errorf("video %q: %w", videoID, domain.ErrVideoNotFound)
}
it := resp.Items[0]
return domain.Video{
UserID: userID,
Provider: domain.ProviderYouTube,
ProviderVideoID: videoID,
Title: it.Snippet.Title,
ChannelTitle: it.Snippet.ChannelTitle,
URL: "https://www.youtube.com/watch?v=" + videoID,
PublishedAt: it.Snippet.PublishedAt,
}, nil
}
// uploadsPlaylistID derives a channel's uploads playlist id at zero API cost: // uploadsPlaylistID derives a channel's uploads playlist id at zero API cost:
// a standard channel id "UCxxxx" maps to uploads playlist "UUxxxx". Returns // a standard channel id "UCxxxx" maps to uploads playlist "UUxxxx". Returns
// ok=false for ids that don't follow this convention (caller falls back to // ok=false for ids that don't follow this convention (caller falls back to
@@ -342,6 +375,16 @@ type playlistItemListResponse struct {
} `json:"items"` } `json:"items"`
} }
type videoListResponse struct {
Items []struct {
Snippet struct {
Title string `json:"title"`
ChannelTitle string `json:"channelTitle"`
PublishedAt time.Time `json:"publishedAt"`
} `json:"snippet"`
} `json:"items"`
}
type channelListResponse struct { type channelListResponse struct {
Items []struct { Items []struct {
ContentDetails struct { ContentDetails struct {
+4 -1
View File
@@ -131,7 +131,7 @@ func TestNewVideos(t *testing.T) {
}`)) }`))
}) })
sub := domain.Subscription{ID: "s1", UserID: "u1", ChannelID: "UC_acme"} sub := domain.Subscription{ID: "s1", UserID: "u1", ChannelID: "UC_acme", ChannelTitle: "Acme Channel"}
vids, err := a.NewVideos(context.Background(), sub) vids, err := a.NewVideos(context.Background(), sub)
if err != nil { if err != nil {
t.Fatalf("NewVideos: %v", err) t.Fatalf("NewVideos: %v", err)
@@ -143,6 +143,9 @@ func TestNewVideos(t *testing.T) {
if v.ProviderVideoID != "vid1" || v.Title != "Designing for Attention" { if v.ProviderVideoID != "vid1" || v.Title != "Designing for Attention" {
t.Errorf("unexpected video: %+v", v) t.Errorf("unexpected video: %+v", v)
} }
if v.ChannelTitle != "Acme Channel" {
t.Errorf("ChannelTitle = %q, want Acme Channel", v.ChannelTitle)
}
if v.Provider != domain.ProviderYouTube || v.URL != "https://www.youtube.com/watch?v=vid1" { if v.Provider != domain.ProviderYouTube || v.URL != "https://www.youtube.com/watch?v=vid1" {
t.Errorf("video not wired correctly: %+v", v) t.Errorf("video not wired correctly: %+v", v)
} }
+34
View File
@@ -13,6 +13,7 @@ import (
"os" "os"
"path/filepath" "path/filepath"
"sort" "sort"
"strconv"
"strings" "strings"
"time" "time"
) )
@@ -79,6 +80,13 @@ type Config struct {
// pre-recency behaviour). Default ~7 days. // pre-recency behaviour). Default ~7 days.
AutoSummarizeWindow time.Duration AutoSummarizeWindow time.Duration
// OnboardSummarizeCount caps how many of a freshly-connected user's newest
// videos are summarized immediately on connect (the onboarding "it works"
// burst). HARD-capped at maxOnboardSummarizeCount so onboarding can never
// bulk-fetch; 0 disables the burst. Every fetch still flows through the shared
// caption rate gate (ADR-014) — the cap bounds count, never the pacing. Default 3.
OnboardSummarizeCount int
// DiscoveryInterval, when > 0, makes `serve` run in-process scheduled discovery // DiscoveryInterval, when > 0, makes `serve` run in-process scheduled discovery
// for ALL users on that cadence (ADR-018). Zero/unset = disabled, so dev and // for ALL users on that cadence (ADR-018). Zero/unset = disabled, so dev and
// tests never auto-fetch. Single-replica assumption — see cmdServe. // tests never auto-fetch. Single-replica assumption — see cmdServe.
@@ -119,6 +127,8 @@ const (
defaultFetchRate = 2 * time.Second defaultFetchRate = 2 * time.Second
defaultPublicURL = "https://tapir.d-ma.be" defaultPublicURL = "https://tapir.d-ma.be"
defaultAutoSummarizeWindow = 7 * 24 * time.Hour defaultAutoSummarizeWindow = 7 * 24 * time.Hour
defaultOnboardSummarizeCount = 3
maxOnboardSummarizeCount = 5
) )
// Load reads the environment into a Config, applying defaults. It does not // Load reads the environment into a Config, applying defaults. It does not
@@ -183,6 +193,18 @@ func Load() (Config, error) {
} }
c.AutoSummarizeWindow = autoWindow c.AutoSummarizeWindow = autoWindow
onboard, err := intOr("TAPIR_ONBOARD_SUMMARIZE_COUNT", defaultOnboardSummarizeCount)
if err != nil {
return Config{}, err
}
if onboard < 0 {
onboard = 0
}
if onboard > maxOnboardSummarizeCount {
onboard = maxOnboardSummarizeCount
}
c.OnboardSummarizeCount = onboard
return c, nil return c, nil
} }
@@ -247,6 +269,18 @@ func envOr(key, fallback string) string {
return fallback return fallback
} }
func intOr(key string, fallback int) (int, error) {
v := os.Getenv(key)
if v == "" {
return fallback, nil
}
n, err := strconv.Atoi(v)
if err != nil {
return 0, fmt.Errorf("config: %s=%q: %w", key, v, err)
}
return n, nil
}
func durationOr(key string, fallback time.Duration) (time.Duration, error) { func durationOr(key string, fallback time.Duration) (time.Duration, error) {
v := os.Getenv(key) v := os.Getenv(key)
if v == "" { if v == "" {
+32
View File
@@ -53,6 +53,38 @@ func TestLoad_AppliesDefaults(t *testing.T) {
} }
} }
func TestLoad_OnboardSummarizeCount(t *testing.T) {
cases := []struct {
name, env string
want int
}{
{"default", "", defaultOnboardSummarizeCount},
{"explicit", "4", 4},
{"zero disables", "0", 0},
{"clamped to hard cap", "50", maxOnboardSummarizeCount},
{"negative clamps to zero", "-3", 0},
}
for _, c := range cases {
t.Run(c.name, func(t *testing.T) {
setEnv(t, map[string]string{"TAPIR_ONBOARD_SUMMARIZE_COUNT": c.env})
cfg, err := Load()
if err != nil {
t.Fatalf("Load: %v", err)
}
if cfg.OnboardSummarizeCount != c.want {
t.Fatalf("OnboardSummarizeCount = %d, want %d", cfg.OnboardSummarizeCount, c.want)
}
})
}
}
func TestLoad_OnboardSummarizeCountInvalid(t *testing.T) {
setEnv(t, map[string]string{"TAPIR_ONBOARD_SUMMARIZE_COUNT": "three"})
if _, err := Load(); err == nil {
t.Fatal("Load: want error for non-numeric TAPIR_ONBOARD_SUMMARIZE_COUNT")
}
}
func TestLoad_ParsesValues(t *testing.T) { func TestLoad_ParsesValues(t *testing.T) {
setEnv(t, map[string]string{ setEnv(t, map[string]string{
"TAPIR_USER_ID": "11111111-1111-1111-1111-111111111111", "TAPIR_USER_ID": "11111111-1111-1111-1111-111111111111",
+7
View File
@@ -3,10 +3,16 @@
package domain package domain
import ( import (
"errors"
"fmt" "fmt"
"time" "time"
) )
// ErrVideoNotFound is returned when a video id resolves to no video (deleted,
// private, or a typo'd paste). Defined in domain so adapters and the web layer
// share one sentinel without coupling to each other.
var ErrVideoNotFound = errors.New("video not found")
// ErrChannelUnavailable is returned by a VideoSource when a channel's upload // ErrChannelUnavailable is returned by a VideoSource when a channel's upload
// playlist returns HTTP 404 — the channel was deleted or made private. The runner // playlist returns HTTP 404 — the channel was deleted or made private. The runner
// stores these so the account page can surface them to the user. // stores these so the account page can surface them to the user.
@@ -67,6 +73,7 @@ type Video struct {
Provider Provider Provider Provider
ProviderVideoID string ProviderVideoID string
Title string Title string
ChannelTitle string
URL string URL string
PublishedAt time.Time PublishedAt time.Time
SeenAt time.Time SeenAt time.Time
+20
View File
@@ -27,6 +27,26 @@ type Summarizer interface {
Summarize(ctx context.Context, v domain.Video, t domain.Transcript) (domain.Summary, error) Summarize(ctx context.Context, v domain.Video, t domain.Transcript) (domain.Summary, error)
} }
// TranscriptStore persists transcripts as shared, video-keyed public content
// (ADR-021). It is keyed by the cross-user dedup key (provider, providerVideoID)
// — the video's public identity, NOT Tapir's per-user videos.id — and holds only
// public caption content, so it is deliberately NOT user-scoped: two users who
// share a video share the one row. The engine reads it before any caption fetch
// so re-analysis never re-touches YouTube (ADR-010/014).
type TranscriptStore interface {
// GetTranscript returns the stored transcript for a video and whether one
// exists. A stored Source == SourceNone (captions permanently absent) is a
// real hit: ok is true and HasText() is false, so callers skip without
// re-fetching. A transient rate-limit is never stored, so it never appears
// here as a false absence.
GetTranscript(ctx context.Context, provider, providerVideoID string) (t domain.Transcript, ok bool, err error)
// SaveTranscript upserts the transcript for (provider, providerVideoID). Only
// terminal outcomes are persisted: SourceCaptions (with text) or SourceNone.
// SourceRateLimited must NOT be passed — it is a per-user retry (ADR-014), not
// a shared terminal state.
SaveTranscript(ctx context.Context, provider, providerVideoID string, t domain.Transcript) error
}
// Sink delivers a summary to a destination (user store, brain, ...). // Sink delivers a summary to a destination (user store, brain, ...).
// Implementations fail independently of one another. // Implementations fail independently of one another.
type Sink interface { type Sink interface {
+42 -2
View File
@@ -27,6 +27,13 @@ type Engine struct {
AI ports.Summarizer AI ports.Summarizer
Sinks []ports.Sink Sinks []ports.Sink
// Transcripts, when set, is the shared transcript cache (ADR-021): the engine
// reads it before any caption fetch and writes resolved transcripts back, so
// re-analysis — the same user re-summarizing, or a second user with the same
// video — never re-touches YouTube (ADR-010/014). Optional: nil disables
// persistence (fetch every time), keeping the pure-core/scaffold wiring valid.
Transcripts ports.TranscriptStore
// processed dedups videos within this engine's lifetime so a video is not // processed dedups videos within this engine's lifetime so a video is not
// summarized twice when the watcher sees it again. Durable cross-restart // summarized twice when the watcher sees it again. Durable cross-restart
// dedup is the store's concern (a resolved TRANSCRIPT / existing SUMMARY, // dedup is the store's concern (a resolved TRANSCRIPT / existing SUMMARY,
@@ -57,9 +64,9 @@ type ProcessResult struct {
// resolve transcript -> (summarize -> deliver) | skip. // resolve transcript -> (summarize -> deliver) | skip.
// See docs/use-cases/summarize_new_video.feature. // See docs/use-cases/summarize_new_video.feature.
func (e *Engine) ProcessNewVideo(ctx context.Context, v domain.Video) (ProcessResult, error) { func (e *Engine) ProcessNewVideo(ctx context.Context, v domain.Video) (ProcessResult, error) {
t, err := e.Source.FetchTranscript(ctx, v) t, err := e.resolveTranscript(ctx, v)
if err != nil { if err != nil {
return ProcessResult{Video: v}, fmt.Errorf("fetch transcript: %w", err) return ProcessResult{Video: v}, err
} }
if !t.HasText() { if !t.HasText() {
// No usable transcript: record the skip, produce no summary, deliver nothing // No usable transcript: record the skip, produce no summary, deliver nothing
@@ -86,6 +93,39 @@ func (e *Engine) ProcessNewVideo(ctx context.Context, v domain.Video) (ProcessRe
return ProcessResult{Video: v, Summary: &sum, TranscriptSource: string(t.Source)}, errors.Join(errs...) return ProcessResult{Video: v, Summary: &sum, TranscriptSource: string(t.Source)}, errors.Join(errs...)
} }
// resolveTranscript returns v's transcript, reading the shared store first
// (ADR-021): a stored transcript — including a stored SourceNone (captions
// permanently absent) — is returned without touching YouTube, so re-analysis
// never re-fetches. On a store miss it fetches through the source (which gates
// the caption call, ADR-014) and persists the terminal outcome so the next
// analysis, for any user, reads from the store. A transient SourceRateLimited is
// returned to the caller (the runner stamps a per-user backoff) but never stored,
// so persistence can never mask a 429 as a permanent "no transcript". When no
// TranscriptStore is wired the engine simply fetches every time.
func (e *Engine) resolveTranscript(ctx context.Context, v domain.Video) (domain.Transcript, error) {
if e.Transcripts != nil {
stored, ok, err := e.Transcripts.GetTranscript(ctx, string(v.Provider), v.ProviderVideoID)
if err != nil {
return domain.Transcript{}, fmt.Errorf("get stored transcript: %w", err)
}
if ok {
return stored, nil
}
}
t, err := e.Source.FetchTranscript(ctx, v)
if err != nil {
return domain.Transcript{}, fmt.Errorf("fetch transcript: %w", err)
}
if e.Transcripts != nil && t.Source != domain.SourceRateLimited {
if err := e.Transcripts.SaveTranscript(ctx, string(v.Provider), v.ProviderVideoID, t); err != nil {
return domain.Transcript{}, fmt.Errorf("save transcript: %w", err)
}
}
return t, nil
}
// ProcessNewVideos walks a user's subscriptions and processes each newly seen // ProcessNewVideos walks a user's subscriptions and processes each newly seen
// video. Only videos surfaced via the user's subscriptions are considered, so a // video. Only videos surfaced via the user's subscriptions are considered, so a
// channel the user is not subscribed to is never processed. A video already // channel the user is not subscribed to is never processed. A video already
+187
View File
@@ -0,0 +1,187 @@
package usecase
import (
"context"
"testing"
"gitea.d-ma.be/mathias/tapir/internal/domain"
)
// These tests pin the ADR-021 read-stored-first behaviour at the engine core:
// a stored transcript is summarized without re-touching the source, a miss
// fetches once and persists, and a transient rate-limit is never cached.
type recordingSource struct {
transcript domain.Transcript
fetchCalls int
}
func (s *recordingSource) ListSubscriptions(context.Context, string) ([]domain.Subscription, error) {
return nil, nil
}
func (s *recordingSource) NewVideos(context.Context, domain.Subscription) ([]domain.Video, error) {
return nil, nil
}
func (s *recordingSource) FetchTranscript(context.Context, domain.Video) (domain.Transcript, error) {
s.fetchCalls++
return s.transcript, nil
}
type fakeTranscriptStore struct {
stored map[string]domain.Transcript
saves int
}
func newFakeTranscriptStore() *fakeTranscriptStore {
return &fakeTranscriptStore{stored: make(map[string]domain.Transcript)}
}
func (f *fakeTranscriptStore) key(provider, id string) string { return provider + "|" + id }
func (f *fakeTranscriptStore) GetTranscript(_ context.Context, provider, id string) (domain.Transcript, bool, error) {
t, ok := f.stored[f.key(provider, id)]
return t, ok, nil
}
func (f *fakeTranscriptStore) SaveTranscript(_ context.Context, provider, id string, t domain.Transcript) error {
f.saves++
f.stored[f.key(provider, id)] = t
return nil
}
type countingSummarizer struct{ calls int }
func (c *countingSummarizer) Summarize(_ context.Context, v domain.Video, _ domain.Transcript) (domain.Summary, error) {
c.calls++
return domain.Summary{VideoID: v.ID, UserID: v.UserID, Summary: "s", AIProvider: "local"}, nil
}
type nopSink struct{}
func (nopSink) Name() string { return "nop" }
func (nopSink) Deliver(context.Context, domain.Summary) error { return nil }
func testVideo() domain.Video {
return domain.Video{ID: "v1", UserID: "u1", Provider: domain.ProviderYouTube, ProviderVideoID: "yt1"}
}
func TestProcessNewVideo_StoredTranscriptSkipsFetch(t *testing.T) {
src := &recordingSource{}
ts := newFakeTranscriptStore()
ts.stored[ts.key("youtube", "yt1")] = domain.Transcript{Source: domain.SourceCaptions, Content: "stored words"}
sum := &countingSummarizer{}
eng := NewEngine(src, sum, nopSink{})
eng.Transcripts = ts
res, err := eng.ProcessNewVideo(context.Background(), testVideo())
if err != nil {
t.Fatalf("ProcessNewVideo: %v", err)
}
if src.fetchCalls != 0 {
t.Fatalf("stored transcript must not re-fetch from source; got %d fetches", src.fetchCalls)
}
if ts.saves != 0 {
t.Fatalf("a store hit must not re-save; got %d saves", ts.saves)
}
if sum.calls != 1 || res.Summary == nil {
t.Fatalf("expected a summary from the stored transcript; calls=%d summary=%v", sum.calls, res.Summary)
}
}
func TestProcessNewVideo_StoreMissFetchesAndPersists(t *testing.T) {
src := &recordingSource{transcript: domain.Transcript{Source: domain.SourceCaptions, Language: "en", Content: "fetched words"}}
ts := newFakeTranscriptStore()
sum := &countingSummarizer{}
eng := NewEngine(src, sum, nopSink{})
eng.Transcripts = ts
if _, err := eng.ProcessNewVideo(context.Background(), testVideo()); err != nil {
t.Fatalf("ProcessNewVideo: %v", err)
}
if src.fetchCalls != 1 {
t.Fatalf("a store miss must fetch exactly once; got %d", src.fetchCalls)
}
if ts.saves != 1 {
t.Fatalf("a fetched transcript must be persisted; got %d saves", ts.saves)
}
got, ok, _ := ts.GetTranscript(context.Background(), "youtube", "yt1")
if !ok || got.Content != "fetched words" {
t.Fatalf("persisted transcript not readable back: ok=%v content=%q", ok, got.Content)
}
}
// The second summarize of the same video reads the persisted transcript and does
// NOT re-fetch — the primary ADR-021 win, proven end to end at the engine.
func TestProcessNewVideo_SecondSummarizeDoesNotRefetch(t *testing.T) {
src := &recordingSource{transcript: domain.Transcript{Source: domain.SourceCaptions, Content: "words"}}
ts := newFakeTranscriptStore()
eng := NewEngine(src, &countingSummarizer{}, nopSink{})
eng.Transcripts = ts
for i := 0; i < 2; i++ {
if _, err := eng.ProcessNewVideo(context.Background(), testVideo()); err != nil {
t.Fatalf("pass %d: %v", i, err)
}
}
if src.fetchCalls != 1 {
t.Fatalf("the second summarize must reuse the stored transcript; got %d fetches", src.fetchCalls)
}
}
// A stored "no captions" outcome short-circuits before both fetch and summarize.
func TestProcessNewVideo_StoredNoneSkipsFetchAndSummarize(t *testing.T) {
src := &recordingSource{}
ts := newFakeTranscriptStore()
ts.stored[ts.key("youtube", "yt1")] = domain.Transcript{Source: domain.SourceNone}
sum := &countingSummarizer{}
eng := NewEngine(src, sum, nopSink{})
eng.Transcripts = ts
res, err := eng.ProcessNewVideo(context.Background(), testVideo())
if err != nil {
t.Fatalf("ProcessNewVideo: %v", err)
}
if !res.Skipped {
t.Fatal("a stored SourceNone must skip")
}
if src.fetchCalls != 0 || sum.calls != 0 {
t.Fatalf("stored none must neither fetch nor summarize; fetches=%d calls=%d", src.fetchCalls, sum.calls)
}
}
// A transient 429 is surfaced (so the runner backs off per-user) but never cached
// as a shared terminal state — otherwise it would mask a rate-limit as permanent.
func TestProcessNewVideo_RateLimitedIsNotPersisted(t *testing.T) {
src := &recordingSource{transcript: domain.Transcript{Source: domain.SourceRateLimited}}
ts := newFakeTranscriptStore()
eng := NewEngine(src, &countingSummarizer{}, nopSink{})
eng.Transcripts = ts
res, err := eng.ProcessNewVideo(context.Background(), testVideo())
if err != nil {
t.Fatalf("ProcessNewVideo: %v", err)
}
if !res.Skipped || res.TranscriptSource != string(domain.SourceRateLimited) {
t.Fatalf("expected a rate-limited skip; skipped=%v source=%q", res.Skipped, res.TranscriptSource)
}
if ts.saves != 0 {
t.Fatalf("a transient rate-limit must not be persisted; got %d saves", ts.saves)
}
}
// With no TranscriptStore wired the engine fetches every time (back-compat).
func TestProcessNewVideo_NilStoreFetchesEveryTime(t *testing.T) {
src := &recordingSource{transcript: domain.Transcript{Source: domain.SourceCaptions, Content: "words"}}
eng := NewEngine(src, &countingSummarizer{}, nopSink{})
for i := 0; i < 2; i++ {
if _, err := eng.ProcessNewVideo(context.Background(), testVideo()); err != nil {
t.Fatalf("pass %d: %v", i, err)
}
}
if src.fetchCalls != 2 {
t.Fatalf("nil store must fetch every time; got %d", src.fetchCalls)
}
}
+100 -10
View File
@@ -3,6 +3,8 @@ package web
import ( import (
"context" "context"
"errors" "errors"
"html/template"
"io"
"log/slog" "log/slog"
"net/http" "net/http"
"time" "time"
@@ -10,6 +12,7 @@ import (
"github.com/a-h/templ" "github.com/a-h/templ"
"gitea.d-ma.be/mathias/tapir/internal/adapters/store" "gitea.d-ma.be/mathias/tapir/internal/adapters/store"
"gitea.d-ma.be/mathias/tapir/internal/domain"
) )
// Store is the read/write surface the web handlers depend on — a narrow port over // Store is the read/write surface the web handlers depend on — a narrow port over
@@ -18,6 +21,9 @@ import (
// fake without a database. // fake without a database.
type Store interface { type Store interface {
ListVideos(ctx context.Context, userID string, limit int) ([]store.SummaryRow, error) ListVideos(ctx context.Context, userID string, limit int) ([]store.SummaryRow, error)
// DistinctChannels lists the user's source channels — the options for the
// feed's channel multi-select filter.
DistinctChannels(ctx context.Context, userID string) ([]string, error)
GetSummaryByVideo(ctx context.Context, userID, videoID string) (*store.SummaryRow, error) GetSummaryByVideo(ctx context.Context, userID, videoID string) (*store.SummaryRow, error)
GetVideoRow(ctx context.Context, userID, videoID string) (*store.SummaryRow, error) GetVideoRow(ctx context.Context, userID, videoID string) (*store.SummaryRow, error)
ActionsFor(ctx context.Context, userID string, videoIDs []string) (map[string][]string, error) ActionsFor(ctx context.Context, userID string, videoIDs []string) (map[string][]string, error)
@@ -30,6 +36,10 @@ type Store interface {
SetAutoSummarize(ctx context.Context, userID string, enabled bool) error SetAutoSummarize(ctx context.Context, userID string, enabled bool) error
RequestSummarize(ctx context.Context, userID, videoID string) error RequestSummarize(ctx context.Context, userID, videoID string) error
// UpsertVideo persists a pasted video (idempotent on user+provider+video id,
// so it also dedups) and returns its durable store id.
UpsertVideo(ctx context.Context, v domain.Video) (string, error)
// Account management (the /account page, disconnect, delete-account). // Account management (the /account page, disconnect, delete-account).
ConnectionsForUser(ctx context.Context, userID string) ([]store.Connection, error) ConnectionsForUser(ctx context.Context, userID string) ([]store.Connection, error)
DeleteConnection(ctx context.Context, userID, provider string) error DeleteConnection(ctx context.Context, userID, provider string) error
@@ -78,6 +88,9 @@ type App struct {
// background goroutine (the "Summarize" button kicks it off). Nil = queue-only: // background goroutine (the "Summarize" button kicks it off). Nil = queue-only:
// the button flips the DB flag and the next `tapir run` does the work. // the button flips the DB flag and the next `tapir run` does the work.
Processor Processor Processor Processor
// Fetcher, when non-nil, resolves an arbitrary YouTube video id to metadata for
// the paste-a-URL flow (Feature 2). Nil = the /paste route is not mounted.
Fetcher VideoFetcher
// Processing tracks in-flight immediate summarizations so the status endpoint // Processing tracks in-flight immediate summarizations so the status endpoint
// shows the animation until the summary lands. The zero value is ready to use. // shows the animation until the summary lands. The zero value is ready to use.
Processing ProcessingSet Processing ProcessingSet
@@ -132,6 +145,9 @@ func (a *App) Router() http.Handler {
app.HandleFunc("POST /v/{videoId}/action", a.handleAction) app.HandleFunc("POST /v/{videoId}/action", a.handleAction)
app.HandleFunc("POST /v/{videoId}/summarize", a.handleRequestSummarize) app.HandleFunc("POST /v/{videoId}/summarize", a.handleRequestSummarize)
app.HandleFunc("POST /v/{videoId}/retry-now", a.handleRetryNow) app.HandleFunc("POST /v/{videoId}/retry-now", a.handleRetryNow)
if a.Fetcher != nil {
app.HandleFunc("POST /paste", a.handlePaste)
}
app.HandleFunc("GET /v/{videoId}/status", a.handleStatus) app.HandleFunc("GET /v/{videoId}/status", a.handleStatus)
app.HandleFunc("GET /register", a.handleRegisterForm) app.HandleFunc("GET /register", a.handleRegisterForm)
app.HandleFunc("POST /register", a.handleRegister) app.HandleFunc("POST /register", a.handleRegister)
@@ -183,7 +199,7 @@ func (a *App) handleList(w http.ResponseWriter, r *http.Request) {
} }
q := r.URL.Query() q := r.URL.Query()
f := Filter{ f := Filter{
Channel: q.Get("channel"), Channels: nonEmptyStrings(q["channel"]),
From: q.Get("from"), From: q.Get("from"),
To: q.Get("to"), To: q.Get("to"),
OnlySummarized: q.Get("summarized") == "1", OnlySummarized: q.Get("summarized") == "1",
@@ -198,25 +214,30 @@ func (a *App) handleList(w http.ResponseWriter, r *http.Request) {
rows := f.apply(allRows) rows := f.apply(allRows)
buckets := bucketRows(rows, a.recencyCutoff()) buckets := bucketRows(rows, a.recencyCutoff())
// hasConnected drives the empty state: a fresh account with a connection but // Channel options for the multi-select filter (the user's source channels).
// no discovery pass yet has zero rows, and we want it to read "connected, channels, err := a.Store.DistinctChannels(r.Context(), userID)
// summaries land gradually" rather than "nothing here". Only needed when the if err != nil {
// list is empty. a.serverError(w, r, "distinct channels", err)
hasConnected := false return
if buckets.empty() { }
// hasConnected drives both the paste box (shown to ANY connected user, #2) and
// the empty-state copy (a fresh account with a connection but no discovery pass
// yet reads "connected, summaries land gradually" rather than "nothing here").
// Computed every render — not only when empty — so a user with videos still
// gets the paste box.
conns, err := a.Store.ConnectionsForUser(r.Context(), userID) conns, err := a.Store.ConnectionsForUser(r.Context(), userID)
if err != nil { if err != nil {
a.serverError(w, r, "connections for user", err) a.serverError(w, r, "connections for user", err)
return return
} }
hasConnected = len(conns) > 0 hasConnected := len(conns) > 0
}
if isHTMX(r) { if isHTMX(r) {
a.render(w, r, summaryList(buckets, hasConnected)) a.render(w, r, summaryList(buckets, hasConnected))
return return
} }
a.render(w, r, ListPage(buckets, f, stats, takeFlash(w, r), hasConnected)) a.render(w, r, ListPage(buckets, f, stats, takeFlash(w, r), hasConnected, channels))
} }
// handleDetail renders one summary in full (highlights, takeaways, action group). // handleDetail renders one summary in full (highlights, takeaways, action group).
@@ -324,6 +345,75 @@ func (a *App) handleRequestSummarize(w http.ResponseWriter, r *http.Request) {
a.render(w, r, VideoCard(*row)) a.render(w, r, VideoCard(*row))
} }
// handlePaste handles "paste a YouTube URL" (Feature 2). It parses the video id,
// fetches metadata (Data API — ungated), upserts a subscription-less video row
// scoped to the user (idempotent, so it also dedups), and — if the video isn't
// already summarized — requests a summary and kicks off immediate processing
// through the SAME rate gate as the Summarize button. An explicit paste is a
// manual request, so it summarizes regardless of the recency window. A video that
// turns out to have no captions resolves to the honest "no transcript" terminal
// state via the engine (ADR-010), not an error here.
func (a *App) handlePaste(w http.ResponseWriter, r *http.Request) {
userID, ok := a.currentUserID(w, r)
if !ok {
return
}
videoID, err := parseYouTubeVideoID(r.FormValue("url"))
if err != nil {
a.pasteFailure(w, http.StatusBadRequest, "That doesn't look like a YouTube video link.")
return
}
v, err := a.Fetcher.FetchVideo(r.Context(), userID, videoID)
if errors.Is(err, domain.ErrVideoNotFound) {
a.pasteFailure(w, http.StatusNotFound, "That video couldn't be found — it may be private or removed.")
return
}
if err != nil {
a.serverError(w, r, "paste fetch", err)
return
}
id, err := a.Store.UpsertVideo(r.Context(), v)
if err != nil {
a.serverError(w, r, "paste upsert", err)
return
}
row, err := a.Store.GetVideoRow(r.Context(), userID, id)
if err != nil {
a.serverError(w, r, "paste get video", err)
return
}
// Dedup: already in the feed with a summary — surface the existing entry,
// don't re-summarize.
if row.Summarized {
a.render(w, r, VideoCard(*row))
return
}
// New or unsummarized: queue + (if a Processor is wired) summarize now, through
// the shared gate. RequestSummarize makes it durable even if the process dies.
if err := a.Store.RequestSummarize(r.Context(), userID, id); err != nil {
a.serverError(w, r, "paste request summarize", err)
return
}
if a.Processor != nil {
a.startProcessing(userID, id)
a.render(w, r, processingCard(*row))
return
}
a.render(w, r, VideoCard(*row))
}
// pasteFailure renders a minimal inline error fragment for the paste form (HTMX
// swaps it in). No templ dependency so it renders even on a bad-input fast path.
func (a *App) pasteFailure(w http.ResponseWriter, status int, msg string) {
w.Header().Set("Content-Type", "text/html; charset=utf-8")
w.WriteHeader(status)
_, _ = io.WriteString(w, `<p class="paste-error" role="alert">`+template.HTMLEscapeString(msg)+`</p>`)
}
// handleRetryNow handles the "Try now" button on rate-limited video cards. It // handleRetryNow handles the "Try now" button on rate-limited video cards. It
// clears the rate_limited_at backoff so the scheduler won't skip the video, then // clears the rate_limited_at backoff so the scheduler won't skip the video, then
// triggers an immediate ProcessVideo — same background path as handleRequestSummarize. // triggers an immediate ProcessVideo — same background path as handleRequestSummarize.
+5 -3
View File
@@ -207,13 +207,15 @@ func TestListChannelFilter(t *testing.T) {
resetDB(t, p) resetDB(t, p)
require.NoError(t, deliver(ctx, app, videoX, "body x")) require.NoError(t, deliver(ctx, app, videoX, "body x"))
seedVideo(t, p, videoX, "X Title", "https://x", time.Time{}) seedVideo(t, p, videoX, "X Title", "https://x", time.Time{})
_, err := p.Exec(ctx, `UPDATE videos SET channel_title = 'Acme Channel' WHERE id = $1`, videoX)
require.NoError(t, err)
// Channel is "youtube" for seeded rows; a non-matching filter hides them. // Selecting a different channel hides the row; selecting its channel shows it.
rec := do(t, app, httptest.NewRequest(http.MethodGet, "/?channel=vimeo", nil)) rec := do(t, app, httptest.NewRequest(http.MethodGet, "/?channel=Other+Channel", nil))
require.Equal(t, http.StatusOK, rec.Code) require.Equal(t, http.StatusOK, rec.Code)
require.NotContains(t, body(t, rec), "X Title") require.NotContains(t, body(t, rec), "X Title")
rec = do(t, app, httptest.NewRequest(http.MethodGet, "/?channel=youtube", nil)) rec = do(t, app, httptest.NewRequest(http.MethodGet, "/?channel=Acme+Channel", nil))
require.Contains(t, body(t, rec), "X Title") require.Contains(t, body(t, rec), "X Title")
} }
+62
View File
@@ -0,0 +1,62 @@
package web
import (
"fmt"
"net/url"
"regexp"
"strings"
)
// youtubeVideoID matches a canonical YouTube video id: exactly 11 URL-safe chars.
var youtubeVideoID = regexp.MustCompile(`^[A-Za-z0-9_-]{11}$`)
// parseYouTubeVideoID extracts the 11-character video id from a pasted YouTube
// URL (watch?v=, youtu.be/, shorts/, embed/) or a bare id. It rejects non-YouTube
// hosts and anything that doesn't yield a valid id, so the paste flow never tries
// to fetch a video that can't exist (Feature 2).
func parseYouTubeVideoID(raw string) (string, error) {
s := strings.TrimSpace(raw)
if s == "" {
return "", fmt.Errorf("empty input")
}
// Bare id (no URL) — accept directly.
if youtubeVideoID.MatchString(s) {
return s, nil
}
// Accept scheme-less URLs (youtube.com/watch?v=...) by giving url.Parse a host.
if !strings.Contains(s, "://") {
s = "https://" + s
}
u, err := url.Parse(s)
if err != nil {
return "", fmt.Errorf("not a URL: %w", err)
}
host := strings.ToLower(u.Hostname())
isYouTube := host == "youtu.be" || host == "youtube.com" || strings.HasSuffix(host, ".youtube.com")
if !isYouTube {
return "", fmt.Errorf("not a YouTube URL: %q", host)
}
var id string
switch {
case host == "youtu.be":
// youtu.be/<id>
id = strings.Trim(u.Path, "/")
case u.Path == "/watch":
id = u.Query().Get("v")
default:
// /shorts/<id>, /embed/<id>
parts := strings.Split(strings.Trim(u.Path, "/"), "/")
if len(parts) == 2 && (parts[0] == "shorts" || parts[0] == "embed") {
id = parts[1]
}
}
if !youtubeVideoID.MatchString(id) {
return "", fmt.Errorf("no YouTube video id in %q", raw)
}
return id, nil
}
+134
View File
@@ -0,0 +1,134 @@
package web_test
import (
"context"
"net/http"
"net/http/httptest"
"net/url"
"strings"
"testing"
"time"
"github.com/stretchr/testify/require"
"gitea.d-ma.be/mathias/tapir/internal/domain"
)
// fakeFetcher is a web.VideoFetcher returning a fixed video (or an error),
// scoped to whatever (userID, videoID) the handler asks for.
type fakeFetcher struct {
title string
err error
calls int
}
func (f *fakeFetcher) FetchVideo(_ context.Context, userID, videoID string) (domain.Video, error) {
f.calls++
if f.err != nil {
return domain.Video{}, f.err
}
return domain.Video{
UserID: userID,
Provider: domain.ProviderYouTube,
ProviderVideoID: videoID,
Title: f.title,
URL: "https://www.youtube.com/watch?v=" + videoID,
}, nil
}
func pasteReq(rawURL string) *http.Request {
req := httptest.NewRequest(http.MethodPost, "/paste",
strings.NewReader("url="+url.QueryEscape(rawURL)))
req.Header.Set("Content-Type", "application/x-www-form-urlencoded")
return req
}
func TestPasteValidURLAddsAndRequests(t *testing.T) {
ctx := context.Background()
app := newApp(t)
resetDB(t, rawPool(t))
app.Fetcher = &fakeFetcher{title: "Pasted Talk"}
p := rawPool(t)
rec := do(t, app, pasteReq("https://youtu.be/dQw4w9WgXcQ"))
require.Equal(t, http.StatusOK, rec.Code)
var (
count, requested int
title string
)
require.NoError(t, p.QueryRow(ctx,
`SELECT count(*), coalesce(max(title),'') FROM videos
WHERE user_id=$1 AND provider_video_id='dQw4w9WgXcQ'`, userID).Scan(&count, &title))
require.Equal(t, 1, count, "pasted video added once, scoped to the user")
require.Equal(t, "Pasted Talk", title)
require.NoError(t, p.QueryRow(ctx,
`SELECT count(*) FROM videos
WHERE user_id=$1 AND provider_video_id='dQw4w9WgXcQ' AND summarize_requested`,
userID).Scan(&requested))
require.Equal(t, 1, requested, "pasted video is queued for summarization (through the gate)")
}
func TestPasteInvalidURLRejected(t *testing.T) {
app := newApp(t)
resetDB(t, rawPool(t))
app.Fetcher = &fakeFetcher{title: "x"}
rec := do(t, app, pasteReq("definitely not a url"))
require.Equal(t, http.StatusBadRequest, rec.Code)
var count int
require.NoError(t, rawPool(t).QueryRow(context.Background(),
`SELECT count(*) FROM videos WHERE user_id=$1`, userID).Scan(&count))
require.Equal(t, 0, count, "invalid input adds nothing")
}
func TestPasteVideoNotFound(t *testing.T) {
app := newApp(t)
resetDB(t, rawPool(t))
app.Fetcher = &fakeFetcher{err: domain.ErrVideoNotFound}
rec := do(t, app, pasteReq("https://youtu.be/dQw4w9WgXcQ"))
require.Equal(t, http.StatusNotFound, rec.Code)
var count int
require.NoError(t, rawPool(t).QueryRow(context.Background(),
`SELECT count(*) FROM videos WHERE user_id=$1`, userID).Scan(&count))
require.Equal(t, 0, count, "a not-found video adds nothing")
}
func TestPasteDedupNoDuplicate(t *testing.T) {
app := newApp(t)
resetDB(t, rawPool(t))
app.Fetcher = &fakeFetcher{title: "Pasted Talk"}
require.Equal(t, http.StatusOK, do(t, app, pasteReq("https://youtu.be/dQw4w9WgXcQ")).Code)
require.Equal(t, http.StatusOK, do(t, app, pasteReq("https://www.youtube.com/watch?v=dQw4w9WgXcQ")).Code)
var count int
require.NoError(t, rawPool(t).QueryRow(context.Background(),
`SELECT count(*) FROM videos WHERE user_id=$1 AND provider_video_id='dQw4w9WgXcQ'`,
userID).Scan(&count))
require.Equal(t, 1, count, "pasting the same video twice must not duplicate the row")
}
func TestListShowsPasteFormForConnectedUserWithVideos(t *testing.T) {
app := newApp(t)
resetDB(t, rawPool(t))
app.Fetcher = &fakeFetcher{title: "x"}
p := rawPool(t)
// Connected user with a non-empty feed (the case the bug missed: hasConnected
// was only computed for an empty feed).
_, err := p.Exec(context.Background(),
`INSERT INTO video_connections (user_id, provider, token_ref, status)
VALUES ($1, 'youtube', 'youtube/x/refresh_token', 'active')`, userID)
require.NoError(t, err)
seedVideo(t, p, "11111111-1111-1111-1111-111111111111", "A talk", "https://youtu.be/aaaaaaaaaaa", time.Now())
rec := do(t, app, httptest.NewRequest(http.MethodGet, "/", nil))
require.Equal(t, http.StatusOK, rec.Code)
require.Contains(t, body(t, rec), `action="/paste"`,
"a connected user must see the paste box even when the feed has videos")
}
+97
View File
@@ -0,0 +1,97 @@
package web
import (
"bytes"
"context"
"gitea.d-ma.be/mathias/tapir/internal/adapters/store"
"strings"
"testing"
)
func TestParseYouTubeVideoID(t *testing.T) {
const id = "dQw4w9WgXcQ"
ok := []struct {
name, in string
}{
{"watch", "https://www.youtube.com/watch?v=" + id},
{"watch no www", "https://youtube.com/watch?v=" + id},
{"watch m", "https://m.youtube.com/watch?v=" + id},
{"watch extra params", "https://www.youtube.com/watch?v=" + id + "&t=42s&list=PLxyz"},
{"watch param after", "https://www.youtube.com/watch?list=PLxyz&v=" + id},
{"short link", "https://youtu.be/" + id},
{"short link param", "https://youtu.be/" + id + "?si=abcd&t=1"},
{"shorts", "https://www.youtube.com/shorts/" + id},
{"embed", "https://www.youtube.com/embed/" + id},
{"bare id", id},
{"http scheme", "http://youtube.com/watch?v=" + id},
{"no scheme", "youtube.com/watch?v=" + id},
{"trailing space", " https://youtu.be/" + id + " "},
}
for _, c := range ok {
t.Run(c.name, func(t *testing.T) {
got, err := parseYouTubeVideoID(c.in)
if err != nil {
t.Fatalf("parseYouTubeVideoID(%q) error: %v", c.in, err)
}
if got != id {
t.Fatalf("parseYouTubeVideoID(%q) = %q, want %q", c.in, got, id)
}
})
}
bad := []struct {
name, in string
}{
{"empty", ""},
{"blank", " "},
{"vimeo", "https://vimeo.com/123456789"},
{"other host", "https://example.com/watch?v=" + id},
{"watch no id", "https://www.youtube.com/watch?v="},
{"short id", "https://youtu.be/abc"},
{"long id", "https://youtu.be/" + id + "extra"},
{"bad chars", "https://youtu.be/dQw4w9Wg!cQ"},
{"not a url", "just some text"},
{"channel url", "https://www.youtube.com/@somechannel"},
}
for _, c := range bad {
t.Run("reject "+c.name, func(t *testing.T) {
if got, err := parseYouTubeVideoID(c.in); err == nil {
t.Fatalf("parseYouTubeVideoID(%q) = %q, want error", c.in, got)
}
})
}
}
func TestListPageShowsPasteFormOnlyWhenConnected(t *testing.T) {
render := func(connected bool) string {
var buf bytes.Buffer
if err := ListPage(listBuckets{}, Filter{}, PipelineStats{}, "", connected, nil).Render(context.Background(), &buf); err != nil {
t.Fatalf("render: %v", err)
}
return buf.String()
}
html := render(true)
if !strings.Contains(html, `name="url"`) || !strings.Contains(html, `action="/paste"`) {
t.Errorf("connected feed must show the paste form")
}
if strings.Contains(render(false), `name="url"`) {
t.Errorf("disconnected feed must not show the paste form")
}
}
func TestFilterMatchesMultipleChannels(t *testing.T) {
f := Filter{Channels: []string{"Acme", "Zeta"}}
row := func(ch string) store.SummaryRow { return store.SummaryRow{ChannelTitle: ch, Summarized: true} }
rows := []store.SummaryRow{row("Acme"), row("Beta"), row("Zeta")}
got := f.apply(rows)
if len(got) != 2 || got[0].ChannelTitle != "Acme" || got[1].ChannelTitle != "Zeta" {
t.Fatalf("multi-channel filter = %+v, want Acme+Zeta only", got)
}
// Empty selection = no channel constraint (all pass).
if n := len(Filter{}.apply(rows)); n != 3 {
t.Fatalf("no channel filter should pass all rows, got %d", n)
}
}
+10
View File
@@ -3,6 +3,8 @@ package web
import ( import (
"context" "context"
"sync" "sync"
"gitea.d-ma.be/mathias/tapir/internal/domain"
) )
// Processor runs the core summarization use case for a single already-discovered // Processor runs the core summarization use case for a single already-discovered
@@ -14,6 +16,14 @@ type Processor interface {
ProcessVideo(ctx context.Context, userID, videoID string) error ProcessVideo(ctx context.Context, userID, videoID string) error
} }
// VideoFetcher resolves an arbitrary YouTube video id to its metadata for the
// paste-a-URL flow (Feature 2). It is a Data API call, NOT the rate-limited
// caption path. Returns domain.ErrVideoNotFound for a deleted/private/typo'd id.
// cmd/tapir wires a per-user YouTube adapter; nil disables the paste route.
type VideoFetcher interface {
FetchVideo(ctx context.Context, userID, videoID string) (domain.Video, error)
}
// ProcessingSet tracks the (user, video) ids currently being summarized in-process // ProcessingSet tracks the (user, video) ids currently being summarized in-process
// so the status endpoint can show the animation until the summary lands. It is // so the status endpoint can show the animation until the summary lands. It is
// ephemeral (single-instance Stage-1): a restart drops it, and the DB holds the // ephemeral (single-instance Stage-1): a restart drops it, and the DB holds the
+27 -5
View File
@@ -2,6 +2,7 @@ package web
import ( import (
"regexp" "regexp"
"slices"
"strings" "strings"
"time" "time"
"unicode/utf8" "unicode/utf8"
@@ -333,7 +334,7 @@ type flashView struct {
// flashMessages maps each flash code to its banner. An unknown code renders no // flashMessages maps each flash code to its banner. An unknown code renders no
// banner (flashFor returns ok=false), so a forged cookie value is inert. // banner (flashFor returns ok=false), so a forged cookie value is inert.
var flashMessages = map[string]flashView{ var flashMessages = map[string]flashView{
flashConnected: {"success", "YouTube account connected."}, flashConnected: {"success", "YouTube account connected — finding your subscriptions. Your newest videos will appear below as they're summarized."},
flashConnectFailed: {"error", "Could not connect your YouTube account. Please try again."}, flashConnectFailed: {"error", "Could not connect your YouTube account. Please try again."},
flashDisconnected: {"success", "Account disconnected."}, flashDisconnected: {"success", "Account disconnected."},
flashDeleted: {"success", "Your account and all its data were deleted."}, flashDeleted: {"success", "Your account and all its data were deleted."},
@@ -469,7 +470,7 @@ func (b listBuckets) empty() bool {
// Dates are kept as the raw YYYY-MM-DD strings so the form re-renders the user's // Dates are kept as the raw YYYY-MM-DD strings so the form re-renders the user's
// input verbatim; parsing happens in matchFilter. // input verbatim; parsing happens in matchFilter.
type Filter struct { type Filter struct {
Channel string Channels []string // selected channel titles; empty = all channels
From string From string
To string To string
OnlySummarized bool // show only videos that have a summary OnlySummarized bool // show only videos that have a summary
@@ -480,7 +481,28 @@ type Filter struct {
// filter) the bar is hidden so the connect CTA stands alone (UX review C1); a // filter) the bar is hidden so the connect CTA stands alone (UX review C1); a
// filter that happens to match nothing still shows the bar so it can be cleared. // filter that happens to match nothing still shows the bar so it can be cleared.
func (f Filter) active() bool { func (f Filter) active() bool {
return f.Channel != "" || f.From != "" || f.To != "" || f.OnlySummarized return len(f.Channels) > 0 || f.From != "" || f.To != "" || f.OnlySummarized
}
// HasChannel reports whether a channel is currently selected (drives the
// multi-select's selected state in the view).
func (f Filter) HasChannel(c string) bool {
return slices.Contains(f.Channels, c)
}
// nonEmptyStrings drops blank entries. A channel multi-select submits real
// channel titles; this guards against a stray empty value reaching the filter.
func nonEmptyStrings(ss []string) []string {
out := ss[:0:0]
for _, s := range ss {
if strings.TrimSpace(s) != "" {
out = append(out, s)
}
}
if len(out) == 0 {
return nil
}
return out
} }
// matches reports whether a row satisfies the filter. Channel is an exact match; // matches reports whether a row satisfies the filter. Channel is an exact match;
@@ -491,7 +513,7 @@ func (f Filter) matches(r store.SummaryRow) bool {
if f.OnlySummarized && !r.Summarized { if f.OnlySummarized && !r.Summarized {
return false return false
} }
if f.Channel != "" && r.Channel != f.Channel { if len(f.Channels) > 0 && !slices.Contains(f.Channels, r.ChannelTitle) {
return false return false
} }
if from, ok := parseDate(f.From); ok { if from, ok := parseDate(f.From); ok {
@@ -521,7 +543,7 @@ func parseDate(s string) (time.Time, bool) {
// apply returns the subset of rows matching the filter, preserving order. // apply returns the subset of rows matching the filter, preserving order.
func (f Filter) apply(rows []store.SummaryRow) []store.SummaryRow { func (f Filter) apply(rows []store.SummaryRow) []store.SummaryRow {
if f.Channel == "" && f.From == "" && f.To == "" && !f.OnlySummarized { if len(f.Channels) == 0 && f.From == "" && f.To == "" && !f.OnlySummarized {
return rows return rows
} }
out := rows[:0:0] out := rows[:0:0]
+38 -4
View File
@@ -104,11 +104,14 @@ templ flashBanner(code string) {
// #summary-list region; a non-HTMX request renders the whole page. flash carries // #summary-list region; a non-HTMX request renders the whole page. flash carries
// a one-shot notification (e.g. "connected", "registered") surfaced on arrival // a one-shot notification (e.g. "connected", "registered") surfaced on arrival
// after a POST→redirect. // after a POST→redirect.
templ ListPage(b listBuckets, f Filter, stats PipelineStats, flash string, hasConnected bool) { templ ListPage(b listBuckets, f Filter, stats PipelineStats, flash string, hasConnected bool, channels []string) {
@Layout("Tapir — Summaries") { @Layout("Tapir — Summaries") {
@flashBanner(flash) @flashBanner(flash)
if hasConnected {
@pasteForm()
}
if !b.empty() || f.active() { if !b.empty() || f.active() {
@filterForm(f) @filterForm(f, channels)
} }
if stats.RateLimited > 0 || stats.Pending > 0 || stats.NoText > 0 { if stats.RateLimited > 0 || stats.Pending > 0 || stats.NoText > 0 {
@pipelineBar(stats) @pipelineBar(stats)
@@ -143,7 +146,29 @@ templ pipelineBar(s PipelineStats) {
</div> </div>
} }
templ filterForm(f Filter) { // pasteForm lets a connected user summarize any YouTube video by pasting its URL
// (Feature 2). The result (a video card, or an inline error) swaps into
// #paste-result; the next list refresh shows it inline. Summarization runs
// through the shared caption rate gate like every other fetch.
templ pasteForm() {
<form
class="paste"
method="post"
action="/paste"
hx-post="/paste"
hx-target="#paste-result"
hx-swap="innerHTML"
>
<label>
Summarize any video
<input type="url" name="url" placeholder="Paste a YouTube link…" required/>
</label>
<button type="submit">Add</button>
</form>
<div id="paste-result"></div>
}
templ filterForm(f Filter, channels []string) {
<form <form
class="filters" class="filters"
method="get" method="get"
@@ -153,7 +178,16 @@ templ filterForm(f Filter) {
hx-swap="innerHTML" hx-swap="innerHTML"
hx-indicator="#filter-indicator" hx-indicator="#filter-indicator"
> >
<label>Channel <input type="text" name="channel" value={ f.Channel } placeholder="any"/></label> if len(channels) > 0 {
<label>
Channels
<select name="channel" multiple size="4">
for _, c := range channels {
<option value={ c } selected?={ f.HasChannel(c) }>{ c }</option>
}
</select>
</label>
}
<label class="filter-check"> <label class="filter-check">
<input type="checkbox" name="summarized" value="1" if f.OnlySummarized { checked }/> <input type="checkbox" name="summarized" value="1" if f.OnlySummarized { checked }/>
Summarized only Summarized only
File diff suppressed because it is too large Load Diff
@@ -39,7 +39,14 @@ var scenarioCoverage = map[string]string{
// connect_account.feature // connect_account.feature
"Connect a YouTube account": "TestCallbackExchangesAndRecordsConnection", "Connect a YouTube account": "TestCallbackExchangesAndRecordsConnection",
"Connecting an account discovers videos immediately": "TestCallbackTriggersDiscovery", "Connecting an account discovers videos immediately": "TestCallbackTriggersDiscovery",
"Connecting summarizes my newest videos right away": "TestNewestUnsummarizedVideoIDs",
"Tokens are never stored in the clear": "TestCallbackExchangesAndRecordsConnection", "Tokens are never stored in the clear": "TestCallbackExchangesAndRecordsConnection",
// paste_url.feature
"Paste a valid YouTube URL": "TestPasteValidURLAddsAndRequests",
"Pasting an invalid link is rejected": "TestPasteInvalidURLRejected",
"Pasting a video that cannot be found is honest": "TestPasteVideoNotFound",
"Pasting the same video twice does not duplicate it": "TestPasteDedupNoDuplicate",
"Revoking a connection stops watching but keeps history": "TestDisconnectRemovesTokenAndConnectionKeepsAccount", "Revoking a connection stops watching but keeps history": "TestDisconnectRemovesTokenAndConnectionKeepsAccount",
// summarize_mode.feature // summarize_mode.feature
@@ -60,6 +67,7 @@ var scenarioCoverage = map[string]string{
"A subscribed channel posts a video with no usable transcript": "TestVideoWithNoTranscriptIsSkipped", "A subscribed channel posts a video with no usable transcript": "TestVideoWithNoTranscriptIsSkipped",
"A channel I am not subscribed to posts a video": "TestUnsubscribedChannelVideoIsNotProcessed", "A channel I am not subscribed to posts a video": "TestUnsubscribedChannelVideoIsNotProcessed",
"The same video is not summarized twice": "TestAlreadySummarizedVideoIsNotReprocessed", "The same video is not summarized twice": "TestAlreadySummarizedVideoIsNotReprocessed",
"Re-analyzing a stored video does not re-fetch its transcript": "TestProcessNewVideo_SecondSummarizeDoesNotRefetch",
} }
var ( var (