Compare commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
5219561a91 | ||
|
|
9db06d8a63 | ||
|
|
1aa8a97f95 | ||
|
|
cc69a912f4 | ||
|
|
f4a0544903 | ||
|
|
e2a52789b9 | ||
|
|
e696b6405b | ||
|
|
e9b5a3f3e7 | ||
|
|
0ba78e8868 | ||
|
|
821d5f99cd | ||
|
|
5c70408e75 | ||
|
|
cb6917ca59 | ||
|
|
099b2d4c68 | ||
|
|
f66c1bcdcc | ||
|
|
1e65c3b413 | ||
|
|
87c978774f | ||
|
|
70a9f1d4cd | ||
|
|
1d5b2c6365 | ||
|
|
62acfee2ed | ||
|
|
2907801aca | ||
|
|
59050c4db6 | ||
|
|
c320ed88aa | ||
|
|
bddd75d92e | ||
|
|
e4c701c6f1 | ||
|
|
b527db9739 | ||
|
|
c5f556d1d6 | ||
|
|
a884e7e9c5 | ||
|
|
27fd33c99c | ||
|
|
2c96926ff7 |
@@ -46,8 +46,9 @@ These caused real mistakes that were caught and corrected; the corrections are l
|
|||||||
(See `DECISIONS.md` for full rationale. Listed here so you don't propose them.)
|
(See `DECISIONS.md` for full rationale. Listed here so you don't propose them.)
|
||||||
|
|
||||||
- **No Supabase** — reuse Dex / ESO+1Password / Postgres (ADR-002).
|
- **No Supabase** — reuse Dex / ESO+1Password / Postgres (ADR-002).
|
||||||
- **No global cross-tenant video/transcript table** — per-user isolation (data-model). Dedup
|
- **No global cross-tenant *video* table** — videos stay per-user (data-model). Transcripts ARE
|
||||||
across users is a Future C concern, not a Stage 0/1 default.
|
shared since ADR-021 (public caption content, keyed by `(provider, provider_video_id)`, non-RLS)
|
||||||
|
so re-analysis never re-fetches; the *videos* half of cross-tenant dedup stays a Future C concern.
|
||||||
- **No audio-download + speech-to-text in the core path** — captions-first (ADR-007). STT is a
|
- **No audio-download + speech-to-text in the core path** — captions-first (ADR-007). STT is a
|
||||||
deferred, bounded optional component.
|
deferred, bounded optional component.
|
||||||
- **No public SaaS / sign-up / billing / Google OAuth verification at scale** — Future C,
|
- **No public SaaS / sign-up / billing / Google OAuth verification at scale** — Future C,
|
||||||
@@ -60,8 +61,14 @@ These caused real mistakes that were caught and corrected; the corrections are l
|
|||||||
stores, and sinks are adapters. Adding a video provider or a sink = a new adapter implementing
|
stores, and sinks are adapters. Adding a video provider or a sink = a new adapter implementing
|
||||||
the interface, nothing in the engine changes. This is what keeps "standalone vs homelab" a
|
the interface, nothing in the engine changes. This is what keeps "standalone vs homelab" a
|
||||||
wiring choice (ADR-003).
|
wiring choice (ADR-003).
|
||||||
- **BDD.** The `docs/use-cases/*.feature` files are the behavior spec. New behavior gets a
|
- **BDD.** The `docs/use-cases/*.feature` files are the behavior spec (design records — there is
|
||||||
scenario; the use-case core is tested through fake adapters, not live YouTube/brain.
|
no godog runner). New behavior gets a scenario; the use-case core is tested through fake
|
||||||
|
adapters, not live YouTube/brain. A name-coverage gate (`test/acceptance/scenario_coverage_test.go`,
|
||||||
|
`TestScenarioCoverage`) keeps the two from drifting: every non-`@pending` scenario must be
|
||||||
|
mapped to an existing Go test in `scenarioCoverage`. When you add a scenario, either map it to
|
||||||
|
its covering test or tag it `@pending` in the `.feature` with a one-line reason. It checks the
|
||||||
|
*link*, not that the test exercises the scenario — that's the deliberate trade for not running
|
||||||
|
godog (see issue #5 / the BDD-runner decision).
|
||||||
|
|
||||||
## Skills (engineering discipline)
|
## Skills (engineering discipline)
|
||||||
|
|
||||||
@@ -79,7 +86,7 @@ Skills live in the canonical library `mathias/skills` and are wired into this re
|
|||||||
|
|
||||||
## Current build state (start here for the first task)
|
## Current build state (start here for the first task)
|
||||||
|
|
||||||
The repo is **green and shipping** — last tag `v0.4.0`. `task check` passes (fmt, vet, lint,
|
The repo is **green and shipping** — last tag `v0.15.0`. `task check` passes (fmt, vet, lint,
|
||||||
`go test -p 1 ./...`). Go is `1.26.1` (see `go.mod`).
|
`go test -p 1 ./...`). Go is `1.26.1` (see `go.mod`).
|
||||||
|
|
||||||
- Clean Architecture core is implemented: `internal/domain` (entities), `internal/ports`
|
- Clean Architecture core is implemented: `internal/domain` (entities), `internal/ports`
|
||||||
@@ -88,13 +95,19 @@ The repo is **green and shipping** — last tag `v0.4.0`. `task check` passes (f
|
|||||||
green against it.
|
green against it.
|
||||||
- Adapters present under `internal/adapters/`: `youtube` (captions-first `VideoSource`,
|
- Adapters present under `internal/adapters/`: `youtube` (captions-first `VideoSource`,
|
||||||
timedtext/InnerTube acquisition per ADR-010), `summarizer` + `llm` (the copied AI router,
|
timedtext/InnerTube acquisition per ADR-010), `summarizer` + `llm` (the copied AI router,
|
||||||
Primary→Fallback per ADR-004), `store` (Postgres, golang-migrate migrations 001–006),
|
now a resilient endpoint chain — local primary → local fallback → external worst-case,
|
||||||
|
parse-failure-aware, ADR-004 + ADR-022), `store` (Postgres, golang-migrate migrations 001–015),
|
||||||
`secrets` (file-backed `SecretStore`). The brain HTTP sink (ADR-005) is the remaining
|
`secrets` (file-backed `SecretStore`). The brain HTTP sink (ADR-005) is the remaining
|
||||||
optional sink.
|
optional sink.
|
||||||
- Stage 1 is open (ADR-012): multi-user with **DB-enforced** isolation — Postgres RLS `FORCE`d
|
- Stage 1 is open (ADR-012): multi-user with **DB-enforced** isolation — Postgres RLS `FORCE`d
|
||||||
on all user-owned tables (migration 003), two-user isolation test in
|
on all user-owned tables (migration 003), two-user isolation test in
|
||||||
`internal/adapters/store/rls_test.go`. Registration gate, per-user YouTube web connect, and
|
`internal/adapters/store/rls_test.go`. Registration gate, per-user YouTube web connect, and
|
||||||
account management (disconnect / delete, ADR-013) all shipped.
|
account management (disconnect / delete, ADR-013) all shipped.
|
||||||
|
- Transcript persistence (ADR-021, migration 015): transcripts are a **shared, non-RLS** store
|
||||||
|
keyed by `(provider, provider_video_id)` — the single exception to the isolation boundary
|
||||||
|
(`TestTranscriptsTableIsSharedNotRLS`). The engine reads stored transcripts before any caption
|
||||||
|
fetch (`usecase.resolveTranscript`), so re-analysis — re-summarize, paste-a-URL, onboarding
|
||||||
|
burst — never re-touches YouTube. Per-user summaries/videos stay RLS-scoped.
|
||||||
- `cmd/tapir` subcommands: `list`, `show`, `auth` (interactive host-side OAuth), `run` (batch
|
- `cmd/tapir` subcommands: `list`, `show`, `auth` (interactive host-side OAuth), `run` (batch
|
||||||
watch→summarize), `serve` (the HTMX+Templ web reader/writer under `internal/web`, a new
|
watch→summarize), `serve` (the HTMX+Templ web reader/writer under `internal/web`, a new
|
||||||
transport over the unchanged engine/ports — ADR-003). `tapir env` prints config.
|
transport over the unchanged engine/ports — ADR-003). `tapir env` prints config.
|
||||||
|
|||||||
+181
-1
@@ -743,6 +743,186 @@ collapse keys off the same window (`App.RecencyWindow=0` → everything inline).
|
|||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
|
## ADR-021 — Persist transcripts as shared, video-keyed public content (re-analysis never re-fetches)
|
||||||
|
|
||||||
|
**Status:** Accepted (2026-06-09). **Reopens the transcripts half of** the "Global cross-tenant
|
||||||
|
`videos`/`transcripts` table" rejection (data-model.md). **Builds on ADR-007** (captions-first),
|
||||||
|
**ADR-010/ADR-014** (the per-IP caption rate gate), and **ADR-012** (per-user RLS isolation).
|
||||||
|
|
||||||
|
**Context.** Every summarization fetches the transcript fresh through the caption path, even when
|
||||||
|
the exact same transcript was fetched moments ago — for the same user re-summarizing, or for a
|
||||||
|
second user who happens to watch the same video. The caption fetch is the one genuinely scarce,
|
||||||
|
genuinely risky operation in the system: YouTube's timedtext endpoint is unofficial and per-IP
|
||||||
|
rate-limited (ADR-010), and tripping it risks the maintainer's Google standing (ADR-014). So the
|
||||||
|
operation we most want to *avoid repeating* is the one we currently repeat unconditionally. A
|
||||||
|
transcript is **public content** — the same words YouTube serves to anyone — and carries nothing
|
||||||
|
user-identifying. The per-user isolation that protects summaries, feeds, and tokens (ADR-012) is
|
||||||
|
the wrong shape for it: it forces a re-fetch per user for data that is identical across users.
|
||||||
|
|
||||||
|
The original rejection ("Global cross-tenant `videos`/`transcripts` table") bundled videos and
|
||||||
|
transcripts together and rejected both on the grounds that "at 1–5 users, re-summarizing is
|
||||||
|
cheaper than the coupling." That reasoning holds for **videos** (per-user feed rows, genuinely
|
||||||
|
user-scoped) but not for **transcripts**: the cost being avoided is not LLM re-summarization, it
|
||||||
|
is a *rate-gated, reputation-risky network fetch*, and that cost is paid per re-fetch regardless
|
||||||
|
of user count. One re-fetch avoided is strictly worth more than the coupling it removes.
|
||||||
|
|
||||||
|
**Decision.**
|
||||||
|
1. **A single shared `transcripts` table, keyed by the cross-user dedup key
|
||||||
|
`(provider, provider_video_id)`** — the stable public identity of the video, not Tapir's
|
||||||
|
internal per-user `videos.id`. Columns: the key, `source` (`captions`/`none`), `language`,
|
||||||
|
`content`, `fetched_at`. It holds **only public caption content + the video's public id** —
|
||||||
|
nothing user-identifying — and is therefore **NOT RLS-scoped**: no `user_id`, no policy, no
|
||||||
|
`FORCE ROW LEVEL SECURITY`. This is the deliberate, single exception to the ADR-012 isolation
|
||||||
|
boundary, and the only one.
|
||||||
|
2. **Summarize path becomes read-stored-first.** Have a stored transcript for this video? →
|
||||||
|
summarize from the stored text, **no caption fetch**. No stored transcript? → fetch *through
|
||||||
|
the unchanged gate* (ADR-014) → store it → summarize. The gate is neither bypassed nor
|
||||||
|
weakened; persistence reduces how *often* we reach it, never how *fast*.
|
||||||
|
3. **De-facto cross-user dedup is the intended behaviour, not a feature with a switch.** Two
|
||||||
|
users who share a video share the one transcript row. A permanent `source = 'none'` (no
|
||||||
|
captions) is stored too, so a known-caption-less video is not re-fetched by anyone. A
|
||||||
|
transient 429 (`SourceRateLimited`) is **never** stored as terminal — it stays a per-user
|
||||||
|
retry via the existing `transcript_status` backoff (ADR-014), so persistence cannot mask a
|
||||||
|
rate-limit into a false "no transcript."
|
||||||
|
4. **Per-user `summaries` stay RLS-scoped (ADR-012 unchanged)** and reference the transcript by
|
||||||
|
video id. Videos stay per-user. Only transcripts go shared.
|
||||||
|
|
||||||
|
**Consequences.** Re-analysis (re-summarize, different model, paste of an already-seen video,
|
||||||
|
onboarding of a second user with overlapping subscriptions) never re-touches YouTube — the
|
||||||
|
primary win, and it *reduces* aggregate caption-gate pressure, reinforcing ADR-010/ADR-014 rather
|
||||||
|
than straining them. The isolation surface gains exactly one non-RLS table; an isolation test
|
||||||
|
asserts the boundary is *exactly* there and has not leaked to any user-owned table (this is the
|
||||||
|
proof the public-content classification was implemented as designed). It also unblocks
|
||||||
|
multi-model / customizable analysis (re-run analysis on stored text for free) — enabling that is
|
||||||
|
this ADR's point; building it is separate.
|
||||||
|
|
||||||
|
**Reversibility.** The read-stored-first check is the only behavioural coupling; removing it
|
||||||
|
restores fetch-every-time. The down-migration recreates the per-user RLS-scoped transcripts shape
|
||||||
|
(001/003). No user-facing surface depends on cross-user sharing — sharing is the *storage shape*,
|
||||||
|
never exposed in the UI.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## ADR-022 — Summarizer is a resilient endpoint chain, not a single model
|
||||||
|
|
||||||
|
**Status:** Accepted (2026-06-10). **Extends ADR-004** (the copied `llm` Primary→Fallback
|
||||||
|
routing). Triggered by the first friendly-pilot live run, where a connected user got **zero**
|
||||||
|
summaries after 12h.
|
||||||
|
|
||||||
|
**Context.** Stage-0 ran a single summarizer model (`koala/phi4-mini`) with no fallback wired
|
||||||
|
(`summarizer.New(primary, nil)`). The live run exposed three independent failure modes, each of
|
||||||
|
which silently produced no summary:
|
||||||
|
|
||||||
|
1. **Context overflow.** `phi4-mini` has an 8k context. Real transcripts (one was 11,602 tokens)
|
||||||
|
exceed it and the gateway returns HTTP 400 — and the request also sent `max_tokens=8192`, so
|
||||||
|
even a short transcript plus the completion budget could overflow the window.
|
||||||
|
2. **Malformed model output.** `phi4-mini` intermittently emits `highlights` as a bare string
|
||||||
|
instead of an array, producing `cannot unmarshal string into []string`. The old code returned
|
||||||
|
the parse error **without** trying any other model — a 200-with-bad-JSON short-circuited.
|
||||||
|
3. **No fallback existed at all** — any primary failure was terminal for that video.
|
||||||
|
|
||||||
|
`phi4-mini` is kept as primary deliberately: it is fast and, on transcripts that fit, correct.
|
||||||
|
The fix is resilience around it, not replacing it.
|
||||||
|
|
||||||
|
**Decision.**
|
||||||
|
1. **Ordered endpoint chain (`summarizer.NewChain`).** Endpoints are tried in order; the first to
|
||||||
|
return a *parseable* summary wins. Default chain:
|
||||||
|
`koala/phi4-mini` (primary, local) → `koala/phi4-14b` (fallback, local) →
|
||||||
|
`berget/mistral-small` (worst-case, external). All three are reached through the **one** LiteLLM
|
||||||
|
gateway by alias — the gateway already fronts both llama-swap and berget — so a fallback is a
|
||||||
|
different alias, not a second client config.
|
||||||
|
2. **A parse failure advances the chain, same as a transport error.** "Reliably summarized" means
|
||||||
|
*parseable summary returned*, not *HTTP 200*. This is the behaviour the old Primary→Fallback
|
||||||
|
shape missed.
|
||||||
|
3. **Tolerant parse.** `highlights`/`takeaways` coerce from a bare string (or a mixed scalar
|
||||||
|
array) to `[]string`, so the most common small-model quirk is absorbed **without** spending a
|
||||||
|
fallback round-trip — keeping the fast path fast.
|
||||||
|
4. **Transcript truncation (`TAPIR_MAX_TRANSCRIPT_CHARS`, default 18000).** Input is bounded
|
||||||
|
up-front to fit a small-context primary, so overflow is prevented rather than recovered-from.
|
||||||
|
5. **Bounded completion budget (`TAPIR_SUMMARY_MAX_TOKENS`, default 1500).** A summary needs few
|
||||||
|
hundred tokens; the old 8192 budget itself contributed to 8k-window overflow.
|
||||||
|
|
||||||
|
**Local-first guarantee preserved.** The chain ordering *is* the guarantee: locals are tried
|
||||||
|
first, so content reaches the external endpoint only after every local endpoint has failed.
|
||||||
|
`TAPIR_CLOUD_FALLBACK_MODEL=""` removes the external endpoint entirely — the lever a
|
||||||
|
**client/NDA deployment** pulls so content never leaves the local stack. With no external endpoint
|
||||||
|
configured the `ai_routing.feature` "content only local" scenarios hold unchanged.
|
||||||
|
|
||||||
|
**Reversibility.** Pure wiring + config. Setting `TAPIR_FALLBACK_MODEL` and
|
||||||
|
`TAPIR_CLOUD_FALLBACK_MODEL` empty collapses the chain back to single-primary behaviour; the
|
||||||
|
tolerant parse and truncation are strict supersets of the old behaviour (a previously-parseable
|
||||||
|
reply still parses; a transcript within budget is unchanged).
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## ADR-023 — Drop Shorts/livestreams at discovery to protect the caption budget
|
||||||
|
|
||||||
|
**Status:** Accepted (2026-06-10). **Builds on ADR-014** (per-IP caption rate limit is the
|
||||||
|
binding constraint) and the ADR-022 live-run findings.
|
||||||
|
|
||||||
|
**Context.** The scarce resource is the unofficial timedtext caption fetch (per-egress-IP
|
||||||
|
429, ~3 successful/pass). The first multi-user run showed the candidate set was mostly noise —
|
||||||
|
Shorts, sub-minute clips, and live broadcasts — each of which still consumes a caption-fetch
|
||||||
|
attempt (and a "none" result is a *completed* fetch, so it costs budget even when it yields
|
||||||
|
nothing). Spending the rate-limited budget on content the user will not read is the waste to
|
||||||
|
cut first; it is cheaper and lower-risk than raising the ceiling (multi-IP, Whisper).
|
||||||
|
|
||||||
|
Duration and live status are NOT in the playlistItems discovery response, but they ARE in the
|
||||||
|
Data API `videos.list` (contentDetails.duration + snippet.liveBroadcastContent) — the official
|
||||||
|
**quota-based** API (1 unit/call, 50 ids/call), which is a *different* limit from the timedtext
|
||||||
|
429. So one cheap quota call buys a filter that saves many expensive throttled fetches.
|
||||||
|
|
||||||
|
**Decision.**
|
||||||
|
1. `NewVideos` enriches its candidates with a single `videos.list` call and drops, before
|
||||||
|
returning: videos shorter than `TAPIR_MIN_VIDEO_SECONDS` (default 60) and any `live`/
|
||||||
|
`upcoming` broadcast. Dropped videos are never persisted, so they also declutter the list.
|
||||||
|
2. The filter is **degrade-open**: `MinVideoSeconds=0` disables it (no quota call), and a
|
||||||
|
`videos.list` error returns the candidates unfiltered — discovery must never break because a
|
||||||
|
metadata call hiccuped (worst case = pre-ADR-023 behaviour).
|
||||||
|
3. The paste-a-URL path (`VideoByID`) is **not** filtered — an explicit user request for a
|
||||||
|
specific video (even a Short) is honoured.
|
||||||
|
|
||||||
|
**Reversibility.** Pure discovery-time filter + config. `TAPIR_MIN_VIDEO_SECONDS=0` restores
|
||||||
|
the old behaviour. No schema change, no effect on already-stored videos.
|
||||||
|
|
||||||
|
**Quota note.** Per-channel enrichment adds ~1 unit/channel/pass. At pilot scale (≤3 users)
|
||||||
|
this is well under the 10k/day cap; at larger scale, batch `videos.list` across channels
|
||||||
|
(50 ids/call) by collecting all discovered ids per pass before enriching.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## ADR-024 — Per-channel caption-availability memory
|
||||||
|
|
||||||
|
**Status:** Accepted (2026-06-10). **Builds on ADR-014** (per-IP caption budget), **ADR-021**
|
||||||
|
(shared transcript cache), **ADR-023** (Shorts filter).
|
||||||
|
|
||||||
|
**Context.** After ADR-021 caches transcripts and ADR-023 drops Shorts, the remaining caption
|
||||||
|
waste is the *first* fetch on every new video of a channel that never publishes English captions
|
||||||
|
(foreign-language news, music, etc.). Each costs one rate-limited fetch to resolve to "none" —
|
||||||
|
and on a throttled IP that fetch may 429 and churn the backoff machinery before it ever gets a
|
||||||
|
verdict. A pilot user's feed had several such channels.
|
||||||
|
|
||||||
|
**Decision.** Remember, per `(user, channel)`, a streak of consecutive no-caption outcomes
|
||||||
|
(`channel_caption_state`, migration 016, RLS-scoped like the rest of the user-owned schema).
|
||||||
|
Once the streak reaches `TAPIR_CHANNEL_CAPTIONLESS_THRESHOLD` (default 5) the channel is
|
||||||
|
suppressed — its videos are discovered/listed but not caption-fetched — for
|
||||||
|
`TAPIR_CHANNEL_CAPTIONLESS_WINDOW` (default 14d), after which one video is re-probed
|
||||||
|
(auto-recovery for a channel that starts adding captions). A successful fetch resets the streak;
|
||||||
|
a fresh 429 does NOT count (transient, not a caption verdict). An explicit manual request
|
||||||
|
bypasses suppression. `threshold = 0` disables the feature.
|
||||||
|
|
||||||
|
**Why per-user, not global.** Caption availability is really a channel property (public), so a
|
||||||
|
global table would let users share the learning. But subscriptions are per-user (ADR-012) and at
|
||||||
|
pilot scale users' channel sets barely overlap, so per-user + RLS keeps it consistent with the
|
||||||
|
existing isolation model with no new non-RLS exception to justify. Promoting to a shared table
|
||||||
|
(like transcripts, ADR-021) is a future optimisation if channel overlap grows.
|
||||||
|
|
||||||
|
**Reversibility.** Migration 016 is a clean drop; `threshold = 0` disables at runtime. The
|
||||||
|
memory only ever *suppresses fetches* — it never deletes content or affects already-stored
|
||||||
|
summaries.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
## Rejected alternatives
|
## Rejected alternatives
|
||||||
|
|
||||||
Approaches considered during the 2026-06-02 planning + grill session and **deliberately not
|
Approaches considered during the 2026-06-02 planning + grill session and **deliberately not
|
||||||
@@ -757,7 +937,7 @@ maps to the ADR that settles it.
|
|||||||
| Lifting shared packages into a `brain-common` module | Couples Tapir's release cycle to the monolith for negligible code savings | ADR-004 |
|
| Lifting shared packages into a `brain-common` module | Couples Tapir's release cycle to the monolith for negligible code savings | ADR-004 |
|
||||||
| Importing/replicating the filesystem `brain` package | Assumes co-location with the brain git checkout; wrong for a standalone networked service | ADR-005 |
|
| Importing/replicating the filesystem `brain` package | Assumes co-location with the brain git checkout; wrong for a standalone networked service | ADR-005 |
|
||||||
| Reusing `ingestion`'s `oauth` package for YouTube/Vimeo | Same name, opposite direction — it's inbound MCP-server auth, not outbound provider OAuth | ADR-006 |
|
| Reusing `ingestion`'s `oauth` package for YouTube/Vimeo | Same name, opposite direction — it's inbound MCP-server auth, not outbound provider OAuth | ADR-006 |
|
||||||
| Global cross-tenant `videos`/`transcripts` table (dedup) | Reintroduces the cross-domain DB coupling the homelab review is removing; at 1–5 users, re-summarizing is cheaper than the coupling | data-model.md |
|
| Global cross-tenant `videos`/`transcripts` table (dedup) | Reintroduces the cross-domain DB coupling the homelab review is removing; at 1–5 users, re-summarizing is cheaper than the coupling. **Transcripts half reopened by ADR-021** — the avoided cost there is a rate-gated, reputation-risky *caption fetch*, not LLM re-summarization, so it outweighs the coupling; **videos stay per-user.** | data-model.md, **ADR-021** (transcripts only) |
|
||||||
| Audio-download + Whisper STT in the core path | ToS-grey, breakage-prone (yt-dlp), contends for koala GPU with the JEPA PoC; captions alone test the core hypothesis | ADR-007 |
|
| Audio-download + Whisper STT in the core path | ToS-grey, breakage-prone (yt-dlp), contends for koala GPU with the JEPA PoC; captions alone test the core hypothesis | ADR-007 |
|
||||||
| Building multi-tenant SaaS / Google OAuth verification now | "Real users soon" was lowered to Future B; SaaS machinery before the Stage 0 self-use gate is the primary documented anti-goal | ADR-008, VISION |
|
| Building multi-tenant SaaS / Google OAuth verification now | "Real users soon" was lowered to Future B; SaaS machinery before the Stage 0 self-use gate is the primary documented anti-goal | ADR-008, VISION |
|
||||||
| Delegating the S5 reuse spike to an agent swarm | A 1-hour sequential read-and-judge with a single coupled conclusion; orchestration overhead exceeds the work, and it's Diamond-1 judgment the maintainer wanted to own | (process note) |
|
| Delegating the S5 reuse spike to an agent swarm | A 1-hour sequential read-and-judge with a single coupled conclusion; orchestration overhead exceeds the work, and it's Diamond-1 judgment the maintainer wanted to own | (process note) |
|
||||||
|
|||||||
@@ -75,8 +75,10 @@ password or Google); a new Dex subject is routed to `/register` to create a Tapi
|
|||||||
Stage 1 is multi-user: each user connects their own YouTube account from the browser and
|
Stage 1 is multi-user: each user connects their own YouTube account from the browser and
|
||||||
manages their own summaries under DB-enforced RLS isolation. When `TAPIR_DISCOVERY_INTERVAL`
|
manages their own summaries under DB-enforced RLS isolation. When `TAPIR_DISCOVERY_INTERVAL`
|
||||||
is set (e.g. `2h`), the serve process runs a scheduled discovery pass for every registered
|
is set (e.g. `2h`), the serve process runs a scheduled discovery pass for every registered
|
||||||
user automatically — no CronJob required. See `docs/homelab-integration.md` for the full
|
user automatically — no CronJob required. In auto mode only videos published within
|
||||||
config reference.
|
`TAPIR_AUTO_SUMMARIZE_WINDOW` (default ~7d, ADR-020) are summarised automatically; older videos
|
||||||
|
are listed and summarised on demand, so a large back-catalogue doesn't keep re-driving the
|
||||||
|
caption rate gate. See `docs/homelab-integration.md` for the full config reference.
|
||||||
|
|
||||||
### Headless on koala
|
### Headless on koala
|
||||||
|
|
||||||
|
|||||||
@@ -0,0 +1,50 @@
|
|||||||
|
package main
|
||||||
|
|
||||||
|
import (
|
||||||
|
"context"
|
||||||
|
"log/slog"
|
||||||
|
"sync"
|
||||||
|
|
||||||
|
"gitea.d-ma.be/mathias/tapir/internal/runner"
|
||||||
|
)
|
||||||
|
|
||||||
|
// discoveryRunner runs one user's discovery pass.
|
||||||
|
type discoveryRunner func(ctx context.Context, userID string) (runner.Stats, error)
|
||||||
|
|
||||||
|
// serialize wraps run so calls never overlap: every discovery pass — scheduled
|
||||||
|
// or connect-triggered (#6) — acquires the same lock, preserving the
|
||||||
|
// one-fetcher-at-a-time invariant the scheduler relies on (ADR-018, the
|
||||||
|
// single-replica assumption). Locking is per-user, so a connect-triggered pass
|
||||||
|
// interleaves between the scheduler's users instead of waiting for a whole pass.
|
||||||
|
func serialize(mu *sync.Mutex, run discoveryRunner) discoveryRunner {
|
||||||
|
return func(ctx context.Context, userID string) (runner.Stats, error) {
|
||||||
|
mu.Lock()
|
||||||
|
defer mu.Unlock()
|
||||||
|
return run(ctx, userID)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// discoveryTrigger fires an out-of-band discovery pass for one user without
|
||||||
|
// blocking the caller (the connect HTTP handler). The pass runs on the server's
|
||||||
|
// long-lived ctx — not the request ctx — so it survives the post-connect
|
||||||
|
// redirect. run is the serialized runner, so a trigger never overlaps the
|
||||||
|
// scheduler. Satisfies web.DiscoveryTrigger.
|
||||||
|
type discoveryTrigger struct {
|
||||||
|
ctx context.Context
|
||||||
|
run discoveryRunner
|
||||||
|
// onboard, when set, runs after the discovery pass to summarize a capped number
|
||||||
|
// of the user's newest videos (Feature 1). Optional.
|
||||||
|
onboard func(ctx context.Context, userID string)
|
||||||
|
log *slog.Logger
|
||||||
|
}
|
||||||
|
|
||||||
|
func (t *discoveryTrigger) Enqueue(userID string) {
|
||||||
|
go func() {
|
||||||
|
if _, err := t.run(t.ctx, userID); err != nil {
|
||||||
|
t.log.Warn("discovery: connect-triggered pass had errors", "user", userID, "err", err)
|
||||||
|
}
|
||||||
|
if t.onboard != nil {
|
||||||
|
t.onboard(t.ctx, userID)
|
||||||
|
}
|
||||||
|
}()
|
||||||
|
}
|
||||||
@@ -0,0 +1,62 @@
|
|||||||
|
package main
|
||||||
|
|
||||||
|
import (
|
||||||
|
"context"
|
||||||
|
"fmt"
|
||||||
|
"sync"
|
||||||
|
"sync/atomic"
|
||||||
|
"testing"
|
||||||
|
"time"
|
||||||
|
|
||||||
|
"github.com/stretchr/testify/require"
|
||||||
|
|
||||||
|
"gitea.d-ma.be/mathias/tapir/internal/runner"
|
||||||
|
)
|
||||||
|
|
||||||
|
// serialize must guarantee at most one discovery pass runs at a time, so a
|
||||||
|
// connect-triggered pass never fetches concurrently with the scheduler.
|
||||||
|
func TestSerializeRunsOneAtATime(t *testing.T) {
|
||||||
|
var active, maxActive int32
|
||||||
|
run := func(_ context.Context, _ string) (runner.Stats, error) {
|
||||||
|
n := atomic.AddInt32(&active, 1)
|
||||||
|
for { // record the high-water mark of concurrent runs
|
||||||
|
m := atomic.LoadInt32(&maxActive)
|
||||||
|
if n <= m || atomic.CompareAndSwapInt32(&maxActive, m, n) {
|
||||||
|
break
|
||||||
|
}
|
||||||
|
}
|
||||||
|
time.Sleep(2 * time.Millisecond)
|
||||||
|
atomic.AddInt32(&active, -1)
|
||||||
|
return runner.Stats{}, nil
|
||||||
|
}
|
||||||
|
|
||||||
|
s := serialize(&sync.Mutex{}, run)
|
||||||
|
var wg sync.WaitGroup
|
||||||
|
for i := 0; i < 20; i++ {
|
||||||
|
wg.Add(1)
|
||||||
|
go func(i int) { defer wg.Done(); _, _ = s(context.Background(), fmt.Sprintf("u%d", i)) }(i)
|
||||||
|
}
|
||||||
|
wg.Wait()
|
||||||
|
|
||||||
|
require.Equal(t, int32(1), atomic.LoadInt32(&maxActive),
|
||||||
|
"serialize must run at most one pass at a time")
|
||||||
|
}
|
||||||
|
|
||||||
|
// Enqueue runs the user's pass out-of-band (non-blocking) on the trigger's ctx.
|
||||||
|
func TestDiscoveryTriggerEnqueueRunsUser(t *testing.T) {
|
||||||
|
done := make(chan string, 1)
|
||||||
|
run := func(_ context.Context, userID string) (runner.Stats, error) {
|
||||||
|
done <- userID
|
||||||
|
return runner.Stats{}, nil
|
||||||
|
}
|
||||||
|
tr := &discoveryTrigger{ctx: context.Background(), run: run, log: quietLog()}
|
||||||
|
|
||||||
|
tr.Enqueue("u1")
|
||||||
|
|
||||||
|
select {
|
||||||
|
case got := <-done:
|
||||||
|
require.Equal(t, "u1", got)
|
||||||
|
case <-time.After(2 * time.Second):
|
||||||
|
t.Fatal("Enqueue did not run the user's pass")
|
||||||
|
}
|
||||||
|
}
|
||||||
+39
-3
@@ -20,6 +20,7 @@ import (
|
|||||||
"net/http"
|
"net/http"
|
||||||
"os"
|
"os"
|
||||||
"os/signal"
|
"os/signal"
|
||||||
|
"sync"
|
||||||
"time"
|
"time"
|
||||||
|
|
||||||
"gitea.d-ma.be/mathias/tapir/internal/adapters/secrets"
|
"gitea.d-ma.be/mathias/tapir/internal/adapters/secrets"
|
||||||
@@ -135,7 +136,8 @@ func cmdRun(ctx context.Context, log *slog.Logger) error {
|
|||||||
|
|
||||||
r := runner.New(engine.Source, st, engine, cfg.UserID, log,
|
r := runner.New(engine.Source, st, engine, cfg.UserID, log,
|
||||||
runner.WithBackoff(cfg.FetchBackoff),
|
runner.WithBackoff(cfg.FetchBackoff),
|
||||||
runner.WithAutoWindow(cfg.AutoSummarizeWindow))
|
runner.WithAutoWindow(cfg.AutoSummarizeWindow),
|
||||||
|
runner.WithCaptionMemory(cfg.ChannelCaptionlessThreshold, cfg.ChannelCaptionlessWindow))
|
||||||
|
|
||||||
log.Info("starting run", "user", cfg.UserID, "model", cfg.SummarizerModel,
|
log.Info("starting run", "user", cfg.UserID, "model", cfg.SummarizerModel,
|
||||||
"gateway", cfg.GatewayURL, "poll_interval", cfg.PollInterval, "fetch_backoff", cfg.FetchBackoff,
|
"gateway", cfg.GatewayURL, "poll_interval", cfg.PollInterval, "fetch_backoff", cfg.FetchBackoff,
|
||||||
@@ -209,7 +211,10 @@ func cmdServe(ctx context.Context, log *slog.Logger) error {
|
|||||||
ClientSecret: cfg.YTClientSecret,
|
ClientSecret: cfg.YTClientSecret,
|
||||||
RedirectURL: cfg.YTConnectRedirectURL,
|
RedirectURL: cfg.YTConnectRedirectURL,
|
||||||
}, secretStore, st, log)
|
}, secretStore, st, log)
|
||||||
log.Info("web youtube connect enabled", "redirect", cfg.YTConnectRedirectURL)
|
// Paste-a-URL (Feature 2): same YouTube credentials, per-user adapter built
|
||||||
|
// per request. Mounting the /paste route keys off app.Fetcher being set.
|
||||||
|
app.Fetcher = videoFetcher{cfg: cfg, secrets: secretStore}
|
||||||
|
log.Info("web youtube connect + paste enabled", "redirect", cfg.YTConnectRedirectURL)
|
||||||
}
|
}
|
||||||
|
|
||||||
// Immediate summarization for the web "Summarize" button. When the engine can
|
// Immediate summarization for the web "Summarize" button. When the engine can
|
||||||
@@ -238,13 +243,44 @@ func cmdServe(ctx context.Context, log *slog.Logger) error {
|
|||||||
if cfg.DiscoveryInterval > 0 {
|
if cfg.DiscoveryInterval > 0 {
|
||||||
log.Info("scheduled discovery enabled", "interval", cfg.DiscoveryInterval, "fetch_rate", cfg.FetchRate)
|
log.Info("scheduled discovery enabled", "interval", cfg.DiscoveryInterval, "fetch_rate", cfg.FetchRate)
|
||||||
log.Warn("scheduled discovery assumes a SINGLE replica — running serve at >1 replica double-runs discovery (ADR-018)")
|
log.Warn("scheduled discovery assumes a SINGLE replica — running serve at >1 replica double-runs discovery (ADR-018)")
|
||||||
runUser := func(ctx context.Context, userID string) (runner.Stats, error) {
|
rawRunUser := func(ctx context.Context, userID string) (runner.Stats, error) {
|
||||||
r, err := buildUserRunner(cfg, st, secretStore, userID, log)
|
r, err := buildUserRunner(cfg, st, secretStore, userID, log)
|
||||||
if err != nil {
|
if err != nil {
|
||||||
return runner.Stats{}, err
|
return runner.Stats{}, err
|
||||||
}
|
}
|
||||||
return r.RunOnce(ctx)
|
return r.RunOnce(ctx)
|
||||||
}
|
}
|
||||||
|
// One lock shared by the scheduler and connect-triggered passes (#6) so
|
||||||
|
// they never fetch concurrently — the single-fetcher invariant (ADR-018).
|
||||||
|
runUser := serialize(&sync.Mutex{}, rawRunUser)
|
||||||
|
// Onboarding burst (Feature 1): after the connect-triggered discovery pass,
|
||||||
|
// summarize up to OnboardSummarizeCount of the user's NEWEST unsummarized
|
||||||
|
// videos so a fresh account gets real summaries in its first session. Hard
|
||||||
|
// cap; explicit, so it bypasses the recency window — but every fetch still
|
||||||
|
// goes through globalFetchGate via the Processor. No-op when disabled
|
||||||
|
// (count 0) or queue-only (no Processor).
|
||||||
|
onboard := func(ctx context.Context, userID string) {
|
||||||
|
if cfg.OnboardSummarizeCount <= 0 || app.Processor == nil {
|
||||||
|
return
|
||||||
|
}
|
||||||
|
ids, err := st.NewestUnsummarizedVideoIDs(ctx, userID, cfg.OnboardSummarizeCount)
|
||||||
|
if err != nil {
|
||||||
|
log.Warn("onboarding: list newest unsummarized", "user", userID, "err", err)
|
||||||
|
return
|
||||||
|
}
|
||||||
|
for _, id := range ids {
|
||||||
|
if err := app.Processor.ProcessVideo(ctx, userID, id); err != nil {
|
||||||
|
log.Warn("onboarding: summarize", "user", userID, "video", id, "err", err)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
if len(ids) > 0 {
|
||||||
|
log.Info("onboarding burst complete", "user", userID, "summarized", len(ids), "cap", cfg.OnboardSummarizeCount)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
if app.Connect != nil {
|
||||||
|
app.Connect.Discovery = &discoveryTrigger{ctx: ctx, run: runUser, onboard: onboard, log: log}
|
||||||
|
log.Info("connect-triggered discovery enabled", "onboard_cap", cfg.OnboardSummarizeCount)
|
||||||
|
}
|
||||||
go runScheduler(ctx, cfg.DiscoveryInterval, st, runUser, log)
|
go runScheduler(ctx, cfg.DiscoveryInterval, st, runUser, log)
|
||||||
} else {
|
} else {
|
||||||
log.Info("scheduled discovery disabled (TAPIR_DISCOVERY_INTERVAL unset or 0)")
|
log.Info("scheduled discovery disabled (TAPIR_DISCOVERY_INTERVAL unset or 0)")
|
||||||
|
|||||||
+63
-8
@@ -3,6 +3,7 @@ package main
|
|||||||
import (
|
import (
|
||||||
"context"
|
"context"
|
||||||
"fmt"
|
"fmt"
|
||||||
|
"strings"
|
||||||
|
|
||||||
"gitea.d-ma.be/mathias/tapir/internal/adapters/llm"
|
"gitea.d-ma.be/mathias/tapir/internal/adapters/llm"
|
||||||
"gitea.d-ma.be/mathias/tapir/internal/adapters/secrets"
|
"gitea.d-ma.be/mathias/tapir/internal/adapters/secrets"
|
||||||
@@ -11,9 +12,29 @@ import (
|
|||||||
"gitea.d-ma.be/mathias/tapir/internal/adapters/youtube"
|
"gitea.d-ma.be/mathias/tapir/internal/adapters/youtube"
|
||||||
"gitea.d-ma.be/mathias/tapir/internal/config"
|
"gitea.d-ma.be/mathias/tapir/internal/config"
|
||||||
"gitea.d-ma.be/mathias/tapir/internal/domain"
|
"gitea.d-ma.be/mathias/tapir/internal/domain"
|
||||||
|
"gitea.d-ma.be/mathias/tapir/internal/ports"
|
||||||
"gitea.d-ma.be/mathias/tapir/internal/usecase"
|
"gitea.d-ma.be/mathias/tapir/internal/usecase"
|
||||||
|
"gitea.d-ma.be/mathias/tapir/internal/web"
|
||||||
)
|
)
|
||||||
|
|
||||||
|
// videoFetcher adapts the YouTube adapter to web.VideoFetcher for the paste flow
|
||||||
|
// (Feature 2). It builds a per-user adapter bound to that user's token ref and
|
||||||
|
// resolves a single video's metadata via the Data API — ungated; only the later
|
||||||
|
// transcript fetch goes through globalFetchGate.
|
||||||
|
type videoFetcher struct {
|
||||||
|
cfg config.Config
|
||||||
|
secrets ports.SecretStore
|
||||||
|
}
|
||||||
|
|
||||||
|
func (f videoFetcher) FetchVideo(ctx context.Context, userID, videoID string) (domain.Video, error) {
|
||||||
|
a := youtube.New(youtube.Config{
|
||||||
|
ClientID: f.cfg.YTClientID,
|
||||||
|
ClientSecret: f.cfg.YTClientSecret,
|
||||||
|
TokenSecretRef: web.YouTubeTokenRef(userID),
|
||||||
|
}, f.secrets)
|
||||||
|
return a.VideoByID(ctx, userID, videoID)
|
||||||
|
}
|
||||||
|
|
||||||
// buildProcessor wires the summarization engine — YouTube source (captions-first),
|
// buildProcessor wires the summarization engine — YouTube source (captions-first),
|
||||||
// AI-router summarizer, store sink — shared by `tapir run` and the web
|
// AI-router summarizer, store sink — shared by `tapir run` and the web
|
||||||
// "Summarize now" path so the wiring lives in one place. It returns (nil, nil) —
|
// "Summarize now" path so the wiring lives in one place. It returns (nil, nil) —
|
||||||
@@ -22,6 +43,40 @@ import (
|
|||||||
// queue-only fallback: the web UI keeps working (the button just queues) and
|
// queue-only fallback: the web UI keeps working (the button just queues) and
|
||||||
// `tapir run` reports the gap via its own ValidateForRun. Missing engine config
|
// `tapir run` reports the gap via its own ValidateForRun. Missing engine config
|
||||||
// is never an error here.
|
// is never an error here.
|
||||||
|
// buildSummarizer wires the summarization endpoint chain (ADR-022) shared by the
|
||||||
|
// web "Summarize now" path and the scheduler's per-user runners. The chain is:
|
||||||
|
// primary (local, fast) → local fallback → cloud fallback (worst case). Each
|
||||||
|
// endpoint reaches the same LiteLLM gateway with a different model alias — the
|
||||||
|
// gateway fronts both llama-swap and berget — so a fallback is just a different
|
||||||
|
// alias, not a second client config. Empty model entries are skipped, so a
|
||||||
|
// client deployment can set the cloud fallback empty to keep content local.
|
||||||
|
func buildSummarizer(cfg config.Config) *summarizer.Summarizer {
|
||||||
|
mk := func(model string) summarizer.Endpoint {
|
||||||
|
return summarizer.Endpoint{
|
||||||
|
Client: llm.New(cfg.GatewayURL, cfg.GatewayKey, model, cfg.SummarizerTimeout, llm.WithMaxTokens(cfg.SummaryMaxTokens)),
|
||||||
|
Provider: providerOf(model),
|
||||||
|
Model: model,
|
||||||
|
}
|
||||||
|
}
|
||||||
|
eps := []summarizer.Endpoint{mk(cfg.SummarizerModel)}
|
||||||
|
if cfg.FallbackModel != "" && cfg.FallbackModel != cfg.SummarizerModel {
|
||||||
|
eps = append(eps, mk(cfg.FallbackModel))
|
||||||
|
}
|
||||||
|
if cfg.CloudFallbackModel != "" && cfg.CloudFallbackModel != cfg.SummarizerModel {
|
||||||
|
eps = append(eps, mk(cfg.CloudFallbackModel))
|
||||||
|
}
|
||||||
|
return summarizer.NewChain(eps, cfg.MaxTranscriptChars)
|
||||||
|
}
|
||||||
|
|
||||||
|
// providerOf maps a model alias to the domain AIProvider recorded on summaries.
|
||||||
|
// A "berget/" alias is an external provider; everything else is the local stack.
|
||||||
|
func providerOf(model string) string {
|
||||||
|
if strings.HasPrefix(model, "berget/") {
|
||||||
|
return "berget"
|
||||||
|
}
|
||||||
|
return "local"
|
||||||
|
}
|
||||||
|
|
||||||
func buildProcessor(cfg config.Config, st *store.Store) (*usecase.Engine, error) {
|
func buildProcessor(cfg config.Config, st *store.Store) (*usecase.Engine, error) {
|
||||||
if cfg.GatewayURL == "" || cfg.YTClientID == "" || cfg.YTClientSecret == "" || cfg.SecretsFile == "" {
|
if cfg.GatewayURL == "" || cfg.YTClientID == "" || cfg.YTClientSecret == "" || cfg.SecretsFile == "" {
|
||||||
return nil, nil
|
return nil, nil
|
||||||
@@ -33,17 +88,17 @@ func buildProcessor(cfg config.Config, st *store.Store) (*usecase.Engine, error)
|
|||||||
ClientSecret: cfg.YTClientSecret,
|
ClientSecret: cfg.YTClientSecret,
|
||||||
TokenSecretRef: cfg.YTTokenRef,
|
TokenSecretRef: cfg.YTTokenRef,
|
||||||
PreferredLanguages: []string{"en"},
|
PreferredLanguages: []string{"en"},
|
||||||
|
MinVideoSeconds: cfg.MinVideoSeconds,
|
||||||
}, secretStore)
|
}, secretStore)
|
||||||
|
|
||||||
// Local Primary only; no BYO fallback for the demo (fallback nil).
|
sum := buildSummarizer(cfg)
|
||||||
primary := summarizer.Endpoint{
|
|
||||||
Client: llm.New(cfg.GatewayURL, cfg.GatewayKey, cfg.SummarizerModel, cfg.SummarizerTimeout),
|
|
||||||
Provider: "local",
|
|
||||||
Model: cfg.SummarizerModel,
|
|
||||||
}
|
|
||||||
sum := summarizer.New(primary, nil)
|
|
||||||
|
|
||||||
return usecase.NewEngine(src, sum, st), nil
|
// The store is both the summary sink and the shared transcript cache (ADR-021):
|
||||||
|
// the engine reads stored transcripts before any caption fetch and writes
|
||||||
|
// resolved ones back, so re-analysis never re-touches YouTube.
|
||||||
|
eng := usecase.NewEngine(src, sum, st)
|
||||||
|
eng.Transcripts = st
|
||||||
|
return eng, nil
|
||||||
}
|
}
|
||||||
|
|
||||||
// engineProcessor adapts the engine (which works in terms of a domain.Video) to
|
// engineProcessor adapts the engine (which works in terms of a domain.Video) to
|
||||||
|
|||||||
+81
-24
@@ -6,9 +6,7 @@ import (
|
|||||||
"log/slog"
|
"log/slog"
|
||||||
"time"
|
"time"
|
||||||
|
|
||||||
"gitea.d-ma.be/mathias/tapir/internal/adapters/llm"
|
|
||||||
"gitea.d-ma.be/mathias/tapir/internal/adapters/store"
|
"gitea.d-ma.be/mathias/tapir/internal/adapters/store"
|
||||||
"gitea.d-ma.be/mathias/tapir/internal/adapters/summarizer"
|
|
||||||
"gitea.d-ma.be/mathias/tapir/internal/adapters/youtube"
|
"gitea.d-ma.be/mathias/tapir/internal/adapters/youtube"
|
||||||
"gitea.d-ma.be/mathias/tapir/internal/config"
|
"gitea.d-ma.be/mathias/tapir/internal/config"
|
||||||
"gitea.d-ma.be/mathias/tapir/internal/ports"
|
"gitea.d-ma.be/mathias/tapir/internal/ports"
|
||||||
@@ -35,24 +33,28 @@ func buildUserRunner(cfg config.Config, st *store.Store, secretStore ports.Secre
|
|||||||
ClientSecret: cfg.YTClientSecret,
|
ClientSecret: cfg.YTClientSecret,
|
||||||
TokenSecretRef: web.YouTubeTokenRef(userID),
|
TokenSecretRef: web.YouTubeTokenRef(userID),
|
||||||
PreferredLanguages: []string{"en"},
|
PreferredLanguages: []string{"en"},
|
||||||
|
MinVideoSeconds: cfg.MinVideoSeconds,
|
||||||
}, secretStore)
|
}, secretStore)
|
||||||
|
|
||||||
primary := summarizer.Endpoint{
|
engine := usecase.NewEngine(src, buildSummarizer(cfg), st)
|
||||||
Client: llm.New(cfg.GatewayURL, cfg.GatewayKey, cfg.SummarizerModel, cfg.SummarizerTimeout),
|
// Share the transcript cache (ADR-021) on the scheduler path too — without
|
||||||
Provider: "local",
|
// this every scheduled pass re-fetches transcripts it already had, burning the
|
||||||
Model: cfg.SummarizerModel,
|
// scarce per-IP caption budget (ADR-014) and starving other users. The web
|
||||||
}
|
// "Summarize now" path already sets this; the scheduler omitting it was a bug.
|
||||||
engine := usecase.NewEngine(src, summarizer.New(primary, nil), st)
|
engine.Transcripts = st
|
||||||
|
|
||||||
return runner.New(src, st, engine, userID, log,
|
return runner.New(src, st, engine, userID, log,
|
||||||
runner.WithBackoff(cfg.FetchBackoff),
|
runner.WithBackoff(cfg.FetchBackoff),
|
||||||
runner.WithAutoWindow(cfg.AutoSummarizeWindow)), nil
|
runner.WithAutoWindow(cfg.AutoSummarizeWindow),
|
||||||
|
runner.WithCaptionMemory(cfg.ChannelCaptionlessThreshold, cfg.ChannelCaptionlessWindow)), nil
|
||||||
}
|
}
|
||||||
|
|
||||||
// userLister enumerates every registered user. *store.Store satisfies it via
|
// userLister enumerates every registered user and reports a user's video
|
||||||
// ListAllUsers. A small local interface keeps the scheduler testable with a fake.
|
// connections. *store.Store satisfies it via ListAllUsers + ConnectionsForUser.
|
||||||
|
// A small local interface keeps the scheduler testable with a fake.
|
||||||
type userLister interface {
|
type userLister interface {
|
||||||
ListAllUsers(ctx context.Context) ([]store.UserIdentity, error)
|
ListAllUsers(ctx context.Context) ([]store.UserIdentity, error)
|
||||||
|
ConnectionsForUser(ctx context.Context, userID string) ([]store.Connection, error)
|
||||||
}
|
}
|
||||||
|
|
||||||
// runDiscoveryPass runs one discovery pass for every user. runUser performs a
|
// runDiscoveryPass runs one discovery pass for every user. runUser performs a
|
||||||
@@ -62,6 +64,7 @@ type userLister interface {
|
|||||||
// isolation). Returns the stats summed across users.
|
// isolation). Returns the stats summed across users.
|
||||||
func runDiscoveryPass(
|
func runDiscoveryPass(
|
||||||
ctx context.Context,
|
ctx context.Context,
|
||||||
|
pass int,
|
||||||
lister userLister,
|
lister userLister,
|
||||||
runUser func(context.Context, string) (runner.Stats, error),
|
runUser func(context.Context, string) (runner.Stats, error),
|
||||||
log *slog.Logger,
|
log *slog.Logger,
|
||||||
@@ -72,9 +75,39 @@ func runDiscoveryPass(
|
|||||||
return runner.Stats{}
|
return runner.Stats{}
|
||||||
}
|
}
|
||||||
|
|
||||||
log.Info("scheduler: starting discovery pass", "users", len(users))
|
// Keep only users with a video connection. A pass for a connectionless user
|
||||||
var total runner.Stats
|
// (e.g. a stale Dex-era orphan identity) only tries to resolve a token that
|
||||||
|
// was never minted, logging a spurious "ref not found" every tick. Filtering
|
||||||
|
// here — BEFORE rotation — also keeps fairness honest: rotation is over the
|
||||||
|
// users that actually consume the caption budget, so a dead identity can't eat
|
||||||
|
// a rotation slot and skew the lead share.
|
||||||
|
var connected []store.UserIdentity
|
||||||
for _, u := range users {
|
for _, u := range users {
|
||||||
|
if ctx.Err() != nil {
|
||||||
|
return runner.Stats{} // shutting down
|
||||||
|
}
|
||||||
|
conns, err := lister.ConnectionsForUser(ctx, u.UserID)
|
||||||
|
if err != nil {
|
||||||
|
log.Warn("scheduler: list connections failed", "user", u.UserID, "err", err)
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
if len(conns) == 0 {
|
||||||
|
log.Debug("scheduler: skipping user with no video connections", "user", u.UserID)
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
connected = append(connected, u)
|
||||||
|
}
|
||||||
|
|
||||||
|
// Rotate who goes first each pass. Caption fetches share one per-egress-IP
|
||||||
|
// rate budget (ADR-014); whoever runs first each pass spends the pre-throttle
|
||||||
|
// window, so a FIXED order permanently starves whoever is last (a new pilot
|
||||||
|
// user got 0 fetches for 12h while the first-listed user got all of them).
|
||||||
|
// Rotation over the connected set gives each real user the lead in turn.
|
||||||
|
connected = rotateUsers(connected, pass)
|
||||||
|
|
||||||
|
log.Info("scheduler: starting discovery pass", "users", len(connected))
|
||||||
|
var total runner.Stats
|
||||||
|
for _, u := range connected {
|
||||||
if ctx.Err() != nil {
|
if ctx.Err() != nil {
|
||||||
break // shutting down: stop enumerating
|
break // shutting down: stop enumerating
|
||||||
}
|
}
|
||||||
@@ -89,6 +122,7 @@ func runDiscoveryPass(
|
|||||||
"skipped_seen", total.SkippedSeen, "skipped_no_text", total.SkippedNoText,
|
"skipped_seen", total.SkippedSeen, "skipped_no_text", total.SkippedNoText,
|
||||||
"skipped_manual", total.SkippedManual, "skipped_too_old", total.SkippedTooOld,
|
"skipped_manual", total.SkippedManual, "skipped_too_old", total.SkippedTooOld,
|
||||||
"skipped_rate_limited", total.SkippedRateLimited,
|
"skipped_rate_limited", total.SkippedRateLimited,
|
||||||
|
"skipped_no_caption_channel", total.SkippedNoCaptionChannel,
|
||||||
"channel_unavailable", total.ChannelUnavailable, "errors", total.Errors)
|
"channel_unavailable", total.ChannelUnavailable, "errors", total.Errors)
|
||||||
return total
|
return total
|
||||||
}
|
}
|
||||||
@@ -109,7 +143,8 @@ func runScheduler(
|
|||||||
return // disabled
|
return // disabled
|
||||||
}
|
}
|
||||||
|
|
||||||
runDiscoveryPass(ctx, lister, runUser, log)
|
pass := 0
|
||||||
|
runDiscoveryPass(ctx, pass, lister, runUser, log)
|
||||||
|
|
||||||
ticker := time.NewTicker(interval)
|
ticker := time.NewTicker(interval)
|
||||||
defer ticker.Stop()
|
defer ticker.Stop()
|
||||||
@@ -118,23 +153,45 @@ func runScheduler(
|
|||||||
case <-ctx.Done():
|
case <-ctx.Done():
|
||||||
return
|
return
|
||||||
case <-ticker.C:
|
case <-ticker.C:
|
||||||
runDiscoveryPass(ctx, lister, runUser, log)
|
pass++
|
||||||
|
runDiscoveryPass(ctx, pass, lister, runUser, log)
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// rotateUsers left-rotates users by pass positions so a different user leads each
|
||||||
|
// pass. With n users, user i leads on every pass where pass ≡ i (mod n). A pass
|
||||||
|
// offset that is negative or exceeds n is normalised. Order within the rotation
|
||||||
|
// is otherwise preserved, so the set of users run is unchanged — only who is
|
||||||
|
// first (and thus wins the scarce caption-fetch budget) rotates.
|
||||||
|
func rotateUsers(users []store.UserIdentity, pass int) []store.UserIdentity {
|
||||||
|
n := len(users)
|
||||||
|
if n <= 1 {
|
||||||
|
return users
|
||||||
|
}
|
||||||
|
off := ((pass % n) + n) % n
|
||||||
|
if off == 0 {
|
||||||
|
return users
|
||||||
|
}
|
||||||
|
out := make([]store.UserIdentity, 0, n)
|
||||||
|
out = append(out, users[off:]...)
|
||||||
|
out = append(out, users[:off]...)
|
||||||
|
return out
|
||||||
|
}
|
||||||
|
|
||||||
// sumStats adds two passes' stats field-wise, so runDiscoveryPass can report a
|
// sumStats adds two passes' stats field-wise, so runDiscoveryPass can report a
|
||||||
// per-tick aggregate across all users.
|
// per-tick aggregate across all users.
|
||||||
func sumStats(a, b runner.Stats) runner.Stats {
|
func sumStats(a, b runner.Stats) runner.Stats {
|
||||||
return runner.Stats{
|
return runner.Stats{
|
||||||
Candidates: a.Candidates + b.Candidates,
|
Candidates: a.Candidates + b.Candidates,
|
||||||
Summarized: a.Summarized + b.Summarized,
|
Summarized: a.Summarized + b.Summarized,
|
||||||
SkippedSeen: a.SkippedSeen + b.SkippedSeen,
|
SkippedSeen: a.SkippedSeen + b.SkippedSeen,
|
||||||
SkippedNoText: a.SkippedNoText + b.SkippedNoText,
|
SkippedNoText: a.SkippedNoText + b.SkippedNoText,
|
||||||
SkippedManual: a.SkippedManual + b.SkippedManual,
|
SkippedManual: a.SkippedManual + b.SkippedManual,
|
||||||
SkippedTooOld: a.SkippedTooOld + b.SkippedTooOld,
|
SkippedTooOld: a.SkippedTooOld + b.SkippedTooOld,
|
||||||
SkippedRateLimited: a.SkippedRateLimited + b.SkippedRateLimited,
|
SkippedRateLimited: a.SkippedRateLimited + b.SkippedRateLimited,
|
||||||
ChannelUnavailable: a.ChannelUnavailable + b.ChannelUnavailable,
|
SkippedNoCaptionChannel: a.SkippedNoCaptionChannel + b.SkippedNoCaptionChannel,
|
||||||
Errors: a.Errors + b.Errors,
|
ChannelUnavailable: a.ChannelUnavailable + b.ChannelUnavailable,
|
||||||
|
Errors: a.Errors + b.Errors,
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -21,19 +21,31 @@ func quietLog() *slog.Logger {
|
|||||||
|
|
||||||
// fakeLister returns a fixed user set (or an error) for the scheduler under test.
|
// fakeLister returns a fixed user set (or an error) for the scheduler under test.
|
||||||
type fakeLister struct {
|
type fakeLister struct {
|
||||||
users []store.UserIdentity
|
users []store.UserIdentity
|
||||||
err error
|
err error
|
||||||
|
noConn map[string]bool // users that have NOT connected a video source
|
||||||
}
|
}
|
||||||
|
|
||||||
func (f fakeLister) ListAllUsers(context.Context) ([]store.UserIdentity, error) {
|
func (f fakeLister) ListAllUsers(context.Context) ([]store.UserIdentity, error) {
|
||||||
return f.users, f.err
|
return f.users, f.err
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// ConnectionsForUser reports a single youtube connection for every user except
|
||||||
|
// those in noConn, which return zero — the connection-less case the scheduler
|
||||||
|
// must skip instead of running (and failing to resolve a token for).
|
||||||
|
func (f fakeLister) ConnectionsForUser(_ context.Context, userID string) ([]store.Connection, error) {
|
||||||
|
if f.noConn[userID] {
|
||||||
|
return nil, nil
|
||||||
|
}
|
||||||
|
return []store.Connection{{Provider: "youtube"}}, nil
|
||||||
|
}
|
||||||
|
|
||||||
// countingRunUser records how many passes each user got, optionally failing for
|
// countingRunUser records how many passes each user got, optionally failing for
|
||||||
// specific users, under a mutex so it is safe across the scheduler goroutine.
|
// specific users, under a mutex so it is safe across the scheduler goroutine.
|
||||||
type countingRunUser struct {
|
type countingRunUser struct {
|
||||||
mu sync.Mutex
|
mu sync.Mutex
|
||||||
calls map[string]int
|
calls map[string]int
|
||||||
|
order []string // userIDs in the order they were run, across all passes
|
||||||
failFor map[string]bool
|
failFor map[string]bool
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -49,12 +61,19 @@ func (c *countingRunUser) run(_ context.Context, userID string) (runner.Stats, e
|
|||||||
c.mu.Lock()
|
c.mu.Lock()
|
||||||
defer c.mu.Unlock()
|
defer c.mu.Unlock()
|
||||||
c.calls[userID]++
|
c.calls[userID]++
|
||||||
|
c.order = append(c.order, userID)
|
||||||
if c.failFor[userID] {
|
if c.failFor[userID] {
|
||||||
return runner.Stats{Errors: 1}, errors.New("boom")
|
return runner.Stats{Errors: 1}, errors.New("boom")
|
||||||
}
|
}
|
||||||
return runner.Stats{Summarized: 1}, nil
|
return runner.Stats{Summarized: 1}, nil
|
||||||
}
|
}
|
||||||
|
|
||||||
|
func (c *countingRunUser) runOrder() []string {
|
||||||
|
c.mu.Lock()
|
||||||
|
defer c.mu.Unlock()
|
||||||
|
return append([]string(nil), c.order...)
|
||||||
|
}
|
||||||
|
|
||||||
func (c *countingRunUser) count(userID string) int {
|
func (c *countingRunUser) count(userID string) int {
|
||||||
c.mu.Lock()
|
c.mu.Lock()
|
||||||
defer c.mu.Unlock()
|
defer c.mu.Unlock()
|
||||||
@@ -83,7 +102,7 @@ func TestDiscoveryPassRunsEveryUserOnce(t *testing.T) {
|
|||||||
lister := fakeLister{users: usersN("a", "b", "c")}
|
lister := fakeLister{users: usersN("a", "b", "c")}
|
||||||
rc := newCountingRunUser()
|
rc := newCountingRunUser()
|
||||||
|
|
||||||
stats := runDiscoveryPass(context.Background(), lister, rc.run, quietLog())
|
stats := runDiscoveryPass(context.Background(), 0, lister, rc.run, quietLog())
|
||||||
|
|
||||||
require.Equal(t, 1, rc.count("a"))
|
require.Equal(t, 1, rc.count("a"))
|
||||||
require.Equal(t, 1, rc.count("b"))
|
require.Equal(t, 1, rc.count("b"))
|
||||||
@@ -91,11 +110,59 @@ func TestDiscoveryPassRunsEveryUserOnce(t *testing.T) {
|
|||||||
require.Equal(t, 3, stats.Summarized, "stats are summed across users")
|
require.Equal(t, 3, stats.Summarized, "stats are summed across users")
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// Caption fetches share one per-IP budget; a fixed user order starves whoever is
|
||||||
|
// last. Each pass must rotate which user leads so the lead slot is shared.
|
||||||
|
func TestDiscoveryPassRotatesLeadUser(t *testing.T) {
|
||||||
|
lister := fakeLister{users: usersN("a", "b", "c")}
|
||||||
|
rc := newCountingRunUser()
|
||||||
|
|
||||||
|
runDiscoveryPass(context.Background(), 0, lister, rc.run, quietLog())
|
||||||
|
runDiscoveryPass(context.Background(), 1, lister, rc.run, quietLog())
|
||||||
|
runDiscoveryPass(context.Background(), 2, lister, rc.run, quietLog())
|
||||||
|
|
||||||
|
require.Equal(t, []string{"a", "b", "c", "b", "c", "a", "c", "a", "b"}, rc.runOrder(),
|
||||||
|
"each pass left-rotates the user order so every user leads in turn")
|
||||||
|
// Fairness: over a full rotation cycle every user ran the same number of times.
|
||||||
|
require.Equal(t, 3, rc.count("a"))
|
||||||
|
require.Equal(t, 3, rc.count("b"))
|
||||||
|
require.Equal(t, 3, rc.count("c"))
|
||||||
|
}
|
||||||
|
|
||||||
|
// A connectionless orphan must not consume a rotation slot: rotation is over the
|
||||||
|
// connected users only, so two real users alternate the lead 50/50 even with a
|
||||||
|
// dead identity listed between them.
|
||||||
|
func TestDiscoveryPassRotationIgnoresConnectionlessUsers(t *testing.T) {
|
||||||
|
lister := fakeLister{users: usersN("a", "orphan", "c"), noConn: map[string]bool{"orphan": true}}
|
||||||
|
rc := newCountingRunUser()
|
||||||
|
|
||||||
|
runDiscoveryPass(context.Background(), 0, lister, rc.run, quietLog())
|
||||||
|
runDiscoveryPass(context.Background(), 1, lister, rc.run, quietLog())
|
||||||
|
|
||||||
|
require.Equal(t, []string{"a", "c", "c", "a"}, rc.runOrder(),
|
||||||
|
"only connected users rotate; the orphan never runs and never holds a slot")
|
||||||
|
require.Equal(t, 0, rc.count("orphan"))
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestDiscoveryPassSkipsUsersWithoutConnections(t *testing.T) {
|
||||||
|
// b never connected a video source (e.g. a stale Dex-era orphan identity).
|
||||||
|
// It must be skipped silently — not run and logged as a token error every pass.
|
||||||
|
lister := fakeLister{users: usersN("a", "b", "c"), noConn: map[string]bool{"b": true}}
|
||||||
|
rc := newCountingRunUser()
|
||||||
|
|
||||||
|
stats := runDiscoveryPass(context.Background(), 0, lister, rc.run, quietLog())
|
||||||
|
|
||||||
|
require.Equal(t, 1, rc.count("a"))
|
||||||
|
require.Equal(t, 0, rc.count("b"), "a user with no connection must be skipped, not run")
|
||||||
|
require.Equal(t, 1, rc.count("c"))
|
||||||
|
require.Equal(t, 2, stats.Summarized, "only connected users contribute")
|
||||||
|
require.Equal(t, 0, stats.Errors, "skipping is silent — no spurious error stat")
|
||||||
|
}
|
||||||
|
|
||||||
func TestDiscoveryPassOneUserFailureDoesNotStopOthers(t *testing.T) {
|
func TestDiscoveryPassOneUserFailureDoesNotStopOthers(t *testing.T) {
|
||||||
lister := fakeLister{users: usersN("a", "b", "c")}
|
lister := fakeLister{users: usersN("a", "b", "c")}
|
||||||
rc := newCountingRunUser("b") // user b's pass errors
|
rc := newCountingRunUser("b") // user b's pass errors
|
||||||
|
|
||||||
stats := runDiscoveryPass(context.Background(), lister, rc.run, quietLog())
|
stats := runDiscoveryPass(context.Background(), 0, lister, rc.run, quietLog())
|
||||||
|
|
||||||
require.Equal(t, 1, rc.count("a"))
|
require.Equal(t, 1, rc.count("a"))
|
||||||
require.Equal(t, 1, rc.count("b"))
|
require.Equal(t, 1, rc.count("b"))
|
||||||
@@ -108,7 +175,7 @@ func TestDiscoveryPassListerErrorIsContained(t *testing.T) {
|
|||||||
lister := fakeLister{err: errors.New("db down")}
|
lister := fakeLister{err: errors.New("db down")}
|
||||||
rc := newCountingRunUser()
|
rc := newCountingRunUser()
|
||||||
|
|
||||||
stats := runDiscoveryPass(context.Background(), lister, rc.run, quietLog())
|
stats := runDiscoveryPass(context.Background(), 0, lister, rc.run, quietLog())
|
||||||
|
|
||||||
require.Equal(t, 0, rc.total(), "no users enumerated → no passes")
|
require.Equal(t, 0, rc.total(), "no users enumerated → no passes")
|
||||||
require.Equal(t, runner.Stats{}, stats)
|
require.Equal(t, runner.Stats{}, stats)
|
||||||
|
|||||||
@@ -147,9 +147,17 @@ graph TB
|
|||||||
rate-limiting — ADR-014).
|
rate-limiting — ADR-014).
|
||||||
- **Summarization mode** — `users.auto_summarize` (migration 006). Default is **true** for new
|
- **Summarization mode** — `users.auto_summarize` (migration 006). Default is **true** for new
|
||||||
users (migration 011, ADR-018); all existing rows were back-filled via migration 012. Auto:
|
users (migration 011, ADR-018); all existing rows were back-filled via migration 012. Auto:
|
||||||
every new video is summarized. Manual: new videos appear unsummarized; the button sets
|
new videos **published within the recency window** (`TAPIR_AUTO_SUMMARIZE_WINDOW`, default ~7d,
|
||||||
`videos.summarize_requested`, which the next `tapir run` processes and clears. Both the click
|
ADR-020) are summarized automatically; older videos are discovered and listed but wait for an
|
||||||
path and the batch `tapir run` drive the same unchanged engine.
|
explicit "Summarize". Manual: new videos appear unsummarized; the button sets
|
||||||
|
`videos.summarize_requested`, which the next `tapir run` processes and clears. A manual request
|
||||||
|
bypasses the recency bound. Both the click path and the batch `tapir run` drive the same
|
||||||
|
unchanged engine.
|
||||||
|
- **List surface (ADR-020)** — the list reads `ListVideos` ordered summarized-first, then
|
||||||
|
`published_at DESC NULLS LAST`. The web layer collapses the noise so summaries are not buried:
|
||||||
|
un-summarized videos older than the recency window fold into one "Show N older videos"
|
||||||
|
disclosure, and caption-less videos collapse to a single count line. Copy surfaces scarcity
|
||||||
|
honestly (queue counts, gradual-fill note) — it never implies the feed is fuller than it is.
|
||||||
|
|
||||||
The engine, ports, and sink adapters are **untouched** by all of the above — the web surface only
|
The engine, ports, and sink adapters are **untouched** by all of the above — the web surface only
|
||||||
reads the store and triggers the existing engine. Adding it changed wiring, not the core (ADR-003).
|
reads the store and triggers the existing engine. Adding it changed wiring, not the core (ADR-003).
|
||||||
@@ -174,6 +182,7 @@ sequenceDiagram
|
|||||||
loop per user
|
loop per user
|
||||||
S->>DB: GetAutoSummarize(userID)
|
S->>DB: GetAutoSummarize(userID)
|
||||||
S->>YT: ListSubscriptions + NewVideos
|
S->>YT: ListSubscriptions + NewVideos
|
||||||
|
Note over S,DB: auto: skip videos published before<br/>TAPIR_AUTO_SUMMARIZE_WINDOW (ADR-020);<br/>older ones listed, await manual request
|
||||||
Note over S,YT: WaitFetchGate(ctx) throttles<br/>all fetches to TAPIR_FETCH_RATE
|
Note over S,YT: WaitFetchGate(ctx) throttles<br/>all fetches to TAPIR_FETCH_RATE
|
||||||
alt transcript available
|
alt transcript available
|
||||||
S->>LLM: Summarize
|
S->>LLM: Summarize
|
||||||
@@ -214,19 +223,22 @@ Both paths share `globalFetchGate` — rate limiting is **respected in both**, n
|
|||||||
|
|
||||||
| Path | Trigger | Order | Rationale |
|
| Path | Trigger | Order | Rationale |
|
||||||
|------|---------|-------|-----------|
|
|------|---------|-------|-----------|
|
||||||
| **Foreground** | User clicks "Summarize now" on any non-summarized card (`POST /v/{id}/retry-now` for rate-limited; `POST /v/{id}/summarize` for pending) | Single chosen video | On-demand value: user picks a specific video to read now |
|
| **Foreground** | User clicks "Summarize" on any non-summarized card (`POST /v/{id}/retry-now` for rate-limited; `POST /v/{id}/summarize` for pending) | Single chosen video | On-demand value: user picks a specific video to read now — bypasses the recency bound |
|
||||||
| **Background batch** | Scheduled discovery pass every `TAPIR_DISCOVERY_INTERVAL` | **Newest-first across all channels** (see below) | Onboarding prioritisation: most recent, relevant videos surface first |
|
| **Background batch** | Scheduled discovery pass every `TAPIR_DISCOVERY_INTERVAL` | **Newest-first across all channels** (see below), **bounded to the recency window** (ADR-020) | Onboarding prioritisation within bounded load: recent videos auto-fill; the older back-catalogue stays on-demand |
|
||||||
|
|
||||||
The rationale for both paths is **onboarding prioritisation** — a new user should get summaries
|
The rationale for both paths is **onboarding prioritisation under an honest, bounded load** — a
|
||||||
of their most recent, relevant videos quickly while the older back-catalogue fills in behind,
|
new user gets summaries of their most recent videos automatically, while the older back-catalogue
|
||||||
all within the honest shared rate limit.
|
is listed but summarised only on demand, so it never re-drives the shared rate gate every cycle.
|
||||||
|
|
||||||
### Newest-first batch ordering (ADR-018)
|
### Newest-first batch ordering (ADR-018)
|
||||||
|
|
||||||
Within each scheduled pass, `RunOnce` uses a three-phase structure:
|
Within each scheduled pass, `RunOnce` uses a three-phase structure:
|
||||||
|
|
||||||
1. **Discover + persist**: walk all channels, `UpsertVideo` every candidate (so it appears in
|
1. **Discover + persist**: walk all channels, `UpsertVideo` every candidate (so it appears in
|
||||||
the list), apply pre-filters (seen/manual/backoff), collect surviving candidates.
|
the list), apply pre-filters (seen/manual/backoff/**recency**), collect surviving candidates.
|
||||||
|
The recency pre-filter (ADR-020) drops auto-mode videos published before
|
||||||
|
`now - TAPIR_AUTO_SUMMARIZE_WINDOW` unless they are explicitly requested; an undated video is
|
||||||
|
never aged out. They remain persisted/listed — only auto-summarisation is skipped.
|
||||||
2. **Sort**: order candidates `published_at DESC, NULLS LAST, discovery_pos ASC`. Videos with
|
2. **Sort**: order candidates `published_at DESC, NULLS LAST, discovery_pos ASC`. Videos with
|
||||||
no publish date (schema 001: nullable) sort after all dated content. The sort is in-memory
|
no publish date (schema 001: nullable) sort after all dated content. The sort is in-memory
|
||||||
(`slices.SortStableFunc`) — at current scale this is fine.
|
(`slices.SortStableFunc`) — at current scale this is fine.
|
||||||
@@ -235,7 +247,8 @@ Within each scheduled pass, `RunOnce` uses a three-phase structure:
|
|||||||
Before (per-channel inline): `[chanA-old, chanA-mid, chanB-new, chanB-null]`
|
Before (per-channel inline): `[chanA-old, chanA-mid, chanB-new, chanB-null]`
|
||||||
After (newest-first): `[chanB-new, chanA-mid, chanA-old, chanB-null]`
|
After (newest-first): `[chanB-new, chanA-mid, chanA-old, chanB-null]`
|
||||||
|
|
||||||
The set of processed videos is identical; only the order within a pass changes.
|
The set of *processed* videos now also excludes auto-mode back-catalogue beyond the recency
|
||||||
|
window (those stay listed, summarised on demand); within the processed set, only order changes.
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
|
|||||||
+25
-16
@@ -10,12 +10,15 @@ only opaque references to them; the secret material lives in ESO/1Password (ADR-
|
|||||||
|
|
||||||
## Design decisions baked into this model
|
## Design decisions baked into this model
|
||||||
|
|
||||||
- **Per-user isolation, not a shared global video table.** The earlier draft proposed a
|
- **Per-user isolation for everything except transcripts.** The earlier draft proposed a
|
||||||
global `videos`/`transcripts` table deduped across tenants. Rejected for Future B: it
|
global `videos`/`transcripts` table deduped across tenants. **Videos** stay per-user and
|
||||||
reintroduces exactly the cross-domain coupling the homelab architecture review is
|
RLS-scoped — a shared video table reintroduces exactly the cross-domain coupling the homelab
|
||||||
removing, and at 1–5 users the cost of occasionally re-summarizing the same video is
|
architecture review is removing. **Transcripts**, however, are now shared (ADR-021): keyed by
|
||||||
trivial compared to the isolation it would cost. Each user's data is self-contained.
|
`(provider, provider_video_id)`, no `user_id`, **not** RLS-scoped. The cost avoided there is
|
||||||
(Revisit only if Future C makes GPU/transcription cost dominate — a new ADR, not a default.)
|
not LLM re-summarization but a rate-gated, reputation-risky caption fetch (ADR-010/014), which
|
||||||
|
is paid per re-fetch regardless of user count — so persisting public caption content once and
|
||||||
|
sharing it strictly beats the coupling it removes. Everything else each user owns is
|
||||||
|
self-contained; `rls_test.go` proves transcripts is the single exception.
|
||||||
- **Secrets by reference only.** Tables hold a `secret_ref` (opaque string/UUID resolved via
|
- **Secrets by reference only.** Tables hold a `secret_ref` (opaque string/UUID resolved via
|
||||||
the `SecretStore` port), never tokens or keys.
|
the `SecretStore` port), never tokens or keys.
|
||||||
- **The brain sink is just a delivery target.** No brain-specific tables. Whether a summary
|
- **The brain sink is just a delivery target.** No brain-specific tables. Whether a summary
|
||||||
@@ -34,7 +37,7 @@ erDiagram
|
|||||||
USER ||--o{ AI_CREDENTIAL : "has (planned)"
|
USER ||--o{ AI_CREDENTIAL : "has (planned)"
|
||||||
VIDEO_CONNECTION ||--o{ SUBSCRIPTION : "exposes (planned)"
|
VIDEO_CONNECTION ||--o{ SUBSCRIPTION : "exposes (planned)"
|
||||||
SUBSCRIPTION ||--o{ VIDEO : "produces (per user)"
|
SUBSCRIPTION ||--o{ VIDEO : "produces (per user)"
|
||||||
VIDEO ||--o| TRANSCRIPT : "has at most one"
|
VIDEO }o--o| TRANSCRIPT : "shares one by (provider, provider_video_id) — not FK (ADR-021)"
|
||||||
VIDEO ||--o| SUMMARY : "has at most one"
|
VIDEO ||--o| SUMMARY : "has at most one"
|
||||||
SUMMARY ||--o{ SINK_DELIVERY : "delivered via"
|
SUMMARY ||--o{ SINK_DELIVERY : "delivered via"
|
||||||
USER ||--o{ CHANNEL_ERROR : "reports unavailable channels"
|
USER ||--o{ CHANNEL_ERROR : "reports unavailable channels"
|
||||||
@@ -92,12 +95,12 @@ erDiagram
|
|||||||
timestamptz rate_limited_at "backoff clock for 429 retries (migration 007)"
|
timestamptz rate_limited_at "backoff clock for 429 retries (migration 007)"
|
||||||
}
|
}
|
||||||
TRANSCRIPT {
|
TRANSCRIPT {
|
||||||
uuid video_id PK_FK
|
text provider PK "part of shared key (ADR-021)"
|
||||||
uuid user_id FK
|
text provider_video_id PK "part of shared key — the cross-user dedup key"
|
||||||
text source "captions | none"
|
text source "captions | none"
|
||||||
text language
|
text language
|
||||||
text content "null when source = none"
|
text content "null when source = none"
|
||||||
timestamptz resolved_at
|
timestamptz fetched_at
|
||||||
}
|
}
|
||||||
SUMMARY {
|
SUMMARY {
|
||||||
uuid id PK
|
uuid id PK
|
||||||
@@ -147,9 +150,10 @@ mechanism.
|
|||||||
|
|
||||||
- **USER** — one row per registered user (Stage 1, ADR-012; no longer single-row). The Tapir-side
|
- **USER** — one row per registered user (Stage 1, ADR-012; no longer single-row). The Tapir-side
|
||||||
profile; the Dex identity is held separately in `USER_IDENTITY`, not on this row. `auto_summarize`
|
profile; the Dex identity is held separately in `USER_IDENTITY`, not on this row. `auto_summarize`
|
||||||
(migration 006) is the per-user mode flag: `TRUE` = auto-summarize every new video. Default is
|
(migration 006) is the per-user mode flag: `TRUE` = auto-summarize new videos **published within
|
||||||
**true** for new users (migration 011, ADR-018); existing rows were back-filled via migration 012
|
the recency window** (`TAPIR_AUTO_SUMMARIZE_WINDOW`, default ~7d, ADR-020); older videos are
|
||||||
with RLS bypass.
|
listed but summarised on demand. Default is **true** for new users (migration 011, ADR-018);
|
||||||
|
existing rows were back-filled via migration 012 with RLS bypass.
|
||||||
- **USER_IDENTITY** (migration 004) — the `dex_subject → user_id` map. `dex_subject` is the PK,
|
- **USER_IDENTITY** (migration 004) — the `dex_subject → user_id` map. `dex_subject` is the PK,
|
||||||
`user_id` a `UNIQUE` FK to `users` with `ON DELETE CASCADE`. This is the bridge resolved at login
|
`user_id` a `UNIQUE` FK to `users` with `ON DELETE CASCADE`. This is the bridge resolved at login
|
||||||
*before* a `user_id` is known, so it is **deliberately not RLS-enabled** (it holds no user data;
|
*before* a `user_id` is known, so it is **deliberately not RLS-enabled** (it holds no user data;
|
||||||
@@ -172,8 +176,12 @@ mechanism.
|
|||||||
`transcript_status` and `rate_limited_at` (migration 007) track caption-fetch outcomes for
|
`transcript_status` and `rate_limited_at` (migration 007) track caption-fetch outcomes for
|
||||||
rate-limit backoff: `NULL` = not attempted; `rate_limited` = 429 seen, skip until
|
rate-limit backoff: `NULL` = not attempted; `rate_limited` = 429 seen, skip until
|
||||||
`NOW() - rate_limited_at > TAPIR_FETCH_BACKOFF`; `fetched` = resolved; `none` = no transcript.
|
`NOW() - rate_limited_at > TAPIR_FETCH_BACKOFF`; `fetched` = resolved; `none` = no transcript.
|
||||||
- **TRANSCRIPT** — at most one per video. `source = none` records "checked, no usable
|
- **TRANSCRIPT** — shared public caption content, one row per `(provider, provider_video_id)`,
|
||||||
transcript" so the watcher doesn't reprocess (ADR-007). `content` null in that case.
|
**not** RLS-scoped and carrying no `user_id` (ADR-021). Two users who watch the same video
|
||||||
|
share the one row; the summarize path reads it before any caption fetch, so re-analysis never
|
||||||
|
re-touches YouTube (ADR-010/014). `source = none` records "checked, no usable transcript" so
|
||||||
|
no one reprocesses (ADR-007); `content` null in that case. A transient 429 is never stored
|
||||||
|
here — it stays a per-user retry via `VIDEO.transcript_status`.
|
||||||
- **SUMMARY** — at most one per video. `fallback_used` + `ai_provider`/`ai_model` make the
|
- **SUMMARY** — at most one per video. `fallback_used` + `ai_provider`/`ai_model` make the
|
||||||
"is local good enough?" question queryable (the Stage 0 quality signal). `highlights`/
|
"is local good enough?" question queryable (the Stage 0 quality signal). `highlights`/
|
||||||
`takeaways` as jsonb to stay schema-flexible while the output format settles.
|
`takeaways` as jsonb to stay schema-flexible while the output format settles.
|
||||||
@@ -230,7 +238,8 @@ queue, doesn't replace it). Deferred until there's a reason.
|
|||||||
|
|
||||||
## Explicitly out of scope (Future C)
|
## Explicitly out of scope (Future C)
|
||||||
|
|
||||||
- Global cross-tenant video/transcript dedup (rejected above).
|
- Global cross-tenant *video* dedup (rejected above). Note: cross-tenant *transcript* sharing
|
||||||
|
is now in scope and shipped (ADR-021); only the videos half stays per-user.
|
||||||
- Sharding / per-tenant physical databases.
|
- Sharding / per-tenant physical databases.
|
||||||
- Soft-delete + full audit trail on connections/credentials (a Stage 2 hardening item; add
|
- Soft-delete + full audit trail on connections/credentials (a Stage 2 hardening item; add
|
||||||
via ADR when Stage 2 work starts).
|
via ADR when Stage 2 work starts).
|
||||||
|
|||||||
@@ -27,6 +27,28 @@ it** — endpoints and aliases drift, and this file is a snapshot (2026-06-06),
|
|||||||
`iguana/deepseek-r1-14b`) is preferred for summary quality if its latency/output is acceptable.
|
`iguana/deepseek-r1-14b`) is preferred for summary quality if its latency/output is acceptable.
|
||||||
The `max_tokens` fix below means thinking models no longer return empty content, so they are now
|
The `max_tokens` fix below means thinking models no longer return empty content, so they are now
|
||||||
viable choices, not blocked ones. Do not assume a coder alias is right for prose.
|
viable choices, not blocked ones. Do not assume a coder alias is right for prose.
|
||||||
|
- **Summarizer fallback chain (ADR-022).** The primary alias is the *first* of an ordered chain;
|
||||||
|
on failure or unparseable output the summarizer advances to the next model. All reached through
|
||||||
|
the same gateway by alias.
|
||||||
|
- `TAPIR_FALLBACK_MODEL` — local fallback. **Default `koala/phi4-14b`.** Empty disables it.
|
||||||
|
- `TAPIR_CLOUD_FALLBACK_MODEL` — worst-case EXTERNAL fallback. **Default `berget/mistral-small`.**
|
||||||
|
**Set this empty (`""`) for any client/NDA deployment** so content never leaves the local
|
||||||
|
stack — the chain then contains only local endpoints.
|
||||||
|
- `TAPIR_SUMMARY_MAX_TOKENS` — per-summary completion budget. **Default `1500`.** Small on
|
||||||
|
purpose: with the old 8192 budget, prompt + completion overflowed `phi4-mini`'s 8k window.
|
||||||
|
- `TAPIR_MAX_TRANSCRIPT_CHARS` — transcript truncation budget sent to the model. **Default
|
||||||
|
`18000`** (~fits an 8k-context model). `0` disables truncation. Prevents the context-overflow
|
||||||
|
HTTP 400 a long transcript caused on `phi4-mini`.
|
||||||
|
- **Discovery low-value filter (ADR-023).** `TAPIR_MIN_VIDEO_SECONDS` — **default `60`**. At
|
||||||
|
discovery, `NewVideos` enriches candidates with one cheap `videos.list` call (quota API, NOT
|
||||||
|
the timedtext 429 path) and drops videos shorter than this plus any live/upcoming broadcast,
|
||||||
|
so the scarce caption-fetch budget isn't spent on Shorts. `0` disables the filter. The
|
||||||
|
paste-a-URL path is never filtered.
|
||||||
|
- **Per-channel caption memory (ADR-024).** `TAPIR_CHANNEL_CAPTIONLESS_THRESHOLD` — **default
|
||||||
|
`5`** consecutive no-caption results before a channel is suppressed (its videos listed but not
|
||||||
|
caption-fetched). `TAPIR_CHANNEL_CAPTIONLESS_WINDOW` — **default `336h`** (14d) suppression
|
||||||
|
before one video is re-probed. `THRESHOLD=0` disables. A successful fetch resets the channel;
|
||||||
|
a 429 does not count; an explicit manual request bypasses suppression.
|
||||||
- **Thinking models need an explicit `max_tokens`.** qwen3 / deepseek-r1 spend the budget on
|
- **Thinking models need an explicit `max_tokens`.** qwen3 / deepseek-r1 spend the budget on
|
||||||
reasoning and return **empty content** if `max_tokens` is too low (or unset). The summarizer's
|
reasoning and return **empty content** if `max_tokens` is too low (or unset). The summarizer's
|
||||||
parser treats an empty summary as an error for exactly this reason. **Done (2026-06-02, Worker F):**
|
parser treats an empty summary as an error for exactly this reason. **Done (2026-06-02, Worker F):**
|
||||||
|
|||||||
@@ -1,5 +1,11 @@
|
|||||||
# Spec — Newest-first batch ordering + honest "Try now" / prioritisation docs
|
# Spec — Newest-first batch ordering + honest "Try now" / prioritisation docs
|
||||||
|
|
||||||
|
> **Extended by ADR-020 (2026-06-08).** This spec covers the *batch processing* order within a
|
||||||
|
> pass. ADR-020 adds (a) a recency pre-filter — auto mode skips videos published before
|
||||||
|
> `TAPIR_AUTO_SUMMARIZE_WINDOW`, listed but summarised on demand — and (b) the same
|
||||||
|
> `published_at DESC NULLS LAST` ordering on the **list read** (`ListVideos`), which previously
|
||||||
|
> sorted by `seen_at`. See `DECISIONS.md` ADR-020.
|
||||||
|
|
||||||
**Repo:** tapir · **Size:** small · **Solo session.**
|
**Repo:** tapir · **Size:** small · **Solo session.**
|
||||||
|
|
||||||
**Why.** Product intent (maintainer, 2026-06-06): a new user should get summaries of their
|
**Why.** Product intent (maintainer, 2026-06-06): a new user should get summaries of their
|
||||||
|
|||||||
@@ -1,5 +1,10 @@
|
|||||||
# Spec — In-process scheduled discovery + auto-summarize + rate-gate finish
|
# Spec — In-process scheduled discovery + auto-summarize + rate-gate finish
|
||||||
|
|
||||||
|
> **Extended by ADR-020 (2026-06-08).** Auto-summarize is no longer "every unseen video": the
|
||||||
|
> scheduler now skips videos published before `TAPIR_AUTO_SUMMARIZE_WINDOW` (default ~7d) unless
|
||||||
|
> explicitly requested, so a back-catalogue does not re-drive the rate gate every cycle. See
|
||||||
|
> `DECISIONS.md` ADR-020.
|
||||||
|
|
||||||
**Repo:** tapir · **Size:** medium · **Solo session** (not a swarm).
|
**Repo:** tapir · **Size:** medium · **Solo session** (not a swarm).
|
||||||
|
|
||||||
**Why this exists.** The Stage-0 gate ("me or a friend returns and reads/acts in ≥2 separate
|
**Why this exists.** The Stage-0 gate ("me or a friend returns and reads/acts in ≥2 separate
|
||||||
|
|||||||
@@ -1,5 +1,12 @@
|
|||||||
# Spec — Unify video-card states: one "Summarize now" verb, honest no-captions state
|
# Spec — Unify video-card states: one "Summarize now" verb, honest no-captions state
|
||||||
|
|
||||||
|
> **Superseded in part by ADR-020 (2026-06-08).** The five card states still hold, but the copy
|
||||||
|
> changed: the nudge verb is now **"Summarize"** (not "Summarize now"), the rate-limited state
|
||||||
|
> reads **"In queue"** (not "Fetching soon…"), and the queued state reads **"summarizing
|
||||||
|
> shortly"** (not "waiting for the next run"). The list also now collapses older un-summarized
|
||||||
|
> and caption-less videos. See `DECISIONS.md` ADR-020 and `views.templ` (`VideoCard`) for the
|
||||||
|
> current copy; this doc is kept as the original design record.
|
||||||
|
|
||||||
**Repo:** tapir · **Size:** small, **view-layer only** (`views.templ` + a little CSS in
|
**Repo:** tapir · **Size:** small, **view-layer only** (`views.templ` + a little CSS in
|
||||||
`view.go`; regenerate `views_templ.go`). No handler, store, or DB change. The two existing
|
`view.go`; regenerate `views_templ.go`). No handler, store, or DB change. The two existing
|
||||||
handlers (`/summarize`, `/retry-now`) stay exactly as they are — only what the card *shows*
|
handlers (`/summarize`, `/retry-now`) stay exactly as they are — only what the card *shows*
|
||||||
|
|||||||
@@ -179,3 +179,4 @@ distinguishable.
|
|||||||
| **"Summarize now" foreground path** | Unified quiet nudge button on actionable non-summarized cards. Five explicit card states — (1) summarized: chip + no button; (2) no captions (`transcript_status = 'none'`): "No transcript available", no button; (3) queued: "Queued" chip, no button; (4) rate-limited: "Fetching soon…" + "Summarize now" → `POST /v/{id}/retry-now` (clears `rate_limited_at`, triggers engine); (5) pending: "Not summarized" + "Summarize now" → `POST /v/{id}/summarize` (queues + triggers engine). One verb, one style (`.btn-quiet`); backend difference invisible to user. Both handlers call `ProcessVideo` through `globalFetchGate`. Rate gate respected, not bypassed — this is onboarding prioritisation. | Fast onboarding value; honest dead-end for no-captions videos (no button that fails). | `internal/web/handlers.go` (`handleRetryNow`, `handleRequestSummarize`); `internal/web/views.templ` (`VideoCard`) |
|
| **"Summarize now" foreground path** | Unified quiet nudge button on actionable non-summarized cards. Five explicit card states — (1) summarized: chip + no button; (2) no captions (`transcript_status = 'none'`): "No transcript available", no button; (3) queued: "Queued" chip, no button; (4) rate-limited: "Fetching soon…" + "Summarize now" → `POST /v/{id}/retry-now` (clears `rate_limited_at`, triggers engine); (5) pending: "Not summarized" + "Summarize now" → `POST /v/{id}/summarize` (queues + triggers engine). One verb, one style (`.btn-quiet`); backend difference invisible to user. Both handlers call `ProcessVideo` through `globalFetchGate`. Rate gate respected, not bypassed — this is onboarding prioritisation. | Fast onboarding value; honest dead-end for no-captions videos (no button that fails). | `internal/web/handlers.go` (`handleRetryNow`, `handleRequestSummarize`); `internal/web/views.templ` (`VideoCard`) |
|
||||||
| **Pipeline stats bar** | A one-line status bar above the video list: `N summarized · M fetching soon · K no captions`. Computed from the unfiltered row set; hidden when all videos are summarized. Gives the user a clear read on pipeline state without any interaction. | Replaces the "why is nothing happening?" confusion when most videos are pending or rate-limited. | `internal/web/view.go` (`PipelineStats`, `pipelineStats`) |
|
| **Pipeline stats bar** | A one-line status bar above the video list: `N summarized · M fetching soon · K no captions`. Computed from the unfiltered row set; hidden when all videos are summarized. Gives the user a clear read on pipeline state without any interaction. | Replaces the "why is nothing happening?" confusion when most videos are pending or rate-limited. | `internal/web/view.go` (`PipelineStats`, `pipelineStats`) |
|
||||||
| **Unavailable channels (account page)** | The `/account` page shows a "Unavailable channels" section when any channels returned HTTP 404 on the last discovery pass. Lists channel name, an "unavailable" badge, and the first-seen date. Data sourced from the `channel_errors` table (migration 013). | Surfaces silent failures so users know why some subscribed channels produce no new videos. | migration 013; `internal/web/account.go`; `internal/adapters/youtube/youtube.go` (`domain.ErrChannelUnavailable`) |
|
| **Unavailable channels (account page)** | The `/account` page shows a "Unavailable channels" section when any channels returned HTTP 404 on the last discovery pass. Lists channel name, an "unavailable" badge, and the first-seen date. Data sourced from the `channel_errors` table (migration 013). | Surfaces silent failures so users know why some subscribed channels produce no new videos. | migration 013; `internal/web/account.go`; `internal/adapters/youtube/youtube.go` (`domain.ErrChannelUnavailable`) |
|
||||||
|
| **Recency window + sparse-state honesty (ADR-020)** | Supersedes the copy/sort in the rows above. Auto-summarize is bounded to videos published within `TAPIR_AUTO_SUMMARIZE_WINDOW` (~7d); older un-summarized videos collapse behind a single "Show N older videos — summarize on demand" disclosure, and caption-less videos collapse to a one-line count (not N cards). List order is now `summarized-first, published_at DESC NULLS LAST`. Copy reframed for honest scarcity: pipeline bar reads "N ready · M in queue · K no captions" (no "fetching soon"); a gradual-fill note explains the rate limit; the nudge verb is "Summarize" (not "Summarize now"); the queued card says "summarizing shortly"; the empty-connected state drops the impossible `tapir run` instruction. Detail leads with Takeaways. Filters slimmed (no date pickers; hidden when empty); watched/skipped segmented; back link on detail; empty terms checkbox removed. | Make the sparse reality legible and honest instead of implying abundance/imminence; bound auto load so the back-catalogue doesn't re-drive the caption gate. Never fetch harder — scarcity is surfaced, not engineered around. | ADR-020; `2384c47`, `3df0459`, `40b703e`, `a1a5217`, `4a0a56e`, `9bf1c31`, `980638d`, `12fb031`, `f775441`, `51aa5d9` |
|
||||||
|
|||||||
@@ -31,5 +31,17 @@ Feature: Local-first AI with optional BYO fallback
|
|||||||
When any transcript is summarized
|
When any transcript is summarized
|
||||||
Then my content is only ever sent to the local AI stack
|
Then my content is only ever sent to the local AI stack
|
||||||
|
|
||||||
# "Reliably" is operationalized as: Primary returned without error within timeout.
|
Scenario: A model returns unparseable output and the next endpoint succeeds
|
||||||
|
Given the local AI stack is available
|
||||||
|
But the primary model returns output that cannot be parsed into a summary
|
||||||
|
And a fallback model is configured
|
||||||
|
When a transcript is summarized
|
||||||
|
Then Tapir falls back to the next model in the chain
|
||||||
|
And the summary records fallback_used as true
|
||||||
|
|
||||||
|
# "Reliably" is operationalized as: an endpoint returned a PARSEABLE summary
|
||||||
|
# within timeout. A 200 with malformed JSON (or highlights emitted as a bare
|
||||||
|
# string) counts as a failure and advances the chain (ADR-022). Endpoints are
|
||||||
|
# tried in order, locals first, so the external worst-case model only ever sees
|
||||||
|
# content after every local endpoint has failed.
|
||||||
# Quality scoring may be added later without changing these scenarios.
|
# Quality scoring may be added later without changing these scenarios.
|
||||||
|
|||||||
@@ -10,12 +10,26 @@ Feature: Connect and manage video accounts
|
|||||||
And my refresh token is stored only as a secret reference
|
And my refresh token is stored only as a secret reference
|
||||||
And my subscriptions are synced
|
And my subscriptions are synced
|
||||||
|
|
||||||
|
@pending
|
||||||
|
# Vimeo connect is not built yet (provider label exists; no connect flow or test).
|
||||||
Scenario: Connect a Vimeo account
|
Scenario: Connect a Vimeo account
|
||||||
Given I have no connected video accounts
|
Given I have no connected video accounts
|
||||||
When I connect my Vimeo account
|
When I connect my Vimeo account
|
||||||
Then the connection is stored with status "active"
|
Then the connection is stored with status "active"
|
||||||
And my subscriptions are synced
|
And my subscriptions are synced
|
||||||
|
|
||||||
|
Scenario: Connecting an account discovers videos immediately
|
||||||
|
Given I have no connected video accounts
|
||||||
|
When I connect my YouTube account
|
||||||
|
Then a discovery pass for my account is triggered right away
|
||||||
|
And I do not have to wait for the next scheduled pass to see my videos
|
||||||
|
|
||||||
|
Scenario: Connecting summarizes my newest videos right away
|
||||||
|
Given I have no connected video accounts
|
||||||
|
When I connect my YouTube account
|
||||||
|
Then up to the onboarding cap of my newest videos are summarized through the rate gate
|
||||||
|
And the rest are left to the scheduled recency-bounded pass
|
||||||
|
|
||||||
Scenario: Tokens are never stored in the clear
|
Scenario: Tokens are never stored in the clear
|
||||||
When I connect any video account
|
When I connect any video account
|
||||||
Then no OAuth token value is stored in the database
|
Then no OAuth token value is stored in the database
|
||||||
@@ -28,6 +42,9 @@ Feature: Connect and manage video accounts
|
|||||||
And no new videos are watched for that connection
|
And no new videos are watched for that connection
|
||||||
And my existing summaries remain readable
|
And my existing summaries remain readable
|
||||||
|
|
||||||
|
@pending
|
||||||
|
# Per-provider BYO credential config is not built as a web flow yet (the summarizer
|
||||||
|
# supports a fallback endpoint, but there is no user-facing BYO setup + its test).
|
||||||
Scenario Outline: BYO AI credential is optional and per-provider
|
Scenario Outline: BYO AI credential is optional and per-provider
|
||||||
When I configure a BYO provider "<provider>"
|
When I configure a BYO provider "<provider>"
|
||||||
Then the credential is stored only as a secret reference
|
Then the credential is stored only as a secret reference
|
||||||
|
|||||||
@@ -19,6 +19,9 @@ Feature: Public landing page
|
|||||||
Then I see a link to my summaries
|
Then I see a link to my summaries
|
||||||
And I see a way to log out
|
And I see a way to log out
|
||||||
|
|
||||||
|
@pending
|
||||||
|
# Behaviour ships (logout redirects to /welcome) but is not unit-tested: logout lives in
|
||||||
|
# the OIDC Auth impl and StubAuth has no routes to exercise it cheaply.
|
||||||
Scenario: Logging out returns to the welcome page
|
Scenario: Logging out returns to the welcome page
|
||||||
Given I am logged in
|
Given I am logged in
|
||||||
When I log out
|
When I log out
|
||||||
|
|||||||
@@ -0,0 +1,30 @@
|
|||||||
|
Feature: Paste a YouTube URL to summarize any video
|
||||||
|
As a user
|
||||||
|
I want to paste a YouTube link and get a summary
|
||||||
|
So that I can pull the specific video I want now, even from channels I don't follow
|
||||||
|
|
||||||
|
Scenario: Paste a valid YouTube URL
|
||||||
|
Given I am connected
|
||||||
|
When I paste a valid YouTube video URL
|
||||||
|
Then the video is added to my feed scoped to me
|
||||||
|
And it is queued for summarization through the shared rate gate
|
||||||
|
|
||||||
|
Scenario: Pasting an invalid link is rejected
|
||||||
|
When I paste something that is not a YouTube video URL
|
||||||
|
Then I get a clear error and nothing is added
|
||||||
|
|
||||||
|
Scenario: Pasting a video that cannot be found is honest
|
||||||
|
When I paste a URL whose video cannot be found
|
||||||
|
Then I am told it couldn't be found and nothing is added
|
||||||
|
|
||||||
|
Scenario: Pasting the same video twice does not duplicate it
|
||||||
|
Given I have pasted a video
|
||||||
|
When I paste the same video again
|
||||||
|
Then my feed still has exactly one entry for it
|
||||||
|
|
||||||
|
@pending
|
||||||
|
# Covered by the engine's ADR-010 no-transcript terminal state (degrade-never-error);
|
||||||
|
# there is no paste-specific test for it.
|
||||||
|
Scenario: A pasted video with no captions resolves honestly
|
||||||
|
When I paste a video that has no captions
|
||||||
|
Then it resolves to the "no transcript available" terminal state
|
||||||
@@ -34,6 +34,9 @@ Feature: Register and manage a multi-user account
|
|||||||
And the other user's data remains intact
|
And the other user's data remains intact
|
||||||
And my Dex identity is left intact
|
And my Dex identity is left intact
|
||||||
|
|
||||||
|
@pending
|
||||||
|
# Re-registration after delete is supported by design (delete leaves the Dex identity,
|
||||||
|
# ADR-013) but has no dedicated end-to-end test yet.
|
||||||
Scenario: A deleted user can register again as a fresh account
|
Scenario: A deleted user can register again as a fresh account
|
||||||
Given I deleted my Tapir account but my Dex identity still exists
|
Given I deleted my Tapir account but my Dex identity still exists
|
||||||
When I sign in again
|
When I sign in again
|
||||||
|
|||||||
@@ -6,15 +6,25 @@ Feature: Choose how new videos get summarized
|
|||||||
Background:
|
Background:
|
||||||
Given I am a registered user with a connected video account
|
Given I am a registered user with a connected video account
|
||||||
|
|
||||||
Scenario: Auto mode summarizes every new video
|
Scenario: Auto mode summarizes recent new videos automatically
|
||||||
Given my summarization mode is "auto"
|
Given my summarization mode is "auto"
|
||||||
When a subscribed channel posts a new video with captions
|
When a subscribed channel posts a new video with captions within the recency window
|
||||||
Then Tapir summarizes it without my asking
|
Then Tapir summarizes it without my asking
|
||||||
And the summary appears in my list
|
And the summary appears in my list
|
||||||
|
|
||||||
Scenario: Manual mode is the default and leaves new videos unsummarized
|
Scenario: Auto mode lists older videos without summarizing them
|
||||||
Given I have not changed my summarization mode
|
Given my summarization mode is "auto"
|
||||||
Then my mode is "manual"
|
When discovery finds a video published before the recency window
|
||||||
|
Then the video appears in my list with no summary
|
||||||
|
And it is not summarized automatically
|
||||||
|
And I can still summarize it on demand with "Summarize"
|
||||||
|
|
||||||
|
Scenario: Automatic is the default for a new user
|
||||||
|
Given I have just registered
|
||||||
|
Then my summarization mode is "auto"
|
||||||
|
|
||||||
|
Scenario: Manual mode leaves new videos unsummarized
|
||||||
|
Given my summarization mode is "manual"
|
||||||
When a subscribed channel posts a new video with captions
|
When a subscribed channel posts a new video with captions
|
||||||
Then the video appears in my list with no summary
|
Then the video appears in my list with no summary
|
||||||
And nothing is summarized until I request it
|
And nothing is summarized until I request it
|
||||||
@@ -30,3 +40,8 @@ Feature: Choose how new videos get summarized
|
|||||||
# auto_summarize is a per-user setting and summarize_requested is a per-video queue
|
# auto_summarize is a per-user setting and summarize_requested is a per-video queue
|
||||||
# flag (migration 006). The web button sets the flag; `tapir run` processes both the
|
# flag (migration 006). The web button sets the flag; `tapir run` processes both the
|
||||||
# auto videos and the manually queued ones, then clears the flag.
|
# auto videos and the manually queued ones, then clears the flag.
|
||||||
|
#
|
||||||
|
# Recency bound (ADR-020): in auto mode the scheduler only summarizes videos published
|
||||||
|
# within TAPIR_AUTO_SUMMARIZE_WINDOW (default ~7d); older videos are discovered and
|
||||||
|
# listed but wait for an explicit "Summarize" — so a back-catalogue does not re-drive
|
||||||
|
# the per-IP caption gate (ADR-014) every cycle. A manual request bypasses the bound.
|
||||||
|
|||||||
@@ -31,5 +31,16 @@ Feature: Summarize new videos from subscribed channels
|
|||||||
When the watcher sees "Designing for Attention" again
|
When the watcher sees "Designing for Attention" again
|
||||||
Then Tapir does not produce a second summary for it
|
Then Tapir does not produce a second summary for it
|
||||||
|
|
||||||
|
Scenario: Re-analyzing a stored video does not re-fetch its transcript
|
||||||
|
Given a transcript for "Designing for Attention" is already stored
|
||||||
|
When the video is summarized again
|
||||||
|
Then Tapir reads the stored transcript
|
||||||
|
And Tapir does not fetch captions from YouTube
|
||||||
|
|
||||||
# Captions-first is the core path (ADR-007). Audio-download + speech-to-text is
|
# Captions-first is the core path (ADR-007). Audio-download + speech-to-text is
|
||||||
# deferred and intentionally has no scenario here yet.
|
# deferred and intentionally has no scenario here yet.
|
||||||
|
#
|
||||||
|
# Transcript persistence (ADR-021): the stored transcript is shared, keyed by
|
||||||
|
# (provider, provider_video_id) and read before any caption fetch, so the
|
||||||
|
# re-analysis scenario above also covers paste-a-URL and the onboarding burst —
|
||||||
|
# both summarize through the same engine chokepoint.
|
||||||
|
|||||||
@@ -34,15 +34,35 @@ type Client struct {
|
|||||||
httpClient *http.Client
|
httpClient *http.Client
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// Option configures a Client at construction. Variadic so the existing 4-arg
|
||||||
|
// call sites stay valid as new knobs are added.
|
||||||
|
type Option func(*Client)
|
||||||
|
|
||||||
|
// WithMaxTokens overrides the per-request completion budget. The summarizer uses
|
||||||
|
// this to cap completion for small-context models (e.g. koala/phi4-mini, 8k):
|
||||||
|
// with the default 8192 budget, prompt + max_tokens overflows an 8k context and
|
||||||
|
// the gateway returns HTTP 400. A non-positive n is ignored (keeps the default).
|
||||||
|
func WithMaxTokens(n int) Option {
|
||||||
|
return func(c *Client) {
|
||||||
|
if n > 0 {
|
||||||
|
c.maxTokens = n
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
// New constructs a Client.
|
// New constructs a Client.
|
||||||
func New(baseURL, apiKey, model string, timeout time.Duration) *Client {
|
func New(baseURL, apiKey, model string, timeout time.Duration, opts ...Option) *Client {
|
||||||
return &Client{
|
c := &Client{
|
||||||
baseURL: strings.TrimRight(baseURL, "/"),
|
baseURL: strings.TrimRight(baseURL, "/"),
|
||||||
apiKey: apiKey,
|
apiKey: apiKey,
|
||||||
model: model,
|
model: model,
|
||||||
maxTokens: defaultMaxTokens,
|
maxTokens: defaultMaxTokens,
|
||||||
httpClient: &http.Client{Timeout: timeout},
|
httpClient: &http.Client{Timeout: timeout},
|
||||||
}
|
}
|
||||||
|
for _, opt := range opts {
|
||||||
|
opt(c)
|
||||||
|
}
|
||||||
|
return c
|
||||||
}
|
}
|
||||||
|
|
||||||
type chatRequest struct {
|
type chatRequest struct {
|
||||||
|
|||||||
@@ -64,6 +64,27 @@ func TestClient_SendsMaxTokens(t *testing.T) {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// TestClient_WithMaxTokens overrides the completion budget — the summarizer caps
|
||||||
|
// it small so prompt + max_tokens fits a small-context model's window (8k).
|
||||||
|
func TestClient_WithMaxTokens(t *testing.T) {
|
||||||
|
var body chatRequest
|
||||||
|
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
|
||||||
|
_ = json.NewDecoder(r.Body).Decode(&body)
|
||||||
|
_ = json.NewEncoder(w).Encode(map[string]any{
|
||||||
|
"choices": []map[string]any{{"message": map[string]any{"content": "ok"}}},
|
||||||
|
})
|
||||||
|
}))
|
||||||
|
defer srv.Close()
|
||||||
|
|
||||||
|
c := New(srv.URL, "", "test-model", 10*time.Second, WithMaxTokens(1500))
|
||||||
|
if _, err := c.Complete(context.Background(), "sys", "user"); err != nil {
|
||||||
|
t.Fatalf("Complete: %v", err)
|
||||||
|
}
|
||||||
|
if body.MaxTokens != 1500 {
|
||||||
|
t.Errorf("max_tokens = %d, want 1500", body.MaxTokens)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
func TestClient_ReturnsErrorOnNon200(t *testing.T) {
|
func TestClient_ReturnsErrorOnNon200(t *testing.T) {
|
||||||
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
|
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
|
||||||
http.Error(w, "overloaded", http.StatusServiceUnavailable)
|
http.Error(w, "overloaded", http.StatusServiceUnavailable)
|
||||||
|
|||||||
@@ -10,10 +10,13 @@ import (
|
|||||||
// DeleteUser permanently removes a user and all of their data. It runs through
|
// DeleteUser permanently removes a user and all of their data. It runs through
|
||||||
// withUser so RLS confines every statement to the calling user's own rows.
|
// withUser so RLS confines every statement to the calling user's own rows.
|
||||||
//
|
//
|
||||||
// Deleting the users row cascades (ON DELETE CASCADE) to videos, transcripts,
|
// Deleting the users row cascades (ON DELETE CASCADE) to videos, summaries
|
||||||
// summaries (→ sink_deliveries), video_connections, and the user_identities map
|
// (→ sink_deliveries), video_connections, and the user_identities map —
|
||||||
// — referential-integrity cascades bypass RLS, so a user's child rows are removed
|
// referential-integrity cascades bypass RLS, so a user's child rows are removed
|
||||||
// even though the deleting connection is scoped. summary_actions and login_events
|
// even though the deleting connection is scoped. Transcripts are NOT removed:
|
||||||
|
// since ADR-021 they are shared public content keyed by (provider,
|
||||||
|
// provider_video_id) with no user_id, so another user may still reference the
|
||||||
|
// same row — a user deletion must not strip shared caption content. summary_actions and login_events
|
||||||
// are the exceptions: each carries a user_id but has NO foreign key to users
|
// are the exceptions: each carries a user_id but has NO foreign key to users
|
||||||
// (migrations 002 and 010), so the cascade does not reach them; they are deleted
|
// (migrations 002 and 010), so the cascade does not reach them; they are deleted
|
||||||
// explicitly in the same scoped transaction. Deleting an absent user is a no-op
|
// explicitly in the same scoped transaction. Deleting an absent user is a no-op
|
||||||
|
|||||||
@@ -0,0 +1,83 @@
|
|||||||
|
package store
|
||||||
|
|
||||||
|
import (
|
||||||
|
"context"
|
||||||
|
"fmt"
|
||||||
|
"time"
|
||||||
|
|
||||||
|
"github.com/jackc/pgx/v5"
|
||||||
|
)
|
||||||
|
|
||||||
|
// CaptionlessChannels returns the set of channel ids currently suppressed for the
|
||||||
|
// user — channels whose recent videos all yielded no captions, within their
|
||||||
|
// suppression window (ADR-024). The runner skips caption fetches for these
|
||||||
|
// channels' videos. A channel whose window has expired is not returned, so its
|
||||||
|
// next video is re-probed (auto-recovery).
|
||||||
|
func (s *Store) CaptionlessChannels(ctx context.Context, userID string) (map[string]bool, error) {
|
||||||
|
out := map[string]bool{}
|
||||||
|
err := s.withUser(ctx, userID, func(tx pgx.Tx) error {
|
||||||
|
rows, err := tx.Query(ctx, `
|
||||||
|
SELECT channel_id FROM channel_caption_state
|
||||||
|
WHERE user_id = $1 AND captionless_until IS NOT NULL AND captionless_until > now()`,
|
||||||
|
userID)
|
||||||
|
if err != nil {
|
||||||
|
return fmt.Errorf("store: caption-less channels: %w", err)
|
||||||
|
}
|
||||||
|
defer rows.Close()
|
||||||
|
for rows.Next() {
|
||||||
|
var ch string
|
||||||
|
if err := rows.Scan(&ch); err != nil {
|
||||||
|
return fmt.Errorf("store: scan caption-less channel: %w", err)
|
||||||
|
}
|
||||||
|
out[ch] = true
|
||||||
|
}
|
||||||
|
return rows.Err()
|
||||||
|
})
|
||||||
|
return out, err
|
||||||
|
}
|
||||||
|
|
||||||
|
// RecordChannelCaptionOutcome updates a channel's caption-availability memory
|
||||||
|
// after a fetch attempt (ADR-024). hadCaptions resets the channel (consecutive
|
||||||
|
// count to 0, suppression cleared). Otherwise the consecutive no-caption count is
|
||||||
|
// incremented; once it reaches threshold the channel is suppressed for window.
|
||||||
|
// threshold <= 0 is a no-op (feature disabled). An empty channelID is ignored
|
||||||
|
// (some sources may not carry one).
|
||||||
|
func (s *Store) RecordChannelCaptionOutcome(ctx context.Context, userID, channelID string, hadCaptions bool, threshold int, window time.Duration) error {
|
||||||
|
if channelID == "" || threshold <= 0 {
|
||||||
|
return nil
|
||||||
|
}
|
||||||
|
return s.withUser(ctx, userID, func(tx pgx.Tx) error {
|
||||||
|
if hadCaptions {
|
||||||
|
_, err := tx.Exec(ctx, `
|
||||||
|
INSERT INTO channel_caption_state (user_id, channel_id, consecutive_none, captionless_until, updated_at)
|
||||||
|
VALUES ($1, $2, 0, NULL, now())
|
||||||
|
ON CONFLICT (user_id, channel_id)
|
||||||
|
DO UPDATE SET consecutive_none = 0, captionless_until = NULL, updated_at = now()`,
|
||||||
|
userID, channelID)
|
||||||
|
if err != nil {
|
||||||
|
return fmt.Errorf("store: reset channel caption state: %w", err)
|
||||||
|
}
|
||||||
|
return nil
|
||||||
|
}
|
||||||
|
// No captions: increment the streak; suppress once it reaches threshold.
|
||||||
|
// captionless_until is set from the NEW count inside the same statement so
|
||||||
|
// the decision is atomic with the increment.
|
||||||
|
until := time.Now().Add(window)
|
||||||
|
_, err := tx.Exec(ctx, `
|
||||||
|
INSERT INTO channel_caption_state (user_id, channel_id, consecutive_none, captionless_until, updated_at)
|
||||||
|
VALUES ($1, $2, 1, CASE WHEN 1 >= $3 THEN $4::timestamptz ELSE NULL END, now())
|
||||||
|
ON CONFLICT (user_id, channel_id)
|
||||||
|
DO UPDATE SET
|
||||||
|
consecutive_none = channel_caption_state.consecutive_none + 1,
|
||||||
|
captionless_until = CASE
|
||||||
|
WHEN channel_caption_state.consecutive_none + 1 >= $3 THEN $4::timestamptz
|
||||||
|
ELSE channel_caption_state.captionless_until
|
||||||
|
END,
|
||||||
|
updated_at = now()`,
|
||||||
|
userID, channelID, threshold, until)
|
||||||
|
if err != nil {
|
||||||
|
return fmt.Errorf("store: record channel no-caption: %w", err)
|
||||||
|
}
|
||||||
|
return nil
|
||||||
|
})
|
||||||
|
}
|
||||||
@@ -0,0 +1,67 @@
|
|||||||
|
package store_test
|
||||||
|
|
||||||
|
import (
|
||||||
|
"context"
|
||||||
|
"testing"
|
||||||
|
"time"
|
||||||
|
|
||||||
|
"github.com/stretchr/testify/require"
|
||||||
|
)
|
||||||
|
|
||||||
|
func TestChannelCaptionMemory_SuppressesAfterThreshold(t *testing.T) {
|
||||||
|
ctx := context.Background()
|
||||||
|
s := newStore(t)
|
||||||
|
super := rawPool(t)
|
||||||
|
resetDB(t, super)
|
||||||
|
seedUser(t, super, userA)
|
||||||
|
|
||||||
|
const threshold = 3
|
||||||
|
window := time.Hour
|
||||||
|
|
||||||
|
// Below threshold: not yet suppressed.
|
||||||
|
require.NoError(t, s.RecordChannelCaptionOutcome(ctx, userA, "chanX", false, threshold, window))
|
||||||
|
require.NoError(t, s.RecordChannelCaptionOutcome(ctx, userA, "chanX", false, threshold, window))
|
||||||
|
got, err := s.CaptionlessChannels(ctx, userA)
|
||||||
|
require.NoError(t, err)
|
||||||
|
require.NotContains(t, got, "chanX", "2 < threshold 3: not suppressed yet")
|
||||||
|
|
||||||
|
// Crossing the threshold suppresses the channel.
|
||||||
|
require.NoError(t, s.RecordChannelCaptionOutcome(ctx, userA, "chanX", false, threshold, window))
|
||||||
|
got, err = s.CaptionlessChannels(ctx, userA)
|
||||||
|
require.NoError(t, err)
|
||||||
|
require.Contains(t, got, "chanX", "3 consecutive no-caption results suppress the channel")
|
||||||
|
|
||||||
|
// A successful caption fetch resets it.
|
||||||
|
require.NoError(t, s.RecordChannelCaptionOutcome(ctx, userA, "chanX", true, threshold, window))
|
||||||
|
got, err = s.CaptionlessChannels(ctx, userA)
|
||||||
|
require.NoError(t, err)
|
||||||
|
require.NotContains(t, got, "chanX", "a captioned video clears suppression")
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestChannelCaptionMemory_WindowExpiryReProbes(t *testing.T) {
|
||||||
|
ctx := context.Background()
|
||||||
|
s := newStore(t)
|
||||||
|
super := rawPool(t)
|
||||||
|
resetDB(t, super)
|
||||||
|
seedUser(t, super, userA)
|
||||||
|
|
||||||
|
// A negative window means captionless_until lands in the past — modelling an
|
||||||
|
// elapsed suppression window, which must make the channel eligible again.
|
||||||
|
require.NoError(t, s.RecordChannelCaptionOutcome(ctx, userA, "chanY", false, 1, -time.Hour))
|
||||||
|
got, err := s.CaptionlessChannels(ctx, userA)
|
||||||
|
require.NoError(t, err)
|
||||||
|
require.NotContains(t, got, "chanY", "an expired window re-enables the channel for a re-probe")
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestChannelCaptionMemory_DisabledThresholdIsNoOp(t *testing.T) {
|
||||||
|
ctx := context.Background()
|
||||||
|
s := newStore(t)
|
||||||
|
super := rawPool(t)
|
||||||
|
resetDB(t, super)
|
||||||
|
seedUser(t, super, userA)
|
||||||
|
|
||||||
|
require.NoError(t, s.RecordChannelCaptionOutcome(ctx, userA, "chanZ", false, 0, time.Hour))
|
||||||
|
got, err := s.CaptionlessChannels(ctx, userA)
|
||||||
|
require.NoError(t, err)
|
||||||
|
require.Empty(t, got, "threshold 0 disables the memory — nothing recorded")
|
||||||
|
}
|
||||||
@@ -53,7 +53,13 @@ func TestMigration010LoginEventsUpDown(t *testing.T) {
|
|||||||
require.True(t, loginEventsExists(t), "login_events must exist at latest migration")
|
require.True(t, loginEventsExists(t), "login_events must exist at latest migration")
|
||||||
|
|
||||||
m := fileMigrator(t)
|
m := fileMigrator(t)
|
||||||
// 011, 012, 013 sit above 010; step them down first so 010 is exercised in isolation.
|
// 011..016 sit above 010; step them down first so 010 is exercised in isolation.
|
||||||
|
require.NoError(t, m.Steps(-1), "down 016 drops channel_caption_state, login_events intact")
|
||||||
|
require.True(t, loginEventsExists(t), "016 down leaves login_events intact")
|
||||||
|
require.NoError(t, m.Steps(-1), "down 015 reshapes transcripts, login_events intact")
|
||||||
|
require.True(t, loginEventsExists(t), "015 down leaves login_events intact")
|
||||||
|
require.NoError(t, m.Steps(-1), "down 014 drops channel_title, login_events intact")
|
||||||
|
require.True(t, loginEventsExists(t), "014 down leaves login_events intact")
|
||||||
require.NoError(t, m.Steps(-1), "down 013 drops channel_errors, login_events intact")
|
require.NoError(t, m.Steps(-1), "down 013 drops channel_errors, login_events intact")
|
||||||
require.True(t, loginEventsExists(t), "013 down leaves login_events intact")
|
require.True(t, loginEventsExists(t), "013 down leaves login_events intact")
|
||||||
require.NoError(t, m.Steps(-1), "down 012 is a no-op, login_events intact")
|
require.NoError(t, m.Steps(-1), "down 012 is a no-op, login_events intact")
|
||||||
@@ -64,7 +70,7 @@ func TestMigration010LoginEventsUpDown(t *testing.T) {
|
|||||||
require.NoError(t, m.Steps(-1), "down 010 must drop login_events")
|
require.NoError(t, m.Steps(-1), "down 010 must drop login_events")
|
||||||
require.False(t, loginEventsExists(t), "login_events must be gone after the down migration")
|
require.False(t, loginEventsExists(t), "login_events must be gone after the down migration")
|
||||||
|
|
||||||
require.NoError(t, m.Steps(4), "up must recreate 010 then re-apply 011, 012, 013")
|
require.NoError(t, m.Steps(7), "up must recreate 010 then re-apply 011..016")
|
||||||
require.True(t, loginEventsExists(t), "login_events must be restored after the up migration")
|
require.True(t, loginEventsExists(t), "login_events must be restored after the up migration")
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -87,6 +93,9 @@ func TestMigration011AutoSummarizeDefaultUpDown(t *testing.T) {
|
|||||||
require.Equal(t, "true", autoSummarizeDefault(t), "011 sets the default to TRUE")
|
require.Equal(t, "true", autoSummarizeDefault(t), "011 sets the default to TRUE")
|
||||||
|
|
||||||
m := fileMigrator(t)
|
m := fileMigrator(t)
|
||||||
|
require.NoError(t, m.Steps(-1), "down 016 drops channel_caption_state")
|
||||||
|
require.NoError(t, m.Steps(-1), "down 015 reshapes transcripts")
|
||||||
|
require.NoError(t, m.Steps(-1), "down 014 drops channel_title")
|
||||||
require.NoError(t, m.Steps(-1), "down 013 drops channel_errors")
|
require.NoError(t, m.Steps(-1), "down 013 drops channel_errors")
|
||||||
require.NoError(t, m.Steps(-1), "down 012 is a no-op")
|
require.NoError(t, m.Steps(-1), "down 012 is a no-op")
|
||||||
require.NoError(t, m.Steps(-1), "down 011 reverts the column default")
|
require.NoError(t, m.Steps(-1), "down 011 reverts the column default")
|
||||||
@@ -96,6 +105,39 @@ func TestMigration011AutoSummarizeDefaultUpDown(t *testing.T) {
|
|||||||
require.Equal(t, "true", autoSummarizeDefault(t))
|
require.Equal(t, "true", autoSummarizeDefault(t))
|
||||||
require.NoError(t, m.Steps(1), "up 012 runs clean (no FORCE RLS on fresh schema)")
|
require.NoError(t, m.Steps(1), "up 012 runs clean (no FORCE RLS on fresh schema)")
|
||||||
require.NoError(t, m.Steps(1), "up 013 creates channel_errors")
|
require.NoError(t, m.Steps(1), "up 013 creates channel_errors")
|
||||||
|
require.NoError(t, m.Steps(1), "up 014 recreates channel_title")
|
||||||
|
require.NoError(t, m.Steps(1), "up 015 reshapes transcripts to shared")
|
||||||
|
require.NoError(t, m.Steps(1), "up 016 recreates channel_caption_state (HEAD)")
|
||||||
|
}
|
||||||
|
|
||||||
|
// channelTitleExists reports whether videos.channel_title is present.
|
||||||
|
func channelTitleExists(t *testing.T) bool {
|
||||||
|
t.Helper()
|
||||||
|
var exists bool
|
||||||
|
require.NoError(t, rawPool(t).QueryRow(context.Background(),
|
||||||
|
`SELECT EXISTS (SELECT 1 FROM information_schema.columns
|
||||||
|
WHERE table_name = 'videos' AND column_name = 'channel_title')`).Scan(&exists))
|
||||||
|
return exists
|
||||||
|
}
|
||||||
|
|
||||||
|
// TestMigration014VideoChannelTitleUpDown proves 014 is reversible: down drops
|
||||||
|
// videos.channel_title, up recreates it.
|
||||||
|
func TestMigration014VideoChannelTitleUpDown(t *testing.T) {
|
||||||
|
newStore(t) // latest (014 applied)
|
||||||
|
require.True(t, channelTitleExists(t), "channel_title exists at latest migration")
|
||||||
|
|
||||||
|
m := fileMigrator(t)
|
||||||
|
require.NoError(t, m.Steps(-1), "down 016 drops channel_caption_state, channel_title intact")
|
||||||
|
require.True(t, channelTitleExists(t), "016 down leaves channel_title intact")
|
||||||
|
require.NoError(t, m.Steps(-1), "down 015 reshapes transcripts, channel_title intact")
|
||||||
|
require.True(t, channelTitleExists(t), "015 down leaves channel_title intact")
|
||||||
|
require.NoError(t, m.Steps(-1), "down 014 must drop channel_title")
|
||||||
|
require.False(t, channelTitleExists(t), "channel_title must be gone after the down migration")
|
||||||
|
|
||||||
|
require.NoError(t, m.Steps(1), "up 014 must recreate channel_title")
|
||||||
|
require.True(t, channelTitleExists(t), "channel_title must be restored after the up migration")
|
||||||
|
require.NoError(t, m.Steps(1), "up 015 restores the shared transcripts shape")
|
||||||
|
require.NoError(t, m.Steps(1), "up 016 recreates channel_caption_state (HEAD)")
|
||||||
}
|
}
|
||||||
|
|
||||||
// TestMigration012FixAutoSummarizeRLS proves 012 runs cleanly and flips any
|
// TestMigration012FixAutoSummarizeRLS proves 012 runs cleanly and flips any
|
||||||
|
|||||||
@@ -0,0 +1 @@
|
|||||||
|
ALTER TABLE videos DROP COLUMN channel_title;
|
||||||
@@ -0,0 +1,8 @@
|
|||||||
|
-- Store the source channel's title per video so the list can offer a real
|
||||||
|
-- channel filter (multi-select of the user's channels) instead of the dead
|
||||||
|
-- free-text field that only ever matched the provider string. Nullable: existing
|
||||||
|
-- rows backfill on the next discovery pass (UpsertVideo writes it); pasted videos
|
||||||
|
-- get it immediately from videos.list. No FK to a channels table at Stage 0 — the
|
||||||
|
-- title is a denormalised display/filter value, consistent with the existing
|
||||||
|
-- subscription_id-stays-NULL stance (data-model.md).
|
||||||
|
ALTER TABLE videos ADD COLUMN channel_title TEXT;
|
||||||
@@ -0,0 +1,19 @@
|
|||||||
|
-- Down 015: restore the per-user RLS-scoped transcripts shape (001 + 003).
|
||||||
|
DROP TABLE transcripts;
|
||||||
|
|
||||||
|
CREATE TABLE transcripts (
|
||||||
|
video_id UUID PRIMARY KEY REFERENCES videos(id) ON DELETE CASCADE,
|
||||||
|
user_id UUID NOT NULL REFERENCES users(id) ON DELETE CASCADE,
|
||||||
|
source TEXT NOT NULL,
|
||||||
|
language TEXT,
|
||||||
|
content TEXT,
|
||||||
|
resolved_at TIMESTAMPTZ NOT NULL DEFAULT NOW()
|
||||||
|
);
|
||||||
|
|
||||||
|
CREATE INDEX idx_transcripts_user_id ON transcripts(user_id);
|
||||||
|
|
||||||
|
ALTER TABLE transcripts ENABLE ROW LEVEL SECURITY;
|
||||||
|
ALTER TABLE transcripts FORCE ROW LEVEL SECURITY;
|
||||||
|
CREATE POLICY transcripts_isolation ON transcripts
|
||||||
|
FOR ALL
|
||||||
|
USING (user_id = current_setting('tapir.current_user_id', true)::uuid);
|
||||||
@@ -0,0 +1,30 @@
|
|||||||
|
-- Migration 015: transcripts become SHARED public-content storage (ADR-021).
|
||||||
|
--
|
||||||
|
-- The per-user transcripts table from 001 (PK videos.id, user_id NOT NULL, RLS
|
||||||
|
-- FORCEd in 003) was dead: no application code ever read or wrote it — only the
|
||||||
|
-- transcript_status columns on `videos` (007) carried fetch outcomes. ADR-021
|
||||||
|
-- repurposes it as the single shared store of public caption content, keyed by
|
||||||
|
-- the cross-user dedup key (provider, provider_video_id) — the video's public
|
||||||
|
-- identity, not Tapir's per-user videos.id — so re-analysis never re-fetches
|
||||||
|
-- from YouTube (ADR-010/014).
|
||||||
|
--
|
||||||
|
-- It holds ONLY public caption content + the video's public id (nothing
|
||||||
|
-- user-identifying), so it is deliberately NOT RLS-scoped: no user_id, no
|
||||||
|
-- policy, no FORCE. This is the single, intentional exception to the ADR-012
|
||||||
|
-- isolation boundary; rls_test.go asserts the boundary is exactly here and
|
||||||
|
-- nowhere else. Dropping the old table drops its RLS policy with it; it held no
|
||||||
|
-- real data, so drop+recreate loses nothing.
|
||||||
|
DROP TABLE transcripts;
|
||||||
|
|
||||||
|
CREATE TABLE transcripts (
|
||||||
|
provider TEXT NOT NULL,
|
||||||
|
provider_video_id TEXT NOT NULL,
|
||||||
|
source TEXT NOT NULL, -- 'captions' (content set) | 'none' (no captions; content NULL)
|
||||||
|
language TEXT,
|
||||||
|
content TEXT,
|
||||||
|
fetched_at TIMESTAMPTZ NOT NULL DEFAULT NOW(),
|
||||||
|
PRIMARY KEY (provider, provider_video_id)
|
||||||
|
);
|
||||||
|
|
||||||
|
COMMENT ON TABLE transcripts IS
|
||||||
|
'Shared public caption content keyed by (provider, provider_video_id). NOT RLS-scoped — public content only, de-facto cross-user dedup (ADR-021).';
|
||||||
@@ -0,0 +1 @@
|
|||||||
|
DROP TABLE channel_caption_state;
|
||||||
@@ -0,0 +1,30 @@
|
|||||||
|
-- Migration 016: per-(user, channel) caption-availability memory (ADR-024).
|
||||||
|
--
|
||||||
|
-- Some channels never publish English captions (foreign-language news, music,
|
||||||
|
-- etc.). Each of their new videos still costs ONE rate-limited caption fetch
|
||||||
|
-- (ADR-014) before resolving to "none" — and on a throttled egress IP that fetch
|
||||||
|
-- may 429 and churn through the backoff machinery first. This table remembers
|
||||||
|
-- channels that repeatedly yield no captions so discovery can stop attempting
|
||||||
|
-- their videos, freeing the scarce fetch budget for channels that do have them.
|
||||||
|
--
|
||||||
|
-- consecutive_none counts no-caption outcomes in a row; a successful fetch resets
|
||||||
|
-- it to 0. Once it crosses the threshold the channel is suppressed until
|
||||||
|
-- captionless_until, after which one video is re-probed (auto-recovery for a
|
||||||
|
-- channel that starts adding captions). Per-user + RLS-scoped, consistent with
|
||||||
|
-- the rest of the user-owned schema (subscriptions are per-user; ADR-012).
|
||||||
|
CREATE TABLE channel_caption_state (
|
||||||
|
user_id UUID NOT NULL REFERENCES users(id) ON DELETE CASCADE,
|
||||||
|
channel_id TEXT NOT NULL,
|
||||||
|
consecutive_none INT NOT NULL DEFAULT 0,
|
||||||
|
captionless_until TIMESTAMPTZ,
|
||||||
|
updated_at TIMESTAMPTZ NOT NULL DEFAULT now(),
|
||||||
|
PRIMARY KEY (user_id, channel_id)
|
||||||
|
);
|
||||||
|
|
||||||
|
CREATE INDEX idx_channel_caption_state_user_id ON channel_caption_state(user_id);
|
||||||
|
|
||||||
|
ALTER TABLE channel_caption_state ENABLE ROW LEVEL SECURITY;
|
||||||
|
ALTER TABLE channel_caption_state FORCE ROW LEVEL SECURITY;
|
||||||
|
CREATE POLICY channel_caption_state_isolation ON channel_caption_state
|
||||||
|
FOR ALL
|
||||||
|
USING (user_id = current_setting('tapir.current_user_id', true)::uuid);
|
||||||
@@ -29,6 +29,7 @@ type SummaryRow struct {
|
|||||||
ProviderVideoID string // videos.provider_video_id; empty when no videos row
|
ProviderVideoID string // videos.provider_video_id; empty when no videos row
|
||||||
Title string // videos.title; empty when no videos row
|
Title string // videos.title; empty when no videos row
|
||||||
Channel string // videos.provider for now; empty when no videos row
|
Channel string // videos.provider for now; empty when no videos row
|
||||||
|
ChannelTitle string // videos.channel_title; the source channel, for display + filtering
|
||||||
URL string // videos.url; empty when no videos row
|
URL string // videos.url; empty when no videos row
|
||||||
PublishedAt time.Time // videos.published_at; zero when absent
|
PublishedAt time.Time // videos.published_at; zero when absent
|
||||||
Summary string
|
Summary string
|
||||||
@@ -137,7 +138,8 @@ const selectVideo = `
|
|||||||
COALESCE(s.created_at, v.seen_at),
|
COALESCE(s.created_at, v.seen_at),
|
||||||
(s.id IS NOT NULL) AS summarized,
|
(s.id IS NOT NULL) AS summarized,
|
||||||
v.summarize_requested,
|
v.summarize_requested,
|
||||||
COALESCE(v.transcript_status, '')
|
COALESCE(v.transcript_status, ''),
|
||||||
|
COALESCE(v.channel_title, '')
|
||||||
FROM videos v
|
FROM videos v
|
||||||
LEFT JOIN summaries s ON s.video_id = v.id AND s.user_id = v.user_id`
|
LEFT JOIN summaries s ON s.video_id = v.id AND s.user_id = v.user_id`
|
||||||
|
|
||||||
@@ -254,6 +256,7 @@ func scanVideoRow(rows pgx.Row) (SummaryRow, error) {
|
|||||||
&row.Summarized,
|
&row.Summarized,
|
||||||
&row.SummarizeRequested,
|
&row.SummarizeRequested,
|
||||||
&row.TranscriptStatus,
|
&row.TranscriptStatus,
|
||||||
|
&row.ChannelTitle,
|
||||||
); err != nil {
|
); err != nil {
|
||||||
return SummaryRow{}, fmt.Errorf("store: scan video: %w", err)
|
return SummaryRow{}, fmt.Errorf("store: scan video: %w", err)
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -22,9 +22,11 @@ import (
|
|||||||
// (no GUC set → zero rows) proves the enforcement path is live, not bypassed.
|
// (no GUC set → zero rows) proves the enforcement path is live, not bypassed.
|
||||||
|
|
||||||
// userIsolatedTables are the tables that carry a user_id and whose policy keys
|
// userIsolatedTables are the tables that carry a user_id and whose policy keys
|
||||||
// directly off the tapir.current_user_id GUC.
|
// directly off the tapir.current_user_id GUC. transcripts is deliberately ABSENT
|
||||||
|
// — ADR-021 made it shared public content (non-RLS); TestTranscriptsTableIsSharedNotRLS
|
||||||
|
// proves that is the only place the isolation boundary moved.
|
||||||
var userIsolatedTables = []string{
|
var userIsolatedTables = []string{
|
||||||
"users", "videos", "transcripts", "summaries", "summary_actions", "login_events", "video_connections",
|
"users", "videos", "summaries", "summary_actions", "login_events", "video_connections",
|
||||||
}
|
}
|
||||||
|
|
||||||
// allIsolatedTables adds sink_deliveries, whose ownership is derived from its
|
// allIsolatedTables adds sink_deliveries, whose ownership is derived from its
|
||||||
@@ -38,9 +40,10 @@ type seeded struct {
|
|||||||
summaryID string
|
summaryID string
|
||||||
}
|
}
|
||||||
|
|
||||||
// seedUser inserts one full chain (user → video → transcript → summary →
|
// seedUser inserts one full chain (user → video → summary → action → delivery)
|
||||||
// action → delivery) as the superuser pool, which bypasses RLS so both users'
|
// as the superuser pool, which bypasses RLS so both users' data lands regardless
|
||||||
// data lands regardless of the GUC.
|
// of the GUC. Transcripts are NOT seeded here: they are shared, non-RLS public
|
||||||
|
// content (ADR-021), so they have no place in a per-user isolation chain.
|
||||||
func seedUser(t *testing.T, p *pgxpool.Pool, userID string) seeded {
|
func seedUser(t *testing.T, p *pgxpool.Pool, userID string) seeded {
|
||||||
t.Helper()
|
t.Helper()
|
||||||
ctx := context.Background()
|
ctx := context.Background()
|
||||||
@@ -54,11 +57,6 @@ func seedUser(t *testing.T, p *pgxpool.Pool, userID string) seeded {
|
|||||||
VALUES ($1, 'youtube', $2, 'title') RETURNING id`,
|
VALUES ($1, 'youtube', $2, 'title') RETURNING id`,
|
||||||
userID, "vid-"+userID).Scan(&videoID))
|
userID, "vid-"+userID).Scan(&videoID))
|
||||||
|
|
||||||
_, err = p.Exec(ctx,
|
|
||||||
`INSERT INTO transcripts (video_id, user_id, source, content)
|
|
||||||
VALUES ($1, $2, 'captions', 'words')`, videoID, userID)
|
|
||||||
require.NoError(t, err)
|
|
||||||
|
|
||||||
var summaryID string
|
var summaryID string
|
||||||
require.NoError(t, p.QueryRow(ctx,
|
require.NoError(t, p.QueryRow(ctx,
|
||||||
`INSERT INTO summaries (user_id, video_id, summary) VALUES ($1, $2, 'sum')
|
`INSERT INTO summaries (user_id, video_id, summary) VALUES ($1, $2, 'sum')
|
||||||
@@ -92,9 +90,16 @@ func appPool(t *testing.T, super *pgxpool.Pool) *pgxpool.Pool {
|
|||||||
t.Helper()
|
t.Helper()
|
||||||
ctx := context.Background()
|
ctx := context.Background()
|
||||||
|
|
||||||
// Idempotent across test runs (schema/role persist for the TestMain PG).
|
// Idempotent across tests AND runs: the role persists for the TestMain PG and
|
||||||
_, _ = super.Exec(ctx, `DROP ROLE IF EXISTS app`)
|
// owns granted privileges, so a plain DROP ROLE fails once any GRANT exists
|
||||||
_, err := super.Exec(ctx, `CREATE ROLE app LOGIN PASSWORD 'app'`)
|
// (and more than one test now builds an app pool). Create only if absent; the
|
||||||
|
// GRANTs below are themselves idempotent.
|
||||||
|
_, err := super.Exec(ctx,
|
||||||
|
`DO $$ BEGIN
|
||||||
|
IF NOT EXISTS (SELECT FROM pg_roles WHERE rolname = 'app') THEN
|
||||||
|
CREATE ROLE app LOGIN PASSWORD 'app';
|
||||||
|
END IF;
|
||||||
|
END $$`)
|
||||||
require.NoError(t, err)
|
require.NoError(t, err)
|
||||||
_, err = super.Exec(ctx, `GRANT USAGE ON SCHEMA public TO app`)
|
_, err = super.Exec(ctx, `GRANT USAGE ON SCHEMA public TO app`)
|
||||||
require.NoError(t, err)
|
require.NoError(t, err)
|
||||||
@@ -186,7 +191,6 @@ func TestRLSEnforcesPerUserIsolation(t *testing.T) {
|
|||||||
{"update users", `UPDATE users SET display_name = 'hacked' WHERE id = $1`, b.userID},
|
{"update users", `UPDATE users SET display_name = 'hacked' WHERE id = $1`, b.userID},
|
||||||
{"update videos", `UPDATE videos SET title = 'hacked' WHERE user_id = $1`, b.userID},
|
{"update videos", `UPDATE videos SET title = 'hacked' WHERE user_id = $1`, b.userID},
|
||||||
{"queue videos summarize", `UPDATE videos SET summarize_requested = TRUE WHERE id = $1`, b.videoID},
|
{"queue videos summarize", `UPDATE videos SET summarize_requested = TRUE WHERE id = $1`, b.videoID},
|
||||||
{"update transcripts", `UPDATE transcripts SET content = 'hacked' WHERE user_id = $1`, b.userID},
|
|
||||||
{"update summaries", `UPDATE summaries SET summary = 'hacked' WHERE user_id = $1`, b.userID},
|
{"update summaries", `UPDATE summaries SET summary = 'hacked' WHERE user_id = $1`, b.userID},
|
||||||
{"update summary_actions", `UPDATE summary_actions SET action = 'skipped' WHERE user_id = $1`, b.userID},
|
{"update summary_actions", `UPDATE summary_actions SET action = 'skipped' WHERE user_id = $1`, b.userID},
|
||||||
{"update login_events", `UPDATE login_events SET seen_at = NOW() WHERE user_id = $1`, b.userID},
|
{"update login_events", `UPDATE login_events SET seen_at = NOW() WHERE user_id = $1`, b.userID},
|
||||||
@@ -235,3 +239,55 @@ func TestRLSEnforcesPerUserIsolation(t *testing.T) {
|
|||||||
|
|
||||||
_ = a // a's ids are seeded for the symmetric read assertions above
|
_ = a // a's ids are seeded for the symmetric read assertions above
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// TestTranscriptsTableIsSharedNotRLS is the ADR-021 isolation proof: transcripts
|
||||||
|
// is the ONE shared, non-RLS surface, and the public-content classification
|
||||||
|
// leaked to nothing else. It is the inverse of TestRLSEnforcesPerUserIsolation —
|
||||||
|
// where that asserts deny-all on every user-owned table, this asserts transcripts
|
||||||
|
// is readable and writable with no user scope at all, holds no user_id, and is
|
||||||
|
// the single table with row-level security switched off.
|
||||||
|
func TestTranscriptsTableIsSharedNotRLS(t *testing.T) {
|
||||||
|
newStore(t)
|
||||||
|
super := rawPool(t)
|
||||||
|
resetDB(t, super)
|
||||||
|
app := appPool(t, super)
|
||||||
|
ctx := context.Background()
|
||||||
|
|
||||||
|
// 1. Shared + non-RLS: with NO GUC set, the app role both writes and reads a
|
||||||
|
// transcript. On an RLS table this would be deny-all (zero rows), exactly as
|
||||||
|
// the main isolation test asserts for every user-owned table.
|
||||||
|
_, err := app.Exec(ctx,
|
||||||
|
`INSERT INTO transcripts (provider, provider_video_id, source, content)
|
||||||
|
VALUES ('youtube', 'shared-vid', 'captions', 'public words')`)
|
||||||
|
require.NoError(t, err, "app role must write shared transcript content with no user scope")
|
||||||
|
require.Equal(t, 1, scopedCount(t, app, "", "transcripts"),
|
||||||
|
"transcripts must be readable with NO user scope — it is shared, non-RLS (ADR-021)")
|
||||||
|
|
||||||
|
// 2. No user_id column: the table holds only public caption content + the
|
||||||
|
// video's public id, nothing user-identifying.
|
||||||
|
var hasUserID bool
|
||||||
|
require.NoError(t, super.QueryRow(ctx,
|
||||||
|
`SELECT EXISTS (SELECT 1 FROM information_schema.columns
|
||||||
|
WHERE table_name = 'transcripts' AND column_name = 'user_id')`).Scan(&hasUserID))
|
||||||
|
require.False(t, hasUserID, "transcripts must carry no user_id (ADR-021 public content)")
|
||||||
|
|
||||||
|
// 3. The boundary is EXACTLY here: every user-owned table still has row-level
|
||||||
|
// security enabled; transcripts alone has it off. This is the proof the
|
||||||
|
// non-RLS classification was applied to transcripts and leaked nowhere else.
|
||||||
|
for _, table := range allIsolatedTables {
|
||||||
|
require.True(t, rlsEnabled(t, super, table),
|
||||||
|
"%s must still enforce row-level security — isolation must not have regressed", table)
|
||||||
|
}
|
||||||
|
require.False(t, rlsEnabled(t, super, "transcripts"),
|
||||||
|
"transcripts must be the single table with row-level security OFF (the one shared surface)")
|
||||||
|
}
|
||||||
|
|
||||||
|
// rlsEnabled reports whether a public table has ROW LEVEL SECURITY enabled.
|
||||||
|
func rlsEnabled(t *testing.T, p *pgxpool.Pool, table string) bool {
|
||||||
|
t.Helper()
|
||||||
|
var enabled bool
|
||||||
|
require.NoError(t, p.QueryRow(context.Background(),
|
||||||
|
`SELECT relrowsecurity FROM pg_class
|
||||||
|
WHERE relname = $1 AND relnamespace = 'public'::regnamespace`, table).Scan(&enabled))
|
||||||
|
return enabled
|
||||||
|
}
|
||||||
|
|||||||
@@ -0,0 +1,66 @@
|
|||||||
|
package store
|
||||||
|
|
||||||
|
import (
|
||||||
|
"context"
|
||||||
|
"errors"
|
||||||
|
"fmt"
|
||||||
|
|
||||||
|
"github.com/jackc/pgx/v5"
|
||||||
|
|
||||||
|
"gitea.d-ma.be/mathias/tapir/internal/domain"
|
||||||
|
)
|
||||||
|
|
||||||
|
// GetTranscript returns the shared, stored transcript for a video keyed by the
|
||||||
|
// cross-user dedup key (provider, providerVideoID), and whether one exists
|
||||||
|
// (ADR-021). It reads via the raw pool, NOT withUser: the table holds public
|
||||||
|
// content with no user_id and no RLS policy, so it is shared across users by
|
||||||
|
// construction. A stored SourceNone is a real hit (ok == true, HasText() ==
|
||||||
|
// false) — a known caption-less video, so the caller skips without re-fetching.
|
||||||
|
func (s *Store) GetTranscript(ctx context.Context, provider, providerVideoID string) (domain.Transcript, bool, error) {
|
||||||
|
var source, lang, content string
|
||||||
|
err := s.pool.QueryRow(ctx,
|
||||||
|
`SELECT source, COALESCE(language, ''), COALESCE(content, '')
|
||||||
|
FROM transcripts WHERE provider = $1 AND provider_video_id = $2`,
|
||||||
|
provider, providerVideoID).Scan(&source, &lang, &content)
|
||||||
|
if errors.Is(err, pgx.ErrNoRows) {
|
||||||
|
return domain.Transcript{}, false, nil
|
||||||
|
}
|
||||||
|
if err != nil {
|
||||||
|
return domain.Transcript{}, false, fmt.Errorf("store: get transcript: %w", err)
|
||||||
|
}
|
||||||
|
return domain.Transcript{
|
||||||
|
Source: domain.TranscriptSource(source),
|
||||||
|
Language: lang,
|
||||||
|
Content: content,
|
||||||
|
}, true, nil
|
||||||
|
}
|
||||||
|
|
||||||
|
// SaveTranscript upserts the shared transcript for (provider, providerVideoID).
|
||||||
|
// Only terminal outcomes belong here: SourceCaptions (with text) or SourceNone
|
||||||
|
// (no captions). A transient SourceRateLimited is rejected so persistence never
|
||||||
|
// masks a 429 as a permanent absence — that stays a per-user retry (ADR-014).
|
||||||
|
// Last write wins on conflict (a later re-fetch may correct an entry). It writes
|
||||||
|
// via the raw pool, NOT withUser — public content, shared, non-RLS (ADR-021).
|
||||||
|
func (s *Store) SaveTranscript(ctx context.Context, provider, providerVideoID string, t domain.Transcript) error {
|
||||||
|
switch t.Source {
|
||||||
|
case domain.SourceCaptions, domain.SourceNone:
|
||||||
|
// terminal — persist
|
||||||
|
case domain.SourceRateLimited:
|
||||||
|
return fmt.Errorf("store: refusing to persist transient rate-limited transcript for %s/%s", provider, providerVideoID)
|
||||||
|
default:
|
||||||
|
return fmt.Errorf("store: invalid transcript source %q", t.Source)
|
||||||
|
}
|
||||||
|
_, err := s.pool.Exec(ctx,
|
||||||
|
`INSERT INTO transcripts (provider, provider_video_id, source, language, content)
|
||||||
|
VALUES ($1, $2, $3, NULLIF($4, ''), NULLIF($5, ''))
|
||||||
|
ON CONFLICT (provider, provider_video_id)
|
||||||
|
DO UPDATE SET source = EXCLUDED.source,
|
||||||
|
language = EXCLUDED.language,
|
||||||
|
content = EXCLUDED.content,
|
||||||
|
fetched_at = NOW()`,
|
||||||
|
provider, providerVideoID, string(t.Source), t.Language, t.Content)
|
||||||
|
if err != nil {
|
||||||
|
return fmt.Errorf("store: save transcript: %w", err)
|
||||||
|
}
|
||||||
|
return nil
|
||||||
|
}
|
||||||
@@ -0,0 +1,87 @@
|
|||||||
|
package store_test
|
||||||
|
|
||||||
|
import (
|
||||||
|
"context"
|
||||||
|
"testing"
|
||||||
|
|
||||||
|
"github.com/stretchr/testify/require"
|
||||||
|
|
||||||
|
"gitea.d-ma.be/mathias/tapir/internal/adapters/store"
|
||||||
|
"gitea.d-ma.be/mathias/tapir/internal/domain"
|
||||||
|
"gitea.d-ma.be/mathias/tapir/internal/ports"
|
||||||
|
)
|
||||||
|
|
||||||
|
// Static check: Store satisfies the shared TranscriptStore port (ADR-021).
|
||||||
|
var _ ports.TranscriptStore = (*store.Store)(nil)
|
||||||
|
|
||||||
|
func TestSaveAndGetTranscript_RoundTrip(t *testing.T) {
|
||||||
|
s := newStore(t)
|
||||||
|
resetDB(t, rawPool(t))
|
||||||
|
ctx := context.Background()
|
||||||
|
|
||||||
|
want := domain.Transcript{Source: domain.SourceCaptions, Language: "en", Content: "the words"}
|
||||||
|
require.NoError(t, s.SaveTranscript(ctx, "youtube", "vid-1", want))
|
||||||
|
|
||||||
|
got, ok, err := s.GetTranscript(ctx, "youtube", "vid-1")
|
||||||
|
require.NoError(t, err)
|
||||||
|
require.True(t, ok, "a saved transcript must be found")
|
||||||
|
require.Equal(t, domain.SourceCaptions, got.Source)
|
||||||
|
require.Equal(t, "en", got.Language)
|
||||||
|
require.Equal(t, "the words", got.Content)
|
||||||
|
require.True(t, got.HasText())
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestGetTranscript_Miss(t *testing.T) {
|
||||||
|
s := newStore(t)
|
||||||
|
resetDB(t, rawPool(t))
|
||||||
|
|
||||||
|
_, ok, err := s.GetTranscript(context.Background(), "youtube", "absent")
|
||||||
|
require.NoError(t, err, "a miss is not an error")
|
||||||
|
require.False(t, ok)
|
||||||
|
}
|
||||||
|
|
||||||
|
// A stored "no captions" outcome is a real hit: callers must skip without
|
||||||
|
// re-fetching, so ok is true even though there is no text (ADR-021 / ADR-007).
|
||||||
|
func TestSaveAndGetTranscript_NoneIsAStoredHit(t *testing.T) {
|
||||||
|
s := newStore(t)
|
||||||
|
resetDB(t, rawPool(t))
|
||||||
|
ctx := context.Background()
|
||||||
|
|
||||||
|
require.NoError(t, s.SaveTranscript(ctx, "youtube", "vid-none", domain.Transcript{Source: domain.SourceNone}))
|
||||||
|
|
||||||
|
got, ok, err := s.GetTranscript(ctx, "youtube", "vid-none")
|
||||||
|
require.NoError(t, err)
|
||||||
|
require.True(t, ok, "a stored SourceNone is a hit, not a miss")
|
||||||
|
require.Equal(t, domain.SourceNone, got.Source)
|
||||||
|
require.False(t, got.HasText())
|
||||||
|
}
|
||||||
|
|
||||||
|
// A transient 429 must never be persisted as a terminal transcript, or a later
|
||||||
|
// read would mask the rate-limit as a permanent "no transcript" (ADR-014).
|
||||||
|
func TestSaveTranscript_RejectsRateLimited(t *testing.T) {
|
||||||
|
s := newStore(t)
|
||||||
|
resetDB(t, rawPool(t))
|
||||||
|
|
||||||
|
err := s.SaveTranscript(context.Background(), "youtube", "vid-429",
|
||||||
|
domain.Transcript{Source: domain.SourceRateLimited})
|
||||||
|
require.Error(t, err)
|
||||||
|
|
||||||
|
_, ok, _ := s.GetTranscript(context.Background(), "youtube", "vid-429")
|
||||||
|
require.False(t, ok, "a rejected rate-limited save must leave nothing stored")
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestSaveTranscript_UpsertLastWriteWins(t *testing.T) {
|
||||||
|
s := newStore(t)
|
||||||
|
resetDB(t, rawPool(t))
|
||||||
|
ctx := context.Background()
|
||||||
|
|
||||||
|
require.NoError(t, s.SaveTranscript(ctx, "youtube", "vid-up", domain.Transcript{Source: domain.SourceNone}))
|
||||||
|
require.NoError(t, s.SaveTranscript(ctx, "youtube", "vid-up",
|
||||||
|
domain.Transcript{Source: domain.SourceCaptions, Language: "en", Content: "now resolved"}))
|
||||||
|
|
||||||
|
got, ok, err := s.GetTranscript(ctx, "youtube", "vid-up")
|
||||||
|
require.NoError(t, err)
|
||||||
|
require.True(t, ok)
|
||||||
|
require.Equal(t, domain.SourceCaptions, got.Source)
|
||||||
|
require.Equal(t, "now resolved", got.Content)
|
||||||
|
}
|
||||||
@@ -46,14 +46,15 @@ func (s *Store) UpsertVideo(ctx context.Context, v domain.Video) (string, error)
|
|||||||
}
|
}
|
||||||
|
|
||||||
if err := tx.QueryRow(ctx,
|
if err := tx.QueryRow(ctx,
|
||||||
`INSERT INTO videos (user_id, provider, provider_video_id, title, url, published_at)
|
`INSERT INTO videos (user_id, provider, provider_video_id, title, url, published_at, channel_title)
|
||||||
VALUES ($1, $2, $3, $4, $5, $6)
|
VALUES ($1, $2, $3, $4, $5, $6, $7)
|
||||||
ON CONFLICT (user_id, provider, provider_video_id) DO UPDATE SET
|
ON CONFLICT (user_id, provider, provider_video_id) DO UPDATE SET
|
||||||
title = EXCLUDED.title,
|
title = EXCLUDED.title,
|
||||||
url = EXCLUDED.url,
|
url = EXCLUDED.url,
|
||||||
published_at = EXCLUDED.published_at
|
published_at = EXCLUDED.published_at,
|
||||||
|
channel_title = COALESCE(NULLIF(EXCLUDED.channel_title, ''), videos.channel_title)
|
||||||
RETURNING id`,
|
RETURNING id`,
|
||||||
v.UserID, provider, v.ProviderVideoID, v.Title, v.URL, nullTime(v.PublishedAt),
|
v.UserID, provider, v.ProviderVideoID, v.Title, v.URL, nullTime(v.PublishedAt), v.ChannelTitle,
|
||||||
).Scan(&id); err != nil {
|
).Scan(&id); err != nil {
|
||||||
return fmt.Errorf("store: upsert video: %w", err)
|
return fmt.Errorf("store: upsert video: %w", err)
|
||||||
}
|
}
|
||||||
@@ -72,3 +73,69 @@ func nullTime(t time.Time) *time.Time {
|
|||||||
}
|
}
|
||||||
return &t
|
return &t
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// NewestUnsummarizedVideoIDs returns up to limit of the user's videos that have
|
||||||
|
// no summary yet, newest first (published_at DESC, NULLS LAST). It caps the
|
||||||
|
// connect-time onboarding burst (Feature 1) at a fixed count: the caller marks
|
||||||
|
// these for summarization through the shared rate gate. RLS-scoped via withUser,
|
||||||
|
// so it only ever sees the requesting user's rows. limit <= 0 returns nil.
|
||||||
|
func (s *Store) NewestUnsummarizedVideoIDs(ctx context.Context, userID string, limit int) ([]string, error) {
|
||||||
|
if limit <= 0 {
|
||||||
|
return nil, nil
|
||||||
|
}
|
||||||
|
var ids []string
|
||||||
|
if err := s.withUser(ctx, userID, func(tx pgx.Tx) error {
|
||||||
|
rows, err := tx.Query(ctx,
|
||||||
|
`SELECT v.id
|
||||||
|
FROM videos v
|
||||||
|
WHERE v.user_id = $1
|
||||||
|
AND NOT EXISTS (
|
||||||
|
SELECT 1 FROM summaries su
|
||||||
|
WHERE su.user_id = v.user_id AND su.video_id = v.id)
|
||||||
|
ORDER BY v.published_at DESC NULLS LAST, v.seen_at DESC
|
||||||
|
LIMIT $2`, userID, limit)
|
||||||
|
if err != nil {
|
||||||
|
return fmt.Errorf("store: newest unsummarized: %w", err)
|
||||||
|
}
|
||||||
|
defer rows.Close()
|
||||||
|
for rows.Next() {
|
||||||
|
var id string
|
||||||
|
if err := rows.Scan(&id); err != nil {
|
||||||
|
return fmt.Errorf("store: scan newest unsummarized: %w", err)
|
||||||
|
}
|
||||||
|
ids = append(ids, id)
|
||||||
|
}
|
||||||
|
return rows.Err()
|
||||||
|
}); err != nil {
|
||||||
|
return nil, err
|
||||||
|
}
|
||||||
|
return ids, nil
|
||||||
|
}
|
||||||
|
|
||||||
|
// DistinctChannels returns the user's distinct, non-empty source channel titles
|
||||||
|
// (the channels they have videos from), alphabetically — the option list for the
|
||||||
|
// feed's channel filter. RLS-scoped via withUser.
|
||||||
|
func (s *Store) DistinctChannels(ctx context.Context, userID string) ([]string, error) {
|
||||||
|
var out []string
|
||||||
|
if err := s.withUser(ctx, userID, func(tx pgx.Tx) error {
|
||||||
|
rows, err := tx.Query(ctx,
|
||||||
|
`SELECT DISTINCT channel_title FROM videos
|
||||||
|
WHERE user_id = $1 AND channel_title IS NOT NULL AND channel_title <> ''
|
||||||
|
ORDER BY channel_title`, userID)
|
||||||
|
if err != nil {
|
||||||
|
return fmt.Errorf("store: distinct channels: %w", err)
|
||||||
|
}
|
||||||
|
defer rows.Close()
|
||||||
|
for rows.Next() {
|
||||||
|
var c string
|
||||||
|
if err := rows.Scan(&c); err != nil {
|
||||||
|
return fmt.Errorf("store: scan channel: %w", err)
|
||||||
|
}
|
||||||
|
out = append(out, c)
|
||||||
|
}
|
||||||
|
return rows.Err()
|
||||||
|
}); err != nil {
|
||||||
|
return nil, err
|
||||||
|
}
|
||||||
|
return out, nil
|
||||||
|
}
|
||||||
|
|||||||
@@ -81,3 +81,61 @@ func TestUpsertVideo_PerUserIsolation(t *testing.T) {
|
|||||||
|
|
||||||
require.NotEqual(t, idA, idB, "same provider video for two users must be two distinct rows")
|
require.NotEqual(t, idA, idB, "same provider video for two users must be two distinct rows")
|
||||||
}
|
}
|
||||||
|
|
||||||
|
func TestNewestUnsummarizedVideoIDs(t *testing.T) {
|
||||||
|
ctx := context.Background()
|
||||||
|
s := newStore(t)
|
||||||
|
resetDB(t, rawPool(t))
|
||||||
|
|
||||||
|
mk := func(user, pid string, day int) string {
|
||||||
|
v := ytVideo(user, pid, pid)
|
||||||
|
v.PublishedAt = time.Date(2026, 6, day, 12, 0, 0, 0, time.UTC)
|
||||||
|
id, err := s.UpsertVideo(ctx, v)
|
||||||
|
require.NoError(t, err)
|
||||||
|
return id
|
||||||
|
}
|
||||||
|
|
||||||
|
_ = mk(userA, "a1vid000001", 1)
|
||||||
|
id2 := mk(userA, "a2vid000002", 2)
|
||||||
|
id3 := mk(userA, "a3vid000003", 3)
|
||||||
|
id4 := mk(userA, "a4vid000004", 4)
|
||||||
|
mk(userB, "b1vid000009", 9) // userB's newest — must never leak via RLS
|
||||||
|
|
||||||
|
// The newest (v4) is summarized, so it's excluded from "unsummarized".
|
||||||
|
require.NoError(t, s.Deliver(ctx, summary(userA, id4, "done")))
|
||||||
|
|
||||||
|
// Cap 2, newest-first unsummarized: v3 then v2 (v4 excluded; userB excluded).
|
||||||
|
got, err := s.NewestUnsummarizedVideoIDs(ctx, userA, 2)
|
||||||
|
require.NoError(t, err)
|
||||||
|
require.Equal(t, []string{id3, id2}, got)
|
||||||
|
|
||||||
|
none, err := s.NewestUnsummarizedVideoIDs(ctx, userA, 0)
|
||||||
|
require.NoError(t, err)
|
||||||
|
require.Empty(t, none, "limit 0 returns nothing")
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestUpsertVideoPersistsChannelAndDistinctChannels(t *testing.T) {
|
||||||
|
ctx := context.Background()
|
||||||
|
s := newStore(t)
|
||||||
|
resetDB(t, rawPool(t))
|
||||||
|
|
||||||
|
mk := func(pid, channel string) {
|
||||||
|
v := ytVideo(userA, pid, pid)
|
||||||
|
v.ChannelTitle = channel
|
||||||
|
_, err := s.UpsertVideo(ctx, v)
|
||||||
|
require.NoError(t, err)
|
||||||
|
}
|
||||||
|
mk("aa11111aaaa", "Acme Talks")
|
||||||
|
mk("bb22222bbbb", "Acme Talks") // same channel
|
||||||
|
mk("cc33333cccc", "Zeta Channel")
|
||||||
|
// userB's channel must not leak.
|
||||||
|
vb := ytVideo(userB, "dd44444dddd", "x")
|
||||||
|
vb.ChannelTitle = "Bravo Only"
|
||||||
|
_, err := s.UpsertVideo(ctx, vb)
|
||||||
|
require.NoError(t, err)
|
||||||
|
|
||||||
|
got, err := s.DistinctChannels(ctx, userA)
|
||||||
|
require.NoError(t, err)
|
||||||
|
require.Equal(t, []string{"Acme Talks", "Zeta Channel"}, got,
|
||||||
|
"distinct, alphabetical, user-scoped (no Bravo Only)")
|
||||||
|
}
|
||||||
|
|||||||
@@ -1,17 +1,21 @@
|
|||||||
// Package summarizer implements ports.Summarizer backed by the copied llm
|
// Package summarizer implements ports.Summarizer backed by the copied llm
|
||||||
// package's local-Primary -> BYO-Fallback routing (ADR-004). It is the only
|
// package's routing (ADR-004, extended by ADR-022). It is the only place content
|
||||||
// place content ever leaves the engine toward an AI model, so it is also the
|
// ever leaves the engine toward an AI model, so it is also the enforcement point
|
||||||
// enforcement point for the local-first guarantee in
|
// for the local-first guarantee in docs/use-cases/ai_routing.feature: endpoints
|
||||||
// docs/use-cases/ai_routing.feature: a user with no BYO provider configured has
|
// are tried in order, locals first, so content only reaches an external model
|
||||||
// their content sent to the local stack and nowhere else.
|
// after every local endpoint has failed — and never at all when no external
|
||||||
|
// endpoint is configured.
|
||||||
package summarizer
|
package summarizer
|
||||||
|
|
||||||
import (
|
import (
|
||||||
|
"bytes"
|
||||||
"context"
|
"context"
|
||||||
"encoding/json"
|
"encoding/json"
|
||||||
|
"errors"
|
||||||
"fmt"
|
"fmt"
|
||||||
"strings"
|
"strings"
|
||||||
"time"
|
"time"
|
||||||
|
"unicode/utf8"
|
||||||
|
|
||||||
"gitea.d-ma.be/mathias/tapir/internal/domain"
|
"gitea.d-ma.be/mathias/tapir/internal/domain"
|
||||||
)
|
)
|
||||||
@@ -29,19 +33,41 @@ type Endpoint struct {
|
|||||||
Model string // resolved alias, e.g. "iguana/deepseek-r1-14b"
|
Model string // resolved alias, e.g. "iguana/deepseek-r1-14b"
|
||||||
}
|
}
|
||||||
|
|
||||||
// Summarizer routes a transcript through the local endpoint first, then the
|
// Summarizer routes a transcript through an ordered chain of endpoints, trying
|
||||||
// optional BYO endpoint. It owns its routing (rather than delegating to
|
// each in turn until one returns a parseable summary. It owns its routing
|
||||||
// llm.Router) so it can record which provider answered and whether the fallback
|
// (rather than delegating to llm.Router) so it can record which provider answered
|
||||||
// was used — information llm.Router collapses away.
|
// and whether a fallback was used — information llm.Router collapses away. The
|
||||||
|
// chain ordering is the local-first guarantee: callers place local endpoints
|
||||||
|
// first and any external endpoint last, so content only reaches an external model
|
||||||
|
// after every local endpoint has failed.
|
||||||
type Summarizer struct {
|
type Summarizer struct {
|
||||||
primary Endpoint
|
endpoints []Endpoint
|
||||||
fallback *Endpoint // nil => no BYO; primary errors are returned, never sent externally
|
maxInputChars int // transcript truncation budget; 0 = no limit
|
||||||
now func() time.Time
|
now func() time.Time
|
||||||
}
|
}
|
||||||
|
|
||||||
// New constructs a Summarizer. fallback may be nil (no BYO provider configured).
|
// New constructs a Summarizer from a primary endpoint and an optional fallback
|
||||||
|
// (the historical local-Primary -> BYO-Fallback shape, ADR-004). A nil fallback
|
||||||
|
// means a single-endpoint chain: errors are returned, content never leaves it.
|
||||||
func New(primary Endpoint, fallback *Endpoint) *Summarizer {
|
func New(primary Endpoint, fallback *Endpoint) *Summarizer {
|
||||||
return &Summarizer{primary: primary, fallback: fallback, now: time.Now}
|
eps := []Endpoint{primary}
|
||||||
|
if fallback != nil {
|
||||||
|
eps = append(eps, *fallback)
|
||||||
|
}
|
||||||
|
return &Summarizer{endpoints: eps, now: time.Now}
|
||||||
|
}
|
||||||
|
|
||||||
|
// NewChain constructs a Summarizer over an ordered endpoint chain (ADR-022).
|
||||||
|
// endpoints are tried in order; the first to return a parseable summary wins, and
|
||||||
|
// FallbackUsed is recorded true for any endpoint past the first. maxInputChars
|
||||||
|
// bounds the transcript text sent to every endpoint (0 = unbounded), so a long
|
||||||
|
// transcript does not overflow a small-context primary model's window. It panics
|
||||||
|
// on an empty chain — a wiring bug, not a runtime condition.
|
||||||
|
func NewChain(endpoints []Endpoint, maxInputChars int) *Summarizer {
|
||||||
|
if len(endpoints) == 0 {
|
||||||
|
panic("summarizer: NewChain requires at least one endpoint")
|
||||||
|
}
|
||||||
|
return &Summarizer{endpoints: endpoints, maxInputChars: maxInputChars, now: time.Now}
|
||||||
}
|
}
|
||||||
|
|
||||||
const systemPrompt = `You are Tapir, a video-summarization assistant.
|
const systemPrompt = `You are Tapir, a video-summarization assistant.
|
||||||
@@ -53,31 +79,35 @@ Respond with ONLY a JSON object, no prose and no code fences:
|
|||||||
- "takeaways": the actionable conclusions a viewer should leave with.
|
- "takeaways": the actionable conclusions a viewer should leave with.
|
||||||
Output the JSON object and nothing else.`
|
Output the JSON object and nothing else.`
|
||||||
|
|
||||||
// Summarize implements ports.Summarizer.
|
// Summarize implements ports.Summarizer. It walks the endpoint chain in order:
|
||||||
|
// the first endpoint whose reply parses into a non-empty summary wins. An
|
||||||
|
// endpoint is considered failed — and the next one tried — when the model call
|
||||||
|
// errors OR when its reply cannot be parsed (a 200 with malformed JSON or a
|
||||||
|
// highlights field the model emitted as a bare string). Truncation is applied
|
||||||
|
// once, up front, so every endpoint sees the same bounded prompt. When the whole
|
||||||
|
// chain fails, the joined error is returned so the engine queues the work for
|
||||||
|
// retry and delivers no summary.
|
||||||
func (s *Summarizer) Summarize(ctx context.Context, v domain.Video, t domain.Transcript) (domain.Summary, error) {
|
func (s *Summarizer) Summarize(ctx context.Context, v domain.Video, t domain.Transcript) (domain.Summary, error) {
|
||||||
if !t.HasText() {
|
if !t.HasText() {
|
||||||
return domain.Summary{}, fmt.Errorf("summarize: transcript for video %s has no text", v.ID)
|
return domain.Summary{}, fmt.Errorf("summarize: transcript for video %s has no text", v.ID)
|
||||||
}
|
}
|
||||||
user := buildUserPrompt(v, t)
|
user := buildUserPrompt(v, t, s.maxInputChars)
|
||||||
|
|
||||||
// Primary = local stack. Only on its failure is anything sent externally,
|
var errs []error
|
||||||
// and only when a BYO fallback is configured.
|
for i, ep := range s.endpoints {
|
||||||
out, err := s.primary.Client.Complete(ctx, systemPrompt, user)
|
out, err := ep.Client.Complete(ctx, systemPrompt, user)
|
||||||
if err == nil {
|
if err != nil {
|
||||||
return s.build(v, s.primary, false, out)
|
errs = append(errs, fmt.Errorf("%s/%s call: %w", ep.Provider, ep.Model, err))
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
sum, perr := s.build(v, ep, i > 0, out)
|
||||||
|
if perr != nil {
|
||||||
|
errs = append(errs, fmt.Errorf("%s/%s output: %w", ep.Provider, ep.Model, perr))
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
return sum, nil
|
||||||
}
|
}
|
||||||
|
return domain.Summary{}, fmt.Errorf("summarize: all %d endpoint(s) failed: %w", len(s.endpoints), errors.Join(errs...))
|
||||||
if s.fallback == nil {
|
|
||||||
// No BYO: content was sent to the local stack only. Surface the error so
|
|
||||||
// the engine can queue the work for retry; deliver no summary.
|
|
||||||
return domain.Summary{}, fmt.Errorf("summarize: local AI failed and no BYO provider configured: %w", err)
|
|
||||||
}
|
|
||||||
|
|
||||||
out, ferr := s.fallback.Client.Complete(ctx, systemPrompt, user)
|
|
||||||
if ferr != nil {
|
|
||||||
return domain.Summary{}, fmt.Errorf("summarize: local AI failed: %w; BYO %s failed: %v", err, s.fallback.Provider, ferr)
|
|
||||||
}
|
|
||||||
return s.build(v, *s.fallback, true, out)
|
|
||||||
}
|
}
|
||||||
|
|
||||||
func (s *Summarizer) build(v domain.Video, ep Endpoint, fallbackUsed bool, raw string) (domain.Summary, error) {
|
func (s *Summarizer) build(v domain.Video, ep Endpoint, fallbackUsed bool, raw string) (domain.Summary, error) {
|
||||||
@@ -89,8 +119,8 @@ func (s *Summarizer) build(v domain.Video, ep Endpoint, fallbackUsed bool, raw s
|
|||||||
UserID: v.UserID,
|
UserID: v.UserID,
|
||||||
VideoID: v.ID,
|
VideoID: v.ID,
|
||||||
Summary: parsed.Summary,
|
Summary: parsed.Summary,
|
||||||
Highlights: parsed.Highlights,
|
Highlights: []string(parsed.Highlights),
|
||||||
Takeaways: parsed.Takeaways,
|
Takeaways: []string(parsed.Takeaways),
|
||||||
AIProvider: ep.Provider,
|
AIProvider: ep.Provider,
|
||||||
AIModel: ep.Model,
|
AIModel: ep.Model,
|
||||||
FallbackUsed: fallbackUsed,
|
FallbackUsed: fallbackUsed,
|
||||||
@@ -98,20 +128,94 @@ func (s *Summarizer) build(v domain.Video, ep Endpoint, fallbackUsed bool, raw s
|
|||||||
}, nil
|
}, nil
|
||||||
}
|
}
|
||||||
|
|
||||||
func buildUserPrompt(v domain.Video, t domain.Transcript) string {
|
func buildUserPrompt(v domain.Video, t domain.Transcript, maxInputChars int) string {
|
||||||
var b strings.Builder
|
var b strings.Builder
|
||||||
fmt.Fprintf(&b, "Title: %s\n", v.Title)
|
fmt.Fprintf(&b, "Title: %s\n", v.Title)
|
||||||
if v.URL != "" {
|
if v.URL != "" {
|
||||||
fmt.Fprintf(&b, "URL: %s\n", v.URL)
|
fmt.Fprintf(&b, "URL: %s\n", v.URL)
|
||||||
}
|
}
|
||||||
fmt.Fprintf(&b, "\nTranscript:\n%s", t.Content)
|
fmt.Fprintf(&b, "\nTranscript:\n%s", truncate(t.Content, maxInputChars))
|
||||||
return b.String()
|
return b.String()
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// truncate caps content to max bytes on a UTF-8 rune boundary, appending a
|
||||||
|
// marker so the model knows the transcript was cut. A non-positive max (or a
|
||||||
|
// content already within budget) returns content unchanged. Bounding the input
|
||||||
|
// keeps a long transcript from overflowing a small-context model's window — the
|
||||||
|
// production failure mode where koala/phi4-mini's 8k context returned HTTP 400 on
|
||||||
|
// a 11.6k-token transcript.
|
||||||
|
func truncate(content string, max int) string {
|
||||||
|
if max <= 0 || len(content) <= max {
|
||||||
|
return content
|
||||||
|
}
|
||||||
|
cut := max
|
||||||
|
for cut > 0 && !utf8.RuneStart(content[cut]) {
|
||||||
|
cut--
|
||||||
|
}
|
||||||
|
return content[:cut] + "\n…[transcript truncated to fit the model context]"
|
||||||
|
}
|
||||||
|
|
||||||
|
// flexStrings is a []string that also unmarshals from a single JSON string or a
|
||||||
|
// JSON array of scalars. Small local models (koala/phi4-mini) sometimes emit
|
||||||
|
// "highlights": "one point" instead of an array, or mix in a number; rather than
|
||||||
|
// fail the whole summary on that quirk, coerce to []string. Empty/whitespace
|
||||||
|
// elements are dropped.
|
||||||
|
type flexStrings []string
|
||||||
|
|
||||||
|
func (f *flexStrings) UnmarshalJSON(b []byte) error {
|
||||||
|
b = bytes.TrimSpace(b)
|
||||||
|
if len(b) == 0 || string(b) == "null" {
|
||||||
|
*f = nil
|
||||||
|
return nil
|
||||||
|
}
|
||||||
|
if b[0] == '[' {
|
||||||
|
var raw []json.RawMessage
|
||||||
|
if err := json.Unmarshal(b, &raw); err != nil {
|
||||||
|
return err
|
||||||
|
}
|
||||||
|
out := make([]string, 0, len(raw))
|
||||||
|
for _, r := range raw {
|
||||||
|
s, err := rawToString(r)
|
||||||
|
if err != nil {
|
||||||
|
return err
|
||||||
|
}
|
||||||
|
if strings.TrimSpace(s) != "" {
|
||||||
|
out = append(out, s)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
*f = out
|
||||||
|
return nil
|
||||||
|
}
|
||||||
|
s, err := rawToString(b)
|
||||||
|
if err != nil {
|
||||||
|
return err
|
||||||
|
}
|
||||||
|
if strings.TrimSpace(s) == "" {
|
||||||
|
*f = nil
|
||||||
|
} else {
|
||||||
|
*f = flexStrings{s}
|
||||||
|
}
|
||||||
|
return nil
|
||||||
|
}
|
||||||
|
|
||||||
|
// rawToString renders a JSON scalar as text: a quoted string is unquoted; any
|
||||||
|
// other scalar (number, bool) is kept as its literal source so no content is lost.
|
||||||
|
func rawToString(r json.RawMessage) (string, error) {
|
||||||
|
r = bytes.TrimSpace(r)
|
||||||
|
if len(r) > 0 && r[0] == '"' {
|
||||||
|
var s string
|
||||||
|
if err := json.Unmarshal(r, &s); err != nil {
|
||||||
|
return "", err
|
||||||
|
}
|
||||||
|
return s, nil
|
||||||
|
}
|
||||||
|
return string(r), nil
|
||||||
|
}
|
||||||
|
|
||||||
type parsedSummary struct {
|
type parsedSummary struct {
|
||||||
Summary string `json:"summary"`
|
Summary string `json:"summary"`
|
||||||
Highlights []string `json:"highlights"`
|
Highlights flexStrings `json:"highlights"`
|
||||||
Takeaways []string `json:"takeaways"`
|
Takeaways flexStrings `json:"takeaways"`
|
||||||
}
|
}
|
||||||
|
|
||||||
// parse extracts the JSON object from a model reply. Thinking models (qwen3,
|
// parse extracts the JSON object from a model reply. Thinking models (qwen3,
|
||||||
|
|||||||
@@ -129,8 +129,8 @@ func TestSummarize_NoBYO_ContentOnlyLocal(t *testing.T) {
|
|||||||
local := &fakeClient{reply: goodReply}
|
local := &fakeClient{reply: goodReply}
|
||||||
s := New(Endpoint{Client: local, Provider: "local", Model: "iguana/deepseek-r1-14b"}, nil)
|
s := New(Endpoint{Client: local, Provider: "local", Model: "iguana/deepseek-r1-14b"}, nil)
|
||||||
|
|
||||||
if s.fallback != nil {
|
if len(s.endpoints) != 1 {
|
||||||
t.Fatal("no BYO configured but fallback endpoint is non-nil")
|
t.Fatalf("no BYO configured but chain has %d endpoints, want 1", len(s.endpoints))
|
||||||
}
|
}
|
||||||
for i := 0; i < 3; i++ {
|
for i := 0; i < 3; i++ {
|
||||||
sum, err := s.Summarize(context.Background(), testVideo(), testTranscript())
|
sum, err := s.Summarize(context.Background(), testVideo(), testTranscript())
|
||||||
@@ -176,3 +176,97 @@ func TestParse_EmptySummaryRejected(t *testing.T) {
|
|||||||
t.Fatal("want error for empty summary (thinking model returned no content)")
|
t.Fatal("want error for empty summary (thinking model returned no content)")
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// parse tolerates a small model emitting "highlights" as a bare string instead
|
||||||
|
// of an array — the production koala/phi4-mini quirk that errored with
|
||||||
|
// "cannot unmarshal string into Go struct field ... highlights of type []string".
|
||||||
|
func TestParse_ToleratesStringHighlights(t *testing.T) {
|
||||||
|
p, err := parse(`{"summary":"s","highlights":"one big point","takeaways":["a","b"]}`)
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("parse: %v", err)
|
||||||
|
}
|
||||||
|
if len(p.Highlights) != 1 || p.Highlights[0] != "one big point" {
|
||||||
|
t.Errorf("highlights = %v, want [\"one big point\"]", p.Highlights)
|
||||||
|
}
|
||||||
|
if len(p.Takeaways) != 2 {
|
||||||
|
t.Errorf("takeaways = %v, want 2", p.Takeaways)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// Chain: an endpoint that returns a 200 with unparseable output is treated as a
|
||||||
|
// failure, and the next endpoint in the chain is tried. This is the case the old
|
||||||
|
// primary->fallback shape missed — a parse error short-circuited instead of
|
||||||
|
// falling back.
|
||||||
|
func TestSummarize_FallsBackOnMalformedOutput(t *testing.T) {
|
||||||
|
bad := &fakeClient{reply: `{"summary": not json`}
|
||||||
|
good := &fakeClient{reply: goodReply}
|
||||||
|
s := NewChain([]Endpoint{
|
||||||
|
{Client: bad, Provider: "local", Model: "koala/phi4-mini"},
|
||||||
|
{Client: good, Provider: "local", Model: "koala/phi4-14b"},
|
||||||
|
}, 0)
|
||||||
|
|
||||||
|
sum, err := s.Summarize(context.Background(), testVideo(), testTranscript())
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("Summarize: %v", err)
|
||||||
|
}
|
||||||
|
if sum.AIModel != "koala/phi4-14b" {
|
||||||
|
t.Errorf("AIModel = %q, want koala/phi4-14b (fell back past malformed primary)", sum.AIModel)
|
||||||
|
}
|
||||||
|
if !sum.FallbackUsed {
|
||||||
|
t.Error("FallbackUsed = false, want true")
|
||||||
|
}
|
||||||
|
if bad.calls != 1 || good.calls != 1 {
|
||||||
|
t.Errorf("calls: bad=%d good=%d, want 1 and 1", bad.calls, good.calls)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// Chain: when every endpoint fails, no summary is produced and the joined error
|
||||||
|
// names each failure so the engine queues the work for retry.
|
||||||
|
func TestSummarize_ChainAllEndpointsFail(t *testing.T) {
|
||||||
|
a := &fakeClient{err: errors.New("context overflow")}
|
||||||
|
b := &fakeClient{reply: "not even json"}
|
||||||
|
s := NewChain([]Endpoint{
|
||||||
|
{Client: a, Provider: "local", Model: "m1"},
|
||||||
|
{Client: b, Provider: "berget", Model: "m2"},
|
||||||
|
}, 0)
|
||||||
|
|
||||||
|
if _, err := s.Summarize(context.Background(), testVideo(), testTranscript()); err == nil {
|
||||||
|
t.Fatal("want error when all endpoints fail")
|
||||||
|
}
|
||||||
|
if a.calls != 1 || b.calls != 1 {
|
||||||
|
t.Errorf("calls: a=%d b=%d, want 1 and 1", a.calls, b.calls)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// A transcript longer than the chain's input budget is truncated before it
|
||||||
|
// reaches any model, so a small-context primary does not overflow its window.
|
||||||
|
func TestSummarize_TruncatesLongTranscript(t *testing.T) {
|
||||||
|
local := &fakeClient{reply: goodReply}
|
||||||
|
const budget = 100
|
||||||
|
s := NewChain([]Endpoint{{Client: local, Provider: "local", Model: "m"}}, budget)
|
||||||
|
|
||||||
|
long := domain.Transcript{
|
||||||
|
VideoID: "vid-1", UserID: "user-1", Source: domain.SourceCaptions,
|
||||||
|
Content: strings.Repeat("word ", 1000), // 5000 bytes, well over budget
|
||||||
|
}
|
||||||
|
if _, err := s.Summarize(context.Background(), testVideo(), long); err != nil {
|
||||||
|
t.Fatalf("Summarize: %v", err)
|
||||||
|
}
|
||||||
|
// The prompt carries title/URL framing plus the truncation marker, so allow
|
||||||
|
// headroom over the raw transcript budget — but it must be far below 5000.
|
||||||
|
if len(local.lastUser) > budget+300 {
|
||||||
|
t.Errorf("prompt length = %d, want <= %d (transcript not truncated)", len(local.lastUser), budget+300)
|
||||||
|
}
|
||||||
|
if !strings.Contains(local.lastUser, "truncated") {
|
||||||
|
t.Error("truncation marker missing from prompt")
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestNewChain_PanicsOnEmptyChain(t *testing.T) {
|
||||||
|
defer func() {
|
||||||
|
if recover() == nil {
|
||||||
|
t.Fatal("want panic on empty endpoint chain")
|
||||||
|
}
|
||||||
|
}()
|
||||||
|
NewChain(nil, 0)
|
||||||
|
}
|
||||||
|
|||||||
@@ -0,0 +1,62 @@
|
|||||||
|
package youtube
|
||||||
|
|
||||||
|
import (
|
||||||
|
"context"
|
||||||
|
"errors"
|
||||||
|
"net/http"
|
||||||
|
"testing"
|
||||||
|
|
||||||
|
"gitea.d-ma.be/mathias/tapir/internal/domain"
|
||||||
|
)
|
||||||
|
|
||||||
|
func TestVideoByID(t *testing.T) {
|
||||||
|
const id = "dQw4w9WgXcQ"
|
||||||
|
a, secrets := newTestAdapter(t, func(w http.ResponseWriter, r *http.Request) {
|
||||||
|
if r.URL.Path != "/videos" {
|
||||||
|
t.Errorf("unexpected path %q (must use videos.list)", r.URL.Path)
|
||||||
|
}
|
||||||
|
if got := r.URL.Query().Get("id"); got != id {
|
||||||
|
t.Errorf("expected id=%s, got %q", id, got)
|
||||||
|
}
|
||||||
|
if got := r.URL.Query().Get("part"); got != "snippet" {
|
||||||
|
t.Errorf("expected part=snippet, got %q", got)
|
||||||
|
}
|
||||||
|
_, _ = w.Write([]byte(`{"items":[{"snippet":{"title":"Never Gonna Give You Up","channelTitle":"Rick Astley","publishedAt":"2026-05-20T09:00:00Z"}}]}`))
|
||||||
|
})
|
||||||
|
|
||||||
|
v, err := a.VideoByID(context.Background(), "u1", id)
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("VideoByID: %v", err)
|
||||||
|
}
|
||||||
|
if v.UserID != "u1" {
|
||||||
|
t.Errorf("UserID = %q, want u1", v.UserID)
|
||||||
|
}
|
||||||
|
if v.ProviderVideoID != id || v.Title != "Never Gonna Give You Up" {
|
||||||
|
t.Errorf("unexpected video: %+v", v)
|
||||||
|
}
|
||||||
|
if v.ChannelTitle != "Rick Astley" {
|
||||||
|
t.Errorf("ChannelTitle = %q, want Rick Astley", v.ChannelTitle)
|
||||||
|
}
|
||||||
|
if v.Provider != domain.ProviderYouTube || v.URL != "https://www.youtube.com/watch?v="+id {
|
||||||
|
t.Errorf("video not wired correctly: %+v", v)
|
||||||
|
}
|
||||||
|
if v.PublishedAt.IsZero() {
|
||||||
|
t.Errorf("expected publishedAt parsed, got zero")
|
||||||
|
}
|
||||||
|
if v.SubscriptionID != "" {
|
||||||
|
t.Errorf("a pasted video must have no subscription, got %q", v.SubscriptionID)
|
||||||
|
}
|
||||||
|
if secrets.byRef == nil {
|
||||||
|
t.Errorf("token must be resolved by reference through the SecretStore")
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestVideoByIDNotFound(t *testing.T) {
|
||||||
|
a, _ := newTestAdapter(t, func(w http.ResponseWriter, _ *http.Request) {
|
||||||
|
_, _ = w.Write([]byte(`{"items":[]}`))
|
||||||
|
})
|
||||||
|
_, err := a.VideoByID(context.Background(), "u1", "missingvid0")
|
||||||
|
if !errors.Is(err, domain.ErrVideoNotFound) {
|
||||||
|
t.Fatalf("VideoByID for missing id = %v, want domain.ErrVideoNotFound", err)
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -66,6 +66,13 @@ type Config struct {
|
|||||||
// poll. Zero means defaultMaxVideos.
|
// poll. Zero means defaultMaxVideos.
|
||||||
MaxVideosPerSubscription int
|
MaxVideosPerSubscription int
|
||||||
|
|
||||||
|
// MinVideoSeconds drops videos shorter than this from discovery (Shorts/clips,
|
||||||
|
// ADR-023). NewVideos enriches candidates with a single cheap videos.list call
|
||||||
|
// (contentDetails.duration + snippet.liveBroadcastContent) and filters before
|
||||||
|
// returning, so the scarce caption-fetch budget is never spent on them. Live
|
||||||
|
// and upcoming broadcasts are dropped too. Zero disables the filter.
|
||||||
|
MinVideoSeconds int
|
||||||
|
|
||||||
// BaseURL overrides the Data API root. Empty means defaultBaseURL.
|
// BaseURL overrides the Data API root. Empty means defaultBaseURL.
|
||||||
BaseURL string
|
BaseURL string
|
||||||
|
|
||||||
@@ -234,6 +241,7 @@ func (a *Adapter) NewVideos(ctx context.Context, sub domain.Subscription) ([]dom
|
|||||||
Provider: domain.ProviderYouTube,
|
Provider: domain.ProviderYouTube,
|
||||||
ProviderVideoID: vid,
|
ProviderVideoID: vid,
|
||||||
Title: item.Snippet.Title,
|
Title: item.Snippet.Title,
|
||||||
|
ChannelTitle: sub.ChannelTitle,
|
||||||
URL: "https://www.youtube.com/watch?v=" + vid,
|
URL: "https://www.youtube.com/watch?v=" + vid,
|
||||||
PublishedAt: item.Snippet.PublishedAt,
|
PublishedAt: item.Snippet.PublishedAt,
|
||||||
})
|
})
|
||||||
@@ -241,7 +249,129 @@ func (a *Adapter) NewVideos(ctx context.Context, sub domain.Subscription) ([]dom
|
|||||||
break
|
break
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
return videos, nil
|
|
||||||
|
// Drop Shorts/sub-minute clips and live/upcoming broadcasts before they ever
|
||||||
|
// reach the rate-limited caption path (ADR-023). One cheap videos.list call
|
||||||
|
// (quota API, not the timedtext throttle) supplies duration + live status.
|
||||||
|
return a.filterLowValue(ctx, client, videos), nil
|
||||||
|
}
|
||||||
|
|
||||||
|
// filterLowValue removes videos shorter than cfg.MinVideoSeconds and any live or
|
||||||
|
// upcoming broadcast, using a single videos.list lookup for duration +
|
||||||
|
// liveBroadcastContent. The filter is best-effort: if MinVideoSeconds is 0 (off)
|
||||||
|
// or the lookup fails, the input is returned unfiltered — discovery must not break
|
||||||
|
// because a metadata call hiccuped; the worst case is the pre-ADR-023 behaviour.
|
||||||
|
func (a *Adapter) filterLowValue(ctx context.Context, client *http.Client, videos []domain.Video) []domain.Video {
|
||||||
|
if a.cfg.MinVideoSeconds <= 0 || len(videos) == 0 {
|
||||||
|
return videos
|
||||||
|
}
|
||||||
|
|
||||||
|
ids := make([]string, 0, len(videos))
|
||||||
|
for _, v := range videos {
|
||||||
|
ids = append(ids, v.ProviderVideoID)
|
||||||
|
}
|
||||||
|
q := url.Values{
|
||||||
|
"part": {"contentDetails,snippet"},
|
||||||
|
"id": {strings.Join(ids, ",")},
|
||||||
|
}
|
||||||
|
var resp videoListResponse
|
||||||
|
if err := a.getJSON(ctx, client, "/videos", q, &resp); err != nil {
|
||||||
|
// Degrade open: keep the candidates rather than lose discovery.
|
||||||
|
return videos
|
||||||
|
}
|
||||||
|
|
||||||
|
type meta struct {
|
||||||
|
seconds int
|
||||||
|
live string
|
||||||
|
}
|
||||||
|
byID := make(map[string]meta, len(resp.Items))
|
||||||
|
for _, it := range resp.Items {
|
||||||
|
byID[it.ID] = meta{seconds: parseISO8601Seconds(it.ContentDetails.Duration), live: it.Snippet.LiveBroadcastContent}
|
||||||
|
}
|
||||||
|
|
||||||
|
kept := videos[:0]
|
||||||
|
for _, v := range videos {
|
||||||
|
m, ok := byID[v.ProviderVideoID]
|
||||||
|
if !ok {
|
||||||
|
kept = append(kept, v) // unknown metadata: keep, let the fetch decide
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
if m.live != "" && m.live != "none" {
|
||||||
|
continue // live or upcoming broadcast
|
||||||
|
}
|
||||||
|
if m.seconds > 0 && m.seconds < a.cfg.MinVideoSeconds {
|
||||||
|
continue // Short / sub-threshold clip
|
||||||
|
}
|
||||||
|
kept = append(kept, v)
|
||||||
|
}
|
||||||
|
return kept
|
||||||
|
}
|
||||||
|
|
||||||
|
// parseISO8601Seconds parses an ISO 8601 duration as returned by the YouTube Data
|
||||||
|
// API (e.g. "PT1H2M3S", "PT45S", "PT3M") into seconds. Only the hour/minute/second
|
||||||
|
// components YouTube emits are handled; an unparseable or zero value returns 0,
|
||||||
|
// which the caller treats as "unknown" (not filtered on duration).
|
||||||
|
func parseISO8601Seconds(d string) int {
|
||||||
|
if !strings.HasPrefix(d, "PT") {
|
||||||
|
return 0
|
||||||
|
}
|
||||||
|
d = d[2:]
|
||||||
|
total, num := 0, 0
|
||||||
|
seen := false
|
||||||
|
for _, r := range d {
|
||||||
|
switch {
|
||||||
|
case r >= '0' && r <= '9':
|
||||||
|
num = num*10 + int(r-'0')
|
||||||
|
seen = true
|
||||||
|
case r == 'H':
|
||||||
|
total += num * 3600
|
||||||
|
num, seen = 0, false
|
||||||
|
case r == 'M':
|
||||||
|
total += num * 60
|
||||||
|
num, seen = 0, false
|
||||||
|
case r == 'S':
|
||||||
|
total += num
|
||||||
|
num, seen = 0, false
|
||||||
|
default:
|
||||||
|
return 0 // unexpected component (days/weeks) — treat as unknown
|
||||||
|
}
|
||||||
|
}
|
||||||
|
if seen {
|
||||||
|
return 0 // trailing digits without a unit: malformed
|
||||||
|
}
|
||||||
|
return total
|
||||||
|
}
|
||||||
|
|
||||||
|
// VideoByID fetches a single video's metadata (videos.list, snippet) for an
|
||||||
|
// arbitrary video id — including channels the user does not follow (paste-a-URL,
|
||||||
|
// Feature 2). This is a Data API call (1 quota unit), NOT the rate-limited
|
||||||
|
// caption path, so it is not gated: only the later transcript fetch goes through
|
||||||
|
// globalFetchGate. UserID is set on the result and SubscriptionID is left empty
|
||||||
|
// (a pasted video has no subscription parent). Returns ErrVideoNotFound when the
|
||||||
|
// id resolves to no video.
|
||||||
|
func (a *Adapter) VideoByID(ctx context.Context, userID, videoID string) (domain.Video, error) {
|
||||||
|
client, err := a.httpClient(ctx, a.cfg.TokenSecretRef)
|
||||||
|
if err != nil {
|
||||||
|
return domain.Video{}, err
|
||||||
|
}
|
||||||
|
q := url.Values{"part": {"snippet"}, "id": {videoID}}
|
||||||
|
var resp videoListResponse
|
||||||
|
if err := a.getJSON(ctx, client, "/videos", q, &resp); err != nil {
|
||||||
|
return domain.Video{}, fmt.Errorf("video by id %q: %w", videoID, err)
|
||||||
|
}
|
||||||
|
if len(resp.Items) == 0 {
|
||||||
|
return domain.Video{}, fmt.Errorf("video %q: %w", videoID, domain.ErrVideoNotFound)
|
||||||
|
}
|
||||||
|
it := resp.Items[0]
|
||||||
|
return domain.Video{
|
||||||
|
UserID: userID,
|
||||||
|
Provider: domain.ProviderYouTube,
|
||||||
|
ProviderVideoID: videoID,
|
||||||
|
Title: it.Snippet.Title,
|
||||||
|
ChannelTitle: it.Snippet.ChannelTitle,
|
||||||
|
URL: "https://www.youtube.com/watch?v=" + videoID,
|
||||||
|
PublishedAt: it.Snippet.PublishedAt,
|
||||||
|
}, nil
|
||||||
}
|
}
|
||||||
|
|
||||||
// uploadsPlaylistID derives a channel's uploads playlist id at zero API cost:
|
// uploadsPlaylistID derives a channel's uploads playlist id at zero API cost:
|
||||||
@@ -342,6 +472,21 @@ type playlistItemListResponse struct {
|
|||||||
} `json:"items"`
|
} `json:"items"`
|
||||||
}
|
}
|
||||||
|
|
||||||
|
type videoListResponse struct {
|
||||||
|
Items []struct {
|
||||||
|
ID string `json:"id"`
|
||||||
|
Snippet struct {
|
||||||
|
Title string `json:"title"`
|
||||||
|
ChannelTitle string `json:"channelTitle"`
|
||||||
|
PublishedAt time.Time `json:"publishedAt"`
|
||||||
|
LiveBroadcastContent string `json:"liveBroadcastContent"`
|
||||||
|
} `json:"snippet"`
|
||||||
|
ContentDetails struct {
|
||||||
|
Duration string `json:"duration"` // ISO 8601, e.g. "PT1M30S"
|
||||||
|
} `json:"contentDetails"`
|
||||||
|
} `json:"items"`
|
||||||
|
}
|
||||||
|
|
||||||
type channelListResponse struct {
|
type channelListResponse struct {
|
||||||
Items []struct {
|
Items []struct {
|
||||||
ContentDetails struct {
|
ContentDetails struct {
|
||||||
|
|||||||
@@ -131,7 +131,7 @@ func TestNewVideos(t *testing.T) {
|
|||||||
}`))
|
}`))
|
||||||
})
|
})
|
||||||
|
|
||||||
sub := domain.Subscription{ID: "s1", UserID: "u1", ChannelID: "UC_acme"}
|
sub := domain.Subscription{ID: "s1", UserID: "u1", ChannelID: "UC_acme", ChannelTitle: "Acme Channel"}
|
||||||
vids, err := a.NewVideos(context.Background(), sub)
|
vids, err := a.NewVideos(context.Background(), sub)
|
||||||
if err != nil {
|
if err != nil {
|
||||||
t.Fatalf("NewVideos: %v", err)
|
t.Fatalf("NewVideos: %v", err)
|
||||||
@@ -143,6 +143,9 @@ func TestNewVideos(t *testing.T) {
|
|||||||
if v.ProviderVideoID != "vid1" || v.Title != "Designing for Attention" {
|
if v.ProviderVideoID != "vid1" || v.Title != "Designing for Attention" {
|
||||||
t.Errorf("unexpected video: %+v", v)
|
t.Errorf("unexpected video: %+v", v)
|
||||||
}
|
}
|
||||||
|
if v.ChannelTitle != "Acme Channel" {
|
||||||
|
t.Errorf("ChannelTitle = %q, want Acme Channel", v.ChannelTitle)
|
||||||
|
}
|
||||||
if v.Provider != domain.ProviderYouTube || v.URL != "https://www.youtube.com/watch?v=vid1" {
|
if v.Provider != domain.ProviderYouTube || v.URL != "https://www.youtube.com/watch?v=vid1" {
|
||||||
t.Errorf("video not wired correctly: %+v", v)
|
t.Errorf("video not wired correctly: %+v", v)
|
||||||
}
|
}
|
||||||
@@ -183,6 +186,91 @@ func TestNewVideosCapsAtMax(t *testing.T) {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// TestNewVideosFiltersShortsAndLive: with MinVideoSeconds set, discovery enriches
|
||||||
|
// candidates via videos.list and drops sub-threshold clips (Shorts) and
|
||||||
|
// live/upcoming broadcasts before they reach the rate-limited caption path.
|
||||||
|
func TestNewVideosFiltersShortsAndLive(t *testing.T) {
|
||||||
|
a, _ := newTestAdapter(t, func(w http.ResponseWriter, r *http.Request) {
|
||||||
|
switch r.URL.Path {
|
||||||
|
case "/playlistItems":
|
||||||
|
_, _ = w.Write([]byte(`{
|
||||||
|
"items": [
|
||||||
|
{"snippet": {"title": "Real Talk", "publishedAt": "2026-06-03T10:00:00Z", "resourceId": {"videoId": "long1"}}},
|
||||||
|
{"snippet": {"title": "A Short", "publishedAt": "2026-06-03T09:00:00Z", "resourceId": {"videoId": "short1"}}},
|
||||||
|
{"snippet": {"title": "Live Now", "publishedAt": "2026-06-03T08:00:00Z", "resourceId": {"videoId": "live1"}}}
|
||||||
|
]
|
||||||
|
}`))
|
||||||
|
case "/videos":
|
||||||
|
if got := r.URL.Query().Get("part"); got != "contentDetails,snippet" {
|
||||||
|
t.Errorf("videos.list part=%q, want contentDetails,snippet", got)
|
||||||
|
}
|
||||||
|
_, _ = w.Write([]byte(`{
|
||||||
|
"items": [
|
||||||
|
{"id": "long1", "contentDetails": {"duration": "PT12M30S"}, "snippet": {"liveBroadcastContent": "none"}},
|
||||||
|
{"id": "short1", "contentDetails": {"duration": "PT45S"}, "snippet": {"liveBroadcastContent": "none"}},
|
||||||
|
{"id": "live1", "contentDetails": {"duration": "PT0S"}, "snippet": {"liveBroadcastContent": "live"}}
|
||||||
|
]
|
||||||
|
}`))
|
||||||
|
default:
|
||||||
|
t.Errorf("unexpected path %q", r.URL.Path)
|
||||||
|
}
|
||||||
|
})
|
||||||
|
a.cfg.MinVideoSeconds = 60
|
||||||
|
|
||||||
|
vids, err := a.NewVideos(context.Background(), domain.Subscription{ID: "s1", UserID: "u1", ChannelID: "UC_acme"})
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("NewVideos: %v", err)
|
||||||
|
}
|
||||||
|
if len(vids) != 1 || vids[0].ProviderVideoID != "long1" {
|
||||||
|
t.Fatalf("expected only long1 to survive the filter, got %+v", vids)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// TestNewVideosNoFilterWhenDisabled: MinVideoSeconds=0 keeps the pre-ADR-023
|
||||||
|
// behaviour — no videos.list call, no filtering.
|
||||||
|
func TestNewVideosNoFilterWhenDisabled(t *testing.T) {
|
||||||
|
a, _ := newTestAdapter(t, func(w http.ResponseWriter, r *http.Request) {
|
||||||
|
if r.URL.Path == "/videos" {
|
||||||
|
t.Errorf("videos.list must not be called when MinVideoSeconds is 0")
|
||||||
|
}
|
||||||
|
_, _ = w.Write([]byte(`{"items": [
|
||||||
|
{"snippet": {"title": "A Short", "publishedAt": "2026-06-03T09:00:00Z", "resourceId": {"videoId": "short1"}}}
|
||||||
|
]}`))
|
||||||
|
})
|
||||||
|
a.cfg.MinVideoSeconds = 0
|
||||||
|
|
||||||
|
vids, err := a.NewVideos(context.Background(), domain.Subscription{ID: "s1", UserID: "u1", ChannelID: "UC_acme"})
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("NewVideos: %v", err)
|
||||||
|
}
|
||||||
|
if len(vids) != 1 {
|
||||||
|
t.Fatalf("filter disabled must keep all videos, got %d", len(vids))
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestParseISO8601Seconds(t *testing.T) {
|
||||||
|
cases := []struct {
|
||||||
|
in string
|
||||||
|
want int
|
||||||
|
}{
|
||||||
|
{"PT45S", 45},
|
||||||
|
{"PT1M30S", 90},
|
||||||
|
{"PT3M", 180},
|
||||||
|
{"PT1H2M3S", 3723},
|
||||||
|
{"PT2H", 7200},
|
||||||
|
{"PT0S", 0},
|
||||||
|
{"", 0},
|
||||||
|
{"garbage", 0},
|
||||||
|
{"P1D", 0}, // days component not handled → unknown
|
||||||
|
{"PT10", 0}, // trailing digits without a unit → malformed
|
||||||
|
}
|
||||||
|
for _, c := range cases {
|
||||||
|
if got := parseISO8601Seconds(c.in); got != c.want {
|
||||||
|
t.Errorf("parseISO8601Seconds(%q) = %d, want %d", c.in, got, c.want)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
// TestUploadsPlaylistID covers the zero-cost UC->UU derivation, including
|
// TestUploadsPlaylistID covers the zero-cost UC->UU derivation, including
|
||||||
// non-standard ids that must fall through unchanged (handled via fallback).
|
// non-standard ids that must fall through unchanged (handled via fallback).
|
||||||
func TestUploadsPlaylistID(t *testing.T) {
|
func TestUploadsPlaylistID(t *testing.T) {
|
||||||
|
|||||||
+135
-12
@@ -13,6 +13,7 @@ import (
|
|||||||
"os"
|
"os"
|
||||||
"path/filepath"
|
"path/filepath"
|
||||||
"sort"
|
"sort"
|
||||||
|
"strconv"
|
||||||
"strings"
|
"strings"
|
||||||
"time"
|
"time"
|
||||||
)
|
)
|
||||||
@@ -27,8 +28,38 @@ type Config struct {
|
|||||||
GatewayURL string
|
GatewayURL string
|
||||||
// GatewayKey authorizes the gateway. Read from env, never committed.
|
// GatewayKey authorizes the gateway. Read from env, never committed.
|
||||||
GatewayKey string
|
GatewayKey string
|
||||||
// SummarizerModel is the alias in host/name form, e.g. "koala/phi4-mini".
|
// SummarizerModel is the primary summarizer alias in host/name form, tried
|
||||||
|
// first on every video, e.g. "koala/phi4-mini".
|
||||||
SummarizerModel string
|
SummarizerModel string
|
||||||
|
// FallbackModel is the LOCAL fallback alias tried when the primary fails or
|
||||||
|
// returns unparseable output (ADR-022). Kept local so content stays on the
|
||||||
|
// homelab stack. Empty disables it. Default a bigger-context local model.
|
||||||
|
FallbackModel string
|
||||||
|
// CloudFallbackModel is the worst-case EXTERNAL fallback alias, tried only
|
||||||
|
// after every local endpoint has failed (ADR-022). For client deployments set
|
||||||
|
// this empty so content never leaves the local stack. Default a berget alias.
|
||||||
|
CloudFallbackModel string
|
||||||
|
// SummaryMaxTokens caps the completion budget per summary call. Small-context
|
||||||
|
// models (koala/phi4-mini, 8k) overflow when prompt + max_tokens exceeds the
|
||||||
|
// window; a summary needs only a few hundred tokens, so the default is small.
|
||||||
|
SummaryMaxTokens int
|
||||||
|
// MaxTranscriptChars bounds the transcript text sent to the model so a long
|
||||||
|
// transcript does not overflow a small-context primary. 0 disables truncation.
|
||||||
|
MaxTranscriptChars int
|
||||||
|
|
||||||
|
// MinVideoSeconds drops videos shorter than this from discovery (Shorts and
|
||||||
|
// other sub-minute clips that are noise and waste the scarce caption-fetch
|
||||||
|
// budget, ADR-014/ADR-023). Enforced via a cheap Data API videos.list lookup at
|
||||||
|
// discovery, never the rate-limited caption path. 0 disables the filter.
|
||||||
|
MinVideoSeconds int
|
||||||
|
|
||||||
|
// ChannelCaptionlessThreshold is how many consecutive no-caption results a
|
||||||
|
// channel may yield before its videos are suppressed from caption fetching
|
||||||
|
// (ADR-024). 0 disables the per-channel caption memory entirely.
|
||||||
|
ChannelCaptionlessThreshold int
|
||||||
|
// ChannelCaptionlessWindow is how long a suppressed channel stays suppressed
|
||||||
|
// before one video is re-probed (auto-recovery for a channel that adds captions).
|
||||||
|
ChannelCaptionlessWindow time.Duration
|
||||||
// SummarizerTimeout bounds a single completion call. Thinking models are
|
// SummarizerTimeout bounds a single completion call. Thinking models are
|
||||||
// slow, so the default is generous.
|
// slow, so the default is generous.
|
||||||
SummarizerTimeout time.Duration
|
SummarizerTimeout time.Duration
|
||||||
@@ -79,6 +110,13 @@ type Config struct {
|
|||||||
// pre-recency behaviour). Default ~7 days.
|
// pre-recency behaviour). Default ~7 days.
|
||||||
AutoSummarizeWindow time.Duration
|
AutoSummarizeWindow time.Duration
|
||||||
|
|
||||||
|
// OnboardSummarizeCount caps how many of a freshly-connected user's newest
|
||||||
|
// videos are summarized immediately on connect (the onboarding "it works"
|
||||||
|
// burst). HARD-capped at maxOnboardSummarizeCount so onboarding can never
|
||||||
|
// bulk-fetch; 0 disables the burst. Every fetch still flows through the shared
|
||||||
|
// caption rate gate (ADR-014) — the cap bounds count, never the pacing. Default 3.
|
||||||
|
OnboardSummarizeCount int
|
||||||
|
|
||||||
// DiscoveryInterval, when > 0, makes `serve` run in-process scheduled discovery
|
// DiscoveryInterval, when > 0, makes `serve` run in-process scheduled discovery
|
||||||
// for ALL users on that cadence (ADR-018). Zero/unset = disabled, so dev and
|
// for ALL users on that cadence (ADR-018). Zero/unset = disabled, so dev and
|
||||||
// tests never auto-fetch. Single-replica assumption — see cmdServe.
|
// tests never auto-fetch. Single-replica assumption — see cmdServe.
|
||||||
@@ -108,17 +146,26 @@ func (c Config) DexConfigured() bool { return strings.TrimSpace(c.OIDCIssuer) !=
|
|||||||
|
|
||||||
// Defaults (see docs/homelab-integration.md). All overridable via env.
|
// Defaults (see docs/homelab-integration.md). All overridable via env.
|
||||||
const (
|
const (
|
||||||
defaultGatewayURL = "http://koala:30401/v1"
|
defaultGatewayURL = "http://koala:30401/v1"
|
||||||
defaultSummarizerModel = "koala/phi4-mini"
|
defaultSummarizerModel = "koala/phi4-mini"
|
||||||
defaultSummarizerTimeout = 5 * time.Minute
|
defaultFallbackModel = "koala/phi4-14b"
|
||||||
defaultYTTokenRef = "youtube/refresh_token"
|
defaultCloudFallbackModel = "berget/mistral-small"
|
||||||
defaultYTConnectRedirectURL = "https://tapir.d-ma.be/oauth/youtube/callback"
|
defaultSummaryMaxTokens = 1500
|
||||||
defaultOAuthRedirectAddr = "localhost:8080"
|
defaultMaxTranscriptChars = 18000
|
||||||
defaultHTTPAddr = ":8080"
|
defaultMinVideoSeconds = 60
|
||||||
defaultFetchBackoff = time.Hour
|
defaultCaptionlessThreshold = 5
|
||||||
defaultFetchRate = 2 * time.Second
|
defaultCaptionlessWindow = 14 * 24 * time.Hour
|
||||||
defaultPublicURL = "https://tapir.d-ma.be"
|
defaultSummarizerTimeout = 5 * time.Minute
|
||||||
defaultAutoSummarizeWindow = 7 * 24 * time.Hour
|
defaultYTTokenRef = "youtube/refresh_token"
|
||||||
|
defaultYTConnectRedirectURL = "https://tapir.d-ma.be/oauth/youtube/callback"
|
||||||
|
defaultOAuthRedirectAddr = "localhost:8080"
|
||||||
|
defaultHTTPAddr = ":8080"
|
||||||
|
defaultFetchBackoff = time.Hour
|
||||||
|
defaultFetchRate = 2 * time.Second
|
||||||
|
defaultPublicURL = "https://tapir.d-ma.be"
|
||||||
|
defaultAutoSummarizeWindow = 7 * 24 * time.Hour
|
||||||
|
defaultOnboardSummarizeCount = 3
|
||||||
|
maxOnboardSummarizeCount = 5
|
||||||
)
|
)
|
||||||
|
|
||||||
// Load reads the environment into a Config, applying defaults. It does not
|
// Load reads the environment into a Config, applying defaults. It does not
|
||||||
@@ -131,6 +178,8 @@ func Load() (Config, error) {
|
|||||||
GatewayURL: envOr("TAPIR_GATEWAY_URL", defaultGatewayURL),
|
GatewayURL: envOr("TAPIR_GATEWAY_URL", defaultGatewayURL),
|
||||||
GatewayKey: os.Getenv("TAPIR_GATEWAY_KEY"),
|
GatewayKey: os.Getenv("TAPIR_GATEWAY_KEY"),
|
||||||
SummarizerModel: envOr("TAPIR_SUMMARIZER_MODEL", defaultSummarizerModel),
|
SummarizerModel: envOr("TAPIR_SUMMARIZER_MODEL", defaultSummarizerModel),
|
||||||
|
FallbackModel: lookupOr("TAPIR_FALLBACK_MODEL", defaultFallbackModel),
|
||||||
|
CloudFallbackModel: lookupOr("TAPIR_CLOUD_FALLBACK_MODEL", defaultCloudFallbackModel),
|
||||||
DBDSN: os.Getenv("TAPIR_DB_DSN"),
|
DBDSN: os.Getenv("TAPIR_DB_DSN"),
|
||||||
YTClientID: os.Getenv("TAPIR_YT_CLIENT_ID"),
|
YTClientID: os.Getenv("TAPIR_YT_CLIENT_ID"),
|
||||||
YTClientSecret: os.Getenv("TAPIR_YT_CLIENT_SECRET"),
|
YTClientSecret: os.Getenv("TAPIR_YT_CLIENT_SECRET"),
|
||||||
@@ -183,6 +232,57 @@ func Load() (Config, error) {
|
|||||||
}
|
}
|
||||||
c.AutoSummarizeWindow = autoWindow
|
c.AutoSummarizeWindow = autoWindow
|
||||||
|
|
||||||
|
summaryTokens, err := intOr("TAPIR_SUMMARY_MAX_TOKENS", defaultSummaryMaxTokens)
|
||||||
|
if err != nil {
|
||||||
|
return Config{}, err
|
||||||
|
}
|
||||||
|
c.SummaryMaxTokens = summaryTokens
|
||||||
|
|
||||||
|
maxChars, err := intOr("TAPIR_MAX_TRANSCRIPT_CHARS", defaultMaxTranscriptChars)
|
||||||
|
if err != nil {
|
||||||
|
return Config{}, err
|
||||||
|
}
|
||||||
|
if maxChars < 0 {
|
||||||
|
maxChars = 0
|
||||||
|
}
|
||||||
|
c.MaxTranscriptChars = maxChars
|
||||||
|
|
||||||
|
minVideo, err := intOr("TAPIR_MIN_VIDEO_SECONDS", defaultMinVideoSeconds)
|
||||||
|
if err != nil {
|
||||||
|
return Config{}, err
|
||||||
|
}
|
||||||
|
if minVideo < 0 {
|
||||||
|
minVideo = 0
|
||||||
|
}
|
||||||
|
c.MinVideoSeconds = minVideo
|
||||||
|
|
||||||
|
captionThreshold, err := intOr("TAPIR_CHANNEL_CAPTIONLESS_THRESHOLD", defaultCaptionlessThreshold)
|
||||||
|
if err != nil {
|
||||||
|
return Config{}, err
|
||||||
|
}
|
||||||
|
if captionThreshold < 0 {
|
||||||
|
captionThreshold = 0
|
||||||
|
}
|
||||||
|
c.ChannelCaptionlessThreshold = captionThreshold
|
||||||
|
|
||||||
|
captionWindow, err := durationOr("TAPIR_CHANNEL_CAPTIONLESS_WINDOW", defaultCaptionlessWindow)
|
||||||
|
if err != nil {
|
||||||
|
return Config{}, err
|
||||||
|
}
|
||||||
|
c.ChannelCaptionlessWindow = captionWindow
|
||||||
|
|
||||||
|
onboard, err := intOr("TAPIR_ONBOARD_SUMMARIZE_COUNT", defaultOnboardSummarizeCount)
|
||||||
|
if err != nil {
|
||||||
|
return Config{}, err
|
||||||
|
}
|
||||||
|
if onboard < 0 {
|
||||||
|
onboard = 0
|
||||||
|
}
|
||||||
|
if onboard > maxOnboardSummarizeCount {
|
||||||
|
onboard = maxOnboardSummarizeCount
|
||||||
|
}
|
||||||
|
c.OnboardSummarizeCount = onboard
|
||||||
|
|
||||||
return c, nil
|
return c, nil
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -247,6 +347,29 @@ func envOr(key, fallback string) string {
|
|||||||
return fallback
|
return fallback
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// lookupOr returns the env value when the key is PRESENT (even if empty), else
|
||||||
|
// fallback. Unlike envOr it lets an explicit empty value override the default —
|
||||||
|
// needed to DISABLE an optional fallback model (e.g. set the cloud fallback empty
|
||||||
|
// for a client deployment so content never leaves the local stack).
|
||||||
|
func lookupOr(key, fallback string) string {
|
||||||
|
if v, ok := os.LookupEnv(key); ok {
|
||||||
|
return v
|
||||||
|
}
|
||||||
|
return fallback
|
||||||
|
}
|
||||||
|
|
||||||
|
func intOr(key string, fallback int) (int, error) {
|
||||||
|
v := os.Getenv(key)
|
||||||
|
if v == "" {
|
||||||
|
return fallback, nil
|
||||||
|
}
|
||||||
|
n, err := strconv.Atoi(v)
|
||||||
|
if err != nil {
|
||||||
|
return 0, fmt.Errorf("config: %s=%q: %w", key, v, err)
|
||||||
|
}
|
||||||
|
return n, nil
|
||||||
|
}
|
||||||
|
|
||||||
func durationOr(key string, fallback time.Duration) (time.Duration, error) {
|
func durationOr(key string, fallback time.Duration) (time.Duration, error) {
|
||||||
v := os.Getenv(key)
|
v := os.Getenv(key)
|
||||||
if v == "" {
|
if v == "" {
|
||||||
|
|||||||
@@ -1,11 +1,29 @@
|
|||||||
package config
|
package config
|
||||||
|
|
||||||
import (
|
import (
|
||||||
|
"os"
|
||||||
"strings"
|
"strings"
|
||||||
"testing"
|
"testing"
|
||||||
"time"
|
"time"
|
||||||
)
|
)
|
||||||
|
|
||||||
|
// unset removes an env key for the duration of the test, restoring it after.
|
||||||
|
// Needed to observe a default for a key read with LookupEnv (where present-empty
|
||||||
|
// means "explicitly disabled", not "use default").
|
||||||
|
func unset(t *testing.T, key string) {
|
||||||
|
t.Helper()
|
||||||
|
if old, ok := os.LookupEnv(key); ok {
|
||||||
|
t.Cleanup(func() {
|
||||||
|
if err := os.Setenv(key, old); err != nil {
|
||||||
|
t.Fatalf("restore %s: %v", key, err)
|
||||||
|
}
|
||||||
|
})
|
||||||
|
}
|
||||||
|
if err := os.Unsetenv(key); err != nil {
|
||||||
|
t.Fatalf("unset %s: %v", key, err)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
// setEnv sets env vars for the test and clears them afterward, so cases don't
|
// setEnv sets env vars for the test and clears them afterward, so cases don't
|
||||||
// leak into one another. t.Setenv handles restoration.
|
// leak into one another. t.Setenv handles restoration.
|
||||||
func setEnv(t *testing.T, kv map[string]string) {
|
func setEnv(t *testing.T, kv map[string]string) {
|
||||||
@@ -53,6 +71,78 @@ func TestLoad_AppliesDefaults(t *testing.T) {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
func TestLoad_SummarizerChainDefaults(t *testing.T) {
|
||||||
|
setEnv(t, map[string]string{
|
||||||
|
"TAPIR_SUMMARIZER_MODEL": "",
|
||||||
|
"TAPIR_SUMMARY_MAX_TOKENS": "",
|
||||||
|
"TAPIR_MAX_TRANSCRIPT_CHARS": "",
|
||||||
|
})
|
||||||
|
unset(t, "TAPIR_FALLBACK_MODEL")
|
||||||
|
unset(t, "TAPIR_CLOUD_FALLBACK_MODEL")
|
||||||
|
|
||||||
|
c, err := Load()
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("Load: %v", err)
|
||||||
|
}
|
||||||
|
if c.FallbackModel != defaultFallbackModel {
|
||||||
|
t.Errorf("FallbackModel = %q, want %q", c.FallbackModel, defaultFallbackModel)
|
||||||
|
}
|
||||||
|
if c.CloudFallbackModel != defaultCloudFallbackModel {
|
||||||
|
t.Errorf("CloudFallbackModel = %q, want %q", c.CloudFallbackModel, defaultCloudFallbackModel)
|
||||||
|
}
|
||||||
|
if c.SummaryMaxTokens != defaultSummaryMaxTokens {
|
||||||
|
t.Errorf("SummaryMaxTokens = %d, want %d", c.SummaryMaxTokens, defaultSummaryMaxTokens)
|
||||||
|
}
|
||||||
|
if c.MaxTranscriptChars != defaultMaxTranscriptChars {
|
||||||
|
t.Errorf("MaxTranscriptChars = %d, want %d", c.MaxTranscriptChars, defaultMaxTranscriptChars)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// An explicitly empty cloud-fallback env disables external routing — the lever a
|
||||||
|
// client deployment pulls so content never leaves the local stack.
|
||||||
|
func TestLoad_EmptyCloudFallbackDisables(t *testing.T) {
|
||||||
|
t.Setenv("TAPIR_CLOUD_FALLBACK_MODEL", "")
|
||||||
|
c, err := Load()
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("Load: %v", err)
|
||||||
|
}
|
||||||
|
if c.CloudFallbackModel != "" {
|
||||||
|
t.Errorf("CloudFallbackModel = %q, want empty (disabled)", c.CloudFallbackModel)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestLoad_OnboardSummarizeCount(t *testing.T) {
|
||||||
|
cases := []struct {
|
||||||
|
name, env string
|
||||||
|
want int
|
||||||
|
}{
|
||||||
|
{"default", "", defaultOnboardSummarizeCount},
|
||||||
|
{"explicit", "4", 4},
|
||||||
|
{"zero disables", "0", 0},
|
||||||
|
{"clamped to hard cap", "50", maxOnboardSummarizeCount},
|
||||||
|
{"negative clamps to zero", "-3", 0},
|
||||||
|
}
|
||||||
|
for _, c := range cases {
|
||||||
|
t.Run(c.name, func(t *testing.T) {
|
||||||
|
setEnv(t, map[string]string{"TAPIR_ONBOARD_SUMMARIZE_COUNT": c.env})
|
||||||
|
cfg, err := Load()
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("Load: %v", err)
|
||||||
|
}
|
||||||
|
if cfg.OnboardSummarizeCount != c.want {
|
||||||
|
t.Fatalf("OnboardSummarizeCount = %d, want %d", cfg.OnboardSummarizeCount, c.want)
|
||||||
|
}
|
||||||
|
})
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestLoad_OnboardSummarizeCountInvalid(t *testing.T) {
|
||||||
|
setEnv(t, map[string]string{"TAPIR_ONBOARD_SUMMARIZE_COUNT": "three"})
|
||||||
|
if _, err := Load(); err == nil {
|
||||||
|
t.Fatal("Load: want error for non-numeric TAPIR_ONBOARD_SUMMARIZE_COUNT")
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
func TestLoad_ParsesValues(t *testing.T) {
|
func TestLoad_ParsesValues(t *testing.T) {
|
||||||
setEnv(t, map[string]string{
|
setEnv(t, map[string]string{
|
||||||
"TAPIR_USER_ID": "11111111-1111-1111-1111-111111111111",
|
"TAPIR_USER_ID": "11111111-1111-1111-1111-111111111111",
|
||||||
|
|||||||
@@ -3,10 +3,16 @@
|
|||||||
package domain
|
package domain
|
||||||
|
|
||||||
import (
|
import (
|
||||||
|
"errors"
|
||||||
"fmt"
|
"fmt"
|
||||||
"time"
|
"time"
|
||||||
)
|
)
|
||||||
|
|
||||||
|
// ErrVideoNotFound is returned when a video id resolves to no video (deleted,
|
||||||
|
// private, or a typo'd paste). Defined in domain so adapters and the web layer
|
||||||
|
// share one sentinel without coupling to each other.
|
||||||
|
var ErrVideoNotFound = errors.New("video not found")
|
||||||
|
|
||||||
// ErrChannelUnavailable is returned by a VideoSource when a channel's upload
|
// ErrChannelUnavailable is returned by a VideoSource when a channel's upload
|
||||||
// playlist returns HTTP 404 — the channel was deleted or made private. The runner
|
// playlist returns HTTP 404 — the channel was deleted or made private. The runner
|
||||||
// stores these so the account page can surface them to the user.
|
// stores these so the account page can surface them to the user.
|
||||||
@@ -67,6 +73,7 @@ type Video struct {
|
|||||||
Provider Provider
|
Provider Provider
|
||||||
ProviderVideoID string
|
ProviderVideoID string
|
||||||
Title string
|
Title string
|
||||||
|
ChannelTitle string
|
||||||
URL string
|
URL string
|
||||||
PublishedAt time.Time
|
PublishedAt time.Time
|
||||||
SeenAt time.Time
|
SeenAt time.Time
|
||||||
|
|||||||
@@ -27,6 +27,26 @@ type Summarizer interface {
|
|||||||
Summarize(ctx context.Context, v domain.Video, t domain.Transcript) (domain.Summary, error)
|
Summarize(ctx context.Context, v domain.Video, t domain.Transcript) (domain.Summary, error)
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// TranscriptStore persists transcripts as shared, video-keyed public content
|
||||||
|
// (ADR-021). It is keyed by the cross-user dedup key (provider, providerVideoID)
|
||||||
|
// — the video's public identity, NOT Tapir's per-user videos.id — and holds only
|
||||||
|
// public caption content, so it is deliberately NOT user-scoped: two users who
|
||||||
|
// share a video share the one row. The engine reads it before any caption fetch
|
||||||
|
// so re-analysis never re-touches YouTube (ADR-010/014).
|
||||||
|
type TranscriptStore interface {
|
||||||
|
// GetTranscript returns the stored transcript for a video and whether one
|
||||||
|
// exists. A stored Source == SourceNone (captions permanently absent) is a
|
||||||
|
// real hit: ok is true and HasText() is false, so callers skip without
|
||||||
|
// re-fetching. A transient rate-limit is never stored, so it never appears
|
||||||
|
// here as a false absence.
|
||||||
|
GetTranscript(ctx context.Context, provider, providerVideoID string) (t domain.Transcript, ok bool, err error)
|
||||||
|
// SaveTranscript upserts the transcript for (provider, providerVideoID). Only
|
||||||
|
// terminal outcomes are persisted: SourceCaptions (with text) or SourceNone.
|
||||||
|
// SourceRateLimited must NOT be passed — it is a per-user retry (ADR-014), not
|
||||||
|
// a shared terminal state.
|
||||||
|
SaveTranscript(ctx context.Context, provider, providerVideoID string, t domain.Transcript) error
|
||||||
|
}
|
||||||
|
|
||||||
// Sink delivers a summary to a destination (user store, brain, ...).
|
// Sink delivers a summary to a destination (user store, brain, ...).
|
||||||
// Implementations fail independently of one another.
|
// Implementations fail independently of one another.
|
||||||
type Sink interface {
|
type Sink interface {
|
||||||
|
|||||||
+63
-12
@@ -27,8 +27,9 @@ import (
|
|||||||
// passCandidate is a video that passed all pre-filters (seen/manual/backoff)
|
// passCandidate is a video that passed all pre-filters (seen/manual/backoff)
|
||||||
// and is queued for transcript fetch + summarization in this pass.
|
// and is queued for transcript fetch + summarization in this pass.
|
||||||
type passCandidate struct {
|
type passCandidate struct {
|
||||||
v domain.Video
|
v domain.Video
|
||||||
pos int // discovery position — used as a stable tiebreak when published_at ties
|
channelID string // owning channel — keys the caption-availability memory (ADR-024)
|
||||||
|
pos int // discovery position — used as a stable tiebreak when published_at ties
|
||||||
}
|
}
|
||||||
|
|
||||||
// compareNewestFirst orders candidates by published_at descending, NULLS LAST,
|
// compareNewestFirst orders candidates by published_at descending, NULLS LAST,
|
||||||
@@ -74,6 +75,14 @@ type VideoStore interface {
|
|||||||
// Called when NewVideos returns domain.ErrChannelUnavailable; best-effort, errors
|
// Called when NewVideos returns domain.ErrChannelUnavailable; best-effort, errors
|
||||||
// are logged and never abort the pass.
|
// are logged and never abort the pass.
|
||||||
UpsertChannelError(ctx context.Context, userID, channelID, channelTitle string) error
|
UpsertChannelError(ctx context.Context, userID, channelID, channelTitle string) error
|
||||||
|
// CaptionlessChannels returns channel ids currently suppressed because their
|
||||||
|
// recent videos all yielded no captions (ADR-024). The loop skips caption
|
||||||
|
// fetches for these channels' (non-requested) videos.
|
||||||
|
CaptionlessChannels(ctx context.Context, userID string) (map[string]bool, error)
|
||||||
|
// RecordChannelCaptionOutcome updates a channel's caption memory after a fetch:
|
||||||
|
// hadCaptions resets it, otherwise the no-caption streak grows and the channel
|
||||||
|
// is suppressed for window once it reaches threshold. A no-op when threshold<=0.
|
||||||
|
RecordChannelCaptionOutcome(ctx context.Context, userID, channelID string, hadCaptions bool, threshold int, window time.Duration) error
|
||||||
}
|
}
|
||||||
|
|
||||||
// Processor runs the core use case for a single video. *usecase.Engine
|
// Processor runs the core use case for a single video. *usecase.Engine
|
||||||
@@ -93,6 +102,9 @@ type Runner struct {
|
|||||||
backoff time.Duration // rate-limit retry window; 0 = always retry
|
backoff time.Duration // rate-limit retry window; 0 = always retry
|
||||||
autoWindow time.Duration // recency bound for auto-summarize; 0 = no bound
|
autoWindow time.Duration // recency bound for auto-summarize; 0 = no bound
|
||||||
now func() time.Time // injectable clock (tests); defaults to time.Now
|
now func() time.Time // injectable clock (tests); defaults to time.Now
|
||||||
|
|
||||||
|
captionThreshold int // consecutive no-caption results before a channel is suppressed; 0 = feature off
|
||||||
|
captionWindow time.Duration // how long a caption-less channel stays suppressed before re-probe
|
||||||
}
|
}
|
||||||
|
|
||||||
// Option configures a Runner at construction. Variadic so existing call sites
|
// Option configures a Runner at construction. Variadic so existing call sites
|
||||||
@@ -114,6 +126,14 @@ func WithClock(now func() time.Time) Option { return func(r *Runner) { r.now = n
|
|||||||
// bypasses the bound. 0 (the default) disables it (summarize every unseen video).
|
// bypasses the bound. 0 (the default) disables it (summarize every unseen video).
|
||||||
func WithAutoWindow(d time.Duration) Option { return func(r *Runner) { r.autoWindow = d } }
|
func WithAutoWindow(d time.Duration) Option { return func(r *Runner) { r.autoWindow = d } }
|
||||||
|
|
||||||
|
// WithCaptionMemory enables per-channel caption-availability suppression
|
||||||
|
// (ADR-024): after threshold consecutive no-caption results a channel's videos
|
||||||
|
// are skipped (no caption fetch) for window, then one is re-probed. threshold<=0
|
||||||
|
// (the default) disables the feature entirely.
|
||||||
|
func WithCaptionMemory(threshold int, window time.Duration) Option {
|
||||||
|
return func(r *Runner) { r.captionThreshold = threshold; r.captionWindow = window }
|
||||||
|
}
|
||||||
|
|
||||||
// New builds a Runner. A nil logger falls back to slog.Default.
|
// New builds a Runner. A nil logger falls back to slog.Default.
|
||||||
func New(src ports.VideoSource, store VideoStore, engine Processor, userID string, log *slog.Logger, opts ...Option) *Runner {
|
func New(src ports.VideoSource, store VideoStore, engine Processor, userID string, log *slog.Logger, opts ...Option) *Runner {
|
||||||
if log == nil {
|
if log == nil {
|
||||||
@@ -131,15 +151,16 @@ func New(src ports.VideoSource, store VideoStore, engine Processor, userID strin
|
|||||||
|
|
||||||
// Stats summarizes one RunOnce pass.
|
// Stats summarizes one RunOnce pass.
|
||||||
type Stats struct {
|
type Stats struct {
|
||||||
Candidates int
|
Candidates int
|
||||||
Summarized int
|
Summarized int
|
||||||
SkippedSeen int
|
SkippedSeen int
|
||||||
SkippedNoText int
|
SkippedNoText int
|
||||||
SkippedManual int // discovered but not queued, in manual mode
|
SkippedManual int // discovered but not queued, in manual mode
|
||||||
SkippedTooOld int // auto mode: published outside the recency window (not requested)
|
SkippedTooOld int // auto mode: published outside the recency window (not requested)
|
||||||
SkippedRateLimited int // 429'd previously and still inside the backoff window
|
SkippedRateLimited int // 429'd previously and still inside the backoff window
|
||||||
Errors int
|
SkippedNoCaptionChannel int // channel suppressed as caption-less (ADR-024)
|
||||||
ChannelUnavailable int // channels that returned HTTP 404 (deleted/private)
|
Errors int
|
||||||
|
ChannelUnavailable int // channels that returned HTTP 404 (deleted/private)
|
||||||
}
|
}
|
||||||
|
|
||||||
// tooOld reports whether a video published at publishedAt falls outside the
|
// tooOld reports whether a video published at publishedAt falls outside the
|
||||||
@@ -213,6 +234,17 @@ func (r *Runner) RunOnce(ctx context.Context) (Stats, error) {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// Per-channel caption memory (ADR-024): channels whose recent videos all
|
||||||
|
// yielded no captions are suppressed so their new videos don't burn the scarce
|
||||||
|
// fetch budget. Loaded only when the feature is enabled (threshold > 0).
|
||||||
|
var captionless map[string]bool
|
||||||
|
if r.captionThreshold > 0 {
|
||||||
|
captionless, err = r.store.CaptionlessChannels(ctx, r.userID)
|
||||||
|
if err != nil {
|
||||||
|
return stats, fmt.Errorf("runner: load caption-less channels: %w", err)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
subs, err := r.src.ListSubscriptions(ctx, r.userID)
|
subs, err := r.src.ListSubscriptions(ctx, r.userID)
|
||||||
if err != nil {
|
if err != nil {
|
||||||
return stats, fmt.Errorf("runner: list subscriptions: %w", err)
|
return stats, fmt.Errorf("runner: list subscriptions: %w", err)
|
||||||
@@ -273,6 +305,14 @@ func (r *Runner) RunOnce(ctx context.Context) (Stats, error) {
|
|||||||
continue
|
continue
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// Caption-less channel (ADR-024): its recent videos all returned no
|
||||||
|
// captions, so skip the fetch entirely. The video is still listed
|
||||||
|
// (UpsertVideo above); an explicit manual request bypasses the skip.
|
||||||
|
if !requested[id] && captionless[sub.ChannelID] {
|
||||||
|
stats.SkippedNoCaptionChannel++
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
|
||||||
// Still inside the rate-limit backoff window: skip without fetching.
|
// Still inside the rate-limit backoff window: skip without fetching.
|
||||||
if at, ok := rateLimited[id]; ok && r.now().Sub(at) < r.backoff {
|
if at, ok := rateLimited[id]; ok && r.now().Sub(at) < r.backoff {
|
||||||
stats.SkippedRateLimited++
|
stats.SkippedRateLimited++
|
||||||
@@ -280,7 +320,7 @@ func (r *Runner) RunOnce(ctx context.Context) (Stats, error) {
|
|||||||
continue
|
continue
|
||||||
}
|
}
|
||||||
|
|
||||||
candidates = append(candidates, passCandidate{v: v, pos: pos})
|
candidates = append(candidates, passCandidate{v: v, channelID: sub.ChannelID, pos: pos})
|
||||||
pos++
|
pos++
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
@@ -320,6 +360,11 @@ func (r *Runner) RunOnce(ctx context.Context) (Stats, error) {
|
|||||||
errs = append(errs, fmt.Errorf("set none status %q: %w", c.v.ProviderVideoID, err))
|
errs = append(errs, fmt.Errorf("set none status %q: %w", c.v.ProviderVideoID, err))
|
||||||
stats.Errors++
|
stats.Errors++
|
||||||
}
|
}
|
||||||
|
// No captions: grow this channel's no-caption streak (ADR-024).
|
||||||
|
if err := r.store.RecordChannelCaptionOutcome(ctx, r.userID, c.channelID, false, r.captionThreshold, r.captionWindow); err != nil {
|
||||||
|
errs = append(errs, fmt.Errorf("record no-caption %q: %w", c.v.ProviderVideoID, err))
|
||||||
|
stats.Errors++
|
||||||
|
}
|
||||||
r.log.Info("skipped video (no transcript)", "video", c.v.ProviderVideoID, "title", c.v.Title)
|
r.log.Info("skipped video (no transcript)", "video", c.v.ProviderVideoID, "title", c.v.Title)
|
||||||
case res.Summary != nil:
|
case res.Summary != nil:
|
||||||
stats.Summarized++
|
stats.Summarized++
|
||||||
@@ -327,6 +372,11 @@ func (r *Runner) RunOnce(ctx context.Context) (Stats, error) {
|
|||||||
errs = append(errs, fmt.Errorf("set fetched status %q: %w", c.v.ProviderVideoID, err))
|
errs = append(errs, fmt.Errorf("set fetched status %q: %w", c.v.ProviderVideoID, err))
|
||||||
stats.Errors++
|
stats.Errors++
|
||||||
}
|
}
|
||||||
|
// Captions present: reset this channel's caption memory (ADR-024).
|
||||||
|
if err := r.store.RecordChannelCaptionOutcome(ctx, r.userID, c.channelID, true, r.captionThreshold, r.captionWindow); err != nil {
|
||||||
|
errs = append(errs, fmt.Errorf("record has-caption %q: %w", c.v.ProviderVideoID, err))
|
||||||
|
stats.Errors++
|
||||||
|
}
|
||||||
// In manual mode the video was explicitly queued; clear the flag so
|
// In manual mode the video was explicitly queued; clear the flag so
|
||||||
// it is not re-summarized and the UI drops the "Queued" chip.
|
// it is not re-summarized and the UI drops the "Queued" chip.
|
||||||
if !auto {
|
if !auto {
|
||||||
@@ -354,6 +404,7 @@ func (r *Runner) Loop(ctx context.Context, interval time.Duration) error {
|
|||||||
"skipped_seen", stats.SkippedSeen, "skipped_no_text", stats.SkippedNoText,
|
"skipped_seen", stats.SkippedSeen, "skipped_no_text", stats.SkippedNoText,
|
||||||
"skipped_manual", stats.SkippedManual, "skipped_too_old", stats.SkippedTooOld,
|
"skipped_manual", stats.SkippedManual, "skipped_too_old", stats.SkippedTooOld,
|
||||||
"skipped_rate_limited", stats.SkippedRateLimited,
|
"skipped_rate_limited", stats.SkippedRateLimited,
|
||||||
|
"skipped_no_caption_channel", stats.SkippedNoCaptionChannel,
|
||||||
"channel_unavailable", stats.ChannelUnavailable, "errors", stats.Errors)
|
"channel_unavailable", stats.ChannelUnavailable, "errors", stats.Errors)
|
||||||
if err != nil {
|
if err != nil {
|
||||||
r.log.Warn("run pass had errors", "err", err)
|
r.log.Warn("run pass had errors", "err", err)
|
||||||
|
|||||||
@@ -51,6 +51,13 @@ type fakeStore struct {
|
|||||||
cleared []string
|
cleared []string
|
||||||
rateLimited map[string]time.Time // id -> when 429'd (seeds the backoff window)
|
rateLimited map[string]time.Time // id -> when 429'd (seeds the backoff window)
|
||||||
statuses map[string]string // id -> last SetTranscriptStatus value
|
statuses map[string]string // id -> last SetTranscriptStatus value
|
||||||
|
captionless map[string]bool // channel ids currently suppressed (ADR-024)
|
||||||
|
captionRecs []captionRec // RecordChannelCaptionOutcome calls, in order
|
||||||
|
}
|
||||||
|
|
||||||
|
type captionRec struct {
|
||||||
|
channelID string
|
||||||
|
had bool
|
||||||
}
|
}
|
||||||
|
|
||||||
func (f *fakeStore) UpsertVideo(_ context.Context, v domain.Video) (string, error) {
|
func (f *fakeStore) UpsertVideo(_ context.Context, v domain.Video) (string, error) {
|
||||||
@@ -93,6 +100,22 @@ func (f *fakeStore) RateLimitedVideoIDs(_ context.Context, _ string) (map[string
|
|||||||
|
|
||||||
func (f *fakeStore) UpsertChannelError(_ context.Context, _, _, _ string) error { return nil }
|
func (f *fakeStore) UpsertChannelError(_ context.Context, _, _, _ string) error { return nil }
|
||||||
|
|
||||||
|
func (f *fakeStore) CaptionlessChannels(_ context.Context, _ string) (map[string]bool, error) {
|
||||||
|
cp := make(map[string]bool, len(f.captionless))
|
||||||
|
for k, v := range f.captionless {
|
||||||
|
cp[k] = v
|
||||||
|
}
|
||||||
|
return cp, nil
|
||||||
|
}
|
||||||
|
|
||||||
|
func (f *fakeStore) RecordChannelCaptionOutcome(_ context.Context, _, channelID string, hadCaptions bool, threshold int, _ time.Duration) error {
|
||||||
|
if threshold <= 0 {
|
||||||
|
return nil
|
||||||
|
}
|
||||||
|
f.captionRecs = append(f.captionRecs, captionRec{channelID: channelID, had: hadCaptions})
|
||||||
|
return nil
|
||||||
|
}
|
||||||
|
|
||||||
func (f *fakeStore) SetTranscriptStatus(_ context.Context, _, videoID, status string) error {
|
func (f *fakeStore) SetTranscriptStatus(_ context.Context, _, videoID, status string) error {
|
||||||
if f.statuses == nil {
|
if f.statuses == nil {
|
||||||
f.statuses = map[string]string{}
|
f.statuses = map[string]string{}
|
||||||
@@ -284,6 +307,55 @@ func TestRunOnce_AutoMode_OldVideoRequestedBypassesWindow(t *testing.T) {
|
|||||||
require.Len(t, sink.delivered, 1)
|
require.Len(t, sink.delivered, 1)
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// TestRunOnce_CaptionlessChannelSkipped: a channel flagged caption-less (ADR-024)
|
||||||
|
// has its videos skipped from fetching but still discovered/listed, while a
|
||||||
|
// normal channel's video is summarized.
|
||||||
|
func TestRunOnce_CaptionlessChannelSkipped(t *testing.T) {
|
||||||
|
src := &fakeSource{
|
||||||
|
subs: []domain.Subscription{sub("dead", "Dead Channel"), sub("live", "Live Channel")},
|
||||||
|
videos: map[string][]domain.Video{
|
||||||
|
"dead": {vid("d1", "Dead One")},
|
||||||
|
"live": {vid("l1", "Live One")},
|
||||||
|
},
|
||||||
|
}
|
||||||
|
st := &fakeStore{seen: map[string]bool{}, auto: true, captionless: map[string]bool{"dead": true}}
|
||||||
|
sink := &recordingSink{}
|
||||||
|
eng := usecase.NewEngine(src, fakeSummarizer{}, sink)
|
||||||
|
r := runner.New(src, st, eng, testUser, quietLogger(),
|
||||||
|
runner.WithCaptionMemory(5, 14*24*time.Hour))
|
||||||
|
|
||||||
|
stats, err := r.RunOnce(context.Background())
|
||||||
|
require.NoError(t, err)
|
||||||
|
require.Equal(t, 1, stats.SkippedNoCaptionChannel, "dead channel's video skipped from fetch")
|
||||||
|
require.Equal(t, 1, stats.Summarized, "live channel's video still summarized")
|
||||||
|
require.Len(t, st.upserted, 2, "both videos are still discovered and listed")
|
||||||
|
}
|
||||||
|
|
||||||
|
// TestRunOnce_RecordsCaptionOutcomes: a no-caption result grows the channel's
|
||||||
|
// streak (had=false); a successful summary resets it (had=true).
|
||||||
|
func TestRunOnce_RecordsCaptionOutcomes(t *testing.T) {
|
||||||
|
src := &fakeSource{
|
||||||
|
subs: []domain.Subscription{sub("c1", "Has Caps"), sub("c2", "No Caps")},
|
||||||
|
videos: map[string][]domain.Video{
|
||||||
|
"c1": {vid("good", "Good")},
|
||||||
|
"c2": {vid("bad", "Bad")},
|
||||||
|
},
|
||||||
|
transcripts: map[string]domain.Transcript{
|
||||||
|
"bad": {Source: domain.SourceNone}, // no usable text → engine skips
|
||||||
|
},
|
||||||
|
}
|
||||||
|
st := &fakeStore{seen: map[string]bool{}, auto: true}
|
||||||
|
sink := &recordingSink{}
|
||||||
|
eng := usecase.NewEngine(src, fakeSummarizer{}, sink)
|
||||||
|
r := runner.New(src, st, eng, testUser, quietLogger(),
|
||||||
|
runner.WithCaptionMemory(5, 14*24*time.Hour))
|
||||||
|
|
||||||
|
_, err := r.RunOnce(context.Background())
|
||||||
|
require.NoError(t, err)
|
||||||
|
require.Contains(t, st.captionRecs, captionRec{channelID: "c1", had: true}, "captioned channel reset")
|
||||||
|
require.Contains(t, st.captionRecs, captionRec{channelID: "c2", had: false}, "no-caption channel streak grown")
|
||||||
|
}
|
||||||
|
|
||||||
// TestRunOnce_AutoWindowZero_SummarizesOld: a zero window disables the bound —
|
// TestRunOnce_AutoWindowZero_SummarizesOld: a zero window disables the bound —
|
||||||
// the pre-recency behaviour (summarize every unseen video) is preserved.
|
// the pre-recency behaviour (summarize every unseen video) is preserved.
|
||||||
func TestRunOnce_AutoWindowZero_SummarizesOld(t *testing.T) {
|
func TestRunOnce_AutoWindowZero_SummarizesOld(t *testing.T) {
|
||||||
|
|||||||
@@ -27,6 +27,13 @@ type Engine struct {
|
|||||||
AI ports.Summarizer
|
AI ports.Summarizer
|
||||||
Sinks []ports.Sink
|
Sinks []ports.Sink
|
||||||
|
|
||||||
|
// Transcripts, when set, is the shared transcript cache (ADR-021): the engine
|
||||||
|
// reads it before any caption fetch and writes resolved transcripts back, so
|
||||||
|
// re-analysis — the same user re-summarizing, or a second user with the same
|
||||||
|
// video — never re-touches YouTube (ADR-010/014). Optional: nil disables
|
||||||
|
// persistence (fetch every time), keeping the pure-core/scaffold wiring valid.
|
||||||
|
Transcripts ports.TranscriptStore
|
||||||
|
|
||||||
// processed dedups videos within this engine's lifetime so a video is not
|
// processed dedups videos within this engine's lifetime so a video is not
|
||||||
// summarized twice when the watcher sees it again. Durable cross-restart
|
// summarized twice when the watcher sees it again. Durable cross-restart
|
||||||
// dedup is the store's concern (a resolved TRANSCRIPT / existing SUMMARY,
|
// dedup is the store's concern (a resolved TRANSCRIPT / existing SUMMARY,
|
||||||
@@ -57,9 +64,9 @@ type ProcessResult struct {
|
|||||||
// resolve transcript -> (summarize -> deliver) | skip.
|
// resolve transcript -> (summarize -> deliver) | skip.
|
||||||
// See docs/use-cases/summarize_new_video.feature.
|
// See docs/use-cases/summarize_new_video.feature.
|
||||||
func (e *Engine) ProcessNewVideo(ctx context.Context, v domain.Video) (ProcessResult, error) {
|
func (e *Engine) ProcessNewVideo(ctx context.Context, v domain.Video) (ProcessResult, error) {
|
||||||
t, err := e.Source.FetchTranscript(ctx, v)
|
t, err := e.resolveTranscript(ctx, v)
|
||||||
if err != nil {
|
if err != nil {
|
||||||
return ProcessResult{Video: v}, fmt.Errorf("fetch transcript: %w", err)
|
return ProcessResult{Video: v}, err
|
||||||
}
|
}
|
||||||
if !t.HasText() {
|
if !t.HasText() {
|
||||||
// No usable transcript: record the skip, produce no summary, deliver nothing
|
// No usable transcript: record the skip, produce no summary, deliver nothing
|
||||||
@@ -86,6 +93,39 @@ func (e *Engine) ProcessNewVideo(ctx context.Context, v domain.Video) (ProcessRe
|
|||||||
return ProcessResult{Video: v, Summary: &sum, TranscriptSource: string(t.Source)}, errors.Join(errs...)
|
return ProcessResult{Video: v, Summary: &sum, TranscriptSource: string(t.Source)}, errors.Join(errs...)
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// resolveTranscript returns v's transcript, reading the shared store first
|
||||||
|
// (ADR-021): a stored transcript — including a stored SourceNone (captions
|
||||||
|
// permanently absent) — is returned without touching YouTube, so re-analysis
|
||||||
|
// never re-fetches. On a store miss it fetches through the source (which gates
|
||||||
|
// the caption call, ADR-014) and persists the terminal outcome so the next
|
||||||
|
// analysis, for any user, reads from the store. A transient SourceRateLimited is
|
||||||
|
// returned to the caller (the runner stamps a per-user backoff) but never stored,
|
||||||
|
// so persistence can never mask a 429 as a permanent "no transcript". When no
|
||||||
|
// TranscriptStore is wired the engine simply fetches every time.
|
||||||
|
func (e *Engine) resolveTranscript(ctx context.Context, v domain.Video) (domain.Transcript, error) {
|
||||||
|
if e.Transcripts != nil {
|
||||||
|
stored, ok, err := e.Transcripts.GetTranscript(ctx, string(v.Provider), v.ProviderVideoID)
|
||||||
|
if err != nil {
|
||||||
|
return domain.Transcript{}, fmt.Errorf("get stored transcript: %w", err)
|
||||||
|
}
|
||||||
|
if ok {
|
||||||
|
return stored, nil
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
t, err := e.Source.FetchTranscript(ctx, v)
|
||||||
|
if err != nil {
|
||||||
|
return domain.Transcript{}, fmt.Errorf("fetch transcript: %w", err)
|
||||||
|
}
|
||||||
|
|
||||||
|
if e.Transcripts != nil && t.Source != domain.SourceRateLimited {
|
||||||
|
if err := e.Transcripts.SaveTranscript(ctx, string(v.Provider), v.ProviderVideoID, t); err != nil {
|
||||||
|
return domain.Transcript{}, fmt.Errorf("save transcript: %w", err)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return t, nil
|
||||||
|
}
|
||||||
|
|
||||||
// ProcessNewVideos walks a user's subscriptions and processes each newly seen
|
// ProcessNewVideos walks a user's subscriptions and processes each newly seen
|
||||||
// video. Only videos surfaced via the user's subscriptions are considered, so a
|
// video. Only videos surfaced via the user's subscriptions are considered, so a
|
||||||
// channel the user is not subscribed to is never processed. A video already
|
// channel the user is not subscribed to is never processed. A video already
|
||||||
|
|||||||
@@ -0,0 +1,187 @@
|
|||||||
|
package usecase
|
||||||
|
|
||||||
|
import (
|
||||||
|
"context"
|
||||||
|
"testing"
|
||||||
|
|
||||||
|
"gitea.d-ma.be/mathias/tapir/internal/domain"
|
||||||
|
)
|
||||||
|
|
||||||
|
// These tests pin the ADR-021 read-stored-first behaviour at the engine core:
|
||||||
|
// a stored transcript is summarized without re-touching the source, a miss
|
||||||
|
// fetches once and persists, and a transient rate-limit is never cached.
|
||||||
|
|
||||||
|
type recordingSource struct {
|
||||||
|
transcript domain.Transcript
|
||||||
|
fetchCalls int
|
||||||
|
}
|
||||||
|
|
||||||
|
func (s *recordingSource) ListSubscriptions(context.Context, string) ([]domain.Subscription, error) {
|
||||||
|
return nil, nil
|
||||||
|
}
|
||||||
|
|
||||||
|
func (s *recordingSource) NewVideos(context.Context, domain.Subscription) ([]domain.Video, error) {
|
||||||
|
return nil, nil
|
||||||
|
}
|
||||||
|
|
||||||
|
func (s *recordingSource) FetchTranscript(context.Context, domain.Video) (domain.Transcript, error) {
|
||||||
|
s.fetchCalls++
|
||||||
|
return s.transcript, nil
|
||||||
|
}
|
||||||
|
|
||||||
|
type fakeTranscriptStore struct {
|
||||||
|
stored map[string]domain.Transcript
|
||||||
|
saves int
|
||||||
|
}
|
||||||
|
|
||||||
|
func newFakeTranscriptStore() *fakeTranscriptStore {
|
||||||
|
return &fakeTranscriptStore{stored: make(map[string]domain.Transcript)}
|
||||||
|
}
|
||||||
|
|
||||||
|
func (f *fakeTranscriptStore) key(provider, id string) string { return provider + "|" + id }
|
||||||
|
|
||||||
|
func (f *fakeTranscriptStore) GetTranscript(_ context.Context, provider, id string) (domain.Transcript, bool, error) {
|
||||||
|
t, ok := f.stored[f.key(provider, id)]
|
||||||
|
return t, ok, nil
|
||||||
|
}
|
||||||
|
|
||||||
|
func (f *fakeTranscriptStore) SaveTranscript(_ context.Context, provider, id string, t domain.Transcript) error {
|
||||||
|
f.saves++
|
||||||
|
f.stored[f.key(provider, id)] = t
|
||||||
|
return nil
|
||||||
|
}
|
||||||
|
|
||||||
|
type countingSummarizer struct{ calls int }
|
||||||
|
|
||||||
|
func (c *countingSummarizer) Summarize(_ context.Context, v domain.Video, _ domain.Transcript) (domain.Summary, error) {
|
||||||
|
c.calls++
|
||||||
|
return domain.Summary{VideoID: v.ID, UserID: v.UserID, Summary: "s", AIProvider: "local"}, nil
|
||||||
|
}
|
||||||
|
|
||||||
|
type nopSink struct{}
|
||||||
|
|
||||||
|
func (nopSink) Name() string { return "nop" }
|
||||||
|
func (nopSink) Deliver(context.Context, domain.Summary) error { return nil }
|
||||||
|
|
||||||
|
func testVideo() domain.Video {
|
||||||
|
return domain.Video{ID: "v1", UserID: "u1", Provider: domain.ProviderYouTube, ProviderVideoID: "yt1"}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestProcessNewVideo_StoredTranscriptSkipsFetch(t *testing.T) {
|
||||||
|
src := &recordingSource{}
|
||||||
|
ts := newFakeTranscriptStore()
|
||||||
|
ts.stored[ts.key("youtube", "yt1")] = domain.Transcript{Source: domain.SourceCaptions, Content: "stored words"}
|
||||||
|
sum := &countingSummarizer{}
|
||||||
|
eng := NewEngine(src, sum, nopSink{})
|
||||||
|
eng.Transcripts = ts
|
||||||
|
|
||||||
|
res, err := eng.ProcessNewVideo(context.Background(), testVideo())
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("ProcessNewVideo: %v", err)
|
||||||
|
}
|
||||||
|
if src.fetchCalls != 0 {
|
||||||
|
t.Fatalf("stored transcript must not re-fetch from source; got %d fetches", src.fetchCalls)
|
||||||
|
}
|
||||||
|
if ts.saves != 0 {
|
||||||
|
t.Fatalf("a store hit must not re-save; got %d saves", ts.saves)
|
||||||
|
}
|
||||||
|
if sum.calls != 1 || res.Summary == nil {
|
||||||
|
t.Fatalf("expected a summary from the stored transcript; calls=%d summary=%v", sum.calls, res.Summary)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestProcessNewVideo_StoreMissFetchesAndPersists(t *testing.T) {
|
||||||
|
src := &recordingSource{transcript: domain.Transcript{Source: domain.SourceCaptions, Language: "en", Content: "fetched words"}}
|
||||||
|
ts := newFakeTranscriptStore()
|
||||||
|
sum := &countingSummarizer{}
|
||||||
|
eng := NewEngine(src, sum, nopSink{})
|
||||||
|
eng.Transcripts = ts
|
||||||
|
|
||||||
|
if _, err := eng.ProcessNewVideo(context.Background(), testVideo()); err != nil {
|
||||||
|
t.Fatalf("ProcessNewVideo: %v", err)
|
||||||
|
}
|
||||||
|
if src.fetchCalls != 1 {
|
||||||
|
t.Fatalf("a store miss must fetch exactly once; got %d", src.fetchCalls)
|
||||||
|
}
|
||||||
|
if ts.saves != 1 {
|
||||||
|
t.Fatalf("a fetched transcript must be persisted; got %d saves", ts.saves)
|
||||||
|
}
|
||||||
|
got, ok, _ := ts.GetTranscript(context.Background(), "youtube", "yt1")
|
||||||
|
if !ok || got.Content != "fetched words" {
|
||||||
|
t.Fatalf("persisted transcript not readable back: ok=%v content=%q", ok, got.Content)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// The second summarize of the same video reads the persisted transcript and does
|
||||||
|
// NOT re-fetch — the primary ADR-021 win, proven end to end at the engine.
|
||||||
|
func TestProcessNewVideo_SecondSummarizeDoesNotRefetch(t *testing.T) {
|
||||||
|
src := &recordingSource{transcript: domain.Transcript{Source: domain.SourceCaptions, Content: "words"}}
|
||||||
|
ts := newFakeTranscriptStore()
|
||||||
|
eng := NewEngine(src, &countingSummarizer{}, nopSink{})
|
||||||
|
eng.Transcripts = ts
|
||||||
|
|
||||||
|
for i := 0; i < 2; i++ {
|
||||||
|
if _, err := eng.ProcessNewVideo(context.Background(), testVideo()); err != nil {
|
||||||
|
t.Fatalf("pass %d: %v", i, err)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
if src.fetchCalls != 1 {
|
||||||
|
t.Fatalf("the second summarize must reuse the stored transcript; got %d fetches", src.fetchCalls)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// A stored "no captions" outcome short-circuits before both fetch and summarize.
|
||||||
|
func TestProcessNewVideo_StoredNoneSkipsFetchAndSummarize(t *testing.T) {
|
||||||
|
src := &recordingSource{}
|
||||||
|
ts := newFakeTranscriptStore()
|
||||||
|
ts.stored[ts.key("youtube", "yt1")] = domain.Transcript{Source: domain.SourceNone}
|
||||||
|
sum := &countingSummarizer{}
|
||||||
|
eng := NewEngine(src, sum, nopSink{})
|
||||||
|
eng.Transcripts = ts
|
||||||
|
|
||||||
|
res, err := eng.ProcessNewVideo(context.Background(), testVideo())
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("ProcessNewVideo: %v", err)
|
||||||
|
}
|
||||||
|
if !res.Skipped {
|
||||||
|
t.Fatal("a stored SourceNone must skip")
|
||||||
|
}
|
||||||
|
if src.fetchCalls != 0 || sum.calls != 0 {
|
||||||
|
t.Fatalf("stored none must neither fetch nor summarize; fetches=%d calls=%d", src.fetchCalls, sum.calls)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// A transient 429 is surfaced (so the runner backs off per-user) but never cached
|
||||||
|
// as a shared terminal state — otherwise it would mask a rate-limit as permanent.
|
||||||
|
func TestProcessNewVideo_RateLimitedIsNotPersisted(t *testing.T) {
|
||||||
|
src := &recordingSource{transcript: domain.Transcript{Source: domain.SourceRateLimited}}
|
||||||
|
ts := newFakeTranscriptStore()
|
||||||
|
eng := NewEngine(src, &countingSummarizer{}, nopSink{})
|
||||||
|
eng.Transcripts = ts
|
||||||
|
|
||||||
|
res, err := eng.ProcessNewVideo(context.Background(), testVideo())
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("ProcessNewVideo: %v", err)
|
||||||
|
}
|
||||||
|
if !res.Skipped || res.TranscriptSource != string(domain.SourceRateLimited) {
|
||||||
|
t.Fatalf("expected a rate-limited skip; skipped=%v source=%q", res.Skipped, res.TranscriptSource)
|
||||||
|
}
|
||||||
|
if ts.saves != 0 {
|
||||||
|
t.Fatalf("a transient rate-limit must not be persisted; got %d saves", ts.saves)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// With no TranscriptStore wired the engine fetches every time (back-compat).
|
||||||
|
func TestProcessNewVideo_NilStoreFetchesEveryTime(t *testing.T) {
|
||||||
|
src := &recordingSource{transcript: domain.Transcript{Source: domain.SourceCaptions, Content: "words"}}
|
||||||
|
eng := NewEngine(src, &countingSummarizer{}, nopSink{})
|
||||||
|
|
||||||
|
for i := 0; i < 2; i++ {
|
||||||
|
if _, err := eng.ProcessNewVideo(context.Background(), testVideo()); err != nil {
|
||||||
|
t.Fatalf("pass %d: %v", i, err)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
if src.fetchCalls != 2 {
|
||||||
|
t.Fatalf("nil store must fetch every time; got %d", src.fetchCalls)
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -23,6 +23,15 @@ type Connections interface {
|
|||||||
UpsertConnection(ctx context.Context, userID string, c store.Connection) error
|
UpsertConnection(ctx context.Context, userID string, c store.Connection) error
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// DiscoveryTrigger requests an out-of-band discovery pass for a user. The connect
|
||||||
|
// flow fires it the moment a YouTube account is linked so videos appear promptly
|
||||||
|
// instead of waiting for the next scheduled pass (#6). Enqueue must be
|
||||||
|
// non-blocking and safe to call from the request goroutine; the implementation
|
||||||
|
// owns serialization with the scheduler (one pass at a time). nil = no trigger.
|
||||||
|
type DiscoveryTrigger interface {
|
||||||
|
Enqueue(userID string)
|
||||||
|
}
|
||||||
|
|
||||||
// connectStateTTL bounds how long a generated CSRF state is valid between the
|
// connectStateTTL bounds how long a generated CSRF state is valid between the
|
||||||
// connect redirect and the provider callback.
|
// connect redirect and the provider callback.
|
||||||
const connectStateTTL = 10 * time.Minute
|
const connectStateTTL = 10 * time.Minute
|
||||||
@@ -43,6 +52,10 @@ type ConnectHandler struct {
|
|||||||
Conns Connections
|
Conns Connections
|
||||||
Log *slog.Logger
|
Log *slog.Logger
|
||||||
|
|
||||||
|
// Discovery, when set, is fired after a successful connect so the new
|
||||||
|
// connection's videos are discovered immediately (#6). Optional.
|
||||||
|
Discovery DiscoveryTrigger
|
||||||
|
|
||||||
states *connectStateStore
|
states *connectStateStore
|
||||||
now func() time.Time
|
now func() time.Time
|
||||||
}
|
}
|
||||||
@@ -134,6 +147,12 @@ func (h *ConnectHandler) handleCallback(w http.ResponseWriter, r *http.Request)
|
|||||||
return
|
return
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// Discover this user's videos now rather than waiting for the next scheduled
|
||||||
|
// pass (#6). Non-blocking; the trigger serializes with the scheduler.
|
||||||
|
if h.Discovery != nil {
|
||||||
|
h.Discovery.Enqueue(userID)
|
||||||
|
}
|
||||||
|
|
||||||
setFlash(w, flashConnected)
|
setFlash(w, flashConnected)
|
||||||
http.Redirect(w, r, "/", http.StatusSeeOther)
|
http.Redirect(w, r, "/", http.StatusSeeOther)
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -45,6 +45,24 @@ func (c *fakeConns) UpsertConnection(_ context.Context, userID string, conn stor
|
|||||||
return nil
|
return nil
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// fakeTrigger records Enqueue calls so a test can assert connect fired discovery.
|
||||||
|
type fakeTrigger struct {
|
||||||
|
mu sync.Mutex
|
||||||
|
users []string
|
||||||
|
}
|
||||||
|
|
||||||
|
func (f *fakeTrigger) Enqueue(userID string) {
|
||||||
|
f.mu.Lock()
|
||||||
|
defer f.mu.Unlock()
|
||||||
|
f.users = append(f.users, userID)
|
||||||
|
}
|
||||||
|
|
||||||
|
func (f *fakeTrigger) seen() []string {
|
||||||
|
f.mu.Lock()
|
||||||
|
defer f.mu.Unlock()
|
||||||
|
return append([]string(nil), f.users...)
|
||||||
|
}
|
||||||
|
|
||||||
// tokenServer fakes Google's token endpoint, returning body for any POST.
|
// tokenServer fakes Google's token endpoint, returning body for any POST.
|
||||||
func tokenServer(t *testing.T, body string) *httptest.Server {
|
func tokenServer(t *testing.T, body string) *httptest.Server {
|
||||||
t.Helper()
|
t.Helper()
|
||||||
@@ -126,6 +144,36 @@ func TestCallbackExchangesAndRecordsConnection(t *testing.T) {
|
|||||||
require.Equal(t, wantRef, conns.conn.TokenRef)
|
require.Equal(t, wantRef, conns.conn.TokenRef)
|
||||||
}
|
}
|
||||||
|
|
||||||
|
func TestCallbackTriggersDiscovery(t *testing.T) {
|
||||||
|
srv := tokenServer(t,
|
||||||
|
`{"access_token":"at","refresh_token":"rt-secret","token_type":"Bearer","expires_in":3600}`)
|
||||||
|
app := newConnectApp(t, srv.URL, &fakeWriter{}, &fakeConns{})
|
||||||
|
trig := &fakeTrigger{}
|
||||||
|
app.Connect.Discovery = trig
|
||||||
|
|
||||||
|
state := connectState(t, app)
|
||||||
|
rec := do(t, app, httptest.NewRequest(http.MethodGet,
|
||||||
|
"/oauth/youtube/callback?state="+state+"&code=the-code", nil))
|
||||||
|
|
||||||
|
require.Equal(t, http.StatusSeeOther, rec.Code)
|
||||||
|
require.Equal(t, []string{userID}, trig.seen(),
|
||||||
|
"a successful connect must trigger discovery for the connecting user")
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestCallbackNoDiscoveryOnFailedConnect(t *testing.T) {
|
||||||
|
srv := tokenServer(t,
|
||||||
|
`{"access_token":"at","refresh_token":"rt","token_type":"Bearer","expires_in":3600}`)
|
||||||
|
app := newConnectApp(t, srv.URL, &fakeWriter{}, &fakeConns{})
|
||||||
|
trig := &fakeTrigger{}
|
||||||
|
app.Connect.Discovery = trig
|
||||||
|
|
||||||
|
// No state → CSRF reject → nothing connected, so no discovery.
|
||||||
|
rec := do(t, app, httptest.NewRequest(http.MethodGet,
|
||||||
|
"/oauth/youtube/callback?code=the-code", nil))
|
||||||
|
require.Equal(t, http.StatusBadRequest, rec.Code)
|
||||||
|
require.Empty(t, trig.seen(), "a failed connect must not trigger discovery")
|
||||||
|
}
|
||||||
|
|
||||||
func TestCallbackRejectsMissingState(t *testing.T) {
|
func TestCallbackRejectsMissingState(t *testing.T) {
|
||||||
srv := tokenServer(t,
|
srv := tokenServer(t,
|
||||||
`{"access_token":"at","refresh_token":"rt","token_type":"Bearer","expires_in":3600}`)
|
`{"access_token":"at","refresh_token":"rt","token_type":"Bearer","expires_in":3600}`)
|
||||||
|
|||||||
+114
-15
@@ -3,6 +3,8 @@ package web
|
|||||||
import (
|
import (
|
||||||
"context"
|
"context"
|
||||||
"errors"
|
"errors"
|
||||||
|
"html/template"
|
||||||
|
"io"
|
||||||
"log/slog"
|
"log/slog"
|
||||||
"net/http"
|
"net/http"
|
||||||
"time"
|
"time"
|
||||||
@@ -10,6 +12,7 @@ import (
|
|||||||
"github.com/a-h/templ"
|
"github.com/a-h/templ"
|
||||||
|
|
||||||
"gitea.d-ma.be/mathias/tapir/internal/adapters/store"
|
"gitea.d-ma.be/mathias/tapir/internal/adapters/store"
|
||||||
|
"gitea.d-ma.be/mathias/tapir/internal/domain"
|
||||||
)
|
)
|
||||||
|
|
||||||
// Store is the read/write surface the web handlers depend on — a narrow port over
|
// Store is the read/write surface the web handlers depend on — a narrow port over
|
||||||
@@ -18,6 +21,9 @@ import (
|
|||||||
// fake without a database.
|
// fake without a database.
|
||||||
type Store interface {
|
type Store interface {
|
||||||
ListVideos(ctx context.Context, userID string, limit int) ([]store.SummaryRow, error)
|
ListVideos(ctx context.Context, userID string, limit int) ([]store.SummaryRow, error)
|
||||||
|
// DistinctChannels lists the user's source channels — the options for the
|
||||||
|
// feed's channel multi-select filter.
|
||||||
|
DistinctChannels(ctx context.Context, userID string) ([]string, error)
|
||||||
GetSummaryByVideo(ctx context.Context, userID, videoID string) (*store.SummaryRow, error)
|
GetSummaryByVideo(ctx context.Context, userID, videoID string) (*store.SummaryRow, error)
|
||||||
GetVideoRow(ctx context.Context, userID, videoID string) (*store.SummaryRow, error)
|
GetVideoRow(ctx context.Context, userID, videoID string) (*store.SummaryRow, error)
|
||||||
ActionsFor(ctx context.Context, userID string, videoIDs []string) (map[string][]string, error)
|
ActionsFor(ctx context.Context, userID string, videoIDs []string) (map[string][]string, error)
|
||||||
@@ -30,6 +36,10 @@ type Store interface {
|
|||||||
SetAutoSummarize(ctx context.Context, userID string, enabled bool) error
|
SetAutoSummarize(ctx context.Context, userID string, enabled bool) error
|
||||||
RequestSummarize(ctx context.Context, userID, videoID string) error
|
RequestSummarize(ctx context.Context, userID, videoID string) error
|
||||||
|
|
||||||
|
// UpsertVideo persists a pasted video (idempotent on user+provider+video id,
|
||||||
|
// so it also dedups) and returns its durable store id.
|
||||||
|
UpsertVideo(ctx context.Context, v domain.Video) (string, error)
|
||||||
|
|
||||||
// Account management (the /account page, disconnect, delete-account).
|
// Account management (the /account page, disconnect, delete-account).
|
||||||
ConnectionsForUser(ctx context.Context, userID string) ([]store.Connection, error)
|
ConnectionsForUser(ctx context.Context, userID string) ([]store.Connection, error)
|
||||||
DeleteConnection(ctx context.Context, userID, provider string) error
|
DeleteConnection(ctx context.Context, userID, provider string) error
|
||||||
@@ -78,6 +88,9 @@ type App struct {
|
|||||||
// background goroutine (the "Summarize" button kicks it off). Nil = queue-only:
|
// background goroutine (the "Summarize" button kicks it off). Nil = queue-only:
|
||||||
// the button flips the DB flag and the next `tapir run` does the work.
|
// the button flips the DB flag and the next `tapir run` does the work.
|
||||||
Processor Processor
|
Processor Processor
|
||||||
|
// Fetcher, when non-nil, resolves an arbitrary YouTube video id to metadata for
|
||||||
|
// the paste-a-URL flow (Feature 2). Nil = the /paste route is not mounted.
|
||||||
|
Fetcher VideoFetcher
|
||||||
// Processing tracks in-flight immediate summarizations so the status endpoint
|
// Processing tracks in-flight immediate summarizations so the status endpoint
|
||||||
// shows the animation until the summary lands. The zero value is ready to use.
|
// shows the animation until the summary lands. The zero value is ready to use.
|
||||||
Processing ProcessingSet
|
Processing ProcessingSet
|
||||||
@@ -132,6 +145,9 @@ func (a *App) Router() http.Handler {
|
|||||||
app.HandleFunc("POST /v/{videoId}/action", a.handleAction)
|
app.HandleFunc("POST /v/{videoId}/action", a.handleAction)
|
||||||
app.HandleFunc("POST /v/{videoId}/summarize", a.handleRequestSummarize)
|
app.HandleFunc("POST /v/{videoId}/summarize", a.handleRequestSummarize)
|
||||||
app.HandleFunc("POST /v/{videoId}/retry-now", a.handleRetryNow)
|
app.HandleFunc("POST /v/{videoId}/retry-now", a.handleRetryNow)
|
||||||
|
if a.Fetcher != nil {
|
||||||
|
app.HandleFunc("POST /paste", a.handlePaste)
|
||||||
|
}
|
||||||
app.HandleFunc("GET /v/{videoId}/status", a.handleStatus)
|
app.HandleFunc("GET /v/{videoId}/status", a.handleStatus)
|
||||||
app.HandleFunc("GET /register", a.handleRegisterForm)
|
app.HandleFunc("GET /register", a.handleRegisterForm)
|
||||||
app.HandleFunc("POST /register", a.handleRegister)
|
app.HandleFunc("POST /register", a.handleRegister)
|
||||||
@@ -183,7 +199,7 @@ func (a *App) handleList(w http.ResponseWriter, r *http.Request) {
|
|||||||
}
|
}
|
||||||
q := r.URL.Query()
|
q := r.URL.Query()
|
||||||
f := Filter{
|
f := Filter{
|
||||||
Channel: q.Get("channel"),
|
Channels: nonEmptyStrings(q["channel"]),
|
||||||
From: q.Get("from"),
|
From: q.Get("from"),
|
||||||
To: q.Get("to"),
|
To: q.Get("to"),
|
||||||
OnlySummarized: q.Get("summarized") == "1",
|
OnlySummarized: q.Get("summarized") == "1",
|
||||||
@@ -198,25 +214,39 @@ func (a *App) handleList(w http.ResponseWriter, r *http.Request) {
|
|||||||
rows := f.apply(allRows)
|
rows := f.apply(allRows)
|
||||||
buckets := bucketRows(rows, a.recencyCutoff())
|
buckets := bucketRows(rows, a.recencyCutoff())
|
||||||
|
|
||||||
// hasConnected drives the empty state: a fresh account with a connection but
|
// Channel options for the multi-select filter (the user's source channels).
|
||||||
// no discovery pass yet has zero rows, and we want it to read "connected,
|
channels, err := a.Store.DistinctChannels(r.Context(), userID)
|
||||||
// summaries land gradually" rather than "nothing here". Only needed when the
|
if err != nil {
|
||||||
// list is empty.
|
a.serverError(w, r, "distinct channels", err)
|
||||||
hasConnected := false
|
return
|
||||||
if buckets.empty() {
|
}
|
||||||
conns, err := a.Store.ConnectionsForUser(r.Context(), userID)
|
|
||||||
if err != nil {
|
// hasConnected drives both the paste box (shown to ANY connected user, #2) and
|
||||||
a.serverError(w, r, "connections for user", err)
|
// the empty-state copy (a fresh account with a connection but no discovery pass
|
||||||
return
|
// yet reads "connected, summaries land gradually" rather than "nothing here").
|
||||||
}
|
// Computed every render — not only when empty — so a user with videos still
|
||||||
hasConnected = len(conns) > 0
|
// gets the paste box.
|
||||||
|
conns, err := a.Store.ConnectionsForUser(r.Context(), userID)
|
||||||
|
if err != nil {
|
||||||
|
a.serverError(w, r, "connections for user", err)
|
||||||
|
return
|
||||||
|
}
|
||||||
|
hasConnected := len(conns) > 0
|
||||||
|
|
||||||
|
// Summarization mode drives the backlog copy: an auto user is told summaries
|
||||||
|
// land gradually; a manual user is told to click Summarize (the first pilot
|
||||||
|
// user sat in manual mode reading "land automatically" and waited forever).
|
||||||
|
autoSummarize, err := a.Store.GetAutoSummarize(r.Context(), userID)
|
||||||
|
if err != nil {
|
||||||
|
a.serverError(w, r, "summarize mode", err)
|
||||||
|
return
|
||||||
}
|
}
|
||||||
|
|
||||||
if isHTMX(r) {
|
if isHTMX(r) {
|
||||||
a.render(w, r, summaryList(buckets, hasConnected))
|
a.render(w, r, summaryList(buckets, hasConnected, autoSummarize))
|
||||||
return
|
return
|
||||||
}
|
}
|
||||||
a.render(w, r, ListPage(buckets, f, stats, takeFlash(w, r), hasConnected))
|
a.render(w, r, ListPage(buckets, f, stats, takeFlash(w, r), hasConnected, channels, autoSummarize))
|
||||||
}
|
}
|
||||||
|
|
||||||
// handleDetail renders one summary in full (highlights, takeaways, action group).
|
// handleDetail renders one summary in full (highlights, takeaways, action group).
|
||||||
@@ -324,6 +354,75 @@ func (a *App) handleRequestSummarize(w http.ResponseWriter, r *http.Request) {
|
|||||||
a.render(w, r, VideoCard(*row))
|
a.render(w, r, VideoCard(*row))
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// handlePaste handles "paste a YouTube URL" (Feature 2). It parses the video id,
|
||||||
|
// fetches metadata (Data API — ungated), upserts a subscription-less video row
|
||||||
|
// scoped to the user (idempotent, so it also dedups), and — if the video isn't
|
||||||
|
// already summarized — requests a summary and kicks off immediate processing
|
||||||
|
// through the SAME rate gate as the Summarize button. An explicit paste is a
|
||||||
|
// manual request, so it summarizes regardless of the recency window. A video that
|
||||||
|
// turns out to have no captions resolves to the honest "no transcript" terminal
|
||||||
|
// state via the engine (ADR-010), not an error here.
|
||||||
|
func (a *App) handlePaste(w http.ResponseWriter, r *http.Request) {
|
||||||
|
userID, ok := a.currentUserID(w, r)
|
||||||
|
if !ok {
|
||||||
|
return
|
||||||
|
}
|
||||||
|
videoID, err := parseYouTubeVideoID(r.FormValue("url"))
|
||||||
|
if err != nil {
|
||||||
|
a.pasteFailure(w, http.StatusBadRequest, "That doesn't look like a YouTube video link.")
|
||||||
|
return
|
||||||
|
}
|
||||||
|
|
||||||
|
v, err := a.Fetcher.FetchVideo(r.Context(), userID, videoID)
|
||||||
|
if errors.Is(err, domain.ErrVideoNotFound) {
|
||||||
|
a.pasteFailure(w, http.StatusNotFound, "That video couldn't be found — it may be private or removed.")
|
||||||
|
return
|
||||||
|
}
|
||||||
|
if err != nil {
|
||||||
|
a.serverError(w, r, "paste fetch", err)
|
||||||
|
return
|
||||||
|
}
|
||||||
|
|
||||||
|
id, err := a.Store.UpsertVideo(r.Context(), v)
|
||||||
|
if err != nil {
|
||||||
|
a.serverError(w, r, "paste upsert", err)
|
||||||
|
return
|
||||||
|
}
|
||||||
|
row, err := a.Store.GetVideoRow(r.Context(), userID, id)
|
||||||
|
if err != nil {
|
||||||
|
a.serverError(w, r, "paste get video", err)
|
||||||
|
return
|
||||||
|
}
|
||||||
|
|
||||||
|
// Dedup: already in the feed with a summary — surface the existing entry,
|
||||||
|
// don't re-summarize.
|
||||||
|
if row.Summarized {
|
||||||
|
a.render(w, r, VideoCard(*row))
|
||||||
|
return
|
||||||
|
}
|
||||||
|
|
||||||
|
// New or unsummarized: queue + (if a Processor is wired) summarize now, through
|
||||||
|
// the shared gate. RequestSummarize makes it durable even if the process dies.
|
||||||
|
if err := a.Store.RequestSummarize(r.Context(), userID, id); err != nil {
|
||||||
|
a.serverError(w, r, "paste request summarize", err)
|
||||||
|
return
|
||||||
|
}
|
||||||
|
if a.Processor != nil {
|
||||||
|
a.startProcessing(userID, id)
|
||||||
|
a.render(w, r, processingCard(*row))
|
||||||
|
return
|
||||||
|
}
|
||||||
|
a.render(w, r, VideoCard(*row))
|
||||||
|
}
|
||||||
|
|
||||||
|
// pasteFailure renders a minimal inline error fragment for the paste form (HTMX
|
||||||
|
// swaps it in). No templ dependency so it renders even on a bad-input fast path.
|
||||||
|
func (a *App) pasteFailure(w http.ResponseWriter, status int, msg string) {
|
||||||
|
w.Header().Set("Content-Type", "text/html; charset=utf-8")
|
||||||
|
w.WriteHeader(status)
|
||||||
|
_, _ = io.WriteString(w, `<p class="paste-error" role="alert">`+template.HTMLEscapeString(msg)+`</p>`)
|
||||||
|
}
|
||||||
|
|
||||||
// handleRetryNow handles the "Try now" button on rate-limited video cards. It
|
// handleRetryNow handles the "Try now" button on rate-limited video cards. It
|
||||||
// clears the rate_limited_at backoff so the scheduler won't skip the video, then
|
// clears the rate_limited_at backoff so the scheduler won't skip the video, then
|
||||||
// triggers an immediate ProcessVideo — same background path as handleRequestSummarize.
|
// triggers an immediate ProcessVideo — same background path as handleRequestSummarize.
|
||||||
|
|||||||
@@ -207,13 +207,15 @@ func TestListChannelFilter(t *testing.T) {
|
|||||||
resetDB(t, p)
|
resetDB(t, p)
|
||||||
require.NoError(t, deliver(ctx, app, videoX, "body x"))
|
require.NoError(t, deliver(ctx, app, videoX, "body x"))
|
||||||
seedVideo(t, p, videoX, "X Title", "https://x", time.Time{})
|
seedVideo(t, p, videoX, "X Title", "https://x", time.Time{})
|
||||||
|
_, err := p.Exec(ctx, `UPDATE videos SET channel_title = 'Acme Channel' WHERE id = $1`, videoX)
|
||||||
|
require.NoError(t, err)
|
||||||
|
|
||||||
// Channel is "youtube" for seeded rows; a non-matching filter hides them.
|
// Selecting a different channel hides the row; selecting its channel shows it.
|
||||||
rec := do(t, app, httptest.NewRequest(http.MethodGet, "/?channel=vimeo", nil))
|
rec := do(t, app, httptest.NewRequest(http.MethodGet, "/?channel=Other+Channel", nil))
|
||||||
require.Equal(t, http.StatusOK, rec.Code)
|
require.Equal(t, http.StatusOK, rec.Code)
|
||||||
require.NotContains(t, body(t, rec), "X Title")
|
require.NotContains(t, body(t, rec), "X Title")
|
||||||
|
|
||||||
rec = do(t, app, httptest.NewRequest(http.MethodGet, "/?channel=youtube", nil))
|
rec = do(t, app, httptest.NewRequest(http.MethodGet, "/?channel=Acme+Channel", nil))
|
||||||
require.Contains(t, body(t, rec), "X Title")
|
require.Contains(t, body(t, rec), "X Title")
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -475,3 +477,36 @@ func postAction(t *testing.T, app *web.App, videoID, action string, htmx bool) *
|
|||||||
}
|
}
|
||||||
return do(t, app, req)
|
return do(t, app, req)
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// TestListManualModeBannerCopy: a manual-mode user with un-summarized videos
|
||||||
|
// sees the manual prompt (click Summarize), NOT the "summaries land
|
||||||
|
// automatically" copy that misled the first pilot user into waiting forever.
|
||||||
|
func TestListManualModeBannerCopy(t *testing.T) {
|
||||||
|
ctx := context.Background()
|
||||||
|
app := newApp(t)
|
||||||
|
p := rawPool(t)
|
||||||
|
resetDB(t, p)
|
||||||
|
seedVideo(t, p, videoX, "X Title", "https://x", time.Time{}) // pending, un-summarized
|
||||||
|
require.NoError(t, app.Store.SetAutoSummarize(ctx, userID, false))
|
||||||
|
|
||||||
|
html := body(t, do(t, app, httptest.NewRequest(http.MethodGet, "/", nil)))
|
||||||
|
require.Contains(t, html, "Manual mode")
|
||||||
|
require.Contains(t, html, "are not summarized automatically")
|
||||||
|
require.NotContains(t, html, "land gradually",
|
||||||
|
"manual-mode user must not be told summaries arrive automatically")
|
||||||
|
}
|
||||||
|
|
||||||
|
// TestListAutoModeBannerCopy: an auto-mode user with a backlog sees the
|
||||||
|
// gradual-delivery copy, not the manual prompt.
|
||||||
|
func TestListAutoModeBannerCopy(t *testing.T) {
|
||||||
|
ctx := context.Background()
|
||||||
|
app := newApp(t)
|
||||||
|
p := rawPool(t)
|
||||||
|
resetDB(t, p)
|
||||||
|
seedVideo(t, p, videoX, "X Title", "https://x", time.Time{})
|
||||||
|
require.NoError(t, app.Store.SetAutoSummarize(ctx, userID, true))
|
||||||
|
|
||||||
|
html := body(t, do(t, app, httptest.NewRequest(http.MethodGet, "/", nil)))
|
||||||
|
require.Contains(t, html, "land gradually")
|
||||||
|
require.NotContains(t, html, "are not summarized automatically")
|
||||||
|
}
|
||||||
|
|||||||
@@ -0,0 +1,62 @@
|
|||||||
|
package web
|
||||||
|
|
||||||
|
import (
|
||||||
|
"fmt"
|
||||||
|
"net/url"
|
||||||
|
"regexp"
|
||||||
|
"strings"
|
||||||
|
)
|
||||||
|
|
||||||
|
// youtubeVideoID matches a canonical YouTube video id: exactly 11 URL-safe chars.
|
||||||
|
var youtubeVideoID = regexp.MustCompile(`^[A-Za-z0-9_-]{11}$`)
|
||||||
|
|
||||||
|
// parseYouTubeVideoID extracts the 11-character video id from a pasted YouTube
|
||||||
|
// URL (watch?v=, youtu.be/, shorts/, embed/) or a bare id. It rejects non-YouTube
|
||||||
|
// hosts and anything that doesn't yield a valid id, so the paste flow never tries
|
||||||
|
// to fetch a video that can't exist (Feature 2).
|
||||||
|
func parseYouTubeVideoID(raw string) (string, error) {
|
||||||
|
s := strings.TrimSpace(raw)
|
||||||
|
if s == "" {
|
||||||
|
return "", fmt.Errorf("empty input")
|
||||||
|
}
|
||||||
|
|
||||||
|
// Bare id (no URL) — accept directly.
|
||||||
|
if youtubeVideoID.MatchString(s) {
|
||||||
|
return s, nil
|
||||||
|
}
|
||||||
|
|
||||||
|
// Accept scheme-less URLs (youtube.com/watch?v=...) by giving url.Parse a host.
|
||||||
|
if !strings.Contains(s, "://") {
|
||||||
|
s = "https://" + s
|
||||||
|
}
|
||||||
|
u, err := url.Parse(s)
|
||||||
|
if err != nil {
|
||||||
|
return "", fmt.Errorf("not a URL: %w", err)
|
||||||
|
}
|
||||||
|
|
||||||
|
host := strings.ToLower(u.Hostname())
|
||||||
|
isYouTube := host == "youtu.be" || host == "youtube.com" || strings.HasSuffix(host, ".youtube.com")
|
||||||
|
if !isYouTube {
|
||||||
|
return "", fmt.Errorf("not a YouTube URL: %q", host)
|
||||||
|
}
|
||||||
|
|
||||||
|
var id string
|
||||||
|
switch {
|
||||||
|
case host == "youtu.be":
|
||||||
|
// youtu.be/<id>
|
||||||
|
id = strings.Trim(u.Path, "/")
|
||||||
|
case u.Path == "/watch":
|
||||||
|
id = u.Query().Get("v")
|
||||||
|
default:
|
||||||
|
// /shorts/<id>, /embed/<id>
|
||||||
|
parts := strings.Split(strings.Trim(u.Path, "/"), "/")
|
||||||
|
if len(parts) == 2 && (parts[0] == "shorts" || parts[0] == "embed") {
|
||||||
|
id = parts[1]
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
if !youtubeVideoID.MatchString(id) {
|
||||||
|
return "", fmt.Errorf("no YouTube video id in %q", raw)
|
||||||
|
}
|
||||||
|
return id, nil
|
||||||
|
}
|
||||||
@@ -0,0 +1,134 @@
|
|||||||
|
package web_test
|
||||||
|
|
||||||
|
import (
|
||||||
|
"context"
|
||||||
|
"net/http"
|
||||||
|
"net/http/httptest"
|
||||||
|
"net/url"
|
||||||
|
"strings"
|
||||||
|
"testing"
|
||||||
|
"time"
|
||||||
|
|
||||||
|
"github.com/stretchr/testify/require"
|
||||||
|
|
||||||
|
"gitea.d-ma.be/mathias/tapir/internal/domain"
|
||||||
|
)
|
||||||
|
|
||||||
|
// fakeFetcher is a web.VideoFetcher returning a fixed video (or an error),
|
||||||
|
// scoped to whatever (userID, videoID) the handler asks for.
|
||||||
|
type fakeFetcher struct {
|
||||||
|
title string
|
||||||
|
err error
|
||||||
|
calls int
|
||||||
|
}
|
||||||
|
|
||||||
|
func (f *fakeFetcher) FetchVideo(_ context.Context, userID, videoID string) (domain.Video, error) {
|
||||||
|
f.calls++
|
||||||
|
if f.err != nil {
|
||||||
|
return domain.Video{}, f.err
|
||||||
|
}
|
||||||
|
return domain.Video{
|
||||||
|
UserID: userID,
|
||||||
|
Provider: domain.ProviderYouTube,
|
||||||
|
ProviderVideoID: videoID,
|
||||||
|
Title: f.title,
|
||||||
|
URL: "https://www.youtube.com/watch?v=" + videoID,
|
||||||
|
}, nil
|
||||||
|
}
|
||||||
|
|
||||||
|
func pasteReq(rawURL string) *http.Request {
|
||||||
|
req := httptest.NewRequest(http.MethodPost, "/paste",
|
||||||
|
strings.NewReader("url="+url.QueryEscape(rawURL)))
|
||||||
|
req.Header.Set("Content-Type", "application/x-www-form-urlencoded")
|
||||||
|
return req
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestPasteValidURLAddsAndRequests(t *testing.T) {
|
||||||
|
ctx := context.Background()
|
||||||
|
app := newApp(t)
|
||||||
|
resetDB(t, rawPool(t))
|
||||||
|
app.Fetcher = &fakeFetcher{title: "Pasted Talk"}
|
||||||
|
p := rawPool(t)
|
||||||
|
|
||||||
|
rec := do(t, app, pasteReq("https://youtu.be/dQw4w9WgXcQ"))
|
||||||
|
require.Equal(t, http.StatusOK, rec.Code)
|
||||||
|
|
||||||
|
var (
|
||||||
|
count, requested int
|
||||||
|
title string
|
||||||
|
)
|
||||||
|
require.NoError(t, p.QueryRow(ctx,
|
||||||
|
`SELECT count(*), coalesce(max(title),'') FROM videos
|
||||||
|
WHERE user_id=$1 AND provider_video_id='dQw4w9WgXcQ'`, userID).Scan(&count, &title))
|
||||||
|
require.Equal(t, 1, count, "pasted video added once, scoped to the user")
|
||||||
|
require.Equal(t, "Pasted Talk", title)
|
||||||
|
|
||||||
|
require.NoError(t, p.QueryRow(ctx,
|
||||||
|
`SELECT count(*) FROM videos
|
||||||
|
WHERE user_id=$1 AND provider_video_id='dQw4w9WgXcQ' AND summarize_requested`,
|
||||||
|
userID).Scan(&requested))
|
||||||
|
require.Equal(t, 1, requested, "pasted video is queued for summarization (through the gate)")
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestPasteInvalidURLRejected(t *testing.T) {
|
||||||
|
app := newApp(t)
|
||||||
|
resetDB(t, rawPool(t))
|
||||||
|
app.Fetcher = &fakeFetcher{title: "x"}
|
||||||
|
|
||||||
|
rec := do(t, app, pasteReq("definitely not a url"))
|
||||||
|
require.Equal(t, http.StatusBadRequest, rec.Code)
|
||||||
|
|
||||||
|
var count int
|
||||||
|
require.NoError(t, rawPool(t).QueryRow(context.Background(),
|
||||||
|
`SELECT count(*) FROM videos WHERE user_id=$1`, userID).Scan(&count))
|
||||||
|
require.Equal(t, 0, count, "invalid input adds nothing")
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestPasteVideoNotFound(t *testing.T) {
|
||||||
|
app := newApp(t)
|
||||||
|
resetDB(t, rawPool(t))
|
||||||
|
app.Fetcher = &fakeFetcher{err: domain.ErrVideoNotFound}
|
||||||
|
|
||||||
|
rec := do(t, app, pasteReq("https://youtu.be/dQw4w9WgXcQ"))
|
||||||
|
require.Equal(t, http.StatusNotFound, rec.Code)
|
||||||
|
|
||||||
|
var count int
|
||||||
|
require.NoError(t, rawPool(t).QueryRow(context.Background(),
|
||||||
|
`SELECT count(*) FROM videos WHERE user_id=$1`, userID).Scan(&count))
|
||||||
|
require.Equal(t, 0, count, "a not-found video adds nothing")
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestPasteDedupNoDuplicate(t *testing.T) {
|
||||||
|
app := newApp(t)
|
||||||
|
resetDB(t, rawPool(t))
|
||||||
|
app.Fetcher = &fakeFetcher{title: "Pasted Talk"}
|
||||||
|
|
||||||
|
require.Equal(t, http.StatusOK, do(t, app, pasteReq("https://youtu.be/dQw4w9WgXcQ")).Code)
|
||||||
|
require.Equal(t, http.StatusOK, do(t, app, pasteReq("https://www.youtube.com/watch?v=dQw4w9WgXcQ")).Code)
|
||||||
|
|
||||||
|
var count int
|
||||||
|
require.NoError(t, rawPool(t).QueryRow(context.Background(),
|
||||||
|
`SELECT count(*) FROM videos WHERE user_id=$1 AND provider_video_id='dQw4w9WgXcQ'`,
|
||||||
|
userID).Scan(&count))
|
||||||
|
require.Equal(t, 1, count, "pasting the same video twice must not duplicate the row")
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestListShowsPasteFormForConnectedUserWithVideos(t *testing.T) {
|
||||||
|
app := newApp(t)
|
||||||
|
resetDB(t, rawPool(t))
|
||||||
|
app.Fetcher = &fakeFetcher{title: "x"}
|
||||||
|
p := rawPool(t)
|
||||||
|
|
||||||
|
// Connected user with a non-empty feed (the case the bug missed: hasConnected
|
||||||
|
// was only computed for an empty feed).
|
||||||
|
_, err := p.Exec(context.Background(),
|
||||||
|
`INSERT INTO video_connections (user_id, provider, token_ref, status)
|
||||||
|
VALUES ($1, 'youtube', 'youtube/x/refresh_token', 'active')`, userID)
|
||||||
|
require.NoError(t, err)
|
||||||
|
seedVideo(t, p, "11111111-1111-1111-1111-111111111111", "A talk", "https://youtu.be/aaaaaaaaaaa", time.Now())
|
||||||
|
|
||||||
|
rec := do(t, app, httptest.NewRequest(http.MethodGet, "/", nil))
|
||||||
|
require.Equal(t, http.StatusOK, rec.Code)
|
||||||
|
require.Contains(t, body(t, rec), `action="/paste"`,
|
||||||
|
"a connected user must see the paste box even when the feed has videos")
|
||||||
|
}
|
||||||
@@ -0,0 +1,97 @@
|
|||||||
|
package web
|
||||||
|
|
||||||
|
import (
|
||||||
|
"bytes"
|
||||||
|
"context"
|
||||||
|
"gitea.d-ma.be/mathias/tapir/internal/adapters/store"
|
||||||
|
"strings"
|
||||||
|
"testing"
|
||||||
|
)
|
||||||
|
|
||||||
|
func TestParseYouTubeVideoID(t *testing.T) {
|
||||||
|
const id = "dQw4w9WgXcQ"
|
||||||
|
ok := []struct {
|
||||||
|
name, in string
|
||||||
|
}{
|
||||||
|
{"watch", "https://www.youtube.com/watch?v=" + id},
|
||||||
|
{"watch no www", "https://youtube.com/watch?v=" + id},
|
||||||
|
{"watch m", "https://m.youtube.com/watch?v=" + id},
|
||||||
|
{"watch extra params", "https://www.youtube.com/watch?v=" + id + "&t=42s&list=PLxyz"},
|
||||||
|
{"watch param after", "https://www.youtube.com/watch?list=PLxyz&v=" + id},
|
||||||
|
{"short link", "https://youtu.be/" + id},
|
||||||
|
{"short link param", "https://youtu.be/" + id + "?si=abcd&t=1"},
|
||||||
|
{"shorts", "https://www.youtube.com/shorts/" + id},
|
||||||
|
{"embed", "https://www.youtube.com/embed/" + id},
|
||||||
|
{"bare id", id},
|
||||||
|
{"http scheme", "http://youtube.com/watch?v=" + id},
|
||||||
|
{"no scheme", "youtube.com/watch?v=" + id},
|
||||||
|
{"trailing space", " https://youtu.be/" + id + " "},
|
||||||
|
}
|
||||||
|
for _, c := range ok {
|
||||||
|
t.Run(c.name, func(t *testing.T) {
|
||||||
|
got, err := parseYouTubeVideoID(c.in)
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("parseYouTubeVideoID(%q) error: %v", c.in, err)
|
||||||
|
}
|
||||||
|
if got != id {
|
||||||
|
t.Fatalf("parseYouTubeVideoID(%q) = %q, want %q", c.in, got, id)
|
||||||
|
}
|
||||||
|
})
|
||||||
|
}
|
||||||
|
|
||||||
|
bad := []struct {
|
||||||
|
name, in string
|
||||||
|
}{
|
||||||
|
{"empty", ""},
|
||||||
|
{"blank", " "},
|
||||||
|
{"vimeo", "https://vimeo.com/123456789"},
|
||||||
|
{"other host", "https://example.com/watch?v=" + id},
|
||||||
|
{"watch no id", "https://www.youtube.com/watch?v="},
|
||||||
|
{"short id", "https://youtu.be/abc"},
|
||||||
|
{"long id", "https://youtu.be/" + id + "extra"},
|
||||||
|
{"bad chars", "https://youtu.be/dQw4w9Wg!cQ"},
|
||||||
|
{"not a url", "just some text"},
|
||||||
|
{"channel url", "https://www.youtube.com/@somechannel"},
|
||||||
|
}
|
||||||
|
for _, c := range bad {
|
||||||
|
t.Run("reject "+c.name, func(t *testing.T) {
|
||||||
|
if got, err := parseYouTubeVideoID(c.in); err == nil {
|
||||||
|
t.Fatalf("parseYouTubeVideoID(%q) = %q, want error", c.in, got)
|
||||||
|
}
|
||||||
|
})
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestListPageShowsPasteFormOnlyWhenConnected(t *testing.T) {
|
||||||
|
render := func(connected bool) string {
|
||||||
|
var buf bytes.Buffer
|
||||||
|
if err := ListPage(listBuckets{}, Filter{}, PipelineStats{}, "", connected, nil, true).Render(context.Background(), &buf); err != nil {
|
||||||
|
t.Fatalf("render: %v", err)
|
||||||
|
}
|
||||||
|
return buf.String()
|
||||||
|
}
|
||||||
|
|
||||||
|
html := render(true)
|
||||||
|
if !strings.Contains(html, `name="url"`) || !strings.Contains(html, `action="/paste"`) {
|
||||||
|
t.Errorf("connected feed must show the paste form")
|
||||||
|
}
|
||||||
|
if strings.Contains(render(false), `name="url"`) {
|
||||||
|
t.Errorf("disconnected feed must not show the paste form")
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestFilterMatchesMultipleChannels(t *testing.T) {
|
||||||
|
f := Filter{Channels: []string{"Acme", "Zeta"}}
|
||||||
|
row := func(ch string) store.SummaryRow { return store.SummaryRow{ChannelTitle: ch, Summarized: true} }
|
||||||
|
|
||||||
|
rows := []store.SummaryRow{row("Acme"), row("Beta"), row("Zeta")}
|
||||||
|
got := f.apply(rows)
|
||||||
|
if len(got) != 2 || got[0].ChannelTitle != "Acme" || got[1].ChannelTitle != "Zeta" {
|
||||||
|
t.Fatalf("multi-channel filter = %+v, want Acme+Zeta only", got)
|
||||||
|
}
|
||||||
|
|
||||||
|
// Empty selection = no channel constraint (all pass).
|
||||||
|
if n := len(Filter{}.apply(rows)); n != 3 {
|
||||||
|
t.Fatalf("no channel filter should pass all rows, got %d", n)
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -3,6 +3,8 @@ package web
|
|||||||
import (
|
import (
|
||||||
"context"
|
"context"
|
||||||
"sync"
|
"sync"
|
||||||
|
|
||||||
|
"gitea.d-ma.be/mathias/tapir/internal/domain"
|
||||||
)
|
)
|
||||||
|
|
||||||
// Processor runs the core summarization use case for a single already-discovered
|
// Processor runs the core summarization use case for a single already-discovered
|
||||||
@@ -14,6 +16,14 @@ type Processor interface {
|
|||||||
ProcessVideo(ctx context.Context, userID, videoID string) error
|
ProcessVideo(ctx context.Context, userID, videoID string) error
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// VideoFetcher resolves an arbitrary YouTube video id to its metadata for the
|
||||||
|
// paste-a-URL flow (Feature 2). It is a Data API call, NOT the rate-limited
|
||||||
|
// caption path. Returns domain.ErrVideoNotFound for a deleted/private/typo'd id.
|
||||||
|
// cmd/tapir wires a per-user YouTube adapter; nil disables the paste route.
|
||||||
|
type VideoFetcher interface {
|
||||||
|
FetchVideo(ctx context.Context, userID, videoID string) (domain.Video, error)
|
||||||
|
}
|
||||||
|
|
||||||
// ProcessingSet tracks the (user, video) ids currently being summarized in-process
|
// ProcessingSet tracks the (user, video) ids currently being summarized in-process
|
||||||
// so the status endpoint can show the animation until the summary lands. It is
|
// so the status endpoint can show the animation until the summary lands. It is
|
||||||
// ephemeral (single-instance Stage-1): a restart drops it, and the DB holds the
|
// ephemeral (single-instance Stage-1): a restart drops it, and the DB holds the
|
||||||
|
|||||||
+27
-5
@@ -2,6 +2,7 @@ package web
|
|||||||
|
|
||||||
import (
|
import (
|
||||||
"regexp"
|
"regexp"
|
||||||
|
"slices"
|
||||||
"strings"
|
"strings"
|
||||||
"time"
|
"time"
|
||||||
"unicode/utf8"
|
"unicode/utf8"
|
||||||
@@ -333,7 +334,7 @@ type flashView struct {
|
|||||||
// flashMessages maps each flash code to its banner. An unknown code renders no
|
// flashMessages maps each flash code to its banner. An unknown code renders no
|
||||||
// banner (flashFor returns ok=false), so a forged cookie value is inert.
|
// banner (flashFor returns ok=false), so a forged cookie value is inert.
|
||||||
var flashMessages = map[string]flashView{
|
var flashMessages = map[string]flashView{
|
||||||
flashConnected: {"success", "YouTube account connected."},
|
flashConnected: {"success", "YouTube account connected — finding your subscriptions. Your newest videos will appear below as they're summarized."},
|
||||||
flashConnectFailed: {"error", "Could not connect your YouTube account. Please try again."},
|
flashConnectFailed: {"error", "Could not connect your YouTube account. Please try again."},
|
||||||
flashDisconnected: {"success", "Account disconnected."},
|
flashDisconnected: {"success", "Account disconnected."},
|
||||||
flashDeleted: {"success", "Your account and all its data were deleted."},
|
flashDeleted: {"success", "Your account and all its data were deleted."},
|
||||||
@@ -469,7 +470,7 @@ func (b listBuckets) empty() bool {
|
|||||||
// Dates are kept as the raw YYYY-MM-DD strings so the form re-renders the user's
|
// Dates are kept as the raw YYYY-MM-DD strings so the form re-renders the user's
|
||||||
// input verbatim; parsing happens in matchFilter.
|
// input verbatim; parsing happens in matchFilter.
|
||||||
type Filter struct {
|
type Filter struct {
|
||||||
Channel string
|
Channels []string // selected channel titles; empty = all channels
|
||||||
From string
|
From string
|
||||||
To string
|
To string
|
||||||
OnlySummarized bool // show only videos that have a summary
|
OnlySummarized bool // show only videos that have a summary
|
||||||
@@ -480,7 +481,28 @@ type Filter struct {
|
|||||||
// filter) the bar is hidden so the connect CTA stands alone (UX review C1); a
|
// filter) the bar is hidden so the connect CTA stands alone (UX review C1); a
|
||||||
// filter that happens to match nothing still shows the bar so it can be cleared.
|
// filter that happens to match nothing still shows the bar so it can be cleared.
|
||||||
func (f Filter) active() bool {
|
func (f Filter) active() bool {
|
||||||
return f.Channel != "" || f.From != "" || f.To != "" || f.OnlySummarized
|
return len(f.Channels) > 0 || f.From != "" || f.To != "" || f.OnlySummarized
|
||||||
|
}
|
||||||
|
|
||||||
|
// HasChannel reports whether a channel is currently selected (drives the
|
||||||
|
// multi-select's selected state in the view).
|
||||||
|
func (f Filter) HasChannel(c string) bool {
|
||||||
|
return slices.Contains(f.Channels, c)
|
||||||
|
}
|
||||||
|
|
||||||
|
// nonEmptyStrings drops blank entries. A channel multi-select submits real
|
||||||
|
// channel titles; this guards against a stray empty value reaching the filter.
|
||||||
|
func nonEmptyStrings(ss []string) []string {
|
||||||
|
out := ss[:0:0]
|
||||||
|
for _, s := range ss {
|
||||||
|
if strings.TrimSpace(s) != "" {
|
||||||
|
out = append(out, s)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
if len(out) == 0 {
|
||||||
|
return nil
|
||||||
|
}
|
||||||
|
return out
|
||||||
}
|
}
|
||||||
|
|
||||||
// matches reports whether a row satisfies the filter. Channel is an exact match;
|
// matches reports whether a row satisfies the filter. Channel is an exact match;
|
||||||
@@ -491,7 +513,7 @@ func (f Filter) matches(r store.SummaryRow) bool {
|
|||||||
if f.OnlySummarized && !r.Summarized {
|
if f.OnlySummarized && !r.Summarized {
|
||||||
return false
|
return false
|
||||||
}
|
}
|
||||||
if f.Channel != "" && r.Channel != f.Channel {
|
if len(f.Channels) > 0 && !slices.Contains(f.Channels, r.ChannelTitle) {
|
||||||
return false
|
return false
|
||||||
}
|
}
|
||||||
if from, ok := parseDate(f.From); ok {
|
if from, ok := parseDate(f.From); ok {
|
||||||
@@ -521,7 +543,7 @@ func parseDate(s string) (time.Time, bool) {
|
|||||||
|
|
||||||
// apply returns the subset of rows matching the filter, preserving order.
|
// apply returns the subset of rows matching the filter, preserving order.
|
||||||
func (f Filter) apply(rows []store.SummaryRow) []store.SummaryRow {
|
func (f Filter) apply(rows []store.SummaryRow) []store.SummaryRow {
|
||||||
if f.Channel == "" && f.From == "" && f.To == "" && !f.OnlySummarized {
|
if len(f.Channels) == 0 && f.From == "" && f.To == "" && !f.OnlySummarized {
|
||||||
return rows
|
return rows
|
||||||
}
|
}
|
||||||
out := rows[:0:0]
|
out := rows[:0:0]
|
||||||
|
|||||||
@@ -104,23 +104,35 @@ templ flashBanner(code string) {
|
|||||||
// #summary-list region; a non-HTMX request renders the whole page. flash carries
|
// #summary-list region; a non-HTMX request renders the whole page. flash carries
|
||||||
// a one-shot notification (e.g. "connected", "registered") surfaced on arrival
|
// a one-shot notification (e.g. "connected", "registered") surfaced on arrival
|
||||||
// after a POST→redirect.
|
// after a POST→redirect.
|
||||||
templ ListPage(b listBuckets, f Filter, stats PipelineStats, flash string, hasConnected bool) {
|
templ ListPage(b listBuckets, f Filter, stats PipelineStats, flash string, hasConnected bool, channels []string, autoSummarize bool) {
|
||||||
@Layout("Tapir — Summaries") {
|
@Layout("Tapir — Summaries") {
|
||||||
@flashBanner(flash)
|
@flashBanner(flash)
|
||||||
|
if hasConnected {
|
||||||
|
@pasteForm()
|
||||||
|
}
|
||||||
if !b.empty() || f.active() {
|
if !b.empty() || f.active() {
|
||||||
@filterForm(f)
|
@filterForm(f, channels)
|
||||||
}
|
}
|
||||||
if stats.RateLimited > 0 || stats.Pending > 0 || stats.NoText > 0 {
|
if stats.RateLimited > 0 || stats.Pending > 0 || stats.NoText > 0 {
|
||||||
@pipelineBar(stats)
|
@pipelineBar(stats)
|
||||||
}
|
}
|
||||||
if stats.RateLimited+stats.Pending > 0 {
|
if (stats.RateLimited+stats.Pending) > 0 && autoSummarize {
|
||||||
<p class="pipeline-note muted">
|
<p class="pipeline-note muted">
|
||||||
Tapir fetches captions slowly on purpose, to respect YouTube's limits —
|
Tapir fetches captions slowly on purpose, to respect YouTube's limits —
|
||||||
new summaries land gradually. Check back tomorrow.
|
new summaries land gradually. Check back tomorrow.
|
||||||
</p>
|
</p>
|
||||||
}
|
}
|
||||||
|
if (stats.RateLimited+stats.Pending) > 0 && !autoSummarize {
|
||||||
|
<p class="pipeline-note muted">
|
||||||
|
You are in Manual mode: new videos appear here but are not summarized
|
||||||
|
automatically. Use the Summarize button on the ones you want.
|
||||||
|
</p>
|
||||||
|
<p class="pipeline-note muted">
|
||||||
|
<a href="/account">Switch to Automatic</a> to have new videos summarized for you.
|
||||||
|
</p>
|
||||||
|
}
|
||||||
<div id="summary-list">
|
<div id="summary-list">
|
||||||
@summaryList(b, hasConnected)
|
@summaryList(b, hasConnected, autoSummarize)
|
||||||
</div>
|
</div>
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
@@ -143,7 +155,29 @@ templ pipelineBar(s PipelineStats) {
|
|||||||
</div>
|
</div>
|
||||||
}
|
}
|
||||||
|
|
||||||
templ filterForm(f Filter) {
|
// pasteForm lets a connected user summarize any YouTube video by pasting its URL
|
||||||
|
// (Feature 2). The result (a video card, or an inline error) swaps into
|
||||||
|
// #paste-result; the next list refresh shows it inline. Summarization runs
|
||||||
|
// through the shared caption rate gate like every other fetch.
|
||||||
|
templ pasteForm() {
|
||||||
|
<form
|
||||||
|
class="paste"
|
||||||
|
method="post"
|
||||||
|
action="/paste"
|
||||||
|
hx-post="/paste"
|
||||||
|
hx-target="#paste-result"
|
||||||
|
hx-swap="innerHTML"
|
||||||
|
>
|
||||||
|
<label>
|
||||||
|
Summarize any video
|
||||||
|
<input type="url" name="url" placeholder="Paste a YouTube link…" required/>
|
||||||
|
</label>
|
||||||
|
<button type="submit">Add</button>
|
||||||
|
</form>
|
||||||
|
<div id="paste-result"></div>
|
||||||
|
}
|
||||||
|
|
||||||
|
templ filterForm(f Filter, channels []string) {
|
||||||
<form
|
<form
|
||||||
class="filters"
|
class="filters"
|
||||||
method="get"
|
method="get"
|
||||||
@@ -153,7 +187,16 @@ templ filterForm(f Filter) {
|
|||||||
hx-swap="innerHTML"
|
hx-swap="innerHTML"
|
||||||
hx-indicator="#filter-indicator"
|
hx-indicator="#filter-indicator"
|
||||||
>
|
>
|
||||||
<label>Channel <input type="text" name="channel" value={ f.Channel } placeholder="any"/></label>
|
if len(channels) > 0 {
|
||||||
|
<label>
|
||||||
|
Channels
|
||||||
|
<select name="channel" multiple size="4">
|
||||||
|
for _, c := range channels {
|
||||||
|
<option value={ c } selected?={ f.HasChannel(c) }>{ c }</option>
|
||||||
|
}
|
||||||
|
</select>
|
||||||
|
</label>
|
||||||
|
}
|
||||||
<label class="filter-check">
|
<label class="filter-check">
|
||||||
<input type="checkbox" name="summarized" value="1" if f.OnlySummarized { checked }/>
|
<input type="checkbox" name="summarized" value="1" if f.OnlySummarized { checked }/>
|
||||||
Summarized only
|
Summarized only
|
||||||
@@ -169,12 +212,16 @@ templ filterForm(f Filter) {
|
|||||||
// and a single disclosure holding the older un-summarized back-catalogue. Cards
|
// and a single disclosure holding the older un-summarized back-catalogue. Cards
|
||||||
// reflow to a single column on mobile; an empty list shows a friendly first-run
|
// reflow to a single column on mobile; an empty list shows a friendly first-run
|
||||||
// state instead of a blank table.
|
// state instead of a blank table.
|
||||||
templ summaryList(b listBuckets, hasConnected bool) {
|
templ summaryList(b listBuckets, hasConnected bool, autoSummarize bool) {
|
||||||
if b.empty() {
|
if b.empty() {
|
||||||
if hasConnected {
|
if hasConnected {
|
||||||
<div class="empty empty-connected">
|
<div class="empty empty-connected">
|
||||||
<strong>Your account is connected</strong>
|
<strong>Your account is connected</strong>
|
||||||
<span>Tapir is finding your subscriptions and fetching captions — summaries appear here gradually. Check back later.</span>
|
if autoSummarize {
|
||||||
|
<span>Tapir is finding your subscriptions and fetching captions — summaries appear here gradually. Check back later.</span>
|
||||||
|
} else {
|
||||||
|
<span>Tapir is finding your subscriptions. You are in Manual mode, so videos appear here with a Summarize button — pick the ones you want, or switch to Automatic in your account.</span>
|
||||||
|
}
|
||||||
</div>
|
</div>
|
||||||
} else {
|
} else {
|
||||||
<div class="empty">
|
<div class="empty">
|
||||||
|
|||||||
+712
-601
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,221 @@
|
|||||||
|
package acceptance
|
||||||
|
|
||||||
|
// This is the name-coverage gate for the BDD spec (see docs/use-cases/*.feature).
|
||||||
|
// There is no godog runner — the .feature files are design records, and the real
|
||||||
|
// behaviour is covered by the hand-written Go tests across the module. This test
|
||||||
|
// keeps the two from drifting in the cheapest honest way: every non-@pending
|
||||||
|
// Scenario must have an entry in scenarioCoverage pointing at a Go test that
|
||||||
|
// actually exists. It does NOT prove the test exercises the scenario (only godog
|
||||||
|
// could); it catches the common drift — "added a scenario, forgot the test", a
|
||||||
|
// renamed/deleted covering test, or a scenario removed without cleaning the map.
|
||||||
|
//
|
||||||
|
// When you add a Scenario: either map it here to its covering test, or tag it
|
||||||
|
// @pending in the .feature with a one-line reason for why it has no test yet.
|
||||||
|
|
||||||
|
import (
|
||||||
|
"os"
|
||||||
|
"path/filepath"
|
||||||
|
"regexp"
|
||||||
|
"strings"
|
||||||
|
"testing"
|
||||||
|
)
|
||||||
|
|
||||||
|
// scenarioCoverage maps each non-@pending Scenario name to the Go test that
|
||||||
|
// covers it. Keep it in sync with docs/use-cases/*.feature — the test below
|
||||||
|
// fails if a scenario is unmapped, a mapped test is missing, or an entry no
|
||||||
|
// longer matches a real non-pending scenario.
|
||||||
|
var scenarioCoverage = map[string]string{
|
||||||
|
// ai_routing.feature
|
||||||
|
"Local AI produces the summary": "TestSummarize_LocalSucceeds",
|
||||||
|
"Local AI fails and the user has a BYO provider configured": "TestSummarize_FallsBackToBYO",
|
||||||
|
"Local AI fails and the user has no BYO provider": "TestSummarize_LocalFailsNoBYO_NoExternalSend",
|
||||||
|
"A user without BYO never has content sent externally": "TestSummarize_NoBYO_ContentOnlyLocal",
|
||||||
|
"A model returns unparseable output and the next endpoint succeeds": "TestSummarize_FallsBackOnMalformedOutput",
|
||||||
|
|
||||||
|
// landing_page.feature
|
||||||
|
"An unauthenticated visit to the root is sent to the welcome page": "TestUnauthenticatedRootRedirectsToWelcome",
|
||||||
|
"The welcome page invites an unauthenticated visitor to start": "TestWelcomeLoggedOut",
|
||||||
|
"An authenticated user on the welcome page sees their way in and out": "TestWelcomeLoggedIn",
|
||||||
|
|
||||||
|
// connect_account.feature
|
||||||
|
"Connect a YouTube account": "TestCallbackExchangesAndRecordsConnection",
|
||||||
|
"Connecting an account discovers videos immediately": "TestCallbackTriggersDiscovery",
|
||||||
|
"Connecting summarizes my newest videos right away": "TestNewestUnsummarizedVideoIDs",
|
||||||
|
"Tokens are never stored in the clear": "TestCallbackExchangesAndRecordsConnection",
|
||||||
|
|
||||||
|
// paste_url.feature
|
||||||
|
"Paste a valid YouTube URL": "TestPasteValidURLAddsAndRequests",
|
||||||
|
"Pasting an invalid link is rejected": "TestPasteInvalidURLRejected",
|
||||||
|
"Pasting a video that cannot be found is honest": "TestPasteVideoNotFound",
|
||||||
|
"Pasting the same video twice does not duplicate it": "TestPasteDedupNoDuplicate",
|
||||||
|
"Revoking a connection stops watching but keeps history": "TestDisconnectRemovesTokenAndConnectionKeepsAccount",
|
||||||
|
|
||||||
|
// summarize_mode.feature
|
||||||
|
"Auto mode summarizes recent new videos automatically": "TestRunOnce_AutoMode_SkipsOldVideos",
|
||||||
|
"Auto mode lists older videos without summarizing them": "TestRunOnce_AutoMode_SkipsOldVideos",
|
||||||
|
"Automatic is the default for a new user": "TestRegisteredUserDefaultsAutoSummarizeOn",
|
||||||
|
"Manual mode leaves new videos unsummarized": "TestRunOnce_ManualMode_SkipsUnrequested",
|
||||||
|
"Requesting a summary in manual mode queues it for the next run": "TestRunOnce_ManualMode_ProcessesRequested",
|
||||||
|
|
||||||
|
// registration.feature
|
||||||
|
"A new Dex subject is routed to registration": "TestUnregisteredSubjectRedirectedToRegister",
|
||||||
|
"Registering creates the account and its identity mapping": "TestRegisterCreatesExactlyOneUserAndIdentity",
|
||||||
|
"A returning subject passes straight through": "TestRegisteredSubjectPassesThrough",
|
||||||
|
"Deleting an account removes only my data and leaves other users untouched": "TestDeleteAccountWipesDataAndSecretsAndLogsOut",
|
||||||
|
|
||||||
|
// summarize_new_video.feature
|
||||||
|
"A subscribed channel posts a video that has captions": "TestSubscribedVideoWithCaptionsIsSummarizedAndDelivered",
|
||||||
|
"A subscribed channel posts a video with no usable transcript": "TestVideoWithNoTranscriptIsSkipped",
|
||||||
|
"A channel I am not subscribed to posts a video": "TestUnsubscribedChannelVideoIsNotProcessed",
|
||||||
|
"The same video is not summarized twice": "TestAlreadySummarizedVideoIsNotReprocessed",
|
||||||
|
"Re-analyzing a stored video does not re-fetch its transcript": "TestProcessNewVideo_SecondSummarizeDoesNotRefetch",
|
||||||
|
}
|
||||||
|
|
||||||
|
var (
|
||||||
|
scenarioRe = regexp.MustCompile(`^\s*Scenario(?: Outline)?:\s*(.+?)\s*$`)
|
||||||
|
testFuncRe = regexp.MustCompile(`^func (Test\w+)\(`)
|
||||||
|
)
|
||||||
|
|
||||||
|
// scenario is one parsed Gherkin scenario and whether it is @pending.
|
||||||
|
type scenario struct {
|
||||||
|
name string
|
||||||
|
pending bool
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestScenarioCoverage(t *testing.T) {
|
||||||
|
root := moduleRoot(t)
|
||||||
|
|
||||||
|
scenarios := parseScenarios(t, filepath.Join(root, "docs", "use-cases"))
|
||||||
|
if len(scenarios) == 0 {
|
||||||
|
t.Fatal("no scenarios parsed from docs/use-cases — wrong path?")
|
||||||
|
}
|
||||||
|
tests := allTestFuncNames(t, root)
|
||||||
|
|
||||||
|
// Index scenario names for the reverse (stale-entry) check.
|
||||||
|
active := map[string]bool{} // non-pending scenario names
|
||||||
|
var pending []string
|
||||||
|
for _, s := range scenarios {
|
||||||
|
if s.pending {
|
||||||
|
pending = append(pending, s.name)
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
active[s.name] = true
|
||||||
|
|
||||||
|
// 1. Every non-pending scenario must be mapped.
|
||||||
|
fn, ok := scenarioCoverage[s.name]
|
||||||
|
if !ok {
|
||||||
|
t.Errorf("scenario %q has no coverage entry — map it in scenarioCoverage to a covering test, or tag it @pending in the .feature", s.name)
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
// 2. The mapped test must actually exist.
|
||||||
|
if !tests[fn] {
|
||||||
|
t.Errorf("scenario %q maps to %q, which is not a Test function anywhere in the module", s.name, fn)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// 3. No stale entries: every map key must be a real, non-pending scenario.
|
||||||
|
for name := range scenarioCoverage {
|
||||||
|
if !active[name] {
|
||||||
|
t.Errorf("scenarioCoverage has entry %q, which is not a current non-pending scenario (renamed, removed, or now @pending?)", name)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
if len(pending) > 0 {
|
||||||
|
t.Logf("%d @pending scenario(s) without a test (tracked, not required): %s",
|
||||||
|
len(pending), strings.Join(pending, "; "))
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// parseScenarios reads every *.feature under dir and returns its scenarios with
|
||||||
|
// their @pending status. A scenario is @pending when a `@pending` tag line
|
||||||
|
// precedes it (tags survive intervening comment lines, the layout these files
|
||||||
|
// use); the flag is consumed at the Scenario line and reset afterward.
|
||||||
|
func parseScenarios(t *testing.T, dir string) []scenario {
|
||||||
|
t.Helper()
|
||||||
|
entries, err := os.ReadDir(dir)
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("read use-cases dir: %v", err)
|
||||||
|
}
|
||||||
|
var out []scenario
|
||||||
|
for _, e := range entries {
|
||||||
|
if e.IsDir() || !strings.HasSuffix(e.Name(), ".feature") {
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
b, err := os.ReadFile(filepath.Join(dir, e.Name()))
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("read %s: %v", e.Name(), err)
|
||||||
|
}
|
||||||
|
pending := false
|
||||||
|
for _, line := range strings.Split(string(b), "\n") {
|
||||||
|
trimmed := strings.TrimSpace(line)
|
||||||
|
if strings.HasPrefix(trimmed, "@") {
|
||||||
|
if strings.Contains(trimmed, "@pending") {
|
||||||
|
pending = true
|
||||||
|
}
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
if m := scenarioRe.FindStringSubmatch(line); m != nil {
|
||||||
|
out = append(out, scenario{name: m[1], pending: pending})
|
||||||
|
pending = false
|
||||||
|
}
|
||||||
|
// comment (#) and step lines leave a set @pending intact until the
|
||||||
|
// scenario consumes it; a blank line between scenarios is harmless.
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return out
|
||||||
|
}
|
||||||
|
|
||||||
|
// allTestFuncNames walks the module for `func TestXxx(` declarations, excluding
|
||||||
|
// this file (whose regex literal would otherwise look like a definition).
|
||||||
|
func allTestFuncNames(t *testing.T, root string) map[string]bool {
|
||||||
|
t.Helper()
|
||||||
|
self := "scenario_coverage_test.go"
|
||||||
|
names := map[string]bool{}
|
||||||
|
err := filepath.WalkDir(root, func(path string, d os.DirEntry, err error) error {
|
||||||
|
if err != nil {
|
||||||
|
return err
|
||||||
|
}
|
||||||
|
if d.IsDir() {
|
||||||
|
if d.Name() == ".git" {
|
||||||
|
return filepath.SkipDir
|
||||||
|
}
|
||||||
|
return nil
|
||||||
|
}
|
||||||
|
if !strings.HasSuffix(path, "_test.go") || filepath.Base(path) == self {
|
||||||
|
return nil
|
||||||
|
}
|
||||||
|
b, err := os.ReadFile(path)
|
||||||
|
if err != nil {
|
||||||
|
return err
|
||||||
|
}
|
||||||
|
for _, line := range strings.Split(string(b), "\n") {
|
||||||
|
if m := testFuncRe.FindStringSubmatch(line); m != nil {
|
||||||
|
names[m[1]] = true
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return nil
|
||||||
|
})
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("walk module: %v", err)
|
||||||
|
}
|
||||||
|
return names
|
||||||
|
}
|
||||||
|
|
||||||
|
// moduleRoot walks up from the working directory to the dir containing go.mod.
|
||||||
|
func moduleRoot(t *testing.T) string {
|
||||||
|
t.Helper()
|
||||||
|
dir, err := os.Getwd()
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("getwd: %v", err)
|
||||||
|
}
|
||||||
|
for {
|
||||||
|
if _, err := os.Stat(filepath.Join(dir, "go.mod")); err == nil {
|
||||||
|
return dir
|
||||||
|
}
|
||||||
|
parent := filepath.Dir(dir)
|
||||||
|
if parent == dir {
|
||||||
|
t.Fatal("go.mod not found walking up from cwd")
|
||||||
|
}
|
||||||
|
dir = parent
|
||||||
|
}
|
||||||
|
}
|
||||||
Reference in New Issue
Block a user