feat(store): shared, video-keyed transcript persistence (ADR-021)

Reshape the dead per-user transcripts table (PK videos.id, user_id,
RLS-FORCEd — never read or written by app code) into the shared public
caption store ADR-021 specifies: keyed by (provider, provider_video_id),
no user_id, NOT RLS-scoped. Migration 015 (reversible). Add
ports.TranscriptStore + Store.GetTranscript/SaveTranscript via the raw
pool (no withUser): public content, shared across users by construction.
SaveTranscript persists only terminal outcomes (captions/none) and
refuses SourceRateLimited so a transient 429 can never be stored as a
false permanent absence (ADR-014).

Flip the isolation proof: transcripts leaves the RLS-scoped set;
TestTranscriptsTableIsSharedNotRLS asserts it is the SINGLE non-RLS
surface (writable/readable with no user scope, no user_id column, RLS off
on it alone, still on every user-owned table) — the proof the
public-content classification was applied exactly here and leaked nowhere.
appPool made idempotent so two tests can build it. Adjust the 010/011/014
up-down migration tests for the new HEAD. account.go: user deletion no
longer strips shared transcripts. Reconcile data-model.md + CLAUDE.md.

Wiring the engine to read-stored-first is the next commit.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-06-09 23:32:12 +02:00
co-authored by Claude Opus 4.8
parent 099b2d4c68
commit cb6917ca59
10 changed files with 332 additions and 35 deletions
+21 -13
View File
@@ -10,12 +10,15 @@ only opaque references to them; the secret material lives in ESO/1Password (ADR-
## Design decisions baked into this model
- **Per-user isolation, not a shared global video table.** The earlier draft proposed a
global `videos`/`transcripts` table deduped across tenants. Rejected for Future B: it
reintroduces exactly the cross-domain coupling the homelab architecture review is
removing, and at 15 users the cost of occasionally re-summarizing the same video is
trivial compared to the isolation it would cost. Each user's data is self-contained.
(Revisit only if Future C makes GPU/transcription cost dominate — a new ADR, not a default.)
- **Per-user isolation for everything except transcripts.** The earlier draft proposed a
global `videos`/`transcripts` table deduped across tenants. **Videos** stay per-user and
RLS-scoped — a shared video table reintroduces exactly the cross-domain coupling the homelab
architecture review is removing. **Transcripts**, however, are now shared (ADR-021): keyed by
`(provider, provider_video_id)`, no `user_id`, **not** RLS-scoped. The cost avoided there is
not LLM re-summarization but a rate-gated, reputation-risky caption fetch (ADR-010/014), which
is paid per re-fetch regardless of user count — so persisting public caption content once and
sharing it strictly beats the coupling it removes. Everything else each user owns is
self-contained; `rls_test.go` proves transcripts is the single exception.
- **Secrets by reference only.** Tables hold a `secret_ref` (opaque string/UUID resolved via
the `SecretStore` port), never tokens or keys.
- **The brain sink is just a delivery target.** No brain-specific tables. Whether a summary
@@ -34,7 +37,7 @@ erDiagram
USER ||--o{ AI_CREDENTIAL : "has (planned)"
VIDEO_CONNECTION ||--o{ SUBSCRIPTION : "exposes (planned)"
SUBSCRIPTION ||--o{ VIDEO : "produces (per user)"
VIDEO ||--o| TRANSCRIPT : "has at most one"
VIDEO }o--o| TRANSCRIPT : "shares one by (provider, provider_video_id) — not FK (ADR-021)"
VIDEO ||--o| SUMMARY : "has at most one"
SUMMARY ||--o{ SINK_DELIVERY : "delivered via"
USER ||--o{ CHANNEL_ERROR : "reports unavailable channels"
@@ -92,12 +95,12 @@ erDiagram
timestamptz rate_limited_at "backoff clock for 429 retries (migration 007)"
}
TRANSCRIPT {
uuid video_id PK_FK
uuid user_id FK
text provider PK "part of shared key (ADR-021)"
text provider_video_id PK "part of shared key — the cross-user dedup key"
text source "captions | none"
text language
text content "null when source = none"
timestamptz resolved_at
timestamptz fetched_at
}
SUMMARY {
uuid id PK
@@ -173,8 +176,12 @@ mechanism.
`transcript_status` and `rate_limited_at` (migration 007) track caption-fetch outcomes for
rate-limit backoff: `NULL` = not attempted; `rate_limited` = 429 seen, skip until
`NOW() - rate_limited_at > TAPIR_FETCH_BACKOFF`; `fetched` = resolved; `none` = no transcript.
- **TRANSCRIPT** — at most one per video. `source = none` records "checked, no usable
transcript" so the watcher doesn't reprocess (ADR-007). `content` null in that case.
- **TRANSCRIPT** — shared public caption content, one row per `(provider, provider_video_id)`,
**not** RLS-scoped and carrying no `user_id` (ADR-021). Two users who watch the same video
share the one row; the summarize path reads it before any caption fetch, so re-analysis never
re-touches YouTube (ADR-010/014). `source = none` records "checked, no usable transcript" so
no one reprocesses (ADR-007); `content` null in that case. A transient 429 is never stored
here — it stays a per-user retry via `VIDEO.transcript_status`.
- **SUMMARY** — at most one per video. `fallback_used` + `ai_provider`/`ai_model` make the
"is local good enough?" question queryable (the Stage 0 quality signal). `highlights`/
`takeaways` as jsonb to stay schema-flexible while the output format settles.
@@ -231,7 +238,8 @@ queue, doesn't replace it). Deferred until there's a reason.
## Explicitly out of scope (Future C)
- Global cross-tenant video/transcript dedup (rejected above).
- Global cross-tenant *video* dedup (rejected above). Note: cross-tenant *transcript* sharing
is now in scope and shipped (ADR-021); only the videos half stays per-user.
- Sharding / per-tenant physical databases.
- Soft-delete + full audit trail on connections/credentials (a Stage 2 hardening item; add
via ADR when Stage 2 work starts).