In automatic mode the scheduler now only summarizes videos published within
TAPIR_AUTO_SUMMARIZE_WINDOW (default ~7d). Older videos are still discovered
and listed — they keep the manual "Summarize" affordance — but are not
auto-processed, so a large back-catalogue (the maintainer's ~256-deep queue)
stops self-inflicting 429s against the per-IP caption gate each cycle (UX
review B1, recency design).
- runner.WithAutoWindow + Stats.SkippedTooOld; tooOld() treats a zero window
as disabled and an undated video as never-aged-out (processed, not stranded).
- An explicit manual request bypasses the bound even in auto mode (requested
videos are loaded in auto mode when a window is active).
- Wired through cmdRun, the scheduler's per-user runner, sumStats, and pass
logging. config: TAPIR_AUTO_SUMMARIZE_WINDOW (default 168h), .env.example.
- Account copy (A7) updated to match: "Automatic summarizes new videos from
about the last week; older videos stay browsable — summarize on demand."
The rate gate is untouched; the manual path still serialises through it. This
bounds auto LOAD, it does not fetch harder.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Restructures RunOnce from per-channel inline processing to collect-sort-process:
Phase 1 — discover, persist (UpsertVideo), apply pre-filters (seen/manual/backoff)
and collect surviving candidates with their discovery position.
Phase 2 — sort candidates by published_at DESC, NULLS LAST, pos ASC tiebreak
so videos with no publish date never jump ahead of dated content.
Phase 3 — process in sorted order through the unchanged globalFetchGate.
Before (per-channel): chanA=[v-old, v-mid], chanB=[v-new, v-null]
→ [v-old, v-mid, v-new, v-null]
After (newest-first): [v-new, v-mid, v-old, v-null]
Same set of videos processed; only the order changes within a pass. All existing
behaviour is preserved: failure isolation, backoff skip, manual mode,
channel-unavailable, stats. In-memory sort; no new table or persisted queue.
The ordering is onboarding prioritisation — new users get summaries of their most
recent, relevant videos first; the back-catalogue fills in behind across subsequent
passes. Both this background batch and the foreground 'Try now' button honour the
shared globalFetchGate: rate limiting is respected, not evaded.
YouTube channels that 404 on playlist discovery (deleted/private) are now:
1. Wrapped in domain.ErrChannelUnavailable by the YouTube adapter (instead of
a generic error), so the runner can identify them without string-matching.
2. Stored per-user in channel_errors (migration 013, RLS-guarded) via runner's
new UpsertChannelError path — removed from the generic Errors counter,
counted separately as ChannelUnavailable.
3. Shown on the account page under "Unavailable channels" with name, chip-warn
badge, and first-seen date, so users know why some subscribed channels
produce no videos.
After ProcessNewVideo the runner records transcript_status per outcome:
rate_limited (stamps the backoff clock), none, or fetched. Before fetching, a
video inside the TAPIR_FETCH_BACKOFF window is skipped (SkippedRateLimited) so a
just-429'd caption endpoint is not re-hit; once the window expires it retries.
Backoff/clock injected via variadic Options (WithBackoff, WithClock) so existing
New call sites and the fake-driven loop tests stay valid. Backoff 0 = always retry.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
RunOnce now reads the user's mode (GetAutoSummarize) at the start of each pass:
- Auto (unchanged): summarize every unseen video.
- Manual: still UpsertVideo for every candidate (discovery — the user sees new
videos in the list), but skip ProcessNewVideo unless the video is queued
(RequestedVideoIDs). A queued video is summarized, then its flag is cleared
(ClearSummarizeRequested) so it is not re-processed and the UI drops "Queued".
New Stats.SkippedManual counts discovered-but-unqueued videos. The VideoStore
port gains the three methods; the existing auto-mode tests set auto:true.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Diagnosis of a live run: caption tracks resolve fine (player + watch-page scrape),
but the timedtext baseUrl fetch returns 429 under back-to-back volume — YouTube
rate-limits the unauthenticated caption-download endpoint per IP. A per-video
delay spaces the fetches. (Follow-up: treat 429 distinctly from genuine
no-caption instead of silently degrading to SourceNone; consider Whisper if the
endpoint stays hostile at any sustainable rate.)
RunOnce walks the user's subscriptions, upserts each candidate video (assigning
its durable store id), skips videos already summarized via the store's
SeenVideoIDs (cross-restart dedup the engine's in-memory map can't provide),
and processes the rest through the engine. Loop adds an optional poll cadence;
per-item errors are collected, not fatal. Tested with fakes — no live deps.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>