Phase-1 live-DB investigation found the connect-time burst (ADR-018) fires but delivers a weak first impression: pure newest-first selection picks junk (a livestream + regional news for pilot user Jonte), and all burst summaries run on the weak koala/phi4-mini instead of the validated iguana/gemma4-26b. Cached-first was investigated and rejected — ~3% cross-user overlap, 0 cached-and-unsummarized, newest-20 all uncached (newest-first and cached-first are structurally incompatible). ADR-028 records: persist duration_s at discovery (ADR-023 already fetches it, just stops discarding), junk-avoiding burst selection (drop known too-short/ too-long), and a stronger burst-only summarizer chain. Not a throughput change.
8.7 KiB
Spec — Onboarding "wow" burst: better picks, stronger model
Repo: tapir · Size: medium · Solo session (not a swarm).
Status: built (v0.25.0, ADR-028). This supersedes the original investigate-first brief (committed as the prior version of this file): Phase 1 was run against the live pilot DB and its findings are folded into "Why this exists" below; Phase 2 was built as described here. The one brief lever NOT built — the honest "the rest fill in over the coming days" framing copy — is listed under Explicitly NOT in this slice.
Why this exists. A new user's first session decides whether they return (the Stage-0 gate,
VISION.md). On connect, Tapir fires a capped burst (≤TAPIR_ONBOARD_SUMMARIZE_COUNT, default 3)
that summarizes the user's newest unsummarized videos so the feed isn't empty (the burst itself
works — wired in cmd/tapir/discovery.go → cmd/tapir/main.go onboard). A Phase-1
investigation of the live pilot DB found the burst fires but delivers a weak first
impression for two concrete reasons, and ruled out a third idea:
- Picks are junk. Selection is pure newest-first (
videos.NewestUnsummarizedVideoIDs,ORDER BY published_at DESC) with zero quality signal. Pilot user "Jonte"'s live burst-3 were a stock-ticker livestream + two regional news clips — the newest, not the best. - Weakest model on the first impression. All of Jonte's summaries ran on
koala/phi4-mini(the documented weak link — ADR-022 was born from its failures). The stronger, brain-validatediguana/gemma4-26bwas never used for the burst. - Cached-first is empty at pilot scale — REJECTED. The idea (summarize already-cached transcripts instantly, zero fetch) dies on the numbers: only 11 videos overlap between the two pilot users (~3% of each library), 0 cached-and-unsummarized, and a new user's newest-20 unsummarized are 20/20 NOT cached — newest-first and cached-first are structurally incompatible (fresh uploads are exactly what nobody has fetched yet). Not built.
This is a curation/latency problem for ~3 videos, NOT a throughput/429 problem — fetching 3 captions is nowhere near the rate limit. Nothing here fetches harder or pressures the rate gate; it picks the right few videos and runs a better model on them.
Read CLAUDE.md, DECISIONS.md (esp. ADR-014, ADR-018, ADR-020, ADR-021, ADR-022, ADR-023,
and the new ADR-028), and VISION.md (the Stage-0 gate) first. TBD — commit directly to
main, one logical change per commit, conventional commits, task check green before each
commit, templ generate if any view changes (none expected).
Decisions already made (do not reopen)
- Not a throughput change. The caption rate gate (ADR-014) is untouched — same pacing, same priority lane (ADR-026). This slice changes which ≤3 videos the burst spends its fetches on and which model summarizes them, never how fast or how many.
- Cached-first is dropped (ADR-028, the 3% overlap). The engine's existing read-stored-first
(ADR-021,
resolveTranscript) stays — it already gives a free instant summary on the rare cache hit, transparently. We do not select for cache hits. - has-captions is not a pre-fetch signal. It is only knowable after a gate fetch (or a cache hit, ~0 for new videos). Selection can only avoid known-junk (Shorts/live/over-long) — it cannot guarantee captions. The spec is honest about this: better odds, not a promise.
- No credentialed caption fetch (ADR-010/ADR-026 dead end). No client extension.
1. Persist duration_s at discovery (the enabling change)
The videos.duration_s column exists (migration 001) but is never written — ADR-023's
filterLowValue (internal/adapters/youtube/youtube.go) already fetches each candidate's
duration via the cheap quota videos.list call, uses it to drop Shorts/live, then discards
it. Stop discarding:
- Add
DurationSeconds inttodomain.Video. - In
filterLowValue, setDurationSecondson each kept video from thevideos.listmeta. UpsertVideowritesduration_s, COALESCE-preserving a known value (never overwrite a real duration with 0/unknown), mirroring thechannel_titlebackfill stance (migration 014).- No new migration — the column is already there.
Consequence: a fresh user's connect-triggered discovery pass runs before the onboard burst
(Enqueue: run() then onboard()), so duration is populated for the burst's candidates at
connect. Existing rows backfill on their next discovery pass; until then their duration_s is
NULL and treated as "unknown" (§2).
2. Junk-avoiding burst selection
New store method, RLS-scoped via withUser:
OnboardBurstVideoIDs(ctx, userID string, limit, minSeconds, maxSeconds int) ([]string, error)
- Same base as the old
NewestUnsummarizedVideoIDs: the user's videos with no summary yet,ORDER BY published_at DESC NULLS LAST, seen_at DESC,LIMIT limit. - Exclude known-junk: a row is dropped only when
duration_s IS NOT NULLand (duration_s < minSecondsORduration_s > maxSeconds). A NULL duration is unknown — kept (degrade-open: never starve the burst because metadata is missing), but ordered after rows with a known-good duration so a freshly-enriched good pick wins when both exist. minSecondsreusesTAPIR_MIN_VIDEO_SECONDS(default 60 — the Shorts floor, ADR-023).maxSecondsis new:TAPIR_ONBOARD_MAX_VIDEO_SECONDS(default 14400 = 4h) — drops the multi-hour livestream VODs that pass the live filter once ended.minSeconds<=0andmaxSeconds<=0each disable that bound (so0/0== the old newest-first behaviour, the reversibility lever).- The burst switches to this method;
NewestUnsummarizedVideoIDsis removed (fully superseded —OnboardBurstVideoIDs(., 0, 0)is identical pure-newest behaviour).
3. Stronger model for the burst
The burst summarizes only ≤3 videos, so a slower, stronger model is affordable exactly here.
- New config
TAPIR_ONBOARD_SUMMARIZER_MODEL(defaultiguana/gemma4-26b— the brain-validated homelab general-purpose model, already the ADR-022 fallback). - Build a burst-specific summarizer chain that puts the onboard model first, then the
standard chain (primary → local fallback → cloud) as resilience, deduped. Wrap it in a
burst-specific
engineProcessorreusing the same store/transcript-cache/sink — a pure wiring choice, engine and ports unchanged (Clean Architecture, ADR-003). - The
onboardclosure uses the burst processor instead ofapp.Processor. - Collapse cleanly: when
OnboardSummarizerModelis empty or equalsSummarizerModel, the onboard path reusesapp.Processor(no separate chain) — the reversibility lever. - Local-first preserved: the onboard model is a local alias; the cloud endpoint stays last in the
chain, so a client/NDA deployment with
TAPIR_CLOUD_FALLBACK_MODEL=""keeps burst content local too.
4. Behaviour spec + docs
- Add scenarios to
docs/use-cases/connect_account.feature(the connect → burst flow): burst skips a too-long/live video in favour of a reasonable-length one; burst summarizes with the stronger model first. Map them inscenarioCoveragesoTestScenarioCoveragestays green. - Update
docs/architecture/architecture.md(the onboarding-burst section) to describe the junk-avoiding selection + the burst model override. - ADR-028 in
DECISIONS.mdrecords the rationale (incl. the rejected cached-first lever).
Success criteria
task checkgreen (fmt, vet, lint,go test -p 1 ./...).- A unit test proves
OnboardBurstVideoIDsdrops a known too-long / sub-min video and keeps a good one, newest-first, RLS-scoped, unsummarized-only. - A test proves discovery persists
duration_sand does not clobber it on re-upsert. - A test proves the burst chain leads with the onboard model (then the standard chain).
- Config defaults + bounds tested (
OnboardMaxVideoSeconds,OnboardSummarizerModel). - No change to the rate gate, fetch pacing, or burst cap.
0/0+ empty model == prior behaviour.
Explicitly NOT in this slice
- Cached-first selection (rejected, ADR-028).
- Any caption-availability guarantee (impossible pre-fetch).
- Honest "taster" framing copy ("summaries of a few of your videos to get you started — the rest fill in over the coming days"). A good lever from the original brief, but it's a UI/copy change with no backend dependency; deferred to a UI pass, tracked as an issue.
- Backfilling
duration_sfor existing rows via a migration (it backfills lazily on discovery). - Return-nudges / digests (ADR-020: poisons the unprompted-return signal).
- Raising fetch throughput, multi-IP, or Whisper (out of scope; the gate is deliberate).