# Spec — Onboarding "wow" burst: better picks, stronger model **Repo:** tapir · **Size:** medium · **Solo session** (not a swarm). > **Status: built (v0.25.0, ADR-028).** This supersedes the original investigate-first brief > (committed as the prior version of this file): Phase 1 was run against the live pilot DB and its > findings are folded into "Why this exists" below; Phase 2 was built as described here. The one > brief lever NOT built — the honest "the rest fill in over the coming days" framing copy — is > listed under *Explicitly NOT in this slice*. **Why this exists.** A new user's first session decides whether they return (the Stage-0 gate, VISION.md). On connect, Tapir fires a capped burst (≤`TAPIR_ONBOARD_SUMMARIZE_COUNT`, default 3) that summarizes the user's newest unsummarized videos so the feed isn't empty (the burst itself works — wired in `cmd/tapir/discovery.go` → `cmd/tapir/main.go` `onboard`). A Phase-1 investigation of the live pilot DB found the burst *fires* but delivers a **weak first impression** for two concrete reasons, and ruled out a third idea: 1. **Picks are junk.** Selection is pure newest-first (`videos.NewestUnsummarizedVideoIDs`, `ORDER BY published_at DESC`) with **zero quality signal**. Pilot user "Jonte"'s live burst-3 were a stock-ticker **livestream** + two regional news clips — the newest, not the best. 2. **Weakest model on the first impression.** All of Jonte's summaries ran on `koala/phi4-mini` (the documented weak link — ADR-022 was born from its failures). The stronger, brain-validated `iguana/gemma4-26b` was never used for the burst. 3. **Cached-first is empty at pilot scale — REJECTED.** The idea (summarize already-cached transcripts instantly, zero fetch) dies on the numbers: only **11 videos** overlap between the two pilot users (~3% of each library), **0** cached-and-unsummarized, and a new user's newest-20 unsummarized are **20/20 NOT cached** — newest-first and cached-first are structurally incompatible (fresh uploads are exactly what nobody has fetched yet). Not built. This is a **curation/latency problem for ~3 videos, NOT a throughput/429 problem** — fetching 3 captions is nowhere near the rate limit. Nothing here fetches harder or pressures the rate gate; it picks the right few videos and runs a better model on them. Read `CLAUDE.md`, `DECISIONS.md` (esp. ADR-014, ADR-018, ADR-020, ADR-021, ADR-022, ADR-023, and the new **ADR-028**), and `VISION.md` (the Stage-0 gate) first. TBD — commit directly to `main`, one logical change per commit, conventional commits, `task check` green before each commit, `templ generate` if any view changes (none expected). ## Decisions already made (do not reopen) - **Not a throughput change.** The caption rate gate (ADR-014) is untouched — same pacing, same priority lane (ADR-026). This slice changes *which* ≤3 videos the burst spends its fetches on and *which model* summarizes them, never how fast or how many. - **Cached-first is dropped** (ADR-028, the 3% overlap). The engine's existing read-stored-first (ADR-021, `resolveTranscript`) stays — it already gives a free instant summary on the rare cache hit, transparently. We do not *select* for cache hits. - **has-captions is not a pre-fetch signal.** It is only knowable after a gate fetch (or a cache hit, ~0 for new videos). Selection can only *avoid known-junk* (Shorts/live/over-long) — it cannot *guarantee* captions. The spec is honest about this: better odds, not a promise. - **No credentialed caption fetch** (ADR-010/ADR-026 dead end). **No client extension.** ## 1. Persist `duration_s` at discovery (the enabling change) The `videos.duration_s` column exists (migration 001) but is **never written** — ADR-023's `filterLowValue` (`internal/adapters/youtube/youtube.go`) already fetches each candidate's duration via the cheap quota `videos.list` call, uses it to drop Shorts/live, then **discards it**. Stop discarding: - Add `DurationSeconds int` to `domain.Video`. - In `filterLowValue`, set `DurationSeconds` on each kept video from the `videos.list` `meta`. - `UpsertVideo` writes `duration_s`, **COALESCE-preserving** a known value (never overwrite a real duration with 0/unknown), mirroring the `channel_title` backfill stance (migration 014). - No new migration — the column is already there. Consequence: a fresh user's connect-triggered discovery pass runs **before** the onboard burst (`Enqueue`: `run()` then `onboard()`), so duration is populated for the burst's candidates at connect. Existing rows backfill on their next discovery pass; until then their `duration_s` is NULL and treated as "unknown" (§2). ## 2. Junk-avoiding burst selection New store method, RLS-scoped via `withUser`: ``` OnboardBurstVideoIDs(ctx, userID string, limit, minSeconds, maxSeconds int) ([]string, error) ``` - Same base as the old `NewestUnsummarizedVideoIDs`: the user's videos with no summary yet, `ORDER BY published_at DESC NULLS LAST, seen_at DESC`, `LIMIT limit`. - **Exclude known-junk**: a row is dropped only when `duration_s IS NOT NULL` **and** (`duration_s < minSeconds` OR `duration_s > maxSeconds`). A NULL duration is **unknown** — kept (degrade-open: never starve the burst because metadata is missing), but ordered *after* rows with a known-good duration so a freshly-enriched good pick wins when both exist. - `minSeconds` reuses `TAPIR_MIN_VIDEO_SECONDS` (default 60 — the Shorts floor, ADR-023). `maxSeconds` is new: `TAPIR_ONBOARD_MAX_VIDEO_SECONDS` (default 14400 = 4h) — drops the multi-hour livestream VODs that pass the live filter once ended. - `minSeconds<=0` and `maxSeconds<=0` each disable that bound (so `0/0` == the old newest-first behaviour, the reversibility lever). - The burst switches to this method; `NewestUnsummarizedVideoIDs` is removed (fully superseded — `OnboardBurstVideoIDs(., 0, 0)` is identical pure-newest behaviour). ## 3. Stronger model for the burst The burst summarizes only ≤3 videos, so a slower, stronger model is affordable exactly here. - New config `TAPIR_ONBOARD_SUMMARIZER_MODEL` (default `iguana/gemma4-26b` — the brain-validated homelab general-purpose model, already the ADR-022 fallback). - Build a **burst-specific summarizer chain** that puts the onboard model **first**, then the standard chain (primary → local fallback → cloud) as resilience, deduped. Wrap it in a burst-specific `engineProcessor` reusing the same store/transcript-cache/sink — a pure wiring choice, engine and ports unchanged (Clean Architecture, ADR-003). - The `onboard` closure uses the burst processor instead of `app.Processor`. - **Collapse cleanly**: when `OnboardSummarizerModel` is empty or equals `SummarizerModel`, the onboard path reuses `app.Processor` (no separate chain) — the reversibility lever. - Local-first preserved: the onboard model is a local alias; the cloud endpoint stays last in the chain, so a client/NDA deployment with `TAPIR_CLOUD_FALLBACK_MODEL=""` keeps burst content local too. ## 4. Behaviour spec + docs - Add scenarios to `docs/use-cases/connect_account.feature` (the connect → burst flow): burst skips a too-long/live video in favour of a reasonable-length one; burst summarizes with the stronger model first. Map them in `scenarioCoverage` so `TestScenarioCoverage` stays green. - Update `docs/architecture/architecture.md` (the onboarding-burst section) to describe the junk-avoiding selection + the burst model override. - ADR-028 in `DECISIONS.md` records the rationale (incl. the rejected cached-first lever). ## Success criteria - `task check` green (fmt, vet, lint, `go test -p 1 ./...`). - A unit test proves `OnboardBurstVideoIDs` drops a known too-long / sub-min video and keeps a good one, newest-first, RLS-scoped, unsummarized-only. - A test proves discovery persists `duration_s` and does not clobber it on re-upsert. - A test proves the burst chain leads with the onboard model (then the standard chain). - Config defaults + bounds tested (`OnboardMaxVideoSeconds`, `OnboardSummarizerModel`). - No change to the rate gate, fetch pacing, or burst cap. `0/0` + empty model == prior behaviour. ## Explicitly NOT in this slice - Cached-first selection (rejected, ADR-028). - Any caption-availability *guarantee* (impossible pre-fetch). - **Honest "taster" framing copy** ("summaries of a few of your videos to get you started — the rest fill in over the coming days"). A good lever from the original brief, but it's a UI/copy change with no backend dependency; deferred to a UI pass, tracked as an issue. - Backfilling `duration_s` for existing rows via a migration (it backfills lazily on discovery). - Return-nudges / digests (ADR-020: poisons the unprompted-return signal). - Raising fetch throughput, multi-IP, or Whisper (out of scope; the gate is deliberate).