docs: spec onboarding-wow-burst + ADR-028 (better picks, stronger model)

Phase-1 live-DB investigation found the connect-time burst (ADR-018) fires
but delivers a weak first impression: pure newest-first selection picks junk
(a livestream + regional news for pilot user Jonte), and all burst summaries
run on the weak koala/phi4-mini instead of the validated iguana/gemma4-26b.
Cached-first was investigated and rejected — ~3% cross-user overlap, 0
cached-and-unsummarized, newest-20 all uncached (newest-first and cached-first
are structurally incompatible).

ADR-028 records: persist duration_s at discovery (ADR-023 already fetches it,
just stops discarding), junk-avoiding burst selection (drop known too-short/
too-long), and a stronger burst-only summarizer chain. Not a throughput change.
This commit is contained in:
2026-06-11 18:26:42 +02:00
parent 19ca4282a8
commit 15926fac0e
2 changed files with 185 additions and 0 deletions
+69
View File
@@ -1060,6 +1060,74 @@ no migration, no stored state to unwind. Spec: `docs/specs/chat-with-transcript.
---
## ADR-028 — Onboarding burst: pick likely-good videos, summarize them with a stronger model
**Status:** Accepted (2026-06-11). **Refines ADR-018** (the connect-time burst) and **ADR-020**
(recency-bounded auto-summarize). **Builds on ADR-022** (the endpoint chain), **ADR-023**
(the discovery-time `videos.list` enrichment), and **ADR-021** (the shared transcript cache).
Triggered by a Phase-1 investigation of the live pilot DB.
**Context.** A new user's first session decides whether they return (the Stage-0 gate, ADR-016).
The connect-time burst (ADR-018: summarize ≤`TAPIR_ONBOARD_SUMMARIZE_COUNT` newest videos so the
feed isn't empty) *fires* in production, but a live-DB investigation of the second pilot user
("Jonte") found it delivers a weak first impression for two reasons, and ruled out a third idea:
1. **Junk picks.** Selection was pure newest-first (`NewestUnsummarizedVideoIDs`,
`ORDER BY published_at DESC`) with **no quality signal**. Jonte's live burst-3 were a
stock-ticker **livestream** + two regional news clips — newest, not best. The cheap signals
that *could* gate this (duration, live status) are fetched by ADR-023's `videos.list`
enrichment at discovery and then **thrown away**: the `videos.duration_s` column (migration
001) was never written.
2. **Weakest model on the first impression.** All burst summaries ran on `koala/phi4-mini` — the
documented weak link (ADR-022 was born from its failures). The stronger, brain-validated
`iguana/gemma4-26b` was never used, even though the burst is only ~3 summaries.
3. **Cached-first instant summaries — REJECTED.** The idea: skip the fetch, summarize
already-cached transcripts (ADR-021) instantly. The pilot numbers kill it — only **11 videos**
overlap between the two users (~3% of each ~350400-video library), **0** cached-and-
unsummarized, and a new user's newest-20 unsummarized are **20/20 NOT cached**. Newest-first
and cached-first are structurally incompatible: fresh uploads are exactly what nobody has
fetched. An empty lever at pilot scale.
**Decision.**
1. **Persist `duration_s` at discovery.** `filterLowValue` (ADR-023) already has each candidate's
duration in hand; carry it onto the kept `domain.Video` and have `UpsertVideo` write it,
COALESCE-preserving a known value (the channel-title backfill stance, migration 014). No new
migration — the column exists. The connect-triggered discovery pass runs *before* the burst,
so a fresh user's candidates are enriched in time.
2. **Junk-avoiding selection.** A new `OnboardBurstVideoIDs(userID, limit, minSeconds, maxSeconds)`
keeps the newest-first order but drops a video when its duration is *known* and outside
`[minSeconds, maxSeconds]``minSeconds` = `TAPIR_MIN_VIDEO_SECONDS` (60, the Shorts floor),
`maxSeconds` = new `TAPIR_ONBOARD_MAX_VIDEO_SECONDS` (default 14400 = 4h, to drop multi-hour
livestream VODs that pass the live filter once ended). A NULL duration is **unknown** — kept
(degrade-open) but ranked after known-good rows. **has-captions stays un-gateable pre-fetch**
(only knowable after a gate fetch or a ~0-probability cache hit); selection only *avoids
known-junk*, it does not *promise* captions.
3. **Stronger model for the burst only.** `TAPIR_ONBOARD_SUMMARIZER_MODEL` (default
`iguana/gemma4-26b`) leads a burst-specific summarizer chain (onboard model first, then the
standard ADR-022 chain as resilience, deduped), wrapped in a burst-specific processor over the
*same* store/cache/sink — a pure wiring choice; the engine and ports are unchanged (ADR-003).
Empty or equal-to-primary collapses the burst back onto the shared processor.
**Not a throughput change.** The caption rate gate (ADR-014) and the foreground priority lane
(ADR-026) are untouched — same pacing, same cap. This changes *which* ≤3 videos the burst spends
its fetches on and *which model* summarizes them, never how fast or how many. The engine's
existing read-stored-first (ADR-021) is unchanged and still yields a free instant summary on the
rare cache hit — we simply do not *select* for cache hits.
**Consequences.** Better odds of a strong first session: the burst avoids the obvious junk and
runs the better model on the one impression that decides return. The selection improvement is
forward-looking — existing rows have NULL `duration_s` until their next discovery pass backfills
it (lazy, like channel_title); a brand-new user benefits immediately because connect-discovery
runs first. `duration_s` becoming live also unblocks future length-aware features (feed sorting,
"long read" badges) for free.
**Reversibility.** Pure config + wiring + one column write + one query, no migration.
`TAPIR_ONBOARD_MAX_VIDEO_SECONDS=0` (and `TAPIR_MIN_VIDEO_SECONDS=0`) restores pure newest-first;
`TAPIR_ONBOARD_SUMMARIZER_MODEL=""` collapses the burst back to the shared processor.
Spec: `docs/specs/onboarding-wow-burst.md`.
---
## Rejected alternatives
Approaches considered during the 2026-06-02 planning + grill session and **deliberately not
@@ -1083,6 +1151,7 @@ maps to the ADR that settles it.
| Feedback-based Stage 0 gate (friends saying it's useful) | Politeness bias makes asked-for feedback the least reliable signal; return-usage is the real test | ADR-016 |
| Reverse the Dex-write invite flow (Google OIDC only) | Some intended Future-B users won't use Google; OIDC-only leaves them with no onboarding path — invite flow is load-bearing | ADR-017 |
| k8s CronJob for scheduled discovery (vs in-process) | At Future-B scale the in-process scheduler is simpler to deploy; CronJob's failure-isolation benefit was weighed and traded away knowingly (revisit if >1 replica or load grows) | ADR-018 |
| Cached-transcript-first onboarding burst (instant, zero-fetch picks) | Live pilot DB: ~3% cross-user video overlap, 0 cached-and-unsummarized, a new user's newest-20 are 20/20 uncached — newest-first and cached-first are structurally incompatible. Empty lever at pilot scale | ADR-028 |
If a future case genuinely reopens one of these, that's a new ADR superseding the relevant one —
not a silent reversal.