Compare commits

...
9 Commits
Author SHA1 Message Date
mathiasandClaude Opus 4.8 36dd182fb5 feat(report): Stage-0 usage gate counts from a baseline date (default 2026-06-11)
CI / Build & Import (push) Successful in 11s
CI / Lint / Test / Vet (push) Successful in 10s
The return-usage gate (ADR-016) counted distinct active weeks over ALL history,
so pre-launch noise — testing churn and the period the pilot sat blocked on zero
summaries — would inflate the signal. Add a baseline: ActiveWeeks(ctx, since)
filters login_events + summary_actions to seen_at/acted_at >= since. The report
command sets it to TAPIR_USAGE_GATE_START (YYYY-MM-DD, default 2026-06-11 — the
morning the pilot was unblocked) and prints the baseline. The gate now measures
whether users RETURN once it genuinely works.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-11 22:52:15 +02:00
mathias 40808f2d4b docs(onboard): BDD scenarios + architecture for the ADR-028 burst
CI / Lint / Test / Vet (push) Successful in 10s
CI / Build & Import (push) Successful in 10s
- connect_account.feature: refine the burst scenario to "best recent" (likely-good
  selection, skips too-short/too-long) and add a scenario for the stronger model.
- scenario_coverage: remap to TestOnboardBurstVideoIDs + add the model scenario
  (the old NewestUnsummarizedVideoIDs test was removed).
- architecture.md: new "Connect-time onboarding burst" subsection — junk-avoiding
  selection over persisted duration, the burst-only stronger-model chain, and why
  has-captions/cached-first are not selection signals.
2026-06-11 19:54:29 +02:00
mathias 9232f49555 feat(onboard): burst picks likely-good videos, summarizes with stronger model (ADR-028)
Wire the onboarding burst to its quality-aware selection and a stronger model:
- main.go onboard uses OnboardBurstVideoIDs (junk-avoiding) instead of pure
  newest-first, bounded by MinVideoSeconds / OnboardMaxVideoSeconds.
- buildBurstProcessor builds a burst-only summarizer chain led by the onboard
  model (burstChainModels: onboard -> standard ADR-022 chain, deduped, NDA lever
  intact), over the same store/cache/sink. Collapses onto the shared Processor
  when the onboard model is empty/equal-to-primary or config is incomplete.
- summarizerEndpoint + newYouTubeSource extracted so standard and burst wiring
  share one definition.
- Remove now-superseded NewestUnsummarizedVideoIDs: OnboardBurstVideoIDs(.,0,0)
  is identical pure-newest behaviour and its test covers RLS + ordering.
2026-06-11 19:54:29 +02:00
mathias 582c1a2065 feat(store): OnboardBurstVideoIDs — junk-avoiding burst selection (ADR-028)
Newest-first but quality-aware: excludes a video when its duration is KNOWN and
outside [minSeconds, maxSeconds], dropping Shorts and multi-hour livestream VODs
that waste a scarce caption fetch on a poor first impression. NULL/unknown
duration is kept (degrade-open) but ranked after known-good rows. 0/0 bounds
disable the filter (pure newest-first, the reversibility lever). RLS-scoped.
NewestUnsummarizedVideoIDs is left in place for callers that want pure-newest.
2026-06-11 19:54:29 +02:00
mathias 9f0d8cf198 feat(discovery): persist video duration_s instead of discarding it (ADR-028)
ADR-023's filterLowValue already fetches each candidate's duration via the
cheap videos.list quota call to drop Shorts/live, then threw it away — the
videos.duration_s column (migration 001) was never written. Carry it onto the
kept domain.Video and have UpsertVideo persist it, COALESCE-preserving a known
value so an unknown (0) re-upsert never clobbers it (the channel_title backfill
stance, migration 014). This is the enabling change for length-aware burst
selection. No new migration — the column already exists.
2026-06-11 19:54:29 +02:00
mathias d21077303d feat(config): add OnboardSummarizerModel + OnboardMaxVideoSeconds (ADR-028)
Two knobs for the onboarding-burst quality work:
- TAPIR_ONBOARD_SUMMARIZER_MODEL (default iguana/gemma4-26b): the stronger model
  the burst leads its chain with; empty collapses the burst onto the shared
  processor (lookupOr, so explicit-empty disables).
- TAPIR_ONBOARD_MAX_VIDEO_SECONDS (default 14400/4h): upper duration bound for
  burst picks; 0 disables, negative clamps to 0.

Table-driven tests cover defaults, explicit, disable, and invalid input.
2026-06-11 19:54:29 +02:00
mathias 9b2ee2e765 docs: spec onboarding-wow-burst + ADR-028 (better picks, stronger model)
Phase-1 live-DB investigation found the connect-time burst (ADR-018) fires
but delivers a weak first impression: pure newest-first selection picks junk
(a livestream + regional news for pilot user Jonte), and all burst summaries
run on the weak koala/phi4-mini instead of the validated iguana/gemma4-26b.
Cached-first was investigated and rejected — ~3% cross-user overlap, 0
cached-and-unsummarized, newest-20 all uncached (newest-first and cached-first
are structurally incompatible).

ADR-028 records: persist duration_s at discovery (ADR-023 already fetches it,
just stops discarding), junk-avoiding burst selection (drop known too-short/
too-long), and a stronger burst-only summarizer chain. Not a throughput change.
2026-06-11 19:54:29 +02:00
mathias b8fbc5a805 docs: spec first-session wow — verify & improve onboarding summary burst (investigate-first)
CI / Lint / Test / Vet (push) Successful in 11s
CI / Build & Import (push) Successful in 10s
The concern is new-user first-contact: see some GOOD summaries fast or they
don't return. Reframed as curation/latency for ~3 videos, NOT a 429/throughput
problem (3 fetches is nowhere near the wall). Phase 1 (report-and-stop) verifies
whether the existing cap-3 onboarding burst even fires today, what it delivers,
and — critically — how much transcript-cache overlap exists between users (drives
the blend). Phase 2 levers: cached-transcript-first (instant, zero-fetch),
likely-good selection (has-captions/good-length, not just newest), and optionally
the stronger model for the burst's few summaries. Blend deferred to the
maintainer post-Phase-1. Explicitly NOT bulk-fetch, NOT credentials (ADR-010/026
dead end), NOT a client extension.
2026-06-11 15:36:09 +00:00
mathiasandClaude Opus 4.8 19ca4282a8 refactor(web): dock chat inline below the summary, integrated view (ADR-027)
CI / Lint / Test / Vet (push) Successful in 10s
CI / Build & Import (push) Successful in 10s
The chat was a separate page — opening it left the summary behind, the very
context you're asking about. Now the chat docks open IN PLACE below the summary:
"Dig deeper" reveals the chat section (HTMX, outerHTML over the closed dock) so
the summary stays on screen above it; no navigation. The no-JS fallback renders
the full summary AND the open chat on one page (the same integrated view), so
progressive enhancement holds.

Extracts a shared summaryBody templ so the detail page and the chat page render
one identical summary, not two divergent ones. GET /v/{id}/chat returns just the
open chat section as a fragment for the inline reveal, or the full summary+chat
page for a no-JS navigation; POST swaps the panel inline or re-renders the whole
page. Tests assert summary+chat coexist on the page and the reveal is a fragment.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-11 14:02:39 +02:00
26 changed files with 1547 additions and 730 deletions
+69
View File
@@ -1060,6 +1060,74 @@ no migration, no stored state to unwind. Spec: `docs/specs/chat-with-transcript.
---
## ADR-028 — Onboarding burst: pick likely-good videos, summarize them with a stronger model
**Status:** Accepted (2026-06-11). **Refines ADR-018** (the connect-time burst) and **ADR-020**
(recency-bounded auto-summarize). **Builds on ADR-022** (the endpoint chain), **ADR-023**
(the discovery-time `videos.list` enrichment), and **ADR-021** (the shared transcript cache).
Triggered by a Phase-1 investigation of the live pilot DB.
**Context.** A new user's first session decides whether they return (the Stage-0 gate, ADR-016).
The connect-time burst (ADR-018: summarize ≤`TAPIR_ONBOARD_SUMMARIZE_COUNT` newest videos so the
feed isn't empty) *fires* in production, but a live-DB investigation of the second pilot user
("Jonte") found it delivers a weak first impression for two reasons, and ruled out a third idea:
1. **Junk picks.** Selection was pure newest-first (`NewestUnsummarizedVideoIDs`,
`ORDER BY published_at DESC`) with **no quality signal**. Jonte's live burst-3 were a
stock-ticker **livestream** + two regional news clips — newest, not best. The cheap signals
that *could* gate this (duration, live status) are fetched by ADR-023's `videos.list`
enrichment at discovery and then **thrown away**: the `videos.duration_s` column (migration
001) was never written.
2. **Weakest model on the first impression.** All burst summaries ran on `koala/phi4-mini` — the
documented weak link (ADR-022 was born from its failures). The stronger, brain-validated
`iguana/gemma4-26b` was never used, even though the burst is only ~3 summaries.
3. **Cached-first instant summaries — REJECTED.** The idea: skip the fetch, summarize
already-cached transcripts (ADR-021) instantly. The pilot numbers kill it — only **11 videos**
overlap between the two users (~3% of each ~350400-video library), **0** cached-and-
unsummarized, and a new user's newest-20 unsummarized are **20/20 NOT cached**. Newest-first
and cached-first are structurally incompatible: fresh uploads are exactly what nobody has
fetched. An empty lever at pilot scale.
**Decision.**
1. **Persist `duration_s` at discovery.** `filterLowValue` (ADR-023) already has each candidate's
duration in hand; carry it onto the kept `domain.Video` and have `UpsertVideo` write it,
COALESCE-preserving a known value (the channel-title backfill stance, migration 014). No new
migration — the column exists. The connect-triggered discovery pass runs *before* the burst,
so a fresh user's candidates are enriched in time.
2. **Junk-avoiding selection.** A new `OnboardBurstVideoIDs(userID, limit, minSeconds, maxSeconds)`
keeps the newest-first order but drops a video when its duration is *known* and outside
`[minSeconds, maxSeconds]``minSeconds` = `TAPIR_MIN_VIDEO_SECONDS` (60, the Shorts floor),
`maxSeconds` = new `TAPIR_ONBOARD_MAX_VIDEO_SECONDS` (default 14400 = 4h, to drop multi-hour
livestream VODs that pass the live filter once ended). A NULL duration is **unknown** — kept
(degrade-open) but ranked after known-good rows. **has-captions stays un-gateable pre-fetch**
(only knowable after a gate fetch or a ~0-probability cache hit); selection only *avoids
known-junk*, it does not *promise* captions.
3. **Stronger model for the burst only.** `TAPIR_ONBOARD_SUMMARIZER_MODEL` (default
`iguana/gemma4-26b`) leads a burst-specific summarizer chain (onboard model first, then the
standard ADR-022 chain as resilience, deduped), wrapped in a burst-specific processor over the
*same* store/cache/sink — a pure wiring choice; the engine and ports are unchanged (ADR-003).
Empty or equal-to-primary collapses the burst back onto the shared processor.
**Not a throughput change.** The caption rate gate (ADR-014) and the foreground priority lane
(ADR-026) are untouched — same pacing, same cap. This changes *which* ≤3 videos the burst spends
its fetches on and *which model* summarizes them, never how fast or how many. The engine's
existing read-stored-first (ADR-021) is unchanged and still yields a free instant summary on the
rare cache hit — we simply do not *select* for cache hits.
**Consequences.** Better odds of a strong first session: the burst avoids the obvious junk and
runs the better model on the one impression that decides return. The selection improvement is
forward-looking — existing rows have NULL `duration_s` until their next discovery pass backfills
it (lazy, like channel_title); a brand-new user benefits immediately because connect-discovery
runs first. `duration_s` becoming live also unblocks future length-aware features (feed sorting,
"long read" badges) for free.
**Reversibility.** Pure config + wiring + one column write + one query, no migration.
`TAPIR_ONBOARD_MAX_VIDEO_SECONDS=0` (and `TAPIR_MIN_VIDEO_SECONDS=0`) restores pure newest-first;
`TAPIR_ONBOARD_SUMMARIZER_MODEL=""` collapses the burst back to the shared processor.
Spec: `docs/specs/onboarding-wow-burst.md`.
---
## Rejected alternatives
Approaches considered during the 2026-06-02 planning + grill session and **deliberately not
@@ -1083,6 +1151,7 @@ maps to the ADR that settles it.
| Feedback-based Stage 0 gate (friends saying it's useful) | Politeness bias makes asked-for feedback the least reliable signal; return-usage is the real test | ADR-016 |
| Reverse the Dex-write invite flow (Google OIDC only) | Some intended Future-B users won't use Google; OIDC-only leaves them with no onboarding path — invite flow is load-bearing | ADR-017 |
| k8s CronJob for scheduled discovery (vs in-process) | At Future-B scale the in-process scheduler is simpler to deploy; CronJob's failure-isolation benefit was weighed and traded away knowingly (revisit if >1 replica or load grows) | ADR-018 |
| Cached-transcript-first onboarding burst (instant, zero-fetch picks) | Live pilot DB: ~3% cross-user video overlap, 0 cached-and-unsummarized, a new user's newest-20 are 20/20 uncached — newest-first and cached-first are structurally incompatible. Empty lever at pilot scale | ADR-028 |
If a future case genuinely reopens one of these, that's a new ADR superseding the relevant one —
not a silent reversal.
+23 -10
View File
@@ -262,23 +262,36 @@ func cmdServe(ctx context.Context, log *slog.Logger) error {
// One lock shared by the scheduler and connect-triggered passes (#6) so
// they never fetch concurrently — the single-fetcher invariant (ADR-018).
runUser := serialize(&sync.Mutex{}, rawRunUser)
// Onboarding burst (Feature 1): after the connect-triggered discovery pass,
// summarize up to OnboardSummarizeCount of the user's NEWEST unsummarized
// videos so a fresh account gets real summaries in its first session. Hard
// cap; explicit, so it bypasses the recency window — but every fetch still
// goes through globalFetchGate via the Processor. No-op when disabled
// (count 0) or queue-only (no Processor).
// Onboarding burst (Feature 1, refined by ADR-028): after the connect-triggered
// discovery pass, summarize up to OnboardSummarizeCount of the user's newest
// LIKELY-GOOD unsummarized videos so a fresh account gets a strong first
// session. Selection avoids known-junk (Shorts/over-long/livestream VODs via
// the persisted duration); the burst leads its chain with the stronger onboard
// model. Hard cap; explicit, so it bypasses the recency window — but every
// fetch still goes through globalFetchGate. No-op when disabled (count 0) or
// queue-only (no processor).
//
// burstProcessor leads with the stronger model (ADR-028); it collapses onto the
// shared Processor when the onboard model is empty/equal-to-primary or the
// engine config is incomplete.
burstProcessor := app.Processor
if burstEngine, berr := buildBurstProcessor(cfg, st); berr != nil {
return berr
} else if burstEngine != nil {
burstProcessor = &engineProcessor{engine: burstEngine, store: st}
log.Info("onboarding burst uses a stronger model", "onboard_model", cfg.OnboardSummarizerModel)
}
onboard := func(ctx context.Context, userID string) {
if cfg.OnboardSummarizeCount <= 0 || app.Processor == nil {
if cfg.OnboardSummarizeCount <= 0 || burstProcessor == nil {
return
}
ids, err := st.NewestUnsummarizedVideoIDs(ctx, userID, cfg.OnboardSummarizeCount)
ids, err := st.OnboardBurstVideoIDs(ctx, userID, cfg.OnboardSummarizeCount, cfg.MinVideoSeconds, cfg.OnboardMaxVideoSeconds)
if err != nil {
log.Warn("onboarding: list newest unsummarized", "user", userID, "err", err)
log.Warn("onboarding: list burst candidates", "user", userID, "err", err)
return
}
for _, id := range ids {
if err := app.Processor.ProcessVideo(ctx, userID, id); err != nil {
if err := burstProcessor.ProcessVideo(ctx, userID, id); err != nil {
log.Warn("onboarding: summarize", "user", userID, "video", id, "err", err)
}
}
+83 -16
View File
@@ -52,13 +52,7 @@ func (f videoFetcher) FetchVideo(ctx context.Context, userID, videoID string) (d
// alias, not a second client config. Empty model entries are skipped, so a
// client deployment can set the cloud fallback empty to keep content local.
func buildSummarizer(cfg config.Config) *summarizer.Summarizer {
mk := func(model string) summarizer.Endpoint {
return summarizer.Endpoint{
Client: llm.New(cfg.GatewayURL, cfg.GatewayKey, model, cfg.SummarizerTimeout, llm.WithMaxTokens(cfg.SummaryMaxTokens)),
Provider: providerOf(model),
Model: model,
}
}
mk := summarizerEndpoint(cfg)
eps := []summarizer.Endpoint{mk(cfg.SummarizerModel)}
if cfg.FallbackModel != "" && cfg.FallbackModel != cfg.SummarizerModel {
eps = append(eps, mk(cfg.FallbackModel))
@@ -69,6 +63,55 @@ func buildSummarizer(cfg config.Config) *summarizer.Summarizer {
return summarizer.NewChain(eps, cfg.MaxTranscriptChars)
}
// summarizerEndpoint returns a constructor for a chain endpoint over the one
// LiteLLM gateway, varying only the model alias (the gateway fronts both
// llama-swap and berget). Shared by the standard and burst chains.
func summarizerEndpoint(cfg config.Config) func(model string) summarizer.Endpoint {
return func(model string) summarizer.Endpoint {
return summarizer.Endpoint{
Client: llm.New(cfg.GatewayURL, cfg.GatewayKey, model, cfg.SummarizerTimeout, llm.WithMaxTokens(cfg.SummaryMaxTokens)),
Provider: providerOf(model),
Model: model,
}
}
}
// burstChainModels is the ordered, deduped model list for the onboarding burst
// (ADR-028): the stronger onboard model leads, then the standard ADR-022 chain
// (primary → local fallback → cloud) follows as resilience. Empty entries are
// dropped and duplicates collapsed, so the NDA lever (empty cloud fallback) keeps
// the burst chain fully local exactly as the standard chain does.
func burstChainModels(cfg config.Config) []string {
var models []string
add := func(m string) {
if m == "" {
return
}
for _, e := range models {
if e == m {
return
}
}
models = append(models, m)
}
add(cfg.OnboardSummarizerModel)
add(cfg.SummarizerModel)
add(cfg.FallbackModel)
add(cfg.CloudFallbackModel)
return models
}
// buildBurstSummarizer builds the onboarding-burst summarizer chain (ADR-028):
// the onboard model first, then the standard chain as fallback, deduped.
func buildBurstSummarizer(cfg config.Config) *summarizer.Summarizer {
mk := summarizerEndpoint(cfg)
var eps []summarizer.Endpoint
for _, m := range burstChainModels(cfg) {
eps = append(eps, mk(m))
}
return summarizer.NewChain(eps, cfg.MaxTranscriptChars)
}
// chatModels is the ordered, local-first set of models offered in the chat
// switcher (ADR-027), reusing the ADR-022 chain: primary → local fallback →
// cloud. Empty entries are dropped and duplicates collapsed, so a client/NDA
@@ -126,15 +169,7 @@ func buildProcessor(cfg config.Config, st *store.Store) (*usecase.Engine, error)
return nil, nil
}
secretStore := secrets.NewFileStore(cfg.SecretsFile)
src := youtube.New(youtube.Config{
ClientID: cfg.YTClientID,
ClientSecret: cfg.YTClientSecret,
TokenSecretRef: cfg.YTTokenRef,
PreferredLanguages: []string{"en"},
MinVideoSeconds: cfg.MinVideoSeconds,
}, secretStore)
src := newYouTubeSource(cfg, secrets.NewFileStore(cfg.SecretsFile))
sum := buildSummarizer(cfg)
// The store is both the summary sink and the shared transcript cache (ADR-021):
@@ -145,6 +180,38 @@ func buildProcessor(cfg config.Config, st *store.Store) (*usecase.Engine, error)
return eng, nil
}
// newYouTubeSource builds the captions-first VideoSource shared by the standard
// and burst processors — same per-process YouTube credentials and ADR-023 Shorts
// filter; only the summarizer chain differs between them.
func newYouTubeSource(cfg config.Config, secretStore ports.SecretStore) ports.VideoSource {
return youtube.New(youtube.Config{
ClientID: cfg.YTClientID,
ClientSecret: cfg.YTClientSecret,
TokenSecretRef: cfg.YTTokenRef,
PreferredLanguages: []string{"en"},
MinVideoSeconds: cfg.MinVideoSeconds,
}, secretStore)
}
// buildBurstProcessor wires a processor whose summarizer leads with the stronger
// onboard model (ADR-028), used only by the connect-time burst over the SAME
// store / transcript cache / sink — a wiring choice; the engine and ports are
// unchanged. Returns (nil, nil) — the collapse lever — when the onboard model is
// empty or equal to the primary (the burst then reuses the shared processor), or
// when the engine config is incomplete (queue-only, same as buildProcessor).
func buildBurstProcessor(cfg config.Config, st *store.Store) (*usecase.Engine, error) {
if cfg.OnboardSummarizerModel == "" || cfg.OnboardSummarizerModel == cfg.SummarizerModel {
return nil, nil
}
if cfg.GatewayURL == "" || cfg.YTClientID == "" || cfg.YTClientSecret == "" || cfg.SecretsFile == "" {
return nil, nil
}
src := newYouTubeSource(cfg, secrets.NewFileStore(cfg.SecretsFile))
eng := usecase.NewEngine(src, buildBurstSummarizer(cfg), st)
eng.Transcripts = st
return eng, nil
}
// engineProcessor adapts the engine (which works in terms of a domain.Video) to
// the web.Processor port (which works in terms of a stored video id): it loads the
// video row, runs the engine, and — on a produced summary — clears the manual
+62
View File
@@ -45,3 +45,65 @@ func TestBuildProcessorNilOnIncompleteConfig(t *testing.T) {
})
}
}
// TestBurstChainModelsLeadsWithOnboardModel: the onboarding burst chain (ADR-028)
// leads with the stronger onboard model, then falls back through the standard
// ADR-022 chain (primary -> local fallback -> cloud), deduped.
func TestBurstChainModelsLeadsWithOnboardModel(t *testing.T) {
got := burstChainModels(config.Config{
OnboardSummarizerModel: "iguana/gemma4-26b",
SummarizerModel: "koala/phi4-mini",
FallbackModel: "iguana/gemma4-26b", // also the onboard model -> dedup
CloudFallbackModel: "berget/mistral-small",
})
want := []string{"iguana/gemma4-26b", "koala/phi4-mini", "berget/mistral-small"}
if len(got) != len(want) {
t.Fatalf("burstChainModels = %v, want %v", got, want)
}
for i := range want {
if got[i] != want[i] {
t.Fatalf("burstChainModels = %v, want %v", got, want)
}
}
}
// TestBurstChainModelsCloudAbsentWhenDisabled: the NDA lever holds for the burst
// too — empty cloud fallback keeps the burst chain fully local.
func TestBurstChainModelsCloudAbsentWhenDisabled(t *testing.T) {
got := burstChainModels(config.Config{
OnboardSummarizerModel: "iguana/gemma4-26b",
SummarizerModel: "koala/phi4-mini",
CloudFallbackModel: "",
})
for _, m := range got {
if m == "" || m == "berget/mistral-small" {
t.Fatalf("cloud model leaked into burst chain: %v", got)
}
}
}
// TestBuildBurstProcessorNilWhenCollapsed: an empty or primary-equal onboard model
// collapses the burst onto the shared processor (buildBurstProcessor returns nil).
func TestBuildBurstProcessorNilWhenCollapsed(t *testing.T) {
base := config.Config{
GatewayURL: "http://gw/v1",
YTClientID: "id",
YTClientSecret: "secret",
SecretsFile: "/tmp/secrets.json",
SummarizerModel: "koala/phi4-mini",
}
t.Run("empty onboard model", func(t *testing.T) {
base.OnboardSummarizerModel = ""
eng, err := buildBurstProcessor(base, nil)
if err != nil || eng != nil {
t.Fatalf("buildBurstProcessor = (%v, %v), want (nil, nil)", eng, err)
}
})
t.Run("onboard model equals primary", func(t *testing.T) {
base.OnboardSummarizerModel = "koala/phi4-mini"
eng, err := buildBurstProcessor(base, nil)
if err != nil || eng != nil {
t.Fatalf("buildBurstProcessor = (%v, %v), want (nil, nil)", eng, err)
}
})
}
+33 -3
View File
@@ -6,6 +6,7 @@ import (
"io"
"os"
"text/tabwriter"
"time"
"gitea.d-ma.be/mathias/tapir/internal/adapters/store"
)
@@ -14,6 +15,27 @@ import (
// weeks. The gate passes when any user reaches it.
const gateThreshold = 2
// defaultGateStart is the date Stage-0 return-usage tracking begins: the morning
// the pilot was actually unblocked and summaries started flowing (2026-06-11).
// Activity before this — testing, the period the pilot was stuck on zero — is
// noise and must not count toward the gate. Override with TAPIR_USAGE_GATE_START
// (YYYY-MM-DD). The gate measures whether users RETURN once it genuinely works.
const defaultGateStart = "2026-06-11"
// gateStart resolves the baseline date from TAPIR_USAGE_GATE_START or the default,
// parsed as a UTC calendar day.
func gateStart() (time.Time, error) {
v := os.Getenv("TAPIR_USAGE_GATE_START")
if v == "" {
v = defaultGateStart
}
t, err := time.Parse("2006-01-02", v)
if err != nil {
return time.Time{}, fmt.Errorf("TAPIR_USAGE_GATE_START=%q: want YYYY-MM-DD: %w", v, err)
}
return t, nil
}
// runReport prints the Stage-0 usage gate: per-user distinct active weeks (reads
// UNION acts) and the pass/fail verdict. Read-only, cross-user — needs only
// TAPIR_DB_DSN (not TAPIR_USER_ID; the report enumerates all users itself).
@@ -28,16 +50,24 @@ func runReport(ctx context.Context, _ []string) error {
}
defer s.Close()
rows, err := s.ActiveWeeks(ctx)
since, err := gateStart()
if err != nil {
return err
}
return formatReport(os.Stdout, rows)
rows, err := s.ActiveWeeks(ctx, since)
if err != nil {
return err
}
return formatReport(os.Stdout, rows, since)
}
// formatReport renders the per-user week counts and the gate verdict. Pure: no DB,
// no env — so the layout and verdict logic are unit-testable without Postgres.
func formatReport(w io.Writer, rows []store.UserActiveWeeks) error {
func formatReport(w io.Writer, rows []store.UserActiveWeeks, since time.Time) error {
if _, err := fmt.Fprintf(w, "Counting usage since %s (Stage-0 gate baseline)\n\n", since.Format("2006-01-02")); err != nil {
return err
}
if len(rows) == 0 {
_, err := fmt.Fprintln(w, "no users yet")
return err
+7 -3
View File
@@ -3,12 +3,15 @@ package main
import (
"strings"
"testing"
"time"
"github.com/stretchr/testify/require"
"gitea.d-ma.be/mathias/tapir/internal/adapters/store"
)
var testSince = time.Date(2026, 6, 11, 0, 0, 0, 0, time.UTC)
func TestFormatReportColumnsAndGatePass(t *testing.T) {
rows := []store.UserActiveWeeks{
{UserID: "user-a", DisplayName: "Ada", ActiveWeeks: 3},
@@ -16,9 +19,10 @@ func TestFormatReportColumnsAndGatePass(t *testing.T) {
}
var b strings.Builder
require.NoError(t, formatReport(&b, rows))
require.NoError(t, formatReport(&b, rows, testSince))
out := b.String()
require.Contains(t, out, "since 2026-06-11", "report states the gate baseline date")
require.Contains(t, out, "USER")
require.Contains(t, out, "ACTIVE_WEEKS")
require.Contains(t, out, "Ada")
@@ -34,12 +38,12 @@ func TestFormatReportGateNotMet(t *testing.T) {
rows := []store.UserActiveWeeks{{UserID: "user-a", ActiveWeeks: 1}}
var b strings.Builder
require.NoError(t, formatReport(&b, rows))
require.NoError(t, formatReport(&b, rows, testSince))
require.Contains(t, b.String(), "NOT YET MET", "no user at >= 2 weeks fails the gate")
}
func TestFormatReportEmpty(t *testing.T) {
var b strings.Builder
require.NoError(t, formatReport(&b, nil))
require.NoError(t, formatReport(&b, nil, testSince))
require.Contains(t, b.String(), "no users yet")
}
+26
View File
@@ -250,6 +250,32 @@ After (newest-first): `[chanB-new, chanA-mid, chanA-old, chanB-null]`
The set of *processed* videos now also excludes auto-mode back-catalogue beyond the recency
window (those stay listed, summarised on demand); within the processed set, only order changes.
### Connect-time onboarding burst (ADR-018 → ADR-028)
On a successful YouTube connect, `ConnectHandler` enqueues a connect-triggered discovery pass;
the `discoveryTrigger` runs that pass and then fires the **onboarding burst** — a third entry path
that summarises up to `TAPIR_ONBOARD_SUMMARIZE_COUNT` (default 3, hard-capped) of the new user's
videos so the first session is not empty. The burst still flows through `globalFetchGate` (it is
not a throughput change); ADR-028 sharpened *which* videos and *which model*:
- **Selection** is `OnboardBurstVideoIDs`, not pure newest-first. It keeps newest-first order but
excludes a video whose **known** duration is outside `[TAPIR_MIN_VIDEO_SECONDS,
TAPIR_ONBOARD_MAX_VIDEO_SECONDS]` (drops Shorts and multi-hour livestream VODs). An unknown
(NULL) duration is degrade-open — kept, but ranked after known-good rows. The connect-triggered
discovery pass runs *before* the burst, and ADR-023's `videos.list` enrichment now **persists**
`duration_s` (instead of discarding it after the Shorts filter), so a fresh user's candidates
carry a duration in time for selection.
- **Model**: the burst runs through a dedicated summarizer chain led by
`TAPIR_ONBOARD_SUMMARIZER_MODEL` (default `iguana/gemma4-26b`, the stronger local model), with
the standard ADR-022 chain following as fallback. This is a wiring choice — a second
`engineProcessor` over the same store / transcript cache / sink; the engine and ports are
unchanged. Empty / equal-to-primary collapses it back onto the shared processor.
`has-captions` is deliberately **not** a selection signal — it is only knowable after a gate fetch
(or a ~0-probability cache hit at pilot scale), so the burst can avoid known-junk but cannot
promise captions. Cached-transcript-first selection was investigated and rejected (ADR-028:
~3% cross-user overlap).
---
## Sequence — core use case: new video summarized
+3
View File
@@ -224,6 +224,9 @@ knobs plus one load-bearing deployment constraint:
- `TAPIR_DISCOVERY_INTERVAL` — Go duration, e.g. `2h`. The cadence the serve process runs a
discovery pass for every registered user (run-once-on-startup, then every interval).
**Unset or `0` = disabled** (dev/tests never auto-fetch).
- `TAPIR_USAGE_GATE_START``YYYY-MM-DD`, default **`2026-06-11`** (the morning the pilot was
unblocked and summaries started flowing). `tapir report` counts return-usage (distinct active
weeks, ADR-016) only from this date, so pre-launch testing and the blocked period are excluded.
- `TAPIR_FETCH_RATE` — Go duration, default `2s`. The **process-wide per-egress-IP caption-fetch
rate gate** (ADR-014 item 2). Every caption fetch — scheduler runners *and* the web "Summarize"
click-path — serialises through this one limiter so the pod cannot collectively trip 429s. `0`
+128
View File
@@ -0,0 +1,128 @@
# Spec — Onboarding "wow" burst: better picks, stronger model
**Repo:** tapir · **Size:** medium · **Solo session** (not a swarm).
> **Status: built (v0.25.0, ADR-028).** This supersedes the original investigate-first brief
> (committed as the prior version of this file): Phase 1 was run against the live pilot DB and its
> findings are folded into "Why this exists" below; Phase 2 was built as described here. The one
> brief lever NOT built — the honest "the rest fill in over the coming days" framing copy — is
> listed under *Explicitly NOT in this slice*.
**Why this exists.** A new user's first session decides whether they return (the Stage-0 gate,
VISION.md). On connect, Tapir fires a capped burst (≤`TAPIR_ONBOARD_SUMMARIZE_COUNT`, default 3)
that summarizes the user's newest unsummarized videos so the feed isn't empty (the burst itself
works — wired in `cmd/tapir/discovery.go``cmd/tapir/main.go` `onboard`). A Phase-1
investigation of the live pilot DB found the burst *fires* but delivers a **weak first
impression** for two concrete reasons, and ruled out a third idea:
1. **Picks are junk.** Selection is pure newest-first (`videos.NewestUnsummarizedVideoIDs`,
`ORDER BY published_at DESC`) with **zero quality signal**. Pilot user "Jonte"'s live burst-3
were a stock-ticker **livestream** + two regional news clips — the newest, not the best.
2. **Weakest model on the first impression.** All of Jonte's summaries ran on
`koala/phi4-mini` (the documented weak link — ADR-022 was born from its failures). The
stronger, brain-validated `iguana/gemma4-26b` was never used for the burst.
3. **Cached-first is empty at pilot scale — REJECTED.** The idea (summarize already-cached
transcripts instantly, zero fetch) dies on the numbers: only **11 videos** overlap between the
two pilot users (~3% of each library), **0** cached-and-unsummarized, and a new user's
newest-20 unsummarized are **20/20 NOT cached** — newest-first and cached-first are
structurally incompatible (fresh uploads are exactly what nobody has fetched yet). Not built.
This is a **curation/latency problem for ~3 videos, NOT a throughput/429 problem** — fetching 3
captions is nowhere near the rate limit. Nothing here fetches harder or pressures the rate gate;
it picks the right few videos and runs a better model on them.
Read `CLAUDE.md`, `DECISIONS.md` (esp. ADR-014, ADR-018, ADR-020, ADR-021, ADR-022, ADR-023,
and the new **ADR-028**), and `VISION.md` (the Stage-0 gate) first. TBD — commit directly to
`main`, one logical change per commit, conventional commits, `task check` green before each
commit, `templ generate` if any view changes (none expected).
## Decisions already made (do not reopen)
- **Not a throughput change.** The caption rate gate (ADR-014) is untouched — same pacing, same
priority lane (ADR-026). This slice changes *which* ≤3 videos the burst spends its fetches on
and *which model* summarizes them, never how fast or how many.
- **Cached-first is dropped** (ADR-028, the 3% overlap). The engine's existing read-stored-first
(ADR-021, `resolveTranscript`) stays — it already gives a free instant summary on the rare
cache hit, transparently. We do not *select* for cache hits.
- **has-captions is not a pre-fetch signal.** It is only knowable after a gate fetch (or a cache
hit, ~0 for new videos). Selection can only *avoid known-junk* (Shorts/live/over-long) — it
cannot *guarantee* captions. The spec is honest about this: better odds, not a promise.
- **No credentialed caption fetch** (ADR-010/ADR-026 dead end). **No client extension.**
## 1. Persist `duration_s` at discovery (the enabling change)
The `videos.duration_s` column exists (migration 001) but is **never written** — ADR-023's
`filterLowValue` (`internal/adapters/youtube/youtube.go`) already fetches each candidate's
duration via the cheap quota `videos.list` call, uses it to drop Shorts/live, then **discards
it**. Stop discarding:
- Add `DurationSeconds int` to `domain.Video`.
- In `filterLowValue`, set `DurationSeconds` on each kept video from the `videos.list` `meta`.
- `UpsertVideo` writes `duration_s`, **COALESCE-preserving** a known value (never overwrite a
real duration with 0/unknown), mirroring the `channel_title` backfill stance (migration 014).
- No new migration — the column is already there.
Consequence: a fresh user's connect-triggered discovery pass runs **before** the onboard burst
(`Enqueue`: `run()` then `onboard()`), so duration is populated for the burst's candidates at
connect. Existing rows backfill on their next discovery pass; until then their `duration_s` is
NULL and treated as "unknown" (§2).
## 2. Junk-avoiding burst selection
New store method, RLS-scoped via `withUser`:
```
OnboardBurstVideoIDs(ctx, userID string, limit, minSeconds, maxSeconds int) ([]string, error)
```
- Same base as the old `NewestUnsummarizedVideoIDs`: the user's videos with no summary yet,
`ORDER BY published_at DESC NULLS LAST, seen_at DESC`, `LIMIT limit`.
- **Exclude known-junk**: a row is dropped only when `duration_s IS NOT NULL` **and**
(`duration_s < minSeconds` OR `duration_s > maxSeconds`). A NULL duration is **unknown** — kept
(degrade-open: never starve the burst because metadata is missing), but ordered *after* rows
with a known-good duration so a freshly-enriched good pick wins when both exist.
- `minSeconds` reuses `TAPIR_MIN_VIDEO_SECONDS` (default 60 — the Shorts floor, ADR-023).
`maxSeconds` is new: `TAPIR_ONBOARD_MAX_VIDEO_SECONDS` (default 14400 = 4h) — drops the
multi-hour livestream VODs that pass the live filter once ended.
- `minSeconds<=0` and `maxSeconds<=0` each disable that bound (so `0/0` == the old
newest-first behaviour, the reversibility lever).
- The burst switches to this method; `NewestUnsummarizedVideoIDs` is removed (fully superseded —
`OnboardBurstVideoIDs(., 0, 0)` is identical pure-newest behaviour).
## 3. Stronger model for the burst
The burst summarizes only ≤3 videos, so a slower, stronger model is affordable exactly here.
- New config `TAPIR_ONBOARD_SUMMARIZER_MODEL` (default `iguana/gemma4-26b` — the brain-validated
homelab general-purpose model, already the ADR-022 fallback).
- Build a **burst-specific summarizer chain** that puts the onboard model **first**, then the
standard chain (primary → local fallback → cloud) as resilience, deduped. Wrap it in a
burst-specific `engineProcessor` reusing the same store/transcript-cache/sink — a pure wiring
choice, engine and ports unchanged (Clean Architecture, ADR-003).
- The `onboard` closure uses the burst processor instead of `app.Processor`.
- **Collapse cleanly**: when `OnboardSummarizerModel` is empty or equals `SummarizerModel`, the
onboard path reuses `app.Processor` (no separate chain) — the reversibility lever.
- Local-first preserved: the onboard model is a local alias; the cloud endpoint stays last in the
chain, so a client/NDA deployment with `TAPIR_CLOUD_FALLBACK_MODEL=""` keeps burst content
local too.
## 4. Behaviour spec + docs
- Add scenarios to `docs/use-cases/connect_account.feature` (the connect → burst flow): burst
skips a too-long/live video in favour of a reasonable-length one; burst summarizes with the
stronger model first. Map them in `scenarioCoverage` so `TestScenarioCoverage` stays green.
- Update `docs/architecture/architecture.md` (the onboarding-burst section) to describe the
junk-avoiding selection + the burst model override.
- ADR-028 in `DECISIONS.md` records the rationale (incl. the rejected cached-first lever).
## Success criteria
- `task check` green (fmt, vet, lint, `go test -p 1 ./...`).
- A unit test proves `OnboardBurstVideoIDs` drops a known too-long / sub-min video and keeps a
good one, newest-first, RLS-scoped, unsummarized-only.
- A test proves discovery persists `duration_s` and does not clobber it on re-upsert.
- A test proves the burst chain leads with the onboard model (then the standard chain).
- Config defaults + bounds tested (`OnboardMaxVideoSeconds`, `OnboardSummarizerModel`).
- No change to the rate gate, fetch pacing, or burst cap. `0/0` + empty model == prior behaviour.
## Explicitly NOT in this slice
- Cached-first selection (rejected, ADR-028).
- Any caption-availability *guarantee* (impossible pre-fetch).
- **Honest "taster" framing copy** ("summaries of a few of your videos to get you started — the
rest fill in over the coming days"). A good lever from the original brief, but it's a UI/copy
change with no backend dependency; deferred to a UI pass, tracked as an issue.
- Backfilling `duration_s` for existing rows via a migration (it backfills lazily on discovery).
- Return-nudges / digests (ADR-020: poisons the unprompted-return signal).
- Raising fetch throughput, multi-IP, or Whisper (out of scope; the gate is deliberate).
+6
View File
@@ -14,6 +14,12 @@ Feature: Chat with a video's stored transcript
Scenario: A summary view offers a deeper-dive into the video
When I view the summary
Then I see a "dig deeper" affordance that opens a chat about this video
And it opens the chat in place, below the summary, without leaving the page
Scenario: The summary and the chat are on one page
When I open the chat
Then the summary stays visible alongside the chat
And I can read the summary while I ask questions
Scenario: Ask a question answered from the stored transcript
When I ask a question in the chat
+9 -2
View File
@@ -24,12 +24,19 @@ Feature: Connect and manage video accounts
Then a discovery pass for my account is triggered right away
And I do not have to wait for the next scheduled pass to see my videos
Scenario: Connecting summarizes my newest videos right away
Scenario: Connecting summarizes my best recent videos right away
Given I have no connected video accounts
When I connect my YouTube account
Then up to the onboarding cap of my newest videos are summarized through the rate gate
Then up to the onboarding cap of my newest likely-good videos are summarized through the rate gate
And videos whose known duration is too short or too long are skipped
And the rest are left to the scheduled recency-bounded pass
Scenario: The onboarding burst summarizes with a stronger model
Given I have no connected video accounts
When I connect my YouTube account
Then the burst summarizes with the stronger onboarding model first
And the standard summarizer chain still follows as a fallback
Scenario: Tokens are never stored in the clear
When I connect any video account
Then no OAuth token value is stored in the database
+12 -6
View File
@@ -4,6 +4,7 @@ import (
"context"
"fmt"
"sort"
"time"
"github.com/jackc/pgx/v5"
)
@@ -32,7 +33,12 @@ type UserActiveWeeks struct {
// Scope note: the enumeration covers users with a Dex identity (the web users the
// gate is about). A CLI-only user created by the store sink without an identity
// row would not appear — out of scope for this gate.
func (s *Store) ActiveWeeks(ctx context.Context) ([]UserActiveWeeks, error) {
// ActiveWeeks counts each user's distinct active weeks from `since` onward. A zero
// `since` means no lower bound (count all history). The Stage-0 gate baseline is
// set by the caller (the report command) to the date real usage tracking began,
// so pre-launch noise — testing, the period the pilot was blocked — does not count
// toward the return-usage signal (ADR-016).
func (s *Store) ActiveWeeks(ctx context.Context, since time.Time) ([]UserActiveWeeks, error) {
userIDs, err := s.identityUserIDs(ctx)
if err != nil {
return nil, err
@@ -40,7 +46,7 @@ func (s *Store) ActiveWeeks(ctx context.Context) ([]UserActiveWeeks, error) {
out := make([]UserActiveWeeks, 0, len(userIDs))
for _, uid := range userIDs {
row, err := s.activeWeeksFor(ctx, uid)
row, err := s.activeWeeksFor(ctx, uid, since)
if err != nil {
return nil, err
}
@@ -85,18 +91,18 @@ func (s *Store) identityUserIDs(ctx context.Context) ([]string, error) {
// activeWeeksFor counts one user's distinct active weeks (reads UNION acts) and
// reads their display name, RLS-scoped via withUser. The UNION dedups a week that
// has both a login and an action so it counts once.
func (s *Store) activeWeeksFor(ctx context.Context, userID string) (UserActiveWeeks, error) {
func (s *Store) activeWeeksFor(ctx context.Context, userID string, since time.Time) (UserActiveWeeks, error) {
res := UserActiveWeeks{UserID: userID}
if err := s.withUser(ctx, userID, func(tx pgx.Tx) error {
if err := tx.QueryRow(ctx,
`WITH weeks AS (
SELECT date_trunc('week', seen_at) AS wk
FROM login_events WHERE user_id = $1
FROM login_events WHERE user_id = $1 AND seen_at >= $2
UNION
SELECT date_trunc('week', acted_at)
FROM summary_actions WHERE user_id = $1
FROM summary_actions WHERE user_id = $1 AND acted_at >= $2
)
SELECT count(DISTINCT wk) FROM weeks`, userID).Scan(&res.ActiveWeeks); err != nil {
SELECT count(DISTINCT wk) FROM weeks`, userID, since).Scan(&res.ActiveWeeks); err != nil {
return fmt.Errorf("store: count active weeks: %w", err)
}
if err := tx.QueryRow(ctx,
+29 -2
View File
@@ -3,6 +3,7 @@ package store_test
import (
"context"
"testing"
"time"
"github.com/jackc/pgx/v5/pgxpool"
"github.com/stretchr/testify/require"
@@ -54,7 +55,7 @@ func TestActiveWeeksCountsDistinctWeeksAcrossReadsAndActs(t *testing.T) {
($1, 'vid-2', 'saved', '2026-01-19T18:00:00Z')`, userA)
require.NoError(t, err)
got, err := s.ActiveWeeks(ctx)
got, err := s.ActiveWeeks(ctx, time.Time{}) // zero since = no lower bound
require.NoError(t, err)
require.Len(t, got, 2, "both identity users must appear")
@@ -72,7 +73,33 @@ func TestActiveWeeksEmptyWhenNoUsers(t *testing.T) {
s := newStore(t)
resetDB(t, rawPool(t))
got, err := s.ActiveWeeks(ctx)
got, err := s.ActiveWeeks(ctx, time.Time{})
require.NoError(t, err)
require.Empty(t, got)
}
// TestActiveWeeksExcludesBeforeGateStart proves the baseline cutoff: activity
// before `since` does not count, so pre-launch noise (testing, the pilot's blocked
// period) is excluded from the Stage-0 return-usage gate (ADR-016).
func TestActiveWeeksExcludesBeforeGateStart(t *testing.T) {
ctx := context.Background()
s := newStore(t)
p := rawPool(t)
resetDB(t, p)
seedReportUser(t, p, userA, "subject-a", "Ada")
// One read well before the baseline, two reads in distinct weeks after it.
_, err := p.Exec(ctx,
`INSERT INTO login_events (user_id, seen_at) VALUES
($1, '2026-05-01T09:00:00Z'),
($1, '2026-06-12T09:00:00Z'),
($1, '2026-06-19T09:00:00Z')`, userA)
require.NoError(t, err)
since := time.Date(2026, 6, 11, 0, 0, 0, 0, time.UTC)
got, err := s.ActiveWeeks(ctx, since)
require.NoError(t, err)
require.Len(t, got, 1)
require.Equal(t, 2, got[0].ActiveWeeks, "only the two post-baseline weeks count; the May read is excluded")
}
+35 -14
View File
@@ -46,15 +46,16 @@ func (s *Store) UpsertVideo(ctx context.Context, v domain.Video) (string, error)
}
if err := tx.QueryRow(ctx,
`INSERT INTO videos (user_id, provider, provider_video_id, title, url, published_at, channel_title)
VALUES ($1, $2, $3, $4, $5, $6, $7)
`INSERT INTO videos (user_id, provider, provider_video_id, title, url, published_at, channel_title, duration_s)
VALUES ($1, $2, $3, $4, $5, $6, $7, $8)
ON CONFLICT (user_id, provider, provider_video_id) DO UPDATE SET
title = EXCLUDED.title,
url = EXCLUDED.url,
published_at = EXCLUDED.published_at,
channel_title = COALESCE(NULLIF(EXCLUDED.channel_title, ''), videos.channel_title)
channel_title = COALESCE(NULLIF(EXCLUDED.channel_title, ''), videos.channel_title),
duration_s = COALESCE(EXCLUDED.duration_s, videos.duration_s)
RETURNING id`,
v.UserID, provider, v.ProviderVideoID, v.Title, v.URL, nullTime(v.PublishedAt), v.ChannelTitle,
v.UserID, provider, v.ProviderVideoID, v.Title, v.URL, nullTime(v.PublishedAt), v.ChannelTitle, nullDuration(v.DurationSeconds),
).Scan(&id); err != nil {
return fmt.Errorf("store: upsert video: %w", err)
}
@@ -74,12 +75,27 @@ func nullTime(t time.Time) *time.Time {
return &t
}
// NewestUnsummarizedVideoIDs returns up to limit of the user's videos that have
// no summary yet, newest first (published_at DESC, NULLS LAST). It caps the
// connect-time onboarding burst (Feature 1) at a fixed count: the caller marks
// these for summarization through the shared rate gate. RLS-scoped via withUser,
// so it only ever sees the requesting user's rows. limit <= 0 returns nil.
func (s *Store) NewestUnsummarizedVideoIDs(ctx context.Context, userID string, limit int) ([]string, error) {
// nullDuration maps an unknown duration (0) to SQL NULL so the upsert's
// COALESCE(EXCLUDED.duration_s, videos.duration_s) preserves a previously-known
// value instead of clobbering it with 0 (ADR-028; the channel_title backfill
// stance, migration 014).
func nullDuration(seconds int) *int {
if seconds <= 0 {
return nil
}
return &seconds
}
// OnboardBurstVideoIDs returns up to limit of the user's unsummarized videos for
// the connect-time onboarding burst (ADR-028), newest-first but quality-aware: a
// video is excluded when its duration is KNOWN and outside [minSeconds, maxSeconds]
// — dropping Shorts (below min) and multi-hour livestream VODs (above max) that
// would waste a scarce caption fetch on a poor first impression. A NULL/unknown
// duration is kept (degrade-open) but ranked AFTER known-good rows, so a freshly
// enriched good pick wins when both exist. minSeconds<=0 / maxSeconds<=0 each
// disable that bound (0/0 == pure newest-first, the reversibility lever).
// RLS-scoped via withUser; limit <= 0 returns nil.
func (s *Store) OnboardBurstVideoIDs(ctx context.Context, userID string, limit, minSeconds, maxSeconds int) ([]string, error) {
if limit <= 0 {
return nil, nil
}
@@ -92,16 +108,21 @@ func (s *Store) NewestUnsummarizedVideoIDs(ctx context.Context, userID string, l
AND NOT EXISTS (
SELECT 1 FROM summaries su
WHERE su.user_id = v.user_id AND su.video_id = v.id)
ORDER BY v.published_at DESC NULLS LAST, v.seen_at DESC
LIMIT $2`, userID, limit)
AND NOT (
v.duration_s IS NOT NULL
AND ( ($3 > 0 AND v.duration_s < $3)
OR ($4 > 0 AND v.duration_s > $4) ))
ORDER BY (v.duration_s IS NOT NULL) DESC,
v.published_at DESC NULLS LAST, v.seen_at DESC
LIMIT $2`, userID, limit, minSeconds, maxSeconds)
if err != nil {
return fmt.Errorf("store: newest unsummarized: %w", err)
return fmt.Errorf("store: onboard burst videos: %w", err)
}
defer rows.Close()
for rows.Next() {
var id string
if err := rows.Scan(&id); err != nil {
return fmt.Errorf("store: scan newest unsummarized: %w", err)
return fmt.Errorf("store: scan onboard burst video: %w", err)
}
ids = append(ids, id)
}
+63 -14
View File
@@ -52,6 +52,36 @@ func TestUpsertVideo_ReturnsStableID(t *testing.T) {
require.Equal(t, 1, count, "must not duplicate the row")
}
func TestUpsertVideo_PersistsAndPreservesDuration(t *testing.T) {
ctx := context.Background()
s := newStore(t)
resetDB(t, rawPool(t))
// First upsert carries a known duration (ADR-028: discovery enriches it).
v := ytVideo(userA, "dur0000001x", "with duration")
v.DurationSeconds = 750
id, err := s.UpsertVideo(ctx, v)
require.NoError(t, err)
p := rawPool(t)
readDuration := func() *int {
var d *int
require.NoError(t, p.QueryRow(ctx, `SELECT duration_s FROM videos WHERE id = $1`, id).Scan(&d))
return d
}
require.NotNil(t, readDuration())
require.Equal(t, 750, *readDuration(), "duration must persist")
// A later upsert that does NOT know the duration (0) must not clobber it —
// the channel_title backfill stance (migration 014): COALESCE-preserve.
v2 := ytVideo(userA, "dur0000001x", "title updated, duration unknown")
v2.DurationSeconds = 0
_, err = s.UpsertVideo(ctx, v2)
require.NoError(t, err)
require.NotNil(t, readDuration(), "a 0/unknown re-upsert must not erase a known duration")
require.Equal(t, 750, *readDuration())
}
func TestUpsertVideo_IDMatchesSummaryDedup(t *testing.T) {
ctx := context.Background()
s := newStore(t)
@@ -82,36 +112,55 @@ func TestUpsertVideo_PerUserIsolation(t *testing.T) {
require.NotEqual(t, idA, idB, "same provider video for two users must be two distinct rows")
}
func TestNewestUnsummarizedVideoIDs(t *testing.T) {
func TestOnboardBurstVideoIDs(t *testing.T) {
ctx := context.Background()
s := newStore(t)
resetDB(t, rawPool(t))
mk := func(user, pid string, day int) string {
mk := func(user, pid string, day, dur int) string {
v := ytVideo(user, pid, pid)
v.PublishedAt = time.Date(2026, 6, day, 12, 0, 0, 0, time.UTC)
v.DurationSeconds = dur // 0 == unknown (NULL)
id, err := s.UpsertVideo(ctx, v)
require.NoError(t, err)
return id
}
_ = mk(userA, "a1vid000001", 1)
id2 := mk(userA, "a2vid000002", 2)
id3 := mk(userA, "a3vid000003", 3)
id4 := mk(userA, "a4vid000004", 4)
mk(userB, "b1vid000009", 9) // userB's newest — must never leak via RLS
summarized := mk(userA, "summ0000001", 6, 600) // newest known-good, but already summarized
good1 := mk(userA, "good0000001", 5, 600) // 10m, newest UNsummarized known-good
tooLong := mk(userA, "toolong0001", 4, 20000) // > maxSeconds -> dropped
tooShort := mk(userA, "tooshort001", 3, 30) // < minSeconds -> dropped
unknown := mk(userA, "unknown0001", 2, 0) // NULL duration -> kept, ranked last
good2 := mk(userA, "good0000002", 1, 800) // known-good but oldest
mk(userB, "bvid0000009", 9, 600) // userB -> must not leak via RLS
// The newest (v4) is summarized, so it's excluded from "unsummarized".
require.NoError(t, s.Deliver(ctx, summary(userA, id4, "done")))
// The newest video is summarized, so it is excluded from the burst.
require.NoError(t, s.Deliver(ctx, summary(userA, summarized, "done")))
// Cap 2, newest-first unsummarized: v3 then v2 (v4 excluded; userB excluded).
got, err := s.NewestUnsummarizedVideoIDs(ctx, userA, 2)
const minSec, maxSec = 60, 14400
// Known-good ranked before unknown, each newest-first within its group; the
// too-long and too-short videos are excluded by their known duration.
got, err := s.OnboardBurstVideoIDs(ctx, userA, 5, minSec, maxSec)
require.NoError(t, err)
require.Equal(t, []string{id3, id2}, got)
require.Equal(t, []string{good1, good2, unknown}, got,
"known-good first (newest-first), then unknown-duration; junk excluded")
none, err := s.NewestUnsummarizedVideoIDs(ctx, userA, 0)
// Cap is honoured.
capped, err := s.OnboardBurstVideoIDs(ctx, userA, 2, minSec, maxSec)
require.NoError(t, err)
require.Empty(t, none, "limit 0 returns nothing")
require.Equal(t, []string{good1, good2}, capped)
// Bounds disabled (0/0) == pure newest-first, nothing excluded.
all, err := s.OnboardBurstVideoIDs(ctx, userA, 10, 0, 0)
require.NoError(t, err)
require.ElementsMatch(t, []string{good1, tooLong, tooShort, unknown, good2}, all,
"0/0 bounds disable the duration filter (prior newest-first behaviour)")
// limit <= 0 returns nothing.
none, err := s.OnboardBurstVideoIDs(ctx, userA, 0, minSec, maxSec)
require.NoError(t, err)
require.Empty(t, none)
}
func TestUpsertVideoPersistsChannelAndDistinctChannels(t *testing.T) {
+4
View File
@@ -302,6 +302,10 @@ func (a *Adapter) filterLowValue(ctx context.Context, client *http.Client, video
if m.seconds > 0 && m.seconds < a.cfg.MinVideoSeconds {
continue // Short / sub-threshold clip
}
// Carry the duration we already fetched onto the kept video so the store
// can persist it (ADR-028) — the burst's length-aware selection depends on
// it. Discarding it here was the gap the onboarding investigation found.
v.DurationSeconds = m.seconds
kept = append(kept, v)
}
return kept
@@ -224,6 +224,11 @@ func TestNewVideosFiltersShortsAndLive(t *testing.T) {
if len(vids) != 1 || vids[0].ProviderVideoID != "long1" {
t.Fatalf("expected only long1 to survive the filter, got %+v", vids)
}
// The duration fetched for the filter is carried onto the kept video so the
// store can persist it (ADR-028) instead of discarding it.
if vids[0].DurationSeconds != 750 {
t.Fatalf("kept video DurationSeconds = %d, want 750 (PT12M30S)", vids[0].DurationSeconds)
}
}
// TestNewVideosNoFilterWhenDisabled: MinVideoSeconds=0 keeps the pre-ADR-023
+27
View File
@@ -120,6 +120,21 @@ type Config struct {
// caption rate gate (ADR-014) — the cap bounds count, never the pacing. Default 3.
OnboardSummarizeCount int
// OnboardSummarizerModel is the summarizer alias the connect-time burst leads
// its chain with (ADR-028) — a stronger model is affordable on the ≤3 summaries
// that form a new user's first impression. It heads a burst-specific chain;
// the standard chain (ADR-022) follows as resilience. Empty (or equal to
// SummarizerModel) collapses the burst back onto the shared processor — the
// reversibility lever. Default iguana/gemma4-26b (the brain-validated model).
OnboardSummarizerModel string
// OnboardMaxVideoSeconds upper-bounds the duration of a video the onboarding
// burst will pick (ADR-028), so the burst does not spend a scarce caption fetch
// on a multi-hour livestream VOD that passed the live filter once it ended. Only
// a KNOWN duration outside [MinVideoSeconds, this] is dropped; a NULL/unknown
// duration is kept (degrade-open). 0 disables the upper bound. Default 14400 (4h).
OnboardMaxVideoSeconds int
// DiscoveryInterval, when > 0, makes `serve` run in-process scheduled discovery
// for ALL users on that cadence (ADR-018). Zero/unset = disabled, so dev and
// tests never auto-fetch. Single-replica assumption — see cmdServe.
@@ -169,6 +184,8 @@ const (
defaultAutoSummarizeWindow = 7 * 24 * time.Hour
defaultOnboardSummarizeCount = 3
maxOnboardSummarizeCount = 5
defaultOnboardSummarizerModel = "iguana/gemma4-26b"
defaultOnboardMaxVideoSeconds = 14400 // 4h
)
// Load reads the environment into a Config, applying defaults. It does not
@@ -181,6 +198,7 @@ func Load() (Config, error) {
GatewayURL: envOr("TAPIR_GATEWAY_URL", defaultGatewayURL),
GatewayKey: os.Getenv("TAPIR_GATEWAY_KEY"),
SummarizerModel: envOr("TAPIR_SUMMARIZER_MODEL", defaultSummarizerModel),
OnboardSummarizerModel: lookupOr("TAPIR_ONBOARD_SUMMARIZER_MODEL", defaultOnboardSummarizerModel),
FallbackModel: lookupOr("TAPIR_FALLBACK_MODEL", defaultFallbackModel),
CloudFallbackModel: lookupOr("TAPIR_CLOUD_FALLBACK_MODEL", defaultCloudFallbackModel),
DBDSN: os.Getenv("TAPIR_DB_DSN"),
@@ -286,6 +304,15 @@ func Load() (Config, error) {
}
c.OnboardSummarizeCount = onboard
onboardMax, err := intOr("TAPIR_ONBOARD_MAX_VIDEO_SECONDS", defaultOnboardMaxVideoSeconds)
if err != nil {
return Config{}, err
}
if onboardMax < 0 {
onboardMax = 0
}
c.OnboardMaxVideoSeconds = onboardMax
return c, nil
}
+59
View File
@@ -214,3 +214,62 @@ func TestValidateForAuth_PassesWhenComplete(t *testing.T) {
t.Errorf("ValidateForAuth: unexpected error %v", err)
}
}
func TestLoad_OnboardSummarizerModel(t *testing.T) {
cases := []struct {
name, env string
set bool
want string
}{
{"default", "", false, defaultOnboardSummarizerModel},
{"explicit", "koala/some-model", true, "koala/some-model"},
{"empty disables (collapses to shared processor)", "", true, ""},
}
for _, c := range cases {
t.Run(c.name, func(t *testing.T) {
env := map[string]string{}
if c.set {
env["TAPIR_ONBOARD_SUMMARIZER_MODEL"] = c.env
}
setEnv(t, env)
cfg, err := Load()
if err != nil {
t.Fatalf("Load: %v", err)
}
if cfg.OnboardSummarizerModel != c.want {
t.Fatalf("OnboardSummarizerModel = %q, want %q", cfg.OnboardSummarizerModel, c.want)
}
})
}
}
func TestLoad_OnboardMaxVideoSeconds(t *testing.T) {
cases := []struct {
name, env string
want int
}{
{"default", "", defaultOnboardMaxVideoSeconds},
{"explicit", "7200", 7200},
{"zero disables", "0", 0},
{"negative clamps to zero", "-9", 0},
}
for _, c := range cases {
t.Run(c.name, func(t *testing.T) {
setEnv(t, map[string]string{"TAPIR_ONBOARD_MAX_VIDEO_SECONDS": c.env})
cfg, err := Load()
if err != nil {
t.Fatalf("Load: %v", err)
}
if cfg.OnboardMaxVideoSeconds != c.want {
t.Fatalf("OnboardMaxVideoSeconds = %d, want %d", cfg.OnboardMaxVideoSeconds, c.want)
}
})
}
}
func TestLoad_OnboardMaxVideoSecondsInvalid(t *testing.T) {
setEnv(t, map[string]string{"TAPIR_ONBOARD_MAX_VIDEO_SECONDS": "long"})
if _, err := Load(); err == nil {
t.Fatal("Load: want error for non-numeric TAPIR_ONBOARD_MAX_VIDEO_SECONDS")
}
}
+5
View File
@@ -77,6 +77,11 @@ type Video struct {
URL string
PublishedAt time.Time
SeenAt time.Time
// DurationSeconds is the video length in seconds, when known (fetched by the
// ADR-023 videos.list enrichment at discovery). 0 means unknown — the store
// preserves a previously-known value rather than overwriting it with 0, and
// the onboarding burst (ADR-028) treats unknown as degrade-open (kept).
DurationSeconds int
}
// Transcript is the text of a video (or a record that none was available).
+16 -7
View File
@@ -64,13 +64,21 @@ func (a *App) handleChat(w http.ResponseWriter, r *http.Request) {
if !ok {
return
}
a.render(w, r, ChatPage(chatView{
view := chatView{
VideoID: row.VideoID,
Title: displayTitle(*row),
Available: hasText,
Models: a.Chat.Models(),
Selected: a.Chat.DefaultModel(row.AIModel),
}))
}
// HTMX (the in-place reveal from the summary) gets just the open chat section,
// swapped over the closed dock so the summary above it stays put. A no-JS
// navigation gets the full page: the whole summary plus the open chat.
if isHTMX(r) {
a.render(w, r, chatSection(view))
return
}
a.render(w, r, ChatPage(*row, view))
}
// handleChatMessage answers one question against the stored transcript (POST).
@@ -129,7 +137,7 @@ func (a *App) handleChatMessage(w http.ResponseWriter, r *http.Request) {
view.Truncated = reply.Truncated
}
}
a.renderChat(w, r, view)
a.renderChatTurn(w, r, *row, view)
}
// loadOwnedSummary fetches the summary for the path's video scoped to userID, or
@@ -166,14 +174,15 @@ func (a *App) readTranscript(w http.ResponseWriter, r *http.Request, row store.S
return t.Content, true, true
}
// renderChat returns the chat panel fragment for an HTMX request, or the full
// chat page otherwise (no-JS POST re-renders the whole page).
func (a *App) renderChat(w http.ResponseWriter, r *http.Request, v chatView) {
// renderChatTurn returns the chat panel fragment for an HTMX answer (swapped in
// place within the open dock), or the full chat page otherwise — the no-JS POST
// re-renders the whole summary + open chat with the new turn.
func (a *App) renderChatTurn(w http.ResponseWriter, r *http.Request, row store.SummaryRow, v chatView) {
if isHTMX(r) {
a.render(w, r, chatPanel(v))
return
}
a.render(w, r, ChatPage(v))
a.render(w, r, ChatPage(row, v))
}
// resolveModel keeps the posted model only when it is an offered option; anything
+35
View File
@@ -137,6 +137,8 @@ func TestChatEntryAffordanceOnSummaryView(t *testing.T) {
html := body(t, do(t, withChat, httptest.NewRequest(http.MethodGet, "/v/"+videoX, nil)))
require.Contains(t, html, "Dig deeper", "the deeper-dive affordance is shown when chat is enabled")
require.Contains(t, html, "/v/"+videoX+"/chat", "it links to this video's chat")
require.Contains(t, html, `id="chat-section"`, "the dock lives on the detail page")
require.Contains(t, html, `hx-get="/v/`+videoX+`/chat"`, "it opens the chat in place (HTMX), not a navigation")
// With no chat backend wired the affordance is absent (routes unmounted).
noChat := newApp(t)
@@ -145,6 +147,39 @@ func TestChatEntryAffordanceOnSummaryView(t *testing.T) {
require.NotContains(t, html, "Dig deeper", "no affordance when chat is disabled")
}
// The summary and the chat live together (the integrated UX): the no-JS chat page
// renders the full summary alongside the chat, and the HTMX reveal returns just
// the open chat section as a fragment so it docks in below the summary already on
// screen — the summary is never navigated away from.
func TestChatIntegratedWithSummaryOnSamePage(t *testing.T) {
ctx := context.Background()
chatter := &fakeChatter{models: []string{"phi4-mini", "gemma4-26b"}}
app := newChatApp(t, chatter, nil)
p := rawPool(t)
resetDB(t, p)
require.NoError(t, deliver(ctx, app, videoX, "SUMMARY-BODY-MARKER"))
seedVideo(t, p, videoX, "X Title", "https://x", time.Time{})
seedTranscript(t, app.Store.(*store.Store), videoX, "the transcript")
// No-JS full page: the summary payload and the chat are on one page.
html := body(t, getChat(t, app, videoX))
require.Contains(t, html, "SUMMARY-BODY-MARKER", "the summary text is shown on the chat page")
require.Contains(t, html, "takeaway one", "takeaways shown alongside the chat")
require.Contains(t, html, "highlight one", "highlights shown alongside the chat")
require.Contains(t, html, "Ask about this video", "the chat sits on the same page as the summary")
require.Contains(t, html, `name="question"`, "the ask form is present")
// HTMX reveal: the open chat section ONLY (a fragment) — no full-page chrome and
// no duplicated summary, so it swaps in below the summary already rendered.
req := httptest.NewRequest(http.MethodGet, "/v/"+videoX+"/chat", nil)
req.Header.Set("HX-Request", "true")
frag := body(t, do(t, app, req))
require.NotContains(t, frag, "<html", "the reveal is a fragment, not a full page")
require.NotContains(t, frag, "SUMMARY-BODY-MARKER", "the reveal does not re-send the summary (it's already on screen)")
require.Contains(t, frag, `id="chat-section"`, "the fragment replaces the dock in place")
require.Contains(t, frag, `name="question"`, "the ask form is in the revealed section")
}
// THE KEY SAFETY ASSERTION (ADR-027): a chat answer is produced entirely from the
// stored transcript — the model receives the stored text, and neither the
// summarize→fetch path nor the YouTube fetch path is ever touched. The tripwire
+5 -3
View File
@@ -731,9 +731,11 @@ a.btn, a.btn:visited { color: var(--accent-fg); }
.detail ul { margin: 0; padding-left: 1.2rem; line-height: 1.6; }
.detail li { margin-bottom: var(--s1); }
/* deeper-dive chat (ADR-027) */
.dig-deeper { margin: var(--s3) 0 0; }
.chat .chat-scope { margin: 0 0 var(--s4); font-size: .9rem; }
/* deeper-dive chat (ADR-027) — docks in place below the summary */
.chat-dock { margin-top: var(--s5); border-top: 1px solid var(--line); padding-top: var(--s4); }
.chat-dock .chat-open { display: inline-block; }
.chat-heading { font-size: 1.1rem; margin: 0 0 var(--s2); }
.chat-scope { margin: 0 0 var(--s3); font-size: .9rem; }
.chat-panel { display: flex; flex-direction: column; gap: var(--s3); }
.chat-log { display: flex; flex-direction: column; gap: var(--s3); }
.chat-turn { border-radius: var(--radius); padding: var(--s2) var(--s3); }
+58 -26
View File
@@ -410,14 +410,11 @@ templ noCaptionsCard(r store.SummaryRow) {
</li>
}
// DetailPage is the full summary view: text, highlights, takeaways, metadata,
// and the action button group. chatEnabled adds the "dig deeper" affordance
// (ADR-027) — a quiet link into the per-video chat over the stored transcript.
templ DetailPage(r store.SummaryRow, chatEnabled bool) {
@Layout("Tapir — " + displayTitle(r)) {
<article class="detail">
<p class="back"><a href="/"> Summaries</a></p>
<h1>{ displayTitle(r) }</h1>
// summaryBody is the summary payload shared by the detail page and the no-JS
// chat page (so the chat page shows the same summary, not a separate view):
// metadata, embed, source, the action toggles, then the attention-saving order
// Takeaways → Highlights → Summary (UX review A8).
templ summaryBody(r store.SummaryRow) {
<p class="meta">
if detailMeta(r) != "" {
<span>{ detailMeta(r) }</span>
@@ -441,16 +438,7 @@ templ DetailPage(r store.SummaryRow, chatEnabled bool) {
if r.URL != "" {
<p class="source"><a href={ externalURL(r.URL) } rel="noopener noreferrer">watch on source </a></p>
}
if chatEnabled {
// Where the "I want more" reaction goes (ADR-027). A quiet link, not a
// loud CTA — it deepens value for a reader already here, never nudges.
<p class="dig-deeper"><a class="btn-secondary" href={ chatURL(r.VideoID) }>Dig deeper ask about this video </a></p>
}
@ActionButtons(r.VideoID, actionSet(r.Actions))
// Lead with the attention-saving payload: Takeaways ("is this worth my
// time?") first, then Highlights, then the full Summary last (UX review
// A8). Takeaways/Highlights are conditional, so a video without them falls
// through to the Summary leading naturally.
if len(r.Takeaways) > 0 {
<section>
<h2>Takeaways</h2>
@@ -475,21 +463,65 @@ templ DetailPage(r store.SummaryRow, chatEnabled bool) {
<h2>Summary</h2>
<p class="body">{ r.Summary }</p>
</section>
}
// DetailPage is the full summary view: the summary payload, then (when chat is
// enabled) the deeper-dive dock (ADR-027) — a reveal that opens the chat IN PLACE
// below the summary, so the summary stays on screen as the context being asked
// about rather than being navigated away from.
templ DetailPage(r store.SummaryRow, chatEnabled bool) {
@Layout("Tapir — " + displayTitle(r)) {
<article class="detail">
<p class="back"><a href="/"> Summaries</a></p>
<h1>{ displayTitle(r) }</h1>
@summaryBody(r)
if chatEnabled {
@chatReveal(r.VideoID)
}
</article>
}
}
// ChatPage is the per-video deeper-dive chat (ADR-027): a question/answer surface
// over the video's ALREADY-STORED transcript. The back link returns to the
// summary it was entered from. All the interaction lives in chatPanel so the HTMX
// answer-swap and the no-JS full-page render share one component.
templ ChatPage(v chatView) {
@Layout("Tapir — Ask — " + v.Title) {
<article class="detail chat">
<p class="back"><a href={ videoURL(v.VideoID) }> { v.Title }</a></p>
<h1>Ask about this video</h1>
// chatReveal is the CLOSED dock at the foot of the summary: a quiet affordance,
// not a loud CTA (it deepens value for a reader already here, never nudges). With
// JS it swaps itself for the open chat section in place (HTMX, summary stays
// above); without JS the same href navigates to the full chat page, which renders
// the summary alongside the chat. Either way the summary is never lost.
templ chatReveal(videoID string) {
<section id="chat-section" class="chat-dock">
<a
class="btn-secondary chat-open"
href={ chatURL(videoID) }
hx-get={ string(chatURL(videoID)) }
hx-target="#chat-section"
hx-swap="outerHTML"
>
Dig deeper ask about this video
</a>
</section>
}
// chatSection is the OPEN dock: heading + scope note + the chat panel, swapped in
// over the closed reveal (same #chat-section id, outerHTML). It is the HTMX reveal
// response AND the inline chat block on the no-JS chat page.
templ chatSection(v chatView) {
<section id="chat-section" class="chat-dock chat-dock-open">
<h2 class="chat-heading">Ask about this video</h2>
<p class="chat-scope muted">Answers come only from this video's stored transcript Tapir never fetches anything new here.</p>
@chatPanel(v)
</section>
}
// ChatPage is the no-JS full-page render of the chat: the whole summary followed
// by the open chat dock, so a visitor without JS sees the same integrated view
// (summary beside the conversation) that JS users get inline via the reveal.
templ ChatPage(r store.SummaryRow, v chatView) {
@Layout("Tapir — " + displayTitle(r)) {
<article class="detail">
<p class="back"><a href="/"> Summaries</a></p>
<h1>{ displayTitle(r) }</h1>
@summaryBody(r)
@chatSection(v)
</article>
}
}
File diff suppressed because it is too large Load Diff
+3 -1
View File
@@ -40,7 +40,8 @@ var scenarioCoverage = map[string]string{
// connect_account.feature
"Connect a YouTube account": "TestCallbackExchangesAndRecordsConnection",
"Connecting an account discovers videos immediately": "TestCallbackTriggersDiscovery",
"Connecting summarizes my newest videos right away": "TestNewestUnsummarizedVideoIDs",
"Connecting summarizes my best recent videos right away": "TestOnboardBurstVideoIDs",
"The onboarding burst summarizes with a stronger model": "TestBurstChainModelsLeadsWithOnboardModel",
"Tokens are never stored in the clear": "TestCallbackExchangesAndRecordsConnection",
// paste_url.feature
@@ -65,6 +66,7 @@ var scenarioCoverage = map[string]string{
// chat_transcript.feature (ADR-027)
"A summary view offers a deeper-dive into the video": "TestChatEntryAffordanceOnSummaryView",
"The summary and the chat are on one page": "TestChatIntegratedWithSummaryOnSamePage",
"Ask a question answered from the stored transcript": "TestChatAnswersFromStoredTranscriptWithoutAnyFetch",
"Chat never fetches captions or reaches YouTube": "TestChatAnswersFromStoredTranscriptWithoutAnyFetch",
"A video with no stored transcript offers no chat": "TestChatUnavailableWhenNoStoredTranscript",