Compare commits

..
8 Commits
Author SHA1 Message Date
mathiasandClaude Opus 4.8 5219561a91 feat(discovery): per-channel caption-availability memory (ADR-024)
CI / Lint / Test / Vet (push) Successful in 30s
CI / Build & Import (push) Successful in 11s
After transcript caching (ADR-021) and the Shorts filter (ADR-023), the remaining
caption waste is the first fetch on every new video of a channel that never has
English captions — each costs one rate-limited fetch to resolve to "none", and on
a throttled IP churns the backoff machinery first.

Remember, per (user, channel), a streak of consecutive no-caption outcomes
(channel_caption_state, migration 016, RLS-scoped). Once it reaches
TAPIR_CHANNEL_CAPTIONLESS_THRESHOLD (default 5) the channel is suppressed — videos
discovered/listed but not caption-fetched — for TAPIR_CHANNEL_CAPTIONLESS_WINDOW
(default 14d), then one is re-probed (auto-recovery). A successful fetch resets
the streak; a 429 does not count; an explicit manual request bypasses suppression.
threshold=0 disables.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-10 20:27:25 +02:00
mathiasandClaude Opus 4.8 9db06d8a63 feat(discovery): drop Shorts and livestreams before the caption fetch (ADR-023)
CI / Lint / Test / Vet (push) Successful in 12s
CI / Build & Import (push) Successful in 10s
The scarce resource is the per-IP timedtext caption fetch (ADR-014); the pilot's
candidate set was mostly Shorts/clips/livestreams, each burning a fetch (a "none"
result is a completed fetch — it costs budget even when it yields nothing).

NewVideos now enriches candidates with one cheap Data API videos.list call
(contentDetails.duration + snippet.liveBroadcastContent — the quota API, a
DIFFERENT limit from the timedtext 429) and drops, before returning: videos
shorter than TAPIR_MIN_VIDEO_SECONDS (default 60) and any live/upcoming
broadcast. Dropped videos are never persisted, so the list declutters too.

Degrade-open: MinVideoSeconds=0 disables it (no quota call); a videos.list error
returns candidates unfiltered so discovery never breaks on a metadata hiccup. The
paste-a-URL path (VideoByID) is not filtered — an explicit request is honoured.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-10 19:34:45 +02:00
mathiasandClaude Opus 4.8 1aa8a97f95 fix(scheduler): rotate over connected users only for true fair share
CI / Lint / Test / Vet (push) Successful in 12s
CI / Build & Import (push) Successful in 10s
The lead-user rotation rotated the full ListAllUsers set, so a connectionless
orphan identity ate a rotation slot — collapsing onto the next real user and
skewing the lead share (two real users got 2/3 vs 1/3 instead of 50/50). Filter
to connected users BEFORE rotating so the rotation is over exactly the users that
consume the caption budget. A dead identity can no longer skew fairness.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-10 18:28:54 +02:00
mathiasandClaude Opus 4.8 cc69a912f4 fix(scheduler): rotate lead user each pass so caption budget is shared
CI / Lint / Test / Vet (push) Successful in 12s
CI / Build & Import (push) Successful in 10s
Caption fetches share one per-egress-IP rate budget; whoever runs first each pass
spends the pre-throttle window before YouTube starts 429ing. ListAllUsers order
is unspecified and was stable, so the last-listed user was permanently starved —
a friendly-pilot user got 0 fetches in 12h (all rate_limited) while the
first-listed user got every successful fetch. rotateUsers left-rotates the user
order by pass index so each user leads 1/N passes and the lead slot is shared.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-10 18:07:01 +02:00
mathiasandClaude Opus 4.8 f4a0544903 fix(scheduler): cache transcripts on the scheduled path (ADR-021 regression)
buildUserRunner built the engine without engine.Transcripts = st, so the
scheduler — unlike the web "Summarize now" path — never read or wrote the shared
transcript cache. Every discovery pass re-fetched transcripts it had already
fetched, burning the scarce per-egress-IP caption budget (ADR-014) on redundant
work and starving other users' first-time fetches. The transcripts table was
empty despite summaries existing. Wire the cache on this path too.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-10 18:07:01 +02:00
mathiasandClaude Opus 4.8 e2a52789b9 feat(web): mode-aware backlog banner — stop telling manual users summaries auto-arrive
CI / Lint / Test / Vet (push) Successful in 12s
CI / Build & Import (push) Successful in 10s
The first pilot user sat in Manual mode reading "new summaries land gradually,
check back tomorrow" — copy that only makes sense in Automatic mode. Manual mode
never auto-summarizes, so the banner promised delivery that would never come.

The list page now reads the user's summarize mode and shows mode-correct copy:
- Auto: unchanged "land gradually" backlog note.
- Manual: "new videos appear here but are not summarized automatically — use the
  Summarize button" plus a "Switch to Automatic" link to /account.
The connected-but-empty first-run state is likewise mode-aware.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-10 09:14:52 +02:00
mathias e696b6405b docs(build-state): v0.15.0, summarizer is now a resilient chain (ADR-022)
CI / Lint / Test / Vet (push) Successful in 12s
CI / Build & Import (push) Successful in 10s
2026-06-10 08:59:26 +02:00
mathiasandClaude Opus 4.8 e9b5a3f3e7 feat(summarizer): resilient endpoint chain with local→cloud fallback (ADR-022)
CI / Lint / Test / Vet (push) Successful in 11s
CI / Build & Import (push) Successful in 10s
The first friendly-pilot live run produced zero summaries: koala/phi4-mini hit
three silent failure modes — 8k context overflow on long transcripts (HTTP 400),
intermittent malformed JSON (highlights as a bare string), and no fallback wired
at all (summarizer.New(primary, nil)).

Keep phi4-mini as the fast primary and add resilience around it:

- Ordered endpoint chain (summarizer.NewChain): phi4-mini → koala/phi4-14b
  (local) → berget/mistral-small (worst-case external). All reached through the
  one LiteLLM gateway by alias.
- A parse failure now advances the chain like a transport error — the old
  Primary→Fallback shape returned the parse error without trying anyone else.
- Tolerant parse: highlights/takeaways coerce string→[]string, absorbing the
  common small-model quirk without spending a fallback round-trip.
- Transcript truncation (TAPIR_MAX_TRANSCRIPT_CHARS=18000) prevents the overflow
  rather than recovering from it; validated to fit phi4-mini's 8k window.
- Bounded completion budget (TAPIR_SUMMARY_MAX_TOKENS=1500) — the old 8192 budget
  itself contributed to the overflow.

Local-first guarantee preserved by ordering: external endpoint is tried only
after every local one fails. TAPIR_CLOUD_FALLBACK_MODEL="" disables it entirely
for client/NDA deployments.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-10 08:36:16 +02:00
29 changed files with 1599 additions and 364 deletions
+3 -2
View File
@@ -86,7 +86,7 @@ Skills live in the canonical library `mathias/skills` and are wired into this re
## Current build state (start here for the first task)
The repo is **green and shipping** — last tag `v0.14.0`. `task check` passes (fmt, vet, lint,
The repo is **green and shipping** — last tag `v0.15.0`. `task check` passes (fmt, vet, lint,
`go test -p 1 ./...`). Go is `1.26.1` (see `go.mod`).
- Clean Architecture core is implemented: `internal/domain` (entities), `internal/ports`
@@ -95,7 +95,8 @@ The repo is **green and shipping** — last tag `v0.14.0`. `task check` passes (
green against it.
- Adapters present under `internal/adapters/`: `youtube` (captions-first `VideoSource`,
timedtext/InnerTube acquisition per ADR-010), `summarizer` + `llm` (the copied AI router,
Primary→Fallback per ADR-004), `store` (Postgres, golang-migrate migrations 001015),
now a resilient endpoint chain — local primary → local fallback → external worst-case,
parse-failure-aware, ADR-004 + ADR-022), `store` (Postgres, golang-migrate migrations 001015),
`secrets` (file-backed `SecretStore`). The brain HTTP sink (ADR-005) is the remaining
optional sink.
- Stage 1 is open (ADR-012): multi-user with **DB-enforced** isolation — Postgres RLS `FORCE`d
+120
View File
@@ -803,6 +803,126 @@ never exposed in the UI.
---
## ADR-022 — Summarizer is a resilient endpoint chain, not a single model
**Status:** Accepted (2026-06-10). **Extends ADR-004** (the copied `llm` Primary→Fallback
routing). Triggered by the first friendly-pilot live run, where a connected user got **zero**
summaries after 12h.
**Context.** Stage-0 ran a single summarizer model (`koala/phi4-mini`) with no fallback wired
(`summarizer.New(primary, nil)`). The live run exposed three independent failure modes, each of
which silently produced no summary:
1. **Context overflow.** `phi4-mini` has an 8k context. Real transcripts (one was 11,602 tokens)
exceed it and the gateway returns HTTP 400 — and the request also sent `max_tokens=8192`, so
even a short transcript plus the completion budget could overflow the window.
2. **Malformed model output.** `phi4-mini` intermittently emits `highlights` as a bare string
instead of an array, producing `cannot unmarshal string into []string`. The old code returned
the parse error **without** trying any other model — a 200-with-bad-JSON short-circuited.
3. **No fallback existed at all** — any primary failure was terminal for that video.
`phi4-mini` is kept as primary deliberately: it is fast and, on transcripts that fit, correct.
The fix is resilience around it, not replacing it.
**Decision.**
1. **Ordered endpoint chain (`summarizer.NewChain`).** Endpoints are tried in order; the first to
return a *parseable* summary wins. Default chain:
`koala/phi4-mini` (primary, local) → `koala/phi4-14b` (fallback, local) →
`berget/mistral-small` (worst-case, external). All three are reached through the **one** LiteLLM
gateway by alias — the gateway already fronts both llama-swap and berget — so a fallback is a
different alias, not a second client config.
2. **A parse failure advances the chain, same as a transport error.** "Reliably summarized" means
*parseable summary returned*, not *HTTP 200*. This is the behaviour the old Primary→Fallback
shape missed.
3. **Tolerant parse.** `highlights`/`takeaways` coerce from a bare string (or a mixed scalar
array) to `[]string`, so the most common small-model quirk is absorbed **without** spending a
fallback round-trip — keeping the fast path fast.
4. **Transcript truncation (`TAPIR_MAX_TRANSCRIPT_CHARS`, default 18000).** Input is bounded
up-front to fit a small-context primary, so overflow is prevented rather than recovered-from.
5. **Bounded completion budget (`TAPIR_SUMMARY_MAX_TOKENS`, default 1500).** A summary needs few
hundred tokens; the old 8192 budget itself contributed to 8k-window overflow.
**Local-first guarantee preserved.** The chain ordering *is* the guarantee: locals are tried
first, so content reaches the external endpoint only after every local endpoint has failed.
`TAPIR_CLOUD_FALLBACK_MODEL=""` removes the external endpoint entirely — the lever a
**client/NDA deployment** pulls so content never leaves the local stack. With no external endpoint
configured the `ai_routing.feature` "content only local" scenarios hold unchanged.
**Reversibility.** Pure wiring + config. Setting `TAPIR_FALLBACK_MODEL` and
`TAPIR_CLOUD_FALLBACK_MODEL` empty collapses the chain back to single-primary behaviour; the
tolerant parse and truncation are strict supersets of the old behaviour (a previously-parseable
reply still parses; a transcript within budget is unchanged).
---
## ADR-023 — Drop Shorts/livestreams at discovery to protect the caption budget
**Status:** Accepted (2026-06-10). **Builds on ADR-014** (per-IP caption rate limit is the
binding constraint) and the ADR-022 live-run findings.
**Context.** The scarce resource is the unofficial timedtext caption fetch (per-egress-IP
429, ~3 successful/pass). The first multi-user run showed the candidate set was mostly noise —
Shorts, sub-minute clips, and live broadcasts — each of which still consumes a caption-fetch
attempt (and a "none" result is a *completed* fetch, so it costs budget even when it yields
nothing). Spending the rate-limited budget on content the user will not read is the waste to
cut first; it is cheaper and lower-risk than raising the ceiling (multi-IP, Whisper).
Duration and live status are NOT in the playlistItems discovery response, but they ARE in the
Data API `videos.list` (contentDetails.duration + snippet.liveBroadcastContent) — the official
**quota-based** API (1 unit/call, 50 ids/call), which is a *different* limit from the timedtext
429. So one cheap quota call buys a filter that saves many expensive throttled fetches.
**Decision.**
1. `NewVideos` enriches its candidates with a single `videos.list` call and drops, before
returning: videos shorter than `TAPIR_MIN_VIDEO_SECONDS` (default 60) and any `live`/
`upcoming` broadcast. Dropped videos are never persisted, so they also declutter the list.
2. The filter is **degrade-open**: `MinVideoSeconds=0` disables it (no quota call), and a
`videos.list` error returns the candidates unfiltered — discovery must never break because a
metadata call hiccuped (worst case = pre-ADR-023 behaviour).
3. The paste-a-URL path (`VideoByID`) is **not** filtered — an explicit user request for a
specific video (even a Short) is honoured.
**Reversibility.** Pure discovery-time filter + config. `TAPIR_MIN_VIDEO_SECONDS=0` restores
the old behaviour. No schema change, no effect on already-stored videos.
**Quota note.** Per-channel enrichment adds ~1 unit/channel/pass. At pilot scale (≤3 users)
this is well under the 10k/day cap; at larger scale, batch `videos.list` across channels
(50 ids/call) by collecting all discovered ids per pass before enriching.
---
## ADR-024 — Per-channel caption-availability memory
**Status:** Accepted (2026-06-10). **Builds on ADR-014** (per-IP caption budget), **ADR-021**
(shared transcript cache), **ADR-023** (Shorts filter).
**Context.** After ADR-021 caches transcripts and ADR-023 drops Shorts, the remaining caption
waste is the *first* fetch on every new video of a channel that never publishes English captions
(foreign-language news, music, etc.). Each costs one rate-limited fetch to resolve to "none" —
and on a throttled IP that fetch may 429 and churn the backoff machinery before it ever gets a
verdict. A pilot user's feed had several such channels.
**Decision.** Remember, per `(user, channel)`, a streak of consecutive no-caption outcomes
(`channel_caption_state`, migration 016, RLS-scoped like the rest of the user-owned schema).
Once the streak reaches `TAPIR_CHANNEL_CAPTIONLESS_THRESHOLD` (default 5) the channel is
suppressed — its videos are discovered/listed but not caption-fetched — for
`TAPIR_CHANNEL_CAPTIONLESS_WINDOW` (default 14d), after which one video is re-probed
(auto-recovery for a channel that starts adding captions). A successful fetch resets the streak;
a fresh 429 does NOT count (transient, not a caption verdict). An explicit manual request
bypasses suppression. `threshold = 0` disables the feature.
**Why per-user, not global.** Caption availability is really a channel property (public), so a
global table would let users share the learning. But subscriptions are per-user (ADR-012) and at
pilot scale users' channel sets barely overlap, so per-user + RLS keeps it consistent with the
existing isolation model with no new non-RLS exception to justify. Promoting to a shared table
(like transcripts, ADR-021) is a future optimisation if channel overlap grows.
**Reversibility.** Migration 016 is a clean drop; `threshold = 0` disables at runtime. The
memory only ever *suppresses fetches* — it never deletes content or affects already-stored
summaries.
---
## Rejected alternatives
Approaches considered during the 2026-06-02 planning + grill session and **deliberately not
+2 -1
View File
@@ -136,7 +136,8 @@ func cmdRun(ctx context.Context, log *slog.Logger) error {
r := runner.New(engine.Source, st, engine, cfg.UserID, log,
runner.WithBackoff(cfg.FetchBackoff),
runner.WithAutoWindow(cfg.AutoSummarizeWindow))
runner.WithAutoWindow(cfg.AutoSummarizeWindow),
runner.WithCaptionMemory(cfg.ChannelCaptionlessThreshold, cfg.ChannelCaptionlessWindow))
log.Info("starting run", "user", cfg.UserID, "model", cfg.SummarizerModel,
"gateway", cfg.GatewayURL, "poll_interval", cfg.PollInterval, "fetch_backoff", cfg.FetchBackoff,
+37 -7
View File
@@ -3,6 +3,7 @@ package main
import (
"context"
"fmt"
"strings"
"gitea.d-ma.be/mathias/tapir/internal/adapters/llm"
"gitea.d-ma.be/mathias/tapir/internal/adapters/secrets"
@@ -42,6 +43,40 @@ func (f videoFetcher) FetchVideo(ctx context.Context, userID, videoID string) (d
// queue-only fallback: the web UI keeps working (the button just queues) and
// `tapir run` reports the gap via its own ValidateForRun. Missing engine config
// is never an error here.
// buildSummarizer wires the summarization endpoint chain (ADR-022) shared by the
// web "Summarize now" path and the scheduler's per-user runners. The chain is:
// primary (local, fast) → local fallback → cloud fallback (worst case). Each
// endpoint reaches the same LiteLLM gateway with a different model alias — the
// gateway fronts both llama-swap and berget — so a fallback is just a different
// alias, not a second client config. Empty model entries are skipped, so a
// client deployment can set the cloud fallback empty to keep content local.
func buildSummarizer(cfg config.Config) *summarizer.Summarizer {
mk := func(model string) summarizer.Endpoint {
return summarizer.Endpoint{
Client: llm.New(cfg.GatewayURL, cfg.GatewayKey, model, cfg.SummarizerTimeout, llm.WithMaxTokens(cfg.SummaryMaxTokens)),
Provider: providerOf(model),
Model: model,
}
}
eps := []summarizer.Endpoint{mk(cfg.SummarizerModel)}
if cfg.FallbackModel != "" && cfg.FallbackModel != cfg.SummarizerModel {
eps = append(eps, mk(cfg.FallbackModel))
}
if cfg.CloudFallbackModel != "" && cfg.CloudFallbackModel != cfg.SummarizerModel {
eps = append(eps, mk(cfg.CloudFallbackModel))
}
return summarizer.NewChain(eps, cfg.MaxTranscriptChars)
}
// providerOf maps a model alias to the domain AIProvider recorded on summaries.
// A "berget/" alias is an external provider; everything else is the local stack.
func providerOf(model string) string {
if strings.HasPrefix(model, "berget/") {
return "berget"
}
return "local"
}
func buildProcessor(cfg config.Config, st *store.Store) (*usecase.Engine, error) {
if cfg.GatewayURL == "" || cfg.YTClientID == "" || cfg.YTClientSecret == "" || cfg.SecretsFile == "" {
return nil, nil
@@ -53,15 +88,10 @@ func buildProcessor(cfg config.Config, st *store.Store) (*usecase.Engine, error)
ClientSecret: cfg.YTClientSecret,
TokenSecretRef: cfg.YTTokenRef,
PreferredLanguages: []string{"en"},
MinVideoSeconds: cfg.MinVideoSeconds,
}, secretStore)
// Local Primary only; no BYO fallback for the demo (fallback nil).
primary := summarizer.Endpoint{
Client: llm.New(cfg.GatewayURL, cfg.GatewayKey, cfg.SummarizerModel, cfg.SummarizerTimeout),
Provider: "local",
Model: cfg.SummarizerModel,
}
sum := summarizer.New(primary, nil)
sum := buildSummarizer(cfg)
// The store is both the summary sink and the shared transcript cache (ADR-021):
// the engine reads stored transcripts before any caption fetch and writes
+69 -26
View File
@@ -6,9 +6,7 @@ import (
"log/slog"
"time"
"gitea.d-ma.be/mathias/tapir/internal/adapters/llm"
"gitea.d-ma.be/mathias/tapir/internal/adapters/store"
"gitea.d-ma.be/mathias/tapir/internal/adapters/summarizer"
"gitea.d-ma.be/mathias/tapir/internal/adapters/youtube"
"gitea.d-ma.be/mathias/tapir/internal/config"
"gitea.d-ma.be/mathias/tapir/internal/ports"
@@ -35,18 +33,20 @@ func buildUserRunner(cfg config.Config, st *store.Store, secretStore ports.Secre
ClientSecret: cfg.YTClientSecret,
TokenSecretRef: web.YouTubeTokenRef(userID),
PreferredLanguages: []string{"en"},
MinVideoSeconds: cfg.MinVideoSeconds,
}, secretStore)
primary := summarizer.Endpoint{
Client: llm.New(cfg.GatewayURL, cfg.GatewayKey, cfg.SummarizerModel, cfg.SummarizerTimeout),
Provider: "local",
Model: cfg.SummarizerModel,
}
engine := usecase.NewEngine(src, summarizer.New(primary, nil), st)
engine := usecase.NewEngine(src, buildSummarizer(cfg), st)
// Share the transcript cache (ADR-021) on the scheduler path too — without
// this every scheduled pass re-fetches transcripts it already had, burning the
// scarce per-IP caption budget (ADR-014) and starving other users. The web
// "Summarize now" path already sets this; the scheduler omitting it was a bug.
engine.Transcripts = st
return runner.New(src, st, engine, userID, log,
runner.WithBackoff(cfg.FetchBackoff),
runner.WithAutoWindow(cfg.AutoSummarizeWindow)), nil
runner.WithAutoWindow(cfg.AutoSummarizeWindow),
runner.WithCaptionMemory(cfg.ChannelCaptionlessThreshold, cfg.ChannelCaptionlessWindow)), nil
}
// userLister enumerates every registered user and reports a user's video
@@ -64,6 +64,7 @@ type userLister interface {
// isolation). Returns the stats summed across users.
func runDiscoveryPass(
ctx context.Context,
pass int,
lister userLister,
runUser func(context.Context, string) (runner.Stats, error),
log *slog.Logger,
@@ -74,15 +75,17 @@ func runDiscoveryPass(
return runner.Stats{}
}
log.Info("scheduler: starting discovery pass", "users", len(users))
var total runner.Stats
// Keep only users with a video connection. A pass for a connectionless user
// (e.g. a stale Dex-era orphan identity) only tries to resolve a token that
// was never minted, logging a spurious "ref not found" every tick. Filtering
// here — BEFORE rotation — also keeps fairness honest: rotation is over the
// users that actually consume the caption budget, so a dead identity can't eat
// a rotation slot and skew the lead share.
var connected []store.UserIdentity
for _, u := range users {
if ctx.Err() != nil {
break // shutting down: stop enumerating
return runner.Stats{} // shutting down
}
// Skip users with no video connection. A discovery pass for them only
// attempts to resolve a token that was never minted, logging a spurious
// "ref not found" every tick (e.g. stale Dex-era orphan identities).
conns, err := lister.ConnectionsForUser(ctx, u.UserID)
if err != nil {
log.Warn("scheduler: list connections failed", "user", u.UserID, "err", err)
@@ -92,6 +95,22 @@ func runDiscoveryPass(
log.Debug("scheduler: skipping user with no video connections", "user", u.UserID)
continue
}
connected = append(connected, u)
}
// Rotate who goes first each pass. Caption fetches share one per-egress-IP
// rate budget (ADR-014); whoever runs first each pass spends the pre-throttle
// window, so a FIXED order permanently starves whoever is last (a new pilot
// user got 0 fetches for 12h while the first-listed user got all of them).
// Rotation over the connected set gives each real user the lead in turn.
connected = rotateUsers(connected, pass)
log.Info("scheduler: starting discovery pass", "users", len(connected))
var total runner.Stats
for _, u := range connected {
if ctx.Err() != nil {
break // shutting down: stop enumerating
}
stats, err := runUser(ctx, u.UserID)
total = sumStats(total, stats)
if err != nil {
@@ -103,6 +122,7 @@ func runDiscoveryPass(
"skipped_seen", total.SkippedSeen, "skipped_no_text", total.SkippedNoText,
"skipped_manual", total.SkippedManual, "skipped_too_old", total.SkippedTooOld,
"skipped_rate_limited", total.SkippedRateLimited,
"skipped_no_caption_channel", total.SkippedNoCaptionChannel,
"channel_unavailable", total.ChannelUnavailable, "errors", total.Errors)
return total
}
@@ -123,7 +143,8 @@ func runScheduler(
return // disabled
}
runDiscoveryPass(ctx, lister, runUser, log)
pass := 0
runDiscoveryPass(ctx, pass, lister, runUser, log)
ticker := time.NewTicker(interval)
defer ticker.Stop()
@@ -132,23 +153,45 @@ func runScheduler(
case <-ctx.Done():
return
case <-ticker.C:
runDiscoveryPass(ctx, lister, runUser, log)
pass++
runDiscoveryPass(ctx, pass, lister, runUser, log)
}
}
}
// rotateUsers left-rotates users by pass positions so a different user leads each
// pass. With n users, user i leads on every pass where pass ≡ i (mod n). A pass
// offset that is negative or exceeds n is normalised. Order within the rotation
// is otherwise preserved, so the set of users run is unchanged — only who is
// first (and thus wins the scarce caption-fetch budget) rotates.
func rotateUsers(users []store.UserIdentity, pass int) []store.UserIdentity {
n := len(users)
if n <= 1 {
return users
}
off := ((pass % n) + n) % n
if off == 0 {
return users
}
out := make([]store.UserIdentity, 0, n)
out = append(out, users[off:]...)
out = append(out, users[:off]...)
return out
}
// sumStats adds two passes' stats field-wise, so runDiscoveryPass can report a
// per-tick aggregate across all users.
func sumStats(a, b runner.Stats) runner.Stats {
return runner.Stats{
Candidates: a.Candidates + b.Candidates,
Summarized: a.Summarized + b.Summarized,
SkippedSeen: a.SkippedSeen + b.SkippedSeen,
SkippedNoText: a.SkippedNoText + b.SkippedNoText,
SkippedManual: a.SkippedManual + b.SkippedManual,
SkippedTooOld: a.SkippedTooOld + b.SkippedTooOld,
SkippedRateLimited: a.SkippedRateLimited + b.SkippedRateLimited,
ChannelUnavailable: a.ChannelUnavailable + b.ChannelUnavailable,
Errors: a.Errors + b.Errors,
Candidates: a.Candidates + b.Candidates,
Summarized: a.Summarized + b.Summarized,
SkippedSeen: a.SkippedSeen + b.SkippedSeen,
SkippedNoText: a.SkippedNoText + b.SkippedNoText,
SkippedManual: a.SkippedManual + b.SkippedManual,
SkippedTooOld: a.SkippedTooOld + b.SkippedTooOld,
SkippedRateLimited: a.SkippedRateLimited + b.SkippedRateLimited,
SkippedNoCaptionChannel: a.SkippedNoCaptionChannel + b.SkippedNoCaptionChannel,
ChannelUnavailable: a.ChannelUnavailable + b.ChannelUnavailable,
Errors: a.Errors + b.Errors,
}
}
+45 -4
View File
@@ -45,6 +45,7 @@ func (f fakeLister) ConnectionsForUser(_ context.Context, userID string) ([]stor
type countingRunUser struct {
mu sync.Mutex
calls map[string]int
order []string // userIDs in the order they were run, across all passes
failFor map[string]bool
}
@@ -60,12 +61,19 @@ func (c *countingRunUser) run(_ context.Context, userID string) (runner.Stats, e
c.mu.Lock()
defer c.mu.Unlock()
c.calls[userID]++
c.order = append(c.order, userID)
if c.failFor[userID] {
return runner.Stats{Errors: 1}, errors.New("boom")
}
return runner.Stats{Summarized: 1}, nil
}
func (c *countingRunUser) runOrder() []string {
c.mu.Lock()
defer c.mu.Unlock()
return append([]string(nil), c.order...)
}
func (c *countingRunUser) count(userID string) int {
c.mu.Lock()
defer c.mu.Unlock()
@@ -94,7 +102,7 @@ func TestDiscoveryPassRunsEveryUserOnce(t *testing.T) {
lister := fakeLister{users: usersN("a", "b", "c")}
rc := newCountingRunUser()
stats := runDiscoveryPass(context.Background(), lister, rc.run, quietLog())
stats := runDiscoveryPass(context.Background(), 0, lister, rc.run, quietLog())
require.Equal(t, 1, rc.count("a"))
require.Equal(t, 1, rc.count("b"))
@@ -102,13 +110,46 @@ func TestDiscoveryPassRunsEveryUserOnce(t *testing.T) {
require.Equal(t, 3, stats.Summarized, "stats are summed across users")
}
// Caption fetches share one per-IP budget; a fixed user order starves whoever is
// last. Each pass must rotate which user leads so the lead slot is shared.
func TestDiscoveryPassRotatesLeadUser(t *testing.T) {
lister := fakeLister{users: usersN("a", "b", "c")}
rc := newCountingRunUser()
runDiscoveryPass(context.Background(), 0, lister, rc.run, quietLog())
runDiscoveryPass(context.Background(), 1, lister, rc.run, quietLog())
runDiscoveryPass(context.Background(), 2, lister, rc.run, quietLog())
require.Equal(t, []string{"a", "b", "c", "b", "c", "a", "c", "a", "b"}, rc.runOrder(),
"each pass left-rotates the user order so every user leads in turn")
// Fairness: over a full rotation cycle every user ran the same number of times.
require.Equal(t, 3, rc.count("a"))
require.Equal(t, 3, rc.count("b"))
require.Equal(t, 3, rc.count("c"))
}
// A connectionless orphan must not consume a rotation slot: rotation is over the
// connected users only, so two real users alternate the lead 50/50 even with a
// dead identity listed between them.
func TestDiscoveryPassRotationIgnoresConnectionlessUsers(t *testing.T) {
lister := fakeLister{users: usersN("a", "orphan", "c"), noConn: map[string]bool{"orphan": true}}
rc := newCountingRunUser()
runDiscoveryPass(context.Background(), 0, lister, rc.run, quietLog())
runDiscoveryPass(context.Background(), 1, lister, rc.run, quietLog())
require.Equal(t, []string{"a", "c", "c", "a"}, rc.runOrder(),
"only connected users rotate; the orphan never runs and never holds a slot")
require.Equal(t, 0, rc.count("orphan"))
}
func TestDiscoveryPassSkipsUsersWithoutConnections(t *testing.T) {
// b never connected a video source (e.g. a stale Dex-era orphan identity).
// It must be skipped silently — not run and logged as a token error every pass.
lister := fakeLister{users: usersN("a", "b", "c"), noConn: map[string]bool{"b": true}}
rc := newCountingRunUser()
stats := runDiscoveryPass(context.Background(), lister, rc.run, quietLog())
stats := runDiscoveryPass(context.Background(), 0, lister, rc.run, quietLog())
require.Equal(t, 1, rc.count("a"))
require.Equal(t, 0, rc.count("b"), "a user with no connection must be skipped, not run")
@@ -121,7 +162,7 @@ func TestDiscoveryPassOneUserFailureDoesNotStopOthers(t *testing.T) {
lister := fakeLister{users: usersN("a", "b", "c")}
rc := newCountingRunUser("b") // user b's pass errors
stats := runDiscoveryPass(context.Background(), lister, rc.run, quietLog())
stats := runDiscoveryPass(context.Background(), 0, lister, rc.run, quietLog())
require.Equal(t, 1, rc.count("a"))
require.Equal(t, 1, rc.count("b"))
@@ -134,7 +175,7 @@ func TestDiscoveryPassListerErrorIsContained(t *testing.T) {
lister := fakeLister{err: errors.New("db down")}
rc := newCountingRunUser()
stats := runDiscoveryPass(context.Background(), lister, rc.run, quietLog())
stats := runDiscoveryPass(context.Background(), 0, lister, rc.run, quietLog())
require.Equal(t, 0, rc.total(), "no users enumerated → no passes")
require.Equal(t, runner.Stats{}, stats)
+22
View File
@@ -27,6 +27,28 @@ it** — endpoints and aliases drift, and this file is a snapshot (2026-06-06),
`iguana/deepseek-r1-14b`) is preferred for summary quality if its latency/output is acceptable.
The `max_tokens` fix below means thinking models no longer return empty content, so they are now
viable choices, not blocked ones. Do not assume a coder alias is right for prose.
- **Summarizer fallback chain (ADR-022).** The primary alias is the *first* of an ordered chain;
on failure or unparseable output the summarizer advances to the next model. All reached through
the same gateway by alias.
- `TAPIR_FALLBACK_MODEL` — local fallback. **Default `koala/phi4-14b`.** Empty disables it.
- `TAPIR_CLOUD_FALLBACK_MODEL` — worst-case EXTERNAL fallback. **Default `berget/mistral-small`.**
**Set this empty (`""`) for any client/NDA deployment** so content never leaves the local
stack — the chain then contains only local endpoints.
- `TAPIR_SUMMARY_MAX_TOKENS` — per-summary completion budget. **Default `1500`.** Small on
purpose: with the old 8192 budget, prompt + completion overflowed `phi4-mini`'s 8k window.
- `TAPIR_MAX_TRANSCRIPT_CHARS` — transcript truncation budget sent to the model. **Default
`18000`** (~fits an 8k-context model). `0` disables truncation. Prevents the context-overflow
HTTP 400 a long transcript caused on `phi4-mini`.
- **Discovery low-value filter (ADR-023).** `TAPIR_MIN_VIDEO_SECONDS`**default `60`**. At
discovery, `NewVideos` enriches candidates with one cheap `videos.list` call (quota API, NOT
the timedtext 429 path) and drops videos shorter than this plus any live/upcoming broadcast,
so the scarce caption-fetch budget isn't spent on Shorts. `0` disables the filter. The
paste-a-URL path is never filtered.
- **Per-channel caption memory (ADR-024).** `TAPIR_CHANNEL_CAPTIONLESS_THRESHOLD` — **default
`5`** consecutive no-caption results before a channel is suppressed (its videos listed but not
caption-fetched). `TAPIR_CHANNEL_CAPTIONLESS_WINDOW`**default `336h`** (14d) suppression
before one video is re-probed. `THRESHOLD=0` disables. A successful fetch resets the channel;
a 429 does not count; an explicit manual request bypasses suppression.
- **Thinking models need an explicit `max_tokens`.** qwen3 / deepseek-r1 spend the budget on
reasoning and return **empty content** if `max_tokens` is too low (or unset). The summarizer's
parser treats an empty summary as an error for exactly this reason. **Done (2026-06-02, Worker F):**
+13 -1
View File
@@ -31,5 +31,17 @@ Feature: Local-first AI with optional BYO fallback
When any transcript is summarized
Then my content is only ever sent to the local AI stack
# "Reliably" is operationalized as: Primary returned without error within timeout.
Scenario: A model returns unparseable output and the next endpoint succeeds
Given the local AI stack is available
But the primary model returns output that cannot be parsed into a summary
And a fallback model is configured
When a transcript is summarized
Then Tapir falls back to the next model in the chain
And the summary records fallback_used as true
# "Reliably" is operationalized as: an endpoint returned a PARSEABLE summary
# within timeout. A 200 with malformed JSON (or highlights emitted as a bare
# string) counts as a failure and advances the chain (ADR-022). Endpoints are
# tried in order, locals first, so the external worst-case model only ever sees
# content after every local endpoint has failed.
# Quality scoring may be added later without changing these scenarios.
+22 -2
View File
@@ -34,15 +34,35 @@ type Client struct {
httpClient *http.Client
}
// Option configures a Client at construction. Variadic so the existing 4-arg
// call sites stay valid as new knobs are added.
type Option func(*Client)
// WithMaxTokens overrides the per-request completion budget. The summarizer uses
// this to cap completion for small-context models (e.g. koala/phi4-mini, 8k):
// with the default 8192 budget, prompt + max_tokens overflows an 8k context and
// the gateway returns HTTP 400. A non-positive n is ignored (keeps the default).
func WithMaxTokens(n int) Option {
return func(c *Client) {
if n > 0 {
c.maxTokens = n
}
}
}
// New constructs a Client.
func New(baseURL, apiKey, model string, timeout time.Duration) *Client {
return &Client{
func New(baseURL, apiKey, model string, timeout time.Duration, opts ...Option) *Client {
c := &Client{
baseURL: strings.TrimRight(baseURL, "/"),
apiKey: apiKey,
model: model,
maxTokens: defaultMaxTokens,
httpClient: &http.Client{Timeout: timeout},
}
for _, opt := range opts {
opt(c)
}
return c
}
type chatRequest struct {
+21
View File
@@ -64,6 +64,27 @@ func TestClient_SendsMaxTokens(t *testing.T) {
}
}
// TestClient_WithMaxTokens overrides the completion budget — the summarizer caps
// it small so prompt + max_tokens fits a small-context model's window (8k).
func TestClient_WithMaxTokens(t *testing.T) {
var body chatRequest
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
_ = json.NewDecoder(r.Body).Decode(&body)
_ = json.NewEncoder(w).Encode(map[string]any{
"choices": []map[string]any{{"message": map[string]any{"content": "ok"}}},
})
}))
defer srv.Close()
c := New(srv.URL, "", "test-model", 10*time.Second, WithMaxTokens(1500))
if _, err := c.Complete(context.Background(), "sys", "user"); err != nil {
t.Fatalf("Complete: %v", err)
}
if body.MaxTokens != 1500 {
t.Errorf("max_tokens = %d, want 1500", body.MaxTokens)
}
}
func TestClient_ReturnsErrorOnNon200(t *testing.T) {
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
http.Error(w, "overloaded", http.StatusServiceUnavailable)
@@ -0,0 +1,83 @@
package store
import (
"context"
"fmt"
"time"
"github.com/jackc/pgx/v5"
)
// CaptionlessChannels returns the set of channel ids currently suppressed for the
// user — channels whose recent videos all yielded no captions, within their
// suppression window (ADR-024). The runner skips caption fetches for these
// channels' videos. A channel whose window has expired is not returned, so its
// next video is re-probed (auto-recovery).
func (s *Store) CaptionlessChannels(ctx context.Context, userID string) (map[string]bool, error) {
out := map[string]bool{}
err := s.withUser(ctx, userID, func(tx pgx.Tx) error {
rows, err := tx.Query(ctx, `
SELECT channel_id FROM channel_caption_state
WHERE user_id = $1 AND captionless_until IS NOT NULL AND captionless_until > now()`,
userID)
if err != nil {
return fmt.Errorf("store: caption-less channels: %w", err)
}
defer rows.Close()
for rows.Next() {
var ch string
if err := rows.Scan(&ch); err != nil {
return fmt.Errorf("store: scan caption-less channel: %w", err)
}
out[ch] = true
}
return rows.Err()
})
return out, err
}
// RecordChannelCaptionOutcome updates a channel's caption-availability memory
// after a fetch attempt (ADR-024). hadCaptions resets the channel (consecutive
// count to 0, suppression cleared). Otherwise the consecutive no-caption count is
// incremented; once it reaches threshold the channel is suppressed for window.
// threshold <= 0 is a no-op (feature disabled). An empty channelID is ignored
// (some sources may not carry one).
func (s *Store) RecordChannelCaptionOutcome(ctx context.Context, userID, channelID string, hadCaptions bool, threshold int, window time.Duration) error {
if channelID == "" || threshold <= 0 {
return nil
}
return s.withUser(ctx, userID, func(tx pgx.Tx) error {
if hadCaptions {
_, err := tx.Exec(ctx, `
INSERT INTO channel_caption_state (user_id, channel_id, consecutive_none, captionless_until, updated_at)
VALUES ($1, $2, 0, NULL, now())
ON CONFLICT (user_id, channel_id)
DO UPDATE SET consecutive_none = 0, captionless_until = NULL, updated_at = now()`,
userID, channelID)
if err != nil {
return fmt.Errorf("store: reset channel caption state: %w", err)
}
return nil
}
// No captions: increment the streak; suppress once it reaches threshold.
// captionless_until is set from the NEW count inside the same statement so
// the decision is atomic with the increment.
until := time.Now().Add(window)
_, err := tx.Exec(ctx, `
INSERT INTO channel_caption_state (user_id, channel_id, consecutive_none, captionless_until, updated_at)
VALUES ($1, $2, 1, CASE WHEN 1 >= $3 THEN $4::timestamptz ELSE NULL END, now())
ON CONFLICT (user_id, channel_id)
DO UPDATE SET
consecutive_none = channel_caption_state.consecutive_none + 1,
captionless_until = CASE
WHEN channel_caption_state.consecutive_none + 1 >= $3 THEN $4::timestamptz
ELSE channel_caption_state.captionless_until
END,
updated_at = now()`,
userID, channelID, threshold, until)
if err != nil {
return fmt.Errorf("store: record channel no-caption: %w", err)
}
return nil
})
}
@@ -0,0 +1,67 @@
package store_test
import (
"context"
"testing"
"time"
"github.com/stretchr/testify/require"
)
func TestChannelCaptionMemory_SuppressesAfterThreshold(t *testing.T) {
ctx := context.Background()
s := newStore(t)
super := rawPool(t)
resetDB(t, super)
seedUser(t, super, userA)
const threshold = 3
window := time.Hour
// Below threshold: not yet suppressed.
require.NoError(t, s.RecordChannelCaptionOutcome(ctx, userA, "chanX", false, threshold, window))
require.NoError(t, s.RecordChannelCaptionOutcome(ctx, userA, "chanX", false, threshold, window))
got, err := s.CaptionlessChannels(ctx, userA)
require.NoError(t, err)
require.NotContains(t, got, "chanX", "2 < threshold 3: not suppressed yet")
// Crossing the threshold suppresses the channel.
require.NoError(t, s.RecordChannelCaptionOutcome(ctx, userA, "chanX", false, threshold, window))
got, err = s.CaptionlessChannels(ctx, userA)
require.NoError(t, err)
require.Contains(t, got, "chanX", "3 consecutive no-caption results suppress the channel")
// A successful caption fetch resets it.
require.NoError(t, s.RecordChannelCaptionOutcome(ctx, userA, "chanX", true, threshold, window))
got, err = s.CaptionlessChannels(ctx, userA)
require.NoError(t, err)
require.NotContains(t, got, "chanX", "a captioned video clears suppression")
}
func TestChannelCaptionMemory_WindowExpiryReProbes(t *testing.T) {
ctx := context.Background()
s := newStore(t)
super := rawPool(t)
resetDB(t, super)
seedUser(t, super, userA)
// A negative window means captionless_until lands in the past — modelling an
// elapsed suppression window, which must make the channel eligible again.
require.NoError(t, s.RecordChannelCaptionOutcome(ctx, userA, "chanY", false, 1, -time.Hour))
got, err := s.CaptionlessChannels(ctx, userA)
require.NoError(t, err)
require.NotContains(t, got, "chanY", "an expired window re-enables the channel for a re-probe")
}
func TestChannelCaptionMemory_DisabledThresholdIsNoOp(t *testing.T) {
ctx := context.Background()
s := newStore(t)
super := rawPool(t)
resetDB(t, super)
seedUser(t, super, userA)
require.NoError(t, s.RecordChannelCaptionOutcome(ctx, userA, "chanZ", false, 0, time.Hour))
got, err := s.CaptionlessChannels(ctx, userA)
require.NoError(t, err)
require.Empty(t, got, "threshold 0 disables the memory — nothing recorded")
}
+10 -3
View File
@@ -53,7 +53,9 @@ func TestMigration010LoginEventsUpDown(t *testing.T) {
require.True(t, loginEventsExists(t), "login_events must exist at latest migration")
m := fileMigrator(t)
// 011..015 sit above 010; step them down first so 010 is exercised in isolation.
// 011..016 sit above 010; step them down first so 010 is exercised in isolation.
require.NoError(t, m.Steps(-1), "down 016 drops channel_caption_state, login_events intact")
require.True(t, loginEventsExists(t), "016 down leaves login_events intact")
require.NoError(t, m.Steps(-1), "down 015 reshapes transcripts, login_events intact")
require.True(t, loginEventsExists(t), "015 down leaves login_events intact")
require.NoError(t, m.Steps(-1), "down 014 drops channel_title, login_events intact")
@@ -68,7 +70,7 @@ func TestMigration010LoginEventsUpDown(t *testing.T) {
require.NoError(t, m.Steps(-1), "down 010 must drop login_events")
require.False(t, loginEventsExists(t), "login_events must be gone after the down migration")
require.NoError(t, m.Steps(6), "up must recreate 010 then re-apply 011..015")
require.NoError(t, m.Steps(7), "up must recreate 010 then re-apply 011..016")
require.True(t, loginEventsExists(t), "login_events must be restored after the up migration")
}
@@ -91,6 +93,7 @@ func TestMigration011AutoSummarizeDefaultUpDown(t *testing.T) {
require.Equal(t, "true", autoSummarizeDefault(t), "011 sets the default to TRUE")
m := fileMigrator(t)
require.NoError(t, m.Steps(-1), "down 016 drops channel_caption_state")
require.NoError(t, m.Steps(-1), "down 015 reshapes transcripts")
require.NoError(t, m.Steps(-1), "down 014 drops channel_title")
require.NoError(t, m.Steps(-1), "down 013 drops channel_errors")
@@ -104,6 +107,7 @@ func TestMigration011AutoSummarizeDefaultUpDown(t *testing.T) {
require.NoError(t, m.Steps(1), "up 013 creates channel_errors")
require.NoError(t, m.Steps(1), "up 014 recreates channel_title")
require.NoError(t, m.Steps(1), "up 015 reshapes transcripts to shared")
require.NoError(t, m.Steps(1), "up 016 recreates channel_caption_state (HEAD)")
}
// channelTitleExists reports whether videos.channel_title is present.
@@ -123,6 +127,8 @@ func TestMigration014VideoChannelTitleUpDown(t *testing.T) {
require.True(t, channelTitleExists(t), "channel_title exists at latest migration")
m := fileMigrator(t)
require.NoError(t, m.Steps(-1), "down 016 drops channel_caption_state, channel_title intact")
require.True(t, channelTitleExists(t), "016 down leaves channel_title intact")
require.NoError(t, m.Steps(-1), "down 015 reshapes transcripts, channel_title intact")
require.True(t, channelTitleExists(t), "015 down leaves channel_title intact")
require.NoError(t, m.Steps(-1), "down 014 must drop channel_title")
@@ -130,7 +136,8 @@ func TestMigration014VideoChannelTitleUpDown(t *testing.T) {
require.NoError(t, m.Steps(1), "up 014 must recreate channel_title")
require.True(t, channelTitleExists(t), "channel_title must be restored after the up migration")
require.NoError(t, m.Steps(1), "up 015 restores the shared transcripts shape (HEAD)")
require.NoError(t, m.Steps(1), "up 015 restores the shared transcripts shape")
require.NoError(t, m.Steps(1), "up 016 recreates channel_caption_state (HEAD)")
}
// TestMigration012FixAutoSummarizeRLS proves 012 runs cleanly and flips any
@@ -0,0 +1 @@
DROP TABLE channel_caption_state;
@@ -0,0 +1,30 @@
-- Migration 016: per-(user, channel) caption-availability memory (ADR-024).
--
-- Some channels never publish English captions (foreign-language news, music,
-- etc.). Each of their new videos still costs ONE rate-limited caption fetch
-- (ADR-014) before resolving to "none" — and on a throttled egress IP that fetch
-- may 429 and churn through the backoff machinery first. This table remembers
-- channels that repeatedly yield no captions so discovery can stop attempting
-- their videos, freeing the scarce fetch budget for channels that do have them.
--
-- consecutive_none counts no-caption outcomes in a row; a successful fetch resets
-- it to 0. Once it crosses the threshold the channel is suppressed until
-- captionless_until, after which one video is re-probed (auto-recovery for a
-- channel that starts adding captions). Per-user + RLS-scoped, consistent with
-- the rest of the user-owned schema (subscriptions are per-user; ADR-012).
CREATE TABLE channel_caption_state (
user_id UUID NOT NULL REFERENCES users(id) ON DELETE CASCADE,
channel_id TEXT NOT NULL,
consecutive_none INT NOT NULL DEFAULT 0,
captionless_until TIMESTAMPTZ,
updated_at TIMESTAMPTZ NOT NULL DEFAULT now(),
PRIMARY KEY (user_id, channel_id)
);
CREATE INDEX idx_channel_caption_state_user_id ON channel_caption_state(user_id);
ALTER TABLE channel_caption_state ENABLE ROW LEVEL SECURITY;
ALTER TABLE channel_caption_state FORCE ROW LEVEL SECURITY;
CREATE POLICY channel_caption_state_isolation ON channel_caption_state
FOR ALL
USING (user_id = current_setting('tapir.current_user_id', true)::uuid);
+144 -40
View File
@@ -1,17 +1,21 @@
// Package summarizer implements ports.Summarizer backed by the copied llm
// package's local-Primary -> BYO-Fallback routing (ADR-004). It is the only
// place content ever leaves the engine toward an AI model, so it is also the
// enforcement point for the local-first guarantee in
// docs/use-cases/ai_routing.feature: a user with no BYO provider configured has
// their content sent to the local stack and nowhere else.
// package's routing (ADR-004, extended by ADR-022). It is the only place content
// ever leaves the engine toward an AI model, so it is also the enforcement point
// for the local-first guarantee in docs/use-cases/ai_routing.feature: endpoints
// are tried in order, locals first, so content only reaches an external model
// after every local endpoint has failed — and never at all when no external
// endpoint is configured.
package summarizer
import (
"bytes"
"context"
"encoding/json"
"errors"
"fmt"
"strings"
"time"
"unicode/utf8"
"gitea.d-ma.be/mathias/tapir/internal/domain"
)
@@ -29,19 +33,41 @@ type Endpoint struct {
Model string // resolved alias, e.g. "iguana/deepseek-r1-14b"
}
// Summarizer routes a transcript through the local endpoint first, then the
// optional BYO endpoint. It owns its routing (rather than delegating to
// llm.Router) so it can record which provider answered and whether the fallback
// was used — information llm.Router collapses away.
// Summarizer routes a transcript through an ordered chain of endpoints, trying
// each in turn until one returns a parseable summary. It owns its routing
// (rather than delegating to llm.Router) so it can record which provider answered
// and whether a fallback was used — information llm.Router collapses away. The
// chain ordering is the local-first guarantee: callers place local endpoints
// first and any external endpoint last, so content only reaches an external model
// after every local endpoint has failed.
type Summarizer struct {
primary Endpoint
fallback *Endpoint // nil => no BYO; primary errors are returned, never sent externally
now func() time.Time
endpoints []Endpoint
maxInputChars int // transcript truncation budget; 0 = no limit
now func() time.Time
}
// New constructs a Summarizer. fallback may be nil (no BYO provider configured).
// New constructs a Summarizer from a primary endpoint and an optional fallback
// (the historical local-Primary -> BYO-Fallback shape, ADR-004). A nil fallback
// means a single-endpoint chain: errors are returned, content never leaves it.
func New(primary Endpoint, fallback *Endpoint) *Summarizer {
return &Summarizer{primary: primary, fallback: fallback, now: time.Now}
eps := []Endpoint{primary}
if fallback != nil {
eps = append(eps, *fallback)
}
return &Summarizer{endpoints: eps, now: time.Now}
}
// NewChain constructs a Summarizer over an ordered endpoint chain (ADR-022).
// endpoints are tried in order; the first to return a parseable summary wins, and
// FallbackUsed is recorded true for any endpoint past the first. maxInputChars
// bounds the transcript text sent to every endpoint (0 = unbounded), so a long
// transcript does not overflow a small-context primary model's window. It panics
// on an empty chain — a wiring bug, not a runtime condition.
func NewChain(endpoints []Endpoint, maxInputChars int) *Summarizer {
if len(endpoints) == 0 {
panic("summarizer: NewChain requires at least one endpoint")
}
return &Summarizer{endpoints: endpoints, maxInputChars: maxInputChars, now: time.Now}
}
const systemPrompt = `You are Tapir, a video-summarization assistant.
@@ -53,31 +79,35 @@ Respond with ONLY a JSON object, no prose and no code fences:
- "takeaways": the actionable conclusions a viewer should leave with.
Output the JSON object and nothing else.`
// Summarize implements ports.Summarizer.
// Summarize implements ports.Summarizer. It walks the endpoint chain in order:
// the first endpoint whose reply parses into a non-empty summary wins. An
// endpoint is considered failed — and the next one tried — when the model call
// errors OR when its reply cannot be parsed (a 200 with malformed JSON or a
// highlights field the model emitted as a bare string). Truncation is applied
// once, up front, so every endpoint sees the same bounded prompt. When the whole
// chain fails, the joined error is returned so the engine queues the work for
// retry and delivers no summary.
func (s *Summarizer) Summarize(ctx context.Context, v domain.Video, t domain.Transcript) (domain.Summary, error) {
if !t.HasText() {
return domain.Summary{}, fmt.Errorf("summarize: transcript for video %s has no text", v.ID)
}
user := buildUserPrompt(v, t)
user := buildUserPrompt(v, t, s.maxInputChars)
// Primary = local stack. Only on its failure is anything sent externally,
// and only when a BYO fallback is configured.
out, err := s.primary.Client.Complete(ctx, systemPrompt, user)
if err == nil {
return s.build(v, s.primary, false, out)
var errs []error
for i, ep := range s.endpoints {
out, err := ep.Client.Complete(ctx, systemPrompt, user)
if err != nil {
errs = append(errs, fmt.Errorf("%s/%s call: %w", ep.Provider, ep.Model, err))
continue
}
sum, perr := s.build(v, ep, i > 0, out)
if perr != nil {
errs = append(errs, fmt.Errorf("%s/%s output: %w", ep.Provider, ep.Model, perr))
continue
}
return sum, nil
}
if s.fallback == nil {
// No BYO: content was sent to the local stack only. Surface the error so
// the engine can queue the work for retry; deliver no summary.
return domain.Summary{}, fmt.Errorf("summarize: local AI failed and no BYO provider configured: %w", err)
}
out, ferr := s.fallback.Client.Complete(ctx, systemPrompt, user)
if ferr != nil {
return domain.Summary{}, fmt.Errorf("summarize: local AI failed: %w; BYO %s failed: %v", err, s.fallback.Provider, ferr)
}
return s.build(v, *s.fallback, true, out)
return domain.Summary{}, fmt.Errorf("summarize: all %d endpoint(s) failed: %w", len(s.endpoints), errors.Join(errs...))
}
func (s *Summarizer) build(v domain.Video, ep Endpoint, fallbackUsed bool, raw string) (domain.Summary, error) {
@@ -89,8 +119,8 @@ func (s *Summarizer) build(v domain.Video, ep Endpoint, fallbackUsed bool, raw s
UserID: v.UserID,
VideoID: v.ID,
Summary: parsed.Summary,
Highlights: parsed.Highlights,
Takeaways: parsed.Takeaways,
Highlights: []string(parsed.Highlights),
Takeaways: []string(parsed.Takeaways),
AIProvider: ep.Provider,
AIModel: ep.Model,
FallbackUsed: fallbackUsed,
@@ -98,20 +128,94 @@ func (s *Summarizer) build(v domain.Video, ep Endpoint, fallbackUsed bool, raw s
}, nil
}
func buildUserPrompt(v domain.Video, t domain.Transcript) string {
func buildUserPrompt(v domain.Video, t domain.Transcript, maxInputChars int) string {
var b strings.Builder
fmt.Fprintf(&b, "Title: %s\n", v.Title)
if v.URL != "" {
fmt.Fprintf(&b, "URL: %s\n", v.URL)
}
fmt.Fprintf(&b, "\nTranscript:\n%s", t.Content)
fmt.Fprintf(&b, "\nTranscript:\n%s", truncate(t.Content, maxInputChars))
return b.String()
}
// truncate caps content to max bytes on a UTF-8 rune boundary, appending a
// marker so the model knows the transcript was cut. A non-positive max (or a
// content already within budget) returns content unchanged. Bounding the input
// keeps a long transcript from overflowing a small-context model's window — the
// production failure mode where koala/phi4-mini's 8k context returned HTTP 400 on
// a 11.6k-token transcript.
func truncate(content string, max int) string {
if max <= 0 || len(content) <= max {
return content
}
cut := max
for cut > 0 && !utf8.RuneStart(content[cut]) {
cut--
}
return content[:cut] + "\n…[transcript truncated to fit the model context]"
}
// flexStrings is a []string that also unmarshals from a single JSON string or a
// JSON array of scalars. Small local models (koala/phi4-mini) sometimes emit
// "highlights": "one point" instead of an array, or mix in a number; rather than
// fail the whole summary on that quirk, coerce to []string. Empty/whitespace
// elements are dropped.
type flexStrings []string
func (f *flexStrings) UnmarshalJSON(b []byte) error {
b = bytes.TrimSpace(b)
if len(b) == 0 || string(b) == "null" {
*f = nil
return nil
}
if b[0] == '[' {
var raw []json.RawMessage
if err := json.Unmarshal(b, &raw); err != nil {
return err
}
out := make([]string, 0, len(raw))
for _, r := range raw {
s, err := rawToString(r)
if err != nil {
return err
}
if strings.TrimSpace(s) != "" {
out = append(out, s)
}
}
*f = out
return nil
}
s, err := rawToString(b)
if err != nil {
return err
}
if strings.TrimSpace(s) == "" {
*f = nil
} else {
*f = flexStrings{s}
}
return nil
}
// rawToString renders a JSON scalar as text: a quoted string is unquoted; any
// other scalar (number, bool) is kept as its literal source so no content is lost.
func rawToString(r json.RawMessage) (string, error) {
r = bytes.TrimSpace(r)
if len(r) > 0 && r[0] == '"' {
var s string
if err := json.Unmarshal(r, &s); err != nil {
return "", err
}
return s, nil
}
return string(r), nil
}
type parsedSummary struct {
Summary string `json:"summary"`
Highlights []string `json:"highlights"`
Takeaways []string `json:"takeaways"`
Summary string `json:"summary"`
Highlights flexStrings `json:"highlights"`
Takeaways flexStrings `json:"takeaways"`
}
// parse extracts the JSON object from a model reply. Thinking models (qwen3,
@@ -129,8 +129,8 @@ func TestSummarize_NoBYO_ContentOnlyLocal(t *testing.T) {
local := &fakeClient{reply: goodReply}
s := New(Endpoint{Client: local, Provider: "local", Model: "iguana/deepseek-r1-14b"}, nil)
if s.fallback != nil {
t.Fatal("no BYO configured but fallback endpoint is non-nil")
if len(s.endpoints) != 1 {
t.Fatalf("no BYO configured but chain has %d endpoints, want 1", len(s.endpoints))
}
for i := 0; i < 3; i++ {
sum, err := s.Summarize(context.Background(), testVideo(), testTranscript())
@@ -176,3 +176,97 @@ func TestParse_EmptySummaryRejected(t *testing.T) {
t.Fatal("want error for empty summary (thinking model returned no content)")
}
}
// parse tolerates a small model emitting "highlights" as a bare string instead
// of an array — the production koala/phi4-mini quirk that errored with
// "cannot unmarshal string into Go struct field ... highlights of type []string".
func TestParse_ToleratesStringHighlights(t *testing.T) {
p, err := parse(`{"summary":"s","highlights":"one big point","takeaways":["a","b"]}`)
if err != nil {
t.Fatalf("parse: %v", err)
}
if len(p.Highlights) != 1 || p.Highlights[0] != "one big point" {
t.Errorf("highlights = %v, want [\"one big point\"]", p.Highlights)
}
if len(p.Takeaways) != 2 {
t.Errorf("takeaways = %v, want 2", p.Takeaways)
}
}
// Chain: an endpoint that returns a 200 with unparseable output is treated as a
// failure, and the next endpoint in the chain is tried. This is the case the old
// primary->fallback shape missed — a parse error short-circuited instead of
// falling back.
func TestSummarize_FallsBackOnMalformedOutput(t *testing.T) {
bad := &fakeClient{reply: `{"summary": not json`}
good := &fakeClient{reply: goodReply}
s := NewChain([]Endpoint{
{Client: bad, Provider: "local", Model: "koala/phi4-mini"},
{Client: good, Provider: "local", Model: "koala/phi4-14b"},
}, 0)
sum, err := s.Summarize(context.Background(), testVideo(), testTranscript())
if err != nil {
t.Fatalf("Summarize: %v", err)
}
if sum.AIModel != "koala/phi4-14b" {
t.Errorf("AIModel = %q, want koala/phi4-14b (fell back past malformed primary)", sum.AIModel)
}
if !sum.FallbackUsed {
t.Error("FallbackUsed = false, want true")
}
if bad.calls != 1 || good.calls != 1 {
t.Errorf("calls: bad=%d good=%d, want 1 and 1", bad.calls, good.calls)
}
}
// Chain: when every endpoint fails, no summary is produced and the joined error
// names each failure so the engine queues the work for retry.
func TestSummarize_ChainAllEndpointsFail(t *testing.T) {
a := &fakeClient{err: errors.New("context overflow")}
b := &fakeClient{reply: "not even json"}
s := NewChain([]Endpoint{
{Client: a, Provider: "local", Model: "m1"},
{Client: b, Provider: "berget", Model: "m2"},
}, 0)
if _, err := s.Summarize(context.Background(), testVideo(), testTranscript()); err == nil {
t.Fatal("want error when all endpoints fail")
}
if a.calls != 1 || b.calls != 1 {
t.Errorf("calls: a=%d b=%d, want 1 and 1", a.calls, b.calls)
}
}
// A transcript longer than the chain's input budget is truncated before it
// reaches any model, so a small-context primary does not overflow its window.
func TestSummarize_TruncatesLongTranscript(t *testing.T) {
local := &fakeClient{reply: goodReply}
const budget = 100
s := NewChain([]Endpoint{{Client: local, Provider: "local", Model: "m"}}, budget)
long := domain.Transcript{
VideoID: "vid-1", UserID: "user-1", Source: domain.SourceCaptions,
Content: strings.Repeat("word ", 1000), // 5000 bytes, well over budget
}
if _, err := s.Summarize(context.Background(), testVideo(), long); err != nil {
t.Fatalf("Summarize: %v", err)
}
// The prompt carries title/URL framing plus the truncation marker, so allow
// headroom over the raw transcript budget — but it must be far below 5000.
if len(local.lastUser) > budget+300 {
t.Errorf("prompt length = %d, want <= %d (transcript not truncated)", len(local.lastUser), budget+300)
}
if !strings.Contains(local.lastUser, "truncated") {
t.Error("truncation marker missing from prompt")
}
}
func TestNewChain_PanicsOnEmptyChain(t *testing.T) {
defer func() {
if recover() == nil {
t.Fatal("want panic on empty endpoint chain")
}
}()
NewChain(nil, 0)
}
+106 -4
View File
@@ -66,6 +66,13 @@ type Config struct {
// poll. Zero means defaultMaxVideos.
MaxVideosPerSubscription int
// MinVideoSeconds drops videos shorter than this from discovery (Shorts/clips,
// ADR-023). NewVideos enriches candidates with a single cheap videos.list call
// (contentDetails.duration + snippet.liveBroadcastContent) and filters before
// returning, so the scarce caption-fetch budget is never spent on them. Live
// and upcoming broadcasts are dropped too. Zero disables the filter.
MinVideoSeconds int
// BaseURL overrides the Data API root. Empty means defaultBaseURL.
BaseURL string
@@ -242,7 +249,97 @@ func (a *Adapter) NewVideos(ctx context.Context, sub domain.Subscription) ([]dom
break
}
}
return videos, nil
// Drop Shorts/sub-minute clips and live/upcoming broadcasts before they ever
// reach the rate-limited caption path (ADR-023). One cheap videos.list call
// (quota API, not the timedtext throttle) supplies duration + live status.
return a.filterLowValue(ctx, client, videos), nil
}
// filterLowValue removes videos shorter than cfg.MinVideoSeconds and any live or
// upcoming broadcast, using a single videos.list lookup for duration +
// liveBroadcastContent. The filter is best-effort: if MinVideoSeconds is 0 (off)
// or the lookup fails, the input is returned unfiltered — discovery must not break
// because a metadata call hiccuped; the worst case is the pre-ADR-023 behaviour.
func (a *Adapter) filterLowValue(ctx context.Context, client *http.Client, videos []domain.Video) []domain.Video {
if a.cfg.MinVideoSeconds <= 0 || len(videos) == 0 {
return videos
}
ids := make([]string, 0, len(videos))
for _, v := range videos {
ids = append(ids, v.ProviderVideoID)
}
q := url.Values{
"part": {"contentDetails,snippet"},
"id": {strings.Join(ids, ",")},
}
var resp videoListResponse
if err := a.getJSON(ctx, client, "/videos", q, &resp); err != nil {
// Degrade open: keep the candidates rather than lose discovery.
return videos
}
type meta struct {
seconds int
live string
}
byID := make(map[string]meta, len(resp.Items))
for _, it := range resp.Items {
byID[it.ID] = meta{seconds: parseISO8601Seconds(it.ContentDetails.Duration), live: it.Snippet.LiveBroadcastContent}
}
kept := videos[:0]
for _, v := range videos {
m, ok := byID[v.ProviderVideoID]
if !ok {
kept = append(kept, v) // unknown metadata: keep, let the fetch decide
continue
}
if m.live != "" && m.live != "none" {
continue // live or upcoming broadcast
}
if m.seconds > 0 && m.seconds < a.cfg.MinVideoSeconds {
continue // Short / sub-threshold clip
}
kept = append(kept, v)
}
return kept
}
// parseISO8601Seconds parses an ISO 8601 duration as returned by the YouTube Data
// API (e.g. "PT1H2M3S", "PT45S", "PT3M") into seconds. Only the hour/minute/second
// components YouTube emits are handled; an unparseable or zero value returns 0,
// which the caller treats as "unknown" (not filtered on duration).
func parseISO8601Seconds(d string) int {
if !strings.HasPrefix(d, "PT") {
return 0
}
d = d[2:]
total, num := 0, 0
seen := false
for _, r := range d {
switch {
case r >= '0' && r <= '9':
num = num*10 + int(r-'0')
seen = true
case r == 'H':
total += num * 3600
num, seen = 0, false
case r == 'M':
total += num * 60
num, seen = 0, false
case r == 'S':
total += num
num, seen = 0, false
default:
return 0 // unexpected component (days/weeks) — treat as unknown
}
}
if seen {
return 0 // trailing digits without a unit: malformed
}
return total
}
// VideoByID fetches a single video's metadata (videos.list, snippet) for an
@@ -377,11 +474,16 @@ type playlistItemListResponse struct {
type videoListResponse struct {
Items []struct {
ID string `json:"id"`
Snippet struct {
Title string `json:"title"`
ChannelTitle string `json:"channelTitle"`
PublishedAt time.Time `json:"publishedAt"`
Title string `json:"title"`
ChannelTitle string `json:"channelTitle"`
PublishedAt time.Time `json:"publishedAt"`
LiveBroadcastContent string `json:"liveBroadcastContent"`
} `json:"snippet"`
ContentDetails struct {
Duration string `json:"duration"` // ISO 8601, e.g. "PT1M30S"
} `json:"contentDetails"`
} `json:"items"`
}
+85
View File
@@ -186,6 +186,91 @@ func TestNewVideosCapsAtMax(t *testing.T) {
}
}
// TestNewVideosFiltersShortsAndLive: with MinVideoSeconds set, discovery enriches
// candidates via videos.list and drops sub-threshold clips (Shorts) and
// live/upcoming broadcasts before they reach the rate-limited caption path.
func TestNewVideosFiltersShortsAndLive(t *testing.T) {
a, _ := newTestAdapter(t, func(w http.ResponseWriter, r *http.Request) {
switch r.URL.Path {
case "/playlistItems":
_, _ = w.Write([]byte(`{
"items": [
{"snippet": {"title": "Real Talk", "publishedAt": "2026-06-03T10:00:00Z", "resourceId": {"videoId": "long1"}}},
{"snippet": {"title": "A Short", "publishedAt": "2026-06-03T09:00:00Z", "resourceId": {"videoId": "short1"}}},
{"snippet": {"title": "Live Now", "publishedAt": "2026-06-03T08:00:00Z", "resourceId": {"videoId": "live1"}}}
]
}`))
case "/videos":
if got := r.URL.Query().Get("part"); got != "contentDetails,snippet" {
t.Errorf("videos.list part=%q, want contentDetails,snippet", got)
}
_, _ = w.Write([]byte(`{
"items": [
{"id": "long1", "contentDetails": {"duration": "PT12M30S"}, "snippet": {"liveBroadcastContent": "none"}},
{"id": "short1", "contentDetails": {"duration": "PT45S"}, "snippet": {"liveBroadcastContent": "none"}},
{"id": "live1", "contentDetails": {"duration": "PT0S"}, "snippet": {"liveBroadcastContent": "live"}}
]
}`))
default:
t.Errorf("unexpected path %q", r.URL.Path)
}
})
a.cfg.MinVideoSeconds = 60
vids, err := a.NewVideos(context.Background(), domain.Subscription{ID: "s1", UserID: "u1", ChannelID: "UC_acme"})
if err != nil {
t.Fatalf("NewVideos: %v", err)
}
if len(vids) != 1 || vids[0].ProviderVideoID != "long1" {
t.Fatalf("expected only long1 to survive the filter, got %+v", vids)
}
}
// TestNewVideosNoFilterWhenDisabled: MinVideoSeconds=0 keeps the pre-ADR-023
// behaviour — no videos.list call, no filtering.
func TestNewVideosNoFilterWhenDisabled(t *testing.T) {
a, _ := newTestAdapter(t, func(w http.ResponseWriter, r *http.Request) {
if r.URL.Path == "/videos" {
t.Errorf("videos.list must not be called when MinVideoSeconds is 0")
}
_, _ = w.Write([]byte(`{"items": [
{"snippet": {"title": "A Short", "publishedAt": "2026-06-03T09:00:00Z", "resourceId": {"videoId": "short1"}}}
]}`))
})
a.cfg.MinVideoSeconds = 0
vids, err := a.NewVideos(context.Background(), domain.Subscription{ID: "s1", UserID: "u1", ChannelID: "UC_acme"})
if err != nil {
t.Fatalf("NewVideos: %v", err)
}
if len(vids) != 1 {
t.Fatalf("filter disabled must keep all videos, got %d", len(vids))
}
}
func TestParseISO8601Seconds(t *testing.T) {
cases := []struct {
in string
want int
}{
{"PT45S", 45},
{"PT1M30S", 90},
{"PT3M", 180},
{"PT1H2M3S", 3723},
{"PT2H", 7200},
{"PT0S", 0},
{"", 0},
{"garbage", 0},
{"P1D", 0}, // days component not handled → unknown
{"PT10", 0}, // trailing digits without a unit → malformed
}
for _, c := range cases {
if got := parseISO8601Seconds(c.in); got != c.want {
t.Errorf("parseISO8601Seconds(%q) = %d, want %d", c.in, got, c.want)
}
}
}
// TestUploadsPlaylistID covers the zero-cost UC->UU derivation, including
// non-standard ids that must fall through unchanged (handled via fallback).
func TestUploadsPlaylistID(t *testing.T) {
+90 -1
View File
@@ -28,8 +28,38 @@ type Config struct {
GatewayURL string
// GatewayKey authorizes the gateway. Read from env, never committed.
GatewayKey string
// SummarizerModel is the alias in host/name form, e.g. "koala/phi4-mini".
// SummarizerModel is the primary summarizer alias in host/name form, tried
// first on every video, e.g. "koala/phi4-mini".
SummarizerModel string
// FallbackModel is the LOCAL fallback alias tried when the primary fails or
// returns unparseable output (ADR-022). Kept local so content stays on the
// homelab stack. Empty disables it. Default a bigger-context local model.
FallbackModel string
// CloudFallbackModel is the worst-case EXTERNAL fallback alias, tried only
// after every local endpoint has failed (ADR-022). For client deployments set
// this empty so content never leaves the local stack. Default a berget alias.
CloudFallbackModel string
// SummaryMaxTokens caps the completion budget per summary call. Small-context
// models (koala/phi4-mini, 8k) overflow when prompt + max_tokens exceeds the
// window; a summary needs only a few hundred tokens, so the default is small.
SummaryMaxTokens int
// MaxTranscriptChars bounds the transcript text sent to the model so a long
// transcript does not overflow a small-context primary. 0 disables truncation.
MaxTranscriptChars int
// MinVideoSeconds drops videos shorter than this from discovery (Shorts and
// other sub-minute clips that are noise and waste the scarce caption-fetch
// budget, ADR-014/ADR-023). Enforced via a cheap Data API videos.list lookup at
// discovery, never the rate-limited caption path. 0 disables the filter.
MinVideoSeconds int
// ChannelCaptionlessThreshold is how many consecutive no-caption results a
// channel may yield before its videos are suppressed from caption fetching
// (ADR-024). 0 disables the per-channel caption memory entirely.
ChannelCaptionlessThreshold int
// ChannelCaptionlessWindow is how long a suppressed channel stays suppressed
// before one video is re-probed (auto-recovery for a channel that adds captions).
ChannelCaptionlessWindow time.Duration
// SummarizerTimeout bounds a single completion call. Thinking models are
// slow, so the default is generous.
SummarizerTimeout time.Duration
@@ -118,6 +148,13 @@ func (c Config) DexConfigured() bool { return strings.TrimSpace(c.OIDCIssuer) !=
const (
defaultGatewayURL = "http://koala:30401/v1"
defaultSummarizerModel = "koala/phi4-mini"
defaultFallbackModel = "koala/phi4-14b"
defaultCloudFallbackModel = "berget/mistral-small"
defaultSummaryMaxTokens = 1500
defaultMaxTranscriptChars = 18000
defaultMinVideoSeconds = 60
defaultCaptionlessThreshold = 5
defaultCaptionlessWindow = 14 * 24 * time.Hour
defaultSummarizerTimeout = 5 * time.Minute
defaultYTTokenRef = "youtube/refresh_token"
defaultYTConnectRedirectURL = "https://tapir.d-ma.be/oauth/youtube/callback"
@@ -141,6 +178,8 @@ func Load() (Config, error) {
GatewayURL: envOr("TAPIR_GATEWAY_URL", defaultGatewayURL),
GatewayKey: os.Getenv("TAPIR_GATEWAY_KEY"),
SummarizerModel: envOr("TAPIR_SUMMARIZER_MODEL", defaultSummarizerModel),
FallbackModel: lookupOr("TAPIR_FALLBACK_MODEL", defaultFallbackModel),
CloudFallbackModel: lookupOr("TAPIR_CLOUD_FALLBACK_MODEL", defaultCloudFallbackModel),
DBDSN: os.Getenv("TAPIR_DB_DSN"),
YTClientID: os.Getenv("TAPIR_YT_CLIENT_ID"),
YTClientSecret: os.Getenv("TAPIR_YT_CLIENT_SECRET"),
@@ -193,6 +232,45 @@ func Load() (Config, error) {
}
c.AutoSummarizeWindow = autoWindow
summaryTokens, err := intOr("TAPIR_SUMMARY_MAX_TOKENS", defaultSummaryMaxTokens)
if err != nil {
return Config{}, err
}
c.SummaryMaxTokens = summaryTokens
maxChars, err := intOr("TAPIR_MAX_TRANSCRIPT_CHARS", defaultMaxTranscriptChars)
if err != nil {
return Config{}, err
}
if maxChars < 0 {
maxChars = 0
}
c.MaxTranscriptChars = maxChars
minVideo, err := intOr("TAPIR_MIN_VIDEO_SECONDS", defaultMinVideoSeconds)
if err != nil {
return Config{}, err
}
if minVideo < 0 {
minVideo = 0
}
c.MinVideoSeconds = minVideo
captionThreshold, err := intOr("TAPIR_CHANNEL_CAPTIONLESS_THRESHOLD", defaultCaptionlessThreshold)
if err != nil {
return Config{}, err
}
if captionThreshold < 0 {
captionThreshold = 0
}
c.ChannelCaptionlessThreshold = captionThreshold
captionWindow, err := durationOr("TAPIR_CHANNEL_CAPTIONLESS_WINDOW", defaultCaptionlessWindow)
if err != nil {
return Config{}, err
}
c.ChannelCaptionlessWindow = captionWindow
onboard, err := intOr("TAPIR_ONBOARD_SUMMARIZE_COUNT", defaultOnboardSummarizeCount)
if err != nil {
return Config{}, err
@@ -269,6 +347,17 @@ func envOr(key, fallback string) string {
return fallback
}
// lookupOr returns the env value when the key is PRESENT (even if empty), else
// fallback. Unlike envOr it lets an explicit empty value override the default —
// needed to DISABLE an optional fallback model (e.g. set the cloud fallback empty
// for a client deployment so content never leaves the local stack).
func lookupOr(key, fallback string) string {
if v, ok := os.LookupEnv(key); ok {
return v
}
return fallback
}
func intOr(key string, fallback int) (int, error) {
v := os.Getenv(key)
if v == "" {
+58
View File
@@ -1,11 +1,29 @@
package config
import (
"os"
"strings"
"testing"
"time"
)
// unset removes an env key for the duration of the test, restoring it after.
// Needed to observe a default for a key read with LookupEnv (where present-empty
// means "explicitly disabled", not "use default").
func unset(t *testing.T, key string) {
t.Helper()
if old, ok := os.LookupEnv(key); ok {
t.Cleanup(func() {
if err := os.Setenv(key, old); err != nil {
t.Fatalf("restore %s: %v", key, err)
}
})
}
if err := os.Unsetenv(key); err != nil {
t.Fatalf("unset %s: %v", key, err)
}
}
// setEnv sets env vars for the test and clears them afterward, so cases don't
// leak into one another. t.Setenv handles restoration.
func setEnv(t *testing.T, kv map[string]string) {
@@ -53,6 +71,46 @@ func TestLoad_AppliesDefaults(t *testing.T) {
}
}
func TestLoad_SummarizerChainDefaults(t *testing.T) {
setEnv(t, map[string]string{
"TAPIR_SUMMARIZER_MODEL": "",
"TAPIR_SUMMARY_MAX_TOKENS": "",
"TAPIR_MAX_TRANSCRIPT_CHARS": "",
})
unset(t, "TAPIR_FALLBACK_MODEL")
unset(t, "TAPIR_CLOUD_FALLBACK_MODEL")
c, err := Load()
if err != nil {
t.Fatalf("Load: %v", err)
}
if c.FallbackModel != defaultFallbackModel {
t.Errorf("FallbackModel = %q, want %q", c.FallbackModel, defaultFallbackModel)
}
if c.CloudFallbackModel != defaultCloudFallbackModel {
t.Errorf("CloudFallbackModel = %q, want %q", c.CloudFallbackModel, defaultCloudFallbackModel)
}
if c.SummaryMaxTokens != defaultSummaryMaxTokens {
t.Errorf("SummaryMaxTokens = %d, want %d", c.SummaryMaxTokens, defaultSummaryMaxTokens)
}
if c.MaxTranscriptChars != defaultMaxTranscriptChars {
t.Errorf("MaxTranscriptChars = %d, want %d", c.MaxTranscriptChars, defaultMaxTranscriptChars)
}
}
// An explicitly empty cloud-fallback env disables external routing — the lever a
// client deployment pulls so content never leaves the local stack.
func TestLoad_EmptyCloudFallbackDisables(t *testing.T) {
t.Setenv("TAPIR_CLOUD_FALLBACK_MODEL", "")
c, err := Load()
if err != nil {
t.Fatalf("Load: %v", err)
}
if c.CloudFallbackModel != "" {
t.Errorf("CloudFallbackModel = %q, want empty (disabled)", c.CloudFallbackModel)
}
}
func TestLoad_OnboardSummarizeCount(t *testing.T) {
cases := []struct {
name, env string
+63 -12
View File
@@ -27,8 +27,9 @@ import (
// passCandidate is a video that passed all pre-filters (seen/manual/backoff)
// and is queued for transcript fetch + summarization in this pass.
type passCandidate struct {
v domain.Video
pos int // discovery position — used as a stable tiebreak when published_at ties
v domain.Video
channelID string // owning channel — keys the caption-availability memory (ADR-024)
pos int // discovery position — used as a stable tiebreak when published_at ties
}
// compareNewestFirst orders candidates by published_at descending, NULLS LAST,
@@ -74,6 +75,14 @@ type VideoStore interface {
// Called when NewVideos returns domain.ErrChannelUnavailable; best-effort, errors
// are logged and never abort the pass.
UpsertChannelError(ctx context.Context, userID, channelID, channelTitle string) error
// CaptionlessChannels returns channel ids currently suppressed because their
// recent videos all yielded no captions (ADR-024). The loop skips caption
// fetches for these channels' (non-requested) videos.
CaptionlessChannels(ctx context.Context, userID string) (map[string]bool, error)
// RecordChannelCaptionOutcome updates a channel's caption memory after a fetch:
// hadCaptions resets it, otherwise the no-caption streak grows and the channel
// is suppressed for window once it reaches threshold. A no-op when threshold<=0.
RecordChannelCaptionOutcome(ctx context.Context, userID, channelID string, hadCaptions bool, threshold int, window time.Duration) error
}
// Processor runs the core use case for a single video. *usecase.Engine
@@ -93,6 +102,9 @@ type Runner struct {
backoff time.Duration // rate-limit retry window; 0 = always retry
autoWindow time.Duration // recency bound for auto-summarize; 0 = no bound
now func() time.Time // injectable clock (tests); defaults to time.Now
captionThreshold int // consecutive no-caption results before a channel is suppressed; 0 = feature off
captionWindow time.Duration // how long a caption-less channel stays suppressed before re-probe
}
// Option configures a Runner at construction. Variadic so existing call sites
@@ -114,6 +126,14 @@ func WithClock(now func() time.Time) Option { return func(r *Runner) { r.now = n
// bypasses the bound. 0 (the default) disables it (summarize every unseen video).
func WithAutoWindow(d time.Duration) Option { return func(r *Runner) { r.autoWindow = d } }
// WithCaptionMemory enables per-channel caption-availability suppression
// (ADR-024): after threshold consecutive no-caption results a channel's videos
// are skipped (no caption fetch) for window, then one is re-probed. threshold<=0
// (the default) disables the feature entirely.
func WithCaptionMemory(threshold int, window time.Duration) Option {
return func(r *Runner) { r.captionThreshold = threshold; r.captionWindow = window }
}
// New builds a Runner. A nil logger falls back to slog.Default.
func New(src ports.VideoSource, store VideoStore, engine Processor, userID string, log *slog.Logger, opts ...Option) *Runner {
if log == nil {
@@ -131,15 +151,16 @@ func New(src ports.VideoSource, store VideoStore, engine Processor, userID strin
// Stats summarizes one RunOnce pass.
type Stats struct {
Candidates int
Summarized int
SkippedSeen int
SkippedNoText int
SkippedManual int // discovered but not queued, in manual mode
SkippedTooOld int // auto mode: published outside the recency window (not requested)
SkippedRateLimited int // 429'd previously and still inside the backoff window
Errors int
ChannelUnavailable int // channels that returned HTTP 404 (deleted/private)
Candidates int
Summarized int
SkippedSeen int
SkippedNoText int
SkippedManual int // discovered but not queued, in manual mode
SkippedTooOld int // auto mode: published outside the recency window (not requested)
SkippedRateLimited int // 429'd previously and still inside the backoff window
SkippedNoCaptionChannel int // channel suppressed as caption-less (ADR-024)
Errors int
ChannelUnavailable int // channels that returned HTTP 404 (deleted/private)
}
// tooOld reports whether a video published at publishedAt falls outside the
@@ -213,6 +234,17 @@ func (r *Runner) RunOnce(ctx context.Context) (Stats, error) {
}
}
// Per-channel caption memory (ADR-024): channels whose recent videos all
// yielded no captions are suppressed so their new videos don't burn the scarce
// fetch budget. Loaded only when the feature is enabled (threshold > 0).
var captionless map[string]bool
if r.captionThreshold > 0 {
captionless, err = r.store.CaptionlessChannels(ctx, r.userID)
if err != nil {
return stats, fmt.Errorf("runner: load caption-less channels: %w", err)
}
}
subs, err := r.src.ListSubscriptions(ctx, r.userID)
if err != nil {
return stats, fmt.Errorf("runner: list subscriptions: %w", err)
@@ -273,6 +305,14 @@ func (r *Runner) RunOnce(ctx context.Context) (Stats, error) {
continue
}
// Caption-less channel (ADR-024): its recent videos all returned no
// captions, so skip the fetch entirely. The video is still listed
// (UpsertVideo above); an explicit manual request bypasses the skip.
if !requested[id] && captionless[sub.ChannelID] {
stats.SkippedNoCaptionChannel++
continue
}
// Still inside the rate-limit backoff window: skip without fetching.
if at, ok := rateLimited[id]; ok && r.now().Sub(at) < r.backoff {
stats.SkippedRateLimited++
@@ -280,7 +320,7 @@ func (r *Runner) RunOnce(ctx context.Context) (Stats, error) {
continue
}
candidates = append(candidates, passCandidate{v: v, pos: pos})
candidates = append(candidates, passCandidate{v: v, channelID: sub.ChannelID, pos: pos})
pos++
}
}
@@ -320,6 +360,11 @@ func (r *Runner) RunOnce(ctx context.Context) (Stats, error) {
errs = append(errs, fmt.Errorf("set none status %q: %w", c.v.ProviderVideoID, err))
stats.Errors++
}
// No captions: grow this channel's no-caption streak (ADR-024).
if err := r.store.RecordChannelCaptionOutcome(ctx, r.userID, c.channelID, false, r.captionThreshold, r.captionWindow); err != nil {
errs = append(errs, fmt.Errorf("record no-caption %q: %w", c.v.ProviderVideoID, err))
stats.Errors++
}
r.log.Info("skipped video (no transcript)", "video", c.v.ProviderVideoID, "title", c.v.Title)
case res.Summary != nil:
stats.Summarized++
@@ -327,6 +372,11 @@ func (r *Runner) RunOnce(ctx context.Context) (Stats, error) {
errs = append(errs, fmt.Errorf("set fetched status %q: %w", c.v.ProviderVideoID, err))
stats.Errors++
}
// Captions present: reset this channel's caption memory (ADR-024).
if err := r.store.RecordChannelCaptionOutcome(ctx, r.userID, c.channelID, true, r.captionThreshold, r.captionWindow); err != nil {
errs = append(errs, fmt.Errorf("record has-caption %q: %w", c.v.ProviderVideoID, err))
stats.Errors++
}
// In manual mode the video was explicitly queued; clear the flag so
// it is not re-summarized and the UI drops the "Queued" chip.
if !auto {
@@ -354,6 +404,7 @@ func (r *Runner) Loop(ctx context.Context, interval time.Duration) error {
"skipped_seen", stats.SkippedSeen, "skipped_no_text", stats.SkippedNoText,
"skipped_manual", stats.SkippedManual, "skipped_too_old", stats.SkippedTooOld,
"skipped_rate_limited", stats.SkippedRateLimited,
"skipped_no_caption_channel", stats.SkippedNoCaptionChannel,
"channel_unavailable", stats.ChannelUnavailable, "errors", stats.Errors)
if err != nil {
r.log.Warn("run pass had errors", "err", err)
+72
View File
@@ -51,6 +51,13 @@ type fakeStore struct {
cleared []string
rateLimited map[string]time.Time // id -> when 429'd (seeds the backoff window)
statuses map[string]string // id -> last SetTranscriptStatus value
captionless map[string]bool // channel ids currently suppressed (ADR-024)
captionRecs []captionRec // RecordChannelCaptionOutcome calls, in order
}
type captionRec struct {
channelID string
had bool
}
func (f *fakeStore) UpsertVideo(_ context.Context, v domain.Video) (string, error) {
@@ -93,6 +100,22 @@ func (f *fakeStore) RateLimitedVideoIDs(_ context.Context, _ string) (map[string
func (f *fakeStore) UpsertChannelError(_ context.Context, _, _, _ string) error { return nil }
func (f *fakeStore) CaptionlessChannels(_ context.Context, _ string) (map[string]bool, error) {
cp := make(map[string]bool, len(f.captionless))
for k, v := range f.captionless {
cp[k] = v
}
return cp, nil
}
func (f *fakeStore) RecordChannelCaptionOutcome(_ context.Context, _, channelID string, hadCaptions bool, threshold int, _ time.Duration) error {
if threshold <= 0 {
return nil
}
f.captionRecs = append(f.captionRecs, captionRec{channelID: channelID, had: hadCaptions})
return nil
}
func (f *fakeStore) SetTranscriptStatus(_ context.Context, _, videoID, status string) error {
if f.statuses == nil {
f.statuses = map[string]string{}
@@ -284,6 +307,55 @@ func TestRunOnce_AutoMode_OldVideoRequestedBypassesWindow(t *testing.T) {
require.Len(t, sink.delivered, 1)
}
// TestRunOnce_CaptionlessChannelSkipped: a channel flagged caption-less (ADR-024)
// has its videos skipped from fetching but still discovered/listed, while a
// normal channel's video is summarized.
func TestRunOnce_CaptionlessChannelSkipped(t *testing.T) {
src := &fakeSource{
subs: []domain.Subscription{sub("dead", "Dead Channel"), sub("live", "Live Channel")},
videos: map[string][]domain.Video{
"dead": {vid("d1", "Dead One")},
"live": {vid("l1", "Live One")},
},
}
st := &fakeStore{seen: map[string]bool{}, auto: true, captionless: map[string]bool{"dead": true}}
sink := &recordingSink{}
eng := usecase.NewEngine(src, fakeSummarizer{}, sink)
r := runner.New(src, st, eng, testUser, quietLogger(),
runner.WithCaptionMemory(5, 14*24*time.Hour))
stats, err := r.RunOnce(context.Background())
require.NoError(t, err)
require.Equal(t, 1, stats.SkippedNoCaptionChannel, "dead channel's video skipped from fetch")
require.Equal(t, 1, stats.Summarized, "live channel's video still summarized")
require.Len(t, st.upserted, 2, "both videos are still discovered and listed")
}
// TestRunOnce_RecordsCaptionOutcomes: a no-caption result grows the channel's
// streak (had=false); a successful summary resets it (had=true).
func TestRunOnce_RecordsCaptionOutcomes(t *testing.T) {
src := &fakeSource{
subs: []domain.Subscription{sub("c1", "Has Caps"), sub("c2", "No Caps")},
videos: map[string][]domain.Video{
"c1": {vid("good", "Good")},
"c2": {vid("bad", "Bad")},
},
transcripts: map[string]domain.Transcript{
"bad": {Source: domain.SourceNone}, // no usable text → engine skips
},
}
st := &fakeStore{seen: map[string]bool{}, auto: true}
sink := &recordingSink{}
eng := usecase.NewEngine(src, fakeSummarizer{}, sink)
r := runner.New(src, st, eng, testUser, quietLogger(),
runner.WithCaptionMemory(5, 14*24*time.Hour))
_, err := r.RunOnce(context.Background())
require.NoError(t, err)
require.Contains(t, st.captionRecs, captionRec{channelID: "c1", had: true}, "captioned channel reset")
require.Contains(t, st.captionRecs, captionRec{channelID: "c2", had: false}, "no-caption channel streak grown")
}
// TestRunOnce_AutoWindowZero_SummarizesOld: a zero window disables the bound —
// the pre-recency behaviour (summarize every unseen video) is preserved.
func TestRunOnce_AutoWindowZero_SummarizesOld(t *testing.T) {
+12 -3
View File
@@ -233,11 +233,20 @@ func (a *App) handleList(w http.ResponseWriter, r *http.Request) {
}
hasConnected := len(conns) > 0
if isHTMX(r) {
a.render(w, r, summaryList(buckets, hasConnected))
// Summarization mode drives the backlog copy: an auto user is told summaries
// land gradually; a manual user is told to click Summarize (the first pilot
// user sat in manual mode reading "land automatically" and waited forever).
autoSummarize, err := a.Store.GetAutoSummarize(r.Context(), userID)
if err != nil {
a.serverError(w, r, "summarize mode", err)
return
}
a.render(w, r, ListPage(buckets, f, stats, takeFlash(w, r), hasConnected, channels))
if isHTMX(r) {
a.render(w, r, summaryList(buckets, hasConnected, autoSummarize))
return
}
a.render(w, r, ListPage(buckets, f, stats, takeFlash(w, r), hasConnected, channels, autoSummarize))
}
// handleDetail renders one summary in full (highlights, takeaways, action group).
+33
View File
@@ -477,3 +477,36 @@ func postAction(t *testing.T, app *web.App, videoID, action string, htmx bool) *
}
return do(t, app, req)
}
// TestListManualModeBannerCopy: a manual-mode user with un-summarized videos
// sees the manual prompt (click Summarize), NOT the "summaries land
// automatically" copy that misled the first pilot user into waiting forever.
func TestListManualModeBannerCopy(t *testing.T) {
ctx := context.Background()
app := newApp(t)
p := rawPool(t)
resetDB(t, p)
seedVideo(t, p, videoX, "X Title", "https://x", time.Time{}) // pending, un-summarized
require.NoError(t, app.Store.SetAutoSummarize(ctx, userID, false))
html := body(t, do(t, app, httptest.NewRequest(http.MethodGet, "/", nil)))
require.Contains(t, html, "Manual mode")
require.Contains(t, html, "are not summarized automatically")
require.NotContains(t, html, "land gradually",
"manual-mode user must not be told summaries arrive automatically")
}
// TestListAutoModeBannerCopy: an auto-mode user with a backlog sees the
// gradual-delivery copy, not the manual prompt.
func TestListAutoModeBannerCopy(t *testing.T) {
ctx := context.Background()
app := newApp(t)
p := rawPool(t)
resetDB(t, p)
seedVideo(t, p, videoX, "X Title", "https://x", time.Time{})
require.NoError(t, app.Store.SetAutoSummarize(ctx, userID, true))
html := body(t, do(t, app, httptest.NewRequest(http.MethodGet, "/", nil)))
require.Contains(t, html, "land gradually")
require.NotContains(t, html, "are not summarized automatically")
}
+1 -1
View File
@@ -65,7 +65,7 @@ func TestParseYouTubeVideoID(t *testing.T) {
func TestListPageShowsPasteFormOnlyWhenConnected(t *testing.T) {
render := func(connected bool) string {
var buf bytes.Buffer
if err := ListPage(listBuckets{}, Filter{}, PipelineStats{}, "", connected, nil).Render(context.Background(), &buf); err != nil {
if err := ListPage(listBuckets{}, Filter{}, PipelineStats{}, "", connected, nil, true).Render(context.Background(), &buf); err != nil {
t.Fatalf("render: %v", err)
}
return buf.String()
+18 -5
View File
@@ -104,7 +104,7 @@ templ flashBanner(code string) {
// #summary-list region; a non-HTMX request renders the whole page. flash carries
// a one-shot notification (e.g. "connected", "registered") surfaced on arrival
// after a POST→redirect.
templ ListPage(b listBuckets, f Filter, stats PipelineStats, flash string, hasConnected bool, channels []string) {
templ ListPage(b listBuckets, f Filter, stats PipelineStats, flash string, hasConnected bool, channels []string, autoSummarize bool) {
@Layout("Tapir — Summaries") {
@flashBanner(flash)
if hasConnected {
@@ -116,14 +116,23 @@ templ ListPage(b listBuckets, f Filter, stats PipelineStats, flash string, hasCo
if stats.RateLimited > 0 || stats.Pending > 0 || stats.NoText > 0 {
@pipelineBar(stats)
}
if stats.RateLimited+stats.Pending > 0 {
if (stats.RateLimited+stats.Pending) > 0 && autoSummarize {
<p class="pipeline-note muted">
Tapir fetches captions slowly on purpose, to respect YouTube's limits
new summaries land gradually. Check back tomorrow.
</p>
}
if (stats.RateLimited+stats.Pending) > 0 && !autoSummarize {
<p class="pipeline-note muted">
You are in Manual mode: new videos appear here but are not summarized
automatically. Use the Summarize button on the ones you want.
</p>
<p class="pipeline-note muted">
<a href="/account">Switch to Automatic</a> to have new videos summarized for you.
</p>
}
<div id="summary-list">
@summaryList(b, hasConnected)
@summaryList(b, hasConnected, autoSummarize)
</div>
}
}
@@ -203,12 +212,16 @@ templ filterForm(f Filter, channels []string) {
// and a single disclosure holding the older un-summarized back-catalogue. Cards
// reflow to a single column on mobile; an empty list shows a friendly first-run
// state instead of a blank table.
templ summaryList(b listBuckets, hasConnected bool) {
templ summaryList(b listBuckets, hasConnected bool, autoSummarize bool) {
if b.empty() {
if hasConnected {
<div class="empty empty-connected">
<strong>Your account is connected</strong>
<span>Tapir is finding your subscriptions and fetching captions summaries appear here gradually. Check back later.</span>
if autoSummarize {
<span>Tapir is finding your subscriptions and fetching captions summaries appear here gradually. Check back later.</span>
} else {
<span>Tapir is finding your subscriptions. You are in Manual mode, so videos appear here with a Summarize button pick the ones you want, or switch to Automatic in your account.</span>
}
</div>
} else {
<div class="empty">
File diff suppressed because it is too large Load Diff
+5 -4
View File
@@ -26,10 +26,11 @@ import (
// longer matches a real non-pending scenario.
var scenarioCoverage = map[string]string{
// ai_routing.feature
"Local AI produces the summary": "TestSummarize_LocalSucceeds",
"Local AI fails and the user has a BYO provider configured": "TestSummarize_FallsBackToBYO",
"Local AI fails and the user has no BYO provider": "TestSummarize_LocalFailsNoBYO_NoExternalSend",
"A user without BYO never has content sent externally": "TestSummarize_NoBYO_ContentOnlyLocal",
"Local AI produces the summary": "TestSummarize_LocalSucceeds",
"Local AI fails and the user has a BYO provider configured": "TestSummarize_FallsBackToBYO",
"Local AI fails and the user has no BYO provider": "TestSummarize_LocalFailsNoBYO_NoExternalSend",
"A user without BYO never has content sent externally": "TestSummarize_NoBYO_ContentOnlyLocal",
"A model returns unparseable output and the next endpoint succeeds": "TestSummarize_FallsBackOnMalformedOutput",
// landing_page.feature
"An unauthenticated visit to the root is sent to the welcome page": "TestUnauthenticatedRootRedirectsToWelcome",