Compare commits

...
57 Commits
Author SHA1 Message Date
mathiasandClaude Opus 5 489b4bb2a1 chore(spike): recover the /bygge toolchain from tmpfs (#28)
CI / Lint / Test / Vet (push) Successful in 37s
CI / Build & Import (push) Failing after 5s
CI / Deploy via GitOps (push) Skipped
analyze_srt.py, build_page.py and transcode.yaml produced the /bygge
prototype. They were sitting in a session scratchpad on tmpfs, one reboot
from gone, while spike issues #29-#31 were written as if they had to be
built from scratch.

Committed as found, defects documented rather than fixed: max_tokens=6000
truncates anything past ~6 minutes of audio, the page title is hardcoded,
and the schema still carries fields no consumer renders.

The four-check validator is the part worth keeping — the coverage-gap and
Swedish-number-grounding checks catch errors that timestamp and citation
checks structurally cannot.

Real transcripts and analyses stay out: this repo is public and the
recordings are a named person discussing a client's project. Only the
synthetic fixture is committed, which exercises all four checks with no
model call.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BeAp5LTscnz2W7Ubt4eJ6m
2026-08-13 11:37:03 +02:00
mathias 21e7d6c74b fix(ci): smoke test hung 14min — buildah run doesn't die under plain timeout
CI / Lint / Test / Vet (push) Successful in 15s
CI / Build & Import (push) Successful in 20s
CI / Deploy via GitOps (push) Successful in 5s
Bare `/tapir` runs the long-running server, same as the old ctr-based smoke
test. ctr's --rm reliably force-killed it; a plain `timeout N buildah run`
does not — it only signals the wrapper, and the container process can
survive that and keep the log pipe open, hanging the whole job (observed
live: run 155, 14min before failure). Backgrounds the run and tears it down
with `buildah rm -f`, which forcibly kills regardless of wrapper state, and
captures output via a file instead of a blocking pipe.

Refs infra#132.
2026-07-26 06:17:18 +00:00
mathias 8814ba6673 ci(smoke): use buildah run instead of sudo k3s ctr for smoke test
CI / Lint / Test / Vet (push) Successful in 12s
CI / Build & Import (push) Failing after 13m42s
CI / Deploy via GitOps (push) Has been skipped
Same fix as cobalt-dingo — removes the sudo/host-containerd dependency so
this still works once the act_runner is containerized (infra#132).

Refs infra#132.
2026-07-26 05:57:37 +00:00
mathias 4b557a4325 ci(deploy): add Flux GitOps deploy job, mirroring cobalt-dingo's pattern
CI / Lint / Test / Vet (push) Successful in 27s
CI / Build & Import (push) Successful in 18s
CI / Deploy via GitOps (push) Successful in 5s
Fixes silent stale-deploy gap: CI built+pushed images but nothing bumped
k3s/apps/tapir/deployment.yaml, so merged features sat CI-green with zero
production effect (infra#111, infra#168). Flux native image-automation
can't scan localhost:5000 from inside k3s pods, so this patches the infra
repo directly via the existing INFRA_DEPLOY_KEY org secret (same key
cobalt-dingo and brain-gardener already use) on every push to main.

Refs infra#111.
2026-07-25 21:21:02 +00:00
mathiasandClaude Opus 4.8 72cb25111f test(store): address migrations by version, not step count (#8)
CI / Lint / Test / Vet (push) Successful in 12s
CI / Build & Import (push) Successful in 12s
The up/down migration tests stepped a hard-coded number of Steps(-N)/Steps(+N)
down from HEAD and back. The counts assumed a specific latest migration, so
adding one shifted every count by one and unrelated tests (010/011/014) went
red with confusing off-by-one symptoms — a papercut on every new migration.

Drive the schema to an exact version with m.Migrate(version) via two helpers
(headVersion, migrateTo). Each test now steps to just below its target by
version, asserts the down effect, steps up to the target, asserts the up
effect, then restores to the captured HEAD. A migration added on top changes
HEAD but shifts no count, so no test needs editing.

Verified by adding a throwaway migration 017 on top: all four tests stayed
green with zero edits.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QbdxXWxLefS5AwLN5eyze
2026-07-02 14:49:02 +02:00
mathiasandClaude Sonnet 4.6 38f222c931 chore: rename Go module path gitea.d-ma.be → git.d-ma.be
CI / Lint / Test / Vet (push) Successful in 11s
CI / Build & Import (push) Successful in 12s
Infra ADR-0004 renamed the Gitea host. Bulk replace across go.mod and
all .go import paths. Build and tests pass unchanged.

Closes #20

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Dt6aHEDWRjkK14Voi6HnGh
2026-07-02 14:37:33 +02:00
mathias e042b6d26d chore: add .mcp.json (brain + gitea) — wire tapir for the single-harness workflow
CI / Lint / Test / Vet (push) Successful in 16s
CI / Build & Import (push) Successful in 12s
tapir had no MCP config at all — a hyperguild session here couldn't reach brain
or Gitea, breaking the report-back path before it could start. Matches the
shape now standard post-consolidation (hyperguild#75/#76, infra#178 template
refresh). Prerequisite for tapir#20 being run as the first real workflow test.
2026-07-02 07:00:53 +00:00
mathias c85a32e770 Merge pull request 'LANGUAGE.md: fix LiteLLM gateway never-say (stale endpoints, not piguard)' (#21) from language-md-litellm-fix into main
CI / Lint / Test / Vet (push) Successful in 12s
CI / Build & Import (push) Successful in 10s
2026-06-16 15:45:45 +00:00
mathias b246c0e688 fix(language): LiteLLM gateway never-say — stale endpoints, not piguard
CI / Lint / Test / Vet (pull_request) Successful in 12s
CI / Build & Import (pull_request) Has been skipped
piguard is legitimately in the request path as the reverse proxy; only the
stale endpoint forms (piguard:4000, koala:4000) are wrong. Per brain#6 review
amendment #3 / tapir#18 follow-on.
2026-06-16 15:35:31 +00:00
mathias c7896cb3cc Merge pull request 'LANGUAGE.md pilot: hand-written vocabulary + caveman rubric (Phase 1)' (#19) from language-md-pilot into main
CI / Lint / Test / Vet (push) Successful in 12s
CI / Build & Import (push) Successful in 10s
2026-06-16 15:22:14 +00:00
mathiasandClaude Opus 4.8 b79fb892c8 docs(claude): point to LANGUAGE.md + caveman rubric
CI / Lint / Test / Vet (pull_request) Successful in 13s
CI / Build & Import (pull_request) Has been skipped
Wires the Phase 1 vocabulary pilot into the canonical agent instructions
(tapir#18). CLAUDE.md is canonical here (no .context/ source, no generation
header).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-12 22:54:57 +02:00
mathias c315e3e003 feat(language): add hand-written vocabulary + caveman rubric pilot
Phase 1 of the ubiquitous-language system (tapir#18). Tapir-specific terms
verified against domain/ports code + README/VISION; homelab-core terms (brain
sink, LiteLLM gateway) derived from brain#5 glossary. Tooling-free pilot.
2026-06-12 20:52:23 +00:00
mathiasandClaude Opus 4.8 64e3368f5f feat(web): charm-reader visual refresh with light/dark theme toggle (ADR-032, #17)
CI / Lint / Test / Vet (push) Successful in 11s
CI / Build & Import (push) Successful in 10s
The UI read flat and boring. Reskin to one charm/TUI-inspired layout in two
palettes (CSS custom properties): a warm "reader" light theme (sketch B) and a
"cozy terminal" dark theme (sketch C).

- Palette chosen in cascade order: :root light default; an OS-preference dark
  block scoped to :root:not([data-theme]) so it applies only absent an explicit
  choice; and :root[data-theme="dark"|"light"] set by a header toggle that
  outranks the media query by specificity and persists in localStorage (guarded,
  degrades to OS default). A <head> init script applies the stored choice before
  paint, so no flash of the wrong palette.
- Charm touches via existing classes (no templ structure churn): monospace meta
  lines, accent uppercase section dividers with a trailing rule, pill buttons, a
  lifted/accent-edged expanded card.
- Error/danger shades become --err-* tokens so they follow the theme, replacing
  three per-block prefers-color-scheme dark overrides.
- Theme toggle wired into Layout and PublicLayout headers.

BDD: docs/use-cases/visual_theme.feature un-pended, mapped in scenarioCoverage.
TDD: internal/web/visual_theme_test.go (palettes, OS default, persisted toggle,
expanded-card embed). Verified light+dark on list/reader/welcome via web-shot.
Sketches kept as the design record. ui-spec.md as-built row added.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-12 12:54:06 +02:00
mathias b01f8d5783 docs(bdd): visual_theme.feature scenarios (@pending until TDD, #17) 2026-06-12 12:29:31 +02:00
mathias 8ca767c144 docs(adr): ADR-032 visual refresh — light(B)+dark(C) charm-reader themes + sketches (issue #17) 2026-06-12 11:54:54 +02:00
mathiasandClaude Opus 4.8 f98b640531 feat(web): inline-expand summary + Q&A in the list (ADR-031, #16)
CI / Lint / Test / Vet (push) Successful in 11s
CI / Build & Import (push) Successful in 11s
Click a summarized card → its full summary (summaryBody) + chat dock (chatReveal)
expand in place via HTMX (GET /v/{id}/expand → expandedCard), collapse back via
GET /v/{id}/card → compact VideoCard. Same <li id>, outerHTML swap — the existing
list-fragment pattern. The card title carries href=/v/{id} as the no-JS fallback
(detail page stays for no-JS + deep links); only summarized cards expand. Reuses
summaryBody + chatReveal so the expanded card never drifts from the detail page.

BDD: inline_expand.feature un-pended + mapped. TDD: 6 handler/fragment tests
(expand/collapse fragments, chat dock, summarized-only, no-JS href, shared body).
Minimal CSS only — the TUI/charm restyle is #17.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-12 10:47:47 +02:00
mathias 607a8cbe8d docs(bdd): inline_expand.feature scenarios (@pending until TDD, #16) 2026-06-12 10:38:40 +02:00
mathias 54d60e53b9 docs(adr): ADR-031 inline-expand summary + Q&A in the list (issue #16) 2026-06-12 10:37:55 +02:00
mathias 9298e0c686 docs(homelab): TAPIR_METRICS_ADDR + key metric series (ADR-030, #15)
CI / Lint / Test / Vet (push) Successful in 10s
CI / Build & Import (push) Successful in 10s
2026-06-12 08:50:36 +02:00
mathiasandClaude Opus 4.8 cd461b95f8 feat(observability): instrument AI + HTTP paths, serve /metrics on a side port (ADR-030, #15)
CI / Build & Import (push) Successful in 11s
CI / Lint / Test / Vet (push) Successful in 10s
Wire the metrics package into the live paths and serve it:
- summarizer: per-endpoint latency by model/outcome(success|error|parse_error)/fallback + slog.
- youtube.FetchTranscript: latency by outcome (captions|none|rate_limited) + slog.
- chat: answer latency by model + slog.
- llm usage hook → token counts (prompt|completion) per model, wired in buildSummarizer/buildChat.
- oidc callback: login counter.
- cmdServe: wrap Router in metrics.HTTPMiddleware (request count + latency by bounded
  route pattern) and serve /metrics on TAPIR_METRICS_ADDR (default :9090), a SEPARATE
  port — never on the public app mux.

BDD: observability.feature scenarios un-pended + mapped. TDD: summarizer wiring tested
black-box via the /metrics scrape; metrics-not-on-public-mux asserted.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-12 08:49:51 +02:00
mathiasandClaude Opus 4.8 9c2a04406b feat(llm): usage hook to surface token counts (ADR-030, #15)
WithUsageHook callback fires with model + prompt/completion tokens parsed from the
response usage block. Keeps the copied stdlib-only llm package decoupled from
metrics (ADR-004) — the caller wires it to internal/metrics. TDD covered.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-12 08:36:02 +02:00
mathiasandClaude Opus 4.8 a5a8cf6f6d feat(metrics): Prometheus collectors + typed API + HTTP middleware (ADR-030, #15)
New internal/metrics package: AI metrics (summarize/caption/chat latency, LLM
tokens), session metrics (requests+latency by bounded route pattern, logins), and
a /metrics Handler. Adapters call a typed API; never touch prometheus types.
Dep: github.com/prometheus/client_golang (standard Go client; cluster runs
prometheus-operator). TDD: collectors + middleware route-pattern cardinality
covered.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-12 08:32:41 +02:00
mathias 64d11af9ef docs(bdd): observability.feature scenarios (@pending until TDD, #15) 2026-06-12 08:28:46 +02:00
mathias b590d2708d docs(adr): ADR-030 observability — slog + Prometheus metrics (issue #15) 2026-06-12 08:27:56 +02:00
mathiasandClaude Opus 4.8 7314895ec4 fix(auth): stateless session cookie — stop logging users out on deploy (ADR-029)
CI / Lint / Test / Vet (push) Successful in 10s
CI / Build & Import (push) Successful in 11s
Pilot feedback: lots of re-logging-in on iPhone. Three causes: sessions lived in
an in-memory map (wiped on every pod restart/deploy), a 1h TTL (idle >1h forced
re-login on a check-back-tomorrow reader), and a session cookie with no Max-Age
(dropped on Safari close). Each re-login is the full IdP redirect dance.

Make sessions stateless: identity + absolute expiry live inside the existing
HMAC-signed cookie (no server table), TTL 1h → 30 days sliding, cookie now
persistent (Max-Age). Survives restarts (test: a cookie from one instance is
accepted by a fresh instance with the same secret), browser-close, and idle.
Trade: no server-side revocation — logout clears the cookie client-side; rotating
tapir-session-secret is the global logout lever. Accepted for the Stage-0 reader.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-11 23:07:13 +02:00
mathiasandClaude Opus 4.8 36dd182fb5 feat(report): Stage-0 usage gate counts from a baseline date (default 2026-06-11)
CI / Build & Import (push) Successful in 11s
CI / Lint / Test / Vet (push) Successful in 10s
The return-usage gate (ADR-016) counted distinct active weeks over ALL history,
so pre-launch noise — testing churn and the period the pilot sat blocked on zero
summaries — would inflate the signal. Add a baseline: ActiveWeeks(ctx, since)
filters login_events + summary_actions to seen_at/acted_at >= since. The report
command sets it to TAPIR_USAGE_GATE_START (YYYY-MM-DD, default 2026-06-11 — the
morning the pilot was unblocked) and prints the baseline. The gate now measures
whether users RETURN once it genuinely works.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-11 22:52:15 +02:00
mathias 40808f2d4b docs(onboard): BDD scenarios + architecture for the ADR-028 burst
CI / Lint / Test / Vet (push) Successful in 10s
CI / Build & Import (push) Successful in 10s
- connect_account.feature: refine the burst scenario to "best recent" (likely-good
  selection, skips too-short/too-long) and add a scenario for the stronger model.
- scenario_coverage: remap to TestOnboardBurstVideoIDs + add the model scenario
  (the old NewestUnsummarizedVideoIDs test was removed).
- architecture.md: new "Connect-time onboarding burst" subsection — junk-avoiding
  selection over persisted duration, the burst-only stronger-model chain, and why
  has-captions/cached-first are not selection signals.
2026-06-11 19:54:29 +02:00
mathias 9232f49555 feat(onboard): burst picks likely-good videos, summarizes with stronger model (ADR-028)
Wire the onboarding burst to its quality-aware selection and a stronger model:
- main.go onboard uses OnboardBurstVideoIDs (junk-avoiding) instead of pure
  newest-first, bounded by MinVideoSeconds / OnboardMaxVideoSeconds.
- buildBurstProcessor builds a burst-only summarizer chain led by the onboard
  model (burstChainModels: onboard -> standard ADR-022 chain, deduped, NDA lever
  intact), over the same store/cache/sink. Collapses onto the shared Processor
  when the onboard model is empty/equal-to-primary or config is incomplete.
- summarizerEndpoint + newYouTubeSource extracted so standard and burst wiring
  share one definition.
- Remove now-superseded NewestUnsummarizedVideoIDs: OnboardBurstVideoIDs(.,0,0)
  is identical pure-newest behaviour and its test covers RLS + ordering.
2026-06-11 19:54:29 +02:00
mathias 582c1a2065 feat(store): OnboardBurstVideoIDs — junk-avoiding burst selection (ADR-028)
Newest-first but quality-aware: excludes a video when its duration is KNOWN and
outside [minSeconds, maxSeconds], dropping Shorts and multi-hour livestream VODs
that waste a scarce caption fetch on a poor first impression. NULL/unknown
duration is kept (degrade-open) but ranked after known-good rows. 0/0 bounds
disable the filter (pure newest-first, the reversibility lever). RLS-scoped.
NewestUnsummarizedVideoIDs is left in place for callers that want pure-newest.
2026-06-11 19:54:29 +02:00
mathias 9f0d8cf198 feat(discovery): persist video duration_s instead of discarding it (ADR-028)
ADR-023's filterLowValue already fetches each candidate's duration via the
cheap videos.list quota call to drop Shorts/live, then threw it away — the
videos.duration_s column (migration 001) was never written. Carry it onto the
kept domain.Video and have UpsertVideo persist it, COALESCE-preserving a known
value so an unknown (0) re-upsert never clobbers it (the channel_title backfill
stance, migration 014). This is the enabling change for length-aware burst
selection. No new migration — the column already exists.
2026-06-11 19:54:29 +02:00
mathias d21077303d feat(config): add OnboardSummarizerModel + OnboardMaxVideoSeconds (ADR-028)
Two knobs for the onboarding-burst quality work:
- TAPIR_ONBOARD_SUMMARIZER_MODEL (default iguana/gemma4-26b): the stronger model
  the burst leads its chain with; empty collapses the burst onto the shared
  processor (lookupOr, so explicit-empty disables).
- TAPIR_ONBOARD_MAX_VIDEO_SECONDS (default 14400/4h): upper duration bound for
  burst picks; 0 disables, negative clamps to 0.

Table-driven tests cover defaults, explicit, disable, and invalid input.
2026-06-11 19:54:29 +02:00
mathias 9b2ee2e765 docs: spec onboarding-wow-burst + ADR-028 (better picks, stronger model)
Phase-1 live-DB investigation found the connect-time burst (ADR-018) fires
but delivers a weak first impression: pure newest-first selection picks junk
(a livestream + regional news for pilot user Jonte), and all burst summaries
run on the weak koala/phi4-mini instead of the validated iguana/gemma4-26b.
Cached-first was investigated and rejected — ~3% cross-user overlap, 0
cached-and-unsummarized, newest-20 all uncached (newest-first and cached-first
are structurally incompatible).

ADR-028 records: persist duration_s at discovery (ADR-023 already fetches it,
just stops discarding), junk-avoiding burst selection (drop known too-short/
too-long), and a stronger burst-only summarizer chain. Not a throughput change.
2026-06-11 19:54:29 +02:00
mathias b8fbc5a805 docs: spec first-session wow — verify & improve onboarding summary burst (investigate-first)
CI / Lint / Test / Vet (push) Successful in 11s
CI / Build & Import (push) Successful in 10s
The concern is new-user first-contact: see some GOOD summaries fast or they
don't return. Reframed as curation/latency for ~3 videos, NOT a 429/throughput
problem (3 fetches is nowhere near the wall). Phase 1 (report-and-stop) verifies
whether the existing cap-3 onboarding burst even fires today, what it delivers,
and — critically — how much transcript-cache overlap exists between users (drives
the blend). Phase 2 levers: cached-transcript-first (instant, zero-fetch),
likely-good selection (has-captions/good-length, not just newest), and optionally
the stronger model for the burst's few summaries. Blend deferred to the
maintainer post-Phase-1. Explicitly NOT bulk-fetch, NOT credentials (ADR-010/026
dead end), NOT a client extension.
2026-06-11 15:36:09 +00:00
mathiasandClaude Opus 4.8 19ca4282a8 refactor(web): dock chat inline below the summary, integrated view (ADR-027)
CI / Lint / Test / Vet (push) Successful in 10s
CI / Build & Import (push) Successful in 10s
The chat was a separate page — opening it left the summary behind, the very
context you're asking about. Now the chat docks open IN PLACE below the summary:
"Dig deeper" reveals the chat section (HTMX, outerHTML over the closed dock) so
the summary stays on screen above it; no navigation. The no-JS fallback renders
the full summary AND the open chat on one page (the same integrated view), so
progressive enhancement holds.

Extracts a shared summaryBody templ so the detail page and the chat page render
one identical summary, not two divergent ones. GET /v/{id}/chat returns just the
open chat section as a fragment for the inline reveal, or the full summary+chat
page for a no-JS navigation; POST swaps the panel inline or re-renders the whole
page. Tests assert summary+chat coexist on the page and the reveal is a fragment.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-11 14:02:39 +02:00
mathiasandClaude Opus 4.8 71df696448 feat(web): per-video chat over the stored transcript (ADR-027)
CI / Lint / Test / Vet (push) Successful in 10s
CI / Build & Import (push) Successful in 10s
A deeper-dive chat entered from the summary view: ask questions about a video
against its already-stored transcript (ADR-021), no caption fetch, ever.

Safety by construction — the load-bearing property. The chat handlers reach the
chat.Service only after reading the SHARED stored transcript via Store.GetTranscript
(a pure DB read); the service holds no VideoSource. So an enabled chat cannot
trigger a caption fetch, touch the rate gate, or reach YouTube. A video with no
stored transcript gets an honest "not available" — no fetch, no model call. The
key web test wires the summarize/fetch collaborators as tripwires that fail the
test if chat ever routes into them, and asserts the model answered from the
stored text.

Model defaults to the summary's own model and is switchable among the ADR-022
chain (phi4-mini → gemma4-26b → mistral-small); switching re-runs against the
same transcript — deliberate model-comparison instrumentation. The cloud model
is absent from the switcher when TAPIR_CLOUD_FALLBACK_MODEL="" (the local-first /
NDA lever), honoured the same way the summarizer honours it. Reuses the existing
LiteLLM gateway client (a chat is a different call, not a new integration) and
the TAPIR_MAX_TRANSCRIPT_CHARS truncation, surfacing an honest bounded-context
note when a long transcript is cut.

Ephemeral v1: the multi-turn conversation rides in hidden request fields; no
table, no migration, nothing persisted. Entry is RLS-scoped through
GetSummaryByVideo, so chat is reachable only from the user's own summary view.
Show-source verification and on-demand fetch are deferred (ADR-027).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-11 09:26:58 +02:00
mathiasandClaude Opus 4.8 cc3cda4ab8 feat(chat): stored-transcript QA service (ADR-027)
A read-only deeper-dive over a video's already-stored transcript (ADR-021).
Safe by construction: the Service has no VideoSource and no caption-fetch
dependency — only a Completer factory over the existing LiteLLM gateway — so it
cannot reach YouTube or the rate gate. Reuses the summarizer's truncation
discipline (TAPIR_MAX_TRANSCRIPT_CHARS), reporting the cut so the UI can be
honest about a bounded transcript. Model defaults to the summary's model and is
switchable among an offered, local-first list; an un-offered alias is forced
back to the default so chat can never call the gateway with an arbitrary model.
Ephemeral: history is carried per-request, nothing persisted.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-11 09:16:45 +02:00
mathiasandClaude Opus 4.8 69a49bc603 docs(decisions): add ADR-027 — chat with a video's stored transcript
Records the deeper-dive chat decision before its build: stored-transcript-only
(safe by construction — no caption fetch, no rate gate, no YouTube), entered
from the summary view, model defaulting to the summary's model and switchable
among the ADR-022 chain. Ephemeral v1; show-source verification deferred to v2.
Inserted before "Rejected alternatives" so the decision precedes the feature.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-11 09:13:31 +02:00
mathias 137804b0b1 docs: spec chat-with-transcript (ADR-027)
CI / Lint / Test / Vet (push) Successful in 11s
CI / Build & Import (push) Successful in 10s
Per-video chat against the ADR-021 stored transcript, entered from the summary
view, born from observed demand (maintainer read real summaries, some made him
want to dig deeper). HARD constraint: stored-transcript-only — never fetches
captions, never touches the rate gate or YouTube, safe by construction. Default
model = the summary's model, user-switchable among the ADR-022 chain models
(doubles as model-comparison instrumentation). Ephemeral v1 (no persisted
history); chat-only/trust-the-model with show-source verification recorded as the
natural v2. ADR-027 to be appended to DECISIONS.md as the first commit.
2026-06-11 06:38:16 +00:00
mathiasandClaude Opus 4.8 a9be5f285b feat(summarizer): move local fallback off koala to iguana/gemma4-26b
CI / Build & Import (push) Successful in 10s
CI / Lint / Test / Vet (push) Successful in 10s
koala now carries other GPU loads, so the first fallback should not run there.
Change the default chain to koala/phi4-mini → iguana/gemma4-26b → berget/mistral-small:
the local fallback now runs on iguana (M2 Ultra headroom, different host = different
egress IP for the rare fallback fetch). gemma4-26b is the brain-validated homelab
general-purpose model (agentsquad H2/H3 executor) and returned valid summary JSON on
the real prompt in a smoke test (~37s incl. cold-load — fine for a fallback path).

Pure config default (TAPIR_FALLBACK_MODEL); chain mechanism (ADR-022) unchanged.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-11 07:54:56 +02:00
mathiasandClaude Opus 4.8 fe56e2fe01 test: make embedded-postgres per-process so concurrent CI runs don't collide
CI / Lint / Test / Vet (push) Successful in 9s
CI / Build & Import (push) Successful in 10s
A push to main and its version tag fire two CI runs for the same commit. Both ran
`go test ./...`, which starts embedded-postgres on a FIXED port (54329/54330) and
a shared data dir. -p 1 serialises packages WITHIN a run, not across two
concurrent runs — so when the two runs overlapped they collided on the port/data
dir and BOTH failed the Lint/Test job (no image built). Prior commits passed only
because their two runs happened not to overlap.

Derive the port and runtime/data dirs from the PID; share only CachePath so the
PG archive downloads once. Proven: two concurrent `go test` of the store package
now both pass. Unblocks the v0.21.0 (Pillar A) build.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-11 07:26:33 +02:00
mathiasandClaude Opus 4.8 beeb5bc31b feat(gate): foreground caption fetches take priority over the background sweep (ADR-026)
CI / Lint / Test / Vet (push) Failing after 9s
CI / Build & Import (push) Has been skipped
A user waiting on a Summarize click shared the per-IP caption gate equally with
the background firehose, so on a busy IP the click was slow or 429'd. Add a
context-marked priority lane: the web path (engineProcessor.ProcessVideo) marks
its context foreground; the gate serves foreground immediately while background
fetches yield until no foreground is pending. Threaded via a context value (no
new signatures) + a process-wide foregroundPending counter. Clicks are rare, so
the background barely loses throughput; the waiting human gets the cleaner slot.

Drops the credentials probe: ADR-010 and captions.go already settle it — the
timedtext/InnerTube path rejects authenticated requests and the OAuth token does
not authenticate it anyway, so auth cannot help and can hurt. Documented in
ADR-026 rather than built.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-10 22:10:15 +02:00
mathiasandClaude Opus 4.8 09eb31d1fe fix(scheduler): derive rotation offset from wall-clock, not a reset-on-restart counter
CI / Lint / Test / Vet (push) Successful in 11s
CI / Build & Import (push) Successful in 14s
The lead-user rotation used an in-memory pass counter reset to 0 on every pod
restart, so the first-listed user always re-took the lead after a restart — a
deploy-heavy session re-starved the last user (the pilot stalled at 2 summaries
because each deploy reset his every-other-pass lead before the 2h tick fired).
Derive the offset from wall-clock (floor(now/interval)) so it advances with real
time and is identical across restarts: rotation stays fair however often the pod
bounces.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-10 21:59:12 +02:00
mathiasandClaude Opus 4.8 1665a1e7c4 feat(web): honest, state-aware summarize status with charm spinner (ADR-025)
CI / Build & Import (push) Successful in 10s
CI / Lint / Test / Vet (push) Successful in 12s
Clicking Summarize polled /status, which knew only "spinning" or "done". The web
ProcessVideo recorded an outcome only on success, so a 429'd or caption-less
click left transcript_status unset and the poll silently reverted to the
Summarize button — the rate limit was invisible and the click felt broken.

- ProcessVideo now records rate_limited / none / fetched (mirrors the runner); a
  rate-limited video keeps its requested flag so the background sweep retries it.
- /status is state-aware: summary card (done), working spinner (in-flight),
  a calm "waiting on rate limit, will retry" card that keeps polling so the
  summary lands on its own (no re-click), and a terminal "no captions" card.
- Charm status text (Claude-Code / Crush inspired): the spinner cycles playful
  tapir-themed gerunds via CSS only (no JS), decorative + an sr-only stable
  status line for a11y.

Pillar A (foreground fetch priority lane) is a separate follow-up.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-10 21:52:33 +02:00
mathiasandClaude Opus 4.8 5219561a91 feat(discovery): per-channel caption-availability memory (ADR-024)
CI / Lint / Test / Vet (push) Successful in 30s
CI / Build & Import (push) Successful in 11s
After transcript caching (ADR-021) and the Shorts filter (ADR-023), the remaining
caption waste is the first fetch on every new video of a channel that never has
English captions — each costs one rate-limited fetch to resolve to "none", and on
a throttled IP churns the backoff machinery first.

Remember, per (user, channel), a streak of consecutive no-caption outcomes
(channel_caption_state, migration 016, RLS-scoped). Once it reaches
TAPIR_CHANNEL_CAPTIONLESS_THRESHOLD (default 5) the channel is suppressed — videos
discovered/listed but not caption-fetched — for TAPIR_CHANNEL_CAPTIONLESS_WINDOW
(default 14d), then one is re-probed (auto-recovery). A successful fetch resets
the streak; a 429 does not count; an explicit manual request bypasses suppression.
threshold=0 disables.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-10 20:27:25 +02:00
mathiasandClaude Opus 4.8 9db06d8a63 feat(discovery): drop Shorts and livestreams before the caption fetch (ADR-023)
CI / Lint / Test / Vet (push) Successful in 12s
CI / Build & Import (push) Successful in 10s
The scarce resource is the per-IP timedtext caption fetch (ADR-014); the pilot's
candidate set was mostly Shorts/clips/livestreams, each burning a fetch (a "none"
result is a completed fetch — it costs budget even when it yields nothing).

NewVideos now enriches candidates with one cheap Data API videos.list call
(contentDetails.duration + snippet.liveBroadcastContent — the quota API, a
DIFFERENT limit from the timedtext 429) and drops, before returning: videos
shorter than TAPIR_MIN_VIDEO_SECONDS (default 60) and any live/upcoming
broadcast. Dropped videos are never persisted, so the list declutters too.

Degrade-open: MinVideoSeconds=0 disables it (no quota call); a videos.list error
returns candidates unfiltered so discovery never breaks on a metadata hiccup. The
paste-a-URL path (VideoByID) is not filtered — an explicit request is honoured.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-10 19:34:45 +02:00
mathiasandClaude Opus 4.8 1aa8a97f95 fix(scheduler): rotate over connected users only for true fair share
CI / Lint / Test / Vet (push) Successful in 12s
CI / Build & Import (push) Successful in 10s
The lead-user rotation rotated the full ListAllUsers set, so a connectionless
orphan identity ate a rotation slot — collapsing onto the next real user and
skewing the lead share (two real users got 2/3 vs 1/3 instead of 50/50). Filter
to connected users BEFORE rotating so the rotation is over exactly the users that
consume the caption budget. A dead identity can no longer skew fairness.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-10 18:28:54 +02:00
mathiasandClaude Opus 4.8 cc69a912f4 fix(scheduler): rotate lead user each pass so caption budget is shared
CI / Lint / Test / Vet (push) Successful in 12s
CI / Build & Import (push) Successful in 10s
Caption fetches share one per-egress-IP rate budget; whoever runs first each pass
spends the pre-throttle window before YouTube starts 429ing. ListAllUsers order
is unspecified and was stable, so the last-listed user was permanently starved —
a friendly-pilot user got 0 fetches in 12h (all rate_limited) while the
first-listed user got every successful fetch. rotateUsers left-rotates the user
order by pass index so each user leads 1/N passes and the lead slot is shared.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-10 18:07:01 +02:00
mathiasandClaude Opus 4.8 f4a0544903 fix(scheduler): cache transcripts on the scheduled path (ADR-021 regression)
buildUserRunner built the engine without engine.Transcripts = st, so the
scheduler — unlike the web "Summarize now" path — never read or wrote the shared
transcript cache. Every discovery pass re-fetched transcripts it had already
fetched, burning the scarce per-egress-IP caption budget (ADR-014) on redundant
work and starving other users' first-time fetches. The transcripts table was
empty despite summaries existing. Wire the cache on this path too.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-10 18:07:01 +02:00
mathiasandClaude Opus 4.8 e2a52789b9 feat(web): mode-aware backlog banner — stop telling manual users summaries auto-arrive
CI / Lint / Test / Vet (push) Successful in 12s
CI / Build & Import (push) Successful in 10s
The first pilot user sat in Manual mode reading "new summaries land gradually,
check back tomorrow" — copy that only makes sense in Automatic mode. Manual mode
never auto-summarizes, so the banner promised delivery that would never come.

The list page now reads the user's summarize mode and shows mode-correct copy:
- Auto: unchanged "land gradually" backlog note.
- Manual: "new videos appear here but are not summarized automatically — use the
  Summarize button" plus a "Switch to Automatic" link to /account.
The connected-but-empty first-run state is likewise mode-aware.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-10 09:14:52 +02:00
mathias e696b6405b docs(build-state): v0.15.0, summarizer is now a resilient chain (ADR-022)
CI / Lint / Test / Vet (push) Successful in 12s
CI / Build & Import (push) Successful in 10s
2026-06-10 08:59:26 +02:00
mathiasandClaude Opus 4.8 e9b5a3f3e7 feat(summarizer): resilient endpoint chain with local→cloud fallback (ADR-022)
CI / Lint / Test / Vet (push) Successful in 11s
CI / Build & Import (push) Successful in 10s
The first friendly-pilot live run produced zero summaries: koala/phi4-mini hit
three silent failure modes — 8k context overflow on long transcripts (HTTP 400),
intermittent malformed JSON (highlights as a bare string), and no fallback wired
at all (summarizer.New(primary, nil)).

Keep phi4-mini as the fast primary and add resilience around it:

- Ordered endpoint chain (summarizer.NewChain): phi4-mini → koala/phi4-14b
  (local) → berget/mistral-small (worst-case external). All reached through the
  one LiteLLM gateway by alias.
- A parse failure now advances the chain like a transport error — the old
  Primary→Fallback shape returned the parse error without trying anyone else.
- Tolerant parse: highlights/takeaways coerce string→[]string, absorbing the
  common small-model quirk without spending a fallback round-trip.
- Transcript truncation (TAPIR_MAX_TRANSCRIPT_CHARS=18000) prevents the overflow
  rather than recovering from it; validated to fit phi4-mini's 8k window.
- Bounded completion budget (TAPIR_SUMMARY_MAX_TOKENS=1500) — the old 8192 budget
  itself contributed to the overflow.

Local-first guarantee preserved by ordering: external endpoint is tried only
after every local one fails. TAPIR_CLOUD_FALLBACK_MODEL="" disables it entirely
for client/NDA deployments.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-10 08:36:16 +02:00
mathiasandClaude Opus 4.8 0ba78e8868 docs: refresh build-state for transcript persistence (ADR-021)
CI / Lint / Test / Vet (push) Successful in 11s
CI / Build & Import (push) Successful in 10s
Update the CLAUDE.md orientation block: last tag v0.14.0, migrations
001–015, and a transcript-persistence bullet (shared non-RLS store,
engine reads stored-first). The stale "v0.9.0" reference is corrected.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-09 23:39:27 +02:00
mathiasandClaude Opus 4.8 821d5f99cd docs(bdd): scenario for transcript reuse — re-analysis never re-fetches
Capture the ADR-021 promise as a mapped BDD scenario: re-analyzing a
stored video reads the stored transcript and does not fetch captions.
Since paste-a-URL and the onboarding burst summarize through the same
engine chokepoint (resolveTranscript, store-first), this one scenario
covers their reuse path too — there is exactly one gated caption entry
point (youtube.FetchTranscript → WaitFetchGate) and one engine caller in
front of it, so the dedup is structural, not per-feature.

Mapped to TestProcessNewVideo_SecondSummarizeDoesNotRefetch.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-09 23:37:54 +02:00
mathiasandClaude Opus 4.8 5c70408e75 feat(usecase): read stored transcript before fetching (ADR-021)
The engine now resolves transcripts store-first: a stored transcript —
including a stored SourceNone — is summarized without touching YouTube,
so re-analysis never re-fetches. On a miss it fetches through the source
(caption call still gated, ADR-014) and persists the terminal outcome for
the next analysis by any user. A transient SourceRateLimited is surfaced
to the runner for per-user backoff but never cached, so persistence can
never mask a 429 as a permanent "no transcript".

The TranscriptStore is optional (nil → fetch every time), keeping the
pure-core and scaffold wiring valid. cmd/tapir wires the store as both
summary sink and transcript cache, so `tapir run` and the web summarize
path (incl. paste + onboarding) all share the dedup.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-09 23:35:04 +02:00
mathiasandClaude Opus 4.8 cb6917ca59 feat(store): shared, video-keyed transcript persistence (ADR-021)
Reshape the dead per-user transcripts table (PK videos.id, user_id,
RLS-FORCEd — never read or written by app code) into the shared public
caption store ADR-021 specifies: keyed by (provider, provider_video_id),
no user_id, NOT RLS-scoped. Migration 015 (reversible). Add
ports.TranscriptStore + Store.GetTranscript/SaveTranscript via the raw
pool (no withUser): public content, shared across users by construction.
SaveTranscript persists only terminal outcomes (captions/none) and
refuses SourceRateLimited so a transient 429 can never be stored as a
false permanent absence (ADR-014).

Flip the isolation proof: transcripts leaves the RLS-scoped set;
TestTranscriptsTableIsSharedNotRLS asserts it is the SINGLE non-RLS
surface (writable/readable with no user scope, no user_id column, RLS off
on it alone, still on every user-owned table) — the proof the
public-content classification was applied exactly here and leaked nowhere.
appPool made idempotent so two tests can build it. Adjust the 010/011/014
up-down migration tests for the new HEAD. account.go: user deletion no
longer strips shared transcripts. Reconcile data-model.md + CLAUDE.md.

Wiring the engine to read-stored-first is the next commit.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-09 23:32:12 +02:00
mathiasandClaude Opus 4.8 099b2d4c68 docs(decisions): ADR-021 — shared, video-keyed transcript persistence
Persist transcripts in a single shared table keyed by
(provider, provider_video_id) — public caption content, NOT RLS-scoped —
so re-analysis (re-summarize, paste of an already-seen video, a second
user with overlapping subs) never re-fetches from YouTube. The avoided
cost is the rate-gated, reputation-risky caption fetch (ADR-010/014), not
LLM re-summarization, which is why this reopens the transcripts half of
the "no global cross-tenant table" rejection while videos stay per-user.
Summaries remain RLS-scoped (ADR-012 unchanged). The gate is neither
bypassed nor weakened — persistence reduces fetch frequency, not pacing.

Annotate the rejected-alternatives row to record the partial reopen.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-09 23:24:11 +02:00
mathiasandClaude Opus 4.8 f66c1bcdcc feat(web): real channel filter — multi-select of the user's channels
CI / Build & Import (push) Successful in 10s
CI / Lint / Test / Vet (push) Successful in 11s
The free-text 'channel' filter was dead: it exact-matched SummaryRow.Channel,
which is just the provider ('youtube'), because videos never stored their source
channel. Now they do.

- migration 014: videos.channel_title (nullable; existing rows backfill on the
  next discovery pass, pasted videos immediately).
- discovery (NewVideos) + paste (VideoByID) populate channel_title; UpsertVideo
  persists it, preserving an existing title when an update arrives empty.
- store.DistinctChannels lists a user's channels (RLS-scoped); SummaryRow carries
  ChannelTitle via the shared projection.
- Filter: single Channel -> Channels []string, matching on ChannelTitle; the feed
  renders a multi-select of DistinctChannels (hidden until channels exist).
- migrate tests: 014 reversibility + fixed the relative-step counts in the 010/011
  up/down tests (014 shifted the topology).

TDD throughout: channel persist + distinct, adapter channel wiring, multi-channel
filter match, handler channel filter, migration up/down.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-09 23:02:54 +02:00
118 changed files with 9031 additions and 1159 deletions
+137 -5
View File
@@ -78,16 +78,148 @@ jobs:
${REGISTRY}/${{ env.IMAGE }}:${{ steps.meta.outputs.version-tag }} || true
echo "Image pushed to ${REF}"
# Run the just-built local image briefly via buildah, not k3s ctr —
# avoids sudo/host-containerd access so this still works once the
# runner itself is containerized (infra#132). Tests the local
# buildah-store image directly, no registry round-trip needed.
#
# Bare `/tapir` (no subcommand) runs the long-running server, same as
# the old ctr-based smoke test -- ctr's --rm reliably force-killed it,
# but a plain `timeout N buildah run` does NOT: it only signals the
# `buildah run` wrapper, and the actual container process can survive
# that and keep the output pipe open, hanging the whole job (hit this
# live: 14min hang, infra#132). Backgrounding the run + `buildah rm -f`
# decouples "is the script blocked" from "did the process exit" --
# rm -f forcibly tears down the container regardless of wrapper state.
- name: Smoke test
run: |
REGISTRY="localhost:5000"
REF="${REGISTRY}/${{ env.IMAGE }}:${{ steps.meta.outputs.sha-tag }}"
CNAME="smoke-${{ steps.meta.outputs.sha-tag }}"
sudo k3s ctr images pull --plain-http ${REF}
OUTPUT=$(timeout 5 sudo k3s ctr run --rm ${REF} ${CNAME} /tapir 2>&1 || true)
sudo k3s ctr containers delete ${CNAME} 2>/dev/null || true
CONTAINER=$(buildah from ${REF})
LOG=$(mktemp)
buildah run "$CONTAINER" -- /tapir > "$LOG" 2>&1 &
RUNPID=$!
sleep 5
kill -9 "$RUNPID" 2>/dev/null || true
wait "$RUNPID" 2>/dev/null || true
buildah rm -f "$CONTAINER" >/dev/null 2>&1 || true
OUTPUT=$(cat "$LOG"); rm -f "$LOG"
echo "$OUTPUT" | grep -q "tapir" \
&& echo "Smoke test passed" \
|| echo "Smoke test inconclusive: $OUTPUT"
# ── 3. Mirror to GitHub — skipped for now (SSH key rotation pending)
# ── 3. Deploy via infra repo + Flux ────────────────────────────────────────
# Flux native image-automation can't scan localhost:5000 from inside k3s pods
# (mathias/infra k3s/flux/flux-system/image-automation.yaml) — this job
# mirrors cobalt-dingo's proven pattern instead: patch the infra repo's
# manifest directly on every push to main, then annotate Flux for a fast
# reconcile. Fixes infra#111 (image built+pushed but manifest bump was
# manual, so merged features silently didn't deploy).
deploy:
name: Deploy via GitOps
needs: build
runs-on: self-hosted
if: github.ref == 'refs/heads/main' && github.event_name == 'push'
environment: staging
steps:
- name: Update image tag in infra repo
env:
IMAGE_TAG: ${{ needs.build.outputs.image-tag }}
DEPLOY_KEY: ${{ secrets.INFRA_DEPLOY_KEY }}
run: |
set -euo pipefail
# INFRA_DEPLOY_KEY is a Gitea org secret (mathias org), already
# configured per docs/cd-pipeline.md in the infra repo — same key
# cobalt-dingo and brain-gardener use, no new secret needed.
mkdir -p ~/.ssh
echo "$DEPLOY_KEY" > ~/.ssh/id_infra
chmod 600 ~/.ssh/id_infra
ssh-keyscan -p 30022 10.0.1.20 >> ~/.ssh/known_hosts 2>/dev/null
export GIT_SSH_COMMAND="ssh -i ~/.ssh/id_infra -o IdentitiesOnly=yes"
rm -rf /tmp/infra
git clone -b main ssh://git@10.0.1.20:30022/mathias/infra.git /tmp/infra
cd /tmp/infra
DEPLOYMENT="k3s/apps/tapir/deployment.yaml"
# In-place update of the image tag. sed (not yq) so we don't
# depend on additional tooling on the runner — same as cobalt-dingo.
sed -i "s|image: localhost:5000/tapir:.*|image: localhost:5000/tapir:${IMAGE_TAG}|" "$DEPLOYMENT"
# Verify the patch took effect.
grep -q "localhost:5000/tapir:${IMAGE_TAG}" "$DEPLOYMENT" \
|| { echo "✗ image tag patch failed"; exit 1; }
if git diff --quiet "$DEPLOYMENT"; then
echo " image tag unchanged — skipping push"
else
git -c user.name="tapir CI" \
-c user.email="ci@tapir.local" \
commit -m "chore(deploy): tapir → ${IMAGE_TAG}" "$DEPLOYMENT"
git push origin main
echo "✓ pushed to infra repo"
fi
shred -u ~/.ssh/id_infra
- name: Trigger Flux reconcile (immediate)
run: |
# Without these annotations, Flux would still pick up the change
# within 30s (the apps Kustomization interval). The annotations
# cut latency to ~1s.
kubectl -n flux-system annotate gitrepository flux-system \
reconcile.fluxcd.io/requestedAt="$(date +%s)" --overwrite
kubectl -n flux-system annotate kustomization apps \
reconcile.fluxcd.io/requestedAt="$(date +%s)" --overwrite
- name: Wait for Flux to apply new image
env:
IMAGE_TAG: ${{ needs.build.outputs.image-tag }}
run: |
# Poll the Deployment spec until it reflects the new tag.
# Bound to 60s so a stuck Flux doesn't hang CI.
EXPECTED="localhost:5000/tapir:${IMAGE_TAG}"
for i in $(seq 1 60); do
CURRENT=$(kubectl get deploy tapir -n tapir \
-o jsonpath='{.spec.template.spec.containers[0].image}' 2>/dev/null || echo "")
if [ "$CURRENT" = "$EXPECTED" ]; then
echo "✓ Flux applied new image after ${i}s"
break
fi
sleep 1
done
# Final assertion (in case the loop exited without matching).
kubectl get deploy tapir -n tapir \
-o jsonpath='{.spec.template.spec.containers[0].image}' \
| grep -qx "$EXPECTED" \
|| { echo "✗ Flux did not apply new image within 60s"; exit 1; }
- name: Verify rollout
run: |
kubectl rollout status deployment/tapir \
--namespace tapir \
--timeout=120s \
|| {
echo "── pod status ──"
kubectl get pods -n tapir -o wide
echo "── events ──"
kubectl get events -n tapir --sort-by='.lastTimestamp' | tail -20
echo "── describe ──"
kubectl describe pods -n tapir -l app=tapir | tail -40
exit 1
}
- name: Confirm pod running new image
env:
IMAGE_TAG: ${{ needs.build.outputs.image-tag }}
run: |
kubectl get pods -n tapir \
-l app=tapir \
--field-selector=status.phase=Running \
-o jsonpath='{.items[*].spec.containers[0].image}' \
| grep -q "localhost:5000/tapir:${IMAGE_TAG}" \
&& echo "✓ pod running new image" \
|| { echo "✗ pod image mismatch"; exit 1; }
# ── 4. Mirror to GitHub — skipped for now (SSH key rotation pending) ─
+11
View File
@@ -11,3 +11,14 @@
.env.*
!.env.example
*.local
# Spike media (#28): real recordings, their transcripts and derived analyses are
# private third-party content and this repo is public. Only synthetic fixtures
# are committed — see scripts/spike-media/README.md.
/scripts/spike-media/*.mov
/scripts/spike-media/*.mp4
/scripts/spike-media/*.wav
/scripts/spike-media/*.srt
/scripts/spike-media/*.json
/scripts/spike-media/*.html
!/scripts/spike-media/fixtures/
+18
View File
@@ -0,0 +1,18 @@
{
"mcpServers": {
"brain": {
"type": "http",
"url": "https://brain-mcp.d-ma.be/mcp",
"headers": {
"Authorization": "Bearer ${BRAIN_MCP_TOKEN}"
}
},
"gitea": {
"type": "http",
"url": "https://git-mcp.d-ma.be/mcp",
"headers": {
"Authorization": "Bearer ${GITEA_MCP_TOKEN}"
}
}
}
}
+12 -4
View File
@@ -12,6 +12,7 @@ docs it indexes.
3. `DECISIONS.md` — the ADRs. Decisions are settled here; do not re-litigate without a new ADR.
4. `docs/architecture/architecture.md`, `docs/data-model.md`, `docs/use-cases/*.feature`.
5. `docs/homelab-integration.md` — the concrete endpoints/conventions you'll need.
6. `LANGUAGE.md` — the project vocabulary. Apply the caveman rubric before destructive operations.
## How to work in this repo
@@ -46,8 +47,9 @@ These caused real mistakes that were caught and corrected; the corrections are l
(See `DECISIONS.md` for full rationale. Listed here so you don't propose them.)
- **No Supabase** — reuse Dex / ESO+1Password / Postgres (ADR-002).
- **No global cross-tenant video/transcript table** — per-user isolation (data-model). Dedup
across users is a Future C concern, not a Stage 0/1 default.
- **No global cross-tenant *video* table** — videos stay per-user (data-model). Transcripts ARE
shared since ADR-021 (public caption content, keyed by `(provider, provider_video_id)`, non-RLS)
so re-analysis never re-fetches; the *videos* half of cross-tenant dedup stays a Future C concern.
- **No audio-download + speech-to-text in the core path** — captions-first (ADR-007). STT is a
deferred, bounded optional component.
- **No public SaaS / sign-up / billing / Google OAuth verification at scale** — Future C,
@@ -85,7 +87,7 @@ Skills live in the canonical library `mathias/skills` and are wired into this re
## Current build state (start here for the first task)
The repo is **green and shipping** — last tag `v0.9.0`. `task check` passes (fmt, vet, lint,
The repo is **green and shipping** — last tag `v0.15.0`. `task check` passes (fmt, vet, lint,
`go test -p 1 ./...`). Go is `1.26.1` (see `go.mod`).
- Clean Architecture core is implemented: `internal/domain` (entities), `internal/ports`
@@ -94,13 +96,19 @@ The repo is **green and shipping** — last tag `v0.9.0`. `task check` passes (f
green against it.
- Adapters present under `internal/adapters/`: `youtube` (captions-first `VideoSource`,
timedtext/InnerTube acquisition per ADR-010), `summarizer` + `llm` (the copied AI router,
Primary→Fallback per ADR-004), `store` (Postgres, golang-migrate migrations 001006),
now a resilient endpoint chain — local primary → local fallback → external worst-case,
parse-failure-aware, ADR-004 + ADR-022), `store` (Postgres, golang-migrate migrations 001015),
`secrets` (file-backed `SecretStore`). The brain HTTP sink (ADR-005) is the remaining
optional sink.
- Stage 1 is open (ADR-012): multi-user with **DB-enforced** isolation — Postgres RLS `FORCE`d
on all user-owned tables (migration 003), two-user isolation test in
`internal/adapters/store/rls_test.go`. Registration gate, per-user YouTube web connect, and
account management (disconnect / delete, ADR-013) all shipped.
- Transcript persistence (ADR-021, migration 015): transcripts are a **shared, non-RLS** store
keyed by `(provider, provider_video_id)` — the single exception to the isolation boundary
(`TestTranscriptsTableIsSharedNotRLS`). The engine reads stored transcripts before any caption
fetch (`usecase.resolveTranscript`), so re-analysis — re-summarize, paste-a-URL, onboarding
burst — never re-touches YouTube. Per-user summaries/videos stay RLS-scoped.
- `cmd/tapir` subcommands: `list`, `show`, `auth` (interactive host-side OAuth), `run` (batch
watch→summarize), `serve` (the HTMX+Templ web reader/writer under `internal/web`, a new
transport over the unchanged engine/ports — ADR-003). `tapir env` prints config.
+573 -1
View File
@@ -743,6 +743,577 @@ collapse keys off the same window (`App.RecencyWindow=0` → everything inline).
---
## ADR-021 — Persist transcripts as shared, video-keyed public content (re-analysis never re-fetches)
**Status:** Accepted (2026-06-09). **Reopens the transcripts half of** the "Global cross-tenant
`videos`/`transcripts` table" rejection (data-model.md). **Builds on ADR-007** (captions-first),
**ADR-010/ADR-014** (the per-IP caption rate gate), and **ADR-012** (per-user RLS isolation).
**Context.** Every summarization fetches the transcript fresh through the caption path, even when
the exact same transcript was fetched moments ago — for the same user re-summarizing, or for a
second user who happens to watch the same video. The caption fetch is the one genuinely scarce,
genuinely risky operation in the system: YouTube's timedtext endpoint is unofficial and per-IP
rate-limited (ADR-010), and tripping it risks the maintainer's Google standing (ADR-014). So the
operation we most want to *avoid repeating* is the one we currently repeat unconditionally. A
transcript is **public content** — the same words YouTube serves to anyone — and carries nothing
user-identifying. The per-user isolation that protects summaries, feeds, and tokens (ADR-012) is
the wrong shape for it: it forces a re-fetch per user for data that is identical across users.
The original rejection ("Global cross-tenant `videos`/`transcripts` table") bundled videos and
transcripts together and rejected both on the grounds that "at 15 users, re-summarizing is
cheaper than the coupling." That reasoning holds for **videos** (per-user feed rows, genuinely
user-scoped) but not for **transcripts**: the cost being avoided is not LLM re-summarization, it
is a *rate-gated, reputation-risky network fetch*, and that cost is paid per re-fetch regardless
of user count. One re-fetch avoided is strictly worth more than the coupling it removes.
**Decision.**
1. **A single shared `transcripts` table, keyed by the cross-user dedup key
`(provider, provider_video_id)`** — the stable public identity of the video, not Tapir's
internal per-user `videos.id`. Columns: the key, `source` (`captions`/`none`), `language`,
`content`, `fetched_at`. It holds **only public caption content + the video's public id**
nothing user-identifying — and is therefore **NOT RLS-scoped**: no `user_id`, no policy, no
`FORCE ROW LEVEL SECURITY`. This is the deliberate, single exception to the ADR-012 isolation
boundary, and the only one.
2. **Summarize path becomes read-stored-first.** Have a stored transcript for this video? →
summarize from the stored text, **no caption fetch**. No stored transcript? → fetch *through
the unchanged gate* (ADR-014) → store it → summarize. The gate is neither bypassed nor
weakened; persistence reduces how *often* we reach it, never how *fast*.
3. **De-facto cross-user dedup is the intended behaviour, not a feature with a switch.** Two
users who share a video share the one transcript row. A permanent `source = 'none'` (no
captions) is stored too, so a known-caption-less video is not re-fetched by anyone. A
transient 429 (`SourceRateLimited`) is **never** stored as terminal — it stays a per-user
retry via the existing `transcript_status` backoff (ADR-014), so persistence cannot mask a
rate-limit into a false "no transcript."
4. **Per-user `summaries` stay RLS-scoped (ADR-012 unchanged)** and reference the transcript by
video id. Videos stay per-user. Only transcripts go shared.
**Consequences.** Re-analysis (re-summarize, different model, paste of an already-seen video,
onboarding of a second user with overlapping subscriptions) never re-touches YouTube — the
primary win, and it *reduces* aggregate caption-gate pressure, reinforcing ADR-010/ADR-014 rather
than straining them. The isolation surface gains exactly one non-RLS table; an isolation test
asserts the boundary is *exactly* there and has not leaked to any user-owned table (this is the
proof the public-content classification was implemented as designed). It also unblocks
multi-model / customizable analysis (re-run analysis on stored text for free) — enabling that is
this ADR's point; building it is separate.
**Reversibility.** The read-stored-first check is the only behavioural coupling; removing it
restores fetch-every-time. The down-migration recreates the per-user RLS-scoped transcripts shape
(001/003). No user-facing surface depends on cross-user sharing — sharing is the *storage shape*,
never exposed in the UI.
---
## ADR-022 — Summarizer is a resilient endpoint chain, not a single model
**Status:** Accepted (2026-06-10). **Extends ADR-004** (the copied `llm` Primary→Fallback
routing). Triggered by the first friendly-pilot live run, where a connected user got **zero**
summaries after 12h.
**Context.** Stage-0 ran a single summarizer model (`koala/phi4-mini`) with no fallback wired
(`summarizer.New(primary, nil)`). The live run exposed three independent failure modes, each of
which silently produced no summary:
1. **Context overflow.** `phi4-mini` has an 8k context. Real transcripts (one was 11,602 tokens)
exceed it and the gateway returns HTTP 400 — and the request also sent `max_tokens=8192`, so
even a short transcript plus the completion budget could overflow the window.
2. **Malformed model output.** `phi4-mini` intermittently emits `highlights` as a bare string
instead of an array, producing `cannot unmarshal string into []string`. The old code returned
the parse error **without** trying any other model — a 200-with-bad-JSON short-circuited.
3. **No fallback existed at all** — any primary failure was terminal for that video.
`phi4-mini` is kept as primary deliberately: it is fast and, on transcripts that fit, correct.
The fix is resilience around it, not replacing it.
**Decision.**
1. **Ordered endpoint chain (`summarizer.NewChain`).** Endpoints are tried in order; the first to
return a *parseable* summary wins. Default chain:
`koala/phi4-mini` (primary, local) → `iguana/gemma4-26b` (fallback, local on a
*different host*) → `berget/mistral-small` (worst-case, external). All three are reached through
the **one** LiteLLM gateway by alias — the gateway already fronts both llama-swap and berget — so
a fallback is a different alias, not a second client config.
**Update 2026-06-11:** the local fallback moved from `koala/phi4-14b` to `iguana/gemma4-26b`.
koala now carries other GPU loads, so keeping the fallback on koala competed with them; iguana
(M2 Ultra) has the headroom, and a different host is also a different egress IP for the rare
fallback fetch. `gemma4-26b` is the brain-validated homelab general-purpose model (agentsquad
H2/H3 executor) and returned valid summary JSON on the real prompt in a smoke test
(~37s incl. cold-load — fine for a path hit only when the fast primary fails). Pure config:
`TAPIR_FALLBACK_MODEL`.
2. **A parse failure advances the chain, same as a transport error.** "Reliably summarized" means
*parseable summary returned*, not *HTTP 200*. This is the behaviour the old Primary→Fallback
shape missed.
3. **Tolerant parse.** `highlights`/`takeaways` coerce from a bare string (or a mixed scalar
array) to `[]string`, so the most common small-model quirk is absorbed **without** spending a
fallback round-trip — keeping the fast path fast.
4. **Transcript truncation (`TAPIR_MAX_TRANSCRIPT_CHARS`, default 18000).** Input is bounded
up-front to fit a small-context primary, so overflow is prevented rather than recovered-from.
5. **Bounded completion budget (`TAPIR_SUMMARY_MAX_TOKENS`, default 1500).** A summary needs few
hundred tokens; the old 8192 budget itself contributed to 8k-window overflow.
**Local-first guarantee preserved.** The chain ordering *is* the guarantee: locals are tried
first, so content reaches the external endpoint only after every local endpoint has failed.
`TAPIR_CLOUD_FALLBACK_MODEL=""` removes the external endpoint entirely — the lever a
**client/NDA deployment** pulls so content never leaves the local stack. With no external endpoint
configured the `ai_routing.feature` "content only local" scenarios hold unchanged.
**Reversibility.** Pure wiring + config. Setting `TAPIR_FALLBACK_MODEL` and
`TAPIR_CLOUD_FALLBACK_MODEL` empty collapses the chain back to single-primary behaviour; the
tolerant parse and truncation are strict supersets of the old behaviour (a previously-parseable
reply still parses; a transcript within budget is unchanged).
---
## ADR-023 — Drop Shorts/livestreams at discovery to protect the caption budget
**Status:** Accepted (2026-06-10). **Builds on ADR-014** (per-IP caption rate limit is the
binding constraint) and the ADR-022 live-run findings.
**Context.** The scarce resource is the unofficial timedtext caption fetch (per-egress-IP
429, ~3 successful/pass). The first multi-user run showed the candidate set was mostly noise —
Shorts, sub-minute clips, and live broadcasts — each of which still consumes a caption-fetch
attempt (and a "none" result is a *completed* fetch, so it costs budget even when it yields
nothing). Spending the rate-limited budget on content the user will not read is the waste to
cut first; it is cheaper and lower-risk than raising the ceiling (multi-IP, Whisper).
Duration and live status are NOT in the playlistItems discovery response, but they ARE in the
Data API `videos.list` (contentDetails.duration + snippet.liveBroadcastContent) — the official
**quota-based** API (1 unit/call, 50 ids/call), which is a *different* limit from the timedtext
429. So one cheap quota call buys a filter that saves many expensive throttled fetches.
**Decision.**
1. `NewVideos` enriches its candidates with a single `videos.list` call and drops, before
returning: videos shorter than `TAPIR_MIN_VIDEO_SECONDS` (default 60) and any `live`/
`upcoming` broadcast. Dropped videos are never persisted, so they also declutter the list.
2. The filter is **degrade-open**: `MinVideoSeconds=0` disables it (no quota call), and a
`videos.list` error returns the candidates unfiltered — discovery must never break because a
metadata call hiccuped (worst case = pre-ADR-023 behaviour).
3. The paste-a-URL path (`VideoByID`) is **not** filtered — an explicit user request for a
specific video (even a Short) is honoured.
**Reversibility.** Pure discovery-time filter + config. `TAPIR_MIN_VIDEO_SECONDS=0` restores
the old behaviour. No schema change, no effect on already-stored videos.
**Quota note.** Per-channel enrichment adds ~1 unit/channel/pass. At pilot scale (≤3 users)
this is well under the 10k/day cap; at larger scale, batch `videos.list` across channels
(50 ids/call) by collecting all discovered ids per pass before enriching.
---
## ADR-024 — Per-channel caption-availability memory
**Status:** Accepted (2026-06-10). **Builds on ADR-014** (per-IP caption budget), **ADR-021**
(shared transcript cache), **ADR-023** (Shorts filter).
**Context.** After ADR-021 caches transcripts and ADR-023 drops Shorts, the remaining caption
waste is the *first* fetch on every new video of a channel that never publishes English captions
(foreign-language news, music, etc.). Each costs one rate-limited fetch to resolve to "none" —
and on a throttled IP that fetch may 429 and churn the backoff machinery before it ever gets a
verdict. A pilot user's feed had several such channels.
**Decision.** Remember, per `(user, channel)`, a streak of consecutive no-caption outcomes
(`channel_caption_state`, migration 016, RLS-scoped like the rest of the user-owned schema).
Once the streak reaches `TAPIR_CHANNEL_CAPTIONLESS_THRESHOLD` (default 5) the channel is
suppressed — its videos are discovered/listed but not caption-fetched — for
`TAPIR_CHANNEL_CAPTIONLESS_WINDOW` (default 14d), after which one video is re-probed
(auto-recovery for a channel that starts adding captions). A successful fetch resets the streak;
a fresh 429 does NOT count (transient, not a caption verdict). An explicit manual request
bypasses suppression. `threshold = 0` disables the feature.
**Why per-user, not global.** Caption availability is really a channel property (public), so a
global table would let users share the learning. But subscriptions are per-user (ADR-012) and at
pilot scale users' channel sets barely overlap, so per-user + RLS keeps it consistent with the
existing isolation model with no new non-RLS exception to justify. Promoting to a shared table
(like transcripts, ADR-021) is a future optimisation if channel overlap grows.
**Reversibility.** Migration 016 is a clean drop; `threshold = 0` disables at runtime. The
memory only ever *suppresses fetches* — it never deletes content or affects already-stored
summaries.
---
## ADR-025 — Honest, state-aware foreground summarization status
**Status:** Accepted (2026-06-10). **Pillar B of the manual-mode UX work** (Pillar A, foreground
fetch priority, is a separate follow-up). Builds on ADR-014 (the rate limit the UX must make
legible).
**Context.** Clicking "Summarize" spawned a background goroutine and polled `/status`, which
returned only two states: the spinner (in-flight) or the normal card (done). But the web
`ProcessVideo` only recorded an outcome on *success* — a 429'd or caption-less click left
`transcript_status` unset, so the next poll silently reverted to the "Summarize" button. The
user saw either an endless spinner or a button that did nothing useful when clicked again. The
binding constraint (YouTube's caption rate limit) was completely invisible.
**Decision.**
1. **Record every outcome on the web path**, mirroring the runner: `ProcessVideo` stamps
`rate_limited` / `none` / `fetched`. A rate-limited video keeps its requested flag so the
background sweep retries it; `none` and `fetched` are terminal.
2. **`/status` is state-aware**: summarized → summary card; in-flight → working spinner;
`rate_limited` → a calm "waiting, will retry" card that keeps polling (every 30s) so the
summary appears on its own when the retry lands — the user never clicks again;
`none` → a terminal "no captions" card with no poll and no dead-end button.
3. **Charm status text** (Claude-Code / Crush inspired): the working spinner cycles playful,
tapir-themed gerunds ("Chewing the cud…", "Munching leaves…", "Distilling the gist…") via
CSS only — no JS, keeping the HTMX/no-JS ethos. Decorative (aria-hidden) with a stable
`role=status` line for assistive tech.
**Principle.** When the system cannot be fast (throttled IP), it is at least honest, and it
self-resolves without making the user retry. Honesty is the load-bearing half — Pillar A's
priority lane only improves the odds of a fast slot; it cannot beat an already-hot IP.
**Reversibility.** Pure transport-layer + view change over the unchanged engine/ports. No
schema change (reuses `transcript_status` from migration 007).
---
## ADR-026 — Foreground caption fetches take priority; the credentials probe is dead
**Status:** Accepted (2026-06-10). **Pillar A of the manual-mode UX work** (Pillar B was
ADR-025). Builds on ADR-014 (the shared per-IP gate).
**Context.** Every caption fetch — the background sweep and the web click-path — shared one
process-wide rate gate equally. So a user waiting on a "Summarize" click competed with the
firehose for both pacing and the scarce pre-429 window; on a busy IP the click was slow or
429'd while the background churned.
**Decision.** A context-marked priority lane. The web path
(`engineProcessor.ProcessVideo`) wraps its context with `ForegroundContext`; the gate gives
foreground fetches a token immediately, while **background fetches yield** — they wait until no
foreground fetch is pending before taking a token. Threaded via a context value (not new
signatures) and a process-wide `foregroundPending` counter. Clicks are rare and bursty, so the
background barely loses throughput; the waiting human gets the next (and cleanest) slot.
**Credentials probe — rejected, not built.** The idea was to fetch captions with the user's
auth in manual mode to dodge 429s. It is a dead end, already settled by ADR-010 and the code:
the caption path is *deliberately anonymous* because the InnerTube/timedtext endpoints **reject
or break on authenticated requests** (`captions.go`: "no OAuth token — it can break the
timedtext endpoint"). The user's OAuth (a Data API credential) does not authenticate InnerTube
at all, and the official `captions.download` is owner-only (403 on third-party). So auth cannot
help here and can actively hurt. No probe needed — building one would only re-confirm the ADR.
**Reversibility.** Context-marker + a yield loop in the gate; removing the marker collapses to
the prior equal-share behaviour. No schema or API change.
---
## ADR-027 — Chat with a video's stored transcript (deeper-dive, on an already-summarized video)
**Status:** Accepted (2026-06-11). **Consumes ADR-021** (the shared, video-keyed transcript
store) for the first time beyond summarization; **uses the ADR-022 chain models**; relates to
ADR-012 (isolation) and ADR-016 (the Stage-0 gate).
**Context — observed demand, not hypothetical.** The maintainer read 10+ real pilot summaries
and reported the reactions: *many good; some he wanted to dig deeper into; some less useful*
(the "less useful" split between weak-*model* output and uninteresting-*video* content). The
middle reaction is the signal: a good summary that makes the reader want *more* is the summary
succeeding at triage and then hitting a wall — there is nowhere to go deeper short of watching
the video. That want is the feature. It is also the cheapest possible feature to satisfy
honestly, because ADR-021 already persists the transcript: the deeper-dive runs entirely on
stored public-content text + local models, touching **no** caption fetch and **no** YouTube.
**Decision.** Add a per-video chat that lets the user ask questions against a video's stored
transcript.
1. **Entry from the summary view only.** A "dig deeper / ask" affordance on a summarized video —
the chat lives exactly where the "I want more" reaction happens. No standalone chat surface.
2. **Stored-transcript-only (load-bearing constraint).** Chat is available **only** for videos
that already have a stored transcript. It never triggers a caption fetch, so it cannot touch
the rate gate, the 429 surface, or YouTube at all — the entire account-safety constraint that
governs the rest of Tapir is satisfied *by construction* here, not by careful gating. (Entry
being "from a summary" guarantees the transcript exists.) On-demand fetch for un-stored videos
is explicitly deferred.
3. **Model = the summary's model by default; user-switchable among the ADR-022 chain models**
(`phi4-mini` / `gemma4-26b` / `mistral-small` to start). This is deliberate: it doubles as
live model-comparison instrumentation — ask the same question of the same transcript under two
models and the difference is directly felt. This is the mechanism by which the maintainer
learns *which* model is worth defaulting to, and it is the multi-model-analysis direction
ADR-021 anticipated, arriving as a user-facing capability.
- Chat is a **read-bounded retrieval/QA task** (the user supplies the focus), which is
*easier* than summarization (the model must decide what matters). So a model that summarizes
mediocrely may chat well — chat is plausibly a partial remedy for the weak-summary case, not
an inheritor of it.
4. **Ephemeral chat (v1).** No persisted history; chat is per-session. Persisting per-user,
RLS-scoped history is deferred until there is evidence anyone wants to revisit a conversation.
5. **Chat-only, trust-the-model (v1) — with a recorded limitation.** The chat does not expose the
raw transcript for verification in v1 (kept simple). **Known limitation:** because some
summaries were weak-model output, the user has reason not to fully trust a chat answer's
fidelity to the transcript, and v1 gives no in-UI way to check. The model-switcher partially
compensates (two models disagreeing on the same question is itself a signal). A
"show source / view transcript" verification path is the natural **v2** and is *not*
foreclosed — ADR-021's stored transcript already makes it cheap. Recorded so v2 is a known
next step, not a rediscovery.
**Why this is safe and in-scope.** It adds no caption-fetch surface (stored-only), no new
non-RLS table (transcripts already shared per ADR-021; ephemeral chat stores nothing), and no
auth change. It is additive to the read path. The one genuine product expansion — Tapir becomes
an interactive transcript-QA tool, not only a summarizer — is justified by *observed* demand from
real reading, which is exactly the kind of evidence the Stage-0 discipline asks for before
building.
**Relation to the Stage-0 gate.** This is **not** a return-nudge (those stay deferred, ADR-020) —
it adds nothing that prompts the user to return; it deepens the value *once they are already
reading*. It does not contaminate the unprompted-return signal. If anything it strengthens the
"useful to me" case the gate measures, by giving a good summary somewhere to lead.
**Reversibility.** Additive read-path feature over the unchanged engine + the ADR-021 store.
Removing the summary-view affordance removes the feature; nothing else depends on it. Ephemeral =
no migration, no stored state to unwind. Spec: `docs/specs/chat-with-transcript.md`.
---
## ADR-028 — Onboarding burst: pick likely-good videos, summarize them with a stronger model
**Status:** Accepted (2026-06-11). **Refines ADR-018** (the connect-time burst) and **ADR-020**
(recency-bounded auto-summarize). **Builds on ADR-022** (the endpoint chain), **ADR-023**
(the discovery-time `videos.list` enrichment), and **ADR-021** (the shared transcript cache).
Triggered by a Phase-1 investigation of the live pilot DB.
**Context.** A new user's first session decides whether they return (the Stage-0 gate, ADR-016).
The connect-time burst (ADR-018: summarize ≤`TAPIR_ONBOARD_SUMMARIZE_COUNT` newest videos so the
feed isn't empty) *fires* in production, but a live-DB investigation of the second pilot user
("Jonte") found it delivers a weak first impression for two reasons, and ruled out a third idea:
1. **Junk picks.** Selection was pure newest-first (`NewestUnsummarizedVideoIDs`,
`ORDER BY published_at DESC`) with **no quality signal**. Jonte's live burst-3 were a
stock-ticker **livestream** + two regional news clips — newest, not best. The cheap signals
that *could* gate this (duration, live status) are fetched by ADR-023's `videos.list`
enrichment at discovery and then **thrown away**: the `videos.duration_s` column (migration
001) was never written.
2. **Weakest model on the first impression.** All burst summaries ran on `koala/phi4-mini` — the
documented weak link (ADR-022 was born from its failures). The stronger, brain-validated
`iguana/gemma4-26b` was never used, even though the burst is only ~3 summaries.
3. **Cached-first instant summaries — REJECTED.** The idea: skip the fetch, summarize
already-cached transcripts (ADR-021) instantly. The pilot numbers kill it — only **11 videos**
overlap between the two users (~3% of each ~350400-video library), **0** cached-and-
unsummarized, and a new user's newest-20 unsummarized are **20/20 NOT cached**. Newest-first
and cached-first are structurally incompatible: fresh uploads are exactly what nobody has
fetched. An empty lever at pilot scale.
**Decision.**
1. **Persist `duration_s` at discovery.** `filterLowValue` (ADR-023) already has each candidate's
duration in hand; carry it onto the kept `domain.Video` and have `UpsertVideo` write it,
COALESCE-preserving a known value (the channel-title backfill stance, migration 014). No new
migration — the column exists. The connect-triggered discovery pass runs *before* the burst,
so a fresh user's candidates are enriched in time.
2. **Junk-avoiding selection.** A new `OnboardBurstVideoIDs(userID, limit, minSeconds, maxSeconds)`
keeps the newest-first order but drops a video when its duration is *known* and outside
`[minSeconds, maxSeconds]``minSeconds` = `TAPIR_MIN_VIDEO_SECONDS` (60, the Shorts floor),
`maxSeconds` = new `TAPIR_ONBOARD_MAX_VIDEO_SECONDS` (default 14400 = 4h, to drop multi-hour
livestream VODs that pass the live filter once ended). A NULL duration is **unknown** — kept
(degrade-open) but ranked after known-good rows. **has-captions stays un-gateable pre-fetch**
(only knowable after a gate fetch or a ~0-probability cache hit); selection only *avoids
known-junk*, it does not *promise* captions.
3. **Stronger model for the burst only.** `TAPIR_ONBOARD_SUMMARIZER_MODEL` (default
`iguana/gemma4-26b`) leads a burst-specific summarizer chain (onboard model first, then the
standard ADR-022 chain as resilience, deduped), wrapped in a burst-specific processor over the
*same* store/cache/sink — a pure wiring choice; the engine and ports are unchanged (ADR-003).
Empty or equal-to-primary collapses the burst back onto the shared processor.
**Not a throughput change.** The caption rate gate (ADR-014) and the foreground priority lane
(ADR-026) are untouched — same pacing, same cap. This changes *which* ≤3 videos the burst spends
its fetches on and *which model* summarizes them, never how fast or how many. The engine's
existing read-stored-first (ADR-021) is unchanged and still yields a free instant summary on the
rare cache hit — we simply do not *select* for cache hits.
**Consequences.** Better odds of a strong first session: the burst avoids the obvious junk and
runs the better model on the one impression that decides return. The selection improvement is
forward-looking — existing rows have NULL `duration_s` until their next discovery pass backfills
it (lazy, like channel_title); a brand-new user benefits immediately because connect-discovery
runs first. `duration_s` becoming live also unblocks future length-aware features (feed sorting,
"long read" badges) for free.
**Reversibility.** Pure config + wiring + one column write + one query, no migration.
`TAPIR_ONBOARD_MAX_VIDEO_SECONDS=0` (and `TAPIR_MIN_VIDEO_SECONDS=0`) restores pure newest-first;
`TAPIR_ONBOARD_SUMMARIZER_MODEL=""` collapses the burst back to the shared processor.
Spec: `docs/specs/onboarding-wow-burst.md`.
---
## ADR-029 — Stateless session cookie (survives restarts, browser-close, idle)
**Status:** Accepted (2026-06-11). Triggered by pilot feedback: "lots of clicking to log in again
on iPhone." Supersedes the in-memory session store in the ADR-011 login.
**Context.** Three compounding causes made users re-login constantly:
1. **In-memory session store** (`sessionStore` map) — wiped on every pod restart, so each deploy
logged everyone out. During the active build period that was ~15 logouts.
2. **1-hour session TTL** — for a "check back tomorrow" reader, idle > 1h forced a re-login on
nearly every visit.
3. **No cookie Max-Age** — a session cookie (deleted on browser/app close); iPhone Safari closing
the tab dropped it.
Each re-login is the full Dex/Authentik redirect dance — many taps on mobile.
**Decision.** Make the session **stateless**: the identity (subject + email) and an absolute
expiry live INSIDE the existing HMAC-signed (HS256) cookie — no server-side table. Plus:
- **30-day sliding TTL** (was 1h), re-signed on each request so an active user never lapses.
- **Persistent cookie** (`Max-Age` set) so it survives browser/app close.
The cookie is HttpOnly + Secure + SameSite=Lax; the HMAC (keyed by the stable ESO
`tapir-session-secret`, which does NOT rotate per deploy) makes it tamper-proof. The payload is
identity, not secrets — the OIDC access/ID tokens are still discarded after callback.
**Consequences.** A deploy/restart no longer logs anyone out (proven by a test: a cookie issued by
one instance is accepted by a fresh instance with the same secret); works across replicas for
free. **Trade:** no server-side revocation — `logout` clears the cookie client-side, but a copied
cookie stays valid until expiry. Accepted for the Stage-0 reader pilot; revisit (server-side
revocation list, or shorter TTL + refresh) if it ever holds sensitive actions. Rotating
`tapir-session-secret` invalidates all sessions — the global logout lever.
**Not addressed here:** the tap-count of the IdP login page itself is Authentik's UX; with
re-login now rare (30-day idle or explicit logout), it matters far less.
---
## ADR-030 — Observability: slog timing + Prometheus metrics (AI-focused)
**Status:** Proposed (2026-06-11). Issue #15. **Draft for review — no code yet.**
**Context / requirements.** Nothing measures the activities that drive Tapir's performance and
UX, and the Stage-0 eval gate (ADR-016) needs a *performance* dimension to sit beside the
return-usage one. We need timing for: caption fetches (the scarce op), summarization (which model
won, how long, fallbacks), Q&A latency, LLM token spend, and basic session/usage (request rate,
latency by route, logins). Requirements:
- R1: structured `slog` timing at each AI call site (human-readable, already the logging stack).
- R2: Prometheus metrics for the same, scrapeable by the cluster's prometheus-operator.
- R3: **AI metrics are the priority** — summarize latency by `model`/`outcome`/`fallback`,
caption-fetch latency by `outcome`, chat latency by `model`, and LLM `tokens` by model+kind.
- R4: HTTP/session metrics via middleware — request count + latency by route, logins.
- R5: bounded label cardinality (no per-user, no raw-path labels).
- R6: `/metrics` must NOT be publicly exposed.
**Decision / architecture.**
1. **New package `internal/metrics`** owns all Prometheus collectors + a typed API
(`ObserveSummarize`, `ObserveCaptionFetch`, `ObserveChat`, `RecordTokens`, `IncLogin`,
`HTTPMiddleware`, `Handler`). Adapters call this API; they never import prometheus types.
2. **New dependency `github.com/prometheus/client_golang`.** Justification: it is *the* standard
Go Prometheus client and the cluster already runs prometheus-operator; hand-rolling exposition
is not worth it. (Needs the dep-justification note in the commit per repo rules.)
3. **The copied `llm` package stays stdlib-only (ADR-004).** It must not import `internal/metrics`.
Token usage is surfaced via an **optional callback** `llm.WithUsageHook(func(model string, prompt, completion int))`
set at wiring time (`buildSummarizer`/`buildChat`) to `metrics.RecordTokens`; `llm.Client` only
gains parsing of the response `usage` block. Our own adapters (`summarizer`, `youtube`, `chat`)
may import `internal/metrics` directly.
4. **HTTP middleware** reads `r.Pattern` AFTER routing (Go 1.22 sets it during ServeMux match), so
the `route` label is the bounded registered pattern (`GET /v/{videoId}`), satisfying R5;
unmatched → `other`.
5. **Dedicated metrics port** (`TAPIR_METRICS_ADDR`, default `:9090`) served by a second
`http.Server` in `cmdServe`; `/metrics` is never on the public app mux (R6). A **PodMonitor**
in `mathias/infra` scrapes it; the deployment exposes the port.
6. **slog** elapsed fields are emitted alongside each metric at the call sites (R1).
**Hook points (where the instrumentation lands).**
- `summarizer.Summarize` — per-endpoint timing + outcome (`success`/`parse_error`/`error`) + fallback flag.
- `youtube.FetchTranscript` — fetch timing + outcome from `domain.Transcript.Source`.
- `chat.Service` answer — timing by model.
- `llm.Client.Complete` — parse `usage`, fire the usage hook.
- `oidc.handleCallback``IncLogin`.
- `cmdServe` — wrap `Router()` in `metrics.HTTPMiddleware`; start the metrics server.
**Out of scope / later.** Persisting per-summary latency into Postgres for `tapir report`
(derive UX latency — publish/discovery → summary — from existing timestamps first; only persist
op-latency if the scrape proves insufficient). SPA view (#16) and visual refresh (#17).
**Reversibility.** Additive: a new package + middleware + a metrics port. Removing the PodMonitor
stops scraping; the app is unaffected. No schema change.
**Next steps (gated):** on approval of this ADR → BDD scenarios (`docs/use-cases/observability.feature`
+ scenario-coverage map) → TDD → implement → SemVer + docs + PodMonitor.
---
## ADR-031 — SPA-like reader: inline-expand summary + Q&A in the list (HTMX, no framework)
**Status:** Proposed (2026-06-12). Issue #16. **Draft for review — no code yet.**
**Context / requirements.** The reader is multi-page: a list of compact cards (`/`), then a
navigation to a separate detail page (`/v/{id}`) for the full summary + the docked chat (ADR-027).
It feels less fluid than a single integrated view. We want the full summary AND the per-video
Q&A to open **in place in the list**, no page hop. Requirements:
- R1: clicking a summarized card expands it in place to the full summary (summary/highlights/
takeaways) + the chat dock; a collapse returns it to the compact card.
- R2: **no SPA framework** — stay HTMX + Templ (ADR-003); reuse existing fragments, not a rewrite.
- R3: **progressive enhancement** — with JS off, the card link still navigates to `/v/{id}`
(the detail page stays as the no-JS + deep-link surface). Nothing becomes JS-only.
- R4: only **summarized** cards expand; pending/rate-limited/no-caption cards keep their current
footer behaviour (Summarize button, waiting/none states).
- R5: chat inside an expanded card works exactly as on the detail page (reuse `chatReveal`/
`chatSection` + the existing `/v/{id}/chat` endpoints, unchanged).
**Decision / architecture.**
1. **Reuse the existing fragments.** `summaryBody(r)` and `chatReveal(videoID)` already exist and
render the detail page; a new `expandedCard(r, chatEnabled)` composes the compact header + a
collapse control + `summaryBody` + `chatReveal`. `DetailPage` is refactored to also compose
`summaryBody` so the two never drift (DRY).
2. **Two fragment endpoints** (mirroring the existing list/status HTMX fragment pattern):
`GET /v/{videoId}/expand``expandedCard`; collapse reuses the existing compact `VideoCard`
via `GET /v/{videoId}/card`. Both are list-card `<li>` fragments with the SAME `id`
(`video-{id}`), swapped `outerHTML` — same mechanism as `processingCard`/`VideoCard` today.
3. **The compact card's title/"Read" affordance** becomes `hx-get=/v/{id}/expand`,
`hx-target=#video-{id}`, `hx-swap=outerHTML`, with `href=/v/{id}` as the no-JS fallback (R3).
The expanded card's collapse control is the inverse (`hx-get=/v/{id}/card`).
4. **Only when `r.Summarized`** does the expand affordance render (R4); the other states are
unchanged.
5. **v1 does NOT push the URL** (`hx-push-url`) — expand/collapse is ephemeral list UI state; the
detail page remains the deep-link/shareable URL. Deep-linking the open state via `hx-push-url`
is noted as a later option (needs list-state restore on back).
**Out of scope / later.** URL push / deep-linkable open state; the visual refresh (#17) — though
the expanded-card markup is where #17's TUI/charm styling will land, so they pair.
**Reversibility.** Additive: two fragment endpoints + one templ + an affordance swap on the
compact card. Removing the affordance reverts to plain list→detail navigation; the detail page is
untouched. No schema change.
**Next steps (gated):** on approval → BDD (`docs/use-cases/inline_expand.feature` + coverage
map) → TDD → implement → SemVer + docs.
---
## ADR-032 — Visual refresh: one charm-reader layout, light + dark themes
**Status:** Proposed (2026-06-12). Issue #17. **Draft for review — no code yet.** Follows the
sketch-first explore step (3 throwaway mockups in `docs/sketches/`, screenshotted for review).
**Context / decision.** The UI is flat. From the mockups, directions **B (light reader + charm)**
and **C (dark cozy terminal)** are the SAME layout — readable sans body, monospace meta, charm
palette accents, lipgloss-style bordered cards — in two palettes. Direction A (full-monospace TUI)
is dropped as too heavy to read long summaries. Decision: ship that one layout with **both a light
theme (B) and a dark theme (C)**, user-toggleable, defaulting to the OS preference.
**Requirements.**
- R1: one set of markup/structure; the two themes are pure palette (CSS variables), no duplicate templates.
- R2: a **theme toggle** persisted across visits; default to `prefers-color-scheme` when no choice stored.
- R3: charm language in both — mint/purple/pink accents, mono meta + section labels, lipgloss
bordered/gradient cards, the (fixed) ASCII tapir; readable sans body.
- R4: style the existing pieces — list, compact card, **expanded card incl. the already-present
video embed (ADR-031/summaryBody)**, detail page, chat dock, the queue note, the charm spinner.
- R5: WCAG-AA contrast for body text in BOTH themes; keep `prefers-reduced-motion` (already honoured).
- R6: stay HTMX+Templ; no CSS framework.
**Architecture.**
1. **Palette as CSS variables.** `:root` holds the light (B) tokens; `:root[data-theme="dark"]`
holds the dark (C) tokens; a `prefers-color-scheme: dark` media block sets the dark tokens when
no explicit `data-theme` is set. All component CSS references variables only (R1). The existing
`CharmMint/Purple/Pink/Cream/Dim` Go consts remain the source for the spinner's inline colours.
2. **Theme toggle** = a small inline script (a dozen lines, no framework) in `Layout`: on load,
apply stored theme (localStorage) or fall through to the media query; a header toggle button
flips `data-theme` on `<html>` and stores it. This is the one new bit of JS; everything else
stays server-rendered + HTMX. (Considered: cookie + server-render — rejected, a full round-trip
per toggle is clunky for a pure presentation flip.)
3. **Scope = the `styleTag` CSS in `view.go`** (the single style source) plus tiny class hooks in
the templ where needed; content/structure are unchanged, so existing view tests keep passing.
**Verification.** The deployed UI is behind auth (web-shot can't log in), so visual review is via
the mockups now + a styled full-set mockup screenshot before merge, then a live eyeball on device.
Automated tests stay structural/behavioural (theme tokens present, toggle persists, expanded card
embeds the video, no-JS still renders a readable default) — colours are not unit-tested.
**Reversibility.** A CSS theme swap + one small script + a few class hooks; revert `styleTag` to
roll back. No schema, no structural change.
**Next steps (gated):** on approval → BDD (`docs/use-cases/visual_theme.feature` + coverage map)
→ TDD → implement → SemVer + docs.
---
## Rejected alternatives
Approaches considered during the 2026-06-02 planning + grill session and **deliberately not
@@ -757,7 +1328,7 @@ maps to the ADR that settles it.
| Lifting shared packages into a `brain-common` module | Couples Tapir's release cycle to the monolith for negligible code savings | ADR-004 |
| Importing/replicating the filesystem `brain` package | Assumes co-location with the brain git checkout; wrong for a standalone networked service | ADR-005 |
| Reusing `ingestion`'s `oauth` package for YouTube/Vimeo | Same name, opposite direction — it's inbound MCP-server auth, not outbound provider OAuth | ADR-006 |
| Global cross-tenant `videos`/`transcripts` table (dedup) | Reintroduces the cross-domain DB coupling the homelab review is removing; at 15 users, re-summarizing is cheaper than the coupling | data-model.md |
| Global cross-tenant `videos`/`transcripts` table (dedup) | Reintroduces the cross-domain DB coupling the homelab review is removing; at 15 users, re-summarizing is cheaper than the coupling. **Transcripts half reopened by ADR-021** — the avoided cost there is a rate-gated, reputation-risky *caption fetch*, not LLM re-summarization, so it outweighs the coupling; **videos stay per-user.** | data-model.md, **ADR-021** (transcripts only) |
| Audio-download + Whisper STT in the core path | ToS-grey, breakage-prone (yt-dlp), contends for koala GPU with the JEPA PoC; captions alone test the core hypothesis | ADR-007 |
| Building multi-tenant SaaS / Google OAuth verification now | "Real users soon" was lowered to Future B; SaaS machinery before the Stage 0 self-use gate is the primary documented anti-goal | ADR-008, VISION |
| Delegating the S5 reuse spike to an agent swarm | A 1-hour sequential read-and-judge with a single coupled conclusion; orchestration overhead exceeds the work, and it's Diamond-1 judgment the maintainer wanted to own | (process note) |
@@ -766,6 +1337,7 @@ maps to the ADR that settles it.
| Feedback-based Stage 0 gate (friends saying it's useful) | Politeness bias makes asked-for feedback the least reliable signal; return-usage is the real test | ADR-016 |
| Reverse the Dex-write invite flow (Google OIDC only) | Some intended Future-B users won't use Google; OIDC-only leaves them with no onboarding path — invite flow is load-bearing | ADR-017 |
| k8s CronJob for scheduled discovery (vs in-process) | At Future-B scale the in-process scheduler is simpler to deploy; CronJob's failure-isolation benefit was weighed and traded away knowingly (revisit if >1 replica or load grows) | ADR-018 |
| Cached-transcript-first onboarding burst (instant, zero-fetch picks) | Live pilot DB: ~3% cross-user video overlap, 0 cached-and-unsummarized, a new user's newest-20 are 20/20 uncached — newest-first and cached-first are structurally incompatible. Empty lever at pilot scale | ADR-028 |
If a future case genuinely reopens one of these, that's a new ADR superseding the relevant one —
not a silent reversal.
+26
View File
@@ -0,0 +1,26 @@
# Vocabulary (hand-written pilot 2026-06 — will be generated by langgen in Phase 2)
| Term | Means | Never say |
|---|---|---|
| transcript | Captions-first text of a video; the source for summarizing. Shared store keyed by `(provider, video_id)`. | "subtitles file", "the audio" |
| video | Per-user record of a seen video; the unit of work. Not globally deduped. | "the shared video" |
| subscription | A watched channel on a user's connection that Tapir polls for new videos. | "feed", "follow" |
| summary | A video's produced output: summary text + highlights + takeaways + AI provenance. | "transcript" |
| highlight | A notable point pulled from a video (`Summary.Highlights`). | "takeaway" |
| takeaway | An actionable conclusion from a video (`Summary.Takeaways`). | "highlight" |
| sink | A delivery destination for a summary (`store`, `brain`). New one = new adapter. | "the database" |
| AI router | Local-first chain: local Primary → local fallback → external worst-case. | "the API" |
| BYO-AI fallback | User's own external AI key, used only when local fails; opt-in. Without it, content is never sent externally. | "the default AI" |
| brain | Persistent homelab knowledge store; in Tapir, one optional HTTP sink (`brain_ingest`), not the filesystem package. | "the database", "the store" |
| LiteLLM gateway | Local AI gateway (`koala:30401/v1`); the Primary in the AI router. | "piguard:4000", "koala:4000", "the cloud" |
| connection | A connected video account a user authorizes via OAuth; subscriptions hang off it. | "login", "session" |
## Caveman rubric
Before any HIGH/CRITICAL operation, output one line:
`caveman: me <verb> <object>, not <excluded thing>`
Valid iff: (1) only vocabulary terms + plain verbs, (2) a stranger could identify
the exact operation, (3) names one thing explicitly NOT being done.
Examples:
- `caveman: me delete user connection, not the shared transcript`
- `caveman: me send transcript to BYO-AI fallback, not the LiteLLM gateway`
+64
View File
@@ -0,0 +1,64 @@
package main
import (
"strings"
"testing"
"git.d-ma.be/mathias/tapir/internal/config"
)
// TestChatModelsReuseTheChainLocalFirst: the switcher offers the ADR-022 chain in
// order, deduped — primary, local fallback, cloud.
func TestChatModelsReuseTheChainLocalFirst(t *testing.T) {
got := chatModels(config.Config{
SummarizerModel: "koala/phi4-mini",
FallbackModel: "iguana/gemma4-26b",
CloudFallbackModel: "berget/mistral-small",
})
want := []string{"koala/phi4-mini", "iguana/gemma4-26b", "berget/mistral-small"}
if strings.Join(got, ",") != strings.Join(want, ",") {
t.Fatalf("chatModels = %v, want %v", got, want)
}
}
// TestChatCloudModelAbsentWhenDisabled: the local-first / NDA lever — with the
// cloud fallback empty (TAPIR_CLOUD_FALLBACK_MODEL=""), no external model is
// offered in the switcher, so chat content never leaves the local stack (ADR-027,
// honouring ADR-022's "content stays local" guarantee).
func TestChatCloudModelAbsentWhenDisabled(t *testing.T) {
got := chatModels(config.Config{
SummarizerModel: "koala/phi4-mini",
FallbackModel: "iguana/gemma4-26b",
CloudFallbackModel: "",
})
for _, m := range got {
if strings.HasPrefix(m, "berget/") || strings.Contains(m, "mistral") {
t.Fatalf("cloud model %q offered though the cloud fallback is disabled", m)
}
}
want := []string{"koala/phi4-mini", "iguana/gemma4-26b"}
if strings.Join(got, ",") != strings.Join(want, ",") {
t.Fatalf("chatModels = %v, want %v", got, want)
}
}
// TestChatModelsDedup: a config that reuses one alias across slots collapses to a
// single switcher entry (no duplicate options).
func TestChatModelsDedup(t *testing.T) {
got := chatModels(config.Config{
SummarizerModel: "koala/phi4-mini",
FallbackModel: "koala/phi4-mini",
CloudFallbackModel: "",
})
if len(got) != 1 || got[0] != "koala/phi4-mini" {
t.Fatalf("chatModels = %v, want a single deduped entry", got)
}
}
// TestBuildChatNilWithoutGateway: no gateway → no chat backend (routes unmounted,
// read path unaffected).
func TestBuildChatNilWithoutGateway(t *testing.T) {
if c := buildChat(config.Config{GatewayURL: ""}); c != nil {
t.Fatal("buildChat must return nil without a gateway URL")
}
}
+1 -1
View File
@@ -5,7 +5,7 @@ import (
"log/slog"
"sync"
"gitea.d-ma.be/mathias/tapir/internal/runner"
"git.d-ma.be/mathias/tapir/internal/runner"
)
// discoveryRunner runs one user's discovery pass.
+1 -1
View File
@@ -10,7 +10,7 @@ import (
"github.com/stretchr/testify/require"
"gitea.d-ma.be/mathias/tapir/internal/runner"
"git.d-ma.be/mathias/tapir/internal/runner"
)
// serialize must guarantee at most one discovery pass runs at a time, so a
+1 -1
View File
@@ -5,7 +5,7 @@ import (
"fmt"
"os"
"gitea.d-ma.be/mathias/tapir/internal/adapters/store"
"git.d-ma.be/mathias/tapir/internal/adapters/store"
)
// Env var names for the read-only CLI. DSN and user id are never hardcoded — the
+1 -1
View File
@@ -7,7 +7,7 @@ import (
"github.com/stretchr/testify/require"
"gitea.d-ma.be/mathias/tapir/internal/adapters/store"
"git.d-ma.be/mathias/tapir/internal/adapters/store"
)
func TestFormatListColumnsAndOrdering(t *testing.T) {
+1 -1
View File
@@ -9,7 +9,7 @@ import (
"strings"
"text/tabwriter"
"gitea.d-ma.be/mathias/tapir/internal/adapters/store"
"git.d-ma.be/mathias/tapir/internal/adapters/store"
)
// runList prints the user's stored summaries as a table, most recent first.
+62 -20
View File
@@ -23,14 +23,15 @@ import (
"sync"
"time"
"gitea.d-ma.be/mathias/tapir/internal/adapters/secrets"
"gitea.d-ma.be/mathias/tapir/internal/adapters/store"
"gitea.d-ma.be/mathias/tapir/internal/adapters/youtube"
"gitea.d-ma.be/mathias/tapir/internal/auth"
"gitea.d-ma.be/mathias/tapir/internal/config"
"gitea.d-ma.be/mathias/tapir/internal/runner"
"gitea.d-ma.be/mathias/tapir/internal/web"
"gitea.d-ma.be/mathias/tapir/internal/web/oidc"
"git.d-ma.be/mathias/tapir/internal/adapters/secrets"
"git.d-ma.be/mathias/tapir/internal/adapters/store"
"git.d-ma.be/mathias/tapir/internal/adapters/youtube"
"git.d-ma.be/mathias/tapir/internal/auth"
"git.d-ma.be/mathias/tapir/internal/config"
"git.d-ma.be/mathias/tapir/internal/metrics"
"git.d-ma.be/mathias/tapir/internal/runner"
"git.d-ma.be/mathias/tapir/internal/web"
"git.d-ma.be/mathias/tapir/internal/web/oidc"
)
func main() {
@@ -136,7 +137,8 @@ func cmdRun(ctx context.Context, log *slog.Logger) error {
r := runner.New(engine.Source, st, engine, cfg.UserID, log,
runner.WithBackoff(cfg.FetchBackoff),
runner.WithAutoWindow(cfg.AutoSummarizeWindow))
runner.WithAutoWindow(cfg.AutoSummarizeWindow),
runner.WithCaptionMemory(cfg.ChannelCaptionlessThreshold, cfg.ChannelCaptionlessWindow))
log.Info("starting run", "user", cfg.UserID, "model", cfg.SummarizerModel,
"gateway", cfg.GatewayURL, "poll_interval", cfg.PollInterval, "fetch_backoff", cfg.FetchBackoff,
@@ -196,6 +198,15 @@ func cmdServe(ctx context.Context, log *slog.Logger) error {
secretStore := secrets.NewFileStore(cfg.SecretsFile)
app := &web.App{Store: st, Identity: st, Auth: authn, Secrets: secretStore, Log: log, RecencyWindow: cfg.AutoSummarizeWindow}
// Per-video deeper-dive chat over the STORED transcript (ADR-027). Enabled
// whenever a gateway is configured — it needs no YouTube credentials because it
// never fetches. Guarded so a typed-nil never lands in the interface field
// (which would mount the routes over a nil backend).
if c := buildChat(cfg); c != nil {
app.Chat = c
log.Info("web chat enabled (stored-transcript only)", "models", chatModels(cfg))
}
// User onboarding is handled by the IdP (Authentik invite flow), not Tapir —
// the Dex local-password provisioning path was removed (ADR-019). An
// authenticated subject with no Tapir user is routed to /register.
@@ -252,23 +263,36 @@ func cmdServe(ctx context.Context, log *slog.Logger) error {
// One lock shared by the scheduler and connect-triggered passes (#6) so
// they never fetch concurrently — the single-fetcher invariant (ADR-018).
runUser := serialize(&sync.Mutex{}, rawRunUser)
// Onboarding burst (Feature 1): after the connect-triggered discovery pass,
// summarize up to OnboardSummarizeCount of the user's NEWEST unsummarized
// videos so a fresh account gets real summaries in its first session. Hard
// cap; explicit, so it bypasses the recency window — but every fetch still
// goes through globalFetchGate via the Processor. No-op when disabled
// (count 0) or queue-only (no Processor).
// Onboarding burst (Feature 1, refined by ADR-028): after the connect-triggered
// discovery pass, summarize up to OnboardSummarizeCount of the user's newest
// LIKELY-GOOD unsummarized videos so a fresh account gets a strong first
// session. Selection avoids known-junk (Shorts/over-long/livestream VODs via
// the persisted duration); the burst leads its chain with the stronger onboard
// model. Hard cap; explicit, so it bypasses the recency window — but every
// fetch still goes through globalFetchGate. No-op when disabled (count 0) or
// queue-only (no processor).
//
// burstProcessor leads with the stronger model (ADR-028); it collapses onto the
// shared Processor when the onboard model is empty/equal-to-primary or the
// engine config is incomplete.
burstProcessor := app.Processor
if burstEngine, berr := buildBurstProcessor(cfg, st); berr != nil {
return berr
} else if burstEngine != nil {
burstProcessor = &engineProcessor{engine: burstEngine, store: st}
log.Info("onboarding burst uses a stronger model", "onboard_model", cfg.OnboardSummarizerModel)
}
onboard := func(ctx context.Context, userID string) {
if cfg.OnboardSummarizeCount <= 0 || app.Processor == nil {
if cfg.OnboardSummarizeCount <= 0 || burstProcessor == nil {
return
}
ids, err := st.NewestUnsummarizedVideoIDs(ctx, userID, cfg.OnboardSummarizeCount)
ids, err := st.OnboardBurstVideoIDs(ctx, userID, cfg.OnboardSummarizeCount, cfg.MinVideoSeconds, cfg.OnboardMaxVideoSeconds)
if err != nil {
log.Warn("onboarding: list newest unsummarized", "user", userID, "err", err)
log.Warn("onboarding: list burst candidates", "user", userID, "err", err)
return
}
for _, id := range ids {
if err := app.Processor.ProcessVideo(ctx, userID, id); err != nil {
if err := burstProcessor.ProcessVideo(ctx, userID, id); err != nil {
log.Warn("onboarding: summarize", "user", userID, "video", id, "err", err)
}
}
@@ -287,16 +311,34 @@ func cmdServe(ctx context.Context, log *slog.Logger) error {
srv := &http.Server{
Addr: cfg.HTTPAddr,
Handler: app.Router(),
Handler: metrics.HTTPMiddleware(app.Router()),
ReadHeaderTimeout: 10 * time.Second,
}
// Prometheus /metrics on a SEPARATE port (ADR-030) — never on the public app
// mux, so a scrape is in-cluster only. Empty TAPIR_METRICS_ADDR disables it.
var metricsSrv *http.Server
if cfg.MetricsAddr != "" {
mmux := http.NewServeMux()
mmux.Handle("GET /metrics", metrics.Handler())
metricsSrv = &http.Server{Addr: cfg.MetricsAddr, Handler: mmux, ReadHeaderTimeout: 10 * time.Second}
go func() {
log.Info("serving metrics", "addr", cfg.MetricsAddr)
if err := metricsSrv.ListenAndServe(); err != nil && !errors.Is(err, http.ErrServerClosed) {
log.Error("metrics server", "err", err)
}
}()
}
// Graceful shutdown on signal: stop accepting, drain in-flight requests.
go func() {
<-ctx.Done()
shutdownCtx, cancel := context.WithTimeout(context.Background(), 10*time.Second)
defer cancel()
_ = srv.Shutdown(shutdownCtx)
if metricsSrv != nil {
_ = metricsSrv.Shutdown(shutdownCtx)
}
}()
log.Info("serving web ui", "addr", cfg.HTTPAddr, "user", cfg.UserID)
+194 -21
View File
@@ -3,17 +3,20 @@ package main
import (
"context"
"fmt"
"strings"
"gitea.d-ma.be/mathias/tapir/internal/adapters/llm"
"gitea.d-ma.be/mathias/tapir/internal/adapters/secrets"
"gitea.d-ma.be/mathias/tapir/internal/adapters/store"
"gitea.d-ma.be/mathias/tapir/internal/adapters/summarizer"
"gitea.d-ma.be/mathias/tapir/internal/adapters/youtube"
"gitea.d-ma.be/mathias/tapir/internal/config"
"gitea.d-ma.be/mathias/tapir/internal/domain"
"gitea.d-ma.be/mathias/tapir/internal/ports"
"gitea.d-ma.be/mathias/tapir/internal/usecase"
"gitea.d-ma.be/mathias/tapir/internal/web"
"git.d-ma.be/mathias/tapir/internal/adapters/chat"
"git.d-ma.be/mathias/tapir/internal/adapters/llm"
"git.d-ma.be/mathias/tapir/internal/adapters/secrets"
"git.d-ma.be/mathias/tapir/internal/adapters/store"
"git.d-ma.be/mathias/tapir/internal/adapters/summarizer"
"git.d-ma.be/mathias/tapir/internal/adapters/youtube"
"git.d-ma.be/mathias/tapir/internal/config"
"git.d-ma.be/mathias/tapir/internal/domain"
"git.d-ma.be/mathias/tapir/internal/metrics"
"git.d-ma.be/mathias/tapir/internal/ports"
"git.d-ma.be/mathias/tapir/internal/usecase"
"git.d-ma.be/mathias/tapir/internal/web"
)
// videoFetcher adapts the YouTube adapter to web.VideoFetcher for the paste flow
@@ -42,28 +45,172 @@ func (f videoFetcher) FetchVideo(ctx context.Context, userID, videoID string) (d
// queue-only fallback: the web UI keeps working (the button just queues) and
// `tapir run` reports the gap via its own ValidateForRun. Missing engine config
// is never an error here.
// buildSummarizer wires the summarization endpoint chain (ADR-022) shared by the
// web "Summarize now" path and the scheduler's per-user runners. The chain is:
// primary (local, fast) → local fallback → cloud fallback (worst case). Each
// endpoint reaches the same LiteLLM gateway with a different model alias — the
// gateway fronts both llama-swap and berget — so a fallback is just a different
// alias, not a second client config. Empty model entries are skipped, so a
// client deployment can set the cloud fallback empty to keep content local.
func buildSummarizer(cfg config.Config) *summarizer.Summarizer {
mk := summarizerEndpoint(cfg)
eps := []summarizer.Endpoint{mk(cfg.SummarizerModel)}
if cfg.FallbackModel != "" && cfg.FallbackModel != cfg.SummarizerModel {
eps = append(eps, mk(cfg.FallbackModel))
}
if cfg.CloudFallbackModel != "" && cfg.CloudFallbackModel != cfg.SummarizerModel {
eps = append(eps, mk(cfg.CloudFallbackModel))
}
return summarizer.NewChain(eps, cfg.MaxTranscriptChars)
}
// summarizerEndpoint returns a constructor for a chain endpoint over the one
// LiteLLM gateway, varying only the model alias (the gateway fronts both
// llama-swap and berget). Shared by the standard and burst chains.
func summarizerEndpoint(cfg config.Config) func(model string) summarizer.Endpoint {
return func(model string) summarizer.Endpoint {
return summarizer.Endpoint{
Client: llm.New(cfg.GatewayURL, cfg.GatewayKey, model, cfg.SummarizerTimeout, llm.WithMaxTokens(cfg.SummaryMaxTokens), llm.WithUsageHook(metrics.RecordTokens)),
Provider: providerOf(model),
Model: model,
}
}
}
// burstChainModels is the ordered, deduped model list for the onboarding burst
// (ADR-028): the stronger onboard model leads, then the standard ADR-022 chain
// (primary → local fallback → cloud) follows as resilience. Empty entries are
// dropped and duplicates collapsed, so the NDA lever (empty cloud fallback) keeps
// the burst chain fully local exactly as the standard chain does.
func burstChainModels(cfg config.Config) []string {
var models []string
add := func(m string) {
if m == "" {
return
}
for _, e := range models {
if e == m {
return
}
}
models = append(models, m)
}
add(cfg.OnboardSummarizerModel)
add(cfg.SummarizerModel)
add(cfg.FallbackModel)
add(cfg.CloudFallbackModel)
return models
}
// buildBurstSummarizer builds the onboarding-burst summarizer chain (ADR-028):
// the onboard model first, then the standard chain as fallback, deduped.
func buildBurstSummarizer(cfg config.Config) *summarizer.Summarizer {
mk := summarizerEndpoint(cfg)
var eps []summarizer.Endpoint
for _, m := range burstChainModels(cfg) {
eps = append(eps, mk(m))
}
return summarizer.NewChain(eps, cfg.MaxTranscriptChars)
}
// chatModels is the ordered, local-first set of models offered in the chat
// switcher (ADR-027), reusing the ADR-022 chain: primary → local fallback →
// cloud. Empty entries are dropped and duplicates collapsed, so a client/NDA
// deployment that sets the cloud fallback empty simply has no external model in
// the switcher — the same local-first lever the summarizer honours.
func chatModels(cfg config.Config) []string {
var models []string
add := func(m string) {
if m == "" {
return
}
for _, e := range models {
if e == m {
return
}
}
models = append(models, m)
}
add(cfg.SummarizerModel)
add(cfg.FallbackModel)
add(cfg.CloudFallbackModel)
return models
}
// buildChat wires the per-video chat service (ADR-027): a Completer factory over
// the SAME LiteLLM gateway the summarizer uses (a different alias per model, not a
// second client config) and the same transcript-truncation budget. It returns nil
// when no gateway is configured — chat is simply not mounted, the read path is
// unaffected. It deliberately takes NO YouTube source: chat is stored-only.
func buildChat(cfg config.Config) *chat.Service {
if cfg.GatewayURL == "" {
return nil
}
models := chatModels(cfg)
if len(models) == 0 {
return nil
}
newClient := func(model string) chat.Completer {
return llm.New(cfg.GatewayURL, cfg.GatewayKey, model, cfg.SummarizerTimeout, llm.WithMaxTokens(cfg.SummaryMaxTokens), llm.WithUsageHook(metrics.RecordTokens))
}
return chat.New(newClient, models, cfg.MaxTranscriptChars)
}
// providerOf maps a model alias to the domain AIProvider recorded on summaries.
// A "berget/" alias is an external provider; everything else is the local stack.
func providerOf(model string) string {
if strings.HasPrefix(model, "berget/") {
return "berget"
}
return "local"
}
func buildProcessor(cfg config.Config, st *store.Store) (*usecase.Engine, error) {
if cfg.GatewayURL == "" || cfg.YTClientID == "" || cfg.YTClientSecret == "" || cfg.SecretsFile == "" {
return nil, nil
}
secretStore := secrets.NewFileStore(cfg.SecretsFile)
src := youtube.New(youtube.Config{
src := newYouTubeSource(cfg, secrets.NewFileStore(cfg.SecretsFile))
sum := buildSummarizer(cfg)
// The store is both the summary sink and the shared transcript cache (ADR-021):
// the engine reads stored transcripts before any caption fetch and writes
// resolved ones back, so re-analysis never re-touches YouTube.
eng := usecase.NewEngine(src, sum, st)
eng.Transcripts = st
return eng, nil
}
// newYouTubeSource builds the captions-first VideoSource shared by the standard
// and burst processors — same per-process YouTube credentials and ADR-023 Shorts
// filter; only the summarizer chain differs between them.
func newYouTubeSource(cfg config.Config, secretStore ports.SecretStore) ports.VideoSource {
return youtube.New(youtube.Config{
ClientID: cfg.YTClientID,
ClientSecret: cfg.YTClientSecret,
TokenSecretRef: cfg.YTTokenRef,
PreferredLanguages: []string{"en"},
MinVideoSeconds: cfg.MinVideoSeconds,
}, secretStore)
}
// Local Primary only; no BYO fallback for the demo (fallback nil).
primary := summarizer.Endpoint{
Client: llm.New(cfg.GatewayURL, cfg.GatewayKey, cfg.SummarizerModel, cfg.SummarizerTimeout),
Provider: "local",
Model: cfg.SummarizerModel,
// buildBurstProcessor wires a processor whose summarizer leads with the stronger
// onboard model (ADR-028), used only by the connect-time burst over the SAME
// store / transcript cache / sink — a wiring choice; the engine and ports are
// unchanged. Returns (nil, nil) — the collapse lever — when the onboard model is
// empty or equal to the primary (the burst then reuses the shared processor), or
// when the engine config is incomplete (queue-only, same as buildProcessor).
func buildBurstProcessor(cfg config.Config, st *store.Store) (*usecase.Engine, error) {
if cfg.OnboardSummarizerModel == "" || cfg.OnboardSummarizerModel == cfg.SummarizerModel {
return nil, nil
}
sum := summarizer.New(primary, nil)
return usecase.NewEngine(src, sum, st), nil
if cfg.GatewayURL == "" || cfg.YTClientID == "" || cfg.YTClientSecret == "" || cfg.SecretsFile == "" {
return nil, nil
}
src := newYouTubeSource(cfg, secrets.NewFileStore(cfg.SecretsFile))
eng := usecase.NewEngine(src, buildBurstSummarizer(cfg), st)
eng.Transcripts = st
return eng, nil
}
// engineProcessor adapts the engine (which works in terms of a domain.Video) to
@@ -78,6 +225,11 @@ type engineProcessor struct {
}
func (p *engineProcessor) ProcessVideo(ctx context.Context, userID, videoID string) error {
// This is the user-initiated (foreground) path — a click on "Summarize",
// "Try now", or a pasted URL. Mark the context so the caption gate gives it
// priority over the background sweep (ADR-026, Pillar A).
ctx = youtube.ForegroundContext(ctx)
row, err := p.store.GetVideoRow(ctx, userID, videoID)
if err != nil {
return fmt.Errorf("load video %q: %w", videoID, err)
@@ -97,7 +249,28 @@ func (p *engineProcessor) ProcessVideo(ctx context.Context, userID, videoID stri
if err != nil {
return fmt.Errorf("process video %q: %w", videoID, err)
}
if res.Summary != nil {
// Record the outcome so the status endpoint can show honest state (ADR-025):
// a 429'd or caption-less click used to leave transcript_status unset, so the
// poll silently reverted to the "Summarize" button. Mirror the runner: stamp
// rate_limited / none / fetched. A rate-limited video keeps its requested flag
// so the background sweep retries it; none and fetched are terminal here.
switch {
case res.Skipped && res.TranscriptSource == string(domain.SourceRateLimited):
if err := p.store.SetTranscriptStatus(ctx, userID, videoID, "rate_limited"); err != nil {
return fmt.Errorf("set rate_limited status %q: %w", videoID, err)
}
case res.Skipped:
if err := p.store.SetTranscriptStatus(ctx, userID, videoID, "none"); err != nil {
return fmt.Errorf("set none status %q: %w", videoID, err)
}
if err := p.store.ClearSummarizeRequested(ctx, userID, videoID); err != nil {
return fmt.Errorf("clear summarize flag %q: %w", videoID, err)
}
case res.Summary != nil:
if err := p.store.SetTranscriptStatus(ctx, userID, videoID, "fetched"); err != nil {
return fmt.Errorf("set fetched status %q: %w", videoID, err)
}
if err := p.store.ClearSummarizeRequested(ctx, userID, videoID); err != nil {
return fmt.Errorf("clear summarize flag %q: %w", videoID, err)
}
+63 -1
View File
@@ -3,7 +3,7 @@ package main
import (
"testing"
"gitea.d-ma.be/mathias/tapir/internal/config"
"git.d-ma.be/mathias/tapir/internal/config"
)
// TestBuildProcessorNilOnIncompleteConfig asserts the queue-only fallback: when a
@@ -45,3 +45,65 @@ func TestBuildProcessorNilOnIncompleteConfig(t *testing.T) {
})
}
}
// TestBurstChainModelsLeadsWithOnboardModel: the onboarding burst chain (ADR-028)
// leads with the stronger onboard model, then falls back through the standard
// ADR-022 chain (primary -> local fallback -> cloud), deduped.
func TestBurstChainModelsLeadsWithOnboardModel(t *testing.T) {
got := burstChainModels(config.Config{
OnboardSummarizerModel: "iguana/gemma4-26b",
SummarizerModel: "koala/phi4-mini",
FallbackModel: "iguana/gemma4-26b", // also the onboard model -> dedup
CloudFallbackModel: "berget/mistral-small",
})
want := []string{"iguana/gemma4-26b", "koala/phi4-mini", "berget/mistral-small"}
if len(got) != len(want) {
t.Fatalf("burstChainModels = %v, want %v", got, want)
}
for i := range want {
if got[i] != want[i] {
t.Fatalf("burstChainModels = %v, want %v", got, want)
}
}
}
// TestBurstChainModelsCloudAbsentWhenDisabled: the NDA lever holds for the burst
// too — empty cloud fallback keeps the burst chain fully local.
func TestBurstChainModelsCloudAbsentWhenDisabled(t *testing.T) {
got := burstChainModels(config.Config{
OnboardSummarizerModel: "iguana/gemma4-26b",
SummarizerModel: "koala/phi4-mini",
CloudFallbackModel: "",
})
for _, m := range got {
if m == "" || m == "berget/mistral-small" {
t.Fatalf("cloud model leaked into burst chain: %v", got)
}
}
}
// TestBuildBurstProcessorNilWhenCollapsed: an empty or primary-equal onboard model
// collapses the burst onto the shared processor (buildBurstProcessor returns nil).
func TestBuildBurstProcessorNilWhenCollapsed(t *testing.T) {
base := config.Config{
GatewayURL: "http://gw/v1",
YTClientID: "id",
YTClientSecret: "secret",
SecretsFile: "/tmp/secrets.json",
SummarizerModel: "koala/phi4-mini",
}
t.Run("empty onboard model", func(t *testing.T) {
base.OnboardSummarizerModel = ""
eng, err := buildBurstProcessor(base, nil)
if err != nil || eng != nil {
t.Fatalf("buildBurstProcessor = (%v, %v), want (nil, nil)", eng, err)
}
})
t.Run("onboard model equals primary", func(t *testing.T) {
base.OnboardSummarizerModel = "koala/phi4-mini"
eng, err := buildBurstProcessor(base, nil)
if err != nil || eng != nil {
t.Fatalf("buildBurstProcessor = (%v, %v), want (nil, nil)", eng, err)
}
})
}
+34 -4
View File
@@ -6,14 +6,36 @@ import (
"io"
"os"
"text/tabwriter"
"time"
"gitea.d-ma.be/mathias/tapir/internal/adapters/store"
"git.d-ma.be/mathias/tapir/internal/adapters/store"
)
// gateThreshold is the Stage-0 gate (VISION/ADR-016): usage in >= 2 distinct
// weeks. The gate passes when any user reaches it.
const gateThreshold = 2
// defaultGateStart is the date Stage-0 return-usage tracking begins: the morning
// the pilot was actually unblocked and summaries started flowing (2026-06-11).
// Activity before this — testing, the period the pilot was stuck on zero — is
// noise and must not count toward the gate. Override with TAPIR_USAGE_GATE_START
// (YYYY-MM-DD). The gate measures whether users RETURN once it genuinely works.
const defaultGateStart = "2026-06-11"
// gateStart resolves the baseline date from TAPIR_USAGE_GATE_START or the default,
// parsed as a UTC calendar day.
func gateStart() (time.Time, error) {
v := os.Getenv("TAPIR_USAGE_GATE_START")
if v == "" {
v = defaultGateStart
}
t, err := time.Parse("2006-01-02", v)
if err != nil {
return time.Time{}, fmt.Errorf("TAPIR_USAGE_GATE_START=%q: want YYYY-MM-DD: %w", v, err)
}
return t, nil
}
// runReport prints the Stage-0 usage gate: per-user distinct active weeks (reads
// UNION acts) and the pass/fail verdict. Read-only, cross-user — needs only
// TAPIR_DB_DSN (not TAPIR_USER_ID; the report enumerates all users itself).
@@ -28,16 +50,24 @@ func runReport(ctx context.Context, _ []string) error {
}
defer s.Close()
rows, err := s.ActiveWeeks(ctx)
since, err := gateStart()
if err != nil {
return err
}
return formatReport(os.Stdout, rows)
rows, err := s.ActiveWeeks(ctx, since)
if err != nil {
return err
}
return formatReport(os.Stdout, rows, since)
}
// formatReport renders the per-user week counts and the gate verdict. Pure: no DB,
// no env — so the layout and verdict logic are unit-testable without Postgres.
func formatReport(w io.Writer, rows []store.UserActiveWeeks) error {
func formatReport(w io.Writer, rows []store.UserActiveWeeks, since time.Time) error {
if _, err := fmt.Fprintf(w, "Counting usage since %s (Stage-0 gate baseline)\n\n", since.Format("2006-01-02")); err != nil {
return err
}
if len(rows) == 0 {
_, err := fmt.Fprintln(w, "no users yet")
return err
+8 -4
View File
@@ -3,12 +3,15 @@ package main
import (
"strings"
"testing"
"time"
"github.com/stretchr/testify/require"
"gitea.d-ma.be/mathias/tapir/internal/adapters/store"
"git.d-ma.be/mathias/tapir/internal/adapters/store"
)
var testSince = time.Date(2026, 6, 11, 0, 0, 0, 0, time.UTC)
func TestFormatReportColumnsAndGatePass(t *testing.T) {
rows := []store.UserActiveWeeks{
{UserID: "user-a", DisplayName: "Ada", ActiveWeeks: 3},
@@ -16,9 +19,10 @@ func TestFormatReportColumnsAndGatePass(t *testing.T) {
}
var b strings.Builder
require.NoError(t, formatReport(&b, rows))
require.NoError(t, formatReport(&b, rows, testSince))
out := b.String()
require.Contains(t, out, "since 2026-06-11", "report states the gate baseline date")
require.Contains(t, out, "USER")
require.Contains(t, out, "ACTIVE_WEEKS")
require.Contains(t, out, "Ada")
@@ -34,12 +38,12 @@ func TestFormatReportGateNotMet(t *testing.T) {
rows := []store.UserActiveWeeks{{UserID: "user-a", ActiveWeeks: 1}}
var b strings.Builder
require.NoError(t, formatReport(&b, rows))
require.NoError(t, formatReport(&b, rows, testSince))
require.Contains(t, b.String(), "NOT YET MET", "no user at >= 2 weeks fails the gate")
}
func TestFormatReportEmpty(t *testing.T) {
var b strings.Builder
require.NoError(t, formatReport(&b, nil))
require.NoError(t, formatReport(&b, nil, testSince))
require.Contains(t, b.String(), "no users yet")
}
+85 -33
View File
@@ -6,15 +6,13 @@ import (
"log/slog"
"time"
"gitea.d-ma.be/mathias/tapir/internal/adapters/llm"
"gitea.d-ma.be/mathias/tapir/internal/adapters/store"
"gitea.d-ma.be/mathias/tapir/internal/adapters/summarizer"
"gitea.d-ma.be/mathias/tapir/internal/adapters/youtube"
"gitea.d-ma.be/mathias/tapir/internal/config"
"gitea.d-ma.be/mathias/tapir/internal/ports"
"gitea.d-ma.be/mathias/tapir/internal/runner"
"gitea.d-ma.be/mathias/tapir/internal/usecase"
"gitea.d-ma.be/mathias/tapir/internal/web"
"git.d-ma.be/mathias/tapir/internal/adapters/store"
"git.d-ma.be/mathias/tapir/internal/adapters/youtube"
"git.d-ma.be/mathias/tapir/internal/config"
"git.d-ma.be/mathias/tapir/internal/ports"
"git.d-ma.be/mathias/tapir/internal/runner"
"git.d-ma.be/mathias/tapir/internal/usecase"
"git.d-ma.be/mathias/tapir/internal/web"
)
// buildUserRunner constructs a runner.Runner for one user, reusing the same
@@ -35,18 +33,20 @@ func buildUserRunner(cfg config.Config, st *store.Store, secretStore ports.Secre
ClientSecret: cfg.YTClientSecret,
TokenSecretRef: web.YouTubeTokenRef(userID),
PreferredLanguages: []string{"en"},
MinVideoSeconds: cfg.MinVideoSeconds,
}, secretStore)
primary := summarizer.Endpoint{
Client: llm.New(cfg.GatewayURL, cfg.GatewayKey, cfg.SummarizerModel, cfg.SummarizerTimeout),
Provider: "local",
Model: cfg.SummarizerModel,
}
engine := usecase.NewEngine(src, summarizer.New(primary, nil), st)
engine := usecase.NewEngine(src, buildSummarizer(cfg), st)
// Share the transcript cache (ADR-021) on the scheduler path too — without
// this every scheduled pass re-fetches transcripts it already had, burning the
// scarce per-IP caption budget (ADR-014) and starving other users. The web
// "Summarize now" path already sets this; the scheduler omitting it was a bug.
engine.Transcripts = st
return runner.New(src, st, engine, userID, log,
runner.WithBackoff(cfg.FetchBackoff),
runner.WithAutoWindow(cfg.AutoSummarizeWindow)), nil
runner.WithAutoWindow(cfg.AutoSummarizeWindow),
runner.WithCaptionMemory(cfg.ChannelCaptionlessThreshold, cfg.ChannelCaptionlessWindow)), nil
}
// userLister enumerates every registered user and reports a user's video
@@ -64,6 +64,7 @@ type userLister interface {
// isolation). Returns the stats summed across users.
func runDiscoveryPass(
ctx context.Context,
pass int,
lister userLister,
runUser func(context.Context, string) (runner.Stats, error),
log *slog.Logger,
@@ -74,15 +75,17 @@ func runDiscoveryPass(
return runner.Stats{}
}
log.Info("scheduler: starting discovery pass", "users", len(users))
var total runner.Stats
// Keep only users with a video connection. A pass for a connectionless user
// (e.g. a stale Dex-era orphan identity) only tries to resolve a token that
// was never minted, logging a spurious "ref not found" every tick. Filtering
// here — BEFORE rotation — also keeps fairness honest: rotation is over the
// users that actually consume the caption budget, so a dead identity can't eat
// a rotation slot and skew the lead share.
var connected []store.UserIdentity
for _, u := range users {
if ctx.Err() != nil {
break // shutting down: stop enumerating
return runner.Stats{} // shutting down
}
// Skip users with no video connection. A discovery pass for them only
// attempts to resolve a token that was never minted, logging a spurious
// "ref not found" every tick (e.g. stale Dex-era orphan identities).
conns, err := lister.ConnectionsForUser(ctx, u.UserID)
if err != nil {
log.Warn("scheduler: list connections failed", "user", u.UserID, "err", err)
@@ -92,6 +95,22 @@ func runDiscoveryPass(
log.Debug("scheduler: skipping user with no video connections", "user", u.UserID)
continue
}
connected = append(connected, u)
}
// Rotate who goes first each pass. Caption fetches share one per-egress-IP
// rate budget (ADR-014); whoever runs first each pass spends the pre-throttle
// window, so a FIXED order permanently starves whoever is last (a new pilot
// user got 0 fetches for 12h while the first-listed user got all of them).
// Rotation over the connected set gives each real user the lead in turn.
connected = rotateUsers(connected, pass)
log.Info("scheduler: starting discovery pass", "users", len(connected))
var total runner.Stats
for _, u := range connected {
if ctx.Err() != nil {
break // shutting down: stop enumerating
}
stats, err := runUser(ctx, u.UserID)
total = sumStats(total, stats)
if err != nil {
@@ -103,6 +122,7 @@ func runDiscoveryPass(
"skipped_seen", total.SkippedSeen, "skipped_no_text", total.SkippedNoText,
"skipped_manual", total.SkippedManual, "skipped_too_old", total.SkippedTooOld,
"skipped_rate_limited", total.SkippedRateLimited,
"skipped_no_caption_channel", total.SkippedNoCaptionChannel,
"channel_unavailable", total.ChannelUnavailable, "errors", total.Errors)
return total
}
@@ -123,7 +143,18 @@ func runScheduler(
return // disabled
}
runDiscoveryPass(ctx, lister, runUser, log)
// Derive the rotation offset from wall-clock, NOT an in-memory counter. A
// counter reset to 0 on every pod restart always hands the lead to the
// first-listed user — so frequent deploys re-starve whoever is last (exactly
// what happened to the first pilot user during a deploy-heavy session). A
// time-based offset advances with real time and is identical across restarts,
// so the lead rotates fairly regardless of how often the pod bounces.
runPass := func() {
pass := int(time.Now().Unix() / int64(interval/time.Second))
runDiscoveryPass(ctx, pass, lister, runUser, log)
}
runPass()
ticker := time.NewTicker(interval)
defer ticker.Stop()
@@ -132,23 +163,44 @@ func runScheduler(
case <-ctx.Done():
return
case <-ticker.C:
runDiscoveryPass(ctx, lister, runUser, log)
runPass()
}
}
}
// rotateUsers left-rotates users by pass positions so a different user leads each
// pass. With n users, user i leads on every pass where pass ≡ i (mod n). A pass
// offset that is negative or exceeds n is normalised. Order within the rotation
// is otherwise preserved, so the set of users run is unchanged — only who is
// first (and thus wins the scarce caption-fetch budget) rotates.
func rotateUsers(users []store.UserIdentity, pass int) []store.UserIdentity {
n := len(users)
if n <= 1 {
return users
}
off := ((pass % n) + n) % n
if off == 0 {
return users
}
out := make([]store.UserIdentity, 0, n)
out = append(out, users[off:]...)
out = append(out, users[:off]...)
return out
}
// sumStats adds two passes' stats field-wise, so runDiscoveryPass can report a
// per-tick aggregate across all users.
func sumStats(a, b runner.Stats) runner.Stats {
return runner.Stats{
Candidates: a.Candidates + b.Candidates,
Summarized: a.Summarized + b.Summarized,
SkippedSeen: a.SkippedSeen + b.SkippedSeen,
SkippedNoText: a.SkippedNoText + b.SkippedNoText,
SkippedManual: a.SkippedManual + b.SkippedManual,
SkippedTooOld: a.SkippedTooOld + b.SkippedTooOld,
SkippedRateLimited: a.SkippedRateLimited + b.SkippedRateLimited,
ChannelUnavailable: a.ChannelUnavailable + b.ChannelUnavailable,
Errors: a.Errors + b.Errors,
Candidates: a.Candidates + b.Candidates,
Summarized: a.Summarized + b.Summarized,
SkippedSeen: a.SkippedSeen + b.SkippedSeen,
SkippedNoText: a.SkippedNoText + b.SkippedNoText,
SkippedManual: a.SkippedManual + b.SkippedManual,
SkippedTooOld: a.SkippedTooOld + b.SkippedTooOld,
SkippedRateLimited: a.SkippedRateLimited + b.SkippedRateLimited,
SkippedNoCaptionChannel: a.SkippedNoCaptionChannel + b.SkippedNoCaptionChannel,
ChannelUnavailable: a.ChannelUnavailable + b.ChannelUnavailable,
Errors: a.Errors + b.Errors,
}
}
+47 -6
View File
@@ -11,8 +11,8 @@ import (
"github.com/stretchr/testify/require"
"gitea.d-ma.be/mathias/tapir/internal/adapters/store"
"gitea.d-ma.be/mathias/tapir/internal/runner"
"git.d-ma.be/mathias/tapir/internal/adapters/store"
"git.d-ma.be/mathias/tapir/internal/runner"
)
func quietLog() *slog.Logger {
@@ -45,6 +45,7 @@ func (f fakeLister) ConnectionsForUser(_ context.Context, userID string) ([]stor
type countingRunUser struct {
mu sync.Mutex
calls map[string]int
order []string // userIDs in the order they were run, across all passes
failFor map[string]bool
}
@@ -60,12 +61,19 @@ func (c *countingRunUser) run(_ context.Context, userID string) (runner.Stats, e
c.mu.Lock()
defer c.mu.Unlock()
c.calls[userID]++
c.order = append(c.order, userID)
if c.failFor[userID] {
return runner.Stats{Errors: 1}, errors.New("boom")
}
return runner.Stats{Summarized: 1}, nil
}
func (c *countingRunUser) runOrder() []string {
c.mu.Lock()
defer c.mu.Unlock()
return append([]string(nil), c.order...)
}
func (c *countingRunUser) count(userID string) int {
c.mu.Lock()
defer c.mu.Unlock()
@@ -94,7 +102,7 @@ func TestDiscoveryPassRunsEveryUserOnce(t *testing.T) {
lister := fakeLister{users: usersN("a", "b", "c")}
rc := newCountingRunUser()
stats := runDiscoveryPass(context.Background(), lister, rc.run, quietLog())
stats := runDiscoveryPass(context.Background(), 0, lister, rc.run, quietLog())
require.Equal(t, 1, rc.count("a"))
require.Equal(t, 1, rc.count("b"))
@@ -102,13 +110,46 @@ func TestDiscoveryPassRunsEveryUserOnce(t *testing.T) {
require.Equal(t, 3, stats.Summarized, "stats are summed across users")
}
// Caption fetches share one per-IP budget; a fixed user order starves whoever is
// last. Each pass must rotate which user leads so the lead slot is shared.
func TestDiscoveryPassRotatesLeadUser(t *testing.T) {
lister := fakeLister{users: usersN("a", "b", "c")}
rc := newCountingRunUser()
runDiscoveryPass(context.Background(), 0, lister, rc.run, quietLog())
runDiscoveryPass(context.Background(), 1, lister, rc.run, quietLog())
runDiscoveryPass(context.Background(), 2, lister, rc.run, quietLog())
require.Equal(t, []string{"a", "b", "c", "b", "c", "a", "c", "a", "b"}, rc.runOrder(),
"each pass left-rotates the user order so every user leads in turn")
// Fairness: over a full rotation cycle every user ran the same number of times.
require.Equal(t, 3, rc.count("a"))
require.Equal(t, 3, rc.count("b"))
require.Equal(t, 3, rc.count("c"))
}
// A connectionless orphan must not consume a rotation slot: rotation is over the
// connected users only, so two real users alternate the lead 50/50 even with a
// dead identity listed between them.
func TestDiscoveryPassRotationIgnoresConnectionlessUsers(t *testing.T) {
lister := fakeLister{users: usersN("a", "orphan", "c"), noConn: map[string]bool{"orphan": true}}
rc := newCountingRunUser()
runDiscoveryPass(context.Background(), 0, lister, rc.run, quietLog())
runDiscoveryPass(context.Background(), 1, lister, rc.run, quietLog())
require.Equal(t, []string{"a", "c", "c", "a"}, rc.runOrder(),
"only connected users rotate; the orphan never runs and never holds a slot")
require.Equal(t, 0, rc.count("orphan"))
}
func TestDiscoveryPassSkipsUsersWithoutConnections(t *testing.T) {
// b never connected a video source (e.g. a stale Dex-era orphan identity).
// It must be skipped silently — not run and logged as a token error every pass.
lister := fakeLister{users: usersN("a", "b", "c"), noConn: map[string]bool{"b": true}}
rc := newCountingRunUser()
stats := runDiscoveryPass(context.Background(), lister, rc.run, quietLog())
stats := runDiscoveryPass(context.Background(), 0, lister, rc.run, quietLog())
require.Equal(t, 1, rc.count("a"))
require.Equal(t, 0, rc.count("b"), "a user with no connection must be skipped, not run")
@@ -121,7 +162,7 @@ func TestDiscoveryPassOneUserFailureDoesNotStopOthers(t *testing.T) {
lister := fakeLister{users: usersN("a", "b", "c")}
rc := newCountingRunUser("b") // user b's pass errors
stats := runDiscoveryPass(context.Background(), lister, rc.run, quietLog())
stats := runDiscoveryPass(context.Background(), 0, lister, rc.run, quietLog())
require.Equal(t, 1, rc.count("a"))
require.Equal(t, 1, rc.count("b"))
@@ -134,7 +175,7 @@ func TestDiscoveryPassListerErrorIsContained(t *testing.T) {
lister := fakeLister{err: errors.New("db down")}
rc := newCountingRunUser()
stats := runDiscoveryPass(context.Background(), lister, rc.run, quietLog())
stats := runDiscoveryPass(context.Background(), 0, lister, rc.run, quietLog())
require.Equal(t, 0, rc.total(), "no users enumerated → no passes")
require.Equal(t, runner.Stats{}, stats)
+1 -1
View File
@@ -8,7 +8,7 @@ import (
"os"
"strings"
"gitea.d-ma.be/mathias/tapir/internal/adapters/store"
"git.d-ma.be/mathias/tapir/internal/adapters/store"
)
// runShow prints the full summary for one video: text, highlights, takeaways,
+26
View File
@@ -250,6 +250,32 @@ After (newest-first): `[chanB-new, chanA-mid, chanA-old, chanB-null]`
The set of *processed* videos now also excludes auto-mode back-catalogue beyond the recency
window (those stay listed, summarised on demand); within the processed set, only order changes.
### Connect-time onboarding burst (ADR-018 → ADR-028)
On a successful YouTube connect, `ConnectHandler` enqueues a connect-triggered discovery pass;
the `discoveryTrigger` runs that pass and then fires the **onboarding burst** — a third entry path
that summarises up to `TAPIR_ONBOARD_SUMMARIZE_COUNT` (default 3, hard-capped) of the new user's
videos so the first session is not empty. The burst still flows through `globalFetchGate` (it is
not a throughput change); ADR-028 sharpened *which* videos and *which model*:
- **Selection** is `OnboardBurstVideoIDs`, not pure newest-first. It keeps newest-first order but
excludes a video whose **known** duration is outside `[TAPIR_MIN_VIDEO_SECONDS,
TAPIR_ONBOARD_MAX_VIDEO_SECONDS]` (drops Shorts and multi-hour livestream VODs). An unknown
(NULL) duration is degrade-open — kept, but ranked after known-good rows. The connect-triggered
discovery pass runs *before* the burst, and ADR-023's `videos.list` enrichment now **persists**
`duration_s` (instead of discarding it after the Shorts filter), so a fresh user's candidates
carry a duration in time for selection.
- **Model**: the burst runs through a dedicated summarizer chain led by
`TAPIR_ONBOARD_SUMMARIZER_MODEL` (default `iguana/gemma4-26b`, the stronger local model), with
the standard ADR-022 chain following as fallback. This is a wiring choice — a second
`engineProcessor` over the same store / transcript cache / sink; the engine and ports are
unchanged. Empty / equal-to-primary collapses it back onto the shared processor.
`has-captions` is deliberately **not** a selection signal — it is only knowable after a gate fetch
(or a ~0-probability cache hit at pilot scale), so the burst can avoid known-junk but cannot
promise captions. Cached-transcript-first selection was investigated and rejected (ADR-028:
~3% cross-user overlap).
---
## Sequence — core use case: new video summarized
+21 -13
View File
@@ -10,12 +10,15 @@ only opaque references to them; the secret material lives in ESO/1Password (ADR-
## Design decisions baked into this model
- **Per-user isolation, not a shared global video table.** The earlier draft proposed a
global `videos`/`transcripts` table deduped across tenants. Rejected for Future B: it
reintroduces exactly the cross-domain coupling the homelab architecture review is
removing, and at 15 users the cost of occasionally re-summarizing the same video is
trivial compared to the isolation it would cost. Each user's data is self-contained.
(Revisit only if Future C makes GPU/transcription cost dominate — a new ADR, not a default.)
- **Per-user isolation for everything except transcripts.** The earlier draft proposed a
global `videos`/`transcripts` table deduped across tenants. **Videos** stay per-user and
RLS-scoped — a shared video table reintroduces exactly the cross-domain coupling the homelab
architecture review is removing. **Transcripts**, however, are now shared (ADR-021): keyed by
`(provider, provider_video_id)`, no `user_id`, **not** RLS-scoped. The cost avoided there is
not LLM re-summarization but a rate-gated, reputation-risky caption fetch (ADR-010/014), which
is paid per re-fetch regardless of user count — so persisting public caption content once and
sharing it strictly beats the coupling it removes. Everything else each user owns is
self-contained; `rls_test.go` proves transcripts is the single exception.
- **Secrets by reference only.** Tables hold a `secret_ref` (opaque string/UUID resolved via
the `SecretStore` port), never tokens or keys.
- **The brain sink is just a delivery target.** No brain-specific tables. Whether a summary
@@ -34,7 +37,7 @@ erDiagram
USER ||--o{ AI_CREDENTIAL : "has (planned)"
VIDEO_CONNECTION ||--o{ SUBSCRIPTION : "exposes (planned)"
SUBSCRIPTION ||--o{ VIDEO : "produces (per user)"
VIDEO ||--o| TRANSCRIPT : "has at most one"
VIDEO }o--o| TRANSCRIPT : "shares one by (provider, provider_video_id) — not FK (ADR-021)"
VIDEO ||--o| SUMMARY : "has at most one"
SUMMARY ||--o{ SINK_DELIVERY : "delivered via"
USER ||--o{ CHANNEL_ERROR : "reports unavailable channels"
@@ -92,12 +95,12 @@ erDiagram
timestamptz rate_limited_at "backoff clock for 429 retries (migration 007)"
}
TRANSCRIPT {
uuid video_id PK_FK
uuid user_id FK
text provider PK "part of shared key (ADR-021)"
text provider_video_id PK "part of shared key — the cross-user dedup key"
text source "captions | none"
text language
text content "null when source = none"
timestamptz resolved_at
timestamptz fetched_at
}
SUMMARY {
uuid id PK
@@ -173,8 +176,12 @@ mechanism.
`transcript_status` and `rate_limited_at` (migration 007) track caption-fetch outcomes for
rate-limit backoff: `NULL` = not attempted; `rate_limited` = 429 seen, skip until
`NOW() - rate_limited_at > TAPIR_FETCH_BACKOFF`; `fetched` = resolved; `none` = no transcript.
- **TRANSCRIPT** — at most one per video. `source = none` records "checked, no usable
transcript" so the watcher doesn't reprocess (ADR-007). `content` null in that case.
- **TRANSCRIPT** — shared public caption content, one row per `(provider, provider_video_id)`,
**not** RLS-scoped and carrying no `user_id` (ADR-021). Two users who watch the same video
share the one row; the summarize path reads it before any caption fetch, so re-analysis never
re-touches YouTube (ADR-010/014). `source = none` records "checked, no usable transcript" so
no one reprocesses (ADR-007); `content` null in that case. A transient 429 is never stored
here — it stays a per-user retry via `VIDEO.transcript_status`.
- **SUMMARY** — at most one per video. `fallback_used` + `ai_provider`/`ai_model` make the
"is local good enough?" question queryable (the Stage 0 quality signal). `highlights`/
`takeaways` as jsonb to stay schema-flexible while the output format settles.
@@ -231,7 +238,8 @@ queue, doesn't replace it). Deferred until there's a reason.
## Explicitly out of scope (Future C)
- Global cross-tenant video/transcript dedup (rejected above).
- Global cross-tenant *video* dedup (rejected above). Note: cross-tenant *transcript* sharing
is now in scope and shipped (ADR-021); only the videos half stays per-user.
- Sharding / per-tenant physical databases.
- Soft-delete + full audit trail on connections/credentials (a Stage 2 hardening item; add
via ADR when Stage 2 work starts).
+33
View File
@@ -27,6 +27,30 @@ it** — endpoints and aliases drift, and this file is a snapshot (2026-06-06),
`iguana/deepseek-r1-14b`) is preferred for summary quality if its latency/output is acceptable.
The `max_tokens` fix below means thinking models no longer return empty content, so they are now
viable choices, not blocked ones. Do not assume a coder alias is right for prose.
- **Summarizer fallback chain (ADR-022).** The primary alias is the *first* of an ordered chain;
on failure or unparseable output the summarizer advances to the next model. All reached through
the same gateway by alias.
- `TAPIR_FALLBACK_MODEL` — local fallback. **Default `iguana/gemma4-26b`** — on iguana, NOT
koala, so the fallback does not compete with koala's other GPU loads (and runs from a different
egress IP). Empty disables it.
- `TAPIR_CLOUD_FALLBACK_MODEL` — worst-case EXTERNAL fallback. **Default `berget/mistral-small`.**
**Set this empty (`""`) for any client/NDA deployment** so content never leaves the local
stack — the chain then contains only local endpoints.
- `TAPIR_SUMMARY_MAX_TOKENS` — per-summary completion budget. **Default `1500`.** Small on
purpose: with the old 8192 budget, prompt + completion overflowed `phi4-mini`'s 8k window.
- `TAPIR_MAX_TRANSCRIPT_CHARS` — transcript truncation budget sent to the model. **Default
`18000`** (~fits an 8k-context model). `0` disables truncation. Prevents the context-overflow
HTTP 400 a long transcript caused on `phi4-mini`.
- **Discovery low-value filter (ADR-023).** `TAPIR_MIN_VIDEO_SECONDS`**default `60`**. At
discovery, `NewVideos` enriches candidates with one cheap `videos.list` call (quota API, NOT
the timedtext 429 path) and drops videos shorter than this plus any live/upcoming broadcast,
so the scarce caption-fetch budget isn't spent on Shorts. `0` disables the filter. The
paste-a-URL path is never filtered.
- **Per-channel caption memory (ADR-024).** `TAPIR_CHANNEL_CAPTIONLESS_THRESHOLD` — **default
`5`** consecutive no-caption results before a channel is suppressed (its videos listed but not
caption-fetched). `TAPIR_CHANNEL_CAPTIONLESS_WINDOW`**default `336h`** (14d) suppression
before one video is re-probed. `THRESHOLD=0` disables. A successful fetch resets the channel;
a 429 does not count; an explicit manual request bypasses suppression.
- **Thinking models need an explicit `max_tokens`.** qwen3 / deepseek-r1 spend the budget on
reasoning and return **empty content** if `max_tokens` is too low (or unset). The summarizer's
parser treats an empty summary as an error for exactly this reason. **Done (2026-06-02, Worker F):**
@@ -200,6 +224,15 @@ knobs plus one load-bearing deployment constraint:
- `TAPIR_DISCOVERY_INTERVAL` — Go duration, e.g. `2h`. The cadence the serve process runs a
discovery pass for every registered user (run-once-on-startup, then every interval).
**Unset or `0` = disabled** (dev/tests never auto-fetch).
- `TAPIR_USAGE_GATE_START``YYYY-MM-DD`, default **`2026-06-11`** (the morning the pilot was
unblocked and summaries started flowing). `tapir report` counts return-usage (distinct active
weeks, ADR-016) only from this date, so pre-launch testing and the blocked period are excluded.
- `TAPIR_METRICS_ADDR` — listen address for the Prometheus `/metrics` endpoint (ADR-030).
**Default `:9090`** — a SEPARATE port from `TAPIR_HTTP_ADDR` so metrics are never on the public
app; scraped in-cluster only (PodMonitor). Empty disables the metrics server. Key series:
`tapir_summarize_duration_seconds{model,outcome,fallback}`, `tapir_caption_fetch_duration_seconds{outcome}`,
`tapir_chat_duration_seconds{model}`, `tapir_llm_tokens_total{model,kind}`,
`tapir_http_request_duration_seconds{method,route}`, `tapir_logins_total`.
- `TAPIR_FETCH_RATE` — Go duration, default `2s`. The **process-wide per-egress-IP caption-fetch
rate gate** (ADR-014 item 2). Every caption fetch — scheduler runners *and* the web "Summarize"
click-path — serialises through this one limiter so the pod cannot collectively trip 429s. `0`
+63
View File
@@ -0,0 +1,63 @@
<!DOCTYPE html><html lang="en"><head><meta charset="utf-8">
<meta name="viewport" content="width=device-width, initial-scale=1">
<title>tapir — A · TUI panels</title>
<style>
:root{--bg:#0d0d12;--panel:#14141c;--line:#2a2a3a;--mint:#0EF9B6;--pink:#F740A0;--purple:#7653fc;--cream:#F5E9D6;--dim:#6b6b86;--text:#d9d9e6}
*{box-sizing:border-box}
body{margin:0;background:var(--bg);color:var(--text);font:14px/1.5 ui-monospace,SFMono-Regular,Menlo,"Cascadia Code",monospace}
.wrap{max-width:760px;margin:0 auto;padding:18px 14px 60px}
header{display:flex;align-items:center;gap:10px;border-bottom:1px solid var(--line);padding-bottom:12px;margin-bottom:18px}
.brand{color:var(--mint);font-weight:700;letter-spacing:.5px}
.brand b{color:var(--cream)}
.tip{margin-left:auto;color:var(--dim);font-size:12px}
.note{color:var(--dim);font-size:12.5px;border-left:2px solid var(--purple);padding:6px 10px;margin:0 0 16px;background:#11111a}
/* lipgloss-style bordered panel */
.card{border:1px solid var(--line);border-radius:8px;background:var(--panel);padding:12px 14px;margin:0 0 12px;position:relative}
.card::before{content:"";position:absolute;left:0;top:10px;bottom:10px;width:3px;border-radius:3px;background:var(--purple)}
.card.ready::before{background:var(--mint)}
.title{color:var(--cream);font-size:15px;font-weight:600;margin:0 0 4px}
.meta{color:var(--dim);font-size:12px}
.chip{display:inline-block;border:1px solid var(--mint);color:var(--mint);border-radius:4px;padding:0 6px;font-size:11px;margin-left:6px}
.chip.q{border-color:var(--pink);color:var(--pink)}
.preview{color:var(--dim);margin:6px 0 0;font-size:13px}
/* expanded */
.exp{border-color:var(--purple)}
.exp .head{display:flex;justify-content:space-between;align-items:baseline}
.collapse{color:var(--mint);font-size:12px;text-decoration:none;border:1px solid var(--line);border-radius:4px;padding:1px 7px}
.sec h2{color:var(--mint);font-size:12px;text-transform:uppercase;letter-spacing:1px;margin:16px 0 6px;border-bottom:1px dashed var(--line);padding-bottom:3px}
.sec ul{margin:0;padding-left:18px}.sec li{margin:3px 0}
.body{color:var(--text)}
.dock{margin-top:16px;border-top:1px solid var(--line);padding-top:12px}
.ask{display:inline-block;background:linear-gradient(90deg,var(--purple),var(--pink));color:#fff;border:0;border-radius:6px;padding:7px 12px;font:inherit;font-size:13px;cursor:pointer}
pre.tapir{margin:0;color:var(--mint);font-size:10px;line-height:1.05}
</style></head><body><div class="wrap">
<header>
<pre class="tapir"> ▄█▓▓█▄ ∩
█▓( ◕ ◕ )▓█──┘</pre>
<span class="brand"><b>tapir</b> · watch less, know more</span>
<span class="tip">261 in queue</span>
</header>
<p class="note">Tapir fetches captions slowly on purpose, to respect YouTube's limits — new summaries land gradually.</p>
<div class="card ready"><div class="title">How the Attention Economy Rewires Your Brain</div>
<div class="meta">youtube · 2026-06-11 · 18 min <span class="chip">ready</span></div>
<div class="preview">A tour of the incentive loops behind infinite feeds and three concrete ways to claw back focus…</div></div>
<article class="card exp ready">
<div class="head"><div class="title">Postgres 18 — What's Actually New</div><a class="collapse" href="#">collapse ↑</a></div>
<div class="meta">youtube · 2026-06-10 · 42 min · phi4-mini</div>
<div class="sec"><h2>Takeaways</h2><ul>
<li>Async I/O cuts cold-cache read latency materially on NVMe.</li>
<li>Skip-scan makes more multicolumn indexes usable without rewrites.</li>
<li>Upgrade path is smooth; test the new planner stats first.</li></ul></div>
<div class="sec"><h2>Highlights</h2><ul>
<li>Async I/O subsystem (effective_io_concurrency now matters more).</li>
<li>B-tree skip scan for leading-column gaps.</li>
<li>Better partition-wise joins.</li></ul></div>
<div class="sec"><h2>Summary</h2><p class="body">Postgres 18 is an incremental but meaningful release: the headline is the new asynchronous I/O path, with skip-scan and planner improvements close behind. For most homelabs the upgrade is low-risk and worth it for the read-latency wins.</p></div>
<div class="dock"><button class="ask">Dig deeper — ask about this video →</button></div>
</article>
<div class="card"><div class="title">RAG is dead, right?? — Conference Talk</div>
<div class="meta">youtube · 2026-06-09 · 31 min <span class="chip q">queued</span></div></div>
</div></body></html>
+62
View File
@@ -0,0 +1,62 @@
<!DOCTYPE html><html lang="en"><head><meta charset="utf-8">
<meta name="viewport" content="width=device-width, initial-scale=1">
<title>tapir — B · reader + charm accents</title>
<style>
:root{--bg:#faf7f2;--card:#fff;--ink:#1c1b22;--soft:#6a6878;--line:#e7e2d8;--mint:#0bbf8c;--purple:#6a4cf0;--pink:#e0379a}
*{box-sizing:border-box}
body{margin:0;background:var(--bg);color:var(--ink);font:16px/1.6 -apple-system,BlinkMacSystemFont,"Segoe UI",Inter,sans-serif}
.mono{font-family:ui-monospace,SFMono-Regular,Menlo,monospace}
.wrap{max-width:680px;margin:0 auto;padding:20px 16px 60px}
header{display:flex;align-items:center;gap:10px;margin-bottom:18px}
.brand{font-weight:800;font-size:18px;letter-spacing:-.3px}
.brand .dot{color:var(--mint)}
.tag{color:var(--soft);font-size:13px}
.tip{margin-left:auto;color:var(--soft);font-size:12px}
.note{color:var(--soft);font-size:13px;background:#fff;border:1px solid var(--line);border-left:3px solid var(--mint);border-radius:8px;padding:8px 12px;margin:0 0 18px}
.card{background:var(--card);border:1px solid var(--line);border-radius:12px;padding:16px 18px;margin:0 0 14px;box-shadow:0 1px 2px rgba(20,18,40,.04)}
.title{font-size:18px;font-weight:700;letter-spacing:-.2px;margin:0 0 4px;line-height:1.3}
.meta{color:var(--soft);font-size:13px}
.meta .mono{font-size:12.5px}
.chip{display:inline-block;background:rgba(11,191,140,.12);color:var(--mint);border-radius:999px;padding:1px 9px;font-size:12px;font-weight:600;margin-left:6px}
.chip.q{background:rgba(224,55,154,.12);color:var(--pink)}
.preview{color:var(--soft);margin:8px 0 0}
.exp{border-color:#d9d0ee;box-shadow:0 6px 24px rgba(106,76,240,.10)}
.head{display:flex;justify-content:space-between;align-items:baseline;gap:10px}
.collapse{color:var(--purple);font-size:13px;font-weight:600;text-decoration:none;white-space:nowrap}
.sec h2{font-size:12px;text-transform:uppercase;letter-spacing:1.2px;color:var(--purple);margin:18px 0 6px}
.sec ul{margin:0;padding-left:20px}.sec li{margin:5px 0}
.body{line-height:1.7}
.dock{margin-top:18px;border-top:1px solid var(--line);padding-top:14px}
.ask{display:inline-flex;align-items:center;gap:6px;background:var(--ink);color:#fff;border:0;border-radius:999px;padding:9px 16px;font:inherit;font-size:14px;font-weight:600;cursor:pointer}
pre.tapir{margin:0;color:var(--mint);font-size:10px;line-height:1.05}
</style></head><body><div class="wrap">
<header>
<pre class="tapir"> ▄█▓▓█▄ ∩
█▓( ◕ ◕ )▓█──┘</pre>
<span class="brand">tapir<span class="dot">.</span></span><span class="tag">watch less, know more</span>
<span class="tip mono">261 in queue</span>
</header>
<p class="note">Tapir fetches captions slowly on purpose, to respect YouTube's limits — new summaries land gradually.</p>
<div class="card"><div class="title">How the Attention Economy Rewires Your Brain</div>
<div class="meta mono">youtube · 2026-06-11 · 18 min <span class="chip">ready</span></div>
<div class="preview">A tour of the incentive loops behind infinite feeds and three concrete ways to claw back focus…</div></div>
<article class="card exp">
<div class="head"><div class="title">Postgres 18 — What's Actually New</div><a class="collapse" href="#">collapse ↑</a></div>
<div class="meta mono">youtube · 2026-06-10 · 42 min · phi4-mini</div>
<div class="sec"><h2>Takeaways</h2><ul>
<li>Async I/O cuts cold-cache read latency materially on NVMe.</li>
<li>Skip-scan makes more multicolumn indexes usable without rewrites.</li>
<li>Upgrade path is smooth; test the new planner stats first.</li></ul></div>
<div class="sec"><h2>Highlights</h2><ul>
<li>Async I/O subsystem (effective_io_concurrency now matters more).</li>
<li>B-tree skip scan for leading-column gaps.</li>
<li>Better partition-wise joins.</li></ul></div>
<div class="sec"><h2>Summary</h2><p class="body">Postgres 18 is an incremental but meaningful release: the headline is the new asynchronous I/O path, with skip-scan and planner improvements close behind. For most homelabs the upgrade is low-risk and worth it for the read-latency wins.</p></div>
<div class="dock"><button class="ask">✦ Ask about this video</button></div>
</article>
<div class="card"><div class="title">RAG is dead, right?? — Conference Talk</div>
<div class="meta mono">youtube · 2026-06-09 · 31 min <span class="chip q">queued</span></div></div>
</div></body></html>
+65
View File
@@ -0,0 +1,65 @@
<!DOCTYPE html><html lang="en"><head><meta charset="utf-8">
<meta name="viewport" content="width=device-width, initial-scale=1">
<title>tapir — C · cozy terminal</title>
<style>
:root{--bg:#16131d;--card:#1e1a28;--ink:#ece7f5;--soft:#9a92b4;--line:#322b44;--mint:#2ee6b6;--purple:#9d7bff;--pink:#ff6bdb;--cream:#f3ead8}
*{box-sizing:border-box}
body{margin:0;background:radial-gradient(1200px 600px at 70% -10%,#231b33,transparent),var(--bg);color:var(--ink);font:15px/1.6 -apple-system,BlinkMacSystemFont,"Segoe UI",Inter,sans-serif}
.mono{font-family:ui-monospace,SFMono-Regular,Menlo,monospace}
.wrap{max-width:700px;margin:0 auto;padding:20px 16px 60px}
header{display:flex;align-items:center;gap:12px;margin-bottom:18px}
.brand{font-weight:800;font-size:18px}.brand .c{color:var(--mint)}
.tag{color:var(--soft);font-size:13px}
.tip{margin-left:auto;color:var(--soft);font-size:12px}
.note{color:var(--soft);font-size:13px;background:#1b1726;border:1px solid var(--line);border-radius:10px;padding:9px 12px;margin:0 0 18px}
.note b{color:var(--mint);font-weight:600}
.card{position:relative;background:var(--card);border:1px solid var(--line);border-radius:12px;padding:14px 16px 14px 18px;margin:0 0 13px;overflow:hidden}
.card::before{content:"";position:absolute;left:0;top:0;bottom:0;width:4px;background:var(--purple)}
.card.ready::before{background:linear-gradient(var(--mint),var(--purple))}
.title{font-size:16px;font-weight:700;margin:0 0 4px;color:var(--cream)}
.meta{color:var(--soft);font-size:12.5px}
.chip{display:inline-block;border-radius:999px;padding:1px 9px;font-size:11px;font-weight:600;margin-left:6px;background:rgba(46,230,182,.14);color:var(--mint)}
.chip.q{background:rgba(157,123,255,.18);color:var(--purple)}
.preview{color:var(--soft);margin:7px 0 0;font-size:14px}
.exp{border-color:#473a63;box-shadow:0 10px 40px rgba(0,0,0,.35)}
.head{display:flex;justify-content:space-between;align-items:baseline;gap:10px}
.collapse{color:var(--mint);font-size:12px;text-decoration:none;border:1px solid var(--line);border-radius:6px;padding:2px 8px;white-space:nowrap}
.sec h2{font-size:11px;text-transform:uppercase;letter-spacing:1.3px;color:var(--mint);margin:18px 0 7px;display:flex;align-items:center;gap:8px}
.sec h2::after{content:"";flex:1;height:1px;background:var(--line)}
.sec ul{margin:0;padding-left:18px}.sec li{margin:5px 0}
.body{line-height:1.7;color:var(--ink)}
.dock{margin-top:18px;border-top:1px dashed var(--line);padding-top:14px;display:flex;align-items:center;gap:10px}
.ask{display:inline-flex;align-items:center;gap:7px;background:linear-gradient(90deg,var(--purple),var(--mint));color:#10101a;border:0;border-radius:999px;padding:9px 16px;font:inherit;font-weight:700;font-size:14px;cursor:pointer}
.dockhint{color:var(--soft);font-size:12px}
pre.tapir{margin:0;color:var(--mint);font-size:11px;line-height:1.05}
</style></head><body><div class="wrap">
<header>
<pre class="tapir"> ▄█▓▓█▄ ∩
█▓( ◕ ◕ )▓█──┘</pre>
<span class="brand">tapir<span class="c">_</span></span><span class="tag">watch less, know more</span>
<span class="tip mono">261 in queue</span>
</header>
<p class="note">Tapir fetches captions slowly on purpose, to respect YouTube's limits — <b>new summaries land gradually</b>.</p>
<div class="card ready"><div class="title">How the Attention Economy Rewires Your Brain</div>
<div class="meta mono">youtube · 2026-06-11 · 18 min <span class="chip">ready</span></div>
<div class="preview">A tour of the incentive loops behind infinite feeds and three concrete ways to claw back focus…</div></div>
<article class="card exp ready">
<div class="head"><div class="title">Postgres 18 — What's Actually New</div><a class="collapse" href="#">collapse ↑</a></div>
<div class="meta mono">youtube · 2026-06-10 · 42 min · phi4-mini</div>
<div class="sec"><h2>Takeaways</h2><ul>
<li>Async I/O cuts cold-cache read latency materially on NVMe.</li>
<li>Skip-scan makes more multicolumn indexes usable without rewrites.</li>
<li>Upgrade path is smooth; test the new planner stats first.</li></ul></div>
<div class="sec"><h2>Highlights</h2><ul>
<li>Async I/O subsystem (effective_io_concurrency now matters more).</li>
<li>B-tree skip scan for leading-column gaps.</li>
<li>Better partition-wise joins.</li></ul></div>
<div class="sec"><h2>Summary</h2><p class="body">Postgres 18 is an incremental but meaningful release: the headline is the new asynchronous I/O path, with skip-scan and planner improvements close behind. For most homelabs the upgrade is low-risk and worth it for the read-latency wins.</p></div>
<div class="dock"><button class="ask">◆ Ask about this video</button><span class="dockhint mono">answers only from this video's transcript</span></div>
</article>
<div class="card"><div class="title">RAG is dead, right?? — Conference Talk</div>
<div class="meta mono">youtube · 2026-06-09 · 31 min <span class="chip q">queued</span></div></div>
</div></body></html>
+89
View File
@@ -0,0 +1,89 @@
# Spec — Chat with a video's stored transcript (ADR-027)
**Repo:** tapir · **Size:** medium · **Solo session.** Implements ADR-027. Read CLAUDE.md,
DECISIONS.md (ADR-021 transcript store, ADR-022 model chain, ADR-012 isolation, ADR-027), and
`docs/ui-spec.md` first. TBD, conventional commits, `task check` green per commit, `templ
generate` after view changes.
## FIRST: append ADR-027 to DECISIONS.md
ADR-027 text is provided separately (planning thread). Insert immediately before the
`## Rejected alternatives` heading, as the first commit, so the decision precedes the build.
## What this is
A per-video chat letting the user ask questions against a video's **already-stored** transcript,
entered from the summary view. Born from observed demand: the maintainer read real summaries and
some made him want to dig deeper — this gives that "I want more" reaction somewhere to go, without
watching the video.
## HARD CONSTRAINT — stored-transcript-only (the safety property)
Chat is available **ONLY** for videos that already have a stored transcript (ADR-021). It must
**never** trigger a caption fetch, never touch the rate gate, never reach YouTube. Entry being
"from a summarized video" guarantees the transcript exists. If somehow invoked on a video with no
stored transcript → show "transcript not available for chat", NO fetch. This is what makes the
feature safe by construction; do not add an on-demand-fetch path (explicitly deferred).
## 1. Entry point
- A "Dig deeper" / "Ask about this" affordance on the **summary view** of a summarized video
(not the list cards — the detail/summary page). Quiet, consistent with the existing card-state
styling.
- Opens a chat panel/view scoped to that one video, with its stored transcript as context.
## 2. The chat
- Read the stored transcript for the video (via the ADR-021 `TranscriptStore`, keyed by
`(provider, provider_video_id)`). No fetch.
- Send transcript + the user's question + minimal system framing to the chosen model via the
**existing LiteLLM gateway** (the same client the summarizer uses — a chat is a different
call, not a new integration).
- Stream or return the answer; render in the chat panel. HTMX/no-JS ethos — match the existing
app (the summarize status uses HTMX polling; chat can use a simple POST-and-render or HTMX
streaming if clean).
- **Transcript truncation:** reuse/respect `TAPIR_MAX_TRANSCRIPT_CHARS` (ADR-022) so a long
transcript fits the model context. If truncated, the chat should be honest that it's working
from a bounded portion (a quiet note), since answers about the tail of a long video may be
incomplete.
## 3. Model selection (the instrumentation win)
- **Default model = the model that produced this video's summary.** (Store/lookup which chain
model summarized it — if not already recorded, this is a small addition; if recording it is
non-trivial, default to the chain primary and note the gap.)
- **User can switch** among the ADR-022 chain models (`phi4-mini`, `gemma4-26b`,
`mistral-small` to start) via a simple selector in the chat panel. Switching re-runs against
the same transcript — this is deliberate model-comparison instrumentation.
- Respect the local-first / NDA posture: if `TAPIR_CLOUD_FALLBACK_MODEL=""` (cloud disabled),
the external model is NOT offered in the switcher — only local models. Chat must honor the same
"content stays local" guarantee as ADR-022.
## 4. Ephemeral (v1)
- No persisted chat history. Conversation lives for the session/page. No new table, no migration.
- (Multi-turn within a session is fine — keep the running messages in the request/page state —
but nothing is written to the DB.)
## 5. Isolation
- The transcript is shared/non-RLS (ADR-021) — fine, it's public content. But the chat is invoked
by a user about a video **in their feed**; confirm the entry path is reachable only for the
requesting user's own videos (the summary view is already RLS-scoped). Chat adds no new
user-data surface (ephemeral), so there's nothing new to RLS — but the test should confirm a
user can only open chat from their own summary view, not arbitrary video ids.
## Tests
- Chat on a video with a stored transcript → answer returned; assert NO caption-fetch / no
YouTube call occurs (the safety property — this is the key assertion).
- Chat invoked on a video with no stored transcript → honest "not available", NO fetch.
- Model switch → re-runs against the same transcript with the selected model; cloud model absent
from the switcher when `TAPIR_CLOUD_FALLBACK_MODEL=""`.
- Truncation honored for a long transcript; the bounded-context note shows.
- Entry is reachable only from the user's own summary view (isolation).
## Out of scope / deferred (record, don't build)
- **Persisted chat history** (per-user, RLS-scoped) — deferred until evidence anyone revisits a
conversation.
- **Show-source / transcript-verification UI** — the natural v2 (ADR-027 records it); v1 is
chat-only/trust-the-model. ADR-021's stored transcript makes v2 cheap when wanted.
- **On-demand fetch** for un-stored videos — would reintroduce the caption-fetch surface the
stored-only constraint removes. Not now.
- Anything that nudges the user to return (ADR-020 — gate contamination).
## Boundaries
Stored-transcript-only (HARD). No rate-gate/fetch surface. No auth changes. No new persisted
state in v1. Reuse the existing gateway client + truncation config; don't build a new model
integration.
+128
View File
@@ -0,0 +1,128 @@
# Spec — Onboarding "wow" burst: better picks, stronger model
**Repo:** tapir · **Size:** medium · **Solo session** (not a swarm).
> **Status: built (v0.25.0, ADR-028).** This supersedes the original investigate-first brief
> (committed as the prior version of this file): Phase 1 was run against the live pilot DB and its
> findings are folded into "Why this exists" below; Phase 2 was built as described here. The one
> brief lever NOT built — the honest "the rest fill in over the coming days" framing copy — is
> listed under *Explicitly NOT in this slice*.
**Why this exists.** A new user's first session decides whether they return (the Stage-0 gate,
VISION.md). On connect, Tapir fires a capped burst (≤`TAPIR_ONBOARD_SUMMARIZE_COUNT`, default 3)
that summarizes the user's newest unsummarized videos so the feed isn't empty (the burst itself
works — wired in `cmd/tapir/discovery.go``cmd/tapir/main.go` `onboard`). A Phase-1
investigation of the live pilot DB found the burst *fires* but delivers a **weak first
impression** for two concrete reasons, and ruled out a third idea:
1. **Picks are junk.** Selection is pure newest-first (`videos.NewestUnsummarizedVideoIDs`,
`ORDER BY published_at DESC`) with **zero quality signal**. Pilot user "Jonte"'s live burst-3
were a stock-ticker **livestream** + two regional news clips — the newest, not the best.
2. **Weakest model on the first impression.** All of Jonte's summaries ran on
`koala/phi4-mini` (the documented weak link — ADR-022 was born from its failures). The
stronger, brain-validated `iguana/gemma4-26b` was never used for the burst.
3. **Cached-first is empty at pilot scale — REJECTED.** The idea (summarize already-cached
transcripts instantly, zero fetch) dies on the numbers: only **11 videos** overlap between the
two pilot users (~3% of each library), **0** cached-and-unsummarized, and a new user's
newest-20 unsummarized are **20/20 NOT cached** — newest-first and cached-first are
structurally incompatible (fresh uploads are exactly what nobody has fetched yet). Not built.
This is a **curation/latency problem for ~3 videos, NOT a throughput/429 problem** — fetching 3
captions is nowhere near the rate limit. Nothing here fetches harder or pressures the rate gate;
it picks the right few videos and runs a better model on them.
Read `CLAUDE.md`, `DECISIONS.md` (esp. ADR-014, ADR-018, ADR-020, ADR-021, ADR-022, ADR-023,
and the new **ADR-028**), and `VISION.md` (the Stage-0 gate) first. TBD — commit directly to
`main`, one logical change per commit, conventional commits, `task check` green before each
commit, `templ generate` if any view changes (none expected).
## Decisions already made (do not reopen)
- **Not a throughput change.** The caption rate gate (ADR-014) is untouched — same pacing, same
priority lane (ADR-026). This slice changes *which* ≤3 videos the burst spends its fetches on
and *which model* summarizes them, never how fast or how many.
- **Cached-first is dropped** (ADR-028, the 3% overlap). The engine's existing read-stored-first
(ADR-021, `resolveTranscript`) stays — it already gives a free instant summary on the rare
cache hit, transparently. We do not *select* for cache hits.
- **has-captions is not a pre-fetch signal.** It is only knowable after a gate fetch (or a cache
hit, ~0 for new videos). Selection can only *avoid known-junk* (Shorts/live/over-long) — it
cannot *guarantee* captions. The spec is honest about this: better odds, not a promise.
- **No credentialed caption fetch** (ADR-010/ADR-026 dead end). **No client extension.**
## 1. Persist `duration_s` at discovery (the enabling change)
The `videos.duration_s` column exists (migration 001) but is **never written** — ADR-023's
`filterLowValue` (`internal/adapters/youtube/youtube.go`) already fetches each candidate's
duration via the cheap quota `videos.list` call, uses it to drop Shorts/live, then **discards
it**. Stop discarding:
- Add `DurationSeconds int` to `domain.Video`.
- In `filterLowValue`, set `DurationSeconds` on each kept video from the `videos.list` `meta`.
- `UpsertVideo` writes `duration_s`, **COALESCE-preserving** a known value (never overwrite a
real duration with 0/unknown), mirroring the `channel_title` backfill stance (migration 014).
- No new migration — the column is already there.
Consequence: a fresh user's connect-triggered discovery pass runs **before** the onboard burst
(`Enqueue`: `run()` then `onboard()`), so duration is populated for the burst's candidates at
connect. Existing rows backfill on their next discovery pass; until then their `duration_s` is
NULL and treated as "unknown" (§2).
## 2. Junk-avoiding burst selection
New store method, RLS-scoped via `withUser`:
```
OnboardBurstVideoIDs(ctx, userID string, limit, minSeconds, maxSeconds int) ([]string, error)
```
- Same base as the old `NewestUnsummarizedVideoIDs`: the user's videos with no summary yet,
`ORDER BY published_at DESC NULLS LAST, seen_at DESC`, `LIMIT limit`.
- **Exclude known-junk**: a row is dropped only when `duration_s IS NOT NULL` **and**
(`duration_s < minSeconds` OR `duration_s > maxSeconds`). A NULL duration is **unknown** — kept
(degrade-open: never starve the burst because metadata is missing), but ordered *after* rows
with a known-good duration so a freshly-enriched good pick wins when both exist.
- `minSeconds` reuses `TAPIR_MIN_VIDEO_SECONDS` (default 60 — the Shorts floor, ADR-023).
`maxSeconds` is new: `TAPIR_ONBOARD_MAX_VIDEO_SECONDS` (default 14400 = 4h) — drops the
multi-hour livestream VODs that pass the live filter once ended.
- `minSeconds<=0` and `maxSeconds<=0` each disable that bound (so `0/0` == the old
newest-first behaviour, the reversibility lever).
- The burst switches to this method; `NewestUnsummarizedVideoIDs` is removed (fully superseded —
`OnboardBurstVideoIDs(., 0, 0)` is identical pure-newest behaviour).
## 3. Stronger model for the burst
The burst summarizes only ≤3 videos, so a slower, stronger model is affordable exactly here.
- New config `TAPIR_ONBOARD_SUMMARIZER_MODEL` (default `iguana/gemma4-26b` — the brain-validated
homelab general-purpose model, already the ADR-022 fallback).
- Build a **burst-specific summarizer chain** that puts the onboard model **first**, then the
standard chain (primary → local fallback → cloud) as resilience, deduped. Wrap it in a
burst-specific `engineProcessor` reusing the same store/transcript-cache/sink — a pure wiring
choice, engine and ports unchanged (Clean Architecture, ADR-003).
- The `onboard` closure uses the burst processor instead of `app.Processor`.
- **Collapse cleanly**: when `OnboardSummarizerModel` is empty or equals `SummarizerModel`, the
onboard path reuses `app.Processor` (no separate chain) — the reversibility lever.
- Local-first preserved: the onboard model is a local alias; the cloud endpoint stays last in the
chain, so a client/NDA deployment with `TAPIR_CLOUD_FALLBACK_MODEL=""` keeps burst content
local too.
## 4. Behaviour spec + docs
- Add scenarios to `docs/use-cases/connect_account.feature` (the connect → burst flow): burst
skips a too-long/live video in favour of a reasonable-length one; burst summarizes with the
stronger model first. Map them in `scenarioCoverage` so `TestScenarioCoverage` stays green.
- Update `docs/architecture/architecture.md` (the onboarding-burst section) to describe the
junk-avoiding selection + the burst model override.
- ADR-028 in `DECISIONS.md` records the rationale (incl. the rejected cached-first lever).
## Success criteria
- `task check` green (fmt, vet, lint, `go test -p 1 ./...`).
- A unit test proves `OnboardBurstVideoIDs` drops a known too-long / sub-min video and keeps a
good one, newest-first, RLS-scoped, unsummarized-only.
- A test proves discovery persists `duration_s` and does not clobber it on re-upsert.
- A test proves the burst chain leads with the onboard model (then the standard chain).
- Config defaults + bounds tested (`OnboardMaxVideoSeconds`, `OnboardSummarizerModel`).
- No change to the rate gate, fetch pacing, or burst cap. `0/0` + empty model == prior behaviour.
## Explicitly NOT in this slice
- Cached-first selection (rejected, ADR-028).
- Any caption-availability *guarantee* (impossible pre-fetch).
- **Honest "taster" framing copy** ("summaries of a few of your videos to get you started — the
rest fill in over the coming days"). A good lever from the original brief, but it's a UI/copy
change with no backend dependency; deferred to a UI pass, tracked as an issue.
- Backfilling `duration_s` for existing rows via a migration (it backfills lazily on discovery).
- Return-nudges / digests (ADR-020: poisons the unprompted-return signal).
- Raising fetch throughput, multi-IP, or Whisper (out of scope; the gate is deliberate).
+1
View File
@@ -179,4 +179,5 @@ distinguishable.
| **"Summarize now" foreground path** | Unified quiet nudge button on actionable non-summarized cards. Five explicit card states — (1) summarized: chip + no button; (2) no captions (`transcript_status = 'none'`): "No transcript available", no button; (3) queued: "Queued" chip, no button; (4) rate-limited: "Fetching soon…" + "Summarize now" → `POST /v/{id}/retry-now` (clears `rate_limited_at`, triggers engine); (5) pending: "Not summarized" + "Summarize now" → `POST /v/{id}/summarize` (queues + triggers engine). One verb, one style (`.btn-quiet`); backend difference invisible to user. Both handlers call `ProcessVideo` through `globalFetchGate`. Rate gate respected, not bypassed — this is onboarding prioritisation. | Fast onboarding value; honest dead-end for no-captions videos (no button that fails). | `internal/web/handlers.go` (`handleRetryNow`, `handleRequestSummarize`); `internal/web/views.templ` (`VideoCard`) |
| **Pipeline stats bar** | A one-line status bar above the video list: `N summarized · M fetching soon · K no captions`. Computed from the unfiltered row set; hidden when all videos are summarized. Gives the user a clear read on pipeline state without any interaction. | Replaces the "why is nothing happening?" confusion when most videos are pending or rate-limited. | `internal/web/view.go` (`PipelineStats`, `pipelineStats`) |
| **Unavailable channels (account page)** | The `/account` page shows a "Unavailable channels" section when any channels returned HTTP 404 on the last discovery pass. Lists channel name, an "unavailable" badge, and the first-seen date. Data sourced from the `channel_errors` table (migration 013). | Surfaces silent failures so users know why some subscribed channels produce no new videos. | migration 013; `internal/web/account.go`; `internal/adapters/youtube/youtube.go` (`domain.ErrChannelUnavailable`) |
| **Visual refresh — charm-reader theme + light/dark toggle (ADR-032)** | One layout in two palettes expressed as CSS custom properties: a warm "reader" light theme (sketch B) and a "cozy terminal" dark theme (sketch C). Palette is chosen in cascade order — `:root` light default, an OS-preference dark block scoped to `:root:not([data-theme])` so it only applies absent an explicit choice, and `:root[data-theme="dark"\|"light"]` set by a header toggle that outranks the media query and persists in `localStorage` (guarded; degrades to OS default). An init script in `<head>` applies the stored choice before paint (no flash). Charm touches: monospace meta lines, accent uppercase section dividers with a trailing rule, pill buttons, a lifted/accent-edged expanded card. Error/danger shades became `--err-*` tokens so they follow the theme without per-block dark overrides. Sketches kept as the design record under `docs/sketches/`. | The UI read "flat and boring"; the charm/TUI aesthetic makes it distinctive and gives a real light/dark choice rather than OS-only. | ADR-032; `docs/use-cases/visual_theme.feature`; `internal/web/visual_theme_test.go`; `internal/web/view.go` (`stylesheet`, `themeScript`), `internal/web/views.templ` (Layout/PublicLayout) |
| **Recency window + sparse-state honesty (ADR-020)** | Supersedes the copy/sort in the rows above. Auto-summarize is bounded to videos published within `TAPIR_AUTO_SUMMARIZE_WINDOW` (~7d); older un-summarized videos collapse behind a single "Show N older videos — summarize on demand" disclosure, and caption-less videos collapse to a one-line count (not N cards). List order is now `summarized-first, published_at DESC NULLS LAST`. Copy reframed for honest scarcity: pipeline bar reads "N ready · M in queue · K no captions" (no "fetching soon"); a gradual-fill note explains the rate limit; the nudge verb is "Summarize" (not "Summarize now"); the queued card says "summarizing shortly"; the empty-connected state drops the impossible `tapir run` instruction. Detail leads with Takeaways. Filters slimmed (no date pickers; hidden when empty); watched/skipped segmented; back link on detail; empty terms checkbox removed. | Make the sparse reality legible and honest instead of implying abundance/imminence; bound auto load so the back-catalogue doesn't re-drive the caption gate. Never fetch harder — scarcity is surfaced, not engineered around. | ADR-020; `2384c47`, `3df0459`, `40b703e`, `a1a5217`, `4a0a56e`, `9bf1c31`, `980638d`, `12fb031`, `f775441`, `51aa5d9` |
+13 -1
View File
@@ -31,5 +31,17 @@ Feature: Local-first AI with optional BYO fallback
When any transcript is summarized
Then my content is only ever sent to the local AI stack
# "Reliably" is operationalized as: Primary returned without error within timeout.
Scenario: A model returns unparseable output and the next endpoint succeeds
Given the local AI stack is available
But the primary model returns output that cannot be parsed into a summary
And a fallback model is configured
When a transcript is summarized
Then Tapir falls back to the next model in the chain
And the summary records fallback_used as true
# "Reliably" is operationalized as: an endpoint returned a PARSEABLE summary
# within timeout. A 200 with malformed JSON (or highlights emitted as a bare
# string) counts as a failure and advances the chain (ADR-022). Endpoints are
# tried in order, locals first, so the external worst-case model only ever sees
# content after every local endpoint has failed.
# Quality scoring may be added later without changing these scenarios.
+69
View File
@@ -0,0 +1,69 @@
Feature: Chat with a video's stored transcript
As a reader whose summary made me want to dig deeper
I want to ask questions about the video without watching it
So that I can go further on the ones worth it, without leaving the reader
# ADR-027. The load-bearing constraint is safety-by-construction: chat runs
# ONLY against an already-stored transcript (ADR-021) and never fetches captions,
# never touches the rate gate, never reaches YouTube. Entry is from the summary
# view of one's OWN video; the conversation is ephemeral (no persisted history).
Background:
Given I have a summarized video with a stored transcript
Scenario: A summary view offers a deeper-dive into the video
When I view the summary
Then I see a "dig deeper" affordance that opens a chat about this video
And it opens the chat in place, below the summary, without leaving the page
Scenario: The summary and the chat are on one page
When I open the chat
Then the summary stays visible alongside the chat
And I can read the summary while I ask questions
Scenario: Ask a question answered from the stored transcript
When I ask a question in the chat
Then the answer is produced from the stored transcript
And no caption fetch and no YouTube call occurs
Scenario: Chat never fetches captions or reaches YouTube
When I ask a question in the chat
Then Tapir reads only the stored transcript
And the caption-fetch and video-fetch paths are never invoked
Scenario: A video with no stored transcript offers no chat
Given a video that has no stored transcript
When I open the chat for it
Then I am told chat is not available
And no fetch is attempted and no model is called
Scenario: The default model is the summary's model and is switchable
When I open the chat
Then the model defaults to the model that produced the summary
And I can switch among the offered chain models
Scenario: Switching models re-runs against the same transcript
When I ask a question with a different chain model selected
Then the chosen model answers
And it answers against the same stored transcript
Scenario: The cloud model is hidden when cloud is disabled
Given the cloud fallback model is disabled
Then the chat switcher offers only local models
Scenario: A long transcript is bounded and the chat says so
Given the stored transcript is longer than the model budget
When I ask a question
Then the answer is produced from a bounded portion
And the chat notes that it worked from a bounded portion
Scenario: A multi-turn conversation is ephemeral
When I ask a follow-up question
Then the prior turn is carried into the answer
And nothing about the conversation is written to the database
Scenario: Chat is reachable only from my own summary view
Given another user has a summarized video with a stored transcript
When I try to open the chat for their video
Then I get a not-found response
And no model is called
+9 -2
View File
@@ -24,12 +24,19 @@ Feature: Connect and manage video accounts
Then a discovery pass for my account is triggered right away
And I do not have to wait for the next scheduled pass to see my videos
Scenario: Connecting summarizes my newest videos right away
Scenario: Connecting summarizes my best recent videos right away
Given I have no connected video accounts
When I connect my YouTube account
Then up to the onboarding cap of my newest videos are summarized through the rate gate
Then up to the onboarding cap of my newest likely-good videos are summarized through the rate gate
And videos whose known duration is too short or too long are skipped
And the rest are left to the scheduled recency-bounded pass
Scenario: The onboarding burst summarizes with a stronger model
Given I have no connected video accounts
When I connect my YouTube account
Then the burst summarizes with the stronger onboarding model first
And the standard summarizer chain still follows as a fallback
Scenario: Tokens are never stored in the clear
When I connect any video account
Then no OAuth token value is stored in the database
+37
View File
@@ -0,0 +1,37 @@
Feature: Inline-expand summary + Q&A in the list (ADR-031, #16)
As a reader skimming my summaries
I want to open a summary and its Q&A in place in the list
So that I get the full read and follow-up without leaving the list (SPA-like, no page hop)
# HTMX inline-expand, no SPA framework (ADR-031). Each scenario maps to a Go test
# in scenario_coverage_test.go (the BDD name-coverage gate).
Scenario: A summarized card expands to the full summary in place
Given a summarized video in my list
When I expand its card
Then the full summary, highlights, and takeaways are returned as an in-place card fragment, not a full page
Scenario: An expanded card collapses back to the compact card
Given an expanded card
When I collapse it
Then the compact card fragment is returned in its place
Scenario: The expanded card offers the Q&A dock
Given chat is enabled
When a summarized card is expanded
Then the expanded card includes the deeper-dive chat affordance for that video
Scenario: Only a summarized card offers expand
Given a discovered but not-yet-summarized card
When the card is rendered
Then it shows its summarize/queue footer and no expand affordance
Scenario: With JS off the card still reaches the full summary
Given a summarized card
When it is rendered
Then its expand affordance carries an href to the detail page as a no-JS fallback
Scenario: The detail page and the expanded card show the same summary
Given a summarized video
When I view it on the detail page and as an expanded card
Then both render the same summary body (one shared fragment, no drift)
+46
View File
@@ -0,0 +1,46 @@
Feature: Observability — timing and metrics for performance and UX (ADR-030, #15)
As the maintainer running Tapir for pilot users
I want timing and Prometheus metrics for the activities that drive performance and UX
So that I can see latency, model behaviour, and usage and feed the Stage-0 eval gate
# AI metrics are the priority (ADR-030 R3). Each scenario maps to a Go test in
# test/acceptance/scenario_coverage_test.go (the BDD name-coverage gate).
Scenario: Summarization latency is recorded per endpoint
Given the summarizer runs a transcript through its endpoint chain
When an endpoint returns a parseable summary
Then the summarize latency is recorded with the model, outcome "success", and whether it was a fallback
Scenario: A failing summarizer endpoint records its failure outcome
Given the summarizer runs a transcript through its endpoint chain
When an endpoint errors or returns unparseable output
Then the summarize latency is recorded with outcome "error" or "parse_error" before the chain advances
Scenario: Caption fetch latency is recorded by outcome
Given a caption fetch is attempted for a video
When it resolves to captions, no captions, or a rate limit
Then the caption-fetch latency is recorded labelled by that outcome
Scenario: LLM token usage is recorded from the completion
Given an LLM completion returns a usage block with prompt and completion tokens
When the client finishes the call
Then the prompt and completion tokens are recorded for that model
Scenario: Q&A answer latency is recorded
Given a user asks a question about a video
When the answer is produced from the stored transcript
Then the chat answer latency is recorded for the answering model
Scenario: HTTP requests are counted by route, method, and status
Given the metrics HTTP middleware wraps the app
When a request is served against a registered route
Then it is counted and timed under the bounded route pattern, not the raw path
Scenario: A successful login is counted
Given a user completes the OIDC callback and a session is established
Then the login counter is incremented
Scenario: The metrics endpoint is not on the public app port
Given the service is running
When the public app mux is inspected
Then it exposes no /metrics route metrics are served on the dedicated metrics port only
@@ -31,5 +31,16 @@ Feature: Summarize new videos from subscribed channels
When the watcher sees "Designing for Attention" again
Then Tapir does not produce a second summary for it
Scenario: Re-analyzing a stored video does not re-fetch its transcript
Given a transcript for "Designing for Attention" is already stored
When the video is summarized again
Then Tapir reads the stored transcript
And Tapir does not fetch captions from YouTube
# Captions-first is the core path (ADR-007). Audio-download + speech-to-text is
# deferred and intentionally has no scenario here yet.
#
# Transcript persistence (ADR-021): the stored transcript is shared, keyed by
# (provider, provider_video_id) and read before any caption fetch, so the
# re-analysis scenario above also covers paste-a-URL and the onboarding burst —
# both summarize through the same engine chokepoint.
+29
View File
@@ -0,0 +1,29 @@
Feature: Visual refresh — one charm-reader layout, light + dark themes (ADR-032, #17)
As a reader who likes a TUI/charm aesthetic
I want a fresh look with a light and a dark theme
So that the app feels distinctive and reads well in either mode
# Direction B (light) + C (dark) are one layout in two palettes (CSS variables),
# plus a persisted toggle and an OS-preference default. Colours are reviewed via
# the mockups in docs/sketches/, not unit-tested; these scenarios cover the
# testable structure. Each maps to a Go test in scenario_coverage_test.go.
Scenario: Light and dark themes share one layout via CSS variables
Given the app stylesheet
When it is rendered
Then it defines a light palette on the root and a dark palette under data-theme="dark", with no duplicate markup
Scenario: Without a stored choice the theme follows the OS preference
Given a visitor with no saved theme
When the page loads
Then a prefers-color-scheme dark block applies the dark palette automatically
Scenario: A persisted toggle switches light and dark
Given any page
When it is rendered
Then it includes a theme-toggle control and a small script that flips data-theme and persists the choice
Scenario: The expanded card embeds the video player
Given a summarized video with a valid provider id
When its card is expanded
Then the expanded card includes the embedded video player
+12 -2
View File
@@ -1,4 +1,4 @@
module gitea.d-ma.be/mathias/tapir
module git.d-ma.be/mathias/tapir
go 1.26.1
@@ -9,23 +9,33 @@ require (
github.com/go-jose/go-jose/v4 v4.1.4
github.com/golang-migrate/migrate/v4 v4.19.1
github.com/jackc/pgx/v5 v5.9.2
github.com/prometheus/client_golang v1.23.2
github.com/prometheus/client_model v0.6.2
github.com/stretchr/testify v1.11.1
golang.org/x/crypto v0.45.0
golang.org/x/oauth2 v0.36.0
golang.org/x/time v0.15.0
)
require (
github.com/beorn7/perks v1.0.1 // indirect
github.com/cespare/xxhash/v2 v2.3.0 // indirect
github.com/davecgh/go-spew v1.1.2-0.20180830191138-d8f796af33cc // indirect
github.com/jackc/pgerrcode v0.0.0-20220416144525-469b46aa5efa // indirect
github.com/jackc/pgpassfile v1.0.0 // indirect
github.com/jackc/pgservicefile v0.0.0-20240606120523-5a60cdf6a761 // indirect
github.com/jackc/puddle/v2 v2.2.2 // indirect
github.com/kylelemons/godebug v1.1.0 // indirect
github.com/lib/pq v1.10.9 // indirect
github.com/munnerz/goautoneg v0.0.0-20191010083416-a7dc8b61c822 // indirect
github.com/pmezard/go-difflib v1.0.1-0.20181226105442-5d4384ee4fb2 // indirect
github.com/prometheus/common v0.66.1 // indirect
github.com/prometheus/procfs v0.16.1 // indirect
github.com/rogpeppe/go-internal v1.15.0 // indirect
github.com/xi2/xz v0.0.0-20171230120015-48954b6210f8 // indirect
go.yaml.in/yaml/v2 v2.4.2 // indirect
golang.org/x/sync v0.18.0 // indirect
golang.org/x/sys v0.41.0 // indirect
golang.org/x/text v0.31.0 // indirect
google.golang.org/protobuf v1.36.8 // indirect
gopkg.in/yaml.v3 v3.0.1 // indirect
)
+26 -6
View File
@@ -4,6 +4,10 @@ github.com/Microsoft/go-winio v0.6.2 h1:F2VQgta7ecxGYO8k3ZZz3RS8fVIXVxONVUPlNERo
github.com/Microsoft/go-winio v0.6.2/go.mod h1:yd8OoFMLzJbo9gZq8j5qaps8bJ9aShtEA8Ipt1oGCvU=
github.com/a-h/templ v0.3.1020 h1:ypAT/L5ySWEnZ6Zft/5yfoWXYYkhFNvEFOeeqecg4tw=
github.com/a-h/templ v0.3.1020/go.mod h1:A2DlK61v+K+NRoGnhmYbNYVmtYHcFO5/AisMvBdDxTM=
github.com/beorn7/perks v1.0.1 h1:VlbKKnNfV8bJzeqoa4cOKqO6bYr3WgKZxO8Z16+hsOM=
github.com/beorn7/perks v1.0.1/go.mod h1:G2ZrVWU2WbWT9wwq4/hrbKbnv/1ERSJQ0ibhJ6rlkpw=
github.com/cespare/xxhash/v2 v2.3.0 h1:UL815xU9SqsFlibzuggzjXhog7bL6oX9BbNZnL2UFvs=
github.com/cespare/xxhash/v2 v2.3.0/go.mod h1:VGX0DQ3Q6kWi7AoAeZDth3/j3BFtOZR5XLFGgcrjCOs=
github.com/containerd/errdefs v1.0.0 h1:tg5yIfIlQIrxYtu9ajqY42W3lpS19XqdxRQeEwYG8PI=
github.com/containerd/errdefs v1.0.0/go.mod h1:+YBYIdtsnF4Iw6nWZhJcqGSg/dwvV7tyJ/kCkyJ2k+M=
github.com/containerd/errdefs/pkg v0.3.0 h1:9IKJ06FvyNlexW690DXuQNx2KA2cUJXx151Xdx3ZPPE=
@@ -37,8 +41,8 @@ github.com/gogo/protobuf v1.3.2 h1:Ov1cvc58UF3b5XjBnZv7+opcTcQFZebYjWzi34vdm4Q=
github.com/gogo/protobuf v1.3.2/go.mod h1:P1XiOD3dCwIKUDQYPy72D8LYyHL2YPYrpS2s69NZV8Q=
github.com/golang-migrate/migrate/v4 v4.19.1 h1:OCyb44lFuQfYXYLx1SCxPZQGU7mcaZ7gH9yH4jSFbBA=
github.com/golang-migrate/migrate/v4 v4.19.1/go.mod h1:CTcgfjxhaUtsLipnLoQRWCrjYXycRz/g5+RWDuYgPrE=
github.com/google/go-cmp v0.6.0 h1:ofyhxvXcZhMsU5ulbFiLKl/XBFqE1GSq7atu8tAmTRI=
github.com/google/go-cmp v0.6.0/go.mod h1:17dUlkBOakJ0+DkrSSNjCkIjxS6bF9zb3elmeNGIjoY=
github.com/google/go-cmp v0.7.0 h1:wk8382ETsv4JYUZwIsn6YpYiWiBsYLSJiTsyBybVuN8=
github.com/google/go-cmp v0.7.0/go.mod h1:pXiqmnSA92OHEEa9HXL2W4E7lf9JzCmGVUdgjX3N/iU=
github.com/jackc/pgerrcode v0.0.0-20220416144525-469b46aa5efa h1:s+4MhCQ6YrzisK6hFJUX53drDT4UsSW3DEhKn0ifuHw=
github.com/jackc/pgerrcode v0.0.0-20220416144525-469b46aa5efa/go.mod h1:a/s9Lp5W7n/DD0VrVoyJ00FbP2ytTPDVOivvn2bMlds=
github.com/jackc/pgpassfile v1.0.0 h1:/6Hmqy13Ss2zCq62VdNG8tM1wchn8zjSGOBJ6icpsIM=
@@ -49,10 +53,14 @@ github.com/jackc/pgx/v5 v5.9.2 h1:3ZhOzMWnR4yJ+RW1XImIPsD1aNSz4T4fyP7zlQb56hw=
github.com/jackc/pgx/v5 v5.9.2/go.mod h1:mal1tBGAFfLHvZzaYh77YS/eC6IX9OWbRV1QIIM0Jn4=
github.com/jackc/puddle/v2 v2.2.2 h1:PR8nw+E/1w0GLuRFSmiioY6UooMp6KJv0/61nB7icHo=
github.com/jackc/puddle/v2 v2.2.2/go.mod h1:vriiEXHvEE654aYKXXjOvZM39qJ0q+azkZFrfEOc3H4=
github.com/kr/pretty v0.3.0 h1:WgNl7dwNpEZ6jJ9k1snq4pZsg7DOEN8hP9Xw0Tsjwk0=
github.com/kr/pretty v0.3.0/go.mod h1:640gp4NfQd8pI5XOwp5fnNeVWj67G7CFk/SaSQn7NBk=
github.com/klauspost/compress v1.18.0 h1:c/Cqfb0r+Yi+JtIEq73FWXVkRonBlf0CRNYc8Zttxdo=
github.com/klauspost/compress v1.18.0/go.mod h1:2Pp+KzxcywXVXMr50+X0Q/Lsb43OQHYWRCY2AiWywWQ=
github.com/kr/pretty v0.3.1 h1:flRD4NNwYAUpkphVc1HcthR4KEIFJ65n8Mw5qdRn3LE=
github.com/kr/pretty v0.3.1/go.mod h1:hoEshYVHaxMs3cyo3Yncou5ZscifuDolrwPKZanG3xk=
github.com/kr/text v0.2.0 h1:5Nx0Ya0ZqY2ygV366QzturHI13Jq95ApcVaJBhpS+AY=
github.com/kr/text v0.2.0/go.mod h1:eLer722TekiGuMkidMxC/pM04lWEeraHUUmBw8l2grE=
github.com/kylelemons/godebug v1.1.0 h1:RPNrshWIDI6G2gRW9EHilWtl7Z6Sb1BR0xunSBf0SNc=
github.com/kylelemons/godebug v1.1.0/go.mod h1:9/0rRGxNHcop5bhtWyNeEfOS8JIWk580+fNqagV/RAw=
github.com/lib/pq v1.10.9 h1:YXG7RB+JIjhP29X+OtkiDnYaXQwpS4JEWq7dtCCRUEw=
github.com/lib/pq v1.10.9/go.mod h1:AlVN5x4E4T544tWzH6hKfbfQvm3HdbOxrmggDNAPY9o=
github.com/moby/docker-image-spec v1.3.1 h1:jMKff3w6PgbfSa69GfNg+zN/XLhfXJGnEx3Nl2EsFP0=
@@ -61,6 +69,8 @@ github.com/moby/term v0.5.0 h1:xt8Q1nalod/v7BqbG21f8mQPqH+xAaC9C3N3wfWbVP0=
github.com/moby/term v0.5.0/go.mod h1:8FzsFHVUBGZdbDsJw/ot+X+d5HLUbvklYLJ9uGfcI3Y=
github.com/morikuni/aec v1.0.0 h1:nP9CBfwrvYnBRgY6qfDQkygYDmYwOilePFkwzv4dU8A=
github.com/morikuni/aec v1.0.0/go.mod h1:BbKIizmSmc5MMPqRYbxO4ZU0S0+P200+tUnFx7PXmsc=
github.com/munnerz/goautoneg v0.0.0-20191010083416-a7dc8b61c822 h1:C3w9PqII01/Oq1c1nUAm88MOHcQC9l5mIlSMApZMrHA=
github.com/munnerz/goautoneg v0.0.0-20191010083416-a7dc8b61c822/go.mod h1:+n7T8mK8HuQTcFwEeznm/DIxMOiR9yIdICNftLE1DvQ=
github.com/opencontainers/go-digest v1.0.0 h1:apOUWs51W5PlhuyGyz9FCeeBIOUDA/6nW8Oi/yOhh5U=
github.com/opencontainers/go-digest v1.0.0/go.mod h1:0JzlMkj0TRzQZfJkVvzbP0HBR3IKzErnv2BNG4W4MAM=
github.com/opencontainers/image-spec v1.1.0 h1:8SG7/vwALn54lVB/0yZ/MMwhFrPYtpEHQb2IpWsCzug=
@@ -70,6 +80,14 @@ github.com/pkg/errors v0.9.1/go.mod h1:bwawxfHBFNV+L2hUp1rHADufV3IMtnDRdf1r5NINE
github.com/pmezard/go-difflib v1.0.0/go.mod h1:iKH77koFhYxTK1pcRnkKkqfTogsbg7gZNVY4sRDYZ/4=
github.com/pmezard/go-difflib v1.0.1-0.20181226105442-5d4384ee4fb2 h1:Jamvg5psRIccs7FGNTlIRMkT8wgtp5eCXdBlqhYGL6U=
github.com/pmezard/go-difflib v1.0.1-0.20181226105442-5d4384ee4fb2/go.mod h1:iKH77koFhYxTK1pcRnkKkqfTogsbg7gZNVY4sRDYZ/4=
github.com/prometheus/client_golang v1.23.2 h1:Je96obch5RDVy3FDMndoUsjAhG5Edi49h0RJWRi/o0o=
github.com/prometheus/client_golang v1.23.2/go.mod h1:Tb1a6LWHB3/SPIzCoaDXI4I8UHKeFTEQ1YCr+0Gyqmg=
github.com/prometheus/client_model v0.6.2 h1:oBsgwpGs7iVziMvrGhE53c/GrLUsZdHnqNwqPLxwZyk=
github.com/prometheus/client_model v0.6.2/go.mod h1:y3m2F6Gdpfy6Ut/GBsUqTWZqCUvMVzSfMLjcu6wAwpE=
github.com/prometheus/common v0.66.1 h1:h5E0h5/Y8niHc5DlaLlWLArTQI7tMrsfQjHV+d9ZoGs=
github.com/prometheus/common v0.66.1/go.mod h1:gcaUsgf3KfRSwHY4dIMXLPV0K/Wg1oZ8+SbZk/HH/dA=
github.com/prometheus/procfs v0.16.1 h1:hZ15bTNuirocR6u0JZ6BAHHmwS1p8B4P6MRqxtzMyRg=
github.com/prometheus/procfs v0.16.1/go.mod h1:teAbpZRB1iIAJYREa1LsoWUXykVXA1KlTmWl8x/U+Is=
github.com/rogpeppe/go-internal v1.15.0 h1:D0RCU5rMAp+SpgkiNdrjfJ+LX4J1M32V2NeCY7EJ6hc=
github.com/rogpeppe/go-internal v1.15.0/go.mod h1:DrUVZyrJU+txYW5/1kwtXQSMFio52ZOxX7yM1VHvnxs=
github.com/stretchr/objx v0.1.0/go.mod h1:HFkY916IF+rwdDfMAkV7OtwuqBVzrE8GR6GFx+wExME=
@@ -91,8 +109,8 @@ go.opentelemetry.io/otel/trace v1.37.0 h1:HLdcFNbRQBE2imdSEgm/kwqmQj1Or1l/7bW6mx
go.opentelemetry.io/otel/trace v1.37.0/go.mod h1:TlgrlQ+PtQO5XFerSPUYG0JSgGyryXewPGyayAWSBS0=
go.uber.org/goleak v1.3.0 h1:2K3zAYmnTNqV73imy9J1T3WC+gmCePx2hEGkimedGto=
go.uber.org/goleak v1.3.0/go.mod h1:CoHD4mav9JJNrW/WLlf7HGZPjdw8EucARQHekz1X6bE=
golang.org/x/crypto v0.45.0 h1:jMBrvKuj23MTlT0bQEOBcAE0mjg8mK9RXFhRH6nyF3Q=
golang.org/x/crypto v0.45.0/go.mod h1:XTGrrkGJve7CYK7J8PEww4aY7gM3qMCElcJQ8n8JdX4=
go.yaml.in/yaml/v2 v2.4.2 h1:DzmwEr2rDGHl7lsFgAHxmNz/1NlQ7xLIrlN2h5d1eGI=
go.yaml.in/yaml/v2 v2.4.2/go.mod h1:081UH+NErpNdqlCXm3TtEran0rJZGxAYx9hb/ELlsPU=
golang.org/x/oauth2 v0.36.0 h1:peZ/1z27fi9hUOFCAZaHyrpWG5lwe0RJEEEeH0ThlIs=
golang.org/x/oauth2 v0.36.0/go.mod h1:YDBUJMTkDnJS+A4BP4eZBjCqtokkg1hODuPjwiGPO7Q=
golang.org/x/sync v0.18.0 h1:kr88TuHDroi+UVf+0hZnirlk8o8T+4MrK6mr60WkH/I=
@@ -103,6 +121,8 @@ golang.org/x/text v0.31.0 h1:aC8ghyu4JhP8VojJ2lEHBnochRno1sgL6nEi9WGFGMM=
golang.org/x/text v0.31.0/go.mod h1:tKRAlv61yKIjGGHX/4tP1LTbc13YSec1pxVEWXzfoeM=
golang.org/x/time v0.15.0 h1:bbrp8t3bGUeFOx08pvsMYRTCVSMk89u4tKbNOZbp88U=
golang.org/x/time v0.15.0/go.mod h1:Y4YMaQmXwGQZoFaVFk4YpCt4FLQMYKZe9oeV/f4MSno=
google.golang.org/protobuf v1.36.8 h1:xHScyCOEuuwZEc6UtSOvPbAT4zRh0xcNRYekJwfqyMc=
google.golang.org/protobuf v1.36.8/go.mod h1:fuxRtAxBytpl4zzqUh6/eyUujkJdNiuEkXntxiD/uRU=
gopkg.in/check.v1 v0.0.0-20161208181325-20d25e280405/go.mod h1:Co6ibVJAznAaIkqp8huTwlJQCZ016jof/cbN4VW5Yz0=
gopkg.in/check.v1 v1.0.0-20201130134442-10cb98267c6c h1:Hei/4ADfdWqJk1ZMxUNpqntNwaWcugrBjAiHlqqRiVk=
gopkg.in/check.v1 v1.0.0-20201130134442-10cb98267c6c/go.mod h1:JHkPIbrfpd72SG/EVd6muEfDQjcINNoR0C8j2r3qZ4Q=
+179
View File
@@ -0,0 +1,179 @@
// Package chat implements the per-video deeper-dive chat (ADR-027): a read-only
// QA over a video's ALREADY-STORED transcript (ADR-021). It is the enforcement
// point for the feature's load-bearing safety property — stored-transcript-only:
// the Service has NO VideoSource and NO caption-fetch dependency, only a
// Completer factory, so it CANNOT reach YouTube or the rate gate by construction.
// The caller supplies the stored transcript text; chat never fetches.
//
// It reuses the same LiteLLM gateway as the summarizer (a chat is a different
// call, not a new integration) and the same transcript-truncation discipline
// (TAPIR_MAX_TRANSCRIPT_CHARS) so a long transcript fits a small-context model.
package chat
import (
"context"
"fmt"
"log/slog"
"strings"
"time"
"unicode/utf8"
"git.d-ma.be/mathias/tapir/internal/metrics"
)
// Completer is the minimal LLM chat surface the Service needs. *llm.Client
// satisfies it; tests use a fake. It is the SAME surface the summarizer uses.
type Completer interface {
Complete(ctx context.Context, system, user string) (string, error)
}
// Turn is one completed exchange in an ephemeral, session-only conversation
// (ADR-027 v1: nothing is persisted).
type Turn struct {
Question string
Answer string
}
// Request is one chat turn: the chosen model, the stored transcript text, the
// prior turns (for multi-turn context within the session), and the new question.
type Request struct {
Model string
Transcript string
History []Turn
Question string
}
// Reply is the model's answer plus whether the transcript was bounded to fit the
// model context (so the UI can be honest that an answer about the tail of a long
// video may be incomplete).
type Reply struct {
Answer string
Truncated bool
}
// Service answers questions against a stored transcript via a switchable set of
// models. models is the ordered, local-first list offered to the user (the cloud
// model is simply absent when disabled — see cmd wiring); maxChars bounds the
// transcript sent to any model (0 = unbounded). newClient builds a Completer for
// a chosen model alias (the same gateway, a different alias).
type Service struct {
newClient func(model string) Completer
models []string
maxChars int
}
// New constructs a Service. models must be non-empty and already filtered to the
// offerable set (cloud excluded when disabled) and de-duplicated by the caller.
func New(newClient func(model string) Completer, models []string, maxChars int) *Service {
return &Service{newClient: newClient, models: models, maxChars: maxChars}
}
// Models returns a copy of the offerable model list (local-first order).
func (s *Service) Models() []string {
out := make([]string, len(s.models))
copy(out, s.models)
return out
}
// offers reports whether model is in the offerable set — the guard that keeps an
// arbitrary, un-offered alias (e.g. a forged form value) from reaching the gateway.
func (s *Service) offers(model string) bool {
for _, m := range s.models {
if m == model {
return true
}
}
return false
}
// DefaultModel resolves the model a fresh chat opens with: the summary's own
// model when it is still an offered option (the ADR-027 default — chat continues
// in the model that produced the summary), otherwise the first offered model.
// Returns "" only when no models are configured.
func (s *Service) DefaultModel(summaryModel string) string {
if summaryModel != "" && s.offers(summaryModel) {
return summaryModel
}
if len(s.models) > 0 {
return s.models[0]
}
return ""
}
// Answer runs one chat turn. The model is forced back to a default if the request
// names an un-offered alias, so chat can never call the gateway with an arbitrary
// model. The transcript is truncated up front (reporting whether it was cut) and
// passed as system context; the running conversation is the user message.
func (s *Service) Answer(ctx context.Context, req Request) (Reply, error) {
if len(s.models) == 0 {
return Reply{}, fmt.Errorf("chat: no models configured")
}
model := req.Model
if !s.offers(model) {
model = s.DefaultModel("")
}
transcript, truncated := truncate(req.Transcript, s.maxChars)
system := buildSystem(transcript, truncated)
user := buildUser(req.History, req.Question)
start := time.Now()
out, err := s.newClient(model).Complete(ctx, system, user)
if err != nil {
return Reply{}, fmt.Errorf("chat: %s: %w", model, err)
}
dur := time.Since(start)
metrics.ObserveChat(model, dur)
slog.Default().Info("chat answer", "model", model, "elapsed_ms", dur.Milliseconds())
answer := strings.TrimSpace(out)
if answer == "" {
return Reply{}, fmt.Errorf("chat: %s returned an empty answer", model)
}
return Reply{Answer: answer, Truncated: truncated}, nil
}
const systemPreamble = `You are Tapir, answering questions about ONE video using ONLY the transcript below.
Ground every answer in the transcript. If the transcript does not contain the answer, say so plainly rather than guessing.`
const truncatedNote = `
The transcript below is truncated to fit the model — if a question seems to concern something missing, note it may be beyond the available portion.`
// buildSystem frames the model as a transcript-grounded QA assistant and embeds
// the (possibly truncated) transcript as context.
func buildSystem(transcript string, truncated bool) string {
var b strings.Builder
b.WriteString(systemPreamble)
if truncated {
b.WriteString(truncatedNote)
}
b.WriteString("\n\nTranscript:\n")
b.WriteString(transcript)
return b.String()
}
// buildUser renders the running conversation as the user message: prior turns as
// Q/A pairs followed by the new question. Folding history into one message keeps
// the Completer surface (a single system+user call) unchanged — no new llm method.
func buildUser(history []Turn, question string) string {
var b strings.Builder
for _, t := range history {
fmt.Fprintf(&b, "Q: %s\nA: %s\n\n", t.Question, t.Answer)
}
fmt.Fprintf(&b, "Q: %s", question)
return b.String()
}
// truncate caps content to max bytes on a UTF-8 rune boundary, reporting whether
// it cut. It mirrors the summarizer's truncation discipline (ADR-022) but returns
// the cut flag so the chat UI can be honest about a bounded transcript. A
// non-positive max (or content already within budget) returns content unchanged.
func truncate(content string, max int) (string, bool) {
if max <= 0 || len(content) <= max {
return content, false
}
cut := max
for cut > 0 && !utf8.RuneStart(content[cut]) {
cut--
}
return content[:cut], true
}
+164
View File
@@ -0,0 +1,164 @@
package chat
import (
"context"
"errors"
"strings"
"testing"
)
// recordingCompleter captures the system+user it was asked with and returns a
// canned answer (or error). It also records which model alias built it.
type recordingCompleter struct {
model string
lastSystem string
lastUser string
answer string
err error
calls *int
}
func (c *recordingCompleter) Complete(_ context.Context, system, user string) (string, error) {
*c.calls++
c.lastSystem = system
c.lastUser = user
if c.err != nil {
return "", c.err
}
return c.answer, nil
}
// factory builds a recordingCompleter per model and records the last one built so
// the test can assert which model alias was actually used for the gateway call.
type factory struct {
answer string
err error
calls int
used *recordingCompleter
}
func (f *factory) make(model string) Completer {
c := &recordingCompleter{model: model, answer: f.answer, err: f.err, calls: &f.calls}
f.used = c
return c
}
func TestModelsAreOfferedLocalFirstAndCopied(t *testing.T) {
f := &factory{answer: "ok"}
s := New(f.make, []string{"koala/phi4-mini", "iguana/gemma4-26b"}, 0)
got := s.Models()
want := []string{"koala/phi4-mini", "iguana/gemma4-26b"}
if len(got) != len(want) || got[0] != want[0] || got[1] != want[1] {
t.Fatalf("Models() = %v, want %v", got, want)
}
// Mutating the returned slice must not corrupt the Service's list.
got[0] = "tampered"
if s.Models()[0] != "koala/phi4-mini" {
t.Fatal("Models() leaked its backing slice")
}
}
func TestDefaultModelIsTheSummarysModelWhenOffered(t *testing.T) {
f := &factory{answer: "ok"}
s := New(f.make, []string{"koala/phi4-mini", "iguana/gemma4-26b", "berget/mistral-small"}, 0)
if got := s.DefaultModel("iguana/gemma4-26b"); got != "iguana/gemma4-26b" {
t.Fatalf("DefaultModel(summary) = %q, want the summary's model", got)
}
// A summary model no longer offered (e.g. cloud disabled) falls back to first.
if got := s.DefaultModel("berget/old-model"); got != "koala/phi4-mini" {
t.Fatalf("DefaultModel(un-offered) = %q, want the first offered model", got)
}
// No summary model recorded → first offered.
if got := s.DefaultModel(""); got != "koala/phi4-mini" {
t.Fatalf("DefaultModel(\"\") = %q, want the first offered model", got)
}
}
func TestAnswerGroundsOnTranscriptAndCarriesHistory(t *testing.T) {
f := &factory{answer: " The video is about attention. "}
s := New(f.make, []string{"koala/phi4-mini"}, 0)
reply, err := s.Answer(context.Background(), Request{
Model: "koala/phi4-mini",
Transcript: "ATTENTION-TRANSCRIPT-MARKER",
History: []Turn{{Question: "who", Answer: "the host"}},
Question: "what is it about",
})
if err != nil {
t.Fatalf("Answer: %v", err)
}
if reply.Answer != "The video is about attention." {
t.Fatalf("answer not trimmed: %q", reply.Answer)
}
if reply.Truncated {
t.Fatal("short transcript must not report truncated")
}
// The transcript rides in the system prompt; the conversation in the user msg.
if !strings.Contains(f.used.lastSystem, "ATTENTION-TRANSCRIPT-MARKER") {
t.Fatal("transcript not grounded into the system prompt")
}
if !strings.Contains(f.used.lastUser, "Q: who") || !strings.Contains(f.used.lastUser, "A: the host") {
t.Fatalf("history not carried into the user message: %q", f.used.lastUser)
}
if !strings.Contains(f.used.lastUser, "what is it about") {
t.Fatal("new question missing from the user message")
}
}
func TestAnswerTruncatesLongTranscriptAndReportsIt(t *testing.T) {
f := &factory{answer: "answer"}
s := New(f.make, []string{"koala/phi4-mini"}, 10)
reply, err := s.Answer(context.Background(), Request{
Model: "koala/phi4-mini",
Transcript: strings.Repeat("x", 500),
Question: "summarize",
})
if err != nil {
t.Fatalf("Answer: %v", err)
}
if !reply.Truncated {
t.Fatal("a transcript past maxChars must report Truncated")
}
if strings.Count(f.used.lastSystem, "x") != 10 {
t.Fatalf("transcript not bounded to maxChars: got %d x's", strings.Count(f.used.lastSystem, "x"))
}
}
func TestAnswerForcesAnUnofferedModelBackToDefault(t *testing.T) {
f := &factory{answer: "answer"}
s := New(f.make, []string{"koala/phi4-mini", "iguana/gemma4-26b"}, 0)
// A forged/un-offered model must never reach the gateway as-is — it is forced
// to the default offered model (the cloud-absent guarantee depends on this).
_, err := s.Answer(context.Background(), Request{
Model: "berget/secret-cloud-model",
Transcript: "t",
Question: "q",
})
if err != nil {
t.Fatalf("Answer: %v", err)
}
if f.used.model != "koala/phi4-mini" {
t.Fatalf("un-offered model reached the gateway as %q, want the default", f.used.model)
}
}
func TestAnswerPropagatesCompleterError(t *testing.T) {
f := &factory{err: errors.New("gateway down")}
s := New(f.make, []string{"koala/phi4-mini"}, 0)
_, err := s.Answer(context.Background(), Request{Model: "koala/phi4-mini", Transcript: "t", Question: "q"})
if err == nil {
t.Fatal("expected the gateway error to propagate")
}
}
func TestAnswerRejectsEmptyModelSet(t *testing.T) {
s := New(func(string) Completer { return nil }, nil, 0)
if _, err := s.Answer(context.Background(), Request{Question: "q"}); err == nil {
t.Fatal("expected an error when no models are configured")
}
}
+38 -2
View File
@@ -32,17 +32,46 @@ type Client struct {
model string
maxTokens int
httpClient *http.Client
usageHook func(model string, prompt, completion int)
}
// Option configures a Client at construction. Variadic so the existing 4-arg
// call sites stay valid as new knobs are added.
type Option func(*Client)
// WithMaxTokens overrides the per-request completion budget. The summarizer uses
// this to cap completion for small-context models (e.g. koala/phi4-mini, 8k):
// with the default 8192 budget, prompt + max_tokens overflows an 8k context and
// the gateway returns HTTP 400. A non-positive n is ignored (keeps the default).
func WithMaxTokens(n int) Option {
return func(c *Client) {
if n > 0 {
c.maxTokens = n
}
}
}
// WithUsageHook registers a callback fired after a successful completion with the
// model and the prompt/completion token counts from the response usage block. It
// keeps this copied, stdlib-only package (ADR-004) decoupled from metrics: the
// caller wires it to internal/metrics, the client imports nothing. nil is ignored.
func WithUsageHook(fn func(model string, prompt, completion int)) Option {
return func(c *Client) { c.usageHook = fn }
}
// New constructs a Client.
func New(baseURL, apiKey, model string, timeout time.Duration) *Client {
return &Client{
func New(baseURL, apiKey, model string, timeout time.Duration, opts ...Option) *Client {
c := &Client{
baseURL: strings.TrimRight(baseURL, "/"),
apiKey: apiKey,
model: model,
maxTokens: defaultMaxTokens,
httpClient: &http.Client{Timeout: timeout},
}
for _, opt := range opts {
opt(c)
}
return c
}
type chatRequest struct {
@@ -61,6 +90,10 @@ type chatResponse struct {
Choices []struct {
Message message `json:"message"`
} `json:"choices"`
Usage struct {
PromptTokens int `json:"prompt_tokens"`
CompletionTokens int `json:"completion_tokens"`
} `json:"usage"`
}
// Complete sends a system + user message and returns the assistant's reply.
@@ -132,5 +165,8 @@ func (c *Client) Complete(ctx context.Context, system, user string) (string, err
if len(cr.Choices) == 0 {
return "", fmt.Errorf("LLM returned no choices")
}
if c.usageHook != nil {
c.usageHook(c.model, cr.Usage.PromptTokens, cr.Usage.CompletionTokens)
}
return cr.Choices[0].Message.Content, nil
}
+45
View File
@@ -64,6 +64,51 @@ func TestClient_SendsMaxTokens(t *testing.T) {
}
}
// TestClient_WithMaxTokens overrides the completion budget — the summarizer caps
// it small so prompt + max_tokens fits a small-context model's window (8k).
func TestClient_WithMaxTokens(t *testing.T) {
var body chatRequest
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
_ = json.NewDecoder(r.Body).Decode(&body)
_ = json.NewEncoder(w).Encode(map[string]any{
"choices": []map[string]any{{"message": map[string]any{"content": "ok"}}},
})
}))
defer srv.Close()
c := New(srv.URL, "", "test-model", 10*time.Second, WithMaxTokens(1500))
if _, err := c.Complete(context.Background(), "sys", "user"); err != nil {
t.Fatalf("Complete: %v", err)
}
if body.MaxTokens != 1500 {
t.Errorf("max_tokens = %d, want 1500", body.MaxTokens)
}
}
// TestClient_UsageHookRecordsTokens: the usage hook fires with the model and the
// prompt/completion token counts parsed from the response usage block.
func TestClient_UsageHookRecordsTokens(t *testing.T) {
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, _ *http.Request) {
_ = json.NewEncoder(w).Encode(map[string]any{
"choices": []map[string]any{{"message": map[string]any{"content": "ok"}}},
"usage": map[string]any{"prompt_tokens": 123, "completion_tokens": 45},
})
}))
defer srv.Close()
var gotModel string
var gotPrompt, gotCompletion int
c := New(srv.URL, "", "test-model", 10*time.Second, WithUsageHook(func(model string, p, comp int) {
gotModel, gotPrompt, gotCompletion = model, p, comp
}))
if _, err := c.Complete(context.Background(), "sys", "user"); err != nil {
t.Fatalf("Complete: %v", err)
}
if gotModel != "test-model" || gotPrompt != 123 || gotCompletion != 45 {
t.Errorf("usage hook got (%q, %d, %d), want (test-model, 123, 45)", gotModel, gotPrompt, gotCompletion)
}
}
func TestClient_ReturnsErrorOnNon200(t *testing.T) {
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
http.Error(w, "overloaded", http.StatusServiceUnavailable)
+1 -1
View File
@@ -18,7 +18,7 @@ import (
"path/filepath"
"sync"
"gitea.d-ma.be/mathias/tapir/internal/ports"
"git.d-ma.be/mathias/tapir/internal/ports"
)
// FileStore is a SecretStore backed by a single 0600 JSON file mapping opaque
+1 -1
View File
@@ -7,7 +7,7 @@ import (
"path/filepath"
"testing"
"gitea.d-ma.be/mathias/tapir/internal/adapters/secrets"
"git.d-ma.be/mathias/tapir/internal/adapters/secrets"
)
func TestPutThenGet(t *testing.T) {
+7 -4
View File
@@ -10,10 +10,13 @@ import (
// DeleteUser permanently removes a user and all of their data. It runs through
// withUser so RLS confines every statement to the calling user's own rows.
//
// Deleting the users row cascades (ON DELETE CASCADE) to videos, transcripts,
// summaries (→ sink_deliveries), video_connections, and the user_identities map
// referential-integrity cascades bypass RLS, so a user's child rows are removed
// even though the deleting connection is scoped. summary_actions and login_events
// Deleting the users row cascades (ON DELETE CASCADE) to videos, summaries
// (→ sink_deliveries), video_connections, and the user_identities map
// referential-integrity cascades bypass RLS, so a user's child rows are removed
// even though the deleting connection is scoped. Transcripts are NOT removed:
// since ADR-021 they are shared public content keyed by (provider,
// provider_video_id) with no user_id, so another user may still reference the
// same row — a user deletion must not strip shared caption content. summary_actions and login_events
// are the exceptions: each carries a user_id but has NO foreign key to users
// (migrations 002 and 010), so the cascade does not reach them; they are deleted
// explicitly in the same scoped transaction. Deleting an absent user is a no-op
@@ -0,0 +1,83 @@
package store
import (
"context"
"fmt"
"time"
"github.com/jackc/pgx/v5"
)
// CaptionlessChannels returns the set of channel ids currently suppressed for the
// user — channels whose recent videos all yielded no captions, within their
// suppression window (ADR-024). The runner skips caption fetches for these
// channels' videos. A channel whose window has expired is not returned, so its
// next video is re-probed (auto-recovery).
func (s *Store) CaptionlessChannels(ctx context.Context, userID string) (map[string]bool, error) {
out := map[string]bool{}
err := s.withUser(ctx, userID, func(tx pgx.Tx) error {
rows, err := tx.Query(ctx, `
SELECT channel_id FROM channel_caption_state
WHERE user_id = $1 AND captionless_until IS NOT NULL AND captionless_until > now()`,
userID)
if err != nil {
return fmt.Errorf("store: caption-less channels: %w", err)
}
defer rows.Close()
for rows.Next() {
var ch string
if err := rows.Scan(&ch); err != nil {
return fmt.Errorf("store: scan caption-less channel: %w", err)
}
out[ch] = true
}
return rows.Err()
})
return out, err
}
// RecordChannelCaptionOutcome updates a channel's caption-availability memory
// after a fetch attempt (ADR-024). hadCaptions resets the channel (consecutive
// count to 0, suppression cleared). Otherwise the consecutive no-caption count is
// incremented; once it reaches threshold the channel is suppressed for window.
// threshold <= 0 is a no-op (feature disabled). An empty channelID is ignored
// (some sources may not carry one).
func (s *Store) RecordChannelCaptionOutcome(ctx context.Context, userID, channelID string, hadCaptions bool, threshold int, window time.Duration) error {
if channelID == "" || threshold <= 0 {
return nil
}
return s.withUser(ctx, userID, func(tx pgx.Tx) error {
if hadCaptions {
_, err := tx.Exec(ctx, `
INSERT INTO channel_caption_state (user_id, channel_id, consecutive_none, captionless_until, updated_at)
VALUES ($1, $2, 0, NULL, now())
ON CONFLICT (user_id, channel_id)
DO UPDATE SET consecutive_none = 0, captionless_until = NULL, updated_at = now()`,
userID, channelID)
if err != nil {
return fmt.Errorf("store: reset channel caption state: %w", err)
}
return nil
}
// No captions: increment the streak; suppress once it reaches threshold.
// captionless_until is set from the NEW count inside the same statement so
// the decision is atomic with the increment.
until := time.Now().Add(window)
_, err := tx.Exec(ctx, `
INSERT INTO channel_caption_state (user_id, channel_id, consecutive_none, captionless_until, updated_at)
VALUES ($1, $2, 1, CASE WHEN 1 >= $3 THEN $4::timestamptz ELSE NULL END, now())
ON CONFLICT (user_id, channel_id)
DO UPDATE SET
consecutive_none = channel_caption_state.consecutive_none + 1,
captionless_until = CASE
WHEN channel_caption_state.consecutive_none + 1 >= $3 THEN $4::timestamptz
ELSE channel_caption_state.captionless_until
END,
updated_at = now()`,
userID, channelID, threshold, until)
if err != nil {
return fmt.Errorf("store: record channel no-caption: %w", err)
}
return nil
})
}
@@ -0,0 +1,67 @@
package store_test
import (
"context"
"testing"
"time"
"github.com/stretchr/testify/require"
)
func TestChannelCaptionMemory_SuppressesAfterThreshold(t *testing.T) {
ctx := context.Background()
s := newStore(t)
super := rawPool(t)
resetDB(t, super)
seedUser(t, super, userA)
const threshold = 3
window := time.Hour
// Below threshold: not yet suppressed.
require.NoError(t, s.RecordChannelCaptionOutcome(ctx, userA, "chanX", false, threshold, window))
require.NoError(t, s.RecordChannelCaptionOutcome(ctx, userA, "chanX", false, threshold, window))
got, err := s.CaptionlessChannels(ctx, userA)
require.NoError(t, err)
require.NotContains(t, got, "chanX", "2 < threshold 3: not suppressed yet")
// Crossing the threshold suppresses the channel.
require.NoError(t, s.RecordChannelCaptionOutcome(ctx, userA, "chanX", false, threshold, window))
got, err = s.CaptionlessChannels(ctx, userA)
require.NoError(t, err)
require.Contains(t, got, "chanX", "3 consecutive no-caption results suppress the channel")
// A successful caption fetch resets it.
require.NoError(t, s.RecordChannelCaptionOutcome(ctx, userA, "chanX", true, threshold, window))
got, err = s.CaptionlessChannels(ctx, userA)
require.NoError(t, err)
require.NotContains(t, got, "chanX", "a captioned video clears suppression")
}
func TestChannelCaptionMemory_WindowExpiryReProbes(t *testing.T) {
ctx := context.Background()
s := newStore(t)
super := rawPool(t)
resetDB(t, super)
seedUser(t, super, userA)
// A negative window means captionless_until lands in the past — modelling an
// elapsed suppression window, which must make the channel eligible again.
require.NoError(t, s.RecordChannelCaptionOutcome(ctx, userA, "chanY", false, 1, -time.Hour))
got, err := s.CaptionlessChannels(ctx, userA)
require.NoError(t, err)
require.NotContains(t, got, "chanY", "an expired window re-enables the channel for a re-probe")
}
func TestChannelCaptionMemory_DisabledThresholdIsNoOp(t *testing.T) {
ctx := context.Background()
s := newStore(t)
super := rawPool(t)
resetDB(t, super)
seedUser(t, super, userA)
require.NoError(t, s.RecordChannelCaptionOutcome(ctx, userA, "chanZ", false, 0, time.Hour))
got, err := s.CaptionlessChannels(ctx, userA)
require.NoError(t, err)
require.Empty(t, got, "threshold 0 disables the memory — nothing recorded")
}
+1 -1
View File
@@ -6,7 +6,7 @@ import (
"github.com/stretchr/testify/require"
"gitea.d-ma.be/mathias/tapir/internal/adapters/store"
"git.d-ma.be/mathias/tapir/internal/adapters/store"
)
// seedUserRow inserts a bare users row (FK target for a connection) as the
+1 -1
View File
@@ -7,7 +7,7 @@ import (
"github.com/stretchr/testify/require"
"gitea.d-ma.be/mathias/tapir/internal/adapters/store"
"git.d-ma.be/mathias/tapir/internal/adapters/store"
)
const (
+70 -17
View File
@@ -3,6 +3,7 @@ package store_test
import (
"context"
"database/sql"
"errors"
"os"
"testing"
@@ -34,6 +35,30 @@ func fileMigrator(t *testing.T) *migrate.Migrate {
return m
}
// headVersion reports the current (HEAD) schema version so a test can restore
// to it after stepping down, without hard-coding what HEAD is. Adding a
// migration on top changes HEAD but no test that uses this needs editing.
func headVersion(t *testing.T, m *migrate.Migrate) uint {
t.Helper()
v, dirty, err := m.Version()
require.NoError(t, err)
require.False(t, dirty, "schema must not be dirty")
return v
}
// migrateTo drives the schema to an exact version *by version number*, not by
// step count. This is the whole point of the migrate-test design: a migration
// added above the target does not shift any count here, so unrelated tests stay
// green (see issue #8). ErrNoChange (already at that version) is not a failure.
func migrateTo(t *testing.T, m *migrate.Migrate, version uint) {
t.Helper()
err := m.Migrate(version)
if errors.Is(err, migrate.ErrNoChange) {
return
}
require.NoError(t, err)
}
// loginEventsExists reports whether the login_events relation is present.
func loginEventsExists(t *testing.T) bool {
t.Helper()
@@ -53,19 +78,15 @@ func TestMigration010LoginEventsUpDown(t *testing.T) {
require.True(t, loginEventsExists(t), "login_events must exist at latest migration")
m := fileMigrator(t)
// 011, 012, 013 sit above 010; step them down first so 010 is exercised in isolation.
require.NoError(t, m.Steps(-1), "down 013 drops channel_errors, login_events intact")
require.True(t, loginEventsExists(t), "013 down leaves login_events intact")
require.NoError(t, m.Steps(-1), "down 012 is a no-op, login_events intact")
require.True(t, loginEventsExists(t), "012 down leaves login_events intact")
require.NoError(t, m.Steps(-1), "down 011 must not touch login_events")
require.True(t, loginEventsExists(t), "011 down leaves login_events intact")
head := headVersion(t, m)
require.NoError(t, m.Steps(-1), "down 010 must drop login_events")
migrateTo(t, m, 9) // just below 010 — everything above steps down
require.False(t, loginEventsExists(t), "login_events must be gone after the down migration")
require.NoError(t, m.Steps(4), "up must recreate 010 then re-apply 011, 012, 013")
migrateTo(t, m, 10) // up 010
require.True(t, loginEventsExists(t), "login_events must be restored after the up migration")
migrateTo(t, m, head) // restore to HEAD for sibling tests
}
// autoSummarizeDefault reads the users.auto_summarize column default as text
@@ -87,15 +108,43 @@ func TestMigration011AutoSummarizeDefaultUpDown(t *testing.T) {
require.Equal(t, "true", autoSummarizeDefault(t), "011 sets the default to TRUE")
m := fileMigrator(t)
require.NoError(t, m.Steps(-1), "down 013 drops channel_errors")
require.NoError(t, m.Steps(-1), "down 012 is a no-op")
require.NoError(t, m.Steps(-1), "down 011 reverts the column default")
head := headVersion(t, m)
migrateTo(t, m, 10) // just below 011 reverts the column default
require.Equal(t, "false", autoSummarizeDefault(t), "default is FALSE after the down migration")
require.NoError(t, m.Steps(1), "up 011 re-applies the TRUE default")
migrateTo(t, m, 11) // up 011 re-applies the TRUE default
require.Equal(t, "true", autoSummarizeDefault(t))
require.NoError(t, m.Steps(1), "up 012 runs clean (no FORCE RLS on fresh schema)")
require.NoError(t, m.Steps(1), "up 013 creates channel_errors")
migrateTo(t, m, head) // restore to HEAD for sibling tests
}
// channelTitleExists reports whether videos.channel_title is present.
func channelTitleExists(t *testing.T) bool {
t.Helper()
var exists bool
require.NoError(t, rawPool(t).QueryRow(context.Background(),
`SELECT EXISTS (SELECT 1 FROM information_schema.columns
WHERE table_name = 'videos' AND column_name = 'channel_title')`).Scan(&exists))
return exists
}
// TestMigration014VideoChannelTitleUpDown proves 014 is reversible: down drops
// videos.channel_title, up recreates it.
func TestMigration014VideoChannelTitleUpDown(t *testing.T) {
newStore(t) // latest (014 applied)
require.True(t, channelTitleExists(t), "channel_title exists at latest migration")
m := fileMigrator(t)
head := headVersion(t, m)
migrateTo(t, m, 13) // just below 014 — drops channel_title
require.False(t, channelTitleExists(t), "channel_title must be gone after the down migration")
migrateTo(t, m, 14) // up 014 recreates channel_title
require.True(t, channelTitleExists(t), "channel_title must be restored after the up migration")
migrateTo(t, m, head) // restore to HEAD for sibling tests
}
// TestMigration012FixAutoSummarizeRLS proves 012 runs cleanly and flips any
@@ -106,7 +155,11 @@ func TestMigration012FixAutoSummarizeRLS(t *testing.T) {
// Round-trip: down 012, then up 012 — must be idempotent.
m := fileMigrator(t)
require.NoError(t, m.Steps(-1), "down 012 must not error")
require.NoError(t, m.Steps(1), "up 012 must re-apply cleanly")
head := headVersion(t, m)
migrateTo(t, m, 11) // down 012 must not error
migrateTo(t, m, 12) // up 012 must re-apply cleanly
require.Equal(t, "true", autoSummarizeDefault(t), "default still TRUE after 012 re-applied")
migrateTo(t, m, head) // restore to HEAD for sibling tests
}
@@ -0,0 +1 @@
ALTER TABLE videos DROP COLUMN channel_title;
@@ -0,0 +1,8 @@
-- Store the source channel's title per video so the list can offer a real
-- channel filter (multi-select of the user's channels) instead of the dead
-- free-text field that only ever matched the provider string. Nullable: existing
-- rows backfill on the next discovery pass (UpsertVideo writes it); pasted videos
-- get it immediately from videos.list. No FK to a channels table at Stage 0 — the
-- title is a denormalised display/filter value, consistent with the existing
-- subscription_id-stays-NULL stance (data-model.md).
ALTER TABLE videos ADD COLUMN channel_title TEXT;
@@ -0,0 +1,19 @@
-- Down 015: restore the per-user RLS-scoped transcripts shape (001 + 003).
DROP TABLE transcripts;
CREATE TABLE transcripts (
video_id UUID PRIMARY KEY REFERENCES videos(id) ON DELETE CASCADE,
user_id UUID NOT NULL REFERENCES users(id) ON DELETE CASCADE,
source TEXT NOT NULL,
language TEXT,
content TEXT,
resolved_at TIMESTAMPTZ NOT NULL DEFAULT NOW()
);
CREATE INDEX idx_transcripts_user_id ON transcripts(user_id);
ALTER TABLE transcripts ENABLE ROW LEVEL SECURITY;
ALTER TABLE transcripts FORCE ROW LEVEL SECURITY;
CREATE POLICY transcripts_isolation ON transcripts
FOR ALL
USING (user_id = current_setting('tapir.current_user_id', true)::uuid);
@@ -0,0 +1,30 @@
-- Migration 015: transcripts become SHARED public-content storage (ADR-021).
--
-- The per-user transcripts table from 001 (PK videos.id, user_id NOT NULL, RLS
-- FORCEd in 003) was dead: no application code ever read or wrote it — only the
-- transcript_status columns on `videos` (007) carried fetch outcomes. ADR-021
-- repurposes it as the single shared store of public caption content, keyed by
-- the cross-user dedup key (provider, provider_video_id) — the video's public
-- identity, not Tapir's per-user videos.id — so re-analysis never re-fetches
-- from YouTube (ADR-010/014).
--
-- It holds ONLY public caption content + the video's public id (nothing
-- user-identifying), so it is deliberately NOT RLS-scoped: no user_id, no
-- policy, no FORCE. This is the single, intentional exception to the ADR-012
-- isolation boundary; rls_test.go asserts the boundary is exactly here and
-- nowhere else. Dropping the old table drops its RLS policy with it; it held no
-- real data, so drop+recreate loses nothing.
DROP TABLE transcripts;
CREATE TABLE transcripts (
provider TEXT NOT NULL,
provider_video_id TEXT NOT NULL,
source TEXT NOT NULL, -- 'captions' (content set) | 'none' (no captions; content NULL)
language TEXT,
content TEXT,
fetched_at TIMESTAMPTZ NOT NULL DEFAULT NOW(),
PRIMARY KEY (provider, provider_video_id)
);
COMMENT ON TABLE transcripts IS
'Shared public caption content keyed by (provider, provider_video_id). NOT RLS-scoped — public content only, de-facto cross-user dedup (ADR-021).';
@@ -0,0 +1 @@
DROP TABLE channel_caption_state;
@@ -0,0 +1,30 @@
-- Migration 016: per-(user, channel) caption-availability memory (ADR-024).
--
-- Some channels never publish English captions (foreign-language news, music,
-- etc.). Each of their new videos still costs ONE rate-limited caption fetch
-- (ADR-014) before resolving to "none" — and on a throttled egress IP that fetch
-- may 429 and churn through the backoff machinery first. This table remembers
-- channels that repeatedly yield no captions so discovery can stop attempting
-- their videos, freeing the scarce fetch budget for channels that do have them.
--
-- consecutive_none counts no-caption outcomes in a row; a successful fetch resets
-- it to 0. Once it crosses the threshold the channel is suppressed until
-- captionless_until, after which one video is re-probed (auto-recovery for a
-- channel that starts adding captions). Per-user + RLS-scoped, consistent with
-- the rest of the user-owned schema (subscriptions are per-user; ADR-012).
CREATE TABLE channel_caption_state (
user_id UUID NOT NULL REFERENCES users(id) ON DELETE CASCADE,
channel_id TEXT NOT NULL,
consecutive_none INT NOT NULL DEFAULT 0,
captionless_until TIMESTAMPTZ,
updated_at TIMESTAMPTZ NOT NULL DEFAULT now(),
PRIMARY KEY (user_id, channel_id)
);
CREATE INDEX idx_channel_caption_state_user_id ON channel_caption_state(user_id);
ALTER TABLE channel_caption_state ENABLE ROW LEVEL SECURITY;
ALTER TABLE channel_caption_state FORCE ROW LEVEL SECURITY;
CREATE POLICY channel_caption_state_isolation ON channel_caption_state
FOR ALL
USING (user_id = current_setting('tapir.current_user_id', true)::uuid);
+4 -1
View File
@@ -29,6 +29,7 @@ type SummaryRow struct {
ProviderVideoID string // videos.provider_video_id; empty when no videos row
Title string // videos.title; empty when no videos row
Channel string // videos.provider for now; empty when no videos row
ChannelTitle string // videos.channel_title; the source channel, for display + filtering
URL string // videos.url; empty when no videos row
PublishedAt time.Time // videos.published_at; zero when absent
Summary string
@@ -137,7 +138,8 @@ const selectVideo = `
COALESCE(s.created_at, v.seen_at),
(s.id IS NOT NULL) AS summarized,
v.summarize_requested,
COALESCE(v.transcript_status, '')
COALESCE(v.transcript_status, ''),
COALESCE(v.channel_title, '')
FROM videos v
LEFT JOIN summaries s ON s.video_id = v.id AND s.user_id = v.user_id`
@@ -254,6 +256,7 @@ func scanVideoRow(rows pgx.Row) (SummaryRow, error) {
&row.Summarized,
&row.SummarizeRequested,
&row.TranscriptStatus,
&row.ChannelTitle,
); err != nil {
return SummaryRow{}, fmt.Errorf("store: scan video: %w", err)
}
+1 -1
View File
@@ -8,7 +8,7 @@ import (
"github.com/jackc/pgx/v5/pgxpool"
"github.com/stretchr/testify/require"
"gitea.d-ma.be/mathias/tapir/internal/adapters/store"
"git.d-ma.be/mathias/tapir/internal/adapters/store"
)
// seedVideo inserts a videos row whose id matches a summary's video_id, so the
+12 -6
View File
@@ -4,6 +4,7 @@ import (
"context"
"fmt"
"sort"
"time"
"github.com/jackc/pgx/v5"
)
@@ -32,7 +33,12 @@ type UserActiveWeeks struct {
// Scope note: the enumeration covers users with a Dex identity (the web users the
// gate is about). A CLI-only user created by the store sink without an identity
// row would not appear — out of scope for this gate.
func (s *Store) ActiveWeeks(ctx context.Context) ([]UserActiveWeeks, error) {
// ActiveWeeks counts each user's distinct active weeks from `since` onward. A zero
// `since` means no lower bound (count all history). The Stage-0 gate baseline is
// set by the caller (the report command) to the date real usage tracking began,
// so pre-launch noise — testing, the period the pilot was blocked — does not count
// toward the return-usage signal (ADR-016).
func (s *Store) ActiveWeeks(ctx context.Context, since time.Time) ([]UserActiveWeeks, error) {
userIDs, err := s.identityUserIDs(ctx)
if err != nil {
return nil, err
@@ -40,7 +46,7 @@ func (s *Store) ActiveWeeks(ctx context.Context) ([]UserActiveWeeks, error) {
out := make([]UserActiveWeeks, 0, len(userIDs))
for _, uid := range userIDs {
row, err := s.activeWeeksFor(ctx, uid)
row, err := s.activeWeeksFor(ctx, uid, since)
if err != nil {
return nil, err
}
@@ -85,18 +91,18 @@ func (s *Store) identityUserIDs(ctx context.Context) ([]string, error) {
// activeWeeksFor counts one user's distinct active weeks (reads UNION acts) and
// reads their display name, RLS-scoped via withUser. The UNION dedups a week that
// has both a login and an action so it counts once.
func (s *Store) activeWeeksFor(ctx context.Context, userID string) (UserActiveWeeks, error) {
func (s *Store) activeWeeksFor(ctx context.Context, userID string, since time.Time) (UserActiveWeeks, error) {
res := UserActiveWeeks{UserID: userID}
if err := s.withUser(ctx, userID, func(tx pgx.Tx) error {
if err := tx.QueryRow(ctx,
`WITH weeks AS (
SELECT date_trunc('week', seen_at) AS wk
FROM login_events WHERE user_id = $1
FROM login_events WHERE user_id = $1 AND seen_at >= $2
UNION
SELECT date_trunc('week', acted_at)
FROM summary_actions WHERE user_id = $1
FROM summary_actions WHERE user_id = $1 AND acted_at >= $2
)
SELECT count(DISTINCT wk) FROM weeks`, userID).Scan(&res.ActiveWeeks); err != nil {
SELECT count(DISTINCT wk) FROM weeks`, userID, since).Scan(&res.ActiveWeeks); err != nil {
return fmt.Errorf("store: count active weeks: %w", err)
}
if err := tx.QueryRow(ctx,
+29 -2
View File
@@ -3,6 +3,7 @@ package store_test
import (
"context"
"testing"
"time"
"github.com/jackc/pgx/v5/pgxpool"
"github.com/stretchr/testify/require"
@@ -54,7 +55,7 @@ func TestActiveWeeksCountsDistinctWeeksAcrossReadsAndActs(t *testing.T) {
($1, 'vid-2', 'saved', '2026-01-19T18:00:00Z')`, userA)
require.NoError(t, err)
got, err := s.ActiveWeeks(ctx)
got, err := s.ActiveWeeks(ctx, time.Time{}) // zero since = no lower bound
require.NoError(t, err)
require.Len(t, got, 2, "both identity users must appear")
@@ -72,7 +73,33 @@ func TestActiveWeeksEmptyWhenNoUsers(t *testing.T) {
s := newStore(t)
resetDB(t, rawPool(t))
got, err := s.ActiveWeeks(ctx)
got, err := s.ActiveWeeks(ctx, time.Time{})
require.NoError(t, err)
require.Empty(t, got)
}
// TestActiveWeeksExcludesBeforeGateStart proves the baseline cutoff: activity
// before `since` does not count, so pre-launch noise (testing, the pilot's blocked
// period) is excluded from the Stage-0 return-usage gate (ADR-016).
func TestActiveWeeksExcludesBeforeGateStart(t *testing.T) {
ctx := context.Background()
s := newStore(t)
p := rawPool(t)
resetDB(t, p)
seedReportUser(t, p, userA, "subject-a", "Ada")
// One read well before the baseline, two reads in distinct weeks after it.
_, err := p.Exec(ctx,
`INSERT INTO login_events (user_id, seen_at) VALUES
($1, '2026-05-01T09:00:00Z'),
($1, '2026-06-12T09:00:00Z'),
($1, '2026-06-19T09:00:00Z')`, userA)
require.NoError(t, err)
since := time.Date(2026, 6, 11, 0, 0, 0, 0, time.UTC)
got, err := s.ActiveWeeks(ctx, since)
require.NoError(t, err)
require.Len(t, got, 1)
require.Equal(t, 2, got[0].ActiveWeeks, "only the two post-baseline weeks count; the May read is excluded")
}
+70 -14
View File
@@ -22,9 +22,11 @@ import (
// (no GUC set → zero rows) proves the enforcement path is live, not bypassed.
// userIsolatedTables are the tables that carry a user_id and whose policy keys
// directly off the tapir.current_user_id GUC.
// directly off the tapir.current_user_id GUC. transcripts is deliberately ABSENT
// — ADR-021 made it shared public content (non-RLS); TestTranscriptsTableIsSharedNotRLS
// proves that is the only place the isolation boundary moved.
var userIsolatedTables = []string{
"users", "videos", "transcripts", "summaries", "summary_actions", "login_events", "video_connections",
"users", "videos", "summaries", "summary_actions", "login_events", "video_connections",
}
// allIsolatedTables adds sink_deliveries, whose ownership is derived from its
@@ -38,9 +40,10 @@ type seeded struct {
summaryID string
}
// seedUser inserts one full chain (user → video → transcript → summary →
// action → delivery) as the superuser pool, which bypasses RLS so both users'
// data lands regardless of the GUC.
// seedUser inserts one full chain (user → video → summary → action → delivery)
// as the superuser pool, which bypasses RLS so both users' data lands regardless
// of the GUC. Transcripts are NOT seeded here: they are shared, non-RLS public
// content (ADR-021), so they have no place in a per-user isolation chain.
func seedUser(t *testing.T, p *pgxpool.Pool, userID string) seeded {
t.Helper()
ctx := context.Background()
@@ -54,11 +57,6 @@ func seedUser(t *testing.T, p *pgxpool.Pool, userID string) seeded {
VALUES ($1, 'youtube', $2, 'title') RETURNING id`,
userID, "vid-"+userID).Scan(&videoID))
_, err = p.Exec(ctx,
`INSERT INTO transcripts (video_id, user_id, source, content)
VALUES ($1, $2, 'captions', 'words')`, videoID, userID)
require.NoError(t, err)
var summaryID string
require.NoError(t, p.QueryRow(ctx,
`INSERT INTO summaries (user_id, video_id, summary) VALUES ($1, $2, 'sum')
@@ -92,9 +90,16 @@ func appPool(t *testing.T, super *pgxpool.Pool) *pgxpool.Pool {
t.Helper()
ctx := context.Background()
// Idempotent across test runs (schema/role persist for the TestMain PG).
_, _ = super.Exec(ctx, `DROP ROLE IF EXISTS app`)
_, err := super.Exec(ctx, `CREATE ROLE app LOGIN PASSWORD 'app'`)
// Idempotent across tests AND runs: the role persists for the TestMain PG and
// owns granted privileges, so a plain DROP ROLE fails once any GRANT exists
// (and more than one test now builds an app pool). Create only if absent; the
// GRANTs below are themselves idempotent.
_, err := super.Exec(ctx,
`DO $$ BEGIN
IF NOT EXISTS (SELECT FROM pg_roles WHERE rolname = 'app') THEN
CREATE ROLE app LOGIN PASSWORD 'app';
END IF;
END $$`)
require.NoError(t, err)
_, err = super.Exec(ctx, `GRANT USAGE ON SCHEMA public TO app`)
require.NoError(t, err)
@@ -186,7 +191,6 @@ func TestRLSEnforcesPerUserIsolation(t *testing.T) {
{"update users", `UPDATE users SET display_name = 'hacked' WHERE id = $1`, b.userID},
{"update videos", `UPDATE videos SET title = 'hacked' WHERE user_id = $1`, b.userID},
{"queue videos summarize", `UPDATE videos SET summarize_requested = TRUE WHERE id = $1`, b.videoID},
{"update transcripts", `UPDATE transcripts SET content = 'hacked' WHERE user_id = $1`, b.userID},
{"update summaries", `UPDATE summaries SET summary = 'hacked' WHERE user_id = $1`, b.userID},
{"update summary_actions", `UPDATE summary_actions SET action = 'skipped' WHERE user_id = $1`, b.userID},
{"update login_events", `UPDATE login_events SET seen_at = NOW() WHERE user_id = $1`, b.userID},
@@ -235,3 +239,55 @@ func TestRLSEnforcesPerUserIsolation(t *testing.T) {
_ = a // a's ids are seeded for the symmetric read assertions above
}
// TestTranscriptsTableIsSharedNotRLS is the ADR-021 isolation proof: transcripts
// is the ONE shared, non-RLS surface, and the public-content classification
// leaked to nothing else. It is the inverse of TestRLSEnforcesPerUserIsolation —
// where that asserts deny-all on every user-owned table, this asserts transcripts
// is readable and writable with no user scope at all, holds no user_id, and is
// the single table with row-level security switched off.
func TestTranscriptsTableIsSharedNotRLS(t *testing.T) {
newStore(t)
super := rawPool(t)
resetDB(t, super)
app := appPool(t, super)
ctx := context.Background()
// 1. Shared + non-RLS: with NO GUC set, the app role both writes and reads a
// transcript. On an RLS table this would be deny-all (zero rows), exactly as
// the main isolation test asserts for every user-owned table.
_, err := app.Exec(ctx,
`INSERT INTO transcripts (provider, provider_video_id, source, content)
VALUES ('youtube', 'shared-vid', 'captions', 'public words')`)
require.NoError(t, err, "app role must write shared transcript content with no user scope")
require.Equal(t, 1, scopedCount(t, app, "", "transcripts"),
"transcripts must be readable with NO user scope — it is shared, non-RLS (ADR-021)")
// 2. No user_id column: the table holds only public caption content + the
// video's public id, nothing user-identifying.
var hasUserID bool
require.NoError(t, super.QueryRow(ctx,
`SELECT EXISTS (SELECT 1 FROM information_schema.columns
WHERE table_name = 'transcripts' AND column_name = 'user_id')`).Scan(&hasUserID))
require.False(t, hasUserID, "transcripts must carry no user_id (ADR-021 public content)")
// 3. The boundary is EXACTLY here: every user-owned table still has row-level
// security enabled; transcripts alone has it off. This is the proof the
// non-RLS classification was applied to transcripts and leaked nowhere else.
for _, table := range allIsolatedTables {
require.True(t, rlsEnabled(t, super, table),
"%s must still enforce row-level security — isolation must not have regressed", table)
}
require.False(t, rlsEnabled(t, super, "transcripts"),
"transcripts must be the single table with row-level security OFF (the one shared surface)")
}
// rlsEnabled reports whether a public table has ROW LEVEL SECURITY enabled.
func rlsEnabled(t *testing.T, p *pgxpool.Pool, table string) bool {
t.Helper()
var enabled bool
require.NoError(t, p.QueryRow(context.Background(),
`SELECT relrowsecurity FROM pg_class
WHERE relname = $1 AND relnamespace = 'public'::regnamespace`, table).Scan(&enabled))
return enabled
}
+1 -1
View File
@@ -24,7 +24,7 @@ import (
_ "github.com/jackc/pgx/v5/stdlib" // register the "pgx" database/sql driver for migrate
"gitea.d-ma.be/mathias/tapir/internal/domain"
"git.d-ma.be/mathias/tapir/internal/domain"
)
//go:embed migrations/*.sql
+18 -5
View File
@@ -4,15 +4,16 @@ import (
"context"
"fmt"
"os"
"path/filepath"
"testing"
embeddedpostgres "github.com/fergusstrange/embedded-postgres"
"github.com/jackc/pgx/v5/pgxpool"
"github.com/stretchr/testify/require"
"gitea.d-ma.be/mathias/tapir/internal/adapters/store"
"gitea.d-ma.be/mathias/tapir/internal/domain"
"gitea.d-ma.be/mathias/tapir/internal/ports"
"git.d-ma.be/mathias/tapir/internal/adapters/store"
"git.d-ma.be/mathias/tapir/internal/domain"
"git.d-ma.be/mathias/tapir/internal/ports"
)
// Static check: Store satisfies the Sink port.
@@ -24,11 +25,22 @@ var _ ports.Sink = (*store.Store)(nil)
var dsn string
func TestMain(m *testing.M) {
const port = 54329
// Port + runtime/data dirs are per-process (PID-derived) so two concurrent
// `go test` invocations — e.g. a push-run and a tag-run firing together in CI —
// don't collide on a fixed port or a shared data dir (which silently failed
// both runs). CachePath is shared so the PG archive is downloaded once, not
// per process. Base 54000 keeps this package's range distinct from web's.
port := uint32(54000 + os.Getpid()%1000)
dsn = fmt.Sprintf("postgres://postgres:postgres@localhost:%d/postgres?sslmode=disable", port)
rt := filepath.Join(os.TempDir(), fmt.Sprintf("tapir-epg-store-%d", os.Getpid()))
pg := embeddedpostgres.NewDatabase(
embeddedpostgres.DefaultConfig().Port(port),
embeddedpostgres.DefaultConfig().
Port(port).
RuntimePath(rt).
DataPath(filepath.Join(rt, "data")).
BinariesPath(filepath.Join(rt, "bin")).
CachePath(filepath.Join(os.TempDir(), "tapir-epg-cache")),
)
if err := pg.Start(); err != nil {
fmt.Fprintf(os.Stderr, "embedded-postgres start: %v\n", err)
@@ -40,6 +52,7 @@ func TestMain(m *testing.M) {
if err := pg.Stop(); err != nil {
fmt.Fprintf(os.Stderr, "embedded-postgres stop: %v\n", err)
}
_ = os.RemoveAll(rt)
os.Exit(code)
}
@@ -8,7 +8,7 @@ import (
"github.com/jackc/pgx/v5/pgxpool"
"github.com/stretchr/testify/require"
"gitea.d-ma.be/mathias/tapir/internal/adapters/store"
"git.d-ma.be/mathias/tapir/internal/adapters/store"
)
// seedBareVideo inserts a videos row with no summary, so the all-videos read and
+66
View File
@@ -0,0 +1,66 @@
package store
import (
"context"
"errors"
"fmt"
"github.com/jackc/pgx/v5"
"git.d-ma.be/mathias/tapir/internal/domain"
)
// GetTranscript returns the shared, stored transcript for a video keyed by the
// cross-user dedup key (provider, providerVideoID), and whether one exists
// (ADR-021). It reads via the raw pool, NOT withUser: the table holds public
// content with no user_id and no RLS policy, so it is shared across users by
// construction. A stored SourceNone is a real hit (ok == true, HasText() ==
// false) — a known caption-less video, so the caller skips without re-fetching.
func (s *Store) GetTranscript(ctx context.Context, provider, providerVideoID string) (domain.Transcript, bool, error) {
var source, lang, content string
err := s.pool.QueryRow(ctx,
`SELECT source, COALESCE(language, ''), COALESCE(content, '')
FROM transcripts WHERE provider = $1 AND provider_video_id = $2`,
provider, providerVideoID).Scan(&source, &lang, &content)
if errors.Is(err, pgx.ErrNoRows) {
return domain.Transcript{}, false, nil
}
if err != nil {
return domain.Transcript{}, false, fmt.Errorf("store: get transcript: %w", err)
}
return domain.Transcript{
Source: domain.TranscriptSource(source),
Language: lang,
Content: content,
}, true, nil
}
// SaveTranscript upserts the shared transcript for (provider, providerVideoID).
// Only terminal outcomes belong here: SourceCaptions (with text) or SourceNone
// (no captions). A transient SourceRateLimited is rejected so persistence never
// masks a 429 as a permanent absence — that stays a per-user retry (ADR-014).
// Last write wins on conflict (a later re-fetch may correct an entry). It writes
// via the raw pool, NOT withUser — public content, shared, non-RLS (ADR-021).
func (s *Store) SaveTranscript(ctx context.Context, provider, providerVideoID string, t domain.Transcript) error {
switch t.Source {
case domain.SourceCaptions, domain.SourceNone:
// terminal — persist
case domain.SourceRateLimited:
return fmt.Errorf("store: refusing to persist transient rate-limited transcript for %s/%s", provider, providerVideoID)
default:
return fmt.Errorf("store: invalid transcript source %q", t.Source)
}
_, err := s.pool.Exec(ctx,
`INSERT INTO transcripts (provider, provider_video_id, source, language, content)
VALUES ($1, $2, $3, NULLIF($4, ''), NULLIF($5, ''))
ON CONFLICT (provider, provider_video_id)
DO UPDATE SET source = EXCLUDED.source,
language = EXCLUDED.language,
content = EXCLUDED.content,
fetched_at = NOW()`,
provider, providerVideoID, string(t.Source), t.Language, t.Content)
if err != nil {
return fmt.Errorf("store: save transcript: %w", err)
}
return nil
}
@@ -6,7 +6,7 @@ import (
"github.com/stretchr/testify/require"
"gitea.d-ma.be/mathias/tapir/internal/adapters/store"
"git.d-ma.be/mathias/tapir/internal/adapters/store"
)
func TestSetTranscriptStatus_RoundTrip(t *testing.T) {
@@ -0,0 +1,87 @@
package store_test
import (
"context"
"testing"
"github.com/stretchr/testify/require"
"git.d-ma.be/mathias/tapir/internal/adapters/store"
"git.d-ma.be/mathias/tapir/internal/domain"
"git.d-ma.be/mathias/tapir/internal/ports"
)
// Static check: Store satisfies the shared TranscriptStore port (ADR-021).
var _ ports.TranscriptStore = (*store.Store)(nil)
func TestSaveAndGetTranscript_RoundTrip(t *testing.T) {
s := newStore(t)
resetDB(t, rawPool(t))
ctx := context.Background()
want := domain.Transcript{Source: domain.SourceCaptions, Language: "en", Content: "the words"}
require.NoError(t, s.SaveTranscript(ctx, "youtube", "vid-1", want))
got, ok, err := s.GetTranscript(ctx, "youtube", "vid-1")
require.NoError(t, err)
require.True(t, ok, "a saved transcript must be found")
require.Equal(t, domain.SourceCaptions, got.Source)
require.Equal(t, "en", got.Language)
require.Equal(t, "the words", got.Content)
require.True(t, got.HasText())
}
func TestGetTranscript_Miss(t *testing.T) {
s := newStore(t)
resetDB(t, rawPool(t))
_, ok, err := s.GetTranscript(context.Background(), "youtube", "absent")
require.NoError(t, err, "a miss is not an error")
require.False(t, ok)
}
// A stored "no captions" outcome is a real hit: callers must skip without
// re-fetching, so ok is true even though there is no text (ADR-021 / ADR-007).
func TestSaveAndGetTranscript_NoneIsAStoredHit(t *testing.T) {
s := newStore(t)
resetDB(t, rawPool(t))
ctx := context.Background()
require.NoError(t, s.SaveTranscript(ctx, "youtube", "vid-none", domain.Transcript{Source: domain.SourceNone}))
got, ok, err := s.GetTranscript(ctx, "youtube", "vid-none")
require.NoError(t, err)
require.True(t, ok, "a stored SourceNone is a hit, not a miss")
require.Equal(t, domain.SourceNone, got.Source)
require.False(t, got.HasText())
}
// A transient 429 must never be persisted as a terminal transcript, or a later
// read would mask the rate-limit as a permanent "no transcript" (ADR-014).
func TestSaveTranscript_RejectsRateLimited(t *testing.T) {
s := newStore(t)
resetDB(t, rawPool(t))
err := s.SaveTranscript(context.Background(), "youtube", "vid-429",
domain.Transcript{Source: domain.SourceRateLimited})
require.Error(t, err)
_, ok, _ := s.GetTranscript(context.Background(), "youtube", "vid-429")
require.False(t, ok, "a rejected rate-limited save must leave nothing stored")
}
func TestSaveTranscript_UpsertLastWriteWins(t *testing.T) {
s := newStore(t)
resetDB(t, rawPool(t))
ctx := context.Background()
require.NoError(t, s.SaveTranscript(ctx, "youtube", "vid-up", domain.Transcript{Source: domain.SourceNone}))
require.NoError(t, s.SaveTranscript(ctx, "youtube", "vid-up",
domain.Transcript{Source: domain.SourceCaptions, Language: "en", Content: "now resolved"}))
got, ok, err := s.GetTranscript(ctx, "youtube", "vid-up")
require.NoError(t, err)
require.True(t, ok)
require.Equal(t, domain.SourceCaptions, got.Source)
require.Equal(t, "now resolved", got.Content)
}
+67 -17
View File
@@ -7,7 +7,7 @@ import (
"github.com/jackc/pgx/v5"
"gitea.d-ma.be/mathias/tapir/internal/domain"
"git.d-ma.be/mathias/tapir/internal/domain"
)
// UpsertVideo persists a video's metadata and returns its durable store id (the
@@ -46,14 +46,16 @@ func (s *Store) UpsertVideo(ctx context.Context, v domain.Video) (string, error)
}
if err := tx.QueryRow(ctx,
`INSERT INTO videos (user_id, provider, provider_video_id, title, url, published_at)
VALUES ($1, $2, $3, $4, $5, $6)
`INSERT INTO videos (user_id, provider, provider_video_id, title, url, published_at, channel_title, duration_s)
VALUES ($1, $2, $3, $4, $5, $6, $7, $8)
ON CONFLICT (user_id, provider, provider_video_id) DO UPDATE SET
title = EXCLUDED.title,
url = EXCLUDED.url,
published_at = EXCLUDED.published_at
title = EXCLUDED.title,
url = EXCLUDED.url,
published_at = EXCLUDED.published_at,
channel_title = COALESCE(NULLIF(EXCLUDED.channel_title, ''), videos.channel_title),
duration_s = COALESCE(EXCLUDED.duration_s, videos.duration_s)
RETURNING id`,
v.UserID, provider, v.ProviderVideoID, v.Title, v.URL, nullTime(v.PublishedAt),
v.UserID, provider, v.ProviderVideoID, v.Title, v.URL, nullTime(v.PublishedAt), v.ChannelTitle, nullDuration(v.DurationSeconds),
).Scan(&id); err != nil {
return fmt.Errorf("store: upsert video: %w", err)
}
@@ -73,12 +75,27 @@ func nullTime(t time.Time) *time.Time {
return &t
}
// NewestUnsummarizedVideoIDs returns up to limit of the user's videos that have
// no summary yet, newest first (published_at DESC, NULLS LAST). It caps the
// connect-time onboarding burst (Feature 1) at a fixed count: the caller marks
// these for summarization through the shared rate gate. RLS-scoped via withUser,
// so it only ever sees the requesting user's rows. limit <= 0 returns nil.
func (s *Store) NewestUnsummarizedVideoIDs(ctx context.Context, userID string, limit int) ([]string, error) {
// nullDuration maps an unknown duration (0) to SQL NULL so the upsert's
// COALESCE(EXCLUDED.duration_s, videos.duration_s) preserves a previously-known
// value instead of clobbering it with 0 (ADR-028; the channel_title backfill
// stance, migration 014).
func nullDuration(seconds int) *int {
if seconds <= 0 {
return nil
}
return &seconds
}
// OnboardBurstVideoIDs returns up to limit of the user's unsummarized videos for
// the connect-time onboarding burst (ADR-028), newest-first but quality-aware: a
// video is excluded when its duration is KNOWN and outside [minSeconds, maxSeconds]
// — dropping Shorts (below min) and multi-hour livestream VODs (above max) that
// would waste a scarce caption fetch on a poor first impression. A NULL/unknown
// duration is kept (degrade-open) but ranked AFTER known-good rows, so a freshly
// enriched good pick wins when both exist. minSeconds<=0 / maxSeconds<=0 each
// disable that bound (0/0 == pure newest-first, the reversibility lever).
// RLS-scoped via withUser; limit <= 0 returns nil.
func (s *Store) OnboardBurstVideoIDs(ctx context.Context, userID string, limit, minSeconds, maxSeconds int) ([]string, error) {
if limit <= 0 {
return nil, nil
}
@@ -91,16 +108,21 @@ func (s *Store) NewestUnsummarizedVideoIDs(ctx context.Context, userID string, l
AND NOT EXISTS (
SELECT 1 FROM summaries su
WHERE su.user_id = v.user_id AND su.video_id = v.id)
ORDER BY v.published_at DESC NULLS LAST, v.seen_at DESC
LIMIT $2`, userID, limit)
AND NOT (
v.duration_s IS NOT NULL
AND ( ($3 > 0 AND v.duration_s < $3)
OR ($4 > 0 AND v.duration_s > $4) ))
ORDER BY (v.duration_s IS NOT NULL) DESC,
v.published_at DESC NULLS LAST, v.seen_at DESC
LIMIT $2`, userID, limit, minSeconds, maxSeconds)
if err != nil {
return fmt.Errorf("store: newest unsummarized: %w", err)
return fmt.Errorf("store: onboard burst videos: %w", err)
}
defer rows.Close()
for rows.Next() {
var id string
if err := rows.Scan(&id); err != nil {
return fmt.Errorf("store: scan newest unsummarized: %w", err)
return fmt.Errorf("store: scan onboard burst video: %w", err)
}
ids = append(ids, id)
}
@@ -110,3 +132,31 @@ func (s *Store) NewestUnsummarizedVideoIDs(ctx context.Context, userID string, l
}
return ids, nil
}
// DistinctChannels returns the user's distinct, non-empty source channel titles
// (the channels they have videos from), alphabetically — the option list for the
// feed's channel filter. RLS-scoped via withUser.
func (s *Store) DistinctChannels(ctx context.Context, userID string) ([]string, error) {
var out []string
if err := s.withUser(ctx, userID, func(tx pgx.Tx) error {
rows, err := tx.Query(ctx,
`SELECT DISTINCT channel_title FROM videos
WHERE user_id = $1 AND channel_title IS NOT NULL AND channel_title <> ''
ORDER BY channel_title`, userID)
if err != nil {
return fmt.Errorf("store: distinct channels: %w", err)
}
defer rows.Close()
for rows.Next() {
var c string
if err := rows.Scan(&c); err != nil {
return fmt.Errorf("store: scan channel: %w", err)
}
out = append(out, c)
}
return rows.Err()
}); err != nil {
return nil, err
}
return out, nil
}
+90 -15
View File
@@ -7,7 +7,7 @@ import (
"github.com/stretchr/testify/require"
"gitea.d-ma.be/mathias/tapir/internal/domain"
"git.d-ma.be/mathias/tapir/internal/domain"
)
func ytVideo(userID, provVideoID, title string) domain.Video {
@@ -52,6 +52,36 @@ func TestUpsertVideo_ReturnsStableID(t *testing.T) {
require.Equal(t, 1, count, "must not duplicate the row")
}
func TestUpsertVideo_PersistsAndPreservesDuration(t *testing.T) {
ctx := context.Background()
s := newStore(t)
resetDB(t, rawPool(t))
// First upsert carries a known duration (ADR-028: discovery enriches it).
v := ytVideo(userA, "dur0000001x", "with duration")
v.DurationSeconds = 750
id, err := s.UpsertVideo(ctx, v)
require.NoError(t, err)
p := rawPool(t)
readDuration := func() *int {
var d *int
require.NoError(t, p.QueryRow(ctx, `SELECT duration_s FROM videos WHERE id = $1`, id).Scan(&d))
return d
}
require.NotNil(t, readDuration())
require.Equal(t, 750, *readDuration(), "duration must persist")
// A later upsert that does NOT know the duration (0) must not clobber it —
// the channel_title backfill stance (migration 014): COALESCE-preserve.
v2 := ytVideo(userA, "dur0000001x", "title updated, duration unknown")
v2.DurationSeconds = 0
_, err = s.UpsertVideo(ctx, v2)
require.NoError(t, err)
require.NotNil(t, readDuration(), "a 0/unknown re-upsert must not erase a known duration")
require.Equal(t, 750, *readDuration())
}
func TestUpsertVideo_IDMatchesSummaryDedup(t *testing.T) {
ctx := context.Background()
s := newStore(t)
@@ -82,34 +112,79 @@ func TestUpsertVideo_PerUserIsolation(t *testing.T) {
require.NotEqual(t, idA, idB, "same provider video for two users must be two distinct rows")
}
func TestNewestUnsummarizedVideoIDs(t *testing.T) {
func TestOnboardBurstVideoIDs(t *testing.T) {
ctx := context.Background()
s := newStore(t)
resetDB(t, rawPool(t))
mk := func(user, pid string, day int) string {
mk := func(user, pid string, day, dur int) string {
v := ytVideo(user, pid, pid)
v.PublishedAt = time.Date(2026, 6, day, 12, 0, 0, 0, time.UTC)
v.DurationSeconds = dur // 0 == unknown (NULL)
id, err := s.UpsertVideo(ctx, v)
require.NoError(t, err)
return id
}
_ = mk(userA, "a1vid000001", 1)
id2 := mk(userA, "a2vid000002", 2)
id3 := mk(userA, "a3vid000003", 3)
id4 := mk(userA, "a4vid000004", 4)
mk(userB, "b1vid000009", 9) // userB's newest — must never leak via RLS
summarized := mk(userA, "summ0000001", 6, 600) // newest known-good, but already summarized
good1 := mk(userA, "good0000001", 5, 600) // 10m, newest UNsummarized known-good
tooLong := mk(userA, "toolong0001", 4, 20000) // > maxSeconds -> dropped
tooShort := mk(userA, "tooshort001", 3, 30) // < minSeconds -> dropped
unknown := mk(userA, "unknown0001", 2, 0) // NULL duration -> kept, ranked last
good2 := mk(userA, "good0000002", 1, 800) // known-good but oldest
mk(userB, "bvid0000009", 9, 600) // userB -> must not leak via RLS
// The newest (v4) is summarized, so it's excluded from "unsummarized".
require.NoError(t, s.Deliver(ctx, summary(userA, id4, "done")))
// The newest video is summarized, so it is excluded from the burst.
require.NoError(t, s.Deliver(ctx, summary(userA, summarized, "done")))
// Cap 2, newest-first unsummarized: v3 then v2 (v4 excluded; userB excluded).
got, err := s.NewestUnsummarizedVideoIDs(ctx, userA, 2)
const minSec, maxSec = 60, 14400
// Known-good ranked before unknown, each newest-first within its group; the
// too-long and too-short videos are excluded by their known duration.
got, err := s.OnboardBurstVideoIDs(ctx, userA, 5, minSec, maxSec)
require.NoError(t, err)
require.Equal(t, []string{id3, id2}, got)
require.Equal(t, []string{good1, good2, unknown}, got,
"known-good first (newest-first), then unknown-duration; junk excluded")
none, err := s.NewestUnsummarizedVideoIDs(ctx, userA, 0)
// Cap is honoured.
capped, err := s.OnboardBurstVideoIDs(ctx, userA, 2, minSec, maxSec)
require.NoError(t, err)
require.Empty(t, none, "limit 0 returns nothing")
require.Equal(t, []string{good1, good2}, capped)
// Bounds disabled (0/0) == pure newest-first, nothing excluded.
all, err := s.OnboardBurstVideoIDs(ctx, userA, 10, 0, 0)
require.NoError(t, err)
require.ElementsMatch(t, []string{good1, tooLong, tooShort, unknown, good2}, all,
"0/0 bounds disable the duration filter (prior newest-first behaviour)")
// limit <= 0 returns nothing.
none, err := s.OnboardBurstVideoIDs(ctx, userA, 0, minSec, maxSec)
require.NoError(t, err)
require.Empty(t, none)
}
func TestUpsertVideoPersistsChannelAndDistinctChannels(t *testing.T) {
ctx := context.Background()
s := newStore(t)
resetDB(t, rawPool(t))
mk := func(pid, channel string) {
v := ytVideo(userA, pid, pid)
v.ChannelTitle = channel
_, err := s.UpsertVideo(ctx, v)
require.NoError(t, err)
}
mk("aa11111aaaa", "Acme Talks")
mk("bb22222bbbb", "Acme Talks") // same channel
mk("cc33333cccc", "Zeta Channel")
// userB's channel must not leak.
vb := ytVideo(userB, "dd44444dddd", "x")
vb.ChannelTitle = "Bravo Only"
_, err := s.UpsertVideo(ctx, vb)
require.NoError(t, err)
got, err := s.DistinctChannels(ctx, userA)
require.NoError(t, err)
require.Equal(t, []string{"Acme Talks", "Zeta Channel"}, got,
"distinct, alphabetical, user-scoped (no Bravo Only)")
}
+154 -41
View File
@@ -1,19 +1,25 @@
// Package summarizer implements ports.Summarizer backed by the copied llm
// package's local-Primary -> BYO-Fallback routing (ADR-004). It is the only
// place content ever leaves the engine toward an AI model, so it is also the
// enforcement point for the local-first guarantee in
// docs/use-cases/ai_routing.feature: a user with no BYO provider configured has
// their content sent to the local stack and nowhere else.
// package's routing (ADR-004, extended by ADR-022). It is the only place content
// ever leaves the engine toward an AI model, so it is also the enforcement point
// for the local-first guarantee in docs/use-cases/ai_routing.feature: endpoints
// are tried in order, locals first, so content only reaches an external model
// after every local endpoint has failed — and never at all when no external
// endpoint is configured.
package summarizer
import (
"bytes"
"context"
"encoding/json"
"errors"
"fmt"
"log/slog"
"strings"
"time"
"unicode/utf8"
"gitea.d-ma.be/mathias/tapir/internal/domain"
"git.d-ma.be/mathias/tapir/internal/domain"
"git.d-ma.be/mathias/tapir/internal/metrics"
)
// Completer is the minimal LLM chat surface the Summarizer needs.
@@ -29,19 +35,41 @@ type Endpoint struct {
Model string // resolved alias, e.g. "iguana/deepseek-r1-14b"
}
// Summarizer routes a transcript through the local endpoint first, then the
// optional BYO endpoint. It owns its routing (rather than delegating to
// llm.Router) so it can record which provider answered and whether the fallback
// was used — information llm.Router collapses away.
// Summarizer routes a transcript through an ordered chain of endpoints, trying
// each in turn until one returns a parseable summary. It owns its routing
// (rather than delegating to llm.Router) so it can record which provider answered
// and whether a fallback was used — information llm.Router collapses away. The
// chain ordering is the local-first guarantee: callers place local endpoints
// first and any external endpoint last, so content only reaches an external model
// after every local endpoint has failed.
type Summarizer struct {
primary Endpoint
fallback *Endpoint // nil => no BYO; primary errors are returned, never sent externally
now func() time.Time
endpoints []Endpoint
maxInputChars int // transcript truncation budget; 0 = no limit
now func() time.Time
}
// New constructs a Summarizer. fallback may be nil (no BYO provider configured).
// New constructs a Summarizer from a primary endpoint and an optional fallback
// (the historical local-Primary -> BYO-Fallback shape, ADR-004). A nil fallback
// means a single-endpoint chain: errors are returned, content never leaves it.
func New(primary Endpoint, fallback *Endpoint) *Summarizer {
return &Summarizer{primary: primary, fallback: fallback, now: time.Now}
eps := []Endpoint{primary}
if fallback != nil {
eps = append(eps, *fallback)
}
return &Summarizer{endpoints: eps, now: time.Now}
}
// NewChain constructs a Summarizer over an ordered endpoint chain (ADR-022).
// endpoints are tried in order; the first to return a parseable summary wins, and
// FallbackUsed is recorded true for any endpoint past the first. maxInputChars
// bounds the transcript text sent to every endpoint (0 = unbounded), so a long
// transcript does not overflow a small-context primary model's window. It panics
// on an empty chain — a wiring bug, not a runtime condition.
func NewChain(endpoints []Endpoint, maxInputChars int) *Summarizer {
if len(endpoints) == 0 {
panic("summarizer: NewChain requires at least one endpoint")
}
return &Summarizer{endpoints: endpoints, maxInputChars: maxInputChars, now: time.Now}
}
const systemPrompt = `You are Tapir, a video-summarization assistant.
@@ -53,31 +81,42 @@ Respond with ONLY a JSON object, no prose and no code fences:
- "takeaways": the actionable conclusions a viewer should leave with.
Output the JSON object and nothing else.`
// Summarize implements ports.Summarizer.
// Summarize implements ports.Summarizer. It walks the endpoint chain in order:
// the first endpoint whose reply parses into a non-empty summary wins. An
// endpoint is considered failed — and the next one tried — when the model call
// errors OR when its reply cannot be parsed (a 200 with malformed JSON or a
// highlights field the model emitted as a bare string). Truncation is applied
// once, up front, so every endpoint sees the same bounded prompt. When the whole
// chain fails, the joined error is returned so the engine queues the work for
// retry and delivers no summary.
func (s *Summarizer) Summarize(ctx context.Context, v domain.Video, t domain.Transcript) (domain.Summary, error) {
if !t.HasText() {
return domain.Summary{}, fmt.Errorf("summarize: transcript for video %s has no text", v.ID)
}
user := buildUserPrompt(v, t)
user := buildUserPrompt(v, t, s.maxInputChars)
// Primary = local stack. Only on its failure is anything sent externally,
// and only when a BYO fallback is configured.
out, err := s.primary.Client.Complete(ctx, systemPrompt, user)
if err == nil {
return s.build(v, s.primary, false, out)
var errs []error
for i, ep := range s.endpoints {
fallback := i > 0
start := time.Now()
out, err := ep.Client.Complete(ctx, systemPrompt, user)
dur := time.Since(start)
if err != nil {
metrics.ObserveSummarize(ep.Model, "error", fallback, dur)
errs = append(errs, fmt.Errorf("%s/%s call: %w", ep.Provider, ep.Model, err))
continue
}
sum, perr := s.build(v, ep, fallback, out)
if perr != nil {
metrics.ObserveSummarize(ep.Model, "parse_error", fallback, dur)
errs = append(errs, fmt.Errorf("%s/%s output: %w", ep.Provider, ep.Model, perr))
continue
}
metrics.ObserveSummarize(ep.Model, "success", fallback, dur)
slog.Default().Info("summarized", "model", ep.Model, "fallback", fallback, "elapsed_ms", dur.Milliseconds())
return sum, nil
}
if s.fallback == nil {
// No BYO: content was sent to the local stack only. Surface the error so
// the engine can queue the work for retry; deliver no summary.
return domain.Summary{}, fmt.Errorf("summarize: local AI failed and no BYO provider configured: %w", err)
}
out, ferr := s.fallback.Client.Complete(ctx, systemPrompt, user)
if ferr != nil {
return domain.Summary{}, fmt.Errorf("summarize: local AI failed: %w; BYO %s failed: %v", err, s.fallback.Provider, ferr)
}
return s.build(v, *s.fallback, true, out)
return domain.Summary{}, fmt.Errorf("summarize: all %d endpoint(s) failed: %w", len(s.endpoints), errors.Join(errs...))
}
func (s *Summarizer) build(v domain.Video, ep Endpoint, fallbackUsed bool, raw string) (domain.Summary, error) {
@@ -89,8 +128,8 @@ func (s *Summarizer) build(v domain.Video, ep Endpoint, fallbackUsed bool, raw s
UserID: v.UserID,
VideoID: v.ID,
Summary: parsed.Summary,
Highlights: parsed.Highlights,
Takeaways: parsed.Takeaways,
Highlights: []string(parsed.Highlights),
Takeaways: []string(parsed.Takeaways),
AIProvider: ep.Provider,
AIModel: ep.Model,
FallbackUsed: fallbackUsed,
@@ -98,20 +137,94 @@ func (s *Summarizer) build(v domain.Video, ep Endpoint, fallbackUsed bool, raw s
}, nil
}
func buildUserPrompt(v domain.Video, t domain.Transcript) string {
func buildUserPrompt(v domain.Video, t domain.Transcript, maxInputChars int) string {
var b strings.Builder
fmt.Fprintf(&b, "Title: %s\n", v.Title)
if v.URL != "" {
fmt.Fprintf(&b, "URL: %s\n", v.URL)
}
fmt.Fprintf(&b, "\nTranscript:\n%s", t.Content)
fmt.Fprintf(&b, "\nTranscript:\n%s", truncate(t.Content, maxInputChars))
return b.String()
}
// truncate caps content to max bytes on a UTF-8 rune boundary, appending a
// marker so the model knows the transcript was cut. A non-positive max (or a
// content already within budget) returns content unchanged. Bounding the input
// keeps a long transcript from overflowing a small-context model's window — the
// production failure mode where koala/phi4-mini's 8k context returned HTTP 400 on
// a 11.6k-token transcript.
func truncate(content string, max int) string {
if max <= 0 || len(content) <= max {
return content
}
cut := max
for cut > 0 && !utf8.RuneStart(content[cut]) {
cut--
}
return content[:cut] + "\n…[transcript truncated to fit the model context]"
}
// flexStrings is a []string that also unmarshals from a single JSON string or a
// JSON array of scalars. Small local models (koala/phi4-mini) sometimes emit
// "highlights": "one point" instead of an array, or mix in a number; rather than
// fail the whole summary on that quirk, coerce to []string. Empty/whitespace
// elements are dropped.
type flexStrings []string
func (f *flexStrings) UnmarshalJSON(b []byte) error {
b = bytes.TrimSpace(b)
if len(b) == 0 || string(b) == "null" {
*f = nil
return nil
}
if b[0] == '[' {
var raw []json.RawMessage
if err := json.Unmarshal(b, &raw); err != nil {
return err
}
out := make([]string, 0, len(raw))
for _, r := range raw {
s, err := rawToString(r)
if err != nil {
return err
}
if strings.TrimSpace(s) != "" {
out = append(out, s)
}
}
*f = out
return nil
}
s, err := rawToString(b)
if err != nil {
return err
}
if strings.TrimSpace(s) == "" {
*f = nil
} else {
*f = flexStrings{s}
}
return nil
}
// rawToString renders a JSON scalar as text: a quoted string is unquoted; any
// other scalar (number, bool) is kept as its literal source so no content is lost.
func rawToString(r json.RawMessage) (string, error) {
r = bytes.TrimSpace(r)
if len(r) > 0 && r[0] == '"' {
var s string
if err := json.Unmarshal(r, &s); err != nil {
return "", err
}
return s, nil
}
return string(r), nil
}
type parsedSummary struct {
Summary string `json:"summary"`
Highlights []string `json:"highlights"`
Takeaways []string `json:"takeaways"`
Summary string `json:"summary"`
Highlights flexStrings `json:"highlights"`
Takeaways flexStrings `json:"takeaways"`
}
// parse extracts the JSON object from a model reply. Thinking models (qwen3,
+121 -4
View File
@@ -7,13 +7,36 @@ package summarizer
import (
"context"
"errors"
"net/http"
"net/http/httptest"
"strings"
"testing"
"gitea.d-ma.be/mathias/tapir/internal/domain"
"gitea.d-ma.be/mathias/tapir/internal/ports"
"git.d-ma.be/mathias/tapir/internal/domain"
"git.d-ma.be/mathias/tapir/internal/metrics"
"git.d-ma.be/mathias/tapir/internal/ports"
)
// TestSummarizerRecordsMetric verifies the summarizer→metrics wiring (ADR-030)
// black-box: after a successful summarize, the public /metrics scrape shows a
// success observation for that endpoint's model.
func TestSummarizerRecordsMetric(t *testing.T) {
const model = "metrics-test-model"
s := New(Endpoint{Client: &fakeClient{reply: goodReply}, Provider: "local", Model: model}, nil)
if _, err := s.Summarize(context.Background(), testVideo(), testTranscript()); err != nil {
t.Fatalf("Summarize: %v", err)
}
rec := httptest.NewRecorder()
metrics.Handler().ServeHTTP(rec, httptest.NewRequest(http.MethodGet, "/metrics", nil))
body := rec.Body.String()
if !strings.Contains(body, `tapir_summarize_duration_seconds`) ||
!strings.Contains(body, `model="`+model+`"`) ||
!strings.Contains(body, `outcome="success"`) {
t.Errorf("metrics scrape missing summarize success for %s", model)
}
}
// compile-time check: Summarizer satisfies the port.
var _ ports.Summarizer = (*Summarizer)(nil)
@@ -129,8 +152,8 @@ func TestSummarize_NoBYO_ContentOnlyLocal(t *testing.T) {
local := &fakeClient{reply: goodReply}
s := New(Endpoint{Client: local, Provider: "local", Model: "iguana/deepseek-r1-14b"}, nil)
if s.fallback != nil {
t.Fatal("no BYO configured but fallback endpoint is non-nil")
if len(s.endpoints) != 1 {
t.Fatalf("no BYO configured but chain has %d endpoints, want 1", len(s.endpoints))
}
for i := 0; i < 3; i++ {
sum, err := s.Summarize(context.Background(), testVideo(), testTranscript())
@@ -176,3 +199,97 @@ func TestParse_EmptySummaryRejected(t *testing.T) {
t.Fatal("want error for empty summary (thinking model returned no content)")
}
}
// parse tolerates a small model emitting "highlights" as a bare string instead
// of an array — the production koala/phi4-mini quirk that errored with
// "cannot unmarshal string into Go struct field ... highlights of type []string".
func TestParse_ToleratesStringHighlights(t *testing.T) {
p, err := parse(`{"summary":"s","highlights":"one big point","takeaways":["a","b"]}`)
if err != nil {
t.Fatalf("parse: %v", err)
}
if len(p.Highlights) != 1 || p.Highlights[0] != "one big point" {
t.Errorf("highlights = %v, want [\"one big point\"]", p.Highlights)
}
if len(p.Takeaways) != 2 {
t.Errorf("takeaways = %v, want 2", p.Takeaways)
}
}
// Chain: an endpoint that returns a 200 with unparseable output is treated as a
// failure, and the next endpoint in the chain is tried. This is the case the old
// primary->fallback shape missed — a parse error short-circuited instead of
// falling back.
func TestSummarize_FallsBackOnMalformedOutput(t *testing.T) {
bad := &fakeClient{reply: `{"summary": not json`}
good := &fakeClient{reply: goodReply}
s := NewChain([]Endpoint{
{Client: bad, Provider: "local", Model: "koala/phi4-mini"},
{Client: good, Provider: "local", Model: "koala/phi4-14b"},
}, 0)
sum, err := s.Summarize(context.Background(), testVideo(), testTranscript())
if err != nil {
t.Fatalf("Summarize: %v", err)
}
if sum.AIModel != "koala/phi4-14b" {
t.Errorf("AIModel = %q, want koala/phi4-14b (fell back past malformed primary)", sum.AIModel)
}
if !sum.FallbackUsed {
t.Error("FallbackUsed = false, want true")
}
if bad.calls != 1 || good.calls != 1 {
t.Errorf("calls: bad=%d good=%d, want 1 and 1", bad.calls, good.calls)
}
}
// Chain: when every endpoint fails, no summary is produced and the joined error
// names each failure so the engine queues the work for retry.
func TestSummarize_ChainAllEndpointsFail(t *testing.T) {
a := &fakeClient{err: errors.New("context overflow")}
b := &fakeClient{reply: "not even json"}
s := NewChain([]Endpoint{
{Client: a, Provider: "local", Model: "m1"},
{Client: b, Provider: "berget", Model: "m2"},
}, 0)
if _, err := s.Summarize(context.Background(), testVideo(), testTranscript()); err == nil {
t.Fatal("want error when all endpoints fail")
}
if a.calls != 1 || b.calls != 1 {
t.Errorf("calls: a=%d b=%d, want 1 and 1", a.calls, b.calls)
}
}
// A transcript longer than the chain's input budget is truncated before it
// reaches any model, so a small-context primary does not overflow its window.
func TestSummarize_TruncatesLongTranscript(t *testing.T) {
local := &fakeClient{reply: goodReply}
const budget = 100
s := NewChain([]Endpoint{{Client: local, Provider: "local", Model: "m"}}, budget)
long := domain.Transcript{
VideoID: "vid-1", UserID: "user-1", Source: domain.SourceCaptions,
Content: strings.Repeat("word ", 1000), // 5000 bytes, well over budget
}
if _, err := s.Summarize(context.Background(), testVideo(), long); err != nil {
t.Fatalf("Summarize: %v", err)
}
// The prompt carries title/URL framing plus the truncation marker, so allow
// headroom over the raw transcript budget — but it must be far below 5000.
if len(local.lastUser) > budget+300 {
t.Errorf("prompt length = %d, want <= %d (transcript not truncated)", len(local.lastUser), budget+300)
}
if !strings.Contains(local.lastUser, "truncated") {
t.Error("truncation marker missing from prompt")
}
}
func TestNewChain_PanicsOnEmptyChain(t *testing.T) {
defer func() {
if recover() == nil {
t.Fatal("want panic on empty endpoint chain")
}
}()
NewChain(nil, 0)
}
+32 -1
View File
@@ -7,10 +7,13 @@ import (
"encoding/xml"
"fmt"
"io"
"log/slog"
"net/http"
"strings"
"time"
"gitea.d-ma.be/mathias/tapir/internal/domain"
"git.d-ma.be/mathias/tapir/internal/domain"
"git.d-ma.be/mathias/tapir/internal/metrics"
)
// defaultPlayerBaseURL is the InnerTube / watch-page host. Overridable via
@@ -43,7 +46,35 @@ const maxCaptionBytes = 16 << 20 // 16 MiB
// fetch, or an unparseable body all yield SourceNone rather than an error. Only
// genuine transport (network) faults return an error. Audio download and
// speech-to-text remain absent (ADR-007).
// FetchTranscript times the caption fetch and records its latency by outcome
// (ADR-030) before returning. Transport errors are surfaced to the caller and not
// recorded as an outcome (logged upstream); the three resolved outcomes
// captions|none|rate_limited are the ones that consume the scarce fetch budget.
func (a *Adapter) FetchTranscript(ctx context.Context, v domain.Video) (domain.Transcript, error) {
start := time.Now()
tr, err := a.fetchTranscript(ctx, v)
if err == nil {
dur := time.Since(start)
outcome := captionOutcome(tr.Source)
metrics.ObserveCaptionFetch(outcome, dur)
slog.Default().Info("caption fetch", "video", v.ProviderVideoID, "outcome", outcome, "elapsed_ms", dur.Milliseconds())
}
return tr, err
}
// captionOutcome maps a transcript source to the metric outcome label.
func captionOutcome(s domain.TranscriptSource) string {
switch s {
case domain.SourceCaptions:
return "captions"
case domain.SourceRateLimited:
return "rate_limited"
default:
return "none"
}
}
func (a *Adapter) fetchTranscript(ctx context.Context, v domain.Video) (domain.Transcript, error) {
client := a.plainClient()
tracks, err := a.captionTracks(ctx, client, v.ProviderVideoID)
+47
View File
@@ -2,6 +2,7 @@ package youtube
import (
"context"
"sync/atomic"
"time"
"golang.org/x/time/rate"
@@ -29,9 +30,55 @@ func SetFetchRate(interval time.Duration) {
globalFetchGate = rate.NewLimiter(rate.Every(interval), 1)
}
// foregroundPending counts in-flight foreground (user-initiated) caption fetches.
// The background sweep yields the gate while this is non-zero so a human waiting
// on a click gets the next slot — and, on a near-throttled IP, the pre-429 window
// — instead of competing equally with the firehose (ADR-026, Pillar A). Clicks are
// rare and bursty, so background barely notices; the win to the click is large.
var foregroundPending atomic.Int64
// fgCtxKey marks a context as foreground (user-initiated). Unexported; set via
// ForegroundContext and read via isForeground so only this package owns the key.
type fgCtxKey struct{}
// ForegroundContext marks ctx as a user-initiated (foreground) fetch so the gate
// gives it priority. The web "Summarize"/paste/retry path wraps its context with
// this; the background scheduler leaves it unset.
func ForegroundContext(ctx context.Context) context.Context {
return context.WithValue(ctx, fgCtxKey{}, true)
}
func isForeground(ctx context.Context) bool {
v, _ := ctx.Value(fgCtxKey{}).(bool)
return v
}
// fgYieldPoll is how often a background waiter re-checks whether a foreground
// fetch is still pending. Short enough to feel immediate, long enough not to spin.
const fgYieldPoll = 200 * time.Millisecond
// WaitFetchGate blocks until the process-wide gate allows one timedtext fetch,
// respecting ctx cancellation. Called from httpDo before every live outbound
// caption fetch so the scheduler and the click-path share the same egress budget.
//
// Foreground (user-initiated) fetches take priority: they register as pending and
// acquire a token immediately. Background fetches first yield — they wait until no
// foreground fetch is pending — so a live click is never stuck behind the
// background sweep and gets the cleaner slot against the per-IP limit (ADR-026).
func WaitFetchGate(ctx context.Context) error {
if isForeground(ctx) {
foregroundPending.Add(1)
defer foregroundPending.Add(-1)
return globalFetchGate.Wait(ctx)
}
// Background: defer to any pending foreground fetch before taking a token.
for foregroundPending.Load() > 0 {
select {
case <-ctx.Done():
return ctx.Err()
case <-time.After(fgYieldPoll):
}
}
return globalFetchGate.Wait(ctx)
}
+34
View File
@@ -82,3 +82,37 @@ func TestSetFetchRateZeroIsUnlimited(t *testing.T) {
require.NoError(t, WaitFetchGate(context.Background()))
}
}
func TestForegroundContextMarker(t *testing.T) {
require.False(t, isForeground(context.Background()), "plain context is background")
require.True(t, isForeground(ForegroundContext(context.Background())), "marked context is foreground")
}
// TestWaitFetchGateForegroundProceedsImmediately: a foreground fetch acquires a
// token without yielding, even when background callers exist.
func TestWaitFetchGateForegroundProceedsImmediately(t *testing.T) {
SetFetchRate(0) // unlimited limiter — isolate the yield logic from pacing
foregroundPending.Store(0)
t.Cleanup(func() { foregroundPending.Store(0) })
ctx, cancel := context.WithTimeout(context.Background(), time.Second)
defer cancel()
require.NoError(t, WaitFetchGate(ForegroundContext(ctx)), "foreground proceeds immediately")
}
// TestWaitFetchGateBackgroundYieldsToForeground: while a foreground fetch is
// pending, a background fetch yields (does not take a token) until the foreground
// clears — proven by a background wait timing out against its own deadline, then
// succeeding once the foreground is done.
func TestWaitFetchGateBackgroundYieldsToForeground(t *testing.T) {
SetFetchRate(0)
foregroundPending.Store(1) // simulate a foreground fetch in flight
t.Cleanup(func() { foregroundPending.Store(0) })
ctx, cancel := context.WithTimeout(context.Background(), 250*time.Millisecond)
defer cancel()
require.Error(t, WaitFetchGate(ctx), "background yields (blocks) while foreground is pending")
foregroundPending.Store(0) // foreground done
require.NoError(t, WaitFetchGate(context.Background()), "background proceeds once foreground clears")
}
+5 -2
View File
@@ -6,7 +6,7 @@ import (
"net/http"
"testing"
"gitea.d-ma.be/mathias/tapir/internal/domain"
"git.d-ma.be/mathias/tapir/internal/domain"
)
func TestVideoByID(t *testing.T) {
@@ -21,7 +21,7 @@ func TestVideoByID(t *testing.T) {
if got := r.URL.Query().Get("part"); got != "snippet" {
t.Errorf("expected part=snippet, got %q", got)
}
_, _ = w.Write([]byte(`{"items":[{"snippet":{"title":"Never Gonna Give You Up","publishedAt":"2026-05-20T09:00:00Z"}}]}`))
_, _ = w.Write([]byte(`{"items":[{"snippet":{"title":"Never Gonna Give You Up","channelTitle":"Rick Astley","publishedAt":"2026-05-20T09:00:00Z"}}]}`))
})
v, err := a.VideoByID(context.Background(), "u1", id)
@@ -34,6 +34,9 @@ func TestVideoByID(t *testing.T) {
if v.ProviderVideoID != id || v.Title != "Never Gonna Give You Up" {
t.Errorf("unexpected video: %+v", v)
}
if v.ChannelTitle != "Rick Astley" {
t.Errorf("ChannelTitle = %q, want Rick Astley", v.ChannelTitle)
}
if v.Provider != domain.ProviderYouTube || v.URL != "https://www.youtube.com/watch?v="+id {
t.Errorf("video not wired correctly: %+v", v)
}
+114 -5
View File
@@ -27,8 +27,8 @@ import (
"golang.org/x/oauth2"
"gitea.d-ma.be/mathias/tapir/internal/domain"
"gitea.d-ma.be/mathias/tapir/internal/ports"
"git.d-ma.be/mathias/tapir/internal/domain"
"git.d-ma.be/mathias/tapir/internal/ports"
)
// defaultBaseURL is the YouTube Data API v3 root. Overridable via Config.BaseURL
@@ -66,6 +66,13 @@ type Config struct {
// poll. Zero means defaultMaxVideos.
MaxVideosPerSubscription int
// MinVideoSeconds drops videos shorter than this from discovery (Shorts/clips,
// ADR-023). NewVideos enriches candidates with a single cheap videos.list call
// (contentDetails.duration + snippet.liveBroadcastContent) and filters before
// returning, so the scarce caption-fetch budget is never spent on them. Live
// and upcoming broadcasts are dropped too. Zero disables the filter.
MinVideoSeconds int
// BaseURL overrides the Data API root. Empty means defaultBaseURL.
BaseURL string
@@ -234,6 +241,7 @@ func (a *Adapter) NewVideos(ctx context.Context, sub domain.Subscription) ([]dom
Provider: domain.ProviderYouTube,
ProviderVideoID: vid,
Title: item.Snippet.Title,
ChannelTitle: sub.ChannelTitle,
URL: "https://www.youtube.com/watch?v=" + vid,
PublishedAt: item.Snippet.PublishedAt,
})
@@ -241,7 +249,101 @@ func (a *Adapter) NewVideos(ctx context.Context, sub domain.Subscription) ([]dom
break
}
}
return videos, nil
// Drop Shorts/sub-minute clips and live/upcoming broadcasts before they ever
// reach the rate-limited caption path (ADR-023). One cheap videos.list call
// (quota API, not the timedtext throttle) supplies duration + live status.
return a.filterLowValue(ctx, client, videos), nil
}
// filterLowValue removes videos shorter than cfg.MinVideoSeconds and any live or
// upcoming broadcast, using a single videos.list lookup for duration +
// liveBroadcastContent. The filter is best-effort: if MinVideoSeconds is 0 (off)
// or the lookup fails, the input is returned unfiltered — discovery must not break
// because a metadata call hiccuped; the worst case is the pre-ADR-023 behaviour.
func (a *Adapter) filterLowValue(ctx context.Context, client *http.Client, videos []domain.Video) []domain.Video {
if a.cfg.MinVideoSeconds <= 0 || len(videos) == 0 {
return videos
}
ids := make([]string, 0, len(videos))
for _, v := range videos {
ids = append(ids, v.ProviderVideoID)
}
q := url.Values{
"part": {"contentDetails,snippet"},
"id": {strings.Join(ids, ",")},
}
var resp videoListResponse
if err := a.getJSON(ctx, client, "/videos", q, &resp); err != nil {
// Degrade open: keep the candidates rather than lose discovery.
return videos
}
type meta struct {
seconds int
live string
}
byID := make(map[string]meta, len(resp.Items))
for _, it := range resp.Items {
byID[it.ID] = meta{seconds: parseISO8601Seconds(it.ContentDetails.Duration), live: it.Snippet.LiveBroadcastContent}
}
kept := videos[:0]
for _, v := range videos {
m, ok := byID[v.ProviderVideoID]
if !ok {
kept = append(kept, v) // unknown metadata: keep, let the fetch decide
continue
}
if m.live != "" && m.live != "none" {
continue // live or upcoming broadcast
}
if m.seconds > 0 && m.seconds < a.cfg.MinVideoSeconds {
continue // Short / sub-threshold clip
}
// Carry the duration we already fetched onto the kept video so the store
// can persist it (ADR-028) — the burst's length-aware selection depends on
// it. Discarding it here was the gap the onboarding investigation found.
v.DurationSeconds = m.seconds
kept = append(kept, v)
}
return kept
}
// parseISO8601Seconds parses an ISO 8601 duration as returned by the YouTube Data
// API (e.g. "PT1H2M3S", "PT45S", "PT3M") into seconds. Only the hour/minute/second
// components YouTube emits are handled; an unparseable or zero value returns 0,
// which the caller treats as "unknown" (not filtered on duration).
func parseISO8601Seconds(d string) int {
if !strings.HasPrefix(d, "PT") {
return 0
}
d = d[2:]
total, num := 0, 0
seen := false
for _, r := range d {
switch {
case r >= '0' && r <= '9':
num = num*10 + int(r-'0')
seen = true
case r == 'H':
total += num * 3600
num, seen = 0, false
case r == 'M':
total += num * 60
num, seen = 0, false
case r == 'S':
total += num
num, seen = 0, false
default:
return 0 // unexpected component (days/weeks) — treat as unknown
}
}
if seen {
return 0 // trailing digits without a unit: malformed
}
return total
}
// VideoByID fetches a single video's metadata (videos.list, snippet) for an
@@ -270,6 +372,7 @@ func (a *Adapter) VideoByID(ctx context.Context, userID, videoID string) (domain
Provider: domain.ProviderYouTube,
ProviderVideoID: videoID,
Title: it.Snippet.Title,
ChannelTitle: it.Snippet.ChannelTitle,
URL: "https://www.youtube.com/watch?v=" + videoID,
PublishedAt: it.Snippet.PublishedAt,
}, nil
@@ -375,10 +478,16 @@ type playlistItemListResponse struct {
type videoListResponse struct {
Items []struct {
ID string `json:"id"`
Snippet struct {
Title string `json:"title"`
PublishedAt time.Time `json:"publishedAt"`
Title string `json:"title"`
ChannelTitle string `json:"channelTitle"`
PublishedAt time.Time `json:"publishedAt"`
LiveBroadcastContent string `json:"liveBroadcastContent"`
} `json:"snippet"`
ContentDetails struct {
Duration string `json:"duration"` // ISO 8601, e.g. "PT1M30S"
} `json:"contentDetails"`
} `json:"items"`
}
+95 -2
View File
@@ -7,7 +7,7 @@ import (
"net/http/httptest"
"testing"
"gitea.d-ma.be/mathias/tapir/internal/domain"
"git.d-ma.be/mathias/tapir/internal/domain"
)
// --- fakes ------------------------------------------------------------------
@@ -131,7 +131,7 @@ func TestNewVideos(t *testing.T) {
}`))
})
sub := domain.Subscription{ID: "s1", UserID: "u1", ChannelID: "UC_acme"}
sub := domain.Subscription{ID: "s1", UserID: "u1", ChannelID: "UC_acme", ChannelTitle: "Acme Channel"}
vids, err := a.NewVideos(context.Background(), sub)
if err != nil {
t.Fatalf("NewVideos: %v", err)
@@ -143,6 +143,9 @@ func TestNewVideos(t *testing.T) {
if v.ProviderVideoID != "vid1" || v.Title != "Designing for Attention" {
t.Errorf("unexpected video: %+v", v)
}
if v.ChannelTitle != "Acme Channel" {
t.Errorf("ChannelTitle = %q, want Acme Channel", v.ChannelTitle)
}
if v.Provider != domain.ProviderYouTube || v.URL != "https://www.youtube.com/watch?v=vid1" {
t.Errorf("video not wired correctly: %+v", v)
}
@@ -183,6 +186,96 @@ func TestNewVideosCapsAtMax(t *testing.T) {
}
}
// TestNewVideosFiltersShortsAndLive: with MinVideoSeconds set, discovery enriches
// candidates via videos.list and drops sub-threshold clips (Shorts) and
// live/upcoming broadcasts before they reach the rate-limited caption path.
func TestNewVideosFiltersShortsAndLive(t *testing.T) {
a, _ := newTestAdapter(t, func(w http.ResponseWriter, r *http.Request) {
switch r.URL.Path {
case "/playlistItems":
_, _ = w.Write([]byte(`{
"items": [
{"snippet": {"title": "Real Talk", "publishedAt": "2026-06-03T10:00:00Z", "resourceId": {"videoId": "long1"}}},
{"snippet": {"title": "A Short", "publishedAt": "2026-06-03T09:00:00Z", "resourceId": {"videoId": "short1"}}},
{"snippet": {"title": "Live Now", "publishedAt": "2026-06-03T08:00:00Z", "resourceId": {"videoId": "live1"}}}
]
}`))
case "/videos":
if got := r.URL.Query().Get("part"); got != "contentDetails,snippet" {
t.Errorf("videos.list part=%q, want contentDetails,snippet", got)
}
_, _ = w.Write([]byte(`{
"items": [
{"id": "long1", "contentDetails": {"duration": "PT12M30S"}, "snippet": {"liveBroadcastContent": "none"}},
{"id": "short1", "contentDetails": {"duration": "PT45S"}, "snippet": {"liveBroadcastContent": "none"}},
{"id": "live1", "contentDetails": {"duration": "PT0S"}, "snippet": {"liveBroadcastContent": "live"}}
]
}`))
default:
t.Errorf("unexpected path %q", r.URL.Path)
}
})
a.cfg.MinVideoSeconds = 60
vids, err := a.NewVideos(context.Background(), domain.Subscription{ID: "s1", UserID: "u1", ChannelID: "UC_acme"})
if err != nil {
t.Fatalf("NewVideos: %v", err)
}
if len(vids) != 1 || vids[0].ProviderVideoID != "long1" {
t.Fatalf("expected only long1 to survive the filter, got %+v", vids)
}
// The duration fetched for the filter is carried onto the kept video so the
// store can persist it (ADR-028) instead of discarding it.
if vids[0].DurationSeconds != 750 {
t.Fatalf("kept video DurationSeconds = %d, want 750 (PT12M30S)", vids[0].DurationSeconds)
}
}
// TestNewVideosNoFilterWhenDisabled: MinVideoSeconds=0 keeps the pre-ADR-023
// behaviour — no videos.list call, no filtering.
func TestNewVideosNoFilterWhenDisabled(t *testing.T) {
a, _ := newTestAdapter(t, func(w http.ResponseWriter, r *http.Request) {
if r.URL.Path == "/videos" {
t.Errorf("videos.list must not be called when MinVideoSeconds is 0")
}
_, _ = w.Write([]byte(`{"items": [
{"snippet": {"title": "A Short", "publishedAt": "2026-06-03T09:00:00Z", "resourceId": {"videoId": "short1"}}}
]}`))
})
a.cfg.MinVideoSeconds = 0
vids, err := a.NewVideos(context.Background(), domain.Subscription{ID: "s1", UserID: "u1", ChannelID: "UC_acme"})
if err != nil {
t.Fatalf("NewVideos: %v", err)
}
if len(vids) != 1 {
t.Fatalf("filter disabled must keep all videos, got %d", len(vids))
}
}
func TestParseISO8601Seconds(t *testing.T) {
cases := []struct {
in string
want int
}{
{"PT45S", 45},
{"PT1M30S", 90},
{"PT3M", 180},
{"PT1H2M3S", 3723},
{"PT2H", 7200},
{"PT0S", 0},
{"", 0},
{"garbage", 0},
{"P1D", 0}, // days component not handled → unknown
{"PT10", 0}, // trailing digits without a unit → malformed
}
for _, c := range cases {
if got := parseISO8601Seconds(c.in); got != c.want {
t.Errorf("parseISO8601Seconds(%q) = %d, want %d", c.in, got, c.want)
}
}
}
// TestUploadsPlaylistID covers the zero-cost UC->UU derivation, including
// non-standard ids that must fall through unchanged (handled via fallback).
func TestUploadsPlaylistID(t *testing.T) {
+1 -1
View File
@@ -12,7 +12,7 @@ import (
"golang.org/x/oauth2"
"gitea.d-ma.be/mathias/tapir/internal/auth"
"git.d-ma.be/mathias/tapir/internal/auth"
)
// fakeWriter is a TokenWriter capturing the persisted (ref, value).
+158 -32
View File
@@ -28,8 +28,41 @@ type Config struct {
GatewayURL string
// GatewayKey authorizes the gateway. Read from env, never committed.
GatewayKey string
// SummarizerModel is the alias in host/name form, e.g. "koala/phi4-mini".
// SummarizerModel is the primary summarizer alias in host/name form, tried
// first on every video, e.g. "koala/phi4-mini".
SummarizerModel string
// FallbackModel is the LOCAL fallback alias tried when the primary fails or
// returns unparseable output (ADR-022). Kept local so content stays on the
// homelab stack. Default is an IGUANA model (not koala) so the fallback runs
// on a different host than the koala primary — koala carries other loads, and
// a different host also means a different egress IP for the (rare) fallback.
// Empty disables it.
FallbackModel string
// CloudFallbackModel is the worst-case EXTERNAL fallback alias, tried only
// after every local endpoint has failed (ADR-022). For client deployments set
// this empty so content never leaves the local stack. Default a berget alias.
CloudFallbackModel string
// SummaryMaxTokens caps the completion budget per summary call. Small-context
// models (koala/phi4-mini, 8k) overflow when prompt + max_tokens exceeds the
// window; a summary needs only a few hundred tokens, so the default is small.
SummaryMaxTokens int
// MaxTranscriptChars bounds the transcript text sent to the model so a long
// transcript does not overflow a small-context primary. 0 disables truncation.
MaxTranscriptChars int
// MinVideoSeconds drops videos shorter than this from discovery (Shorts and
// other sub-minute clips that are noise and waste the scarce caption-fetch
// budget, ADR-014/ADR-023). Enforced via a cheap Data API videos.list lookup at
// discovery, never the rate-limited caption path. 0 disables the filter.
MinVideoSeconds int
// ChannelCaptionlessThreshold is how many consecutive no-caption results a
// channel may yield before its videos are suppressed from caption fetching
// (ADR-024). 0 disables the per-channel caption memory entirely.
ChannelCaptionlessThreshold int
// ChannelCaptionlessWindow is how long a suppressed channel stays suppressed
// before one video is re-probed (auto-recovery for a channel that adds captions).
ChannelCaptionlessWindow time.Duration
// SummarizerTimeout bounds a single completion call. Thinking models are
// slow, so the default is generous.
SummarizerTimeout time.Duration
@@ -87,6 +120,21 @@ type Config struct {
// caption rate gate (ADR-014) — the cap bounds count, never the pacing. Default 3.
OnboardSummarizeCount int
// OnboardSummarizerModel is the summarizer alias the connect-time burst leads
// its chain with (ADR-028) — a stronger model is affordable on the ≤3 summaries
// that form a new user's first impression. It heads a burst-specific chain;
// the standard chain (ADR-022) follows as resilience. Empty (or equal to
// SummarizerModel) collapses the burst back onto the shared processor — the
// reversibility lever. Default iguana/gemma4-26b (the brain-validated model).
OnboardSummarizerModel string
// OnboardMaxVideoSeconds upper-bounds the duration of a video the onboarding
// burst will pick (ADR-028), so the burst does not spend a scarce caption fetch
// on a multi-hour livestream VOD that passed the live filter once it ended. Only
// a KNOWN duration outside [MinVideoSeconds, this] is dropped; a NULL/unknown
// duration is kept (degrade-open). 0 disables the upper bound. Default 14400 (4h).
OnboardMaxVideoSeconds int
// DiscoveryInterval, when > 0, makes `serve` run in-process scheduled discovery
// for ALL users on that cadence (ADR-018). Zero/unset = disabled, so dev and
// tests never auto-fetch. Single-replica assumption — see cmdServe.
@@ -95,6 +143,11 @@ type Config struct {
// HTTPAddr is the listen address for `tapir serve` (the Stage-0 web UI).
HTTPAddr string
// MetricsAddr is the listen address for the Prometheus /metrics endpoint
// (ADR-030). A SEPARATE port from HTTPAddr so /metrics is never exposed on the
// public app — only scraped in-cluster. Empty disables the metrics server.
MetricsAddr string
// PublicURL is the externally-reachable base URL of the deployed service,
// e.g. "https://tapir.d-ma.be". Used to build absolute links handed to humans
// (the `tapir invite` URL). No trailing slash is assumed — callers trim it.
@@ -116,19 +169,29 @@ func (c Config) DexConfigured() bool { return strings.TrimSpace(c.OIDCIssuer) !=
// Defaults (see docs/homelab-integration.md). All overridable via env.
const (
defaultGatewayURL = "http://koala:30401/v1"
defaultSummarizerModel = "koala/phi4-mini"
defaultSummarizerTimeout = 5 * time.Minute
defaultYTTokenRef = "youtube/refresh_token"
defaultYTConnectRedirectURL = "https://tapir.d-ma.be/oauth/youtube/callback"
defaultOAuthRedirectAddr = "localhost:8080"
defaultHTTPAddr = ":8080"
defaultFetchBackoff = time.Hour
defaultFetchRate = 2 * time.Second
defaultPublicURL = "https://tapir.d-ma.be"
defaultAutoSummarizeWindow = 7 * 24 * time.Hour
defaultOnboardSummarizeCount = 3
maxOnboardSummarizeCount = 5
defaultGatewayURL = "http://koala:30401/v1"
defaultSummarizerModel = "koala/phi4-mini"
defaultFallbackModel = "iguana/gemma4-26b"
defaultCloudFallbackModel = "berget/mistral-small"
defaultSummaryMaxTokens = 1500
defaultMaxTranscriptChars = 18000
defaultMinVideoSeconds = 60
defaultCaptionlessThreshold = 5
defaultCaptionlessWindow = 14 * 24 * time.Hour
defaultSummarizerTimeout = 5 * time.Minute
defaultYTTokenRef = "youtube/refresh_token"
defaultYTConnectRedirectURL = "https://tapir.d-ma.be/oauth/youtube/callback"
defaultOAuthRedirectAddr = "localhost:8080"
defaultHTTPAddr = ":8080"
defaultMetricsAddr = ":9090"
defaultFetchBackoff = time.Hour
defaultFetchRate = 2 * time.Second
defaultPublicURL = "https://tapir.d-ma.be"
defaultAutoSummarizeWindow = 7 * 24 * time.Hour
defaultOnboardSummarizeCount = 3
maxOnboardSummarizeCount = 5
defaultOnboardSummarizerModel = "iguana/gemma4-26b"
defaultOnboardMaxVideoSeconds = 14400 // 4h
)
// Load reads the environment into a Config, applying defaults. It does not
@@ -137,24 +200,28 @@ const (
// it needs.
func Load() (Config, error) {
c := Config{
UserID: os.Getenv("TAPIR_USER_ID"),
GatewayURL: envOr("TAPIR_GATEWAY_URL", defaultGatewayURL),
GatewayKey: os.Getenv("TAPIR_GATEWAY_KEY"),
SummarizerModel: envOr("TAPIR_SUMMARIZER_MODEL", defaultSummarizerModel),
DBDSN: os.Getenv("TAPIR_DB_DSN"),
YTClientID: os.Getenv("TAPIR_YT_CLIENT_ID"),
YTClientSecret: os.Getenv("TAPIR_YT_CLIENT_SECRET"),
YTTokenRef: envOr("TAPIR_YT_TOKEN_REF", defaultYTTokenRef),
YTConnectRedirectURL: envOr("TAPIR_YT_CONNECT_REDIRECT_URL", defaultYTConnectRedirectURL),
SecretsFile: envOr("TAPIR_SECRETS_FILE", defaultSecretsFile()),
OAuthRedirectAddr: envOr("TAPIR_OAUTH_REDIRECT_ADDR", defaultOAuthRedirectAddr),
HTTPAddr: envOr("TAPIR_HTTP_ADDR", defaultHTTPAddr),
PublicURL: envOr("TAPIR_PUBLIC_URL", defaultPublicURL),
OIDCIssuer: os.Getenv("TAPIR_OIDC_ISSUER"),
DexClientID: os.Getenv("TAPIR_DEX_CLIENT_ID"),
DexClientSecret: os.Getenv("TAPIR_DEX_CLIENT_SECRET"),
OIDCRedirectURL: os.Getenv("TAPIR_OIDC_REDIRECT_URL"),
SessionSecret: os.Getenv("TAPIR_SESSION_SECRET"),
UserID: os.Getenv("TAPIR_USER_ID"),
GatewayURL: envOr("TAPIR_GATEWAY_URL", defaultGatewayURL),
GatewayKey: os.Getenv("TAPIR_GATEWAY_KEY"),
SummarizerModel: envOr("TAPIR_SUMMARIZER_MODEL", defaultSummarizerModel),
OnboardSummarizerModel: lookupOr("TAPIR_ONBOARD_SUMMARIZER_MODEL", defaultOnboardSummarizerModel),
FallbackModel: lookupOr("TAPIR_FALLBACK_MODEL", defaultFallbackModel),
CloudFallbackModel: lookupOr("TAPIR_CLOUD_FALLBACK_MODEL", defaultCloudFallbackModel),
DBDSN: os.Getenv("TAPIR_DB_DSN"),
YTClientID: os.Getenv("TAPIR_YT_CLIENT_ID"),
YTClientSecret: os.Getenv("TAPIR_YT_CLIENT_SECRET"),
YTTokenRef: envOr("TAPIR_YT_TOKEN_REF", defaultYTTokenRef),
YTConnectRedirectURL: envOr("TAPIR_YT_CONNECT_REDIRECT_URL", defaultYTConnectRedirectURL),
SecretsFile: envOr("TAPIR_SECRETS_FILE", defaultSecretsFile()),
OAuthRedirectAddr: envOr("TAPIR_OAUTH_REDIRECT_ADDR", defaultOAuthRedirectAddr),
HTTPAddr: envOr("TAPIR_HTTP_ADDR", defaultHTTPAddr),
MetricsAddr: lookupOr("TAPIR_METRICS_ADDR", defaultMetricsAddr),
PublicURL: envOr("TAPIR_PUBLIC_URL", defaultPublicURL),
OIDCIssuer: os.Getenv("TAPIR_OIDC_ISSUER"),
DexClientID: os.Getenv("TAPIR_DEX_CLIENT_ID"),
DexClientSecret: os.Getenv("TAPIR_DEX_CLIENT_SECRET"),
OIDCRedirectURL: os.Getenv("TAPIR_OIDC_REDIRECT_URL"),
SessionSecret: os.Getenv("TAPIR_SESSION_SECRET"),
}
timeout, err := durationOr("TAPIR_SUMMARIZER_TIMEOUT", defaultSummarizerTimeout)
@@ -193,6 +260,45 @@ func Load() (Config, error) {
}
c.AutoSummarizeWindow = autoWindow
summaryTokens, err := intOr("TAPIR_SUMMARY_MAX_TOKENS", defaultSummaryMaxTokens)
if err != nil {
return Config{}, err
}
c.SummaryMaxTokens = summaryTokens
maxChars, err := intOr("TAPIR_MAX_TRANSCRIPT_CHARS", defaultMaxTranscriptChars)
if err != nil {
return Config{}, err
}
if maxChars < 0 {
maxChars = 0
}
c.MaxTranscriptChars = maxChars
minVideo, err := intOr("TAPIR_MIN_VIDEO_SECONDS", defaultMinVideoSeconds)
if err != nil {
return Config{}, err
}
if minVideo < 0 {
minVideo = 0
}
c.MinVideoSeconds = minVideo
captionThreshold, err := intOr("TAPIR_CHANNEL_CAPTIONLESS_THRESHOLD", defaultCaptionlessThreshold)
if err != nil {
return Config{}, err
}
if captionThreshold < 0 {
captionThreshold = 0
}
c.ChannelCaptionlessThreshold = captionThreshold
captionWindow, err := durationOr("TAPIR_CHANNEL_CAPTIONLESS_WINDOW", defaultCaptionlessWindow)
if err != nil {
return Config{}, err
}
c.ChannelCaptionlessWindow = captionWindow
onboard, err := intOr("TAPIR_ONBOARD_SUMMARIZE_COUNT", defaultOnboardSummarizeCount)
if err != nil {
return Config{}, err
@@ -205,6 +311,15 @@ func Load() (Config, error) {
}
c.OnboardSummarizeCount = onboard
onboardMax, err := intOr("TAPIR_ONBOARD_MAX_VIDEO_SECONDS", defaultOnboardMaxVideoSeconds)
if err != nil {
return Config{}, err
}
if onboardMax < 0 {
onboardMax = 0
}
c.OnboardMaxVideoSeconds = onboardMax
return c, nil
}
@@ -269,6 +384,17 @@ func envOr(key, fallback string) string {
return fallback
}
// lookupOr returns the env value when the key is PRESENT (even if empty), else
// fallback. Unlike envOr it lets an explicit empty value override the default —
// needed to DISABLE an optional fallback model (e.g. set the cloud fallback empty
// for a client deployment so content never leaves the local stack).
func lookupOr(key, fallback string) string {
if v, ok := os.LookupEnv(key); ok {
return v
}
return fallback
}
func intOr(key string, fallback int) (int, error) {
v := os.Getenv(key)
if v == "" {
+117
View File
@@ -1,11 +1,29 @@
package config
import (
"os"
"strings"
"testing"
"time"
)
// unset removes an env key for the duration of the test, restoring it after.
// Needed to observe a default for a key read with LookupEnv (where present-empty
// means "explicitly disabled", not "use default").
func unset(t *testing.T, key string) {
t.Helper()
if old, ok := os.LookupEnv(key); ok {
t.Cleanup(func() {
if err := os.Setenv(key, old); err != nil {
t.Fatalf("restore %s: %v", key, err)
}
})
}
if err := os.Unsetenv(key); err != nil {
t.Fatalf("unset %s: %v", key, err)
}
}
// setEnv sets env vars for the test and clears them afterward, so cases don't
// leak into one another. t.Setenv handles restoration.
func setEnv(t *testing.T, kv map[string]string) {
@@ -53,6 +71,46 @@ func TestLoad_AppliesDefaults(t *testing.T) {
}
}
func TestLoad_SummarizerChainDefaults(t *testing.T) {
setEnv(t, map[string]string{
"TAPIR_SUMMARIZER_MODEL": "",
"TAPIR_SUMMARY_MAX_TOKENS": "",
"TAPIR_MAX_TRANSCRIPT_CHARS": "",
})
unset(t, "TAPIR_FALLBACK_MODEL")
unset(t, "TAPIR_CLOUD_FALLBACK_MODEL")
c, err := Load()
if err != nil {
t.Fatalf("Load: %v", err)
}
if c.FallbackModel != defaultFallbackModel {
t.Errorf("FallbackModel = %q, want %q", c.FallbackModel, defaultFallbackModel)
}
if c.CloudFallbackModel != defaultCloudFallbackModel {
t.Errorf("CloudFallbackModel = %q, want %q", c.CloudFallbackModel, defaultCloudFallbackModel)
}
if c.SummaryMaxTokens != defaultSummaryMaxTokens {
t.Errorf("SummaryMaxTokens = %d, want %d", c.SummaryMaxTokens, defaultSummaryMaxTokens)
}
if c.MaxTranscriptChars != defaultMaxTranscriptChars {
t.Errorf("MaxTranscriptChars = %d, want %d", c.MaxTranscriptChars, defaultMaxTranscriptChars)
}
}
// An explicitly empty cloud-fallback env disables external routing — the lever a
// client deployment pulls so content never leaves the local stack.
func TestLoad_EmptyCloudFallbackDisables(t *testing.T) {
t.Setenv("TAPIR_CLOUD_FALLBACK_MODEL", "")
c, err := Load()
if err != nil {
t.Fatalf("Load: %v", err)
}
if c.CloudFallbackModel != "" {
t.Errorf("CloudFallbackModel = %q, want empty (disabled)", c.CloudFallbackModel)
}
}
func TestLoad_OnboardSummarizeCount(t *testing.T) {
cases := []struct {
name, env string
@@ -156,3 +214,62 @@ func TestValidateForAuth_PassesWhenComplete(t *testing.T) {
t.Errorf("ValidateForAuth: unexpected error %v", err)
}
}
func TestLoad_OnboardSummarizerModel(t *testing.T) {
cases := []struct {
name, env string
set bool
want string
}{
{"default", "", false, defaultOnboardSummarizerModel},
{"explicit", "koala/some-model", true, "koala/some-model"},
{"empty disables (collapses to shared processor)", "", true, ""},
}
for _, c := range cases {
t.Run(c.name, func(t *testing.T) {
env := map[string]string{}
if c.set {
env["TAPIR_ONBOARD_SUMMARIZER_MODEL"] = c.env
}
setEnv(t, env)
cfg, err := Load()
if err != nil {
t.Fatalf("Load: %v", err)
}
if cfg.OnboardSummarizerModel != c.want {
t.Fatalf("OnboardSummarizerModel = %q, want %q", cfg.OnboardSummarizerModel, c.want)
}
})
}
}
func TestLoad_OnboardMaxVideoSeconds(t *testing.T) {
cases := []struct {
name, env string
want int
}{
{"default", "", defaultOnboardMaxVideoSeconds},
{"explicit", "7200", 7200},
{"zero disables", "0", 0},
{"negative clamps to zero", "-9", 0},
}
for _, c := range cases {
t.Run(c.name, func(t *testing.T) {
setEnv(t, map[string]string{"TAPIR_ONBOARD_MAX_VIDEO_SECONDS": c.env})
cfg, err := Load()
if err != nil {
t.Fatalf("Load: %v", err)
}
if cfg.OnboardMaxVideoSeconds != c.want {
t.Fatalf("OnboardMaxVideoSeconds = %d, want %d", cfg.OnboardMaxVideoSeconds, c.want)
}
})
}
}
func TestLoad_OnboardMaxVideoSecondsInvalid(t *testing.T) {
setEnv(t, map[string]string{"TAPIR_ONBOARD_MAX_VIDEO_SECONDS": "long"})
if _, err := Load(); err == nil {
t.Fatal("Load: want error for non-numeric TAPIR_ONBOARD_MAX_VIDEO_SECONDS")
}
}
+6
View File
@@ -73,9 +73,15 @@ type Video struct {
Provider Provider
ProviderVideoID string
Title string
ChannelTitle string
URL string
PublishedAt time.Time
SeenAt time.Time
// DurationSeconds is the video length in seconds, when known (fetched by the
// ADR-023 videos.list enrichment at discovery). 0 means unknown — the store
// preserves a previously-known value rather than overwriting it with 0, and
// the onboarding burst (ADR-028) treats unknown as degrade-open (kept).
DurationSeconds int
}
// Transcript is the text of a video (or a record that none was available).
+139
View File
@@ -0,0 +1,139 @@
// Package metrics is Tapir's Prometheus instrumentation (ADR-030, issue #15). It
// owns the collectors and a small typed API the rest of the app calls — adapters
// never touch prometheus types directly. Two themes:
//
// - HTTP/session: request count + latency by route (the matched pattern, so
// cardinality stays bounded), and logins.
// - AI (the priority): summarization latency by model/outcome/fallback, caption
// fetch latency by outcome, chat latency by model, and LLM token usage.
//
// Handler() is served on a dedicated port (never the public app port) so a scrape
// is in-cluster only. slog timing lines are emitted at the call sites too.
package metrics
import (
"net/http"
"strconv"
"time"
"github.com/prometheus/client_golang/prometheus"
"github.com/prometheus/client_golang/prometheus/promauto"
"github.com/prometheus/client_golang/prometheus/promhttp"
)
// latencyBuckets spans sub-second UI calls up to multi-minute model calls (a cold
// local model load is tens of seconds; the cloud fallback can be longer).
var latencyBuckets = []float64{0.05, 0.1, 0.25, 0.5, 1, 2, 5, 10, 20, 30, 60, 120, 300}
var (
httpRequests = promauto.NewCounterVec(prometheus.CounterOpts{
Name: "tapir_http_requests_total",
Help: "HTTP requests by method, matched route pattern, and status code.",
}, []string{"method", "route", "code"})
httpDuration = promauto.NewHistogramVec(prometheus.HistogramOpts{
Name: "tapir_http_request_duration_seconds",
Help: "HTTP request latency by method and matched route pattern.",
Buckets: []float64{0.005, 0.01, 0.025, 0.05, 0.1, 0.25, 0.5, 1, 2, 5},
}, []string{"method", "route"})
logins = promauto.NewCounter(prometheus.CounterOpts{
Name: "tapir_logins_total",
Help: "Successful OIDC logins (session established).",
})
summarizeDuration = promauto.NewHistogramVec(prometheus.HistogramOpts{
Name: "tapir_summarize_duration_seconds",
Help: "Per-endpoint summarization latency by model, outcome (success|parse_error|error), and whether it was a fallback.",
Buckets: latencyBuckets,
}, []string{"model", "outcome", "fallback"})
captionFetchDuration = promauto.NewHistogramVec(prometheus.HistogramOpts{
Name: "tapir_caption_fetch_duration_seconds",
Help: "Caption fetch latency by outcome (captions|none|rate_limited).",
Buckets: latencyBuckets,
}, []string{"outcome"})
chatDuration = promauto.NewHistogramVec(prometheus.HistogramOpts{
Name: "tapir_chat_duration_seconds",
Help: "Per-video Q&A answer latency by model.",
Buckets: latencyBuckets,
}, []string{"model"})
llmTokens = promauto.NewCounterVec(prometheus.CounterOpts{
Name: "tapir_llm_tokens_total",
Help: "LLM tokens consumed by model and kind (prompt|completion).",
}, []string{"model", "kind"})
)
// Handler serves the Prometheus exposition format. Mount on the dedicated metrics
// port, never the public app mux.
func Handler() http.Handler { return promhttp.Handler() }
// IncLogin records a successful login.
func IncLogin() { logins.Inc() }
// ObserveSummarize records one summarization endpoint attempt.
func ObserveSummarize(model, outcome string, fallback bool, d time.Duration) {
summarizeDuration.WithLabelValues(model, outcome, strconv.FormatBool(fallback)).Observe(d.Seconds())
}
// ObserveCaptionFetch records one caption fetch by outcome.
func ObserveCaptionFetch(outcome string, d time.Duration) {
captionFetchDuration.WithLabelValues(outcome).Observe(d.Seconds())
}
// ObserveChat records one Q&A answer latency.
func ObserveChat(model string, d time.Duration) {
chatDuration.WithLabelValues(model).Observe(d.Seconds())
}
// RecordTokens records LLM token usage from a completion's usage block. Zero
// counts are skipped so a provider that omits usage adds nothing.
func RecordTokens(model string, prompt, completion int) {
if prompt > 0 {
llmTokens.WithLabelValues(model, "prompt").Add(float64(prompt))
}
if completion > 0 {
llmTokens.WithLabelValues(model, "completion").Add(float64(completion))
}
}
// HTTPMiddleware records request count + latency. It reads r.Pattern AFTER the
// inner handler routes (Go 1.22 sets it during ServeMux matching), so the label is
// the bounded registered pattern (e.g. "GET /v/{videoId}"), never the raw path
// with its high-cardinality ids. Unmatched requests bucket as "other".
func HTTPMiddleware(next http.Handler) http.Handler {
return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
start := time.Now()
sw := &statusWriter{ResponseWriter: w, code: http.StatusOK}
next.ServeHTTP(sw, r)
route := r.Pattern
if route == "" {
route = "other"
}
httpRequests.WithLabelValues(r.Method, route, strconv.Itoa(sw.code)).Inc()
httpDuration.WithLabelValues(r.Method, route).Observe(time.Since(start).Seconds())
})
}
// statusWriter captures the response status for the request-count label.
type statusWriter struct {
http.ResponseWriter
code int
wroteHeader bool
}
func (s *statusWriter) WriteHeader(code int) {
if !s.wroteHeader {
s.code = code
s.wroteHeader = true
}
s.ResponseWriter.WriteHeader(code)
}
func (s *statusWriter) Write(b []byte) (int, error) {
s.wroteHeader = true // an implicit 200
return s.ResponseWriter.Write(b)
}
+92
View File
@@ -0,0 +1,92 @@
package metrics
import (
"net/http"
"net/http/httptest"
"testing"
"time"
"github.com/prometheus/client_golang/prometheus"
"github.com/prometheus/client_golang/prometheus/testutil"
dto "github.com/prometheus/client_model/go"
"github.com/stretchr/testify/require"
)
// histCount reads a histogram child's observation count (testutil.ToFloat64 only
// works on counters/gauges; a histogram's WithLabelValues child is an Observer).
func histCount(t *testing.T, o prometheus.Observer) uint64 {
t.Helper()
m, ok := o.(prometheus.Metric)
require.True(t, ok, "histogram child must be a prometheus.Metric")
var d dto.Metric
require.NoError(t, m.Write(&d))
return d.GetHistogram().GetSampleCount()
}
// TestObserveSummarizeRecordsModelOutcomeFallback: a success observation lands on
// the right model/outcome/fallback series.
func TestObserveSummarizeRecordsModelOutcomeFallback(t *testing.T) {
before := histCount(t, summarizeDuration.WithLabelValues("koala/phi4-mini", "success", "false"))
ObserveSummarize("koala/phi4-mini", "success", false, 1200*time.Millisecond)
after := histCount(t, summarizeDuration.WithLabelValues("koala/phi4-mini", "success", "false"))
require.Equal(t, before+1, after, "one success observation recorded for the model")
}
// TestObserveSummarizeRecordsFailureOutcomes: error and parse_error are distinct
// series so a fallback chain's failures are visible.
func TestObserveSummarizeRecordsFailureOutcomes(t *testing.T) {
e0 := histCount(t, summarizeDuration.WithLabelValues("m", "error", "false"))
p0 := histCount(t, summarizeDuration.WithLabelValues("m", "parse_error", "false"))
ObserveSummarize("m", "error", false, time.Second)
ObserveSummarize("m", "parse_error", false, time.Second)
require.Equal(t, e0+1, histCount(t, summarizeDuration.WithLabelValues("m", "error", "false")))
require.Equal(t, p0+1, histCount(t, summarizeDuration.WithLabelValues("m", "parse_error", "false")))
}
func TestObserveCaptionFetchByOutcome(t *testing.T) {
b := histCount(t, captionFetchDuration.WithLabelValues("captions"))
ObserveCaptionFetch("captions", 3*time.Second)
require.Equal(t, b+1, histCount(t, captionFetchDuration.WithLabelValues("captions")))
}
func TestChatAnswerLatencyRecorded(t *testing.T) {
b := histCount(t, chatDuration.WithLabelValues("iguana/gemma4-26b"))
ObserveChat("iguana/gemma4-26b", 2*time.Second)
require.Equal(t, b+1, histCount(t, chatDuration.WithLabelValues("iguana/gemma4-26b")))
}
// TestRecordTokens: prompt + completion land on their kind series; zero is skipped.
func TestRecordTokens(t *testing.T) {
p0 := testutil.ToFloat64(llmTokens.WithLabelValues("m", "prompt"))
c0 := testutil.ToFloat64(llmTokens.WithLabelValues("m", "completion"))
RecordTokens("m", 100, 40)
RecordTokens("m", 0, 0) // skipped, no panic
require.Equal(t, p0+100, testutil.ToFloat64(llmTokens.WithLabelValues("m", "prompt")))
require.Equal(t, c0+40, testutil.ToFloat64(llmTokens.WithLabelValues("m", "completion")))
}
func TestLoginCounted(t *testing.T) {
b := testutil.ToFloat64(logins)
IncLogin()
require.Equal(t, b+1, testutil.ToFloat64(logins))
}
// TestHTTPMiddlewareRecordsByRoutePattern: the request is counted under the bounded
// registered pattern (r.Pattern after routing), not the raw path with its ids.
func TestHTTPMiddlewareRecordsByRoutePattern(t *testing.T) {
mux := http.NewServeMux()
mux.HandleFunc("GET /v/{videoId}", func(w http.ResponseWriter, _ *http.Request) {
w.WriteHeader(http.StatusTeapot)
})
h := HTTPMiddleware(mux)
before := testutil.ToFloat64(httpRequests.WithLabelValues("GET", "GET /v/{videoId}", "418"))
rec := httptest.NewRecorder()
h.ServeHTTP(rec, httptest.NewRequest(http.MethodGet, "/v/abc-123", nil))
require.Equal(t, http.StatusTeapot, rec.Code)
after := testutil.ToFloat64(httpRequests.WithLabelValues("GET", "GET /v/{videoId}", "418"))
require.Equal(t, before+1, after, "counted under the pattern, not /v/abc-123")
require.Equal(t, float64(0), testutil.ToFloat64(httpRequests.WithLabelValues("GET", "/v/abc-123", "418")),
"raw path must never be a label value")
}
+21 -1
View File
@@ -7,7 +7,7 @@ package ports
import (
"context"
"gitea.d-ma.be/mathias/tapir/internal/domain"
"git.d-ma.be/mathias/tapir/internal/domain"
)
// VideoSource is a video platform Tapir watches (YouTube, Vimeo).
@@ -27,6 +27,26 @@ type Summarizer interface {
Summarize(ctx context.Context, v domain.Video, t domain.Transcript) (domain.Summary, error)
}
// TranscriptStore persists transcripts as shared, video-keyed public content
// (ADR-021). It is keyed by the cross-user dedup key (provider, providerVideoID)
// — the video's public identity, NOT Tapir's per-user videos.id — and holds only
// public caption content, so it is deliberately NOT user-scoped: two users who
// share a video share the one row. The engine reads it before any caption fetch
// so re-analysis never re-touches YouTube (ADR-010/014).
type TranscriptStore interface {
// GetTranscript returns the stored transcript for a video and whether one
// exists. A stored Source == SourceNone (captions permanently absent) is a
// real hit: ok is true and HasText() is false, so callers skip without
// re-fetching. A transient rate-limit is never stored, so it never appears
// here as a false absence.
GetTranscript(ctx context.Context, provider, providerVideoID string) (t domain.Transcript, ok bool, err error)
// SaveTranscript upserts the transcript for (provider, providerVideoID). Only
// terminal outcomes are persisted: SourceCaptions (with text) or SourceNone.
// SourceRateLimited must NOT be passed — it is a per-user retry (ADR-014), not
// a shared terminal state.
SaveTranscript(ctx context.Context, provider, providerVideoID string, t domain.Transcript) error
}
// Sink delivers a summary to a destination (user store, brain, ...).
// Implementations fail independently of one another.
type Sink interface {
+66 -15
View File
@@ -19,16 +19,17 @@ import (
"slices"
"time"
"gitea.d-ma.be/mathias/tapir/internal/domain"
"gitea.d-ma.be/mathias/tapir/internal/ports"
"gitea.d-ma.be/mathias/tapir/internal/usecase"
"git.d-ma.be/mathias/tapir/internal/domain"
"git.d-ma.be/mathias/tapir/internal/ports"
"git.d-ma.be/mathias/tapir/internal/usecase"
)
// passCandidate is a video that passed all pre-filters (seen/manual/backoff)
// and is queued for transcript fetch + summarization in this pass.
type passCandidate struct {
v domain.Video
pos int // discovery position — used as a stable tiebreak when published_at ties
v domain.Video
channelID string // owning channel — keys the caption-availability memory (ADR-024)
pos int // discovery position — used as a stable tiebreak when published_at ties
}
// compareNewestFirst orders candidates by published_at descending, NULLS LAST,
@@ -74,6 +75,14 @@ type VideoStore interface {
// Called when NewVideos returns domain.ErrChannelUnavailable; best-effort, errors
// are logged and never abort the pass.
UpsertChannelError(ctx context.Context, userID, channelID, channelTitle string) error
// CaptionlessChannels returns channel ids currently suppressed because their
// recent videos all yielded no captions (ADR-024). The loop skips caption
// fetches for these channels' (non-requested) videos.
CaptionlessChannels(ctx context.Context, userID string) (map[string]bool, error)
// RecordChannelCaptionOutcome updates a channel's caption memory after a fetch:
// hadCaptions resets it, otherwise the no-caption streak grows and the channel
// is suppressed for window once it reaches threshold. A no-op when threshold<=0.
RecordChannelCaptionOutcome(ctx context.Context, userID, channelID string, hadCaptions bool, threshold int, window time.Duration) error
}
// Processor runs the core use case for a single video. *usecase.Engine
@@ -93,6 +102,9 @@ type Runner struct {
backoff time.Duration // rate-limit retry window; 0 = always retry
autoWindow time.Duration // recency bound for auto-summarize; 0 = no bound
now func() time.Time // injectable clock (tests); defaults to time.Now
captionThreshold int // consecutive no-caption results before a channel is suppressed; 0 = feature off
captionWindow time.Duration // how long a caption-less channel stays suppressed before re-probe
}
// Option configures a Runner at construction. Variadic so existing call sites
@@ -114,6 +126,14 @@ func WithClock(now func() time.Time) Option { return func(r *Runner) { r.now = n
// bypasses the bound. 0 (the default) disables it (summarize every unseen video).
func WithAutoWindow(d time.Duration) Option { return func(r *Runner) { r.autoWindow = d } }
// WithCaptionMemory enables per-channel caption-availability suppression
// (ADR-024): after threshold consecutive no-caption results a channel's videos
// are skipped (no caption fetch) for window, then one is re-probed. threshold<=0
// (the default) disables the feature entirely.
func WithCaptionMemory(threshold int, window time.Duration) Option {
return func(r *Runner) { r.captionThreshold = threshold; r.captionWindow = window }
}
// New builds a Runner. A nil logger falls back to slog.Default.
func New(src ports.VideoSource, store VideoStore, engine Processor, userID string, log *slog.Logger, opts ...Option) *Runner {
if log == nil {
@@ -131,15 +151,16 @@ func New(src ports.VideoSource, store VideoStore, engine Processor, userID strin
// Stats summarizes one RunOnce pass.
type Stats struct {
Candidates int
Summarized int
SkippedSeen int
SkippedNoText int
SkippedManual int // discovered but not queued, in manual mode
SkippedTooOld int // auto mode: published outside the recency window (not requested)
SkippedRateLimited int // 429'd previously and still inside the backoff window
Errors int
ChannelUnavailable int // channels that returned HTTP 404 (deleted/private)
Candidates int
Summarized int
SkippedSeen int
SkippedNoText int
SkippedManual int // discovered but not queued, in manual mode
SkippedTooOld int // auto mode: published outside the recency window (not requested)
SkippedRateLimited int // 429'd previously and still inside the backoff window
SkippedNoCaptionChannel int // channel suppressed as caption-less (ADR-024)
Errors int
ChannelUnavailable int // channels that returned HTTP 404 (deleted/private)
}
// tooOld reports whether a video published at publishedAt falls outside the
@@ -213,6 +234,17 @@ func (r *Runner) RunOnce(ctx context.Context) (Stats, error) {
}
}
// Per-channel caption memory (ADR-024): channels whose recent videos all
// yielded no captions are suppressed so their new videos don't burn the scarce
// fetch budget. Loaded only when the feature is enabled (threshold > 0).
var captionless map[string]bool
if r.captionThreshold > 0 {
captionless, err = r.store.CaptionlessChannels(ctx, r.userID)
if err != nil {
return stats, fmt.Errorf("runner: load caption-less channels: %w", err)
}
}
subs, err := r.src.ListSubscriptions(ctx, r.userID)
if err != nil {
return stats, fmt.Errorf("runner: list subscriptions: %w", err)
@@ -273,6 +305,14 @@ func (r *Runner) RunOnce(ctx context.Context) (Stats, error) {
continue
}
// Caption-less channel (ADR-024): its recent videos all returned no
// captions, so skip the fetch entirely. The video is still listed
// (UpsertVideo above); an explicit manual request bypasses the skip.
if !requested[id] && captionless[sub.ChannelID] {
stats.SkippedNoCaptionChannel++
continue
}
// Still inside the rate-limit backoff window: skip without fetching.
if at, ok := rateLimited[id]; ok && r.now().Sub(at) < r.backoff {
stats.SkippedRateLimited++
@@ -280,7 +320,7 @@ func (r *Runner) RunOnce(ctx context.Context) (Stats, error) {
continue
}
candidates = append(candidates, passCandidate{v: v, pos: pos})
candidates = append(candidates, passCandidate{v: v, channelID: sub.ChannelID, pos: pos})
pos++
}
}
@@ -320,6 +360,11 @@ func (r *Runner) RunOnce(ctx context.Context) (Stats, error) {
errs = append(errs, fmt.Errorf("set none status %q: %w", c.v.ProviderVideoID, err))
stats.Errors++
}
// No captions: grow this channel's no-caption streak (ADR-024).
if err := r.store.RecordChannelCaptionOutcome(ctx, r.userID, c.channelID, false, r.captionThreshold, r.captionWindow); err != nil {
errs = append(errs, fmt.Errorf("record no-caption %q: %w", c.v.ProviderVideoID, err))
stats.Errors++
}
r.log.Info("skipped video (no transcript)", "video", c.v.ProviderVideoID, "title", c.v.Title)
case res.Summary != nil:
stats.Summarized++
@@ -327,6 +372,11 @@ func (r *Runner) RunOnce(ctx context.Context) (Stats, error) {
errs = append(errs, fmt.Errorf("set fetched status %q: %w", c.v.ProviderVideoID, err))
stats.Errors++
}
// Captions present: reset this channel's caption memory (ADR-024).
if err := r.store.RecordChannelCaptionOutcome(ctx, r.userID, c.channelID, true, r.captionThreshold, r.captionWindow); err != nil {
errs = append(errs, fmt.Errorf("record has-caption %q: %w", c.v.ProviderVideoID, err))
stats.Errors++
}
// In manual mode the video was explicitly queued; clear the flag so
// it is not re-summarized and the UI drops the "Queued" chip.
if !auto {
@@ -354,6 +404,7 @@ func (r *Runner) Loop(ctx context.Context, interval time.Duration) error {
"skipped_seen", stats.SkippedSeen, "skipped_no_text", stats.SkippedNoText,
"skipped_manual", stats.SkippedManual, "skipped_too_old", stats.SkippedTooOld,
"skipped_rate_limited", stats.SkippedRateLimited,
"skipped_no_caption_channel", stats.SkippedNoCaptionChannel,
"channel_unavailable", stats.ChannelUnavailable, "errors", stats.Errors)
if err != nil {
r.log.Warn("run pass had errors", "err", err)
+75 -3
View File
@@ -9,9 +9,9 @@ import (
"github.com/stretchr/testify/require"
"gitea.d-ma.be/mathias/tapir/internal/domain"
"gitea.d-ma.be/mathias/tapir/internal/runner"
"gitea.d-ma.be/mathias/tapir/internal/usecase"
"git.d-ma.be/mathias/tapir/internal/domain"
"git.d-ma.be/mathias/tapir/internal/runner"
"git.d-ma.be/mathias/tapir/internal/usecase"
)
const testUser = "11111111-1111-1111-1111-111111111111"
@@ -51,6 +51,13 @@ type fakeStore struct {
cleared []string
rateLimited map[string]time.Time // id -> when 429'd (seeds the backoff window)
statuses map[string]string // id -> last SetTranscriptStatus value
captionless map[string]bool // channel ids currently suppressed (ADR-024)
captionRecs []captionRec // RecordChannelCaptionOutcome calls, in order
}
type captionRec struct {
channelID string
had bool
}
func (f *fakeStore) UpsertVideo(_ context.Context, v domain.Video) (string, error) {
@@ -93,6 +100,22 @@ func (f *fakeStore) RateLimitedVideoIDs(_ context.Context, _ string) (map[string
func (f *fakeStore) UpsertChannelError(_ context.Context, _, _, _ string) error { return nil }
func (f *fakeStore) CaptionlessChannels(_ context.Context, _ string) (map[string]bool, error) {
cp := make(map[string]bool, len(f.captionless))
for k, v := range f.captionless {
cp[k] = v
}
return cp, nil
}
func (f *fakeStore) RecordChannelCaptionOutcome(_ context.Context, _, channelID string, hadCaptions bool, threshold int, _ time.Duration) error {
if threshold <= 0 {
return nil
}
f.captionRecs = append(f.captionRecs, captionRec{channelID: channelID, had: hadCaptions})
return nil
}
func (f *fakeStore) SetTranscriptStatus(_ context.Context, _, videoID, status string) error {
if f.statuses == nil {
f.statuses = map[string]string{}
@@ -284,6 +307,55 @@ func TestRunOnce_AutoMode_OldVideoRequestedBypassesWindow(t *testing.T) {
require.Len(t, sink.delivered, 1)
}
// TestRunOnce_CaptionlessChannelSkipped: a channel flagged caption-less (ADR-024)
// has its videos skipped from fetching but still discovered/listed, while a
// normal channel's video is summarized.
func TestRunOnce_CaptionlessChannelSkipped(t *testing.T) {
src := &fakeSource{
subs: []domain.Subscription{sub("dead", "Dead Channel"), sub("live", "Live Channel")},
videos: map[string][]domain.Video{
"dead": {vid("d1", "Dead One")},
"live": {vid("l1", "Live One")},
},
}
st := &fakeStore{seen: map[string]bool{}, auto: true, captionless: map[string]bool{"dead": true}}
sink := &recordingSink{}
eng := usecase.NewEngine(src, fakeSummarizer{}, sink)
r := runner.New(src, st, eng, testUser, quietLogger(),
runner.WithCaptionMemory(5, 14*24*time.Hour))
stats, err := r.RunOnce(context.Background())
require.NoError(t, err)
require.Equal(t, 1, stats.SkippedNoCaptionChannel, "dead channel's video skipped from fetch")
require.Equal(t, 1, stats.Summarized, "live channel's video still summarized")
require.Len(t, st.upserted, 2, "both videos are still discovered and listed")
}
// TestRunOnce_RecordsCaptionOutcomes: a no-caption result grows the channel's
// streak (had=false); a successful summary resets it (had=true).
func TestRunOnce_RecordsCaptionOutcomes(t *testing.T) {
src := &fakeSource{
subs: []domain.Subscription{sub("c1", "Has Caps"), sub("c2", "No Caps")},
videos: map[string][]domain.Video{
"c1": {vid("good", "Good")},
"c2": {vid("bad", "Bad")},
},
transcripts: map[string]domain.Transcript{
"bad": {Source: domain.SourceNone}, // no usable text → engine skips
},
}
st := &fakeStore{seen: map[string]bool{}, auto: true}
sink := &recordingSink{}
eng := usecase.NewEngine(src, fakeSummarizer{}, sink)
r := runner.New(src, st, eng, testUser, quietLogger(),
runner.WithCaptionMemory(5, 14*24*time.Hour))
_, err := r.RunOnce(context.Background())
require.NoError(t, err)
require.Contains(t, st.captionRecs, captionRec{channelID: "c1", had: true}, "captioned channel reset")
require.Contains(t, st.captionRecs, captionRec{channelID: "c2", had: false}, "no-caption channel streak grown")
}
// TestRunOnce_AutoWindowZero_SummarizesOld: a zero window disables the bound —
// the pre-recency behaviour (summarize every unseen video) is preserved.
func TestRunOnce_AutoWindowZero_SummarizesOld(t *testing.T) {
+44 -4
View File
@@ -14,8 +14,8 @@ import (
"fmt"
"sync"
"gitea.d-ma.be/mathias/tapir/internal/domain"
"gitea.d-ma.be/mathias/tapir/internal/ports"
"git.d-ma.be/mathias/tapir/internal/domain"
"git.d-ma.be/mathias/tapir/internal/ports"
)
// ErrNotImplemented marks scaffold methods awaiting implementation.
@@ -27,6 +27,13 @@ type Engine struct {
AI ports.Summarizer
Sinks []ports.Sink
// Transcripts, when set, is the shared transcript cache (ADR-021): the engine
// reads it before any caption fetch and writes resolved transcripts back, so
// re-analysis — the same user re-summarizing, or a second user with the same
// video — never re-touches YouTube (ADR-010/014). Optional: nil disables
// persistence (fetch every time), keeping the pure-core/scaffold wiring valid.
Transcripts ports.TranscriptStore
// processed dedups videos within this engine's lifetime so a video is not
// summarized twice when the watcher sees it again. Durable cross-restart
// dedup is the store's concern (a resolved TRANSCRIPT / existing SUMMARY,
@@ -57,9 +64,9 @@ type ProcessResult struct {
// resolve transcript -> (summarize -> deliver) | skip.
// See docs/use-cases/summarize_new_video.feature.
func (e *Engine) ProcessNewVideo(ctx context.Context, v domain.Video) (ProcessResult, error) {
t, err := e.Source.FetchTranscript(ctx, v)
t, err := e.resolveTranscript(ctx, v)
if err != nil {
return ProcessResult{Video: v}, fmt.Errorf("fetch transcript: %w", err)
return ProcessResult{Video: v}, err
}
if !t.HasText() {
// No usable transcript: record the skip, produce no summary, deliver nothing
@@ -86,6 +93,39 @@ func (e *Engine) ProcessNewVideo(ctx context.Context, v domain.Video) (ProcessRe
return ProcessResult{Video: v, Summary: &sum, TranscriptSource: string(t.Source)}, errors.Join(errs...)
}
// resolveTranscript returns v's transcript, reading the shared store first
// (ADR-021): a stored transcript — including a stored SourceNone (captions
// permanently absent) — is returned without touching YouTube, so re-analysis
// never re-fetches. On a store miss it fetches through the source (which gates
// the caption call, ADR-014) and persists the terminal outcome so the next
// analysis, for any user, reads from the store. A transient SourceRateLimited is
// returned to the caller (the runner stamps a per-user backoff) but never stored,
// so persistence can never mask a 429 as a permanent "no transcript". When no
// TranscriptStore is wired the engine simply fetches every time.
func (e *Engine) resolveTranscript(ctx context.Context, v domain.Video) (domain.Transcript, error) {
if e.Transcripts != nil {
stored, ok, err := e.Transcripts.GetTranscript(ctx, string(v.Provider), v.ProviderVideoID)
if err != nil {
return domain.Transcript{}, fmt.Errorf("get stored transcript: %w", err)
}
if ok {
return stored, nil
}
}
t, err := e.Source.FetchTranscript(ctx, v)
if err != nil {
return domain.Transcript{}, fmt.Errorf("fetch transcript: %w", err)
}
if e.Transcripts != nil && t.Source != domain.SourceRateLimited {
if err := e.Transcripts.SaveTranscript(ctx, string(v.Provider), v.ProviderVideoID, t); err != nil {
return domain.Transcript{}, fmt.Errorf("save transcript: %w", err)
}
}
return t, nil
}
// ProcessNewVideos walks a user's subscriptions and processes each newly seen
// video. Only videos surfaced via the user's subscriptions are considered, so a
// channel the user is not subscribed to is never processed. A video already
+187
View File
@@ -0,0 +1,187 @@
package usecase
import (
"context"
"testing"
"git.d-ma.be/mathias/tapir/internal/domain"
)
// These tests pin the ADR-021 read-stored-first behaviour at the engine core:
// a stored transcript is summarized without re-touching the source, a miss
// fetches once and persists, and a transient rate-limit is never cached.
type recordingSource struct {
transcript domain.Transcript
fetchCalls int
}
func (s *recordingSource) ListSubscriptions(context.Context, string) ([]domain.Subscription, error) {
return nil, nil
}
func (s *recordingSource) NewVideos(context.Context, domain.Subscription) ([]domain.Video, error) {
return nil, nil
}
func (s *recordingSource) FetchTranscript(context.Context, domain.Video) (domain.Transcript, error) {
s.fetchCalls++
return s.transcript, nil
}
type fakeTranscriptStore struct {
stored map[string]domain.Transcript
saves int
}
func newFakeTranscriptStore() *fakeTranscriptStore {
return &fakeTranscriptStore{stored: make(map[string]domain.Transcript)}
}
func (f *fakeTranscriptStore) key(provider, id string) string { return provider + "|" + id }
func (f *fakeTranscriptStore) GetTranscript(_ context.Context, provider, id string) (domain.Transcript, bool, error) {
t, ok := f.stored[f.key(provider, id)]
return t, ok, nil
}
func (f *fakeTranscriptStore) SaveTranscript(_ context.Context, provider, id string, t domain.Transcript) error {
f.saves++
f.stored[f.key(provider, id)] = t
return nil
}
type countingSummarizer struct{ calls int }
func (c *countingSummarizer) Summarize(_ context.Context, v domain.Video, _ domain.Transcript) (domain.Summary, error) {
c.calls++
return domain.Summary{VideoID: v.ID, UserID: v.UserID, Summary: "s", AIProvider: "local"}, nil
}
type nopSink struct{}
func (nopSink) Name() string { return "nop" }
func (nopSink) Deliver(context.Context, domain.Summary) error { return nil }
func testVideo() domain.Video {
return domain.Video{ID: "v1", UserID: "u1", Provider: domain.ProviderYouTube, ProviderVideoID: "yt1"}
}
func TestProcessNewVideo_StoredTranscriptSkipsFetch(t *testing.T) {
src := &recordingSource{}
ts := newFakeTranscriptStore()
ts.stored[ts.key("youtube", "yt1")] = domain.Transcript{Source: domain.SourceCaptions, Content: "stored words"}
sum := &countingSummarizer{}
eng := NewEngine(src, sum, nopSink{})
eng.Transcripts = ts
res, err := eng.ProcessNewVideo(context.Background(), testVideo())
if err != nil {
t.Fatalf("ProcessNewVideo: %v", err)
}
if src.fetchCalls != 0 {
t.Fatalf("stored transcript must not re-fetch from source; got %d fetches", src.fetchCalls)
}
if ts.saves != 0 {
t.Fatalf("a store hit must not re-save; got %d saves", ts.saves)
}
if sum.calls != 1 || res.Summary == nil {
t.Fatalf("expected a summary from the stored transcript; calls=%d summary=%v", sum.calls, res.Summary)
}
}
func TestProcessNewVideo_StoreMissFetchesAndPersists(t *testing.T) {
src := &recordingSource{transcript: domain.Transcript{Source: domain.SourceCaptions, Language: "en", Content: "fetched words"}}
ts := newFakeTranscriptStore()
sum := &countingSummarizer{}
eng := NewEngine(src, sum, nopSink{})
eng.Transcripts = ts
if _, err := eng.ProcessNewVideo(context.Background(), testVideo()); err != nil {
t.Fatalf("ProcessNewVideo: %v", err)
}
if src.fetchCalls != 1 {
t.Fatalf("a store miss must fetch exactly once; got %d", src.fetchCalls)
}
if ts.saves != 1 {
t.Fatalf("a fetched transcript must be persisted; got %d saves", ts.saves)
}
got, ok, _ := ts.GetTranscript(context.Background(), "youtube", "yt1")
if !ok || got.Content != "fetched words" {
t.Fatalf("persisted transcript not readable back: ok=%v content=%q", ok, got.Content)
}
}
// The second summarize of the same video reads the persisted transcript and does
// NOT re-fetch — the primary ADR-021 win, proven end to end at the engine.
func TestProcessNewVideo_SecondSummarizeDoesNotRefetch(t *testing.T) {
src := &recordingSource{transcript: domain.Transcript{Source: domain.SourceCaptions, Content: "words"}}
ts := newFakeTranscriptStore()
eng := NewEngine(src, &countingSummarizer{}, nopSink{})
eng.Transcripts = ts
for i := 0; i < 2; i++ {
if _, err := eng.ProcessNewVideo(context.Background(), testVideo()); err != nil {
t.Fatalf("pass %d: %v", i, err)
}
}
if src.fetchCalls != 1 {
t.Fatalf("the second summarize must reuse the stored transcript; got %d fetches", src.fetchCalls)
}
}
// A stored "no captions" outcome short-circuits before both fetch and summarize.
func TestProcessNewVideo_StoredNoneSkipsFetchAndSummarize(t *testing.T) {
src := &recordingSource{}
ts := newFakeTranscriptStore()
ts.stored[ts.key("youtube", "yt1")] = domain.Transcript{Source: domain.SourceNone}
sum := &countingSummarizer{}
eng := NewEngine(src, sum, nopSink{})
eng.Transcripts = ts
res, err := eng.ProcessNewVideo(context.Background(), testVideo())
if err != nil {
t.Fatalf("ProcessNewVideo: %v", err)
}
if !res.Skipped {
t.Fatal("a stored SourceNone must skip")
}
if src.fetchCalls != 0 || sum.calls != 0 {
t.Fatalf("stored none must neither fetch nor summarize; fetches=%d calls=%d", src.fetchCalls, sum.calls)
}
}
// A transient 429 is surfaced (so the runner backs off per-user) but never cached
// as a shared terminal state — otherwise it would mask a rate-limit as permanent.
func TestProcessNewVideo_RateLimitedIsNotPersisted(t *testing.T) {
src := &recordingSource{transcript: domain.Transcript{Source: domain.SourceRateLimited}}
ts := newFakeTranscriptStore()
eng := NewEngine(src, &countingSummarizer{}, nopSink{})
eng.Transcripts = ts
res, err := eng.ProcessNewVideo(context.Background(), testVideo())
if err != nil {
t.Fatalf("ProcessNewVideo: %v", err)
}
if !res.Skipped || res.TranscriptSource != string(domain.SourceRateLimited) {
t.Fatalf("expected a rate-limited skip; skipped=%v source=%q", res.Skipped, res.TranscriptSource)
}
if ts.saves != 0 {
t.Fatalf("a transient rate-limit must not be persisted; got %d saves", ts.saves)
}
}
// With no TranscriptStore wired the engine fetches every time (back-compat).
func TestProcessNewVideo_NilStoreFetchesEveryTime(t *testing.T) {
src := &recordingSource{transcript: domain.Transcript{Source: domain.SourceCaptions, Content: "words"}}
eng := NewEngine(src, &countingSummarizer{}, nopSink{})
for i := 0; i < 2; i++ {
if _, err := eng.ProcessNewVideo(context.Background(), testVideo()); err != nil {
t.Fatalf("pass %d: %v", i, err)
}
}
if src.fetchCalls != 2 {
t.Fatalf("nil store must fetch every time; got %d", src.fetchCalls)
}
}
+1 -1
View File
@@ -3,7 +3,7 @@ package web
import (
"net/http"
"gitea.d-ma.be/mathias/tapir/internal/adapters/store"
"git.d-ma.be/mathias/tapir/internal/adapters/store"
)
// handleAccount renders the account page: the user's display name, the
+2 -2
View File
@@ -9,8 +9,8 @@ import (
"github.com/stretchr/testify/require"
"gitea.d-ma.be/mathias/tapir/internal/adapters/store"
"gitea.d-ma.be/mathias/tapir/internal/web"
"git.d-ma.be/mathias/tapir/internal/adapters/store"
"git.d-ma.be/mathias/tapir/internal/web"
)
// fakeSecrets is a SecretRemover that records the refs it was asked to delete, so
+227
View File
@@ -0,0 +1,227 @@
package web
import (
"context"
"errors"
"net/http"
"strings"
"git.d-ma.be/mathias/tapir/internal/adapters/chat"
"git.d-ma.be/mathias/tapir/internal/adapters/store"
)
// Chatter is the per-video chat backend (ADR-027). *chat.Service satisfies it;
// tests substitute a fake. It carries NO caption-fetch dependency — the chat
// handlers reach it only after reading an already-stored transcript, so an
// enabled chat cannot trigger a fetch, touch the rate gate, or reach YouTube.
type Chatter interface {
// Models returns the offerable models, local-first (cloud absent when disabled).
Models() []string
// DefaultModel resolves the model a fresh chat opens with given the summary's
// model (the ADR-027 default), falling back to the first offered model.
DefaultModel(summaryModel string) string
// Answer runs one chat turn against the supplied stored transcript text.
Answer(ctx context.Context, req chat.Request) (chat.Reply, error)
}
// maxHistoryTurns bounds the ephemeral conversation carried per request, so a long
// back-and-forth cannot grow the prompt without limit (the transcript already
// dominates the budget). Older turns drop off the front.
const maxHistoryTurns = 8
// chatView is everything the chat templates render: the video identity for links
// and titles, the model switcher state, the running (ephemeral) conversation, and
// the honest flags — Available is false when no usable transcript is stored
// (ADR-027: honest "not available", never a fetch), Truncated when the transcript
// was bounded to fit the model, Error for a transient model failure.
type chatView struct {
VideoID string
Title string
Available bool
Models []string
Selected string
History []chat.Turn
Truncated bool
Error string
}
// handleChat renders the chat page for a summarized video (GET). Entry is scoped
// through GetSummaryByVideo, which is RLS/user-scoped: a video that is not the
// requesting user's own resolves to ErrNotFound → 404, so chat is reachable only
// from the user's own summary view (ADR-027 isolation). The transcript is read
// from the SHARED store (ADR-021) — a pure DB read, never a caption fetch; an
// absent/text-less transcript renders the honest "not available" state, no fetch.
func (a *App) handleChat(w http.ResponseWriter, r *http.Request) {
userID, ok := a.currentUserID(w, r)
if !ok {
return
}
row, ok := a.loadOwnedSummary(w, r, userID)
if !ok {
return
}
_, hasText, ok := a.readTranscript(w, r, *row)
if !ok {
return
}
view := chatView{
VideoID: row.VideoID,
Title: displayTitle(*row),
Available: hasText,
Models: a.Chat.Models(),
Selected: a.Chat.DefaultModel(row.AIModel),
}
// HTMX (the in-place reveal from the summary) gets just the open chat section,
// swapped over the closed dock so the summary above it stays put. A no-JS
// navigation gets the full page: the whole summary plus the open chat.
if isHTMX(r) {
a.render(w, r, chatSection(view))
return
}
a.render(w, r, ChatPage(*row, view))
}
// handleChatMessage answers one question against the stored transcript (POST).
// It reads the transcript from the store (no fetch), runs the chosen model over
// it plus the prior turns, appends the answer, and returns the refreshed chat
// panel (HTMX) or the whole page (no-JS). A model failure is surfaced inline,
// not as a 500 — the conversation and the question are preserved for a retry.
func (a *App) handleChatMessage(w http.ResponseWriter, r *http.Request) {
userID, ok := a.currentUserID(w, r)
if !ok {
return
}
row, ok := a.loadOwnedSummary(w, r, userID)
if !ok {
return
}
if err := r.ParseForm(); err != nil {
http.Error(w, "bad form", http.StatusBadRequest)
return
}
transcript, hasText, ok := a.readTranscript(w, r, *row)
if !ok {
return
}
model := a.resolveModel(r.FormValue("model"), row.AIModel)
history := parseHistory(r.Form["hq"], r.Form["ha"])
question := strings.TrimSpace(r.FormValue("question"))
view := chatView{
VideoID: row.VideoID,
Title: displayTitle(*row),
Available: hasText,
Models: a.Chat.Models(),
Selected: model,
History: history,
}
switch {
case !hasText:
// Honest "not available" — no fetch, no model call (ADR-027).
case question == "":
// A model switch with no text just re-renders — no wasted round-trip.
default:
reply, err := a.Chat.Answer(r.Context(), chat.Request{
Model: model,
Transcript: transcript,
History: history,
Question: question,
})
if err != nil {
a.logger().Error("chat answer", "video", row.VideoID, "model", model, "err", err)
view.Error = "That model couldn't answer just now. Try again, or switch models."
} else {
view.History = appendTurn(history, chat.Turn{Question: question, Answer: reply.Answer})
view.Truncated = reply.Truncated
}
}
a.renderChatTurn(w, r, *row, view)
}
// loadOwnedSummary fetches the summary for the path's video scoped to userID, or
// writes the right response (404 on not-found/not-owned, 500 on error) and reports
// false. It is the single isolation gate for both chat handlers.
func (a *App) loadOwnedSummary(w http.ResponseWriter, r *http.Request, userID string) (*store.SummaryRow, bool) {
videoID := r.PathValue("videoId")
row, err := a.Store.GetSummaryByVideo(r.Context(), userID, videoID)
if errors.Is(err, store.ErrNotFound) {
http.NotFound(w, r)
return nil, false
}
if err != nil {
a.serverError(w, r, "chat get summary", err)
return nil, false
}
return row, true
}
// readTranscript reads the shared stored transcript for a row (ADR-021) and
// reports whether it carries usable text. It is a pure DB read — NO caption fetch,
// the property the whole feature's safety rests on. ok is false only on a store
// error (after a 500 is written); a missing/text-less transcript is (",", false,
// true) — the honest "not available" case, handled by the caller, not an error.
func (a *App) readTranscript(w http.ResponseWriter, r *http.Request, row store.SummaryRow) (content string, hasText, ok bool) {
t, found, err := a.Store.GetTranscript(r.Context(), row.Channel, row.ProviderVideoID)
if err != nil {
a.serverError(w, r, "chat get transcript", err)
return "", false, false
}
if !found || !t.HasText() {
return "", false, true
}
return t.Content, true, true
}
// renderChatTurn returns the chat panel fragment for an HTMX answer (swapped in
// place within the open dock), or the full chat page otherwise — the no-JS POST
// re-renders the whole summary + open chat with the new turn.
func (a *App) renderChatTurn(w http.ResponseWriter, r *http.Request, row store.SummaryRow, v chatView) {
if isHTMX(r) {
a.render(w, r, chatPanel(v))
return
}
a.render(w, r, ChatPage(row, v))
}
// resolveModel keeps the posted model only when it is an offered option; anything
// else (a forged value, or a model dropped because cloud is disabled) falls back
// to the default. The chat.Service enforces the same guard before the gateway;
// this keeps the rendered switcher honest too.
func (a *App) resolveModel(posted, summaryModel string) string {
for _, m := range a.Chat.Models() {
if m == posted {
return posted
}
}
return a.Chat.DefaultModel(summaryModel)
}
// parseHistory zips the parallel hidden hq/ha fields back into ordered turns,
// keeping only the most recent maxHistoryTurns. net/url preserves the submission
// order of repeated fields, so the pairing is stable.
func parseHistory(qs, as []string) []chat.Turn {
n := len(qs)
if len(as) < n {
n = len(as)
}
turns := make([]chat.Turn, 0, n)
for i := 0; i < n; i++ {
turns = append(turns, chat.Turn{Question: qs[i], Answer: as[i]})
}
return capHistory(turns)
}
// appendTurn adds a completed exchange and re-bounds the conversation.
func appendTurn(history []chat.Turn, t chat.Turn) []chat.Turn {
return capHistory(append(history, t))
}
// capHistory keeps the last maxHistoryTurns turns (drops the oldest).
func capHistory(turns []chat.Turn) []chat.Turn {
if len(turns) <= maxHistoryTurns {
return turns
}
return turns[len(turns)-maxHistoryTurns:]
}
+375
View File
@@ -0,0 +1,375 @@
package web_test
import (
"context"
"net/http"
"net/http/httptest"
"net/url"
"strings"
"testing"
"time"
"github.com/jackc/pgx/v5/pgxpool"
"github.com/stretchr/testify/require"
"git.d-ma.be/mathias/tapir/internal/adapters/chat"
"git.d-ma.be/mathias/tapir/internal/adapters/store"
"git.d-ma.be/mathias/tapir/internal/domain"
"git.d-ma.be/mathias/tapir/internal/web"
)
// videoZ is a video id used by the isolation test for a DIFFERENT user's video.
const videoZ = "33333333-3333-3333-3333-333333333333"
// --- chat test doubles -----------------------------------------------------
// fakeChatter is a web.Chatter that records the request it received and returns a
// canned reply. It performs NO network and NO fetch — it stands in for the real
// chat.Service so the handler behaviour is what's under test.
type fakeChatter struct {
models []string
reply chat.Reply
err error
gotReq chat.Request
callCount int
}
func (f *fakeChatter) Models() []string { return f.models }
func (f *fakeChatter) DefaultModel(summaryModel string) string {
for _, m := range f.models {
if m == summaryModel {
return summaryModel
}
}
if len(f.models) > 0 {
return f.models[0]
}
return ""
}
func (f *fakeChatter) Answer(_ context.Context, req chat.Request) (chat.Reply, error) {
f.callCount++
f.gotReq = req
if f.err != nil {
return chat.Reply{}, f.err
}
return f.reply, nil
}
// tripwireProcessor and tripwireFetcher are the YouTube-reaching collaborators
// (summarize → caption fetch, and the paste metadata fetch). Wired into the App
// for the safety test, they fail it the instant chat routes into either — the
// behavioural proof that chat never triggers a fetch (ADR-027).
type tripwireProcessor struct{ t *testing.T }
func (p tripwireProcessor) ProcessVideo(context.Context, string, string) error {
p.t.Fatal("chat triggered summarization (→ caption fetch) — must never happen (ADR-027)")
return nil
}
type tripwireFetcher struct{ t *testing.T }
func (f tripwireFetcher) FetchVideo(context.Context, string, string) (domain.Video, error) {
f.t.Fatal("chat triggered a YouTube video fetch — must never happen (ADR-027)")
return domain.Video{}, nil
}
// --- helpers ---------------------------------------------------------------
// newChatApp builds the App under test as the registered stub user, with a Chat
// backend wired. Mirrors newApp but adds chat (and any extra wiring via mutate).
func newChatApp(t *testing.T, chatter web.Chatter, mutate func(*web.App)) *web.App {
t.Helper()
s := newStore(t)
app := &web.App{
Store: s,
Identity: s,
Auth: web.StubAuth{U: web.User{Subject: stubSubject}},
Chat: chatter,
}
if mutate != nil {
mutate(app)
}
return app
}
// seededProviderVideoID mirrors seedVideo's derivation so the chat path's
// (provider, providerVideoID) transcript key matches the seeded video row.
func seededProviderVideoID(videoID string) string { return "pv-" + videoID[:8] }
// seedTranscript stores a shared (provider, providerVideoID) transcript — the
// ADR-021 stored content the chat reads. Captions source = usable text.
func seedTranscript(t *testing.T, s *store.Store, videoID, content string) {
t.Helper()
err := s.SaveTranscript(context.Background(), "youtube", seededProviderVideoID(videoID),
domain.Transcript{Source: domain.SourceCaptions, Language: "en", Content: content})
require.NoError(t, err)
}
func getChat(t *testing.T, app *web.App, videoID string) *httptest.ResponseRecorder {
t.Helper()
return do(t, app, httptest.NewRequest(http.MethodGet, "/v/"+videoID+"/chat", nil))
}
func postChat(t *testing.T, app *web.App, videoID string, form url.Values, htmx bool) *httptest.ResponseRecorder {
t.Helper()
req := httptest.NewRequest(http.MethodPost, "/v/"+videoID+"/chat", strings.NewReader(form.Encode()))
req.Header.Set("Content-Type", "application/x-www-form-urlencoded")
if htmx {
req.Header.Set("HX-Request", "true")
}
return do(t, app, req)
}
// --- scenarios -------------------------------------------------------------
// The "dig deeper" affordance appears on a summary detail view only when chat is
// enabled, and links to that video's chat — the single entry point (ADR-027 §1).
func TestChatEntryAffordanceOnSummaryView(t *testing.T) {
ctx := context.Background()
withChat := newChatApp(t, &fakeChatter{models: []string{"phi4-mini"}}, nil)
p := rawPool(t)
resetDB(t, p)
require.NoError(t, deliver(ctx, withChat, videoX, "body x"))
seedVideo(t, p, videoX, "X Title", "https://x", time.Time{})
html := body(t, do(t, withChat, httptest.NewRequest(http.MethodGet, "/v/"+videoX, nil)))
require.Contains(t, html, "Dig deeper", "the deeper-dive affordance is shown when chat is enabled")
require.Contains(t, html, "/v/"+videoX+"/chat", "it links to this video's chat")
require.Contains(t, html, `id="chat-section"`, "the dock lives on the detail page")
require.Contains(t, html, `hx-get="/v/`+videoX+`/chat"`, "it opens the chat in place (HTMX), not a navigation")
// With no chat backend wired the affordance is absent (routes unmounted).
noChat := newApp(t)
require.NoError(t, deliver(ctx, noChat, videoX, "body x"))
html = body(t, do(t, noChat, httptest.NewRequest(http.MethodGet, "/v/"+videoX, nil)))
require.NotContains(t, html, "Dig deeper", "no affordance when chat is disabled")
}
// The summary and the chat live together (the integrated UX): the no-JS chat page
// renders the full summary alongside the chat, and the HTMX reveal returns just
// the open chat section as a fragment so it docks in below the summary already on
// screen — the summary is never navigated away from.
func TestChatIntegratedWithSummaryOnSamePage(t *testing.T) {
ctx := context.Background()
chatter := &fakeChatter{models: []string{"phi4-mini", "gemma4-26b"}}
app := newChatApp(t, chatter, nil)
p := rawPool(t)
resetDB(t, p)
require.NoError(t, deliver(ctx, app, videoX, "SUMMARY-BODY-MARKER"))
seedVideo(t, p, videoX, "X Title", "https://x", time.Time{})
seedTranscript(t, app.Store.(*store.Store), videoX, "the transcript")
// No-JS full page: the summary payload and the chat are on one page.
html := body(t, getChat(t, app, videoX))
require.Contains(t, html, "SUMMARY-BODY-MARKER", "the summary text is shown on the chat page")
require.Contains(t, html, "takeaway one", "takeaways shown alongside the chat")
require.Contains(t, html, "highlight one", "highlights shown alongside the chat")
require.Contains(t, html, "Ask about this video", "the chat sits on the same page as the summary")
require.Contains(t, html, `name="question"`, "the ask form is present")
// HTMX reveal: the open chat section ONLY (a fragment) — no full-page chrome and
// no duplicated summary, so it swaps in below the summary already rendered.
req := httptest.NewRequest(http.MethodGet, "/v/"+videoX+"/chat", nil)
req.Header.Set("HX-Request", "true")
frag := body(t, do(t, app, req))
require.NotContains(t, frag, "<html", "the reveal is a fragment, not a full page")
require.NotContains(t, frag, "SUMMARY-BODY-MARKER", "the reveal does not re-send the summary (it's already on screen)")
require.Contains(t, frag, `id="chat-section"`, "the fragment replaces the dock in place")
require.Contains(t, frag, `name="question"`, "the ask form is in the revealed section")
}
// THE KEY SAFETY ASSERTION (ADR-027): a chat answer is produced entirely from the
// stored transcript — the model receives the stored text, and neither the
// summarize→fetch path nor the YouTube fetch path is ever touched. The tripwire
// collaborators t.Fatal the test if chat reaches them.
func TestChatAnswersFromStoredTranscriptWithoutAnyFetch(t *testing.T) {
ctx := context.Background()
const transcript = "STORED-TRANSCRIPT-MARKER: the host explains attention budgets."
chatter := &fakeChatter{
models: []string{"phi4-mini"},
reply: chat.Reply{Answer: "It is about attention budgets."},
}
app := newChatApp(t, chatter, func(a *web.App) {
a.Processor = tripwireProcessor{t} // fails the test if chat summarizes/fetches
a.Fetcher = tripwireFetcher{t} // fails the test if chat fetches video metadata
})
p := rawPool(t)
resetDB(t, p)
require.NoError(t, deliver(ctx, app, videoX, "summary body"))
seedVideo(t, p, videoX, "X Title", "https://x", time.Time{})
seedTranscript(t, app.Store.(*store.Store), videoX, transcript)
form := url.Values{"question": {"what is it about?"}, "model": {"phi4-mini"}}
rec := postChat(t, app, videoX, form, true)
require.Equal(t, http.StatusOK, rec.Code)
require.Contains(t, body(t, rec), "It is about attention budgets.", "the answer is rendered")
require.Equal(t, 1, chatter.callCount, "the model was asked exactly once")
require.Equal(t, transcript, chatter.gotReq.Transcript,
"the model answered from the STORED transcript, not a fetched one")
// The tripwires never firing IS the no-fetch / no-YouTube proof.
}
// A video with no stored transcript yields an honest "not available" — and never
// a fetch, never a model call (ADR-027 §2: do not add an on-demand-fetch path).
func TestChatUnavailableWhenNoStoredTranscript(t *testing.T) {
ctx := context.Background()
chatter := &fakeChatter{models: []string{"phi4-mini"}}
app := newChatApp(t, chatter, func(a *web.App) {
a.Processor = tripwireProcessor{t}
a.Fetcher = tripwireFetcher{t}
})
p := rawPool(t)
resetDB(t, p)
require.NoError(t, deliver(ctx, app, videoX, "summary body"))
seedVideo(t, p, videoX, "X Title", "https://x", time.Time{})
// NB: no seedTranscript — the transcript is absent.
rec := getChat(t, app, videoX)
require.Equal(t, http.StatusOK, rec.Code)
require.Contains(t, body(t, rec), "isn't available", "honest not-available copy")
// Asking anyway still triggers nothing: no model call, no fetch.
rec = postChat(t, app, videoX, url.Values{"question": {"hi"}, "model": {"phi4-mini"}}, true)
require.Equal(t, http.StatusOK, rec.Code)
require.Equal(t, 0, chatter.callCount, "no model call without a stored transcript")
}
// Chat is reachable ONLY from the user's own summary view: opening chat for a
// video that belongs to another user 404s (the summary read is RLS-scoped), so a
// guessed/arbitrary video id is not a chat surface (ADR-027 §5 isolation).
func TestChatOnlyReachableForOwnVideo(t *testing.T) {
chatter := &fakeChatter{models: []string{"phi4-mini"}}
app := newChatApp(t, chatter, func(a *web.App) {
a.Processor = tripwireProcessor{t}
a.Fetcher = tripwireFetcher{t}
})
p := rawPool(t)
resetDB(t, p)
// A summarized video with a stored transcript owned by ANOTHER user.
const other = "22222222-2222-2222-2222-222222222222"
seedForeignSummaryWithTranscript(t, p, other, videoZ, "Foreign Title")
rec := getChat(t, app, videoZ)
require.Equal(t, http.StatusNotFound, rec.Code, "cannot open chat for another user's video")
require.Equal(t, 0, chatter.callCount, "no model call for a non-owned video")
rec = postChat(t, app, videoZ, url.Values{"question": {"hi"}, "model": {"phi4-mini"}}, true)
require.Equal(t, http.StatusNotFound, rec.Code, "cannot post chat to another user's video")
}
// The switcher offers exactly the backend's models, defaults to the summary's own
// model, and a switch re-runs against the SAME transcript with the chosen model
// (ADR-027 §3 — model-comparison instrumentation). A model dropped from the offer
// set (e.g. cloud disabled) is not rendered.
func TestChatModelSwitcherDefaultAndSwitch(t *testing.T) {
ctx := context.Background()
// The seeded summary's model is "phi4-mini" (see handlers_test summary()).
chatter := &fakeChatter{
models: []string{"phi4-mini", "gemma4-26b"}, // note: no cloud model offered
reply: chat.Reply{Answer: "answer"},
}
app := newChatApp(t, chatter, nil)
p := rawPool(t)
resetDB(t, p)
require.NoError(t, deliver(ctx, app, videoX, "summary body"))
seedVideo(t, p, videoX, "X Title", "https://x", time.Time{})
seedTranscript(t, app.Store.(*store.Store), videoX, "the transcript text")
html := body(t, getChat(t, app, videoX))
require.Contains(t, html, "gemma4-26b", "every offered model is in the switcher")
require.NotContains(t, html, "mistral", "a non-offered (cloud-disabled) model is absent")
require.Contains(t, html, `value="phi4-mini" selected`, "defaults to the summary's own model")
// Switching to gemma re-runs against the same stored transcript.
form := url.Values{"question": {"q"}, "model": {"gemma4-26b"}}
rec := postChat(t, app, videoX, form, true)
require.Equal(t, http.StatusOK, rec.Code)
require.Equal(t, "gemma4-26b", chatter.gotReq.Model, "the chosen model answers")
require.Equal(t, "the transcript text", chatter.gotReq.Transcript, "against the same transcript")
}
// A truncated transcript surfaces the honest bounded-context note (ADR-027 §2).
func TestChatTruncationNoteShown(t *testing.T) {
ctx := context.Background()
chatter := &fakeChatter{
models: []string{"phi4-mini"},
reply: chat.Reply{Answer: "answer", Truncated: true},
}
app := newChatApp(t, chatter, nil)
p := rawPool(t)
resetDB(t, p)
require.NoError(t, deliver(ctx, app, videoX, "summary body"))
seedVideo(t, p, videoX, "X Title", "https://x", time.Time{})
seedTranscript(t, app.Store.(*store.Store), videoX, "long transcript")
rec := postChat(t, app, videoX, url.Values{"question": {"q"}, "model": {"phi4-mini"}}, true)
require.Contains(t, body(t, rec), "bounded portion", "the truncation note is shown")
}
// A multi-turn conversation is carried in the request (hidden fields), not the DB:
// prior turns ride back into the next answer, and nothing is persisted (ADR-027 §4).
func TestChatMultiTurnHistoryIsEphemeral(t *testing.T) {
ctx := context.Background()
chatter := &fakeChatter{models: []string{"phi4-mini"}, reply: chat.Reply{Answer: "second answer"}}
app := newChatApp(t, chatter, nil)
p := rawPool(t)
resetDB(t, p)
require.NoError(t, deliver(ctx, app, videoX, "summary body"))
seedVideo(t, p, videoX, "X Title", "https://x", time.Time{})
seedTranscript(t, app.Store.(*store.Store), videoX, "the transcript")
// Second turn posts the prior exchange as hidden history fields.
form := url.Values{
"question": {"follow-up question"},
"model": {"phi4-mini"},
"hq": {"first question"},
"ha": {"first answer"},
}
rec := postChat(t, app, videoX, form, true)
require.Equal(t, http.StatusOK, rec.Code)
require.Len(t, chatter.gotReq.History, 1, "the prior turn was carried into the request")
require.Equal(t, "first question", chatter.gotReq.History[0].Question)
require.Equal(t, "first answer", chatter.gotReq.History[0].Answer)
html := body(t, rec)
require.Contains(t, html, "second answer", "the new answer renders")
require.Contains(t, html, "first question", "the conversation persists in page state")
// Nothing was written to the DB — there is no chat table; the summaries/actions
// are unchanged by a chat turn.
row, err := app.Store.GetSummaryByVideo(ctx, userID, videoX)
require.NoError(t, err)
require.Equal(t, "summary body", row.Summary, "chat never mutates stored data")
}
// --- foreign-user seeding (isolation) --------------------------------------
// seedForeignSummaryWithTranscript creates a fully-summarized, transcript-backed
// video owned by a DIFFERENT user via the raw pool (which bypasses RLS), so the
// isolation test can confirm the requesting user cannot open chat for it.
func seedForeignSummaryWithTranscript(t *testing.T, p *pgxpool.Pool, otherUserID, videoID, title string) {
t.Helper()
ctx := context.Background()
_, err := p.Exec(ctx, `INSERT INTO users (id) VALUES ($1) ON CONFLICT (id) DO NOTHING`, otherUserID)
require.NoError(t, err)
_, err = p.Exec(ctx,
`INSERT INTO videos (id, user_id, provider, provider_video_id, title, url)
VALUES ($1, $2, 'youtube', $3, $4, 'https://z')`,
videoID, otherUserID, seededProviderVideoID(videoID), title)
require.NoError(t, err)
_, err = p.Exec(ctx,
`INSERT INTO summaries (user_id, video_id, summary, ai_provider, ai_model)
VALUES ($1, $2, 'foreign summary', 'local', 'phi4-mini')`,
otherUserID, videoID)
require.NoError(t, err)
_, err = p.Exec(ctx,
`INSERT INTO transcripts (provider, provider_video_id, source, language, content)
VALUES ('youtube', $1, 'captions', 'en', 'foreign transcript')`,
seededProviderVideoID(videoID))
require.NoError(t, err)
}
+2 -2
View File
@@ -12,8 +12,8 @@ import (
"sync"
"time"
"gitea.d-ma.be/mathias/tapir/internal/adapters/store"
"gitea.d-ma.be/mathias/tapir/internal/auth"
"git.d-ma.be/mathias/tapir/internal/adapters/store"
"git.d-ma.be/mathias/tapir/internal/auth"
)
// Connections is the narrow write port the connect flow depends on (Clean
+3 -3
View File
@@ -11,9 +11,9 @@ import (
"github.com/stretchr/testify/require"
"golang.org/x/oauth2"
"gitea.d-ma.be/mathias/tapir/internal/adapters/store"
"gitea.d-ma.be/mathias/tapir/internal/auth"
"gitea.d-ma.be/mathias/tapir/internal/web"
"git.d-ma.be/mathias/tapir/internal/adapters/store"
"git.d-ma.be/mathias/tapir/internal/auth"
"git.d-ma.be/mathias/tapir/internal/web"
)
// fakeWriter is a TokenWriter capturing the persisted (ref, value).
+99 -10
View File
@@ -11,8 +11,8 @@ import (
"github.com/a-h/templ"
"gitea.d-ma.be/mathias/tapir/internal/adapters/store"
"gitea.d-ma.be/mathias/tapir/internal/domain"
"git.d-ma.be/mathias/tapir/internal/adapters/store"
"git.d-ma.be/mathias/tapir/internal/domain"
)
// Store is the read/write surface the web handlers depend on — a narrow port over
@@ -21,8 +21,15 @@ import (
// fake without a database.
type Store interface {
ListVideos(ctx context.Context, userID string, limit int) ([]store.SummaryRow, error)
// DistinctChannels lists the user's source channels — the options for the
// feed's channel multi-select filter.
DistinctChannels(ctx context.Context, userID string) ([]string, error)
GetSummaryByVideo(ctx context.Context, userID, videoID string) (*store.SummaryRow, error)
GetVideoRow(ctx context.Context, userID, videoID string) (*store.SummaryRow, error)
// GetTranscript reads the shared, stored transcript keyed by (provider,
// providerVideoID) — ADR-021. It is a pure DB read: it never fetches captions,
// so the chat path (ADR-027) reaches it without any caption-fetch surface.
GetTranscript(ctx context.Context, provider, providerVideoID string) (domain.Transcript, bool, error)
ActionsFor(ctx context.Context, userID string, videoIDs []string) (map[string][]string, error)
SetAction(ctx context.Context, userID, videoID, action string) error
ClearAction(ctx context.Context, userID, videoID, action string) error
@@ -88,6 +95,11 @@ type App struct {
// Fetcher, when non-nil, resolves an arbitrary YouTube video id to metadata for
// the paste-a-URL flow (Feature 2). Nil = the /paste route is not mounted.
Fetcher VideoFetcher
// Chat, when non-nil, answers per-video questions against a video's STORED
// transcript (ADR-027). Nil = the /v/{id}/chat routes are not mounted and the
// summary view shows no "dig deeper" affordance. It holds no caption-fetch
// dependency, so an enabled chat cannot reach YouTube or the rate gate.
Chat Chatter
// Processing tracks in-flight immediate summarizations so the status endpoint
// shows the animation until the summary lands. The zero value is ready to use.
Processing ProcessingSet
@@ -139,6 +151,8 @@ func (a *App) Router() http.Handler {
app := http.NewServeMux()
app.HandleFunc("GET /{$}", a.handleList)
app.HandleFunc("GET /v/{videoId}", a.handleDetail)
app.HandleFunc("GET /v/{videoId}/expand", a.handleExpand)
app.HandleFunc("GET /v/{videoId}/card", a.handleCard)
app.HandleFunc("POST /v/{videoId}/action", a.handleAction)
app.HandleFunc("POST /v/{videoId}/summarize", a.handleRequestSummarize)
app.HandleFunc("POST /v/{videoId}/retry-now", a.handleRetryNow)
@@ -146,6 +160,12 @@ func (a *App) Router() http.Handler {
app.HandleFunc("POST /paste", a.handlePaste)
}
app.HandleFunc("GET /v/{videoId}/status", a.handleStatus)
// Per-video deeper-dive chat over the STORED transcript (ADR-027). Mounted only
// when a Chat backend is wired; it never fetches captions.
if a.Chat != nil {
app.HandleFunc("GET /v/{videoId}/chat", a.handleChat)
app.HandleFunc("POST /v/{videoId}/chat", a.handleChatMessage)
}
app.HandleFunc("GET /register", a.handleRegisterForm)
app.HandleFunc("POST /register", a.handleRegister)
@@ -196,7 +216,7 @@ func (a *App) handleList(w http.ResponseWriter, r *http.Request) {
}
q := r.URL.Query()
f := Filter{
Channel: q.Get("channel"),
Channels: nonEmptyStrings(q["channel"]),
From: q.Get("from"),
To: q.Get("to"),
OnlySummarized: q.Get("summarized") == "1",
@@ -211,6 +231,13 @@ func (a *App) handleList(w http.ResponseWriter, r *http.Request) {
rows := f.apply(allRows)
buckets := bucketRows(rows, a.recencyCutoff())
// Channel options for the multi-select filter (the user's source channels).
channels, err := a.Store.DistinctChannels(r.Context(), userID)
if err != nil {
a.serverError(w, r, "distinct channels", err)
return
}
// hasConnected drives both the paste box (shown to ANY connected user, #2) and
// the empty-state copy (a fresh account with a connection but no discovery pass
// yet reads "connected, summaries land gradually" rather than "nothing here").
@@ -223,11 +250,20 @@ func (a *App) handleList(w http.ResponseWriter, r *http.Request) {
}
hasConnected := len(conns) > 0
if isHTMX(r) {
a.render(w, r, summaryList(buckets, hasConnected))
// Summarization mode drives the backlog copy: an auto user is told summaries
// land gradually; a manual user is told to click Summarize (the first pilot
// user sat in manual mode reading "land automatically" and waited forever).
autoSummarize, err := a.Store.GetAutoSummarize(r.Context(), userID)
if err != nil {
a.serverError(w, r, "summarize mode", err)
return
}
a.render(w, r, ListPage(buckets, f, stats, takeFlash(w, r), hasConnected))
if isHTMX(r) {
a.render(w, r, summaryList(buckets, hasConnected, autoSummarize))
return
}
a.render(w, r, ListPage(buckets, f, stats, takeFlash(w, r), hasConnected, channels, autoSummarize))
}
// handleDetail renders one summary in full (highlights, takeaways, action group).
@@ -246,7 +282,49 @@ func (a *App) handleDetail(w http.ResponseWriter, r *http.Request) {
a.serverError(w, r, "get summary", err)
return
}
a.render(w, r, DetailPage(*row))
a.render(w, r, DetailPage(*row, a.Chat != nil))
}
// handleExpand returns the inline-expanded card fragment — the full summary +
// chat dock swapped into the list card in place (ADR-031). Only summarized videos
// have a summary to expand; a non-summarized id is a 404 (the compact card never
// offers expand for it).
func (a *App) handleExpand(w http.ResponseWriter, r *http.Request) {
userID, ok := a.currentUserID(w, r)
if !ok {
return
}
videoID := r.PathValue("videoId")
row, err := a.Store.GetSummaryByVideo(r.Context(), userID, videoID)
if errors.Is(err, store.ErrNotFound) {
http.NotFound(w, r)
return
}
if err != nil {
a.serverError(w, r, "get summary", err)
return
}
a.render(w, r, expandedCard(*row, a.Chat != nil))
}
// handleCard returns the compact card fragment — the collapse target that returns
// an expanded card to its compact form in the list (ADR-031).
func (a *App) handleCard(w http.ResponseWriter, r *http.Request) {
userID, ok := a.currentUserID(w, r)
if !ok {
return
}
videoID := r.PathValue("videoId")
row, err := a.Store.GetVideoRow(r.Context(), userID, videoID)
if errors.Is(err, store.ErrNotFound) {
http.NotFound(w, r)
return
}
if err != nil {
a.serverError(w, r, "get video", err)
return
}
a.render(w, r, VideoCard(*row))
}
// handleAction toggles one action: re-clicking an active verb clears it, else it
@@ -476,11 +554,22 @@ func (a *App) handleStatus(w http.ResponseWriter, r *http.Request) {
return
}
if row.Summarized || !a.Processing.Has(processingKey(userID, videoID)) {
// Honest, state-aware status (ADR-025). Order matters: a finished summary wins;
// an in-flight goroutine shows the working spinner; a recorded rate-limit shows
// the calm "waiting, will retry" card that keeps polling; a recorded "none" is
// terminal; anything else falls back to the normal card.
switch {
case row.Summarized:
a.render(w, r, VideoCard(*row))
case a.Processing.Has(processingKey(userID, videoID)):
a.render(w, r, processingCard(*row))
case row.TranscriptStatus == "rate_limited":
a.render(w, r, waitingCard(*row))
case row.TranscriptStatus == "none":
a.render(w, r, noCaptionsCard(*row))
default:
a.render(w, r, VideoCard(*row))
return
}
a.render(w, r, processingCard(*row))
}
// handleSummarizeMode toggles the user's auto/manual summarization mode. The form
+66 -8
View File
@@ -7,6 +7,7 @@ import (
"net/http"
"net/http/httptest"
"os"
"path/filepath"
"strings"
"testing"
"time"
@@ -15,9 +16,9 @@ import (
"github.com/jackc/pgx/v5/pgxpool"
"github.com/stretchr/testify/require"
"gitea.d-ma.be/mathias/tapir/internal/adapters/store"
"gitea.d-ma.be/mathias/tapir/internal/domain"
"gitea.d-ma.be/mathias/tapir/internal/web"
"git.d-ma.be/mathias/tapir/internal/adapters/store"
"git.d-ma.be/mathias/tapir/internal/domain"
"git.d-ma.be/mathias/tapir/internal/web"
)
// dsn points at the in-process Postgres started in TestMain. Handler tests run
@@ -26,10 +27,22 @@ import (
var dsn string
func TestMain(m *testing.M) {
const port = 54330 // distinct from the store package's embedded PG (54329)
// Per-process port + dirs so concurrent `go test` runs (e.g. a push-run and a
// tag-run in CI) never collide on a fixed port or shared data dir. Base 55000
// keeps web's range distinct from the store package (54000). Shared CachePath
// downloads the PG archive once.
port := uint32(55000 + os.Getpid()%1000)
dsn = fmt.Sprintf("postgres://postgres:postgres@localhost:%d/postgres?sslmode=disable", port)
pg := embeddedpostgres.NewDatabase(embeddedpostgres.DefaultConfig().Port(port))
rt := filepath.Join(os.TempDir(), fmt.Sprintf("tapir-epg-web-%d", os.Getpid()))
pg := embeddedpostgres.NewDatabase(
embeddedpostgres.DefaultConfig().
Port(port).
RuntimePath(rt).
DataPath(filepath.Join(rt, "data")).
BinariesPath(filepath.Join(rt, "bin")).
CachePath(filepath.Join(os.TempDir(), "tapir-epg-cache")),
)
if err := pg.Start(); err != nil {
fmt.Fprintf(os.Stderr, "embedded-postgres start: %v\n", err)
os.Exit(1)
@@ -38,6 +51,7 @@ func TestMain(m *testing.M) {
if err := pg.Stop(); err != nil {
fmt.Fprintf(os.Stderr, "embedded-postgres stop: %v\n", err)
}
_ = os.RemoveAll(rt)
os.Exit(code)
}
@@ -207,13 +221,15 @@ func TestListChannelFilter(t *testing.T) {
resetDB(t, p)
require.NoError(t, deliver(ctx, app, videoX, "body x"))
seedVideo(t, p, videoX, "X Title", "https://x", time.Time{})
_, err := p.Exec(ctx, `UPDATE videos SET channel_title = 'Acme Channel' WHERE id = $1`, videoX)
require.NoError(t, err)
// Channel is "youtube" for seeded rows; a non-matching filter hides them.
rec := do(t, app, httptest.NewRequest(http.MethodGet, "/?channel=vimeo", nil))
// Selecting a different channel hides the row; selecting its channel shows it.
rec := do(t, app, httptest.NewRequest(http.MethodGet, "/?channel=Other+Channel", nil))
require.Equal(t, http.StatusOK, rec.Code)
require.NotContains(t, body(t, rec), "X Title")
rec = do(t, app, httptest.NewRequest(http.MethodGet, "/?channel=youtube", nil))
rec = do(t, app, httptest.NewRequest(http.MethodGet, "/?channel=Acme+Channel", nil))
require.Contains(t, body(t, rec), "X Title")
}
@@ -475,3 +491,45 @@ func postAction(t *testing.T, app *web.App, videoID, action string, htmx bool) *
}
return do(t, app, req)
}
// TestListManualModeBannerCopy: a manual-mode user with un-summarized videos
// sees the manual prompt (click Summarize), NOT the "summaries land
// automatically" copy that misled the first pilot user into waiting forever.
func TestListManualModeBannerCopy(t *testing.T) {
ctx := context.Background()
app := newApp(t)
p := rawPool(t)
resetDB(t, p)
seedVideo(t, p, videoX, "X Title", "https://x", time.Time{}) // pending, un-summarized
require.NoError(t, app.Store.SetAutoSummarize(ctx, userID, false))
html := body(t, do(t, app, httptest.NewRequest(http.MethodGet, "/", nil)))
require.Contains(t, html, "Manual mode")
require.Contains(t, html, "are not summarized automatically")
require.NotContains(t, html, "land gradually",
"manual-mode user must not be told summaries arrive automatically")
}
// TestListAutoModeBannerCopy: an auto-mode user with a backlog sees the
// gradual-delivery copy, not the manual prompt.
func TestListAutoModeBannerCopy(t *testing.T) {
ctx := context.Background()
app := newApp(t)
p := rawPool(t)
resetDB(t, p)
seedVideo(t, p, videoX, "X Title", "https://x", time.Time{})
require.NoError(t, app.Store.SetAutoSummarize(ctx, userID, true))
html := body(t, do(t, app, httptest.NewRequest(http.MethodGet, "/", nil)))
require.Contains(t, html, "land gradually")
require.NotContains(t, html, "are not summarized automatically")
}
// TestMetricsNotOnPublicMux: the public app router exposes no /metrics route —
// Prometheus is served on the dedicated metrics port only (ADR-030, security R6).
func TestMetricsNotOnPublicMux(t *testing.T) {
app := newApp(t)
resetDB(t, rawPool(t))
rec := do(t, app, httptest.NewRequest(http.MethodGet, "/metrics", nil))
require.Equal(t, http.StatusNotFound, rec.Code, "/metrics must not be on the public mux")
}
+107
View File
@@ -0,0 +1,107 @@
package web_test
import (
"context"
"net/http"
"net/http/httptest"
"testing"
"time"
"github.com/stretchr/testify/require"
)
// TestExpandReturnsSummaryBodyFragment: GET /v/{id}/expand returns the full
// summary as an in-place card fragment (not a full page) — ADR-031.
func TestExpandReturnsSummaryBodyFragment(t *testing.T) {
ctx := context.Background()
app := newApp(t)
p := rawPool(t)
resetDB(t, p)
require.NoError(t, deliver(ctx, app, videoX, "the full summary text"))
seedVideo(t, p, videoX, "X Title", "https://x", time.Time{})
html := body(t, do(t, app, httptest.NewRequest(http.MethodGet, "/v/"+videoX+"/expand", nil)))
require.Contains(t, html, "the full summary text")
require.Contains(t, html, "Takeaways")
require.Contains(t, html, "highlight one")
require.Contains(t, html, "card-expanded", "rendered as the expanded card")
require.Contains(t, html, "/v/"+videoX+"/card", "carries a collapse affordance")
require.NotContains(t, html, "<html", "fragment, not a full page")
}
// TestCollapseReturnsCompactCard: GET /v/{id}/card returns the compact card with
// the expand affordance — the collapse target.
func TestCollapseReturnsCompactCard(t *testing.T) {
ctx := context.Background()
app := newApp(t)
p := rawPool(t)
resetDB(t, p)
require.NoError(t, deliver(ctx, app, videoX, "summary text"))
seedVideo(t, p, videoX, "X Title", "https://x", time.Time{})
html := body(t, do(t, app, httptest.NewRequest(http.MethodGet, "/v/"+videoX+"/card", nil)))
require.Contains(t, html, `class="card"`, "compact card")
require.Contains(t, html, "/v/"+videoX+"/expand", "compact card offers expand")
require.NotContains(t, html, "card-expanded")
require.NotContains(t, html, "<html", "fragment, not a full page")
}
// TestExpandedCardOffersChatDock: with chat enabled, the expanded card includes
// the deeper-dive chat affordance.
func TestExpandedCardOffersChatDock(t *testing.T) {
ctx := context.Background()
app := newChatApp(t, &fakeChatter{models: []string{"m"}}, nil)
p := rawPool(t)
resetDB(t, p)
require.NoError(t, deliver(ctx, app, videoX, "summary text"))
seedVideo(t, p, videoX, "X Title", "https://x", time.Time{})
html := body(t, do(t, app, httptest.NewRequest(http.MethodGet, "/v/"+videoX+"/expand", nil)))
require.Contains(t, html, "/v/"+videoX+"/chat", "expanded card wires the chat dock")
}
// TestCompactCardExpandOnlyWhenSummarized: a not-yet-summarized card shows its
// summarize footer and no expand affordance.
func TestCompactCardExpandOnlyWhenSummarized(t *testing.T) {
app := newApp(t)
p := rawPool(t)
resetDB(t, p)
seedVideo(t, p, videoX, "Pending Title", "https://x", time.Time{}) // no summary
html := body(t, do(t, app, httptest.NewRequest(http.MethodGet, "/v/"+videoX+"/card", nil)))
require.Contains(t, html, "Not summarized")
require.NotContains(t, html, "/v/"+videoX+"/expand", "pending card offers no expand")
}
// TestCompactCardHasNoJSDetailFallback: the expand affordance carries an href to
// the detail page, so JS-off users still reach the full summary.
func TestCompactCardHasNoJSDetailFallback(t *testing.T) {
ctx := context.Background()
app := newApp(t)
p := rawPool(t)
resetDB(t, p)
require.NoError(t, deliver(ctx, app, videoX, "summary text"))
seedVideo(t, p, videoX, "X Title", "https://x", time.Time{})
html := body(t, do(t, app, httptest.NewRequest(http.MethodGet, "/v/"+videoX+"/card", nil)))
require.Contains(t, html, `href="/v/`+videoX+`"`, "no-JS fallback to the detail page")
require.Contains(t, html, "/v/"+videoX+"/expand", "and the HTMX expand for JS users")
}
// TestDetailAndExpandShareSummaryBody: the detail page and the expanded card render
// the same summary body (one shared fragment, no drift).
func TestDetailAndExpandShareSummaryBody(t *testing.T) {
ctx := context.Background()
app := newApp(t)
p := rawPool(t)
resetDB(t, p)
require.NoError(t, deliver(ctx, app, videoX, "shared summary text"))
seedVideo(t, p, videoX, "X Title", "https://x", time.Time{})
detail := body(t, do(t, app, httptest.NewRequest(http.MethodGet, "/v/"+videoX, nil)))
expand := body(t, do(t, app, httptest.NewRequest(http.MethodGet, "/v/"+videoX+"/expand", nil)))
for _, want := range []string{"shared summary text", "Takeaways", "highlight one"} {
require.Contains(t, detail, want)
require.Contains(t, expand, want)
}
}
+39 -45
View File
@@ -7,10 +7,11 @@
// Authentication is real (Dex OIDC) and is the only gate: any Dex-authenticated
// subject may sign in (ADR-012 dropped ADR-011's single-subject allowlist).
// Authorization/registration is layered on top in internal/web (an authenticated
// subject with no tapir user is routed to registration). Sessions are server-side
// (in-memory, fine for the single Stage-1 replica) addressed by an HMAC-signed
// (HS256) HttpOnly Secure SameSite=Lax cookie with a short TTL and sliding
// refresh. Tokens are never logged.
// subject with no tapir user is routed to registration). Sessions are STATELESS
// (ADR-029): the identity + expiry live inside an HMAC-signed (HS256) HttpOnly
// Secure SameSite=Lax persistent cookie with a long sliding TTL — no server-side
// table, so a deploy/restart never logs anyone out and the cookie also survives
// browser-close. Tokens are never logged; logout clears the cookie client-side.
//
// This is mcp-chassis's cousin but NOT the same code: mcp-chassis validates
// inbound Bearer JWTs for MCP APIs; this is a browser session login.
@@ -26,7 +27,8 @@ import (
"github.com/coreos/go-oidc/v3/oidc"
"golang.org/x/oauth2"
"gitea.d-ma.be/mathias/tapir/internal/web"
"git.d-ma.be/mathias/tapir/internal/metrics"
"git.d-ma.be/mathias/tapir/internal/web"
)
// Config is the OIDC + session configuration. cmd/tapir maps these from
@@ -47,7 +49,11 @@ type Config struct {
}
const (
defaultSessionTTL = time.Hour
// defaultSessionTTL is generous and sliding: Tapir is a "check back tomorrow"
// reader, so a short TTL meant a re-login (full IdP redirect dance) on almost
// every visit. 30 days, slid forward on each request, keeps a regular user
// logged in indefinitely while an abandoned session still lapses.
defaultSessionTTL = 30 * 24 * time.Hour
pendingTTL = 10 * time.Minute
sessionCookie = "tapir_session"
loginPath = "/auth/login"
@@ -59,7 +65,6 @@ type DexAuth struct {
oauth *oauth2.Config
verifier *oidc.IDTokenVerifier
sessions *sessionStore
pending *pendingStore
secret []byte
sessionTTL time.Duration
@@ -127,7 +132,6 @@ func New(ctx context.Context, cfg Config, opts ...Option) (*DexAuth, error) {
RedirectURL: cfg.RedirectURL,
Scopes: []string{oidc.ScopeOpenID, "profile", "email"},
},
sessions: newSessionStore(),
pending: newPendingStore(),
secret: []byte(cfg.SessionSecret),
sessionTTL: defaultSessionTTL,
@@ -159,31 +163,31 @@ func (d *DexAuth) Middleware(h http.Handler) http.Handler {
h.ServeHTTP(w, r)
return
}
sid, ok := d.sessionID(r)
c, err := r.Cookie(sessionCookie)
if err != nil {
d.redirectUnauthenticated(w, r)
return
}
user, _, ok := d.decodeSession(c.Value, d.now())
if !ok {
d.redirectUnauthenticated(w, r)
return
}
if _, ok := d.sessions.get(sid, d.now()); !ok {
d.redirectUnauthenticated(w, r)
return
}
d.sessions.refresh(sid, d.now().Add(d.sessionTTL)) // sliding refresh
// Sliding refresh: re-issue the cookie with a fresh expiry so an active
// user never lapses (the expiry lives in the cookie, so sliding = re-sign).
d.setSessionCookie(w, d.encodeSession(user, d.now().Add(d.sessionTTL)))
h.ServeHTTP(w, r)
})
}
// CurrentUser resolves the authenticated principal from the session cookie.
// CurrentUser resolves the authenticated principal from the stateless cookie.
func (d *DexAuth) CurrentUser(r *http.Request) (web.User, bool) {
sid, ok := d.sessionID(r)
if !ok {
c, err := r.Cookie(sessionCookie)
if err != nil {
return web.User{}, false
}
data, ok := d.sessions.get(sid, d.now())
if !ok {
return web.User{}, false
}
return data.user, true
user, _, ok := d.decodeSession(c.Value, d.now())
return user, ok
}
func (d *DexAuth) handleLogin(w http.ResponseWriter, r *http.Request) {
@@ -248,23 +252,16 @@ func (d *DexAuth) handleCallback(w http.ResponseWriter, r *http.Request) {
}
_ = idToken.Claims(&claims) // email is best-effort; subject is the identity
sid, err := randToken()
if err != nil {
http.Error(w, "internal error", http.StatusInternalServerError)
return
}
d.sessions.put(sid, sessionData{
user: web.User{Subject: idToken.Subject, Email: claims.Email},
expiry: d.now().Add(d.sessionTTL),
})
d.setSessionCookie(w, sid)
user := web.User{Subject: idToken.Subject, Email: claims.Email}
d.setSessionCookie(w, d.encodeSession(user, d.now().Add(d.sessionTTL)))
metrics.IncLogin()
http.Redirect(w, r, "/", http.StatusFound)
}
func (d *DexAuth) handleLogout(w http.ResponseWriter, r *http.Request) {
if sid, ok := d.sessionID(r); ok {
d.sessions.delete(sid)
}
// Stateless sessions: clearing the cookie logs the browser out. There is no
// server-side record to delete (ADR-029); a copy of the cookie stays valid
// until its expiry — an accepted trade for the Stage-0 reader app.
d.clearSessionCookie(w)
// Land on the public landing page, not the login endpoint: a just-logged-out
// visitor should see /welcome, not be bounced straight back into a Dex login.
@@ -287,22 +284,19 @@ func (d *DexAuth) redirectToLogin(w http.ResponseWriter, r *http.Request) {
http.Redirect(w, r, loginPath, http.StatusFound)
}
func (d *DexAuth) sessionID(r *http.Request) (string, bool) {
c, err := r.Cookie(sessionCookie)
if err != nil {
return "", false
}
return d.unsign(c.Value)
}
func (d *DexAuth) setSessionCookie(w http.ResponseWriter, sid string) {
// setSessionCookie writes the signed session value as a PERSISTENT cookie
// (Max-Age set), so it survives the browser/app being closed — a session cookie
// (no Max-Age) was dropped on iPhone Safari close, forcing re-login. value is the
// already-signed payload from encodeSession.
func (d *DexAuth) setSessionCookie(w http.ResponseWriter, value string) {
http.SetCookie(w, &http.Cookie{
Name: sessionCookie,
Value: d.sign(sid),
Value: value,
Path: "/",
HttpOnly: true,
Secure: !d.insecure,
SameSite: http.SameSiteLaxMode,
MaxAge: int(d.sessionTTL.Seconds()),
})
}
+35 -6
View File
@@ -14,8 +14,8 @@ import (
josev4 "github.com/go-jose/go-jose/v4"
"github.com/stretchr/testify/require"
"gitea.d-ma.be/mathias/tapir/internal/web"
"gitea.d-ma.be/mathias/tapir/internal/web/oidc"
"git.d-ma.be/mathias/tapir/internal/web"
"git.d-ma.be/mathias/tapir/internal/web/oidc"
)
const (
@@ -306,13 +306,42 @@ func TestLogoutClearsSession(t *testing.T) {
require.Equal(t, http.StatusFound, rec.Code)
require.Equal(t, "/welcome", rec.Header().Get("Location"), "logout lands on the public page")
cleared := sessionCookie(t, rec.Result())
require.Less(t, cleared.MaxAge, 0, "logout expires the cookie")
require.Less(t, cleared.MaxAge, 0, "logout expires the cookie so the browser drops it")
require.Empty(t, cleared.Value, "logout blanks the cookie value")
// The server-side session is gone: the original cookie no longer resolves.
// Sessions are stateless (ADR-029): logout clears the cookie client-side, so a
// request carrying the cleared (empty) cookie is unauthenticated. The original
// signed cookie remains technically valid until its expiry — the accepted
// trade for no server-side store; the browser no longer holds it.
check := httptest.NewRequest(http.MethodGet, "/", nil)
check.AddCookie(cookie)
check.AddCookie(cleared)
_, ok := auth.CurrentUser(check)
require.False(t, ok)
require.False(t, ok, "the cleared cookie does not authenticate")
}
// TestSessionSurvivesRestart is the core of ADR-029: a cookie issued by one
// process is accepted by a FRESH instance with the same session secret — so a
// deploy/pod-restart no longer logs users out (the old in-memory store did).
func TestSessionSurvivesRestart(t *testing.T) {
f := newFakeIssuer(t)
auth1 := newAuth(t, f)
cookie := authenticate(t, auth1, f)
auth2 := newAuth(t, f) // simulate a redeploy: new process, same SessionSecret
req := httptest.NewRequest(http.MethodGet, "/", nil)
req.AddCookie(cookie)
user, ok := auth2.CurrentUser(req)
require.True(t, ok, "a session must survive a restart (stateless signed cookie)")
require.Equal(t, testSubject, user.Subject)
}
// TestSessionCookieIsPersistent: the cookie carries a positive Max-Age so it
// survives the browser/app being closed (a session cookie was dropped on iOS).
func TestSessionCookieIsPersistent(t *testing.T) {
f := newFakeIssuer(t)
auth := newAuth(t, f)
cookie := authenticate(t, auth, f)
require.Greater(t, cookie.MaxAge, 0, "session cookie must be persistent (Max-Age set)")
}
func TestExpiredSessionRejected(t *testing.T) {
+32 -43
View File
@@ -6,12 +6,13 @@ import (
"crypto/sha256"
"encoding/base64"
"encoding/hex"
"encoding/json"
"fmt"
"strings"
"sync"
"time"
"gitea.d-ma.be/mathias/tapir/internal/web"
"git.d-ma.be/mathias/tapir/internal/web"
)
// randToken returns a URL-safe 256-bit random string for session IDs, OIDC
@@ -53,56 +54,44 @@ func (d *DexAuth) unsign(signed string) (string, bool) {
return value, true
}
// sessionData is the server-side session record.
type sessionData struct {
user web.User
expiry time.Time
// sessionClaims is the self-contained session payload carried INSIDE the signed
// cookie — there is no server-side session table. This is deliberate (ADR-029):
// an in-memory store was wiped on every pod restart, logging every user out on
// each deploy, and a stateless cookie also survives browser-close and works
// across replicas. It holds only the identity (subject + email, not secret) and
// an absolute expiry; the HMAC tag (sign/unsign) makes it tamper-proof.
type sessionClaims struct {
Sub string `json:"s"`
Email string `json:"e"`
Exp int64 `json:"x"` // unix seconds; absolute expiry
}
// sessionStore is an in-memory session table. Single replica at Stage 0, so an
// in-process map is sufficient; it is safe for concurrent use.
type sessionStore struct {
mu sync.Mutex
m map[string]sessionData
// encodeSession produces the signed cookie value for a user with the given expiry.
func (d *DexAuth) encodeSession(u web.User, exp time.Time) string {
b, _ := json.Marshal(sessionClaims{Sub: u.Subject, Email: u.Email, Exp: exp.Unix()})
return d.sign(base64.RawURLEncoding.EncodeToString(b))
}
func newSessionStore() *sessionStore { return &sessionStore{m: make(map[string]sessionData)} }
func (s *sessionStore) put(id string, d sessionData) {
s.mu.Lock()
defer s.mu.Unlock()
s.m[id] = d
}
// get returns the session if present and unexpired; expired entries are evicted.
func (s *sessionStore) get(id string, now time.Time) (sessionData, bool) {
s.mu.Lock()
defer s.mu.Unlock()
d, ok := s.m[id]
// decodeSession verifies the cookie's HMAC, parses the claims, and checks expiry.
// It returns the user and the absolute expiry on success.
func (d *DexAuth) decodeSession(cookieValue string, now time.Time) (web.User, time.Time, bool) {
payload, ok := d.unsign(cookieValue)
if !ok {
return sessionData{}, false
return web.User{}, time.Time{}, false
}
if !now.Before(d.expiry) {
delete(s.m, id)
return sessionData{}, false
raw, err := base64.RawURLEncoding.DecodeString(payload)
if err != nil {
return web.User{}, time.Time{}, false
}
return d, true
}
// refresh slides an existing session's expiry forward; a no-op for unknown ids.
func (s *sessionStore) refresh(id string, expiry time.Time) {
s.mu.Lock()
defer s.mu.Unlock()
if d, ok := s.m[id]; ok {
d.expiry = expiry
s.m[id] = d
var c sessionClaims
if err := json.Unmarshal(raw, &c); err != nil {
return web.User{}, time.Time{}, false
}
}
func (s *sessionStore) delete(id string) {
s.mu.Lock()
defer s.mu.Unlock()
delete(s.m, id)
exp := time.Unix(c.Exp, 0)
if !now.Before(exp) {
return web.User{}, time.Time{}, false // expired
}
return web.User{Subject: c.Sub, Email: c.Email}, exp, true
}
// pendingData holds the nonce bound to an in-flight authorization request.

Some files were not shown because too many files have changed in this diff Show More