39 Commits
Author SHA1 Message Date
mathias b8fbc5a805 docs: spec first-session wow — verify & improve onboarding summary burst (investigate-first)
CI / Lint / Test / Vet (push) Successful in 11s
CI / Build & Import (push) Successful in 10s
The concern is new-user first-contact: see some GOOD summaries fast or they
don't return. Reframed as curation/latency for ~3 videos, NOT a 429/throughput
problem (3 fetches is nowhere near the wall). Phase 1 (report-and-stop) verifies
whether the existing cap-3 onboarding burst even fires today, what it delivers,
and — critically — how much transcript-cache overlap exists between users (drives
the blend). Phase 2 levers: cached-transcript-first (instant, zero-fetch),
likely-good selection (has-captions/good-length, not just newest), and optionally
the stronger model for the burst's few summaries. Blend deferred to the
maintainer post-Phase-1. Explicitly NOT bulk-fetch, NOT credentials (ADR-010/026
dead end), NOT a client extension.
2026-06-11 15:36:09 +00:00
mathias 137804b0b1 docs: spec chat-with-transcript (ADR-027)
CI / Lint / Test / Vet (push) Successful in 11s
CI / Build & Import (push) Successful in 10s
Per-video chat against the ADR-021 stored transcript, entered from the summary
view, born from observed demand (maintainer read real summaries, some made him
want to dig deeper). HARD constraint: stored-transcript-only — never fetches
captions, never touches the rate gate or YouTube, safe by construction. Default
model = the summary's model, user-switchable among the ADR-022 chain models
(doubles as model-comparison instrumentation). Ephemeral v1 (no persisted
history); chat-only/trust-the-model with show-source verification recorded as the
natural v2. ADR-027 to be appended to DECISIONS.md as the first commit.
2026-06-11 06:38:16 +00:00
mathias 58cd68c1eb docs: spec video-card state unification — one "Summarize now" verb + honest no-captions state
CI / Lint / Test / Vet (push) Successful in 12s
CI / Build & Import (push) Successful in 10s
The card shows two verbs ("Try now" for rate-limited, "Summarize" for pending)
for one user intent, and — the real bug — a no-captions video falls into the
pending branch and wrongly shows a Summarize button that can only fail. Spec
collapses to one quiet "Summarize now" verb wherever a nudge is possible (both
handlers unchanged underneath), adds an honest no-button "No transcript
available" state, and keeps the card status-first (buttons are exceptions in auto
mode). View-layer only.
2026-06-06 19:58:45 +00:00
mathias 8403e8e524 docs: spec newest-first batch ordering + honest Try-now/prioritisation docs
CI / Lint / Test / Vet (push) Successful in 12s
CI / Build & Import (push) Successful in 11s
Product intent: new users get summaries of their newest videos fast while the
back-catalogue fills behind, within the shared rate gate. The background batch
currently processes in subscription/channel order, not newest-first — this spec
closes that gap (collect candidates, sort published_at DESC NULLS LAST, process
in order, gate unchanged). Also corrects the docs to describe Try-now as
onboarding prioritisation, explicitly removing the prior "looks organic to
YouTube" traffic-disguising framing — rate limiting is respected, not evaded.
2026-06-06 19:20:27 +00:00
mathias 88294d38bc docs: add ADR-018 — in-process scheduled discovery, auto-summarize, gate-clock reset
CI / Lint / Test / Vet (push) Successful in 12s
CI / Build & Import (push) Successful in 10s
Records the Stage-0 usability decision: tapir serve runs discovery for all users
on an interval (reusing runner.Loop), auto-summarize defaults ON so the list fills
itself, and the gate clock resets to when this ships (unprompted use was
impossible before, so the prior window measured nothing — framed as starting the
clock when the experiment can run, not dodging a failing gate). Records the
in-process-vs-CronJob tradeoff, the load-bearing single-replica constraint, and
that it makes ADR-014 item 2 (process-wide rate gate) a hard requirement folded
into the build. Cross-referenced ADR-014/016; added CronJob to rejected-alts.
2026-06-05 13:03:18 +00:00
mathias ca9bf2657f docs: spec in-process scheduled discovery + auto-summarize + rate-gate finish
CI / Lint / Test / Vet (push) Successful in 12s
CI / Build & Import (push) Successful in 13s
The Stage-0 usability fix: tapir serve runs discovery for all users on an
interval (reusing the existing Runner.Loop, enumerate-users-then-withUser),
auto-summarize defaults ON for Future-B users so the list fills itself, and the
ADR-014 process-wide per-egress-IP rate gate is confirmed/finished in the same
slice because in-process + auto + multi-user makes it load-bearing. Records the
single-replica constraint as load-bearing, and a fallback (auto-summarize OFF
until the gate exists) so the dangerous combination never ships half-built.
2026-06-05 12:59:50 +00:00
mathias 4706c508e9 docs: ratify ADR-017 KEEP — invite flow is load-bearing (some users can't use Google OIDC)
CI / Lint / Test / Vet (push) Successful in 12s
CI / Build & Import (push) Successful in 10s
Records the deliberate keep-or-reverse decision on the Dex-write invite flow.
Decision: KEEP. Deciding fact: not all intended Future-B users will use Google
accounts, so Google OIDC alone can't onboard them — the invite flow is
load-bearing, not redundant. Trust-surface cost accepted deliberately, explicitly
NOT as a precedent for widening further, and explicitly NOT by adding delete RBAC
to fix the orphan gap. Open items reframed as tracked follow-ups (verify RBAC
against the real manifest; accept orphan for Future B, revisit before Future C).
Added the reversal to rejected-alternatives.
2026-06-05 12:47:39 +00:00
mathias b4da97e0e2 docs: resolve ADR-014 migration-008 false alarm (cols shipped in 007)
CI / Lint / Test / Vet (push) Successful in 13s
CI / Build & Import (push) Successful in 10s
Reconciliation finding: the v0.6.0 report's "migration 008 videos.rate_limited_at"
was a mislabel. Both transcript_status and rate_limited_at shipped in migration
007; the sequence legitimately skips 008, nothing was lost, and the runner's
column reads are sound. Updated the ADR-014 implementation note to state this
(was flagged as an unresolved discrepancy). The item-2 (shared per-egress-IP rate
gate vs per-video backoff) flag stays open — still unconfirmed.
2026-06-04 18:54:56 +00:00
mathias 99743af182 docs: add ADR-017 (Dex-write invite flow) retroactively; annotate ADR-002/013/014
CI / Lint / Test / Vet (push) Successful in 12s
CI / Build & Import (push) Successful in 10s
Reconciliation pass after parallel agent sessions shipped v0.6.0/v0.7.0.

ADR-017 documents the v0.7.0 invite flow, which gave Tapir scoped create+get on
passwords.dex.coreos.com in the auth namespace — Tapir now WRITES to the shared
identity provider. This shipped with no ADR; recorded retroactively with the
principle-reversal named (partially supersedes ADR-002/013), the security analysis
(bounded RBAC, but a real larger trust surface), and the open gaps (orphaned Dex
accounts on delete; plaintext invite tokens). Cross-referenced in ADR-002 and
ADR-013 status lines so a future reader isn't misled.

ADR-014 annotated: 429 handling shipped in v0.6.0 but decision item 2 (shared
per-egress-IP rate gate) appears realised as per-VIDEO backoff, not a process-wide
IP gate — flagged not-confirmed-done. Also flags the missing migration 008 /
rate_limited_at discrepancy (v0.6.0 report cited 008; tree jumps 007->009).

No code changed in this commit — audit trail only.
2026-06-03 21:38:49 +00:00
mathias 1cf58768ed docs: spec the Stage 0 usage-measurement build (login events)
CI / Lint / Test / Vet (push) Successful in 10s
CI / Build & Import (push) Successful in 10s
Small tapir slice to make the gate measurable as written: append-only
login_events (RLS, per-user-per-day throttle) + a union query over reads
(login_events) and acts (summary_actions) for distinct-active-weeks. Carries the
honesty caveats (unprompted not measurable; data accrues from deploy; week-bucket
noise at low N) and the delete-cascade footgun (no FK, needs explicit delete +
test) from the prior delete work. Out of scope: analytics, prompt-tracking,
dashboards.
2026-06-03 20:59:49 +00:00
mathias f45ba35e25 docs: VISION Stage 0 — keep "unprompted" as ideal, note measurement gap
CI / Lint / Test / Vet (push) Has been cancelled
CI / Build & Import (push) Has been cancelled
Reframes "unprompted" from an enforced criterion to a named measurement
limitation: organic-vs-prompted returns aren't distinguishable from any data
Tapir holds, so in practice all returns are counted and the result read with that
caveat (a nudged return is a weaker signal). Adds a "how it's measured" note
pointing at summary_actions (acts) + a new append-only login-events table
(read-returns), which accrue from deploy onward. Honest about the gap rather than
silently dropping the word.
2026-06-03 20:59:14 +00:00
mathias e6f508824b docs: add ADR-016 — Stage 0 gate revised to "me or a friend", behavioural
CI / Lint / Test / Vet (push) Successful in 19s
CI / Build & Import (push) Successful in 11s
Records the gate change: Stage 0 now passes when either the maintainer or an
onboarded friend returns unprompted in >=2 separate weeks. Behavioural (return
usage), not feedback-based, to resist politeness bias. Includes an honest
self-scrutiny note that this is a guardrail edit made while the original gate was
unmet — examined on that basis and proceeding because it broadens who supplies the
signal without softening the kind of signal required. Adds feedback-based-gate to
rejected alternatives.
2026-06-03 20:46:40 +00:00
mathias 477701fea2 docs: revise Stage 0 gate to "useful to me or a friend" (behavioural)
CI / Lint / Test / Vet (push) Successful in 11s
CI / Build & Import (push) Successful in 10s
Replaces the original "useful to me, specifically" gate with "me OR a friend
returns unprompted in >=2 separate weeks" — friendly-user signal counts, but the
test stays behavioural (return usage) not feedback-based, to resist politeness
bias. Folds the old Stage 1 ("a trusted user returns") into the new Stage 0 (they
were near-identical), renumbers hardening to Stage 1, and updates the drift
signals (the gate can be softened by mistaking polite feedback for evidence;
multi-user shipping ahead of the gate was a recorded exception per ADR-012, not a
precedent). Rationale recorded in ADR-016.
2026-06-03 20:44:19 +00:00
mathias 17fad140a6 docs: add ADR-015 — per-user credentials envelope-encrypted in PG18
CI / Lint / Test / Vet (push) Successful in 10s
CI / Build & Import (push) Successful in 10s
Records the infra#88 spike decision: per-user OAuth tokens are runtime app-state,
not config, so they live envelope-encrypted in PG18 under RLS (key from 1P via the
existing read-only SA) rather than in the vault. Infra creds stay ESO/1Password —
two mechanisms because they're two different things. Includes the falsification
conditions (frequent rotation; estate audit policy; key-rotation cost) so the
choice is earned not assumed. Adds the two rejected candidates (vault-write SA;
Supabase) to the rejected-alternatives table. Full reasoning in the #88 decision
doc; build + reboot-validation in #89.
2026-06-03 20:29:50 +00:00
mathias 6b817f11b9 docs: add landing-page + doc-reconciliation build spec
CI / Lint / Test / Vet (push) Successful in 11s
CI / Mirror to GitHub (push) Failing after 3s
CI / Build & Import (push) Successful in 10s
Workstream A: public /welcome landing page (bubbletea aesthetic, one Dex login
flow, logged-in shortcuts) — with the oidc.go facts verified against main, incl.
the two corrections that only surface from reading the code (logout must redirect
to /welcome not /auth/login; bare-/ vs deep-link redirect split).

Workstream B: reconcile the guardrail docs against deployed reality (v0.4.0) —
auth.go comments, data-model isolation status + migrations 002-006 schema,
architecture web surface, use-case scenarios for the Stage-1 features, ADR
ordering, and a requirements-vs-shipped deviation check. Structured as two
parallel workstreams so the doc audit isn't done cursorily alongside the build.
2026-06-03 19:18:40 +00:00
mathias 672a0c8580 docs: add ADR-013 (delete semantics) and ADR-014 (429 handling + UX)
CI / Lint / Test / Vet (push) Successful in 10s
CI / Mirror to GitHub (push) Failing after 3s
CI / Build & Import (push) Successful in 10s
ADR-013 records the deliberate choice that account deletion is Tapir-side only
(cascade + secret purge), leaving the shared Dex identity intact — clean
re-registration, but a noted GDPR-shaped gap if Future C ever arrives.

ADR-014 specifies timedtext 429 handling: Retry-After-aware backoff, a single
per-egress-IP rate gate shared by the batch and click paths, and honest in-flight
UX (summarizing / queued-waiting / no-transcript) so a rate-limited fetch never
presents as a stuck spinner or error. Whisper stays deferred pending measurement
of the sustainable rate, which this work finally makes measurable.
2026-06-03 19:03:52 +00:00
mathias b211cd08f2 docs: add build + runtime network egress requirements
CI / Lint / Test / Vet (push) Failing after 2s
CI / Build & Import (push) Has been skipped
CI / Mirror to GitHub (push) Has been skipped
Lists the egress the build assumes (Go module proxy, toolchain download,
raw.githubusercontent for golangci-lint, github for non-proxied modules) and the
runtime assumes (LiteLLM gateway, brain-mcp, YouTube/Vimeo APIs, per-user BYO-AI
hosts only on opt-in), so a locked-down koala act_runner or dev env knows what to
allow or which GOPROXY to set. Notes the claude.ai sandbox allowlist is separate
and unrelated.
2026-06-02 11:40:50 +00:00
mathias ff1fa982be docs: add skills wiring and current-build-state to CLAUDE.md
CI / Lint / Test / Vet (push) Failing after 2s
CI / Build & Import (push) Has been skipped
CI / Mirror to GitHub (push) Has been skipped
Tells an agent to wire skills via `task skills` (gitignored symlinks, never
committed) and which skills matter for Tapir; documents the scaffolded-and-RED
state with the first build task spelled out; flags the unverified setup items
(Go version, brain-mcp URL, secret-ref naming, model alias) to resolve against
the live cluster.
2026-06-02 11:07:30 +00:00
mathias 70e6f77fc8 chore: add .gitignore
CI / Lint / Test / Vet (push) Failing after 6s
CI / Build & Import (push) Has been skipped
CI / Mirror to GitHub (push) Has been skipped
Excludes build artifacts and the .claude/skills symlink (wired by the skills
installer, never committed — matches the mathias/skills convention) and local
env files.
2026-06-02 11:06:13 +00:00
mathias c7250fc493 ci: add Gitea Actions workflow (gitea-ci skill conventions)
CI / Lint / Test / Vet (push) Has been cancelled
CI / Mirror to GitHub (push) Has been cancelled
CI / Build & Import (push) Has been cancelled
check -> build -> mirror, self-hosted runner, buildah to localhost:5000, k3s
smoke test. Follows the gitea-ci skill template and its act_runner gotchas
(secrets inlined in run:, no heredocs). check will be RED until the engine is
implemented (acceptance suite). Deploy job omitted until k3s manifests exist in
infra. GH_DEPLOY_KEY secret must be set before mirror succeeds.
2026-06-02 11:06:09 +00:00
mathias 5c5b10f46b chore: add Taskfile with the check quality gate
Defines `task check` (fmt-check + vet + lint + test) — the gate ADR-009 and CI
both invoke. Also `task skills` to wire the engineering skills library via the
canonical installer (symlinks, gitignored). lint no-ops locally when
golangci-lint is absent; CI installs it.
2026-06-02 11:05:48 +00:00
mathias 26e37fbbb3 feat: add cmd/tapir entrypoint stub
Minimal main that identifies the binary (gives the CI smoke test something to
grep). HTTP server, watcher, adapter wiring, and config come with the build.
2026-06-02 11:05:37 +00:00
mathias 0064e34841 test: add acceptance tests for summarize-new-video (RED)
Executable translation of docs/use-cases/summarize_new_video.feature: captioned
video is summarized and delivered to the store sink; no-transcript video is
skipped with reason "no transcript" and no delivery. Drives the engine through
fake adapters (no live YouTube/brain). Intentionally RED until the engine is
implemented — this is the swarm's target.
2026-06-02 11:05:32 +00:00
mathias 6fd31e197c feat: add use-case engine scaffold (RED)
Engine wires the ports and exposes ProcessNewVideo, the core use case. Returns
ErrNotImplemented on purpose so the acceptance suite fails RED — implementing it
to make those tests pass is the first build task. Depends only on ports + domain
(dependencies point inward).
2026-06-02 11:05:12 +00:00
mathias c26e29ac4d feat: add ports (VideoSource, Summarizer, Sink, SecretStore)
The hexagonal interfaces the engine depends on. Keeps the engine provider- and
sink-agnostic: YouTube/Vimeo implement VideoSource, the AI router implements
Summarizer, store/brain implement Sink, ESO implements SecretStore. This is what
makes standalone-vs-homelab a wiring choice (ADR-003).
2026-06-02 11:05:01 +00:00
mathias 31251a41c3 feat: add domain entities (Clean Architecture core)
Pure domain types matching docs/data-model.md: User, Subscription, Video,
Transcript, Summary, plus Provider and TranscriptSource enums. Stdlib-only,
no outward dependencies — the innermost layer. Video carries user_id per the
per-user-isolation decision (no global dedup).
2026-06-02 11:04:48 +00:00
mathias 9268c99f5e chore: add go.mod (module gitea.d-ma.be/mathias/tapir)
Go module root for Tapir. Go 1.23 — confirm against the koala act_runner
toolchain; bump to match the estate (ingestion uses 1.26.1) if the runner has it.
2026-06-02 11:04:35 +00:00
mathias bc79167dfe docs: record rejected alternatives in DECISIONS.md
Adds a consolidated table of approaches considered and deliberately not taken
(Python, Supabase, living in the monolith, shared-lib lift, filesystem brain
package, reusing inbound oauth, global dedup table, Whisper-in-core, SaaS-now,
swarm-delegating the spike), each mapped to the ADR that settles it. Prevents a
later session from re-proposing settled rejections as fresh ideas.
2026-06-02 10:38:41 +00:00
mathias 32e06843f1 docs: add homelab integration facts reference
Pins the concrete endpoints/conventions Tapir depends on so independent sessions
don't rediscover them: LiteLLM gateway (http://koala:4000/v1/, sk-local-123, host/name
alias format), BYO fallback, brain-mcp HTTP sink, Dex identity, ESO+1Password secrets,
hosts, and Flux GitOps deploy location. Verified facts recorded as such; unconfirmed
items (brain-mcp URL, exact secret-ref naming, post-relocation LiteLLM location,
summarization model alias) explicitly flagged 'confirm' rather than invented.
2026-06-02 10:37:55 +00:00
mathias deb8523d75 docs: add CLAUDE.md agent operating instructions
Operational context for independent agent sessions: orientation order, TBD +
conventional-commit workflow, the three "looks reusable but isn't" traps (llm is
copied not imported, brain sink is HTTP not filesystem, OAuth is fresh), the
settled decisions not to reopen, the Clean Architecture/BDD stance, and a
provenance trail back to the S5 spike and the llm source. Distills the operating
knowledge surfaced during the 2026-06-02 planning+grill session.
2026-06-02 10:37:22 +00:00
mathias 367c28f0e9 docs: expand README as repo orientation and guardrail index
Points anyone (or any agent) landing cold at the vision, decisions, architecture,
data model, and BDD feature specs; states the Clean Architecture / TDD-BDD / TBD
approach and the homelab conventions reused. Notes the repo is pre-code and the
docs are the version-controlled design intent.
2026-06-02 10:33:22 +00:00
mathias 9848156cb3 docs(bdd): add account-connection feature
Gherkin spec for connecting YouTube/Vimeo accounts and configuring optional
per-provider BYO AI credentials: connections sync subscriptions, tokens/keys are
stored only as secret references (never in the clear), revocation stops watching
but preserves history. Encodes the secrets-by-reference and data-isolation
guardrails (ADR-002, ADR-006, data-model).
2026-06-02 10:32:48 +00:00
mathias 5c99c49a50 docs(bdd): add AI-routing feature (local-first, BYO fallback)
Gherkin spec for the routing guardrail: local produces the summary by default;
on local failure, fall back only to a user-configured BYO provider; with no BYO,
queue for retry and never send content to a third-party model. Encodes the
local-first/user-owned principle (VISION) as executable behavior.
2026-06-02 10:32:37 +00:00
mathias 0b124a648b docs(bdd): add summarize-new-video feature
Gherkin spec for the core use case: captioned video is summarized and delivered;
no-transcript video is recorded as skipped; unsubscribed channels are ignored;
already-summarized videos are not reprocessed. These scenarios seed the use-case
test suite (Clean Architecture core tested through fake adapters).
2026-06-02 10:32:27 +00:00
mathias 4456d25651 docs: add data model (Stage 0 / Stage 1 scope)
Per-user-isolated entities (no global cross-tenant video table per the S5/grill
correction), secrets stored by reference only (ESO/1Password, never the token),
brain delivery modelled as a sink_delivery row rather than brain-specific tables.
fallback_used recorded per summary as the Stage 0 quality signal. Future C dedup
and sharding explicitly out of scope.
2026-06-02 10:32:14 +00:00
mathias af38999ee2 docs: add C4 architecture + sequence diagrams (Mermaid)
Context and container diagrams, sequence diagrams for the core summarize-new-video
use case and the local-first/BYO AI routing fallback, and the Clean Architecture
layering. Ports & adapters keep the engine provider- and sink-agnostic, making
standalone-vs-homelab a wiring choice (ADR-003), not two codebases.
2026-06-02 10:31:37 +00:00
mathias 7090fb1e40 docs: add architecture decision records (ADR-001..009)
Records the decisions from the S5 spike and the Full Grill as append-only ADRs:
Go not Python; no Supabase; standalone-first with brain as one sink; copy the
llm package; brain sink via HTTP brain-mcp; fresh outbound OAuth; captions-first
with STT deferred; Future C deferred behind the Stage 0 gate; trunk-based dev.
2026-06-02 10:31:00 +00:00
mathias d1a0b49fa9 docs: add product vision and staged definition of success
The top-level guardrail for Tapir: the problem, the product, the principles
(local-first, standalone-first, attention-as-scarce-resource, data isolation),
who it's for (now / Future B / deferred Future C), and a staged, falsifiable
definition of success with Stage 0 ("useful to me") as the gate.
2026-06-02 10:30:15 +00:00
mathias 36fe5ba1b7 Initial commit 2026-06-02 10:29:39 +00:00