Files
tapir/VISION.md
T
mathiasandClaude Opus 4.8 f35c2a85a5
CI / Lint / Test / Vet (push) Successful in 13s
CI / Build & Import (push) Successful in 11s
docs: scheduled-discovery env + single-replica constraint; VISION gate-clock reset
homelab-integration.md gains a "Scheduled discovery" section documenting
TAPIR_DISCOVERY_INTERVAL and TAPIR_FETCH_RATE and the load-bearing
single-replica constraint (in-process scheduler → replicas: 1 is required;
>1 double-runs discovery). VISION Stage 0 carries a pointer to ADR-018's
gate-clock reset so nothing in docs implies the window started before
unprompted use was possible.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-05 23:43:17 +02:00

140 lines
8.1 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Tapir — Product Vision
> **One line:** Tapir quietly watches the channels you already follow and hands you
> the substance of each new video — highlights and takeaways — so you can decide what
> deserves your full attention without watching everything.
## The problem
People subscribe to far more YouTube/Vimeo channels than they can watch. Valuable
videos go unwatched; time is spent watching videos that turn out not to be worth it.
The signal is buried in hours of runtime. Existing "summary" tools are one-off,
paste-a-URL affairs — they don't *watch on your behalf* and they don't respect where
your attention and data should live.
## The product
Tapir connects to a user's YouTube/Vimeo account, learns their subscriptions, and
when a subscribed channel posts a new video, it:
1. Detects the new video.
2. Fetches its transcript (captions first).
3. Summarizes it into highlights and takeaways relevant to the user.
4. Delivers the summary to the user's store (and, optionally, to other sinks).
The analysis runs on a **local-first AI stack**. If the local stack cannot do the job
reliably, the user may connect their own Claude / ChatGPT / Gemini account as a
fallback — their key, their choice.
## Principles (the guardrails)
- **Local-first, user-owned.** The default processor is the self-hosted stack. External
AI is opt-in, per-user, with the user's own credentials. Tapir never silently ships a
user's content to a third-party model.
- **Standalone is the product.** Tapir is a standalone service first. Feeding a personal
knowledge base ("brain") is *one optional sink*, not the reason Tapir exists.
- **Attention is the scarce resource, not compute.** Every feature is judged by whether it
helps the user spend less time deciding what to watch. Summaries exist to protect
attention.
- **Data isolation is a promise, not a feature flag.** Each user's connected accounts,
credentials, and summaries are separated. This holds from the first user, not "later."
- **Respect the source.** Captions where available; no fragile or ToS-hostile scraping in
the core path. Where richer transcription is added later, it is a clearly-bounded,
optional component.
## Who it is for
- **Now (the first customers):** the maintainer and a small number of known, trusted
friends — each with their own account, isolated data, optional BYO-AI. The maintainer is
the first customer; friendly users provide the earliest real-world signal.
- **Maybe (Future C, explicitly not built yet):** a public multi-tenant service. Deferred
until there is evidence of sustained use **and** real demand. Building for C before that
evidence is a known anti-goal.
## Definition of Success
Success is staged. Each stage has a single, falsifiable headline test. We do not advance
to the next stage's ambition until the current stage's test passes.
### Stage 0 — Useful to me or a friend (the gate)
> **Headline test:** Over a 34 week window, *either* the maintainer *or* at least one
> onboarded friend returns to Tapir and reads/acts on summaries in **≥2 separate weeks**.
> The test is *return usage* (behavioural), not stated approval. The ideal signal is an
> **unprompted** return (organic, not because the maintainer nudged them) — but see the
> measurement note below: we currently cannot distinguish prompted from organic returns, so
> in practice we count all returns and read the result with that caveat.
- Captions-first summarization works end-to-end for real subscriptions (the maintainer's
and onboarded friends').
- Summaries land in each user's own store and (optionally) brain.
- Local-first AI produces summaries of acceptable quality without manual intervention
most of the time.
- **Why behavioural, not feedback.** Friend *feedback* is gathered and genuinely valuable —
but it is **not** the gate. Asked-for feedback from friendly users is the least reliable
signal in product development (politeness bias); whether they *come back* is the thing we
actually care about. So the gate measures returns, not nice words.
- **Measurement note — "unprompted" is an ideal we can't yet measure.** Whether a return was
organic or prompted by a nudge is not captured by any data Tapir holds (it's context only
the maintainer has). Rather than waive the standard, we name the gap: *unprompted* return
is the signal we genuinely want; *returns* (prompted or not) is what the data can show. A
return that needed a nudge is a weaker signal than one that didn't, and the result is read
with that in mind. If distinguishing them ever matters enough, the maintainer tracks nudges
manually or a future build records prompt events — neither is in scope now.
- **Why "me OR a friend".** This replaces the original "useful to *me*, specifically" gate
(2026-06-03 decision, recorded in DECISIONS.md ADR-016). Getting signal from friendly
users is valuable enough to count — but the bar stays behavioural so it can't be cleared
by a polite reaction. (Ties to the 2026-07-01 check-in.)
- **How it's measured.** Return usage is read from two sources: `summary_actions` (timestamped
watch/skip/save per user) answers "acted in ≥2 distinct weeks"; an append-only login-events
table (see infra/Tapir build) answers "returned/read in ≥2 distinct weeks" even without an
action click — the honest signal for a *reading* product. Login events accrue only from their
deploy date onward, so the gate window's data begins then.
- **Gate-clock reset (ADR-018).** The 34 week window starts when in-process scheduled discovery
+ auto-summarize ship — before that, unprompted use was impossible, so the prior window
measured nothing (this is starting the clock when the experiment can actually run, not a reset
to dodge a failing gate). The 2026-07-01 check-in referenced above moves accordingly to ~34
weeks after this deploys. See DECISIONS.md ADR-018.
- **This is the gate.** Hardening (Stage 1) and any SaaS ambition stay deferred until this
behavioural signal exists. Note: multi-user machinery was deliberately built *ahead* of
this gate (ADR-012) with isolation enforced — that was an explicit, recorded call, not a
sign the gate had passed. The gate is about *evidence of use*, which is still open.
### Stage 1 — Trustworthy at rest (hardening, Future B)
> **Headline test:** Credentials (OAuth tokens, BYO-AI keys) are encrypted at rest via the
> homelab's existing secrets convention; a documented, rehearsed recovery path exists; and
> a deliberate isolation test (user A cannot read user B's data) passes in CI or a
> documented manual drill.
- Per-user data isolation is enforced and tested (delivered early via ADR-012 RLS).
- Per-user credentials are encrypted at rest (ADR-015 envelope encryption; build in infra#89).
- A new user can self-connect a YouTube/Vimeo account and get summaries with no code change.
- Optional BYO-AI works per-user.
### Non-goals (current)
- Public sign-up / billing / a marketing surface.
- Google OAuth app verification at production scale.
- Audio-download + speech-to-text transcription in the core path (a bounded optional
component at most, deferred).
- Becoming a general-purpose video archive, player, or recommendation engine.
## How we will know we are drifting
- We declare the Stage 0 gate "passed" on the strength of polite feedback rather than
behavioural return-usage (the politeness-bias trap the gate is designed to resist).
- We build Stage 1 hardening or Future C machinery while the Stage 0 use-evidence is still
absent. (Multi-user machinery already shipped ahead of the gate via ADR-012 — a recorded,
deliberate exception, not a precedent for more.)
- A user's content reaches a third-party model without that user's explicit, per-user opt-in.
- "Brain ingestion" starts dictating the architecture instead of being one sink behind an
interface.
- The codebase acquires a second language or a parallel auth/secrets system that duplicates
the homelab's existing Dex / ESO conventions.
---
_This document is a guardrail. Changes to the principles, the staged definition of success,
or the non-goals are architecture decisions and must be recorded in `DECISIONS.md`._