# spike-media — the `/bygge` toolchain, recovered Throwaway spike tooling for #28 (upload → transcript → takeaways). Committed because it was found in `tmpfs` one reboot from gone, not because it is finished. Nothing here is wired into tapir. No database, no HTTP handler, no ADR. ## What produced `/bygge` The prototype at `tapir.d-ma.be/bygge` was **not** hand-built, which the epic and the first round of spike issues both got wrong. It came from these three scripts: ``` IMG_1233.mov --transcode.yaml--> IMG_1233.web.mp4 (k3s Job, h264, +faststart) IMG_1233.wav --whisper---------> IMG_1233.sv.srt (interactive — the one manual step) IMG_1233.sv.srt --analyze_srt.py--> analys.json (berget/mistral-medium, temperature=0) analys.json + video --build_page.py--> page HTML ``` The reference analysis was produced by **`berget/mistral-medium`** — a mid-tier cloud model, not a frontier one. That matters when judging what a local model has to clear: the bar is mistral-medium, and it should be re-run as the control arm rather than excluded. ## The validator is the valuable part `analyze_srt.py` grades its own output on four checks. Two of them catch classes of error that timestamp- and citation-checking cannot: | Check | Catches | |---|---| | citation exact / **partial** / absent | Paraphrase presented as a quote. Three tiers, because `alternativ` citations legitimately stitch non-contiguous passages — a strict substring test fails the reference output. | | `TALET … FINNS INTE I CITATET` | A wrong number carried by a *verbatim* citation and a *real* timestamp. Converts Swedish number-words to digits first. It caught `45×230` where the transcript said `220`. | | `TACKNINGSLUCKA` | Spans of the clip no section covers. **The recall term** — without it an empty analysis scores perfectly on every other check. | | timestamp exists in the SRT | Fabricated seek targets. Near-solved: 74/75 on the reference, because the model copies the cue starts it was handed. | ## Known defects — read before running - **`max_tokens: 6000` is too low.** The 4-minute reference analysis is ~4,700 output tokens, i.e. ~78% of budget. A 10-minute clip truncates for every model, and truncation reads as a JSON-discipline failure if you are not counting tokens. Record `finish_reason` and separate `truncated` from `malformed`. - **`build_page.py` hardcodes the title** (`"Jonas – genomgång av tillbyggnad"`). - **The schema carries fields nothing renders** — `projekt.sammanfattning`, `projekt.skede`, `avsnitt.bildtid`, `avsnitt.bildmotiv`, `material[]`, `oklart[]` — left over from an earlier still-frame/markdown output that `build_page.py` replaced. They are paid for in output tokens on every run. - **Speaker attribution is inferred, not diarized.** The output attributes statements by name from an SRT with no speaker labels. All four checks above are blind to a misattribution. - **The web encode does not downscale.** The source is 720×1280; `-preset veryfast -crf 24` on handheld motion is what makes it ~5 Mbps, not the resolution. The lever is the preset. ## Fixtures `fixtures/fake.sv.srt` + `fixtures/fake.analys.json` — a synthetic 7-cue Swedish construction walkthrough with planted numbers. No real people, no real project. It exercises the parser and all four validator checks **without a model call**, which is what makes it usable in a test. ## Why the real artifacts are not here **This repo is public.** The prototype's transcript and analysis are a verbatim recording of a named person discussing a client's building project — private third-party content, and exactly the class of data #28 calls out as the reason uploaded transcripts cannot live in the shared `transcripts` table. So `IMG_1233.sv.srt`, `payload.json` and the model-comparison runs stay out of git here. That leaves `/bygge` without a versioned source of truth, which is a real gap and a decision for #28: a private repo, an encrypted blob, or accepting that the reference artifacts live outside version control.