From 525811bc1abc816fcb114f32a02e31eab727b67f Mon Sep 17 00:00:00 2001 From: mathias Date: Thu, 28 May 2026 10:30:39 +0000 Subject: [PATCH] docs: add hypothesis statement and harness boundary clarification --- README.md | 125 ++++++++++++++++++++++++++++++++---------------------- 1 file changed, 74 insertions(+), 51 deletions(-) diff --git a/README.md b/README.md index 12a0d8f..b168937 100644 --- a/README.md +++ b/README.md @@ -5,14 +5,31 @@ Instead of letting Claude Code do whatever it wants, hyperguild enforces structu workflows (TDD red/green/refactor), logs every session, and accumulates learnings into a searchable brain. +## Hypothesis + +> We believe routing skill tasks through local models, backed by brain context, +> produces measurably better outcomes than raw Claude Code alone — +> measurable by per-skill pass rate over rolling 30-day windows +> (available at `GET /pass-rate?skill=&window=30d` on the brain pod). + +This is the falsifiable claim the routing pod and pass-rate infrastructure exist to test. +If per-skill pass rates don't improve over baseline (all-cloud) after 30 days of real +usage, the fast-model routing path should be reconsidered. + +## Harness + +**hyperguild = Claude Code + MCP.** This is a supervisor for Claude Code sessions specifically. +For multi-agent orchestration (OpenCode + LiteLLM, executor/reviewer pipelines), see +[agentsquad](http://gitea.d-ma.be/mathias/agentsquad) — a separate harness for a different +orchestration model. Skills (mathias/skills) are shared between both. + ## How it works ``` Your Claude Code session (in any project) │ │ MCP over HTTP (Tailscale) - ├──▶ supervisor :3200 (NodePort 30320 on koala) — skill workers: tdd, debug, spec, … - ├──▶ routing :3210 (NodePort 30310 on koala) — Mode 2 only: review, debug, retrospective, trainer + ├──▶ routing :3210 (NodePort 30310 on koala) — review, debug, retrospective, trainer └──▶ brain :3300 (NodePort 30330 on koala) — brain_query, brain_write, brain_ingest, session_log │ └─ also serves the legacy REST endpoints (/query, /write, /ingest, …) @@ -20,34 +37,28 @@ Your Claude Code session (in any project) ▼ brain/ ├── sessions/ — JSONL log, one file per session_id -├── wiki/ — searchable knowledge (full-text) -│ ├── concepts/ -│ ├── entities/ -│ └── sources/ -├── raw/ — retrospective output, staged for review -└── training-data/ — SFT/DPO/RL data (Phase 2) +├── wiki/ — searchable knowledge (wing/hall layout) +│ ├── homelab/ +│ ├── claude-sessions/ +│ └── ... +└── knowledge/ — legacy flat notes (migration pending: hyperguild#22) ``` ## Phase 1 tools (available now) | Tool | What it does | |------|-------------| -| `tdd_red` | Writes a failing test for a spec, verifies it fails | -| `tdd_green` | Writes the minimal implementation to make tests pass | -| `tdd_refactor` | Cleans up implementation while keeping tests green | | `session_log` | Appends a structured entry to the session JSONL log | -| `retrospective` | Reads the session log, identifies novel learnings, writes to brain/raw/ | +| `retrospective` | Reads the session log, identifies novel learnings, writes to brain | +| `review` | Structured code review via local model, brain-context injected | +| `debug` | Hypothesis-driven debugging via local model | | `brain_query` | Full-text search over brain/wiki/ | -| `brain_write` | Writes a note to brain/raw/ (with optional YAML frontmatter) | +| `brain_write` | Writes a note to brain (with wing/hall routing) | +| `brain_answer` | BM25 + LLM synthesis — Q&A over brain corpus | | `tier` | Returns the current connectivity tier (1=cloud, 2=LAN, 3=offline) | -## Start the servers - -```bash -# Requires goreman: go install github.com/mattn/goreman@latest -task start # starts ingestion (:3300) + supervisor (:3200) via goreman -task stop # kills both by port -``` +> **Note:** `tdd_red/green/refactor` and `spec` were retired in Plan 7 (2026-05-12). +> They are now SKILL.md files in [mathias/skills](http://gitea.d-ma.be/mathias/skills). ## Connect a project @@ -56,9 +67,9 @@ Create `.mcp.json` in your project root: ```json { "mcpServers": { - "supervisor": { + "routing": { "type": "http", - "url": "http://koala:30320/mcp" + "url": "http://koala:30310/mcp" }, "brain": { "type": "http", @@ -68,65 +79,77 @@ Create `.mcp.json` in your project root: } ``` -Two MCP servers are exposed today, both reachable over Tailscale: +Two MCP servers are exposed, both reachable over Tailscale: -- **`supervisor`** at `koala:30320` — skill workers (`tdd_red/green/refactor`, - `review`, `debug`, `spec`, `retrospective`, `trainer`, `tier`). +- **`routing`** at `koala:30310` — skill workers (`review`, `debug`, `retrospective`, `trainer`). + Routes each call to fast local model or thinking model based on per-skill pass rate. - **`brain`** at `koala:30330` — knowledge access (`brain_query`, `brain_write`, - `brain_ingest`, `brain_ingest_raw`) and `session_log`. Hosted by the ingestion - service directly, no separate pod. + `brain_ingest`, `brain_ingest_raw`, `brain_answer`, `brain_classify`) and `session_log`. No local binary or stdio shim is required — Claude Code talks to both via HTTP. Open Claude Code in your project — run `/mcp` to confirm both servers are listed. -## A typical TDD session +## A typical session ``` -1. Call tdd_red → spec in, failing test file out -2. Call tdd_green → test path in, implementation out -3. Call tdd_refactor → impl + test in, cleaned code out -4. Call session_log → log each phase result -5. Call retrospective → extracts learnings → brain/raw/ -6. Review brain/raw/, move worthy notes to brain/wiki/concepts/ -7. Future sessions: call brain_query to retrieve relevant context +1. Call review → brain context injected + local model review → findings +2. Call session_log → log each phase result +3. Call retrospective → extracts learnings → brain +4. Future sessions: call brain_query / brain_answer to retrieve relevant context ``` ## Tier detection -The supervisor probes connectivity at call time: +The routing pod probes connectivity at call time: | Tier | Label | Condition | -|------|-------|-----------| +|------|-------|-----------| | 1 | full-online | Can reach api.anthropic.com | | 2 | lan-only | Can reach LiteLLM but not Anthropic | | 3 | airplane | No external connectivity | +## Model routing + +The routing pod selects models per skill call based on historical pass rate: + +| Pass rate | Decision | +|-----------|----------| +| ≥ 0.90 (FLOOR) | Fast model (`HYPERGUILD_FAST_MODEL`) | +| ≤ 0.70 (CEIL) | Thinking model (`HYPERGUILD_THINKING_MODEL`) | +| between CEIL and FLOOR | Sample band — probabilistic routing | +| nil (no history yet) | Defaults to thinking model | + +> **Bootstrap note:** With no session history, all calls route to the thinking model. +> The fast-model path activates only after real pass-rate data accumulates at `/pass-rate`. +> Seed with real usage — don't try to pre-populate. + ## Key env vars | Variable | Default | Purpose | -|----------|---------|---------| +|----------|---------|---------| | `INGEST_BRAIN_DIR` | `../brain` | Brain directory for ingestion server | | `INGEST_PORT` | `3300` | Ingestion server port | -| `SUPERVISOR_CONFIG_DIR` | `./config/supervisor` | Skill discipline files | -| `SUPERVISOR_SESSIONS_DIR` | `./brain/sessions` | JSONL session logs | -| `INGEST_BASE_URL` | `http://localhost:3300` | Supervisor → ingestion | +| `INGEST_BASE_URL` | `http://localhost:3300` | Routing pod → brain | | `LITELLM_BASE_URL` | — | LiteLLM proxy for Tier 2 model routing | -| `SUPERVISOR_MCP_TOKEN` | — | Optional bearer token for the supervisor MCP HTTP endpoint; when empty, no auth is enforced | | `ROUTING_PORT` | `3210` | Routing pod's listen port | -| `ROUTING_MCP_TOKEN` | — | Optional bearer token for the routing MCP HTTP endpoint | +| `ROUTING_MCP_TOKEN` | — | Optional bearer token; when empty, no auth enforced | | `BRAIN_URL` | `http://ingestion.supervisor:3300` | Routing pod → brain (in-cluster) | | `HYPERGUILD_FAST_MODEL` | `koala/qwen35-9b-fast` | Fast model for high-pass-rate skill calls | | `HYPERGUILD_THINKING_MODEL` | `iguana/gemma4-26b` | Thinking model for low-pass-rate skill calls | -| `HYPERGUILD_ROUTE_LOCAL_FLOOR` | `0.90` | At/above pass rate, route to fast model | -| `HYPERGUILD_ROUTE_LOCAL_CEIL` | `0.70` | Below pass rate, route to thinking model. Between CEIL and FLOOR is the sample band. | +| `HYPERGUILD_ROUTE_LOCAL_FLOOR` | `0.90` | Fast model threshold | +| `HYPERGUILD_ROUTE_LOCAL_CEIL` | `0.70` | Thinking model threshold | | `HYPERGUILD_PASS_RATE_TTL_SECONDS` | `60` | Per-skill pass-rate cache TTL | -> **Operator note:** LiteLLM at `LITELLM_BASE_URL` must register both `HYPERGUILD_FAST_MODEL` and `HYPERGUILD_THINKING_MODEL` for routing to do useful work. If a model is missing, LiteLLM returns 4xx, the routing pod's fast route fails, the fail-open retry on the thinking model likely also fails (since both are missing), and the only signal is `final_status: "fail"` on `_routing` entries in the brain. +> **Operator note:** LiteLLM at `LITELLM_BASE_URL` must register both `HYPERGUILD_FAST_MODEL` +> and `HYPERGUILD_THINKING_MODEL`. If a model is missing, the fail-open retry also fails and +> the only signal is `final_status: "fail"` on `_routing` entries in the brain. -## Phase 2 (planned) +## Open issues -- `review` skill — structured code review with iron law enforcement -- `debug` skill — hypothesis-driven debugging sessions -- `spec` skill — generates specs from conversations -- `trainer` — extracts SFT/DPO pairs from session logs for fine-tuning +See [issues](http://gitea.d-ma.be/mathias/hyperguild/issues) — key open items: + +- **#25** — skills platform overhaul (audit first, then lazy loading + brain feedback loop) +- **#24** — reduce context burn from skill listing +- **#22** — migrate legacy brain notes to wing/hall layout (one-shot script, low risk) +- **#31** — connect routing-mcp to claude.ai as custom connector