docs: add hypothesis statement and harness boundary clarification
This commit is contained in:
@@ -5,14 +5,31 @@ Instead of letting Claude Code do whatever it wants, hyperguild enforces structu
|
|||||||
workflows (TDD red/green/refactor), logs every session, and accumulates learnings
|
workflows (TDD red/green/refactor), logs every session, and accumulates learnings
|
||||||
into a searchable brain.
|
into a searchable brain.
|
||||||
|
|
||||||
|
## Hypothesis
|
||||||
|
|
||||||
|
> We believe routing skill tasks through local models, backed by brain context,
|
||||||
|
> produces measurably better outcomes than raw Claude Code alone —
|
||||||
|
> measurable by per-skill pass rate over rolling 30-day windows
|
||||||
|
> (available at `GET /pass-rate?skill=<name>&window=30d` on the brain pod).
|
||||||
|
|
||||||
|
This is the falsifiable claim the routing pod and pass-rate infrastructure exist to test.
|
||||||
|
If per-skill pass rates don't improve over baseline (all-cloud) after 30 days of real
|
||||||
|
usage, the fast-model routing path should be reconsidered.
|
||||||
|
|
||||||
|
## Harness
|
||||||
|
|
||||||
|
**hyperguild = Claude Code + MCP.** This is a supervisor for Claude Code sessions specifically.
|
||||||
|
For multi-agent orchestration (OpenCode + LiteLLM, executor/reviewer pipelines), see
|
||||||
|
[agentsquad](http://gitea.d-ma.be/mathias/agentsquad) — a separate harness for a different
|
||||||
|
orchestration model. Skills (mathias/skills) are shared between both.
|
||||||
|
|
||||||
## How it works
|
## How it works
|
||||||
|
|
||||||
```
|
```
|
||||||
Your Claude Code session (in any project)
|
Your Claude Code session (in any project)
|
||||||
│
|
│
|
||||||
│ MCP over HTTP (Tailscale)
|
│ MCP over HTTP (Tailscale)
|
||||||
├──▶ supervisor :3200 (NodePort 30320 on koala) — skill workers: tdd, debug, spec, …
|
├──▶ routing :3210 (NodePort 30310 on koala) — review, debug, retrospective, trainer
|
||||||
├──▶ routing :3210 (NodePort 30310 on koala) — Mode 2 only: review, debug, retrospective, trainer
|
|
||||||
└──▶ brain :3300 (NodePort 30330 on koala) — brain_query, brain_write, brain_ingest, session_log
|
└──▶ brain :3300 (NodePort 30330 on koala) — brain_query, brain_write, brain_ingest, session_log
|
||||||
│
|
│
|
||||||
└─ also serves the legacy REST endpoints (/query, /write, /ingest, …)
|
└─ also serves the legacy REST endpoints (/query, /write, /ingest, …)
|
||||||
@@ -20,34 +37,28 @@ Your Claude Code session (in any project)
|
|||||||
▼
|
▼
|
||||||
brain/
|
brain/
|
||||||
├── sessions/ — JSONL log, one file per session_id
|
├── sessions/ — JSONL log, one file per session_id
|
||||||
├── wiki/ — searchable knowledge (full-text)
|
├── wiki/ — searchable knowledge (wing/hall layout)
|
||||||
│ ├── concepts/
|
│ ├── homelab/
|
||||||
│ ├── entities/
|
│ ├── claude-sessions/
|
||||||
│ └── sources/
|
│ └── ...
|
||||||
├── raw/ — retrospective output, staged for review
|
└── knowledge/ — legacy flat notes (migration pending: hyperguild#22)
|
||||||
└── training-data/ — SFT/DPO/RL data (Phase 2)
|
|
||||||
```
|
```
|
||||||
|
|
||||||
## Phase 1 tools (available now)
|
## Phase 1 tools (available now)
|
||||||
|
|
||||||
| Tool | What it does |
|
| Tool | What it does |
|
||||||
|------|-------------|
|
|------|-------------|
|
||||||
| `tdd_red` | Writes a failing test for a spec, verifies it fails |
|
|
||||||
| `tdd_green` | Writes the minimal implementation to make tests pass |
|
|
||||||
| `tdd_refactor` | Cleans up implementation while keeping tests green |
|
|
||||||
| `session_log` | Appends a structured entry to the session JSONL log |
|
| `session_log` | Appends a structured entry to the session JSONL log |
|
||||||
| `retrospective` | Reads the session log, identifies novel learnings, writes to brain/raw/ |
|
| `retrospective` | Reads the session log, identifies novel learnings, writes to brain |
|
||||||
|
| `review` | Structured code review via local model, brain-context injected |
|
||||||
|
| `debug` | Hypothesis-driven debugging via local model |
|
||||||
| `brain_query` | Full-text search over brain/wiki/ |
|
| `brain_query` | Full-text search over brain/wiki/ |
|
||||||
| `brain_write` | Writes a note to brain/raw/ (with optional YAML frontmatter) |
|
| `brain_write` | Writes a note to brain (with wing/hall routing) |
|
||||||
|
| `brain_answer` | BM25 + LLM synthesis — Q&A over brain corpus |
|
||||||
| `tier` | Returns the current connectivity tier (1=cloud, 2=LAN, 3=offline) |
|
| `tier` | Returns the current connectivity tier (1=cloud, 2=LAN, 3=offline) |
|
||||||
|
|
||||||
## Start the servers
|
> **Note:** `tdd_red/green/refactor` and `spec` were retired in Plan 7 (2026-05-12).
|
||||||
|
> They are now SKILL.md files in [mathias/skills](http://gitea.d-ma.be/mathias/skills).
|
||||||
```bash
|
|
||||||
# Requires goreman: go install github.com/mattn/goreman@latest
|
|
||||||
task start # starts ingestion (:3300) + supervisor (:3200) via goreman
|
|
||||||
task stop # kills both by port
|
|
||||||
```
|
|
||||||
|
|
||||||
## Connect a project
|
## Connect a project
|
||||||
|
|
||||||
@@ -56,9 +67,9 @@ Create `.mcp.json` in your project root:
|
|||||||
```json
|
```json
|
||||||
{
|
{
|
||||||
"mcpServers": {
|
"mcpServers": {
|
||||||
"supervisor": {
|
"routing": {
|
||||||
"type": "http",
|
"type": "http",
|
||||||
"url": "http://koala:30320/mcp"
|
"url": "http://koala:30310/mcp"
|
||||||
},
|
},
|
||||||
"brain": {
|
"brain": {
|
||||||
"type": "http",
|
"type": "http",
|
||||||
@@ -68,65 +79,77 @@ Create `.mcp.json` in your project root:
|
|||||||
}
|
}
|
||||||
```
|
```
|
||||||
|
|
||||||
Two MCP servers are exposed today, both reachable over Tailscale:
|
Two MCP servers are exposed, both reachable over Tailscale:
|
||||||
|
|
||||||
- **`supervisor`** at `koala:30320` — skill workers (`tdd_red/green/refactor`,
|
- **`routing`** at `koala:30310` — skill workers (`review`, `debug`, `retrospective`, `trainer`).
|
||||||
`review`, `debug`, `spec`, `retrospective`, `trainer`, `tier`).
|
Routes each call to fast local model or thinking model based on per-skill pass rate.
|
||||||
- **`brain`** at `koala:30330` — knowledge access (`brain_query`, `brain_write`,
|
- **`brain`** at `koala:30330` — knowledge access (`brain_query`, `brain_write`,
|
||||||
`brain_ingest`, `brain_ingest_raw`) and `session_log`. Hosted by the ingestion
|
`brain_ingest`, `brain_ingest_raw`, `brain_answer`, `brain_classify`) and `session_log`.
|
||||||
service directly, no separate pod.
|
|
||||||
|
|
||||||
No local binary or stdio shim is required — Claude Code talks to both via HTTP.
|
No local binary or stdio shim is required — Claude Code talks to both via HTTP.
|
||||||
|
|
||||||
Open Claude Code in your project — run `/mcp` to confirm both servers are listed.
|
Open Claude Code in your project — run `/mcp` to confirm both servers are listed.
|
||||||
|
|
||||||
## A typical TDD session
|
## A typical session
|
||||||
|
|
||||||
```
|
```
|
||||||
1. Call tdd_red → spec in, failing test file out
|
1. Call review → brain context injected + local model review → findings
|
||||||
2. Call tdd_green → test path in, implementation out
|
2. Call session_log → log each phase result
|
||||||
3. Call tdd_refactor → impl + test in, cleaned code out
|
3. Call retrospective → extracts learnings → brain
|
||||||
4. Call session_log → log each phase result
|
4. Future sessions: call brain_query / brain_answer to retrieve relevant context
|
||||||
5. Call retrospective → extracts learnings → brain/raw/
|
|
||||||
6. Review brain/raw/, move worthy notes to brain/wiki/concepts/
|
|
||||||
7. Future sessions: call brain_query to retrieve relevant context
|
|
||||||
```
|
```
|
||||||
|
|
||||||
## Tier detection
|
## Tier detection
|
||||||
|
|
||||||
The supervisor probes connectivity at call time:
|
The routing pod probes connectivity at call time:
|
||||||
|
|
||||||
| Tier | Label | Condition |
|
| Tier | Label | Condition |
|
||||||
|------|-------|-----------|
|
|------|-------|-----------|
|
||||||
| 1 | full-online | Can reach api.anthropic.com |
|
| 1 | full-online | Can reach api.anthropic.com |
|
||||||
| 2 | lan-only | Can reach LiteLLM but not Anthropic |
|
| 2 | lan-only | Can reach LiteLLM but not Anthropic |
|
||||||
| 3 | airplane | No external connectivity |
|
| 3 | airplane | No external connectivity |
|
||||||
|
|
||||||
|
## Model routing
|
||||||
|
|
||||||
|
The routing pod selects models per skill call based on historical pass rate:
|
||||||
|
|
||||||
|
| Pass rate | Decision |
|
||||||
|
|-----------|----------|
|
||||||
|
| ≥ 0.90 (FLOOR) | Fast model (`HYPERGUILD_FAST_MODEL`) |
|
||||||
|
| ≤ 0.70 (CEIL) | Thinking model (`HYPERGUILD_THINKING_MODEL`) |
|
||||||
|
| between CEIL and FLOOR | Sample band — probabilistic routing |
|
||||||
|
| nil (no history yet) | Defaults to thinking model |
|
||||||
|
|
||||||
|
> **Bootstrap note:** With no session history, all calls route to the thinking model.
|
||||||
|
> The fast-model path activates only after real pass-rate data accumulates at `/pass-rate`.
|
||||||
|
> Seed with real usage — don't try to pre-populate.
|
||||||
|
|
||||||
## Key env vars
|
## Key env vars
|
||||||
|
|
||||||
| Variable | Default | Purpose |
|
| Variable | Default | Purpose |
|
||||||
|----------|---------|---------|
|
|----------|---------|---------|
|
||||||
| `INGEST_BRAIN_DIR` | `../brain` | Brain directory for ingestion server |
|
| `INGEST_BRAIN_DIR` | `../brain` | Brain directory for ingestion server |
|
||||||
| `INGEST_PORT` | `3300` | Ingestion server port |
|
| `INGEST_PORT` | `3300` | Ingestion server port |
|
||||||
| `SUPERVISOR_CONFIG_DIR` | `./config/supervisor` | Skill discipline files |
|
| `INGEST_BASE_URL` | `http://localhost:3300` | Routing pod → brain |
|
||||||
| `SUPERVISOR_SESSIONS_DIR` | `./brain/sessions` | JSONL session logs |
|
|
||||||
| `INGEST_BASE_URL` | `http://localhost:3300` | Supervisor → ingestion |
|
|
||||||
| `LITELLM_BASE_URL` | — | LiteLLM proxy for Tier 2 model routing |
|
| `LITELLM_BASE_URL` | — | LiteLLM proxy for Tier 2 model routing |
|
||||||
| `SUPERVISOR_MCP_TOKEN` | — | Optional bearer token for the supervisor MCP HTTP endpoint; when empty, no auth is enforced |
|
|
||||||
| `ROUTING_PORT` | `3210` | Routing pod's listen port |
|
| `ROUTING_PORT` | `3210` | Routing pod's listen port |
|
||||||
| `ROUTING_MCP_TOKEN` | — | Optional bearer token for the routing MCP HTTP endpoint |
|
| `ROUTING_MCP_TOKEN` | — | Optional bearer token; when empty, no auth enforced |
|
||||||
| `BRAIN_URL` | `http://ingestion.supervisor:3300` | Routing pod → brain (in-cluster) |
|
| `BRAIN_URL` | `http://ingestion.supervisor:3300` | Routing pod → brain (in-cluster) |
|
||||||
| `HYPERGUILD_FAST_MODEL` | `koala/qwen35-9b-fast` | Fast model for high-pass-rate skill calls |
|
| `HYPERGUILD_FAST_MODEL` | `koala/qwen35-9b-fast` | Fast model for high-pass-rate skill calls |
|
||||||
| `HYPERGUILD_THINKING_MODEL` | `iguana/gemma4-26b` | Thinking model for low-pass-rate skill calls |
|
| `HYPERGUILD_THINKING_MODEL` | `iguana/gemma4-26b` | Thinking model for low-pass-rate skill calls |
|
||||||
| `HYPERGUILD_ROUTE_LOCAL_FLOOR` | `0.90` | At/above pass rate, route to fast model |
|
| `HYPERGUILD_ROUTE_LOCAL_FLOOR` | `0.90` | Fast model threshold |
|
||||||
| `HYPERGUILD_ROUTE_LOCAL_CEIL` | `0.70` | Below pass rate, route to thinking model. Between CEIL and FLOOR is the sample band. |
|
| `HYPERGUILD_ROUTE_LOCAL_CEIL` | `0.70` | Thinking model threshold |
|
||||||
| `HYPERGUILD_PASS_RATE_TTL_SECONDS` | `60` | Per-skill pass-rate cache TTL |
|
| `HYPERGUILD_PASS_RATE_TTL_SECONDS` | `60` | Per-skill pass-rate cache TTL |
|
||||||
|
|
||||||
> **Operator note:** LiteLLM at `LITELLM_BASE_URL` must register both `HYPERGUILD_FAST_MODEL` and `HYPERGUILD_THINKING_MODEL` for routing to do useful work. If a model is missing, LiteLLM returns 4xx, the routing pod's fast route fails, the fail-open retry on the thinking model likely also fails (since both are missing), and the only signal is `final_status: "fail"` on `_routing` entries in the brain.
|
> **Operator note:** LiteLLM at `LITELLM_BASE_URL` must register both `HYPERGUILD_FAST_MODEL`
|
||||||
|
> and `HYPERGUILD_THINKING_MODEL`. If a model is missing, the fail-open retry also fails and
|
||||||
|
> the only signal is `final_status: "fail"` on `_routing` entries in the brain.
|
||||||
|
|
||||||
## Phase 2 (planned)
|
## Open issues
|
||||||
|
|
||||||
- `review` skill — structured code review with iron law enforcement
|
See [issues](http://gitea.d-ma.be/mathias/hyperguild/issues) — key open items:
|
||||||
- `debug` skill — hypothesis-driven debugging sessions
|
|
||||||
- `spec` skill — generates specs from conversations
|
- **#25** — skills platform overhaul (audit first, then lazy loading + brain feedback loop)
|
||||||
- `trainer` — extracts SFT/DPO pairs from session logs for fine-tuning
|
- **#24** — reduce context burn from skill listing
|
||||||
|
- **#22** — migrate legacy brain notes to wing/hall layout (one-shot script, low risk)
|
||||||
|
- **#31** — connect routing-mcp to claude.ai as custom connector
|
||||||
|
|||||||
Reference in New Issue
Block a user