A doc with no headings and no blank-line paragraphs (JSON-lines, e.g. wiki/telos/decisions/human-intent-column.md) survived both chunk passes whole and was sent to nomic-embed over its context window → 'input length exceeds the context length' (400, the steady embed errors=1). Add a final hard-split pass (line then UTF-8 rune boundaries) so no chunk exceeds maxBytes. TDD. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
hyperguild
An MCP server that acts as a disciplined AI supervisor for Claude Code sessions. Instead of letting Claude Code do whatever it wants, hyperguild enforces structured workflows (TDD red/green/refactor), logs every session, and accumulates learnings into a searchable brain.
Hypothesis
We believe routing skill tasks through local models, backed by brain context, produces measurably better outcomes than raw Claude Code alone — measurable by per-skill pass rate over rolling 30-day windows (available at
GET /pass-rate?skill=<name>&window=30don the brain pod).
This is the falsifiable claim the routing pod and pass-rate infrastructure exist to test. If per-skill pass rates don't improve over baseline (all-cloud) after 30 days of real usage, the fast-model routing path should be reconsidered.
Harness
hyperguild = Claude Code + MCP. This is a supervisor for Claude Code sessions specifically. For multi-agent orchestration (OpenCode + LiteLLM, executor/reviewer pipelines), see agentsquad — a separate harness for a different orchestration model. Skills (mathias/skills) are shared between both.
How it works
Your Claude Code session (in any project)
│
│ MCP over HTTP (Tailscale)
├──▶ routing :3210 (NodePort 30310 on koala) — review, debug, retrospective, trainer
└──▶ brain :3300 (NodePort 30330 on koala) — brain_query, brain_write, brain_ingest, session_log
│
└─ also serves the legacy REST endpoints (/query, /write, /ingest, …)
│
▼
brain/
├── sessions/ — JSONL log, one file per session_id
├── wiki/ — searchable knowledge (wing/hall layout)
│ ├── homelab/
│ ├── claude-sessions/
│ └── ...
└── knowledge/ — legacy flat notes (migration pending: hyperguild#22)
Phase 1 tools (available now)
| Tool | What it does |
|---|---|
session_log |
Appends a structured entry to the session JSONL log |
retrospective |
Reads the session log, identifies novel learnings, writes to brain |
review |
Structured code review via local model, brain-context injected |
debug |
Hypothesis-driven debugging via local model |
brain_query |
Full-text search over brain/wiki/ |
brain_write |
Writes a note to brain (with wing/hall routing) |
brain_answer |
BM25 + LLM synthesis — Q&A over brain corpus |
tier |
Returns the current connectivity tier (1=cloud, 2=LAN, 3=offline) |
Note:
tdd_red/green/refactorandspecwere retired in Plan 7 (2026-05-12). They are now SKILL.md files in mathias/skills.
Connect a project
Create .mcp.json in your project root:
{
"mcpServers": {
"routing": {
"type": "http",
"url": "http://koala:30310/mcp"
},
"brain": {
"type": "http",
"url": "http://koala:30330/mcp"
}
}
}
Two MCP servers are exposed, both reachable over Tailscale:
routingatkoala:30310— skill workers (review,debug,retrospective,trainer). Routes each call to fast local model or thinking model based on per-skill pass rate.brainatkoala:30330— knowledge access (brain_query,brain_write,brain_ingest,brain_ingest_raw,brain_answer,brain_classify) andsession_log.
No local binary or stdio shim is required — Claude Code talks to both via HTTP.
Open Claude Code in your project — run /mcp to confirm both servers are listed.
A typical session
1. Call review → brain context injected + local model review → findings
2. Call session_log → log each phase result
3. Call retrospective → extracts learnings → brain
4. Future sessions: call brain_query / brain_answer to retrieve relevant context
Tier detection
The routing pod probes connectivity at call time:
| Tier | Label | Condition |
|---|---|---|
| 1 | full-online | Can reach api.anthropic.com |
| 2 | lan-only | Can reach LiteLLM but not Anthropic |
| 3 | airplane | No external connectivity |
Model routing
The routing pod selects models per skill call based on historical pass rate:
| Pass rate | Decision |
|---|---|
| ≥ 0.90 (FLOOR) | Fast model (HYPERGUILD_FAST_MODEL) |
| ≤ 0.70 (CEIL) | Thinking model (HYPERGUILD_THINKING_MODEL) |
| between CEIL and FLOOR | Sample band — probabilistic routing |
| nil (no history yet) | Defaults to thinking model |
Bootstrap note: With no session history, all calls route to the thinking model. The fast-model path activates only after real pass-rate data accumulates at
/pass-rate. Seed with real usage — don't try to pre-populate.
Key env vars
| Variable | Default | Purpose |
|---|---|---|
INGEST_BRAIN_DIR |
../brain |
Brain directory for ingestion server |
INGEST_PORT |
3300 |
Ingestion server port |
INGEST_BASE_URL |
http://localhost:3300 |
Routing pod → brain |
LITELLM_BASE_URL |
— | LiteLLM proxy for Tier 2 model routing |
ROUTING_PORT |
3210 |
Routing pod's listen port |
ROUTING_MCP_TOKEN |
— | Optional bearer token; when empty, no auth enforced |
BRAIN_URL |
http://ingestion.supervisor:3300 |
Routing pod → brain (in-cluster) |
HYPERGUILD_FAST_MODEL |
koala/qwen35-9b-fast |
Fast model for high-pass-rate skill calls |
HYPERGUILD_THINKING_MODEL |
iguana/gemma4-26b |
Thinking model for low-pass-rate skill calls |
HYPERGUILD_ROUTE_LOCAL_FLOOR |
0.90 |
Fast model threshold |
HYPERGUILD_ROUTE_LOCAL_CEIL |
0.70 |
Thinking model threshold |
HYPERGUILD_PASS_RATE_TTL_SECONDS |
60 |
Per-skill pass-rate cache TTL |
Operator note: LiteLLM at
LITELLM_BASE_URLmust register bothHYPERGUILD_FAST_MODELandHYPERGUILD_THINKING_MODEL. If a model is missing, the fail-open retry also fails and the only signal isfinal_status: "fail"on_routingentries in the brain.
Open issues
See issues — key open items:
- #25 — skills platform overhaul (audit first, then lazy loading + brain feedback loop)
- #24 — reduce context burn from skill listing
- #22 — migrate legacy brain notes to wing/hall layout (one-shot script, low risk)
- #31 — connect routing-mcp to claude.ai as custom connector