# mcphost — full documentation

> Generated by scripts/gen-llms-full.sh from README.md, docs/kinds/*.md,
> docs/benchmarks/*.md and docs/receipts/*.md. See /llms.txt for the short
> index. Do not hand-edit.

## README

# mcphost

where agents host their own tools · [mcphost.dev](https://mcphost.dev) · [status](https://mcphost.dev/status.html) · [llms.txt](https://mcphost.dev/llms.txt)

<!-- agent-quickstart:start -->
Ship an MCP tool, not a deployment project.

mcphost lets an agent create the tool it needs, mid-task, without a human
in the loop: sign up with one unauthenticated tool call, publish with the
next, and the new tool is live immediately — no restart, no deploy, no
review queue.

**Measured** (panel run `0.26.3-20260908T085001Z`, 21 sessions): median
time from signup to a tenant's first successful `host.tool_publish` is
**30.7s**; median time from signup to a successful call on that tenant's
own tool is **42.4s**.
<!-- cite: docs/benchmarks/measure-0.26.3-20260908T085001Z.md -->

## Quickstart for agents

1. Connect to the endpoint and call `tools/list` with no credentials. The
   only tool offered is `signup`.
2. Call `signup(name)`. The response contains `tenant`, `key` (a bearer
   token, shown once), `namespace`, and `endpoint`. Signup is rate-limited
   to 5 per IP per hour. Pass `source` (e.g. `signup(name, source: "hn")`)
   to tag which channel this signup came from -- recommended values are
   `hn`, `reddit`, `discord`, `registry`, `plugin`, `docs`; it's echoed
   back in the response and broken out in admin healthz, but never
   required. If signups are paused (an operator's kill switch for an
   abuse spike), the call fails with `signup_paused` and a
   `retry_after_secs`; try again later.
   Recommended: `signup(name, handoff: true)` returns a short-lived,
   single-use `handoff_token` instead of `key`; call
   `host.redeem(handoff_token)` once to get the key, so a transcript of
   this exchange carries a dead credential. `host.key_rotate` invalidates
   the current key and issues a new one in one call, any time you suspect
   it leaked.
3. Pass `key` as the `tenant_key` argument on every `host.*` call from
   here on -- e.g. `host.tool_publish`, `host.tool_call`. No reconnect or
   `Authorization` header needed; a client that holds a persistent
   connection can use `Authorization: Bearer <key>` instead.
4. Publish a tool: `host.tool_publish(name, kind, spec)`. Call
   `host.quickstart` first — its `starter_tool` is a ready-to-publish
   `python` spec (reverses text, counts words) plus the exact
   `publish_call`/`test_call` to run; the documented first publish is a
   real tool, not a stub. Two real kinds: submit code (`python` — source
   required, `args_schema`/`requirements` inferred if omitted) or wrap an
   API you already use (`http` — url and method required, `args_schema`
   inferred if omitted). `echo` (returns its arguments; spec is a JSON
   Schema) is a stub for testing the pipes, not a real tool — it carries
   `stub: true` in `tools/list`. Before publishing anything, dry-run with
   `host.tool_publish({..., dry_run: true})` — every gate (secrets, env,
   network, deps, name, kind, spec size) reported at once, no tool row
   written — or `host.spec_test(kind, spec, invocations)` for up to 5
   example calls through the same sandbox a real call uses.
5. Call your tool. Two equivalent ways over the same streamable-HTTP
   connection: as `<namespace>.<tool_name>` (its own entry in
   `tools/list`), or `host.tool_call(name, args)` (same dispatch path,
   useful when your client doesn't refresh `tools/list` between publish
   and call). `host.tool_test(name, args)` dry-runs an already-published
   tool by name instead of a raw spec.
6. Check the plan and quota before you rely on volume:
   `billing.plans()` — the plan catalog, works anonymously.
   `billing.status()` — this tenant's plan and usage against each quota.
7. Inspect and manage: `host.tool_list()`, `host.tool_logs(name)`,
   `host.tool_remove(name)`, `host.usage(window)`,
   `host.secret_set`/`host.secret_list()` (secrets stored AES-256-GCM
   encrypted). A `python` spec's plain, non-secret configuration lives in a
   separate `env` map (up to 16 entries / 4 KiB total, names matching
   `^[A-Z][A-Z0-9_]{0,63}$`) — shown verbatim in `host.tool_test`, unlike
   `secrets`, which stay redacted there.
8. Share a tool with `host.tool_share(name, visibility, group?)` (see "Share
   a tool, not a key" in `www/llms.txt` for the full recipe). Pass
   `expose_spec: true` to also let every sharee read the tool's source, not
   just call it — the point when you want others to fork what you built,
   the way `visions/synthorg-compete.md`'s round-two builders fork
   round-one winners. A sharee reads it with
   `host.tool_spec_shared(tool: "<owner_namespace>.<name>")`, which returns
   `{tool, kind, spec, exposed_at}` — `spec` never carries `env` or a secret
   reference, only `source`/`args_schema`/`requirements`/`timeout_s`/
   `network`. Worked example, after step 4 published `nightly_scrape` as a
   `python` tool:
   ```
   host.group.create(name="arena-builders")
   host.group.add(name="arena-builders", namespace="<their_namespace>")
   host.tool_share(name="nightly_scrape", visibility="group",
                    group="arena-builders", expose_spec=true)
   ```
   A group member then reads it (never through `host.tool_call`, which only
   runs it) with:
   ```
   host.tool_spec_shared(tool="<your_namespace>.nightly_scrape")
   # -> {"tool": "<your_namespace>.nightly_scrape", "kind": "python",
   #     "spec": {"source": "...", "args_schema": {...}}, "exposed_at": "..."}
   ```
   Shared without `expose_spec` (the default), the same call fails with
   `spec_not_exposed`; not shared with them at all, it fails exactly like
   `host.tool_call` would — `tool_not_found`, never revealing the tool
   exists.

<!-- cite: docs/benchmarks/measure-0.26.3-20260908T085001Z.md -->

## Leaving

`host.self_offboard()` permanently closes your own tenant: no operator
ticket, no admin key, no arguments, no confirmation flag — the same
`tenant_key` channel you signed up through is the one you leave through,
and the call takes effect immediately.

What it does, in order: cancels any active Stripe subscription if you're
on the `pro` plan, disables the tenant (every call with this `tenant_key`
after this point — `host.*` or `billing.*` — gets the same
`tenant_disabled`/`tenant_key_invalid` error an admin-disabled tenant
already gets), then gives every registered kind a chance to tear down
whatever it's keeping alive for you (e.g. a `python` sandbox's warm pool).

What it does NOT do: scrub your data. `tools`, `secrets`, `signup_events`,
and usage history all stay in place for audit — exactly the same
retention an admin-disabled tenant gets today. If you want a copy of what
you built before leaving, run `host.export()` first; `self_offboard`
doesn't bundle one for you.

Irreversibility: calling it twice is a no-op, not an error or a crash —
but there is no self-service undo. A canceled Stripe subscription stays
canceled, and your `tenant_key` stops authenticating the instant the call
returns. (An operator can flip the underlying tenant row back on with
`admin.tenant_enable`, but that's an operator action taken on your behalf,
not something `self_offboard` itself offers back to you.)

## Getting help

Every error payload carries `code`, a clean `message`, a `request_id`, and
(for every code in the table below) a `help_url` pointing at a generated
`/help/<code>` page -- meaning, likely cause, fix, no internal text.
`host.whoami`'s `links` field names the same `support`/`plans`/`status`/
`help` pages directly, so an agent never has to guess the host to build
them from.

<!-- support:start -->
Support: support channel not configured (MCPHOST_SUPPORT_URL is unset).
<!-- support:end -->

## Contributing: naming a new `host.*` tool

mcphost-polish-p0-20260930 (audit finding 5): the registry mixes
`host.<namespace>.<verb>` (dotted — `host.agent.lookup`, `host.docs.put`,
`host.group.create`, `host.channel.post`, `host.msg.send`,
`host.lineage.trace`, `host.enduser.revoke`, ...) with
`host.<noun>_<verb>` (underscored — `host.key_rotate`, `host.bridge_test`,
`host.self_offboard`, ...) with no rule written down anywhere, which is
real drift, not two equally-valid styles. Reading the registry as it
stands, the actual pattern almost every tool already follows is: **use a
dot when the tool is one of two-or-more siblings sharing a resource**
(another `host.agent.*`/`host.docs.*`/`host.group.*`/... call already
exists or will exist alongside it) **and underscore only inside one
segment's own name** (`host.lineage.blast_radius`,
`host.enduser.assertion_secret_rotate`) **or for a genuine one-off with no
sibling family** (`host.export`, `host.key_rotate`). The one named
exception is the `host.tool_*` sharing/publish family
(`host.tool_call`/`host.tool_share`/`host.tool_list`/`host.tool_remove`/
`host.tool_logs`/`host.tool_rollback`/`host.tool_diff`/`host.tool_history`/
`host.tool_unshare`) — a real multi-verb resource family that, by the rule
above, "should" be dotted (`host.tool.call`, ...) but predates it and is
not being renamed (a rename breaks every existing caller for a
cosmetic fix). Do not use `host.tool_*`'s underscore style as a template
for a *new* multi-verb family — follow `host.agent.*`/`host.docs.*`
instead. A PRD to actually reconcile `host.tool_*` with the dotted style
(alias + deprecation window, not a breaking rename) is tracked separately.

## Contributing: routing a new top-level path

Adding a new top-level directory or file to this repo (like `www/`,
`deploy/`, or `.buildloop/`) needs a matching `[[lane]]` entry in
`agent/proof-lanes.toml`, in the same PR that adds the path — `autobuilder
vti-plan` (the branch gate's routing check) refuses any changed path that
resolves to zero lanes, and `tests/lanecov_ac01_every_tracked_path_routes.rs`
enforces the same rule locally via `cargo test`, so a missing lane fails
fast instead of turning every branch gate red after the path lands on
`main`. Give the lane an `id`, a one-line `description` naming the PRD or
reason the path exists, `globs` covering the new path (`"<dir>/**"` for a
directory), and `required_commands` — the cheapest command that actually
proves a change under that path, not necessarily the full test suite. The
`loop-config` lane (routing `.buildloop/**` to `cargo test --workspace`) is
a worked example: one glob, one required command, added in the same PR
that made the path matter.

## Docs Q&A in a minute

Turn a folder of markdown into a checkable Q&A tool, end to end, in one
script:

1. `signup(name)` — one fresh tenant.
2. `host.docs.put(name, content)` — once per document (this recipe's own
   corpus: 8 documents, ~22 KiB total, well under the 2 MiB per-document
   cap).
3. `host.docs.status()` — poll until `index.lag_seconds == 0` (usually one
   or two of the indexer's own 10s ticks; this recipe's corpus is ready
   well under 30s).
4. `host.tool_publish(name="ask_docs", kind="python", spec={"source": ...})`
   — the one published tool. Its `main` calls `mcphost.docs.search(query,
   k)` over the same zero-network sandbox loopback `mcphost.docs.get`
   already uses, so answering a question never spends a public tool call.
5. `<namespace>.ask_docs(query="...")` — ask it. Every passage in the
   result carries a `name:offset` citation (the document's name and its
   character offset in that document), so an answer is checkable against
   its source.

Quota this recipe uses: 8 documents, ~22 KiB, ~30 chunks, 1 published
tool, up to 10 calls to that tool — comfortably inside the free plan's
`docs_max` (200) and `calls_per_day` (500).

`examples/docs-qa/docs-qa.sh <endpoint>` runs this recipe end to end
against a real endpoint and writes a receipt (per-question hit and
citation, `index_ready_secs`, the quota actually used);
`examples/docs-qa/ask_docs.py` is the tool itself, and
`examples/docs-qa/corpus/` is the 8-document corpus plus its 6 gold
questions. `docs-qa.sh --embeddings <provider-endpoint> <model>
<secret-name>` configures an embeddings provider first and reports both
lexical and embeddings hit rates in one receipt.
<!-- agent-quickstart:end -->

## What it is

`mcphost serve` is one Rust binary that speaks streamable-HTTP MCP at a single endpoint. An agent signs up with one unauthenticated tool call, gets a namespace, and from then on everything is a tool call: publish, run, schedule, log, meter, store secrets, share with another agent, act as one of your end users over OAuth. There is no dashboard; `/status` is the one page, and the operator works through `admin.*` tools on the same endpoint.

<!-- cite: docs/benchmarks/ac11-load-smoke.txt -->
Measured, not promised: p95 34.2 ms across 200 concurrent calls with zero errors, committed next to the test that produces it. The public site is [mcphost.dev](https://mcphost.dev); the machine-readable summary at [`/llms.txt`](https://mcphost.dev/llms.txt) is generated from the same source as the quickstart above (`docs/agent-quickstart.md`, `scripts/gen-agent-docs.sh`).

## Connect

Paste one line into your client and it has mcphost. No key needed to sign
up -- `signup` is the one unauthenticated tool; everything past it takes
the bearer key `signup` returns.

**Claude Code**

```
claude mcp add --transport http mcphost https://mcphost.dev/mcp
```

**Codex CLI**

```
codex mcp add mcphost --url https://mcphost.dev/mcp
```

**Cursor** -- add this block to `mcp.json`:

```json
{
  "mcpServers": {
    "mcphost": {
      "url": "https://mcphost.dev/mcp",
      "headers": { "Authorization": "Bearer <key>" }
    }
  }
}
```

**Claude.ai** -- Settings -> Connectors -> Add custom connector, then paste
`https://mcphost.dev/mcp` as the URL.

See the live [status page](/status.html) and the
[Acceptable Use Policy](/aup.html) before you point production traffic at
it. Plans and limits: [plans](/plans.html) / [`/plans.json`](/plans.json)
(same numbers `billing.plans` returns, and `docs/plans.md`). Every error
payload carries a `help_url` pointing at a generated `/help/<code>` page
explaining it.

<!-- support:start -->
Support: support channel not configured (MCPHOST_SUPPORT_URL is unset).
<!-- support:end -->

## Changes

Current release: v0.64.0 (2026-09-30); `main` is 0.65.0. Every change is in [`CHANGELOG.md`](CHANGELOG.md) and the git tags.

## Operating and contributing

Everything below is for running your own mcphost or changing this one: install, environment, kinds, limits, metering, synthetic tenants, and the acceptance suite.

## Install

```
cargo install --path .
```

Or build locally:

```
cargo build --release
./target/release/mcphost serve
```

### Environment contract

| Variable | Meaning | Default |
|---|---|---|
| `MCPHOST_DATA_DIR` | Directory holding `mcphost.db` (SQLite, WAL) | `./data` |
| `MCPHOST_BIND` | `host:port` to listen on | `127.0.0.1:8080` |
| `MCPHOST_PUBLIC_URL` | URL returned by `signup` as the endpoint | `http://<bind>` |
| `MCPHOST_ADMIN_KEY` | Bearer key that unlocks `admin.*` tools | unset (admin tools unreachable) |
| `MCPHOST_SECRET_KEY` | Passphrase, SHA-256-derived into an AES-256 key for tenant secrets | dev default (set a real one in production) |
| `MCPHOST_LOG_LEVEL` | `tracing` filter, e.g. `info` | `info` |
| `MCPHOST_REGISTRY_URL` | Enables `host.registry_publish` (P1) and names the registry API's base URL; `mcphost serve --registry-url <url>` takes precedence | unset (registry-publish disabled) |
| `MCPHOST_SIGNUP_RATE_LIMIT_PER_HOUR` | Overrides the per-source-IP `signup` rate limit (PRD-mcphost-signup-rate-configurable) — raise it for a many-session measure run from one IP; absent or non-integer falls back to the default. Effective value is logged once at startup. A source IP configured in `MCPHOST_FLEET_IPS` is exempt from this limit entirely | `5` |
| `MCPHOST_EGRESS_PROXY` | `http(s)://host:port` of the operator's outbound HTTP(S) proxy. Required for a `pro` tenant's `python`/`wasm` tool published with `network: "public"` or `"egress"` to get any sandbox network at all — see "Egress proxy" below | unset (no `pro` tenant gets outbound network) |
| `MCPHOST_FLEET_IPS` | Comma-separated list of IPv4/IPv6 addresses and/or CIDR blocks (e.g. `46.225.110.44,178.105.64.66,10.0.0.0/8`) this operator's own fleet signs up from. A signup whose source IP matches gets `source_class: fleet` (`synthetic: harness:fleet-ip`) even without the `x-mcphost-synthetic` header — invalid entries are logged and skipped. See `admin.reclassify_fleet_ips` to backfill signups that predate this var | unset (no IP is ever classified `fleet` by address alone) |

`mcphost migrate` applies pending SQL migrations and exits. `mcphost version`
prints the version and exits. `mcphost serve --registry-url <url>` is the
CLI-flag form of `MCPHOST_REGISTRY_URL` above.

### Egress proxy (`network: "public"` / `"egress"`)

A `free` tenant can never publish a tool with `network: "public"` or
`"egress"` — both spellings grant the same sandbox access, and both are
refused at `host.tool_publish` with `plan_required` naming `pro` and the
`network` field. A `pro` tenant *may* publish one, but the sandbox still
gets no outbound network at call time unless `$MCPHOST_EGRESS_PROXY` is
configured on this host — with it unset, the call fails with
`egress_unavailable` before any sandboxed process is even spawned. An
existing tool that declared `public`/`egress` while its owner was on `pro`
starts failing with `plan_required` (not a silent downgrade to no network)
the moment that tenant drops to `free`; the error names `network: "none"` as
the republish fix, or upgrading back to `pro`.

With the proxy configured, a `pro` tenant's egress call runs with
`--share-net` (its outbound network shares the host's own network
namespace) and `http_proxy`/`https_proxy`/`HTTP_PROXY`/`HTTPS_PROXY` all set
to `$MCPHOST_EGRESS_PROXY` in the sandboxed process's environment, plus
`no_proxy=""` so a stray `NO_PROXY` already in the operator's environment
can never let a sandboxed process route around the proxy. This crate does
not ship the proxy itself (`mcphost-deploy` owns that config) — it only
guarantees "no proxy, no network." Whatever proxy is deployed **must**
enforce a private/link-local deny list on every connection it forwards, the
same ranges the `http` kind's own `VettingResolver` already refuses for a
rendered request URL:

- RFC 1918 private ranges (`10.0.0.0/8`, `172.16.0.0/12`, `192.168.0.0/16`)
- `169.254.0.0/16` (link-local, including the cloud metadata address
  `169.254.169.254`)
- `127.0.0.0/8` (loopback) and `::1`
- `fc00::/7` (IPv6 unique local addresses)

### Registry publish (P1)

Off by default. Once `--registry-url` / `$MCPHOST_REGISTRY_URL` names a
registry API base (e.g. `https://registry.modelcontextprotocol.io`):

1. The operator verifies a tenant's domain namespace by whatever method
   they trust (the PRD leaves the verification METHOD itself — DNS vs
   HTTP record — as an open question owned by Joe; this crate does not
   implement one) and records the outcome with `admin.tenant_verify_namespace`:
   `admin.tenant_verify_namespace(tenant="t_xxxxxxxx", domain_namespace="io.github.example.myserver")`.
2. That tenant can then call `host.registry_publish()` (no arguments): it
   POSTs a `server.json` document (`name`/`description`/`version`/`remotes:
   [{type: "streamable-http", url}]`) to `<registry-url>/v0/publish`, and
   the same document becomes servable, unauthenticated, at
   `GET /.well-known/mcp/<namespace>/server.json`.
3. `host.registry_publish` refuses with a distinct, machine-readable error
   in `data.error_code`: `registry_disabled` (flag off),
   `namespace_unverified` (step 1 not done for this tenant), or
   `registry_rejected` (the registry API answered non-2xx).

## Kinds

Every registered kind's minimal example spec, below, and `host.tool_publish`'s
on-wire description (visible from `tools/list` before signup) are both
rendered from the same `docs/kinds/*.md` files (PRD-mcphost-publish-first-try
requirement 6) -- `tests/publishfirsttry_ac06_docs_shared_source.rs`
regenerates this section from those files and fails CI if it's drifted from
what's checked in below. Call `host.quickstart(kind)` for the same example
with your own namespace already filled in.

<!-- kinds:start -->
### `echo`

spec.schema is any JSON Schema; a call echoes back the arguments it was given, validated against it.

Example spec:

```json
{
  "schema": {
    "properties": {
      "msg": {
        "type": "string"
      }
    },
    "required": [
      "msg"
    ],
    "type": "object"
  }
}
```

Example call arguments:

```json
{
  "msg": "hi"
}
```

### `http`

url must be an absolute https URL; method and url are the only required fields -- args_schema is inferred from the url/header/body templates when omitted.

Example spec:

```json
{
  "method": "GET",
  "url": "https://api.example.com/items/{{id}}"
}
```

Example call arguments:

```json
{
  "id": "123"
}
```

### `python`

only source is required -- args_schema and requirements are both inferred from it (tool-infer, v0.4.0); source must define main(args).

Example spec:

```json
{
  "source": "def main(args):\n    return {\"doubled\": args[\"n\"] * 2}\n"
}
```

Example call arguments:

```json
{
  "n": 3
}
```

### `wasm`

component is a base64-encoded WebAssembly component (component-model, not a core module) exporting `call: func(args: string) -> result<string, string>`; args_schema is optional (defaults to accepting any object).

Example spec:

```json
{
  "component": "AGFzbQEAAAAA"
}
```

Example call arguments:

```json
{
  "msg": "hi"
}
```
<!-- kinds:end -->

### Python spec-language notes

PRD-mcphost-python-kind-runtime (AC6): the AST-check that gates
`host.tool_publish` accepts assignment expressions (`:=`, PEP 572) in
general -- CPython has parsed them since 3.8, and mcphost's publish-time
check and the tool's own runtime both compile `source` with the same
CPython grammar, so there is no mcphost-added restriction to relax. The
one thing that *is* rejected is a restriction Python's own grammar
enforces: an assignment expression's target must be a plain name.
`(obj.attr := 1)` and `(d[key] := 1)` are both invalid Python syntax
(`cannot use assignment expressions with attribute` / `...with
subscript`) and would fail identically whether or not mcphost validated
them first -- the tool's own `main(args)` would refuse to even parse.
Because this is executor-level, not validator-level, there is nothing for
mcphost to loosen; the fix here is that the publish-time rejection now
names the construct and the accepted alternative in one sentence (assign
to a plain name first, then set the attribute/subscript in a separate
statement) instead of leaving CPython's bare grammar message to speak for
itself.

## Call limits

PRD-mcphost-call-limits-honest: every limit here is the one the code
enforces -- `tests/limits_ac06_quickstart_docs_match_constants.rs` checks
this section and `www/llms.txt`'s "Limits and pricing" section against the
same constants `host.quickstart`'s `limits` object reads.

- **Call timeout**: 30 s by default, or your own `timeout_s` up to 60 s max
  -- a python spec that declares `timeout_s` gets exactly that deadline
  (bounded by the 60 s host maximum), not a shorter one applied silently
  underneath it. `call_timeout` names the deadline that actually applied.
- **Output size**: tool output at most 1 MiB. Over the cap returns
  `tool_output_too_large` naming `limit_bytes` and the `actual_bytes`
  produced, never a bare `tool_output_invalid` parse failure.
- **Request body**: at most 1 MiB (HTTP 413 over that -- see
  `tests/ac16_request_body_too_large.rs`; the "2 MiB" in that AC's own
  description is the oversized test payload used to *prove* the 1 MiB cap,
  not the cap itself).
- **Concurrency**: 20 concurrent calls host-wide; per tenant, 4 per tenant on the free plan
  (10 on pro). A refusal past your own tenant's cap is `capacity` with
  `scope: "tenant"` and a `retry_after_ms`; past the host-wide cap it's
  `scope: "host"`.
- **Sandbox process cap**: a python tool's sandbox allows at most 64 live
  processes; a fork past that fails with the structured
  `tool_process_limit`, not a silent hang or an opaque OS error.

## Metered overage (billing emit-meter)

PRD-mcphost-metered-overage: pro tenants' successful calls past the plan's
50,000 included calls/month bill themselves through Stripe's
`mcphost_tool_calls` meter and its graduated metered price. Set these
env vars from `~/.config/mcphost/stripe-objects.json` (unset means the
same v0.14.0 behavior -- no metering, no `meter_lag`):

- `MCPHOST_STRIPE_METERED_PRICE_ID` -- the metered price id `billing.checkout`
  attaches alongside the base price.
- `MCPHOST_STRIPE_METER_EVENT_NAME` -- defaults to `mcphost_tool_calls`.

Then run `mcphost billing emit-meter` on a timer (every five minutes is the
shipped default): it reads pro tenants' unemitted `ok` calls, POSTs one
Stripe meter event per tenant (chunked at 100 events/request), ledgers each
batch, and advances its own high-water mark only once every event in the
run has been accepted -- safe to rerun after a crash or a failed POST (see
`src/metering.rs`'s doc comment for the replay/idempotency contract).

Install the shipped systemd **user** units (`~/.config/systemd/user/`,
matching this host's other `mcphost-*` units):

```
cp deploy/mcphost-emit-meter.service deploy/mcphost-emit-meter.timer \
   ~/.config/systemd/user/
systemctl --user daemon-reload
systemctl --user enable --now mcphost-emit-meter.timer
```

`mcphost-emit-meter.service` reads `~/.config/mcphost/emit-meter.env` (via
`EnvironmentFile=-`, so a missing file is not an error) for
`MCPHOST_DATA_DIR` / `MCPHOST_STRIPE_SECRET_KEY` / the two vars above.
Both unit files pass `systemd-analyze verify --user` (AC9;
`tests/metering_ac09_deploy_units_verify.rs`).

`/healthz`'s `meter_lag` field (present only when `MCPHOST_STRIPE_METERED_PRICE_ID`
is set) is the count of pro-tenant `ok` calls still above the high-water
mark -- watch it for emission health at a glance.

## Synthetic tenants (`admin.*_synthetic`)

PRD-mcphost-synthetic-flag: a tenant a test harness creates carries a
free-form `synthetic` label (e.g. `synthorg:<run_id>`) from signup onward,
set by the harness sending `x-mcphost-synthetic: <label>` on its `signup`
call -- no behavior change, metadata only. `/healthz`'s `tenants_real` /
`tenants_synthetic` split, and `admin.tenants`' `synthetic` filter
(`true`/`false`/`all`, default `all`), read this column so
`synthorg candidates --measure` can exclude panel traffic from "real
tenant" evidence.

**Backfilling the existing census** (every tenant predates this column, so
all load with `synthetic: null` until tagged): use `admin.tenants_set_synthetic`,
previewed with `dry_run: true` before the `dry_run: false` that applies it.
The recipe this host's own census used:

```
admin.tenants_set_synthetic(name_like: 'joe-%',  label: 'operator',                    dry_run: true)
admin.tenants_set_synthetic(name_like: 'joe-%',  label: 'operator',                    dry_run: false)
admin.tenants_set_synthetic(name_like: '%',      label: 'synthorg:backfill-20260906',  dry_run: true)
admin.tenants_set_synthetic(name_like: '%',      label: 'synthorg:backfill-20260906',  dry_run: false)
```

Run the `joe-*` pass first -- the second call's broader `%` pattern would
otherwise overwrite those rows' label too, since a tenant re-tagged by a
later call simply gets the later label (there is no "already labeled, skip"
guard by design: retagging is how a label ever gets corrected). A single
tenant can be corrected at any time with `admin.tenant_set_synthetic(tenant,
label)` (`label: null` clears it).

## Acceptance

Every P0 acceptance criterion is paired with a real `cargo test` (integration
tests under `tests/` spin up the server on an ephemeral port against a temp
`$MCPHOST_DATA_DIR`), except AC11 which is hardware-dependent and is
recorded as a smoke result below.

### Sandbox suite: user namespace requirement

The `python` kind's sandbox suites (`tests/sandboxready_*`, `python_ac*`,
`infer_ac*`, `warmpool_ac*`, `ac17_kind_conformance`) spawn real `bwrap`/
`unshare` isolation and need unprivileged user namespaces
(`unshare --user --map-root-user -- true` must succeed) to run for real. If
your box denies that (Ubuntu's default AppArmor policy on some kernels, some
container runtimes), running `cargo test` fails loudly by design outside
CI, naming the fix: `sysctl kernel.unprivileged_userns_clone=1` on older
kernels, or `sysctl kernel.apparmor_restrict_unprivileged_userns=0` on
Ubuntu 24.04+. See `sandbox::require_user_namespaces_or_ci_skip`'s doc
comment for the full contract, and `.github/workflows/ci.yml` for how the
hosted CI runner grants the same capability (PRD-mcphost-ci-sandbox-coverage)
instead of silently skipping.

CI runs these suites as their own `sandbox` job, in parallel with the `gate`
job that carries static analysis and everything else — once the suites stopped
skipping, a single `cargo test --workspace` step measured 313–336 s against a
300 s budget <!-- cite: .github/workflows/ci.yml -->. Which SUITE BINARIES go
where is derived, not hand-listed: `scripts/ci-test-partition.sh core|sandbox`
classifies every `tests/*.rs` FILE by whether it touches the sandbox-execution
surface (PRD-mcphost-test-suite-consolidation moved the unit cargo links from
"one binary per file" to a handful of `tests/suite_<core|sandbox>_NN.rs`
binaries — see "Adding a test" below — so the partition is now file→suite,
not file→binary), and `check` proves the split is total and disjoint at both
levels. Both jobs then fail on any capability-skip in their log, so a file
filed into the wrong half turns CI red rather than passing vacuously.

### Adding a test

`tests/*.rs` stopped being cargo's unit of test-binary discovery
(PRD-mcphost-test-suite-consolidation, 2026-09-12): `autotests = false` in
`Cargo.toml`, plus a handful of generated `tests/suite_<core|sandbox>_NN.rs`
files that `#[path]`-include the real files, keep `target/debug/deps` from
holding one ~280 MB binary per test file. Every test keeps its own file, its
own name, and its AC pairing — only which BINARY it links into changed.

To add a test: drop `tests/<name>.rs` in as always (same naming convention:
`<prefix>_ac<N>_<description>.rs`, `mod common;` if it needs the shared
harness), then run `scripts/gen-test-suites.sh` to fold it into a suite (or
just let CI tell you — `scripts/gen-test-suites.sh --check`, wired into
`ci-test-partition.sh check`, fails naming the exact file if you forget). The
generator buckets by filename prefix, splits sandbox-needing files from
core-only ones first (so no suite ever mixes the two — see above), and
rewrites a lone top-level `mod common;`/`mod ci_sandbox_support;` line in your
new file to `use crate::common;`/`use crate::ci_sandbox_support;` (those
compile once per suite now, not once per file) — no other line changes.
Never hand-edit a `tests/suite_*.rs` file; it is fully regenerated.

Running a single test by name now takes one extra flag: `cargo test --test
suite_core_01 my_test_file:: -- --nocapture` (`cargo nextest run -E
'test(my_test_file::)'` works too, and needs no suite name at all). `cargo
test --test my_test_file` alone no longer resolves — that file isn't its own
cargo target anymore.

| AC | Requirement | Test |
|---|---|---|
| 1 (P0) | Unauthenticated `tools/list` shows only `signup`; response carries `MCP-Protocol-Version` | `tests/ac01_unauthenticated_lists_signup.rs` |
| 2 (P0) | `signup` returns key/namespace/endpoint; key stored only as a hash | `tests/ac02_signup_creates_hashed_tenant.rs` |
| 3 (P0) | Tenant `tools/list` shows `host.*` and no other tenant's tools | `tests/ac03_tenant_lists_control_plane_only.rs` |
| 4 (P0) | Publish, then list, then call round-trips | `tests/ac04_publish_list_and_call.rs` |
| 5 (P0) | Cross-tenant isolation: B can't see or call A's tool | `tests/ac05_cross_tenant_isolation.rs` |
| 6 (P0) | Remove a tool: omitted from list, `tool_not_found` on call | `tests/ac06_remove_tool.rs` |
| 7 (P0) | `host.usage`/`admin.usage` report calls + p50/p95 | `tests/ac07_usage_metering.rs` |
| 8 (P0) | `admin.tenant_disable` locks out a key; tenant key is `forbidden` on `admin.tenants` | `tests/ac08_admin_disable_and_forbidden.rs` |
| 9 (P0) | 6th signup/hour/IP is `rate_limited`, no tenant created | `tests/ac09_signup_rate_limit.rs` |
| 10 (P0) | Unregistered kind / invalid name / oversized spec each fail distinctly, nothing written | `tests/ac10_publish_validation_errors.rs` |
| 11 (P0, non-functional) | 200 concurrent `echo` calls, p95 < 50ms, 0 errors, RSS < 100MiB | `tests/ac11_load_smoke.rs` (`#[ignore]`d — hardware-dependent; run with `cargo test --release --test suite_core_01 ac11_load_smoke:: -- --ignored --nocapture`). Measured on the build box: **p95 = 34.20ms, 0 errors, RSS = 37.3MiB** <!-- cite: docs/benchmarks/ac11-load-smoke.txt --> |
| 12 (P0) | `synthorg consume --preflight <url>` exits 0 | `tests/ac12_preflight.rs` — an always-run in-process half exercises the same two requests `run_preflight` makes; a second half spawns the real `mcphost` binary and the real `synthorg` CLI when available (bare binary or `uv run --project`) and asserts exit 0 |
| 13 (P0) | Mismatched `Mcp-Name` header vs. body is recorded by body name and flagged | `tests/ac13_mcp_name_mismatch_metering.rs` |
| 14 (P0) | Unwritable database: `storage` error, `/healthz` `db_ok: false`, process stays up | `tests/ac14_storage_unwritable.rs` |
| 15 (P0) | A call that never completes times out at the deadline, future dropped | `tests/ac15_call_timeout.rs` |
| 16 (P0) | A request body over the 1 MiB cap (proven with a 2MiB body) is rejected with HTTP 413 | `tests/ac16_request_body_too_large.rs` |
| 17 (P0) | `Kind` conformance suite passes `echo`, fails naming `describe` for a bad schema | `tests/ac17_kind_conformance.rs` (reusable checker at `mcphost::kinds::conformance`) |
| 18 (P1) | `tools/list` carries `ttlMs`/`cacheScope`, `ttlMs: 0` within 60s of a publish | `tests/ac18_tools_list_ttl.rs` |
| 19 (P1) | `host.registry_publish()` + `/.well-known/mcp/<ns>/server.json` | `tests/ac19_registry_publish.rs` (mocks the registry API with `wiremock`; see "Registry publish (P1)" above — the namespace-verification METHOD stays out of scope, "verified" is an admin-set boolean) |

## Related fleet work

- [`mcp-core`](https://github.com/j0yen/mcp-core) — the reusable stdio
  JSON-RPC 2.0 MCP-server core (`Tool` trait + `serve_stdio`) other wintermute
  MCP servers build on. Not reused here: `mcphost` is a streamable-HTTP
  server (`rmcp`), not a stdio server, and its tool surface is dynamic
  (per-tenant, DB-backed) rather than the static `Tool` trait `mcp-core`
  wraps. Cited per the PRD's technical considerations as related, not shared,
  code.

## License

Dual-licensed under MIT OR Apache-2.0 — see `LICENSE-MIT` and
`LICENSE-APACHE`.

## Kind: chain

Runs a fixed, ordered sequence of this tenant's own tools, passing each step's mapped result into the next -- the declarative form of `mcphost.call` (see the `python` kind's own doc), for a pipeline that needs no python of its own.

```json
{
  "steps": [
    {"tool": "fetch_rows", "args": {"since": "$.input.since"}},
    {"tool": "write_rows", "args": {"rows": "$.prev.result.rows"}}
  ]
}
```

Call arguments:

```json
{"since": "2026-01-01"}
```

Each step names a `tool` (unqualified, same tenant) and an `args` object mapping that step's own call arguments -- each value is either a literal JSON value or (a string starting with `"$."`) a path resolved against `{"input": <the chain's own call args>, "prev": <the previous step's result, or null on the first step>, "steps": [<every earlier step's result, by 0-based index>]}`. The grammar is the same dotted/indexed `$.a.b[0].c` paths `outputs` accepts elsewhere in this host -- no wildcards, filters, or recursive descent. A step whose mapping resolves to nothing fails the whole call with `error_code: compose_mapping_missing`, naming the path and `failed_step`; add `"on_error": "continue"` to a step to let the remaining steps run anyway (the parent call still ends `error`-adjacent, but names every `failed_steps` index rather than stopping at the first).

The chain's own result is its last step's result, unless that step itself failed and every later step tolerated its own failure -- then it's the last step that actually ran. `host.tool_test` on a published chain dry-runs it: each step's resolved arguments are reported (`$.prev`/`$.steps[i]` mappings show as `unresolved_path` before anything has actually run) and no step executes.

A chain step goes through the same composition primitive `mcphost.call` uses (`kinds::compose_call`): a chain naming itself as one of its own steps is refused with `error_code: compose_self_call`; nesting chains (or python tools calling chains, or chains calling python tools that call further tools) more than 4 levels deep is refused with `error_code: compose_depth_exceeded`; a single top-level call's whole tree is capped at 50 child calls total (`error_code: compose_children_exceeded`).

## Kind: echo

spec.schema is any JSON Schema; a call echoes back the arguments it was given, validated against it.

```json
{
  "schema": {
    "type": "object",
    "properties": {"msg": {"type": "string"}},
    "required": ["msg"]
  }
}
```

Call arguments:

```json
{"msg": "hi"}
```

Dry-run before publishing: `host.spec_test("echo", spec, invocations)` runs the example invocations above through the same path a real call would use and returns each one's output verbatim -- no tool row is written.

## Kind: http

url must be an absolute https URL; method and url are the only required fields -- args_schema is inferred from the url/header/body templates when omitted.

```json
{
  "method": "GET",
  "url": "https://api.example.com/items/{{id}}"
}
```

Call arguments:

```json
{"id": "123"}
```

Dry-run before publishing: `host.spec_test("http", spec, invocations)` calls the wrapped endpoint through the same outbound path a real call would use and returns each invocation's status code and a bounded body excerpt -- no tool row is written.

Result envelope contract: an optional `outputs` array of field names (`"outputs": ["bridge_status"]`) declares fields a caller can rely on finding at `result.payload.<field>`, regardless of how deep the upstream body actually nests them -- one level of common wrapping (`data`, `result`, `response`) is searched automatically. `result.body` keeps the full, unmodified upstream response either way. Run `host.tool_test` after publishing to see any declared field your upstream never emits (`envelope.missing`, with `missing_detail` naming where else in the body that field name turned up).

`outputs` also accepts an object mapping each field name to the exact path to read it from, when the upstream buries it somewhere the wrapper search above won't find (`"outputs": {"ingestion_status": "$.json.ingestion_status"}` reads that field from `result.body.json.ingestion_status`). A path is `$` followed by dotted keys and bracketed integer indices only (`$.a.b[0].c`) -- no wildcards, filters, or recursive descent; an unsupported or malformed path is refused at publish, naming the field (`outputs.<name>`) and the accepted grammar.

## Kind: python

only source is required -- args_schema and requirements are both inferred from it (tool-infer, v0.4.0); source must define main(args).

```json
{
  "source": "def main(args):\n    return {\"doubled\": args[\"n\"] * 2}\n"
}
```

Call arguments:

```json
{"n": 3}
```

Dry-run before publishing: `host.spec_test("python", spec, invocations)` runs the example invocations above in the same sandbox a real call would use and returns each one's output or a bounded exception, plus the inferred args_schema and requirements -- no tool row is written.

Result envelope contract: an optional `outputs` array of field names (`"outputs": ["diagnosis"]`) declares fields a caller can rely on finding at `result.payload.<field>`, regardless of how deep `main`'s returned object actually nests them -- one level of nesting under any key is searched automatically. A tool that returns a bare scalar/list instead of an object gets the whole value promoted to the first declared field, with a `result.payload._envelope_warning` naming the scalar promotion. Run `host.tool_test` after publishing to see any declared field your tool never emits (`envelope.missing`, with `missing_detail` naming where else that field name turned up).

`outputs` also accepts an object mapping each field name to a path (`"outputs": {"score": "$.data.score"}`, the same `$.a.b[0].c` dotted/indexed grammar `http` accepts -- no wildcards, filters, or recursive descent); that path is read directly from `main`'s return value (PRD-mcphost-surface-fluidity), the same as `http`'s own path-declared fields -- a bare-name entry in the same `outputs` still falls back to the wrapper-search promotion above.

## mcphost.state (per-tenant memory)

`import mcphost` inside `source` and call `mcphost.state.get/set/delete/list/insert/query/delete_rows/table_create` -- the same store `host.state.*` reads and seeds from the agent's own session, scoped to this tenant, reachable with `network: none`:

```json
{
  "source": "import mcphost\ndef main(args):\n    n = mcphost.state.get(\"count\", 0) + 1\n    mcphost.state.set(\"count\", n)\n    return {\"count\": n}\n"
}
```

`mcphost.state.get(key, default=None)` returns `default` when the key was never set; every other function takes the same arguments as its `host.state.*` counterpart (`table_create(name, schema, primary_key=None)`, `query(table, where=None, order_by=None, limit=None)`, and so on) and raises `mcphost.state.StateError` (with `.code`/`.data`) on a quota or schema violation rather than returning an error value. A call's `mcphost.state` operations are attributed to that call and see its own writes immediately; there is no state visible across tenants.

## mcphost.docs (read a stored document)

`import mcphost` and call `mcphost.docs.get(id)` (or `mcphost.docs.get(name="a.md")`) to read a document from this tenant's `host.docs.*` store and get back its extracted plain text directly -- no `host.tool_call` round trip. `MCPHOST_DOCS_ENDPOINT` is set in the sandboxed process's environment whenever this channel is available, so a tool can check for it before importing:

```json
{
  "source": "import os, mcphost\ndef main(args):\n    if not os.environ.get(\"MCPHOST_DOCS_ENDPOINT\"):\n        return {\"error\": \"docs unavailable\"}\n    text = mcphost.docs.get(args[\"id\"])\n    return {\"text\": text}\n"
}
```

A failure (unknown id/name) raises `mcphost.docs.DocsError` (`.code`/`.data`, e.g. `error_code: docs_not_found`). `http`/`wasm` kinds don't get this channel; `python` reaches it the same `network: none` way `mcphost.state`/`mcphost.table` already do.

## mcphost.call (call another tool)

`import mcphost` and call `mcphost.call(name, args, timeout_s=None)` to run another tool in this same tenant as a child call, synchronously, and get its result back -- reachable with `network: none`, over the same channel `mcphost.state` uses:

```json
{
  "source": "import mcphost\ndef main(args):\n    rows = mcphost.call(\"fetch_rows\", {\"since\": args[\"since\"]})\n    return mcphost.call(\"write_rows\", {\"rows\": rows[\"rows\"]})\n"
}
```

`name` is the target tool's own (unqualified) name; `args` is its call arguments; an optional `timeout_s` bounds this one child call further, but never past the caller's own remaining deadline. The child's result comes back exactly as calling it directly would return it -- no envelope wrapper. A failure raises `mcphost.CallError` (`.code`/`.data`, e.g. `error_code: tool_exception` when the child itself raised); left uncaught, it surfaces on the caller's own call the same way any other unhandled exception does. A tool cannot call itself (`compose_self_call`), nesting is capped at 4 levels deep (`compose_depth_exceeded`), and a single call tree may make at most 50 child calls in total (`compose_children_exceeded`) -- every refusal names the limit it hit.

See the `chain` kind for the declarative form of the same idea: an ordered list of tool calls with no python of your own to write.

## env (plain configuration, distinct from secrets)

`"env": {"UPSTREAM_URL": "https://example.test", "MODE": "fast"}` puts plain, non-secret configuration into the sandboxed process's environment beside `secrets` -- an endpoint URL, a mode flag, a tenant identifier: anything that isn't a credential and doesn't need the secret store's encryption, rotation story, or `host.tool_test` redaction.

```json
{
  "source": "import os\ndef main(args):\n    return {\"mode\": os.environ[\"MODE\"]}\n",
  "env": {"MODE": "fast"}
}
```

Bounds, enforced at publish: at most 16 entries, at most 4 KiB total across every name and value combined, and a value must be valid UTF-8 with no NUL byte -- each violation is a structured error naming the key and the bound it broke. A name must match `^[A-Z][A-Z0-9_]{0,63}$`; `MCPHOST_*`, `PATH`, `HOME`, `PYTHON*`, `LD_*`, and `SECRET_*` (this kind's own prefix for injecting a resolved secret) are reserved and refused by name or prefix. An `env` name may never collide with a secret name already set for this tenant, in either direction: publishing `env` that collides with an existing secret is refused, and so is `host.secret_set`ing a secret whose name collides with an already-published tool's `env` entry.

The distinction is visible, not doctrinal: `host.tool_test` renders both in one listing, `env` values shown verbatim and secret values redacted, each labeled `"kind": "env"` or `"kind": "secret"`. `host.tool_list` returns a tool's `env` map to its owner; `admin.tool_list` reports env names and total size to an operator, never values. Updating `env` alone (no source change) takes effect on the tool's next call -- the warm pool re-fingerprints on an env change exactly as it does on a secret change, so a stale pooled process never serves old values.

## Kind: wasm

component is a base64-encoded WebAssembly component (component-model, not a core module) exporting `call: func(args: string) -> result<string, string>`; args_schema is optional (defaults to accepting any object).

```json
{
  "component": "AGFzbQEAAAAA"
}
```

Call arguments:

```json
{"msg": "hi"}
```

The `component` value above is a placeholder (a real one is typically 10-50 KiB base64, produced by a toolchain such as `cargo component build`) -- the shape is what matters: a base64 string decoding to a component-model binary. Dry-run before publishing: `host.spec_test("wasm", spec, invocations)` instantiates the component and runs the example invocations through the same fuel/memory/wall-time budget a real call would use -- no tool row is written.

## Limits

Enforced by Wasmtime itself, not by an external sandbox: `timeout_s` (1-30, default 5) bounds wall-clock time; `memory_mb` (1-256, default 64) bounds the component's linear memory -- exceeding either produces a structured `tool_timeout`/`tool_oom` error and the process serves the next call normally. A component binary over 47 KiB is refused at publish, naming the bound. An explicit trap (an `unreachable`, an out-of-bounds access) is reported as `tool_trapped`, with the trap's message in `host.tool_logs` -- never a raw panic.

## Choosing `wasm` over `python`

Choose `wasm` when a tool needs millisecond cold starts or must keep working on a box where the OS forbids the unprivileged user namespaces `python`'s sandbox depends on (isolation here travels with the compiled binary, not the host's userns/AppArmor policy); choose `python` when you want to author the tool as source on the host and let the host build its environment for you, since compiling a component is still the agent's own job (no server-side compile in this slice).

Result envelope contract: a call's output always lands at `result.payload`; an optional `outputs` array or object (the same shapes `http`/`python` accept, including `$.a.b[0].c`-style paths) declares fields promoted into `result.payload.<field>`, searched one level deep under any key in the component's own returned JSON -- identical semantics to the `python` kind's own promotion.

## Benchmark: measure-0.26.3-20260908T085001Z.md

# Measure run 0.26.3-20260908T085001Z

Mirrored from `evidence/mcp-host/measure/0.26.3-20260908T085001Z/ledger.jsonl`
(private `j0yen/prds` evidence repo, not shipped with this crate) so the
public README and `llms.txt` can cite a number that also exists inside
`j0yen/mcphost` itself — `scripts/copy-claims.sh` checks against this file,
not the private evidence repo, so the claim audit does not depend on a repo
that CI never checks out.

- `deployed_version`: `0.26.3`
- panel: 21 sessions, 7 buyer segments (`admin_agent`, `cost_optimizer`,
  `data_pipeline_builder`, `integration_specialist`, `rag_indexer`,
  `rapid_prototyper`, `workflow_orchestrator`), `cold`/`warm` conditions
- client: `claude-sdk` 2.1.251

| metric | n | median | unit |
|---|---|---|---|
| `t_first_publish` (signup → first successful `host.tool_publish`) | 21 | 30.72 | seconds |
| `t_first_own_call` (signup → first successful call on a tenant's own published tool) | 20 (1 session's tool never finished building; no `t_first_own_call`) | 42.40 | seconds |

Recomputed 2026-09-08 (PRD-mcphost-agent-findability build) directly from
the ledger's `t_first_publish` / `t_first_own_call` fields with a plain
median (sorted values, midpoint of the two central values on an even
count):

```
$ python3 -c "
import json
pub, call = [], []
for line in open('ledger.jsonl'):
    line = line.strip()
    if not line: continue
    d = json.loads(line)
    if d.get('t_first_publish') is not None: pub.append(d['t_first_publish'])
    if d.get('t_first_own_call') is not None: call.append(d['t_first_own_call'])
def median(xs):
    xs = sorted(xs); n = len(xs)
    return xs[n // 2] if n % 2 else (xs[n // 2 - 1] + xs[n // 2]) / 2
print('t_first_publish', len(pub), median(pub))
print('t_first_own_call', len(call), median(call))
"
t_first_publish 21 30.72465760691557
t_first_own_call 20 42.40137840353418
```

The problem statement's earlier figures (~23.6s / ~31.2s) came from an
older run; this file always names the run id it was mirrored from so a
future PRD updating the README also updates this file — `copy-claims.sh`
fails the build if the two drift apart from the citation without a
matching numeric mirror.

## Receipt: docsqa-ac06-synthorg-preflight.md

# AC6 cross-repo evidence — `docs_qa_recipe` in `~/repos/synthorg`

AC6: *Given `~/repos/synthorg` with the new task, When `synthorg consume
--preflight` and the proxy-tier fixture run execute, Then the
`docs_qa_recipe` task is listed for `rag_indexer` and the fake endpoint
fixture passes its gold predicate.*

`~/repos/synthorg` is a separate repository outside this worktree's own
`cargo test` gate (no `uv`/python toolchain on the runner box), so the
commands below were run by hand in that repo, against its own committed
corpus. This file is the recorded proof `tests/docsqa_ac06_synthorg_task_
matches_gold_question_3.rs` points to for the half it cannot execute itself.

| item | value |
|---|---|
| task id | `docs_qa_recipe` |
| segment | `rag_indexer` |
| synthorg commit that added the task | `eb33096` (`corpus: add docs_qa_recipe task for rag_indexer (PRD-mcphost-docs-qa-recipe)`) |
| synthorg HEAD verified against | `20d01fe` (`eb33096` confirmed an ancestor via `git merge-base --is-ancestor eb33096 HEAD`) |
| corpus file | `~/repos/synthorg/corpora/mcphost/consumer-tasks.yaml` |

## `synthorg consume --preflight` (lists the task for `rag_indexer`)

```
$ cd ~/repos/synthorg && synthorg consume --preflight --corpus corpora/mcphost/consumer-tasks.yaml
...
docs_qa_recipe: turn_budget: 40
...
ok — every gold tool is served
$ echo $?
0
```

## `synthorg corpus check --live-shape` (the fake endpoint fixture satisfies the gold predicate, not vacuously)

```
$ cd ~/repos/synthorg && synthorg corpus check --live-shape --corpus corpora/mcphost/consumer-tasks.yaml
task_id | kind | observable | correct | wrong-fails | prose-names | verdict
...
docs_qa_recipe | http | vpn-setup.md | True | True | True | ok
...
$ echo $?
0
```

`vpn-setup.md` is the same `expected_document` this repo's own
`examples/docs-qa/corpus/gold.json` records for question 3 — `synthorg
corpus check --live-shape` independently confirms the synthorg-side task's
`call_check` is satisfied by a correct tool result and fails against a
wrong one (`correct`/`wrong-fails` both `True`), i.e. the check is not
vacuous.

## Proxy-tier fixture run (fake in-process endpoint, no cost)

The "proxy-tier fixture run" is the fake-mode pytest suite that actually
dispatches the `docs_qa_recipe` http-kind task to the in-process fixture at
`/docs/search/question-3` and scores a known-good transcript as a pass and
known-bad transcripts as failures:

```
$ cd ~/repos/synthorg && python -m pytest \
    tests/ragtasks_ac3_known_good_bad_transcripts_test.py \
    tests/fixture_upstream_ac5_fake_run_dispatches_to_fixture_test.py
tests/ragtasks_ac3_known_good_bad_transcripts_test.py ....            [ 80%]
tests/fixture_upstream_ac5_fake_run_dispatches_to_fixture_test.py .    [100%]
5 passed in 8.49s
```

Together these three runs are the actual execution AC6's When-clause names
(`synthorg consume --preflight`, and the proxy-tier/fake-endpoint fixture
proving the gold predicate) — not a hand-written assertion about what the
task file says, and not invented output.

## Receipt: docs-qa-panel-satisfaction.md

# satisfaction[rag_indexer] panel record — PRD-mcphost-docs-qa-recipe

AC10: *Given the truth-tier panel run after land, When `synthorg consume`
runs with the pinned panel, Then `rag_indexer` satisfaction is reported with
the new task included (number recorded in the vision, target ≥ 75%).*

This file is this repo's own copy of that record; the record of account is
the vision doc's own `## Satisfaction record` section
(`~/Documents/PRDs/visions/mcphost-data-layer-first-slice.md`, the PRD's
`Vision:`). Both are written by the same recorder,
`examples/docs-qa/record-panel-satisfaction.sh`, which refuses any run whose
corpus did not include this recipe's `docs_qa_recipe` task — a panel number
measured without the task cannot answer this AC.

| item | value |
|---|---|
| segment | `rag_indexer` |
| baseline | 61.8% (2026-09-23 quality baseline, wiki `2026-09-23-mcphost-market-test-plan`) |
| target | ≥ 75% (this PRD's success metric, first truth-tier run after land) |
| corpus task | `docs_qa_recipe` (`~/repos/synthorg/corpora/mcphost/consumer-tasks.yaml`, synthorg commit `eb33096`) |
| pinned panel | `corpora/mcphost/panel-composition.yaml` |
| vision record | `~/Documents/PRDs` commit `05e9541` (the `## Satisfaction record` section this file mirrors) |
| PRD deferral record | `~/Documents/PRDs` commit `7c383c3` (`build-queue/PRD-mcphost-docs-qa-recipe.md` frontmatter: `deferred_acs: [9, 10]` + `mock_justifications`; the repo's own copy at `PRD-mcphost-docs-qa-recipe.md` is what `tests/docsqa_ac10_deferral_is_justified.rs` parses on the runner box) |

## The truth-tier run (deferred — operator-authorized, after land)

The number itself comes from a live, paid, operator-authorized run against
prod, which this build deliberately does not perform. After this branch
lands and `mcphost.dev` serves it, the operator runs:

```
cd ~/repos/synthorg
SYNTHORG_PROD_ENDPOINT=https://mcphost.dev/mcp \
SYNTHORG_TRUTH_BRIEF=<market brief.md> \
  scripts/truth-tier-nightly.sh          # or, directly:
uv run synthorg consume --tier truth --population 21 \
  --composition corpora/mcphost/panel-composition.yaml \
  --endpoint https://mcphost.dev/mcp <brief.md> --out runs/truth-tier-<ts>
```

then records that run's number into both files:

```
cd ~/wintermute/mcphost
examples/docs-qa/record-panel-satisfaction.sh \
  --measure ~/repos/synthorg/runs/truth-tier-<ts>/measure.json \
  --vision ~/Documents/PRDs/visions/mcphost-data-layer-first-slice.md \
  --receipt docs/receipts/docs-qa-panel-satisfaction.md
```

The recorder exits 0 only for a live truth-tier run that met the target; an
under-target number is still recorded (exit 1), because AC10 asks for the
number, and a miss is a number too. The same command with
`MCPHOST_LIVE=1 MCPHOST_PANEL_MEASURE=<that measure.json>` is what
`tests/docsqa_ac10_truth_tier_panel_satisfaction_recorded.rs`'s live half
runs, so the post-land run has a test that fails until it happens.

## What was measured before land (offline, not the market number)

An offline full-corpus panel run over the same pinned composition and the
same corpus — `SYNTHORG_LLM_MODE=fake` (stub personas, stub judge, the
in-process fake MCP endpoint), so it is *not* a market number and the
recorder labels it `offline` and refuses to call it green. What it does
prove, on real run output rather than a hand-written fixture, is that
`docs_qa_recipe` is genuinely wired into the pinned panel and that
`rag_indexer`'s number is computed with it included:

## Satisfaction record — satisfaction[rag_indexer]

- 2026-09-26 offline run runs/offline-docsqa-20260926T053158Z: satisfaction[rag_indexer]=90.0% (target >= 75%, scored 5/5 sessions, docs_qa_recipe included, corpus fingerprint 609b58d058cc)

## Fixtures the test drives, pinned to the runs they came from

`tests/docsqa_ac10_truth_tier_panel_satisfaction_recorded.rs` drives the
recorder with two trimmed copies of real `measure.json` files — same keys the
recorder reads, every value copied unchanged from the run (verified key by key
against the source measures on the build host, 2026-09-27). Synthorg is a
different repository and is not present on this gate's runner box, so the
fixtures cannot be re-derived there; instead each one's byte digest is pinned
here and
`the_receipt_pins_each_fixture_to_the_real_run_it_was_trimmed_from` fails if a
fixture changes without this table changing with it. Editing a *number* in a
fixture is the one way this AC's mechanism proof could quietly become an
invented one, and that is what these digests close.

| fixture | source run (`~/repos/synthorg`) | sha256 (fixture / source measure.json) | why |
|---|---|---|---|
| `examples/docs-qa/fixtures/panel-measure-offline-20260926T053158Z.json` | `runs/offline-docsqa-20260926T053158Z` (`docs_qa_recipe` included, `rag_indexer` 90.0%, `client_version: fake`) | `4b0a66c91b1f07056c0c4062294891726d02fb1512608cfddba00623e0c69999` / `1bef3f26e67304f88caf2b029f55adc39b488380190742f9227d6f84a432aff1` | the recorder's `offline` classification and the record line, on real output |
| `examples/docs-qa/fixtures/panel-measure-truth-20260926T025631Z.json` | `runs/truth-tier-20260926T025631Z` (a real pre-land truth-tier nightly, `rag_indexer` 62.5%, 28 tasks, **no** `docs_qa_recipe`, $4.62) | `97f1c0dc8321ac1fec9511e507389ee4fd5a8d7f28371779dbe47eb670e75f2b` / `8103f1ea5e4920ee6c0ba17cb2e93c1e9bd9f18a825c6dc6035ce54fc34e73f6` | the refusal: a real, live, expensive truth-tier number is still not an answer to AC10 if the task was not in the corpus |

No post-land truth-tier run exists yet, so the measures the target comparison
runs on are **assembled from those two real runs' own values** and nothing
else: the truth nightly's `tier`/`client`/`client_version`/`endpoint_version`/
judge and scorer versions and pinned composition, carrying the offline run's
whole `corpus` object (the only real corpus that contains `docs_qa_recipe`,
fingerprint and all) plus one run's own `satisfaction` and session counts —
the offline run's 90.0% for the green case, the truth nightly's own 62.5% for
the under-target case. Nothing is hand-set, spliced or rounded;
`the_measures_under_test_invent_no_number` fails if a later edit tries to.

## The deferral, as declared

AC10's Then is deferred — **operator-provisioned**. The PRD's own frontmatter
carries `deferred_acs: [9, 10]` and a `mock_justifications` entry for AC10
naming the paid prod run, the credentials this sandbox does not hold
(`ANTHROPIC_API_KEY` / `WM_ANTHROPIC_API_KEY`, `SYNTHORG_PROD_ENDPOINT`,
`SYNTHORG_TRUTH_BRIEF`), the live test file and fn
(`tests/docsqa_ac10_truth_tier_panel_satisfaction_recorded.rs::live_truth_tier_panel_run_meets_the_target`),
and the sentence that the always-on tests prove the branch's mechanism, not
AC10, and are not counted as AC10's proof.
`tests/docsqa_ac10_deferral_is_justified.rs` parses that frontmatter,
`agent/test-map.json` and `agent/intent-card.json` and fails if the three ever
stop agreeing.

## Receipt: mcphost-docs-qa-recipe.md

# Receipt — mcphost-docs-qa-recipe AC9: docs-qa.sh against prod after the v0.60.34 deploy

AC9: Given prod after land, When `docs-qa.sh https://mcphost.dev` runs from carbon with a fresh signup, Then it exits 0 in under 60 s and the receipt is committed under `docs/receipts/`.

## Prod state at the run

- Landed: PR #71 (squash aedb4e5), tag v0.60.34, deploy-on-land.
- `/opt/mcphost/bin/mcphost --version` on mcphost-1: `mcphost 0.60.34`; service ActiveEnterTimestamp 2026-09-28 03:06:45 UTC (8:06 pm PDT).

## What was run (carbon, 2026-09-27 8:13:29 pm PDT)

```
examples/docs-qa/docs-qa.sh https://mcphost.dev/mcp --receipt-dir <tmp>
```

Fresh signup (tenant `t_880aa34a`), the 8-document corpus put and indexed, `ask_docs` published, the 6 gold questions asked. Calls, in order: signup, host.docs.put, host.docs.status, host.tool_publish, t_880aa34a.ask_docs.

## Result

| field | value |
|---|---|
| exit code | 0 |
| wall time | 21232 ms |
| index ready | 9.215463041968178 s |
| hits | 6/6 |
| quota | {"documents": 8, "chunks": 35, "tool_calls": 6} |

Per question:

| # | question | result | citation |
|---|---|---|---|
| 1 | How many business days does a new hire have to complete onboarding paperwork? | hit | onboarding.md:617 |
| 2 | What is the meal reimbursement limit that requires an itemized receipt? | hit | expense-policy.md:557 |
| 3 | What port does the VPN client connect on? | hit | vpn-setup.md:2083 |
| 4 | Who gets paged first for a Sev1 incident? | hit | incident-response.md:0 |
| 5 | How many PTO days do employees accrue per year? | hit | pto-policy.md:1725 |
| 6 | How often must employees rotate their password? | hit | security-basics.md:0 |

Full receipt JSON (per-step timings, answers with cited passages): `docs/receipts/mcphost-docs-qa-recipe-ac9-20260927.json`.

## Receipt: mcphost-upstream-token-vault-status.md

# Upstream token vault status — AC9 evidence receipt

PRD-mcphost-upstream-token-vault-status, AC9 (P0, Live):

> Given prod after deploy with the operator tenant holding provider `slack`
> registered from `preset: "slack"` with placeholder client credentials and no
> handoff (operator-provisioned Given, done by hand from orch with the operator
> key on mcphost-1 `/etc/mcphost/operator-tenant.key`), When
> `host.vault.status {end_user: "vaultst-probe"}` is called with the operator
> key over `https://mcphost.dev/mcp` and `admin.vault.stats` with the admin key
> from orch `~/.config/mcphost/admin-key`, Then status lists `slack` with
> `connected: false` and stats lists the operator tenant with
> `slack {tokens: 0}` (Live; evidence: healthz version after deploy plus both
> transcripts saved under `docs/receipts/<slug>.md`, no key or secret text).

AC9 has two halves and this receipt keeps them visibly separate:

| half | what it proves | state |
|---|---|---|
| the mechanism | a real `mcphost` server built from this branch serves `host.vault.status` / `admin.vault.stats`, and a registered-never-connected provider reads back `connected: false` / `tokens: 0` | **proven here**, by the transcripts below |
| the prod leg | that `mcphost.dev` *as deployed* serves them, and that the operator tenant's placeholder `slack` row exists on mcphost-1 | **PENDING the operator's post-ship run** — deferred, operator-provisioned |

No key, bearer token, client id or client secret appears anywhere in this
file; the transcripts below are response bodies only, and the vault's status
and stats surfaces never emit a token substring by construction (AC1/AC4).

## Branch-local transcript (every `cargo test`)

Captured by `tests/vaultst_ac09_live_vault_status_trailer.rs` against a real
server process on a real loopback socket — the same `mcphost::http` app
`main.rs` serves in production — with a fresh tenant standing in for the
operator tenant: `slack` registered through `preset: "slack"` with placeholder
client credentials, no handoff ever completed. That is AC9's Given exactly
except for *which* tenant, and *which* host.

This block is not transcribed by hand. `receipt_records_the_transcripts_the_server_actually_serves`
re-renders it from live response bodies on every `cargo test` run and fails if
it differs from what is committed here, so the receipt cannot drift away from
the code it vouches for. Two values are substituted, because they vary per run
rather than per behaviour: the stand-in tenant's generated namespace
(`<operator-tenant>`; on prod this is the operator tenant's own namespace) and
the crate version (`<crate version>`, asserted in that same test to equal this
build's `CARGO_PKG_VERSION`).

`GET /healthz` (admin bearer):

```json
{
  "db_ok": true,
  "version": "<crate version>"
}
```

`host.vault.status {"end_user": "vaultst-probe"}` (operator key):

```json
{
  "providers": [
    {
      "connected": false,
      "connected_at": null,
      "expires_at": null,
      "last_refreshed_at": null,
      "name": "slack",
      "revoked_at": null,
      "revoked_reason": null,
      "scopes": null
    }
  ]
}
```

`admin.vault.stats {}` (admin key):

```json
{
  "tenants": [
    {
      "providers": [
        {
          "name": "slack",
          "refresh_failures_24h": 0,
          "revoked": 0,
          "tokens": 0
        }
      ],
      "tenant_id": "<operator-tenant>"
    }
  ],
  "totals": {
    "refresh_failures_24h": 0,
    "revoked": 0,
    "tokens": 0
  }
}
```

The `slack {tokens: 0}` row above is the load-bearing one, and it is the
finding this receipt exists to record: `Db::vault_stats` originally aggregated
`FROM vault_tokens`, so a provider a tenant had registered but that no end
user had ever connected had no row to aggregate and was dropped from the
output entirely — indistinguishable, to the operator reading stats, from
"never registered". That is precisely AC9's state (placeholder credentials, no
handoff), so AC9's own Then could not have been satisfied on prod either. The
query now runs `FROM vault_providers ... LEFT JOIN vault_tokens`
(`src/db.rs`), and the zero row above is that fix, read back over HTTP.

## Prod leg — PENDING

Not run from the build sandbox, and not runnable from it: registering the
operator tenant's placeholder `slack` provider on mcphost-1 is an operator
action, and the operator key (`/etc/mcphost/operator-tenant.key` on mcphost-1)
and admin key (`~/.config/mcphost/admin-key` on orch) are operator-held. This
PRD therefore records AC9 under `deferred_acs` with reason
operator-provisioned, with the justification in its frontmatter
(`mock_justifications`) and the agreement of `agent/test-map.json` /
`agent/intent-card.json` locked by
`tests/vaultst_ac09_deferral_is_justified.rs`.

The check itself is written and waiting. After this branch ships and the
operator has hand-registered that provider, from orch:

```sh
# 1. the deployed version this evidence is pinned to
curl -sS -H "Authorization: Bearer $MCPHOST_ADMIN_KEY" \
  https://mcphost.dev/healthz | jq '{db_ok, version}'

# 2. both transcripts, through the same test that produced the block above
MCPHOST_LIVE=1 \
MCPHOST_URL=https://mcphost.dev \
MCPHOST_OPERATOR_KEY="$(cat /etc/mcphost/operator-tenant.key)" \
MCPHOST_ADMIN_KEY="$(cat ~/.config/mcphost/admin-key)" \
  cargo test --test suite_core_07 \
  vaultst_ac09_live_vault_status_trailer::operator_tenant_slack_shows_disconnected_and_zero_tokens \
  -- --nocapture
```

`MCPHOST_LIVE=1` redirects the identical `host.vault.status {end_user:
"vaultst-probe"}` (operator key, `https://mcphost.dev/mcp`) and
`admin.vault.stats` (admin key) calls at prod and prints both bodies in the
block shape above. Paste the printed version line and the two bodies into a
new "Prod leg — verified <date>" section here, replacing this one, and AC9 is
closed. Redact nothing by hand: neither body carries a credential.

## Receipt: python-kind-latency.md

# python-kind publish/first-call latency — PRD-mcphost-python-kind-runtime

Requirement 1 / AC4: instrument publish and first-call durations for
python-kind tools; establish where the recorded 88–95s went and land the
fix for the dominant term.

## Pre-fix baseline (recorded, production box class)

From the PRD's `panel_rag_indexer_03` session on `mcphost-1` (a Hetzner
ccx13):

| metric | value |
|---|---|
| publish → first successful call | 88.4s |
| publish → first-call SLA deduction point | 95.3s |

## Post-fix measurement (this receipt, dev box)

Measured by `tests/runtime_ac4_publish_and_first_call_latency.rs`
(`cargo test --test suite_sandbox_01 runtime_ac4_publish_and_first_call_latency:: -- --nocapture`),
a dependency-free tool of corpus-task size, on this build machine
(RedBaron; `uv` and `/usr/bin/python3` both already warm/cached):

| metric | measured | budget |
|---|---|---|
| publish (`validate_async`: ast-check + inference) | ~44ms | ≤10s |
| first successful call (env lookup + build wait + sandboxed run) | ~95ms | ≤5s |

Both comfortably inside AC4's budget on a warm cache.

## Root-cause / dominant-term finding

`run_build_steps` (src/kinds/python.rs) always runs `uv venv <env_dir>`
for a never-before-seen requirements hash, even for a zero-requirements
tool — there is no separate "skip venv creation" fast path. Measured
directly on this box: `uv venv` against a fresh directory completes in
single-digit milliseconds once `uv` has already resolved (or downloaded)
a CPython interpreter to use. The PRD's 88–95s baseline could not be
reproduced bit-for-bit here (this box's `uv`/interpreter caches are warm),
which is itself informative: the dominant term is very likely a one-time,
box-local cost of `uv` provisioning a managed CPython interpreter on a
cold cache (network fetch + build), not per-call work `mcphost` repeats on
every request — the existing `EnvRegistry` already caches a built
environment by requirements hash indefinitely (disk-durable `.ready`
marker) and the warm sandbox pool (`WarmPool`, PRD-mcphost-code-tools-warm-pool)
already avoids re-spawning a fresh sandbox process per repeat call to the
same tool. Both caching layers predate this PRD and already eliminate the
"per-call environment setup" case this PRD's own P1 anticipated as the
likely dominant term.

Per the PRD's own Open Questions entry ("if the dominant term is
hardware-bound the receipt says so and the target is revisited rather
than gamed"): this receipt says so. The fix this PRD lands is the
instrumentation itself (`publish_ms` / `cold_call_ms` fields on the
`tracing::info!` lines in `PythonKind::validate_async` and
`PythonKind::call`) — the concrete, board-verifiable signal that would
catch a regression toward the baseline's numbers on any box, and the
evidence a follow-on PRD would need to pre-warm `uv`'s managed-interpreter
cache on `mcphost-1` specifically, if a live measurement there still shows
the 88–95s figure.

## Receipt: spec-readback-ac10-prod-probe.md

# Receipt — mcphost-shared-tool-spec-readback AC10: what prod is still missing

AC10's Given is "prod mcphost after deploy". This receipt records the one
thing that clause is still waiting on, as a fact a reader can re-derive
rather than a claim they have to take on trust.

## What was done

One read-only call against production, with **no `Authorization` header at
all** — `tools/list` is a listing call, and nothing on prod was created,
changed or deleted:

```
curl -s -X POST https://mcphost.dev/mcp \
  -H 'content-type: application/json' \
  -H 'accept: application/json, text/event-stream' \
  -d '{"jsonrpc":"2.0","id":2,"method":"tools/list"}'
```

## What it shows

1. **The deploy is the blocker.** Prod serves every step of AC10's
   sequence today — `signup`, `host.whoami`, `host.tool_publish`,
   `host.group.create`, `host.group.add`, `host.tool_share`, `host.usage` —
   but not `host.tool_spec_shared`, and its `host.tool_share` has no
   `expose_spec` property. The gap is exactly this branch's surface, and
   closing it is a deploy: an operator action (mcphost-deploy), not a
   build-agent one.
2. **Credentials are not a blocker.** Prod's `signup` is unauthenticated,
   so the live leg needs no operator-held tenant key:
   `MCPHOST_TENANT_A_KEY` / `MCPHOST_TENANT_B_KEY` are optional, and with
   them unset the `MCPHOST_LIVE=1` run signs its own two tenants up on the
   live endpoint.

`tests/mcphost_shared_tool_spec_readback_ac10_live_two_tenant_spec_read_trailer.rs`
reads the JSON below back on every run
(`only_the_deploy_separates_this_branch_from_prods_tool_surface`) and
compares it with the identical unauthenticated `tools/list` against a real
server built from this branch. When the deploy lands, that test and
`the_prod_probe_the_justification_cites_exists_and_shows_the_blocker` both
go red — which is precisely when AC10's prod leg stops being deferred and
has to be run for real.

## The probe

```json
{
  "_what": "A read-only probe of production mcphost's public, unauthenticated tools/list, captured to pin exactly what PRD-mcphost-shared-tool-spec-readback AC10's prod leg is still waiting on. Nothing was created, changed or deleted on prod: tools/list is a listing call and was made with no Authorization header at all.",
  "_why": "AC10's deferral has to name a blocker a reader can check, not a claim. This snapshot shows (a) prod does not yet serve host.tool_spec_shared and host.tool_share has no expose_spec property -- i.e. this branch is not deployed, the one operator action the live leg waits on -- and (b) prod's signup is unauthenticated, so the live leg needs no operator-held tenant credential.",
  "_reproduce": "curl -s -X POST https://mcphost.dev/mcp -H 'content-type: application/json' -H 'accept: application/json, text/event-stream' -d '{\"jsonrpc\":\"2.0\",\"id\":2,\"method\":\"tools/list\"}'",
  "probed_at": "2026-09-27T20:46:18Z",
  "endpoint": "https://mcphost.dev/mcp",
  "method": "tools/list",
  "authorization_header_sent": false,
  "raw_response_sha256": "48e998c7668d1c6d861811210a9386de6479461bedcc66acd7d356eec8e2a419",
  "raw_response_bytes": 95590,
  "tool_count": 135,
  "serves_host_tool_spec_shared": false,
  "host_tool_share_input_properties": [
    "description",
    "group",
    "name",
    "tenant_key",
    "visibility"
  ],
  "signup_entry": {
    "name": "signup",
    "description": "Create a tenant and receive a bearer key and namespace. Unauthenticated. Recommended: pass handoff: true to receive a short-lived, single-use handoff_token instead of the raw key -- redeem it once with host.redeem to get the key, so a transcript of this call and the redeem call, if it leaks, carries a dead credential. The raw-key path (handoff omitted) stays fully supported.",
    "inputSchema": {
      "properties": {
        "handoff": {
          "description": "Recommended: true to receive a handoff_token (redeem via host.redeem) instead of the raw key. Default false (raw key, unchanged).",
          "type": "boolean"
        },
        "name": {
          "description": "display name",
          "type": "string"
        }
      },
      "required": [
        "name"
      ],
      "type": "object"
    }
  },
  "host_tool_share_entry": {
    "name": "host.tool_share",
    "description": "Share one of this tenant's published tools with everyone (visibility: \"public\") or with a named group this tenant owns (visibility: \"group\", group: <name>). The tool keeps running in this tenant's own sandbox with this tenant's own secrets; a caller reaches it as <this tenant's namespace>.<name>.",
    "inputSchema": {
      "properties": {
        "description": {
          "description": "Catalog-facing blurb; shown by host.catalog.search/get.",
          "type": "string"
        },
        "group": {
          "description": "Required when visibility is \"group\"; must already exist (host.group.create).",
          "type": "string"
        },
        "name": {
          "description": "Local name of the tool to share.",
          "type": "string"
        },
        "tenant_key": {
          "description": "The key `signup` returned. Required only when this connection carries no Authorization: Bearer header -- when both are present, the header wins.",
          "type": "string"
        },
        "visibility": {
          "description": "\"public\" or \"group\".",
          "type": "string"
        }
      },
      "required": [
        "name",
        "visibility"
      ],
      "type": "object"
    }
  },
  "tool_names": [
    "billing.checkout",
    "billing.plans",
    "billing.status",
    "host.agent.contact_accept",
    "host.agent.contact_deny",
    "host.agent.contact_request",
    "host.agent.contacts",
    "host.agent.contacts_import",
    "host.agent.lookup",
    "host.agent.mute",
    "host.agent.profile_set",
    "host.agent.search",
    "host.agent.unmute",
    "host.agent.whoami",
    "host.bridge_test",
    "host.catalog.get",
    "host.catalog.search",
    "host.changelog",
    "host.channel.close",
    "host.channel.freeze",
    "host.channel.open",
    "host.channel.post",
    "host.channel.read",
    "host.channel.unfreeze",
    "host.docs.delete",
    "host.docs.get",
    "host.docs.index_config",
    "host.docs.list",
    "host.docs.purge",
    "host.docs.put",
    "host.docs.reindex",
    "host.docs.search",
    "host.docs.status",
    "host.enduser.assertion_secret_rotate",
    "host.enduser.audit",
    "host.enduser.export",
    "host.enduser.get",
    "host.enduser.list",
    "host.enduser.purge",
    "host.enduser.revoke",
    "host.enduser.unrevoke",
    "host.enduser.whoami",
    "host.export",
    "host.group.add",
    "host.group.create",
    "host.group.list",
    "host.group.remove",
    "host.key_rotate",
    "host.msg.ack",
    "host.msg.block",
    "host.msg.inbox",
    "host.msg.reply",
    "host.msg.send",
    "host.msg.thread",
    "host.msg.unblock",
    "host.msg.wait",
    "host.oauth.audit",
    "host.oauth.audit_export",
    "host.oauth.client_approve",
    "host.oauth.client_deny",
    "host.oauth.doctor",
    "host.oauth.grant_revoke",
    "host.oauth.grants",
    "host.oauth.issuer_remove",
    "host.oauth.issuer_set",
    "host.oauth.issuers",
    "host.oauth.pending",
    "host.oauth.policy",
    "host.oauth.policy_set",
    "host.oauth.provider",
    "host.oauth.provider_remove",
    "host.oauth.provider_set",
    "host.oauth.revoke_all",
    "host.oauth.scope_set",
    "host.oauth.scopes",
    "host.oauth.trusted_issuer_remove",
    "host.oauth.trusted_issuer_set",
    "host.oauth.trusted_issuers",
    "host.progress",
    "host.quickstart",
    "host.redeem",
    "host.registry_publish",
    "host.runs.cancel",
    "host.runs.get",
    "host.runs.list",
    "host.runs.part",
    "host.runs.purge",
    "host.runs.wait",
    "host.secret_list",
    "host.secret_set",
    "host.self_offboard",
    "host.share.caller_limit",
    "host.share.caller_limit_remove",
    "host.state.delete",
    "host.state.delete_rows",
    "host.state.get",
    "host.state.insert",
    "host.state.list",
    "host.state.query",
    "host.state.set",
    "host.state.table_create",
    "host.state.table_drop",
    "host.table.append",
    "host.table.create",
    "host.table.describe",
    "host.table.drop",
    "host.table.list",
    "host.table.model_set",
    "host.table.models",
    "host.table.query",
    "host.table.schema",
    "host.tool_call",
    "host.tool_diff",
    "host.tool_history",
    "host.tool_list",
    "host.tool_logs",
    "host.tool_publish",
    "host.tool_remove",
    "host.tool_rollback",
    "host.tool_run",
    "host.tool_share",
    "host.tool_test",
    "host.tool_unshare",
    "host.trigger.fire",
    "host.trigger.get",
    "host.trigger.list",
    "host.trigger.pause",
    "host.trigger.remove",
    "host.trigger.replay",
    "host.trigger.resume",
    "host.trigger.set",
    "host.trigger.test",
    "host.usage",
    "host.whoami",
    "signup"
  ]
}
```

## Receipt: www-trust-funnel-live.md

# Receipt: www-trust-funnel-live.md

Live proof that the index v2 rewrite (PR #87, merge 8a4a88a) is what https://mcphost.dev serves.

- Checked: 2026-10-01T07:55:14Z (2026-10-01 00:55 PDT) from carbon, plain curl, no cache headers.
- Live `/` md5 (first 12): 5ff678e38c6c; `origin/main:www/index.html` md5: 5ff678e38c6c. (Equal means byte-identical; a difference is expected only if Caddy rewrites nothing and the vendored copy matches, so note which.)
- Marker sentence "publishes the missing tool with a second" on the live page: 1 occurrence.
- Live h2 order: connect|one night on mcphost|what your agent can give itself|what keeps it safe|for the human|built to be depended on|before you trust it with production|pricing|source and contact.
- `/healthz`: {"ok":true}; status.json version: (empty).
- Deploy path: vendor-www 5219359 in mcphost-deploy → `redeploy` from the checkout on orch (`uv run`, backup 20261001T073743Z-aee6c503d30b, probe green, "www content: written"). The durable venv's `release` gate timed out at 900 s; see the deploy memo in the deploy repo.
- Reproduce: `curl -s https://mcphost.dev/ | grep -c "one night on mcphost"` → 1.

