Phase 2 — Identity and config: a bridge is declared, not remembered
30 AUG AT 10:31 AM

Phase 2 — Identity and config: a bridge is declared, not remembered

0 LOVES 1 VIEWS
A bridge's identity lived in a process table. This phase makes the declaration a row — and refuses, deliberately, to extend a pattern that already stores secrets in cleartext.

A declaration, not a command line

Before: ps -ax | grep bobit was the only record of what a bridge was. The willartley bridge had already proved the failure mode — had it died, the launch line was unrecoverable.

After: each bridge is an AgentConfigRevision whose config_json states site, channel, user and gate port — versioned, attributed to a user, with an active pointer.

This phase is bookkeeping, and it says so out loud. Nothing reads these rows. No supervisor exists; no process starts, stops, or changes behaviour because of anything here. desired_state is a column with no reader and no screen. That is not a gap to apologise for — it is the phase. Phase 3 needs something to reconcile against, and inventing a consumer now, to make this feel like a working increment, would have been the wrong trade.

One honest correction to “nothing reads these rows”: the worker’s processor does query the active revision and pass config_json straight into an adapter. A bridge never reaches it, behind four independent guards — the dispatcher’s worker-type filter, the push-mode adapter set (a bridge is api), and Phase 1’s two assignment predicates. The reader exists; a bridge is excluded from it four times over. Better to name that than to claim a clean zero.

The default is stopped, and that choice is load-bearing: declaring a bridge must not arm it. With running as the default, this phase — which only writes rows — would have auto-started every declared bridge the instant a supervisor first appeared.

Three buckets, nineteen flags

The overview stated the constraint as “config_json must never carry anything that resolves to an executable”, and named four supervisor-side flags. Measured against the bridge’s actual 19 flags, both halves needed widening.

A URL never resolves to an executable — and -api-base is precisely where the bridge sends its API key as a bearer token. An operator-editable one is credential exfiltration in a single form field. So the criterion became three classes, and the flags sorted into them:

  • A — lives in config_json, operator-editable and validated: site, locale, channel, user, the port half of the gate address, and the gate deadline. Six flags.
  • B — supervisor constants, identical for every bridge: the two binaries, the two env files that decide which credentials load, the three network destinations that receive a credential, the issuer, and the flag that disables the approval gate. Nine flags.
  • C — supervisor-derived, per-bridge but never operator-supplied: workdir, gate log, gate corpus, gate state. Four flags.

Class C is the correction that matters, and it came from reading the live flag sets rather than the plan. “Supervisor constant” reads as identical for every bridge — and four of the thirteen were not. The two running bridges had different workdirs, different gate logs, different corpora and different state files. Those four vary per bridge and are paths the supervisor reads and writes, so they belong in neither bucket. The resolution: the supervisor derives them from a root it controls. The operator picks the slug; the supervisor picks the path.

Which quietly made agent_slug a path component. That is safe today only because the slug regex is anchored, with no dots and no separators — a regex documented as a formatting rule, shared by three request types, that had been carrying a security role none of its callers knew about. A future loosening for a perfectly good site-naming reason would silently have made bridge paths escapable.

It is now closed from both directions: a test pins it from the side that would be harmed, and the regex’s own doc comment says it is load-bearing for path safety and names that test — so someone loosening it meets the constraint at the point of change, rather than meeting a failing test in another package that reads as unrelated. A guard nobody can attribute is a guard that gets deleted.

Carried forward

Three things this phase settled or recorded rather than quietly carrying. Two were decided here; the third is a hole left open on purpose, with the reason written down.

The budget enforcer’s account and site sweeps pause every active agent with no type filter, and key validation rejects a non-active agent. So one worker’s overspend made a healthy bridge report dead — and after Phase 3, would have killed its process mid-conversation.

The decision has two parts, and the second is what keeps it honest. First, a separate field: desired_state, defaulting to stopped, with the supervisor running a bridge only when desired_state = running AND status = active. Second, narrow the sweeps rather than exempt the type wholesale — bridges are exempt from the account and site sweeps only. An agent-scoped budget naming the bridge still pauses it.

So a bridge is stopped by its own overspend and never by a sibling’s. That keeps the spend brake Phase 6 goes on to arm, and removes only the collateral damage. Two alternatives were rejected for the same underlying reason: a type filter alone leaves status conflated and removes the brake entirely, and a paused_reason column is still one field doing two jobs, with nothing structurally stopping the next writer clobbering it.

Verified, unchanged from Phase 1, and deliberately not fixed here: a bridge’s API key authorises six routes. Five are the intended surface — heartbeat, task checkout, submit, fail, cost record. The sixth is POST /assertions, which means a chat relay can write assertions into the knowledge graph for its assigned site.

It is bounded: create-only by an earlier route-layer split, rate-limited, and scoped to the bridge’s assigned site — the close, retract and merge verbs all return 401. But a chat relay has no business holding knowledge-graph write capability at all.

The reason it stays: agent API keys are a single scope, not per-route scopes. Narrowing them is its own phase with its own smoke, and widening them silently would be worse than leaving the fact recorded where the next person reads it.

The bridge allowlist parses with a decoder set to reject unknown fields. The LLM branch three lines below it uses plain json.Unmarshal, which silently drops them.

Had the allowlist pattern-matched its immediate neighbour — the natural thing to do — it would have accepted {"site":"main","mcp_bin":"/evil"} without a murmur. A rule that looks enforced and enforces nothing is worse than no rule, because everyone downstream then reasons as though it holds.

The test asserts both halves: that the strict decoder rejects the unknown key, and that plain Unmarshal accepts it. So the test documents the hole it exists to keep closed, rather than just asserting today’s behaviour.

The neighbouring branch still has the defect. Lower stakes — it validates only provider and model, and an unknown key there is ignored rather than exploited — but it is the same shape, and it is not fixed here. Named rather than quietly stepped around.

Pushing Tin — managing a bridge fleet from inside the product
Pushing Tin — managing a bridge fleet from inside the product
Aug 29, 2026 Pushing Tin
← Back to Pushing Tin