Phase 10 — The audit backlog, and the fixes it found
31 AUG AT 11:21 AM

Phase 10 — The audit backlog, and the fixes it found

0 LOVES 3 VIEWS
An audit of the program's own commits: 32 findings across 14 slices. Its strongest finding turned out to be about itself.

An audit of the program by the program

This phase exists because of a question: had the four pre-commit checks actually been applied to this program’s own commits? The honest answer required reading them rather than asserting it, so the phase became an audit of the program by the program.

Thirty-two findings, fourteen slices, each with its own commit. The method was a sweep by defect class first, then reading the hits — because a judgement-based review only ever catches what it thinks to look at.

⚠ The audit’s own first two passes were wrong, and that is recorded at the top of its doc rather than buried. A finding list assembled without re-measuring is just a memory of a codebase.

What is not closed is stated plainly, because a closing phase that quietly drops its residuals is worse than one that never opened them:

  • The B.V.S.R. finding in the older admin tiers was split out into its own program — it predates this one by seven months and was never this phase’s to carry.
  • The metering half of one finding is BLOCKED, not deferred. The remedy rests on a wire-format fact that costs real money to establish. Decided: do not pay for it, and do not build on it unconfirmed.
  • One finding is a decision, not a fix — 31.6 MB of a 113 MB database with no retention policy, growing about 350 MB a year. Both schemas now carry the census; the answer belongs to a human.
  • One test stays red on purpose. It pins a hand-kept list against a fleet designed to grow, so every successful provisioning breaks it. That is the test being wrong, not the fleet.
  • Two findings were filed and never absorbed, and one is a house-style decision deliberately not taken on this phase’s authority.

⭐ And one gate was released: the launchd work had been blocked on two of these findings, and closing them turned it from a defect into an operator’s choice.

The strongest finding is about the phase

⭐ The phase’s strongest finding is about the phase. Across twelve slices, five instances of the very defect class it was auditing for were committed by the work removing that class — each by someone who had documented the class hours earlier and was actively hunting it.

The first. While fixing a finding about a surface that cannot express a third state, the new predicate read an edge directly and returned true on an empty slice. Empty meant two things — nobody loaded it and none is active — and it resolved the ambiguity toward the alarming reading. That is precisely the defect being fixed, reproduced inside its own fix.

The second, in the very next slice. A function was split so a caller holding a parsed record could skip a second decode, and the test written to protect the split asserted that the wrapper equalled the helper. Those were the same function. A deliberate one-minute offset was injected — a change that shifts every staleness comparison on the status surface — and the test stayed green, because both sides of the equality moved together. A self-referential test, written in a slice about redundancy.

What this establishes is not a resolution to be more careful. These were reproduced by someone looking straight at the class. Vigilance is measurably not a control here. What caught them was blast radius in one case and mutation in the other — and mutation testing is the only thing that has ever caught this class in this program.

The corrective is structural or it is nothing: make the ambiguous value unrepresentable, rather than resolving to remember which way it points.

Two sharper sub-lessons came out of the same work. A mutation that stays green is itself a finding — it means the test is measuring agreement rather than a value, and the repair is to assert against an independently constructed expectation. And a red test is not self-evidently a good test: read whether its failure message names the actual cause, because a test that goes red for the wrong reason will mislead exactly when it is needed.

Three findings overtaken by events

Three findings changed underneath the phase while it ran, and only re-measuring caught it. A finding list assembled without re-measuring is a memory of a codebase, not a description of one.

One finding rested on the observation that production had zero agent-scoped budgets, which made a whole class of enforcement latent rather than live.

By the time the slice ran, production had fifteen of fifteen agents fed. The premise had been overtaken by the program’s own progress — the fleet grew, and the budgets came with it.

The half that still stood was fixed. The half that had dissolved was recorded as dissolved rather than fixed for appearances, which is the only honest outcome when the world moves under a finding.

A second finding was closed not by any slice here but by the orchestrator phase, which had landed in between and removed the condition the finding described.

It is listed because a closed finding with the wrong attribution is a small lie that compounds: the next person reading the audit would look for a fix in this phase’s commits and find nothing, and conclude the record was unreliable.

⭐ This is the one that justifies re-measuring rather than re-reading.

The finding originally said the launchd plist named files that did not exist — an obvious, loud failure: load it and it fails immediately. The recorded remedy was to create the missing directory.

By the time the slice ran, the plist named files that did exist and were wrong. The binaries it pointed at had come into being as a side effect of ordinary builds, and every one of them differed by checksum from the deployed set.

So the recorded remedy would have “closed” the finding while leaving production one command away from executing binaries built in the development tree. A fix that satisfies the ticket and creates the incident. The finding had inverted from fails loudly to succeeds wrongly, and nothing but re-measuring could have caught that.

Pushing Tin — managing a bridge fleet from inside the product
Pushing Tin — managing a bridge fleet from inside the product
Aug 29, 2026 Pushing Tin
← Back to Pushing Tin