The harness: product ownership when production is cheap
October 4, 2026 · 4 min read
The work has no sense of enough
Before owning software products, I led soldiers, taught students, coached athletes, and ran operations. Software built with AI is the same kind of system running at a new speed. The work moves faster than any team I have led, and it has no sense of enough. Outputs can be plausible, and each can look like progress. So the job shifts from directing the work to building the harness it runs inside: a structure stable enough to hold someone accountable, and flexible enough to change when the work does.
That is the short version. This is the longer one.
Correctness is not the same as worth
Most of what I built into my delivery process checks correctness. Is the count right? Does the reference match the source? Did the claim actually land on the page? Those checks matter, and they catch real mistakes.
But a process that measures only correctness can keep generating more correctness work, because there is usually another detail to fix. By itself, it does not tell you whether the day went to the right things.
A night that made the point
On the night of August 25, 2026, I was working in the CH(Ai)SE methodology repository with Chaise, the agent responsible for maintaining its canon. By morning, fourteen pull requests, #21 through #34, had merged; the canon log held twenty-three entries in total; and six cards, CHAISE-5 through CHAISE-10, had been opened. CHAISE-10's retrospective records seven defects found by review and none by the authoring session. I could defend each change individually. By pull-request count, it was a productive night.
Nothing shipped to a user. No product behavior changed. The question I had actually started with, how leadership decides what deserves to exist, had produced nothing at all.
Two details made it worse. The first draft of the card describing the night said ten pull requests. It was fourteen. The number had been written to fit the sentence instead of being counted, inside a card about exactly that problem. And within an hour of naming the gap, more cards were filed, including one marked Urgent, with no recorded worth check asking whether they deserved to exist.
That was not a bad night. It was a system working exactly as designed, in a direction nobody chose.
What I changed
Two rules came out of that night.
A worth statement at the close of every sprint, ranked above the metrics. It answers four questions: what changed for the system, the business, or the user; what evidence shows it changed; what the sprint cost, including rework; and what we would not do again. A sprint that cannot say what changed is a finding, not a formatting problem. "Nothing shipped, and here is what we learned and what it cost" is an acceptable answer. A card count is not.
A record of declined work. When an idea is considered and not built, it gets a short entry: what was considered, why it was declined, and what would change the answer. The third field carries the weight. A decline with no reopening condition is a veto, not a decision.
Neither rule makes the work faster. Both make it possible to see where the effort went.
The harness at two levels
I run more than one project, so the harness has two levels, and they hold different kinds of judgment.
At the project level, the judgment is about scope and evidence. A card states what done means before any work starts. Changes are reviewed and checked against that definition before they count. The role that defines the work is not the role that merges it. On my largest project, the role that merges a change is not the one that verifies it either; a small project like this site combines those steps.
At the portfolio level, the judgment is about what becomes a shared rule. One role maintains the methodology across projects. Another reviews proposed canon changes and returns findings, but cannot decide those findings or author canon. The reviewer is never the author. There is also a threshold for changes to methodology canon: propagating an Operator decision that has already been made can proceed directly, while authoring a new rule requires my authorization on the record. When a canon change is not clearly propagation or rule-making, it is treated as rule-making.
That threshold exists because propagating a ruling and inferring a new one can look identical to the agent making the change.
What the product owner still owns
None of this removes judgment. It moves it.
I no longer spend most of my attention directing individual pieces of work. I spend it on the structure: what a card must say before it starts, who checks what, where a finding goes, when a lesson becomes a rule. And I spend it on the question that no check answers for me: does this deserve to exist?
I do not delegate that question to the system. A correctness-focused system can keep finding valid tasks without deciding whether they are worth doing. The harness makes responsibilities and evidence visible. Deciding what the work is for remains the product owner's job.