The Method and the Verdicts

Giovanni Leonardi·September 2026·6 min read

A verdict is dated evidence, not settled law.

What this page is for

Take any piece of work done every week. Some of it, an AI agent could probably do. This page is the procedure for finding out which parts — without self-deception, and without finding out the expensive way.

The procedure is six steps, run on one piece of work at a time, on real work, with real stakes. At the end comes an answer — the Lab calls it a verdict — to one question: can this piece of work be handed to an agent, or not? Then the steps repeat, because the answer changes as the tools change.

Every experiment in this Lab runs these six steps, every time, with no exceptions. An experiment that changes its rules is a story.

The six steps

The steps are a loop, not a straight line. Steps one to five take a piece of work to its first verdict — and, if it earns it, into real use early. Step six is where the loop lives: measure, improve, go back, test again. Ship early. Improve continuously. On the record.

  1. Isolate. Name the piece of work in one sentence, and say why it is worth handing over. Not “programme management” — that is a job, not an act. This is an act: sorting incoming issues into escalate, absorb, or ignore. Work that cannot be named in a sentence cannot be tested yet.
  2. Decompose and predict. Break the act into four parts: what comes in, what gets judged, what goes out, and what happens as a result. Then write down — dated, before anything runs — which parts the agent is expected to handle and where it is expected to fail. For the issue-sorting example: the agent will likely apply the sorting rules consistently; it will likely miss the issue that is small on paper but political in fact. Written predictions are the difference between an experiment and a demo.
  3. Build and delegate. Build the agent and hand the act over — on live work, where the output matters. Keep a log of what breaks during the build, because the breakage is evidence too. How a piece of work becomes something an agent can run is a craft of its own; it has its own page.
  4. Judge. Two questions, both required. First, as the owner of the work: is this signable? Second, on behalf of the people the work answers to: would they accept it? Where the output failed, name which kind of failure it was: the agent couldn’t do it, or the agent did it and no one would trust it. Those are different results.
  5. Record the verdict. Give the answer one of three labels — delegable, assisted, or human-only. Put it on the grid. Note what did not have to be thought about this run. State the lesson in one line.
  6. Operate. An instrument that earns its verdict goes into real, repeated use — and use has rules: a declared rhythm; a check that what it produces matches what gets released; a visible human gate; a record of where every output came from; and a periodic review that improves it, slows it down, scales it up, or retires it. The verdict sets the width of the gate: delegable means the output is sampled; assisted means every output crosses the human before release; human-only does not operate at all. What the review finds goes back into the build — and can re-open the verdict.

This sixth step was not in the method as first published. The first experiment to complete the loop kept producing after its verdict, and the method as written had no vocabulary for what came next. This step is that vocabulary — and the revision is itself a finding.

The verdicts

Every experiment ends by answering one question: can this piece of work be handed to an agent? The verdict is that answer, made formal so it can sit on the grid and be compared, challenged, and revisited. There are three, and only three:

Verdict What it means
Delegable Signable almost as-is. Cosmetic edits only. No change to the recommendation, the framing, or the risk posture.
Assisted Signable after minor edits — and the recommendation, conclusions, and risk posture survived untouched.
Human-only Everything else — including output that saved time but had to be overturned.

The life of a verdict

A verdict is dated evidence, not settled law. It holds only for the conditions it was earned under — this model, this tooling, this act — and any cell on the grid can be re-entered when those conditions change. A verdict from a single live run is provisional: real use confirms it or overturns it in the ordinary course. Human-only is the exception — it produces no operation, so nothing re-tests it by itself. It stands only until it is deliberately challenged. Human-only is not a wall. It is a standing challenge.

The dual gate

Behind every verdict sits one discipline, applied in a fixed order.

  1. The judgment gate — a yes or no: did the thesis of the work survive? The recommendation, the conclusions, the risk posture. If the call had to be overturned, the verdict is human-only, whatever the stopwatch says.
  2. The effort gate — asked only after the first: what did signing actually cost in corrections?

Assisted does not mean faster.
Assisted means the agent carried the judgment load to an output signable with minor corrections, without rewriting the thesis. Speed alone earns nothing.

The order matters. The easiest self-deception in this field is counting the time saved on typing while ignoring the time spent deciding whether to sign.

One boundary is drawn on purpose: this Lab measures capability and trust, not return on investment. The effort gate counts the cost of signing, not the cost of building. Money enters the picture only through the operating decisions of step six — scale, slow down, retire. If the evidence one day demands an economic gate, it will arrive the way the sixth step arrived: as a published change to the method.

The prediction rule

Every published note carries its prediction: what was expected to hold, what was expected to break — and what actually happened. A failed prediction is reported as flatly as a success. And the rule reaches all the way up: if these definitions — or the method itself — prove wrong against the evidence, they are corrected in the open, and the correction is itself a note. This page is the second proof of that.


More from Uncategorized

The Grid3 min read
The Lab Opens2 min read