Skip to content
← Posts

In defense of the control arm

Nobody is running the control arm, including me. Every productivity number in this field is therefore unfalsifiable, and we should say so.

5 min read

I believe agents have made me several times faster. I am labeling that as a belief on purpose. I have never done the same work without them as a control, so I do not have a number I would defend in a hostile reading.

That is not a small omission. Every productivity claim in this field is a claim about a difference. A difference needs two measurements. Almost nobody has two. I do not have two either.

On August 12, 2026 I audited the scheduled estate I actually run. The inventory was 24 tasks, 22 of them enabled, about 41 invocations a day. That is a real measurement. It is also a measurement of throughput. It counts how often the schedule woke something up and asked it to produce an artifact. It does not count whether any of those artifacts were the right work, or whether I would have been better off doing less.

The audit could not answer what the estate costs. There is no per-task token figure, no API figure, and no dollar figure in that document. I am not going to invent one here so the essay has a punchy number. I can tell you how often the machine runs. I cannot tell you what I paid, and I cannot tell you what I got in return that I would still want if I had to pay it twice.

I feel faster than I was two years ago. Feeling faster is easy to confuse with being faster. The work I do now is also different work. I spend more time specifying, reviewing, and interrupting, and less time typing the first draft. If I compared today's output to 2024's output, I would be comparing two jobs that only look like the same job from far away. That is not a control arm. That is a vibe with a calendar attached.

A control arm would be boring. Take a piece of work with a definition of done you already trust. Do it with agents. Do the same piece of work without them, or have someone else do it without them, under the same definition of done. Time both. Judge both. Publish the comparison, including the cases where the agent path was slower or worse. I have not run that. I have not asked anyone to run it for me.

The reasons are not mysterious. Nobody volunteers to do the work the slow way on purpose. Nobody funds it. The slow way is the thing we are all trying to stop paying for. Asking a team to redo a week of tickets by hand, so that a write-up can have a denominator, is a research project dressed up as engineering. Inside a company it is worse than that. If the mandated number is 100x, or 30 percent, or whatever this quarter's slide says, contradicting it is a commercial risk. It is not treated as a measurement problem. People who like their jobs do not volunteer to be the control group.

So we get two piles of numbers, both unfalsifiable, both drawn from the same missing comparison.

One pile says 100x. Demo videos, launch posts, internal surveys where people estimate their own speedup after the tool has already been purchased. The other pile says it makes you slower. Careful task studies, the essays that get passed around as a corrective. I am not saying both piles are lying. I am saying both piles are mostly arguing from designs that do not include a control the reader can inspect, or from studies whose tasks are not the work people are actually doing. You can believe either pile. You cannot adjudicate them with the evidence that is usually attached.

My own honest accounting is narrower than either pile.

I can measure artifact throughput. I know the schedule. I know roughly how many times a day something runs. I can open a directory and see whether files are still being written. That is not nothing. An estate that cannot even say whether a run happened is not ready to talk about productivity.

I cannot measure whether it was the right work. Forty-one invocations is a busy day or a wasted one, and the count will not tell me which. A research file still waiting to be read, a monitor that wrote a healthy timestamp and nothing else: those look like throughput. Some of that output is useful. Some of it is the machine keeping itself occupied.

I also cannot measure the thing the field wants most, which is the speedup. I have no paired trial. I have a belief, a calendar, and a lot of output. That is the same evidence base I am criticizing.

There is a subsidy question sitting under this, and I do not have the numbers for that either. A lot of this work is cheap because a flat plan hides the meter, or because the unit cost is not something I have recorded. When the meter arrives, the honest question is not "was I 100x." It is "which of these daily invocations was worth the new price." I will not be able to answer that from the August audit. The audit did not capture cost. It captured motion.

I know what an honest control arm would cost me. It would cost a week I would rather spend shipping, on work I would be doing twice on purpose, with a write-up at the end that might show I am not as fast as I think. That is why I have not run it. The reason is not that the design is hard. The reason is that I have not been willing to be the slow group.

I would rather say that than add another unsupported number to either pile.

I am still working out how to talk about this without performing certainty I have not earned. The belief is real. The control arm is not. Until I run one, "several times faster" is a sentence I should keep attaching to the word believe.

Related: The control arm, Verification