Skip to content
← Posts

Nobody volunteers to be the control group

Nobody is running the control arm, including me. Every productivity number in this field is therefore unfalsifiable, and we should say so.

4 min read

I believe agents have made me several times faster. I am labeling that as a belief on purpose. I have never done the same work without them as a control, so I do not have a number I would defend in a hostile reading.

That is not a small omission. Every productivity claim in this field is a claim about a difference. A difference needs two measurements. Almost nobody has two. I do not have two either.

I recently audited the scheduled estate I actually run: 24 tasks, 22 enabled, about 41 invocations a day. That is a real measurement, and it is a measurement of throughput. It counts how often the schedule woke something up and asked it to produce an artifact. It does not count whether any of those artifacts were the right work, or whether I would have been better off doing less. The audit could not answer what the estate costs, either. There is no per-task token figure, no API figure, no dollar figure in it, and I am not going to invent one so the essay has a punchy number.

I feel faster than I was two years ago. Feeling faster is easy to confuse with being faster. The work I do now is also different work: more specifying, reviewing, and interrupting, less typing the first draft. Comparing today's output to 2024's is comparing two jobs that only look like the same job from far away. That is not a control arm. That is a vibe with a calendar attached.

A control arm would be boring. Take a piece of work with a definition of done you already trust. Do it with agents. Do the same work without them, under the same definition of done. Time both. Judge both. Publish the comparison, including the cases where the agent path was slower or worse. I have not run that, and I have not asked anyone to run it for me.

The reasons are not mysterious. Nobody volunteers to do the work the slow way on purpose. Nobody funds it. Asking a team to redo a week of tickets by hand so a write-up can have a denominator is a research project dressed up as engineering, and inside a company it is worse: if the mandated number is 100x, or 30 percent, or whatever this quarter's slide says, contradicting it is a commercial risk. People who like their jobs do not volunteer to be the control group.

So we get two piles of numbers, both unfalsifiable. One pile says 100x: demo videos, launch posts, internal surveys where people estimate their own speedup after the tool has already been purchased. The other pile says it makes you slower: careful task studies, the essays that get passed around as a corrective. I am not saying either pile is lying. I am saying both argue from designs that do not include a control the reader can inspect, or from tasks that are not the work people are actually doing. You can believe either pile. You cannot adjudicate them with the evidence attached.

My own honest accounting is narrower than either pile. I can measure artifact throughput: the schedule, the run counts, whether files are still being written. That is not nothing. An estate that cannot even say whether a run happened is not ready to talk about productivity. I cannot measure whether it was the right work. Forty-one invocations is a busy day or a wasted one, and the count will not tell me which. And I cannot measure the thing the field wants most, the speedup. I have a belief, a calendar, and a lot of output. That is the same evidence base I am criticizing.

There is a subsidy question under this too. A lot of this work is cheap because a flat plan hides the meter. When the meter arrives, the honest question is not "was I 100x." It is "which of these daily invocations was worth the new price." The audit cannot answer that. It captured motion, not cost.

I know what an honest control arm would cost me: a week I would rather spend shipping, on work done twice on purpose, with a write-up at the end that might show I am not as fast as I think. That is why I have not run it. Not because the design is hard. Because I have not been willing to be the slow group.

I would rather say that than add another unsupported number to either pile. Until I run the arm, "several times faster" keeps the word believe attached.

Related: The control arm, Verification