Series · 14Tokyo · Incoming Cambridge HSPS
Five months of running an AI coding agent, measured.
14 posts, published on 17 September 2026. Each one starts from a number I counted in my own transcripts, logs or ledgers while running small businesses with one coding agent. Where a measurement contradicted something I had said, the post says so.
S-01
2026-09-17
Prose rules did not bend the curveFive months of running a coding agent against a rulebook that grew to 2,159 rules and 54 hooks, and the measurement that says only the rules a guard refuses in
→
S-02
2026-09-17
Assume the instrument is wrong before the world isFive months of running twelve web properties and a coding agent taught me that a zero, a round number and a clean series are claims about the meter, and each on
→
S-03
2026-09-17
A score the agent gives itself measures its detectorsI told my coding agent to punish itself for repeating corrected mistakes, it built a 0 to 100 score, and eleven days of the ledger's own data showed the score t
→
S-04
2026-09-17
The complaint distribution has no headI logged 5,270 complaints about my coding agent over five months, found 2,720 distinct behaviours with no dominant one, and learned that per-behaviour guards ca
→
S-05
2026-09-17
Well-formed and wrong: shape improved, relevance did notI spent two months fixing the format of my coding agent's replies, first-attempt compliance went from zero to 43 per cent, and over the same months the share of
→
S-06
2026-09-17
Seven hours a day with a coding agent, and never a usage limitI measured my own agent use from the transcripts: nearly eight hours on a weekday median, 4,175 distinct prompts of my own since July, no limit hit, and the rea
→
S-07
2026-09-17
The sessions that feel hostile are the ones that compactAcross 21 days of coding-agent sessions on my machine, the ones I described as hostile and forgetful were the ones with the most context compactions and the mos
→
S-08
2026-09-17
A cap is a floor. Raise it and re-run.A number that lands exactly on a limit is measuring the limit, so raise the cap and run it again before anything downstream of it gets believed.
→
S-09
2026-09-17
A 404 on a guessed URL is not evidenceMy agent invented four API paths, got four routing errors, told me the product had no API, and so missed a crawler setting that was wrong on twelve of twelve zo
→
S-10
2026-09-17
Merged is not deployed is not servingThree links stand between a commit and the page a customer loads, each one fails in its own way, and each one needs its own proof.
→
S-11
2026-09-17
The gate refused its own author three times before it workedA guard my agent wrote to stop unplanned sprawl fired on its own author three times in an hour, and each refusal named a different thing I had assumed about how
→
S-12
2026-09-17
A filter cannot report what it excludedA date window compared in the wrong timezone made a family business thread read as empty, and my coding agent reported the empty result as a finding, which is w
→
S-13
2026-09-17
Fifty-nine per cent of the log was the tool talking to itselfSixteen reader agents went through five months of my coding-agent transcripts one window each, and found that 9,039 of 15,270 logged prompts were machine-fired,
→
S-14
2026-09-17
Ten changes after five months with a coding agentSixteen reader agents went through five months of my transcripts, and the ten rules they produced each carry the count that forced it.
→