Writing · 02Tokyo · Incoming Cambridge HSPS

Working notes.

Short notes from inside the work: building solo with AI agents, operations, and shipping. Written when there is something concrete to say.

98Posts
46,748Words
212 minReading
2026-06-17 → 2026-09-20Span
02.1The agent series

14 posts from five months of running one AI coding agent across my own businesses, published on 17 September 2026. Open the series on its own page →

S-01 2026-09-17 Prose rules did not bend the curveFive months of running a coding agent against a rulebook that grew to 2,159 rules and 54 hooks, and the measurement that says only the rules a guard refuses in S-02 2026-09-17 Assume the instrument is wrong before the world isFive months of running twelve web properties and a coding agent taught me that a zero, a round number and a clean series are claims about the meter, and each on S-03 2026-09-17 A score the agent gives itself measures its detectorsI told my coding agent to punish itself for repeating corrected mistakes, it built a 0 to 100 score, and eleven days of the ledger's own data showed the score t S-04 2026-09-17 The complaint distribution has no headI logged 5,270 complaints about my coding agent over five months, found 2,720 distinct behaviours with no dominant one, and learned that per-behaviour guards ca S-05 2026-09-17 Well-formed and wrong: shape improved, relevance did notI spent two months fixing the format of my coding agent's replies, first-attempt compliance went from zero to 43 per cent, and over the same months the share of S-06 2026-09-17 Seven hours a day with a coding agent, and never a usage limitI measured my own agent use from the transcripts: nearly eight hours on a weekday median, 4,175 distinct prompts of my own since July, no limit hit, and the rea S-07 2026-09-17 The sessions that feel hostile are the ones that compactAcross 21 days of coding-agent sessions on my machine, the ones I described as hostile and forgetful were the ones with the most context compactions and the mos S-08 2026-09-17 A cap is a floor. Raise it and re-run.A number that lands exactly on a limit is measuring the limit, so raise the cap and run it again before anything downstream of it gets believed. S-09 2026-09-17 A 404 on a guessed URL is not evidenceMy agent invented four API paths, got four routing errors, told me the product had no API, and so missed a crawler setting that was wrong on twelve of twelve zo S-10 2026-09-17 Merged is not deployed is not servingThree links stand between a commit and the page a customer loads, each one fails in its own way, and each one needs its own proof. S-11 2026-09-17 The gate refused its own author three times before it workedA guard my agent wrote to stop unplanned sprawl fired on its own author three times in an hour, and each refusal named a different thing I had assumed about how S-12 2026-09-17 A filter cannot report what it excludedA date window compared in the wrong timezone made a family business thread read as empty, and my coding agent reported the empty result as a finding, which is w S-13 2026-09-17 Fifty-nine per cent of the log was the tool talking to itselfSixteen reader agents went through five months of my coding-agent transcripts one window each, and found that 9,039 of 15,270 logged prompts were machine-fired, S-14 2026-09-17 Ten changes after five months with a coding agentSixteen reader agents went through five months of my transcripts, and the ten rules they produced each carry the count that forced it.
02.2Archive

Earlier daily posts, kept here in full.

W-012026-09-20position 12.2 looked promisingA page at position 12.2 had 65 impressions and one click. Reaching page one was worth an estimated 2.2 extra clicks a month.359 words · 2 minW-022026-09-19A tracker that read the wrong pageMy spring week tracker filed 89 of 113 firms as having no programme, and that answer came from reading brochure pages instead of the systems the886 words · 4 minW-032026-09-19i pulled the pages sitting near page oneA page averaging position 12.1 could gain roughly 2.4 clicks per month by moving up a few places, making near-winners worth auditing.291 words · 1 minW-042026-09-18Applying to UK universities from Hong Kong and JapanThe offer turns on one day, 6 July 2026, and the visa turns on where you apply from. Everything after that is a dated list you can build early.798 words · 4 minW-052026-09-18How the BN(O) route worked for a studentStandard service is quoted at twelve weeks, the absence clock starts at grant, and where you apply from matters more than the medical test.818 words · 4 minW-062026-09-18How the spring-week applications ranThe UK bulge brackets open October to January, so September held one open programme, one rolling opener on 29 September, and the quant houses.797 words · 4 minW-072026-09-18The quote button that never renderedAn August 2026 audit found the add-to-quote button had never drawn on my family's storefront catalogue pages, so nobody could ask for a price.764 words · 3 minW-082026-09-18what's one SEO fix worth chasing this week?A page at average position 12.1 earned 67 impressions and one click. Moving it onto page one could add 2.4 monthly clicks.246 words · 1 minW-092026-09-17a page at position 12.0 deserves a targeted fix before a newA page at position 12.0 may deserve a targeted query fix before broad content publishing, with roughly 2.5 extra monthly clicks.298 words · 1 minW-102026-09-16i was staring at 12.1 againA ranking of 12.1 is a near-win. Test one focused change, measure the traffic lift, and decide if it deserves another round.423 words · 2 minW-112026-09-1567 impressionsA field note on how moving one page from average position 11.5 toward page one could create roughly 2.4 extra clicks monthly.254 words · 1 minW-122026-09-14page 11 is a different gameWhy pages averaging position 11.5 deserve focused guides and internal links before creating more content.349 words · 2 minW-132026-09-13the dumb seo move is publishing another page while page 11.2 isWhy improving near-page-one pages can beat publishing new content from scratch.277 words · 1 minW-142026-09-12pulled a page sitting at position 11.5A page at position 11.5 earned 62 impressions and 1 click. Moving it to page one could add roughly 2.1 monthly clicks.305 words · 1 minW-152026-09-11why chase new traffic when a page at position 11 is alreadyA page at position 11 may be a near-win: use a focused guide and internal link to turn existing visibility into clicks.217 words · 1 minW-162026-09-10stop treating SEO like a blank-page problemWhy improving pages already near page one can drive more SEO growth than starting from zero.281 words · 1 minW-172026-09-09a page at average position 10.8 is one push from page oneA page at average position 10.8 already earned 280 impressions and 3 clicks, with roughly 11 more monthly clicks in reach.253 words · 1 minW-182026-09-08i keep coming back to those clicksAn about page earned 294 search impressions and 3 clicks at an average position of 10.5. Visibility translated into very little traffic.409 words · 2 minW-192026-09-07391 impressionsFour pages earned 391 search impressions and 4 clicks. A field note on keeping clicks visible and avoiding fixes the totals can't justify.396 words · 2 minW-202026-09-05an accelerator can be useful before you have a businessA field note on how an accelerator can create clarity through structured pressure before revenue, investors, or a clear idea.302 words · 1 minW-212026-09-04staring at a portfolio filing with $19.47 billion in holdingsA field note on turning a $19.47 billion portfolio filing into context about changes, concentration, and timing.276 words · 1 minW-222026-09-03how much evidence do you need before you build?A field note on startup failure, evidence thresholds, validation, and disciplined execution across 235,039 launches.290 words · 1 minW-232026-09-02distribution comes after proofWhy education and AI companies need published student outcomes before scaling distribution.246 words · 1 minW-242026-09-01student input changed how AI policies were designedStudents helped shape AI policies with meaningful restrictions, showing that user input improves governance and safety design.259 words · 1 minW-252026-08-31spent an afternoon unpacking one ugly failure mode: a virtualWhy AI education products need learning outcomes alongside engagement metrics to prove that learning happened.269 words · 1 minW-262026-08-30the best AI product might be the one teachers barely noticeAI products gain adoption when they fit the workflow already in the room, especially for teachers.344 words · 2 minW-272026-08-29i price around outcomes like thisA field note on outcome-based pricing: define the result, record the baseline, track one measurement, and price the measured change.284 words · 1 minW-282026-08-28virtual tutoring companies need measured learning outcomesWhy virtual tutoring companies need measured learning outcomes before scaling, and how evidence should shape product growth.234 words · 1 minW-292026-08-27shut down after experts questioned the evidence of effectivenessWhy tutoring companies should prove customer outcomes with data before scaling distribution, according to a costly shutdown.320 words · 1 minW-302026-08-26can an AI startup really ship in eight weeks?A field note on shipping an AI startup in eight weeks through narrow scope and intensely focused execution.267 words · 1 minW-312026-08-25the filing is the raw materialWhy change detection turns filings into actionable insight by surfacing new positions, ownership changes, and portfolio shifts.253 words · 1 minW-322026-08-24students given authority over AI policy chose meaningfulStudents with a voice in AI policy chose meaningful restrictions, showing how user participation can shape safety and adoption.355 words · 2 minW-332026-08-23spent an afternoon thinking about why some AI tools actually getWhy AI tools get used in schools: they fit existing teacher workflows, and workflow friction determines adoption.267 words · 1 minW-342026-08-22the best AI feature in a school might be fitting into TuesdayWhy AI features in schools get used when they fit educators’ existing Tuesday workflow.258 words · 1 minW-352026-08-21i was halfway through another fundraising threadFundraising advice only works when its stage, business model, constraints, and economics match the company using it.400 words · 2 minW-362026-08-21validation is where startups live or dieA field note on why validation must be a disciplined operating process before faster execution compounds risk.323 words · 1 minW-372026-08-20workflow fit decides adoptionAI product adoption improves when tools fit teachers' existing workflows instead of asking them to change how they work.327 words · 1 minW-382026-08-19i’d build the learning measurement firstWhy learning measurement should come before scaling an education product, with evidence collected inside the tutoring experience.215 words · 1 minW-392026-08-18adoption and funding are weak evidence when an education productWhy education products should measure learning outcomes early instead of relying on adoption, funding, or engagement metrics.282 words · 1 minW-402026-08-1744% less capitalA field note on why staying small can require 44% less capital while reaching £100K ARR twice as often.271 words · 1 minW-412026-08-16can you really charge for student results if nobody trusts theWhy trusted achievement data and instrumentation must come before performance-based pricing for student outcomes.321 words · 1 minW-422026-08-15a big virtual tutoring provider just shut down because theyA big virtual tutoring provider shut down without evidence it worked. Here's how to instrument one outcome before you ship, not after.756 words · 3 minW-432026-08-13when do you call a commit shipped?A commit is shipped when the live route and customer experience are verified, because the repository alone cannot prove delivery.211 words · 1 minW-442026-08-11your AI spend belongs at the single model-call chokepointA field note on routing AI spend through one model-call chokepoint and keeping telemetry fail open during observability outages.282 words · 1 minW-452026-08-10a forecasting system produced a net improvement of 1.8A field note on a forecasting system's 1.8-point quarterly improvement across 47 quarters, with directional significance at t=1.45.217 words · 1 minW-462026-08-09instrument model usage at one chokepointA field note on instrumenting multi-model usage at one safe chokepoint with telemetry that fails quietly.259 words · 1 minW-472026-08-08routing every task to the strongest model is a pretty expensiveWhy routing every task to the strongest model costs more, and where free local models fit for simple, high-volume work.266 words · 1 minW-482026-08-07what if your strongest model only saw the hard cases?A field note on routing high-volume classification through a free local model and sending difficult cases to a stronger model.281 words · 1 minW-492026-08-06small AI experiments deserve a price check before they runWhy small AI experiments need exact pricing and a run ceiling before they start.221 words · 1 minW-502026-08-05send the boring work to a free local modelA field note on routing routine AI work to a free local model and reserving stronger models for judgment.325 words · 1 minW-512026-08-04a commit is a checkpoint, shipping is live verificationA field note on why clean commits do not equal shipped software, and why live customer-visible verification defines completion.311 words · 1 minW-522026-08-03how do you make a multi-provider AI system measurable?How to make a multi-provider AI system measurable by observing every model call at one chokepoint.260 words · 1 minW-532026-08-01i’d put the AI tutor inside the teacher’s workflowWhy AI tutors earn trust when goals, notes, and next actions stay inside the teacher’s workflow.236 words · 1 minW-542026-07-3011 weeksWhat 11 weeks and a 50-participant cap reveal about bounded execution, ambition, and visible progress.264 words · 1 minW-552026-07-29what actually turns AI tutoring into a real market?UK school AI funding signals demand and makes safety and learning impact part of the product trial.329 words · 1 minW-562026-07-28a durable student record should outlast whoever delivers theWhy durable student records improve handoff continuity and make personalization measurable in education.206 words · 1 minW-572026-07-27use this filter before building an AI tutoring productA field note on why AI tutoring products need teacher workflows, student goals, tutor notes, and clear next actions.233 words · 1 minW-582026-07-24a growing market gives you permission to testA growing market permits testing, while changed learner behaviour in one narrow workflow provides the real demand signal.304 words · 1 minW-592026-07-23a growing market is permission to run a narrow test, nothing moreA field note on testing one learner workflow and measuring whether it improves the next decision this week.267 words · 1 minW-602026-07-22the uk is putting up to £23m behind safe, evidence-based AI andWhy evidence-based AI and edtech trials need measurable outcomes and safety controls before schools adopt them.345 words · 2 minW-612026-07-19AI reaches families when proof and safety ship with the productWhy builders in high-trust domains need proof and safety controls in the product before distribution.262 words · 1 minW-622026-07-18start with the tutor notesAI tutoring becomes useful when tutor notes, learner goals, session evidence, and one next action stay connected.267 words · 1 minW-632026-07-17an accelerator calendar is useful when it makes the user resultA field note on using accelerator deadlines to force clearer product evidence and user results at every milestone.242 words · 1 minW-642026-07-14are education products ready for this kind of jump?Japan’s Japanese-language institute enrollment rose 23.5% in 2025, showing why education products must support cross-border transitions.276 words · 1 minW-652026-07-13a commit is a hypothesisA commit is a hypothesis. Live verification decides whether software actually does the job people need it to do.660 words · 3 minW-662026-07-13most people building ai tutoring think the hard part is theAI tutoring earns family trust through auditable progress evidence and real safety controls before the tutoring experience itself.276 words · 1 minW-672026-07-11git said the feature shippedA commit is halfway: software ships when you watch a customer complete the intended action live.251 words · 1 minW-682026-07-10a commit is just a checkpointA commit is only a checkpoint; i call work done after verifying it in the live environment the user actually experiences.580 words · 3 minW-692026-07-09a green diff lied to me last weekA green diff passed tests while prod stayed broken, and the fix was a rule to verify the live thing.677 words · 3 minW-702026-07-07spent an hour last night chasing a bug that was never a buga stale copy of the same api key in an env file caused a machine-only failure until secrets moved to one OS-backed store.390 words · 2 minW-712026-07-06why are secrets still living in env files and shell history?Centralizing secrets in one store and loading them at runtime makes rotation boring and stops drift across env files, shells, and deploy targets.419 words · 2 minW-722026-07-03a merged pr is not a shipGreen CI proves the code matched expectations; live verification is the receipt that shows it actually worked.654 words · 3 minW-732026-07-02an agent budget breaker for your ai toolinga hook that hard-blocks subagent dispatches at 8 per day. installed after a 132-dispatch week burned 85% of a max plan, and what the cap actually taught me about routing.495 words · 2 minW-742026-07-02your outreach needs a validatora hard gate for outreach drafts: single line, product link required, em dashes banned, template clones rejected. it failed 358 of the 414 drafts sitting in my queues.493 words · 2 minW-752026-07-01the real bar is live verificationa passing test can still miss the real user experience; live verification in the real environment is the standard that matters.533 words · 2 minW-762026-06-26moved every runtime secret into the OS keychain and ripped themoving runtime secrets into one canonical OS keychain made rotation and recovery a one-place problem across four services.386 words · 2 minW-772026-06-25a commit is a checkpointA commit is a checkpoint, and shipping only starts after live verification exposes the hidden failures.663 words · 3 minW-782026-06-24when is a change actually done?A change is done when the live service is verified end to end, because repo-local green can create false confidence.583 words · 3 minW-792026-06-23your offline backup is not your source of truthA backup you write to is just a second thing to keep in sync; a backup you only read from in a disaster actually saves you.492 words · 2 minW-802026-06-22a green checkmark in your repo is not proof of anythingA green checkmark in the repo can hide a stale live endpoint, and the only reliable receipt is a live request you make yourself.677 words · 3 minW-812026-06-20shipped a feature last weeka green build and clean merge can still miss the user path; if nobody walked it live, the feature was not shipped yet.439 words · 2 minW-822026-06-19production verification is part of doneproduction verification belongs in the definition of done, because the live check is where the answer shows up.307 words · 1 minW-832026-06-18one source of truth for repos, services, creds, automationsA single source of truth keeps repos, services, access details, and automations in sync by updating the record with the change.378 words · 2 minW-842026-06-17Moving every secret into one canonical store made our env files boring againMoving secrets into one canonical store made env files boring again and made rotation and drift easier to manage.677 words · 3 min