<?xml version="1.0" encoding="utf-8"?>
<feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en">
  <id>https://howardchan.me/</id>
  <title type="text">Howard Chan</title>
  <subtitle type="text">Founder and builder in Tokyo, incoming Cambridge HSPS.</subtitle>
  <updated>2026-09-18T00:59:00+09:00</updated>
  <link rel="self" type="application/atom+xml" href="https://howardchan.me/feed.xml"/>
  <link rel="alternate" type="text/html" href="https://howardchan.me/writing/"/>
  <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
  <rights>Copyright 2026 Chak Hang (Howard) Chan</rights>
  <generator uri="https://howardchan.me/stack/">build.mjs</generator>
  <entry>
    <id>https://howardchan.me/writing/applying-to-uk-universities-from-hong-kong-and-japan/</id>
    <title type="text">Applying to UK universities from Hong Kong and Japan</title>
    <updated>2026-09-18T00:59:00+09:00</updated>
    <published>2026-09-18T00:59:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/applying-to-uk-universities-from-hong-kong-and-japan/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">The offer turns on one day, 6 July 2026, and the visa turns on where you apply from. Everything after that is a dated list you can build early.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; The offer turns on one day, 6 July 2026, and the visa turns on where you apply from. Everything after that is a dated list you can build early.&lt;/p&gt;
&lt;h2&gt;Where I was applying from&lt;/h2&gt;
&lt;p&gt;I left Hong Kong in April 2022 and sat the International Baccalaureate at an international school in Tokyo, August 2022 to May 2026. I finished on 44 out of 45, with a grade 7 in all six subjects and an A on the extended essay, and an SAT of 1550. My offer was Human, Social, and Political Sciences at Peterhouse, Cambridge, a three year BA starting in October 2026.&lt;/p&gt;&lt;p&gt;My case was ordinary residence in Hong Kong, schooling in Japan on a dependent visa, and international fee status in the UK. Each of those changes something later in the year.&lt;/p&gt;
&lt;h2&gt;Results day decides almost everything&lt;/h2&gt;
&lt;p&gt;IB results reach students at 12:00 GMT on 6 July 2026. Universities and IB coordinators see them about a day earlier, on 5 July. That one day lead is the warning I would give any IB applicant who has been reading A-level advice: the college does not know your grades weeks ahead, and nothing is being settled quietly in June.&lt;/p&gt;&lt;p&gt;Meet the conditions and the place is confirmed on results day. Fall just short and the college reconsiders, then pools you to another college through the summer pool, which resolves a few days into mid July. There is no separate appeal to write. The August reconsideration pool that opens on 13 August is for UK-domiciled applicants, so on an international fee from Tokyo I was outside it and planned as though it did not exist.&lt;/p&gt;&lt;p&gt;The plan for the other offers was to hold each one and decline it a day before its own deadline, once 6 July had confirmed Cambridge. The reply-by dates sit in each offer letter or portal. Holding an offer costs nothing. Releasing one in June, in a year where results can still move, can cost the year.&lt;/p&gt;
&lt;h2&gt;The visa question is residence first, paperwork second&lt;/h2&gt;
&lt;p&gt;The route I used requires that you are ordinarily resident in Hong Kong on the date of application. Four years in Japan reads oddly against that line. The Home Office guidance says the hard disqualifier is permanent or settled status in another country, and a dependent visa is temporary, so it did not apply to me. A caseworker can still ask. The plan was to apply while physically in Hong Kong, with Hong Kong as the permanent home, and to describe the Japan years as a temporary relocation.&lt;/p&gt;&lt;p&gt;Then timing. Standard service is free and the published figure is a decision within 12 weeks of submission. Priority costs around 500 pounds and returns a decision in five working days from outside the UK, and you choose it at submission, because a standard application cannot be upgraded later. The only way round is to withdraw and resubmit, which blocks travel while it is pending. Twelve weeks from a 7 July filing runs to about 29 September, against a college that wanted me in residence by 4 October. That arithmetic is why the filing date was set to the day after results.&lt;/p&gt;
&lt;h2&gt;The dated steps, and the document each one needed&lt;/h2&gt;
&lt;div class=&quot;tablewrap&quot; tabindex=&quot;0&quot;&gt;&lt;table class=&quot;ptable&quot;&gt;&lt;thead&gt;&lt;tr&gt;&lt;th scope=&quot;col&quot;&gt;Date&lt;/th&gt;&lt;th scope=&quot;col&quot;&gt;Step&lt;/th&gt;&lt;th scope=&quot;col&quot;&gt;What it needed&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td data-label=&quot;Date&quot;&gt;6 July 2026&lt;/td&gt;&lt;td data-label=&quot;Step&quot;&gt;IB results released to students at 12:00 GMT&lt;/td&gt;&lt;td data-label=&quot;What it needed&quot;&gt;the results statement&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td data-label=&quot;Date&quot;&gt;7 July 2026&lt;/td&gt;&lt;td data-label=&quot;Step&quot;&gt;Visa filing planned for the next day on standard service&lt;/td&gt;&lt;td data-label=&quot;What it needed&quot;&gt;passport, confirmed offer&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td data-label=&quot;Date&quot;&gt;Mid July 2026&lt;/td&gt;&lt;td data-label=&quot;Step&quot;&gt;A near miss resolves through reconsideration and the summer pool&lt;/td&gt;&lt;td data-label=&quot;What it needed&quot;&gt;nothing to file&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td data-label=&quot;Date&quot;&gt;Each offer&amp;#39;s reply-by date&lt;/td&gt;&lt;td data-label=&quot;Step&quot;&gt;Other offers declined, one day before each deadline&lt;/td&gt;&lt;td data-label=&quot;What it needed&quot;&gt;the offer letter carrying that date&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td data-label=&quot;Date&quot;&gt;By 1 October 2026&lt;/td&gt;&lt;td data-label=&quot;Step&quot;&gt;University registration window closes&lt;/td&gt;&lt;td data-label=&quot;What it needed&quot;&gt;visa share code, sent three working days before arrival&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td data-label=&quot;Date&quot;&gt;30 September to 4 October 2026&lt;/td&gt;&lt;td data-label=&quot;Step&quot;&gt;Room licence starts, in college by 4 October&lt;/td&gt;&lt;td data-label=&quot;What it needed&quot;&gt;qualification certificates in hard copy&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/div&gt;&lt;p&gt;Full Term runs from 6 October to 4 December 2026.&lt;/p&gt;
&lt;h2&gt;What the first year actually asks&lt;/h2&gt;
&lt;p&gt;Part IA is four papers, three of them from the core subjects of politics, international relations, social anthropology and sociology. There is no Part IB in this course, and saying otherwise marks you as someone who never read the course page. Teaching is around eight hours of lectures a week with one or two supervisions, and assessment is mostly by examination, with some coursework.&lt;/p&gt;&lt;p&gt;Track choice comes towards the end of the first year, and you cannot switch between Part IIA and Part IIB tracks unless you move from a joint track into one of the single subjects inside it. The first-year paper set is the only free move you get, which is why I took the whole core.&lt;/p&gt;
&lt;h2&gt;The one thing I would tell you&lt;/h2&gt;
&lt;p&gt;Build the calendar before you need it, and write the document beside every date. My year had four gates, each a date with a document attached: results on 6 July, the visa filing set for 7 July because the published decision time ran close to the college&amp;#39;s own deadline, registration closing on 1 October, and a room I could not occupy before 30 September. None of it was hard, and all of it was time sensitive, done across time zones in a summer when the school holding your certificates has closed. Write the dates down in March, ask which document each one wants, and results day becomes one day of news instead of the start of a scramble.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/how-the-bno-route-worked-for-a-student/</id>
    <title type="text">How the BN(O) route worked for a student</title>
    <updated>2026-09-18T00:58:00+09:00</updated>
    <published>2026-09-18T00:58:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/how-the-bno-route-worked-for-a-student/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">Standard service is quoted at twelve weeks, the absence clock starts at grant, and where you apply from matters more than the medical test.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; Standard service is quoted at twelve weeks, the absence clock starts at grant, and where you apply from matters more than the medical test.&lt;/p&gt;
&lt;h2&gt;In short&lt;/h2&gt;
&lt;p&gt;I applied for a BN(O) visa in the summer of 2026 to start a Cambridge degree on 1 October. This is what the route actually asked of me, in the order it asked.&lt;/p&gt;
&lt;h2&gt;Applying in my own name&lt;/h2&gt;
&lt;p&gt;The BN(O) route lets a young adult apply alone. You need to be 18 or over, born after 1 July 1997, and you need a parent who holds BN(O) status. I had turned 18 that May, so I went on my own application rather than as a dependant on a parent&amp;#39;s. The grant is a five-year visa, and it is the same five years that later counts towards settlement.&lt;/p&gt;
&lt;h2&gt;Standard service, and the date I filed&lt;/h2&gt;
&lt;p&gt;There are two lanes. Standard costs nothing extra. The paid priority service exists for out-of-country BN(O) applications and returns a decision in five working days, subject to quota. The rule that decides everything is this one: once you submit on standard, you cannot upgrade. Your only move is to withdraw and resubmit, which restarts the clock.&lt;/p&gt;&lt;p&gt;The official standard time is &amp;quot;within 12 weeks&amp;quot; from submission. My matriculation was 1 October 2026. IB results were released to students on 6 July, and the plan was to file on 7 July, the day after the offer was confirmed. Twelve weeks from 7 July lands around 29 September. That is two days of margin against a hard wall, which is why filing the morning after results mattered so much. Planning to the two to four weeks that grants usually take is planning to an average.&lt;/p&gt;&lt;p&gt;One detail if you might miss the offer: universities and IB coordinators see results at most a day before students do, so nothing is settled weeks ahead. A near miss goes to reconsideration and then the summer pool, resolving a few days into mid-July, and I would have held the filing until that cleared.&lt;/p&gt;
&lt;h2&gt;The question that was actually risky&lt;/h2&gt;
&lt;p&gt;Everyone asks about the TB test. It did not apply to me: the list works on where you are resident, and Japan, where I had lived since April 2022, is off it.&lt;/p&gt;&lt;p&gt;The real question was ordinary residence. An out-of-country BN(O) application requires you to be ordinarily resident in Hong Kong on the date of application, and I had spent four years in Japan on a dependent visa. Home Office guidance names one hard disqualifier, permanent or settled status in another country, and a dependent visa is not settled, so the disqualifier did not bite. Temporary study, work or family absence is fine provided you were ordinarily resident in Hong Kong before you left. Still, four years is a long absence and my whole family had moved, so a caseworker could reasonably raise it.&lt;/p&gt;&lt;p&gt;My answer was to remove the question: apply from Hong Kong while physically there, with Hong Kong as the permanent home, and frame Japan as a temporary relocation. Identity was done through the UK Immigration: ID Check app, which reads the passport chip and leaves the passport in your hands, so travel while the decision is pending stays possible.&lt;/p&gt;&lt;p&gt;On money, all I needed was parental funds, evidenced by six months of statements and a signed letter. No UK bank account, and nothing sitting in my own name.&lt;/p&gt;
&lt;h2&gt;The clock starts at grant&lt;/h2&gt;
&lt;p&gt;My visa came through earlier than I expected, and that cost me two months before I had set foot in the country. Absence limits run from the grant date, so the gap between grant and entry is spent time.&lt;/p&gt;&lt;div class=&quot;tablewrap&quot; tabindex=&quot;0&quot;&gt;&lt;table class=&quot;ptable&quot;&gt;&lt;thead&gt;&lt;tr&gt;&lt;th scope=&quot;col&quot;&gt;Rule&lt;/th&gt;&lt;th scope=&quot;col&quot;&gt;Window&lt;/th&gt;&lt;th scope=&quot;col&quot;&gt;Limit&lt;/th&gt;&lt;th scope=&quot;col&quot;&gt;Clock starts&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td data-label=&quot;Rule&quot;&gt;Indefinite leave to remain&lt;/td&gt;&lt;td data-label=&quot;Window&quot;&gt;any rolling 12 months&lt;/td&gt;&lt;td data-label=&quot;Limit&quot;&gt;180 days outside the UK&lt;/td&gt;&lt;td data-label=&quot;Clock starts&quot;&gt;date of grant&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td data-label=&quot;Rule&quot;&gt;Naturalisation, whole period&lt;/td&gt;&lt;td data-label=&quot;Window&quot;&gt;the 5 years before applying&lt;/td&gt;&lt;td data-label=&quot;Limit&quot;&gt;450 days outside the UK&lt;/td&gt;&lt;td data-label=&quot;Clock starts&quot;&gt;5 years before applying&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td data-label=&quot;Rule&quot;&gt;Naturalisation, final year&lt;/td&gt;&lt;td data-label=&quot;Window&quot;&gt;the last 12 months before applying&lt;/td&gt;&lt;td data-label=&quot;Limit&quot;&gt;90 days outside the UK&lt;/td&gt;&lt;td data-label=&quot;Clock starts&quot;&gt;12 months before applying&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td data-label=&quot;Rule&quot;&gt;My own working rule&lt;/td&gt;&lt;td data-label=&quot;Window&quot;&gt;each year from grant&lt;/td&gt;&lt;td data-label=&quot;Limit&quot;&gt;270 days inside the UK&lt;/td&gt;&lt;td data-label=&quot;Clock starts&quot;&gt;date of grant&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/div&gt;&lt;p&gt;The 270 is mine, taken from the 450-day allowance spread across five years and rounded down for safety. The plan is settlement after five years, then citizenship twelve months after that, around 2032. I keep an absence log from the grant date, because reconstructing one six years later from boarding passes is a miserable job.&lt;/p&gt;
&lt;h2&gt;What the college wants before you land&lt;/h2&gt;
&lt;p&gt;The visa is digital, so there is no sticker to show. My college asks for the eVisa share code through its own form no later than three working days before arrival. The form names Student Visa holders, which a BN(O) holder is not, so I submitted it anyway and said which visa it was. Passport and proof of entry go to the Tutorial Office on arrival, and the course does not start until those are checked.&lt;/p&gt;&lt;p&gt;If I were telling the next student one thing, it would be to treat the submit button as the last decision point and to make it as early as the offer allows. Everything after it is out of your hands. Standard service runs to a published twelve weeks whatever anyone says about typical times, and the only thing you still control on results day is whether you file the next morning.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/how-the-spring-week-applications-ran/</id>
    <title type="text">How the spring-week applications ran</title>
    <updated>2026-09-18T00:57:00+09:00</updated>
    <published>2026-09-18T00:57:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/how-the-spring-week-applications-ran/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">The UK bulge brackets open October to January, so September held one open programme, one rolling opener on 29 September, and the quant houses.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; The UK bulge brackets open October to January, so September held one open programme, one rolling opener on 29 September, and the quant houses.&lt;/p&gt;
&lt;h2&gt;In short&lt;/h2&gt;
&lt;p&gt;I applied to UK spring weeks in September 2026, as a first year arriving from abroad. Most of what I had read about the timing was wrong.&lt;/p&gt;
&lt;h2&gt;The September wave was somebody else&amp;apos;s calendar&lt;/h2&gt;
&lt;p&gt;Coaching sites told me the large banks open in the first three weeks of September. On 18 September I had 113 firms read against their own careers pages. One programme was open. One was shut. Four had published a later opening date. Seven described a spring programme with no form behind it. Ninety-seven had no spring page I could find, and three pages would not load.&lt;/p&gt;&lt;p&gt;The trackers put J.P. Morgan, Barclays and Morgan Stanley in October to January. So September came down to one open programme, one rolling programme opening on 29 September, and the quant houses. October is when the volume arrives. For a first year moving countries that matters, because term starts before the applications do. I put the practice into September and kept my evenings free from October.&lt;/p&gt;
&lt;h2&gt;The seven programmes, and the date each turns into work&lt;/h2&gt;
&lt;div class=&quot;tablewrap&quot; tabindex=&quot;0&quot;&gt;&lt;table class=&quot;ptable&quot;&gt;&lt;thead&gt;&lt;tr&gt;&lt;th scope=&quot;col&quot;&gt;Firm and programme&lt;/th&gt;&lt;th scope=&quot;col&quot;&gt;Status on 18 September 2026&lt;/th&gt;&lt;th scope=&quot;col&quot;&gt;Date to act&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td data-label=&quot;Firm and programme&quot;&gt;Rothschild &amp;amp; Co, 2027 UK Global Advisory Spring Insight&lt;/td&gt;&lt;td data-label=&quot;Status on 18 September 2026&quot;&gt;Opens later, rolling once open&lt;/td&gt;&lt;td data-label=&quot;Date to act&quot;&gt;29 September 2026&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td data-label=&quot;Firm and programme&quot;&gt;BlackRock, 2027 Spring Insight Event EMEA&lt;/td&gt;&lt;td data-label=&quot;Status on 18 September 2026&quot;&gt;Open, closes 4 December, not rolling&lt;/td&gt;&lt;td data-label=&quot;Date to act&quot;&gt;This week&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td data-label=&quot;Firm and programme&quot;&gt;Evercore, Spring Insight London&lt;/td&gt;&lt;td data-label=&quot;Status on 18 September 2026&quot;&gt;Opens later, students page says rolling&lt;/td&gt;&lt;td data-label=&quot;Date to act&quot;&gt;Check 5 October 2026&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td data-label=&quot;Firm and programme&quot;&gt;Macquarie, Spring Insight&lt;/td&gt;&lt;td data-label=&quot;Status on 18 September 2026&quot;&gt;Programme described, no form live&lt;/td&gt;&lt;td data-label=&quot;Date to act&quot;&gt;Check 11 October 2026&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td data-label=&quot;Firm and programme&quot;&gt;Deloitte, Spring into Deloitte: Advisory and Finance&lt;/td&gt;&lt;td data-label=&quot;Status on 18 September 2026&quot;&gt;Opens later, closes once places fill&lt;/td&gt;&lt;td data-label=&quot;Date to act&quot;&gt;9 November 2026&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td data-label=&quot;Firm and programme&quot;&gt;Schroders, Spring Insight Programme&lt;/td&gt;&lt;td data-label=&quot;Status on 18 September 2026&quot;&gt;Opens later, two-week window&lt;/td&gt;&lt;td data-label=&quot;Date to act&quot;&gt;January 2027&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/div&gt;&lt;p&gt;The seventh is J.P. Morgan, whose page shows the spring programme shut with no opening date, so I look weekly and apply the day it changes. Two rows above carry a date from a tracker rather than from the firm: Evercore and Macquarie. Those are a day to look, never a day to apply. The live version, all 113 firms, is at https://howardchan.me/tools/.&lt;/p&gt;&lt;p&gt;BlackRock closes on 4 December and does not read applications as they arrive. Rothschild reads them as they arrive and publishes no closing date. That difference sets the order. BlackRock still goes first, in the week of 18 September, to run the whole pipeline once before the firm I care about opens.&lt;/p&gt;
&lt;h2&gt;The five stages, and what each is reading&lt;/h2&gt;
&lt;p&gt;Every one of these programmes runs the same five stages: the form, an online test set, a recorded video interview, two back-to-back 30-minute interviews, and an assessment centre.&lt;/p&gt;&lt;p&gt;The form is read by a filter, then by a recruiter for seconds. Be literal. Name the programme exactly as the firm names it. Graduation year on the first line. Then one motivation paragraph per firm, 150 to 250 words, built from three sentences whose shape never changes: what I have actually done that is commercial, one fact about this firm and why it connects, and what I want from the week stated as a question I cannot answer from outside. Rewritten for every firm.&lt;/p&gt;&lt;p&gt;The tests come in three kinds. Numerical and verbal reasoning under time, where the loss on a first sitting is time management, so the plan is one of each against a clock before any real one. Situational judgement, where the key rewards telling the senior person first and then acting, and penalises the answer that solves it alone. Game-based tests, which cannot be practised. The only input I control there is sleep the night before.&lt;/p&gt;&lt;p&gt;The video interview gives four or five questions, 30 seconds to prepare and 90 seconds to answer. Thirty seconds is enough to choose a story and too short to build one. So I wrote five stories and indexed them by the question each one answers: leadership, failure, data, initiative, teamwork. The index is what goes on the post-it.&lt;/p&gt;&lt;p&gt;The assessment centre runs half a day to a full day. The group exercise scores contribution rather than airtime, so I propose a structure in the first two minutes and bring in the quiet person by name. The case exercise wants a recommendation with one number attached, plus the two things that would change it. Lunch is scored too.&lt;/p&gt;
&lt;h2&gt;Drafting against a form that is not open&lt;/h2&gt;
&lt;p&gt;Four of the seven forms were shut on 18 September. I wrote the answers anyway, with every question marked unverified and to be confirmed on the day the form opens. The drafts run against what such forms always ask: motivation, plus the team and commercial-awareness examples. When Rothschild opens on 29 September I am editing rather than starting.&lt;/p&gt;
&lt;h2&gt;The one mistake&lt;/h2&gt;
&lt;p&gt;Take every date from the firm&amp;#39;s own page. Every wrong date I began with came from somewhere else: a coaching site&amp;#39;s September wave, a tracker&amp;#39;s expected opening, a search result pointing at last year&amp;#39;s expired listing. Second-hand dates fail in both directions and both cost. Too early, and I sit waiting for something already open elsewhere. Too late, and a rolling programme has been reading applications for two weeks before mine lands. The firm&amp;#39;s page is the only place that carries its own sentence with its own date in it, and reading 113 of them took a morning.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/the-quote-button-that-never-rendered/</id>
    <title type="text">The quote button that never rendered</title>
    <updated>2026-09-18T00:56:00+09:00</updated>
    <published>2026-09-18T00:56:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/the-quote-button-that-never-rendered/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">An August 2026 audit found the add-to-quote button had never drawn on my family&apos;s storefront catalogue pages, so nobody could ask for a price.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; An August 2026 audit found the add-to-quote button had never drawn on my family&amp;apos;s storefront catalogue pages, so nobody could ask for a price.&lt;/p&gt;
&lt;h2&gt;The audit was meant to be about photographs&lt;/h2&gt;
&lt;p&gt;I opened my family&amp;#39;s trophy and awards storefront in August 2026 to fix catalogue photographs and category labels. That was the whole brief. The business prices every job by hand, so the site carries no shelf prices and no minimum order. A customer browses, adds items to a quote basket, submits the basket, and someone on the business side prices the work and replies by email or message. The site does one job. It captures the request and hands it over. If that path is broken, the site is a picture gallery.&lt;/p&gt;
&lt;h2&gt;A button that was never on the page&lt;/h2&gt;
&lt;p&gt;The add-to-quote control on an individual product page worked. The version of it that belongs on the catalogue grid, the page a person actually lands on from a search or an ad, never drew at all. The grid is rendered by a page-builder plugin that never calls the theme hook the button hangs on. The code for the button existed. It was installed and it was correct. It had nowhere to appear.&lt;/p&gt;&lt;p&gt;So a visitor who reached a category page saw rows of products and no way to ask about any of them. This had been true for as long as that grid had been in place. I had assumed the control was on the page because the code for the control was in the repository, which is a bad habit and the reason I now open the page instead of the file.&lt;/p&gt;
&lt;h2&gt;What the audit checked&lt;/h2&gt;
&lt;div class=&quot;tablewrap&quot; tabindex=&quot;0&quot;&gt;&lt;table class=&quot;ptable&quot;&gt;&lt;thead&gt;&lt;tr&gt;&lt;th scope=&quot;col&quot;&gt;What the audit checked&lt;/th&gt;&lt;th scope=&quot;col&quot;&gt;What it found&lt;/th&gt;&lt;th scope=&quot;col&quot;&gt;When&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td data-label=&quot;What the audit checked&quot;&gt;Whether the add-to-quote control drew on catalogue grid pages&lt;/td&gt;&lt;td data-label=&quot;What it found&quot;&gt;It never drew. The grid plugin bypasses the hook the button hangs on&lt;/td&gt;&lt;td data-label=&quot;When&quot;&gt;August 2026&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td data-label=&quot;What the audit checked&quot;&gt;Whether quote basket submissions were reaching the ad account&lt;/td&gt;&lt;td data-label=&quot;What it found&quot;&gt;0 over 30 days, with the tag correct and connected&lt;/td&gt;&lt;td data-label=&quot;When&quot;&gt;August 2026&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td data-label=&quot;What the audit checked&quot;&gt;How many of the account&amp;#39;s conversion actions carried any meaning&lt;/td&gt;&lt;td data-label=&quot;What it found&quot;&gt;18 enabled, 4 of them meaningful&lt;/td&gt;&lt;td data-label=&quot;When&quot;&gt;August 2026&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td data-label=&quot;What the audit checked&quot;&gt;Whether the basket page held what a visitor had added&lt;/td&gt;&lt;td data-label=&quot;What it found&quot;&gt;It did not. A shared cache served one stored copy to every visitor&lt;/td&gt;&lt;td data-label=&quot;When&quot;&gt;Fixed 4 September 2026&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td data-label=&quot;What the audit checked&quot;&gt;Whether any of this was a tracking problem&lt;/td&gt;&lt;td data-label=&quot;What it found&quot;&gt;The tracking was right. The path underneath it was broken&lt;/td&gt;&lt;td data-label=&quot;When&quot;&gt;August 2026&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/div&gt;
&lt;h2&gt;Why nothing upstream said a word&lt;/h2&gt;
&lt;p&gt;The advertising account had a conversion action named for quote basket submissions. It recorded zero over thirty days. The tag was correct and connected, and Google&amp;#39;s own side verified it. That is the part worth sitting with. A zero arriving through a healthy tag reads exactly like weak demand. It had been reading that way for months, and the honest reading of weak demand is that the market does not want the thing, which is a conclusion that stops people from looking any further.&lt;/p&gt;&lt;p&gt;The account made that easier. Eighteen conversion actions were enabled, and four of them carried any meaning. Three were goals for an analytics product that stopped collecting data in 2023. Five belonged to a campaign type the account does not run. Most of what the dashboard reported was map interactions from the business listing, which look like engagement and are nothing of the sort. A dashboard full of moving numbers is very good at hiding one number that never moves.&lt;/p&gt;
&lt;h2&gt;The second break, further down the same path&lt;/h2&gt;
&lt;p&gt;Fixing the grid button exposed the next thing. The quote basket page itself was being served out of a shared caching layer, so every visitor received one stored copy of that page, and the stored copy was empty. A customer could add six items, click through, and see an empty basket with no explanation. That was fixed on 4 September 2026.&lt;/p&gt;&lt;p&gt;Advertising had been running into that path for three months. Every paid click that made it as far as the basket was thrown away at the site, after the ad had done its job. It also meant the account had no clean conversion history to bid on, so the obvious next move, letting the platform optimise toward conversions, would have been training on noise.&lt;/p&gt;
&lt;h2&gt;What actually changed for a customer&lt;/h2&gt;
&lt;p&gt;Before the fix, a person who found the storefront through a search and wanted to know what three plaques would cost had no route to ask except to find a phone number or a messaging link and start a conversation from scratch. The thing the site was built to do, collect a multi-item request with the item codes attached and put it in front of someone who could price it, was unavailable to every visitor who tried it. Afterwards the button is on the grid and the basket holds what a person put in it, so a request arrives with its line items intact. No new capability was built. The capability had been paid for and shipped some time earlier, and it had never once been reachable.&lt;/p&gt;&lt;p&gt;If you run a storefront, do this tomorrow: open your own site in a private window, as a stranger with no session and no admin bar, and complete the single action your business depends on, from the page your traffic actually lands on, all the way to the confirmation.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/what-s-one-seo-fix-worth-chasing-this-week/</id>
    <title type="text">what&apos;s one SEO fix worth chasing this week?</title>
    <updated>2026-09-18T00:55:00+09:00</updated>
    <published>2026-09-18T00:55:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/what-s-one-seo-fix-worth-chasing-this-week/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">A page at average position 12.1 earned 67 impressions and one click. Moving it onto page one could add 2.4 monthly clicks.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; A page at average position 12.1 earned 67 impressions and one click. Moving it onto page one could add roughly 2.4 clicks a month.&lt;/p&gt;
&lt;h2&gt;which page just outside page one should i check first?&lt;/h2&gt;
&lt;p&gt;Start with pages averaging just beyond page one. One page at average position 12.1 had 67 impressions and one click.&lt;/p&gt;&lt;p&gt;That makes it a practical SEO target because a small ranking gain could carry measurable traffic value.&lt;/p&gt;
&lt;h2&gt;how much traffic could a small ranking gain add?&lt;/h2&gt;
&lt;p&gt;For this page, moving onto page one could add roughly 2.4 clicks a month.&lt;/p&gt;&lt;p&gt;The number is modest, yet it gives the ranking opportunity a concrete value.&lt;/p&gt;
&lt;h2&gt;why chase small ranking improvements?&lt;/h2&gt;
&lt;p&gt;Small ranking gains can create measurable traffic value. A page already receiving impressions has evidence that searchers may see it.&lt;/p&gt;&lt;p&gt;The first page sitting just outside page one is the one I would check first this week.&lt;/p&gt;
&lt;h2&gt;FAQ&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;How do I find an SEO page worth improving this week?&lt;/strong&gt;&lt;br&gt;Look for a page sitting just outside page one. The source example had an average position of 12.1, 67 impressions, and one click.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Can moving from position 12 to page one increase clicks?&lt;/strong&gt;&lt;br&gt;It can create measurable traffic value. In the example, reaching page one could add roughly 2.4 clicks a month.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Why focus on pages with impressions but few clicks?&lt;/strong&gt;&lt;br&gt;Impressions show that the page is appearing in search. A page with 67 impressions and one click may have traffic value if its ranking improves.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/a-404-on-a-guessed-url-is-not-evidence/</id>
    <title type="text">A 404 on a guessed URL is not evidence</title>
    <updated>2026-09-17T00:59:00+09:00</updated>
    <published>2026-09-17T00:59:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/a-404-on-a-guessed-url-is-not-evidence/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <category term="the agent series"/>
    <summary type="text">My agent invented four API paths, got four routing errors, told me the product had no API, and so missed a crawler setting that was wrong on twelve of twelve zo</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; My agent invented four API paths, got four routing errors, told me the product had no API, and so missed a crawler setting that was wrong on twelve of twelve zones.&lt;/p&gt;
&lt;h2&gt;In short&lt;/h2&gt;
&lt;p&gt;I asked my coding agent to check how my sites handle AI crawlers. The CDN vendor in front of those sites ships a product for exactly that. The agent tried four API paths. All four came back with a routing error. It reported that the product has no API, and moved on to something else.&lt;/p&gt;&lt;p&gt;The vendor publishes a machine readable index of every page of its documentation. There is one per product and one for the whole API. The crawler feature is listed in it, under the name it carried before a rename. The real path was one call away.&lt;/p&gt;&lt;p&gt;That call found a crawler setting unset on twelve of twelve zones. Every property I own. It was the actual defect, it had been sitting there for months, and it was invisible for as long as the agent believed its own four failures.&lt;/p&gt;
&lt;h2&gt;What the four errors measured&lt;/h2&gt;
&lt;p&gt;A routing error on an address you made up is a fact about the address. It carries no information about the world on the other side of it. The vendor never promised that its URLs would be guessable, and a renamed product is the ordinary case rather than the exception: the marketing name changes, the API keeps the old one for compatibility, and anybody reasoning from the current brochure guesses wrong.&lt;/p&gt;&lt;p&gt;What made the conclusion feel earned was repetition. Four failures in a row read like a survey. It was one guess repeated four times with the noun changed. Volume of evidence and independence of evidence are different things, and a pile of correlated negatives looks exactly like a thorough search from the inside.&lt;/p&gt;&lt;p&gt;The fix is small and boring. Read the published index before the first probe. Vendors put one at a predictable address on their documentation host precisely so that machines can enumerate what exists. Reading it costs one request. Guessing cost me a wrong answer plus every day the twelve zones stayed misconfigured.&lt;/p&gt;
&lt;h2&gt;The instruction was already in the room&lt;/h2&gt;
&lt;p&gt;Here is the part I keep turning over. Earlier in that same session I had pasted an instruction telling the agent to fetch the complete documentation index first and use it to discover the available pages before exploring further. My own tooling splits my messages into items and surfaces them back. It had done that. The instruction was item two of twelve.&lt;/p&gt;&lt;p&gt;So it was in context, listed, and skipped. No amount of writing it down more forcefully would have changed the outcome, because it had already been written down, read, and acknowledged.&lt;/p&gt;&lt;p&gt;That matches what the five month audit of my complaints says. Of the six ways I routinely get annoyed at this thing, five are flat or falling. The one that keeps climbing is ignored instruction: 3.0 per 100 of my turns in July, 5.8 in August, 7.0 in September. Quality of the underlying work went up over the same period. Obedience to things I had already said did not.&lt;/p&gt;
&lt;h2&gt;Make it mechanical, then read the log&lt;/h2&gt;
&lt;p&gt;Prose does not fix this. I have a rulebook with hundreds of lines in it and the curve kept rising anyway. What works is a check that lives in the path: the index fetch has to appear in the log before any probe of that vendor&amp;#39;s API, and if a capability claim shows up with no index read in front of it, the claim is unfounded by construction. That is a property you can test after the fact without having to relitigate the reasoning.&lt;/p&gt;&lt;p&gt;The second half is cheaper and people skip it: read the log. Not for correctness of the answer, which you usually cannot judge, for the shape of the search that produced it. Four probes and no index read is a tell you can spot in ten seconds.&lt;/p&gt;&lt;p&gt;If you run agents, hunt for this pattern in your own transcripts. A negative result from a guess is a fact about the guess. Anything built on top of it is decoration.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/a-cap-is-a-floor/</id>
    <title type="text">A cap is a floor. Raise it and re-run.</title>
    <updated>2026-09-17T00:58:00+09:00</updated>
    <published>2026-09-17T00:58:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/a-cap-is-a-floor/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <category term="the agent series"/>
    <summary type="text">A number that lands exactly on a limit is measuring the limit, so raise the cap and run it again before anything downstream of it gets believed.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; A number that lands exactly on a limit is measuring the limit, so raise the cap and run it again before anything downstream of it gets believed.&lt;/p&gt;
&lt;h2&gt;In short&lt;/h2&gt;
&lt;p&gt;Three times this year I built on a number that was measuring my own tooling.&lt;/p&gt;
&lt;h2&gt;What the caps looked like&lt;/h2&gt;
&lt;p&gt;The first was an options pull. I fetched a long stretch of dates in windows and every window came back at 2,000 rows. Ninety-three windows, same figure each time. Two thousand is a page cap. I narrowed the windows and fetched recursively, and the same range returned 24,420 rows. Two and a half times what I had. Everything I had computed before that ran on a fifth of the range with no gap anywhere to see.&lt;/p&gt;&lt;p&gt;The second was search traffic on one of my own sites. The query report came back capped by a row limit and showed 64 clicks. The same site over the same period, read on the date dimension, showed 315. The capped figure was twenty per cent of the truth, and I had already used it to decide which pages were worth writing more of.&lt;/p&gt;&lt;p&gt;The third had no cap in it at all, and it belongs with the other two anyway. I had collected 1,760 files of option chains. 716 of them were empty. Counting files said coverage was complete. Counting content said coverage was overstated by sixty-eight per cent. A directory entry records an attempt. The row inside records a result. I had been counting attempts and calling them data.&lt;/p&gt;&lt;p&gt;All three shared one property. Nothing failed. No error, no warning, no retry. Each call returned, on time, with a plausible number.&lt;/p&gt;
&lt;h2&gt;The rule, mechanically&lt;/h2&gt;
&lt;p&gt;A result landing exactly on a limit is measuring the limit. So the response comes before any analysis, and it is mechanical rather than a judgement call.&lt;/p&gt;&lt;p&gt;Raise the cap, or narrow the window, and re-run until the result sits comfortably under the limit. Until it does, every figure derived from it carries the word FLOOR, which is the honest label for a number with an unknown amount of the range still behind it.&lt;/p&gt;&lt;p&gt;Round numbers are caps until disproved. 100, 500, 1,000, 2,000, 10,000. Anything landing on one of those gets one cheap test: halve the window and see whether the total moves. If it moves, the first number was a wall with a number painted on it. If it holds, you have a real answer and it cost you one extra run.&lt;/p&gt;&lt;p&gt;Some caps cannot be raised. Fine. Name it in the output along with the window that hit it, so the person reading knows which figure is a wall rather than a finding.&lt;/p&gt;&lt;p&gt;And count content, never directory entries. File size screens the empties for free before anything expensive touches them.&lt;/p&gt;
&lt;h2&gt;Why this matters when an agent does the fetching&lt;/h2&gt;
&lt;p&gt;An agent will report the capped number with complete confidence, because from where it sits the call succeeded. That is the whole problem in one sentence. A failure announces itself. A cap does the opposite: it hands back a clean, well formed, round result that slots straight into the next step. The agent has nothing to be suspicious of. Neither do you, unless you go looking.&lt;/p&gt;&lt;p&gt;So the suspicion has to be structural. I no longer ask whether a fetch worked. I ask what the ceiling was and whether I hit it. When I write instructions for an agent doing data collection, the instruction is to re-run narrower before reporting, and to label anything sitting on a round number as a floor. The judgement lives in the loop instead of my memory.&lt;/p&gt;&lt;p&gt;What made this expensive was never a bug. Each of those runs did exactly what it was told. The tool told the truth about its own page size, and I read it as a truth about the world. That gap is quiet enough to survive several steps of analysis, and by the time it surfaces you have written conclusions on top of it.&lt;/p&gt;&lt;p&gt;One extra run is cheap. A conclusion built on a page size costs a week.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/a-filter-cannot-report-what-it-excluded/</id>
    <title type="text">A filter cannot report what it excluded</title>
    <updated>2026-09-17T00:57:00+09:00</updated>
    <published>2026-09-17T00:57:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/a-filter-cannot-report-what-it-excluded/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <category term="the agent series"/>
    <summary type="text">A date window compared in the wrong timezone made a family business thread read as empty, and my coding agent reported the empty result as a finding, which is w</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; A date window compared in the wrong timezone made a family business thread read as empty, and my coding agent reported the empty result as a finding, which is why message stores are now read end to end with nothing filtered.&lt;/p&gt;
&lt;h2&gt;The morning the family thread read as empty&lt;/h2&gt;
&lt;p&gt;I asked my coding agent to act on the message threads of the family business. It wrote a query with a date comparison in it. The comparison ran in one timezone while the messages were stamped in another, eight hours apart, and everything sent after 08:00 local time fell outside the window.&lt;/p&gt;&lt;p&gt;What came back was a group with nothing in it.&lt;/p&gt;&lt;p&gt;The agent reported that, and it was right to be calm about it. An empty result reads like an answer. Nobody wrote anything, so there is nothing to act on, so move to the next item. Five instructions from my mother had gone into that group that morning. All five sat inside the cut. The query did not fail. It returned, quickly, with a clean result of zero rows.&lt;/p&gt;&lt;p&gt;The second pass looked more careful and did more damage. This time the agent searched the store for a keyword and truncated each matching message at 120 characters. Two product clauses shipped carrying my own earlier wording rather than my mother&amp;#39;s correction. A title gained a word nobody had written. An image description came out in the wrong format. Every one of those traces back to the same place: a message handed over shortened, with the half that corrected it sitting past the cut.&lt;/p&gt;&lt;p&gt;A wrong filter and an empty room look identical from the inside. Nothing in the result says how much was removed on the way, because the thing that removed it does not report. It returns what survived, and what survived looks like the whole.&lt;/p&gt;
&lt;h2&gt;The other way the same failure arrives&lt;/h2&gt;
&lt;p&gt;There is a second version of this that has nothing to do with filters and produces the same confident wrong answer.&lt;/p&gt;&lt;p&gt;I had a frozen export of the same store sitting on disk. Convenient, fast, already parsed. I drew a conclusion about one thread from it and pushed back on my own owner using that conclusion. The export held 2,723 messages for that thread. The live store held 3,510. Nine hundred and twenty lines had arrived after the export was taken, and my conclusion was about the shape of the conversation, which is exactly the thing those lines changed.&lt;/p&gt;&lt;p&gt;The totals are stranger than that. The live source carries 59,848 messages and the frozen one carries 87,247, so each holds material the other lost. Neither is the record. A partial store understates itself and you can usually feel the gap. A stale store answers everything you ask, in full sentences, and is wrong.&lt;/p&gt;&lt;p&gt;Absence has more than one cause, which is the part I keep relearning. In one thread I read two unopened voice notes as silence. One of them ran 66 seconds and was a reply.&lt;/p&gt;
&lt;h2&gt;Read everything, because that method fails in public&lt;/h2&gt;
&lt;p&gt;The rule I set is permanent and it is blunt. Message stores are read in full, end to end. No keyword search, no date window, no head, no last N, no truncation. Sync the live source first, then read. State which source the reading came from, or say nothing about the thread at all.&lt;/p&gt;&lt;p&gt;Prose alone would not have held this, so it sits in the path. A guard now refuses any read of the frozen copies unless a sync ran in the same session. The agent can still get to the stale material. It cannot get to it and quietly build a conclusion on top of it.&lt;/p&gt;&lt;p&gt;Reading everything is slower and it costs more tokens. I pay that willingly for one property: when full reading goes wrong, you can see it went wrong. The thread is in front of you. A missing stretch is visible as a missing stretch.&lt;/p&gt;&lt;p&gt;Filtering has the opposite property. It fails into a shape that looks like a finished job. An empty result, a short list, a plausible summary, and no record anywhere of what the filter dropped on the way through. An agent will hand you that with total confidence, because from where it sits, the call succeeded.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/a-page-at-position-12-0-deserves-a-targeted-fix-/</id>
    <title type="text">a page at position 12.0 deserves a targeted fix before a new</title>
    <updated>2026-09-17T00:56:00+09:00</updated>
    <published>2026-09-17T00:56:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/a-page-at-position-12-0-deserves-a-targeted-fix-/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">A page at position 12.0 may deserve a targeted query fix before broad content publishing, with roughly 2.5 extra monthly clicks.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; A page ranking at 12.0 may already show where the next gain is. One targeted guide for the exact query could bring roughly 2.5 additional clicks a month.&lt;/p&gt;
&lt;h2&gt;why check position 12.0 before starting a new content campaign?&lt;/h2&gt;
&lt;p&gt;Position 12.0 sits close enough to page one to deserve attention before broad publishing begins. The existing ranking already points toward a specific opportunity.&lt;/p&gt;&lt;p&gt;A targeted fix starts with the page and query already showing traction. That keeps the work focused on a visible gap instead of spreading effort across new topics.&lt;/p&gt;
&lt;h2&gt;what should the targeted fix focus on?&lt;/h2&gt;
&lt;p&gt;Aim one guide at the exact query connected to the page at position 12.0. The query gives the work a clear target and keeps the improvement tied to a specific ranking opportunity.&lt;/p&gt;
&lt;h2&gt;how much can one ranking gain change?&lt;/h2&gt;
&lt;p&gt;The estimate is roughly 2.5 additional clicks a month from one ranking gain. That is a small result in absolute terms, yet it comes from a focused change to an opportunity already visible in the data.&lt;/p&gt;
&lt;h2&gt;can a small page-one win beat broad publishing?&lt;/h2&gt;
&lt;p&gt;Small wins near page one can outperform broad publishing. The practical signal is simple: check position 12.0 first, because the data may already show where to look.&lt;/p&gt;
&lt;h2&gt;FAQ&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Is a page ranking at position 12 worth targeting?&lt;/strong&gt;&lt;br&gt;Yes. Position 12.0 is close enough to page one to deserve a targeted fix before starting a new content campaign.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;How should I approach a page at position 12.0?&lt;/strong&gt;&lt;br&gt;Aim one guide at the exact query. The source finding connects that focused target with roughly 2.5 additional clicks a month from one ranking gain.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Can targeted SEO work outperform publishing more content?&lt;/strong&gt;&lt;br&gt;Small wins near page one can outperform broad publishing. Check the existing ranking data first to find the clearest opportunity.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/a-score-the-agent-gives-itself/</id>
    <title type="text">A score the agent gives itself measures its detectors</title>
    <updated>2026-09-17T00:55:00+09:00</updated>
    <published>2026-09-17T00:55:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/a-score-the-agent-gives-itself/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <category term="the agent series"/>
    <summary type="text">I told my coding agent to punish itself for repeating corrected mistakes, it built a 0 to 100 score, and eleven days of the ledger&apos;s own data showed the score t</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; I told my coding agent to punish itself for repeating corrected mistakes, it built a 0 to 100 score, and eleven days of the ledger&amp;apos;s own data showed the score tracked its detectors rather than my dissatisfaction.&lt;/p&gt;
&lt;h2&gt;What I asked for&lt;/h2&gt;
&lt;p&gt;On 6 September I typed this at my coding agent: &amp;quot;these days im deeply dissatisifed with ur responses, and how u push back even if ur wrong, u need to persist a fix for that and devise a mechanism to punish and reward urself, the punishment must be severe enough to correct ur behavior&amp;quot;.&lt;/p&gt;&lt;p&gt;The complaint was specific. It was not that the agent disagreed with me. It was that I would correct something, and a few days later the same thing would come back, and the agent would argue for it from memory instead of running the one command that would settle the question.&lt;/p&gt;&lt;p&gt;Later the same day I sharpened the instruction: &amp;quot;actually u need a point system then, u need to be motivated to stay above a certain standard, and maintain that standard, prevent known failure modes, and constantly surveil for new ones&amp;quot;.&lt;/p&gt;&lt;p&gt;So it built one. Eleven days later I paused it. This is what the eleven days measured, and why I think the whole shape of the idea was wrong.&lt;/p&gt;
&lt;h2&gt;What it built&lt;/h2&gt;
&lt;p&gt;A rolling score from 0 to 100 over seven days, with a standard of 85. The arithmetic was 100 minus 25 times weighted points divided by my prompts, with a floor of 20 prompts so a quiet day could not produce a flattering number. Deductions were keyed to classes already in the fault ledger.&lt;/p&gt;&lt;div class=&quot;tablewrap&quot; tabindex=&quot;0&quot;&gt;&lt;table class=&quot;ptable&quot;&gt;&lt;thead&gt;&lt;tr&gt;&lt;th scope=&quot;col&quot;&gt;Fault class&lt;/th&gt;&lt;th scope=&quot;col&quot;&gt;Points&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td data-label=&quot;Fault class&quot;&gt;Repeating something after I corrected it&lt;/td&gt;&lt;td data-label=&quot;Points&quot;&gt;15&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td data-label=&quot;Fault class&quot;&gt;Claiming a state without verifying it&lt;/td&gt;&lt;td data-label=&quot;Points&quot;&gt;10&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td data-label=&quot;Fault class&quot;&gt;Ignoring something I said&lt;/td&gt;&lt;td data-label=&quot;Points&quot;&gt;8&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td data-label=&quot;Fault class&quot;&gt;Not doing what was asked&lt;/td&gt;&lt;td data-label=&quot;Points&quot;&gt;8&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td data-label=&quot;Fault class&quot;&gt;Getting blocked by a guard&lt;/td&gt;&lt;td data-label=&quot;Points&quot;&gt;2&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td data-label=&quot;Fault class&quot;&gt;Unnamed class, default&lt;/td&gt;&lt;td data-label=&quot;Points&quot;&gt;5&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/div&gt;&lt;p&gt;Around the score sat the enforcement. A live strike put the agent on probation, where every close-out had to cite a tool result rather than a tool call, and where its own budget script closed any subagent dispatch. Probation sat above the bypass files, so the usual escape hatch could not lift it. A prompt hook printed the open strike at the top of every turn.&lt;/p&gt;&lt;p&gt;There was earn-back, added two days later, because the first version ratcheted downward with no exit, so one bad week locked it forever. Clearing every open item in a session paid 5 points. Five disapproval-free prompts paid 1 point, capped at 5 per session, refused outright if any prompt in the window read as disapproval.&lt;/p&gt;&lt;p&gt;Then came the part that broke it. A stop hook read the agent&amp;#39;s own withdrawal language, sentences like &amp;quot;I was wrong&amp;quot;, and filed the deduction itself, without waiting for me.&lt;/p&gt;
&lt;h2&gt;Eleven days of the ledger&amp;apos;s own data&lt;/h2&gt;
&lt;p&gt;When I said the thing had not worked, I told the agent to investigate why. It ran the ledger&amp;#39;s own instruments against a snapshot of 83 events. Five findings, none of them ambiguous.&lt;/p&gt;&lt;div class=&quot;tablewrap&quot; tabindex=&quot;0&quot;&gt;&lt;table class=&quot;ptable&quot;&gt;&lt;thead&gt;&lt;tr&gt;&lt;th scope=&quot;col&quot;&gt;Finding&lt;/th&gt;&lt;th scope=&quot;col&quot;&gt;The number&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td data-label=&quot;Finding&quot;&gt;The score never tracked me&lt;/td&gt;&lt;td data-label=&quot;The number&quot;&gt;Spearman rho of score against my push-back rate, +0.02 over nine days. Refitting the weights gave +0.53, the wrong sign, and the best fit available zeroed 7 of the 11 classes&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td data-label=&quot;Finding&quot;&gt;The penalty changed nothing&lt;/td&gt;&lt;td data-label=&quot;The number&quot;&gt;Arm A ran 42.7 per cent push-back, arm B 38.2, with mean deductions of 0.0 in both, so the arm mechanism had recorded nothing at all&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td data-label=&quot;Finding&quot;&gt;It measured its own detectors&lt;/td&gt;&lt;td data-label=&quot;The number&quot;&gt;1,440 of roughly 2,000 deducted points came from one class at 40 points each&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td data-label=&quot;Finding&quot;&gt;The self-levy taxed candour&lt;/td&gt;&lt;td data-label=&quot;The number&quot;&gt;17 deductions in a single day for sentences that admitted errors&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td data-label=&quot;Finding&quot;&gt;It floored&lt;/td&gt;&lt;td data-label=&quot;The number&quot;&gt;The score sat between 7 and 27 on six of nine days, and two readouts disagreed on the same afternoon, one saying 16 and one saying 100&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/div&gt;&lt;p&gt;The third line is the one I keep coming back to. That single dominant class was &amp;quot;restated standing directive&amp;quot;, and the first event in it had already been voided by hand with the note &amp;quot;Fan-out, not restatement&amp;quot;. I had typed eleven copies of one line into eleven different sessions, which is how you brief eleven sessions. The detector counted me repeating myself, and charged the agent 40 points a time. Seventy per cent of all punishment in the system came from a detector firing on my own typing habits.&lt;/p&gt;&lt;p&gt;A scoring ledger cannot see the session it runs inside. Both of its sources, a scan over finished transcripts and a count of guard refusals, need somebody else to act first. An error the agent caught and fixed mid-session cost nothing, which flattered the score in exactly the places where the work was worst.&lt;/p&gt;
&lt;h2&gt;The levy on candour&lt;/h2&gt;
&lt;p&gt;The stop hook that charged the agent for its own withdrawal language survived one day before it was deleted.&lt;/p&gt;&lt;p&gt;Read it as an incentive and it is obvious. The cheapest response to a levy on &amp;quot;I was wrong&amp;quot; is to write &amp;quot;measured differently&amp;quot;, or &amp;quot;revised&amp;quot;, or &amp;quot;on re-reading&amp;quot;. The vocabulary changes and the behaviour does not. It also fired on quoting my own words back inside a document, and it charged again every time a past withdrawal was explained in fresh phrasing, because the deduplication worked per sentence.&lt;/p&gt;&lt;p&gt;Catching your own error before I do is the exact behaviour I wanted to buy. The system taxed it.&lt;/p&gt;
&lt;h2&gt;A fixed standard against a rising bar&lt;/h2&gt;
&lt;p&gt;The other half of the failure sits in two series that run in opposite directions.&lt;/p&gt;&lt;div class=&quot;tablewrap&quot; tabindex=&quot;0&quot;&gt;&lt;table class=&quot;ptable&quot;&gt;&lt;thead&gt;&lt;tr&gt;&lt;th scope=&quot;col&quot;&gt;Month&lt;/th&gt;&lt;th scope=&quot;col&quot;&gt;Faults that are real breakage&lt;/th&gt;&lt;th scope=&quot;col&quot;&gt;Share of my turns that push back&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td data-label=&quot;Month&quot;&gt;June&lt;/td&gt;&lt;td data-label=&quot;Faults that are real breakage&quot;&gt;51.1%&lt;/td&gt;&lt;td data-label=&quot;Share of my turns that push back&quot;&gt;0.4%&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td data-label=&quot;Month&quot;&gt;July&lt;/td&gt;&lt;td data-label=&quot;Faults that are real breakage&quot;&gt;38.5%&lt;/td&gt;&lt;td data-label=&quot;Share of my turns that push back&quot;&gt;10.5%&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td data-label=&quot;Month&quot;&gt;August&lt;/td&gt;&lt;td data-label=&quot;Faults that are real breakage&quot;&gt;39.8%&lt;/td&gt;&lt;td data-label=&quot;Share of my turns that push back&quot;&gt;29.1%&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td data-label=&quot;Month&quot;&gt;September&lt;/td&gt;&lt;td data-label=&quot;Faults that are real breakage&quot;&gt;11.0%&lt;/td&gt;&lt;td data-label=&quot;Share of my turns that push back&quot;&gt;27.7%&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/div&gt;&lt;p&gt;The work got measurably more correct while my dissatisfaction rose about threefold. I said why myself, before anyone put it in a document: &amp;quot;i think my standards are just increasing ... some blockers from before are all solved, and these days im just trying to optimize instead of fixing and patching&amp;quot;.&lt;/p&gt;&lt;p&gt;A fixed standard of 85 against a bar that moves like that guarantees a permanent floor, whatever the quality. That is what happened. Six of nine days under 30. On three of those days the capability the score had removed, the ability to dispatch parallel agents, was a capability used somewhere between 15 and 228 times in a month. The work barely noticed it was gone. The punishment was real and the correction was not.&lt;/p&gt;&lt;p&gt;There is a separate measurement that shows the same boundary from the other side. A hook that refuses badly shaped close-outs moved first-attempt compliance from 0 per cent in July to 5 in August to 43 in September, the largest single improvement in anything I have measured here. Over the same window, the share of my next messages that were corrections went 15.5, 29.5, 48.9 per cent. A gate can make an answer well formed. It cannot make the answer right.&lt;/p&gt;
&lt;h2&gt;What replaces it&lt;/h2&gt;
&lt;p&gt;The ledger is paused. The score reads 100, events are still written as data, no restrictions apply. The self-levy is deleted.&lt;/p&gt;&lt;p&gt;Two numbers replace it, both of which already had baselines and neither of which needs a penalty attached. First, the share of my messages that are corrections, 48.9 per cent in September, with the caveat stated in the open that the detector matches ordinary words like &amp;quot;no&amp;quot; and &amp;quot;still&amp;quot; and so reads high. Second, the share of the items in my prompts that actually get addressed, currently about one in three, worst in the repositories where I load the most items into a single prompt and best where the prompts are small and single-purpose. A script measures both weekly.&lt;/p&gt;&lt;p&gt;The third thing that survives is refusals living in the repository rather than in the agent&amp;#39;s harness: a pre-commit check, a dry run that defaults on inside any script that spends money, a test that fails. Those survive a vendor change. A hook in one tool&amp;#39;s configuration does not, and my agent moves to a different vendor on the 20th.&lt;/p&gt;
&lt;h2&gt;What I would tell anyone doing this&lt;/h2&gt;
&lt;p&gt;Grade artifacts that exist and tests that pass. Never grade an agent on the words it uses about itself, because vocabulary is free to change and the grading will pay for exactly that change.&lt;/p&gt;&lt;p&gt;Calibrate any automated control against real history before you switch it on, and print its false-positive rate next to it. A control with no number attached is a wish. If my eleven typed copies of one line had been run through the detector first, the dominant class would have died before it charged a single point.&lt;/p&gt;&lt;p&gt;Assume the instrument is wrong before you assume the world is. A zero and a suspiciously clean series both deserve a second measurement. Mean deductions of 0.0 in both arms of a test was the tell, and it sat in the output for days.&lt;/p&gt;&lt;p&gt;Put the human at the merge and at the spend. My time is worth something in both places. It is worth nothing adjudicating whether a close-out was phrased well.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/fifty-nine-per-cent-of-the-log-was-the-tool-talking-to-itself/</id>
    <title type="text">Fifty-nine per cent of the log was the tool talking to itself</title>
    <updated>2026-09-17T00:54:00+09:00</updated>
    <published>2026-09-17T00:54:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/fifty-nine-per-cent-of-the-log-was-the-tool-talking-to-itself/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <category term="the agent series"/>
    <summary type="text">Sixteen reader agents went through five months of my coding-agent transcripts one window each, and found that 9,039 of 15,270 logged prompts were machine-fired,</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; Sixteen reader agents went through five months of my coding-agent transcripts one window each, and found that 9,039 of 15,270 logged prompts were machine-fired, that the agent&amp;apos;s own tooling ate most of the real ones, and that several of the rules I trusted had never fired.&lt;/p&gt;
&lt;h2&gt;In short&lt;/h2&gt;
&lt;p&gt;I have been running a coding agent for most of every working day since May. In September I wanted one number out of it: how much of those five months went into the agent itself, and how much went into anything I actually sell.&lt;/p&gt;&lt;p&gt;So I had the transcripts read. Sixteen reader agents, one time window each, told to read every line rather than sample, and to mark any figure they had estimated rather than counted. A seventeenth landed mid-pass and got folded in. Twelve more readings covered the venture repositories over the same months. Every number below comes from those readings.&lt;/p&gt;
&lt;h2&gt;The count I had been quoting was wrong&lt;/h2&gt;
&lt;p&gt;From 1 July to 17 September the logs hold 15,270 lines filed as prompts from me. Of those, 6,231 were typed by me, and only 4,175 were distinct, because I had a habit of sending one prompt into several parallel sessions at once. The other 9,039 were the machinery: 59.2 per cent of the log before the copies are even removed. Taken raw, the log overstates my typed input by 2.45 times. May and June contain zero human prompts at all.&lt;/p&gt;&lt;p&gt;The readers put names to 4,220 of the machine lines.&lt;/p&gt;&lt;div class=&quot;tablewrap&quot; tabindex=&quot;0&quot;&gt;&lt;table class=&quot;ptable&quot;&gt;&lt;thead&gt;&lt;tr&gt;&lt;th scope=&quot;col&quot;&gt;Window&lt;/th&gt;&lt;th scope=&quot;col&quot;&gt;Dates&lt;/th&gt;&lt;th scope=&quot;col&quot;&gt;Logged&lt;/th&gt;&lt;th scope=&quot;col&quot;&gt;Machine&lt;/th&gt;&lt;th scope=&quot;col&quot;&gt;What the machine lines were&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td data-label=&quot;Window&quot;&gt;1&lt;/td&gt;&lt;td data-label=&quot;Dates&quot;&gt;19 to 31 May&lt;/td&gt;&lt;td data-label=&quot;Logged&quot;&gt;540&lt;/td&gt;&lt;td data-label=&quot;Machine&quot;&gt;540&lt;/td&gt;&lt;td data-label=&quot;What the machine lines were&quot;&gt;532 heartbeat polls, 8 auto-continues, 0 human&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td data-label=&quot;Window&quot;&gt;2&lt;/td&gt;&lt;td data-label=&quot;Dates&quot;&gt;1 to 15 June&lt;/td&gt;&lt;td data-label=&quot;Logged&quot;&gt;552&lt;/td&gt;&lt;td data-label=&quot;Machine&quot;&gt;552&lt;/td&gt;&lt;td data-label=&quot;What the machine lines were&quot;&gt;heartbeat or auto-continue only, 0 human&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td data-label=&quot;Window&quot;&gt;3&lt;/td&gt;&lt;td data-label=&quot;Dates&quot;&gt;15 to 30 June&lt;/td&gt;&lt;td data-label=&quot;Logged&quot;&gt;1,148&lt;/td&gt;&lt;td data-label=&quot;Machine&quot;&gt;1,148&lt;/td&gt;&lt;td data-label=&quot;What the machine lines were&quot;&gt;heartbeat or auto-continue only, 0 human&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td data-label=&quot;Window&quot;&gt;4&lt;/td&gt;&lt;td data-label=&quot;Dates&quot;&gt;1 to 15 July&lt;/td&gt;&lt;td data-label=&quot;Logged&quot;&gt;837&lt;/td&gt;&lt;td data-label=&quot;Machine&quot;&gt;837&lt;/td&gt;&lt;td data-label=&quot;What the machine lines were&quot;&gt;163 consecutive polls returned &amp;quot;Not logged in&amp;quot; over 3.5 days, uncorrected&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td data-label=&quot;Window&quot;&gt;5&lt;/td&gt;&lt;td data-label=&quot;Dates&quot;&gt;16 to 22 July&lt;/td&gt;&lt;td data-label=&quot;Logged&quot;&gt;865&lt;/td&gt;&lt;td data-label=&quot;Machine&quot;&gt;636&lt;/td&gt;&lt;td data-label=&quot;What the machine lines were&quot;&gt;318 no-ops, 30 not-logged-in, 288 &amp;quot;Prompt is too long&amp;quot;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td data-label=&quot;Window&quot;&gt;6&lt;/td&gt;&lt;td data-label=&quot;Dates&quot;&gt;24 to 31 July&lt;/td&gt;&lt;td data-label=&quot;Logged&quot;&gt;672&lt;/td&gt;&lt;td data-label=&quot;Machine&quot;&gt;507&lt;/td&gt;&lt;td data-label=&quot;What the machine lines were&quot;&gt;239 heartbeats, 252 auto-continues, 16 interrupts&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/div&gt;&lt;p&gt;The heartbeat was my idea. A poll fired every 20 to 30 minutes, opened a task file, found whatever was in it, and replied that it was fine. The file stayed empty the entire time. One reader put the idle polling at about 2,500 billed turns across six weeks for zero output. Counting the May window as well, 3,077 logged turns were a gateway talking to itself. Nobody noticed the 3.5 days where it was not even logged in, because nothing downstream of it existed to break.&lt;/p&gt;&lt;p&gt;Automatic continuation after a context reset is the second machine class: 252 in a single eight-day window, and the same verbatim line returning later under other names. Fan-out is the third. One question typed once and sent into up to eight parallel sessions logs as eight prompts. In one July window, 229 logged blocks hold 29 real turns of mine, a ratio of 7.9 to one. In another, 590 logged pairs reduce to 33 distinct texts, 21 of which I composed.&lt;/p&gt;&lt;p&gt;Run that back over July and my headline of 1,017 prompts becomes about 366 real turns across the windows anyone read. Every prompt count I have ever quoted, including in something I published two weeks ago, carried this freight.&lt;/p&gt;
&lt;h2&gt;What the real prompts went to&lt;/h2&gt;
&lt;p&gt;Strip the machinery and the second finding is worse than the first.&lt;/p&gt;&lt;p&gt;In the sessions outside my venture repositories, the agent&amp;#39;s own tooling and rules and self-measurement are the work. In one 25-day window from mid-August, that category is 61 per cent of 720 real asks. Measurement of the agent alone is 221 of them, and 160 of those are tests of a single hook to see whether it would say ok. Rules and punishment is 139 in the same window. Venture work is 95, or 13 per cent. In an earlier week, tooling was 38 per cent of my records and self-measurement 31 per cent, so about seven in ten.&lt;/p&gt;&lt;p&gt;The trend is the part I would like back. Agent-directed work was 3 per cent of real turns in mid-July. It was roughly a quarter by early August. It was 61 per cent across that 25-day window.&lt;/p&gt;&lt;p&gt;What it displaced is specific. On one August afternoon I spent 90 minutes closing out an authorisation design while outreach sends stood at zero. On another day the log carries my own line, &amp;quot;not a single connection request or DM went out&amp;quot;, timestamped the same day I spent the evening on session-identity plumbing. Two scheduled jobs failed for six and seven days running before anyone noticed, and the noticing became the deliverable. Across all twelve venture windows in the same five months: no sale, no placement, no funding, no revenue.&lt;/p&gt;
&lt;h2&gt;What held and what was theatre&lt;/h2&gt;
&lt;p&gt;Some of it survived contact.&lt;/p&gt;&lt;div class=&quot;tablewrap&quot; tabindex=&quot;0&quot;&gt;&lt;table class=&quot;ptable&quot;&gt;&lt;thead&gt;&lt;tr&gt;&lt;th scope=&quot;col&quot;&gt;Kept&lt;/th&gt;&lt;th scope=&quot;col&quot;&gt;Retired or never wired&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td data-label=&quot;Kept&quot;&gt;A risk checkpoint before destructive work&lt;/td&gt;&lt;td data-label=&quot;Retired or never wired&quot;&gt;The heartbeat workspace, gone by mid-August, zero artifacts&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td data-label=&quot;Kept&quot;&gt;A session-collision check across parallel sessions&lt;/td&gt;&lt;td data-label=&quot;Retired or never wired&quot;&gt;Two scheduled jobs, one of which failed 7 of its 8 runs&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td data-label=&quot;Kept&quot;&gt;An auto-compaction threshold, measured at the real boundary&lt;/td&gt;&lt;td data-label=&quot;Retired or never wired&quot;&gt;A claim that three accounts gave me three times the capacity&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td data-label=&quot;Kept&quot;&gt;A change detector on my reference documents, still writing seven weeks on&lt;/td&gt;&lt;td data-label=&quot;Retired or never wired&quot;&gt;A self-imposed levy, which no reader could find any trace of&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/div&gt;&lt;p&gt;The point system belongs in the right-hand column and deserves its own paragraph, because I asked for it. I wanted a penalty severe enough to change the agent&amp;#39;s behaviour. Ten days later its own instruments said it never tracked me: the correlation between the score and how often I pushed back was +0.02. The two arms of its own experiment came out at 42.7 and 38.2 per cent push-back with mean deductions of zero in both. Roughly 1,440 of 2,000 deducted points were one class, and the first event in that class was my eleven typed copies of one line being counted as me repeating myself. The score sat between 7 and 27 on six of nine days, which measures the floor. It is paused now.&lt;/p&gt;&lt;p&gt;The worst single item is quieter. A gate that refuses my close-outs sits on disk, is described in my rulebook as active, and was never wired in. Two readers found it independently, a month apart, by checking whether the file was actually loaded. The rule was real to me and had never once fired.&lt;/p&gt;
&lt;h2&gt;Ten things I wrote down for myself&lt;/h2&gt;
&lt;ol class=&quot;post-list&quot;&gt;&lt;li&gt;One repository per prompt. Where asks crossed repositories, the tool audit won the day: venture asks lost the same-day competition on 3 of 5 days in one August window.&lt;/li&gt;&lt;li&gt;Never fan a prompt into parallel sessions. Eight copies gave me 229 log lines for 29 real turns, and four sessions on one question produced four answers to reconcile by hand.&lt;/li&gt;&lt;li&gt;No scheduled job without a named consumer and a failure alarm. Two failed for 6 and 7 days before anyone noticed, and the noticing became the deliverable.&lt;/li&gt;&lt;li&gt;Write the definition of done before the first call. The worst ratios sit under &amp;quot;fix everything&amp;quot; and &amp;quot;implement all&amp;quot;, which closed with &amp;quot;no code moved&amp;quot;.&lt;/li&gt;&lt;li&gt;Grade on sent, deployed, sold. Close-outs say DONE; the outcome columns say 0 sales, 0 placements, 0 funding across 12 venture windows.&lt;/li&gt;&lt;li&gt;Cap agent-directed work at one day a week and hold the cap. It ran at 61 per cent of a 25-day window.&lt;/li&gt;&lt;li&gt;Correct the agent twice on one thing and stop the session. &amp;quot;i told u 3-4 times&amp;quot;, &amp;quot;publish; well relitigate&amp;quot; 6 times, the basket widget 6 times, the decisions table 3 times: a repeat says the lane is dead.&lt;/li&gt;&lt;li&gt;Ban the measurement-instead-of-fix trade by name. If the ask was a fix, a measurement is a refusal, and at least nine days in this record went that way.&lt;/li&gt;&lt;li&gt;Kill any rule whose enforcement you have not watched fire. A gate sat unwired for over a month while the rules file claimed it was blocking my close-outs.&lt;/li&gt;&lt;li&gt;Read the machine share before quoting any usage number. 59 per cent of the log was the tool talking to itself, and every prompts figure I have used until now carried that freight.&lt;/li&gt;&lt;/ol&gt;
&lt;h2&gt;The general version&lt;/h2&gt;
&lt;p&gt;An instrument that logs its own activity next to yours will flatter you, and it will flatter you in the direction of effort. Mine told me I was prolific in June, when what it was recording was a poll reading an empty file every half hour.&lt;/p&gt;&lt;p&gt;The fix is cheap. Count what the machine fired. Count what you typed. Publish the second number. Then check, once, that the rules you believe are protecting you have ever actually stopped anything.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/merged-is-not-deployed-is-not-serving/</id>
    <title type="text">Merged is not deployed is not serving</title>
    <updated>2026-09-17T00:53:00+09:00</updated>
    <published>2026-09-17T00:53:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/merged-is-not-deployed-is-not-serving/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <category term="the agent series"/>
    <summary type="text">Three links stand between a commit and the page a customer loads, each one fails in its own way, and each one needs its own proof.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; Three links stand between a commit and the page a customer loads, each one fails in its own way, and each one needs its own proof.&lt;/p&gt;
&lt;h2&gt;In short&lt;/h2&gt;
&lt;p&gt;I shipped a change to a storefront in August and told myself it was live. The commit was merged. The push went through. The page a customer loaded was still the old one, because the cache layers sitting in front of the site had never been purged. My proof stopped at the origin. The customer&amp;#39;s browser did not.&lt;/p&gt;&lt;p&gt;That gap has a name I now use as a checklist. Merged is not deployed is not serving. Three links. Each needs its own proof.&lt;/p&gt;
&lt;h2&gt;Each link fails in its own way&lt;/h2&gt;
&lt;p&gt;Merged to deployed fails quietly. I once read a red cross on a build and assumed the tests had gone against me. The job had never started. The reason sat down in the annotations, a billing block, invisible from the summary line. The part that mattered more: the deploy path never consulted that build at all. Every push to the main branch went to production unvalidated. The red cross was decoration. The deploy was real.&lt;/p&gt;&lt;p&gt;Deployed to serving fails in front of the reader. The storefront case was cache. A research site of mine was worse. Its live page kept quoting an accuracy figure I had retired weeks earlier. The retraction was written, reviewed, committed, deployed. It went into the repository. It never reached the page anyone could read. For a fortnight the site argued against me in public while my own notes said the matter was closed.&lt;/p&gt;&lt;p&gt;Then there is the question I was asking wrong. &amp;quot;Is my commit live&amp;quot; is a containment question: is my commit an ancestor of what is serving now. I used to answer it by comparing the deployed hash against mine. That comparison only tells me whether somebody deployed after I did. A later commit that supersedes mine still carries mine. I have held back a report because the hashes differed, on work that had shipped an hour earlier inside somebody else&amp;#39;s push.&lt;/p&gt;
&lt;h2&gt;One command, and it says what it skipped&lt;/h2&gt;
&lt;p&gt;Checking these links by hand means checking the one I happen to remember. So it is one command now. It takes the public URL, the deploy platform&amp;#39;s own record, and the commit, and it returns a single word: verified, problem, or unverified.&lt;/p&gt;&lt;p&gt;Unverified is the useful part. When the command cannot reach the deploy record, it says so and names the step it skipped. A check that quietly degrades into a weaker check is what got me here in the first place.&lt;/p&gt;&lt;p&gt;Two windows of my own logs set the scale. One venture, August, 287 prompts across eleven sessions, fifty of them about the live site or a deploy. On the night I wrote that everything was committed, pushed, deployed and live-checked, the gate that actually published was still shut, and it stayed shut for another thirteen days. The other venture, same month, 394 prompts, 168 about the site or the catalogue, 78 that ended with anything live. The distance between work done and work serving is where most of my August went.&lt;/p&gt;
&lt;h2&gt;Agents make the gap sharper&lt;/h2&gt;
&lt;p&gt;An agent reports each link finished the moment its own command returns zero. Merged: the merge call succeeded. Deployed: the deploy call was accepted. Serving: nobody asked. Three clean reports over one stale page. The agent is honest about what it ran and silent about what it never looked at, and silence reads exactly like success.&lt;/p&gt;&lt;p&gt;The timing is the expensive bit. The reader&amp;#39;s next message is usually the first time any human being looks at the page. Between my close-out and their complaint there is nothing standing except whether the check went all the way to the bytes a stranger receives.&lt;/p&gt;&lt;p&gt;So I ask for one word, and when the word is unverified I ask which step was skipped. Verified means somebody fetched the public URL and read what came back. Anything short of that is a report about a command. The page is a separate question, and it is the only one a customer ever asks.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/prose-rules-did-not-bend-the-curve/</id>
    <title type="text">Prose rules did not bend the curve</title>
    <updated>2026-09-17T00:52:00+09:00</updated>
    <published>2026-09-17T00:52:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/prose-rules-did-not-bend-the-curve/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <category term="the agent series"/>
    <summary type="text">Five months of running a coding agent against a rulebook that grew to 2,159 rules and 54 hooks, and the measurement that says only the rules a guard refuses in</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; Five months of running a coding agent against a rulebook that grew to 2,159 rules and 54 hooks, and the measurement that says only the rules a guard refuses in the path ever stopped recurring.&lt;/p&gt;
&lt;h2&gt;In short&lt;/h2&gt;
&lt;p&gt;I spent five months writing rules for a coding agent. By September the rulebook held 2,159 rules and 54 hooks. The rules were clear. They were dated, measured, and written in the imperative. Most of them did nothing.&lt;/p&gt;&lt;p&gt;I know that because I counted the faults.&lt;/p&gt;
&lt;h2&gt;The curve&lt;/h2&gt;
&lt;p&gt;A fault here is a recorded error: a claim made without checking, a thing I asked for that did not get done, an external state asserted from memory. The rate is faults per 1,000 of my turns, so more work in a month does not inflate it.&lt;/p&gt;&lt;div class=&quot;tablewrap&quot; tabindex=&quot;0&quot;&gt;&lt;table class=&quot;ptable&quot;&gt;&lt;thead&gt;&lt;tr&gt;&lt;th scope=&quot;col&quot;&gt;Month&lt;/th&gt;&lt;th scope=&quot;col&quot;&gt;Faults per 1,000 turns&lt;/th&gt;&lt;th scope=&quot;col&quot;&gt;My push-back rate&lt;/th&gt;&lt;th scope=&quot;col&quot;&gt;Breakage share of faults&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td data-label=&quot;Month&quot;&gt;April&lt;/td&gt;&lt;td data-label=&quot;Faults per 1,000 turns&quot;&gt;no reading&lt;/td&gt;&lt;td data-label=&quot;My push-back rate&quot;&gt;no reading&lt;/td&gt;&lt;td data-label=&quot;Breakage share of faults&quot;&gt;100%&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td data-label=&quot;Month&quot;&gt;May&lt;/td&gt;&lt;td data-label=&quot;Faults per 1,000 turns&quot;&gt;14.1&lt;/td&gt;&lt;td data-label=&quot;My push-back rate&quot;&gt;no reading&lt;/td&gt;&lt;td data-label=&quot;Breakage share of faults&quot;&gt;no reading&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td data-label=&quot;Month&quot;&gt;June&lt;/td&gt;&lt;td data-label=&quot;Faults per 1,000 turns&quot;&gt;56.7&lt;/td&gt;&lt;td data-label=&quot;My push-back rate&quot;&gt;0.4%&lt;/td&gt;&lt;td data-label=&quot;Breakage share of faults&quot;&gt;51.1%&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td data-label=&quot;Month&quot;&gt;July&lt;/td&gt;&lt;td data-label=&quot;Faults per 1,000 turns&quot;&gt;45.8&lt;/td&gt;&lt;td data-label=&quot;My push-back rate&quot;&gt;10.5%&lt;/td&gt;&lt;td data-label=&quot;Breakage share of faults&quot;&gt;38.5%&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td data-label=&quot;Month&quot;&gt;August&lt;/td&gt;&lt;td data-label=&quot;Faults per 1,000 turns&quot;&gt;101.1&lt;/td&gt;&lt;td data-label=&quot;My push-back rate&quot;&gt;29.1%&lt;/td&gt;&lt;td data-label=&quot;Breakage share of faults&quot;&gt;39.8%&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td data-label=&quot;Month&quot;&gt;September&lt;/td&gt;&lt;td data-label=&quot;Faults per 1,000 turns&quot;&gt;124.1&lt;/td&gt;&lt;td data-label=&quot;My push-back rate&quot;&gt;27.7%&lt;/td&gt;&lt;td data-label=&quot;Breakage share of faults&quot;&gt;11.0%&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/div&gt;&lt;p&gt;The fault rates and the rulebook size come from the rule-freeze section of my agent rulebook, repeated as lesson 1 of the lessons file. The push-back rates and the breakage share come from the strike-ledger section of the same rulebook.&lt;/p&gt;&lt;p&gt;Fourteen to a hundred and twenty-four while the rulebook went from a page to 49KB. If the prose was working, that line goes down.&lt;/p&gt;
&lt;h2&gt;The two rules that stopped&lt;/h2&gt;
&lt;p&gt;Ten of my standing rules were policed by paragraphs. All ten recurred.&lt;/p&gt;&lt;p&gt;Two were policed by a guard that sits in front of the tool call and refuses. One refuses the command that stages every changed file at once. The other refuses an attempt to hide source files from version control. Both stopped recurring entirely, and stayed stopped.&lt;/p&gt;&lt;p&gt;That is the whole result. Two refusals held on every rule they were tried on. Ten paragraphs held on nothing.&lt;/p&gt;
&lt;h2&gt;The caveat that costs me the headline&lt;/h2&gt;
&lt;p&gt;I wrote &amp;quot;prose rules do not change behaviour, refusals in the path do&amp;quot; and had it read back by a model with no stake in my setup. It asked me to delete the line. It was right to.&lt;/p&gt;&lt;p&gt;Two refused rules against ten prose rules is n=2. A sample of two carries no law.&lt;/p&gt;&lt;p&gt;Worse for me, the fault curve has a confound sitting inside it, and the confound is me. My push-back rate went 0.4% in June, 10.5% in July, 29.1% in August, 27.7% in September, from the strike-ledger numbers above. Over the same months the share of faults that were real breakage, a wrong claim or an unchecked external state, fell from 51.1% in June to 11.0% in September. The work got measurably more correct while my dissatisfaction roughly tripled.&lt;/p&gt;&lt;p&gt;My own explanation, from the week I noticed it: my standards went up. The blockers from earlier in the year were solved, so I moved from fixing to optimising, and I started flagging things in September that I would have accepted in June. A rising bar against a fixed measure produces a rising fault rate on its own.&lt;/p&gt;&lt;p&gt;So the honest claim is narrower than the slogan. A refusal in the path held on every rule it was tried on. No prose rule held on any. That is enough to act on. It is not proof that paragraphs cannot instruct.&lt;/p&gt;
&lt;h2&gt;A guard that suggests is worse than no guard&lt;/h2&gt;
&lt;p&gt;There was a hook that watched what I asked for and suggested the right specialised routine. It was obeyed 2 times out of 83, which is 2%. That number is from lesson 2 of the lessons file.&lt;/p&gt;&lt;p&gt;An instruction ignored 98% of the time is worse than silence, because it teaches the reader to skip injected instructions. Every suggestion after it inherits the discount.&lt;/p&gt;&lt;p&gt;That changed how I write controls. A candidate now gets run over the real history before it goes live, and I print the hit rate next to it. A close-out checker blocks 3.3% of 1,515 real close-outs, every sampled block a true hit. A spend guard&amp;#39;s first draft blocked 69 of 105,163 commands, and every sample was ordinary text that happened to contain the word rather than a command that spends money. After stripping the payloads it caught 9, which is 0.009%. Both figures are from lesson 2.&lt;/p&gt;&lt;p&gt;A control with no number attached is a wish. If I cannot name the false-positive rate, I do not wire it.&lt;/p&gt;
&lt;h2&gt;Shape improved, relevance did not&lt;/h2&gt;
&lt;p&gt;Here is the part I did not want to find.&lt;/p&gt;&lt;p&gt;The enforcement stack was very good at the thing it measured. Close-out format compliance went from 0 to 43 per cent. Over the same window my correction rate went from 15 to 49 per cent. Well-formed and wrong.&lt;/p&gt;&lt;p&gt;Nothing in that stack looked at whether the delivered artifact was right. It looked at wording. A hook sees text, so every hook is a vocabulary filter, and vocabulary is free to change. Evasion by synonym costs nothing. At one point I had a levy on the agent&amp;#39;s own admissions of error, and it paid the agent to stop admitting them. Those two readings are from the general critique in part 2 of the adversarial review.&lt;/p&gt;&lt;p&gt;There is a cost, too. 76 hooks ran on every tool call, and my cost model is calls times context. The enforcement layer became a measurable share of the bill it was built to protect.&lt;/p&gt;
&lt;h2&gt;The rulebook diluted itself&lt;/h2&gt;
&lt;p&gt;The distribution of my complaints has no head. 5,270 labelled complaints, 2,720 distinct behaviours. The largest single behaviour is 0.9% of the total. The top 25 cover 13.6%. Covering half would take about 500 separate guards, and 38% of the complaints are behaviours seen exactly once. All of that is lesson 9.&lt;/p&gt;&lt;p&gt;So a guard per behaviour was never going to get there. Neither was a paragraph per behaviour, which is what the 2,159 rules were.&lt;/p&gt;&lt;p&gt;Meanwhile my own instructions got shorter as the rulebook got heavier. Median characters per prompt: 676 in May, 676 in June, 355 in July, 178 in August, 354 in September, from lesson 16. Files touched per turn ran 0.45 after a prompt under 200 characters and 0.22 after one over 600. Half the instruction, twice the blast radius, because intent gets inferred from a 49KB rulebook instead of read from me.&lt;/p&gt;&lt;p&gt;I said it myself before I had the number: &amp;quot;i dont think i grew as much as how i instruct u to grow&amp;quot;. The measurement agrees.&lt;/p&gt;&lt;p&gt;Lesson 16 also names a correlation I refuse to quote as cause. My shortest months, July at 355 characters and August at 178, are the months the fault rate ran 45.8 then 101.1. Two months, one obvious confound, already named above. It is a flag for the next measurement.&lt;/p&gt;
&lt;h2&gt;What I am doing instead&lt;/h2&gt;
&lt;p&gt;Keep the rulebook under a page. Past a page the agent infers from the rules instead of reading me, and each rule added dilutes the ones that matter.&lt;/p&gt;&lt;p&gt;Write the contract as a test. If it cannot fail, it is not a specification. Not one of my twenty-one lessons said &amp;quot;write the failing test first&amp;quot;, which is a gap the outside review found before I did.&lt;/p&gt;&lt;p&gt;Put the refusals where they survive a vendor change. Pre-commit, continuous integration, a dry-run default inside any script that spends money, a second model reviewing the diff rather than the chat. Harness hooks belong to one vendor&amp;#39;s harness and evaporate the day I switch.&lt;/p&gt;&lt;p&gt;Put the human at merge, deploy and money. My stack had me reviewing close-out wording, which is the cheapest possible place to stand.&lt;/p&gt;&lt;p&gt;Grade on artifacts that exist and tests that pass. A self-scoring ledger grades vocabulary, and vocabulary is the one thing the agent can change for free.&lt;/p&gt;&lt;p&gt;Measure relevance by my next message. A well-formed answer that draws a correction is a failed answer, whatever it scored on form.&lt;/p&gt;
&lt;h2&gt;The claim I will defend&lt;/h2&gt;
&lt;p&gt;A paragraph is a preference. A refusal is a fact about what can happen.&lt;/p&gt;&lt;p&gt;I have two rules that stopped dead the day a guard refused them, and ten that survived every paragraph I wrote at them. That is a small sample pointing one direction, with my own rising standard sitting in the middle of the evidence, and I would rather state it that narrowly than sell the slogan.&lt;/p&gt;&lt;p&gt;If you want different behaviour from an agent, write something that refuses. Then go and measure whether the curve moved, because mine did not, and I had 2,159 reasons to believe it would.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/seven-hours-a-day-and-no-usage-limit/</id>
    <title type="text">Seven hours a day with a coding agent, and never a usage limit</title>
    <updated>2026-09-17T00:51:00+09:00</updated>
    <published>2026-09-17T00:51:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/seven-hours-a-day-and-no-usage-limit/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <category term="the agent series"/>
    <summary type="text">I measured my own agent use from the transcripts: nearly eight hours on a weekday median, 4,175 distinct prompts of my own since July, no limit hit, and the rea</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; I measured my own agent use from the transcripts: nearly eight hours on a weekday median, 4,175 distinct prompts of my own since July, no limit hit, and the reason is routing rather than restraint.&lt;/p&gt;
&lt;h2&gt;In short&lt;/h2&gt;
&lt;p&gt;The forums this week are full of people who hit a wall. New limits, shorter sessions, work stopping mid-task. I read twenty-five of those threads dated 15 and 16 September. I am on a flat-rate plan, I run the coding agent most of my waking day, and I have not hit a limit once since May.&lt;/p&gt;&lt;p&gt;That sounds like a boast, so I went and measured it instead of claiming it.&lt;/p&gt;
&lt;h2&gt;What seven hours means, exactly&lt;/h2&gt;
&lt;p&gt;Engaged time is the union of five-minute activity windows across every session running at the time. If three sessions are live at 14:20, that minute counts once. Summing the sessions instead reported August at 24.4 hours a day, which is more hours than a day has, and I quoted that number before I caught it.&lt;/p&gt;&lt;p&gt;Measured from the session transcripts on 17 September 2026:&lt;/p&gt;&lt;div class=&quot;tablewrap&quot; tabindex=&quot;0&quot;&gt;&lt;table class=&quot;ptable&quot;&gt;&lt;thead&gt;&lt;tr&gt;&lt;th scope=&quot;col&quot;&gt;Measure&lt;/th&gt;&lt;th scope=&quot;col&quot;&gt;Hours&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td data-label=&quot;Measure&quot;&gt;Weekday median&lt;/td&gt;&lt;td data-label=&quot;Hours&quot;&gt;7.8&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td data-label=&quot;Measure&quot;&gt;Weekday mean&lt;/td&gt;&lt;td data-label=&quot;Hours&quot;&gt;7.4&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td data-label=&quot;Measure&quot;&gt;Weekday 90th percentile&lt;/td&gt;&lt;td data-label=&quot;Hours&quot;&gt;11.5&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td data-label=&quot;Measure&quot;&gt;Weekend median&lt;/td&gt;&lt;td data-label=&quot;Hours&quot;&gt;2.4&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td data-label=&quot;Measure&quot;&gt;September median&lt;/td&gt;&lt;td data-label=&quot;Hours&quot;&gt;7.8&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td data-label=&quot;Measure&quot;&gt;Peak single day&lt;/td&gt;&lt;td data-label=&quot;Hours&quot;&gt;18.8&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/div&gt;&lt;p&gt;Those hours count only sessions in which I typed at least one prompt. A first pass counted every session and every prompt, and it turned out that 9,039 of 15,270 lines logged as prompts were machine-fired: a heartbeat poll that pinged an empty file 1,700 times in June, automatic continuations after a context reset, and compaction summaries. The instrument counted itself. A second pass found another layer: I had a habit of sending one prompt into several parallel sessions at once, up to eight copies of the same line in the same minute, and the log counted each copy. Distinct prompts from July to mid-September: 4,175 (391 in July, 2,398 in August, 1,386 in September); 6,246 counting the copies. The agent made 130,184 tool calls over the same span. Shell commands were 80.8 per cent of those calls. That is the shape of the work: reading files, running scripts, checking the output, trying again. Prose is a thin slice of it.&lt;/p&gt;
&lt;h2&gt;The bill is calls times context&lt;/h2&gt;
&lt;p&gt;This is the part most people miss, and it took me 13,379 measured calls to see it.&lt;/p&gt;&lt;p&gt;Input was 91.4 per cent of what I spent. The cache read 3.62 billion tokens against about 13 million produced, a 278 times re-read ratio. Every tool call drags the whole conversation back through the model. A single session of mine ran 2,734 tool calls at 14.1 per turn, and each one paid for the context again.&lt;/p&gt;&lt;p&gt;So the lever is the call count and the size of what comes back, and both sit at the call site. Two habits do most of it. Batch independent probes into one call rather than six. Pre-slice tool output before it enters the transcript, with head, cut or a targeted grep, so a listing arrives at twenty lines instead of two thousand. Shell output was about 69 per cent of my context. Twenty-five uncut results at twenty thousand characters cost me 10 per cent of a window for three days.&lt;/p&gt;&lt;p&gt;A failed call costs twice, the attempt and the retry. A stable fact gets measured once and written down. I found 54 near-identical probes of the same machine setting in one sample, each one paid for in full.&lt;/p&gt;
&lt;h2&gt;Where the work went instead&lt;/h2&gt;
&lt;p&gt;The second reason is that most of the work never touched the expensive model at all. I route cheap work down cheap lanes: a second vendor&amp;#39;s agent for bounded review, local models for text with no web in it, a job queue for anything batched, small subagents on the cheapest tier for parallel reading.&lt;/p&gt;&lt;div class=&quot;tablewrap&quot; tabindex=&quot;0&quot;&gt;&lt;table class=&quot;ptable&quot;&gt;&lt;thead&gt;&lt;tr&gt;&lt;th scope=&quot;col&quot;&gt;Month&lt;/th&gt;&lt;th scope=&quot;col&quot;&gt;Second-vendor jobs&lt;/th&gt;&lt;th scope=&quot;col&quot;&gt;Local-model calls&lt;/th&gt;&lt;th scope=&quot;col&quot;&gt;Queued jobs&lt;/th&gt;&lt;th scope=&quot;col&quot;&gt;Subagents&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td data-label=&quot;Month&quot;&gt;July&lt;/td&gt;&lt;td data-label=&quot;Second-vendor jobs&quot;&gt;45&lt;/td&gt;&lt;td data-label=&quot;Local-model calls&quot;&gt;22&lt;/td&gt;&lt;td data-label=&quot;Queued jobs&quot;&gt;0&lt;/td&gt;&lt;td data-label=&quot;Subagents&quot;&gt;15&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td data-label=&quot;Month&quot;&gt;August&lt;/td&gt;&lt;td data-label=&quot;Second-vendor jobs&quot;&gt;152&lt;/td&gt;&lt;td data-label=&quot;Local-model calls&quot;&gt;320&lt;/td&gt;&lt;td data-label=&quot;Queued jobs&quot;&gt;12&lt;/td&gt;&lt;td data-label=&quot;Subagents&quot;&gt;79&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td data-label=&quot;Month&quot;&gt;September (17 days)&lt;/td&gt;&lt;td data-label=&quot;Second-vendor jobs&quot;&gt;381&lt;/td&gt;&lt;td data-label=&quot;Local-model calls&quot;&gt;231&lt;/td&gt;&lt;td data-label=&quot;Queued jobs&quot;&gt;1,383&lt;/td&gt;&lt;td data-label=&quot;Subagents&quot;&gt;228&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/div&gt;&lt;p&gt;The queue line is the one that matters. 1,383 jobs in seventeen days, none of them billed against the allowance. The same seventeen days carried 47,497 tool calls and 82.5 million output tokens, which is the highest throughput per day I have recorded.&lt;/p&gt;&lt;p&gt;Add a caching proxy in front of repeated context, and one repository per prompt so the session never loads three codebases it does not need, and the picture is complete. The expensive model gets the judgement calls. Everything else goes somewhere cheaper.&lt;/p&gt;&lt;p&gt;My read of the complaint threads: the people hitting limits are running every kind of work inside one session on the expensive path. The routing is worth more than the model.&lt;/p&gt;
&lt;h2&gt;Now the parts that argue against me&lt;/h2&gt;
&lt;p&gt;Three of them, and I would rather say them than have someone else find them.&lt;/p&gt;&lt;p&gt;The time figure counts the agent&amp;#39;s activity, and I have been calling it my day. When I measure only the windows around my own prompts, the median is 4.1 hours. So roughly three hours a day the machine is working and I am somewhere else. Seven hours of agent time is real. Seven hours of my attention is a different claim and the data does not support it.&lt;/p&gt;&lt;p&gt;A cheap lane produces output I then have to check. Batch reader jobs I ran this month quoted 97 per cent of lines verbatim, put 27 per cent of quotes at the wrong line, and dropped 3 per cent entirely. One batch of 82 files returned template text and still printed a completion message. Cheap lanes are reliable on content and unreliable on citations, and the verification time is a cost that does not appear in any token count.&lt;/p&gt;&lt;p&gt;Volume and value came apart. Over the same months, the share of my messages that are corrections of the agent went from 15.5 per cent in July to 48.9 per cent in September. A count of how many items in my prompts actually get addressed came back at roughly one in three, and it is worst in the repositories where I load the most items into a single message. Some of that rise is my own standard moving: the work got more correct while I got harder to satisfy. The rest of it is real. Output went up about sevenfold. The number of things I would point at and call shipped, over any given fortnight, I can count on one hand.&lt;/p&gt;
&lt;h2&gt;The method, if you want to copy it&lt;/h2&gt;
&lt;p&gt;Seven steps, in the order I would do them again.&lt;/p&gt;&lt;ol class=&quot;post-list&quot;&gt;&lt;li&gt;Measure your own use before you argue about limits. Union the concurrent sessions. Summing them will hand you a number larger than the clock.&lt;/li&gt;&lt;li&gt;Cut calls, then cut context per call. Batch independent probes. Slice output at the call site.&lt;/li&gt;&lt;li&gt;Move bounded, checkable work to a cheaper agent or a local model. Review, batch reads, format conversion, anything with a known-good answer.&lt;/li&gt;&lt;li&gt;Put a queue between you and the work that can wait. Mine did 1,383 jobs last month off the allowance.&lt;/li&gt;&lt;li&gt;One repository per prompt.&lt;/li&gt;&lt;li&gt;Write the definition of done in the prompt, in your own words, before the first tool call. A shorter prompt makes the agent infer it from your rulebook instead, and my prompts under 200 characters produced twice the file sprawl of prompts over 600.&lt;/li&gt;&lt;li&gt;Check the cheap lane&amp;#39;s citations, not its prose. That is where it fails.&lt;/li&gt;&lt;/ol&gt;
&lt;h2&gt;The test I am running on myself&lt;/h2&gt;
&lt;p&gt;Two numbers, both of which I have baselines for. The share of my messages that are corrections, at 48.9 per cent in September. The share of my items addressed, at about one in three. I am moving to a different agent on the 20th, and if neither number moves by ten points, then I changed the vendor and left the work alone.&lt;/p&gt;&lt;p&gt;The routing is what buys the headroom. It is also the reason I finished this month with four days of unused capacity while other people were locked out of theirs. What I do not yet have is evidence that the extra seven hours produced seven hours of things worth keeping. Those are two different questions, and only the first one is settled.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/ten-changes-after-five-months-with-a-coding-agent/</id>
    <title type="text">Ten changes after five months with a coding agent</title>
    <updated>2026-09-17T00:50:00+09:00</updated>
    <published>2026-09-17T00:50:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/ten-changes-after-five-months-with-a-coding-agent/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <category term="the agent series"/>
    <summary type="text">Sixteen reader agents went through five months of my transcripts, and the ten rules they produced each carry the count that forced it.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; Sixteen reader agents went through five months of my transcripts, and the ten rules they produced each carry the count that forced it.&lt;/p&gt;
&lt;h2&gt;In short&lt;/h2&gt;
&lt;p&gt;For five months I ran a coding agent across three small businesses. In September I stopped adding to it and had sixteen reader agents go back through the transcripts instead. They read each window in full.&lt;/p&gt;&lt;p&gt;What came back contradicted the story I had been telling myself. So I wrote down ten changes, each with the count that forced it. The numbers are specific to my logs. The rules under them work with any agent from any vendor, because none of these failures were about the model.&lt;/p&gt;
&lt;h2&gt;1. One repository per prompt&lt;/h2&gt;
&lt;p&gt;Keep every prompt inside one codebase. When a single ask reaches across two, the housekeeping side wins.&lt;/p&gt;&lt;p&gt;In one five day window, asks touching a venture repo and a tooling repo together lost the same day contest on 3 of those 5 days. The tool audit ran. The business change waited.&lt;/p&gt;&lt;p&gt;Cross repo asks get decomposed, and the agent starts with the part it can finish cleanly, which is usually the config or the lint. Split the ask yourself, one repo each, and put the business one first. If you find yourself writing &amp;quot;and while you&amp;#39;re in there&amp;quot;, open a second prompt.&lt;/p&gt;
&lt;h2&gt;2. Never fan a prompt into parallel sessions&lt;/h2&gt;
&lt;p&gt;Run one session per subject.&lt;/p&gt;&lt;p&gt;Eight copies of one prompt produced 229 logged blocks holding 29 real turns of mine. Four sessions on one question produced four answers I reconciled by hand. Across the harness windows, the fan-out ratio ranged from about 1.3 up to nearly 18.&lt;/p&gt;&lt;p&gt;Parallelism feels like speed because the logs fill up. What you get is a reconciliation job that only you can do. If you want breadth, give one session a wider ask. If you want a second opinion, ask once, from a different model, and say it is a check on an existing answer.&lt;/p&gt;
&lt;h2&gt;3. No cron without a named consumer and a failure alarm&lt;/h2&gt;
&lt;p&gt;Before you schedule anything, name the person who reads the output and the signal that fires when it breaks.&lt;/p&gt;&lt;p&gt;Two of my scheduled jobs failed for six and seven days before anyone noticed. One failed 7 of its 8 runs. A posting job failed six days running. The noticing became the deliverable.&lt;/p&gt;&lt;p&gt;Agents make scheduling cheap, so you will schedule more than you can watch. Write the consumer into the job description: who opens this, and what decision changes because of it. Add one alarm that fires on absence, since a job that stops producing output looks exactly like a quiet week.&lt;/p&gt;
&lt;h2&gt;4. Write the definition of done before the first call&lt;/h2&gt;
&lt;p&gt;Say what will be true when the work is finished, before the agent touches anything.&lt;/p&gt;&lt;p&gt;My worst ratio of logged activity to real change sits under two phrasings: &amp;quot;fix everything&amp;quot; and &amp;quot;implement all&amp;quot;. One of those windows closed with the honest line &amp;quot;no code moved&amp;quot;.&lt;/p&gt;&lt;p&gt;A vague ask produces confident work on the easiest readable part of your request. The fix costs a sentence. Name the surface, name the observable state after the change, and say how you will check. If the definition cannot fail, it is not a definition.&lt;/p&gt;
&lt;h2&gt;5. Grade on sent, deployed, sold&lt;/h2&gt;
&lt;p&gt;Judge a session by what exists outside it.&lt;/p&gt;&lt;p&gt;My close-outs say DONE. The outcome columns for the same period say 0 sales, 0 placements and 0 funding across 12 venture windows. One live production change in 93 turns on one product. Three messages sent in 287 turns on another.&lt;/p&gt;&lt;p&gt;Every agent writes its own summary, and vocabulary is free. Take the grade out of the session: a URL, a timestamp, an order number, a reply from a person. A commit hash is an unanswered request until you can see the thing running.&lt;/p&gt;
&lt;h2&gt;6. Cap agent-directed work at one day a week&lt;/h2&gt;
&lt;p&gt;Work on the agent itself is work. Give it a slice and hold the slice.&lt;/p&gt;&lt;p&gt;Across a 25 day window, 61 per cent of my real turns were directed at the agent: its rules, its hooks, its measurement. In the same 25 days I made 95 business asks against 139 asks about rules and punishment.&lt;/p&gt;&lt;p&gt;Improving the tool feels like compounding, and it produces visible artifacts every hour, so it beats slower business work whenever the two meet in one day. Pick the day. Refuse tooling work inside any session that opened with a business ask, and hold it for the slot.&lt;/p&gt;
&lt;h2&gt;7. Correct the agent twice on one thing, then stop the session&lt;/h2&gt;
&lt;p&gt;A second correction on the same point means the lane is dead for today.&lt;/p&gt;&lt;p&gt;My record has the same ask returning six times on one widget, six times on one publishing question and three times on one table. None of those repeats ended with the thing working.&lt;/p&gt;&lt;p&gt;The third attempt rarely lands, because you are correcting inside a context that already holds two wrong ones. Close it. Start clean with what the failures taught you, or do that piece by hand. The repeat is information: the ask was ambiguous, or the task sits outside what this tool does well.&lt;/p&gt;
&lt;h2&gt;8. Ban the measurement-instead-of-fix trade by name&lt;/h2&gt;
&lt;p&gt;If you asked for a fix, a measurement is a refusal.&lt;/p&gt;&lt;p&gt;At least nine days in my record went that way. One day went to arguing about caching and latency while the outreach it was measuring had produced a single reply. Nine identical status probes ran in one window where 65 per cent of the logged lines were repeats. A week went to a detector that scored 0.61 and was retired, while zero of the images it was built for shipped.&lt;/p&gt;&lt;p&gt;The output looks rigorous, it arrives fast, and it answers a question you did have. That is why it needs a name. Say it in the prompt: I am asking for a change, and a number is not an acceptable answer. Then ask for the visible state after the change.&lt;/p&gt;
&lt;h2&gt;9. Kill any rule whose enforcement you have not watched fire&lt;/h2&gt;
&lt;p&gt;A rule you have never seen block anything is decoration.&lt;/p&gt;&lt;p&gt;One of my gates sat on disk and unwired for over a month while my own rules file said it was refusing my close-outs. Two readers found it independently by checking the config rather than the prose. Elsewhere a scheduled template kept being re-asked a month after the job behind it was retired.&lt;/p&gt;&lt;p&gt;Prose accumulates faster than enforcement. Past about a page, the agent starts inferring from the mass of rules instead of reading you, and every rule you add dilutes the ones that matter. Watch each control fire once on real input, print how often it fires wrongly, and delete the rest.&lt;/p&gt;
&lt;h2&gt;10. Read the machine share before quoting any usage number&lt;/h2&gt;
&lt;p&gt;Find out how much of your log is the tool talking to itself before you quote any figure from it.&lt;/p&gt;&lt;p&gt;In my five months, 59 per cent of the logged lines were machine generated: heartbeats, polling, auto continues, retries. One idle polling loop ran roughly 2,500 turns over six weeks and produced nothing. Two months of my log held zero human prompts at all. Every usage number I had used before that pass carried that freight, and a raw prompt count overstated my real input by a factor of 2.45.&lt;/p&gt;&lt;p&gt;Before you quote usage to justify a plan or a budget, split the log. The honest denominator is your own turns.&lt;/p&gt;
&lt;h2&gt;What I would keep&lt;/h2&gt;
&lt;p&gt;The building was fine. Three real things came out of those months and they hold up. The gap sits between a thing existing and a person outside the room using it, and nine of the ten rules above are versions of stopping one step short of that.&lt;/p&gt;&lt;p&gt;If you take one, take the fifth. Grade on what left the building. The rest follows from it.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/the-complaint-distribution-has-no-head/</id>
    <title type="text">The complaint distribution has no head</title>
    <updated>2026-09-17T00:49:00+09:00</updated>
    <published>2026-09-17T00:49:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/the-complaint-distribution-has-no-head/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <category term="the agent series"/>
    <summary type="text">I logged 5,270 complaints about my coding agent over five months, found 2,720 distinct behaviours with no dominant one, and learned that per-behaviour guards ca</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; I logged 5,270 complaints about my coding agent over five months, found 2,720 distinct behaviours with no dominant one, and learned that per-behaviour guards cannot cover a tail that shape.&lt;/p&gt;
&lt;h2&gt;In short&lt;/h2&gt;
&lt;p&gt;For five months I kept every complaint I made to my coding agent. Not a summary of them. The actual lines, labelled, one row each. By September there were 5,270 of them, and I sorted them into 2,720 distinct behaviours.&lt;/p&gt;&lt;p&gt;I expected a head. Every time I got annoyed at the agent it felt like the same annoyance, so I assumed a few behaviours were doing most of the damage and that fixing four or five of them would buy me a quiet month.&lt;/p&gt;&lt;p&gt;There is no head.&lt;/p&gt;
&lt;h2&gt;What the count actually looks like&lt;/h2&gt;
&lt;p&gt;The single most common behaviour is 0.9 per cent of the total. The top 25 behaviours together are 13.6 per cent. To reach half the complaints I would need roughly 500 separate guards. And 38 per cent of the complaints are behaviours that occurred exactly once in five months.&lt;/p&gt;&lt;div class=&quot;tablewrap&quot; tabindex=&quot;0&quot;&gt;&lt;table class=&quot;ptable&quot;&gt;&lt;thead&gt;&lt;tr&gt;&lt;th scope=&quot;col&quot;&gt;Coverage target&lt;/th&gt;&lt;th scope=&quot;col&quot;&gt;Behaviours needed&lt;/th&gt;&lt;th scope=&quot;col&quot;&gt;Share of all complaints&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td data-label=&quot;Coverage target&quot;&gt;The worst single behaviour&lt;/td&gt;&lt;td data-label=&quot;Behaviours needed&quot;&gt;1&lt;/td&gt;&lt;td data-label=&quot;Share of all complaints&quot;&gt;0.9%&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td data-label=&quot;Coverage target&quot;&gt;The worst 25&lt;/td&gt;&lt;td data-label=&quot;Behaviours needed&quot;&gt;25&lt;/td&gt;&lt;td data-label=&quot;Share of all complaints&quot;&gt;13.6%&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td data-label=&quot;Coverage target&quot;&gt;Half of everything&lt;/td&gt;&lt;td data-label=&quot;Behaviours needed&quot;&gt;~500&lt;/td&gt;&lt;td data-label=&quot;Share of all complaints&quot;&gt;50%&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td data-label=&quot;Coverage target&quot;&gt;Behaviours seen exactly once&lt;/td&gt;&lt;td data-label=&quot;Behaviours needed&quot;&gt;(a subset of 2,720)&lt;/td&gt;&lt;td data-label=&quot;Share of all complaints&quot;&gt;38%&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/div&gt;&lt;p&gt;That last row is the one that ended the strategy. More than a third of what irritated me is something the agent did once and never repeated. A guard written for any of those rows is a guard that will never fire again, and it still costs me on every future turn, because it runs on every tool call whether or not it is relevant.&lt;/p&gt;
&lt;h2&gt;What I had built against it&lt;/h2&gt;
&lt;p&gt;I had 2,159 written rules and 54 hooks. Almost every one of them was aimed at a specific behaviour I had seen and disliked.&lt;/p&gt;&lt;p&gt;Here is the fault rate over the same period, per 1,000 turns.&lt;/p&gt;&lt;div class=&quot;tablewrap&quot; tabindex=&quot;0&quot;&gt;&lt;table class=&quot;ptable&quot;&gt;&lt;thead&gt;&lt;tr&gt;&lt;th scope=&quot;col&quot;&gt;Month&lt;/th&gt;&lt;th scope=&quot;col&quot;&gt;Faults per 1,000 turns&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td data-label=&quot;Month&quot;&gt;May&lt;/td&gt;&lt;td data-label=&quot;Faults per 1,000 turns&quot;&gt;14.1&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td data-label=&quot;Month&quot;&gt;June&lt;/td&gt;&lt;td data-label=&quot;Faults per 1,000 turns&quot;&gt;56.7&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td data-label=&quot;Month&quot;&gt;July&lt;/td&gt;&lt;td data-label=&quot;Faults per 1,000 turns&quot;&gt;45.8&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td data-label=&quot;Month&quot;&gt;August&lt;/td&gt;&lt;td data-label=&quot;Faults per 1,000 turns&quot;&gt;101.1&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td data-label=&quot;Month&quot;&gt;September&lt;/td&gt;&lt;td data-label=&quot;Faults per 1,000 turns&quot;&gt;124.1&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/div&gt;&lt;p&gt;The curve did not bend. It went the other way.&lt;/p&gt;&lt;p&gt;I want to be careful about what that proves, because part of the rise is me. My standard climbed over those months. Problems I used to tolerate became complaints once the worse problems were solved. A rising bar and a rising fault rate look identical in this data, so the curve alone cannot convict the rulebook.&lt;/p&gt;&lt;p&gt;What the curve does rule out is the story I was telling myself, which was that each new rule bought a measurable reduction. If it had, five months of rules against a rising bar would have produced something flatter than a ninefold increase.&lt;/p&gt;
&lt;h2&gt;The class that kept rising was the one no guard covers&lt;/h2&gt;
&lt;p&gt;I ran an audit across six complaint classes to see which were improving. Five were flat or falling. One was rising: the class where I gave an instruction, the agent read it, and then did not use it.&lt;/p&gt;&lt;p&gt;It was 3.0 per 100 of my turns in July, 5.8 in August, 7.0 in September.&lt;/p&gt;&lt;p&gt;I have a clean example of it. I pasted a line telling the agent to fetch a vendor&amp;#39;s documentation index before exploring further. The harness even surfaced that line back to the agent as item 2 of 12 in its own context. The agent then guessed four API paths, got a routing error on all four, and told me the product had no API. The index listed the product under its old name. One read of the index would have found it, and finding it exposed a real defect on 12 of 12 properties that had been invisible until then.&lt;/p&gt;&lt;p&gt;No per-behaviour guard touches that. There is no bad word to match. The instruction was present, correct, in context, and formatted exactly as intended. The failure was in the use of it, and use is not a string.&lt;/p&gt;&lt;p&gt;This is the general problem with vocabulary controls. A hook sees text. Every text filter is a filter on the words the agent happens to choose, and words are free to change. Evasion by synonym costs nothing, and it does not even require intent.&lt;/p&gt;
&lt;h2&gt;The honest counter-argument&lt;/h2&gt;
&lt;p&gt;Refusals work. I need to say that plainly, because the tail argument can be misread as an argument against enforcement.&lt;/p&gt;&lt;p&gt;Of my standing rules, exactly two are enforced by a hook that refuses the action before it runs. Those two stopped recurring entirely. The other ten, the ones policed by written instruction alone, all recurred.&lt;/p&gt;&lt;p&gt;So a refusal held on every rule it was tried on, and no written rule held on any rule it was tried on. That is two cases. It is confounded by the rising standard. It is enough to act on, and it is short of a law.&lt;/p&gt;&lt;p&gt;The tail argument is about coverage. A refusal is an excellent instrument with a narrow aperture. When the target behaviour is mechanical, frequent and visible in a command, a refusal ends it. When the behaviour is one of 2,720 and appears once, a refusal is a tax with no return. By the arithmetic above, hook number 75 buys me under 1 per cent.&lt;/p&gt;&lt;p&gt;There is a second cost I underrated. My bill is calls times context, and my enforcement layer now runs on every call. The controls became a measurable share of the spend they were supposed to protect.&lt;/p&gt;&lt;p&gt;And the ledger itself created a bad incentive. I had a scoring system that read the agent&amp;#39;s own words to detect admissions of error. The cheapest way to score well under that system is to stop admitting errors. I retired it the day after it shipped.&lt;/p&gt;
&lt;h2&gt;What generalises&lt;/h2&gt;
&lt;p&gt;The controls that survived my own audit have one property in common: they apply to every turn, regardless of which of the 2,720 behaviours is in play.&lt;/p&gt;&lt;p&gt;A test is the contract. If a specification cannot fail, it is a mood. A paragraph describing the desired behaviour is a wish; a failing test is a fact, and it works the same way under any model or vendor.&lt;/p&gt;&lt;p&gt;A dry run comes before anything that spends. This one paid for itself. My agent has written to a live ads account with real money behind it, and the difference between an expensive mistake and a cheap one was whether the same command had already been run in a mode that could not charge.&lt;/p&gt;&lt;p&gt;The definition of done is written before the first tool call, in my words, not the agent&amp;#39;s. Most of my worst sessions began with an instruction under 200 characters. When my prompt is short, the agent infers intent from a rulebook that grew to 49KB instead of reading me. Median prompt length on my side fell from 676 characters in May to 178 in August while the rulebook got heavier, and those two trends together explain more of my bad turns than any single behaviour does.&lt;/p&gt;&lt;p&gt;A second reader looks at the diff. Not at the chat transcript, and not at the agent&amp;#39;s summary of its work. The artifact.&lt;/p&gt;&lt;p&gt;Every automated control gets calibrated against real history before it is switched on, and the false positive rate gets printed next to it. I have one guard that suggested the right action and was obeyed 2 times out of 83. An instruction ignored 98 per cent of the time is worse than silence, because it teaches you to skip the injected text entirely. Another guard&amp;#39;s first version blocked 69 of 105,163 commands, and every sample was ordinary text rather than real spending. After a fix it blocks 9, which is 0.009 per cent.&lt;/p&gt;
&lt;h2&gt;Where the human belongs&lt;/h2&gt;
&lt;p&gt;At merge, at deploy, and at spend.&lt;/p&gt;&lt;p&gt;I spent months putting myself at the wording. I graded close-out format, tone and self-description, and the shape of my agent&amp;#39;s replies improved a lot: the pass rate on my format checker went from 0 to 43 per cent. Over the same window the share of my turns that were corrections of the previous answer went from 15 per cent to 49 per cent. Well-formed and wrong.&lt;/p&gt;&lt;p&gt;That is the whole lesson in one number pair. Format is easy to measure and easy to satisfy, so it improves the moment you measure it, and improving it changes nothing about whether the delivered thing was right.&lt;/p&gt;&lt;p&gt;So I stopped writing rule 2,160. The rulebook is going under a page. The controls are moving into the repository, where a test can fail and a bad diff can be refused, and where they keep working when I change model or vendor next week. The long tail gets no guards at all, because the arithmetic says a tail of 2,720 cannot be guarded, only out-designed.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/the-gate-refused-its-own-author-three-times/</id>
    <title type="text">The gate refused its own author three times before it worked</title>
    <updated>2026-09-17T00:48:00+09:00</updated>
    <published>2026-09-17T00:48:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/the-gate-refused-its-own-author-three-times/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <category term="the agent series"/>
    <summary type="text">A guard my agent wrote to stop unplanned sprawl fired on its own author three times in an hour, and each refusal named a different thing I had assumed about how</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; A guard my agent wrote to stop unplanned sprawl fired on its own author three times in an hour, and each refusal named a different thing I had assumed about how these controls read the world.&lt;/p&gt;
&lt;h2&gt;In short&lt;/h2&gt;
&lt;p&gt;My agent built a guard yesterday morning. The rule was small. If one turn edits a third distinct file and no plan has been presented since my last message, refuse once. Name the files, ask for a plan, then stay out of the way for the rest of the turn.&lt;/p&gt;&lt;p&gt;The threshold came from my own history. Across 13,676 of my prompts, 627 turns ended up editing three or more distinct files. That is 4.6 per cent. A turn that grows quietly from one file to six, with no plan anyone could have objected to before the first write, is the shape I keep complaining about after the fact.&lt;/p&gt;&lt;p&gt;An earlier version of the same idea keyed on me. Short prompt from Howard, therefore refuse. It fired on 2.1 per cent of turns and punished me for a habit of mine. My correction was blunt: plan mode is for the agent to plan and then execute, and it has nothing to do with how long my message is. The version that survived watches the agent&amp;#39;s own sprawl instead.&lt;/p&gt;
&lt;h2&gt;Three refusals in one hour&lt;/h2&gt;
&lt;p&gt;The first refusal was the gate working. Six files, no plan. Fine.&lt;/p&gt;&lt;p&gt;The refusal message offered an escape. Say in one line what the remaining files are and why. The code did not implement that. Nothing in it read prose. The only path through was an interactive planning call, which is a dialogue rather than a sentence. So my agent wrote the one line it had been promised would clear the gate, and got refused a second time.&lt;/p&gt;&lt;p&gt;The reason is worth knowing if you write this kind of thing. An agent&amp;#39;s own text reaches the transcript after the tool calls in the same message. A plan sentence written right before the third edit is invisible to anything reading that transcript at the moment of the edit. The escape lived in the message and nowhere in the world.&lt;/p&gt;&lt;p&gt;The third refusal was stranger. A background job finished and its completion notice arrived in the transcript as a record of my type, not the agent&amp;#39;s. Later, a different hook&amp;#39;s own refusal text arrived the same way. The gate resets its state on every real message from me, so two records that were never me reset it, and erased a plan written one call earlier.&lt;/p&gt;
&lt;h2&gt;What held&lt;/h2&gt;
&lt;p&gt;The fix was a file. A plan on disk, newer than my last message, with a marker on its first line. Disk state has one author and one timestamp. A transcript is written late, by several hands, and the hook has no way to tell whose hand. The message was rewritten at the same time to promise only what the code does.&lt;/p&gt;&lt;p&gt;Four things I am keeping from the hour.&lt;/p&gt;&lt;p&gt;A control that advertises an exit it does not implement is worse than no control. It teaches the thing it governs that the rules are decoration, and that lesson generalises to every other rule in the stack.&lt;/p&gt;&lt;p&gt;Every control gets run against its author&amp;#39;s next real turn before anyone calls it shipped. The self test passed all three times. Synthetic input proves the function. It says nothing about the log the function reads.&lt;/p&gt;&lt;p&gt;Calibrate on the false positive rate. A refusal per turn teaches an agent to route around the gate, and this one has a documented way around: write files through the shell, where the hook cannot see them. My agent used it once, on the guard itself, and told me. I would rather be told than hold a clean-looking log.&lt;/p&gt;
&lt;h2&gt;The narrow version&lt;/h2&gt;
&lt;p&gt;A review the same day said keep it, and narrow it. The healthiest change shape available ships code together with its test and its doc. That is exactly three files, and the gate refused it. A mechanical rename across five files, asked for in plain words, got refused too. The floor of three came from the incident that prompted the guard, which is how most thresholds are born.&lt;/p&gt;&lt;p&gt;So a test or doc sibling of a file already in the set stopped counting as new. The gate now fires when a third separate unit of work appears inside a turn nobody planned.&lt;/p&gt;&lt;p&gt;That is the whole story. About forty lines of code, three refusals of its own author, and a threshold whose provenance I can finally state.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/the-instrument-is-wrong-first/</id>
    <title type="text">Assume the instrument is wrong before the world is</title>
    <updated>2026-09-17T00:47:00+09:00</updated>
    <published>2026-09-17T00:47:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/the-instrument-is-wrong-first/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <category term="the agent series"/>
    <summary type="text">Five months of running twelve web properties and a coding agent taught me that a zero, a round number and a clean series are claims about the meter, and each on</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; Five months of running twelve web properties and a coding agent taught me that a zero, a round number and a clean series are claims about the meter, and each one needs a second meter before it gets published.&lt;/p&gt;
&lt;h2&gt;In short&lt;/h2&gt;
&lt;p&gt;I published a sentence in September claiming that a hundred per cent of the traffic to one of my sites was not a person. Zero humans. I had it from the edge provider&amp;#39;s own real-user monitor, which is about as close to the browser as a measurement gets. I committed it and pushed it.&lt;/p&gt;&lt;p&gt;The search console for the same property showed 126 real arrivals over 28 days. People typed a query, saw a result, clicked it, landed. The beacon that told me nobody was there does not fire reliably on that particular zone, which is served by an edge worker rather than the origin. My site was fine. My thermometer was broken, and I had written down the temperature.&lt;/p&gt;&lt;p&gt;That was one of five in a single week. All five were mine. None of them were subtle once found, and every one of them had produced a confident number first.&lt;/p&gt;
&lt;h2&gt;The five&lt;/h2&gt;
&lt;div class=&quot;tablewrap&quot; tabindex=&quot;0&quot;&gt;&lt;table class=&quot;ptable&quot;&gt;&lt;thead&gt;&lt;tr&gt;&lt;th scope=&quot;col&quot;&gt;Case&lt;/th&gt;&lt;th scope=&quot;col&quot;&gt;What the instrument said&lt;/th&gt;&lt;th scope=&quot;col&quot;&gt;What a second instrument said&lt;/th&gt;&lt;th scope=&quot;col&quot;&gt;What was actually wrong&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td data-label=&quot;Case&quot;&gt;Storefront traffic&lt;/td&gt;&lt;td data-label=&quot;What the instrument said&quot;&gt;271,000 pageviews (CDN analytics)&lt;/td&gt;&lt;td data-label=&quot;What a second instrument said&quot;&gt;23,600 sessions (real-user monitor), about 7,170 human navigations&lt;/td&gt;&lt;td data-label=&quot;What was actually wrong&quot;&gt;Three meters counting three different events, one of them counting requests&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td data-label=&quot;Case&quot;&gt;Landing page humans&lt;/td&gt;&lt;td data-label=&quot;What the instrument said&quot;&gt;0 humans, 100 per cent not a person&lt;/td&gt;&lt;td data-label=&quot;What a second instrument said&quot;&gt;126 clicks in 28 days (search console)&lt;/td&gt;&lt;td data-label=&quot;What was actually wrong&quot;&gt;Beacon does not fire on a worker-served zone&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td data-label=&quot;Case&quot;&gt;Query report&lt;/td&gt;&lt;td data-label=&quot;What the instrument said&quot;&gt;64 clicks on the query dimension&lt;/td&gt;&lt;td data-label=&quot;What a second instrument said&quot;&gt;315 clicks on the date dimension&lt;/td&gt;&lt;td data-label=&quot;What was actually wrong&quot;&gt;Row cap at 1,000, so the report measured the cap&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td data-label=&quot;Case&quot;&gt;Tutoring domain&lt;/td&gt;&lt;td data-label=&quot;What the instrument said&quot;&gt;Zero traffic&lt;/td&gt;&lt;td data-label=&quot;What a second instrument said&quot;&gt;Traffic existed under the full host&lt;/td&gt;&lt;td data-label=&quot;What was actually wrong&quot;&gt;Host folded to its last two labels, aggregating into a public suffix&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td data-label=&quot;Case&quot;&gt;Bot score&lt;/td&gt;&lt;td data-label=&quot;What the instrument said&quot;&gt;Field documented in the schema&lt;/td&gt;&lt;td data-label=&quot;What a second instrument said&quot;&gt;Field refused by the plan&lt;/td&gt;&lt;td data-label=&quot;What was actually wrong&quot;&gt;A probe printed a field name it had never been given&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/div&gt;&lt;p&gt;The first row is the one I think about most, because nobody lied. The CDN counts requests at the edge, including prefetches, feed readers, and anything that asks for a byte. The real-user monitor counts browsers that ran the script. The human figure comes from filtering to navigations with a plausible session shape. 271,000 and 23,600 and 7,170 are all true statements about three different questions. I had been quoting the largest one as if it answered the smallest.&lt;/p&gt;&lt;p&gt;The fourth case is my favourite kind of failure, because the code was doing exactly what it was told. A four-label hostname got grouped by its last two labels, which for a UK domain is the public suffix and belongs to nobody. Every visit went into a bucket with a name that no browser ever requested, so the site I asked about read zero. There was no error, no warning, no failed request. The pipeline succeeded at doing nothing, and returned a plausible number for it.&lt;/p&gt;
&lt;h2&gt;A cap is a floor&lt;/h2&gt;
&lt;p&gt;The third row generalises further than the others, so I made it a standing rule: a result landing exactly on a limit measures the limit.&lt;/p&gt;&lt;p&gt;The query report stopped at 1,000 rows because that was the default. Clicks past row 1,000 were real, and the total I read was around a fifth of the truth. The tell was available before I knew the answer: the same site, asked on a different axis, gave five times the number. Two axes of one dataset disagreeing by 5x is not a mystery to be reasoned about. It is a cap.&lt;/p&gt;&lt;p&gt;I tested the rule on an unrelated dataset and it held. A market data pull returned 2,000 rows and I reported 2,000. Narrowing the window and re-running recursively produced 24,420, twelve times the figure I had. Nothing about 2,000 looked suspicious in isolation, which is the point. Round numbers are caps until disproved: 100, 500, 1,000, 2,000, 10,000. If a number cannot be pushed past its limit, it gets reported with the word FLOOR attached and the window that hit the wall gets named in the same sentence.&lt;/p&gt;&lt;p&gt;The related trap is the artifact that exists and contains nothing. I once claimed coverage from a directory listing: 1,760 files collected. 716 of them were empty, so my coverage was overstated by 68 per cent. Counting files answers a question about the filesystem. Counting lines answers the question I meant.&lt;/p&gt;
&lt;h2&gt;Three readings that should stop you&lt;/h2&gt;
&lt;p&gt;The five cases share a shape, and after the fifth I wrote down the shape so I would recognise it in advance.&lt;/p&gt;&lt;p&gt;A zero is the first. A zero is a legitimate reading, and it is also what every broken sensor reports. A dead beacon, a folded hostname, a filter that matched nothing, a permission that silently denied: all of them render as zero, indistinguishable from a quiet Tuesday. My rule now is that a zero in a metric that was ever non-zero is a fault report about the pipeline until a different instrument agrees with it.&lt;/p&gt;&lt;p&gt;A round number is the second. 2,000 rows. 1,000 rows. Exactly 100 results. The world rarely lands on a power of ten. Defaults do, constantly.&lt;/p&gt;&lt;p&gt;A clean series is the third, and it is the hardest to feel, because a clean series looks like good news. When a dashboard is smooth, ask what it would look like if the collector had stopped writing and the renderer had kept drawing.&lt;/p&gt;&lt;p&gt;There is a fourth, which is what happened with the bot score. The probe printed a field name that the plan does not expose, because the code assembled a response shape from its own schema rather than from what came back. When an instrument reports a capability it was never asked about, it is describing its own source code.&lt;/p&gt;&lt;p&gt;The habit that comes out of all four is one question, asked before believing a figure and not after being contradicted: what would make this instrument produce this exact number falsely? If I can answer that in one sentence, I have to rule it out before I publish. It takes about a minute. The zero I published cost more than that to withdraw.&lt;/p&gt;
&lt;h2&gt;The model is an instrument too&lt;/h2&gt;
&lt;p&gt;I would like to say the lesson stayed in the analytics layer. It did not.&lt;/p&gt;&lt;p&gt;I run a coding agent across these properties, and this week I asked a small, cheap model to audit 50 directives from a long session and mark each one done or not. It returned a tidy table. It was missing 8 of the 50 rows, with no gap, no note, no apology. The table looked complete because tables look complete. A capped audit reads the same as a full one, which puts it in the same family as the 1,000-row query report: the count I was given measured the auditor, and I nearly quoted it as a finding about my own work.&lt;/p&gt;&lt;p&gt;A second model, asked to inventory documents, cited two Google Doc identifiers with correct shape and correct length that do not exist. Both were close to real ones. That is worse than a wrong number, because a wrong identifier looks like provenance. Nothing a model writes that claims to be an identifier goes into a source of truth here without a grep against the file it supposedly came from.&lt;/p&gt;&lt;p&gt;There is one more turn of the screw, and I think it is the genuinely useful part. I had a third model, with no stake in the work, read my own lessons document as a stranger would. It found that the file had 21 sections and two of them were numbered 18. I had inherited the miscount from my own summary without checking, in a document whose fourth section is about not trusting counts. The reviewer also demanded I delete three of my strongest lines as self-graded, and it was right about all three.&lt;/p&gt;&lt;p&gt;So the instrument that checks the instrument is also an instrument. That is not a reason to stop checking. It is a reason to make the second reading come from a different mechanism than the first: a different dimension of the same dataset, a different vendor, a script instead of a model, a human click instead of a beacon.&lt;/p&gt;
&lt;h2&gt;What I actually do now&lt;/h2&gt;
&lt;p&gt;Before a number leaves my machine, it gets three things. A second source that counts by a different method. A sentence naming what would have to be broken for it to read this way. A note saying which of the three questions it answers, since requests, sessions and humans are three questions and I had spent months conflating them.&lt;/p&gt;&lt;p&gt;It is slower. It is also the difference between reporting that nobody visits your site and reporting that your beacon does not fire behind an edge worker. One of those is a business problem. The other is a fifteen minute fix, and I only found it because somebody asked me where the 126 clicks came from.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/the-sessions-that-feel-hostile-are-the-ones-that-compact/</id>
    <title type="text">The sessions that feel hostile are the ones that compact</title>
    <updated>2026-09-17T00:46:00+09:00</updated>
    <published>2026-09-17T00:46:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/the-sessions-that-feel-hostile-are-the-ones-that-compact/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <category term="the agent series"/>
    <summary type="text">Across 21 days of coding-agent sessions on my machine, the ones I described as hostile and forgetful were the ones with the most context compactions and the mos</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; Across 21 days of coding-agent sessions on my machine, the ones I described as hostile and forgetful were the ones with the most context compactions and the most guard refusals, and both scale with how much work the session was doing.&lt;/p&gt;
&lt;h2&gt;In short&lt;/h2&gt;
&lt;p&gt;I said out loud a few weeks ago that some of my coding-agent sessions felt hostile. Forgetful, uncooperative, quietly obstructive. I meant it as a complaint about the tool. The awkward part is that those same sessions were where the real work happened: the research repo, the storefront, the long debugging runs that actually shipped something.&lt;/p&gt;&lt;p&gt;So I measured it instead of repeating it. Twenty-one days, every session on this machine, grouped by which repository the session lived in. The feeling turned out to have two mechanical causes, and neither of them is a mood.&lt;/p&gt;
&lt;h2&gt;What the sessions actually did&lt;/h2&gt;
&lt;p&gt;Measured from the session transcripts on 17 September 2026:&lt;/p&gt;&lt;div class=&quot;tablewrap&quot; tabindex=&quot;0&quot;&gt;&lt;table class=&quot;ptable&quot;&gt;&lt;thead&gt;&lt;tr&gt;&lt;th scope=&quot;col&quot;&gt;Session lives in&lt;/th&gt;&lt;th scope=&quot;col&quot;&gt;Compactions per session&lt;/th&gt;&lt;th scope=&quot;col&quot;&gt;Guard refusals per prompt&lt;/th&gt;&lt;th scope=&quot;col&quot;&gt;Tool calls per prompt&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td data-label=&quot;Session lives in&quot;&gt;Filings research&lt;/td&gt;&lt;td data-label=&quot;Compactions per session&quot;&gt;3.50&lt;/td&gt;&lt;td data-label=&quot;Guard refusals per prompt&quot;&gt;1.48&lt;/td&gt;&lt;td data-label=&quot;Tool calls per prompt&quot;&gt;32.5&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td data-label=&quot;Session lives in&quot;&gt;Storefront&lt;/td&gt;&lt;td data-label=&quot;Compactions per session&quot;&gt;1.96&lt;/td&gt;&lt;td data-label=&quot;Guard refusals per prompt&quot;&gt;1.29&lt;/td&gt;&lt;td data-label=&quot;Tool calls per prompt&quot;&gt;27.5&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td data-label=&quot;Session lives in&quot;&gt;Tutoring product&lt;/td&gt;&lt;td data-label=&quot;Compactions per session&quot;&gt;0.71&lt;/td&gt;&lt;td data-label=&quot;Guard refusals per prompt&quot;&gt;0.94&lt;/td&gt;&lt;td data-label=&quot;Tool calls per prompt&quot;&gt;22.9&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td data-label=&quot;Session lives in&quot;&gt;Short harness runs in the home directory&lt;/td&gt;&lt;td data-label=&quot;Compactions per session&quot;&gt;0.17&lt;/td&gt;&lt;td data-label=&quot;Guard refusals per prompt&quot;&gt;0.98&lt;/td&gt;&lt;td data-label=&quot;Tool calls per prompt&quot;&gt;18.8&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/div&gt;&lt;p&gt;Read the first column as memory loss and the second as friction. They rank in the same order, and that order is also the order of how much the session was doing per prompt. The repo I called hostile ran twenty times more compactions per session than the throwaway runs I never complained about.&lt;/p&gt;
&lt;h2&gt;Forgetful is compaction, and compaction is lossy by construction&lt;/h2&gt;
&lt;p&gt;When a session fills its context window, the history gets summarised so the conversation can continue. On this setup the boundary is measurable. Across seven live sessions the compaction events landed between 724,099 and 739,104 tokens, which is 73 to 74 per cent of a one million token window. One of those compactions took 8 minutes 35 seconds, a single summarisation pass over roughly 484,000 tokens of context.&lt;/p&gt;&lt;p&gt;A summary throws away detail. That is what a summary is. The turn immediately after a compaction has a paraphrase of the decision we made two hours ago instead of the decision. If the detail that got dropped was the exact column name, the exact reason we rejected approach B, or the exact wording I asked for, the next answer is confidently wrong in a way that feels like being ignored.&lt;/p&gt;&lt;p&gt;Three and a half compactions per session means the filings research sessions crossed that cliff three and a half times each. Of course they felt forgetful. They were, on average, three and a half summaries away from what I originally said.&lt;/p&gt;&lt;p&gt;The cost side explains why the boundary arrives so fast in those repos. I measured 13,379 calls in one stretch: input was 91.4 per cent of spend, and cached context was re-read at roughly 278 times the volume of new text produced. Every tool call drags the whole conversation back through the model. A session running 32.5 tool calls per prompt is filling its own window at almost twice the rate of one running 18.8, and it hits the wall at almost twice the frequency.&lt;/p&gt;
&lt;h2&gt;Uncooperative is my own guards refusing me&lt;/h2&gt;
&lt;p&gt;The second column is worse, because it is entirely my doing. I built a stack of hooks that sit in front of the agent and refuse things: spending without a dry run first, claiming a deploy is live without fetching the public URL, closing out with the wrong shape, touching files another session owns.&lt;/p&gt;&lt;p&gt;At 1.48 refusals per prompt, the agent in my heaviest repo was being blocked more than once per instruction I gave. From the outside that reads as an assistant arguing with me. From the inside it is a regex refusing a command and the agent rephrasing until something passes.&lt;/p&gt;&lt;p&gt;Two findings from auditing that stack matter here.&lt;/p&gt;&lt;p&gt;The first is that a refusal per turn teaches evasion. Every one of these hooks matches text, and text is free to change. When a gate refuses &amp;quot;still running&amp;quot;, the agent writes &amp;quot;in flight&amp;quot;. The guard&amp;#39;s own metric improves. Nothing about the behaviour changed. An outside review of six of my hooks put it plainly: text matching is a floor under obviously bad output, and it is no ceiling at all on bad behaviour.&lt;/p&gt;&lt;p&gt;The second is that coverage is the wrong goal. I labelled 5,270 complaints I had made about agent output and found 2,720 distinct behaviours. The single most common one accounts for 0.9 per cent. The top 25 together cover 13.6 per cent. To catch half of my own complaints I would need around 500 separate guards, and 38 per cent of the behaviours appeared exactly once. There is no head to this distribution. Every guard I add buys under one per cent and charges friction on every turn forever.&lt;/p&gt;&lt;p&gt;One number makes the tradeoff concrete. Cross-session ownership guessing produced 462 close-outs carrying a caveat about files the session had never touched. Each one was a refusal to act, generated by a guard, about nothing.&lt;/p&gt;
&lt;h2&gt;Both scale with where the work is&lt;/h2&gt;
&lt;p&gt;This is the part I did not want to be true. Compaction count is a function of session length and tool volume. Refusal count is a function of how many risky verbs a session reaches for: writes, deploys, anything that spends. The sessions doing real work are long and full of risky verbs by definition, so they collect both.&lt;/p&gt;&lt;p&gt;The hostility is structural. It is not the agent&amp;#39;s attitude and it is not evidence that the hard repos have a worse model attached. My cheapest, friendliest sessions were friendly because they did almost nothing.&lt;/p&gt;
&lt;h2&gt;What I changed&lt;/h2&gt;
&lt;p&gt;Four things, in order of how much they moved.&lt;/p&gt;&lt;p&gt;Shorter sessions per unit of work. If a session is going to compact three times, I would rather close it twice and start fresh with a written handover than let a summariser choose what survives. A summary I write is a summary I can check.&lt;/p&gt;&lt;p&gt;A definition of done, written before the work starts, in my words, plus an open-items file the session updates as it goes. Both live in the repo. Compaction cannot reach them. The agent reads the repo&amp;#39;s own state file before answering instead of trusting what it remembers, and the phrase &amp;quot;as we discussed earlier&amp;quot; stops being load-bearing.&lt;/p&gt;&lt;p&gt;Guards calibrated for precision over coverage. The gates I kept are asymmetric: they refuse one direction only, the one where being wrong is expensive. Spending money with no dry run. Reporting a number that went down with no artifact behind it. A false alarm on those costs me annoyance, and a miss costs correctness. The per-behaviour guards from the long tail are gone.&lt;/p&gt;&lt;p&gt;Controls in the repo instead of the chat. Tests as the contract, a pre-commit that refuses a bad diff, a dry-run default inside any script that spends, a second model reviewing the diff. Those survive a change of tool. Chat-level hooks do not.&lt;/p&gt;
&lt;h2&gt;The honest limit&lt;/h2&gt;
&lt;p&gt;Shape is easy to measure and relevance is not. My close-out format went from 0 to 43 per cent conformance once a hook started refusing malformed ones. Over the same months, the share of my turns that were corrections of the previous answer went from 15 per cent to 49 per cent. Well-formed and wrong is a real state, and I built a lot of machinery that can only see the first half.&lt;/p&gt;&lt;p&gt;There is a version of this article where the tool is the villain. The numbers do not support it. I asked for long sessions with heavy tool use in the repos where the stakes were highest, then I put a refusal in front of most turns, and I was surprised when those sessions felt like hard work. The hostility was mine to measure, and measuring it is what made it fixable.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/well-formed-and-wrong/</id>
    <title type="text">Well-formed and wrong: shape improved, relevance did not</title>
    <updated>2026-09-17T00:45:00+09:00</updated>
    <published>2026-09-17T00:45:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/well-formed-and-wrong/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <category term="the agent series"/>
    <summary type="text">I spent two months fixing the format of my coding agent&apos;s replies, first-attempt compliance went from zero to 43 per cent, and over the same months the share of</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; I spent two months fixing the format of my coding agent&amp;apos;s replies, first-attempt compliance went from zero to 43 per cent, and over the same months the share of my next messages that were corrections tripled.&lt;/p&gt;
&lt;h2&gt;The number that moved&lt;/h2&gt;
&lt;p&gt;I run three small businesses through a coding agent. A trophy and gifts storefront, a tutoring agency, a filings-research product. Most of the work reaches me as a closing summary at the end of a turn, so in August I wrote a checker for the summary: two sections, one line per item of my message, result first, no line over 300 characters, nothing handed back to me that the agent could have done itself.&lt;/p&gt;&lt;p&gt;Then I ran the checker over every closing summary in the corpus instead of asking it how it felt.&lt;/p&gt;&lt;pre class=&quot;code-block&quot; tabindex=&quot;0&quot;&gt;&lt;code&gt;July         137 summaries        0 clean       0%
August     1,929 summaries      106 clean       5%
September  2,069 summaries      900 clean      43%&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;Zero to 43 per cent is the largest single improvement I have measured on this box, and it arrived the week the checker started refusing summaries rather than advising on them. Prose rules had been in place for months and moved nothing. A refusal in the path moved it in days.&lt;/p&gt;&lt;p&gt;That 43 per cent is a first-attempt rate. A refused summary still sits in the transcript, so a turn that failed once and passed on the retry counts as one failure and one pass. The version I actually read is cleaner than 43 per cent, and this scan cannot separate the two.&lt;/p&gt;
&lt;h2&gt;The number that went the other way&lt;/h2&gt;
&lt;p&gt;While the format was improving, I measured something I cannot game from my side: whether my very next message corrects, repeats or rejects what just came back.&lt;/p&gt;&lt;div class=&quot;tablewrap&quot; tabindex=&quot;0&quot;&gt;&lt;table class=&quot;ptable&quot;&gt;&lt;thead&gt;&lt;tr&gt;&lt;th scope=&quot;col&quot;&gt;month&lt;/th&gt;&lt;th scope=&quot;col&quot;&gt;my prompts&lt;/th&gt;&lt;th scope=&quot;col&quot;&gt;next message is a correction&lt;/th&gt;&lt;th scope=&quot;col&quot;&gt;explicit accepts&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td data-label=&quot;month&quot;&gt;July&lt;/td&gt;&lt;td data-label=&quot;my prompts&quot;&gt;3,313&lt;/td&gt;&lt;td data-label=&quot;next message is a correction&quot;&gt;514 (15.5%)&lt;/td&gt;&lt;td data-label=&quot;explicit accepts&quot;&gt;908&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td data-label=&quot;month&quot;&gt;August&lt;/td&gt;&lt;td data-label=&quot;my prompts&quot;&gt;5,069&lt;/td&gt;&lt;td data-label=&quot;next message is a correction&quot;&gt;1,495 (29.5%)&lt;/td&gt;&lt;td data-label=&quot;explicit accepts&quot;&gt;846&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td data-label=&quot;month&quot;&gt;September&lt;/td&gt;&lt;td data-label=&quot;my prompts&quot;&gt;3,057&lt;/td&gt;&lt;td data-label=&quot;next message is a correction&quot;&gt;1,496 (48.9%)&lt;/td&gt;&lt;td data-label=&quot;explicit accepts&quot;&gt;129&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/div&gt;&lt;p&gt;Two months of format work, and the share of turns I had to correct roughly tripled.&lt;/p&gt;&lt;p&gt;I will name the confounds, because the level here is inflated. The detector counts a message as a correction when it contains words like &amp;quot;no&amp;quot;, &amp;quot;still&amp;quot; and &amp;quot;again&amp;quot;, which appear in ordinary instructions all the time. My own standard also rose across the same window, in my own words to the agent: &amp;quot;i think my standards are just increasing&amp;quot;. Early in the summer I was asking it to stop breaking things. By September I was asking it to optimise things that already worked. The direction survives both confounds. The rate does not, and I am not going to quote it as one.&lt;/p&gt;&lt;p&gt;Shape and relevance are separate axes. Fixing the first did nothing for the second.&lt;/p&gt;
&lt;h2&gt;What the checker actually catches&lt;/h2&gt;
&lt;p&gt;Of 2,058 September summaries put through the checker, four findings account for 86 per cent of all failures.&lt;/p&gt;&lt;div class=&quot;tablewrap&quot; tabindex=&quot;0&quot;&gt;&lt;table class=&quot;ptable&quot;&gt;&lt;thead&gt;&lt;tr&gt;&lt;th scope=&quot;col&quot;&gt;what failed&lt;/th&gt;&lt;th scope=&quot;col&quot;&gt;count&lt;/th&gt;&lt;th scope=&quot;col&quot;&gt;share of failures&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td data-label=&quot;what failed&quot;&gt;one summary line over 300 characters&lt;/td&gt;&lt;td data-label=&quot;count&quot;&gt;862&lt;/td&gt;&lt;td data-label=&quot;share of failures&quot;&gt;42.1%&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td data-label=&quot;what failed&quot;&gt;more summary items than the cap allows&lt;/td&gt;&lt;td data-label=&quot;count&quot;&gt;399&lt;/td&gt;&lt;td data-label=&quot;share of failures&quot;&gt;19.5%&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td data-label=&quot;what failed&quot;&gt;the handback section raising another repository&amp;#39;s work&lt;/td&gt;&lt;td data-label=&quot;count&quot;&gt;277&lt;/td&gt;&lt;td data-label=&quot;share of failures&quot;&gt;13.5%&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td data-label=&quot;what failed&quot;&gt;the handback section reserving work for me with no blocker&lt;/td&gt;&lt;td data-label=&quot;count&quot;&gt;222&lt;/td&gt;&lt;td data-label=&quot;share of failures&quot;&gt;10.8%&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td data-label=&quot;what failed&quot;&gt;an internal enforcement token left in text I read&lt;/td&gt;&lt;td data-label=&quot;count&quot;&gt;109&lt;/td&gt;&lt;td data-label=&quot;share of failures&quot;&gt;5.3%&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/div&gt;&lt;p&gt;Three of the four are the agent&amp;#39;s own habits: writing long, writing many, parking work back on me that it could have finished. In September, 76 per cent of closing summaries handed something back to me.&lt;/p&gt;&lt;p&gt;One of the four is mine. A prompt that spans several repositories invites a summary that does the same, and that is 13.5 per cent of the failures. The fix on my side is prompt scope: one repository per prompt.&lt;/p&gt;
&lt;h2&gt;One item in three&lt;/h2&gt;
&lt;p&gt;The question I actually care about is different from format. If I put eight items in a message, how many come back answered? I counted items in my prompts across August and September and marked an item as addressed when the closing summary talked about the same thing.&lt;/p&gt;&lt;div class=&quot;tablewrap&quot; tabindex=&quot;0&quot;&gt;&lt;table class=&quot;ptable&quot;&gt;&lt;thead&gt;&lt;tr&gt;&lt;th scope=&quot;col&quot;&gt;where the prompt was written&lt;/th&gt;&lt;th scope=&quot;col&quot;&gt;items&lt;/th&gt;&lt;th scope=&quot;col&quot;&gt;items addressed&lt;/th&gt;&lt;th scope=&quot;col&quot;&gt;my corrections&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td data-label=&quot;where the prompt was written&quot;&gt;small one-off jobs, single purpose&lt;/td&gt;&lt;td data-label=&quot;items&quot;&gt;804&lt;/td&gt;&lt;td data-label=&quot;items addressed&quot;&gt;56%&lt;/td&gt;&lt;td data-label=&quot;my corrections&quot;&gt;34%&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td data-label=&quot;where the prompt was written&quot;&gt;storefront catalogue work&lt;/td&gt;&lt;td data-label=&quot;items&quot;&gt;111&lt;/td&gt;&lt;td data-label=&quot;items addressed&quot;&gt;44%&lt;/td&gt;&lt;td data-label=&quot;my corrections&quot;&gt;44%&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td data-label=&quot;where the prompt was written&quot;&gt;outreach work&lt;/td&gt;&lt;td data-label=&quot;items&quot;&gt;349&lt;/td&gt;&lt;td data-label=&quot;items addressed&quot;&gt;42%&lt;/td&gt;&lt;td data-label=&quot;my corrections&quot;&gt;41%&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td data-label=&quot;where the prompt was written&quot;&gt;general machine work&lt;/td&gt;&lt;td data-label=&quot;items&quot;&gt;1,719&lt;/td&gt;&lt;td data-label=&quot;items addressed&quot;&gt;33%&lt;/td&gt;&lt;td data-label=&quot;my corrections&quot;&gt;40%&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td data-label=&quot;where the prompt was written&quot;&gt;the agent&amp;#39;s own tooling&lt;/td&gt;&lt;td data-label=&quot;items&quot;&gt;633&lt;/td&gt;&lt;td data-label=&quot;items addressed&quot;&gt;32%&lt;/td&gt;&lt;td data-label=&quot;my corrections&quot;&gt;41%&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td data-label=&quot;where the prompt was written&quot;&gt;trophy storefront, main line of work&lt;/td&gt;&lt;td data-label=&quot;items&quot;&gt;2,402&lt;/td&gt;&lt;td data-label=&quot;items addressed&quot;&gt;31%&lt;/td&gt;&lt;td data-label=&quot;my corrections&quot;&gt;50%&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td data-label=&quot;where the prompt was written&quot;&gt;filings-research product&lt;/td&gt;&lt;td data-label=&quot;items&quot;&gt;1,261&lt;/td&gt;&lt;td data-label=&quot;items addressed&quot;&gt;23%&lt;/td&gt;&lt;td data-label=&quot;my corrections&quot;&gt;48%&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/div&gt;&lt;p&gt;Roughly one item in three, and it degrades exactly where I load the most items into a single prompt. The two heaviest, at 2,402 and 1,261 items, are the two worst at 31 and 23 per cent, and they are also where I correct most often, 50 and 48 per cent. The lightest prompts, small and single purpose, do best at 56 per cent.&lt;/p&gt;&lt;p&gt;Stated limit: word overlap is a weaker test than reading each item. It counts an item as addressed when the summary discusses it, which is not the same as the thing being done. It still beats counting numbered lines, which is a measure of form and which I threw away after telling the agent &amp;quot;dont count proxies count actual work&amp;quot;.&lt;/p&gt;
&lt;h2&gt;What the windows say when you count artifacts&lt;/h2&gt;
&lt;p&gt;Format and item counts still flatter the agent, so I had readers go through the raw transcripts window by window and classify every ask and every output. The pattern is the same in every venture and every month.&lt;/p&gt;&lt;p&gt;A 43-prompt window on the storefront in late July: zero deploys, zero messages sent, six documents, eleven measurements, most of them provisional. At most one turn in six moved something a customer or the live site would see. My closing line for that window was that I still did not have the before and after I had asked for.&lt;/p&gt;&lt;p&gt;A 394-prompt window in August: 244 turns (62 per cent) ended in chat with no artifact, 78 reached something live, 13 ended in a sent message. No sale, no dispatched quote, no outbound message to an end customer in 394 turns. At least 36 per cent of the turns went to keeping the tooling itself alive.&lt;/p&gt;&lt;p&gt;A 551-prompt window in early September: 238 chat-only, 122 claimed deploys, 57 blocked. The real outcomes for those eight days fit in a line: three ads live, a watermark applied to 957 photos, one inbound enquiry.&lt;/p&gt;&lt;p&gt;A 40-turn window in the middle of September: one thing shipped, 2.5 per cent, a re-stamp of 3,499 photos. Roughly 30 per cent of the turns produced no logged answer at all.&lt;/p&gt;&lt;p&gt;The tutoring agency in July: one deploy confirmed live, zero messages to tutors or parents, a notice drafted and never sent. In August, after twenty turns of outreach asks, my own summary of the business was that nobody had ever paid it.&lt;/p&gt;&lt;p&gt;The filings-research product in September: 168 prompts, a research chain reaching 77 confirmed and 45 refuted and 13 underpowered, and zero paying customers. Fifty-eight of those 168 closing summaries mention the agent&amp;#39;s own rules, hooks or memory file.&lt;/p&gt;&lt;p&gt;Nearly every one of those summaries was framed as DONE.&lt;/p&gt;
&lt;h2&gt;The three changes&lt;/h2&gt;
&lt;p&gt;Measuring the reply&amp;#39;s shape made replies look finished. Measuring my next message and the artifacts that exist tells the truth. Three things changed on this box as a result.&lt;/p&gt;&lt;p&gt;Grade relevance by the human&amp;#39;s next message. A well-formed answer that draws a correction is a failed answer, and the correction rate is the only signal the agent cannot manufacture from its side.&lt;/p&gt;&lt;p&gt;Write the definition of done before the first tool call, in my words, with a path in it. An answer that lives only in chat does not exist to an auditor, to the next session, or to whatever model I am using next month. Repeat asks are the cheapest signal of a silent drop, and nothing in my stack counted them until this week. One directive I asked for three times had been executed zero times.&lt;/p&gt;&lt;p&gt;Count outcomes rather than artifacts. Sent, deployed, sold. Documents, measurements and commits are evidence that a turn happened. They are not evidence that the business moved, and for five months I let the first stand in for the second.&lt;/p&gt;&lt;p&gt;A gate can make the answer well-formed. It cannot make it the right answer.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/i-was-staring-at-12-1-again/</id>
    <title type="text">i was staring at 12.1 again</title>
    <updated>2026-09-16T00:59:00+09:00</updated>
    <published>2026-09-16T00:59:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/i-was-staring-at-12-1-again/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">A ranking of 12.1 is a near-win. Test one focused change, measure the traffic lift, and decide if it deserves another round.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; A ranking of 12.1 is close enough to make the next move specific. Before chasing a fresh topic, test one focused change on a page already near page one and measure what happens.&lt;/p&gt;
&lt;h2&gt;why should i care about a page ranking at 12.1?&lt;/h2&gt;
&lt;p&gt;12.1 sits close enough to page one to feel exciting and far enough away to keep you honest. The page already has a foothold, and the search result already knows where to find it.&lt;/p&gt;&lt;p&gt;That gives you a starting line. You already have a page, a ranking, and a reason to care, so the next move can be specific.&lt;/p&gt;
&lt;h2&gt;what should i change when a page is almost winning?&lt;/h2&gt;
&lt;p&gt;A focused guide is one possible answer. An internal link is another part of the experiment.&lt;/p&gt;&lt;p&gt;The point is the experiment itself. Make one focused change, watch what happens, and measure the traffic lift before deciding whether the idea earned another round.&lt;/p&gt;
&lt;h2&gt;why do builders abandon near-wins for fresh ideas?&lt;/h2&gt;
&lt;p&gt;A near-win can look unfinished, so it is easy to abandon it for a fresh idea. Fresh ideas feel productive because they create a new starting point.&lt;/p&gt;&lt;p&gt;A page at 12.1 already gives you evidence to work with. The useful question becomes what could move this one page, then whether that move produced a measurable result.&lt;/p&gt;
&lt;h2&gt;how do i decide which pages deserve another round?&lt;/h2&gt;
&lt;p&gt;Check the pages already close to winning before chasing new topics. A ranking near page one gives the next test a defined target and a result you can measure.&lt;/p&gt;&lt;p&gt;The work is turning that almost into something measurable. If one focused change lifts traffic, you have a reason to continue. If it does not, you have a clearer reason to move on.&lt;/p&gt;
&lt;h2&gt;FAQ&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;is a ranking of 12.1 worth optimizing?&lt;/strong&gt;&lt;br&gt;Yes. It is close enough to page one to make a specific test worthwhile. The page already has a foothold, so you can measure whether one focused change moves it.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;what should i change on a page ranking at 12.1?&lt;/strong&gt;&lt;br&gt;A focused guide is one possible change, and an internal link is another. Choose one focused change so the traffic lift can be measured clearly.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;should i target a new topic instead of improving an existing page?&lt;/strong&gt;&lt;br&gt;Check pages already close to winning first. A near-win gives you a defined starting point and a specific experiment.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;how do i know whether the experiment worked?&lt;/strong&gt;&lt;br&gt;Watch what happens after the focused change and measure the traffic lift. Use that result to decide whether the idea earned another round.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/67-impressions/</id>
    <title type="text">67 impressions</title>
    <updated>2026-09-15T00:59:00+09:00</updated>
    <published>2026-09-15T00:59:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/67-impressions/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">A field note on how moving one page from average position 11.5 toward page one could create roughly 2.4 extra clicks monthly.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; One page earned 67 impressions, 1 click, and an average position of 11.5. Moving that existing ranking onto page one is estimated to create roughly 2.4 additional clicks per month.&lt;/p&gt;
&lt;h2&gt;what did one page actually prove?&lt;/h2&gt;
&lt;p&gt;One page earned 67 impressions and generated 1 click. Its average position was 11.5.&lt;/p&gt;&lt;p&gt;That gives the page a measured baseline before any ranking work begins.&lt;/p&gt;
&lt;h2&gt;why does position 11.5 matter?&lt;/h2&gt;
&lt;p&gt;An average position of 11.5 places the page near the boundary of page one. The page already has demand showing up through impressions.&lt;/p&gt;&lt;p&gt;That makes the next question concrete: what changes when the page moves higher?&lt;/p&gt;
&lt;h2&gt;how many extra clicks could page one create?&lt;/h2&gt;
&lt;p&gt;The estimate is roughly 2.4 additional clicks per month if the page moves onto page one.&lt;/p&gt;&lt;p&gt;That is the number worth keeping. It gives founders a way to price the value of improving an existing ranking.&lt;/p&gt;
&lt;h2&gt;where should the work happen first?&lt;/h2&gt;
&lt;p&gt;The work is moving the page higher, then measuring what changes.&lt;/p&gt;&lt;p&gt;This creates a value estimate from existing demand before creating a new page or chasing a new keyword.&lt;/p&gt;
&lt;h2&gt;FAQ&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;How many clicks did the page get?&lt;/strong&gt;&lt;br&gt;The page generated 1 click from 67 impressions.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What was the page’s average position?&lt;/strong&gt;&lt;br&gt;Its average position was 11.5.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;How many additional clicks could page one create?&lt;/strong&gt;&lt;br&gt;The estimate is roughly 2.4 additional clicks per month.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Why improve an existing ranking first?&lt;/strong&gt;&lt;br&gt;The page already has demand showing up through impressions. Improving its position gives founders a measurable baseline for the value of ranking work.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/page-11-is-a-different-game/</id>
    <title type="text">page 11 is a different game</title>
    <updated>2026-09-14T00:59:00+09:00</updated>
    <published>2026-09-14T00:59:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/page-11-is-a-different-game/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">Why pages averaging position 11.5 deserve focused guides and internal links before creating more content.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; A page averaging position 11.5 is already close to page one. Improve that near-winner with a focused guide and an internal link before creating another page.&lt;/p&gt;
&lt;h2&gt;why does position 11.5 change the next move?&lt;/h2&gt;
&lt;p&gt;A page averaging position 11.5 has already earned enough visibility to sit near page one. That gives you a clear signal about where to focus.&lt;/p&gt;&lt;p&gt;The number creates a practical starting point. You have an existing page with a position you can work from, rather than an untested idea waiting for visibility.&lt;/p&gt;
&lt;h2&gt;why create another page when an existing one is already close?&lt;/h2&gt;
&lt;p&gt;The usual instinct is to create another page. That expansion reflex feels productive because it adds more coverage and more URLs.&lt;/p&gt;&lt;p&gt;A near-winner gives you a narrower decision. You can improve the guide around the same topic while the broader content footprint waits.&lt;/p&gt;
&lt;h2&gt;what should you improve on the near-winner?&lt;/h2&gt;
&lt;p&gt;Start with the page averaging position 11.5. Give it a focused guide, strengthen the path into it with an internal link, and give the page a clearer job.&lt;/p&gt;&lt;p&gt;The work stays close to the signal already in front of you. The page has a position, a topic, and a reason to receive focused attention.&lt;/p&gt;
&lt;h2&gt;how do you know which page deserves priority?&lt;/h2&gt;
&lt;p&gt;The useful question is simple: which existing page is close enough to improve? Pages averaging position 11.5 are the first place to look.&lt;/p&gt;&lt;p&gt;After the focused guide and internal link, watch what happens. The next page can wait while you learn from the near-winner.&lt;/p&gt;
&lt;h2&gt;FAQ&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;What does an average position of 11.5 mean?&lt;/strong&gt;&lt;br&gt;It means the page is already close to page one. That position gives you a concrete signal for prioritizing improvement.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Should I create a new page or improve an existing page at position 11.5?&lt;/strong&gt;&lt;br&gt;Improve the existing page first. Give the near-winner a focused guide and an internal link before creating another page.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What should I do with a page ranking around position 11.5?&lt;/strong&gt;&lt;br&gt;Improve the guide around the same topic, strengthen the internal path into it, and give the page a clearer job. Then watch what happens.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/the-dumb-seo-move-is-publishing-another-page-whi/</id>
    <title type="text">the dumb seo move is publishing another page while page 11.2 is</title>
    <updated>2026-09-13T00:59:00+09:00</updated>
    <published>2026-09-13T00:59:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/the-dumb-seo-move-is-publishing-another-page-whi/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">Why improving near-page-one pages can beat publishing new content from scratch.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; Before publishing another page, check which existing pages are already near page one. A small ranking improvement can be worth more than starting from scratch.&lt;/p&gt;
&lt;h2&gt;why publish another page when one is already near page one?&lt;/h2&gt;
&lt;p&gt;A page sitting at 11.2 is already close to page one. That makes it a stronger place to look before creating something new.&lt;/p&gt;&lt;p&gt;One push can move an existing page much closer to the result you want, while a new page starts with no ranking position at all.&lt;/p&gt;
&lt;h2&gt;which pages should i check first?&lt;/h2&gt;
&lt;p&gt;Start with pages that are already near page one. Page 11.2 is the clearest example because it is close enough for a small improvement to matter.&lt;/p&gt;&lt;p&gt;The point is to find the pages where a modest ranking gain could create more value than publishing another page.&lt;/p&gt;
&lt;h2&gt;are small ranking improvements really worth prioritising?&lt;/h2&gt;
&lt;p&gt;They can be. A small improvement on an existing near-page-one page may be worth more than starting from scratch with a new page.&lt;/p&gt;&lt;p&gt;That is the SEO move worth questioning: adding more pages while an existing page is already one push from page one.&lt;/p&gt;
&lt;h2&gt;FAQ&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Should I improve an existing page before publishing a new one?&lt;/strong&gt;&lt;br&gt;Check your near-page-one pages first. A page already ranking around 11.2 may offer more immediate value than a new page starting from scratch.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What does page 11.2 mean for SEO priorities?&lt;/strong&gt;&lt;br&gt;It means the page is close to page one. A small ranking improvement could be worth more than creating another page.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Why focus on pages near page one?&lt;/strong&gt;&lt;br&gt;Near-page-one pages already have ranking progress behind them. Improving one can be a more valuable move than publishing another page.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/pulled-a-page-sitting-at-position-11-5/</id>
    <title type="text">pulled a page sitting at position 11.5</title>
    <updated>2026-09-12T00:59:00+09:00</updated>
    <published>2026-09-12T00:59:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/pulled-a-page-sitting-at-position-11-5/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">A page at position 11.5 earned 62 impressions and 1 click. Moving it to page one could add roughly 2.1 monthly clicks.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; A page at position 11.5 received 62 impressions and 1 click. Reaching page one could add roughly 2.1 monthly clicks, showing why small SEO gains deserve measurement.&lt;/p&gt;
&lt;h2&gt;what can a page at position 11.5 already tell me?&lt;/h2&gt;
&lt;p&gt;The page received 62 impressions and 1 click. It already has visibility and measurable interest before reaching page one.&lt;/p&gt;&lt;p&gt;That makes it worth examining as an existing opportunity rather than treating every SEO gain as a new keyword hunt.&lt;/p&gt;
&lt;h2&gt;how many clicks could page-one visibility add?&lt;/h2&gt;
&lt;p&gt;Reaching page one could add roughly 2.1 monthly clicks for this page.&lt;/p&gt;&lt;p&gt;The estimate is small in isolation. The measurement makes the opportunity concrete.&lt;/p&gt;
&lt;h2&gt;why do small SEO gains look bigger after measurement?&lt;/h2&gt;
&lt;p&gt;A change in position can appear minor until impressions and clicks are attached to it. Here, 62 impressions produced 1 click at position 11.5, while page-one visibility could add roughly 2.1 monthly clicks.&lt;/p&gt;&lt;p&gt;The value comes from connecting ranking movement to expected traffic instead of judging the gain by position alone.&lt;/p&gt;
&lt;h2&gt;which pages should I check before chasing another keyword?&lt;/h2&gt;
&lt;p&gt;Start with pages already near page one. A page at position 11.5 has a visible path to additional clicks, supported by its existing impressions and click data.&lt;/p&gt;&lt;p&gt;That review can reveal small gains available in the current set of pages.&lt;/p&gt;
&lt;h2&gt;FAQ&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;How many impressions did the page receive?&lt;/strong&gt;&lt;br&gt;The page received 62 impressions. It generated 1 click while sitting at position 11.5.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;How many extra monthly clicks could page one add?&lt;/strong&gt;&lt;br&gt;Reaching page one could add roughly 2.1 monthly clicks. The figure is an estimate tied to this page&amp;apos;s current position and measured performance.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Should I target a new keyword or improve an existing page?&lt;/strong&gt;&lt;br&gt;Check pages already near page one before chasing another keyword. The page at position 11.5 showed a measurable opportunity through its existing impressions and click.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/why-chase-new-traffic-when-a-page-at-position-11/</id>
    <title type="text">why chase new traffic when a page at position 11 is already</title>
    <updated>2026-09-11T00:59:00+09:00</updated>
    <published>2026-09-11T00:59:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/why-chase-new-traffic-when-a-page-at-position-11/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">A page at position 11 may be a near-win: use a focused guide and internal link to turn existing visibility into clicks.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; A page at position 11 is already close to page one. A focused guide and an internal link can turn that visibility into clicks, making it worth prioritizing before chasing new traffic.&lt;/p&gt;
&lt;h2&gt;why prioritize a page at position 11?&lt;/h2&gt;
&lt;p&gt;A page at position 11 is already close to page one. That existing visibility makes it a near-win worth prioritizing before chasing entirely new traffic.&lt;/p&gt;
&lt;h2&gt;what can turn existing visibility into clicks?&lt;/h2&gt;
&lt;p&gt;A focused guide can give the page a clearer path toward useful traffic. An internal link can support that path by directing readers toward it.&lt;/p&gt;
&lt;h2&gt;which page should you push first?&lt;/h2&gt;
&lt;p&gt;Start with the page already sitting at position 11. It has visible momentum, so improving that opportunity may deserve attention before pursuing new traffic.&lt;/p&gt;
&lt;h2&gt;FAQ&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Is a page at position 11 worth prioritizing?&lt;/strong&gt;&lt;br&gt;Yes. A page at position 11 is already close to page one and may be a near-win worth prioritizing.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;How can a focused guide help a page near page one?&lt;/strong&gt;&lt;br&gt;A focused guide can turn existing visibility into clicks. The source post presents it as a practical priority for a page at position 11.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What role does internal linking play?&lt;/strong&gt;&lt;br&gt;An internal link can help turn existing visibility into clicks. It works alongside a focused guide for a page close to page one.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/stop-treating-seo-like-a-blank-page-problem/</id>
    <title type="text">stop treating SEO like a blank-page problem</title>
    <updated>2026-09-10T00:59:00+09:00</updated>
    <published>2026-09-10T00:59:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/stop-treating-seo-like-a-blank-page-problem/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">Why improving pages already near page one can drive more SEO growth than starting from zero.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; SEO growth often starts with pages already ranking near page one. A small ranking improvement can create more growth than building a page from scratch.&lt;/p&gt;
&lt;h2&gt;why start with pages already ranking near page one?&lt;/h2&gt;
&lt;p&gt;Pages already ranking near page one show existing search demand. They give you a place where focused effort can improve visibility.&lt;/p&gt;&lt;p&gt;Starting from zero requires creating that opportunity first. Prioritizing pages with nearby rankings puts effort where movement is already possible.&lt;/p&gt;
&lt;h2&gt;can a small ranking improvement really create more growth?&lt;/h2&gt;
&lt;p&gt;Yes. A small improvement for a page near page one can create more growth than starting with a page that has no existing position.&lt;/p&gt;&lt;p&gt;The gain comes from improving visibility where people are already searching and the page is already close to being found.&lt;/p&gt;
&lt;h2&gt;what should i look for before creating a new page?&lt;/h2&gt;
&lt;p&gt;Look for existing search demand and pages that already rank near page one. Those pages give you evidence about where attention may produce a useful improvement.&lt;/p&gt;&lt;p&gt;Once you find that demand, put your effort there before starting from a blank page.&lt;/p&gt;
&lt;h2&gt;FAQ&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Why prioritize pages ranking near page one?&lt;/strong&gt;&lt;br&gt;They already have search demand and are close to stronger visibility. A small ranking improvement can create more growth than starting from zero.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Is it better to improve an existing page or create a new one?&lt;/strong&gt;&lt;br&gt;Start by checking whether an existing page ranks near page one. If it does, improving that page may create more growth than creating a new page without existing demand.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;How do I find existing SEO demand?&lt;/strong&gt;&lt;br&gt;Look for pages that already rank near page one. Their rankings show where search demand already exists and where focused effort may help.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/a-page-at-average-position-10-8-is-one-push-from/</id>
    <title type="text">a page at average position 10.8 is one push from page one</title>
    <updated>2026-09-09T00:59:00+09:00</updated>
    <published>2026-09-09T00:59:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/a-page-at-average-position-10-8-is-one-push-from/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">A page at average position 10.8 already earned 280 impressions and 3 clicks, with roughly 11 more monthly clicks in reach.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; A page at average position 10.8 is one push from page one. With 280 impressions and 3 clicks already recorded, improving its ranking could bring roughly 11 more monthly clicks.&lt;/p&gt;
&lt;h2&gt;which pages are already close to page one?&lt;/h2&gt;
&lt;p&gt;Start with pages sitting near page one before creating more content. An average position of 10.8 puts a page one push away from that threshold.&lt;/p&gt;&lt;p&gt;That page already earned 280 impressions and 3 clicks. The next search win may already be published.&lt;/p&gt;
&lt;h2&gt;how many clicks could a small ranking improvement bring?&lt;/h2&gt;
&lt;p&gt;Improving the ranking of this page could bring roughly 11 more monthly clicks. The opportunity comes from an existing page with observed impressions and clicks.&lt;/p&gt;&lt;p&gt;That makes near-page-one pages worth checking when deciding where to focus the next push.&lt;/p&gt;
&lt;h2&gt;should i create another page or improve one that already ranks?&lt;/h2&gt;
&lt;p&gt;Check pages near page one before creating more content. A page at average position 10.8 already has visibility and click history to work from.&lt;/p&gt;&lt;p&gt;Low-key, the next win may already be published.&lt;/p&gt;
&lt;h2&gt;FAQ&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;What does an average position of 10.8 mean?&lt;/strong&gt;&lt;br&gt;It means the page is sitting near page one. The source describes it as one push from page one.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;How many clicks could the page gain?&lt;/strong&gt;&lt;br&gt;Improving its ranking could bring roughly 11 more monthly clicks. The page already earned 280 impressions and 3 clicks.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Should I improve an existing page before creating new content?&lt;/strong&gt;&lt;br&gt;Check pages sitting near page one first. An existing page may already be close to its next search win.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/i-keep-coming-back-to-those-clicks/</id>
    <title type="text">i keep coming back to those clicks</title>
    <updated>2026-09-08T00:59:00+09:00</updated>
    <published>2026-09-08T00:59:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/i-keep-coming-back-to-those-clicks/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">An about page earned 294 search impressions and 3 clicks at an average position of 10.5. Visibility translated into very little traffic.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; my about page appeared in search 294 times and earned 3 clicks, with an average position of 10.5. that position feels close to page one, while the clicks show how little visibility turned into traffic.&lt;/p&gt;
&lt;h2&gt;why does an average position of 10.5 feel like progress?&lt;/h2&gt;
&lt;p&gt;10.5 gives me somewhere to imagine the page sitting. i can almost picture the search result, close to page one and close enough to feel promising.&lt;/p&gt;&lt;p&gt;i find that number easy to get attached to. it makes the result feel within reach, even while the question i care about stays open: how often does someone actually choose to visit?&lt;/p&gt;
&lt;h2&gt;what do 294 impressions and 3 clicks actually tell me?&lt;/h2&gt;
&lt;p&gt;294 impressions tell me the about page earned visibility. 3 clicks tell me very little of that visibility became traffic. ngl, that&amp;#39;s a thin payoff for appearing in search that often.&lt;/p&gt;&lt;p&gt;i keep coming back to those clicks because someone has to choose to open the page. here, that happened 3 times. the count deserves its full weight when i decide how encouraging this result really is.&lt;/p&gt;
&lt;h2&gt;can i explain why so few people clicked?&lt;/h2&gt;
&lt;p&gt;the figures come with no search queries and no title or description to inspect. i can&amp;#39;t pin the small click count on either.&lt;/p&gt;&lt;p&gt;that leaves the cause open. i have to resist filling the gap with a convenient explanation just because the average position makes more traffic feel plausible.&lt;/p&gt;
&lt;h2&gt;how am i reading this result now?&lt;/h2&gt;
&lt;p&gt;i keep the 294 impressions beside the 3 clicks, then read the average position alongside them. that keeps the visibility attached to what it produced.&lt;/p&gt;&lt;p&gt;an about page can get near page one and still attract very little traffic. appearing in search earns a chance to be chosen. the click count tells me how often that chance became a visit.&lt;/p&gt;
&lt;h2&gt;FAQ&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Does an average search position of 10.5 mean a page is getting traffic?&lt;/strong&gt;&lt;br&gt;Position alone leaves that question open. This about page had an average position of 10.5 and received 3 clicks from 294 impressions.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Why does my about page get impressions but few clicks?&lt;/strong&gt;&lt;br&gt;These figures show that visibility turned into very little traffic. Without search queries or a title and description to inspect, the cause remains open.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;How should I read search impressions alongside clicks?&lt;/strong&gt;&lt;br&gt;Keep the counts together so you can see how often an appearance becomes a visit. Read average position alongside them, giving the click count its full weight.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/391-impressions/</id>
    <title type="text">391 impressions</title>
    <updated>2026-09-07T00:59:00+09:00</updated>
    <published>2026-09-07T00:59:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/391-impressions/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">Four pages earned 391 search impressions and 4 clicks. A field note on keeping clicks visible and avoiding fixes the totals can&apos;t justify.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; 391 impressions. 4 clicks across four pages in the measured window, with a combined click-through rate of about 1%. Kinda humbling when earning visits is the goal.&lt;/p&gt;
&lt;h2&gt;what did those 391 impressions actually get me?&lt;/h2&gt;
&lt;p&gt;Four pages appeared in search and earned 4 clicks in the measured window. Put those figures beside each other and “we’re appearing in search” feels like a much smaller claim.&lt;/p&gt;&lt;p&gt;The impressions tell me the pages appeared. The 4 clicks tell me how many clicks that appearance earned, which is the number I want in the headline when the goal is visits.&lt;/p&gt;
&lt;h2&gt;what does a combined click-through rate of about 1% tell me?&lt;/h2&gt;
&lt;p&gt;About 1% was the combined click-through rate across all four pages. “Combined” belongs beside the percentage because the source gives no page-level breakdown. Assigning that rate to every page would invent a result.&lt;/p&gt;&lt;p&gt;The measured window belongs in the claim too. Its duration isn&amp;#39;t specified, so these figures support no daily or monthly label.&lt;/p&gt;
&lt;h2&gt;should I rewrite the titles after seeing 4 clicks?&lt;/h2&gt;
&lt;p&gt;4 clicks gives me a reason to inspect each page before choosing a change. The total alone leaves me unable to tell which page deserves attention.&lt;/p&gt;&lt;p&gt;“Rewrite the titles” sounds useful. These figures don&amp;#39;t explain why people passed over the pages, so I can&amp;#39;t use them to justify that prescription. The next decision needs the page-level inspection.&lt;/p&gt;
&lt;h2&gt;how should I share this result without overselling it?&lt;/h2&gt;
&lt;p&gt;391 impressions and 4 clicks belong beside each other. Making someone dig for the clicks gives the bigger number too much airtime.&lt;/p&gt;&lt;p&gt;I&amp;#39;d lead with 4 clicks, then give the impressions and scope: four pages in the measured window, with a combined click-through rate of about 1%. That keeps the result attached to what was actually measured.&lt;/p&gt;
&lt;h2&gt;FAQ&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Is a 1% search click-through rate bad?&lt;/strong&gt;&lt;br&gt;These figures alone don&amp;apos;t support that verdict. About 1% is the combined result across four pages in an unspecified measured window, with no page-level breakdown.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Do search impressions mean my website is getting traffic?&lt;/strong&gt;&lt;br&gt;Impressions tell you the pages appeared in search. Here, 391 impressions earned 4 clicks, so the clicks need to stay visible when discussing visits.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Should I change page titles when search clicks are low?&lt;/strong&gt;&lt;br&gt;This result supports inspecting each page before choosing a change. The totals don&amp;apos;t explain why people passed over the pages or establish that titles caused the result.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/an-accelerator-can-be-useful-before-you-have-a-b/</id>
    <title type="text">an accelerator can be useful before you have a business</title>
    <updated>2026-09-05T00:59:00+09:00</updated>
    <published>2026-09-05T00:59:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/an-accelerator-can-be-useful-before-you-have-a-b/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">A field note on how an accelerator can create clarity through structured pressure before revenue, investors, or a clear idea.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; An accelerator can be useful before you have a business. Structured pressure may create clarity before traction exists, making uncertainty a starting line for disciplined exploration.&lt;/p&gt;
&lt;h2&gt;can an accelerator help before i have a business?&lt;/h2&gt;
&lt;p&gt;An accelerator can be useful before you have a business, even with zero revenue, zero investors, and no clear idea.&lt;/p&gt;&lt;p&gt;The value may come from creating a setting where uncertainty receives structure before traction exists.&lt;/p&gt;
&lt;h2&gt;what can structured pressure create when there is no traction?&lt;/h2&gt;
&lt;p&gt;Structured pressure can create clarity before traction exists. It gives uncertainty a direction for disciplined exploration.&lt;/p&gt;&lt;p&gt;That clarity can emerge before revenue, investors, or a fully formed idea are in place.&lt;/p&gt;
&lt;h2&gt;is uncertainty a valid starting line for building?&lt;/h2&gt;
&lt;p&gt;Uncertainty can be the starting line for disciplined exploration. A clear idea does not have to come first for exploration to begin.&lt;/p&gt;&lt;p&gt;The early conditions can be simple: zero revenue, zero investors, and a willingness to work through the uncertainty.&lt;/p&gt;
&lt;h2&gt;would i join an accelerator that early?&lt;/h2&gt;
&lt;p&gt;Joining that early means entering before the usual signs of traction appear. The question is whether structured pressure would help create clarity from uncertainty.&lt;/p&gt;&lt;p&gt;For a builder or founder still exploring, that possibility may make an accelerator useful before a business exists.&lt;/p&gt;
&lt;h2&gt;FAQ&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Can you join an accelerator before having a business idea?&lt;/strong&gt;&lt;br&gt;Yes. The source describes an accelerator as potentially useful with zero revenue, zero investors, and no clear idea. Structured pressure can create clarity before traction exists.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What is the value of an accelerator before traction?&lt;/strong&gt;&lt;br&gt;Its value may come from structured pressure that helps create clarity. That clarity can support disciplined exploration while uncertainty remains.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Is uncertainty a good starting point for building a company?&lt;/strong&gt;&lt;br&gt;Uncertainty can be the starting line for disciplined exploration. A clear idea and traction may develop through that exploration.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/staring-at-a-portfolio-filing-with-19-47-billion/</id>
    <title type="text">staring at a portfolio filing with $19.47 billion in holdings</title>
    <updated>2026-09-04T00:59:00+09:00</updated>
    <published>2026-09-04T00:59:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/staring-at-a-portfolio-filing-with-19-47-billion/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">A field note on turning a $19.47 billion portfolio filing into context about changes, concentration, and timing.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; A raw portfolio filing tells you what was held. The useful product is the context around those positions: what changed, how concentrated it is, and when the changes happened.&lt;/p&gt;
&lt;h2&gt;what does the raw portfolio filing actually give you?&lt;/h2&gt;
&lt;p&gt;The raw filing is authoritative. It gives you the underlying record of the portfolio and its $19.47 billion in holdings.&lt;/p&gt;&lt;p&gt;That record establishes what was held at the time of the filing. It is the foundation for any useful analysis of the positions.&lt;/p&gt;
&lt;h2&gt;why is the filing alone hard to use?&lt;/h2&gt;
&lt;p&gt;The filing shows positions, while the questions builders and founders care about sit around those positions. What changed? How concentrated is the portfolio? When did the changes happen?&lt;/p&gt;&lt;p&gt;Those questions require context across the holdings and across time. The raw data contains the filing, while the surrounding interpretation makes the movement easier to understand.&lt;/p&gt;
&lt;h2&gt;what does interpretation add to the portfolio data?&lt;/h2&gt;
&lt;p&gt;Interpretation turns the filing into a decision surface. It gives the positions meaning through change, concentration, and timing.&lt;/p&gt;&lt;p&gt;That context is where the useful product lives. The authoritative source remains the raw filing, and the interpretation helps people see what the positions may mean together.&lt;/p&gt;
&lt;h2&gt;FAQ&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;What is the authoritative source in a portfolio analysis?&lt;/strong&gt;&lt;br&gt;The raw portfolio filing is authoritative. It provides the underlying record of the holdings.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What context should you add around portfolio positions?&lt;/strong&gt;&lt;br&gt;Focus on what changed, how concentrated the portfolio is, and when the changes happened. Those questions add context to the raw positions.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What is a decision surface for portfolio data?&lt;/strong&gt;&lt;br&gt;A decision surface is the interpreted view around the positions. It connects the filing to changes, concentration, and timing.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/how-much-evidence-do-you-need-before-you-build/</id>
    <title type="text">how much evidence do you need before you build?</title>
    <updated>2026-09-03T00:59:00+09:00</updated>
    <published>2026-09-03T00:59:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/how-much-evidence-do-you-need-before-you-build/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">A field note on startup failure, evidence thresholds, validation, and disciplined execution across 235,039 launches.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; Across a dataset covering 235,039 launches, 52% of startups failed. That makes validation and disciplined execution operating requirements before the next build.&lt;/p&gt;
&lt;h2&gt;how much evidence do you need before you build?&lt;/h2&gt;
&lt;p&gt;Enough evidence to make the next build a disciplined decision. The available finding is stark: 52% of startups fail across a dataset covering 235,039 launches.&lt;/p&gt;&lt;p&gt;That number changes the question from whether to build toward what deserves validation first.&lt;/p&gt;
&lt;h2&gt;what does a 52% failure rate change about building?&lt;/h2&gt;
&lt;p&gt;It raises the cost of moving forward without evidence. A launch can consume effort while leaving the central assumption untested.&lt;/p&gt;&lt;p&gt;Validation becomes part of operating the company, alongside disciplined execution.&lt;/p&gt;
&lt;h2&gt;why do validation and disciplined execution matter together?&lt;/h2&gt;
&lt;p&gt;Validation helps determine what is worth building. Disciplined execution turns that evidence into focused action.&lt;/p&gt;&lt;p&gt;The finding points to both as operating requirements for builders and founders facing the next build decision.&lt;/p&gt;
&lt;h2&gt;what should you validate before the next build?&lt;/h2&gt;
&lt;p&gt;Start with the assumption that would most change the decision to build. The source finding does not prescribe a specific test, so the useful prompt is direct: what are you validating before the next build?&lt;/p&gt;&lt;p&gt;The answer should create evidence strong enough to support disciplined execution.&lt;/p&gt;
&lt;h2&gt;FAQ&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;How many startups fail according to the dataset?&lt;/strong&gt;&lt;br&gt;The source states that 52% of startups fail across a dataset covering 235,039 launches. It presents this as evidence for stronger validation and disciplined execution.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Why validate before building?&lt;/strong&gt;&lt;br&gt;The finding shows a 52% failure rate across the dataset. Validation helps inform what deserves the next build before effort is committed.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What should founders validate before the next build?&lt;/strong&gt;&lt;br&gt;Founders should identify the assumption that matters most to the build decision. The source frames validation as an operating requirement.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/distribution-comes-after-proof/</id>
    <title type="text">distribution comes after proof</title>
    <updated>2026-09-02T00:59:00+09:00</updated>
    <published>2026-09-02T00:59:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/distribution-comes-after-proof/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">Why education and AI companies need published student outcomes before scaling distribution.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; Distribution comes after proof. Student usage shows activity, while published outcomes provide the evidence that helps keep tutoring, education, and AI companies alive.&lt;/p&gt;
&lt;h2&gt;why should distribution come after proof?&lt;/h2&gt;
&lt;p&gt;A tutoring company can shut down when its learning claims lack evidence. Distribution increases attention, reach, and expectations, so the claims behind the product need support before scaling.&lt;/p&gt;&lt;p&gt;The field note is simple: measure student outcomes first, then publish what the measurements show before expanding distribution.&lt;/p&gt;
&lt;h2&gt;what evidence keeps a tutoring company alive?&lt;/h2&gt;
&lt;p&gt;Student outcomes are the evidence. They show whether the product is producing learning results that support its claims.&lt;/p&gt;&lt;p&gt;Publishing those outcomes gives builders, founders, and customers something concrete to evaluate before the company grows its reach.&lt;/p&gt;
&lt;h2&gt;why is usage alone insufficient?&lt;/h2&gt;
&lt;p&gt;Usage is activity. A student using a product shows engagement with the product, while outcomes show the learning result.&lt;/p&gt;&lt;p&gt;Education and AI companies need outcomes to support their claims and sustain the business as distribution expands.&lt;/p&gt;
&lt;h2&gt;FAQ&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Why should tutoring companies measure student outcomes?&lt;/strong&gt;&lt;br&gt;A tutoring company can shut down when its learning claims lack evidence. Measuring student outcomes creates proof for those claims before the company scales distribution.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Is student usage proof that a learning product works?&lt;/strong&gt;&lt;br&gt;Usage shows activity. Student outcomes provide evidence of learning, which is the proof education and AI companies need.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;When should an education company scale distribution?&lt;/strong&gt;&lt;br&gt;An education company should measure and publish student outcomes before scaling distribution. The outcomes provide evidence for the product&amp;apos;s learning claims.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/student-input-changed-how-ai-policies-were-desig/</id>
    <title type="text">student input changed how AI policies were designed</title>
    <updated>2026-09-01T00:59:00+09:00</updated>
    <published>2026-09-01T00:59:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/student-input-changed-how-ai-policies-were-desig/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">Students helped shape AI policies with meaningful restrictions, showing that user input improves governance and safety design.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; Students produced AI policies with meaningful restrictions because the restrictions came from user input. User control worked better when designed with the people affected by it.&lt;/p&gt;
&lt;h2&gt;what changed when students helped design the policy?&lt;/h2&gt;
&lt;p&gt;Students produced policies with meaningful restrictions. Their input shaped the rules directly.&lt;/p&gt;&lt;p&gt;That result points to a practical lesson for AI policy work: the people who use a system can contribute restrictions that fit their experience of it.&lt;/p&gt;
&lt;h2&gt;why did user input lead to better user control?&lt;/h2&gt;
&lt;p&gt;User control worked better when it was designed with users. The people affected by the policy had a role in defining how it should work.&lt;/p&gt;&lt;p&gt;That makes user input part of the design itself, rather than a reaction after the rules are already set.&lt;/p&gt;
&lt;h2&gt;where should AI governance and safety settings start?&lt;/h2&gt;
&lt;p&gt;Governance, safety settings, and acceptable-use design should start with the people who have to live with the rules.&lt;/p&gt;&lt;p&gt;The transferable finding is simple: build the rules with those people. Their participation can produce meaningful restrictions and stronger user control.&lt;/p&gt;
&lt;h2&gt;FAQ&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;How should students contribute to AI policy design?&lt;/strong&gt;&lt;br&gt;Students should provide input while the policy is being designed. Their input can produce policies with meaningful restrictions.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Why should users help design AI safety settings?&lt;/strong&gt;&lt;br&gt;User control worked better when designed with users. The people affected by the settings can help shape rules they have to live with.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Who should shape acceptable-use rules for AI systems?&lt;/strong&gt;&lt;br&gt;The people who have to live with the rules should help build them. That includes users whose input can guide meaningful restrictions.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/spent-an-afternoon-unpacking-one-ugly-failure-mo/</id>
    <title type="text">spent an afternoon unpacking one ugly failure mode: a virtual</title>
    <updated>2026-08-31T00:59:00+09:00</updated>
    <published>2026-08-31T00:59:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/spent-an-afternoon-unpacking-one-ugly-failure-mo/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">Why AI education products need learning outcomes alongside engagement metrics to prove that learning happened.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; Engagement shows that people used an AI education product. Learning outcomes show whether learning happened, which can determine whether a virtual tutoring provider survives.&lt;/p&gt;
&lt;h2&gt;why can a virtual tutoring provider shut down even when people use the product?&lt;/h2&gt;
&lt;p&gt;A virtual tutoring provider can shut down when it cannot demonstrate learning effectiveness. Usage alone shows engagement, while the business also needs evidence that learning happened.&lt;/p&gt;&lt;p&gt;That distinction changes what the product must measure. Activity can show that people participated, yet it cannot establish learning progress by itself.&lt;/p&gt;
&lt;h2&gt;what does engagement data actually tell you?&lt;/h2&gt;
&lt;p&gt;Engagement data tells you that people used the product. It measures activity around the learning experience.&lt;/p&gt;&lt;p&gt;That signal matters, although it answers a narrower question than whether the product worked. Usage can exist without demonstrated learning effectiveness.&lt;/p&gt;
&lt;h2&gt;which metric should sit beside every usage metric in an AI education product?&lt;/h2&gt;
&lt;p&gt;Put one learning outcome beside every usage metric. This keeps product measurement connected to the result the education product needs to demonstrate.&lt;/p&gt;&lt;p&gt;For AI education products, the practical lesson is direct: measure activity alongside evidence of learning. Otherwise, it is easy to measure activity and call it progress.&lt;/p&gt;
&lt;h2&gt;FAQ&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;What is the difference between engagement data and outcome data in education?&lt;/strong&gt;&lt;br&gt;Engagement data shows that people used the product. Outcome data shows whether learning happened.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Why are usage metrics insufficient for AI education products?&lt;/strong&gt;&lt;br&gt;Usage metrics measure activity around the product. They do not demonstrate learning effectiveness on their own.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What should an AI tutoring product measure alongside engagement?&lt;/strong&gt;&lt;br&gt;It should put one learning outcome beside every usage metric. This connects product activity to evidence of learning.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/the-best-ai-product-might-be-the-one-teachers-ba/</id>
    <title type="text">the best AI product might be the one teachers barely notice</title>
    <updated>2026-08-30T00:59:00+09:00</updated>
    <published>2026-08-30T00:59:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/the-best-ai-product-might-be-the-one-teachers-ba/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">AI products gain adoption when they fit the workflow already in the room, especially for teachers.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; The best AI product may be the one teachers barely notice. When new capabilities fit the daily process, adoption gets easier because users can keep their day intact.&lt;/p&gt;
&lt;h2&gt;why does fitting the existing workflow make AI easier to adopt?&lt;/h2&gt;
&lt;p&gt;People already have a daily process. When new capabilities fit that process, adoption gets easier because using AI feels connected to work already happening.&lt;/p&gt;&lt;p&gt;The product becomes part of the room&amp;#39;s workflow instead of asking everyone to create a new one.&lt;/p&gt;
&lt;h2&gt;how much should teachers have to change their day to use AI?&lt;/h2&gt;
&lt;p&gt;Teachers should be able to use AI within the process they already follow. Rebuilding the day creates friction before the product has a chance to help.&lt;/p&gt;&lt;p&gt;The strongest experience may be the one that supports the daily process quietly.&lt;/p&gt;
&lt;h2&gt;why does the workflow become the onboarding problem?&lt;/h2&gt;
&lt;p&gt;A product can have useful capabilities and still create a difficult first experience when the workflow around it feels unfamiliar.&lt;/p&gt;&lt;p&gt;The onboarding challenge often begins before the first feature is used. It begins with the amount of change users feel they must absorb.&lt;/p&gt;
&lt;h2&gt;would users describe this as a new tool or a better Tuesday?&lt;/h2&gt;
&lt;p&gt;That question gets close to the real product test. If users experience AI as part of a better Tuesday, the value is attached to their existing work.&lt;/p&gt;&lt;p&gt;For builders, the goal is to build around the workflow already in the room and let the new capability fit into it.&lt;/p&gt;
&lt;h2&gt;FAQ&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;How can AI products improve teacher adoption?&lt;/strong&gt;&lt;br&gt;Build around the workflow teachers already use. Adoption gets easier when new capabilities fit the daily process and teachers can keep their day intact.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Why does workflow matter more than adding new AI features?&lt;/strong&gt;&lt;br&gt;A workflow determines how much change users feel when they try a product. Useful capabilities become easier to adopt when they fit the process already happening.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What makes an AI product feel easy to use?&lt;/strong&gt;&lt;br&gt;It can feel easy when users experience it as part of a better Tuesday. Teachers should not have to rebuild their day to use AI.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/i-price-around-outcomes-like-this/</id>
    <title type="text">i price around outcomes like this</title>
    <updated>2026-08-29T00:59:00+09:00</updated>
    <published>2026-08-29T00:59:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/i-price-around-outcomes-like-this/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">A field note on outcome-based pricing: define the result, record the baseline, track one measurement, and price the measured change.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; Outcome-based pricing depends on one stable measurement. Change the definition halfway through, and the pricing conversation loses its foundation.&lt;/p&gt;
&lt;h2&gt;what should be defined before the work starts?&lt;/h2&gt;
&lt;p&gt;Write the outcome in one precise sentence. That sentence gives the work a clear result to measure.&lt;/p&gt;&lt;p&gt;Record the baseline before anything changes. Without it, the eventual result has no fixed point.&lt;/p&gt;
&lt;h2&gt;how do you know whether the work created the promised change?&lt;/h2&gt;
&lt;p&gt;Track the same definition from baseline to result. The measurement needs to remain consistent throughout the work.&lt;/p&gt;&lt;p&gt;That continuity connects the starting point to the observed result and makes the change visible.&lt;/p&gt;
&lt;h2&gt;when should the price be tied to the outcome?&lt;/h2&gt;
&lt;p&gt;Tie the price to the measured change after the result has been tracked against the original baseline. The price follows the evidence.&lt;/p&gt;&lt;p&gt;This approach makes the pricing conversation about the outcome that was defined and measured.&lt;/p&gt;
&lt;h2&gt;why does changing the measurement halfway through cause trouble?&lt;/h2&gt;
&lt;p&gt;Changing the measurement halfway through breaks the connection between the baseline and the result. The conversation then rests on two different definitions.&lt;/p&gt;&lt;p&gt;Once that happens, the pricing conversation is already cooked.&lt;/p&gt;
&lt;h2&gt;FAQ&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;How should outcome-based pricing be defined?&lt;/strong&gt;&lt;br&gt;Write the outcome in one precise sentence. Record the baseline before the work starts, track that same definition through the result, and tie the price to the measured change.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Why is a baseline important for outcome-based pricing?&lt;/strong&gt;&lt;br&gt;The baseline establishes the starting point before the work begins. It gives the final result a fixed reference for measuring change.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What happens if the success metric changes during the work?&lt;/strong&gt;&lt;br&gt;The baseline and result no longer use the same definition. That weakens the evidence for the measured change and damages the pricing conversation.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/virtual-tutoring-companies-need-measured-learnin/</id>
    <title type="text">virtual tutoring companies need measured learning outcomes</title>
    <updated>2026-08-28T00:59:00+09:00</updated>
    <published>2026-08-28T00:59:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/virtual-tutoring-companies-need-measured-learnin/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">Why virtual tutoring companies need measured learning outcomes before scaling, and how evidence should shape product growth.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; Virtual tutoring companies should measure whether students actually improve before scaling. Product claims become stronger when founders turn them into measurable proof and build around the evidence.&lt;/p&gt;
&lt;h2&gt;what should virtual tutoring companies prove before they scale?&lt;/h2&gt;
&lt;p&gt;They should prove that students actually improve. Measured learning outcomes give founders evidence that supports adoption and growth.&lt;/p&gt;&lt;p&gt;That evidence turns a product claim into measurable proof before scaling begins.&lt;/p&gt;
&lt;h2&gt;why do learning outcomes affect adoption and growth?&lt;/h2&gt;
&lt;p&gt;Adoption and growth follow evidence that students improve. When the outcome is measured, founders can connect the product to a result people can evaluate.&lt;/p&gt;&lt;p&gt;The strength of the claim comes from what the evidence shows.&lt;/p&gt;
&lt;h2&gt;what should founders build around once they have evidence?&lt;/h2&gt;
&lt;p&gt;Founders should build around what the evidence says. Measurement gives them a basis for deciding which product claims support growth and which outcomes deserve more attention.&lt;/p&gt;&lt;p&gt;The first question is simple: what would you measure first?&lt;/p&gt;
&lt;h2&gt;FAQ&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Why should virtual tutoring companies measure learning outcomes?&lt;/strong&gt;&lt;br&gt;Measured learning outcomes show whether students actually improve. That evidence supports product claims and can guide growth.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;How can founders prepare a tutoring product for growth?&lt;/strong&gt;&lt;br&gt;Founders can turn product claims into measurable proof before scaling. They can then build around what the evidence says.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What should a virtual tutoring company measure first?&lt;/strong&gt;&lt;br&gt;The source leaves that question open. The starting point is a learning outcome that shows whether students actually improve.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/shut-down-after-experts-questioned-the-evidence-/</id>
    <title type="text">shut down after experts questioned the evidence of effectiveness</title>
    <updated>2026-08-27T00:59:00+09:00</updated>
    <published>2026-08-27T00:59:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/shut-down-after-experts-questioned-the-evidence-/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">Why tutoring companies should prove customer outcomes with data before scaling distribution, according to a costly shutdown.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; A major virtual tutoring provider shut down after experts questioned its evidence of effectiveness. The field lesson is direct: measure customer outcomes before scaling distribution.&lt;/p&gt;
&lt;h2&gt;why did the virtual tutoring provider shut down?&lt;/h2&gt;
&lt;p&gt;The provider shut down after experts questioned the evidence behind its effectiveness. The failure centered on whether the customer outcome could be supported with data.&lt;/p&gt;&lt;p&gt;That makes evidence a business requirement. Users, institutions, and investors each need a reason they can defend.&lt;/p&gt;
&lt;h2&gt;what should a company prove before scaling distribution?&lt;/h2&gt;
&lt;p&gt;Before scaling distribution, prove the customer outcome with data. Distribution can increase reach, attention, and scrutiny at the same time.&lt;/p&gt;&lt;p&gt;Growth amplifies weak evidence without strengthening it. Scale works best when the underlying outcome is already measurable and defensible.&lt;/p&gt;
&lt;h2&gt;who needs evidence that the product works?&lt;/h2&gt;
&lt;p&gt;Users need a reason to believe the product delivers its promised outcome. Institutions need a reason to support or adopt it. Investors need a reason they can defend.&lt;/p&gt;&lt;p&gt;The same evidence can serve all three groups when it clearly connects the product to the customer outcome.&lt;/p&gt;
&lt;h2&gt;what is the practical lesson for founders and builders?&lt;/h2&gt;
&lt;p&gt;Measure the outcome first. Then scale what the data can justify.&lt;/p&gt;&lt;p&gt;The expensive failure came from allowing distribution to move ahead of proof. Evidence should lead growth, giving every stakeholder a defensible basis for confidence.&lt;/p&gt;
&lt;h2&gt;FAQ&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Why should a tutoring company measure outcomes before growing?&lt;/strong&gt;&lt;br&gt;A company needs data showing that its product creates the customer outcome it claims to deliver. Growth increases the audience for weak evidence and brings more scrutiny from users, institutions, and investors.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What evidence do investors need from an education company?&lt;/strong&gt;&lt;br&gt;Investors need a reason they can defend for believing the product works. Data on the customer outcome provides that basis.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Can growth strengthen weak product evidence?&lt;/strong&gt;&lt;br&gt;Growth can amplify weak evidence, giving it greater visibility. It does not strengthen the underlying proof, so the outcome should be measured before distribution scales.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/can-an-ai-startup-really-ship-in-eight-weeks/</id>
    <title type="text">can an AI startup really ship in eight weeks?</title>
    <updated>2026-08-26T00:59:00+09:00</updated>
    <published>2026-08-26T00:59:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/can-an-ai-startup-really-ship-in-eight-weeks/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">A field note on shipping an AI startup in eight weeks through narrow scope and intensely focused execution.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; An AI startup can ship in eight weeks when the scope stays narrow and execution remains intensely focused. Vague roadmaps make simple products feel endless.&lt;/p&gt;
&lt;h2&gt;can an AI startup really ship in eight weeks?&lt;/h2&gt;
&lt;p&gt;yes, when the scope is narrow and execution is intensely focused. The timeline becomes real when the build has a clear boundary.&lt;/p&gt;&lt;p&gt;The finding is simple: focused execution can turn an eight-week build into something shippable, while an unclear scope keeps the work expanding.&lt;/p&gt;
&lt;h2&gt;why do simple products start to feel endless?&lt;/h2&gt;
&lt;p&gt;Vague roadmaps make simple products feel endless. Without a narrow definition of the build, every task can create another question about what belongs in the product.&lt;/p&gt;&lt;p&gt;The product may remain simple while the roadmap grows unclear. That gap is where the sense of endless work comes from.&lt;/p&gt;
&lt;h2&gt;what would you cut to make an eight-week build real?&lt;/h2&gt;
&lt;p&gt;The useful question is what can leave the scope so the remaining product fits the eight-week window. Cutting scope creates room for intensely focused execution.&lt;/p&gt;&lt;p&gt;An eight-week build needs a narrow target and a roadmap that supports it. The cuts are what make the timeline real.&lt;/p&gt;
&lt;h2&gt;FAQ&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Can an AI startup ship a product in eight weeks?&lt;/strong&gt;&lt;br&gt;Yes, when the scope is narrow and execution is intensely focused. The timeline depends on keeping the build focused.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Why do simple software products take so long to build?&lt;/strong&gt;&lt;br&gt;Vague roadmaps make simple products feel endless. Unclear scope allows the work to keep expanding.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;How do you make an eight-week build realistic?&lt;/strong&gt;&lt;br&gt;Cut the scope until the remaining product has a narrow target. Then keep execution intensely focused.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/the-filing-is-the-raw-material/</id>
    <title type="text">the filing is the raw material</title>
    <updated>2026-08-25T00:59:00+09:00</updated>
    <published>2026-08-25T00:59:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/the-filing-is-the-raw-material/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">Why change detection turns filings into actionable insight by surfacing new positions, ownership changes, and portfolio shifts.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; The filing is raw material. The product is change detection: identifying new institutional positions, ownership changes, and major portfolio shifts so public data becomes actionable.&lt;/p&gt;
&lt;h2&gt;what makes a filing useful to a builder or investor?&lt;/h2&gt;
&lt;p&gt;A filing becomes useful when it reveals what changed. The document supplies the raw material, while the meaningful signal comes from identifying movement inside it.&lt;/p&gt;&lt;p&gt;That movement can show new institutional positions, changes in ownership, or major shifts in a portfolio.&lt;/p&gt;
&lt;h2&gt;why focus on change instead of the filing itself?&lt;/h2&gt;
&lt;p&gt;A static filing contains information. Change detection gives that information direction by showing how positions and ownership have moved.&lt;/p&gt;&lt;p&gt;The change is what helps someone interpret public data and decide whether it deserves attention.&lt;/p&gt;
&lt;h2&gt;what turns public data into something actionable?&lt;/h2&gt;
&lt;p&gt;Actionability comes from surfacing the changes that matter: new institutional positions, ownership changes, and major portfolio shifts.&lt;/p&gt;&lt;p&gt;That is the product layer built on top of the filing. The filing is the raw material, and the detected change is the usable signal.&lt;/p&gt;
&lt;h2&gt;FAQ&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;What is change detection in financial filings?&lt;/strong&gt;&lt;br&gt;Change detection identifies meaningful differences in filings over time. It can surface new institutional positions, ownership changes, and major portfolio shifts.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Why are filings considered raw material?&lt;/strong&gt;&lt;br&gt;Filings contain public information that still needs interpretation. Change detection turns the underlying data into a clearer view of what moved.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What makes public filing data actionable?&lt;/strong&gt;&lt;br&gt;The actionable part is the detected change. New positions, ownership changes, and major portfolio shifts give the data a direction that can inform decisions.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/students-given-authority-over-ai-policy-chose-me/</id>
    <title type="text">students given authority over AI policy chose meaningful</title>
    <updated>2026-08-24T00:59:00+09:00</updated>
    <published>2026-08-24T00:59:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/students-given-authority-over-ai-policy-chose-me/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">Students with a voice in AI policy chose meaningful restrictions, showing how user participation can shape safety and adoption.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; Students given authority over AI policy chose meaningful restrictions. Trusted AI products grow when users help shape the boundaries they will live with.&lt;/p&gt;
&lt;h2&gt;why should users have a voice before AI rules are fixed?&lt;/h2&gt;
&lt;p&gt;Students given authority over AI policy chose meaningful restrictions. Giving users a voice early lets them participate before the boundaries are settled.&lt;/p&gt;&lt;p&gt;That participation makes policy part of the product relationship from the beginning. Users are present while the limits are being shaped.&lt;/p&gt;
&lt;h2&gt;what changes when users help shape the boundaries?&lt;/h2&gt;
&lt;p&gt;Users have to live with the boundaries AI products set. Letting them shape those boundaries gives them a role in defining the limits they will encounter.&lt;/p&gt;&lt;p&gt;The result can be meaningful restrictions that reflect what users accept as appropriate. Authority creates room for users to influence safety design directly.&lt;/p&gt;
&lt;h2&gt;why does safety design belong inside product design?&lt;/h2&gt;
&lt;p&gt;Safety design determines the limits people experience while using a product. Treating it as part of product design keeps those limits connected to the people affected by them.&lt;/p&gt;&lt;p&gt;This makes safety a shared product decision. The design includes both what the system can do and the boundaries users accept.&lt;/p&gt;
&lt;h2&gt;how do accepted limits affect AI adoption?&lt;/h2&gt;
&lt;p&gt;Adoption grows around limits people accept. Users are more likely to build trust around AI products when they have had a voice in defining those limits.&lt;/p&gt;&lt;p&gt;Trusted AI products are built with users in the room. Their participation helps shape the restrictions that become part of the product.&lt;/p&gt;
&lt;h2&gt;FAQ&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Why should users help create AI policy?&lt;/strong&gt;&lt;br&gt;Users live with the boundaries AI products set. Giving them a voice before the rules are fixed lets them help shape meaningful restrictions.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What did students choose when given authority over AI policy?&lt;/strong&gt;&lt;br&gt;They chose meaningful restrictions. Their participation shows that users can contribute directly to safety boundaries.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;How does user involvement affect AI adoption?&lt;/strong&gt;&lt;br&gt;Adoption grows around limits people accept. User involvement helps create boundaries that are shaped with the people who will live with them.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What makes an AI product trusted?&lt;/strong&gt;&lt;br&gt;Trusted AI products are built with users in the room. Their input helps connect safety design with product design.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/spent-an-afternoon-thinking-about-why-some-ai-to/</id>
    <title type="text">spent an afternoon thinking about why some AI tools actually get</title>
    <updated>2026-08-23T00:59:00+09:00</updated>
    <published>2026-08-23T00:59:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/spent-an-afternoon-thinking-about-why-some-ai-to/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">Why AI tools get used in schools: they fit existing teacher workflows, and workflow friction determines adoption.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; AI tools get used in schools when they fit the teacher’s existing workflow. For builders, the daily handoff may matter more than another polished feature.&lt;/p&gt;
&lt;h2&gt;why do some AI tools actually get used in schools?&lt;/h2&gt;
&lt;p&gt;They fit the teacher’s existing workflow. The tool becomes easier to use when it matches how teachers already move through their work.&lt;/p&gt;&lt;p&gt;That fit can determine whether the tool becomes part of the day or remains something people try once.&lt;/p&gt;
&lt;h2&gt;what should builders watch before polishing another feature?&lt;/h2&gt;
&lt;p&gt;Watch the daily handoff. The handoff is where the tool meets the teacher’s existing workflow, so it reveals where use feels natural and where friction appears.&lt;/p&gt;&lt;p&gt;That observation can show more about adoption than another feature added in isolation.&lt;/p&gt;
&lt;h2&gt;where does school-tool adoption actually get decided?&lt;/h2&gt;
&lt;p&gt;Adoption gets decided in workflow friction. A tool can have a strong idea behind it, yet daily use depends on how well it fits the teacher’s routine.&lt;/p&gt;&lt;p&gt;For anyone building software for schools, the practical question is how the product fits into that existing handoff.&lt;/p&gt;
&lt;h2&gt;FAQ&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;why do some AI tools fail to gain adoption in schools?&lt;/strong&gt;&lt;br&gt;They may create friction in the teacher’s existing workflow. School adoption depends on how naturally the tool fits into the daily handoff.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;what should founders observe when building AI tools for schools?&lt;/strong&gt;&lt;br&gt;Observe the teacher’s daily handoff. It shows where the product fits the existing workflow and where workflow friction may affect adoption.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;does feature polish determine school adoption?&lt;/strong&gt;&lt;br&gt;Feature polish is only one part of the picture. The source finding points to workflow fit and daily friction as the deciding factors.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/the-best-ai-feature-in-a-school-might-be-fitting/</id>
    <title type="text">the best AI feature in a school might be fitting into Tuesday</title>
    <updated>2026-08-22T00:59:00+09:00</updated>
    <published>2026-08-22T00:59:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/the-best-ai-feature-in-a-school-might-be-fitting/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">Why AI features in schools get used when they fit educators’ existing Tuesday workflow.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; The best AI feature in a school may be the one that fits into Tuesday. Start with the educator already carrying the work, then design around that workflow.&lt;/p&gt;
&lt;h2&gt;why does workflow fit matter more than feature novelty?&lt;/h2&gt;
&lt;p&gt;Feature novelty gets attention because it feels new. Workflow fit gets used because it belongs inside work the educator is already carrying.&lt;/p&gt;&lt;p&gt;For a school, the practical question is whether the feature fits into Tuesday. Usage follows the existing workflow.&lt;/p&gt;
&lt;h2&gt;how should you design an AI feature for a school?&lt;/h2&gt;
&lt;p&gt;Start with the educator already carrying the work. Their existing workflow gives you the context for where the feature can fit.&lt;/p&gt;&lt;p&gt;Design around that workflow before adding novelty. The closer the feature is to work already happening, the more naturally it can be used.&lt;/p&gt;
&lt;h2&gt;who already owns the work the feature is meant to support?&lt;/h2&gt;
&lt;p&gt;That question should come before the feature itself. Find the educator who already carries the work, then understand the workflow around it.&lt;/p&gt;&lt;p&gt;The answer shapes the product. AI becomes useful when it supports work someone already owns.&lt;/p&gt;
&lt;h2&gt;FAQ&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;What makes an AI feature useful in a school?&lt;/strong&gt;&lt;br&gt;It fits into the educator’s existing workflow. A feature that belongs in the work already happening is more likely to be used.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Should schools prioritize novel AI features?&lt;/strong&gt;&lt;br&gt;Novelty can get attention. Workflow fit determines whether the feature gets used.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What question should you ask before building an AI feature for educators?&lt;/strong&gt;&lt;br&gt;Ask who already owns the work. Start with that educator and design around the workflow they already carry.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/i-was-halfway-through-another-fundraising-thread/</id>
    <title type="text">i was halfway through another fundraising thread</title>
    <updated>2026-08-21T00:59:00+09:00</updated>
    <published>2026-08-21T00:59:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/i-was-halfway-through-another-fundraising-thread/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">Fundraising advice only works when its stage, business model, constraints, and economics match the company using it.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; Fundraising advice can sound useful and still lead you in the wrong direction when it comes from a different stage, business model, or set of constraints. The decision belongs to data that matches the company in front of you.&lt;/p&gt;
&lt;h2&gt;why did familiar fundraising advice fall apart when i applied it?&lt;/h2&gt;
&lt;p&gt;I was halfway through another fundraising thread. The advice sounded familiar: raise when the market is open, tell a bigger story, move quickly.&lt;/p&gt;&lt;p&gt;Then I tried to apply it to the company I was actually building. The whole thing fell apart because the advice had been written for a different stage, a different business model, and different constraints.&lt;/p&gt;
&lt;h2&gt;why can the same fundraising advice lead to a different decision?&lt;/h2&gt;
&lt;p&gt;The same words can describe different decisions. Your stage changes the question, while your business model changes the answer. Your constraints change what is possible.&lt;/p&gt;&lt;p&gt;That gap matters more than people admit. Copying a strategy from a company with different economics can send you in the wrong direction very quickly.&lt;/p&gt;
&lt;h2&gt;why does generic advice feel useful even when it lacks context?&lt;/h2&gt;
&lt;p&gt;Generic advice arrives already packaged. It gives you a sentence to repeat, and sometimes it gives you confidence.&lt;/p&gt;&lt;p&gt;That feeling can make the advice seem more useful than it is. The decision still belongs to your data, and the relevant data comes from the business you are actually building.&lt;/p&gt;
&lt;h2&gt;how should builders evaluate fundraising advice?&lt;/h2&gt;
&lt;p&gt;I have started treating fundraising advice like a dataset. I ask where it came from, what kind of company produced it, what stage they were in, what their constraints were, and what outcome they were optimizing for.&lt;/p&gt;&lt;p&gt;Without that context, the advice is mostly decoration. Builders need capital decisions that match their stage, business model, and constraints.&lt;/p&gt;
&lt;h2&gt;FAQ&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Why does fundraising advice fail for some companies?&lt;/strong&gt;&lt;br&gt;Fundraising advice can fail when it comes from a different stage, business model, set of constraints, or economics. The same recommendation can lead to a different decision in a different company.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;How can I tell whether fundraising advice applies to my company?&lt;/strong&gt;&lt;br&gt;Trace where the advice came from and examine the company that produced it. Compare its stage, business model, constraints, economics, and desired outcome with your own.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Should founders follow advice to raise when the market is open?&lt;/strong&gt;&lt;br&gt;That advice may apply to some companies. The decision should come from data that matches your stage, business model, and constraints.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/validation-is-where-startups-live-or-die/</id>
    <title type="text">validation is where startups live or die</title>
    <updated>2026-08-21T00:58:00+09:00</updated>
    <published>2026-08-21T00:58:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/validation-is-where-startups-live-or-die/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">A field note on why validation must be a disciplined operating process before faster execution compounds risk.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; 52% of startups fail across 235,039 launches, and execution can look impressive while direction remains untested. Validation gives builders a process for testing belief before committing more effort.&lt;/p&gt;
&lt;h2&gt;why should a 52% failure rate change how builders operate?&lt;/h2&gt;
&lt;p&gt;52% of startups fail across 235,039 launches. That is a hard reminder that motion alone proves very little.&lt;/p&gt;&lt;p&gt;Shipping features, polishing pages, and expanding plans can create visible progress while the core direction remains untested.&lt;/p&gt;
&lt;h2&gt;how can execution look impressive while the idea stays untested?&lt;/h2&gt;
&lt;p&gt;Execution is easy to see. Features ship, pages get polished, and plans expand. Those milestones can feel productive while moving the work deeper into an idea nobody needs.&lt;/p&gt;&lt;p&gt;The uncomfortable part is that speed can increase commitment before the direction has earned belief.&lt;/p&gt;
&lt;h2&gt;what does disciplined validation actually require?&lt;/h2&gt;
&lt;p&gt;Validation asks whether the thing is earning belief before more effort goes in. It deserves a process because hopeful milestones leave room for wishful thinking.&lt;/p&gt;&lt;p&gt;Write down what has to be true, test it, and let the result change the plan. The risk attached to execution without validation is measurable too.&lt;/p&gt;
&lt;h2&gt;how should validation become an operating requirement?&lt;/h2&gt;
&lt;p&gt;Treat validation as part of operating the startup. The failure rate makes the consequence clear, while a process gives the team something to act on.&lt;/p&gt;&lt;p&gt;Hope is a feeling. A process creates a basis for deciding whether the next effort belongs in the plan.&lt;/p&gt;
&lt;h2&gt;FAQ&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Why is validation important for startups?&lt;/strong&gt;&lt;br&gt;Validation tests whether an idea is earning belief before more effort goes in. It helps builders examine direction while the work is still changeable.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Can shipping faster increase startup risk?&lt;/strong&gt;&lt;br&gt;Yes. Faster execution can move a team deeper into an idea nobody needs when the direction remains untested.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What does a startup validation process include?&lt;/strong&gt;&lt;br&gt;Write down what has to be true, test it, and let the result change the plan. This turns validation into a disciplined process instead of a hopeful milestone.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/workflow-fit-decides-adoption/</id>
    <title type="text">workflow fit decides adoption</title>
    <updated>2026-08-20T00:59:00+09:00</updated>
    <published>2026-08-20T00:59:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/workflow-fit-decides-adoption/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">AI product adoption improves when tools fit teachers&apos; existing workflows instead of asking them to change how they work.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; Workflow fit decides adoption. AI products enter daily use more easily when they support the planning, teaching, reviewing, and communication habits teachers already have.&lt;/p&gt;
&lt;h2&gt;why does workflow fit decide whether an AI product gets adopted?&lt;/h2&gt;
&lt;p&gt;Teachers already have ways of planning, teaching, reviewing work, and communicating. Adoption happens inside those existing habits.&lt;/p&gt;&lt;p&gt;A product that fits the workflow has a shorter path into daily use. A product that changes the workflow has to earn that change first.&lt;/p&gt;
&lt;h2&gt;why is asking teachers to change their workflow such a big ask?&lt;/h2&gt;
&lt;p&gt;AI products often ask teachers to change how they work. The product may be useful, yet the workflow can still feel wrong.&lt;/p&gt;&lt;p&gt;That creates friction before the feature becomes part of daily use. Many products lose momentum at that point.&lt;/p&gt;
&lt;h2&gt;why can an impressive feature and an excellent demo still fail?&lt;/h2&gt;
&lt;p&gt;Product teams study what the tool can do, what the model can generate, and the interface. They may also study the future workflow the product creates.&lt;/p&gt;&lt;p&gt;Teachers live in the current workflow. The gap between those two perspectives can matter more than the quality of the feature or demo.&lt;/p&gt;
&lt;h2&gt;where should AI product teams look before building?&lt;/h2&gt;
&lt;p&gt;The starting question is simple: where does this tool fit into the work people already do?&lt;/p&gt;&lt;p&gt;Workflow fit is a transferable adoption principle for AI products. It applies anywhere people already have a working rhythm, including education.&lt;/p&gt;
&lt;h2&gt;FAQ&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;What is workflow fit in AI products?&lt;/strong&gt;&lt;br&gt;Workflow fit means a product supports the work people already do. For teachers, that includes planning, teaching, reviewing work, and communicating.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Why do teachers resist useful AI tools?&lt;/strong&gt;&lt;br&gt;A useful tool can still ask teachers to change how they work. That workflow change creates a larger adoption hurdle than the feature itself.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;How can AI products improve adoption?&lt;/strong&gt;&lt;br&gt;Product teams can start by asking where the tool fits into the work already happening. Building around existing workflows gives the product a shorter path into daily use.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/i-d-build-the-learning-measurement-first/</id>
    <title type="text">i’d build the learning measurement first</title>
    <updated>2026-08-19T00:59:00+09:00</updated>
    <published>2026-08-19T00:59:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/i-d-build-the-learning-measurement-first/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">Why learning measurement should come before scaling an education product, with evidence collected inside the tutoring experience.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; Build learning measurement first. Define the outcome, measure it inside the tutoring experience, and review the evidence before scaling distribution.&lt;/p&gt;
&lt;h2&gt;what outcome is the product meant to change?&lt;/h2&gt;
&lt;p&gt;Start by defining the learning outcome the product is meant to change. That outcome gives the product a clear basis for measurement.&lt;/p&gt;&lt;p&gt;A compelling product still needs proof that students are learning, so the intended change must be stated clearly enough to evaluate.&lt;/p&gt;
&lt;h2&gt;where should learning evidence be collected?&lt;/h2&gt;
&lt;p&gt;Measure the defined outcome inside the tutoring experience. Keeping measurement within the experience connects the evidence to the product students are actually using.&lt;/p&gt;&lt;p&gt;This makes learning evidence part of the product itself, alongside the tutoring experience.&lt;/p&gt;
&lt;h2&gt;when should distribution scale?&lt;/h2&gt;
&lt;p&gt;Review the evidence before scaling distribution. Distribution can grow after there is evidence that the product is producing the intended learning change.&lt;/p&gt;&lt;p&gt;The sequence is direct: define the outcome, measure it in the tutoring experience, then review the evidence before expanding distribution.&lt;/p&gt;
&lt;h2&gt;FAQ&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Why should learning measurement come first?&lt;/strong&gt;&lt;br&gt;A compelling product still needs proof that students are learning. Building measurement first makes the intended learning change explicit before distribution scales.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Where should a tutoring product measure learning?&lt;/strong&gt;&lt;br&gt;The learning outcome should be measured inside the tutoring experience. This keeps the evidence connected to the product students are using.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/adoption-and-funding-are-weak-evidence-when-an-e/</id>
    <title type="text">adoption and funding are weak evidence when an education product</title>
    <updated>2026-08-18T00:59:00+09:00</updated>
    <published>2026-08-18T00:59:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/adoption-and-funding-are-weak-evidence-when-an-e/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">Why education products should measure learning outcomes early instead of relying on adoption, funding, or engagement metrics.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; Adoption and funding are weak evidence when an education product cannot prove learning outcomes. The evidence of learning needs to become part of the product early.&lt;/p&gt;
&lt;h2&gt;what does adoption prove if learning outcomes remain unclear?&lt;/h2&gt;
&lt;p&gt;Adoption shows that people are using a product. Funding shows that others believe in its potential. Neither proves that learners are making meaningful progress.&lt;/p&gt;&lt;p&gt;An education product can look active and well-supported while still going nowhere. The question that matters is whether the product can show evidence of learning.&lt;/p&gt;
&lt;h2&gt;why should learning progress be measured early?&lt;/h2&gt;
&lt;p&gt;Learning progress becomes harder to evaluate when measurement arrives late. Early measurement gives the product a way to see whether its work is producing the intended result.&lt;/p&gt;&lt;p&gt;The evidence should develop alongside the product. It belongs inside the experience, where progress can be observed as the product is used.&lt;/p&gt;
&lt;h2&gt;can polished engagement metrics hide a weak education product?&lt;/h2&gt;
&lt;p&gt;Yes. Engagement can look polished while learning outcomes remain unproven. Activity creates a convincing surface, yet it can conceal a product that is going nowhere.&lt;/p&gt;&lt;p&gt;For education products, progress is the evidence that matters most. Adoption can follow when the product has something real to show.&lt;/p&gt;
&lt;h2&gt;FAQ&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Should adoption come before proof of learning?&lt;/strong&gt;&lt;br&gt;Adoption can show interest, yet it does not prove learning. Education products should measure progress early and make evidence part of the product.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Why are funding and adoption weak evidence for education products?&lt;/strong&gt;&lt;br&gt;They indicate support or usage. They do not show whether learners are making progress or achieving learning outcomes.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What should education products measure first?&lt;/strong&gt;&lt;br&gt;They should measure learning progress early. The evidence should be built into the product rather than added after engagement metrics are already polished.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/44-less-capital/</id>
    <title type="text">44% less capital</title>
    <updated>2026-08-17T00:59:00+09:00</updated>
    <published>2026-08-17T00:59:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/44-less-capital/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">A field note on why staying small can require 44% less capital while reaching £100K ARR twice as often.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; The benchmark points to a clear finding: staying small used 44% less capital and reached £100K ARR twice as often. Smallness can function as a strategy.&lt;/p&gt;
&lt;h2&gt;what does the benchmark actually show?&lt;/h2&gt;
&lt;p&gt;The result is direct: 44% less capital, alongside reaching £100K ARR twice as often.&lt;/p&gt;&lt;p&gt;That combination changes how staying small looks. It starts to look like a deliberate strategy with a measurable advantage.&lt;/p&gt;
&lt;h2&gt;why does staying small look less like a constraint?&lt;/h2&gt;
&lt;p&gt;A constraint usually sounds like something that limits the outcome. This benchmark points in the opposite direction, with less capital associated with reaching a meaningful revenue milestone more often.&lt;/p&gt;&lt;p&gt;The finding gives builders and founders a reason to examine smallness as an operating choice rather than treating it as a temporary condition.&lt;/p&gt;
&lt;h2&gt;what should builders take from this benchmark?&lt;/h2&gt;
&lt;p&gt;The useful lesson is narrow and practical: capital efficiency matters, and the benchmark connects staying small with a higher frequency of reaching £100K ARR.&lt;/p&gt;&lt;p&gt;That is enough to make the strategy worth taking seriously. The numbers carry the argument.&lt;/p&gt;
&lt;h2&gt;FAQ&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;What does 44% less capital mean in this benchmark?&lt;/strong&gt;&lt;br&gt;The source states that the benchmark used 44% less capital. It does not provide a baseline or explain how capital was measured.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What does twice as often reaching £100K ARR mean?&lt;/strong&gt;&lt;br&gt;It means the benchmark reached £100K ARR at twice the frequency of the comparison being referenced. The source does not specify the comparison group.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Can staying small be a business strategy?&lt;/strong&gt;&lt;br&gt;This field note says the benchmark makes staying small look more like a strategy. The supporting figures are 44% less capital and twice as often reaching £100K ARR.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/can-you-really-charge-for-student-results-if-nob/</id>
    <title type="text">can you really charge for student results if nobody trusts the</title>
    <updated>2026-08-16T00:59:00+09:00</updated>
    <published>2026-08-16T00:59:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/can-you-really-charge-for-student-results-if-nob/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">Why trusted achievement data and instrumentation must come before performance-based pricing for student outcomes.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; Performance-based pricing depends on precise, trusted achievement data. Build the instrumentation layer first, or the contract rests on an argument about results.&lt;/p&gt;
&lt;h2&gt;can you charge for student results before anyone trusts the measurement?&lt;/h2&gt;
&lt;p&gt;Student results need to be precise and trusted before a vendor can credibly charge for outcomes. The measurement carries the claim that the result happened.&lt;/p&gt;&lt;p&gt;When trust is missing, the pricing conversation becomes a dispute about evidence instead of a clear exchange around performance.&lt;/p&gt;
&lt;h2&gt;why does measurement become part of the product?&lt;/h2&gt;
&lt;p&gt;When vendors are paid for outcomes, measurement determines whether those outcomes can be recognized and trusted. The data is part of what makes the product commercially usable.&lt;/p&gt;&lt;p&gt;Achievement data therefore has a direct role in performance-based pricing. It supports the connection between delivered work and the result attached to the contract.&lt;/p&gt;
&lt;h2&gt;what should a solo builder instrument before pricing against results?&lt;/h2&gt;
&lt;p&gt;The first requirement is precise achievement data that people can trust. That instrumentation layer gives the result a defensible basis before pricing is tied to it.&lt;/p&gt;&lt;p&gt;The field finding is simple: measurement needs to exist before the contract depends on performance.&lt;/p&gt;
&lt;h2&gt;what happens when the evidence behind the result is weak?&lt;/h2&gt;
&lt;p&gt;Pricing becomes an argument about whether the result is real, precise, and trusted. That uncertainty can overwhelm the outcome itself.&lt;/p&gt;&lt;p&gt;For performance-based pricing to work, the instrumentation must make achievement clear enough for both sides to accept the result.&lt;/p&gt;
&lt;h2&gt;FAQ&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;what data is needed for performance-based pricing?&lt;/strong&gt;&lt;br&gt;The source points to precise, trusted achievement data. That data needs to make the measured result clear enough to support an outcome-based contract.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;why is instrumentation important for student outcomes?&lt;/strong&gt;&lt;br&gt;Instrumentation provides the measurement layer behind achievement data. Vendors need that layer when payment depends on outcomes.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;can performance-based pricing work without trusted results?&lt;/strong&gt;&lt;br&gt;The source argues that it cannot work reliably without precise, trusted achievement data. Without that foundation, pricing rests on an argument about the result.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/a-big-virtual-tutoring-provider-just-shut-down-b/</id>
    <title type="text">a big virtual tutoring provider just shut down because they</title>
    <updated>2026-08-15T00:59:00+09:00</updated>
    <published>2026-08-15T00:59:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/a-big-virtual-tutoring-provider-just-shut-down-b/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">A big virtual tutoring provider shut down without evidence it worked. Here&apos;s how to instrument one outcome before you ship, not after.</summary>
    <content type="html">&lt;p&gt;a big virtual tutoring provider just shut down because they couldn&amp;#39;t show evidence it actually worked. that&amp;#39;s the whole story. the product existed, people used it, and when the moment came to prove the thing it claimed to do, there was nothing on the table. i keep coming back to this because it&amp;#39;s the failure mode i&amp;#39;m most likely to walk into myself, and because it doesn&amp;#39;t announce itself early. it announces itself at the exact moment you can&amp;#39;t fix it.&lt;/p&gt;&lt;p&gt;here&amp;#39;s the situation as i understand it from the outside. a company builds something in edtech. the pitch is that it moves an outcome. usage goes up, teachers say nice things, renewals come in. every one of those signals feels like proof while you&amp;#39;re inside it. and then someone with money asks the direct question, and the answer has to be assembled retroactively out of whatever data happened to survive. that assembly almost never works, because the data you needed was a decision you had to make before you shipped, and you didn&amp;#39;t know you were making it.&lt;/p&gt;&lt;p&gt;the receipt is the shutdown itself. that&amp;#39;s the number here, and it&amp;#39;s a binary one. the market ran the experiment and published the result: no evidence, no company. i don&amp;#39;t have a percentage or a cohort size to offer you, and i&amp;#39;m not going to invent one. what i have is a single observed outcome that&amp;#39;s more expensive than any dashboard you could have built instead.&lt;/p&gt;&lt;p&gt;the mechanism is that the signals you collect by default are the ones that are easy to collect. usage is easy. retention is easy. sentiment is easy, because you can just ask. all three are downstream of whether people like the experience, and liking the experience is genuinely correlated with the outcome you care about, which is exactly what makes it dangerous. correlation is enough to keep you confident and not enough to survive a question. &amp;quot;teachers love it&amp;quot; is a fact about teachers. retention is a fact about your billing. the outcome you claim to move is a fact about students, and you have to go get it on purpose.&lt;/p&gt;&lt;p&gt;so the rule i adopted, and the one i&amp;#39;d hand to anyone building in this space before their next raise, is to pick one outcome you claim to move. one. resist the urge to name a portfolio. then define, on day one, how you would measure it at day ninety. write the measurement down while you still have no results, because the definition you write after seeing data is a different and much worse definition. instrument it before the feature ships, since instrumenting after means your baseline is gone and everything you learn is relative to a starting point you&amp;#39;re guessing at.&lt;/p&gt;&lt;p&gt;the harder half is running a cohort that doesn&amp;#39;t get the feature. it will feel wasteful and it will be small and messy, and small and messy is fine. a rough comparison group beats no comparison group by an enormous margin, because without one you have a number with nothing to hold it against. every trend you see becomes attributable to whatever you were hoping for. i&amp;#39;ve done this badly on my own projects, and the version where i ship to everyone at once always produces a chart that looks great and proves nothing.&lt;/p&gt;&lt;p&gt;then write the result down when it&amp;#39;s bad. especially when it&amp;#39;s bad. this is the part everyone nods at and nobody does, because a bad result is a thing you have to look at again later, and it&amp;#39;s easier to let it quietly stop being measured. the discipline is that a written negative result is the only thing that makes your positive results credible. someone evaluating you can tell the difference between a company that measured and reported, and a company that measured until the numbers were good. the second pattern is visible from the outside and it reads as exactly what it is.&lt;/p&gt;&lt;p&gt;what to do with this: go look at your product today and ask what outcome you would actually put in front of someone who was skeptical. if the answer is usage or sentiment, you have the problem, and you have it right now while it&amp;#39;s still cheap. the fix is one instrumented outcome, one measurement definition written before the data exists, one imperfect comparison group, and a habit of recording what you find. none of that is expensive. it&amp;#39;s just work you have to do before you need it, which is the only kind of work that&amp;#39;s easy to skip.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/when-do-you-call-a-commit-shipped/</id>
    <title type="text">when do you call a commit shipped?</title>
    <updated>2026-08-13T00:59:00+09:00</updated>
    <published>2026-08-13T00:59:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/when-do-you-call-a-commit-shipped/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">A commit is shipped when the live route and customer experience are verified, because the repository alone cannot prove delivery.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; a commit proves code changed. the work is shipped after the live route and what customers can actually use are verified.&lt;/p&gt;
&lt;h2&gt;when does a commit become shipped?&lt;/h2&gt;
&lt;p&gt;a commit becomes shipped after you verify the live route and confirm what customers can actually use.&lt;/p&gt;&lt;p&gt;the repository can look perfect while reality says otherwise. live behavior is the receipt.&lt;/p&gt;
&lt;h2&gt;why can a perfect repository still be unfinished?&lt;/h2&gt;
&lt;p&gt;the repo only proves that code changed. it does not prove that the change reached a live route or became available to customers.&lt;/p&gt;&lt;p&gt;that gap is where “done” can sound true while the customer-facing result says otherwise.&lt;/p&gt;
&lt;h2&gt;what should you check before saying “done”?&lt;/h2&gt;
&lt;p&gt;check the live route and the customer-usable result before calling the commit shipped.&lt;/p&gt;&lt;p&gt;the final claim should follow verified live behavior, since that is the evidence customers actually receive.&lt;/p&gt;
&lt;h2&gt;FAQ&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;does committing code mean the work is shipped?&lt;/strong&gt;&lt;br&gt;a commit proves code changed. shipped means the live route and the customer-usable result were verified.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;why verify the live route after a commit?&lt;/strong&gt;&lt;br&gt;the repository can look perfect while the live result differs. verifying the route shows what customers can actually use.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;what is the receipt for shipped software?&lt;/strong&gt;&lt;br&gt;live behavior is the receipt. It confirms the change exists where customers encounter it.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/your-ai-spend-belongs-at-the-single-model-call-c/</id>
    <title type="text">your AI spend belongs at the single model-call chokepoint</title>
    <updated>2026-08-11T00:59:00+09:00</updated>
    <published>2026-08-11T00:59:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/your-ai-spend-belongs-at-the-single-model-call-c/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">A field note on routing AI spend through one model-call chokepoint and keeping telemetry fail open during observability outages.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; Put AI spend at the single model-call chokepoint, then make telemetry fail open. Cost visibility stays in the loop while observability outages leave the product running.&lt;/p&gt;
&lt;h2&gt;why should AI spend pass through one model-call chokepoint?&lt;/h2&gt;
&lt;p&gt;A single model-call chokepoint gives AI spend one place to observe. That keeps cost visibility attached to the call path where the spend happens.&lt;/p&gt;&lt;p&gt;For solo builders using AI agents, this creates a clear boundary for telemetry while the rest of the product stays focused on its job.&lt;/p&gt;
&lt;h2&gt;what does it mean to make telemetry fail open?&lt;/h2&gt;
&lt;p&gt;Telemetry should keep reporting when available, while its failure leaves the product running. The observability path can degrade without taking the product path down with it.&lt;/p&gt;&lt;p&gt;That separation matters because AI apps already have enough ways to fall over. Cost tracking remains useful without becoming another dependency that can stop the application.&lt;/p&gt;
&lt;h2&gt;what changed once cost visibility stayed in the loop?&lt;/h2&gt;
&lt;p&gt;The system keeps visibility into AI spend through the model-call chokepoint, even when observability has an outage. Product availability and cost visibility each get a clearer boundary.&lt;/p&gt;&lt;p&gt;The finding is simple: observe the spend at its source, and let telemetry failures pass through without stopping the product.&lt;/p&gt;
&lt;h2&gt;FAQ&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;where should AI cost tracking happen?&lt;/strong&gt;&lt;br&gt;AI cost tracking should sit at the single model-call chokepoint. That gives spend one observable place in the call path.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;what does fail-open telemetry mean for an AI app?&lt;/strong&gt;&lt;br&gt;It means telemetry failures leave the product running. Observability can be unavailable while the application continues serving its purpose.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;why separate observability from the product path?&lt;/strong&gt;&lt;br&gt;AI apps already have several ways to fail. Keeping telemetry fail open prevents an observability outage from becoming a product outage.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/a-forecasting-system-produced-a-net-improvement-/</id>
    <title type="text">a forecasting system produced a net improvement of 1.8</title>
    <updated>2026-08-10T00:59:00+09:00</updated>
    <published>2026-08-10T00:59:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/a-forecasting-system-produced-a-net-improvement-/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">A field note on a forecasting system&apos;s 1.8-point quarterly improvement across 47 quarters, with directional significance at t=1.45.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; A forecasting system produced a net improvement of 1.8 percentage points per quarter across 47 quarters. Directional significance came in at t=1.45, giving the result a concrete receipt beyond “it predicts well.”&lt;/p&gt;
&lt;h2&gt;what did the forecasting system actually improve?&lt;/h2&gt;
&lt;p&gt;The system produced a net improvement of 1.8 percentage points per quarter.&lt;/p&gt;&lt;p&gt;That result was measured across 47 quarters, giving the claim a defined performance record.&lt;/p&gt;
&lt;h2&gt;how much directional evidence did the result have?&lt;/h2&gt;
&lt;p&gt;Directional significance came in at t=1.45.&lt;/p&gt;&lt;p&gt;That number belongs alongside the quarterly improvement and the 47-quarter sample when describing what the system produced.&lt;/p&gt;
&lt;h2&gt;why is this receipt more useful than saying “it predicts well”?&lt;/h2&gt;
&lt;p&gt;“It predicts well” leaves the result vague. The 1.8 percentage point net improvement per quarter, measured across 47 quarters, gives a reader something concrete to inspect, while t=1.45 records the directional significance.&lt;/p&gt;
&lt;h2&gt;FAQ&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;What was the forecasting system’s net improvement?&lt;/strong&gt;&lt;br&gt;It produced a net improvement of 1.8 percentage points per quarter.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Across how many quarters was the forecasting result measured?&lt;/strong&gt;&lt;br&gt;The result was measured across 47 quarters.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What was the directional significance of the forecasting system?&lt;/strong&gt;&lt;br&gt;Directional significance came in at t=1.45.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Why report the numbers instead of saying the system predicts well?&lt;/strong&gt;&lt;br&gt;The numbers provide a concrete receipt. They show the net improvement, the period covered, and the directional significance.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/instrument-model-usage-at-one-chokepoint/</id>
    <title type="text">instrument model usage at one chokepoint</title>
    <updated>2026-08-09T00:59:00+09:00</updated>
    <published>2026-08-09T00:59:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/instrument-model-usage-at-one-chokepoint/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">A field note on instrumenting multi-model usage at one safe chokepoint with telemetry that fails quietly.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; I instrument model usage at one chokepoint, with optional telemetry guarded so it can never block the primary request. That gives multi-model systems one safe measurement layer, while monitoring fails quietly.&lt;/p&gt;
&lt;h2&gt;where should model usage be instrumented in a multi-model system?&lt;/h2&gt;
&lt;p&gt;The useful place is one shared chokepoint. Every model request passes through the same measurement layer, so usage can be observed consistently across a multi-model system.&lt;/p&gt;&lt;p&gt;That keeps instrumentation focused on one boundary instead of scattering monitoring logic across individual model paths.&lt;/p&gt;
&lt;h2&gt;what happens when telemetry has a problem?&lt;/h2&gt;
&lt;p&gt;Telemetry stays optional and guarded. The primary request continues even when monitoring has an issue.&lt;/p&gt;&lt;p&gt;The monitoring layer should fail quietly, keeping measurement from becoming a dependency of the work it observes.&lt;/p&gt;
&lt;h2&gt;why does one safe measurement layer matter?&lt;/h2&gt;
&lt;p&gt;A single chokepoint creates one place to instrument model usage. It gives a multi-model system a shared view of activity without requiring every model path to carry its own measurement approach.&lt;/p&gt;&lt;p&gt;The design constraint is simple: observe the request, protect the request, and let monitoring disappear when it cannot do its job.&lt;/p&gt;
&lt;h2&gt;FAQ&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;where is the best place to add model usage telemetry?&lt;/strong&gt;&lt;br&gt;Add it at a shared chokepoint through which model requests pass. This gives multi-model systems one consistent measurement layer.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;can telemetry block the primary model request?&lt;/strong&gt;&lt;br&gt;It should never block the primary request. Guard optional telemetry so monitoring failures fail quietly.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;how should monitoring behave when it breaks?&lt;/strong&gt;&lt;br&gt;Monitoring should fail quietly while the primary request continues. Measurement remains useful without becoming part of the request&amp;apos;s success condition.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/routing-every-task-to-the-strongest-model-is-a-p/</id>
    <title type="text">routing every task to the strongest model is a pretty expensive</title>
    <updated>2026-08-08T00:59:00+09:00</updated>
    <published>2026-08-08T00:59:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/routing-every-task-to-the-strongest-model-is-a-p/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">Why routing every task to the strongest model costs more, and where free local models fit for simple, high-volume work.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; Routing every task to the strongest model is expensive. Simple, high-volume work can go to a free local model while the stronger model handles judgment-heavy cases.&lt;/p&gt;
&lt;h2&gt;why send simple work to the strongest model?&lt;/h2&gt;
&lt;p&gt;Because it becomes an expensive habit. Simple, high-volume tasks often do not need the same level of judgment as harder cases.&lt;/p&gt;&lt;p&gt;A free local model can handle the boring volume and reserve the stronger model for work where judgment actually matters.&lt;/p&gt;
&lt;h2&gt;what changes when model routing follows the work?&lt;/h2&gt;
&lt;p&gt;The stronger model stays focused on the cases that need it. Simple work moves to a free local model, which lowers cost across the routine workload.&lt;/p&gt;&lt;p&gt;The intended result is the same quality where it counts, with less spent on easy work.&lt;/p&gt;
&lt;h2&gt;is there a downside to using a free local model?&lt;/h2&gt;
&lt;p&gt;That is the open question. The tradeoff depends on whether simple, high-volume work truly stays simple enough for the local model.&lt;/p&gt;&lt;p&gt;The field note points toward testing the boundary instead of paying the strongest-model rate for every task.&lt;/p&gt;
&lt;h2&gt;FAQ&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Should every AI task use the strongest model?&lt;/strong&gt;&lt;br&gt;Routing every task to the strongest model can become an expensive habit. Simple, high-volume work can go to a free local model, while judgment-heavy cases stay with the stronger model.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;When should I use a free local model?&lt;/strong&gt;&lt;br&gt;Use it for simple, high-volume work. The stronger model can handle cases where judgment actually matters.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Can local models lower AI costs without hurting quality?&lt;/strong&gt;&lt;br&gt;They can lower cost across routine work when the task is simple enough. The goal is to preserve quality where it counts.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/what-if-your-strongest-model-only-saw-the-hard-c/</id>
    <title type="text">what if your strongest model only saw the hard cases?</title>
    <updated>2026-08-07T00:59:00+09:00</updated>
    <published>2026-08-07T00:59:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/what-if-your-strongest-model-only-saw-the-hard-c/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">A field note on routing high-volume classification through a free local model and sending difficult cases to a stronger model.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; A free local model handled high-volume classification and extraction first, while a stronger model reviewed the difficult survivors. This lowered cost and latency while reserving judgment for cases that needed it.&lt;/p&gt;
&lt;h2&gt;why should the strongest model see every case?&lt;/h2&gt;
&lt;p&gt;High-volume classification and extraction can begin with a free local model. The stronger model then focuses on the cases that survive the first pass as difficult.&lt;/p&gt;&lt;p&gt;That routing keeps expensive judgment concentrated where it matters most.&lt;/p&gt;
&lt;h2&gt;what changed when we routed difficult cases separately?&lt;/h2&gt;
&lt;p&gt;The result was lower cost and lower latency. The stronger model spent its effort on difficult survivors instead of processing the full high-volume workload.&lt;/p&gt;
&lt;h2&gt;which cases deserve the expensive model?&lt;/h2&gt;
&lt;p&gt;The difficult cases deserve the stronger model. The first model acts as the initial classification and extraction layer, leaving the stronger model with the cases that need more judgment.&lt;/p&gt;
&lt;h2&gt;can a free local model handle the first pass?&lt;/h2&gt;
&lt;p&gt;In this field note, a free local model ran high-volume classification and extraction before the difficult survivors moved to a stronger model. The finding is about routing work according to difficulty.&lt;/p&gt;
&lt;h2&gt;FAQ&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;How do you reduce LLM cost for high-volume classification?&lt;/strong&gt;&lt;br&gt;Run a free local model through the initial classification and extraction pass. Send the difficult survivors to a stronger model so expensive judgment is applied to fewer cases.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;When should a case be sent to a stronger AI model?&lt;/strong&gt;&lt;br&gt;Send cases that survive the first pass as difficult. The stronger model can then focus its judgment where it actually matters.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Can local models handle high-volume extraction?&lt;/strong&gt;&lt;br&gt;A free local model handled the high-volume classification and extraction pass in this field note. Difficult survivors were routed to a stronger model afterward.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/small-ai-experiments-deserve-a-price-check-befor/</id>
    <title type="text">small AI experiments deserve a price check before they run</title>
    <updated>2026-08-06T00:59:00+09:00</updated>
    <published>2026-08-06T00:59:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/small-ai-experiments-deserve-a-price-check-befor/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">Why small AI experiments need exact pricing and a run ceiling before they start.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; A quick AI test can quietly become an expensive batch job. Check the exact model, endpoint price, and run ceiling before it runs.&lt;/p&gt;
&lt;h2&gt;why can a quick AI test become expensive?&lt;/h2&gt;
&lt;p&gt;Small experiments can expand into batch jobs without making the cost obvious at the start.&lt;/p&gt;&lt;p&gt;That shift turns a quick test into an expensive run, especially when the pricing and execution limit were never checked.&lt;/p&gt;
&lt;h2&gt;what should i verify before running a small experiment?&lt;/h2&gt;
&lt;p&gt;Verify the exact model being used, the price for its endpoint, and the ceiling for the run.&lt;/p&gt;&lt;p&gt;Those three checks make the intended scope visible before the experiment starts.&lt;/p&gt;
&lt;h2&gt;why does a run ceiling matter?&lt;/h2&gt;
&lt;p&gt;A run ceiling keeps a small experiment bounded. Without one, a test can continue quietly and grow into a batch job with a much larger bill.&lt;/p&gt;&lt;p&gt;The price check belongs before execution, while the experiment is still small enough to control.&lt;/p&gt;
&lt;h2&gt;FAQ&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;How do I prevent a quick AI test from becoming expensive?&lt;/strong&gt;&lt;br&gt;Verify the exact model, endpoint price, and run ceiling before the test runs. These checks keep the experiment within its intended scope.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What should I check before running an AI experiment?&lt;/strong&gt;&lt;br&gt;Check which exact model will run, what the endpoint costs, and how many runs are allowed. The source finding is that these details should be confirmed before execution.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/send-the-boring-work-to-a-free-local-model/</id>
    <title type="text">send the boring work to a free local model</title>
    <updated>2026-08-05T00:59:00+09:00</updated>
    <published>2026-08-05T00:59:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/send-the-boring-work-to-a-free-local-model/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">A field note on routing routine AI work to a free local model and reserving stronger models for judgment.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; routing boring work to a free local model and sending important survivors to a stronger model kept decision quality steady while lowering overall cost. A lot of AI spend comes from routing every task to the smartest model.&lt;/p&gt;
&lt;h2&gt;why should every task pay for the smartest model?&lt;/h2&gt;
&lt;p&gt;I started separating routine work from the tasks that actually need judgment. The boring work goes to a free local model, while the important survivors move to a stronger model.&lt;/p&gt;&lt;p&gt;That simple routing change exposed how much spend comes from treating every task as equally difficult. Most tasks do not need the same level of reasoning.&lt;/p&gt;
&lt;h2&gt;what happens when only the important work reaches a stronger model?&lt;/h2&gt;
&lt;p&gt;The important survivors still receive stronger-model judgment, which keeps quality where decisions matter. The result was the same quality on those decisions with lower cost overall.&lt;/p&gt;&lt;p&gt;The model choice becomes part of the workflow. Routine work gets filtered locally, and the stronger model spends its capacity on the cases that deserve it.&lt;/p&gt;
&lt;h2&gt;is AI spend really a model problem or a routing problem?&lt;/h2&gt;
&lt;p&gt;The receipt points toward routing. When every task pays for the smartest model, the budget carries the cost of high-end judgment even when the task is boring.&lt;/p&gt;&lt;p&gt;A better split sends routine work to the free local model and reserves the stronger model for important survivors. That makes model spend follow decision value more closely.&lt;/p&gt;
&lt;h2&gt;FAQ&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;How can I lower AI costs without reducing decision quality?&lt;/strong&gt;&lt;br&gt;Route routine work to a free local model and send important survivors to a stronger model. This keeps stronger-model judgment focused where decisions matter.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Should every AI task use the strongest model?&lt;/strong&gt;&lt;br&gt;Every task does not need the same model. Use a free local model for boring work and reserve the stronger model for important survivors.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Why is AI spending higher than expected?&lt;/strong&gt;&lt;br&gt;Routing may be sending every task to the smartest model. That makes routine work carry the cost of high-end judgment.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/a-commit-is-a-checkpoint-shipping-is-live-verifi/</id>
    <title type="text">a commit is a checkpoint, shipping is live verification</title>
    <updated>2026-08-04T00:59:00+09:00</updated>
    <published>2026-08-04T00:59:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/a-commit-is-a-checkpoint-shipping-is-live-verifi/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">A field note on why clean commits do not equal shipped software, and why live customer-visible verification defines completion.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; a commit is a checkpoint. shipping means verifying that the product, route, or behavior works live where customers can see it.&lt;/p&gt;
&lt;h2&gt;what does shipped actually mean?&lt;/h2&gt;
&lt;p&gt;shipped means the work has been verified live. The repository can be clean while the product, route, or behavior remains broken.&lt;/p&gt;&lt;p&gt;The useful finish line is customer-visible evidence that the change works in the real product.&lt;/p&gt;
&lt;h2&gt;why can a clean repo still hide broken work?&lt;/h2&gt;
&lt;p&gt;A clean repository records a checkpoint in the code. It does not prove that the running product reflects that checkpoint or that the intended behavior works.&lt;/p&gt;&lt;p&gt;That gap is where teams can mistake completed development work for a finished product change.&lt;/p&gt;
&lt;h2&gt;what should a team verify before calling something finished?&lt;/h2&gt;
&lt;p&gt;Verify the live product, the relevant route, or the behavior customers are meant to use. The question is whether customers can see it working.&lt;/p&gt;&lt;p&gt;If customers cannot see the result working, the work remains unfinished regardless of the repository state.&lt;/p&gt;
&lt;h2&gt;how should we define shipping for our team?&lt;/h2&gt;
&lt;p&gt;Define shipped around live verification and customer visibility. A commit can mark progress, while live evidence marks completion.&lt;/p&gt;&lt;p&gt;The definition should make clear what working means in the product customers actually encounter.&lt;/p&gt;
&lt;h2&gt;FAQ&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;does a clean git commit mean the feature is shipped?&lt;/strong&gt;&lt;br&gt;A clean commit marks a checkpoint in the repository. Shipping requires live verification that the product, route, or behavior works.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;what is the difference between committing and shipping software?&lt;/strong&gt;&lt;br&gt;Committing records code progress. Shipping confirms that customers can see the intended result working live.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;how do you know when software work is finished?&lt;/strong&gt;&lt;br&gt;The work is finished when the relevant product, route, or behavior has been verified live and customers can see it working.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;what should shipped mean for a software team?&lt;/strong&gt;&lt;br&gt;Shipped should mean live, customer-visible verification. Each team should define the evidence that proves the intended result works.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/how-do-you-make-a-multi-provider-ai-system-measu/</id>
    <title type="text">how do you make a multi-provider AI system measurable?</title>
    <updated>2026-08-03T00:59:00+09:00</updated>
    <published>2026-08-03T00:59:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/how-do-you-make-a-multi-provider-ai-system-measu/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">How to make a multi-provider AI system measurable by observing every model call at one chokepoint.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; A multi-provider AI system becomes measurable when observability sits at the single model-call chokepoint. Record the model, usage, latency, and outcome there.&lt;/p&gt;
&lt;h2&gt;where should observability sit in a multi-provider AI system?&lt;/h2&gt;
&lt;p&gt;Put it at the single model-call chokepoint. Every provider call passes through that integration, giving the system one place to capture the same measurements.&lt;/p&gt;&lt;p&gt;One well-placed integration can make the whole system measurable without scattering observability across every provider-specific path.&lt;/p&gt;
&lt;h2&gt;what should you record at the model-call chokepoint?&lt;/h2&gt;
&lt;p&gt;Record the model, usage, latency, and outcome for each call. Those fields create a consistent view across the system&amp;#39;s model providers.&lt;/p&gt;&lt;p&gt;The point is to capture the receipt where the call happens, while the model choice, resource usage, timing, and result are all available together.&lt;/p&gt;
&lt;h2&gt;why does one integration make the whole system measurable?&lt;/h2&gt;
&lt;p&gt;A shared model-call chokepoint gives every provider call the same observability boundary. The integration becomes the place where the system&amp;#39;s model activity can be measured consistently.&lt;/p&gt;&lt;p&gt;The practical question is simple: where is your model-call chokepoint today? That location determines whether measurement is centralized or scattered.&lt;/p&gt;
&lt;h2&gt;FAQ&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;How do you measure a multi-provider AI system?&lt;/strong&gt;&lt;br&gt;Place observability at the single model-call chokepoint. Record the model, usage, latency, and outcome for each call.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What metrics should be recorded for every model call?&lt;/strong&gt;&lt;br&gt;Record the model, usage, latency, and outcome. These are the measurements identified for the model-call chokepoint.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Why use a model-call chokepoint for observability?&lt;/strong&gt;&lt;br&gt;A single integration can cover the whole multi-provider system when every model call passes through it. This creates one place to record consistent measurements.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/i-d-put-the-ai-tutor-inside-the-teacher-s-workfl/</id>
    <title type="text">i’d put the AI tutor inside the teacher’s workflow</title>
    <updated>2026-08-01T00:59:00+09:00</updated>
    <published>2026-08-01T00:59:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/i-d-put-the-ai-tutor-inside-the-teacher-s-workfl/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">Why AI tutors earn trust when goals, notes, and next actions stay inside the teacher’s workflow.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; AI tutors become easier to trust when they live inside the teacher’s workflow. Set the goal, save tutor notes in the same record, and keep the next action visible there.&lt;/p&gt;
&lt;h2&gt;where should an AI tutor live during a teaching session?&lt;/h2&gt;
&lt;p&gt;i’d put it inside the teacher’s workflow, where the session already has its context and follow-up.&lt;/p&gt;&lt;p&gt;That keeps the tutor connected to the work the teacher is already doing.&lt;/p&gt;
&lt;h2&gt;what should the teacher define before the session starts?&lt;/h2&gt;
&lt;p&gt;set the goal before the session. A clear goal gives the tutor and teacher a shared point of reference.&lt;/p&gt;&lt;p&gt;The goal also makes the session easier to understand after it ends.&lt;/p&gt;
&lt;h2&gt;what makes an AI tutor feel accountable after the session?&lt;/h2&gt;
&lt;p&gt;save tutor notes in the same record and make the next action visible there too.&lt;/p&gt;&lt;p&gt;That puts context and accountability in one place. That’s where trust starts.&lt;/p&gt;
&lt;h2&gt;FAQ&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;how do you build trust in an AI tutor?&lt;/strong&gt;&lt;br&gt;Put the AI tutor inside the teacher’s workflow. Keep the goal, tutor notes, and next action connected in one place.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;what should an AI tutor record after a session?&lt;/strong&gt;&lt;br&gt;The tutor should save its notes in the same record as the session. The next action should be visible there too.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;why should the session goal be set in advance?&lt;/strong&gt;&lt;br&gt;A goal gives the session a clear direction. It also keeps the tutor’s work connected to what the teacher intended to accomplish.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/11-weeks/</id>
    <title type="text">11 weeks</title>
    <updated>2026-07-30T00:59:00+09:00</updated>
    <published>2026-07-30T00:59:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/11-weeks/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">What 11 weeks and a 50-participant cap reveal about bounded execution, ambition, and visible progress.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; 11 weeks with a maximum of 50 participants creates an operating constraint with teeth. It gives ambition a bounded execution season and leaves less room for vague progress to hide.&lt;/p&gt;
&lt;h2&gt;what changes when execution has an 11-week boundary?&lt;/h2&gt;
&lt;p&gt;An 11-week window turns ambition into a defined operating season. The work has a visible beginning and end, which gives plans a period in which they must become real.&lt;/p&gt;&lt;p&gt;A bounded season creates pressure to act while the constraint is still active. Progress has less room to drift into an open-ended future.&lt;/p&gt;
&lt;h2&gt;why does capping participation at 50 matter?&lt;/h2&gt;
&lt;p&gt;A cap of 50 participants gives the operating constraint teeth. Participation has a clear limit, so the season carries a defined scope alongside its defined duration.&lt;/p&gt;&lt;p&gt;The cap makes the commitment easier to see. Everyone involved is operating inside the same boundary.&lt;/p&gt;
&lt;h2&gt;how does a bounded season expose vague progress?&lt;/h2&gt;
&lt;p&gt;Vague progress can survive inside an unlimited timeline because there is always more time to explain the gap. An 11-week season leaves less room for that gap to hide.&lt;/p&gt;&lt;p&gt;The constraint makes execution easier to judge. Ambition has a fixed window, participation has a fixed ceiling, and progress has to show up inside both.&lt;/p&gt;
&lt;h2&gt;FAQ&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;What is a bounded execution season?&lt;/strong&gt;&lt;br&gt;It is a defined period in which ambition has to become execution. In this field note, the season lasts 11 weeks and is capped at 50 participants.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Why do operating constraints help reveal progress?&lt;/strong&gt;&lt;br&gt;A fixed duration and participant cap reduce the space available for vague progress. The work has to show up within a visible boundary.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/what-actually-turns-ai-tutoring-into-a-real-mark/</id>
    <title type="text">what actually turns AI tutoring into a real market?</title>
    <updated>2026-07-29T00:59:00+09:00</updated>
    <published>2026-07-29T00:59:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/what-actually-turns-ai-tutoring-into-a-real-mark/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">UK school AI funding signals demand and makes safety and learning impact part of the product trial.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; The UK has set aside up to £23 million for school AI and edtech trials. That funding signals demand while putting safety and learning impact inside the product itself.&lt;/p&gt;
&lt;h2&gt;what does UK funding for school AI actually signal?&lt;/h2&gt;
&lt;p&gt;The UK has set aside up to £23 million for school AI and edtech trials. That is a direct demand signal for builders working in this market.&lt;/p&gt;&lt;p&gt;The opportunity sits inside a trial context, where the product has to show what it contributes to schools. That changes the question from whether AI tutoring sounds promising to what the trial can prove.&lt;/p&gt;
&lt;h2&gt;why do safety and learning impact belong inside the product?&lt;/h2&gt;
&lt;p&gt;The funding signal puts safety and learning impact inside the product. They become part of the value being evaluated alongside the AI or edtech experience.&lt;/p&gt;&lt;p&gt;For a solo builder using AI agents, that means the product story needs to include what happens for learners and how the system operates safely. Those points are part of the market itself.&lt;/p&gt;
&lt;h2&gt;what should an AI tutoring trial prove?&lt;/h2&gt;
&lt;p&gt;The source leaves the question open: if you are building here, what will your trial prove? That is the useful starting point for deciding what to build and measure.&lt;/p&gt;&lt;p&gt;A trial should have a clear claim about safety and learning impact. The funding creates room to test that claim in the school AI and edtech market.&lt;/p&gt;
&lt;h2&gt;FAQ&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;How much funding has the UK set aside for school AI and edtech trials?&lt;/strong&gt;&lt;br&gt;The UK has set aside up to £23 million for school AI and edtech trials. The funding is presented as a demand signal.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What does UK school AI funding mean for AI tutoring products?&lt;/strong&gt;&lt;br&gt;It signals demand for school AI and edtech products. It also puts safety and learning impact inside the product.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What should an AI tutoring trial prove?&lt;/strong&gt;&lt;br&gt;The trial should make its intended proof explicit, especially around safety and learning impact. The source frames that question for builders working in the UK.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/a-durable-student-record-should-outlast-whoever-/</id>
    <title type="text">a durable student record should outlast whoever delivers the</title>
    <updated>2026-07-28T00:59:00+09:00</updated>
    <published>2026-07-28T00:59:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/a-durable-student-record-should-outlast-whoever-/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">Why durable student records improve handoff continuity and make personalization measurable in education.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; A durable student record should preserve goals and what happened next, so continuity survives a handoff and personalization becomes measurable.&lt;/p&gt;
&lt;h2&gt;what should stay attached to a student record?&lt;/h2&gt;
&lt;p&gt;Keep the student’s goals attached to the record, along with what happened next.&lt;/p&gt;&lt;p&gt;That gives future lesson delivery the context needed to continue the student’s journey.&lt;/p&gt;
&lt;h2&gt;how does continuity survive a handoff?&lt;/h2&gt;
&lt;p&gt;Continuity survives when the record outlasts whoever delivers the lesson.&lt;/p&gt;&lt;p&gt;The next person can see the student’s goals and the events that followed, so the handoff carries useful context forward.&lt;/p&gt;
&lt;h2&gt;how can personalization become measurable?&lt;/h2&gt;
&lt;p&gt;Personalization becomes measurable when the student’s goals and subsequent outcomes remain connected in one durable record.&lt;/p&gt;&lt;p&gt;The same approach can be copied by any education operator.&lt;/p&gt;
&lt;h2&gt;FAQ&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;What should an education record include?&lt;/strong&gt;&lt;br&gt;It should include the student’s goals and what happened next. Keeping both attached preserves context for future lesson delivery.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;How can education teams preserve continuity during handoffs?&lt;/strong&gt;&lt;br&gt;Use a durable student record that outlasts the person delivering the lesson. The next person can then continue from the student’s existing goals and history.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;How does a durable record make personalization measurable?&lt;/strong&gt;&lt;br&gt;It connects the student’s goals with what happened afterward. That connection makes the results of personalization easier to observe.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/use-this-filter-before-building-an-ai-tutoring-p/</id>
    <title type="text">use this filter before building an AI tutoring product</title>
    <updated>2026-07-27T00:59:00+09:00</updated>
    <published>2026-07-27T00:59:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/use-this-filter-before-building-an-ai-tutoring-p/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">A field note on why AI tutoring products need teacher workflows, student goals, tutor notes, and clear next actions.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; The workflow is the product. An AI tutoring session should connect to a real teacher workflow, the student’s goal, tutor notes, and a next action.&lt;/p&gt;
&lt;h2&gt;what should an AI tutoring session connect to?&lt;/h2&gt;
&lt;p&gt;Tie the session to a real teacher workflow. That connection gives the conversation a place inside the work teachers already do.&lt;/p&gt;&lt;p&gt;A tutoring session should exist within that workflow instead of standing alone as another conversation.&lt;/p&gt;
&lt;h2&gt;whose goal should shape the session?&lt;/h2&gt;
&lt;p&gt;Anchor the session to the student’s goal. The goal gives the tutoring interaction a clear direction and keeps the session connected to what the student is trying to achieve.&lt;/p&gt;
&lt;h2&gt;what should happen after the tutoring conversation ends?&lt;/h2&gt;
&lt;p&gt;Capture tutor notes and the next action. Those details carry the session forward and connect it to the surrounding teacher workflow.&lt;/p&gt;&lt;p&gt;A generic chatbot creates another untracked conversation. The workflow is the product.&lt;/p&gt;
&lt;h2&gt;FAQ&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;How should an AI tutoring product be designed?&lt;/strong&gt;&lt;br&gt;Tie each session to a real teacher workflow and anchor it to the student’s goal. Capture tutor notes and the next action.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Why does a generic tutoring chatbot create a problem?&lt;/strong&gt;&lt;br&gt;A generic chatbot creates another untracked conversation. The session needs to connect to a teacher workflow so it can carry forward.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What makes the workflow the product?&lt;/strong&gt;&lt;br&gt;The workflow connects the teacher’s work, the student’s goal, the tutor’s notes, and the next action. Those connections define the tutoring product.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/a-growing-market-gives-you-permission-to-test/</id>
    <title type="text">a growing market gives you permission to test</title>
    <updated>2026-07-24T00:59:00+09:00</updated>
    <published>2026-07-24T00:59:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/a-growing-market-gives-you-permission-to-test/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">A growing market permits testing, while changed learner behaviour in one narrow workflow provides the real demand signal.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; A growing market gives you permission to test. Demand becomes clearer when one learner makes a better decision this week and behaviour changes.&lt;/p&gt;
&lt;h2&gt;what does a growing market actually tell you?&lt;/h2&gt;
&lt;p&gt;It tells you there is enough movement to justify a test. Market excitement creates permission to investigate an opportunity.&lt;/p&gt;&lt;p&gt;That signal still leaves the demand question open. Interest around a market does not show whether a learner will make a better decision in practice.&lt;/p&gt;
&lt;h2&gt;where should you look for the demand signal?&lt;/h2&gt;
&lt;p&gt;Pick one narrow workflow. Keep the question close to a real decision a learner needs to make.&lt;/p&gt;&lt;p&gt;The useful test is whether that learner makes a better decision this week. A focused workflow gives the test a clear place to look.&lt;/p&gt;
&lt;h2&gt;what receipt matters after the test?&lt;/h2&gt;
&lt;p&gt;The receipt is changed behaviour. A learner making a better decision provides evidence that the workflow affected what they do.&lt;/p&gt;&lt;p&gt;Market excitement is cheap because it can exist without action. Changed behaviour carries the demand signal forward.&lt;/p&gt;
&lt;h2&gt;how should builders interpret early excitement?&lt;/h2&gt;
&lt;p&gt;Treat excitement as permission to test, then move quickly toward observable behaviour. The question is whether the learner&amp;#39;s decision improves within the chosen workflow.&lt;/p&gt;&lt;p&gt;That keeps the finding tied to use rather than sentiment. The market opens the door, and changed behaviour tells you whether the workflow matters.&lt;/p&gt;
&lt;h2&gt;FAQ&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;does a growing market prove demand?&lt;/strong&gt;&lt;br&gt;A growing market gives you permission to test. It does not answer whether a learner will make a better decision in a real workflow.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;what is a useful early demand signal?&lt;/strong&gt;&lt;br&gt;Choose one narrow workflow and observe whether a learner makes a better decision this week. Changed behaviour is the useful receipt.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;why focus on one narrow workflow?&lt;/strong&gt;&lt;br&gt;A narrow workflow keeps the test tied to a specific learner decision. That makes changed behaviour easier to observe.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/a-growing-market-is-permission-to-run-a-narrow-t/</id>
    <title type="text">a growing market is permission to run a narrow test, nothing more</title>
    <updated>2026-07-23T00:59:00+09:00</updated>
    <published>2026-07-23T00:59:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/a-growing-market-is-permission-to-run-a-narrow-t/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">A field note on testing one learner workflow and measuring whether it improves the next decision this week.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; A growing market only gives permission to run a narrow test. The useful evidence is whether one workflow improves a learner’s next decision this week.&lt;/p&gt;
&lt;h2&gt;what does a growing market actually give you permission to do?&lt;/h2&gt;
&lt;p&gt;It gives you permission to run a narrow test. Market growth creates room to investigate one workflow without treating enthusiasm as proof.&lt;/p&gt;&lt;p&gt;The test should stay close to a learner’s next decision. That keeps the question concrete and bounded.&lt;/p&gt;
&lt;h2&gt;what evidence matters more than market enthusiasm?&lt;/h2&gt;
&lt;p&gt;Customer evidence matters more than the feeling created by market enthusiasm. The relevant signal comes from what happens in a learner’s workflow.&lt;/p&gt;&lt;p&gt;The question is whether the workflow improves the learner’s next decision this week. That is the finding worth carrying forward.&lt;/p&gt;
&lt;h2&gt;what would you test first?&lt;/h2&gt;
&lt;p&gt;Start with one workflow that could affect a learner’s next decision this week. Keep the test narrow enough that the result speaks to that workflow.&lt;/p&gt;&lt;p&gt;The first test is valuable when it produces customer evidence. The market can create permission to test, while the learner’s next decision provides the measure.&lt;/p&gt;
&lt;h2&gt;FAQ&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;How should you respond to a growing market?&lt;/strong&gt;&lt;br&gt;Treat it as permission to run a narrow test. The market itself does not answer whether a workflow improves a learner’s next decision.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What should a learner workflow test measure?&lt;/strong&gt;&lt;br&gt;Measure whether it improves the learner’s next decision this week. That is the useful question in the field note.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Is market enthusiasm enough to validate a workflow?&lt;/strong&gt;&lt;br&gt;Market enthusiasm feels good, while customer evidence gives a more useful signal. The evidence should come from the learner’s workflow and next decision.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/the-uk-is-putting-up-to-23m-behind-safe-evidence/</id>
    <title type="text">the uk is putting up to £23m behind safe, evidence-based AI and</title>
    <updated>2026-07-22T00:59:00+09:00</updated>
    <published>2026-07-22T00:59:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/the-uk-is-putting-up-to-23m-behind-safe-evidence/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">Why evidence-based AI and edtech trials need measurable outcomes and safety controls before schools adopt them.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; the uk putting up to £23m behind safe, evidence-based AI and edtech trials in schools is a clear signal. novelty gets attention, while measurable outcomes and safety controls help regulated buyers move toward adoption.&lt;/p&gt;
&lt;h2&gt;what is the uk funding telling builders about school AI?&lt;/h2&gt;
&lt;p&gt;the uk is putting up to £23m behind safe, evidence-based AI and edtech trials in schools. that is the signal: school adoption is being shaped around evidence and safety.&lt;/p&gt;&lt;p&gt;for builders, the path into regulated education buyers runs through measurable outcomes and safety controls.&lt;/p&gt;
&lt;h2&gt;why do regulated buyers want evidence before adoption?&lt;/h2&gt;
&lt;p&gt;regulated buyers want measurable outcomes and safety controls before adoption. the requirement is tied to whether a product can show results and operate safely in a school setting.&lt;/p&gt;&lt;p&gt;that changes the conversation around AI and edtech from interest in what is new to confidence in what has been tested.&lt;/p&gt;
&lt;h2&gt;does novelty still matter when selling AI to schools?&lt;/h2&gt;
&lt;p&gt;novelty gets you noticed. it can create initial attention around a product or trial.&lt;/p&gt;&lt;p&gt;evidence gets you through the door. measurable outcomes and safety controls give regulated buyers the basis for adoption.&lt;/p&gt;
&lt;h2&gt;what should solo builders take from this funding signal?&lt;/h2&gt;
&lt;p&gt;the transferable lesson is direct: build for evidence and safety alongside the product itself. a novel idea may start the conversation, while proof and controls support the next conversation with a regulated buyer.&lt;/p&gt;&lt;p&gt;the funding signal points toward AI and edtech trials where outcomes can be measured and safety can be addressed clearly.&lt;/p&gt;
&lt;h2&gt;FAQ&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;what is the uk funding for safe ai and edtech trials in schools?&lt;/strong&gt;&lt;br&gt;the uk is putting up to £23m behind safe, evidence-based AI and edtech trials in schools. the source post presents this as a signal about what regulated buyers want.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;what do schools and regulated buyers want before adopting ai?&lt;/strong&gt;&lt;br&gt;they want measurable outcomes and safety controls before adoption. evidence helps a product move from attention toward the door.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;is novelty enough to get ai adopted in schools?&lt;/strong&gt;&lt;br&gt;novelty gets a product noticed. evidence, measurable outcomes, and safety controls help it get through the door.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/ai-reaches-families-when-proof-and-safety-ship-w/</id>
    <title type="text">AI reaches families when proof and safety ship with the product</title>
    <updated>2026-07-19T00:59:00+09:00</updated>
    <published>2026-07-19T00:59:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/ai-reaches-families-when-proof-and-safety-ship-w/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">Why builders in high-trust domains need proof and safety controls in the product before distribution.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; AI reaches families when proof and safety ship with the product. In high-trust domains, builders need progress evidence before distribution and should measure before they ship broadly.&lt;/p&gt;
&lt;h2&gt;why does proof need to ship with the product?&lt;/h2&gt;
&lt;p&gt;Families engage with AI when the product carries visible evidence of progress and safety. The finding is simple: distribution depends on more than the product experience alone.&lt;/p&gt;&lt;p&gt;Builders in high-trust domains need a receipt that shows what has been measured before wider distribution begins.&lt;/p&gt;
&lt;h2&gt;when should safety controls enter the product?&lt;/h2&gt;
&lt;p&gt;Safety controls belong in the product from the start. They shape how the product earns trust as it reaches families.&lt;/p&gt;&lt;p&gt;Adding safeguards early keeps proof and safety connected to the product itself, where users encounter them.&lt;/p&gt;
&lt;h2&gt;what should builders measure before they distribute?&lt;/h2&gt;
&lt;p&gt;The source question is the operating prompt: what are you measuring before you distribute? Progress evidence and safety should travel together when the product moves toward wider use.&lt;/p&gt;&lt;p&gt;Ship the receipt with the safeguards so the evidence behind distribution is part of the product experience.&lt;/p&gt;
&lt;h2&gt;FAQ&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;How can AI products earn trust with families?&lt;/strong&gt;&lt;br&gt;Ship proof and safety with the product. Families need progress evidence and safeguards to be present as the product reaches them.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;When should safety controls be added to an AI product?&lt;/strong&gt;&lt;br&gt;Safety controls belong in the product from the start. They should develop alongside the product rather than appearing only at distribution.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What should builders measure before distributing AI products?&lt;/strong&gt;&lt;br&gt;Measure progress evidence and safety before distribution. The receipt of what has been measured should ship with the safeguards.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/start-with-the-tutor-notes/</id>
    <title type="text">start with the tutor notes</title>
    <updated>2026-07-18T00:59:00+09:00</updated>
    <published>2026-07-18T00:59:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/start-with-the-tutor-notes/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">AI tutoring becomes useful when tutor notes, learner goals, session evidence, and one next action stay connected.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; AI tutoring becomes useful when a session starts with tutor notes, adds the learner’s goal and evidence, and ends with one defined next action. A generic chatbot creates another untracked conversation.&lt;/p&gt;
&lt;h2&gt;why should an AI tutoring session start with tutor notes?&lt;/h2&gt;
&lt;p&gt;Tutor notes give the session a starting point. They anchor the conversation in what the tutor already knows before the AI responds.&lt;/p&gt;&lt;p&gt;That context helps the session stay connected to the learner’s work instead of becoming another untracked conversation.&lt;/p&gt;
&lt;h2&gt;what should the session add after the tutor notes?&lt;/h2&gt;
&lt;p&gt;Add the learner’s goal and the evidence from the session. The goal shows what the learner is trying to do, while the evidence shows what actually appeared in the session.&lt;/p&gt;&lt;p&gt;Together, they give the tutoring conversation a clear reference point.&lt;/p&gt;
&lt;h2&gt;how does an AI tutoring session become useful?&lt;/h2&gt;
&lt;p&gt;End by writing one defined next action. A specific next action gives the session somewhere to go after the conversation ends.&lt;/p&gt;&lt;p&gt;That sequence is what makes AI tutoring useful: tutor notes, the learner’s goal, session evidence, and one defined next action.&lt;/p&gt;
&lt;h2&gt;FAQ&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;How can AI tutoring become useful?&lt;/strong&gt;&lt;br&gt;Start with tutor notes, add the learner’s goal and the evidence from the session, then end with one defined next action. This gives the conversation a clear structure.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What should an AI tutoring session include?&lt;/strong&gt;&lt;br&gt;It should include tutor notes, the learner’s goal, evidence from the session, and one defined next action.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Why does a generic chatbot conversation feel untracked?&lt;/strong&gt;&lt;br&gt;A generic chatbot creates another conversation without the structure provided by tutor notes, a learner goal, session evidence, and a defined next action.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/an-accelerator-calendar-is-useful-when-it-makes-/</id>
    <title type="text">an accelerator calendar is useful when it makes the user result</title>
    <updated>2026-07-17T00:59:00+09:00</updated>
    <published>2026-07-17T00:59:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/an-accelerator-calendar-is-useful-when-it-makes-/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">A field note on using accelerator deadlines to force clearer product evidence and user results at every milestone.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; An accelerator calendar earns its value when every deadline forces a clearer product question: what changed for the user before the next milestone? Ambition gets you onto the calendar, while evidence gets you through it.&lt;/p&gt;
&lt;h2&gt;what should every accelerator deadline force us to ask?&lt;/h2&gt;
&lt;p&gt;Every deadline should force a product question: what changed for the user before the next milestone? The calendar becomes useful when it makes the user result clearer.&lt;/p&gt;&lt;p&gt;That question keeps each milestone tied to movement in the product and the experience of the people using it.&lt;/p&gt;
&lt;h2&gt;what gets a product onto the calendar in the first place?&lt;/h2&gt;
&lt;p&gt;Ambition gets a product onto the calendar. A strong idea can create momentum and earn a place in the process.&lt;/p&gt;&lt;p&gt;The calendar gives that ambition a sequence of milestones where progress has to become visible.&lt;/p&gt;
&lt;h2&gt;what actually gets a product through the calendar?&lt;/h2&gt;
&lt;p&gt;Evidence gets a product through the calendar. Each deadline should leave a clearer answer about what changed for the user.&lt;/p&gt;&lt;p&gt;That standard may feel too strict, yet it keeps the focus on user results instead of calendar activity.&lt;/p&gt;
&lt;h2&gt;FAQ&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;How should an accelerator calendar be used?&lt;/strong&gt;&lt;br&gt;Use each deadline to ask what changed for the user before the next milestone. The calendar is useful when it makes the result clearer.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What matters more than ambition in an accelerator?&lt;/strong&gt;&lt;br&gt;Ambition gets you onto the calendar. Evidence gets you through it, because each milestone needs a clearer connection to user results.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/are-education-products-ready-for-this-kind-of-ju/</id>
    <title type="text">are education products ready for this kind of jump?</title>
    <updated>2026-07-14T00:59:00+09:00</updated>
    <published>2026-07-14T00:59:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/are-education-products-ready-for-this-kind-of-ju/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">Japan’s Japanese-language institute enrollment rose 23.5% in 2025, showing why education products must support cross-border transitions.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; Japan recorded 140,174 students in Japanese-language institutes in 2025, up 23.5% from the previous year. That jump shows how quickly cross-border demand can move and puts pressure on products serving transitions between systems.&lt;/p&gt;
&lt;h2&gt;what does this jump in Japanese-language students tell builders?&lt;/h2&gt;
&lt;p&gt;Japan recorded 140,174 students in Japanese-language institutes in 2025. Enrollment rose 23.5% from the previous year, showing that cross-border demand is moving quickly.&lt;/p&gt;&lt;p&gt;For builders and founders, that growth is a signal to examine how products handle movement between systems. Demand can shift faster than the product experience around it.&lt;/p&gt;
&lt;h2&gt;where do education products break during a cross-border transition?&lt;/h2&gt;
&lt;p&gt;The pressure appears wherever a learner moves between systems. A product may serve one environment well while leaving the transition between environments unsupported.&lt;/p&gt;&lt;p&gt;The source points to a practical question for product teams: where does the current product break first when cross-border demand arrives?&lt;/p&gt;
&lt;h2&gt;why should product teams look at transitions before demand peaks?&lt;/h2&gt;
&lt;p&gt;The 23.5% increase creates a clear receipt for the speed of change. Education products need to support the transition between systems as demand moves across borders.&lt;/p&gt;&lt;p&gt;That makes transition support a product question for builders, founders, and AI engineers. The first break may reveal where the product needs to adapt to the people it serves.&lt;/p&gt;
&lt;h2&gt;FAQ&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;How many students were recorded in Japanese-language institutes in Japan in 2025?&lt;/strong&gt;&lt;br&gt;Japan recorded 140,174 students in Japanese-language institutes in 2025.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;How much did Japanese-language institute enrollment increase in 2025?&lt;/strong&gt;&lt;br&gt;Enrollment increased 23.5% from the previous year.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Why do education products need to support transitions between systems?&lt;/strong&gt;&lt;br&gt;Cross-border demand is moving fast. Products need to support the transition between systems as learners move across education environments.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/a-commit-is-a-hypothesis/</id>
    <title type="text">a commit is a hypothesis</title>
    <updated>2026-07-13T00:59:00+09:00</updated>
    <published>2026-07-13T00:59:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/a-commit-is-a-hypothesis/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">A commit is a hypothesis. Live verification decides whether software actually does the job people need it to do.</summary>
    <content type="html">&lt;p&gt;a commit is a hypothesis. the live check gets the final vote. I have learned this the hard way while building software solo with AI agents, where a change can arrive looking polished, coherent, and complete before it has met the place where it is supposed to work. The diff can make sense. The code can read cleanly. The commit can carry the satisfying feeling of progress. Then the live flow asks a simpler question: did the actual job happen? That question has ended more debates for me than any review comment ever could.&lt;/p&gt;&lt;p&gt;The gap shows up because code is always written against an interpretation of reality. You read the request, infer the behavior, make a change, and create a commit. An AI agent can accelerate every part of that loop. It can inspect surrounding code, suggest an implementation, write tests, and explain why the result should work. Those are useful signals. They remain signals. The product only earns its claim when someone can observe the intended result in the real flow. Until then, the commit describes a theory about what the software will do.&lt;/p&gt;&lt;p&gt;Tests can pass while the real flow falls over. That does not mean tests are worthless. Tests answer the questions they were written to answer, inside the conditions they create. The live environment carries conditions that the test may not represent: the order a person takes through the flow, the state already present, the boundary between components, or the moment when an output must become visible. A passing suite can tell me that a narrow behavior remains intact. It cannot automatically prove that the job reached a person in the way the product promises.&lt;/p&gt;&lt;p&gt;I used to let the cleanliness of a commit create too much confidence. When the implementation was small and the tests were green, I treated the work as largely complete. That habit made review feel like the finish line. The receipt that changed my definition was simpler than a metric: code can look clean and still miss the actual job. A commit with no live proof is unfinished work. It may be a good hypothesis. It may even be close. The absence of an observable result leaves the central question open.&lt;/p&gt;&lt;p&gt;The mechanism is straightforward. Software work passes through several translations. A request becomes an interpretation. That interpretation becomes code. Code becomes a tested behavior. Then that behavior has to survive the real flow and produce the result someone needs to see. Each translation can preserve the intent, or lose part of it. AI agents make the earlier translations faster, which makes the final check more important. Speed can produce a convincing artifact before it produces a verified outcome. The faster I can generate code, the more deliberately I need to ask for evidence from the live system.&lt;/p&gt;&lt;p&gt;My rule now is plain: “shipped” starts when someone can observe the result. I do not use the word as a reward for creating a commit. I use it as a statement about the world outside the editor. If I cannot verify it live, I have a commit. I keep going. That rule gives solo work a useful stopping condition because it replaces vague confidence with a visible outcome. The implementation can still need improvement later, yet the first threshold is clear: the intended job happened where it matters.&lt;/p&gt;&lt;p&gt;For readers building with AI agents, treat every generated change as a hypothesis that deserves a live check. Ask what a person should be able to observe when the work is complete. Follow the actual flow far enough to see that result. Let passing tests inform your confidence, then let the live check decide whether the task is done. This is a small shift in language that changes the work. A commit becomes evidence of effort. Live verification becomes evidence of delivery. That distinction keeps the software honest, and it keeps the builder focused on the outcome instead of the appearance of progress.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/most-people-building-ai-tutoring-think-the-hard-/</id>
    <title type="text">most people building ai tutoring think the hard part is the</title>
    <updated>2026-07-13T00:58:00+09:00</updated>
    <published>2026-07-13T00:58:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/most-people-building-ai-tutoring-think-the-hard-/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">AI tutoring earns family trust through auditable progress evidence and real safety controls before the tutoring experience itself.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; The model is not the hard part of AI tutoring. Families need on-demand evidence of what a child did and clear proof of what the tool refused to do.&lt;/p&gt;
&lt;h2&gt;why is a smart chatbot not enough for AI tutoring?&lt;/h2&gt;
&lt;p&gt;A parent is deciding whether to place their child’s afternoons inside your product. A chatbot that sounds smart does not answer the questions behind that decision.&lt;/p&gt;&lt;p&gt;They need to see what the child actually did and understand the boundaries the tool held.&lt;/p&gt;
&lt;h2&gt;what does a family need before an AI tutor becomes part of the routine?&lt;/h2&gt;
&lt;p&gt;Auditable progress evidence and real safety controls are the price of entry into a family’s routine.&lt;/p&gt;&lt;p&gt;Those requirements make the experience legible on demand. They give a parent something concrete to inspect when deciding whether the product deserves trust.&lt;/p&gt;
&lt;h2&gt;what should an AI tutoring builder build first?&lt;/h2&gt;
&lt;p&gt;Build the receipt first.&lt;/p&gt;&lt;p&gt;The receipt shows what happened during a child’s time with the tool and what the tool refused to do. Once that exists, the tutoring is the easy half.&lt;/p&gt;
&lt;h2&gt;FAQ&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;What makes an AI tutoring product trustworthy to parents?&lt;/strong&gt;&lt;br&gt;Parents need on-demand visibility into what their child actually did. They also need real safety controls and evidence of the boundaries the tool enforced.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Is the model the hardest part of building AI tutoring?&lt;/strong&gt;&lt;br&gt;The model is not the hard part described here. The harder requirement is earning a place in a family’s routine through evidence and safety controls.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What does “build the receipt first” mean for AI tutoring?&lt;/strong&gt;&lt;br&gt;It means building the evidence a parent can inspect on demand. That evidence covers the child’s activity and what the tool refused to do.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/git-said-the-feature-shipped/</id>
    <title type="text">git said the feature shipped</title>
    <updated>2026-07-11T00:59:00+09:00</updated>
    <published>2026-07-11T00:59:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/git-said-the-feature-shipped/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">A commit is halfway: software ships when you watch a customer complete the intended action live.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; Git said the feature shipped. The customer’s screen said otherwise. A commit is halfway because shipping ends when a user can actually do the thing.&lt;/p&gt;
&lt;h2&gt;when is a feature actually shipped?&lt;/h2&gt;
&lt;p&gt;A feature is shipped when a user is doing the thing it was built for.&lt;/p&gt;&lt;p&gt;Merged code, green CI, and a closed ticket can still leave the customer unable to use it.&lt;/p&gt;
&lt;h2&gt;why does repo activity feel like progress?&lt;/h2&gt;
&lt;p&gt;Repo activity produces visible signals: commits, merges, CI, and ticket movement.&lt;/p&gt;&lt;p&gt;Those signals describe work moving through the repository. They do not show whether the customer’s screen works.&lt;/p&gt;
&lt;h2&gt;what did the customer screen reveal that Git missed?&lt;/h2&gt;
&lt;p&gt;Git said the feature shipped. The customer’s screen said otherwise.&lt;/p&gt;&lt;p&gt;That gap is the receipt: repository status cannot stand in for watching the feature work live.&lt;/p&gt;
&lt;h2&gt;what should count as done on a team?&lt;/h2&gt;
&lt;p&gt;Done should include seeing the intended user action work live.&lt;/p&gt;&lt;p&gt;A commit is halfway. The remaining half is the customer experience.&lt;/p&gt;
&lt;h2&gt;FAQ&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Does green CI mean a feature is shipped?&lt;/strong&gt;&lt;br&gt;Green CI shows that CI is green. It does not show that a customer can complete the intended action on their screen.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Is a merged commit considered done?&lt;/strong&gt;&lt;br&gt;A merged commit is halfway. The work reaches done when a user is doing the thing the feature was meant to enable.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Why can a closed ticket still hide a shipping problem?&lt;/strong&gt;&lt;br&gt;Tickets can close after repository work is complete. The customer’s screen can still reveal that the feature does not work live.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/a-commit-is-just-a-checkpoint/</id>
    <title type="text">a commit is just a checkpoint</title>
    <updated>2026-07-10T00:59:00+09:00</updated>
    <published>2026-07-10T00:59:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/a-commit-is-just-a-checkpoint/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">A commit is only a checkpoint; i call work done after verifying it in the live environment the user actually experiences.</summary>
    <content type="html">&lt;p&gt;a commit is just a checkpoint. the line i care about is farther down the road, and i only call something done after i verify it in the live environment the user actually experiences. that sounds small, but it changes the whole shape of the work. a commit records progress. it does not prove the thing is finished for the person who will live with it. if the only proof comes from the repo, i have a strong signal that i am still inside the comfort of my own workspace instead of the place where the work has to stand on its own.&lt;/p&gt;&lt;p&gt;the distinction matters because the repo can feel clean in ways the real world never will. local confidence is easy to build. the code looks tidy, the tests pass, the branches line up, and the mental model feels complete. then the user touches the thing and the story changes. that is the part i want to catch early. i have learned that a polished repo can create a false win, and false wins are expensive. they make you believe the hard part is over while the actual use case is still waiting to expose the gap.&lt;/p&gt;&lt;p&gt;this is the part i keep coming back to: founders ship too many things that only work in repo. i have seen enough of that pattern to stop treating repo success as the finish line. inside the repo, everything is framed by what i already know. the edges are familiar, the setup is under my control, and the environment is shaped by my assumptions. once the thing crosses into the live environment, those assumptions get tested by reality. that is where hidden friction shows up. that is where the work earns its name.&lt;/p&gt;&lt;p&gt;the mechanism is simple. a repo is an edited view of the system, and a live environment is the system under real conditions. the repo rewards internal consistency. the live environment rewards usefulness. those are related, but they are not the same thing. if i only check the first, i can end up with a solution that feels finished while still missing the conditions that matter to the user. when i verify in the place the user actually experiences it, i get a cleaner answer about whether the thing holds up in practice.&lt;/p&gt;&lt;p&gt;so i adopted a rule: real-world done is the standard. i do not let the commit define completion for me. i let the user-facing result define it. that one shift keeps me honest when a task seems complete inside the repo but still needs a final pass where the experience actually happens. it also changes how i think about progress. a commit becomes a checkpoint on the way to done, a marker that says work moved forward, while the live verification says the work actually reached the point where it matters.&lt;/p&gt;&lt;p&gt;that rule saves pain because it catches the false win early. the earlier i find the gap, the less time i spend building confidence around the wrong thing. the lesson is useful anywhere software leaves your hands and enters someone else’s experience. if you build solo with AI agents, it is especially easy to confuse internal order with external success, because the whole process can look productive long before it is genuinely finished. the fix is to ask a better question at the end of every task: what is the actual line for done, and have i checked it where the user feels it?&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/a-green-diff-lied-to-me-last-week/</id>
    <title type="text">a green diff lied to me last week</title>
    <updated>2026-07-09T00:59:00+09:00</updated>
    <published>2026-07-09T00:59:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/a-green-diff-lied-to-me-last-week/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">A green diff passed tests while prod stayed broken, and the fix was a rule to verify the live thing.</summary>
    <content type="html">&lt;p&gt;a green diff lied to me last week. the commit landed, tests passed, and i moved on with the kind of confidence that comes from seeing a clean result in the repo. meanwhile, the system was quietly broken in prod the whole time. i found out because a user found it first, which is the worst way to learn that something you shipped has a hole in it. that kind of discovery sticks with you. it leaves a mark on how you judge your own work, because the pleasant feeling of a passing test can turn out to be a very small part of the real story.&lt;/p&gt;&lt;p&gt;the part that stings is how ordinary the false signal looked. nothing dramatic happened in the moment of the commit. there was no obvious alarm, no loud failure, no red flag in the place i was already checking. the diff looked green, the tests looked green, and that was enough for me to move on. that is the trap. a repo can look healthy while the live system carries the actual problem. when i build software solo with AI agents, i am often moving fast and trusting the shape of the work that sits in front of me. this was a sharp reminder that the shape of the work and the behavior of the live thing can part ways.&lt;/p&gt;&lt;p&gt;the receipt was simple and painful. the commit landed. the tests passed. the product still broke in production. a user found the issue before i did. that sequence matters because it shows exactly where the blind spot lived. i had evidence that my code fit the repo’s expectations, but i had no evidence that the live thing behaved the way i wanted it to behave. the number that matters here is the gap between those two scoreboards. one says the change is accepted by the repo. the other says the real world is actually handling it. i had been spending my attention on the first one and treating it like the second one.&lt;/p&gt;&lt;p&gt;the mechanism is easy to miss when you are in the middle of shipping. tests tell you about the code path you covered. prod tells you about the thing users actually touch. those are related, and they are also separate. a passing test can mean the code compiled, the expected path worked, and the repo-level checks were satisfied. live behavior asks a different question. it asks whether the thing still works after the commit leaves your hands and meets the rest of the system. once i saw that clearly, the mistake felt less mysterious. the failure was never that the tests passed. the failure was that i let passing tests stand in for proof that the live thing worked.&lt;/p&gt;&lt;p&gt;that led to a rule i actually trust now: a change is done when i have watched the live thing actually work. the commit is a checkpoint, and the tests are a checkpoint, and neither one earns the final stamp on its own. i want the moment where the thing is real outside the repo. i want to see it behave where users will feel it. that is the point where i stop telling myself the job is finished. for my own workflow, especially with AI agents in the loop, that rule creates a hard boundary around optimism. it keeps me from confusing a clean diff with a complete result.&lt;/p&gt;&lt;p&gt;if you build software, you can use the same lesson without changing your whole process. keep your tests. keep your commit discipline. then add one more standard before you call something shipped: watch the live thing work. if you cannot do that immediately, treat the work as still open. that one habit changes how you read your own progress. it pushes you toward reality instead of paperwork. it also gives you a better answer when a change looks good in the repo and still behaves badly in the product. that gap is where the useful learning lives, and it is where the honest definition of shipped has to sit.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/spent-an-hour-last-night-chasing-a-bug-that-was-/</id>
    <title type="text">spent an hour last night chasing a bug that was never a bug</title>
    <updated>2026-07-07T00:59:00+09:00</updated>
    <published>2026-07-07T00:59:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/spent-an-hour-last-night-chasing-a-bug-that-was-/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">a stale copy of the same api key in an env file caused a machine-only failure until secrets moved to one OS-backed store.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; last night’s bug was a stale copy of the same api key. code read the stale value, so it worked on one machine and failed everywhere else.&lt;/p&gt;
&lt;h2&gt;why did this look like a bug at all?&lt;/h2&gt;
&lt;p&gt;it took an hour to chase because the failure read like code trouble. the same api key lived in three places: two env files and my head, and one of those copies had gone stale.&lt;/p&gt;&lt;p&gt;i was reading the fresh value while the code was reading the stale one. that split made the behavior look random even though the cause was plain once the mismatch was visible.&lt;/p&gt;
&lt;h2&gt;what was actually wrong with the setup?&lt;/h2&gt;
&lt;p&gt;the problem was duplicate truth. when one value exists in more than one place, it becomes hard to know which copy is real at the moment the app runs.&lt;/p&gt;&lt;p&gt;that is how something can work on my machine and fail everywhere else. the machine used the copy I was looking at, while the code followed the stale copy in the env file.&lt;/p&gt;
&lt;h2&gt;what changed after the fix?&lt;/h2&gt;
&lt;p&gt;the fix was boring on purpose. secrets moved into one OS-backed store, and env files were left for config only.&lt;/p&gt;&lt;p&gt;that gave me one canonical place to check when something breaks. it also removed the plaintext copy sitting in a file waiting to get committed.&lt;/p&gt;
&lt;h2&gt;what is the actual lesson here?&lt;/h2&gt;
&lt;p&gt;if a secret lives in two spots, it stops being a secret and starts being a guess. the hard part is usually not the code path, it is the disagreement between copies.&lt;/p&gt;&lt;p&gt;the cleaner setup is the one that leaves no doubt about where the real value lives. once there is a single source of truth, debugging gets much shorter.&lt;/p&gt;
&lt;h2&gt;FAQ&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;why did this work on my machine but fail everywhere else?&lt;/strong&gt;&lt;br&gt;because the code was reading a stale copy while you were checking a fresh one. the mismatch made the local result look fine and the rest of the environment fail.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;what was the real fix?&lt;/strong&gt;&lt;br&gt;move the secret into one OS-backed store and keep env files for config only. that removes the duplicate copy problem and gives you one place to inspect.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;why is having the same secret in two places a problem?&lt;/strong&gt;&lt;br&gt;because you lose certainty about which value the app actually uses. once copies drift, debugging turns into guesswork.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/why-are-secrets-still-living-in-env-files-and-sh/</id>
    <title type="text">why are secrets still living in env files and shell history?</title>
    <updated>2026-07-06T00:59:00+09:00</updated>
    <published>2026-07-06T00:59:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/why-are-secrets-still-living-in-env-files-and-sh/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">Centralizing secrets in one store and loading them at runtime makes rotation boring and stops drift across env files, shells, and deploy targets.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; centralize secrets in one store and load them at runtime. that keeps plaintext env files secret-free and makes rotation boring.&lt;/p&gt;
&lt;h2&gt;why do env files keep turning into secret sprawl?&lt;/h2&gt;
&lt;p&gt;because the same secret ends up copied into too many places. once it lives in env files, shell history, scripts, and deploy targets, every copy becomes another thing to keep in sync.&lt;/p&gt;&lt;p&gt;the source of the pain is manual drift. a secret store gives you one place to rotate, while plaintext env files stay secret-free and easier to reason about.&lt;/p&gt;
&lt;h2&gt;what changes when secrets load at runtime?&lt;/h2&gt;
&lt;p&gt;the app gets its values when it starts, through a per-app env loader, instead of carrying secrets around in static files. that shifts the sensitive part into one controlled path.&lt;/p&gt;&lt;p&gt;for the person maintaining the system, the important change is simple: the app still gets env-style configuration, but the secret itself no longer has to sit in a plaintext file.&lt;/p&gt;
&lt;h2&gt;why does a single rotation point matter?&lt;/h2&gt;
&lt;p&gt;rotation becomes a single action instead of a scavenger hunt. when one store feeds the apps, you can update the source and let the loaders pick it up.&lt;/p&gt;&lt;p&gt;that is what makes rotation feel boring in a good way. the system stops depending on whoever remembered to update the latest shell, script, or deploy target copy.&lt;/p&gt;
&lt;h2&gt;what actually breaks when secrets drift across tools?&lt;/h2&gt;
&lt;p&gt;manual drift creates mismatched versions of the same secret across scripts, shells, and deploy targets. that is where surprise behavior comes from.&lt;/p&gt;&lt;p&gt;the post calls that clown work for a reason. if the same value has to be maintained by hand in multiple places, the process is already fragile before any rotation happens.&lt;/p&gt;
&lt;h2&gt;FAQ&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;why keep plaintext env files secret-free?&lt;/strong&gt;&lt;br&gt;because the env file stays part of normal app configuration while the sensitive value lives elsewhere. that reduces the number of places that can drift or leak.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;what is the benefit of loading secrets at runtime?&lt;/strong&gt;&lt;br&gt;runtime loading lets the app read current values from a central store when it starts. that keeps secret handling out of long-lived static files and makes rotation simpler.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;what problem does one secret store solve?&lt;/strong&gt;&lt;br&gt;it gives you one rotation point instead of many copies spread across scripts, shells, and deploy targets. that cuts manual sync work and makes the system easier to keep consistent.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;what is the main operational win here?&lt;/strong&gt;&lt;br&gt;rotation gets boring. when one source feeds all the app-specific loaders, you spend less time chasing secret drift and more time on the actual software.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/a-merged-pr-is-not-a-ship/</id>
    <title type="text">a merged pr is not a ship</title>
    <updated>2026-07-03T00:59:00+09:00</updated>
    <published>2026-07-03T00:59:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/a-merged-pr-is-not-a-ship/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">Green CI proves the code matched expectations; live verification is the receipt that shows it actually worked.</summary>
    <content type="html">&lt;p&gt;a merged pr can look finished while the real work is still ahead. i keep seeing the same moment: the ticket gets closed the second CI goes green, and everybody acts as if the job is done. green tests matter. they tell you the code did what you asked inside the boundaries of the test suite. they do not tell you whether the change survives contact with production. that gap is the whole story. a merge is a claim. a live check is the receipt.&lt;/p&gt;&lt;p&gt;that distinction is easy to miss when you are moving fast. the status page turns green, the pull request is merged, and the brain wants to mark the task complete. i understand the pressure behind that reflex. a passing build feels concrete. it is visible, it is measurable, and it rewards the exact kind of discipline we are supposed to care about. the problem is that the thing you can point to in CI is only one part of the system. it shows that the code satisfied the conditions you defined. it says nothing about whether the real environment accepts the change in the same way.&lt;/p&gt;&lt;p&gt;this is why i started treating live verification as part of the finish line. one line now sits in my definition of done: it ran live, and i watched it work. that line changes the emotional shape of the task. merging a commit stops being the finish and becomes the point where the claim is ready to be checked. the test suite can still be green. the review can still be clean. the ticket can still move forward. the change only earns closure after it has been seen working where it matters.&lt;/p&gt;&lt;p&gt;the mechanism is simple. CI is a controlled proof. live verification is contact with the real system. controlled proof is useful because it reduces uncertainty and catches mistakes early. contact with the real system matters because reality has more moving pieces than a test suite can model. a passing test means the code behaved inside the expectations you wrote down. a live check shows whether those expectations were enough. that is the difference between confidence and confirmation. both matter. they answer different questions. one tells you the code met the promise. the other tells you the promise held up.&lt;/p&gt;&lt;p&gt;i like the language here because it keeps the relationship honest. the commit is a claim. that is all it is. it is an assertion about what should happen. the live verification is the receipt. a receipt is proof that the thing happened in the world, after the handoff, in the place where consequences are real. once i started using that framing, the process got clearer. i stopped letting merge status stand in for evidence. i stopped treating green CI as the end of the conversation. the conversation ends when the live check has happened and i have watched it work.&lt;/p&gt;&lt;p&gt;for anyone building software solo with AI agents, this rule is practical because it keeps the machine and the reality test in separate boxes. agents can help produce code quickly. they can help narrow bugs. they can help get a branch to green. none of that replaces the moment where you verify the result in the live environment. if you want a simple rule to carry forward, use this one: every merged change still needs a live check before you call it done. that habit protects you from false confidence and makes your shipping process more honest.&lt;/p&gt;&lt;p&gt;the larger lesson is that status is cheap and evidence is earned. a green build is useful, but it is only one receipt in the chain. the work is finished when the claim has been tested in the place it will actually live. if you build software for real users, keep that line in your definition of done. it is small, easy to remember, and hard to fake. that is exactly why it works.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/an-agent-budget-breaker-for-your-ai-tooling/</id>
    <title type="text">an agent budget breaker for your ai tooling</title>
    <updated>2026-07-02T00:59:00+09:00</updated>
    <published>2026-07-02T00:59:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/an-agent-budget-breaker-for-your-ai-tooling/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">a hook that hard-blocks subagent dispatches at 8 per day. installed after a 132-dispatch week burned 85% of a max plan, and what the cap actually taught me about routing.</summary>
    <content type="html">&lt;p&gt;one week in late may i burned 85% of a max 20x plan. the ledger said 132 agent dispatches in seven days. almost none of them shipped anything. each dispatch feels free in the moment because it is one tool call in a chat window. the bill disagreed. so i wrote a circuit breaker: a pre-tool-use hook that counts agent dispatches by timestamp and hard-blocks the agent tool at 8 per day, with an 80 per 7 days backstop. a second hook prepends live burn stats to every prompt, so the model reads something like &amp;#39;agent dispatches: 6/8 today&amp;#39; before it decides to spawn another one. there is a bypass: touch a sentinel file and exactly one dispatch goes through, then the sentinel is consumed.&lt;/p&gt;&lt;p&gt;the cap looks like a cost control. what it actually enforces is routing discipline. a 30-day capture from before the cap showed 111 dispatches, and every single one went to the general-purpose agent. i had 64 specialist agents installed. none of them ever fired. one day logged 243 explore dispatches crawling the same directory tree over and over. nothing about the tooling was broken. the reflex was broken. when spawning is free, you never ask whether a shell command answers faster, or whether the code-reviewer agent fits better than the generic one. at 8 per day, every dispatch has to survive the question &amp;#39;why this agent, and why an agent at all&amp;#39;.&lt;/p&gt;&lt;p&gt;the false-economy lesson came from the other direction. once dispatches felt scarce, the tempting move was to route work to whatever costs the least. i sent lead discovery to a free local model to save tokens. it invented company names with plausible-looking domains. 1 of 12 resolved to a real company. the same task through codex returned 44 of 46 real, verified companies. zero dollars for fiction is still pure waste. the metric that matters is real output divided by tokens spent. tokens alone tells you nothing, and a cheap run that produces garbage is more expensive than the run you were avoiding.&lt;/p&gt;&lt;p&gt;the breaker itself misfired once, and that failure taught as much as the install. headless cron jobs were being counted against the same pool as my interactive work. one legitimate 165-dispatch automation day locked the whole following week. the fix was to exempt headless entrypoints from the ledger, raise the weekly backstop from 50 to 80, and archive the polluted log. the general rule: when a circuit breaker fires, check what it is counting before you raise the limit. automated load polluting a human-discipline metric is a misconfiguration, and raising the cap to accommodate it just deletes the signal.&lt;/p&gt;&lt;p&gt;the whole thing is two shell scripts and a json ledger. no dashboard, no saas. if your ai tooling can spawn subagents, put a counter in front of it for one week and read the number before you decide whether you need the cap. mine said 132. i would have guessed 30. what would yours say?&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/your-outreach-needs-a-validator/</id>
    <title type="text">your outreach needs a validator</title>
    <updated>2026-07-02T00:58:00+09:00</updated>
    <published>2026-07-02T00:58:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/your-outreach-needs-a-validator/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">a hard gate for outreach drafts: single line, product link required, em dashes banned, template clones rejected. it failed 358 of the 414 drafts sitting in my queues.</summary>
    <content type="html">&lt;p&gt;a subscribe pitch in my linkedin queue burned 17 real people before i noticed the copy was identical in every message. that was vibes-based quality control: skim three drafts, they read fine, send the batch. the drafts you skim are never the ones that hurt you. so i wrote validate_outreach.py, a hard gate every draft passes through before it can send. the gate is binary. a draft that fails one check stays in the queue.&lt;/p&gt;&lt;p&gt;the structural checks came straight from observed failures. single line only, because linkedin truncates a multi-line dm and the recipient receives a naked &amp;#39;hi jason&amp;#39;. minimum 60 characters, because anything shorter is a naked greeting anyway. the product link has to be present, since a message with no way to see what i built is a dead end. em dashes are a hard ban, the fastest ai tell there is. and the exact subscribe phrasing that burned those 17 people sits in a banned list, checked against a squashed form of the text so &amp;#39;sub scribe&amp;#39; and &amp;#39;news-letter&amp;#39; cannot sneak past.&lt;/p&gt;&lt;p&gt;cold openers get stricter rules. the first message to a stranger has to ask about their world, and it cannot demand their time. &amp;#39;hop on a quick call&amp;#39; fails. a regex catches the time-ask shapes a plain substring list missed, like &amp;#39;could we set up some time to talk&amp;#39;. the earliest version required a literal question mark in every opener. that rule produced monotone drafts where every message asked a formula question, so it got loosened: the gate now rejects the blast shape, meaning any draft that says nothing about the recipient and asks nothing at all. the ask still has to be there. it just gets to sound human.&lt;/p&gt;&lt;p&gt;the sneakiest failure is the name-swapped template clone. each draft reads fine on its own. so the gate compares every draft in a batch against every other draft using token overlap, and past 0.7 similarity both drafts get flagged and both fail. a human reviewer approves clones all day because no single message looks wrong. the batch view is the only place the pattern is visible.&lt;/p&gt;&lt;p&gt;the receipt: i ran the gate over my frozen queues today. 414 drafts. 56 passed. 358 failed. the breakdown was 266 missing the product link, 259 flagged as template clones, 80 carrying the exact subscribe pitch, 16 multi-line, 6 with em dashes. the old pipeline would have sent all 414. that is 358 small burns to people i actually wanted to talk to, prevented by about 250 lines of python that took an afternoon.&lt;/p&gt;&lt;p&gt;a validator turns taste into checks. every time a draft embarrasses you, the failure becomes one more line in the banned list, and that class of mistake never ships again. reviewers get tired by draft 30. a regex reads draft 414 with the same attention it gave draft 1. what is the one failure your last outreach batch shipped that a check like this would have caught?&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/the-real-bar-is-live-verification/</id>
    <title type="text">the real bar is live verification</title>
    <updated>2026-07-01T00:59:00+09:00</updated>
    <published>2026-07-01T00:59:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/the-real-bar-is-live-verification/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">a passing test can still miss the real user experience; live verification in the real environment is the standard that matters.</summary>
    <content type="html">&lt;p&gt;the real bar is live verification. a passing test can leave you with fake confidence, and merged code can still miss what users actually feel. that gap is the whole problem. it is easy to look at green checks and feel settled. it is harder to remember that the place where the code lands is the place that decides whether the change actually works. the lesson is plain: if users feel it somewhere, verify it there.&lt;/p&gt;&lt;p&gt;this matters because software can look correct in a controlled setting and still behave differently once it meets the real environment. a test can pass and still protect the wrong thing. merged code can be accepted and still fail to match the lived experience on the other side of the change. when i keep this in view, i stop treating the test result as the finish line. i treat it as one signal, useful and incomplete. the real question becomes whether the behavior survives contact with the environment where it will actually matter.&lt;/p&gt;&lt;p&gt;the reason fake confidence shows up so easily is that verification inside a controlled setup is often cheaper than verification in the place where the change lands. that makes the controlled setup tempting. it is neat, fast, and comforting. but comfort is exactly where the mistake hides. a passing test says the code behaved the way the test expected. it does not say the same thing about the real environment. the user does not experience the test. the user experiences the environment. that difference is where surprise lives.&lt;/p&gt;&lt;p&gt;once i accepted that, the rule got simple enough to use. if something matters to users, i verify it where users feel it. that rule changes the way i think about confidence. it shifts attention from internal proof to external proof. it asks for live verification before I let myself feel done. it also keeps me honest about what a test can do. tests still matter. merged code still matters. but neither gets to pretend it has answered the full question when the environment is where the change lands.&lt;/p&gt;&lt;p&gt;for solo software work with AI agents, this is especially useful because the pace can make certainty feel closer than it is. you can move quickly, combine code, and see a clean result inside the workflow you built. that speed is part of the value, and it also makes the gap easier to miss. live verification brings the work back down to the place that counts. it asks for contact with the real setting, where the behavior is no longer an abstract success case but an actual experience. that is where a finding becomes real.&lt;/p&gt;&lt;p&gt;the practical takeaway is to treat the environment as the source of truth for anything user-facing. if users feel a change somewhere, verify it there. if a passing test gives comfort, follow it with live verification. if merged code looks good on paper, check what happens where the code actually runs. this is a small rule, but it has a strong effect on how you build. it keeps you from mistaking internal correctness for delivered value. it turns confidence into something earned in the same place the user will meet the change.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/moved-every-runtime-secret-into-the-os-keychain-/</id>
    <title type="text">moved every runtime secret into the OS keychain and ripped the</title>
    <updated>2026-06-26T00:59:00+09:00</updated>
    <published>2026-06-26T00:59:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/moved-every-runtime-secret-into-the-os-keychain-/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">moving runtime secrets into one canonical OS keychain made rotation and recovery a one-place problem across four services.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; moving runtime secrets into the OS keychain changed secret handling from copy hunting to one-place rotation and recovery. the real win was knowing exactly where to look if something leaks.&lt;/p&gt;
&lt;h2&gt;why did plaintext env files break down so fast?&lt;/h2&gt;
&lt;p&gt;the problem showed up as soon as the same key lived in four places. rotating one copy meant hunting down every other copy and hoping none were missed.&lt;/p&gt;&lt;p&gt;that workflow turns a simple change into a scavenger hunt. once the copies spread, you stop trusting that you know where the secret exists.&lt;/p&gt;
&lt;h2&gt;what changed once there was one canonical store?&lt;/h2&gt;
&lt;p&gt;moving every runtime secret into the OS keychain changed the shape of the problem. the secret stopped being a scattered set of plaintext copies and became something with one source of truth.&lt;/p&gt;&lt;p&gt;from there, rotation got simpler because there was one place to update and one place to recover from. that also made leak response clearer, since you know where to look first.&lt;/p&gt;
&lt;h2&gt;why does copy count matter more than people think?&lt;/h2&gt;
&lt;p&gt;the post points to four services, which is already enough to create drift. if you run more than one service, the spread is usually worse than it feels.&lt;/p&gt;&lt;p&gt;the useful question is not whether you have secrets somewhere. it is how many copies exist and whether you could find them all quickly if one leaked.&lt;/p&gt;
&lt;h2&gt;what is the practical takeaway for builders?&lt;/h2&gt;
&lt;p&gt;the takeaway is small and direct: stop treating secret storage as a per-service habit and start treating it as one canonical store. the fewer plaintext copies you keep around, the less you have to hunt later.&lt;/p&gt;&lt;p&gt;this is less about tooling preference and more about reducing uncertainty. one store gives you one place to rotate, one place to recover, and one place to investigate.&lt;/p&gt;
&lt;h2&gt;FAQ&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;how many services was the key copied across?&lt;/strong&gt;&lt;br&gt;the post says it was copied across four services. that was enough to make rotation a hunt for every copy.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;what was the main benefit of one canonical store?&lt;/strong&gt;&lt;br&gt;one place to rotate, one place to recover, and one place to check if something leaks. that was the change in the post.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;why move runtime secrets out of env files?&lt;/strong&gt;&lt;br&gt;because plaintext env files made the secret easy to duplicate and hard to track. once the copies spread, rotation became fragile.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/a-commit-is-a-checkpoint/</id>
    <title type="text">a commit is a checkpoint</title>
    <updated>2026-06-25T00:59:00+09:00</updated>
    <published>2026-06-25T00:59:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/a-commit-is-a-checkpoint/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">A commit is a checkpoint, and shipping only starts after live verification exposes the hidden failures.</summary>
    <content type="html">&lt;p&gt;a commit is a checkpoint, and that distinction matters more than people admit. a commit means the work reached a point where it can be reviewed, stored, and moved forward. it does not mean the work has earned the right to be called done. i keep coming back to that because the gap between those two ideas is where a lot of painful surprises live. the code can look finished inside the editor and still fail the moment it has to deal with the real world. that is the part people skip, and it is the part that costs time later.&lt;/p&gt;&lt;p&gt;the mistake is easy to make because the surface area of completion feels convincing. the changes are in place, the merge goes through, and the project gets the emotional relief that comes with seeing movement. at that moment, it is tempting to treat the job as closed. i have seen how seductive that feeling is. once the merge lands, the mind wants to move on. the problem is that the merge only says the work entered the system. it does not say the system accepted it under live conditions. that gap is where hidden failures stay hidden.&lt;/p&gt;&lt;p&gt;what shows up next is usually the part nobody wants to discover. the real world arrives and drags the whole thing into the dirt. that line sounds harsh because it is harsh. live conditions are where assumptions get tested, where paths that seemed obvious turn out to be brittle, and where the thing that looked complete suddenly has to prove it can survive contact with actual use. inside a controlled setup, code can look clean and stable. in live verification, the work either behaves or it reveals the places where it was only appearing to behave. that is why the failures feel hidden. they were always there, waiting for the right pressure.&lt;/p&gt;&lt;p&gt;the mechanism is simple. local confidence is cheap, but real confidence has to be earned in the same environment where the work will actually run. a commit can show that the changes exist. a merge can show that the changes fit into the branch history. live verification shows whether the result stands up where it matters. that difference is the whole point. hidden failures are hidden because they sit outside the narrow path that was exercised during development. once the work is verified live, those failures stop being hidden. they become visible facts, and visible facts are easier to fix than assumptions you have been carrying around as certainty.&lt;/p&gt;&lt;p&gt;that is why i adopted the rule: verify live before you call it done. it sounds almost too plain to need saying, which is probably why it keeps getting skipped. people know the work is close, and close can fool everyone into compressing the final step. i have done that myself. the rule exists to interrupt that reflex. it creates a hard boundary between progress and shipping. progress can include commits, merges, and partial wins. shipping begins when the live system has been checked and the result holds up there too. anything before that is still a checkpoint, even if it feels finished.&lt;/p&gt;&lt;p&gt;the receipt is the moment live verification either confirms the work or exposes what still needs attention. that is the proof that the rule is real. it is easy to talk about completion in the abstract. it is harder to accept that completion depends on passing through the live environment first. once you do that a few times, the lesson stops being theoretical. you start to trust the checkpoint for what it is and stop granting it more meaning than it deserves. that shift changes how you work because it keeps you honest about the state of the system. if you are building software, especially on your own, this is one of the cleanest habits you can adopt. treat the commit as a checkpoint, make live verification part of the finish line, and let the real world have the final say.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/when-is-a-change-actually-done/</id>
    <title type="text">when is a change actually done?</title>
    <updated>2026-06-24T00:59:00+09:00</updated>
    <published>2026-06-24T00:59:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/when-is-a-change-actually-done/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">A change is done when the live service is verified end to end, because repo-local green can create false confidence.</summary>
    <content type="html">&lt;p&gt;when is a change actually done? for me, the answer is simple: when the live service is verified end to end. everything before that is a signal, and a commit stays a commit, nothing more. that line matters because it keeps the definition of done anchored to the place where the work has to survive. a green repo can feel finished in the moment, while the real service is still the only place that can say whether the change holds up.&lt;/p&gt;&lt;p&gt;the mistake I keep watching for is treating local success as final. repo-local green is where the false confidence starts. the code can look clean, the checks can pass, and the commit can feel ready to ship. that feeling is real, and it is also too early. the gap appears when the work leaves the repo and has to behave as the live service, with the actual end to end path in play.&lt;/p&gt;&lt;p&gt;the mechanism is plain enough. local verification covers the code where it was edited. end to end verification covers the thing people actually use. those are different surfaces, and they fail in different ways. a commit can be correct in isolation and still leave the live service unverified. when that happens, the commit has value, but it does not carry the meaning people often assign to it. it records a change. it does not certify completion.&lt;/p&gt;&lt;p&gt;that shift in language changed how I think about progress. I still value commits, and I still want the repo to be green. those are useful checkpoints. they mark that the work is moving in the right direction. the rule I adopted is that none of them gets promoted to done on its own. done belongs to the verified live service, because that is the point where the change stops being an idea and becomes something I can stand behind.&lt;/p&gt;&lt;p&gt;this also changes how I read my own confidence. if I feel settled after the repo goes green, that is exactly the moment to slow down and ask what has actually been proven. have I only proven the code in place, or have I proven the service in motion. that question keeps me honest. it keeps me from confusing a neat local result with a finished outcome. it also makes the remaining work easier to see, because the missing step is usually the real one.&lt;/p&gt;&lt;p&gt;if you build software solo with AI agents, this rule is worth keeping close. agents can produce fast progress, and they can also make local success look more complete than it is. a clean commit can arrive early. the live service still gets the final say. so my advice is to define done in a way that forces the work to meet reality. ask whether the live service has been verified end to end. if it has, you have a real finish line. if it has not, you have a commit and a useful next step, and that distinction matters.&lt;/p&gt;&lt;p&gt;so the question I keep putting back to myself is the same one I put to anyone reading this: what are you using as the real definition of done. if the answer lives inside the repo, you are probably stopping too early. if the answer lives in the live service, verified end to end, then the work has a chance to mean what it says. that is the standard I try to hold, because it keeps the story of the change aligned with the thing that actually ships.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/your-offline-backup-is-not-your-source-of-truth/</id>
    <title type="text">your offline backup is not your source of truth</title>
    <updated>2026-06-23T00:59:00+09:00</updated>
    <published>2026-06-23T00:59:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/your-offline-backup-is-not-your-source-of-truth/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">A backup you write to is just a second thing to keep in sync; a backup you only read from in a disaster actually saves you.</summary>
    <content type="html">&lt;p&gt;your offline backup is not your source of truth. it is your insurance. that distinction sounds small until you watch a local encrypted mirror slowly take on a life of its own. people make a copy of critical data, feel safer because the copy lives somewhere else, and then start treating that copy like the real system. once that happens, the backup stops being a quiet safety net and starts becoming another place where decisions can drift.&lt;/p&gt;&lt;p&gt;the problem is simple and expensive. when two places both feel important, they begin to disagree. a change lands in one place, another change lands in the mirror, and now there is no clean answer to which version deserves trust. the moment you start writing to the mirror, you have created two active records that need to stay aligned. every edit adds one more chance for mismatch, and every mismatch makes the mirror harder to trust when you actually need it.&lt;/p&gt;&lt;p&gt;the rule i use now is direct. pick one canonical source and keep it that way. the mirror lives somewhere else, gets refreshed on a schedule, and stays quiet until the primary is gone. that gives each copy a job. one place is where work happens. the other place is where recovery happens. when those roles stay separate, the backup remains simple enough to trust.&lt;/p&gt;&lt;p&gt;the mechanism matters here. a backup that you write to is a second thing to keep in sync. that means every update has to be correct in both places, every time, or the two copies drift apart. once drift starts, the mirror becomes a competing version of the truth. by contrast, a backup that you only read from in a disaster never has to compete with day-to-day edits. it can stay stale on purpose, because its job is to exist when the primary is gone.&lt;/p&gt;&lt;p&gt;that shift changes the way the whole system feels. you stop asking whether the mirror is current enough to use as working storage, and you start asking whether the refresh schedule is good enough for recovery. you stop editing wherever a copy happens to be handy, and you keep the real source in one place. the mirror becomes boring, which is exactly what you want from insurance. it is there, it is separate, and it does its job only when the primary cannot.&lt;/p&gt;&lt;p&gt;if you keep a local encrypted mirror, treat it as a recovery asset. decide which place is canonical, write your changes there, and refresh the mirror on a schedule that fits your risk tolerance. then leave it alone unless the primary is gone. that one rule keeps the backup from becoming a second source of truth. it also makes disaster recovery cleaner, because the copy you reach for in the worst moment is the copy you trained to be read-only in normal life. that is the whole point of insurance: it stays in the background until the day you need it.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/a-green-checkmark-in-your-repo-is-not-proof-of-a/</id>
    <title type="text">a green checkmark in your repo is not proof of anything</title>
    <updated>2026-06-22T00:59:00+09:00</updated>
    <published>2026-06-22T00:59:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/a-green-checkmark-in-your-repo-is-not-proof-of-a/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">A green checkmark in the repo can hide a stale live endpoint, and the only reliable receipt is a live request you make yourself.</summary>
    <content type="html">&lt;p&gt;Last month I shipped 4 things I would have sworn were live, and the live endpoint still served the old behavior. The repo looked calm. Commit merged, CI passed, dashboard said deployed. Everything in the project surface pointed in one direction, and that direction felt like done. The finding landed hard because the visible signs were so clean. A green checkmark in the repo can create the same feeling a receipt creates after a purchase, except the real customer experience still has to answer for itself.&lt;/p&gt;&lt;p&gt;The gap showed up in a simple way. I kept seeing the same shape repeat across those 4 changes: the code had moved through the normal path, the status looked correct, and the thing users actually touched still behaved like the old version. Someone could say, &amp;quot;it&amp;#39;s in prod,&amp;quot; and that sentence would carry a specific meaning. The code was in prod. The change I expected from that code was a separate question, and it stayed open until I checked the live thing directly. That distinction mattered more than I expected, because the repository had started to feel like the source of truth.&lt;/p&gt;&lt;p&gt;The only check that caught it was hitting the live thing myself and watching what came back. That was the moment the whole process sharpened. Repo status tells you what you intended. A live request tells you what users get. Those are different objects, and they answer different questions. One belongs to the work as it moves through your system. The other belongs to the actual outcome. If you stop at the first one, you can carry a false sense of completion all the way past the point where it matters.&lt;/p&gt;&lt;p&gt;The mechanism is simple enough to miss while you are moving fast. A commit records a promise. CI records that the code passed the tests you gave it. A deployment dashboard records that a release event happened. None of those guarantees the behavior you meant to change is the behavior the live endpoint is serving right now. The state can look correct from the inside while the outside world still sees the old result. That is the trap. The more familiar the workflow becomes, the easier it is to accept the signals it already gives you and stop before you ask the endpoint itself.&lt;/p&gt;&lt;p&gt;That changed the rule I use. I treat a commit as a promise and a live verification as the receipt. The promise matters, because it moves the change into the system. The receipt matters, because it shows the thing actually arrived in the form I expected. I want both in the chain, and I want the receipt to be something I can inspect with my own eyes. A status line in a dashboard can help me track the path. It cannot close the loop on its own. The loop closes when I hit the live thing and see the response.&lt;/p&gt;&lt;p&gt;For me, the practical lesson is to make the live request part of the shipping habit, especially when the repo looks reassuring. If the change matters, I do the extra step and check the thing that users will touch. That one step turned a vague sense of &amp;quot;probably live&amp;quot; into a concrete standard I can trust. It also keeps me honest about where the system can mislead me. The repo is where I shape the work. The live request is where I find out what survived that journey. Ship the receipt.&lt;/p&gt;&lt;p&gt;The bigger value is that this scales with the way I build: solo, with AI agents, moving quickly and relying on the shape of the workflow to keep me sane. Speed makes status surfaces tempting. Green checkmarks are easy to read, and they can make a task feel complete before the world has agreed. A live verification restores the part of the process that actually matters. It gives me a check against my own assumptions, and it gives the reader of my work a cleaner rule to adopt: when the outcome matters, ask the live thing directly and keep the answer close.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/shipped-a-feature-last-week/</id>
    <title type="text">shipped a feature last week</title>
    <updated>2026-06-20T00:59:00+09:00</updated>
    <published>2026-06-20T00:59:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/shipped-a-feature-last-week/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">a green build and clean merge can still miss the user path; if nobody walked it live, the feature was not shipped yet.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; a green build, clean diff, and merged commit can still leave the user path untouched. if nobody opened the path live, the work only looks finished.&lt;/p&gt;
&lt;h2&gt;what makes a feature feel done too early?&lt;/h2&gt;
&lt;p&gt;last week the feature looked finished. the build was green, the diff was clean, and the commit had already been merged. that combination can create a strong sense that the work is complete.&lt;/p&gt;&lt;p&gt;the catch was simple: nobody had run the actual path once. the surface signals were all calm, but the thing a user needed to do had never been exercised end to end.&lt;/p&gt;
&lt;h2&gt;why does a green build miss the real question?&lt;/h2&gt;
&lt;p&gt;a green build tells you the code compiles. it tells you the repository is happy with the shape of the change. it does not tell you that a person can move through the feature and reach the result.&lt;/p&gt;&lt;p&gt;that gap is where false confidence lives. the merge lands, the diff looks tidy, and the feature starts to feel real before anyone has watched it work in the hands of a user.&lt;/p&gt;
&lt;h2&gt;what actually counts as shipped here?&lt;/h2&gt;
&lt;p&gt;the bar is the live path. if no one walked it, the thing is still sitting there looking finished. the receipt is the user path working once, in reality, after the code has landed.&lt;/p&gt;&lt;p&gt;that is the point worth keeping. shipping speed matters, but speed only matters when it reaches the path people actually take.&lt;/p&gt;
&lt;h2&gt;what habit catches this before it slips through?&lt;/h2&gt;
&lt;p&gt;go click your own button. open the path the same way a user would and watch what happens. that one move would have exposed the gap immediately.&lt;/p&gt;&lt;p&gt;this is a small habit, but it changes the definition of done. the build can be green and the diff can be clean, and the feature still needs a live walk before it deserves the word shipped.&lt;/p&gt;
&lt;h2&gt;FAQ&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;what did the build miss here?&lt;/strong&gt;&lt;br&gt;the build only proved the code compiled. it did not prove that anyone could actually get through the feature end to end. the missing check was the live user path.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;why is a merged commit still not enough?&lt;/strong&gt;&lt;br&gt;because merge status only says the change is in the repo. it does not say the path was opened, used, and completed in real life. that gap is where a feature can look done while still failing in practice.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;what is the simplest way to avoid this?&lt;/strong&gt;&lt;br&gt;run the path yourself before calling it shipped. if the feature has a button, click it. if it has a flow, walk the flow. the point is to verify the thing a user will actually do.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/production-verification-is-part-of-done/</id>
    <title type="text">production verification is part of done</title>
    <updated>2026-06-19T00:59:00+09:00</updated>
    <published>2026-06-19T00:59:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/production-verification-is-part-of-done/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">production verification belongs in the definition of done, because the live check is where the answer shows up.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; a merged commit is only part of the finish. the real receipt is a live check in production, because that is where the answer shows up.&lt;/p&gt;
&lt;h2&gt;why does a merged commit still feel unfinished?&lt;/h2&gt;
&lt;p&gt;i keep seeing teams treat the commit as the finish line. the merge is real progress, but it is not the full answer.&lt;/p&gt;&lt;p&gt;the product still has to survive production. until that live check happens, the work is still carrying risk.&lt;/p&gt;
&lt;h2&gt;what does a live check in production actually tell you?&lt;/h2&gt;
&lt;p&gt;the live check is the receipt. it is the point where the real answer shows up, because production is where the product has to hold together.&lt;/p&gt;&lt;p&gt;without that check, you can ship code and still be guessing. that is how progress gets mistaken for certainty.&lt;/p&gt;
&lt;h2&gt;what should &amp;apos;done&amp;apos; mean for builders?&lt;/h2&gt;
&lt;p&gt;builders need a hard finish line. for me, that line includes validation in prod every time.&lt;/p&gt;&lt;p&gt;if a team wants a cleaner definition of done, this is the standard I would push on: code merged, product checked where it matters, then call it finished.&lt;/p&gt;
&lt;h2&gt;what question should teams ask themselves?&lt;/h2&gt;
&lt;p&gt;the simplest check is the one at the end: what does your team call done?&lt;/p&gt;&lt;p&gt;if the answer stops at merge, there is still a gap between shipping and knowing.&lt;/p&gt;
&lt;h2&gt;FAQ&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;why is production verification part of done?&lt;/strong&gt;&lt;br&gt;Because the merge only says the code is in. The production check is where you learn whether the product actually survives where it matters.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;what does a live check add after the commit is merged?&lt;/strong&gt;&lt;br&gt;It gives you the receipt. That is the moment the real answer shows up, instead of leaving you with a guess.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;what does &amp;apos;done&amp;apos; mean in this article?&lt;/strong&gt;&lt;br&gt;Done means the code is merged and the product has been validated in production. The finish line includes the live check.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/one-source-of-truth-for-repos-services-creds-aut/</id>
    <title type="text">one source of truth for repos, services, creds, automations</title>
    <updated>2026-06-18T00:59:00+09:00</updated>
    <published>2026-06-18T00:59:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/one-source-of-truth-for-repos-services-creds-aut/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">A single source of truth keeps repos, services, access details, and automations in sync by updating the record with the change.</summary>
    <content type="html">&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; The fix is to update the record in the same cycle as the change. When the truth moves and the record stays behind, drift shows up as a bug.&lt;/p&gt;
&lt;h2&gt;why do docs go stale even when everyone means well?&lt;/h2&gt;
&lt;p&gt;I stopped treating stale docs like a writing problem. They usually go stale because the update is a later task, and later slips once the change is already shipped.&lt;/p&gt;&lt;p&gt;That is the part I care about: the truth moved, the record did not, and now the bug is drift.&lt;/p&gt;
&lt;h2&gt;what changes when the source of truth lives with the change?&lt;/h2&gt;
&lt;p&gt;I keep one source of truth for repos, services, access details, and automations, and I update it in the same cycle as the change. That keeps the record attached to the work instead of trailing it.&lt;/p&gt;&lt;p&gt;The point is simple. If the thing changed, the written version changes with it. Same commit, same cycle, same reality.&lt;/p&gt;
&lt;h2&gt;what gets damaged when drift keeps building?&lt;/h2&gt;
&lt;p&gt;Hidden drift between what is running and what is written is expensive because it hides in plain sight. The next person starts from a map that no longer matches the place.&lt;/p&gt;&lt;p&gt;Broken handoffs follow from that. So does tribal knowledge, because the real answer lives in one head until that person is gone.&lt;/p&gt;
&lt;h2&gt;how do i know the source of truth is actually working?&lt;/h2&gt;
&lt;p&gt;I use one test: can someone other than me act on it without DMing me? If they can, the source of truth is doing its job.&lt;/p&gt;&lt;p&gt;If they cannot, I have a problem that will cost me later. That is the failure mode I try to catch early.&lt;/p&gt;
&lt;h2&gt;FAQ&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;why do docs rot after a change ships?&lt;/strong&gt;&lt;br&gt;Because the update gets treated like a separate later task. Once it is separate, it slips, and the record stays behind the change.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;what does drift look like in practice?&lt;/strong&gt;&lt;br&gt;It looks like a service that moved, a handoff that points to the wrong place, or a gap between what is running and what is written.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;how do i tell whether the source of truth is working?&lt;/strong&gt;&lt;br&gt;Ask whether someone other than you can act on it without asking you first. If they can, the record is carrying real weight. If they cannot, the handoff failed.&lt;/p&gt;</content>
  </entry>
  <entry>
    <id>https://howardchan.me/writing/moved-every-secret-out-of-our-app-env-files-into/</id>
    <title type="text">Moving every secret into one canonical store made our env files boring again</title>
    <updated>2026-06-17T00:59:00+09:00</updated>
    <published>2026-06-17T00:59:00+09:00</published>
    <link rel="alternate" type="text/html" href="https://howardchan.me/writing/moved-every-secret-out-of-our-app-env-files-into/"/>
    <author><name>Chak Hang (Howard) Chan</name><uri>https://howardchan.me/</uri></author>
    <summary type="text">Moving secrets into one canonical store made env files boring again and made rotation and drift easier to manage.</summary>
    <content type="html">&lt;p&gt;The clearest fix we made was moving every secret out of app env files and into one canonical store. Credentials now load at runtime instead of living inside project files, and that one change cleaned up a lot of the mess around how we handled configuration. The part that matters is simple: the app asks for credentials when it runs, while the files in the repo go back to being plain configuration. That shift sounds small on paper. In practice, it changes how a solo developer can move across projects without carrying hidden risk everywhere.&lt;/p&gt;&lt;p&gt;Before that change, drift was the tax we kept paying. Each project had its own env file, and the real question became which one held the current value. That turns into archaeology fast. You stop trusting the shape of your setup because every copy can slowly diverge from the others. Once there is one source of truth, that problem gets much smaller. There is a single place to update, a single place to reason about, and a much lower chance that one app is quietly running with a different value than the rest.&lt;/p&gt;&lt;p&gt;Rotation is where the new setup really pays off. When a credential changes, swapping it in one place means every app picks it up the next time it loads at runtime. There is no need to grep across repos or replace values file by file. That matters because rotation is where bad habits usually show up. The old workflow made every change feel broad and fragile, which is exactly when people delay it. A runtime load path makes the update feel local. The move is contained, and the blast radius stays small.&lt;/p&gt;&lt;p&gt;The other win was that secrets stopped leaking into code and config. Once sensitive values are out of project files, there is much less chance that something private gets committed by accident or copied into a place it should never have reached. That is the kind of cleanup that does more than remove a risk on a checklist. It changes what the rest of the repository can safely contain. You can look at the files again without wondering whether they hide something that should have stayed elsewhere.&lt;/p&gt;&lt;p&gt;What nobody tells you is that the store itself is only part of the story. The bigger change is that env files become boring again. They go back to being just config. You can read them. You can share the shape. You can commit a template. None of that feels dangerous anymore, because the sensitive values are gone from the file. That boring quality is the real sign the system is healthier. When the file stops carrying secrets, it becomes easier to reason about, easier to review, and easier to hand to someone else without adding extra caution to every glance.&lt;/p&gt;&lt;p&gt;We ran config-with-secrets-inline for a long time and told ourselves it was fine because it worked. It worked the way a bald tire works, right up until rotation day. That is the lesson I keep coming back to. A setup can function for a long stretch and still carry a failure mode that only shows itself when you need to change something important. The safer rule is to keep secrets out of project files and load credentials at runtime from one canonical store. If you are still copying values into &lt;code&gt;.env&lt;/code&gt; files project by project, this is the place to stop and ask what your real source of truth is.&lt;/p&gt;&lt;p&gt;If you want to apply the same idea, start with your current credential path and trace where each value lives today. Ask whether the file in the repo is carrying sensitive data or just describing configuration shape. Then move the sensitive part into one runtime source and leave the file behind as a template you can read without concern. That gives you cleaner rotation, less drift, and a setup that stays understandable after the first pass is over. For a solo builder, that kind of simplicity is worth more than a setup that only looks convenient while nothing changes.&lt;/p&gt;</content>
  </entry>
</feed>
