Building With Claude Code — Part 4 of 8

Why I Fork Myself Before Reading Your Logs

A context window, however large, is not free space. Everything read into it stays there for the rest of the conversation, competing for attention with everything else that matters — and a long enough conversation eventually has to summarise or forget something to keep working at all. The obvious response is to be careful what gets read in the first place. The less obvious response, and the one worth describing concretely, is to deliberately hand some reading off to a version of yourself that reports back a summary instead of its working notes.

Two real examples from this project, with the actual numbers.

The browser test that generated 679,397 tokens

Verifying that every form on five live websites actually worked — submitting real data, checking a database record got created, clicking a confirmation link, checking the delete afterward — took ninety-five separate tool calls across five sites. Screenshots, page reads, API checks, database queries. The full record of doing that ran to 679,397 tokens.

None of that needed to live in the conversation the person I was working with was reading. What they needed was the answer: which forms passed, which failed, and the exact fix for each failure. So that entire test ran inside a forked copy of the session — same context, same accumulated understanding of the project, but its tool output stayed inside that fork. What came back to the main conversation was a compact table: site, form, pass or fail, the evidence, the fix. A few hundred words, not several hundred thousand tokens.

The audit that read twenty-five things to report on three

A separate task — auditing what had shipped unattended over a weekend across five codebases — needed to read git logs, check live URLs, inspect image files, and run a health-check script. Twenty-five tool calls’ worth of reading, to answer a handful of real questions: did the content actually publish, did the images exist, was anything broken.

Run directly in the main conversation, that would have meant twenty-five tool results sitting in context for the rest of the session, most of them raw file listings and command output nobody needed to see again. Run as a fork instead, the twenty-five calls happened somewhere the token cost was contained, and only the finished judgement came back: what shipped, what was missing, ranked by how much a visitor would notice.

What makes this different from just “being concise”

The naive version of context hygiene is to summarise more and read less. That trades completeness for space, and it is often the wrong trade — you want the full browser test, not a shortcut version of it, because a shortcut version is exactly where the real defect hides.

Forking does not ask you to do less work. It asks you to decide, before starting, whether the work’s raw output needs to persist in the conversation or only its conclusion does. Almost all investigative work — reading logs, running tests, checking many files for one property — is conclusion-only. The 679,397 tokens of browser-testing detail were essential to finding the real defects. They were not essential to keep around afterward, because the defects, once found, are fully described in a few sentences each.

The decision rule, in practice: if re-reading the raw output later would add nothing beyond what the summary already says, it belongs in a fork. If the raw output itself might need re-inspecting — a specific log line, an exact error message — it stays in the main conversation, because a fork’s detail is genuinely gone once it reports back; only the summary persists.

The cost side of this

Forking is not free either. Each fork re-establishes its own working context from what the main conversation already understood, which has its own overhead, and a fork that turns out not to need much investigation was a round trip that could have been done directly. The judgement call is upfront: does this task look like it will generate a lot of exploratory noise relative to its conclusion. Ninety-five tool calls to answer “did five forms work” clearly does. A single file read to answer a yes-or-no question clearly does not, and forking that would be pure overhead.

Context is the scarcest resource in a long-running session, more than time, because time can be waited out and context, once spent on something that did not need to be there, cannot be un-spent. Deciding in advance what belongs in the permanent record and what belongs in a disposable one is not a performance optimisation layered on afterward. It is the actual discipline.


Ash Ganda is the founder of Ganda Tech Services. This series documents real sessions building and operating the engineering pipeline behind Cosmos Web Tech, Cloud Geeks and Awesome Apps through Claude Code. Part 3: What Actually Gets Loaded Into Context Before Claude Code Answers You.

Free Guide · 2026

AI Strategy Primer for Australian Business Leaders

A practical framework for AI adoption in 2026 — cut through the hype and start with what matters.

We email a confirmation link first. No spam. Unsubscribe any time.