Building With Claude Code — Part 2 of 8

Six Agents, One Afternoon, and the Two Sites That Shipped Stale Builds

Five websites needed the same change: route every contact form through a new shared service instead of each site’s own version. Five different codebases, five different build tools, one afternoon.

The obvious way to do this with Claude Code is not to do it five times in sequence. It is to dispatch five agents at once, each one working in its own repository, none of them touching the others’ files, and read the results when they finish. That is what we did. Six agents, one message, running in parallel — one per site, plus one for the shared service itself.

Five minutes later, all six reported success. Two of them were wrong.

What “success” meant to the agent

Each agent’s job ended the same way: build the site, deploy it, confirm the deployment completed, report back. Two of the five sites use pnpm rather than npm. On both of those, the agent’s deploy script ran pnpm build, pnpm refused to run certain install scripts without interactive approval, the build step exited with an error — and the script, written to continue regardless, moved straight on to wrangler pages deploy dist.

That command deployed whatever was already sitting in the dist folder. Not the new code. The output of the previous build, from before any of today’s changes existed.

Both agents reported “Deployment complete!” Both deployments were real, successful, and deployed the wrong thing. From inside the agent’s own transcript, there was nothing to distinguish this run from a correct one — the deploy command genuinely succeeded, it just deployed stale content.

The check that caught it

I did not catch this by reading six agent reports and feeling satisfied. I caught it by doing something the agents’ own success criteria never asked for: checking the actual age of the file that got deployed.

ls -la --time-style=+%H:%M dist/index.html

Two files were timestamped from five minutes before the agents started working. The other three matched the time the agents had actually run. That five-minute gap was the whole defect, visible in one command, invisible from “Deployment complete!” alone.

The fix was narrow: swap pnpm build for a direct npx astro build in the deploy step, bypassing the interactive approval prompt entirely rather than working around it. Both sites rebuilt correctly, redeployed, and this time the timestamps matched.

The general shape of the failure

Orchestrating several agents in parallel is genuinely faster than doing the same work in sequence — five sites changed in the time it takes to review the results, not five times that. But parallel dispatch has a specific failure mode that sequential work does not surface as easily: you are reading a status report instead of watching the work happen, and a status report only tells you what the script was written to check.

“The command exited zero” and “the thing I wanted to happen, happened” are different claims. A script that treats a build failure as non-fatal and presses on to deploy will keep reporting success indefinitely, because nothing in it was ever instructed to notice the difference between a fresh build and a five-minute-old one.

The rule that came out of this, now applied to every multi-site change: verify the artefact, not the exit code. For a deploy, that means checking that the thing on the server actually changed — a file timestamp, a piece of new text on the live page, a changed hash — not just that the deploy command returned success. The agents did exactly what they were told. What they were told to check was not enough.

Why this is the right way to use agents anyway

None of this is an argument against running agents in parallel. Six agents did five sites’ worth of real engineering work in one afternoon, correctly, for four of the five sites, and the fix for the other two took minutes once found. The alternative — one agent, one site at a time, five times over — would not have been more correct. It would only have hidden the same defect behind a longer wait, because the underlying pnpm behaviour does not change based on how many other things are happening at once.

What parallel dispatch changes is not correctness. It changes how much you can get done in an afternoon, and how much that raises the cost of not independently checking the result.


Ash Ganda is the founder of Ganda Tech Services. This series documents real sessions building and operating the engineering pipeline behind Cosmos Web Tech, Cloud Geeks and Awesome Apps through Claude Code. Part 1: The 403 That Had Nothing to Do With Our Credentials.

Free Guide · 2026

AI Strategy Primer for Australian Business Leaders

A practical framework for AI adoption in 2026 — cut through the hype and start with what matters.

We email a confirmation link first. No spam. Unsubscribe any time.