One Record for Shared Third-Party Spend

One Record for Shared Third-Party Spend

Most organisations discover shared third-party spend the same way: a quota is exhausted, something stops working, and nobody can say which of six internal consumers used it up.

The obvious fix is a ledger. One record, every consumer writes to it before spending, and now you can answer the question. We built exactly that for a shared API key used across several of our pipelines.

Then the first team to adopt it immediately named the hole, and the hole is the interesting part.

The number the ledger could not see

The key is shared. One of the pipelines using it runs a monthly refresh — and that refresh does not run on the machine holding the ledger. It runs as a scheduled job on a cloud runner, which has no access to a local file.

That job is, by a wide margin, the largest single consumer: roughly 37,200 units in a run. It can execute end to end, spend more than everything else combined, and leave no trace in the record built to track it.

This is worse than it first sounds, and the reason is the read path.

The ledger was written to fail open. If it cannot be read, callers proceed rather than halting — a reasonable default, because a cost-tracking file should never be the thing that takes a production pipeline down. But fail-open has a consequence nobody states when they choose it: unrecorded spend is indistinguishable from no spend.

A total of zero now has two possible meanings. Nothing ran. Or 37,200 units ran somewhere the ledger cannot see.

Those two readings demand opposite decisions, and the file presents them identically.

Why a confidently wrong total is worse than none

If we had shipped that, the ledger would have been confidently wrong about the single number it exists to protect.

That is a harder failure than having no ledger at all, and the difference is behavioural rather than technical. Without a ledger, everyone knows they are guessing, and they act cautiously — they check before a big run, they ask, they over-provision. With a ledger, nobody guesses. They read the number and proceed, because the number is there and it is precise and somebody built it on purpose.

A precise number carries an implicit claim of completeness that a shrug does not. Introducing one without establishing that claim converts careful behaviour into confident behaviour, using evidence that does not support the confidence.

This shows up well beyond API quotas. The same shape sits behind:

  • A cost dashboard that covers the cloud accounts you own and silently omits the department paying a vendor directly on a card.
  • A vendor register that lists every supplier procurement onboarded and none of the tools bought as monthly subscriptions by individual teams.
  • A headcount figure that counts employees and not the contractors doing the same work under a statement of work.
  • A licence count derived from your identity provider, where the licences bought outside it simply do not appear.

In each case the number is accurate about everything it measures. It is wrong about the thing the reader believes it measures, and nothing in how it is presented distinguishes those two.

★ Insight ───────────────────────────────────── The failure mode is created by the combination of fail-open reads and an aggregate presented as a total. Either alone is fine. Fail-open with a total labelled “of what we can see” is honest. A hard-fail read with an unqualified total is also honest, because an unreadable ledger stops the work rather than reporting zero. It is specifically the pairing — proceed quietly on a read failure, then present the result as complete — that manufactures false confidence, and it is a pairing most teams arrive at one reasonable decision at a time. ─────────────────────────────────────────────────

What we changed, and what we deliberately did not

Three changes, all additive, because two pipelines had already wired against the existing interface and breaking it would have stranded them.

Blind spots are declared, per resource. The ledger now holds an explicit list of known consumers that cannot write to it. The cloud job is named in that list. This is the load-bearing change and it is barely any code: the system’s knowledge of its own incompleteness is now data rather than something a particular person happens to remember.

Callers may hold back headroom. A budget check can now reserve capacity against invisible spend. Critically this is opt-in rather than a default, and that was a deliberate decision worth explaining. Only the caller knows whether it is about to spend four units or thirty-seven thousand. A reserve large enough to protect against the big consumer would needlessly refuse every small one, so a system-wide default would either be too small to matter or too large to use.

The number tells you it is a floor, at the point of decision. When a check refuses, the message says the real total may be higher and names which consumers cannot report in. When a check passes but is above half the limit, it warns that the figure is a lower bound. Both appear where the decision is being made, rather than as a comment in a file nobody reads at 7am.

Seventeen probes cover it now, four of them added with this change, including the ones that matter most: that a resource with a declared blind spot behaves differently from one without, that a reserve refuses a request the bare total would have approved, and that the same request without a reserve proceeds.

What we did not do is make the ledger mandatory, or make the read path fail closed. Both were tempting. Both would have made a cost-tracking file capable of stopping production work, which trades a visibility problem for an availability problem — and availability problems are the ones that get a control switched off entirely.

The executive version

Strip out the API and this is a question about every aggregate number your organisation acts on.

Does this figure know what it cannot see?

Not “is it accurate” — most of these figures are perfectly accurate about their own scope. The question is whether the scope is written down anywhere the reader of the number will encounter it, and whether anything happens differently when the uncounted portion could be large.

Three follow-ups that make it concrete:

  1. Name one number your leadership team treats as complete. Total cloud spend, total licences, total suppliers, total headcount.
  2. Ask who could add to that number without appearing in it. Not hypothetically — name the team, the card, the contract, the scheduled job. If nobody can answer, that is the answer.
  3. Ask what the number does when its source is unavailable. If it quietly reports the last known value, or zero, or omits a region, then on the day the source breaks the figure will look normal and be wrong. That is the day it will be used for something.

The useful output of that exercise is rarely a better number. It is usually a sentence attached to the number, stating what it excludes — which costs nothing, survives staff turnover, and converts a false claim of completeness into a true claim about a floor.

When you cannot close the blind spot

Sometimes the uncounted consumer genuinely cannot be brought into the record. Ours is a case of that: a job on infrastructure we do not control, reaching a resource we do.

The reflex is to treat that as a temporary state and plan the integration. Often that is right. But it is worth naming the three positions available, because organisations drift into the worst one by default.

Close it. The consumer writes to the same record as everything else. Best, and frequently more work than it is worth for a single monthly job.

Declare it. The record states that this consumer exists, cannot report, and is large. The total becomes an explicit floor. Cheap, honest, and enough for most decisions — you cannot say what was spent, but you can say what was spent at least, and you know which direction the error runs.

Ignore it. The consumer is absent, the total is presented as complete, and everyone downstream acts on a number nobody has established.

The third is not a decision anyone makes out loud. It is what happens when the first is too expensive and the second never occurs to anyone, and it is the default state of most aggregate reporting in most organisations.

The second option is almost always available and almost never taken, because declaring a gap feels like admitting a failure. It is the opposite: a figure that knows its own limits is a stronger artefact than one that does not, and it is the only one of the three that stays true when circumstances change.

Why the fix came from outside the team that built it

One last observation, and it is the most transferable part.

We did not find this. The team adopting the ledger found it, within a day, because they were the ones who knew where their own workload actually ran. From inside, the ledger looked complete — every consumer we could think of wrote to it, and the ones we could think of were the ones we had built.

That is the recurring shape of a blind spot: it is not a thing you failed to check, it is a thing you had no reason to consider, and the person with the reason is usually the person about to depend on it. The practical version is that a new control should be reviewed by whoever will have to live with its output, before it is treated as authoritative — not because they will review the code more carefully, but because they know something you do not.


This is the kind of governance work we do through Ganda Tech Services — with infrastructure and cost visibility through Cloud Geeks.

Free Guide · 2026

AI Strategy Primer for Australian Business Leaders

A practical framework for AI adoption in 2026 — cut through the hype and start with what matters.

We email a confirmation link first. No spam. Unsubscribe any time.