When a Gate Passes Because It Cannot Fail

When a Gate Passes Because It Cannot Fail

We had uploaded the same video twice before. Twenty-one uploads for twelve videos, on one occasion.

So we built a check: before uploading, work out whether this exact video has already gone up, and refuse if it has. It was written, wired into the upload path, deployed, and it ran on every upload from that point on.

It blocked nothing. Not one item, ever.

That was the correct result — there were no duplicates in the period — and it was also completely uninformative, because the check had never been capable of blocking anything.

The mechanism

The check identified a video by reading a small companion file written alongside each render, containing a fingerprint of the output.

Short-form videos never write that file.

So for every short, the check looked for the fingerprint, found nothing, concluded it had not seen this video before, and approved the upload. Not an error. Not a warning. A confident, well-formed pass, produced by a component that had no information at all.

The failure has a particular shape worth naming: the control was not broken, it was blind. A broken control throws. A blind one answers.

Zero is the tell, and it reads as good news

The number that should have triggered the investigation was zero blocked.

That is the difficulty. A newly deployed check that blocks nothing has two possible explanations, and they demand opposite responses:

  1. Nothing is wrong. The check is working and there is nothing to catch.
  2. The check cannot see. It would approve anything.

These are indistinguishable from the result alone. And only one of them is safe to assume, which means the default reading — the one everybody takes, including us — is the unsafe one.

Worse, the safe-looking reading is reinforced every day it runs. A month of zero blocks feels like a month of evidence that the pipeline is clean. It is a month of evidence about nothing.

★ Insight ───────────────────────────────────── The reason this is hard to catch by review is that the code is correct. Read the check on its own and it does exactly what it claims: fetch the fingerprint, compare, refuse on a match. The defect is not in the logic, it is in an assumption the logic makes about its input — that the fingerprint exists. Nothing in the file states that assumption, and nothing downstream fails when it does not hold. Code review finds wrong logic; it does not reliably find correct logic operating on an input that never arrives. ─────────────────────────────────────────────────

Three in one fortnight, which makes it a habit

The reason this got written up rather than quietly fixed is that it was not the only one.

In the same fortnight, two other controls in unrelated systems turned out to have the same property. One was a check on whether a summary file for AI systems was present and useful — it was present, ten lines long, and said nothing a machine could use, so the check that confirmed its existence was confirming the wrong thing. Another was a screen behaviour that could never have worked on one platform because the application had not declared the capability it required, and the platform refuses such requests silently.

Three separate systems. Three components that ran, reported success, and had never been capable of reporting anything else.

That is not three bugs. Three instances of one pattern in two weeks is a habit, and the habit is this: we were building controls and then confirming they were connected, rather than confirming they could act. Connection is easy to verify and feels like verification. It is not the same question.

The distinction is worth naming precisely because it is so easy to satisfy the wrong one:

QuestionHow it is usually answeredWhat it proves
Is the check wired in?It appears in the logs / it runs on every itemThe code path executes
Is the check working?It has not reported any problemsNothing
Can the check refuse?Feed it something it should refuseThe control exists

Only the third produces evidence. The first two are the two that get asked.

The rule we adopted

A new check is not finished when it passes. It is finished when it has been shown to refuse something it should refuse.

One deliberately broken input, fed through once, before the check is trusted. That is the whole discipline. It costs minutes, and it is the only thing that distinguishes a control from a decoration.

In our case the test is trivial to construct: take a video that has already been uploaded and try to upload it again. If the check does not stop it, the check does not work. We had never run that, because the check’s purpose was to prevent something we were no longer doing.

The general version:

  • A validation rule — submit something invalid and confirm it is rejected.
  • An access control — attempt the access you expect to be denied, with an account that should be denied it.
  • An alert — trigger the condition and confirm the alert arrives, in the channel someone actually reads.
  • A backup — restore it.
  • A reconciliation — introduce a discrepancy and confirm the report shows it.

Every one of those is a five-minute exercise that most organisations have never performed on most of their controls.

The cost of finding out late

Worth quantifying, because “we should test our controls” is advice everybody agrees with and nobody schedules.

The check ran for roughly six weeks before this surfaced. In that period every upload was approved by a component that would have approved anything, and the pipeline’s status reports listed duplicate protection as active.

Nothing bad happened. No duplicate was uploaded, because the behaviour that caused the original incident had already been fixed elsewhere.

That is the genuinely uncomfortable part. The control was unnecessary during the exact period it was broken, which is why nothing surfaced it — and it is also why we had no way of knowing that. A control that is both blind and unneeded is indistinguishable from one that is working, for as long as the luck holds.

The cost is therefore not measured in damage. It is measured in the decisions made on the strength of a green light: work that was not scheduled, a risk that was closed, attention that went elsewhere. Those are real and they are invisible, and they are the reason the five-minute negative test is worth insisting on even when the control is protecting against something that feels unlikely.

Why it applies well beyond software

Strip out the technical detail and the question is about every assurance your organisation relies on.

A control produces a signal. The signal is used to decide something — to proceed, to sign off, to stop worrying. The value of the signal depends entirely on whether it could have come out differently.

A check that has never refused anything has produced no evidence. It has produced a sequence of outputs that are consistent with everything being fine and equally consistent with the check being disconnected. You cannot tell which from the outputs, and the longer the sequence runs, the more confident everyone becomes on the basis of the same absent information.

Three questions worth asking about anything in your business that provides assurance:

  1. When did this last flag something? If the answer is never, or nobody remembers, that is a finding rather than a reassurance.
  2. What would it take to make it flag? If nobody can describe the input that would trip it, nobody knows what it does.
  3. Has anyone tried? Not reviewed the design. Fed it something bad and watched.

What we did about it

The fingerprint is now derived from the video file itself rather than from a companion file that may or may not exist, so it works for every category. The check was then given a duplicate on purpose and confirmed to refuse it.

And the standing rule went into our release practice: any new gate ships with one recorded negative case. Not a unit test asserting the logic — an actual run, on the actual path, refusing an actual bad input, with the output kept.

That is a small amount of ceremony. It replaces a much larger amount of unfounded confidence.

The uncomfortable version of the question

The reason this one stung is that the check was built in response to an incident. We had duplicated uploads, we recognised the problem, we built the control, and we recorded the problem as addressed.

For the whole period between building it and finding this, the honest state of affairs was: the incident was unaddressed, and we had stopped looking, because we believed it was handled.

A control you trust and have not tested is worse than no control, for exactly that reason. Without it you stay alert. With it you stop.


Engineering governance, control design and the question of whether your assurance is actually assuring anything is part of the work we do through Ganda Tech Services, with infrastructure through Cloud Geeks.

Free Roadmap · 2026

Digital Transformation Roadmap 2026

A 12-month framework for Australian SMBs ready to modernise — phases, tools, and milestones.

We email a confirmation link first. No spam. Unsubscribe any time.