Skip to main content

Why “did that change break anything?” is such a hard question to answer

Every other question about an ETRM/CTRM change can be answered late. You can review the approval trail months afterward, reconstruct who signed what, and pull the release notes years later if anyone asks. One question doesn't work that way.

How do you know the change didn't alter the numbers you report?

To answer that, you need to know what those numbers were before anyone touched the system. And by the time the question arrives, the moment to find out has usually gone.

You can't compare against something you never recorded

What you need is a record: this is how the system behaved, on this data, before the change. Not roughly, and not in the sense that someone remembers it looking fine, but recorded.

That record is a baseline. It's an unglamorous thing, and it's the most important artifact in the whole question, because without it there's nothing to compare against. Asking whether a change broke anything isn't merely difficult without a baseline. It's unanswerable in principle. You're being asked to spot a difference with only one side of the comparison in front of you.

Most testing doesn't produce one. It asks whether the system works: does the trade book, does the report generate, does the screen show what it should. Those are reasonable questions and they catch real problems. They're not the same as asking whether the regulated outputs you're accountable for behave exactly as they did last week. A system can work perfectly well and still settle on a different answer than it gave before.

What gets used instead

Without a baseline, firms fall back on things that feel like evidence and aren't quite.

There's the sign-off, a named person confirming the testing was done, which is a record of process rather than of outcome. There's the spot check, where someone compares a handful of values before and after, usually the ones they judge most likely to move. That one is genuinely useful, and it's also a sample chosen by instinct on a system with thousands of values in play. And there's the folder of screenshots, which captures what a screen looked like at a moment without a structured way to compare it with anything.

None of this is bad practice. It's what careful people do when nobody has handed them a better instrument. But when the question comes from outside, each one turns into an assertion rather than a record, and someone ends up vouching for a system's behavior from memory and inference.

The window closes

A baseline can only be taken before the change. Once the upgrade has landed, the old behavior is gone. You can reason about it, reconstruct it and argue about it, but you can't observe it anymore. The one moment when that evidence was there for the taking has passed, and it passed without anyone noticing, because nothing appeared to go wrong.

Which is why this tends to surface at the worst possible time. Not during the change, when everyone is paying attention, but months later, when a report is queried or someone asks how you know. By then, you're reconstructing history instead of producing a record.

It's also why the regulatory direction of travel matters here, without needing to open any rulebook. Across the regimes that energy and commodity trading firms answer to, the expectation has been shifting from give us the right number toward show us the systems producing it were controlled as they changed. The second question is far harder to answer after the fact. It rewards the firms that took a baseline and offers little to the ones that meant to.

What a baseline gives you

The useful thing about a baseline is that it turns a judgment into a comparison.

With one in hand, the question stops being whether you believe the change was safe, and becomes a much plainer one: what moved? You can point at the outputs you care about and say that these behave as they did before, and these three don't. Two of those we intended. The third we didn't, and we caught it before it went anywhere.

That comparison works best at the level where the numbers are formed, meaning the logic the system applies to a trade rather than the figure it prints on a screen. It's a technical point with a practical consequence: a screen can look identical while the calculation behind it has shifted. Colleagues in trading IT will recognize that immediately. For everyone else, the short version is that the check has to sit underneath the presentation layer, or it isn't really a check.

To be clear about where the line sits. A baseline comparison tells you whether your system's behavior has changed. It doesn't tell you that the original behavior was right. If a value was being calculated in a way you'd take issue with before the upgrade, a baseline will faithfully confirm it's still being calculated that way afterward. Judging whether a difference was intended, and whether the underlying treatment is sound, stays with the people who know the regulation and the book. What the baseline removes is the guesswork about what actually moved.

Take it before you need it

The awkward thing about the baseline you never took is that you only miss it once it's too late to take.

So, the question to consider is not whether your last upgrade was tested. It almost certainly was. It's whether, if someone asked tomorrow, you could show what your reported numbers looked like before the change and what they looked like after. If the answer depends on someone's memory, a folder of screenshots, or a sample of values chosen at the time, the gap is not in your testing. It is in what your testing leaves behind.

There's a change coming in the next 12 months. There always is. A baseline is simple to take before the change. Afterward, all you're left with is inference.

If that question is sitting with you, the whitepaper goes further into it. The cost of finding out later looks at how routine system change reaches regulated outputs, and what evidence holds up when someone asks.

Tags:

Trinitatum
Trinitatum
7 Sep 2026