← Back to all posts

Founder Notes

Three Uncomfortable Truths About Your Last Risk Assessment.

September 3, 2026 · Jesse McKenna

A patient sits with a thermometer in each side of his mouth, one reading 104.7°F and the other 94.0°F. One doctor says "We need to get you in an ice bath ASAP!" while another says "Let's get you a blanket and some cocoa."

Send the same institution to two different firms for a traditional risk assessment and you'll get two different reports back. Not slightly different - different risk ratings, different priority findings, sometimes a different read on the same exposure entirely. Nobody objects to this in the moment, because there's no baseline to object against. A thermometer gets checked against a known temperature. A traditional risk assessment gets checked against nothing. Whatever the assessor produced is just what the risk assessment says now, until someone commissions another one.

That's the part that doesn't hold up. Call the output a "risk assessment" and everyone assumes it behaves like a measurement - something with a defined method, that produces the same answer on the same inputs no matter who's holding the instrument, and that you can trace back to the evidence behind it. Almost none of that is true in practice. Three specific things are missing, and none of them get fixed by a more careful version of the same process.

Two firms, two answers, and no way to adjudicate between them

Most traditional risk assessments are built the same way: workshops, interviews, a scoring rubric someone applies by feel after listening to a room of people talk about their own risk. That's not a criticism of the people running the workshop. It's a description of the instrument. The instrument is a person's judgment, elicited under time pressure, and every person's judgment differs even holding the underlying facts constant. Swap the analyst and you've swapped the instrument.

Anywhere software gets built, there's a name for what happens when the same input produces two different outputs: a bug, or at minimum a signal that something isn't being accounted for. Doing threat research at Silver Tail Systems and RSA meant living inside that standard - if a test case passed on Monday and failed on Tuesday with nothing else having changed, nobody shrugged and moved on. Somebody tracked down the variable that wasn't being controlled for, because a result you can't reproduce isn't a result. It's noise masquerading as a result.

A traditional risk assessment produces exactly that kind of noise, and nobody treats it as the red flag it would be anywhere else. Bring the same firm back next year and the correspondent banking rating moved from moderate to high - was that because the risk actually changed, or because a different analyst ran the interviews this time? The report can't tell you. Both explanations fit equally well.

None of this shows up in the document itself. It only shows up when you commission a second assessment and the two don't agree.

The instrument is a conversation, and conversations aren't reliable

Most of what feeds a risk assessment doesn't come from a system of record. It comes from someone asking another person a question in a room, and writing down an impression of the answer. Change how the question gets asked - the order it comes in, whether it's phrased "walk me through what happens when a vendor's SOC 2 comes in late" versus "do you feel like your vendor controls are adequate" - and you change the answer, before anyone's actual view of the risk has moved at all.

The respondent isn't a fixed input either. Someone interviewed the week after a near-miss rates their own control environment differently than someone interviewed during a quiet quarter, and neither one is lying. Someone interviewed at 4pm after five straight workshop sessions answers differently than the same person would at 9am. None of that context gets written down next to the answer. It just becomes the answer.

Then there's the interviewer's side of it, which is just as noisy. An interviewer isn't a tape recorder - they're synthesizing in real time, filtering what they hear through their own sense of what's normal and what's alarming. Even a detailed transcript doesn't fix that. You've just moved the filtering to whoever reads the transcript. Two interviewers listening to the identical answer can walk away having "heard" two different findings, and only one version - the interviewer's - ever makes it into the report. The person who actually gave the answer never gets a chance to confirm that's what they meant.

None of this is a problem risk assessment discovered on its own. A measurement needs a fixed method - same inputs, same procedure, a comparable output every time. An open-ended conversation doesn't have one: change the question, the day, or the person, and the answer changes with it, which means what comes out is an impression, not a data point. That's the exact reason unstructured interviews have such a well-documented reputation for unreliability in hiring, and why anyone actually trying to measure something instead of just describing it ends up at structured, scripted questions with a fixed rubric. Risk assessment never made that shift. It's still running on the version of interviewing everyone else moved past once they had something real riding on the answer.

The rating rarely traces back to the data that produced it

Try to trace why a category is rated the way it is in most traditional risk assessments, and what you find in the write-up is a paragraph of narrative summary, not a chain of evidence. "High risk, correspondent banking" reads as a conclusion. It's much harder to find out which specific relationships, transaction patterns, or geographic exposures actually drove that number, because the workshop that produced it synthesized a room's worth of impressions into a rating and didn't keep the wiring diagram.

That should be a low bar to clear. A finding that can't be traced back to the specific evidence behind it isn't a finding - it's an opinion with a page number. Anyone who has to defend a conclusion to an examiner, an auditor, or a court eventually gets asked the same question: what is this based on? A traditional risk assessment usually can't answer it below the level of the narrative itself, because the rating was never built additively from individually inspectable pieces of evidence.

And because nobody can trace the rating back to the evidence, nobody can argue with a specific piece of it either. You can't dispute one data point and expect the rating to move, because the rating was never built from data points to begin with. You can only wait for the next full cycle and hope the new narrative reads differently.

None of this gets fixed by trying harder at the same process

A thermometer earns trust for a specific reason: same input, same reading, traceable back to a known reference, usable on anything you point it at without having to reinvent the instrument each time. None of the three gaps above - reproducibility, the reliability of how the underlying data gets collected, and traceability back to evidence - get closed by a longer workshop, a more experienced facilitator, or a bigger interview list. They get closed by treating risk assessment as something built to a method, the way a measurement is, instead of something assembled by a person applying judgment under a deadline, however good that person's judgment happens to be.

Jesse McKenna has over 20 years of experience in fraud detection, financial crime, and enterprise risk - building detection systems at PayPal and eBay, leading threat research at Silver Tail Systems and RSA, and building SAR prediction models at Refine Intelligence. He is the Founder and CEO of VeloRisk, a risk strategy platform for regulated industries.
← Back to all posts