I've sat through a lot of fraud tool demos over the years, as a buyer, an evaluator, and eventually just as the person who has to actually use the thing once procurement is done. They almost always look great. Clean dashboards, satisfying flag counts, a workflow diagram that makes fraud detection look like a solved problem with a nice UI on top of it. Then the tool ships, and within a week the team has built a shadow process around it: a spreadsheet here, a personal notebook there, a Slack channel where the real judgment calls actually happen. Not because the tool is broken. Because it was built for the pitch, not for the desk.

That gap is the subject of this piece. I'm not writing this as a product person, I'm writing it as the person on the other side of the tool, the one who has spent years working around gaps that a slightly different design decision would have closed. Fraud is one lane within the broader Trust & Safety world, alongside things like content moderation and platform abuse, and I'll leave those to people who actually live in them. This is specifically about fraud investigation tooling, the tools built for people like me. If you're building or buying it, this is the version of the feedback you don't usually get in a sales call.

The pattern

A few misses show up again and again, across different tools, different companies, different budgets. None of these are exotic. They're just the kind of thing that's easy to miss if you've never had to sit in the seat.

01. The buying decision skips the people who'll actually use it

On more than one occasion, I've watched a product or senior leader choose a tool based on cost and a feature checklist, factors that make complete sense at a high level and can still produce a tool that's genuinely painful to use day to day. It's not that those factors don't matter. Budget and feature scope are real constraints, and someone has to own them. It's that they're rarely balanced against the people who'll be living inside the tool eight hours a day.

Involving actual investigators in evaluation, not just a courtesy demo at the end but real hands-on testing, and genuinely listening to pain points and wishlist items, is one of the most overlooked steps in the buying process. The goal isn't to let operations override every business constraint. It's to strike an honest balance between what makes sense on a budget spreadsheet and what makes sense on the desk of the person who has to hit quota with the thing.

02. Confidence gets treated as binary, because that's all that's on offer

Most tools force a decision: flag or don't flag, escalate or close. This is worth sitting with for a second, because it's not that investigators secretly want a spectrum and tools ignore the request. It's that binary is the default response across nearly every tool on the market, and there isn't much out there offering anything else.

Here's the thing that makes this a real gap and not just a preference: fraud investigations may never hit a 100% guaranteed answer. The job isn't about certainty, it's about arriving at a reasonable conclusion, backed by enough data points and critical thinking that you can defend it if someone questions it later. A tool that only understands "flag" or "don't flag" has no way to represent that reasoning, which means the actual thinking, the part that makes a decision defensible, ends up living somewhere the tool can't see: a notebook, a comment thread, a memory you're hoping doesn't fade before the case gets reviewed. I wrote about calibrating certainty to stakes in more detail here if you want the investigator-side version of this same idea.

The fix isn't complicated, even if nobody's building it yet. Let the interface represent a range, not a switch, and give investigators somewhere to attach the reasoning behind wherever they land on that range.

03. There's no good place to park a half-formed thread, and no real audit trail once you do

Case management tools are almost always built around a linear model: open a case, investigate, close a case. But real investigations don't work that way. You notice something, you're not ready to act on it, and you need somewhere to put it down without losing it. I still keep a physical notebook next to my desk for exactly this reason, because no tool I've used has given me a good digital equivalent.

This sounds like a small feature request. It isn't. The absence of it is why so many patterns get missed, not because nobody noticed the first instance, but because nobody had anywhere to write it down, and by the time the third instance showed up, the first two were already forgotten.

It's also bigger than personal note-taking. Solid case management within the tool itself, not bolted on, not offloaded to a notebook or a spreadsheet, is fundamental to how a fraud program actually functions. You need a place to document the decisions you made and why, and you need an audit trail that shows how a case moved from first thread to final call. Whether you're buying a tool or building one in-house, this isn't a nice-to-have. It's the thing that lets you defend a decision six months later when someone asks why, and it's often what regulators and auditors are actually looking for when they review your process.

04. Absence isn't queryable

Every tool I've used is built to search for what's present: a transaction, a flag, a matching record. Almost none of them make it easy to search for what's conspicuously missing, even though that's often the stronger signal. A nonprofit with a polished website and zero press mentions. A donor history with suspiciously round numbers and nothing else. An account with a normal volume of activity and an abnormal absence of the kind of friction real users generate.

Building for absence is harder than building for presence, which is probably why it rarely gets prioritized. But it's exactly the kind of feature that separates a tool investigators trust from one they route around.

05. Impressive new data points, duplicated old work

Sales pitches love to lead with the fancy stuff: proprietary data points, extra fraud filters, signals you can't get anywhere else. Paying for genuinely new data is often a good investment, and I'm not knocking it. But the pitch rarely asks a more basic question: what happens to the data and workflows you already have?

If a tool doesn't integrate with your existing systems, you end up doing the same investigative work twice, once inside the shiny new platform and once again in whatever system you were already using to reach your actual conclusion. That's not a productivity nice-to-have. That's duplicated effort, on every single case, forever, and it quietly erodes whatever time the new data points were supposed to save you.

06. Escalation paths follow the org chart, not the case

Routing logic in most tools mirrors the reporting structure: this case type goes to this queue, which reports to this manager, who escalates to this director. That makes sense on a slide. It falls apart the moment a case needs Legal, Product, and a senior investigator in the same room within the hour, and the tool has no concept of "who actually needs to see this right now" that isn't just "whoever sits above me."

The best escalation systems I've helped build were the ones that let the case dictate the path, not the hierarchy. A front-door intake process that routes based on severity and case type, not seniority, gets urgent things in front of the right eyes faster, and that speed is often the entire ballgame.

07. False positives get treated as a tuning problem, not a business problem

Every dismissed false positive is logged somewhere as a metric to improve. What rarely gets tracked is everything that metric is quietly costing.

There's the human cost first: alert fatigue. After enough false positives, people stop trusting the flag, and once that trust is gone, they start second-guessing everything the system tells them, including the flags that were actually right. That's a much bigger problem than a slightly noisy model, and it almost never shows up on a roadmap because it doesn't show up in a dashboard.

But it doesn't stop there. Every false positive that reaches a legitimate user is friction on someone who did nothing wrong, an extra verification step, a held transaction, a declined donation, and enough of that friction is exactly how you lose good users and the revenue that comes with them. And if the tool charges per check, which plenty of vendor pricing models do, a high false positive rate means you're paying full price to flag people who were never the problem in the first place. That's not a tuning inefficiency. That's money you're actively burning, case by case, on top of everything else it's costing you.

What actually helps

None of this is a complaint without a counterpoint. The tools and frameworks I've had a hand in building were designed specifically around these gaps, because I was building them for people who'd feel the same friction I did, and because I made a point of testing them with the people who'd actually use them, not just presenting the finished version.

The NPO risk scoring framework I built maps findings to a tier and a recommended action instead of a flat flag, so the output actually reflects a spectrum of confidence rather than a single binary call, and the reasoning behind that tier travels with it. The investigation agent I designed has a full analyst sign-off step built in on purpose, not as an afterthought, because automation that removes the human's ability to say "wait, that doesn't feel right" is automation that will eventually be wrong in a way nobody catches in time. Both were built to work with our existing systems rather than beside them, because a tool that makes you repeat your own work isn't actually saving anyone time.

Good fraud tooling, in my experience, tends to share a few things in common. It's evaluated with input from the people who'll use it daily, not just the people who'll approve the invoice. It represents confidence as a range, with room to attach the reasoning behind it. It gives investigators somewhere to park a thread they're not ready to pull. It treats the absence of a signal as seriously as the presence of one. It integrates with what you already have instead of asking you to duplicate the work. It routes based on urgency, not org chart. And it tracks trust in the tool as its own metric, not just accuracy on paper.

Why this matters for whoever's building the next one

The cost of getting this wrong isn't abstract. It's slower cases, because investigators are working around the tool instead of with it, or doing the same analysis twice because the tool doesn't talk to their existing systems. It's missed patterns, because there was nowhere to write down the half-formed thread that would have connected them. And eventually, it's good people leaving, because nothing burns out a sharp investigator faster than a tool that doesn't trust their judgment enough to represent it accurately.

If you're building or buying fraud investigation tooling, or Trust & Safety tooling more broadly, the best thing you can do is spend real time with the people who'll actually live in it, not just the people who'll sign off on buying it. Budget and feature scope are legitimate constraints, but they're only half the equation. The gap between the pitch and the desk is almost always visible from the desk. It just rarely gets asked about from the other side of the table.

Key takeaways