The situation
Fraud detection at volume is rarely limited by modeling capability. It is limited by what the organization can operationally absorb. Any team can raise detection rates by loosening thresholds; the resulting review queue and the customer friction that follows are what make that approach unusable.
Across several banking and insurance engagements the presenting problem was consistent: chargeback losses were running higher than leadership wanted, and previous attempts to address it had produced review volumes the operations teams could not staff.
The approach
Post-screen transaction fraud. Scoring after authorization rather than during it changes the problem usefully, because the model gains context that was not available in the authorization window and loses the hard latency constraint. That allowed richer features and more careful calibration, tuned explicitly to hold the false-positive rate under 8% rather than to maximize raw detection.
Internal adjuster fraud. Insurance fraud committed inside the organization typically runs below 0.1% of claims, which is a base rate at which conventional classification produces noise. Hierarchical models that pooled signal across adjusters, claim types, and geographies produced high-probability flags for targeted review rather than a scored list nobody could prioritize.
The outcome
Annual chargeback reduction ranged from $2 million to $10 million depending on the size of the book, with false-positive review held below 8%. The internal fraud work produced a defensible flagging process for auto and property claims, which mattered to the audit function as much as to the loss numbers.
What made it work
The threshold conversation happened before the modeling, not after. Agreeing with fraud operations on what review volume they could sustain, and designing backwards from that constraint, is the difference between a model that ships and a model that gets quietly disabled two months after launch.
