Skip to content
CDO On-Demand
← All case studies

Cutting chargebacks by millions while keeping false positives under 8%

High-volume transaction fraud is a balancing problem rather than a detection problem. Catching more fraud is straightforward if you accept a review queue nobody can staff and a customer experience nobody can defend. This work optimized post-screen detection to reduce annual chargebacks by $2M to $10M while holding false-positive review below 8%.

Client
Banking and insurance clients
Engagement
Analytics leadership across multiple institutions

$2M–$10M

Annual chargeback reduction, depending on book size

<8%

False positive rate on flagged transactions

<0.1%

Internal adjuster fraud detected through hierarchical models

The situation

Fraud detection at volume is rarely limited by modeling capability. It is limited by what the organization can operationally absorb. Any team can raise detection rates by loosening thresholds; the resulting review queue and the customer friction that follows are what make that approach unusable.

Across several banking and insurance engagements the presenting problem was consistent: chargeback losses were running higher than leadership wanted, and previous attempts to address it had produced review volumes the operations teams could not staff.

The approach

Post-screen transaction fraud. Scoring after authorization rather than during it changes the problem usefully, because the model gains context that was not available in the authorization window and loses the hard latency constraint. That allowed richer features and more careful calibration, tuned explicitly to hold the false-positive rate under 8% rather than to maximize raw detection.

Internal adjuster fraud. Insurance fraud committed inside the organization typically runs below 0.1% of claims, which is a base rate at which conventional classification produces noise. Hierarchical models that pooled signal across adjusters, claim types, and geographies produced high-probability flags for targeted review rather than a scored list nobody could prioritize.

The outcome

Annual chargeback reduction ranged from $2 million to $10 million depending on the size of the book, with false-positive review held below 8%. The internal fraud work produced a defensible flagging process for auto and property claims, which mattered to the audit function as much as to the loss numbers.

What made it work

The threshold conversation happened before the modeling, not after. Agreeing with fraud operations on what review volume they could sustain, and designing backwards from that constraint, is the difference between a model that ships and a model that gets quietly disabled two months after launch.

Why this is relevant to you

Business, technical, and program together.

The business lens

Fraud teams are measured on losses prevented, and they are constrained by review headcount and customer friction. A model that doubles detection while tripling the queue is a failure in operational terms. The framing throughout was the trade-off curve between recovery, review cost, and customer experience, which is the conversation the business actually needed to have.

The technical work

Post-screen scoring operates after authorization, so the model has more context and less time pressure than an authorization-time decision. Internal adjuster fraud is a rarer problem still, running below 0.1% of claims, which makes naive classification useless. Hierarchical models that pool information across adjusters and claim types were necessary to get usable signal at that base rate.

Program and organization

These programs touched fraud operations, claims, customer service, and compliance simultaneously. Getting them into production meant agreeing thresholds with the people who owned the review queue and documenting the decisions for audit before anything went live.

Services

Fraud detectionModel designRisk analytics

Stack

Hierarchical modelsPost-screen scoringPythonSQL

Have a similar problem?

If that resembles the situation in your own organization, a short call is the quickest way to establish whether the same approach would apply to you.