Function 4 · the fraud desk

At what score should the desk decline, and what can it honestly say about its own precision?

The desk declines at 0.0503, the threshold that minimised expected cost on the validation months. On the test months that costs $24,927 against $1,573,327 for approving everything, a saving of 98.4%. Labels arrive late: on the as of date 200,569 of 472,370 decisions were still inside the 60 day chargeback window, so the desk reports precision over labelled decisions only and says how many it cannot yet judge.

Operating point
0.0503
Chosen on validation months by expected cost
Queue point
0.2759
Validation precision floor 90%
Precision at the operating point
74.8%
Test months
Recall at the operating point
97.6%
Test months

The desk, as it runs

The API replays the scored test stream on a simulated clock: an asyncio task reading a file, not a message bus. When it is asleep this page plays the same stream, recorded, in your browser, and says so.

Connecting to the desk…
Replay speed
Approved
…
0 transactions
Sent to review
…
0 transactions
Declined
…
0 transactions
Precision and recall, labelled decisions only
…

The feed

TimeCardCategoryAmountkmScoreDecision

The alert queue

Declines and reviews, the queue point (…) first, then by score.

  • No alerts in the window yet.

Move the threshold

Expected cost of fraud by decline threshold, test months

$0.0M$0.5M$1.0M$1.5M$2.0M1e-40.0010.010.11Decline threshold (score, log scale)operating pointTotalMissed fraudFalse declines

Expected cost

False declines cost $12.00 each and missed fraud costs its amount plus $35.00. The threshold that minimised the sum on the validation months, 0.0503, costs $24,927 on the test months against $1,573,327 for approving everything. Accuracy would have chosen 0.6376 and cost $89,153.

Source: replayed, Sparkov card transactions, June to August 2026, as of 2026-08-31.

Precision and recall across thresholds, test months

0%20%40%60%80%100%1e-40.0010.010.1Decline threshold (score, log scale)operatingqueuePrecisionRecall

Share

At the operating point the desk declines with precision 74.8% and recall 97.6%; the queue point narrows to precision 89.7% for the analyst queue.

Source: replayed, Sparkov card transactions, June to August 2026, as of 2026-08-31.

What the desk could know, day by day

050,000100,000150,000200,000250,000300,0002026-06-012026-06-132026-06-252026-07-072026-07-192026-07-312026-08-122026-08-24DayLabelledUnlabelled

Scored decisions

On 2026-08-31, 200,569 of 472,370 scored decisions were still inside the 60 day chargeback window. Precision over the labelled ones was 73.4%. A figure that treated the unlabelled declines as mistakes would read 27.0%.

Source: replayed, Sparkov card transactions, June to August 2026, as of 2026-08-31.

The fairness audit

The card model's errors by age band and gender, beside application fraud on the BAF suite, which was published with its protected attributes kept so this question can be asked. Neither model sees age; both are measured on it.

Who the card model declines wrongly, by age band and gender

Aged under 250.20%Aged 25 to 390.29%Aged 40 to 590.33%Aged 60 and over0.34%Female0.36%Male0.26%
At the operating point, genuine transactions of cardholders 60 and over are declined at 0.34% and those of cardholders under 25 at 0.20%, a ratio of 0.60. Sparkov generates age and gender, so they can be measured here; the model never sees either.

Source: replayed, Sparkov card transactions, June to August 2026, as of 2026-08-31.

Application fraud under the BAF suite's fairness protocol

Base, 50 and over9.19%Base, under 504.19%Variant II, 50 and over6.96%Variant II, under 503.04%
At a 5% false positive rate the model finds 49.0% of fraudulent applications in the base variant and 49.3% in variant II. Genuine applicants aged 50 and over are flagged at 9.19% against 4.19% for the rest, a predictive equality of 0.46 (0.44 in variant II), though the model never sees age.

Source: real:baf, Bank account applications in the BAF suite, base and variant II, 300,000 each, as of 2026-08-31.

What parity costs: the search for a less discriminatory alternative

0.000.200.400.600.801.0044%45%46%47%48%49%Recall at a five percent false positive rateBaseVariant II

Predictive equality (one is parity)

The model's own features predict whether an applicant is 50 or over with an AUROC of 0.782. Removing the strongest carriers of age one at a time moves the base variant's predictive equality from 0.46 to 0.57 and its recall from 49.0% to 44.2%. The model is kept and bounded to asking for documents, never to declining.

Source: real:baf, Bank account applications in the BAF suite, base and variant II, 300,000 each, as of 2026-08-31.

Method and limitations

  • The stream is Sparkov's generated card transactions (303,401 in the test months), not real card traffic. The model's AUROC of 1.000 is a property of the generator, which makes fraud easier to see than any real desk would find it.
  • Cost assumes $12.00 per false decline and each missed fraud's amount plus $35.00. Both are stated assumptions in the configuration, not measured costs.
  • Labels arrive exactly at the end of the chargeback window for every transaction. Real chargebacks arrive over the window, so a real desk would have some labels earlier.
  • The live desk is a replay on a simulated clock. The API scores nothing new; it reads the scored test stream and reproduces its timing.