Function 4 · the fraud desk
At what score should the desk decline, and what can it honestly say about its own precision?
The desk declines at 0.0503, the threshold that minimised expected cost on the validation months. On the test months that costs $24,927 against $1,573,327 for approving everything, a saving of 98.4%. Labels arrive late: on the as of date 200,569 of 472,370 decisions were still inside the 60 day chargeback window, so the desk reports precision over labelled decisions only and says how many it cannot yet judge.
The desk, as it runs
The API replays the scored test stream on a simulated clock: an asyncio task reading a file, not a message bus. When it is asleep this page plays the same stream, recorded, in your browser, and says so.
The feed
| Time | Card | Category | Amount | km | Score | Decision |
|---|
The alert queue
Declines and reviews, the queue point (…) first, then by score.
- No alerts in the window yet.
Move the threshold
Expected cost of fraud by decline threshold, test months
Expected cost
Source: replayed, Sparkov card transactions, June to August 2026, as of 2026-08-31.
Precision and recall across thresholds, test months
Share
Source: replayed, Sparkov card transactions, June to August 2026, as of 2026-08-31.
What the desk could know, day by day
Scored decisions
Source: replayed, Sparkov card transactions, June to August 2026, as of 2026-08-31.
The fairness audit
The card model's errors by age band and gender, beside application fraud on the BAF suite, which was published with its protected attributes kept so this question can be asked. Neither model sees age; both are measured on it.
Who the card model declines wrongly, by age band and gender
Source: replayed, Sparkov card transactions, June to August 2026, as of 2026-08-31.
Application fraud under the BAF suite's fairness protocol
Source: real:baf, Bank account applications in the BAF suite, base and variant II, 300,000 each, as of 2026-08-31.
What parity costs: the search for a less discriminatory alternative
Predictive equality (one is parity)
Source: real:baf, Bank account applications in the BAF suite, base and variant II, 300,000 each, as of 2026-08-31.
Method and limitations
- The stream is Sparkov's generated card transactions (303,401 in the test months), not real card traffic. The model's AUROC of 1.000 is a property of the generator, which makes fraud easier to see than any real desk would find it.
- Cost assumes $12.00 per false decline and each missed fraud's amount plus $35.00. Both are stated assumptions in the configuration, not measured costs.
- Labels arrive exactly at the end of the chargeback window for every transaction. Real chargebacks arrive over the window, so a real desk would have some labels earlier.
- The live desk is a replay on a simulated clock. The API scores nothing new; it reads the scored test stream and reproduces its timing.