Ringfence

Abuse-ring detection for card-not-present fraud, priced in rupees rather than ranked by AUC.

A merchant does not lose money to fraudsters one at a time. They lose it to rings — one operator running many cards through a handful of shared addresses, devices and email domains. Scored individually each transaction is unremarkable; scored as a group the operator is obvious.

Card fraud is not a rounding error and it is not shrinking. What makes it worse than the sticker price is that a reversed payment costs a merchant several times the disputed amount: the goods are gone, the money is clawed back, a penalty is levied, and staff time is spent contesting it.

What follows is the measured case: §1 lays out the pipeline and where its leakage boundary sits; §2 reports held-out detection performance against an ablation with the ring layer removed; §3 turns the score into an operating point priced in rupees; §4 breaks each honesty guarantee in turn to show what it cost; §5 shows the review queue an analyst would work. §6 describes the dispute side, which is parameterised rather than measured, and says so. (§ is the section sign — §3 reads “section 3”.)

This page reports held-out performance for a ring-aware detector on the IEEE-CIS card-not-present dataset, whose ground truth is reported chargebacks. Every figure is generated from the artefacts in reports/ by scripts/build_site.py, so this page cannot drift from the experiment that produced it.

$0 $25 bn $50 bn $33.41 bn 2024 actual $41.06 bn 2030 projected $403.88 bn cumulative over ten years
Figure 1 Card fraud losses worldwide. Only the two published points are plotted; the connector is dashed because the intervening years are a projection, not a measured series. Source: Nilson Report.
ordinary traffic card P card Q card R own device own device own device One card each, weeks apart. Nothing links them. ring 76963 — observed card A card B card C card D Z965 Build/NMF26V one shared device 5 transactions of exactly ₹44,000 — amount variability 0.000 4 cards, 0.37 days, ₹220,000 total ordinary traffic card P card Q card R own device own device own device One card each, weeks apart. Nothing links them. ring 76963 — observed card A card B card C card D Z965 Build/NMF26V one shared device 5 transactions of exactly ₹44,000 amount variability 0.000 4 cards, 0.37 days, ₹220,000 total
Figure 2 The detection premise, drawn from an actual ring in the held-out period. Taken one at a time, every transaction in the ring is unremarkable: a plausible amount, a valid card, a real address. Only the shape gives the operator away — and once one member is flagged, the other three are known before they cost anything.

1 How it works

Two execution planes, not one pipeline. Training runs on a schedule over history; scoring runs per transaction the moment it arrives. They share state and, critically, share the same feature code — features/causal.py is the identical module in both, which is why performance measured offline means anything online.

every node links to the file that implements it

loading diagram…

Figure 3 System architecture. Two execution planes — offline training over history, and online scoring per transaction — share one entity graph, one model artifact, and the same feature module, which is what makes the offline metrics predictive of online behaviour. The red path is the feedback loop: a confirmed chargeback re-enters the ring's fraud history, but only after the reporting lag, which is why that feature can be used honestly. The dispute branch is triggered by the issuer, not by the model, and terminates at a human.

2 Held-out performance

Evaluated on transactions from a period strictly later than all training data, with a 7-day embargo at the boundary. The ablation column retrains the identical pipeline with every ring and entity feature removed.

Table 1 Detection performance with and without the ring layer. precision@100 is the one measure where the ring model does slightly worse; both are saturated at the top of the ranking.

3 Choosing the operating point

Precision and recall trade against each other, so the threshold is not a modelling choice — it is an economic one. Both mistakes are priced: a missed fraud costs the goods, the amount and a penalty fee; a blocked legitimate customer costs that order's margin, scaled for the share who never return.

Drag the threshold. Flagging everything catches all the fraud and is catastrophic.

Net saving
Figure 4a Net saving against the share of traffic flagged for review, in rupees, relative to a no-model baseline of approving everything. The vertical axis is clipped; the curve continues to when every transaction is flagged. Log horizontal axis.
Precision Recall
Figure 4b Precision and recall against the same horizontal axis. Both are unitless proportions, so they share one scale.

4 What the honesty costs

Claiming a leakage boundary is worth nothing unless the gap is measured. Each guarantee was broken in turn, changing nothing else. The variants below are not results — they are the size of the overstatement avoided.

Variant
Valid measurement Invalid — assumption cannot hold
Figure 5 PR-AUC and net saving under each variant, as separate panels because the two quantities share no scale. Either shortcut lifts PR-AUC from 0.60 to roughly 0.96 and nearly triples the reported saving.

Getting this wrong, first time

The first version of this ablation replaced the causal features rather than leaking the same ones. That also dropped the lag-gated fraud-history feature, so the “leaky” variant scored 10.5% worse — proving nothing except that it had lost an important feature. A leakage test must vary causality and nothing else.

5 Ring review queue

A ranked list of transaction identifiers is not a product; the reviewer still has to reconstruct what happened. Each ring below carries a generated case file stating what links the accounts, what the model reacted to, a graded and reversible action, and — mandatory — the most plausible innocent explanation.

Table 2 Highest-scoring multi-card rings in the held-out period. Select a row for its case file.

6 Disputes that land anyway

Detection is never perfect. When a chargeback does arrive, a merchant can contest it — and published figures put the gross win rate at 44.6% but net recovery at 10.7%. That gap is the whole argument for triage: contesting everything is what produces it.

There is no public dataset of dispute outcomes, so rather than train a win model on invented labels, the detector is inverted. Conditioned on a chargeback having been filed:

Table 3 How the fraud score is read once a dispute exists.

Fraud scoreReadingAction
HighThe card really was stolen; the cardholder is a victim telling the truthAccept — unwinnable, and fighting a real victim is wrong
LowThe transaction looked legitimate on every signal yet is disputed — the friendly-fraud signatureContest

So a model validated on real labels does the work, and the ethical constraint falls out of the same signal. Where the decision is to contest, an agent drafts the Razorpay evidence payload; every document id is checked against the merchant's actual inventory in code, because a hallucinated id is not a formatting bug but a fabricated evidence claim sent to a bank. The payload is hard-wired to draft.

This section carries no measured results, and should not be read as if it did. The triage economics are parameterised from the published figures above; the evidence documents used in testing are illustrative, since IEEE-CIS contains no merchant document store.

7 Limitations

    8 Reproduction

    Every number on this page is regenerated by the commands below. The page reads docs/data.json, which build_site.py writes from reports/.

    uv venv --python 3.11 && uv pip install -e .
    ringfence data download
    ringfence data prepare --tag main
    ringfence train --tag main
    ringfence train --tag noring --data-tag main --no-ring-features
    ringfence ablate --force
    python scripts/build_site.py --with-casefiles