How Stripe Radar catches fraud in under 100 milliseconds without blocking good customers
Catching fraud is easy if you block everything. The hard problem is catching it without turning away good customers, in the fraction of a second a checkout allows. Here is how Stripe Radar does both.

Fraud is rare. Stripe puts it at roughly 1 in 1,000 payments. That rarity is exactly what makes it hard: a system that simply declares every payment safe is right 99.9 percent of the time and completely useless. And the obvious fix, blocking anything suspicious, has its own cost, because every good customer you wrongly decline is a lost sale and often a lost customer. Stripe cites that a third of shoppers will not come back after a single false decline. So the real job of Stripe Radar, the fraud system Stripe's engineers described in 2023, is a balancing act performed in the time a checkout button takes to respond: score every payment against more than a thousand signals, in under 100 milliseconds, and block the fraud without blocking the business.
The problem: rare, fast, and expensive to get wrong
Three constraints shape everything. First, imbalance: because fraud is 1 in 1,000, ordinary accuracy is meaningless and you have to think in terms of how many of your blocks are actually fraud (precision) and how much of the real fraud you catch (recall). Second, the cost of a false positive is real money: Stripe walks through an example where, once you count lost margin and chargeback fees, a single fraudulent sale can wipe out the profit of nearly nineteen good ones, which means even a small false-decline rate can cost more than the fraud does. Third, speed: the score has to be computed live, during the payment, so every feature has to be fetched and updated in milliseconds.
Rules alone cannot square this. A rule like "block every foreign card" is easy to write and blocks a mountain of good customers along with the bad. Machine learning earns its place by finding the more nuanced patterns a human rule misses.
How they built it
The picture at the top of this article is the intuition: a model reads a payment's features, the amount, the card's country, how many countries the card has touched, and maps them to a fraud probability. A hand-drawn decision tree like that is only a teaching aid. The real Radar is a deep neural network over hundreds of signals, but the idea, features in, probability out, is the same.
Features from a network, not just a transaction. Radar uses hundreds of signals, and most are aggregates computed across Stripe's whole network rather than facts about the single payment in front of it. Because a large majority of cards on the network have been seen before, Stripe has context a lone merchant never would. Categorical facts like the issuing bank or the country get turned into learned "embeddings," positions in a space where similar entities sit close together, so a fraud pattern first seen in one country can be recognized in another without retraining from scratch.
A model that got simpler on the surface and deeper underneath. Radar started as an ensemble: gradient-boosted decision trees to memorize known patterns, paired with a neural network to generalize to new ones. In 2022 Stripe moved to a single deep neural network. Stripe reported this cut training time by more than 85 percent, to under two hours, and, more importantly, unlocked techniques like transfer learning and embeddings that the older setup could not support.

Since then Stripe has gone further, training a large transformer model on billions of transactions that compresses each payment into a reusable embedding, and in 2025 announcing a "payments foundation model" built on that idea.
Where to draw the line
A fraud score is only half the system. The other half is the threshold: above what score do you block? This is a genuine trade-off, and Stripe makes it visible rather than hiding it.

Push the threshold up and you wrongly decline fewer good customers but let more fraud through. Push it down and you catch more fraud but annoy more real buyers. There is no universally correct point on that curve, only the point that fits a given business's margins and risk appetite, which is why Stripe shows businesses the curve, benchmarks them against similar companies, and lets them layer their own rules, manual review, or an extra verification step on top. For fast-moving attacks like card testing, the threshold is not even fixed: separate models estimate how big an attack is and where it is concentrated, and tighten the screws only on the precise slice of traffic under attack, then ease off when it subsides. Every new model is also checked against guardrail metrics per business before launch, because a model that is better on average can still be worse for your particular customers.
What Stripe reports, and what has changed since
The numbers here are Stripe's own, not independently audited, so read them as the company's reporting. Stripe says Radar wrongly blocks only about 0.1 percent of legitimate payments, that successful card-testing attacks on its network fell about 80 percent over two years, and its current materials cite roughly a 32 percent average reduction in fraud for businesses that turn Radar on. Recent direction:
- 2024: Radar Assistant, which lets a user write a fraud rule in plain language and backtest it before turning it on.
- 2025: Radar extended to bank debits (ACH and SEPA), and the unveiling of the transformer-based payments foundation model, which Stripe said sharply improved detection of attacks on large businesses.
- 2026: Radar expanded beyond cards to wallets, buy-now-pay-later, and more, with custom per-business models and cross-method network effects, so an IP and device flagged on one payment type are flagged across others.
Why it matters for your business
You do not need Stripe's scale to steal the thinking. Any business taking online payments, or fighting fake signups, free-trial abuse, and bot traffic, faces the same shape of problem:
- Stop measuring accuracy. When the bad thing is rare, accuracy lies. Track precision and recall, and put a rupiah value on each: what a missed fraud costs you, and what a wrongly blocked customer costs you.
- Decide where the line goes on purpose. The block threshold is a business decision, not a technical default. Make it visible and tune it to your margins.
- Shared signal beats isolated rules. Patterns across many customers, devices, and time windows catch what a single hand-written rule cannot.
How we would build it today
For an Indonesian marketplace, SaaS, or online store, we would start with a gradient-boosted model over your own history of transactions and signups, with engineered velocity features: how many attempts came from this device, card, or IP in the last hour and day, which is where a lot of fraud and abuse actually shows up. We would turn the score into a three-way decision, allow, review, or block, with thresholds you set against your own margins, and send the middle band to a human review queue rather than auto-blocking it, so a real customer is never silently turned away. We would measure precision and recall in money, not percentages, and, as Stripe now does, use a language model to help draft and backtest rules in plain language. Same balancing act, sized to your traffic.