MaShop/Journal/Industry/The Model That Scores You, Not Just Your Customer
● IndustrySeptember 21, 2026
Read · 5 min
merchant risk score · payments

The Model That Scores You, Not Just Your Customer

Your processor runs one model on each payment and another on your business. The second one holds your money, and the numbers behind it are published.

Key takeaways
  • Two different models judge you. One scores each payment for fraud, the other scores your account for risk, and only the first one shows up in your dashboard.
  • Stripe Radar scores payments from 0 to 99, with 65 and above treated as elevated risk and 75 and above as high risk by default.
  • Stripe says it deliberately alters the reported score on a small subset of payments to measure model performance, so an inexplicable score is sometimes exactly that.
  • Visa's merchant threshold under VAMP moved to 1.5 percent on 1 April 2026, down from 2.2 percent, and enrolled merchants are assessed 8 dollars per disputed or fraudulent transaction.
  • Mastercard's excessive tier needs 100 to 299 chargebacks and a ratio of 1.50 to 2.99 percent together, and the fee ladder reaches 100,000 dollars a month at 19 months.
  • Leaving a programme takes three consecutive clean months, so a bad November can still be costing you in March.

Most sellers learn about merchant risk scoring on the day a payout does not arrive. The dashboard shows the money. The bank shows nothing. Support replies that the account is under review and cannot say more, which is true and unhelpful in equal measure.

The thing worth understanding is that you were being scored long before that morning, by two separate systems with different jobs. Confusing them is why so many sellers spend a week fixing the wrong problem.

Which model is actually judging you?

Both of them, at different levels. The first scores individual payments for fraud and is visible to you. The second scores your business as a counterparty and is not.

Stripe's documentation on risk evaluation describes the first one plainly: adaptive AI models that use hundreds of risk factors about each payment plus data across its network of businesses, learning from new purchase patterns and from feedback when you mark a payment fraudulent. Each payment gets a score from 0 to 99. By default 65 or above indicates elevated risk and 75 or above indicates high risk, with high risk payments blocked automatically.

The second system is the one that holds your money. It is not scoring whether this buyer is a fraudster. It is scoring whether your business will still be solvent and present when the disputes land in ninety days, because the processor carries that liability if you are not.

Note

These are different models with different aims, and a shop can look clean on one and risky on the other. A sudden tripling of legitimate revenue lowers your dispute ratio and raises your counterparty risk at the same time.

Why would a legitimate spike look bad?

Because the processor's exposure grows before your track record does. A shop that normally settles 8,000 dollars a month and suddenly settles 90,000 has just increased the money at risk elevenfold without any new evidence that it can deliver at that volume.

Stripe's own guidance for platforms managing connected accounts describes the toolkit. Its managed risk documentation explains that in response to risk signals it might slow or pause payouts, pause an account's ability to process charges, or hold a reserve on the account balance. Those are three distinct interventions and sellers routinely describe all of them as being frozen.

Which one you got matters. A paused payout with processing still live is a documentation problem. A paused ability to charge is an existential one, because your storefront is taking orders it cannot collect on.

Sequence diagram showing how a risk signal escalates through review, held payouts, a rolling reserve, programme enrolment and finally exit or closure

What are the numbers the networks actually watch?

Ratios, counted monthly, against thresholds that changed recently and in the stricter direction. This is the part sellers can measure themselves, and almost none do until it is too late, which is also why winning a chargeback representment matters beyond the disputed amount.

Visa replaced two older programmes with a single combined measure. The Merchant Risk Council's account of the current thresholds describes the ratio as reported fraudulent transactions plus total disputes divided by total settled transactions, counting card not present transactions only, with fraud arriving as TC40 reports and disputes as TC15. The merchant threshold became 1.5 percent on 1 April 2026, down from 2.2 percent. Acquirers sit against 0.5 percent and 0.7 percent. Enrolled merchants are assessed 8 dollars per fraudulent or disputed transaction.

Mastercard counts differently, and the difference is worth internalising. Braintree's documentation of the excessive chargeback programme gives the excessive tier as 100 to 299 chargebacks with a ratio of 1.50 to 2.99 percent, and the high excessive tier as 300 or more chargebacks with a ratio of 3.00 percent or higher. Both the count and the ratio have to be met together. The ratio uses this month's chargeback count divided by last month's sales count, so a month where sales fell and disputes from the busy month arrived can push you over on arithmetic alone.

MeasureVisa VAMPMastercard excessive tierWhat it means for you
What is countedFraud reports plus disputes, card not present onlyChargebacks onlyFraud reports can hurt you at Visa without a single dispute
DenominatorTotal settled transactionsPrior month sales countA falling sales month raises your Mastercard ratio on its own
Merchant trigger1.5 percent from 1 April 2026100 to 299 chargebacks and 1.50 to 2.99 percentMastercard needs both conditions, Visa needs one
Cost when enrolled8 dollars per fraudulent or disputed transactionMonthly ladder, 1,000 dollars at month two rising to 100,000 at 19 monthsTime in the programme costs more than the ratio does
Getting outReturn below the thresholdThree consecutive months below the tierRecovery is measured in quarters, not weeks

Can you see this coming before your processor calls?

Usually yes, and the tooling already exists inside accounts sellers log into every day. This is the single most useful thing in this article.

Adyen's documentation on monitoring risk performance describes a card monitoring programmes tab where you can see your monthly chargeback and fraud count, your rate, and your projected VAMP ratio against the programme thresholds, filtered by Visa or Mastercard. A projection is exactly what you want, because the ratio that matters is the one at month end and the month is not over.

If your provider surfaces something similar, put it in a monthly routine. If it does not, you can compute the Mastercard shape yourself from two numbers you already have: this month's chargeback count and last month's order count. That is a spreadsheet cell, not a project.

Card listing the four numbers a seller should check monthly, covering dispute count, dispute ratio, fraud reports and projected VAMP ratio

Why does an unexplainable score sometimes have no explanation?

Because a small number of scores are altered on purpose. Buried in Stripe's risk evaluation documentation is a sentence that deserves more attention than it gets: for a small subset of payments, Stripe modifies the reported risk score so it can measure model performance and gather data for later model development, keeping metrics such as false positive rate and recall within desirable ranges.

That is a defensible engineering practice and it has a consequence for you. If you are staring at one transaction whose score makes no sense against a hundred similar ones, you may be looking at a measurement sample rather than a signal. Chasing it will teach you nothing.

The lesson generalises. Investigate patterns across many payments, never a single score. Individual model outputs are noisy by design in a way that dispute ratios are not, which is the same reason we argued against reading one false decline as a system failure in what false declines really cost a small shop.

What do the risk levels actually mean day to day?

There are five of them, not three, and the two extra ones cause the most confusion. Stripe's documentation lists high risk, elevated risk, normal risk, not evaluated and unknown risk, each surfacing on the charge as a risk_level value.

High risk is blocked by default and appears as highest. Elevated risk is allowed by default and can be routed to a review queue depending on your plan. Normal risk is authorised, with the documentation still warning that normal risk payments can turn out to be fraudulent. So far this matches intuition.

The other two do not. Not evaluated, which appears as not_assessed, covers all non card payments other than ACH and SEPA Direct Debit, card payments that predate the public assignment of risk levels, and any business that has opted out of Radar risk assessment. Unknown risk means the evaluation itself failed through an error. Neither is a judgement about the buyer, and reading either as a green light is a mistake: the payment simply was not scored.

A seller taking payment by bank transfer, wallet or local method in several markets can therefore have a large share of volume that no fraud model ever looked at. That is not a reason to avoid those methods. It is a reason to know which slice of your revenue is unscored, because your own checks are the only checks operating on it.

Why does a subscription behave differently?

Because the model does not rescore every renewal. The documentation states that for recurring billing the fraud models score only the initial payment of a subscription, while rules are still evaluated for all payments.

That distinction repays a moment's thought if you sell memberships, boxes or retainers. The machine learning judgement is made once, at signup, on the day your customer is least known to you. Every renewal after that passes through your rules but not through a fresh model score. A card that was clean in January and has since been reported stolen is caught by other machinery, not by a rescore of the subscription.

The practical consequence is that subscription sellers get less protection from the model than single purchase sellers, on exactly the payments that make up most of their volume. If your dispute ratio is drifting up and your product is a subscription, look at renewals and at cancellation friction before you look at fraud at all. A customer who could not find the cancel button and called the bank instead produces a chargeback that no fraud model was ever going to prevent.

How much does the network already know?

More than your own history can tell it, which is both reassuring and a little uncomfortable. Stripe publishes the odds that a payment instrument it sees is already familiar: a 92 percent chance it has seen the card before, 82 percent for a SEPA account and 71 percent for an ACH account.

Read as a seller, that explains why a brand new shop with no transaction history still gets usable fraud scoring from day one. The model is not learning from you, it is recognising the instrument from everybody else. It also explains why your feedback matters less than the marketing suggests and more than nothing: marking a payment fraudulent adds the email and card fingerprint to your block lists and feeds the shared models, but you are one voice in a very large chorus.

The asymmetry cuts the other way too. Your counterparty score, unlike the fraud score, is mostly about you, because nobody else's data says whether your shop ships. That is why documentation closes an account review and no amount of clean payment history does it on its own.

What does a reserve do to your cash?

It converts a margin problem into a liquidity problem, and the arithmetic is worth doing before it happens rather than during. Take a shop settling 20,000 a month with a 10 percent rolling reserve released after 90 days.

In the first month you receive 18,000 instead of 20,000. In the second and third, the same. From the fourth month onward you receive 20,000 again, because the reserve released from month one arrives while the current month's reserve is withheld. The steady state cost is not 10 percent of revenue. It is a one off working capital hit of roughly 6,000, held permanently while the reserve remains in place.

That framing changes the decision. A 10 percent reserve is survivable for a business with 6,000 of slack and fatal for one operating on a fortnight of cash, regardless of how profitable either is on paper. It also belongs in any comparison of providers, alongside the fee layers a blended card rate hides, because a reserve is a cost that never appears as a fee. If a reserve is imposed, the first call is to your suppliers about payment terms, not to the processor about fairness.

What should you do in the first hour of a hold?

Establish which intervention you are under, then supply evidence rather than argument. Ask three questions in writing: are payouts paused or is processing paused, is a reserve being applied and at what percentage, and what specific documents will close the review.

Then send what underwriting actually needs, which is proof that you deliver. Fulfilment records with tracking numbers for recent orders. Your refund and delivery policies as published. Supplier invoices showing you can source what you are selling. Evidence that the volume change has a cause, such as a campaign or a press mention or a wholesale order.

What does not help is a long message about how unfair this is, or a threat to leave. The reviewer is comparing your documents against a risk position, and the fastest route through is a complete pack sent once. Where disputes are the trigger, the evidence standards are the same ones we set out in assembling chargeback evidence that actually wins.

How do you lower the ratio rather than the symptom?

By attacking the disputes you cause rather than the fraud you suffer, because the first category is larger than most sellers believe. Unrecognised billing descriptors, slow delivery with no updates, refund requests that went unanswered until the customer gave up and called the bank. None of that is fraud and all of it lands in the same numerator.

Three concrete moves, in order of return. Make your billing descriptor recognisable, since a customer who cannot identify a charge disputes it. Answer refund requests within a day, because a dispute is what happens when a refund is too slow. Send delivery updates unprompted, as a chargeback for goods not received is often a shipment that arrived late without news. We covered the support side of that loop in which parts of returns are safe to automate.

One structural point. Your dispute ratio depends on infrastructure you often do not control, including how your checkout names itself on a statement and how quickly order status reaches a customer. That is a reason to care about owning the storefront rather than renting one, which is the argument behind building a store whose checkout and code stay yours.

The uncomfortable summary

You are a counterparty to your processor, not only a customer of it. The models in the middle are optimising the processor's exposure, which is a legitimate aim that does not coincide with yours. Nobody is going to explain the account level score to you, and the appeal channel for it barely exists.

What you do get is the numbers. The ratios are published, the thresholds are published, and at least one major provider will project your position before month end. A seller who checks four figures monthly is not at the mercy of any of this. A seller who learns the thresholds from a held payout is, and the recovery clock runs in months.

Discussion 0

0 / 4000Your email address is not displayed with your comment.
No comments are published yet.

Explore — related articles.

Build something. Move your work forward.

Start with a software project or an agent task. Describe the result you need, review the work and keep control of your connected accounts.

Open the workspace →