Decisions a second
Between about 750 and 1,250 through the day. The traffic follows a daily cycle, quietest around 08:00 UTC and busiest around 20:00 UTC, with a little noise each hour, because real card traffic is not a flat line.
Live, measured, and reproducible from the repository
When you tap a card, the bank has a few dozen milliseconds to decide: approve it, send it to a person to check, or decline it. Getting that decision fast is the easy half. The hard half is that fraud keeps changing shape, and a model that was right last month quietly stops being right, usually because what it learned from no longer looks like what it sees. This platform is the machinery around the model that keeps it honest: it decides every transaction in milliseconds, notices when the traffic shifts, retrains, and will not promote a new model until a person has seen the evidence.
The live stream is synthetic: a generator plays card transactions from an entity graph of 200,000 cards, 150,000 devices and 4,000 merchants, with three kinds of fraud acting on it. No real payment data is involved. A public fraud competition's data runs through the same pipeline offline, and every number on this page says which of the two it came from and carries a 95 percent confidence interval, because a number without one is an opinion.
Each transaction takes the same five steps. The times are the 99th percentile of each step, measured on the live machine with the live configuration: the champion model deciding, the challenger scoring beside it in shadow, and every decision written to history. Pick a load to see where the time goes as it rises.
The model is the smallest part.
A feature is a fact about the recent past: how busy this card has been, how many different cards this device has carried. Every one is computed once, by one piece of code, and written to the store training reads and the store serving reads in the same call, so the model is never trained on a number it will not be given in production. A test recomputes them from the raw log and fails if any one of them could have seen the future.
Five runs at each load, on the live machine (4 virtual CPUs). The lines are the typical decision (p50), the slowest one in twenty (p95) and the slowest one in a hundred (p99), from the moment a transaction is sent to the moment its decision is on the stream. The whiskers are the 95 percent interval across runs. Point at a load to read it off.
A load test is the cleanest measurement, not the whole story. It times every decision exactly, over twenty seconds. Running for days, the platform also shares the machine with its own housekeeping, and the dashboard estimates its percentiles from buckets rather than timing each decision, so its p99 reads higher. The next section has the three-day number, measured the way the live window will be reported.
The platform runs on a spot instance: spare cloud capacity at a fraction of the price, which AWS can take back with two minutes' notice. That is a cost decision with a real consequence, so it was rehearsed. For 72 hours before go-live the whole stack ran at the live rate, and AWS took the machine away - times. Each mark below is one of them; its length is how long until the platform was deciding on time again.
Nothing is lost when the machine goes, and nothing starts from scratch. Transactions wait in the stream while a new machine boots, and the scorer saves its feature state to a disk that survives the machine, so a replacement restores what every card, device and merchant has been doing instead of treating them all as strangers for a day. The minutes of catch-up are counted in uptime and reported beside the latency, never hidden inside it.
Fifty days of generated stream, 86 million transactions, with the fraud changing three times on days the monitors were never told about. Each bar is one day: how many of the monitored quantities had moved beyond a fixed threshold, compared with the week the model was trained on. The monitors watch the model's own score as well as its inputs.
Every change is flagged the day it starts, and the quiet week flags nothing. The second shift is the one that matters most: amounts move and the amount of fraud does not, so an alarm on the fraud rate would see nothing at all while the model is shown a different world. The thresholds are published conventions, fixed before any drift was seen, and were not revisited after these results.
The alarm opens a retraining request. The job fits a new candidate, compares it with the current champion on transactions neither has seen, writes the evidence into a pull request, and stops. The score is PR-AUC, how well a model puts fraud above everything else; 1.0 is perfect, and 0.03 is what guessing would score on this stream.
During the live window the pull requests are opened by the platform itself, on the repository's pull requests page, where anyone can read the evidence the day it is written.
The most expensive mistake in fraud modelling is a model trained on information it will not have when it matters, which scores beautifully in testing and badly in production. So the test for it was written a week before the first feature existed. The day the features were written, it failed.
Two payments on one card can carry the same timestamp. The engine counted the first as history when it scored the second, which a model in production could never know at that instant. On the public competition data it served a wrong value on 65 of 590,540 rows, about one in nine thousand: too rare for any accuracy figure to show, and permanent once trained on.
The dashboard is the platform's own monitoring, open to anyone and read-only. It is built for the person running the system, so here is what each part is saying and what normal looks like.
Between about 750 and 1,250 through the day. The traffic follows a daily cycle, quietest around 08:00 UTC and busiest around 20:00 UTC, with a little noise each hour, because real card traffic is not a flat line.
The headline latency. It reads in the low twenties of milliseconds, where the load test reads 8: it is estimated from histogram buckets, and the one it lands in runs from 10 to 25 ms, so it cannot say where inside that range the true figure lies. Either way it is well inside the 50 ms budget.
That is AWS taking the machine back. Decisions stop, "since the last decision" climbs, a new machine boots and restores, and the backlog is worked off faster than it arrives. Every transaction is still decided; the late ones are counted as late.
The same latency with every catch-up left out, so a reclaim does not hide how the platform performs when it is current. The live report leaves out only the catch-ups AWS announced, and says how many.
The scorer's own work, split into features, model, rules and writing the decision. Each is well under a millisecond; a hop that climbs is a regression in the code, not the network.
Which model version made the decisions. It changes only when a person promotes a model or rolls one back, and that change is a merged pull request you can read.
A drift monitor that has been tuned to the shifts it will face proves nothing. So the live window's fraud schedule, when each shift comes and what it does, is derived from a secret that was sealed before the window opened. Its fingerprints were committed to the repository; the secret is kept where the platform's builder cannot read it, and is published the day after the window closes, so anyone can check the schedule was never changed.