Peter Parker

Project 09: Real-Time Fraud and Risk Decisioning Platform

Running live for sixty days from 28 September 2026
Open the live dashboard Read the code The dashboard shows the platform right now. This page is what it means.

Live, measured, and reproducible from the repository

Every card payment gets a decision before the money moves. Does it stay right when fraud changes?

When you tap a card, the bank has a few dozen milliseconds to decide: approve it, send it to a person to check, or decline it. Getting that decision fast is the easy half. The hard half is that fraud keeps changing shape, and a model that was right last month quietly stops being right, usually because what it learned from no longer looks like what it sees. This platform is the machinery around the model that keeps it honest: it decides every transaction in milliseconds, notices when the traffic shifts, retrains, and will not promote a new model until a person has seen the evidence.

- to decide 99 transactions in every 100 at 1,000 a second on the live machine, against a 50 ms budget
- transactions a second, and it kept up four times the live rate, one scorer, every run
Same day every fraud shift flagged on the day it began and nothing flagged in the clean week before the first one
- more fraud caught per analyst-hour by ranking the review queue on expected loss instead of score

The live stream is synthetic: a generator plays card transactions from an entity graph of 200,000 cards, 150,000 devices and 4,000 merchants, with three kinds of fraud acting on it. No real payment data is involved. A public fraud competition's data runs through the same pipeline offline, and every number on this page says which of the two it came from and carries a 95 percent confidence interval, because a number without one is an opinion.

What happens in those few milliseconds

Each transaction takes the same five steps. The times are the 99th percentile of each step, measured on the live machine with the live configuration: the champion model deciding, the challenger scoring beside it in shadow, and every decision written to history. Pick a load to see where the time goes as it rises.

    The model is the smallest part.

    The sixteen features the model sees, in plain words

    A feature is a fact about the recent past: how busy this card has been, how many different cards this device has carried. Every one is computed once, by one piece of code, and written to the store training reads and the store serving reads in the same call, so the model is never trained on a number it will not be given in production. A test recomputes them from the raw log and fails if any one of them could have seen the future.

    How fast, and how much it can take

    Five runs at each load, on the live machine (4 virtual CPUs). The lines are the typical decision (p50), the slowest one in twenty (p95) and the slowest one in a hundred (p99), from the moment a transaction is sent to the moment its decision is on the stream. The whiskers are the 95 percent interval across runs. Point at a load to read it off.

    A load test is the cleanest measurement, not the whole story. It times every decision exactly, over twenty seconds. Running for days, the platform also shares the machine with its own housekeeping, and the dashboard estimates its percentiles from buckets rather than timing each decision, so its p99 reads higher. The next section has the three-day number, measured the way the live window will be reported.

    It keeps running when the machine is taken away

    The platform runs on a spot instance: spare cloud capacity at a fraction of the price, which AWS can take back with two minutes' notice. That is a cost decision with a real consequence, so it was rehearsed. For 72 hours before go-live the whole stack ran at the live rate, and AWS took the machine away - times. Each mark below is one of them; its length is how long until the platform was deciding on time again.

    Nothing is lost when the machine goes, and nothing starts from scratch. Transactions wait in the stream while a new machine boots, and the scorer saves its feature state to a disk that survives the machine, so a replacement restores what every card, device and merchant has been doing instead of treating them all as strangers for a day. The minutes of catch-up are counted in uptime and reported beside the latency, never hidden inside it.

    Fraud changes. Does anyone notice?

    Fifty days of generated stream, 86 million transactions, with the fraud changing three times on days the monitors were never told about. Each bar is one day: how many of the monitored quantities had moved beyond a fixed threshold, compared with the week the model was trained on. The monitors watch the model's own score as well as its inputs.

    Every change is flagged the day it starts, and the quiet week flags nothing. The second shift is the one that matters most: amounts move and the amount of fraud does not, so an alarm on the fraud rate would see nothing at all while the model is shown a different world. The thresholds are published conventions, fixed before any drift was seen, and were not revisited after these results.

    How a day is judged, and why two days are needed

    What retraining can and cannot do about it

    The alarm opens a retraining request. The job fits a new candidate, compares it with the current champion on transactions neither has seen, writes the evidence into a pull request, and stops. The score is PR-AUC, how well a model puts fraud above everything else; 1.0 is perfect, and 0.03 is what guessing would score on this stream.

    Nothing promotes itself

    1. 1Drift on two days running opens a request, and an email says so.
    2. 2A candidate is fitted once enough of the new days' labels have arrived, and compared on rows neither model saw.
    3. 3A pull request carries the model file and its evidence. Merging it is the approval to run it in shadow.
    4. 4A week in shadow: it scores live traffic beside the champion and decides nothing.
    5. 5The promotion gate needs its whole confidence interval to clear a non-inferiority margin, then opens a second pull request.
    6. 6A person flips the flag. Rolling back is the same flag: - from the flip to the old champion deciding again.

    During the live window the pull requests are opened by the platform itself, on the repository's pull requests page, where anyone can read the evidence the day it is written.

    Who should the analysts look at first?

    The bug it caught on the first day

    The most expensive mistake in fraud modelling is a model trained on information it will not have when it matters, which scores beautifully in testing and badly in production. So the test for it was written a week before the first feature existed. The day the features were written, it failed.

    Two payments on one card can carry the same timestamp. The engine counted the first as history when it scored the second, which a model in production could never know at that instant. On the public competition data it served a wrong value on 65 of 590,540 rows, about one in nine thousand: too rare for any accuracy figure to show, and permanent once trained on.

    How to read the live dashboard

    The dashboard is the platform's own monitoring, open to anyone and read-only. It is built for the person running the system, so here is what each part is saying and what normal looks like.

    Decisions a second

    Between about 750 and 1,250 through the day. The traffic follows a daily cycle, quietest around 08:00 UTC and busiest around 20:00 UTC, with a little noise each hour, because real card traffic is not a flat line.

    Event to decision, 99th percentile

    The headline latency. It reads in the low twenties of milliseconds, where the load test reads 8: it is estimated from histogram buckets, and the one it lands in runs from 10 to 25 ms, so it cannot say where inside that range the true figure lies. Either way it is well inside the 50 ms budget.

    A spike to minutes, then a burst

    That is AWS taking the machine back. Decisions stop, "since the last decision" climbs, a new machine boots and restores, and the backlog is worked off faster than it arrives. Every transaction is still decided; the late ones are counted as late.

    While caught up

    The same latency with every catch-up left out, so a reclaim does not hide how the platform performs when it is current. The live report leaves out only the catch-ups AWS announced, and says how many.

    Scorer hops

    The scorer's own work, split into features, model, rules and writing the decision. Each is well under a millisecond; a hop that climbs is a regression in the code, not the network.

    Model deciding

    Which model version made the decisions. It changes only when a person promotes a model or rolls one back, and that change is a merged pull request you can read.

    Sealed before it started

    A drift monitor that has been tuned to the shifts it will face proves nothing. So the live window's fraud schedule, when each shift comes and what it does, is derived from a secret that was sealed before the window opened. Its fingerprints were committed to the repository; the secret is kept where the platform's builder cannot read it, and is published the day after the window closes, so anyone can check the schedule was never changed.

    What this page does not show