Full-StackLive demo · 2026

Fraud Radar — Real-time card-fraud detection

A real-time card-fraud detection platform built end to end. A FastAPI service scores each transaction with six deterministic rules and an XGBoost model in single-digit milliseconds, attaches a SHAP explanation to every decision, and routes it to an analyst review queue with an append-only audit trail; a React 19 + TypeScript dashboard sits on top. An offline evaluation programme then tested whether the model generalises beyond the synthetic data it was trained on. It does not transfer across data generators without retraining, and the project reports that in full.

Measured, not claimed

0.9327 PR-AUC

In-distribution

Synthetic held-out fold · 93 frauds in 7,502

0.8653 PR-AUC

Retrained on an independent generator

Sparkov fold · 924 frauds in 277,860

0.0087 PR-AUC

Carried across generators, no retraining

Same Sparkov fold · 0.0033 prevalence

0.7670 PR-AUC

Real anonymised card data, isolated track

ULB seed 42 · 52 frauds in 42,722

3.7 ms

Service-layer scoring, p50

n = 500, developer laptop

1,407 tests

Green in CI

ruff · mypy --strict · tsc · eslint

The Fraud Radar transactions view: scored transactions with decisions, amounts converted to a reporting currency, and filters.

Technology stack

Python 3.11FastAPIXGBoostSHAPPydantic v2SQLAlchemy 2.0Alembicscikit-learnReact 19TypeScriptTanStack QueryTailwind CSSVite
01

The problem

Fraud decisions sit on the payment authorisation path. They have to be made in milliseconds, explained to the customer and eventually a regulator, survive network retries without double-charging or double-blocking, keep a human in the loop, and leave an audit trail that still makes sense years later.

Underneath all of that is a model trained on data that resembles production only as closely as whoever built the dataset managed. Most of this project’s effort goes into measuring that gap.

02

What I built

  • A FastAPI service layered api → services → repositories → models, with Pydantic v2 wire contracts, SQLAlchemy 2.0 and five Alembic migrations whose CHECK constraints encode real invariants.
  • Stripe-pattern idempotent ingestion: the same key and body replays the cached response; the same key with a different body returns 409. The key is a SHA-256 of the normalised payload, not the raw bytes.
  • A scoring pipeline of six pure rules (a hard block short-circuits the model) → a 17-feature extractor over a 180-day history window → XGBoost → SHAP TreeExplainer → a conservative-wins decision → an append-only audit log → an analyst review queue.
  • An analyst loop that records human verdicts without overwriting the model’s original decision, so the model’s call stays clean for evaluation and retraining.
  • Keyset-paginated transaction and alert APIs, Decimal money end to end, and multi-currency FX enrichment that degrades gracefully when the rate provider is unavailable.
  • A React 19 + TypeScript dashboard (TanStack Query) with live KPIs, a filterable transaction feed, a per-transaction page showing the stored SHAP attribution, and an alerts worklist.
03

Train/serve parity by construction

The offline batch feature builder calls the same FeatureExtractor the API uses, behind a session that raises if anything tries to reach a database. A golden test compares batch and serving output element by element with exact float equality, on a fixture where no feature is constant.

04

How it was evaluated

Four evaluations, never merged or averaged. Each splits its corpus chronologically 70/15/15: hyperparameters are searched on the training fold, the operating threshold is chosen on validation, and the test fold is scored once. Both benchmark methodologies were frozen in decision records before any test fold was scored, and the published benchmark cards are generated from committed run records.

  • In-house synthetic: 50,010 transactions, 500 customers, 200 merchants, 1.52% fraud, six injected fraud patterns.
  • Sparkov: an independent transaction generator, 1,852,394 rows. Simulated, not real card data.
  • Cross-generator transfer: the synthetic-trained model, unchanged, scored on Sparkov’s test fold.
  • ULB: real, publisher-anonymised card data (284,807 rows, 492 frauds) on an isolated featureset that can be neither promoted nor served.
05

Results, each read against its prevalence

EvaluationPopulationPR-AUCROC-AUCRecall @ 1% FPR
In-distribution (synthetic)93 frauds in 7,502 · prevalence 0.01240.93270.99890.9785
Retrained on Sparkov924 frauds in 277,860 · prevalence 0.00330.86530.99630.9459
Transferred, no retrainingSame Sparkov fold · prevalence 0.00330.00870.73540.0390
ULB real data, isolated track52 frauds in 42,722 · prevalence 0.0012, seed 420.76700.97510.8269
Four evaluations, never merged or averaged. A scorer with no signal scores about the prevalence, so read each PR-AUC against the column beside it.
  • At the threshold chosen on the synthetic validation fold, in-distribution precision is 0.61 and recall 0.96. Applied unchanged to Sparkov, that same threshold flags 19,215 legitimate transactions to catch 58 frauds.
  • The ULB repeats (seeds 43 and 44) score 0.7569 and 0.7751, which shows how much the figure moves with the fit’s randomness on a fold holding 52 frauds.
  • Temporal drift: with the threshold selected once on late-2019 data, monthly recall across 2020 stays between 0.906 and 0.970 while precision falls from about 0.40 to 0.16 as prevalence drops.
  • Rules audit: across all 1.85M Sparkov rows the production rules would approve 9,385 of 9,651 frauds; two rules never fire and one cannot be evaluated on that data.
  • Latency: service-layer scoring p50 3.7 ms / p95 5.8 ms; full HTTP round trip p50 16 ms / p95 20 ms (n = 500, developer laptop).
06

What the evaluation found

In-distribution performance said almost nothing about generalisation. The same 17-feature pipeline scores 0.9327 on its own generator, 0.8653 when retrained on an independent one, and 0.0087 when carried across without retraining.

  • Five of the 17 features are constant on Sparkov, so the model leans on signal the target data does not carry.
  • The operating threshold does not transfer: chosen for a 1% false-positive rate on its own validation fold, it produces 6.94% on Sparkov.
  • The synthetic generator contains structural shortcuts. Six countries enter the training data only through a fraud pattern, so the model never saw a legitimate transaction from them. A closed-loop evaluation hides exactly this kind of shortcut.
07

Limitations

  • The served model is trained on synthetic data. Its 0.9327 measures how learnable the generator is, not how detectable real fraud is.
  • Sparkov is simulated; ULB is real but isolated on its own featureset and not comparable with the other tracks.
  • Scores are not calibrated probabilities. Calibration is measured, never fitted.
  • Label delay (chargebacks arriving weeks later) is not modelled, so every result is optimistic in a way the benchmark does not measure.
  • The public demo reads a static snapshot of a few hundred synthetic transactions; it is not production traffic, and there is no authentication or role-based access yet.
08

Status

Phases 1–5 are complete: the stack runs end to end, the demo is deployed, the benchmarks are executed and written up, and FX enrichment is in the live path. Phase 6 (observability, model governance, production drift monitoring and authentication) is planned; none of it exists yet.

Sources

Checked against the repositories on 22 September 2026.