Research note · 6 Aug 2026 · synthetic PaySim data only · English

From raw events to a deployed decision system

This is a public, synthetic demonstration of the delivery method: define a pre-decision data contract, split evidence chronologically, compare models at a fixed review framing, freeze an artifact, deploy it, and expose versioned live inference. The reported performance belongs only to the simulator; the engineering and evaluation pattern is what MIHZI would rebuild on approved customer data.

Synthetic data only. Nothing here is your production traffic, and nothing here is a promise about accuracy on a real wallet, bank, insurer, clinical, or logistics system. When we work with you, we rebuild and recheck on your data.

What this proves: chronological evaluation; fixed-workload thinking; simple features before complex history; a frozen deployable artifact; live versioned inference with observable identity.

What this does not prove: production performance for any operator; coverage of SIM-swap, device, agent, or social-engineering signals; calibrated real-world probability; clinical or industrial validation; that thresholds transfer to another organization.

Traffic changes over time

Synthetic PaySim chart: transaction volume and fraud rate over time with a train and test cut
Synthetic PaySim: volume (black), fraud rate (rust), train/test cut at step 594

In this simulator, volume falls and the fraud rate gets noisier later on. So we train on earlier hours and test on later hours. That is closer to real life than shuffling the whole file.

On your data, the same idea applies: train on the past, check on what came next. If future weeks look different, a shuffled test can make a model look better than it is.

Simple fields already carry signal

Synthetic charts comparing fraud and genuine transaction types, amounts, and balances
Synthetic training window: type mix, amount shape, origin balance shape

Before any fancy history features, a few basic fields already separate many fraud rows from genuine ones. In this simulator, fraud shows up in cash-out and transfer. Amounts and balances look different too.

Synthetic boxplots of amount for fraud versus genuine
Synthetic amounts: fraud tends higher than genuine on this data

That does not mean “large equals fraud.” It means your first model can often start with fields you already store, then grow as your labels and history improve.

Some tricks look great and still mislead

Synthetic model comparison with smooth score histograms
Synthetic bakeoff: normal setups keep a spread of scores
Synthetic model comparison where one setup collapses scores into spikes
Synthetic bakeoff: one setup collapses scores into a few spikes

One transform made the model look much stronger because many fraud rows in PaySim drain the whole balance. The model basically learned a yes/no flag from the simulator.

On your data, “moved most of the balance” can still matter. It is rarely that clean. We keep the useful idea and reject the easy win that only works because the fake data is too neat.

From model choice to model explanation

We compared two boosted-tree base learners on the same three pre-transaction features and a later chronological hold-out. CatBoost produced higher average precision and recall at the illustrative 0.5 threshold, while XGBoost produced slightly higher precision. We then used CatBoost's native SHAP calculation to inspect which inputs moved its score. This is synthetic experimental evidence, not a claim about production performance or a customer's data.

Synthetic comparison of CatBoost and XGBoost on the same three pre-transaction features
Same inputs, different categorical handling. CatBoost uses transaction type natively; XGBoost uses one-hot encoding. On PaySim steps 595–743, CatBoost achieved 0.956 average precision versus 0.944 for XGBoost. At a 0.5 threshold, CatBoost recovered more synthetic fraud with a small precision trade-off. Thresholds for a real service would be chosen against review capacity, error cost, and policy—not copied from this experiment.
Native CatBoost SHAP summary for amount, origin balance, and transaction type
What moved the CatBoost score. Native SHAP values on a stratified 10,000-row chronological-test slice show amount and origin balance with the largest global attribution magnitudes (mean |SHAP| about 4.69 and 4.41; type about 1.37). Values are log-odds contributions inside this fitted synthetic model; they are not percentages, direct probability changes, causal effects, or proof of a transferable real-world fraud mechanism.

Live demo reason_codes are separate heuristic explanations, not these SHAP values. See the model card for the deployed paysim-catboost-demo-v1 digest and limits.

See the method, then test one decision with your data

The demo workbench shows a simulated review queue beside a live synthetic scorer. You can confirm fraud or not fraud the way an analyst would; that feedback does not retrain the model.

Tell us one repeated decision. We will respond with the smallest useful data checklist, the evaluation boundary, and the measure we would use. No production data is needed for the first conversation. Payments and fraud are one example. The same pattern fits operations, forecasting, claims, logistics exceptions, governed clinical research partners, and data foundations.