03Fraud Model to Production — Databricks MLOps, Months to Weeks
Fraud model moved from notebook to live scoring — releases cut from months to weeks
Role: MLOps Solution Engineer
Executive summary
Built an end-to-end MLOps pipeline on Azure Databricks to productionize a PyTorch fraud-detection model with Feature Store, real-time serving and drift monitoring—cutting deployment time from months to weeks.
- Azure Databricks
- PyTorch
- MLflow
- Feature Store
- Model Serving
- Spark Structured Streaming
- Azure DevOps
A payments provider needed to deploy a PyTorch fraud-detection model in production on Azure Databricks, integrated with a Feature Store, real-time inference and continuous monitoring. Their existing ML process was ad-hoc and lacked an MLOps framework, making deployment and monitoring difficult. The false-positive stakes are well documented in industry research (Javelin Strategy & Research estimated banking false declines at $118B vs $9B of actual card fraud, with only 1 in 5 fraud alerts turning out correct; Wedge et al., 2018 document false-positive rates as high as 10–15% in deployed fraud pipelines, cutting to ~3% with stronger engineering in their published bank case study), so the pipeline was required to make precision/FPR a first-class, monitored outcome.
- Refactor the PyTorch model to run efficiently as a distributed Spark job.
- Set up Databricks Feature Store for consistent training/inference features.
- Automate training and evaluation with Databricks Jobs and MLflow, tracking model versions.
- Deploy via Databricks Model Serving for secure real-time scoring.
- Implement monitoring and drift detection (e.g., population stability index on features and scores).
- Track model precision / false-positive rate as first-class monitored metrics, not afterthoughts.
Working with the client's data scientists, I productionized notebook code with MLflow for versioned tracking, populated Feature Store tables (average transaction size, device frequency) updated daily, and configured Databricks Repos with Azure DevOps CI/CD for automated testing. I containerized the model behind a real-time endpoint and built a Spark Structured Streaming job that logs predictions and alerts on significant drift. I considered leaving monitoring to the ops team's generic APM tooling and rejected it—fraud drift shows up in feature distributions first, so I kept monitoring inside the Databricks pipeline where PSI could run on the same tables the model reads. I aligned ops staff on incident response and documented the pipeline and runbooks.
The client moved from experiment to a production-grade fraud service with sub-second latency that scales on autoscaling clusters, processing live transaction scoring behind a governed endpoint. Every model version is tracked in MLflow with audit trails, making each release reproducible and reviewable. Within the first month, drift metrics flagged a behavior shift, prompting proactive retraining—the alert-to-retrain loop the project was built for, working before the first incident proved it necessary. Release cycles dropped from months to weeks. Model precision at the chosen operating threshold was monitored in the same pipeline that logged predictions, keeping the false-positive ratio visible to review rather than buried in a model card.
The drift alert was the pivot moment — it was ambiguous at first between a genuine behavior shift and a data-quality fault, and telling those apart is now the first check I build into any drift monitor. MLOps discipline (CI/CD, Feature Store, model registry) converted a one-off deployment into a repeatable release process. A feedback loop with drift detection and human review is vital for fraud, where tactics evolve. Treating PyTorch and Spark as one unified platform achieved scale without rewrites.