Resume
All systems

PLATFORM

ML Platform and MLOps

The Azure Databricks platform, ingestion layer and MLOps standard that PromotionsAI, ClearanceAI and the rest of the team's models run on.

  • Azure Databricks
  • Delta Lake
  • MLflow
  • Terraform
  • PySpark
  • Docker
  • Kubernetes
  • Kafka
  • Azure Event Hubs
  • Azure Pipelines
  • Power BI

Context

Two pricing systems and the rest of the team needed the same things: clean ingested data, features that are computed the same way in training and serving, a registry that gates promotion, and a release path that can be rolled back.

Built once as a platform, those are leverage. Built per project, they are four different ways to break production.

Constraints

Infrastructure is provisioned in Terraform, so environments have to be reproducible rather than clicked together.

The platform carries live pricing traffic, which puts a 99.9 percent service SLA on every change to it.

Merchandising, finance and engineering all need to see the same numbers, so KPI impact, data lineage and model health had to be surfaced rather than explained on request.

What I built

Automated feature pipelines and model registry promotion, Dockerized inference, unified MLflow lineage, blue-green releases through Azure Pipelines, and an automated A/B testing harness.

PySpark and Delta Lake ingestion with validation and enrichment, plus reusable connector and transformation libraries.

Cookiecutter project templates, Databricks job scaffolds and distributed training playbooks for Spark clusters.

Power BI and Azure-native dashboards surfacing KPI impact, data lineage and model health to merchandising, finance and engineering.

The runbooks and reproducible notebooks that carry new engineers through onboarding and hold up under audit review.

Architecture

Source systems land through PySpark ingestion with validation and enrichment, against reusable connector libraries rather than a new script per feed.

Delta Lake is the versioned single source of truth, so training data and serving features are drawn from the same tables.

Feature pipelines compute and register automatically, distributed training runs on Spark clusters from templated jobs, and MLflow holds lineage and gates promotion.

Inference ships as Docker images through blue-green releases on Azure Pipelines, with monitoring and drift detection feeding back into retraining.

Results

Over 200,000 predictions served per month.

Research-to-production lead time cut 40 percent, experiment turnaround cut 60 percent.

A 99.9 percent service SLA held, with incident MTTR halved.

New source onboarding 70 percent faster on the reusable connector libraries.

What I would do differently

Ship the dashboards that show model health and data lineage at the same time as the first pipeline, not after it. The platform work is easier to fund and easier to defend once the people outside the team can see the same numbers you can.