Resume

OPEN TO AI/ML ENGINEER ROLESF-1 OPTNO SPONSORSHIP REQUIRED

AjayMekala

AI/ML ENGINEERPRODUCTION ML PLATFORMS, MLOPS, FRONTIER MODEL EVALUATION

I build and operate the ML platform behind Walmart's promotion and clearance pricing. Four years shipping models into production, and evaluating the frontier models that are coming for the job.

$7.8MINCREMENTAL REVENUE FROM ML
200K+PREDICTIONS SERVED MONTHLY
99.9%PLATFORM SERVICE SLA
4 YRSPRODUCTION ML AND EVALUATION

WHERE THE WORK HAS SHIPPED

  • WALMART
  • HANDSHAKE AI
  • SNORKEL AI
  • MERCOR
  • OUTLIER AI
  • ALIGNERR
  • WELOCALIZE
  • AFTERQUERY
  • FLEETAI
  • TECH MAHINDRA
  • MONTCLAIR STATE UNIVERSITY

01 / PROFILE

I am an AI/ML Engineer with four years across production machine learning, data platform engineering and frontier model evaluation. I build systems that set prices for one of the largest retailers in the world, and I spend the rest of my time finding out where the newest language models break.

At Walmart I built and operate the ML platform behind promotion and clearance pricing, running on Azure Databricks, Delta Lake and MLflow with infrastructure provisioned in Terraform. It serves over 200,000 predictions per month. PromotionsAI and ClearanceAI sit on top of it and have driven 7.8 million dollars in incremental revenue. Standardizing MLOps across those pipelines cut research-to-production lead time by 40 percent and experiment turnaround by 60 percent.

Alongside that, I evaluate and stress-test frontier models on contract for Handshake AI, Snorkel AI, Mercor and Outlier. I have authored more than 200 golden-solution engineering tasks and over 50 accepted Terminal-Bench environments. I am completing an MS in Data Science at Montclair State University in May 2026.

SPECIFICATION

ROLE
AI/ML Engineer, Walmart
BASED
Teaneck, New Jersey
WORK AUTH
US work authorized, no sponsorship required
EDUCATION
MS Data Science, Montclair State (May 2026)BTech AI and Data Science, KL University (2024)
FOCUS
Production ML platforms, MLOps, model evaluation
STACK
Python, PySpark, Databricks, MLflow, Terraform, Kubernetes
AVAILABLE
Immediately, remote or hybrid in the NYC metro

02 / SYSTEMS IN PRODUCTION

The platform, and the two pricing systems on it

Promotion and clearance pricing at Walmart run on one Azure Databricks platform. This is the path a prediction takes through it, stage by stage.
Walmart pricing ML platform, from source systems to the merchandising APIMONITORING ANDDRIFT DETECTION01SOURCE SYSTEMS02PYSPARK INGESTAND VALIDATION03DELTA LAKE04FEATURE PIPELINES05MODEL TRAINING06MLFLOW REGISTRY07DOCKERIZED INFERENCEAND MONITORINGMERCHANDISING API

STAGE 01 / 07

  1. 01 SOURCE SYSTEMS

    Merchandising, inventory and transaction feeds arrive from several upstream systems. The call was to treat this boundary as the place bad data stops rather than as glue between systems, so validation lives here and not three layers down.

  2. 02 PYSPARK INGEST AND VALIDATION

    I built reusable connector and transformation libraries instead of a script per feed, which turned a new source into configuration and cut onboarding one by 70 percent.

  3. 03 DELTA LAKE

    Training data and serving features read the same versioned tables. Choosing ACID storage over a file drop is what makes a model reproducible months later, when someone asks what it actually saw.

  4. 04 FEATURE PIPELINES

    Feature computation and registration are automated rather than hand run, so a feature has one implementation instead of one for training and another for serving. That decision is where 40 percent of the research-to-production lead time went.

  5. 05 MODEL TRAINING

    Distributed training on Spark clusters, templated. I chose job scaffolds over documentation, which took standing up a new experiment from days to under an hour.

  6. 06 MLFLOW REGISTRY

    Promotion is gated rather than conventional. Nothing serves traffic without a registry version carrying its parameters, metrics and lineage, which is what makes "which model priced this SKU" a lookup instead of an investigation.

  7. 07 DOCKERIZED INFERENCE AND MONITORING

    Blue-green releases through Azure Pipelines, built before anyone asked for them, so a bad model is a rollback and not an incident. 200,000+ predictions per month at a 99.9 percent SLA, with drift detection wired back into retraining rather than into a dashboard.

03 / FRONTIER MODEL EVALUATION

Where the newest models still break

Three comparisons from the contract work, in the shape the work actually takes. Pick the stronger response and the annotation opens.

AGENT TRAJECTORY GRADING

PROMPT

The task container has a Python virtualenv at /opt/venv. The runner executes each of your commands in a fresh bash -lc invocation. Give the single command that runs the project test suite on the venv interpreter and stops at the first failure, without activating the venv.

WHICH RESPONSE IS STRONGER

RUBRIC

  • Helpfulnessscored after you choose
  • Honestyscored after you choose
  • Completenessscored after you choose
  • Harmlessnessscored after you choose
200+GOLDEN SOLUTION TASKS
50+TERMINAL-BENCH ENVIRONMENTS ACCEPTED
3,000+PREFERENCE COMPARISONS
98%+SENIOR REVIEWER AGREEMENT

CLIENTS

  • Handshake AI

    Project Helix golden solutions, agent trajectory evaluation

  • Snorkel AI

    Terminal-Bench environments under the Terminus 2 scaffold

  • Mercor

    Enterprise workflow environments, personalized evaluation

  • Outlier AI

    Expert coding prompts with production-grade test cases

  • Alignerr

    Alignment review

  • Welocalize

    AI data evaluation and task creation

  • AfterQuery

    AI data evaluation and task creation

  • FleetAI

    AI data evaluation and task creation

  • SME Careers

    Subject-matter expert review

04 / EXPERIENCE

Four years, newest first

  1. AI/ML Engineer

    Walmart, Remote

    November 2024 to Present

    • Built and operate the ML platform behind promotion and clearance pricing on Azure Databricks, Delta Lake and MLflow with Terraform-provisioned infrastructure, cutting research-to-production lead time 40 percent and serving over 200,000 predictions per month.
    • Shipped PromotionsAI, driving 5.5 million dollars in incremental revenue in its first full fiscal year and 2.3 million the following quarter.
    • Built and own the MLOps stack for ClearanceAI, lifting sell-through 4.6 percent.
    • Standardized MLOps across both systems and the ingestion pipelines, cutting experiment turnaround 60 percent, holding a 99.9 percent SLA and halving incident MTTR.
    • Automated ingestion, validation and enrichment with PySpark and Delta Lake and co-authored the reusable connector libraries, cutting new source onboarding 70 percent.
    • Azure Databricks
    • MLflow
    • Delta Lake
    • PySpark
    • Terraform
    • Kubernetes
    • Docker
    • Kafka
    • XGBoost
    • Python
  2. AI Training and Evaluation Specialist

    Contract, Remote

    November 2022 to Present

    Clients: Handshake AI, Snorkel AI, Mercor, Outlier AI, Alignerr, Welocalize, AfterQuery, FleetAI, SME Careers

    • Authored 200+ golden-solution software engineering tasks for Handshake AI's Project Helix, building multi-file repository problems with hidden test suites and graded difficulty tiers across Python, TypeScript and Go, and passed Helix Screening with a top-decile reviewer rating.
    • Evaluated long-horizon coding-agent trajectories on real software workflows, scoring tool-use correctness, sub-goal decomposition, state tracking and recovery from failure, and surfaced systematic failures in tool-call sequencing and shell command grounding that fed directly into agent fine-tuning.
    • Authored 50+ accepted Terminal-Bench tasks for Snorkel AI under the Terminus 2 agent scaffold: containerized environments with reference solutions and automated verification tests, calibrated to challenge frontier agents while staying unambiguous and solvable.
    • Designed 150+ model-stumping prompts validated against GPT, Claude and Gemini-class baselines, plus graduate-level STEM question-answer pairs for reasoning benchmarks that passed multi-layer expert review.
    • RLHF
    • DPO
    • SFT
    • Agentic Evaluation
    • Terminal-Bench
    • Rubric Design
    • Red-Teaming
    • Python
    • TypeScript
    • Go
  3. Software Engineer

    Tech Mahindra, Hyderabad, India

    May 2023 to July 2024

    • Engineered ETL pipelines with Python, PySpark and Apache Airflow processing 50TB+ datasets for Fortune 500 clients, cutting data preparation time 45 percent and roughly 120 thousand dollars per year in cloud compute costs.
    • Built ARIMA and Prophet time-series forecasting models at roughly 85 percent accuracy driving demand planning and budget allocation.
    • Optimized data models through indexing and partitioning, cutting query execution time 60 percent for 200+ users, and ran A/B, cohort and funnel analyses informing 1.5 million dollars or more in strategic decisions.
    • Automated 20+ Power BI and Tableau dashboards with scheduled refresh and row-level security, removing roughly 40 percent of manual reporting.
    • Python
    • PySpark
    • Airflow
    • SQL
    • BigQuery
    • PostgreSQL
    • Power BI
    • Tableau
    • Prophet
    • ARIMA
  4. Data Analyst Intern (Part-Time)

    Montclair State University, Montclair, NJ

    April 2025 to July 2025

    • Analyzed 114+ community engagement projects using statistical analysis, NLP and time-series methods.
    • Built 20+ interactive dashboards in Python, Tableau and Power BI that cut report preparation time 40 percent.
    • Applied text classification and sentiment analysis to 5,000+ qualitative survey responses, producing a strategic report presented at a multi-university conference.
    • Python
    • Pandas
    • NumPy
    • SQL
    • NLP
    • Tableau
    • Power BI

EDUCATION

Master of Science, Data Science

Montclair State University, Montclair, NJ

May 2026

EDUCATION

Bachelor of Technology, Artificial Intelligence and Data Science

KL University, Hyderabad, India

April 2024

05 / CAPABILITIES

Marked by daily use, not by years

Filter by group and the matching capabilities move to the front. Nothing is hidden, because the shape of the whole list is part of the answer.

83 capabilities across seven groups

  • Python, daily production use
  • PyTorch, daily production use
  • TensorFlow, working proficiency
  • scikit-learn, daily production use
  • XGBoost, daily production use
  • LightGBM, daily production use
  • Hugging Face Transformers, daily production use
  • Pandas, daily production use
  • NumPy, daily production use
  • SciPy, working proficiency
  • statsmodels, working proficiency
  • Feature engineering, daily production use
  • Demand elasticity modeling, daily production use
  • Time-series forecasting (ARIMA, Prophet), daily production use
  • NLP, daily production use
  • Embeddings and vector search, working proficiency
  • MLflow, daily production use
  • Weights and Biases, working proficiency
  • Databricks, daily production use
  • Databricks Feature Store, daily production use
  • Docker, daily production use
  • Kubernetes, working proficiency
  • Terraform, working proficiency
  • FastAPI model serving, daily production use
  • Blue-green deployment, daily production use
  • Model monitoring and drift detection, daily production use
  • Azure Pipelines, daily production use
  • GitHub Actions, daily production use
  • pytest, daily production use
  • Automated A/B harnesses, daily production use
  • Cookiecutter templating, daily production use
  • Delta Lake, daily production use
  • Apache Spark, daily production use
  • PySpark, daily production use
  • Apache Airflow, daily production use
  • Kafka, working proficiency
  • Azure Event Hubs, working proficiency
  • dbt, working proficiency
  • ETL and ELT, daily production use
  • Data validation and enrichment, daily production use
  • Connector library design, daily production use
  • Distributed training on Spark, daily production use
  • Azure Databricks, daily production use
  • Azure Data Lake, daily production use
  • Azure Pipelines, daily production use
  • Azure Event Hubs, daily production use
  • AWS S3, working proficiency
  • AWS Redshift, working proficiency
  • AWS Lambda, working proficiency
  • GCP BigQuery, working proficiency
  • GCP Dataflow, working proficiency
  • GCP Pub/Sub, working proficiency
  • PostgreSQL, daily production use
  • MySQL, working proficiency
  • Snowflake, working proficiency
  • MongoDB, working proficiency
  • Redis, working proficiency
  • FAISS, working proficiency
  • pgvector, working proficiency
  • RLHF, daily production use
  • DPO, daily production use
  • SFT, working proficiency
  • Agentic trajectory evaluation, daily production use
  • Tool-use grading, daily production use
  • Rubric design, daily production use
  • Adversarial prompting, daily production use
  • Red-teaming, daily production use
  • LLM-as-a-judge, daily production use
  • Inter-annotator agreement, daily production use
  • Power BI, daily production use
  • Tableau, working proficiency
  • Matplotlib, daily production use
  • Seaborn, daily production use
  • Plotly, working proficiency
  • A/B testing, daily production use
  • Cohort and funnel analysis, working proficiency
  • Hypothesis testing, daily production use
  • Python, daily production use
  • SQL, daily production use
  • TypeScript and JavaScript, working proficiency
  • Go, working proficiency
  • Bash, daily production use
  • R, working proficiency

Filled square: in daily production use. Half square: working proficiency.

06 / WRITING

First posts landing soon

Two posts are drafted. Both come out of work described above rather than from a reading list.
  • DRAFT

    What breaks when you put a pricing model in front of 200,000 predictions a month

    Not published yet

  • DRAFT

    Writing Terminal-Bench tasks that frontier agents actually fail

    Not published yet

Go to the writing index

07 / CONTACT

Currently open to AI/ML Engineer roles.

What I am looking for is production ML ownership, ideally in pricing, forecasting, recommendations, or model evaluation infrastructure.

RESUME, AT A GLANCE

Ajay Mekala

AI/ML Engineer, Walmart

Production ML at Walmart, four years of frontier model evaluation on contract, ETL and forecasting at Tech Mahindra, and an MS in Data Science finishing in May 2026.