Data Engineer
Listed on 2026-07-27
-
IT/Tech
Machine Learning/ ML Engineer, AI Engineer (Applied/Software), Data Engineering
Staff Data Engineer
The Staff Data Engineer, MLOps leads the design, build, and optimization of Hershey’s machine learning operations platform—enabling data science and AI teams to develop, deploy, monitor, and govern ML models at enterprise scale. Sitting within Platform Engineering, this role owns the infrastructure, tooling, and automation that move models from experimentation to production with speed and confidence.
Job Location: Dallas, TX
Posted Date:
Jul 22, 2026
Requisition Number: 129110
What We Are Building for HersheyHershey is building an AI-driven enterprise platform that transforms how we compete across retail, supply chain, and commercial. We are standing up a unified MLOps foundation on Azure and Databricks that will power demand forecasting models that sharpen inventory and production planning, real‑time pricing and promotion optimization engines for our retail and commercial partners, computer vision and quality‑detection models on manufacturing lines, and next‑generation consumer analytics that personalize how we reach millions of households.
This role is at the center of that transformation—engineering the platform that turns breakthrough data science into production AI at Hershey scale.
Duties & Responsibilities
- ML Platform Engineering & Infrastructure
- Design and maintain the end-to-end MLOps platform on Azure and Databricks: model training infrastructure, feature stores, experiment tracking, model registries, and serving endpoints.
- Build and optimize CI/CD pipelines for automated model training, validation, packaging, and deployment across environments.
- Model Deployment, Monitoring & Lifecycle Management
- Implement model serving patterns (batch, real‑time, edge) with blue‑green and canary deployment strategies for safe rollouts.
- Build monitoring frameworks for data drift, concept drift, and prediction quality; automate alerting and retraining triggers.
- Governance, Reproducibility & Responsible AI
- Enforce ML governance: model versioning, experiment lineage, artifact management, approval workflows, and audit trails.
- Embed responsible AI practices including explainability tooling, bias detection, and documentation standards.
- Infrastructure as Code & Cost Optimization
- Author IaC (Terraform/Bicep) for Azure ML work spaces, Databricks clusters, networking, and compute; optimize costs through autoscaling, spot instances, and GPU scheduling.
- Collaboration & Enablement
- Partner with Data Scientists to product ionize models; develop self‑service templates and documentation for platform onboarding; mentor junior engineers.
- MLOps & ML Engineering:
Experience taking ML models from experimentation to production, including training automation, model packaging, deployment, and monitoring. Our environment uses MLflow, Databricks Model Serving, and Azure Machine Learning. - Cloud & Platforms:
Strong hands‑on experience with Azure Cloud and Databricks. Familiarity with services such as Azure ML, AKS, Azure Dev Ops, Data Factory, Unity Catalog, Workflows, and Model Registry. - Programming & Development:
Strong Python and SQL; experience with ML frameworks (PyTorch, Scikit‑learn, XGBoost); comfort building APIs and writing modular, testable code. - Collaboration & Communication:
Proven ability to partner across Data Science, Architecture, and business teams; experience mentoring engineers and driving technical standards.
- CI/CD & IaC: ML‑specific CI/CD pipelines (Azure Dev Ops, Git Hub Actions);
Terraform or Bicep for infrastructure provisioning. - Containerization & Orchestration:
Experience with Docker and Kubernetes for model serving and workload management. - Monitoring & Observability:
Drift detection, prediction quality tracking, and observability tooling (Evidently AI, Azure Monitor, Grafana). - Certifications:
Azure Data Engineer (DP-203), Azure AI Engineer (AI-102), or Databricks ML Professional.
- Bachelor’s degree in Computer Science, Engineering, Data Science, or related field;
Master’s preferred. - 5–10 years in software, ML, data platform, or infrastructure engineering with 3+ years building or operating ML pipelines, model serving…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).