Principal, Machine Learning Engineer
Listed on 2026-09-04
-
Software Development
Machine Learning/ ML Engineer, AI Engineer (Applied/Software)
Position Summary...
The Principal, Machine Learning Engineer will lead the design and development of scalable, production-grade ML systems that address complex business challenges. This role involves architecting end-to-end ML pipelines, driving innovation in advanced modeling techniques, and ensuring robust deployment and monitoring practices. The position requires collaboration with cross-functional teams to align ML solutions with strategic objectives while mentoring technical staff and promoting engineering excellence.
The successful candidate will apply sound technical judgment, foster continuous improvement, and contribute to Walmart’s reputation through thought leadership and technical communication.
The Person alization team at Walmart is committed to enhancing customer experiences by delivering seamless, tailored journeys across all engagement channels. Operating at the nexus of vast product assortments, millions of customers, and thousands of stores, the team develops advanced AI-driven solutions that empower both customers and associates. Comprising data scientists, engineers, and product experts, the team designs and builds innovative machine learning models and systems using cutting-edge techniques such as deep learning, reinforcement learning, and natural language processing.
Their work shapes the future of e-commerce by enabling personalized, efficient, and trusted interactions.
- Architect and lead the development of scalable, production-grade ML systems powering customer-facing personalization and recommendation experiences.
- Own technical direction across the end-to-end ML lifecycle, including data and feature pipelines, model training, evaluation, deployment, serving, monitoring, and continuous improvement.
- Design high-performance batch and real-time ML architectures capable of operating reliably at large scale.
- Partner with Data Scientists and other ML practitioners to translate model prototypes and experimentation into scalable, maintainable production systems.
- Establish engineering patterns and best practices for model deployment, feature management, reproducibility, automated testing, observability, retraining, and model lifecycle management.
- Develop and optimize ML solutions using techniques such as recommendation and ranking, deep learning, representation learning, NLP, and other advanced machine learning approaches.
- Design systems that balance model quality with production requirements such as latency, throughput, scalability, reliability, maintainability, and cost.
- Build reusable ML capabilities, services, and infrastructure that accelerate development and adoption across teams.
- Define technical strategies and architecture for complex and ambiguous ML problems and influence engineering decisions across multiple teams.
- Partner cross-functionally with ML, Software Engineering, Data Science, Product, and Platform teams to align ML capabilities with customer and business objectives.
- Mentor senior technical talent, raise engineering standards, and provide technical leadership across the broader ML engineering community.
- Extensive experience designing, building, deploying, and operating large-scale machine learning systems in production.
- Strong software engineering fundamentals and proficiency in Python and production software development practices.
- Deep understanding of the end-to-end ML lifecycle, including data preparation, feature engineering, model development, training, evaluation, deployment, inference, monitoring, and retraining.
- Experience building production ML systems using frameworks such as PyTorch, Tensor Flow, Scikit-learn, XGBoost, or comparable technologies.
- Expertise in one or more relevant ML domains such as recommendation and ranking, deep learning, representation learning, NLP, reinforcement learning, or related advanced ML techniques.
- Experience designing scalable batch and/or real-time model training, inference, and feature-serving architectures.
- Strong understanding of distributed systems, APIs and services, cloud-native architectures, and scalable data processing.
- Experience with production ML engineering practices…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).