Machine Learning Engineer – ML Frameworks
Job in
California, Moniteau County, Missouri, 65018, USA
Listed on 2026-09-05
Listing for:
Jobtailor
Full Time
position Listed on 2026-09-05
Job specializations:
-
Software Development
Machine Learning/ ML Engineer, AI Engineer (Applied/Software), Cloud Engineer - Software
Job Description & How to Apply Below
Location: California
- Design, develop, and maintain robust AI/ML infrastructure solutions supporting training and deployment of large-scale AI models using Kubernetes and Python on AWS cloud
- Implement and improve distributed training frameworks leveraging GPUs to improve performance and scalability
- Improve resiliency, elasticity, data loading, and out-of-the-box support for FSDP and model parallelism
- Improve orchestration and scheduling to train better models
- Scale the number of jobs and enable faster experimentation with AutoML and similar tools
- Collaborate with data scientists and ML researchers to streamline model training pipelines and ensure efficient resource utilization
- Drive innovation in infrastructure practices supporting machine learning research and development
- PhD or Master’s in computer science or related field and 5+ years of hands-on industry experience
- Proven proficiency with Python and developing systems, frameworks and SDKs
- Experience with infrastructure and understanding of model serving, training, orchestration, and management of GPU resources
- Experience with machine learning and distributed Py Torch
- Strong critical thinking, analytical and quantitative problem-solving ability
- Excellent communication, relationship skills and a strong teammate
- Experience with Kube Flow, MLFlow, Ray, Sage Maker, or similar (added plus)
- Experience with PyTorch distributed, MPI, Megatron, Horovod and other AI training frameworks (added plus)
Demonstrates expertise in designing and maintaining AI/ML infrastructure solutions, with a strong focus on Python, Kubernetes, and AWS. Proven ability to enhance distributed training frameworks and optimize resource utilization for large-scale AI model deployment.
Highest-signal resume keywords- Python Proficiency
- Kubernetes Experience
- Distributed PyTorch Knowledge
- AI/ML Infrastructure Development
- GPU Resource Management
- AI/ML Infrastructure Solutions
- Distributed Training Frameworks
- Model Serving
- Orchestration and Scheduling
- AutoML Tools
- Python Development
- Critical Thinking
- Analytical Problem-Solving
- Quantitative Analysis
- Machine Learning
- Excellent Communication
- Relationship Skills
- Team Collaboration
- PhD in Computer Science
- Master’s in Computer Science
- AI Models
- Machine Learning Research
- Infrastructure Practices
- Resource Utilization
- Scalability
- Kube Flow
- MLFlow
- Ray
- Sage Maker
- PyTorch Distributed
- MPI
- Megatron
- Horovod
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×