×
Register Here to Apply for Jobs or Post Jobs. X

MLops Engineer

Job in Malden, Middlesex County, Massachusetts, 02148, USA
Listing for: Arrayo
Full Time position
Listed on 2026-10-02
Job specializations:
  • Software Development
    Machine Learning/ ML Engineer, AI Engineer (Applied/Software)
Salary/Wage Range or Industry Benchmark: 130000 - 185000 USD Yearly USD 130000.00 185000.00 YEAR
Job Description & How to Apply Below

MLops Engineer (Training Scalability & Workflow Optimization) Overview

We are seeking an MLops Engineer to lead the scaling of machine learning training pipelines and ensure the robustness and efficiency of our end-to-end ML workflows. This role focuses on leveraging Flyte
, Kubernetes (GPU optimization),
Docker
, and distributed training frameworks such as Ray to optimize and streamline our ML infrastructure.

Responsibilities
  • Workflow Orchestration: Develop and maintain ML workflows using Flyte to manage complex ML pipelines for training, testing, and deployment.
  • Training Scalability: Architect and scale large-scale ML training systems on GPU-backed Kubernetes clusters
    , including auto-scaling and performance tuning for multi-node/multi-GPU workloads.
  • Distributed Computing: Implement distributed model training pipelines using frameworks like Ray for parallelization and resource efficiency.
  • Containerization: Design, build, and optimize Docker images for ML workloads with a focus on reproducibility and security.
  • Resource Optimization: Debug and optimize GPU utilization, memory, and compute bottlenecks during training and inference phases.
  • Monitoring & Maintenance: Integrate monitoring for ML jobs, track resource consumption, and enforce cost-efficient resource utilization.
  • Collaboration: Work closely with data scientists and ML engineers to productize and scale ML experiments.
Qualifications
  • Strong proficiency with Kubernetes (GPU scheduling, Helm, cluster autoscaling).
  • Hands‑on experience with Flyte or similar workflow orchestration tools (Airflow, Prefect).
  • Deep knowledge of distributed ML training (e.g., PyTorch DDP, Ray, Horovod).
  • Expertise in Docker and container lifecycle management.
  • Solid understanding of GPU hardware/software stack (CUDA, NCCL).
  • Familiarity with CI/CD for ML (MLops pipelines using tools like Git Hub Actions, ArgoCD).
  • Bonus:
    Familiarity with observability tools for ML systems (Prometheus, Grafana).
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary