Senior ML Infra Engineer
Listed on 2026-07-22
-
Software Development
Machine Learning/ ML Engineer
What You Will Own
Design, build, and evolve infrastructure for ML training workflows, training deployment, experiment execution, and production handoff.
Build and maintain deployment paths for models, jobs, services, and supporting infrastructure across development and production environments.
Improve reliability, scalability, observability, and developer experience for ML workflows and platform tools.
Define interfaces, automation, metadata, artifacts, configuration, environment management, and lifecycle boundaries for ML systems.
Collaborate with research, product, data, and engineering partners to translate incomplete ML workflow needs into maintainable systems.
Support production usage by building clear operational tooling, debugging paths, and safe rollout mechanisms.
Strong software engineering and infrastructure fundamentals, with experience owning production or near-production systems.
Practical experience with PyTorch and ML training workflows, including job orchestration, compute environments, artifact management, and deployment automation.
Solid understanding of heterogeneous computing and high-performance computing, especially for ML training or serving workloads.
Good understanding of model lifecycle concerns: data, configs, checkpoints, artifacts, reproducibility, rollout, rollback, and observability.
Ability to build reliable platform abstractions without hiding the important details ML practitioners need to control.
Clear technical and product sense: you can prioritize platform work that unlocks real training or deployment velocity.
High standards for engineering quality, including tests, documentation, debugging tools, and maintainable system design.
Python
Py Torch
Heterogeneous computing and high-performance computing
ML training pipelines, job orchestration, compute scheduling, containers, and deployment automation
Model artifacts, metadata, storage, experiment tracking, and configuration systems
Model serving, inference deployment, APIs, queues, and observability tools
Docker, CI/CD, cloud infrastructure, GPUs, and internal platform tooling
Experience building training deployment systems, model release workflows, or ML platform tooling for research and production teams.
Experience with inference deployment, model serving, online/offline evaluation, performance tuning, or rollout safety.
Experience with distributed training, GPU infrastructure, workload scheduling, artifact/version management, or reproducibility tooling.
Understanding of CUDA, GPU architecture, or low-level performance optimization.
Experience with open-source inference and serving frameworks such as vLLM, TensorRT, Triton, or similar systems.
Experience migrating ad hoc notebooks, scripts, or manual ML processes into reliable platform workflows.
You mainly want to train models personally and do not enjoy building infrastructure for others to use.
You are comfortable with manual ML workflows and do not care about reproducibility, deployment, or operational quality.
You prefer narrow implementation tasks and do not want to reason about system boundaries, platform UX, or long-term maintenance.
You over-abstract ML workflows without understanding where researchers and engineers need control and visibility.
ML teams move faster when training, deployment, and production usage are supported by reliable infrastructure instead of scattered scripts and manual processes. This role will directly shape how models move from experimentation to production, how safely they are deployed, and how efficiently the team can iterate. For the right person, it is a high-ownership platform role with deep impact on both engineering quality and ML velocity.
WhatWe Would Like to See When You Apply
ML infrastructure, training platforms, deployment systems, or model serving systems you have owned.
Examples of how you improved training reliability, deployment velocity, reproducibility, observability, or operational safety.
Cases where you turned messy ML workflows into maintainable tools, services, or platform abstractions.
Examples that show your technical judgment, communication, and ability to work across research and engineering needs.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).