Senior Machine Learning Engineer, ML Infrastructure- Online
Seattle, King County, Washington, 98127, USA
Listed on 2026-07-26
-
Software Development
Machine Learning/ ML Engineer, AI Engineer (Applied/Software)
The Role
We are seeking a Senior ML engineer to design and evolve Unity Vector’s online model inference platform. This role focuses on building reliable infrastructure for serving machine learning models in production, optimizing inference performance, and enabling safe, efficient experimentation across high-traffic online systems.
You will work closely with ML engineers, platform teams, and product stakeholders to ensure models can be deployed, scaled, monitored, and iterated on efficiently. You will play a key role in shaping how models are packaged, served, validated, monitored, and optimized in production environments.
This role requires strong systems thinking, deep experience with production ML infrastructure, and the ability to drive architectural improvements across teams.
What you'll be doingDesign and operate large-scale online inference infrastructure that serves production ML models with low latency and high reliability, such as PyTorch, Triton Inference Server, Kubernetes, GKE, Ray, or similar distributed serving frameworks.
Develop infrastructure that supports distributed training workflows using technologies such as Pytorch, Ray Data, and Ray Train, etc.
Integrate ML pipelines with workflow orchestration systems (e.g., Flyte, Airflow, or similar) to enable reliable multi-stage training workflows
Optimize model performance through model compilation, GPU/CPU utilization improvements, request scheduling, kernel fusion, and runtime-level tuning.
Improve observability of ML systems through latency, throughput, error-rate, cost, saturation, and model-health monitoring.
Partner closely with ML engineers to support faster model iteration while maintaining production safety, scalability, and cost efficiency.
Improve the reliability and reproducibility of model serving workflows, including model packaging, artifact validation, compatibility testing, and deployment automation.
Lead architectural improvements that make the online ML platform more robust, user-friendly, scalable, and cost-efficient.
Experience building and operating production-grade online ML inference systems, such asNVIDIA Triton Inference Server, Torch Serve, Ray Serve, Tensor Flow Serving, or similar systems.
Experience with model serving frameworks such as NVIDIA Triton Inference Server, Torch Serve, Ray Serve, Tensor Flow Serving, or similar systems.
Experience optimizing inference workloads using techniques such as dynamic batching, model compilation, quantization, GPU acceleration, GPU kernel optimization, caching, or runtime tuning.
Strong experience with distributed systems, Kubernetes, autoscaling, service reliability, and production observability.
Strong programming skills in Python, with practical experience working on production ML systems and high-scale services.
Experience with PyTorch and modern model deployment workflows, including model packaging, validation, and serving lifecycle management.
Experience designing infrastructure for safe model rollout, canary testing, A/B experimentation, and automated rollback.
Strong systems thinking, with the ability to reason about latency, throughput, reliability, scalability, and cost tradeoffs in online systems.
Proven ability to lead technical direction and influence architectural decisions across teams without formal authority.
Relocation support is not available for this position
Work visa/immigration sponsorship is not available for this position
$187,200$-$243,300
This range reflects the anticipated base salary for this position. Beyond base salary, this role may be eligible for equity awards and participation in our company incentive plans (such as annual discretionary bonuses or sales commissions). The final offer amount will depend on several factors, including geographic location and the candidate’s relevant experience, professional background, and skill set.
BenefitsAt Unity, we want our team members to thrive. We offer a wide range of benefits designed to support well-being and work-life balance.
Please note:
Benefits eligibility, specific offerings, and coverage vary based on the country and employment status.
While specific benefits…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).