Tech Lead, Machine Learning Infrastructure Engineer
Job in
Seattle, King County, Washington, 98127, USA
Listed on 2026-08-17
Listing for:
TikTok USDS Joint Venture
Full Time
position Listed on 2026-08-17
Job specializations:
-
Software Development
Machine Learning/ ML Engineer, AI Engineer (Applied/Software)
Job Description & How to Apply Below
Responsibilities
We are a group of applied machine learning engineers that focus on Tik Tok recommendations and search engineers powering multiple product areas such as For-You-Page (FYP), Live Streaming, Global E-commerce, Local Services and more. We are developing innovative algorithms and techniques to improve user engagement and satisfaction, converting creative ideas into business-impacting solutions. We are interested in and excited about pushing the envelope of State-of-the-Art (SOTA) large scale machine learning to solve various real-world problems.
WhatYou'll Do
- Technical Leadership:
Drive the technical roadmap for our large-scale (in billions parameters) distributed real-time ML training and inferencing platforms that power the Tik Tok recommendation and search engines, Short Form Video (SFV) ecosystem. Have a direct business impact on Live, Global E-Commerce, Local Services and many businesses domains. - Large-Scale Parallelism Architecture:
Architect and scale multi-node distributed training systems, implementing advanced 3D parallelism strategies (Data, Tensor, Pipeline) to maximize compute efficiency and model scalability. Lead the architecture, scale-testing, and maintenance of massive distributed computing foundations, GPU cluster configurations, and orchestration pipelines, establish robust SLI/SLO frameworks while maximizing hardware utilization and cluster efficiency to expedite innovation. - Production Inference & Serving:
Build and scale low-latency, high-throughput model serving infrastructure, optimizing inference pipelines and leveraging low-level execution paths to handle massive live traffic under strict boundary isolation. - Algorithm-Infra Co-Design:
Partner closely with Applied ML Research teams to co-design and pioneer next-generation Generative Recommendation systems, abstracting general-purpose components to support advanced generative paradigms in production. - Resiliency & Fault Tolerance:
Design robust, automated fault-detection systems and asynchronous checkpointing mechanisms to gracefully handle hardware drops, silent data corruption (SDC), or network-switch failures in multi-thousand GPU clusters. - Cross-Functional Collaboration:
Partner with Applied ML researchers and data platform teams to engineer high-throughput, secure multi-modal data processing and storage engines that prevent compliance friction. - Security & Compliance Hardening:
Implement and enforce strict encryption-at-rest/in-transit controls, access control lists (ACLs), and secure tenant isolation protocols across the entire compute stack to meet compliance objectives.
Minimum Qualifications
- Bachelor’s or Master's degree in Computer Science, Computer Engineering, or a related technical discipline.
- 5+ years of professional software engineering experience with deep expertise in Python, C++/Java and a proven track record of designing large-scale distributed systems.
- 3+ years of direct experience building and maintaining machine learning infrastructure at enterprise scale (managing large GPU clusters, Kubernetes, or native Slurm environments).
- Deep technical familiarity with the internals of core ML frameworks (PyTorch / Tensorflow) and a strong understanding of low-level GPU memory management, CUDA interactions, and networking topologies (Infini Band/RoCE).
- Solid understanding of production-grade LLM training and inference tools, with hands-on profiling skills to eliminate I/O, compute, or network bottlenecks.
- Strong system-level troubleshooting and debugging skills, with experience profiling and eliminating I/O, compute, or network bottlenecks.
- Experience optimizing high-performance training loops to maximize Model Flops Utilization (MFU) through advanced communication-computation overlap and zero-bubble pipeline scheduling across large-scale distributed clusters using industry-standard frameworks (e.g., Megatron, Deep Speed).
- Proven track record of scaling LLM training or inference workloads across hundreds of GPUs, with deep familiarity in advanced serving techniques such as KV Cache management, Prefill-Decoding (PD) separation, and model quantization.
- Experience working in…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×