×
Register Here to Apply for Jobs or Post Jobs. X

Tech Lead, Machine Learning Infrastructure Engineer

Job in Seattle, King County, Washington, 98127, USA
Listing for: TikTok USDS Joint Venture
Full Time position
Listed on 2026-08-17
Job specializations:
  • Software Development
    Machine Learning/ ML Engineer, AI Engineer (Applied/Software)
Salary/Wage Range or Industry Benchmark: 198000 - 416000 USD Yearly USD 198000.00 416000.00 YEAR
Job Description & How to Apply Below

Responsibilities

We are a group of applied machine learning engineers that focus on Tik Tok recommendations and search engineers powering multiple product areas such as For-You-Page (FYP), Live Streaming, Global E-commerce, Local Services and more. We are developing innovative algorithms and techniques to improve user engagement and satisfaction, converting creative ideas into business-impacting solutions. We are interested in and excited about pushing the envelope of State-of-the-Art (SOTA) large scale machine learning to solve various real-world problems.

What

You'll Do
  • Technical Leadership:
    Drive the technical roadmap for our large-scale (in billions parameters) distributed real-time ML training and inferencing platforms that power the Tik Tok recommendation and search engines, Short Form Video (SFV) ecosystem. Have a direct business impact on Live, Global E-Commerce, Local Services and many businesses domains.
  • Large-Scale Parallelism Architecture:
    Architect and scale multi-node distributed training systems, implementing advanced 3D parallelism strategies (Data, Tensor, Pipeline) to maximize compute efficiency and model scalability. Lead the architecture, scale-testing, and maintenance of massive distributed computing foundations, GPU cluster configurations, and orchestration pipelines, establish robust SLI/SLO frameworks while maximizing hardware utilization and cluster efficiency to expedite innovation.
  • Production Inference & Serving:
    Build and scale low-latency, high-throughput model serving infrastructure, optimizing inference pipelines and leveraging low-level execution paths to handle massive live traffic under strict boundary isolation.
  • Algorithm-Infra Co-Design:
    Partner closely with Applied ML Research teams to co-design and pioneer next-generation Generative Recommendation systems, abstracting general-purpose components to support advanced generative paradigms in production.
  • Resiliency & Fault Tolerance:
    Design robust, automated fault-detection systems and asynchronous checkpointing mechanisms to gracefully handle hardware drops, silent data corruption (SDC), or network-switch failures in multi-thousand GPU clusters.
  • Cross-Functional Collaboration:

    Partner with Applied ML researchers and data platform teams to engineer high-throughput, secure multi-modal data processing and storage engines that prevent compliance friction.
  • Security & Compliance Hardening:
    Implement and enforce strict encryption-at-rest/in-transit controls, access control lists (ACLs), and secure tenant isolation protocols across the entire compute stack to meet compliance objectives.
Qualifications

Minimum Qualifications
  • Bachelor’s or Master's degree in Computer Science, Computer Engineering, or a related technical discipline.
  • 5+ years of professional software engineering experience with deep expertise in Python, C++/Java and a proven track record of designing large-scale distributed systems.
  • 3+ years of direct experience building and maintaining machine learning infrastructure at enterprise scale (managing large GPU clusters, Kubernetes, or native Slurm environments).
  • Deep technical familiarity with the internals of core ML frameworks (PyTorch / Tensorflow) and a strong understanding of low-level GPU memory management, CUDA interactions, and networking topologies (Infini Band/RoCE).
  • Solid understanding of production-grade LLM training and inference tools, with hands-on profiling skills to eliminate I/O, compute, or network bottlenecks.
  • Strong system-level troubleshooting and debugging skills, with experience profiling and eliminating I/O, compute, or network bottlenecks.
Preferred Qualifications
  • Experience optimizing high-performance training loops to maximize Model Flops Utilization (MFU) through advanced communication-computation overlap and zero-bubble pipeline scheduling across large-scale distributed clusters using industry-standard frameworks (e.g., Megatron, Deep Speed).
  • Proven track record of scaling LLM training or inference workloads across hundreds of GPUs, with deep familiarity in advanced serving techniques such as KV Cache management, Prefill-Decoding (PD) separation, and model quantization.
  • Experience working in…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary