Senior Machine Learning Infrastructure Engineer, Recommendations and Search
Listed on 2026-09-09
-
Software Development
Machine Learning/ ML Engineer, AI Engineer (Applied/Software)
Responsibilities
About the Team
We are a group of applied machine learning engineers that focus on Tik Tok recommendations and search engineers powering multiple product areas such as For-You-Page (FYP), Live Streaming, Global E-commerce, Local Services and more. We are developing innovative algorithms and techniques to improve user engagement and satisfaction, converting creative ideas into business-impacting solutions. We are interested in and excited about pushing the envelope of State-of-the-Art (SOTA) large scale machine learning to solve various real-world problems.
WhatYou'll Do
- - Technical Execution & System Ownership:
Contribute to the technical roadmap and hands-on implementation of our large-scale (in billions parameters) distributed real-time ML training and inferencing platforms that power the Tik Tok recommendation and search engines, Short Form Video (SFV) ecosystem. Have a direct business impact on Live, Global E-Commerce, Local Services and many business domains. - - Large-Scale Parallelism Architecture:
Help design and scale multi-node distributed training systems, implementing advanced 3D parallelism strategies (Data, Tensor, Pipeline) to maximize compute efficiency and model scalability. Participate in the architecture, scale-testing, and maintenance of massive distributed computing foundations, GPU cluster configurations, and orchestration pipelines to achieve high hardware utilization and cluster efficiency under established SLI/SLO frameworks. - - Production Inference & Serving:
Build and scale low-latency, high-throughput model serving infrastructure, optimizing inference pipelines and leveraging low-level execution paths to handle massive live traffic under strict boundary isolation. - - Algorithm-Infra Co-Design:
Partner closely with Applied ML Research teams to co-design and pioneer next-generation Generative Recommendation systems, abstracting general-purpose components to support advanced generative paradigms in production. - - Resiliency & Fault Tolerance:
Build robust, automated fault-detection systems and asynchronous checkpointing mechanisms to gracefully handle hardware drops, silent data corruption (SDC), or network-switch failures in multi-thousand GPU clusters. - - Cross-Functional Collaboration:
Partner with Applied ML researchers and data platform teams to engineer high-throughput, secure multi-modal data processing and storage engines that prevent compliance friction. - - Security & Compliance Hardening:
Implement and enforce strict encryption-at-rest/in-transit controls, access control lists (ACLs), and secure tenant isolation protocols across the entire compute stack to meet compliance objectives.
Minimum Qualifications
- - Bachelor’s or Master's degree in Computer Science, Computer Engineering, or a related technical discipline.
- - 4+ years of professional software engineering experience with deep expertise in Python, C++/Java
- - 2+ years of direct experience building and maintaining machine learning infrastructure at enterprise scale
- - Deep technical familiarity with the internals of core ML frameworks (PyTorch / Tensor Flow) and a strong understanding of low-level GPU memory management, CUDA interactions, and networking topologies (Infini Band/RoCE).
- - Solid understanding of production-grade LLM training and inference tools, with hands-on profiling skills to eliminate I/O, compute, or network bottlenecks.
- - Strong system-level troubleshooting and debugging skills, with experience profiling and eliminating I/O, compute, or network bottlenecks.
- - Experience optimizing high-performance training loops to maximize Model Flops Utilization (MFU) through advanced communication-computation overlap and zero-bubble pipeline scheduling across large-scale distributed clusters using industry-standard frameworks (e.g., Megatron, Deep Speed).
- - Proven track record of scaling LLM training or inference workloads across hundreds of GPUs, with deep familiarity in advanced serving techniques such as KV Cache management, Prefill-Decoding (PD) separation, and model quantization.
- - Experience working in highly regulated industries, sovereign cloud environments, or dealing with federal data security compliance frameworks.
- - Strong flavor in low-level kernel development and graph compilation, with experience in CUDA, Triton, Cutlass, TensorRT, or Triton Inference Server being a huge plus.
- - Active background or interest in keeping up with the latest industry breakthroughs in MLOps, MoE (Mixture of Experts) routing infrastructure, and specialized hardware optimization.
- - Strong communication skills, with the ability to collaborate effectively across team boundaries, draft clear technical design docs, and mentor team members.
Tik Tok USDS Joint Venture LLC is dedicated to the safety and security of millions of Americans who create, discover, and connect with what they love on the apps we operate. The Joint Venture has been established in compliance with the Executive Order…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).