×
Register Here to Apply for Jobs or Post Jobs. X

Senior Machine Learning Infrastructure Engineer, Recommendations and Search

Job in San Jose, Santa Clara County, California, 95199, USA
Listing for: TikTok USDS Joint Venture
Full Time position
Listed on 2026-09-09
Job specializations:
  • Software Development
    Machine Learning/ ML Engineer, AI Engineer (Applied/Software)
Salary/Wage Range or Industry Benchmark: 187000 - 360000 USD Yearly USD 187000.00 360000.00 YEAR
Job Description & How to Apply Below

Responsibilities

About the Team

We are a group of applied machine learning engineers that focus on Tik Tok recommendations and search engineers powering multiple product areas such as For-You-Page (FYP), Live Streaming, Global E-commerce, Local Services and more. We are developing innovative algorithms and techniques to improve user engagement and satisfaction, converting creative ideas into business-impacting solutions. We are interested in and excited about pushing the envelope of State-of-the-Art (SOTA) large scale machine learning to solve various real-world problems.

What

You'll Do
  • - Technical Execution & System Ownership:
    Contribute to the technical roadmap and hands-on implementation of our large-scale (in billions parameters) distributed real-time ML training and inferencing platforms that power the Tik Tok recommendation and search engines, Short Form Video (SFV) ecosystem. Have a direct business impact on Live, Global E-Commerce, Local Services and many business domains.
  • - Large-Scale Parallelism Architecture:
    Help design and scale multi-node distributed training systems, implementing advanced 3D parallelism strategies (Data, Tensor, Pipeline) to maximize compute efficiency and model scalability. Participate in the architecture, scale-testing, and maintenance of massive distributed computing foundations, GPU cluster configurations, and orchestration pipelines to achieve high hardware utilization and cluster efficiency under established SLI/SLO frameworks.
  • - Production Inference & Serving:
    Build and scale low-latency, high-throughput model serving infrastructure, optimizing inference pipelines and leveraging low-level execution paths to handle massive live traffic under strict boundary isolation.
  • - Algorithm-Infra Co-Design:
    Partner closely with Applied ML Research teams to co-design and pioneer next-generation Generative Recommendation systems, abstracting general-purpose components to support advanced generative paradigms in production.
  • - Resiliency & Fault Tolerance:
    Build robust, automated fault-detection systems and asynchronous checkpointing mechanisms to gracefully handle hardware drops, silent data corruption (SDC), or network-switch failures in multi-thousand GPU clusters.
  • - Cross-Functional Collaboration:

    Partner with Applied ML researchers and data platform teams to engineer high-throughput, secure multi-modal data processing and storage engines that prevent compliance friction.
  • - Security & Compliance Hardening:
    Implement and enforce strict encryption-at-rest/in-transit controls, access control lists (ACLs), and secure tenant isolation protocols across the entire compute stack to meet compliance objectives.
Qualifications

Minimum Qualifications
  • - Bachelor’s or Master's degree in Computer Science, Computer Engineering, or a related technical discipline.
  • - 4+ years of professional software engineering experience with deep expertise in Python, C++/Java
  • - 2+ years of direct experience building and maintaining machine learning infrastructure at enterprise scale
  • - Deep technical familiarity with the internals of core ML frameworks (PyTorch / Tensor Flow) and a strong understanding of low-level GPU memory management, CUDA interactions, and networking topologies (Infini Band/RoCE).
  • - Solid understanding of production-grade LLM training and inference tools, with hands-on profiling skills to eliminate I/O, compute, or network bottlenecks.
  • - Strong system-level troubleshooting and debugging skills, with experience profiling and eliminating I/O, compute, or network bottlenecks.
Preferred Qualifications
  • - Experience optimizing high-performance training loops to maximize Model Flops Utilization (MFU) through advanced communication-computation overlap and zero-bubble pipeline scheduling across large-scale distributed clusters using industry-standard frameworks (e.g., Megatron, Deep Speed).
  • - Proven track record of scaling LLM training or inference workloads across hundreds of GPUs, with deep familiarity in advanced serving techniques such as KV Cache management, Prefill-Decoding (PD) separation, and model quantization.
  • - Experience working in highly regulated industries, sovereign cloud environments, or dealing with federal data security compliance frameworks.
  • - Strong flavor in low-level kernel development and graph compilation, with experience in CUDA, Triton, Cutlass, TensorRT, or Triton Inference Server being a huge plus.
  • - Active background or interest in keeping up with the latest industry breakthroughs in MLOps, MoE (Mixture of Experts) routing infrastructure, and specialized hardware optimization.
  • - Strong communication skills, with the ability to collaborate effectively across team boundaries, draft clear technical design docs, and mentor team members.
About USDS

Tik Tok USDS Joint Venture LLC is dedicated to the safety and security of millions of Americans who create, discover, and connect with what they love on the apps we operate. The Joint Venture has been established in compliance with the Executive Order…

Position Requirements
10+ Years work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary