×
Register Here to Apply for Jobs or Post Jobs. X

AI Infrastructure — Training Engineer; Model

Job in Menlo Park, San Mateo County, California, 94029, USA
Listing for: Stealth Startup
Apprenticeship/Internship position
Listed on 2026-07-23
Job specializations:
  • Software Development
    Machine Learning/ ML Engineer, AI Engineer (Applied/Software), Software Engineer
Salary/Wage Range or Industry Benchmark: 180000 - 280000 USD Yearly USD 180000.00 280000.00 YEAR
Job Description & How to Apply Below
Position: AI Infrastructure — Training Engineer (Large Model) [33251]
  • Distributed training framework optimization. Own the R&D and tuning of distributed training frameworks for large models (LLMs, multimodal), resolving scalability bottlenecks at the scale of 10k–100k GPU clusters.
  • Kernel & performance tuning. Work close to the underlying hardware (NVIDIA GPU / NPU) on kernel acceleration, memory optimization, and communication optimization (tensor parallelism, pipeline parallelism, ZeRO, and related techniques).
  • System resilience & scheduling. Build stable large-scale training clusters; design high-availability fault-tolerance mechanisms (checkpoint/resume, automatic recovery) and compute scheduling strategies to raise overall cluster throughput and resource utilization.
  • Training pipeline engineering. Build an end-to-end MLOps platform spanning data preprocessing, distributed training, model fine-tuning (RLHF / DPO, etc.), and automated evaluation.
Qualifications
  • Education. Bachelor's degree or above in Computer Science, Software Engineering, Electrical Engineering, or a related field.
  • Programming. Very strong engineering implementation skills; proficient in C/C++ and Python, with a solid foundation in data structures and algorithms.
  • Distributed & parallel computing. Hands-on mastery of mainstream distributed training frameworks such as PyTorch, Megatron-LM, Deep Speed, Deep Speed-Chat, or Horovod.
  • Low-level systems & communication. Familiar with Linux internals, the network stack (RoCE/RDMA), GPU communication primitives (e.g., NCCL), and common storage systems.
  • Tuning & debugging. Skilled with profiling and debugging tools such as Nsight, GDB, and PyTorch Profiler; able to quickly diagnose cluster deadlocks, performance bottlenecks, and out-of-memory (OOM) issues.
#J-18808-Ljbffr
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary