×
Register Here to Apply for Jobs or Post Jobs. X

ML Systems Engineer, -Scale Model Training & RL Infrastructure

Remote / Online - Candidates ideally in
Palo Alto, Santa Clara County, California, 94301, USA
Listing for: Nebius
Apprenticeship/Internship, Remote/Work from Home position
Listed on 2026-07-26
Job specializations:
  • Software Development
    Machine Learning/ ML Engineer, AI Engineer (Applied/Software)
Salary/Wage Range or Industry Benchmark: 195200 USD Yearly USD 195200.00 YEAR
Job Description & How to Apply Below
Position: ML Systems Engineer, Large-Scale Model Training & RL Infrastructure

About Nebius:

Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure.

Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI.

Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D.

The role

Nebius Token Factory is building an AI training and model post-training capability for frontier model improvement. This role owns the infrastructure that makes large-scale training and RL experiments possible, reliable, reproducible, and efficient. The work sits at the intersection of distributed systems, GPU performance, model training frameworks, RL pipelines, and production engineering.

A Senior ML Systems Engineer owns substantial training or RL infrastructure components end to end. They are deeply hands-on, can debug difficult distributed training failures independently, and can deliver measurable improvements in experiment throughput, stability, and GPU utilization.

Your responsibilities:

  • Build and maintain distributed training infrastructure for SFT, continued pretraining, preference optimization, and RL workloads.

  • Integrate and extend frameworks such as Megatron-LM, Deep Speed, PyTorch FSDP/DTensor, Ray, verl, slime, AReaL, OpenRLHF, or equivalent internal systems.

  • Implement and debug parallelism strategies including tensor, pipeline, sequence/context, expert, and data parallelism.

  • Build reliable rollout, reward model serving, replay/data buffer, checkpointing, evaluation, and experiment orchestration components for RL training.

  • Profile and improve GPU utilization, communication efficiency, memory usage, and training throughput.

  • Diagnose failures across NCCL, CUDA, PyTorch, Ray, schedulers, storage, networking, and checkpointing layers.

  • Create reproducible training runs, launch scripts, dashboards, runbooks, and operational tooling for research users.

  • Partner with research scientists to turn algorithmic training recipes into scalable, debuggable systems.

  • Write clear design docs, incident reports, benchmark reports, and operating guides.

Must-haves:

  • Strong Python and PyTorch engineering skills.

  • Hands-on experience with distributed model training, large-scale ML systems, or GPU cluster workloads.

  • Practical understanding of transformer training bottlenecks, memory pressure, gradient/optimizer state, communication overhead, and checkpointing.

  • Experience debugging production or research training jobs across multiple GPUs or nodes.

  • Ability to reason quantitatively about throughput, utilization, memory, reliability, cost, and research velocity.

  • Strong communication skills and ability to collaborate with researchers, ML engineers, platform engineers, and leadership.

Nice-to-have:

  • Experience with Megatron-LM, Deep Speed, PyTorch FSDP/DTensor, Ray, Slurm, Kubernetes, or large internal training platforms.

  • Experience with RL infrastructure frameworks such as verl, slime, AReaL, OpenRLHF, TRL, or custom PPO/GRPO/RLHF systems.

  • Familiarity with NCCL, CUDA, Triton, Nsight, Infini Band, RDMA, RoCE, H100/H200/B200 clusters, or storage/network bottlenecks.

  • Experience supporting SFT, DPO, PPO, GRPO, RLAIF, reward model serving, rollout generation, or agent training workloads.

  • Open-source contributions to distributed training, RL infrastructure, PyTorch, Ray, Megatron, Deep Speed, or related systems.

Key employee benefits in the US:

  • Health insurance:100% company-paid medical, dental, and vision coverage for employees and families.

  • 401(k) plan:Up to 4% company match with immediate vesting.

  • Parental leave:20 weeks paid for primary caregivers, 12 weeks for secondary caregivers.

  • Remote work reimbursement:Up to $85/month for mobile and internet.

  • Disability & life insurance
    :
    Company-paid short-term, long-term and life insurance coverage.

Pay Transparency

We offer competitive compensation and benefits packages. Actual compensation will be determined based on job-related factors, including experience, skills, qualifications, the level at which the candidate is hired, and geographic location, consistent with applicable law.

Base Compensation Range

$195,200 — $262,200 USD

Benefits & Perks:

  • Competitive compensation
  • Career growth and learning opportunities
  • Flexibility and ownership
  • Collaborative and innovative culture
  • Opportunity to work on impactful AI projects
  • International environment and talented teams

What's it like to work at Nebius:

Fast moving
- Bold thinking
- Constant growth
- Meaningful impact
- Trust and real ownership
- Opportunity to shape the future of AI

Equal Opportunity Statement:

Nebius is an equal opportunity…

To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary