×
Register Here to Apply for Jobs or Post Jobs. X

AI Training Infrastructure Engineer

Job in Bellevue, King County, Washington, 98009, USA
Listing for: Designworks Talent
Apprenticeship/Internship position
Listed on 2026-07-24
Job specializations:
  • Software Development
    AI Engineer (Applied/Software), Machine Learning/ ML Engineer
Salary/Wage Range or Industry Benchmark: 180000 - 240000 USD Yearly USD 180000.00 240000.00 YEAR
Job Description & How to Apply Below

AI Training Infrastructure Engineer

Location: Hybrid | Bellevue, WA Area
Titles: Senior and Staff (multiple roles available)

Build the Training Infrastructure Powering Next-Generation AI Models
About the Opportunity

A well-funded, rapidly growing AI infrastructure company is building a next-generation cloud platform designed to power the full lifecycle of artificial intelligence. The organization is developing a comprehensive AI infrastructure, platform, and services portfolio that supports the full spectrum of AI workloads—including large-scale compute, model training, fine-tuning, inference, and emerging agentic AI applications.

Backed by significant long-term investment, the company combines the speed, ownership, and innovation of a startup with the stability and resources of an established parent organization. Engineering teams are intentionally lean, highly collaborative, and AI-native, leveraging modern tooling and automation to build infrastructure capable of supporting the industry s most demanding AI workloads.

We re seeking AI Training Infrastructure Engineers to build and scale the distributed systems that power large-scale AI model training. This team focuses on reliability, efficiency, and operational excellence across GPU clusters, enabling researchers and engineers to train and deploy advanced AI models at scale.

The Opportunity

This is a foundational engineering role focused on building the infrastructure layer behind large-scale AI training workloads. You ll work on distributed training systems, GPU clusters, model pipelines, and the tooling required to make AI development more reliable, efficient, and scalable.

You ll collaborate closely with infrastructure, orchestration, performance, and machine learning teams to solve complex challenges around distributed computing, fault tolerance, training efficiency, and production readiness.

This opportunity is ideal for engineers who enjoy building highly scalable systems and working at the intersection of AI research, infrastructure engineering, and distributed computing.

What You ll Do
  • Build and scale distributed training infrastructure supporting large AI models across large GPU clusters.
  • Design and improve systems that increase training reliability, efficiency, and resource utilization.
  • Develop solutions for fault tolerance, checkpointing, recovery, and large-scale training operations.
  • Integrate AI models into production training pipelines in partnership with platform, orchestration, and performance engineering teams.
  • Diagnose and resolve issues impacting training throughput, stability, reliability, and cost efficiency.
  • Build tools and automation that improve the developer experience for AI researchers and engineers.
  • Establish best practices for training infrastructure, operational processes, and platform reliability.
  • Contribute to the evolution of the AI infrastructure platform as an early member of the engineering team.
What We re Looking For
  • Hands-on experience building and operating distributed training systems or large-scale machine learning infrastructure.
  • Experience supporting large AI models, foundation models, post-training workflows, or similar ML systems.
  • Strong understanding of the reliability, scalability, and efficiency challenges associated with multi-node GPU training.
  • Experience integrating training systems with production machine learning pipelines.
  • Strong programming skills and experience working with complex distributed systems.
  • Ability to independently own technically challenging projects in a fast-moving engineering environment.
  • Comfortable operating with high ownership and limited process overhead.
Preferred Qualifications
  • Experience with distributed training frameworks such as PyTorch Distributed, Deep Speed, Megatron-LM, Ray, or similar technologies.
  • Experience with supervised fine-tuning (SFT), reinforcement learning from human feedback (RLHF), or other post-training workflows.
  • Background operating AI training infrastructure at scale within a hyperscaler, AI research organization, cloud provider, or GPU cloud environment.
  • Experience optimizing GPU utilization, training performance, or distributed system reliability.
  • Familiarity with…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary