×
Register Here to Apply for Jobs or Post Jobs. X

Senior ML Infra Engineer

Job in Santa Clara, Santa Clara County, California, 95053, USA
Listing for: Maxinsights Corporation
Full Time position
Listed on 2026-07-22
Job specializations:
  • Software Development
    Machine Learning/ ML Engineer
Salary/Wage Range or Industry Benchmark: 180000 - 240000 USD Yearly USD 180000.00 240000.00 YEAR
Job Description & How to Apply Below

What You Will Own

  • Design, build, and evolve infrastructure for ML training workflows, training deployment, experiment execution, and production handoff.

  • Build and maintain deployment paths for models, jobs, services, and supporting infrastructure across development and production environments.

  • Improve reliability, scalability, observability, and developer experience for ML workflows and platform tools.

  • Define interfaces, automation, metadata, artifacts, configuration, environment management, and lifecycle boundaries for ML systems.

  • Collaborate with research, product, data, and engineering partners to translate incomplete ML workflow needs into maintainable systems.

  • Support production usage by building clear operational tooling, debugging paths, and safe rollout mechanisms.

What We Look For
  • Strong software engineering and infrastructure fundamentals, with experience owning production or near-production systems.

  • Practical experience with PyTorch and ML training workflows, including job orchestration, compute environments, artifact management, and deployment automation.

  • Solid understanding of heterogeneous computing and high-performance computing, especially for ML training or serving workloads.

  • Good understanding of model lifecycle concerns: data, configs, checkpoints, artifacts, reproducibility, rollout, rollback, and observability.

  • Ability to build reliable platform abstractions without hiding the important details ML practitioners need to control.

  • Clear technical and product sense: you can prioritize platform work that unlocks real training or deployment velocity.

  • High standards for engineering quality, including tests, documentation, debugging tools, and maintainable system design.

Tech Stack You May Work With
  • Python

  • Py Torch

  • Heterogeneous computing and high-performance computing

  • ML training pipelines, job orchestration, compute scheduling, containers, and deployment automation

  • Model artifacts, metadata, storage, experiment tracking, and configuration systems

  • Model serving, inference deployment, APIs, queues, and observability tools

  • Docker, CI/CD, cloud infrastructure, GPUs, and internal platform tooling

Bonus Points
  • Experience building training deployment systems, model release workflows, or ML platform tooling for research and production teams.

  • Experience with inference deployment, model serving, online/offline evaluation, performance tuning, or rollout safety.

  • Experience with distributed training, GPU infrastructure, workload scheduling, artifact/version management, or reproducibility tooling.

  • Understanding of CUDA, GPU architecture, or low-level performance optimization.

  • Experience with open-source inference and serving frameworks such as vLLM, TensorRT, Triton, or similar systems.

  • Experience migrating ad hoc notebooks, scripts, or manual ML processes into reliable platform workflows.

This Role May Not Be a Fit If
  • You mainly want to train models personally and do not enjoy building infrastructure for others to use.

  • You are comfortable with manual ML workflows and do not care about reproducibility, deployment, or operational quality.

  • You prefer narrow implementation tasks and do not want to reason about system boundaries, platform UX, or long-term maintenance.

  • You over-abstract ML workflows without understanding where researchers and engineers need control and visibility.

Why This Role Matters

ML teams move faster when training, deployment, and production usage are supported by reliable infrastructure instead of scattered scripts and manual processes. This role will directly shape how models move from experimentation to production, how safely they are deployed, and how efficiently the team can iterate. For the right person, it is a high-ownership platform role with deep impact on both engineering quality and ML velocity.

What

We Would Like to See When You Apply
  • ML infrastructure, training platforms, deployment systems, or model serving systems you have owned.

  • Examples of how you improved training reliability, deployment velocity, reproducibility, observability, or operational safety.

  • Cases where you turned messy ML workflows into maintainable tools, services, or platform abstractions.

  • Examples that show your technical judgment, communication, and ability to work across research and engineering needs.

#J-18808-Ljbffr
Position Requirements
10+ Years work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary