Member Technical Staff - Systems ML Engineer
Listed on 2026-09-01
-
Software Development
Machine Learning/ ML Engineer, AI Engineer (Applied/Software)
Member of the Technical Staff - Systems ML Engineer About Transfyr
Transfyr is building physical AI for science.
Why is it that a professional athlete has dramatically more information about every play they make than a scientist has about the cause of any experimental failure? Science has no film room, no instant replay. Instead, a protocol says what was meant to happen. A publication is a lossy record of what might have worked. But all the small decisions and invisible actions that determine whether an experiment succeeds, fails, or transfers to the next lab often disappear the moment the work is done or a scientist leaves.
That missing record is why it has been so hard to automate the physical work of science. It’s why training still remains dependent on scarce, one-to-one apprenticeship. It’s why tech transfer typically requires expensive troubleshooting and is one of the biggest causes of drug launch delays. It’s why scientists struggle to distinguish between biological noise and process variability.
We’re changing that. Transfyr builds physical AI systems that capture real scientific work and turn it into a high-fidelity, machine-readable record of execution and analysis of where process variability is impacting results. In doing so, we are also building the world’s largest commercial dataset on real-world scientific execution. The result is infrastructure that helps teams learn from failures, transfer hard-won know-how, train the next generation of scientists, and give models and robots the grounded data they need to be useful in the real world.
We’re tackling some of the hardest problems at the intersection of frontier science, perception, machine learning, and robotics and have significant traction. We’re backed by a $25M seed round, are collaborating with the largest frontier AI labs, and our advisors include Chris Ré (Stanford), David Baker (Nobel winning UW professor), Kevin Weil (fmr CPO at OpenAI), Steve Quake (Stanford biophysicist), Ken Frazier (fmr CEO of Merck), and Jakob Uszkoreit (CEO of Inceptive and author of “Attention Is All You Need”).
We’re unapologetically ambitious and pragmatic. If you want to work on the hardest problems in the most important industry on earth, join us.
The RoleSystems ML Engineers at Transfyr ensure our models train and run efficiently across their full lifecycle, from large-scale training through production inference. You will own performance optimization across our ML stack, working embedded with the research team to make training faster and more efficient, and with our perception and production systems to get models running well at inference, across both cloud and edge infrastructure.
We are looking for engineers who combine deep understanding of ML systems with hands‑on performance engineering skill. The ideal Systems ML Engineer has experience profiling and optimizing large-scale models for both training and inference, is comfortable writing custom GPU kernels when off‑the‑shelf ops aren't fast enough, and can manage cloud and edge infrastructure that holds up in real lab environments.
This role spans deep ML and performance engineering and cloud/Dev Ops responsibilities. You will profile and optimize training and inference workloads, manage cloud and edge infrastructure, optimize cloud spend, ensure security and compliance, and work closely with the perception and research teams to squeeze more performance out of every model we ship.
We're building a team, and we have needs across levels, from hands‑on builders early in their careers to senior engineers who enjoy shaping training infrastructure and technical direction.
This role is in‑person in Cambridge, MA.
What you'll accomplish with us:Profile and Optimize Performance: Use profiling tools (e.g., Nsight, PyTorch Profiler) to identify bottlenecks in data loading, gradient computation, and communication, and implement optimizations like kernel fusion, sharding, and tiling to improve step time.
Optimize Distributed Training: Improve the efficiency of distributed training pipelines using frameworks like PyTorch Distributed, working closely with the research team on training performance.
Develop…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).