Systems ML Engineer; Member Technical Staff
Listed on 2026-07-26
-
Software Development
Machine Learning/ ML Engineer, AI Engineer (Applied/Software)
Member of the Technical Staff - Systems ML Engineer About Transfyr
Transfyr is building physical AI for science.
Why is it that a professional athlete has dramatically more information about every play they make than a scientist has about the cause of any experimental failure? At Transfyr, we are building the infrastructure to make real-world scientific work legible, transferable, and reproducible.
Modern science is capable of extraordinary outcomes, but much of the most important insights never become explicit: how experiments are actually executed, protocols drift, how experts make gametime decisions on the fly, why experiments fail on Tuesdays. This tacit knowledge is rarely captured, making it difficult to reliably reproduce results, much less hand off protocols to new team members or collaborators.
We believe our systematic failure to capture tacit knowledge is holding back the entire industry.
We're building systems that operate directly in real laboratory environments to elucidate, capture, and interpret this missing information. Our platform records and analyzes multimodal data about how scientific work is performed and turns it into durable, operational knowledge. In doing so, we are also building the world's largest commercial dataset on real-world scientific execution.
This foundation is critical not only for driving elite human performance today, but for enabling meaningful automation tomorrow. Physical AI systems cannot learn from outcomes alone; they require rich, grounded records of how work is actually done in the real world.
Want to learn more? You can read some of our writings here.
The RoleSystems ML Engineers at Transfyr ensure our models train and run efficiently across their full lifecycle – from large-scale training through production inference. You will own performance optimization across our ML stack, working embedded with the research team to make training faster and more efficient, and with our perception and production systems to get models running well at inference, across both cloud and edge infrastructure.
We are looking for engineers who combine deep understanding of ML systems with hands‑on performance engineering skill. The ideal Systems ML Engineer has experience profiling and optimizing large-scale models for both training and inference, is comfortable writing custom GPU kernels when off-the-shelf ops aren't fast enough, and can manage cloud and edge infrastructure that holds up in real lab environments.
This role spans deep ML and performance engineering and cloud/Dev Ops responsibilities. You will profile and optimize training and inference workloads, manage cloud and edge infrastructure, optimize cloud spend, ensure security and compliance, and work closely with the perception and research teams to squeeze more performance out of every model we ship.
We're building a team, and we have needs across levels, from hands‑on builders early in their careers to senior engineers who enjoy shaping training infrastructure and technical direction.
This role is in‑person in Cambridge, MA (other locations may open in the future, feel free to reach out even if Boston is not currently an option for you).
What you’ll accomplish with us:Profile and Optimize Performance: Use profiling tools (e.g., Nsight, PyTorch Profiler) to identify bottlenecks in data loading, gradient computation, and communication, and implement optimizations like kernel fusion, sharding, and tiling to improve step time.
Optimize Distributed Training: Improve the efficiency of distributed training pipelines using frameworks like PyTorch Distributed, working closely with the research team on training performance.
Develop Custom Kernels: Design and maintain high-performance GPU kernels in Triton or CUDA for performance-critical ML workloads.
Build Data Pipelines: Design and optimize data loading pipelines that maximize training throughput, and inference pipelines that reliably serve models on real-world, multimodal lab data.
Handle Cloud and Edge: Manage deployment across both cloud infrastructure and edge devices running in active lab environments, where compute and connectivity are more limited.
Keep Production…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).