Principal Machine Learning Engineer
Listed on 2026-07-27
-
Software Development
Machine Learning/ ML Engineer, AI Engineer (Applied/Software)
Remote (United States) or San Francisco, California, USA (Hybrid)
An exciting opportunity has arisen to join a rapidly growing AI company building the next generation of intelligent personal productivity software. The organisation is developing an AI-native assistant capable of understanding context, orchestrating complex workflows, interacting with external tools, and completing real-world tasks with minimal user input. Backed by an ambitious vision to transform how billions of people work, the company combines cutting-edge large language models with world-class engineering to deliver highly reliable AI systems at scale.
As a Principal Machine Learning Engineer, you will play a pivotal role in defining the architecture of the company's core machine learning platform, working alongside exceptional engineers and researchers. This Principal Machine Learning Engineer position offers deep ownership across training, inference, evaluation and deployment, making it an outstanding opportunity for a Principal Machine Learning Engineer who thrives on solving large-scale technical challenges.
Key Responsibilities- Architect and evolve scalable machine learning systems spanning training, inference, evaluation and deployment.
- Design and optimise distributed GPU training pipelines for large-scale foundation models.
- Build high-performance inference platforms that balance latency, throughput, reliability and infrastructure cost.
- Develop robust evaluation and data pipelines to improve model quality, safety and production performance.
- Collaborate closely with product, backend and infrastructure teams to integrate AI capabilities into customer-facing applications.
- Drive continuous improvements across production systems through monitoring, optimisation and rapid iteration.
- Strong industry experience with deep learning and transformer-based architectures.
- Hands-on experience training, fine-tuning or deploying large-scale machine learning models into production.
- Expert knowledge of PyTorch, JAX, or similar modern machine learning frameworks.
- Experience with distributed training and inference technologies such as Deep Speed, FSDP, Megatron, ZeRO, or Ray.
- Strong software engineering skills with experience building production-grade, maintainable systems.
- Experience optimising GPU workloads through mixed precision, quantisation and memory-efficient techniques.
- Ability to design and deliver complex ML systems from concept through production in fast-paced environments.
- Experience with LLM inference frameworks including vLLM, TensorRT-LLM or Faster Transformer.
- Experience with RLHF techniques including PPO, DPO or ORPO.
- Background in CUDA, GPU kernels, scientific computing or compiler optimisation.
- Experience with multimodal models, diffusion models or large-scale synthetic data generation.
- Knowledge of distributed data processing technologies such as Apache Arrow, Spark or Ray.
- Contributions to open-source machine learning or systems software.
- Join a highly talented, engineering-led team building an AI product with global ambitions.
- Work on technically challenging problems across large-scale distributed AI infrastructure.
- Take ownership of core machine learning architecture with significant technical autonomy.
- Influence the direction of a next-generation AI platform from an early stage.
- Collaborate with experienced researchers and engineers in a high-performance, low-ego environment.
- Flexible working with the option to work remotely across the United States or from San Francisco.
This role would suit a Principal, Staff or Senior Staff Machine Learning Engineer, AI Infrastructure Engineer, LLM Platform Engineer, or AI Systems Engineer with extensive experience building production machine learning platforms rather than purely conducting research, those who are developing large-scale AI infrastructure will be particularly well aligned.
By applying to this role you understand that we may collect your personal data and store and process it on our systems. For more information please see our Privacy Notice ().
#J-18808-Ljbffr(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).