×
Register Here to Apply for Jobs or Post Jobs. X

Senior Machine Learning Systems Engineer; Frameworks & Tooling

Remote / Online - Candidates ideally in
Greater London, London, Greater London, W1B, England, UK
Listing for: Cohere
Remote/Work from Home position
Listed on 2026-09-10
Job specializations:
  • Software Development
    DevOps, Software Engineer, Cloud Engineer - Software
Salary/Wage Range or Industry Benchmark: 90000 - 150000 GBP Yearly GBP 90000.00 150000.00 YEAR
Job Description & How to Apply Below
Position: Senior Machine Learning Systems Engineer (Frameworks & Tooling)
Location: Greater London

  • We're looking for a senior engineer to help build, maintain and evolve the training framework that powers our frontier-scale language models. This role sits at the intersection of large-scale training, distributed systems, and HPC infrastructure
  • You will design and maintain the core components that enable fast, reliable, and scalable model training - and build the tooling that connects research ideas to thousands of GPUs
  • If you enjoy working across the full stack of ML systems, this role gives you the opportunity and autonomy to have massive impact
  • Build and own the training framework responsible for large-scale LLM training
  • Design distributed training abstractions (data/tensor/pipeline parallelism, FSDP/ZeRO strategies, memory management, checkpointing)
  • Improve training throughput and stability on multi-node clusters (e.g., GB200/300, AMD, H200/100)
  • Develop and maintain tooling for monitoring, logging, debugging, and developer ergonomics
  • Collaborate closely with infra teams to ensure our cluster, container environments, and hardware configurations support high-performance training
  • Investigate and resolve performance bottlenecks across the ML systems stack
  • Build robust systems that ensure reproducible, debuggable, large-scale runs
  • You'll work on some of the most challenging and consequential ML systems problems today
  • You'll collaborate with a world-class team working fast and at scale
  • You'll have end-to-end ownership over critical components of the training stack
  • You'll shape the next generation of infrastructure for frontier-scale models
  • You'll build tools and systems that directly accelerate research and model quality
  • Sample Projects:
  • Build a high-performance data loading and caching pipeline
  • Implement performance profiling across the ML systems stack
  • Develop internal metrics and monitoring for training runs
  • Build reproducibility and regression testing infrastructure
  • Develop a performant fault-tolerant distributed checkpointing system
Benefits
  • Six weeks' paid vacation
  • Equity / stock options
  • RRSP, 401(k), and Pension Scheme contributions
  • Coverage for 100% of your insurance premiums across health, dental, vision, and travel
  • Additional coverage for accessing mental health providers/services
  • Six months of fully paid parental leave, including adoption and surrogacy
  • Financial support for egg freezing and IVF in Canada and the UK
  • A monthly fitness and wellness allowance
  • Globally dispersed company that supports a remote work culture
  • A $2,000 annual education benefit for professional development
  • A weekly stipend for meals when working remotely and catered lunch when working from one of our global offices
  • A monthly arts and culture allowance
  • A monthly quality time allowance
  • A track record of building tools that increase developer velocity for ML teams

    Experience with multi-node cluster orchestration (Slurm, Ray, Kubernetes, or similar)
    Bonus: paper at top-tier venues (such as NeurIPS, ICML, ICLR, AIStats, MLSys, JAX, AAAI, Nature, COLING, ACL, EMNLP)

    Experience with training LLMs or other large transformer architectures

    Deep familiarity with JAX internals, distributed training libraries, or custom kernels/fused ops Strong engineering experience in large-scale distributed training or HPC systems

    Comfort debugging performance issues across CUDA/NCCL, networking, IO, and data pipelines

    Excellent judgment around trade-offs: performance vs complexity, research velocity vs maintainability

    Experience with data pipeline optimization, sharded datasets, or caching strategies

    Contributions to ML frameworks (PyTorch, JAX, Deep Speed, Megatron, xFormers, etc.)Experience working with containerized environments (Docker, Singularity/Apptainer)
    Background in performance engineering, profiling, or low-level systems

    If some of the above doesn't line up perfectly with your experience, we still encourage you to apply!

    Strong collaboration skills - you'll work closely with infra, research, and deployment teams

    Familiarity with evaluation and serving frameworks (vLLM, TensorRT-LLM, custom KV caches)
#J-18808-Ljbffr
Position Requirements
10+ Years work experience
Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary