Sr. SDE, Edge AI ML Platform, Edge AI and Science
Job in
Vancouver, BC, Canada
Listed on 2026-08-11
Listing for:
Amazon
Full Time
position Listed on 2026-08-11
Job specializations:
-
Software Development
Machine Learning/ ML Engineer, AI Engineer (Applied/Software), Software Engineer, DevOps
Job Description & How to Apply Below
DESCRIPTION:
Amazon Devices (Lab
126) builds products and services that delight millions of customers globally. The Edge AI ML Platform and Infrastructure team is building the platform that enables Amazon teams to train, optimize, evaluate, and deploy generative AI models on devices and in the cloud.
Today, optimizing a large model for a new hardware target requires experts to connect model onboarding, distributed training, compression, evaluation, compilation, and deployment systems by hand. We are turning that work into a repeatable, self-service workflow. Our platform supports large language, vision, audio, multimodal, and mixture-of-experts models. It gives scientists and engineers the tools to move new optimization techniques from research code into reliable production workflows.
We are looking for a Senior Software Development Engineer to lead the architecture and delivery of core ML platform capabilities. You will solve problems across distributed training on multi-node GPU clusters, model onboarding, compression pipelines, evaluation, GPU performance, artifact management, CI/CD, observability, and operational reliability. You will work with applied scientists, ML engineers, GPU kernel engineers, compiler and runtime teams, hardware teams, and product teams to deliver systems for models with hundreds of billions of parameters.
This role combines hands-on software development with technical leadership. You will write and review code, define architecture, resolve ambiguous requirements, lead projects that span multiple engineers and teams, and raise the engineering bar for an evolving ML platform.
Key job responsibilities
- Lead the design and delivery of distributed ML platform services and libraries across model ingestion, optimization, training, evaluation, packaging, and deployment.
- Define stable APIs and architecture boundaries that allow scientists to add algorithms without coupling research code to training, infrastructure, or deployment implementations.
- Design distributed training capabilities across data, tensor, pipeline, and model parallelism for large language and multimodal models.
- Scale workflows on multi-node GPU clusters while improving training throughput, GPU utilization, memory efficiency, communication performance, failure recovery, and developer iteration time.
- Develop infrastructure that connects distributed training with distillation, quantization, pruning, and other model optimization techniques.
- Build evaluation and artifact workflows that measure model quality and system performance, then carry validated models through deployment on target hardware.
- Build automated validation, CI/CD, regression testing, observability, and release mechanisms for GPU-intensive ML workloads.
- Profile and optimize end-to-end system performance with applied scientists and GPU kernel engineers. Translate bottlenecks into durable platform improvements.
- Establish operational mechanisms, including metrics, alarms, runbooks, on-call practices, and root-cause correction for production platform services.
- Partner with model, compiler, runtime, hardware, security, and infrastructure teams to clarify requirements, manage technical dependencies, and deliver multi-team programs.
- Write technical designs, evaluate trade-offs, and build consensus when the customer need is clear but the technology strategy is not.
- Mentor engineers, improve code and design review practices, and help recruit and develop a strong engineering team in Vancouver.
A day in the life
You will move between architecture and implementation. Your work will include reviewing designs for model onboarding interfaces, investigating failures in distributed training runs, profiling GPU workloads with scientists, leading cross-team reviews of end-to-end deployment paths, simplifying platform abstractions, and improving the release and regression mechanisms used by multiple model teams.
You will use performance, reliability, and developer productivity…
Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
Search for further Jobs Here:
×