×
Register Here to Apply for Jobs or Post Jobs. X

Senior research Scientist - Machine Learning Systems & Efficiency Engineer

Job in Seattle, King County, Washington, 98127, USA
Listing for: Adobe
Full Time position
Listed on 2026-07-02
Job specializations:
  • Software Development
    AI Engineer (Applied/Software), Machine Learning/ ML Engineer, AI Reliability/ Performance Engineer
Salary/Wage Range or Industry Benchmark: 168600 - 244200 USD Yearly USD 168600.00 244200.00 YEAR
Job Description & How to Apply Below

Overview

Photoshop ART is seeking a Senior researcher - Machine Learning Systems & Efficiency Engineer to join our R&D team focused on delivering practical, production-ready improvements in inference performance, latency, and cost efficiency across image editing applications. This role sits at the intersection of model architecture, systems, inference runtimes, and services, with a clear mandate: deliver high‑quality ML systems at substantially lower cost and higher efficiency.

Individuals in this role are expected to have deep expertise in AI, ML systems, and computer vision. Strong preference will be given to candidates with experience in distributed inference, multimodal model profiling, and performance optimization. You will work closely with research, product, and infrastructure teams to influence model design decisions, improve GPU utilization, and build scalable, cost‑aware ML systems deployed in production.

This is a hands‑on, high‑leverage role where a single engineer can drive outsized impact, potentially saving millions of dollars in compute costs. The ideal candidate will have a strong interest in developing practical innovations that advance Adobe products.

Responsibilities
  • Inference & Serving Optimization:
    Design and optimize high‑throughput, low‑latency inference systems. Optimize model architectures to improve deployment and runtime efficiency using techniques such as distillation, pruning, quantization, and Mixture‑of‑Experts (MoE). Implement advanced serving strategies including batching, caching (KV, semantic, embedding), quantization (FP8/INT8), and distributed inference strategies including data, tensor, pipeline, expert, and hybrid parallelism, with a focus on balancing computation and communication efficiency. Explore training or fine‑tuning approaches when they directly lead to more efficient inference, simpler deployment, or improved runtime performance.
  • Kernel Development & System Acceleration:
    Write and maintain high‑performance GPU kernels using Triton or CUDA to accelerate custom model layers and critical workloads. Improve GPU utilization through kernel fusion, asynchronous pipelines, and optimized scheduling strategies.
  • Performance Profiling & System Optimization:
    Conduct deep performance analysis using tools such as PyTorch Profiler and NVIDIA Nsight to identify bottlenecks in compute, memory, and communication. Optimize end‑to‑end system performance across inference workloads.
  • Distributed Systems & Infrastructure

    Collaboration:

    Partner with infrastructure teams to design scalable and reliable distributed serving systems across heterogeneous hardware environments (e.g., A100, H100, B200, CPU). Contribute to resource scheduling, GPU pooling, and elastic workload management.
  • Cost‑Aware ML Engineering:
    Establish and track efficiency metrics such as cost per million inferences. Build benchmarking frameworks and dashboards to guide tradeoffs among quality, latency, and compute cost, enabling data‑driven system and product decisions.
  • Technical Leadership & Best Practices:
    Serve as a trusted technical advisor to research and product teams on efficiency tradeoffs. Define best practices for scalable and cost‑efficient ML development and mentor engineers on performance‑oriented systems design.
Qualifications
  • Education:

    Master’s or PhD in Computer Science, Electrical Engineering, or related field focusing on machine learning systems, distributed systems, or high‑performance computing.
  • Distributed Inference & Serving Expertise:
    Hands‑on experience implementing and scaling large‑scale inference or serving workloads using distributed frameworks and runtime systems (e.g., Triton, vLLM, SGLang, xDiT, or similar). Experience applying inference compilation and optimization tools (e.g., TensorRT, ONNX Runtime, AOTi), including techniques such as operator fusion and graph‑level optimization, with a strong understanding of system‑level performance tradeoffs.
  • GPU & Performance Engineering

    Skills:

    Strong understanding of GPU architecture (e.g., memory hierarchy, compute throughput, communication bandwidth) and practical experience diagnosing performance bottlenecks across compute, memory, and I/O…
Position Requirements
10+ Years work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary