Senior Director, Inference Products and Optimizations
Job in
Seattle, King County, Washington, 98113, USA
Listed on 2026-08-05
Listing for:
DigitalOcean
Full Time
position Listed on 2026-08-05
Job specializations:
-
Software Development
AI Engineer (Applied/Software), Machine Learning/ ML Engineer
Job Description & How to Apply Below
We value winning together-while learning, having fun, and making a profound difference for the dreamers and builders in the world.
Our Inference Engine organization is seeking an experienced Senior Director of Engineering to lead a high-performing engineering team building and scaling our Large Language Model (LLM) inference products across the control plane, model optimization, and model architecture layers. This organization sits at the heart of Digital Ocean's mission to bring our signature simplicity to optimized LLM inference.
In this role, you will own Digital Ocean's inference product suite - Serverless Inference, Dedicated Inference, Inference Router, Batch Inference, and Multimodal Inference - along with the model optimization and architecture stack that underpins them. You'll be responsible for delivering robust, cost-efficient systems that serve millions of users globally at scale and high performance.
What You'll Do:
* Team Leadership & Development:
Recruit, mentor, and coach engineers on the team, fostering a culture of ownership, technical excellence, and continuous improvement.
* Build Performant and Scalable Inference Products:
Work with Product teams to define and execute on the Product roadmap for all of Digital Ocean's Inference Products - including Serverless Inference, Dedicated Inference, Inference Router, Batch Inference and Multimodal Inference
* Inference Optimizations and Model Architecture:
Lead the design and evolution of our inference serving stack, driving deep technical strategy across vLLM, SGLang, and LLM-D to optimize throughput, latency, and GPU utilization hitect the model-serving and optimization layer - spanning quantization, KV-cache management, speculative decoding, and disaggregated serving - to deliver best-in-class performance-per-dollar across our LLM inference products.
* Cross-Functional Partnership:
Collaborate with Product Management, other engineering teams, and key stakeholders to align priorities, manage dependencies, and communicate progress and risks.
* Operational Health:
Ensure the production health, stability, and on-call rotation of all to maintain the customer SLAs.
* Champion Best Practices:
Institutionalize benchmarking frameworks, observability, and auto-tuning capabilities to guide system and infrastructure tuning efforts. Encourage contributions to open-source inference engines to advance our capabilities.
Indicators of a Good Fit:
* Experience:
10+ years of software engineering experience, with 6+ years in a technical leadership or management role, ideally within Inference Systems or AI/ML systems.
* Technical Depth:
Deep expertise in distributed systems design, modern AI/ML technologies, Kubernetes at scale, and LLM inference, and AI workload orchestration, scheduling, and resource management. Ability to engage in deep technical discussions with your team regarding highly scalable control plane design, inference engines (vLLM, SGLang), and model architectures. Strong understanding of cloud-native multi region architectures, microservices, and distributed systems fundamentals.
* Hardware-Aware Optimization:
Strategic knowledge of GPU architectures (NVIDIA and/or AMD), interconnects (like NVLink), and hardware topology and their direct impact on AI training and inference performance.
* Systems Engineering & Security:
Familiarity with concepts in container runtime internals, system isolation, and security contexts to manage risk in shared infrastructure.
* Observability and SLOs:
Expertise in defining, tracking, and operationalizing deep infrastructure and inference metrics (e.g., TTFT, TPOT) to drive performance improvements and meet service level objectives.
* Product Mindset:
Demonstrated ability to translate complex technical requirements into user-focused product features. Understanding of the balance between innovation and reliability.
* Communication:
Excellent communication skills, with the ability to explain technical decisions to non-technical stakeholders and align diverse teams around a shared vision.
* Ownership: A strong sense of ownership and a proactive drive to identify and resolve issues preventing your team from delivering value.
Compensation Range:
* $274,400 - $343,000
* This is a Hybrid role
JR:
#LI-Hybrid
Why You'll Like Working for Digital Ocean
* We innovate with purpose. You'll be a part of a cutting-edge technology company with an upward trajectory, who are proud to simplify cloud and AI so builders can spend more time creating software that changes the world. As a member of the team, you will be a Shark who thinks big, bold, and scrappy, like an owner with a bias for action and a powerful sense of…
Position Requirements
10+ Years
work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×