×
Register Here to Apply for Jobs or Post Jobs. X

Senior Engineer, Inference Data Plane

Job in Denver, Denver County, Colorado, 80202, USA
Listing for: DigitalOcean
Full Time position
Listed on 2026-07-26
Job specializations:
  • Software Development
    Cloud Engineer - Software, DevOps, AI Engineer (Applied/Software), Backend Developer
Salary/Wage Range or Industry Benchmark: 139200 - 174000 USD Yearly USD 139200.00 174000.00 YEAR
Job Description & How to Apply Below

Senior Engineer, Inference Data Plane

Dive in and do the best work of your career rney alongside a strong community of top talent who are relentless in their drive to build the simplest scalable cloud. If you have a growth mindset, naturally like to think big and bold, and are energized by the fast-paced environment of a true industry disruptor, you'll find your place here. We value winning together—while learning, having fun, and making a profound difference for the dreamers and builders in the world.

Digital Ocean is expanding its AI Infrastructure layer to support the next generation of AI-driven applications. We are seeking a Senior Engineer to join our AI Inference Data Plane team. In this role, you will be a key technical leader responsible for designing, developing, and delivering high-scale, resilient data plane services that power our "Inference as a Service" offering. You will work at the intersection of distributed systems and specialized AI hardware to ensure our customers can deploy and scale their models with industry-leading performance and reliability.

This is a hands-on role, requiring you to be able to develop high quality software while availing of all the productivity boosts granted by the latest AI coding agents.

What You'll Do:
  • Technical Leadership:
    Act as a technical leader on the team, driving the end-to-end design, development, and delivery of critical data plane components hosting large generative AI models.
  • System Design:
    Architect and refine system design proposals for our high-scale, multi-tenant AI inference cloud ecosystem, ensuring they meet rigorous availability and resiliency standards.
  • Performance Optimization:
    Implement and optimize distributed inference hosting using techniques like tensor/data parallelism, KV cache optimizations, and smart routing.
  • Collaboration:

    Work cross-functionally with Product Managers, customer-facing teams, and other engineering teams to align technical roadmaps with customer needs.
  • Distributed Serving at Scale:
    Build on Kubernetes-native distributed inference frameworks like llm-d (or alternatives such as NVIDIA Dynamo, Ray Serve, KServe) to deliver prefill/decode disaggregation, KV-cache-aware routing, tiered prefix caching, and wide expert parallelism for MoE models.
  • Flow Control & Load Balancing:
    Solve the distributed-systems problems unique to LLM serving — inference-aware load balancing on queue depth, cache locality, and predicted latency; flow control and fairness across tenants; autoscaling inference pools; and moving gigabytes of KV-cache between prefill and decode instances with negligible overhead.
  • Open Source Contributions:
    Contribute upstream to llm-d, vLLM, and the inference gateway ecosystem, and represent Digital Ocean in these communities.
  • Mentorship:
    Coach and mentor junior engineers, fostering a culture of technical excellence and continuous improvement.
  • Operational Excellence:
    Maintain and operate critical, high-scale services, utilizing observability tools and defining SLOs to ensure superior platform health.
What You'll Bring to Digital Ocean:
  • AI/ML Domain Knowledge:
    Hands-on experience hosting large language or multimodal models using inference engines like vLLM, SGLang, or TensorRT.
  • Inference Frameworks:
    Familiarity with distributed inference serving frameworks such as llm-d, NVIDIA Dynamo, or Ray Serve.
  • Inference Engine Depth:
    Hands-on experience with vLLM or alternatives (SGLang, TensorRT-LLM, TGI, Modular MAX), including internals like continuous batching, paged attention, and prefix caching.
  • Distributed Inference Fluency:
    Understanding of why cluster-scale serving is hard: KV-cache locality is partitioned across workers, naive round-robin routing destroys cache hit rates and tail latency, and disaggregated prefill/decode requires fast cross-pod KV transfer (e.g., NIXL).
  • Upstream Track Record:
    Merged contributions to vLLM, llm-d, SGLang, or similar projects strongly preferred.
  • Architecture Proficiency:
    Knowledge of common LLM architectures and optimization techniques (e.g., continuous batching, quantization).
  • Software Engineering:
    Expert-level proficiency in GoLang or Python and familiarity with gRPC.
  • Cloud Operations:
    Proven experience shipping customer-facing software products and running critical services in a high-scale environment similar to Digital Ocean.
  • Open Source Mindset:
    Experience integrating and building with open-source software.
Compensation Range:
  • $139,200 - $174,000

* This is a remote role

Why You'll Like Working for Digital Ocean:
  • We innovate with purpose. You'll be a part of a cutting-edge technology company with an upward trajectory, who are proud to simplify cloud and AI so builders can spend more time creating software that changes the world. As a member of the team, you will be a Shark who thinks big, bold, and scrappy, like an owner with a bias for action and a powerful sense of responsibility for customers, products, employees, and decisions.
  • We prioritize career development. At DO, you'll do the best work of your…
Position Requirements
10+ Years work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary