More jobs:
Engineering Manager, Agentic GenAI Platform
Job in
Santa Clara, Santa Clara County, California, 95053, USA
Listed on 2026-07-19
Listing for:
NVIDIA Corporation
Full Time
position Listed on 2026-07-19
Job specializations:
-
Software Development
AI Engineer (Applied/Software), Backend Developer, DevOps, AI Reliability/ Performance Engineer
Job Description & How to Apply Below
NVIDIA is seeking an Engineering Manager to lead the development of an agentic platform for observing, debugging, and optimizing GenAI models deployed at scale.
Responsibilities- Lead, mentor, and grow a team building an agentic platform for monitoring and improving large-scale LLMs and VLMs in production.
- Build systems that collect, correlate, and analyze telemetry across inference servers, GPUs, schedulers, model runtimes, and customer-facing APIs.
- Develop agentic workflows that help engineers identify root causes, explain regressions, and recommend performance optimizations.
- Collaborate with internal customers and business units to align priorities and deliver production‑grade platform capabilities.
- BSc, MS, or PhD in Computer Science, Computer Engineering, or equivalent experience.
- 8+ years of relevant software engineering experience, including 3+ years in engineering management or technical leadership.
- Experience leading software engineering teams building large‑scale distributed systems, observability platforms, ML infrastructure, or production AI systems.
- Strong understanding of LLM/VLM inference systems, deployment patterns, and production performance challenges.
- Experience with logs, metrics, traces, profiling, alerting, dashboards, or incident/debugging workflows.
- Strong programming, debugging, performance analysis, and test design skills.
- Ability to work across organizations and align technical priorities with product and business goals.
- Excellent communication and collaboration skills.
- Background in GPU performance analysis, distributed inference, model serving optimization, or reliability engineering.
- Experience building observability or telemetry platforms for AI, ML, cloud, or distributed infrastructure.
- Experience with Open Telemetry, Prometheus, Grafana, Jaeger, Click House, Elastic, or similar observability tools.
- Experience building agentic systems that reason over logs, traces, performance data, incidents, or operational workflows.
- Hands‑on experience with production GenAI serving systems and metrics such as TTFT, TPOT, throughput, queueing delay, GPU utilization, KV cache pressure, error rates, and cost per token.
- Base salary range: $224,000
USD–$356,500
USD for Level3; $272,000
USD–$431,250
USD for Level
4. - Equity and benefits.
NVIDIA is committed to fostering an inclusive work environment and is proud to be an equal‑opportunity employer. NVIDIA does not discriminate on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status, or any other characteristic protected by law.
#J-18808-LjbffrTo View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×