×
Register Here to Apply for Jobs or Post Jobs. X

Engineering Manager, Agentic GenAI Platform

Job in Santa Clara, Santa Clara County, California, 95053, USA
Listing for: NVIDIA Corporation
Full Time position
Listed on 2026-07-19
Job specializations:
  • Software Development
    AI Engineer (Applied/Software), Backend Developer, DevOps, AI Reliability/ Performance Engineer
Salary/Wage Range or Industry Benchmark: 224000 - 431250 USD Yearly USD 224000.00 431250.00 YEAR
Job Description & How to Apply Below

NVIDIA is seeking an Engineering Manager to lead the development of an agentic platform for observing, debugging, and optimizing GenAI models deployed at scale.

Responsibilities
  • Lead, mentor, and grow a team building an agentic platform for monitoring and improving large-scale LLMs and VLMs in production.
  • Build systems that collect, correlate, and analyze telemetry across inference servers, GPUs, schedulers, model runtimes, and customer-facing APIs.
  • Develop agentic workflows that help engineers identify root causes, explain regressions, and recommend performance optimizations.
  • Collaborate with internal customers and business units to align priorities and deliver production‑grade platform capabilities.
Qualifications
  • BSc, MS, or PhD in Computer Science, Computer Engineering, or equivalent experience.
  • 8+ years of relevant software engineering experience, including 3+ years in engineering management or technical leadership.
  • Experience leading software engineering teams building large‑scale distributed systems, observability platforms, ML infrastructure, or production AI systems.
  • Strong understanding of LLM/VLM inference systems, deployment patterns, and production performance challenges.
  • Experience with logs, metrics, traces, profiling, alerting, dashboards, or incident/debugging workflows.
  • Strong programming, debugging, performance analysis, and test design skills.
  • Ability to work across organizations and align technical priorities with product and business goals.
  • Excellent communication and collaboration skills.
Preferred Qualifications
  • Background in GPU performance analysis, distributed inference, model serving optimization, or reliability engineering.
  • Experience building observability or telemetry platforms for AI, ML, cloud, or distributed infrastructure.
  • Experience with Open Telemetry, Prometheus, Grafana, Jaeger, Click House, Elastic, or similar observability tools.
  • Experience building agentic systems that reason over logs, traces, performance data, incidents, or operational workflows.
  • Hands‑on experience with production GenAI serving systems and metrics such as TTFT, TPOT, throughput, queueing delay, GPU utilization, KV cache pressure, error rates, and cost per token.
Benefits
  • Base salary range: $224,000

    USD–$356,500

    USD for Level3; $272,000

    USD–$431,250

    USD for Level
    4.
  • Equity and benefits.

NVIDIA is committed to fostering an inclusive work environment and is proud to be an equal‑opportunity employer. NVIDIA does not discriminate on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status, or any other characteristic protected by law.

#J-18808-Ljbffr
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary