×
Register Here to Apply for Jobs or Post Jobs. X

Software Engineer L5 - AI Observability & Agent Evaluation

Job in Los Gatos, Santa Clara County, California, 95032, USA
Listing for: Netflix
Full Time position
Listed on 2026-10-05
Job specializations:
  • Software Development
    AI Engineer (Applied/Software), Machine Learning/ ML Engineer, AI Reliability/ Performance Engineer, Cloud Engineer - Software
Salary/Wage Range or Industry Benchmark: 388000 USD Yearly USD 388000.00 YEAR
Job Description & How to Apply Below

At Netflix, our mission is to entertain the world. Together, we are writing the next episode - pushing the boundaries of storytelling, global fandom and making the unimaginable a reality. We are a dream team obsessed with the uncomfortable excitement of discovering what happens when you merge creativity, intuition and cutting-edge technology. Come be a part of what’s next.

AI and ML power innovation in all areas of the business, including helping members choose the right title for them through personalization, better understanding our audience and our content slate, creating high-quality subtitles, dubbings, images, trailers, and other assets, optimizing our payment processing, and much more. AI Platform (AIP) organization builds highly scalable, differentiated AI infrastructure to maximize the business impact of all AI/ML practitioners at Netflix, which is key to accelerating this innovation.

The

Opportunity

The AI Observability team makes AI, ML, and Agentic systems transparent, reliable, and production-ready  build end-to-end observability for ML and GenAI workloads, capturing model inputs, features, predictions, outcomes, and behavior across online and batch systems. Our platform enables teams to monitor model performance, data quality, drift, latency, and failures, turning the ML system from a black box into an explainable, debuggable system.

We provide developer-friendly libraries, dashboards, and alerts so teams can debug issues, respond to incidents, and ship AI-powered products with confidence.

In This Role, You Will
  • Build the observability framework and platform capabilities that give ML and GenAI systems metrics, logs, and distributed traces across online inference, batch scoring, feature pipelines, and agent orchestration, so teams can instrument their systems consistently.
  • Build the primitives that let teams monitor model performance (accuracy, calibration, error rates), data quality, drift, and degradation on their own systems, rather than monitoring individual models yourself.
  • Build and extend evaluation frameworks for LLM and agentic systems that support response quality, grounding and hallucination, task success, tool-use and trajectory correctness, and LLM-as-a-judge and human-in-the-loop scoring, giving teams reusable building blocks to define and run their own evals.
  • Lead build-vs-buy evaluations for observability and eval tooling, and own the SDKs, connectors, and APIs that integrate vendor platforms into a consistent, well-supported interface, so ML and product teams can onboard models and agents with minimal friction.
  • Build reusable libraries, SDKs, and templates that make observability and evaluation the default for new systems ("observability-by-default"), lowering the barrier for teams to instrument and evaluate their work.
  • Provide the dashboarding, alerting, and SLO/SLI building blocks (plus sensible out-of-the-box templates) that teams use to track model performance, latency, cost, and reliability.
To Succeed In This Role, You Will Need
  • Experience in software, AI/ML, or platform engineering, with hands-on time in production observability, monitoring, or ML/LLM evaluation
  • Proven track record of designing standards, libraries, SDKs, or frameworks that other teams build on and adopt, reflecting a platform and enablement mindset rather than shipping features for a single use case.
  • Strong coding skills in Python and at least one of Java, Go, or Scala, with experience building production services
  • Practical experience operating ML models in production (online serving and/or batch), including drift, model performance, and data quality.
  • Hands-on experience with modern observability stacks (e.g., Prometheus/Grafana, Datadog, Open Telemetry, ELK/Open Search, Jaeger/Tempo, or similar).
  • Solid understanding of distributed systems, microservices, and at least one major cloud platform (AWS, GCP, or Azure).
  • Ability to work cross-functionally with ML, data, infra, and product teams, and to communicate clearly about system behavior, quality, and risk.
  • Experience with vendor integration and VPC deployment.
  • AI-Native Engineering Mindset who uses AI tools as a core part of their own workflow to accelerate design, development, testing, and code review.
Nice To Have
  • Hands-on experience with ML/LLM observability and evaluation tools (e.g., Arize, Braintrust, Lang Fuse, Weights & Biases, Galileo, Vertex AI Model Monitoring, Sage Maker Model Monitor).
  • Experience building or shipping LLM/GenAI applications and evaluating them:…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary