×
Register Here to Apply for Jobs or Post Jobs. X

Sr. Product Manager - Observability

Job in Pleasanton, Alameda County, California, 94566, USA
Listing for: Workday, Inc.
Full Time position
Listed on 2026-09-05
Job specializations:
  • IT/Tech
Salary/Wage Range or Industry Benchmark: 168000 - 252000 USD Yearly USD 168000.00 252000.00 YEAR
Job Description & How to Apply Below

Your work days are brighter here. We’re obsessed with making hard work pay off, for our people, our customers, and the world around us. As a Fortune 500 company and a leading AI platform for managing people, money, and agents, we’re shaping the future of work so teams can reach their potential and focus on what matters most. The minute you join, you’ll feel it.

Not just in the products we build, but in how we show up for each other. Our culture is rooted in integrity, empathy, and shared enthusiasm. We’re in this together, tackling big challenges with bold ideas and genuine care. We look for curious minds and courageous collaborators who bring sun-drenched optimism and drive. Whether you're building smarter solutions, supporting customers, or creating a space where everyone belongs, you’ll do meaningful work with Workmates who’ve got your back.

In return, we’ll give you the trust to take risks, the tools to grow, the skills to develop and the support of a company invested in you for the long haul. So, if you want to inspire a brighter work day for everyone, including yourself, you’ve found a match in Workday, and we hope to be a match for you too.

About

the Team

The Data Platform and Observability Engineering (DPOE) team is building Workday's next-generation, multi-petabyte scale Observability Platform. We own the libraries, distributed services, and infrastructure that power ingestion, storage, and query across the observability stack — Iceberg, Click House, Tempo, Mimir, Grafana, S3, Kafka, and Elasticsearch — serving traces, metrics, and logs for every workload at Workday.

How we help our customers

Every engineering, SRE, and product team at Workday depends on us to see inside their systems — from a single service's latency spike to a cross-service cascading failure impacting thousands of customers. By owning the full observability data lifecycle at petabyte scale, we give internal teams the speed and confidence to find and fix issues before they affect Workday's customers. Our roadmap directly shapes how the company detects, diagnoses, and eventually predicts operational issues at scale — turning observability from a reactive debugging tool into a proactive, AI-assisted safety net for every workload running on Workday's platform.

About

the Role

We are looking for a hands‑on, technical Senior Product Manager (P4) to own and drive our Observability strategy, with a strong emphasis on Distributed Tracing. You will define the product vision for how engineers understand, debug, and optimize complex distributed systems, with a particular focus on building AI-powered detection and triage capabilities that reduce time-to-detect and time-to-resolve production issues. This role requires someone comfortable diving deep into technical architecture discussions, reading code/traces, and partnering closely with engineering and applied ML teams to ship technically sound, high-impact products.

What

You’ll Do
  • Own the product vision, strategy, and roadmap for Observability, with a primary focus on Distributed Tracing capabilities (trace context propagation, sampling strategies, span analysis, service maps, latency/error analysis, etc.)
  • Define and drive the roadmap for AI-enabled anomaly detection, including specifying requirements for statistical and ML-based detection methods (e.g., time-series forecasting, seasonality-aware baselining, change-point detection, multivariate anomaly detection across correlated metrics/traces/logs)
  • Partner with ML engineering to define model requirements, evaluation metrics (precision/recall, false-positive rate, alert-to-incident ratio), and feedback loops for continuous model improvement
  • Define requirements for LLM-based root cause summarization and triage assistance — e.g., generating human-readable incident summaries from raw trace/log/metric data, ranking probable root causes, suggesting remediation runbooks based on historical incident patterns
  • Specify how confidence scores, explainability, and human-in-the-loop review are surfaced in the triage workflow so on‑call engineers can trust and act on AI-generated recommendations
  • Partner closely with engineering teams to define…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary