AVP, Monitoring Engineer
Listed on 2026-08-18
-
IT/Tech
SRE/Site Reliability, AWS
Where Ambition Meets Innovation Build a career that matches all your initiative with an impressive dose of innovation. From cutting‑edge resources and a collaborative environment to the freedom to make an impact and more, you’ll find the ingredients you need at LPL Financial to shape your success while helping clients pursue their financial goals. Observability at LPL only works if a senior engineer is willing to live in the dashboards, the noisy alerts, and the post‑incident reviews.
As AVP, Monitoring, you'll be that engineer — hands‑on across Cloud Watch and X‑Ray, Dynatrace, Grafana / Prometheus, Open Search (with the legacy ELK stack where it still exists), and Service Now — raising the bar on SLOs and alert quality without owning a team. If you'd rather pair on a noisy alert chain than collect status, this is your seat.
Job Overview:
As the AVP, Monitoring Engineer, you are a hands‑on senior cloud observability engineer in the Monitoring pod within the Foundations team in LPL's Cloud Center of Excellence (CCOE). You partner with the VP, Monitoring and with every other CCOE team and pod — the peer Foundations pods (Security & Governance, Fin Ops, Functional Design Engineering & Strategy, Network Engineering), plus the Platforms, Containers, Support, and Delivery teams — to raise the quality of LPL's observability across the multi‑account landing zone.
The stack spans AWS‑native services (Cloud Watch, X‑Ray, Open Search), Grafana / Prometheus (including Amazon Managed Prometheus and Managed Grafana where appropriate), Dynatrace, the legacy ELK stack where it still exists, and Service Now for ITSM and incident ticketing. You partner closely with the Support team on incident response and with the Functional Design Engineering & Strategy pod on the observability paved road for application teams.
LPL is an AWS‑first CCOE: a multi‑account landing zone with 100+ private reusable Terraform modules that enable 60+ AWS services, all delivered through Terraform Cloud and Git Hub Actions. You spend the majority of your time hands‑on in dashboards, alerts, Terraform, and incident response across LPL's US offices and India Global Capability Center (GCC) — your impact comes from technical depth, code review, and peer mentorship rather than positional authority.
- Hands‑on author and curate the CCOE observability paved road: opinionated Terraform modules, Helm charts, and reference dashboards for Cloud Watch, Grafana, Dynatrace, and Open Search — so every workload starts with credible observability
- Raise the bar on SLOs, golden signals, and alert quality across the multi‑account landing zone — kill noisy alerts, surface missing ones, and partner with the Support team on what should and should not page a human
- Operate and continuously improve the CCOE observability stack:
Cloud Watch and X‑Ray, Grafana / Prometheus (including Amazon Managed Prometheus and Managed Grafana), Dynatrace, Open Search (and the legacy ELK stack where it still exists), and Service Now for ITSM and incident ticketing - Partner with the Support team on incident response: alert routing into Service Now, runbook execution, major incident participation, and the feedback loop from incident review back into durable monitoring improvement
- Mentor Engineer 2 and Senior Engineers in the Monitoring pod through code review, design partnership, and pairing on noisy alert chains — uplift the pod without direct reports
- Embed agentic AI capabilities into the team’s engineering practice (e.g., Cursor, Claude Code, Bedrock, MCP servers, agentic IaC and review workflows) and into the platform’s self‑service experience for internal customers
- Use agentic AI capabilities in day‑to‑day observability work: AI‑assisted alert triage and noise reduction, dashboard authoring from natural‑language intent, on‑call copilots that summarize signal during incidents, and MCP‑backed agents over telemetry
- Operate as a hands‑on senior cloud engineer: spend the majority of your time in Terraform code, security tooling configuration, vulnerability remediation, design reviews, peer reviews, and incident response — hands‑on engineering is the primary leverage…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).