×
Register Here to Apply for Jobs or Post Jobs. X

AIOps & Observability Lead

Job in Woodbridge Township, Middlesex County, New Jersey, 07064, USA
Listing for: Bessemer Trust
Full Time position
Listed on 2026-10-05
Job specializations:
  • IT/Tech
    Systems Engineer, Security Management & Operations, Cybersecurity, SRE/Site Reliability
Salary/Wage Range or Industry Benchmark: 140000 - 180000 USD Yearly USD 140000.00 180000.00 YEAR
Job Description & How to Apply Below

100 Woodbridge Center Drive, Woodbridge , New Jersey 07095 , United States

Job Description Role Summary

We are building a new Operations function and need a hands‑on, technical leader to set the direction for our AIOps and observability strategy. This is a player‑coach role, weighted toward player: someone who works directly in the tools, not just the roadmap.

Our new Splunk Observability implementation is a foundation of a broader automated operational intelligence ecosystem. We will lean on managed-service partners to act as first-line "eyes on glass" for the alerts and events it generates — but this role owns the strategy, architecture, and quality behind what those partners are watching. The goal is not just detection; it's ensuring issues can be discovered, triaged, and resolved through tooling and automation, reducing L3 engineer escalations wherever possible.

The role could expand over time into desktop/endpoint observability, including a Nexthink implementation.

Key Responsibilities
  • Define and own the observability and event management strategy, stack, and processes, including how major incidents are detected, triaged, and resolved.
  • Lead the rollout of Splunk Observability across our application and infrastructure portfolio, driving onboarding, instrumentation, and coverage expansion.
  • Lead implementation of event management tooling — such as, Big Panda, Splunk ITSI, or Pager Duty — to correlate, prioritize, and route alerts from Splunk Observability.
  • Coach and enable the Automation team to become proficient in the observability and event intelligence platform, including alert logic, dashboard creation, event enrichment, correlation workflows, triggered automations, MCP‑style integrations, and runbook automation.
  • Partner with the Infrastructure Architect to build a roadmap toward first‑class, enterprise‑grade observability.
  • Develop and execute an AIOps strategy that turns telemetry into HITL‑automated, self‑healing operations — evaluating and building on tooling in Splunk, AWS, or elsewhere as appropriate.
  • Stand up and mature an AIOps/SRE capability: automated remediation, intelligent alert correlation, and reduced manual triage.
  • Direct managed‑service partners providing "eyes on glass" monitoring — defining what they watch, how they escalations, and holding them accountable to SLAs.
  • Design escalation paths so that issues resolvable via tooling are handled there first, protecting L3 engineering time for what truly needs it.
  • Establish standards for event classification, severity, ownership, suppression, and closure across the environment.
  • Assess and help scope the expansion into desktop/endpoint observability including a Nexthink implementation.
  • Ensure dashboards, alerts, and event workflows are built around operational decisions, not just technical visibility.
  • Align observability, event management, incident, problem, change, knowledge, and escalation workflows with ITIL/ITSM practices and Service Now processes.
Required Qualifications
  • Strong, hands‑on experience with observability, monitoring, and event management platforms (e.g., Splunk, Splunk ITSI, Big Panda, Pager Duty).
  • Proven experience designing and implementing AIOps or event‑intelligence capabilities — alert correlation, noise reduction, and automation.
  • Experience defining incident management processes, including major incident response and escalation design.
  • Comfortable operating as both strategic lead and hands‑on technical contributor — able to shape direction and work directly in the tools.
  • Scripting or automation experience (e.g., Python, Power Shell, REST APIs) to support event enrichment and automated remediation.
  • Strong communication skills, with the ability to align infrastructure, application, and operations teams around a…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary