More jobs:
Senior Problem, Incident and Event Management Engineer
Remote / Online - Candidates ideally in
Nashville, Davidson County, Tennessee, 37230, USA
Listed on 2026-07-30
Nashville, Davidson County, Tennessee, 37230, USA
Listing for:
Humana
Remote/Work from Home
position Listed on 2026-07-30
Job specializations:
-
IT/Tech
IT Support, Cybersecurity
Job Description & How to Apply Below
* The Senior Problem, Incident, and Event Management Engineer is responsible for advancing enterprise event correlation, observability, and incident detection capabilities to proactively identify and mitigate service disruptions before user impact. This role specializes in leveraging platforms such as Splunk, Dynatrace, and Service Now to aggregate, correlate, and operationalize data across multiple systems to drive timely escalation and resolution of critical incidents.
This position plays a key role in evolving from reactive incident response to data-driven, predictive operations, utilizing CMDB-driven context, criticality tiering, and advanced analytics to prioritize and escalate issues with significant business impact.
** Key Responsibilities*
* ** Event Correlation & Observability Engineering*
* + Design, configure, and continuously improve
** event correlation rules and alerting strategies
** across platforms such as
** Splunk ITSI and Dynatrace*
* + Integrate data from multiple monitoring, application, and infrastructure sources to create
** meaningful, actionable events*
* + Normalize and enrich event data using standardized fields and metadata to improve correlation accuracy and reduce noise
+ Drive reduction of false positives and duplicate alerts through correlation, aggregation, and suppression strategies
** Dashboarding & Data Visualization*
* + Develop and maintain
** operational and executive dashboards
** in Splunk and other reporting tools
+ Translate technical telemetry into
** clear, business-aligned insights** , highlighting service health, degradation, and emerging risks
+ Partner with command center, TOC, and incident teams to ensure dashboards support
** real-time decision making and escalation*
* ** Incident Detection & Escalation*
* + Leverage correlated event data and observability insights to
** trigger proactive incident identification
** prior to user-reported impact
+ Apply
** criticality tiering and CMDB data
** to assess business impact and drive proper prioritization and escalation paths
** Service Now Integration & ITSM Enablement*
* + Partner with Service Now stakeholders to improve workflows, reporting, and automation capabilities
** CMDB & Data-Driven Decisioning*
* + Leverage
** CMDB relationships and service mapping
** where available to enrich event data with application, infrastructure, and business context
+ Utilize service ownership, business criticality, and operational hours data to inform prioritization decisions
+ Partner with CMDB and service mapping teams to improve
** data quality and completeness*
* ** Trend Analysis & Continuous Improvement*
* + Analyze patterns across incidents, alerts, and events to identify systemic issues and opportunities for improvement
+ Partner with Problem Management to eliminate recurring issues through structural fixes
+ Drive improvements in monitoring coverage, alert quality, and detection speed
+ Contribute to a shift toward
** predictive, AIOps-driven operations*
* ** Use your skills to make an impact*
* ** Required Qualifications*
* + 3-5+ years of experience in
** Incident, Event, or Problem Management*
* + Hands-on experience with
** Splunk (preferably ITSI) and Dynatrace or similar observability platforms*
* + Experience building
** dashboards, reports, and analytics
** to support operational decision-making
+
Experience with
** Service Now ITSM** , including incident lifecycle management and reporting
+ Strong analytical skills with the ability to
** correlate data across multiple systems and platforms*
* + Experience working with
** event correlation, alerting strategies, or AIOps concepts*
* + Ability to assess business impact using
** priority models, criticality tiers, and service context*
* + Strong communication skills with the ability to translate technical findings into actionable insights
** Preferred Qualifications*
* +
Experience with
** CMDB, service mapping, or application dependency mapping*
* + Exposure to
** enterprise monitoring ecosystems** (e.g., APM, synthetic monitoring, infrastructure monitoring)
+ Experience supporting
** command center, TOC, or major incident management environments*
* + Knowledge of ITIL frameworks and service management best practices
+
Experience with automation or scripting (Python, Power Shell, or similar)
+ Bachelor's Degree in Business, Computer Science, or a related field or equal experience
+ ITIL v5 certification
+ Previous experience in the health care industry
** Additional Information:*
* Limited Geography Remote - This is a remote position but located within a specific geography.
To ensure Home or Hybrid Home/Office employees' ability to work effectively, the self-provided internet service of Home or Hybrid Home/Office employees must meet the following criteria:
At minimum, a download speed of 25 Mbps and an upload speed of 10 Mbps is required; wireless, wired cable or DSL connection is suggested.
Satellite, cellular and microwave connection can be used…
Position Requirements
10+ Years
work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×