Technical Analyst - SRE
Listed on 2026-06-28
-
IT/Tech
SRE/Site Reliability
At Blue Cross and Blue Shield of Nebraska, we are a mission‑driven organization dedicated to championing the health and well‑being of our members and the communities we serve.
Our team is the power behind that promise. As the industry rapidly evolves, we are seeking forward‑thinking professionals to help optimize business processes and customer experiences, building a meaningful career and having a powerful impact in our community.
This Technical Analyst role focuses on Site Reliability and is positioned between engineering and analytics. You will build and maintain an observability framework across vendor‑managed and internal platforms, define SLO/SLA measurement strategies, drive MTTR reduction, lead incident response, and ensure platform reliability at enterprise scale. Your work bridges observability engineering, internal development teams, and business stakeholders by pulling telemetry together, establishing metrics, and driving improvement processes that hold system owners accountable for resilience.
AI is a core part of the operating model; you will leverage AI and AIOps tooling to automate incident triage, correlate alerts, generate runbooks, accelerate root‑cause analysis, and drive self‑healing remediation so engineers are not paged at 2 a.m. for problems AI can resolve.
The ideal candidate will live within driving distance of the Omaha, Nebraska office. This position allows remote flexibility but will require one day per week in the office.
What You’ll Do- Integrate and align enterprise observability across vendor‑managed and internal platforms, ensuring monitoring conforms to reliability policies, SLO/SLA targets, and MTTR goals.
- Own the service dependency map and asset criticality model—identify tier‑1 assets, their fail‑over behavior, connections, and single points of failure.
- Establish end‑user experience monitoring—detect when the member or provider experience breaks, not just infrastructure telemetry alerts.
- Leverage AI and AIOps tooling for incident triage, root‑cause analysis, automated runbook execution, predictive anomaly detection, alert correlation, and self‑healing remediation.
- Define and track SLOs for platform reliability—availability, latency, error rates—into business‑aligned targets and error budgets.
- Build and maintain MTTR and reliability dashboards—platform health, incident trends, perfect‑day streaks, and vendor SLA compliance.
- 3+ years of experience in Site Reliability Engineering (SRE), infrastructure engineering, platform operations, Dev Ops, or a related technical operations role supporting enterprise systems.
- Hands‑on experience with enterprise observability and monitoring platforms, ticketing systems, operational dashboards, and at least one scripting language such as Python, Power Shell, or Bash. Experience with AI/AIOps tools, telemetry, logs, traces, alert correlation, and incident response processes is strongly preferred.
- Bachelor’s degree in Computer Science, Information Technology, or a related field (or equivalent combination of education and experience), strong communication skills, analytical problem‑solving ability, and working knowledge of ITIL fundamentals including incident, change, and problem management.
- Certifications in SRE, Dev Ops, cloud, or observability; experience in healthcare payer or health‑plan environments;
Agile/Scrum delivery; and partnership with engineering teams to translate reliability work into backlog items and sprint‑ready user stories. - Experience with AI/ML operations tooling, CI/CD pipelines, infrastructure‑as‑code, CMDB or service dependency mapping, log aggregation or SIEM platforms, distributed systems reliability, chaos engineering, and report/dashboard authoring or data modeling.
To perform this job successfully, an individual must be able to perform each essential duty satisfactorily. The requirements listed are representative of the knowledge, skill, and or ability required. Reasonable accommodations may be made to enable individuals with disabilities to perform the essential functions. Other duties may be assigned.
#J-18808-Ljbffr(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).