Systems Reliability Engineer
Listed on 2026-08-05
-
IT/Tech
IT Support, SRE/Site Reliability
Systems Reliability Engineer
When you join the Members 1st team, you become part of something much bigger than a credit union. You become part of our faM1ly—a tight-knit bunch with big dreams and even bigger values. It is an exciting time for us as we continue to grow, and we hope that you will choose to grow along with us. Wanting the absolute best for our associates means more than just competitive pay.
It means fantastic healthcare, paid benefits, opportunities for professional advancement and work-life balance - and best of all, a place where you are accepted and respected for your individuality.
The Systems Reliability Engineer provides advanced operational support for platform, application, security, and infrastructure issues escalated from Tier 1 support. This role leverages AI-assisted operational tooling to receive and interpret event monitoring alerts, visualize system health and operational planes, and accelerate troubleshooting, impact analysis, and runbook execution. The engineer is responsible for triaging and resolving complex incidents, executing system-level support tasks, and maintaining IT operational runbooks using both traditional documentation and AI-assisted querying and revision techniques.
Key Responsibilities include but are not limited to:
- Triage and resolve Tier II support issues escalated by the Help Desk, including application access issues, system performance degradation, service outages, and infrastructure-level incidents.
- Perform system-level operational tasks to restore service, including IIS resets, application pool recycling, service restarts, and execution of approved Power Shell scripts in accordance with security and change controls.
- Monitor system health and operational telemetry, review logs, and proactively identify and remediate recurring or emerging issues.
- Research, diagnose, and resolve complex user issues across applications, integrations, and infrastructure components.
- Collaborate with Tier III, infrastructure, security, and application teams to escalate, coordinate, and resolve complex technical incidents.
- Manage incident, request, change, and problem tickets within the Service Now ITSM platform in alignment with established SLAs, escalation procedures, and governance standards.
- Participate in problem management activities by identifying root causes, documenting known errors, and contributing to permanent corrective actions.
- Support change management efforts, including review, testing coordination, execution of approved changes, and post-change validation.
- Analyze trends in incidents and requests to provide insights that support service reliability, operational maturity, and continuous improvement.
- Create, execute, and maintain IT operational runbooks to ensure consistent, repeatable support and response procedures.
- Keep runbooks current based on changes to systems, processes, or technologies, coordinating updates with engineering, platform, and support teams.
- Document support procedures, troubleshooting steps, and resolutions in Service Now Knowledge Base articles.
- Capture successful support outcomes and lessons learned in shared knowledge repositories to improve Tier I efficiency and enable self-service.
- Identify and analyze business needs, gather requirements, and help define scope and objectives for system enhancements or operational improvements.
- Research business requirements and document relationships between users, business processes, data, applications, and devices to support impact analysis.
- Translate business requirements into clear, actionable application or operational requirements.
- Make recommendations for technology-based solutions or process improvements using new or existing tools.
- Assist with product refinement and backlog prioritization by contributing operational insights and support data.
- Use AI-enabled monitoring and event intelligence tools to receive, interpret, and prioritize operational alerts across applications, infrastructure, security, and integrations.
- Leverage AI-assisted dashboards and visualizations to assess system health, service dependencies, and operational risk.
- Query AI tools to identify upstream and…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).