×
Register Here to Apply for Jobs or Post Jobs. X

Site Reliability Engineer

Job in Phoenix, Maricopa County, Arizona, 85027, USA
Listing for: OneAZ Credit Union
Full Time position
Listed on 2026-08-24
Job specializations:
  • IT/Tech
    Disaster Recovery IT, Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, IT Infrastructure
Job Description & How to Apply Below

Site Reliability Engineer

Join Us in Making an Impact At OneAZ Credit Union, our success is measured only by yours. We're here to create lasting change in the lives of our members, our communities, and our team. If you're looking for a career with purpose, where your work truly matters—you've found it!

The Site Reliability Engineer is responsible for ensuring the reliability, availability, performance, and recoverability of OneAZ's technology platforms and infrastructure. This position combines infrastructure engineering, automation, monitoring, and resiliency practices to maintain highly available systems supporting associates and members. The engineer designs and implements automation solutions using Power Shell and other scripting technologies, administers enterprise monitoring platforms, leads disaster recovery testing activities, and partners with technology teams to improve operational resilience, service reliability, and recovery readiness across on-premises and cloud environments.

This role serves as a key contributor to incident response, infrastructure modernization, and continuous improvement initiatives focused on reducing operational risk and improving system uptime.

The Site Reliability Engineer works closely with Infrastructure, Information Security, Application Support, Enterprise Architecture, and business teams to identify operational risks, strengthen recovery capabilities, and improve the overall resilience of technology services that support associates and members.

Essential Functions

  • Administer, maintain, and optimize Solar Winds and other enterprise monitoring platforms.
  • Develop and maintain monitoring dashboards, alerts, reports, and performance metrics.
  • Proactively monitor and improve infrastructure health, availability, capacity, and performance, identifying and mitigating potential issues before they impact service reliability or user experience.
  • Proactively identify infrastructure risks, reliability concerns, and opportunities for improvement.
  • Lead disaster recovery planning, testing, failover exercises, and recovery validation activities.
  • Develop and maintain disaster recovery documentation, runbooks, recovery procedures, and test plans.
  • Coordinate and execute system failover and failback activities for critical applications and infrastructure.
  • Partner with Infrastructure, Security, and Application teams to improve system resiliency and operational readiness.
  • Support business continuity planning initiatives and technology recovery efforts.
  • Lead the investigation and resolution of complex infrastructure and service reliability incidents, conduct root cause analyses (RCAs), identify underlying systemic issues, and drive corrective and preventive actions to improve availability, performance, resiliency and operational excellence.
  • Track, document, and report on disaster recovery testing results, remediation activities, and recovery readiness metrics.
  • Maintain infrastructure diagrams, recovery documentation, and operational procedures.
  • Support audits, examinations, and compliance activities related to disaster recovery, resiliency, and infrastructure operations.
  • Research and recommend technologies and best practices that improve reliability, monitoring, and recoverability.
  • Participate in infrastructure maintenance activities, upgrades, and projects.
  • Develop and maintain automation scripts using Power Shell, Python, Bash, or similar tools to streamline infrastructure operations, automate remediation of common issues, and enhance service availability and performance.

What You Bring

  • High School Diploma Required
  • Bachelor's Degree in Information Technology, Computer Science, Information Systems, Engineering, or a related technical field; or equivalent combination of education and experience. Required
  • 5-8 years similar or related experience of experience supporting enterprise infrastructure environments and administering infrastructure monitoring platforms.. Required
  • Deep knowledge of networking and cloud technologies, Windows Server, virtualization, and storage with hands on experience supporting and optimizing enterprise infrastructure environments.
  • Experience with disaster recovery testing,…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary