Principal Site Reliability Engineer
Job in
Yuma, Yuma County, Arizona, 85365, USA
Listed on 2026-09-17
Listing for:
Jobtailor
Full Time
position Listed on 2026-09-17
Job specializations:
-
Software Development
Job Description & How to Apply Below
- Apply software engineering, automation, and Dev Ops principles to improve how services are built, tested, deployed, observed, operated, and recovered
- Use data, evidence, experimentation, and rigorous engineering analysis to identify reliability risks, test assumptions, and guide technical decisions
- Define, implement, or improve SLIs, SLOs, error budgets, and service-health measures
- Improve observability through metrics, logging, tracing, monitoring, alerting, dashboards, and service-health instrumentation
- Drive continuous improvement across CI/CD, observability, deployment practices, Infrastructure as Code, automation, testing, incident response, capacity management, resilience and operational readiness
- Identify recurring or systemic production issues and translate operational experience into improvements in code, architecture, automation, tooling, and engineering practices
- Partner with Software Engineering teams to incorporate reliability, resiliency, scalability, performance, observability, recoverability, and operational readiness throughout the development lifecycle
- Participate in or lead incident response, troubleshooting, service restoration, and blameless post-incident learning
- Provide enterprise-level technical leadership for critical production incidents and influence engineering practices that improve incident response, escalation, service restoration, and sustainable on-call operations
- Reduce operational toil and unnecessary manual intervention through software, automation, reusable patterns, and better engineering practices
- Establish enterprise technical direction, develop senior technical leaders, and multiply the capability of the broader engineering organization
- Typically 15+ years of relevant professional experience in Software Engineering, Site Reliability Engineering, Systems Engineering, Cloud/Platform Engineering, Dev Ops, Infrastructure Engineering, Architecture, or a comparable technical discipline
- Experience with software development or scripting using one or more modern programming languages
- Experience with software engineering principles, distributed systems, production troubleshooting, automation, and observability appropriate to the level
- Experience with public cloud technologies and architectures, preferably AWS, along with infrastructure, networking, Linux/Unix, and modern application architectures
- Demonstrated analytical, problem-solving, communication, and collaboration skills appropriate to the scope of the role
- Preferred hands-on experience with AWS or comparable experience with Microsoft Azure, Google Cloud Platform (GCP), or Oracle Cloud Infrastructure (OCI)
- Experience developing, deploying, operating, or improving highly available production software or distributed systems
- Experience with CI/CD, Infrastructure as Code, containers or orchestration, observability, monitoring, alerting, and software-delivery automation
- Experience with SLIs, SLOs, error budgets, incident management, performance analysis, capacity management, resilience testing, disaster recovery, or operational readiness
- Experience creating reusable automation, tooling, platforms, patterns, or practices that improve engineering effectiveness
- Bachelor's degree in Computer Science, Software Engineering, Computer Engineering, Information Systems, or a related technical field, or equivalent practical experience
- Must independently possess eligibility to work in the United States at the date of hire
- Position is ineligible for employment Visa sponsorship
Demonstrates extensive experience in Software Engineering, Site Reliability Engineering, and Dev Ops principles, with a strong focus on automation, observability, and incident management. Proficient in leveraging cloud technologies, particularly AWS, to enhance system reliability and operational readiness.
Highest-signal resume keywords- 15+ Years Experience in Software Engineering
- Expertise in AWS or Comparable Cloud Technologies
- Proficient in CI/CD and Infrastructure as Code
- Experience with SLIs, SLOs, and Incident Management
- Strong Analytical and Problem-Solving Skills
- Software Development
- Scripting
- Distributed Systems
- Automation
- Observability
- Incident Management
- Performance Analysis
- Capacity Management
- Disaster Recovery
- Resilience Testing
- Analytical Skills
- Problem-Solving
- Communication
- Collaboration
- Site Reliability Engineering
- Dev Ops
- Infrastructure Engineering
- Cloud/Platform Engineering
- Software Engineering Principles
- AWS
- Microsoft…
Position Requirements
5+ Years
work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×