Software Engineer - Resiliency and Platform Engineering
Listed on 2026-07-20
-
Software Development
Cloud Engineer - Software, DevOps, Software Engineer
Overview
This role is not eligible for sponsorship. Four days onsite hybrid at our N. Scottsdale office.
Choice Hotels has an exciting opportunity as our Staff Software Engineer, Resiliency & Platform Engineering in the Sky Touch Technology division. Sky Touch Technology provides the most widely used cloud-based (SaaS) hotel property management system. As a key member of Sky Touch, you will strengthen the resiliency, safety, and operability of a large-scale, multi-tenant SaaS platform by improving foundational platform capabilities, runtime behavior, and the developer experience used to build and operate our systems.
This role sits at the intersection of software engineering, platform engineering, and resiliency, focusing on building shared capabilities, libraries, frameworks, tooling, guardrails, and standards used by dozens of engineers across the organization. These capabilities make resilient behavior the default for application teams and reduce operational risk through better system design rather than reactive response. This is not a traditional Site Reliability Engineering (SRE) role.
Resiliency and platform engineering are proactive, year-round engineering disciplines aimed at preventing failures and enabling teams to build and operate services safely emphasis is on durable, systemic improvements and developer enablement rather than pager-driven operations or feature delivery. You may apply AI-assisted tools and techniques pragmatically to reduce engineering toil, improve diagnostics, and accelerate resiliency and platform outcomes, prioritizing durability, correctness, and adoption over experimentation.
- Design and implement platform-level capabilities including shared libraries, frameworks, tooling, automation, and guardrails that improve application resiliency, runtime safety, and developer experience across the ecosystem, prioritizing durability over short-term delivery.
- Strengthen foundational platform and runtime behavior by identifying and eliminating systemic failure modes such as memory leaks, unsafe defaults, brittle error handling, poor failure propagation, and resource exhaustion.
- Define, standardize, and evolve logging, monitoring, alerting, and observability practices to improve signal quality, reduce noise, and enable faster diagnosis and recovery.
- Partner with Principal Software Engineers, Solution Architects, and Engineering Managers to identify systemic risks and translate them into well-scoped platform and resiliency initiatives.
- Work across software engineering resiliency, data engineering resiliency, and platform engineering teams to design shared solutions and raise the technical bar.
- Engage directly in application codebases, particularly during ramp-up, to understand real-world system behavior, identify failure patterns, and validate resiliency improvements.
- Participate in incident postmortems and operational reviews to translate lessons learned into durable platform or resiliency improvements.
- Evaluate, prototype, and introduce tools and technologies that measurably improve developer productivity, platform safety, and application resiliency, prioritizing adoption and long-term impact.
- Apply AI-assisted development, diagnostics, and operational tools to improve engineering productivity and resiliency outcomes.
- Influence engineering practices and technical direction through design reviews, reference implementations, mentorship, and technical leadership.
- Bachelor’s degree in computer science or a related field, or equivalent practical experience.
- Typically 8–10+ years designing, building, and supporting large-scale software systems in production.
- Hands-on experience with Java-based services, including Spring Boot, in virtualized and containerized environments.
- Experience with cloud-native and serverless workloads, including Python-based services and event-driven architectures.
- Strong practical experience in AWS and an understanding of how cloud-managed services impact reliability and operability.
- Working knowledge of relational and non-relational data stores and how data characteristics influence system design.
- Experience using application monitoring…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).