Site Reliability Engineer - Paze
Listed on 2026-07-08
-
IT/Tech
Cloud Computing: Infrastructure & Operations, Systems Engineer
Overview
At Early Warning, we’ve powered and protected the U.S. financial system for over thirty years with solutions like Zelle®, Paze℠, and more. As a trusted name in payments, we partner with thousands of institutions to increase access to financial services and protect transactions for hundreds of millions of consumers and small businesses.
Positions located in Scottsdale, San Francisco, Chicago, or New York follow a hybrid work model to allow for a more collaborative working environment.
Candidates responding to this posting must independently possess the eligibility to work in the United States, for any employer, at the date of hire. This position is ineligible for employment Visa sponsorship.
RoleOverall Purpose
The Staff Site Reliability Engineer partners with development teams by defining availability standards and implementing availability and resiliency patterns in applications and infrastructure.
Essential Functions- Design and implement software and tools to improve performance, availability, scalability, and latency, delivering end products to customers with high efficiency while meeting security standards.
- Support the company’s commitment to risk management and protecting the integrity and confidentiality of systems and data.
- Build automation and tooling around application management, including deployments, configuration changes, and disaster recovery scenarios.
- Design, implement, and evangelize observability and monitoring systems to proactively detect problems and identify root causes.
- Continuously evaluate the capacity of applications and provide statistics to Product/Business teams; recommend scalable paths for future needs.
- Identify performance bottlenecks and collaborate with cross-functional teams to troubleshoot and resolve issues.
- Serve as a technical liaison for the application and provide documents and runbooks to Level 1 and Level 2 teams.
- Participate in a 24x7 on-call rotation.
- Champion excellent processes by developing repeatable patterns and standard, reusable work across teams.
- Work with application development teams to provide feedback and technical requirements to the software development lifecycle, implementing best-practice microservice design patterns and modern software development approaches.
- Support the adoption of best-practice microservice design patterns and other modern reliability techniques.
- Be a thought leader: act as a senior point of expertise on site reliability engineering issues, industry trends, and developing technologies; coach and mentor team members.
- Support the company’s commitment to risk management and protecting the integrity and confidentiality of systems and data.
- Education and experience typically obtained through completion of a Bachelor’s Degree in Business and/or Computer Science or related field.
- Typically 8+ years of related progressive experience managing large complex projects in a technical or software development environment, including post-graduate degree.
- Proven ability to lead a team through high-priority incidents and improve the RCA process.
- Excellent troubleshooting skills and proven experience resolving technical issues in complex environments.
- Hands-on experience with one or more of the following technologies:
Python, Go, Java;
Docker. - Experience in Microservices Architecture and messaging frameworks such as Kafka, SQS or JMS.
- Database technologies such as Oracle, DynamoDB, Aurora; caching layers like Redis and Memcached.
- Strong understanding of Linux administration.
- Experience with CI/CD pipeline implementation including Git, Chef, Maven, Jenkins, etc.
- Strong understanding of TCP/UDP/IP protocols.
- Experience in leading cross-functional teams to create technical solutions and a track record of designing and building complex end-to-end systems (full stack).
- Background and drug screen.
- Good programming skills in Java, Ruby, Python, JavaScript, and Go.
- Hands-on experience supporting applications in a 24x7 customer-facing production environment.
- Working knowledge of AWS, Docker, Kubernetes, Swarm.
Employee must be able to perform essential functions and physical requirements of the position with or…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).