Senior SRE
Listed on 2026-10-03
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Disaster Recovery IT
Banyan Software is the best permanent home for software businesses that serve specialized industries, their employees, and their customers. With a buy-grow-and-hold-for-life approach and a permanent capital base, Banyan acquires and grows companies worldwide, honoring founder legacies and helping portfolio companies modernize through shared AI expertise and operational discipline. Founded in 2016, Banyan operates more than 120 portfolio companies across North and South America, Europe, and APAC, and has appeared on the Inc.
5000 list for six consecutive years. The Banyan Software Foundation, endowed with $100 million in Banyan stock, leverages technology to build a greener and more equitable world.
We are seeking a highly experienced and hands-on SRE to own the operational excellence of the modernized SaaS applications produced by the Banyan AI Factory. This is not a role focused on building the factory itself; instead, you will run the reliability of the modernized applications the factory delivers to our Operating Companies (OpCos).
You will join a team that provides 24x7 coverage with rotating on-call responsibilities, serving as Tier 1 Site Reliability Engineering (SRE) for our OpCos’ distributed applications. Day to day this will include: automated deployments, cloud service integration, application performance and availability monitoring/observability, and security incident response across our two target clouds — Amazon Web Services (AWS) and Microsoft Azure. The ideal candidate has a track record of keeping secure, highly available production systems running at scale.
Key Responsibilities- 24x7 Operations & On-Call:
Operate as part of a team providing round-the-clock coverage of OpCo containerized applications, participating in a rotating on-call schedule to ensure continuous availability and rapid response. - Tier 1 SRE & Operations:
Serve as Tier 1 SRE for the modernized applications, managing day-to-day cloud integrations across our two target clouds — AWS and Azure — to keep production systems healthy, performant, and secure. - Performance & Availability Monitoring/Observability:
Implement and maintain robust application observability tooling (monitoring, logging, tracing) to track performance and availability, proactively detect degradation, and drive down mean-time-to-detect and mean-time-to-resolve. - Disaster Recovery and Service Restoration:
Develop, maintain, test, and execute disaster recovery and business continuity procedures. Ensure the timely recovery and restoration of services following geographic disruptions, cyber incidents, infrastructure failures, or other disaster events. - Security Incident Response:
Respond to security incidents and operational events affecting OpCo SaaS platforms, executing established runbooks, coordinating remediation - Automation & Infrastructure-as-Code :
Use Infrastructure-as-Code (Terraform) and CI/CD pipelines (e.g., Git Hub Actions, Git Lab CI) to manage, deploy, and automate the operational environments of modernized applications, reducing toil and improving consistency. - AI Agents & Dev Sec Ops Scale:
Build scale in our Dev Sec Ops practice by designing, building, and operating AI agents that automate SRE tasks and incident response, reducing toil and accelerating detection, triage, and remediation. - Hands-on Problem Solving:
Serve as a technical escalation point for operational challenges, applying strong analytical skills to resolve infrastructure, network, and automation issues across distributed, multi-tenant SaaS environments while navigating technical ambiguity.
- Experience:
5–7 years of progressive…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).