Manager of Site Reliability Engineer
Job in
San Francisco, San Francisco County, California, 94199, USA
Listed on 2026-09-12
Listing for:
Forge
Full Time
position Listed on 2026-09-12
Job specializations:
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer
Job Description & How to Apply Below
- As an engineering organization, we pride ourselves on engineering as a creative activity
- Engineering managers enable engineers to do their best work by maintaining a culture and environment where engineers can achieve autonomy, mastery, and purpose
- The Manager, Site Reliability Engineering will lead Forge’s SRE team responsible for keeping Forge systems highly available for customers, while partnering closely with Platform, Engineering, Security, Compliance, and Product teams to improve reliability, observability, incident response, and operational maturity
- This is an opportunity for a hands‑on technical leader who can coach engineers, improve production operations, and help Forge build and run secure, scalable, and highly reliable products
- Partner with Security, Compliance, and Risk partners to ensure reliability and infrastructure practices meet the needs of a regulated business
- Manage Forge’s Site Reliability Engineering team responsible for keeping Forge systems highly available for customers
- Drive strong incident management practices in partnership with engineering teams, including response, mitigation, follow-up, and post‑incident learning
- Build, improve, and manage observability infrastructure in partnership with Platform Engineering, including monitoring, alerting, dashboards, and operational metrics
- Improve monitoring coverage and alert quality to reduce noise, shorten time to detect, and support faster response and mitigation
- Champion reliability best practices across engineering, including service ownership, operational readiness, disaster recovery, and production support standards
- Contribute to technical design, architecture, automation, infrastructure, and overall team delivery
- Collaborate with engineering teams to troubleshoot production issues, identify recurring problems, and improve system reliability
- Hire, coach, mentor, and manage performance for SRE team members while supporting career development and team health
- Health and wellness:
Medical, dental, and vision are available to all employees. - Smart savings: 401k plan, life insurance, HSA and FSA, and pre‑tax mass transit and parking benefits.
- Ownership:
Equity is offered to all qualifying employees. - Life balance:
New parent leave and generous time off. - We grow with you :
We encourage and support your growth through education, training, and mentorship. - Relationship building:
Team lunches, employee interest groups, and company celebrations. - Community and culture:
Employee‑led resource groups, volunteering, and community service. - Rewards and recognition:
Continuous recognition of our employees’ milestones, wins, and hard work. - Food and snacks:
Healthy snacks, drinks, and weekly lunches when in the office.
- 10+ years of total software engineering, infrastructure, platform, cloud, or production operations experience
- Strong technical judgment, communication skills, and ability to influence across engineering and non‑engineering stakeholders
- Experience building, operating, and maintaining large‑scale cloud infrastructure and distributed systems
- Bachelor’s degree in Computer Science, Engineering, or a closely related field, or equivalent practical experience
- Hands‑on experience with observability, monitoring, alerting, incident response, troubleshooting, and production support
- 5+ years of experience leading a Site Reliability Engineering, Dev Ops, Cloud Operations, or similar reliability‑focused function
- Experience with CI/CD, infrastructure automation, cloud platforms, and operational tooling
- Experience supporting growth‑stage companies where speed, scale, reliability, and operational discipline must be balanced
- Experience improving developer experience through paved‑road platforms, standardization, and self‑service infrastructure capabilities
- Familiarity with Kubernetes, container platforms, infrastructure‑as‑code, Terraform, Ansible, or similar automation tooling
- Experience with observability platforms such as Datadog, Cloud Watch, or similar tools
- Experience with AWS and/or Azure cloud platforms
- Experience in Fin Tech, financial services, or another regulated industry
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×