×
Register Here to Apply for Jobs or Post Jobs. X

Lead Site Reliability Engineer

Job in San Francisco, San Francisco County, California, 94199, USA
Listing for: Federal Reserve Bank of New York
Full Time position
Listed on 2026-08-31
Job specializations:
  • IT/Tech
    Cloud Computing: Infrastructure & Operations, Systems Engineer, SRE/Site Reliability
Salary/Wage Range or Industry Benchmark: 147000 - 234000 USD Yearly USD 147000.00 234000.00 YEAR
Job Description & How to Apply Below

Company Federal Reserve Bank of San Francisco When you join the Federal Reserve—the nation's central bank—you’ll play a key role, collaborating with leading tech professionals to strengthen and protect our economic, financial and payments systems. We invest in contemporary and emerging technology each year to support the Federal Reserve and our economy, and we’re building a dynamic and diverse team for our future.

We are seeking an experienced Lead Site Reliability Engineer to join our engineering team and drive the reliability, scalability, and performance of our critical systems. This role combines deep technical expertise in software engineering, cloud infrastructure, and Dev Ops practices to ensure our services meet the highest standards of availability and operational excellence.

Responsibilities System Reliability & Performance
  • Design, implement, and maintain highly available, scalable, and resilient systems across cloud infrastructure
  • Establish and monitor SLIs, SLOs, and SLAs to ensure optimal system performance
  • Lead incident response, conduct root cause analysis, and implement preventive measures
  • Develop and maintain disaster recovery and business continuity plans
Infrastructure & Automation
  • Architect and manage cloud infrastructure on AWS using Infrastructure as Code (Terraform)
  • Automate deployment pipelines, monitoring, and operational workflows
  • Optimize cloud resource utilization and cost management
Engineering & Development
  • Build and maintain internal tools and services to improve operational efficiency
  • Collaborate with development teams to implement reliability best practices
  • Conduct code reviews and provide technical guidance on system design
  • Develop monitoring solutions, alerting systems, and observability frameworks
Security & Compliance
  • Integrate security practices into CI/CD pipelines (SAST/DAST)
  • Implement and maintain security controls across infrastructure and applications
  • Ensure compliance with industry standards and regulatory requirements
  • Conduct security assessments and vulnerability management
Leadership & Collaboration
  • Mentor junior SRE team members and promote SRE culture across the organization
  • Partner with software engineering teams to improve system reliability
  • Drive technical initiatives and contribute to architectural decisions
  • Document processes, runbooks, and technical specifications
Software Engineering
  • Strong proficiency in Java, Python, and Node.js
  • Experience with microservices architecture and distributed systems
  • Solid understanding of data structures, algorithms, and design patterns
  • Proficiency in writing clean, maintainable, and testable code
Cloud Infrastructure (AWS)
  • Compute:
    Lambda, ECS, EC2, Fargate
  • Storage: S3, EBS, EFS
  • Database: RDS, DynamoDB, Aurora
  • Networking: VPC, Route
    53, Cloud Front, API Gateway
  • Monitoring:
    Cloud Watch, X-Ray
  • AWS certifications (Solutions Architect, Dev Ops Engineer) preferred
Dev Ops & CI/CD
  • Expert-level knowledge of Git Lab (CI/CD pipelines, runners, Git Ops)
  • Advanced Terraform skills for infrastructure provisioning and management
  • Experience with containerization (Docker) and orchestration (Kubernetes/ECS)
  • Proficiency with configuration management tools
Security
  • Hands-on experience with SAST (Static Application Security Testing) tools
  • Knowledge of DAST (Dynamic Application Security Testing) methodologies
  • Understanding of security best practices, OWASP Top 10, and compliance frameworks
  • Experience with secrets management and identity access management (IAM)
Monitoring & Observability
  • Experience with monitoring tools (Grafana, Datadog, New Relic, or similar)
  • Log aggregation and analysis (Cloud Watch Logs, Splunk)
  • Distributed tracing with AWS X-Ray
Qualifications
  • Bachelor's degree in Computer Science, Engineering, or related field, or equivalent practical experience
  • 7+ years of experience in Site Reliability Engineering, Dev Ops, or related roles
  • 3+ years in a lead or senior technical position
  • Proven track record of managing large-scale production systems
  • Experience with on-call rotations and incident management
  • GenAI based Applications:
    Working knowledge of LLMs and agentic applications a plus
  • Experience with serverless architectures and event-driven systems
  • Familiarity with…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary