Global Manager of Site Reliability Engineering; Hybrid
Job in
Albany, Albany County, New York, 12201, USA
Listed on 2026-10-02
Listing for:
Broadridge
Full Time, Part Time
position Listed on 2026-10-02
Job specializations:
-
Software Development
Job Description & How to Apply Below
At Broadridge, we've built a culture where the highest goal is to empower others to accomplish more. If you’re passionate about developing your career, while helping others along the way, come join the Broadridge team.
Broadridge is Growing. We are seeking a strategic, hands-on Global Manager of Site Reliability Engineering to lead the reliability, release engineering, and operational evolution of its Integrated Platform. This event-driven platform connects enterprise applications with client-facing experiences through a common ontology, standardized APIs, and scalable event-driven capabilities across Capital Markets, Wealth Management, Investor Communications, and other business domains.
This leader will build and develop a high-performing SRE team while remaining deeply engaged in architecture, software engineering, automation, and production performance. Partnering with application engineering, product, architecture, security, and operations, the Senior Manager will ensure that reliability is engineered into services from design through deployment and ongoing operation.
The role requires strong technical depth in Java, Spring Boot, AWS, Kafka, and PostgreSQL, combined with demonstrated people leadership and delivery accountability. Dev Ops practices, release automation, and responsible AI enablement will be central to improving engineering productivity and production outcomes.
Job responsibilities:
Build and lead a high-performing SRE team. Own hiring, coaching, performance management, career development, and succession planning. Establish clear technical expectations, strengthen engineering judgment, and maintain sustainable operational and on-call responsibilities.
Own the SRE strategy and execution roadmap. Translate platform and business priorities into measurable improvements in reliability, release performance, automation, and scalability. Manage capacity, dependencies, and delivery commitments while protecting time for engineering improvements.
Remain hands-on in engineering decisions. Lead architecture and code reviews, guide complex troubleshooting, and contribute to critical automation and tooling. Ensure operational software meets the same standards for testing, security, maintainability, and documentation as application code.
Advance AWS infrastructure and platform automation. Develop reusable infrastructure-as-code patterns, consistent environment configurations, and automated provisioning and recovery. Strengthen IAM, networking, compute resilience, and capacity management across the platform’s AWS services.
Improve Kafka reliability and event-processing performance. Address partitioning, consumer-group behavior, consumer lag, schema evolution, delivery semantics, replay, and failure recovery. Validate that recovery approaches preserve intended processing behavior and data integrity.
Strengthen PostgreSQL performance and resilience. Partner with application and database engineers on schema design, indexing, query optimization, transaction behavior, connection pooling, safe migrations, and recovery testing.
Lead release engineering and deployment readiness. Improve CI/CD pipelines through automated testing, security checks, artifact traceability, and production validation. Establish safe deployment and rollback or roll-forward patterns that account for API compatibility, database changes, and event-schema dependencies.
Make service health measurable. Define service-level indicators, service-level objectives, and error-budget practices with service owners. Use metrics, logs, and distributed traces to improve detection, reduce alert noise, and prioritize engineering work based on business impact.
Reduce operational toil through automation. Build reusable tooling, self-service capabilities, and controlled remediation for well-understood failure scenarios. Measure reductions in manual effort, recurring incidents, and recovery time.
Enable practical, responsible AI adoption. Integrate approved AI tools into code and test development, infrastructure review, knowledge retrieval, and incident investigation. Validate outputs, measure benefits, protect sensitive information, and retain appropriate human review and authorization for production changes.
Lead incident resolution and resilience improvement. Coordinate technical response to complex incidents, communicate impact and recovery progress, and drive blameless reviews and permanent corrective actions. Validate capacity, failover, backup restoration, and disaster…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×