Sr. Site Reliability Engineer
Listed on 2026-07-28
-
IT/Tech
Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, Systems Engineer
Select how often (in days) to receive an alert:
Join a dynamic team at the pulse of global markets, where we deliver innovative software and service solutions for essential financial reporting and capital markets transactions. At DFIN, we are a values-driven organization that empowers you to build a fulfilling career while bringing your authentic self to work every day. Our “Win as One” mentality ensures that our team’s success is directly linked to Client, Shareholder and Employee Satisfaction.
In 2026, DFIN was named #1 on the 2026 Top 100 Global Most Loved Workplaces® by Best Practice Institute. We have also been recognized as one of America’s Most Loved Workplaces® for five consecutive years and a Built In Best Place to Work for six years, reflecting our continued commitment to supporting employees’ total well-being. Enjoy competitive compensation, a flexible workplace, comprehensive benefits, and opportunities for professional growth.
Bring your passion and talents to DFIN – because being YOU thrives here.
We are looking for technical team members at all levels who want to push themselves to deliver best in market SaaS solutions. We offer a challenging environment where you will have to grow, adapt and use your skills consistently. Our customers rely on us in the moments that matter. Engineering delivers on that promise.
The Senior Site Reliability Engineer is responsible for ensuring our SaaS products are fast, stable and optimized for our customers. SRE’s at DFIN take on availability, performance,managing change, monitoring, response and are guardians of non-functional requirements.
You either havean
SaaS infrastructure background with a programmatic, automated mindset or are someone that comes with a software engineering background with SaaS infrastructure experience. The SRE goal is to build automated systems that reduce or eliminate manual work to keep our products up and running and performing optimally. We are looking for someone who thrives on collaboration within the team and across other groups and can operate independently to deliver solutions.
- Champion and implement a culture of SRE to maintain a high-quality platform infrastructure in DFIN SaaS products
- Leverage AI tools to enhance system reliability, including intelligent observability, incident prediction and automated remediation across cloud infrastructure
- Evaluate and implement emerging AI powered operations and observability solutions to proactively improve system performance, reliability and scalability
- Champion and implement application and infrastructure monitoring and alerting to prevent client impacting issues by ensuring system availability, performance and scalability to maintain SLOs and SLAs
- Optimize application performance at scale
- Automate everything including system operational runbooks
- Define and support continuous integration and deployment pipelines (CI/CD) aligned to branching and quality assurance strategies
- Dive deep into technology and stay on the forefront of the latest tools, technologies, and strategies; help evaluate, prototype, and integrate them into work processes
- Perform with broad independence and deliver on project milestones and tasks on schedule while communicating progress regularly
- Build strong relationships with SRE team members and software engineering teams to hold each other accountable for quality expectations
- Learn continuously and apply lessons learned
- Evangelize best practices, eliminate bottlenecks, and improve process
- Participate in on-call duties 365/24/7 and lead the triage and RCA of production incidents
- 5+years experience designing, building, securing, monitoring and maintaining cloud infrastructure in Azure or AWS
- Experience applying AI capabilities within
Cloud Ops operations - Relevant certifications or training in AI, Cloud AI services or AIOps platforms are a plus
- 5+years experience writing software in any modern software language such as C# .NET, Java
- 5+years experience creating automated deployments with tools such as Harness, Azure Dev Ops, Ansible or Jenkins to manage Infrastructure as Code and software build and deployment in a continuous integration…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).