Senior Application Support Engineer/Site Reliability Engineer; SRE
Listed on 2026-07-24
-
IT/Tech
IT Support, SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer
Senior Application Support Engineer
Are you ready to make an impact at DTCC? Do you want to work on innovative projects, collaborate with a dynamic and supportive team, and receive investment in your professional development? At DTCC, we are at the forefront of innovation in the financial markets. We are committed to helping our employees grow and succeed. We believe that you have the skills and drive to make a real impact.
We foster a thriving internal community and are committed to creating a workplace that looks like the world that we serve. The Information Technology group delivers secure, reliable technology solutions that enable DTCC to be the trusted infrastructure of the global capital markets. The team delivers high-quality information through activities that include development of essential, building infrastructure capabilities to meet client needs and implementing data standards and governance.
Pay and Benefits:
- Competitive compensation, including base pay and annual incentive
- Comprehensive health and life insurance and well-being benefits, based on location
- Pension / Retirement benefits
- Paid Time Off and Personal/Family Care, and other leaves of absence when needed to support your physical, financial, and emotional well-being.
- DTCC offers a flexible/hybrid model of 3 days onsite and 2 days remote (onsite Tuesdays, Wednesdays and a third day unique to each team or employee).
The Impact You Will Have in This Role
As a Senior Application Support Engineer, you will help power DTCC's global financial markets infrastructure by ensuring the reliability, availability, and performance of Institutional Trade Processing (ITP) platforms that support cross-border equity and debt trade processing and settlement.
Leveraging Site Reliability Engineering (SRE) principles, you will support a portfolio of 40+ mission-critical applications across a modern ecosystem of AWS, Open Shift Container Platform (OCP), Kafka, IBM MQ, and distributed systems. You will play a key role in driving operational excellence, improving resiliency, reducing operational risk, and advancing automation across critical trade processing platforms.
Working closely with globally distributed teams across Application Development, Infrastructure, Cloud, Network, and Operations, you will help deliver stable, scalable, and highly available services that support DTCC's mission-critical business functions.
Your
Primary Responsibilities:
Production Reliability & Incident Management
- Ensure the availability, stability, and performance of mission-critical applications.
- Lead incident response, troubleshooting, and root cause analysis (RCA) activities.
- Drive preventative solutions that improve resiliency and reduce recurring issues.
Operational Excellence & Resiliency
- Support application recovery, failover, disaster recovery (DR), and business continuity activities.
- Maintain operational readiness across a portfolio of enterprise applications.
- Support critical processing schedules and service-level commitments.
Change, Release & Automation
- Support application deployments, releases, and vendor upgrades.
- Drive automation initiatives that improve efficiency and reduce operational risk.
- Enhance monitoring, alerting, observability, and operational tooling.
Collaboration & Governance
- Partner with Development, Infrastructure, Cloud, Database, Network, Security, and Product teams to ensure platform reliability.
- Support audit, risk, compliance, and operational control requirements.
- Build strong partnerships across globally distributed teams to drive continuous improvement.
Qualifications
- Bachelor's degree preferred or equivalent practical experience
- 6–8 years of experience supporting enterprise applications in complex production environments
Talents Needed for Success:
Core Technologies
- Java/J2EE
- Oracle, SQL, DB2
- IBM MQ, Kafka
- Splunk, Grafana, Auto Sys
- Service Now and ITIL
- Linux/Unix
- AWS
Required Experience
- Strong understanding of Site Reliability Engineering (SRE) principles, including reliability, observability, automation, and incident prevention.
- Experience supporting enterprise-scale, distributed applications in production environments.
- Experience with cloud technologies and modern…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).