Cloud Development Engineering Senior Principal; Jersey , Tampa, Dallas
Listed on 2026-08-29
-
Software Development
Cloud Engineer - Software, DevOps, Backend Developer
Are you ready to make an impact at DTCC?
Do you want to work on innovative projects, collaborate with a dynamic and supportive team, and receive investment in your professional development? At DTCC, we are at the forefront of innovation in the financial markets. We are committed to helping our employees grow and succeed. We believe that you have the skills and drive to make a real impact. We foster a thriving internal community and are committed to creating a workplace that looks like the world that we serve.
The Information Technology group delivers secure, reliable technology solutions that enable DTCC to be the trusted infrastructure of the global capital markets. The team delivers high-quality information through activities that include development of essential, building infrastructure capabilities to meet client needs and implementing data standards and governance.
Pay and Benefits:- Competitive compensation, including base pay and annual incentive
- Comprehensive health and life insurance and well-being benefits, based on location
- Paid Time Off and Personal/Family Care, and other leaves of absence when needed to support your physical, financial, and emotional well-being.
- DTCC offers a flexible/hybrid model of 3 days onsite and 2 days remote (onsite Tuesdays, Wednesdays and a third day unique to each team or employee).
Primary Responsibilities:
- Build and migrate the cloud-native platform. Design stateless, horizontally scalable OCP services and move the estate incrementally toward full cloud adoption without a big-bang rewrite.
- Remove delivery bottlenecks through automation. Create disposable, reproducible, self-service environments that shorten testing cycles and turn constrained test capacity into an engineering capability.
- Engineer production resilience, observability, and cost discipline. Make systems diagnosable by design with actionable traces, metrics, SLIs, and cost controls across compute, storage, telemetry, and cloud consumption.
- Minimum of 10 years of related experience
- Bachelor's degree preferred or equivalent experience
- 10+ years of experience building and operating production systems, with the majority of that experience being hand-on and current coding responsibilities.
- Deep cloud-native development expertise, including containers, Kubernetes/Open Shift, stateless service design, and resilience patterns applied appropriately — retry with backoff, circuit breaking, idempotency, graceful degradation, and safe rollback.
- Strong proficiency in at least one of Java/Python with meaningful understanding of infrastructure-as-code and CI/CD practices — including Terraform, Helm, and pipeline design as core engineering responsibilities rather than ancillary scripting.
- Proven distributed systems troubleshooting experience, demonstrated through specific production scenarios such as consumer lag and rebalance storms, partition skew, GC pauses and OOMKills, connection pool exhaustion, object-store throttling, and network partitions — including how these issues were diagnosed under time pressure.
- Strong instrumentation discipline, including Open Telemetry, distributed tracing across service boundaries, custom metrics, structured logging with correlation IDs, RED/USE methodology, and SLI/SLO definition. This includes cardinality discipline and an understanding that poor instrumentation can create both reliability and cost challenges.
- Hands-on observability tooling depth across platforms such as Prometheus, Grafana, Datadog, Dynatrace, or Splunk, with evidence of building the underlying signals rather than only consuming dashboards created by others.
- Cloud cost engineering and Fin Ops experience, including rightsizing and autoscaling policy, spot and reserved capacity, storage lifecycle and tiering, egress awareness, Kafka retention economics, Snowflake warehouse sizing and auto-suspend, telemetry ingestion volume, tagging and showback discipline, and the ability to express cost per unit of work rather than only as a monthly total.
- Streaming and data platform exposure, including Kafka operations, object storage, lakehouse table formats, and Snowflake. The role…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).