Senior Engineer, Platform & Site Reliability
Listed on 2026-08-22
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer, AWS
Overview
As a key player within Intercontinental Exchange's (ICE) innovative servicing technology division, our team is dedicated to delivering cutting-edgemortgage processing solutions on a resilient, cloud-native platform. This role is pivotal to our platform engineering and site reliability initiatives, owning the Amazon EKS (Kubernetes) foundation on which our product teams build and run. The Senior Engineer, Platform & Site Reliability will apply deepexpertisein AWS,Git Opswith Argo CD,Crossplane, the Istioservice mesh, and modern observability to keep our systems secure, scalable, andhighly availableacross multiple regions and environments.
Java (Spring) and React (Type Script) development remain part of the role, ensuring you can build platform tooling and contribute to the applications you operate. By joining our team, you will directly shape the reliability, performance, and operability of the platform that powers our business.
Designs, builds, and operates the cloud-native platform and reliability tooling for the MSP DX (IMT) with an emphasis on a secure, scalable, andhighly availablefoundation. Our engineers manage Amazon EKS clusters across multiple AWS regions and environments in an Agile SDLC, delivered entirely through
GitOps. Responsible for the provisioning and lifecycle of Kubernetes clusters and cloud infrastructure, CI/CD pipelines, service mesh, observability frameworks, and secrets management, while also contributing to the Java microservices and Reactmicro frontends that run on the platform.
Provisions, upgrades, and operates Amazon EKS (Kubernetes) clusters across multiple AWS regions and environments (UAT, stable, production, and chaos), ensuring a secure, scalable, andhighly availableplatform.
Manages cloud infrastructure declaratively through Kubernetes using
Crossplane, and implements
GitOpspractices with
ArgoCDto deliver all infrastructure as code.Operates and upgrades the Istio service mesh, including canary rollouts, along with Envoy and ingress, for traffic management, routing, and mutual TLS.
Designs andoperatesmonitoring, logging, and observability solutions using Prometheus, Grafana, Jaeger,Open Telemetry(OTEL), Kiali, and Fluent Bit, and defines SLOs, alerting, and dashboards.
Improves system reliability through capacity planning, autoscaling, incident response, on-call practices, and chaos engineering.
Administers platform services such as cert-manager, external-dns, external-secrets, sealed-secrets, and the AWS Load Balancer Controller, and integrates with AWS services including DMS and MSK.
Builds andmaintainsCI/CD pipelines (Azure Dev Ops) and automated security scanning (for example, Sonar Qube) to deliver changes safely and repeatably.
Provides full-stack Java (Spring) and React (Type Script) development for platform tooling and for the microservices and micro frontends that run on the platform.
Designs and develops APIs and automation that support platform capabilities and self-service for product teams.
Participates in reliability and architecture design ceremonies and analyzes system needs todeterminetechnical requirements.
Writes technical specifications and operational runbooks based on conceptual design and stated business and reliability requirements.
Develops and/or reviews automated tests and reliability validation before release, with an emphasis on Unit,Component, and Scenario tests.
Troubleshootsoperational failures in both test and production environments andleads rootcause analysis.
Mentors or guides the work of less experienced site reliability and software engineers.
Remains current on industry standards in cloud, Dev Ops, SRE, and web development disciplines.
Performsadditionalrelated duties as assigned.
Bachelor's Degree or the equivalent combination of education, training, or work experience.
5+ years of site reliability, platform, Dev Ops, or software engineering work experience.
Hands-on experience operating Kubernetes and cloud-native technologies in production, preferably Amazon EKS on AWS.
Experience with infrastructure as code andGitOpspractices and tools.
Experience developing andmaintainingCI/CD pipelines.
Workingproficiencyin at least one general-purpose language, such as Java or Type Script/JavaScript,sufficient to build automation and platform tooling.
Experience with Kubernetes,ArgoCD,Crossplane, Istio, Envoy, Jaeger, Prometheus, Grafana, Kiali, or similar technologies.
Experience running workloads with cloud providers (preferably AWS) and/or Open Shift, and with the Java JVM.
Experience with modern observability frameworks, including SLO/SLI definition and alerting.
Experienceoperatinga service mesh, including upgrades and canary rollouts.
Experience with secrets management and certificate automation, such as external-secrets, sealed-secrets, and cert-manager.
Experience with chaos engineering and resilience testing.
Experience with RESTful service development and working with microservices…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).