Senior Lead Infrastructure Engineer — Prem OpenShift Platform
Listed on 2026-09-06
-
IT/Tech
Systems Engineer, SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, IT Infrastructure
We're looking for a talented, senior engineering professional ready to take their career to new heights at one of the world's most influential companies.
As a Senior Lead Infrastructure Engineer at JPMorgan Chase within Enterprise Technology Compute Infrastructure Platforms team, you will design, engineer, and operate an on-premises Open Shift-based platform that hosts both containers and virtual machines (e.g., Open Shift Virtualization / Kube Virt). You will own the underlying infrastructure and Open Shift cluster foundations, compute, hypervisor platforms, operating systems, networking, storage, security controls, and reliability enabling product teams to deploy workloads safely, consistently, and efficiently.
This is an infrastructure engineering role focused on platform lifecycle and reliability, partnering closely with hardware engineering, networking, storage, security, identity/access, and platform consumers.
- Design Open Shift cluster architectures (on-premises and/or hybrid) to meet availability, scalability, security, and operability requirements.
- Buildandoperate Open Shift clusters end-to-end, including install, upgrade, patching, scaling, resilience testing as well as virtualization capabilities supporting VM lifecycle (images/templates and relevant hardware acceleration concepts such as SR-IOV/DPDK/GPU passthrough, where applicable).
- Engineer the platform to support both Kubernetes workloads and VM workloads, including capacity planning, performance management, placement strategies, and HA/DR considerations.
- Own Linux platform fundamentals (e.g., RHEL/CoreOS concepts, kernel/sysctl tuning, certificates, and identity integration basics) required for reliable cluster operations.
- Implementandtroubleshootcore networking capabilities (routing/switching fundamentals, DNS, L4/L7 concepts, load balancing, firewalling, segmentation, and packet-path analysis).
- Integrate storage services for container and VM workloads (block/file/object concepts, CSI drivers, performance and failure modes, and backup/restore approaches).
- Drive reliability practices including observability, incident response, root-cause analysis, and preventative improvements across the platform lifecycle.
- Developrunbooks, standards, and reference architectures;lead operational readiness reviews and post-incident actions.
- Uses enterprise-authorized AI capabilities within the work environment to accelerate analysis of complex infrastructure signals and documentation of mitigation options, validating outputs and handling operational data according to sensitivity and security requirements.
- Leads reuse-first adoption of AI-assisted practices across delivery and automation routines to reduce recurring issues, ensuring changes are validated, traceable and auditable, and aligned to resiliency and security expectations.
- Formal training or certification on infrastructure engineering concepts and 5+ years applied experience.
- Hands-on administration ofRed Hat Open Shift and Kubernetesin production environments (e.g., install/upgrade patterns such as IPI/UPI, Operators, SCC/RBAC, cluster operators, ingress).
- Experience operating Kubernetes/container platforms through day-2 operations (upgrades, scaling, troubleshooting, and platform lifecycle management).
- Experience with
OpenShift Virtualization / Kube Virt(or equivalent VM-on-Kubernetes) and VM/container co-tenancy design considerations. - Strong Linux administration and troubleshooting experience (e.g., system performance, certificates, OS configuration, and cluster node operations).
- Demonstrated troubleshooting depth inat least twoof the following domains: networking, storage, virtualization.
- Experience designing and operating highly available platforms, including capacity management, incident handling, root-cause analysis, and remediation tracking.
- Experience partnering with security/compliance stakeholders on hardening, RBAC, secrets/certificates handling, and vulnerability management aligned to secure-by-default configurations.
- Experience implementing infrastructure automation and configuration-as-code practices (e.g.,…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).