IT/Tech, Cloud Computing: Infrastructure & Operations, SRE/Site Reliability
Listed on 2026-08-12
-
IT/Tech
Cloud Computing: Infrastructure & Operations, SRE/Site Reliability
If you are looking for a game-changing career, working for one of the world's leading financial institutions, you've come to the right place.
As a Principal Engineer at JPMorgan
Chase within the Commercial and Investment Bank Digital Channels and Connectivity team, you will own the technical strategy and delivery of the firm's next-generation test environment platform — enabling every developer across a 10,000-engineer organization to spin up an isolated, composable, production-representative virtual environment per pull request, run integration tests independently and simultaneously, and merge with confidence without contention on shared environments.
You will solve the fundamental constraint that limits flow of code to production: environment bottlenecks. Your approach will be composable — reusing most unchanged components while instantiating only the services under test, with intelligent dependency resolution, traffic isolation, and automatic teardown. You will design and build the MVP hands‑on, then shift to hardening, socializing, and driving adoption at scale across multiple organizations.
You will partner as a peer to engineering delivery leadership and collaborate closely with performance engineering, SRE, and quality engineering to ensure the platform meets the demands of a large, hybrid (cloud and on-premises) service estate.
- Architect the composable virtual environment platform — enabling per-PR, per-developer isolated test sessions that reuse unchanged shared services and instantiate only the components under test/change, with correct dependency resolution across the service graph
- Design the traffic routing and isolation layer that directs requests to the correct version of each service within a virtual environment session, leveraging header-based routing, service mesh capabilities, and service discovery without cross-session contamination
- Build the dependency graph intelligence that determines which services must be instantiated live vs. reused from baseline vs. virtualized/stubbed for a given test scenario — making composable environments feasible at 1,000+ service scale
- Deliver a self-service developer experience through CLI, API, and developer portal integration that enables teams to declare environment needs, provision in minutes, execute integration tests, and exit — with zero dependency on a shared environment or manual coordination
- Define and implement data strategies for composable environments, including synthetic data generation, production data masking, snapshot/restore workflows, schema versioning, and repeatable dataset management across heterogeneous data stores (Oracle, Kafka, Cassandra, MongoDB, Cockroach DB, DynamoDB, PostgreSQL)
- Implement lifecycle management, cost governance, and resource reclamation — including time-to-live policies, idle detection, automatic teardown, quota management, chargeback attribution, and Fin Ops reporting across cloud and on-premises infrastructure
- Integrate the platform into CI/CD pipelines as a native stage — environments created on PR open, tests executed, results reported, environments destroyed on merge/close — with support for Jenkins, Spinnaker, Git Hub Actions, and Argo CD
- Ensure the platform operates reliably across hybrid infrastructure (Kubernetes on-premises and EKS, VMs, and managed cloud services) with consistent provisioning semantics, cross-network connectivity, and acceptable spin-up latency regardless of hosting tier
- Establish platform reliability SLOs, observability, and operational runbooks — measuring time-to-environment, environment success rate, isolation correctness, and developer satisfaction as first-class product metrics
- Architects and governs agentic AI-enabled engineering workflows (using enterprise-authorized tools within the work environment) to improve delivery speed, code quality, and operational outcomes at scale (e.g., AI-driven PR review assistance, test generation/maintenance, release readiness checks, incident triage and root-cause acceleration), while defining guardrails for validation, security, resiliency, and reuse across teams
- Applies knowledge of tools within the Software…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).