×
Register Here to Apply for Jobs or Post Jobs. X

Sr. Lead Infrastructure Engineer - Storage SRA

Job in Plano, Collin County, Texas, 75086, USA
Listing for: JPMorganChase
Full Time position
Listed on 2026-08-22
Job specializations:
  • IT/Tech
    Systems Engineer, SRE/Site Reliability, IT Infrastructure
Salary/Wage Range or Industry Benchmark: 180000 - 240000 USD Yearly USD 180000.00 240000.00 YEAR
Job Description & How to Apply Below

Become a member of a team where you can contribute significantly to shaping the future of a world‑renowned and influential company. Among top performers, you can make a direct and meaningful impact.

Senior Lead Infrastructure Engineer – Corporate Sector Infrastructure Platforms

As a Senior Lead Infrastructure Engineer at JPMorgan Chase within the Corporate Sector – Infrastructure Platforms, you exhibit both depth and breadth of knowledge regarding software, applications, and technical processes across multiple technical disciplines. You also have a specialization in a specific domain within infrastructure engineering to drive programs or initiatives consisting of multiple technologies and applications.

Job Responsibilities
  • Applies deep technical expertise and problem‑solving methodologies focused on analyzing complex data and systems, anticipating issues, and finding ways to mitigate risk
  • Works with other platforms to architect and implement changes required to resolve issues and modernize the organization and its technology processes
  • Be responsible for infrastructure engineering in accordance with business requirements
  • Executes work according to compliance standards, risk and security, and business objectives
  • Own end‑to‑end problem detection, resolution, and prevention; lead efforts to improve MTTD, MTTM, and MTTR through better observability, triage, and remediation practices
  • Drive a workstream or project spanning one or more infrastructure engineering technologies, including technology lifecycle management and dependency management across network, compute, database, and application platforms
  • Partner with adjacent platform teams to architect and implement changes that resolve systemic issues, reduce operational risk, and modernize technology and operational processes
  • Design and deliver creative, scalable solutions for high‑complexity engineering challenges, including development of automation, tooling, and repeatable operational patterns
  • Evaluate upstream/downstream impacts across systems, data flows, and integrations; proactively identify risks and provide clear mitigation and contingency recommendations
  • Expertise to continuously improve telemetry and observability of storage products with tools like Grafana, Dynatrace, Prometheus, Splunk, Netcool, etc.
Required Qualifications , Capabilities, and Skills
  • Formal training or certification on site reliability engineering concepts and 5+ years applied experience
  • Own reliability and operational excellence for IP‑based storage services, including availability, performance, capacity, and operational risk reduction
  • Act as the SME for enterprise storage platforms (e.g., Dell EMC Power Flex, Net App Solid Fire, Pure Storage, PMAX) including design input, troubleshooting, upgrades, and lifecycle management; provide 24x7 support coverage, incident response, and escalations; lead restoration activities during high‑severity events
  • Drive SRE best practices; define, measure, and improve SLIs/SLOs, error budgets, and operational KPIs for storage services. Build and maintain observability: metrics, logs, traces (where applicable), dashboards, alerts, and runbooks
  • Reduce operational load by identifying and eliminating toil through automation; automate repetitive tasks (provisioning, health checks, failover validation, reporting, hygiene tasks); implement safe automation with guardrails, change controls, and rollback strategies. Expert knowledge of AAAS automation
  • Perform root cause analysis (RCA) and problem management; implement corrective and preventive actions to prevent recurrence
  • Maintain and continuously improve runbooks, standard operating procedures, on‑call playbooks, and knowledge articles; collaborate with engineering, network, compute, and platform teams on architecture reviews, change planning, and reliability improvements
  • Support capacity management and performance engineering: forecasting, trending, saturation analysis, and proactive remediation; ensure compliance with operational standards (change management, risk controls, documentation, audit readiness)
  • Role participates in a 24x7 on‑call rotation and is expected to respond to production incidents within…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary