Lead Systems Operations Engineer
Job in
Iselin, Middlesex County, New Jersey, 08830, USA
Listed on 2026-08-28
Listing for:
Wells Fargo
Full Time
position Listed on 2026-08-28
Job specializations:
-
IT/Tech
Systems Engineer, SRE/Site Reliability, Cloud Computing: Infrastructure & Operations
Job Description & How to Apply Below
Lead Systems Operations Engineer
Wells Fargo is seeking a highly skilled and forward-thinking Lead Systems Operations Engineer to join our Quality & Test Engineering Operations team within CTO Platform Services team. This role is ideal for someone passionate about building scalable, resilient, and intelligent infrastructure solutions. You will play a key role in driving automation, reducing operational toil, and enabling self-service capabilities through cutting-edge technologies including Generative AI and Agent development.
In this role, you will:
- Lead complex, broad impact initiatives including provision of high level systems consultation for the technology teams
- Work as key participant in large scale planning of computer systems and network infrastructure for Systems Operations functional area
- Review and analyze complex technical challenges, as well as escalated support issues related to core business solutions that require in depth evaluation of multiple factors, such as alternatives, enhancements, periodic systems reviews, or improvements to existing systems
- Make decisions on technical changes and enhancements
- Consult with engineering team on change design requiring solid understanding of technical process controls or standards that influence and drive new initiatives
- Collaborate and consult with technical peers, colleagues, and mid to more experienced level managers to resolve systems support issues and achieve goals
- Collaborate and partner across platform, application, and engineering teams.
- Ability to manage multiple priorities in a fast-paced, high-impact production environment.
- Consistent delivery of high-quality reliability outcomes within expected timelines.
- High attention to detail, data-driven problem-solving, and operational rigor.
- Prior project or initiative leadership experience is highly desirable.
Required Qualifications:
- 5+ years of Systems Engineering, Technology Architecture experience, or equivalent demonstrated through one or a combination of the following: work experience, training, military experience, education
- 3+ years of Systems Engineering, Technology Architecture experience, or equivalent demonstrated through one or a combination of the following: work experience, training, military experience, education
- 3+ years of Proficiency in leveraging observability platforms such as Big Panda, Thousand Eyes, Grafana, Prometheus, ELK, Splunk Observability, and App Dynamics to enhance service reliability and performance monitoring
- 3+ years of experience in IT Service Management (ITSM), with a strong background in incident, problem, and change management processes.
- 3+ years of experience working with Red Hat Enterprise Linux and Kubernetes, with a strong focus on Red Hat Open Shift Container Platform (OCP).
- 3+ years of experience with Site Reliability Engineering and supporting production grade.
- 3+ years of experience with cloud-native architectures, high-availability systems, Cloud & Container Technologies like GCP or Azure and familiarity with Kubernetes.
- 3+ years of experience with Automation & Scripting: including developing and maintaining playbooks.
Desired
Qualifications:
- Strong hands-on experience applying SRE practices, including SLI/SLO definition, error budgets, and reliability metrics.
- Proven experience troubleshooting and resolving large-scale, distributed production systems.
- Hands-on experience with observability and monitoring tools such as Grafana, Splunk, Prometheus, Cribl, Thousand Eyes, App Dynamics, or equivalent, including dashboards, alerting, logs, and metrics.
- Strong scripting and automation skills using Python, Bash, and/or Power Shell to reduce operational toil.
- Experience building automation or reliability tooling using APIs, Git-based workflows, and modern engineering practices.
- Solid understanding of incident, problem, and change management in enterprise production environments.
- Strong communication and influencing skills across engineering teams and senior leadership.
- Experience with capacity management, performance engineering, and resiliency design (HA, fault tolerance, RTO/RPO).
- Experience operating in hybrid environments (on‑prem + cloud) with complex enterprise…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×