More jobs:
Principal Network Engineer - Reliability
Job in
Charlotte, Mecklenburg County, North Carolina, 28201, USA
Listed on 2026-08-05
Listing for:
Wells Fargo
Full Time
position Listed on 2026-08-05
Job specializations:
-
IT/Tech
Systems Engineer, SRE/Site Reliability, Network Engineer
Job Description & How to Apply Below
In this role, you will help shape the reliability strategy for critical network services at enterprise scale. You will influence how network reliability is measured, engineered, automated, and continuously improved while partnering across infrastructure, cloud, security, application, architecture, and business teams. This is an opportunity to drive modernization, reduce operational risk, and improve the stability of services that support critical business operations.
Role Details
* Employment Type:
Full time
* Role Type:
Individual contributor; this is not a people-manager role.
* Work Model:
Hybrid work model with three days per week in the office.
* Eligible Locations:
Dallas, TX metro;
Charlotte, NC metro; or Chandler, AZ metro.
* Escalation Expectations:
Provides senior technical escalation for critical incidents and high-risk changes. This role is expected to support critical incident response as needed; any recurring on-call rotation will be clearly defined before offer acceptance.
* Travel Expectations: 5% or less.
Core Technology Stack
This role is network-centric, with primary focus on routing and switching technologies, data center fabrics, WAN and campus networking, cloud and hybrid connectivity, load balancing, DNS, NTP, firewalls, automation, telemetry, monitoring, and observability tooling.
Key Responsibilities
You will provide senior technical leadership, hands-on engineering guidance, and cross-functional influence across the following areas:
* Network Reliability Strategy:
Define and drive reliability targets, SLOs/SLIs, risk measures, service health indicators, and improvement plans for critical network services.
* Incident Leadership:
Lead major incident response and service restoration, and drive root cause analysis, corrective action planning, and prevention of repeat incidents.
* Architecture Influence:
Review, challenge, and guide network designs to improve resiliency, scalability, security, capacity, failure isolation, and operational simplicity.
* Automation and Observability:
Expand telemetry, alerting, service health reporting, configuration management, automated validation, and self-service capabilities to reduce manual toil.
* Operational Excellence:
Strengthen standards, runbooks, documentation, change quality, failure readiness, compliance-aligned practices, and production readiness reviews.
* Cross-Functional Leadership:
Partner with network, cloud, security, infrastructure, application, architecture, and vendor teams while mentoring engineers and communicating complex technical issues clearly to senior stakeholders.
Required Qualifications
You should demonstrate principal-level technical depth, operational judgment, and cross-functional influence, including:
* 7+ years of experience in network engineering, reliability engineering, or network operations supporting business-critical enterprise network services.
* 5+ years of expert knowledge of routing, switching, TCP/IP, BGP, OSPF, high-availability design, and related network services including load balancing, DNS, NTP, firewalls, and network security.
* 5+ years of Proven ability to lead complex incident response, guide troubleshooting, restore service, and communicate clearly under pressure.
* 5 plus years of practical experience applying SRE or reliability engineering practices, including SLOs/SLIs, problem management, service health measurement, operational health metrics, and continuous improvement.
* 5 plus years experience improving operational stability and reducing repeat incidents through root cause analysis, post-incident reviews, corrective action planning, systemic remediation, and measurable recurrence prevention.
* 5 plus years experience owning corrective actions from post-incident review through implementation, validation, and closure.
* 5 plus years experience improving change success rates through risk assessment, peer review, validation testing, rollback planning, and post-change verification.
* 5 plus years experience conducting operational readiness reviews, production readiness assessments, failure-mode reviews, or service acceptance reviews.
* 5 plus years experience managing vendor escalations, product defects, support cases, and platform lifecycle risks that impact service stability.
* 5 plus years experience with capacity planning, performance analysis, traffic engineering, and resiliency planning for large-scale enterprise networks.
* 5 plus years of hands-on automation or scripting experience with tools such as Python, Power Shell, Bash, Ansible, Terraform, or equivalent technologies.
* 5 plus years of experience improving monitoring, telemetry, alerting, dashboards, metrics, service health reporting, or performance visibility for critical infrastructure services.
Preferred Qualifications
The following qualifications are beneficial and will help a candidate be successful in this role:
* Experience influencing enterprise network architecture and aligning designs with infrastructure, cloud, security,…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×