×
Register Here to Apply for Jobs or Post Jobs. X

Principal Site Reliability Engineer

Job in Durham, Durham County, North Carolina, 27701, USA
Listing for: Fidelity Investments
Full Time position
Listed on 2026-08-05
Job specializations:
  • IT/Tech
    Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, Systems Engineer, IT Support
Job Description & How to Apply Below

Job Title

Facilitates and orchestrates Data Recovery (DR) events including organizational cloud and on-premise routing, failovers, and evidence captures. Improves observability and resilience of software and infrastructure through incident root cause analysis using tools including Datadog and Splunk. Performs SSL certificate management work and renew certificates in non-prod and prod environments before expiry to continue workflow. Deploys and supports highly distributed multitiered systems lds and operates highly resilient platforms in Amazon Web Services (AWS) Cloud environments.

Designs, develops, and executes performance tests using Java, JMeter, Cloud-test, Rush-hour, and other performance testing tools to ensure comprehensive performance testing. Automates with scripting languages -- Python and Shell scripting. Works with Cloud Computing and Dev Ops concepts including Continuous Integration and Continuous Delivery (CI/CD) pipelines and Kubernetes. Builds and improves standard methodologies for performance, load, stress, and chaos testing, along with analytics and reports based on business requirements.

Primary

Responsibilities
  • Defines and leads enterprise-level reliability strategies.
  • Architects resilient systems and infrastructure.
  • Creates and publishes performance test results report with recommendations on quality improvement.
  • Maintains scalability and resiliency of complex environment.
  • Implements advanced observability practices and techniques at scale.
  • Manages and interprets large datasets using query languages and visualization tools.
  • Advises senior leadership on reliability engineering best practices.
  • Mentors junior engineers.
  • Performs independent and complex technical and functional analysis for multiple divisional initiatives.
  • Develops innovative solutions to improve system availability, scalability, and performance.
  • Designs, implements, and maintains performance test frameworks.
Education and Experience

Bachelor's degree in Computer Science, Applied Computer Science, Engineering, Information Technology, Information Systems, or a closely related field (or foreign education equivalent) and five (5) years of experience as a Principal Site Reliability Engineer (or closely related occupation) managing mission-critical applications and administering resilient platform infrastructure across testing and production.

Or, alternatively, Master's degree in Computer Science, Applied Computer Science, Engineering, Information Technology, Information Systems, or a closely related field (or foreign education equivalent) and three (3) years of experience as a Principal Site Reliability Engineer (or closely related occupation) managing mission-critical applications and administering resilient platform infrastructure across testing and production.

Skills and Knowledge

Candidate must also possess:

  • Demonstrated expertise enabling an end-to-end, Continuous Integration and Continuous Delivery (CI/CD) platform using uDeploy, Jenkins Core, Ansible AWX, and Terraform; and modifying programming language scripts in Shell and Python for software application deployments and targeted utility purposes through Deployment as a Service (DAAS).
  • DE enabling Dev Ops and Site Reliability Engineering (SRE) practices and principles in multi-Cloud environments, using automations and proactive monitoring for Azure and AWS services.
  • DE configuring internet-facing network, traffic-routing, firewall, and web servers within a complex environment, using F5, AVI, AWS Route
    53, and Azure Load Balancer; and troubleshooting complex problems that span multiple component tiers and operating systems, using different Splunk and Datadog dashboards.
  • DE creating and operating monitors and dashboards for the tracking, alerting, and presentation of application metrics, traces, and logs, using Datadog, Splunk, and Grafana.
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary